Optical remote sensing cloud and cloud shadow removal method, system and device and storage medium

By pre-processing and feature fusion processing of optical images and SAR images, combined with structural reconstruction of semantic complementary feature images, the shortcomings of cloud and cloud shadow removal in optical remote sensing images are solved, and high-quality image recovery effect is achieved.

CN120298250AActive Publication Date: 2025-07-11SHANDONG UNIV OF SCI & TECH

Patent Information

Application Number
CN202510756828.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-11
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

The prior art has insufficient effect in the cloud and cloud shadow removal method in optical remote sensing images, especially in extreme cloud coverage and complex geographic structure areas, restoring spectrality and retaining structural edges, low cloud removal quality and poor image recovery effect.

Method used

Optical remote sensing cloud and cloud shadow removal methods are adopted to pre-process the optical image and SAR image, including dynamic logarithmic compression, adaptive non-local mean filtering, deep denoising and structural restoration, combined with fast Fourier transform and wavelet enhancement, followed by feature fusion based on mutual attention mechanism and channel space coupled attention processing, and finally structural reconstruction of semantic complementary feature images.

Benefits of technology

It effectively improves the structural consistency and clarity of the image after cloud removal, retains high-frequency details and low-frequency semantics, improves the image reconstruction quality, has good removal effect and high recovery quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298250A_ABST
    Figure CN120298250A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image enhancement, in particular to an optical remote sensing cloud and cloud shadow removal method, system and device and a storage medium. In order to solve the technical problem of poor image cloud removal effect in the prior art, the method comprises the following steps: firstly, respectively preprocessing an optical image and an SAR image, solving the problem of speckle noise existing in the SAR image, and strengthening high-frequency information in the optical image; then, fusing the penetrating power of the SAR enhanced image and the spectral information of the optical enhanced image, and carrying out channel and space coupling attention mechanism processing on a fusion result to obtain a semantic complementary feature image; and finally, performing structure reconstruction processing on the semantic complementary feature image for several times, retaining high-frequency details and low-frequency semantics, improving image reconstruction quality, and performing convolution processing on the obtained structure reconstruction feature image to obtain a cloud noise removed reconstruction image. The method is applied to the field of image cloud and cloud shadow removal, the removal effect is good, and the image recovery quality is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image enhancement, and specifically to an optical remote sensing cloud and cloud shadow removal method, system, device and storage medium. Background Art

[0002] In the field of remote sensing image cloud and cloud shadow removal, the existing technology reduces clouds and cloud shadows by adopting hierarchical design and stacked residual groups, thereby improving the spectral and structural fidelity of the reconstructed image; in addition, the residual block with channel attention mechanism can adaptively emphasize useful features while suppressing speckle noise in synthetic aperture radar images (abbreviated as SAR images) or cloud distortion in optical data, thereby improving the quality of multimodal fusion. However, there are fluctuations between different spectral bands, resulting in insufficient performance in restoring spectral properties and retaining structural edges in areas with extreme cloud cover and complex ground object structures, especially in dense clouds or clouds with complex textures. The cloud removal quality is low and the image restoration effect is poor. Summary of the Invention

[0003] The purpose of the present invention is to provide an optical remote sensing cloud and cloud shadow removal method, system, device and storage medium.

[0004] The technical solution of the present invention is as follows: An optical remote sensing cloud and cloud shadow removal method includes the following operations: S1. Obtain the optical image and SAR image of the area to be processed; the SAR image is processed by dynamic logarithmic compression to obtain a dynamically compressed SAR image; the dynamically compressed SAR image is processed by adaptive non-local mean filtering to obtain a filtered SAR image; the filtered SAR image is processed by deep denoising and structure restoration to obtain a SAR enhanced image; the optical image is processed by fast Fourier transform to map the optical image spatial domain to the frequency domain, and wavelet enhancement processing is performed in the frequency domain to obtain an optical texture edge enhanced image; after the optical texture edge enhanced image is processed by inverse Fourier transform, it is added element by element to the optical image to obtain an optical enhanced image; S2. Perform feature fusion processing on the SAR enhanced image and the optical enhanced image based on the mutual attention mechanism to obtain a feature fusion image; the feature fusion image is processed by a channel and spatial coupling attention mechanism to obtain a semantic complementary feature image; the operation of the channel and spatial coupling attention mechanism is: obtain the pixel gradient feature map of the feature fusion image, after convolution processing, perform spatial attention processing and channel attention processing respectively, and after element-by-element multiplication, obtain an attention mechanism feature image; extract the context information of the attention mechanism feature image, and perform pyramid processing by a multimodal dynamic mechanism to obtain a semantic complementary feature image; S3. The semantically complementary feature image is processed by several structure reconstruction processes to obtain a structure reconstruction feature image; the structure reconstruction feature image is processed by convolution to obtain a cloud noise-removed reconstruction image.

[0005] The operations of depth denoising and structure restoration processing in S1 are as follows: The filtered SAR image is processed by convolution to obtain a filtered SAR convolution image; the filtered SAR convolution image is processed by several residual connection processes based on depthwise separable convolution and residual connection processes based on nonlinear processing to obtain a filtered SAR residual connection image; the filtered SAR residual connection image is processed by convolution to obtain a SAR enhanced image.

[0006] The operations of feature fusion processing based on the mutual attention mechanism in S2 are as follows: The SAR enhanced image is processed by convolution, multiplied element-wise with the optical enhanced image, and then added element-wise to the optical enhanced image to obtain an initial fusion image; the query feature, value feature, and key feature of the initial fusion image are respectively processed by convolution and then processed by the attention mechanism to obtain an attention fusion image; the attention fusion image is processed by global pooling, fully connected, and feature reshaping, and then multiplied element-wise with the attention fusion image to obtain an attention fusion enhanced image; the SAR enhanced image is processed by global pooling, fully connected, and feature reshaping, and then multiplied element-wise with the SAR enhanced image to obtain a SAR enhanced feature image; the SAR enhanced feature image and the attention fusion enhanced image are added element-wise to obtain a feature fusion image.

[0007] The operations of the pyramid processing of the multimodal dynamic mechanism in S2 are as follows: The context information of the attention mechanism feature image is processed by convolution to obtain context convolution features; the context convolution features are upsampled and added element-wise to the context convolution features to obtain initial semantically complementary features; the initial semantically complementary features are upsampled and added element-wise to the context convolution features to obtain a semantically complementary feature image.

[0008] The operation of structural reconstruction processing on the semantic complementary feature image in S3 is as follows: The semantic complementary feature image is processed through convolution, ReLU activation function, and convolution to obtain a semantic complementary convolutional feature image; the semantic complementary feature image is processed through a channel and spatial coupling attention mechanism to obtain a semantic complementary attention feature image; the semantic complementary feature image is processed through a multi-layer perceptron to obtain a semantic complementary perception feature image; the semantic complementary perception feature image and the semantic complementary feature image are multiplied element-wise to obtain a first semantic complementary fusion image; the semantic complementary convolutional feature image and the semantic complementary attention feature image are added element-wise to obtain a second semantic complementary fusion image; the first semantic complementary fusion image and the second semantic complementary fusion image are added element-wise to obtain an initial structural reconstruction image; the structural gradient information of the initial structural reconstruction image is fused with the optical reference image, and after convolution processing, it is added element-wise to the initial structural reconstruction image to obtain a first structural reconstruction feature image, which is used to perform the operation of the second structural reconstruction processing.

[0009] The dynamic compression of the SAR image in S1 is obtained based on the maximum brightness, standard deviation, and image information entropy of the SAR normalized image; the SAR normalized image is obtained by normalizing the SAR image.

[0010] The smoothing parameter in the adaptive non-local mean filtering process in S1 is obtained based on the pixel information at each position in the dynamic compression SAR image.

[0011] An optical remote sensing cloud and cloud shadow removal system for implementing the above optical remote sensing cloud and cloud shadow removal method includes: An image preprocessing module for obtaining an optical image and a SAR image of the area to be processed; the SAR image is processed through dynamic logarithmic compression to obtain a dynamic compression SAR image; the dynamic compression SAR image is processed through adaptive non-local mean filtering to obtain a filtered SAR image; the filtered SAR image is processed through depth denoising and structure restoration to obtain a SAR enhanced image; the optical image is processed through fast Fourier transform to map the optical image from the spatial domain to the frequency domain, and wavelet enhancement processing is performed in the frequency domain to obtain an optically textured edge enhanced image; after the optically textured edge enhanced image is processed through inverse Fourier transform, it is added element-wise to the optical image to obtain an optical enhanced image; The semantic complementary feature image generation module is used to perform feature fusion processing on the SAR enhanced image and the optical enhanced image based on the mutual attention mechanism to obtain a feature fusion image; the feature fusion image is processed by a channel and spatial coupled attention mechanism to obtain a semantic complementary feature image; the operation of the channel and spatial coupled attention mechanism is as follows: obtain the pixel gradient feature map of the feature fusion image, after convolution processing, perform spatial attention processing and channel attention processing respectively, and after element-wise multiplication, obtain an attention mechanism feature image; extract the context information of the attention mechanism feature image, and after pyramid processing by a multi-modal dynamic mechanism, obtain a semantic complementary feature image; The cloud noise removal and reconstruction image generation module is used to perform several structural reconstruction processes on the semantic complementary feature image to obtain a structural reconstruction feature image; the structural reconstruction feature image is processed by convolution to obtain a cloud noise removal and reconstruction image.

[0012] An optical remote sensing cloud and cloud shadow removal device includes a processor and a memory. Among them, when the processor executes the computer program stored in the memory, the above-mentioned optical remote sensing cloud and cloud shadow removal method is implemented.

[0013] A computer-readable storage medium is used to store a computer program. Among them, when the computer program is executed by a processor, the above-mentioned optical remote sensing cloud and cloud shadow removal method is implemented.

[0014] The beneficial effects of the present invention are as follows: The present invention provides an optical remote sensing cloud and cloud shadow removal method. First, the optical image and the SAR image of the area to be processed are preprocessed respectively to solve the problems of high dynamic range and multiplicative speckle noise in the SAR image, and enhance the high-frequency information in the optical image. Then, the penetration ability of the SAR enhanced image and the spectral information of the optical enhanced image are subjected to feature fusion processing based on the mutual attention mechanism, which solves the problems of image blur and information loss in the optical remote sensing image under cloud occlusion, and obtains a feature fusion image. The feature fusion image is processed by a channel and spatial coupled attention mechanism to enhance the complementarity between the SAR image and the optical image, so as to accurately fuse multi-source information, effectively improve the structural consistency and clarity of the image after cloud removal, and obtain a semantic complementary feature image. Finally, the semantic complementary feature image is subjected to several structure reconstruction processes. During the structure reconstruction process, the semantic complementary convolution features reflecting the underlying texture edge information of the image, the semantic complementary attention features reflecting the key area information of the image channels and spatial dimensions, and the semantic complementary perception features reflecting the high-level semantic abstraction information of the image are fused to enhance the structural details, and the gradient information is used to provide a prior structure constraint for the reconstruction, which can retain both high-frequency details and low-frequency semantics during the reconstruction, improve the image reconstruction quality, and obtain a structure reconstruction feature image. The structure reconstruction feature image is subjected to convolution processing to obtain a cloud noise removal reconstruction image. It is applied in the field of image cloud and cloud shadow removal, with good removal effect and high image restoration quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] By reading the detailed description of the preferred embodiments below, the solutions and advantages of the present application will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered as a limitation of the present invention.

[0016] In the drawings: Figure 1 FIG. is a schematic flowchart of the method of this embodiment in the embodiment; Figure 2 FIG. is a qualitative comparison diagram of the visual processing capabilities of the method of this embodiment and three existing methods in three types of cloud cover scenarios in the embodiment; Figure 3 FIG. is a quantitative comparison diagram of the visual processing capabilities of the method of this embodiment and three existing methods in three types of cloud cover scenarios in the embodiment; Figure 4 FIG. is a comparison diagram of the determination coefficients between the predicted NDVI and the ground truth of the method of this embodiment and three existing methods in the embodiment; Figure 5 FIG. is the effect diagram of cloud and cloud shadow removal of the method of this embodiment in different large-scale scenarios in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings.

[0018] This embodiment provides an optical remote sensing cloud and cloud shadow removal method. Refer to Figure 1 , including the following operations: S1. Obtain the optical image and SAR image of the area to be processed; the SAR image is subjected to dynamic logarithmic compression processing to obtain a dynamically compressed SAR image; the dynamically compressed SAR image is subjected to adaptive non-local mean filtering processing to obtain a filtered SAR image; the filtered SAR image is subjected to deep denoising and structure restoration processing to obtain a SAR enhanced image; the optical image is subjected to fast Fourier transform processing to map the optical image spatial domain to the frequency domain, and wavelet enhancement processing is performed in the frequency domain to obtain an optical texture edge enhanced image; after the optical texture edge enhanced image is subjected to inverse Fourier transform processing, it is added element by element to the optical image to obtain an optical enhanced image; S2. Perform feature fusion processing on the SAR enhanced image and the optical enhanced image based on a mutual attention mechanism to obtain a feature fusion image; the feature fusion image is subjected to channel and spatial coupled attention mechanism processing to obtain a semantic complementary feature image; The operation of the channel and spatial coupled attention mechanism processing is as follows: obtain the pixel gradient feature map of the feature fusion image, after convolution processing, perform spatial attention processing and channel attention processing respectively, and after element-by-element multiplication, obtain an attention mechanism feature image; extract the context information of the attention mechanism feature image, and after pyramid processing by a multi-modal dynamic mechanism, obtain a semantic complementary feature image; S3. The semantic complementary feature image is subjected to several structure reconstruction processes to obtain a structure reconstruction feature image; the structure reconstruction feature image is subjected to convolution processing to obtain a cloud noise removal reconstruction image. The specific operation steps are as follows.

[0019] S1. Obtain the optical image and SAR image of the area to be processed; the SAR image is subjected to dynamic logarithmic compression processing to obtain a dynamically compressed SAR image; the dynamically compressed SAR image is subjected to adaptive non-local mean filtering processing to obtain a filtered SAR image; the filtered SAR image is subjected to deep denoising and structure restoration processing to obtain a SAR enhanced image; the optical image is subjected to fast Fourier transform processing to map the optical image spatial domain to the frequency domain, and wavelet enhancement processing is performed in the frequency domain to obtain an optical texture edge enhanced image; after the optical texture edge enhanced image is subjected to inverse Fourier transform processing, it is added element by element to the optical image to obtain an optical enhanced image.

[0020] Preprocess the optical image and SAR image of the area to be processed respectively, solve the problems of high dynamic range and multiplicative speckle noise existing in the SAR image, and strengthen the high-frequency information in the optical image, improve the quality of the image to be processed, so as to improve the subsequent cloud removal and reconstruction effect.

[0021] First, obtain the optical image and SAR image of the area to be processed containing cloud noise, which are used as the objects for cloud and cloud shadow processing, and perform image reconstruction.

[0022] Then, preprocess the optical image and SAR image respectively to solve the problems of high dynamic range and multiplicative speckle noise existing in the SAR image, and enhance the high-frequency information in the optical image.

[0023] The process of preprocessing the optical image is as follows: The SAR image is subjected to dynamic logarithmic compression processing to map its original exponential distribution to the linear domain, effectively reducing the image dynamic range and improving the stability of subsequent processing, and obtaining a dynamically compressed SAR image; The dynamically compressed SAR image is subjected to adaptive non-local mean filtering processing to remove part of the multiplicative noise while keeping the key structures and edge information undamaged, and obtaining a filtered SAR image; The filtered SAR image is subjected to deep denoising and structure restoration processing to enhance the long-range dependence and detail texture retention ability of the image, improving the usability and expression ability of the SAR image, and obtaining a SAR enhanced image.

[0024] The dynamically compressed SAR image in S1 is obtained based on the brightness maximum value, standard deviation, and image information entropy of the SAR normalized image; The SAR normalized image is obtained by normalizing the SAR image; That is, the operation of the above dynamic logarithmic compression processing is realized by the following formula: , , is the dynamically compressed SAR image, C is the contrast adjustment constant (which varies dynamically based on the pixel characteristics of the SAR image to be processed), is the SAR normalized image (obtained by normalizing the SAR image), , , are respectively the brightness maximum value, standard deviation, and image information entropy of the SAR normalized image, , are respectively the first empirical coefficient and the second empirical coefficient, is the denominator compensation amount. Compared with the existing logarithmic compression processing method, the dynamic logarithmic compression processing method in this embodiment can adaptively adjust the logarithmic compression intensity according to the brightness distribution of the image, improving the image compression accuracy.

[0025] The smoothing parameter in the above adaptive non-local mean filtering process is dynamically changed compared to the fixed smoothing parameter in the traditional non-local mean filtering process. It is obtained based on the pixel information at each position in the dynamically compressed SAR image, enabling enhanced sensitivity to similarity differences and protection of structural information at image edges or high-texture regions, and also enhancing the smoothing effect in flat image regions. The smoothing parameter in this embodiment is calculated by the following formula: , is the smoothing parameter at the position (x, y) of the dynamically compressed SAR image, is the initial smoothing coefficient, is a hyperparameter, 、 are the variances of neighboring pixels in the x-direction and y-direction at the position (x, y) of the dynamically compressed SAR image, respectively.

[0026] The operations of the above deep denoising and structure restoration processing can be implemented by the Swin Transformer network. However, to further improve the ability of the image for long-range dependence and maintain detailed texture, and enhance the usability and expressiveness of the SAR image, this embodiment designs an efficient deep denoising and structure restoration processing method. The operation process is as follows: The filtered SAR image is subjected to convolution processing to obtain a filtered SAR convolution image; the filtered SAR convolution image is subjected to several residual connection processes based on depthwise separable convolution and residual connection processes based on nonlinear processing to obtain a filtered SAR residual connection image; the filtered SAR residual connection image is subjected to convolution processing to obtain a SAR enhanced image.

[0027] Among them, the operation of the residual connection process based on depthwise separable convolution is as follows: The input is subjected to convolution, depthwise separable convolution, and convolution processing, and then added element-wise to the input to obtain the output, which is used to perform the operation of the residual connection process based on nonlinear processing. The operation of the residual connection process based on nonlinear processing is specifically as follows: The input (the output of the residual connection process based on depthwise separable convolution with the same number of times) is subjected to convolution processing to obtain a convolution input; the convolution input is subjected to nonlinear processing (which can be implemented by the sigmoid function), then multiplied element-wise with the convolution input, and then added element-wise to the input (the output of the residual connection process based on depthwise separable convolution with the same number of times) to obtain the output, which is used to perform the next residual connection process based on depthwise separable convolution operation or serve as the filtered SAR residual connection image. The number of times of the above residual connection process based on depthwise separable convolution and the residual connection process based on nonlinear processing is preferably 4 times.

[0028] The operation of preprocessing the optical image is as follows: The optical image is processed by fast Fourier transform to map the spatial domain of the optical image to the frequency domain for separating different frequency components. Then, wavelet enhancement processing is performed in the frequency domain to specifically enhance the middle and high frequency components, thereby highlighting the texture details and edge structures in the image and obtaining an optically texture edge-enhanced image. The optically texture edge-enhanced image is processed by inverse Fourier transform to restore the enhanced frequency feature map to the spatial domain. After obtaining an enhanced image with richer detail information, it is added element-wise to the optical image to achieve the complementarity between the frequency-domain enhanced information and the original optical image, obtaining an optically enhanced image, effectively alleviating the problems of edge blurring and detail loss caused by cloud occlusion, and further improving the quality of subsequent cloud removal and image reconstruction.

[0029] S2. Perform feature fusion processing on the SAR enhanced image and the optical enhanced image based on the mutual attention mechanism to obtain a feature fusion image; the feature fusion image is processed by the channel and spatial coupled attention mechanism to obtain a semantically complementary feature image.

[0030] Perform feature fusion processing on the penetration ability of the SAR enhanced image and the spectral information of the optical enhanced image based on the mutual attention mechanism, solve the problems of image blurring and information loss in the optical remote sensing image under cloud occlusion, obtain a feature fusion image, and process the feature fusion image by the channel and spatial coupled attention mechanism to enhance the complementarity between the SAR image and the optical image, thereby accurately fusing multi-source information, effectively improving the structural consistency and clarity of the image after cloud removal, and obtaining a semantically complementary feature image.

[0031] First, perform feature fusion processing on the SAR enhanced image and the optical enhanced image based on the mutual attention mechanism, and at the same time introduce the prior information of the optical image and the SAR image to enhance the features in the cloud-contaminated image, obtaining a feature fusion image.

[0032] The operation of the feature fusion processing based on the mutual attention mechanism is as follows: The SAR enhanced image is convolved and then element-wise multiplied with the optical enhanced image, and then added element-wise to the optical enhanced image to obtain an initial fusion image; the query feature of the initial fusion image, the value feature and the key feature of the SAR enhanced image are respectively convolved and then processed by the attention mechanism to obtain an attention fusion image; the attention fusion image is globally pooled, fully connected, and feature reshaped (which can be achieved by the reshape function), and then element-wise multiplied with the attention fusion image to obtain an attention fusion enhanced image; the SAR enhanced image is globally pooled, fully connected, and feature reshaped (which can be achieved by the reshape function), and then element-wise multiplied with the SAR enhanced image to obtain a SAR enhanced feature image; the SAR enhanced feature image and the attention fusion enhanced image are added element-wise to obtain a feature fusion image.

[0033] Then, the feature fusion image is processed by a channel and spatial coupled attention mechanism to obtain a semantically complementary feature image. The operations of the channel and spatial coupled attention mechanism are as follows: Obtain the pixel gradient feature map of the feature fusion image (which can be achieved through edge detection processing). After (two-dimensional) convolution processing, perform spatial attention processing and channel attention processing respectively. After element-wise multiplication, enhance the recognition ability of the image target edge and cloud structure, improve the perception ability of the target edge and structural details in complex scenes, and obtain the attention mechanism feature image; Extract the context information of the attention mechanism feature image to enhance the understanding of the ground object boundary and cloud morphology. After pyramid processing by the multi-modal dynamic mechanism, enhance the context and detail representation, and enhance the understanding of the ground object boundary and cloud morphology to obtain the semantically complementary feature image.

[0034] The above-mentioned extraction of the context information of the attention mechanism feature image can be achieved by performing convolution processing on the attention mechanism feature image at different scales and then performing fusion in the channel dimension.

[0035] The operations of the above-mentioned pyramid processing of the multi-modal dynamic mechanism are as follows: The context information of the attention mechanism feature image is processed by convolution to obtain the context convolution feature; The context convolution feature is upsampled and added element-wise to the context convolution feature to obtain the initial semantically complementary feature; The initial semantically complementary feature is upsampled and added element-wise to the context convolution feature to obtain the semantically complementary feature image.

[0036] S3. The semantically complementary feature image is subjected to several structure reconstruction processes to obtain a structure reconstruction feature image; The structure reconstruction feature image is processed by convolution to obtain a cloud noise removal reconstruction image.

[0037] Perform several structure reconstruction processes on the semantically complementary feature image. During the structure reconstruction process, fuse the semantically complementary convolution features reflecting the underlying texture edge information of the image, the semantically complementary attention features reflecting the key region information of the image channel and spatial dimensions, and the semantically complementary perception features reflecting the high-level semantic abstract information of the image to strengthen the structural details; And use the gradient information to provide a prior structure constraint for the reconstruction, which can retain both high-frequency details and low-frequency semantics during the reconstruction, avoid the problems of "over-smoothing" or "semantic disorder", improve the image reconstruction quality, and obtain the structure reconstruction feature image; The structure reconstruction feature image is processed by convolution to obtain a cloud noise removal reconstruction image.

[0038] Taking the first structure reconstruction process as an example, the operation steps of the structure reconstruction process are as follows.

[0039] Step 1: The semantic complementary feature image is processed by convolution, ReLU activation function, and convolution to obtain a semantic complementary convolutional feature image; the semantic complementary feature image is processed by a channel and spatial coupling attention mechanism to obtain a semantic complementary attention feature image; the semantic complementary feature image is processed by a multi-layer perceptron to obtain a semantic complementary perceptual feature image.

[0040] Step 2: The semantic complementary perceptual feature image and the semantic complementary feature image are multiplied element-wise to obtain a first semantic complementary fusion image; the semantic complementary convolutional feature image and the semantic complementary attention feature image are added element-wise to obtain a second semantic complementary fusion image; the first semantic complementary fusion image and the second semantic complementary fusion image are added element-wise to obtain an initial structure reconstruction image.

[0041] Step 3: The structural gradient information of the initial structure reconstruction image (which can be obtained by edge detection of the initial structure reconstruction image) is fused with the optical reference image, and after convolution processing, it is added element-wise to the initial structure reconstruction image to obtain a first structural reconstruction feature image; it is used to perform the second structural reconstruction process.

[0042] The number of times of the above structural reconstruction process is preferably 4 times.

[0043] To verify the removal effect of the method in this embodiment, the following experiments were conducted.

[0044] In the experiment, the method in this embodiment (hereinafter referred to as the MSGRF-CR method) was compared and evaluated with 3 most advanced methods in the prior art (including the Dsen2-CR method, the GLF-CR method, and the HS 2 P method). The comparison results are shown in Table 1. It can be clearly seen from Table 1 that the peak signal-to-noise ratio (hereinafter referred to as PSNR) and mean square error (MSE) of the Dsen2-CR method are the lowest, while the structural similarity index (SSIM) and spectral angle mapper (SAM) are only slightly higher than those of the GLF-CR model, and the SSIM and SAM of the latter are the lowest. On the contrary, the MSE performance of the GLF-CR is the best, and the PSNR value is second only to the Dsen2-CR model. The HS 2 P model ranks second in terms of PSNR, SSIM, and SAM, but ranks third in terms of MSE. In contrast, the method in this embodiment (the MSGRF-CR method) has always been superior to the existing 3 advanced methods in all four indicators, indicating the excellent cloud removal performance of the method in this embodiment (the MSGRF-CR method).

[0045] Table 1 Summary of the reconstruction effects of different reconstruction methods 。

[0046] Meanwhile, in the experiment, the qualitative visual processing capabilities of the method of this embodiment (MSGRF-CR method) and three existing methods under three types of cloud cover (large scale, medium scale, and small scale) were also compared. See Figure 2 . Figure 2 In the middle figure, from top to bottom are the cloudy optical image, the SAR image, the cloudless optical image, and the result images processed by the DSen2-CR method, the GLF-CR method, the HS 2 P method, and the method of this embodiment (MSGRF-CR method). The size of each image is 128×128; Figure 2 When counting from left to right in the middle figure, the first and second columns are different scenes under heavy cloud cover (large-scale cloud cover), with a cloud cover rate exceeding 80%. The third and fourth columns are different scenes under medium cloud cover rate (medium-scale cloud cover), with the coverage range between 40% and 80%. The third and fourth columns are different scenes under low cloud cover rate (small-scale cloud cover), with the cloud-covered area less than 40%. From Figure 2 it can be seen that the GLF-CR method produces the worst visual effect, and the generated image is too blurred to restore clear structures, accurate colors, or textures, especially in the case of large-area cloud cover. In contrast, the Dsen2-CR method has improved in the restoration of color, texture, and boundaries compared to the GLF-CR model, but it still shows obvious color and texture distortion, obvious artifacts, and blurred boundaries. When the cloud cover range increases from the smallest to medium and large scales, the image quality deteriorates severely. HS 2 P method provides a more faithful restoration of structures, details, and textures, and its performance is better than that of the GLF-CR method and the Dsen2-CR method; however, the HS 2 P method's restoration of structures is only partial and accompanied by considerable noise. In contrast, the method of this embodiment (MSGRF-CR method) produces results that are closest to the cloudless image, can effectively outline clear object boundaries and retain high-fidelity details. In the case of medium cloud cover, it can accurately reconstruct the cloud-obscured areas and provide rich details; in the case of minimal cloud cover, it can retain texture and edge information, thus showing higher clarity. Even in the case of large-area cloud cover, the method of this embodiment (MSGRF-CR method) can generate images closest to cloudless conditions, performing better than other models, but it is still necessary to further improve the structural details and boundary refinement.

[0047] Next, to evaluate the performance of the method in this embodiment (MSGRF-CR method) and three existing methods under various cloud coverage conditions, in the experiment, datasets in five intervals with cloud coverage rates of 0%-20%, 20%-40%, 40%-60%, 60%-80%, and 80%-100% were also used to evaluate each model. These cloud coverage rates range from light cloud coverage to heavy cloud coverage, which helps to systematically evaluate the robustness and accuracy of each method. The experimental evaluation uses quantitative metrics - MSE, PSNR, SSIM, and SAM to measure the effectiveness of each method in dealing with different cloud densities. See the experimental results in Figure 3 . In Figure 3 , but when the cloud cover rate is lower than 60%, the method in this embodiment (MSGRF-CR method) is superior to all other models. At the same time, in terms of the SAM metric, the method in this embodiment (MSGRF-CR method) is superior to the three existing methods under all conditions. That is to say, the overall performance of the method in this embodiment (MSGRF-CR method) is the best.

[0048] In addition, to further evaluate the effectiveness of the method in this embodiment (MSGRF-CR method) in reconstructing vegetation information under cloud cover, 720 representative multi-cloud remote sensing images and their corresponding cloud-free prediction images were randomly selected in the experiment, and the normalized difference vegetation index (NDVI) was calculated using four methods respectively. The coefficient of determination (R²) was calculated to quantify the consistency of the normalized difference vegetation index between the original image and the predicted image, and the R² value distributions of the four models were statistically analyzed and compared. In addition, to visually display the vegetation reconstruction performance of the method, three representative test images with the lowest, median, and highest R² values were randomly selected in the experiment, and NDVI scatter plots with fitted regression lines were generated to visually illustrate the correlation between the ground truth and the predicted NDVI of the four methods, as shown in Figure 4 . Figure 4 In, the blue area represents the interquartile range, the middle red line represents the median, and the black line represents the range. It can be clearly seen from Figure 4 that the method in this embodiment (MSGRF-CR method) has the highest median R² and the interquartile range is reduced, with the best performance, reflecting its excellent accuracy and stability.

[0049] Finally, to evaluate the robustness and generalization ability of the method in this embodiment (the MSGRF-CR method), the dataset provided by Ebel et al. (2022c) was used in the experiment. Non-overlapping 256 × 256 pixel image patches were recombined into panoramic images. On this basis, experiments on large-scale scene images (4096 × 3072 pixels) were designed in the experiment, and a sliding window preprocessing method was adopted to divide the original image into 1024 × 1024 pixel sub-image patches with an overlap of 128 pixels between adjacent image patches, resulting in approximately 425 sub-images for each scene. Each method independently processes each sub-image patch, and then reconstructs the complete scene output by seamlessly merging the processed image patches. The experimental results of large-scale cloud reconstruction are as Figure 5 shown, Figure 5 From top to bottom in [Figure], there are cloudy optical images, SAR images, cloud-free optical images, the processing result map of the MSGRF-CR method, a locally enlarged area map of the cloudy optical image (the enlarged area is from the cloudy optical image, and the image name is simplified to cloudy optical image (enlarged version)), a locally enlarged area map of the SAR image (the enlarged area is from the SAR image, and the image name is simplified to SAR image (enlarged version)), a locally enlarged area map of the cloud-free optical image (the enlarged area is from the cloud-free optical image, and the image name is simplified to cloud-free optical image (enlarged version)), and a locally enlarged area map of the processing result of the MSGRF-CR method (the enlarged area is from the processing result map of the MSGRF-CR method, and the image name is simplified to MSGRF-CR method (enlarged version)). Among them, the locally enlarged area maps of the cloudy optical image, the SAR image, the cloud-free optical image, and the processing result of the MSGRF-CR method all contain 8 small images. Every 2 adjacent small images form a set of scene images, and the scenes are mountain, plain, forest, and town, respectively. Moreover, the positions of the terrain enlarged areas correspond to the yellow and red framed areas of each scene in the processing result map of the MSGRF-CR method. As can be seen from Figure 5 this, the method in this embodiment (the MSGRF-CR method) has excellent overall cloud removal performance in large-scale cloud-covered images. Compared with the original cloud-contaminated image ( Figure 5 the cloudy optical image in [Figure]), the method in this embodiment (the MSGRF-CR method) can effectively remove clouds and restore the ground surface with clear ground structure and boundary details ( Figure 5 the processing result map of the MSGRF-CR method in [Figure]). In addition, Figure 5 the enlarged details of the areas marked by the red and yellow frames in the processing result map of the MSGRF-CR method can be seen in Figure 5Local enlarged area map of the processing result of the MSGRF-CR method, that is, the MSGRF-CR method (enlarged version) map. These enlarged detailed results show that the size of the finally reassembled image is 4096 × 3072, with high global consistency. Although the image is processed in 1024 × 1024 patches, no obvious stitching artifacts or boundary discontinuities are observed. In addition, in areas with dense clouds, the method of this embodiment (MSGRF-CR method) can also accurately restore surface details such as ground texture, vegetation cover, and urban boundaries.

[0050] This embodiment also provides an optical remote sensing cloud and cloud shadow removal system for implementing the above-mentioned optical remote sensing cloud and cloud shadow removal method, including: An image preprocessing module for obtaining an optical image and a SAR image of the area to be processed; the SAR image is processed by dynamic logarithmic compression to obtain a dynamically compressed SAR image; the dynamically compressed SAR image is processed by adaptive non-local mean filtering to obtain a filtered SAR image; the filtered SAR image is processed by deep denoising and structure restoration to obtain a SAR enhanced image; the optical image is processed by fast Fourier transform to map the optical image spatial domain to the frequency domain, and wavelet enhancement processing is performed in the frequency domain to obtain an optically texture edge enhanced image; after the optically texture edge enhanced image is processed by inverse Fourier transform, it is added element by element to the optical image to obtain an optically enhanced image; A semantic complementary feature image generation module for performing feature fusion processing on the SAR enhanced image and the optically enhanced image based on a mutual attention mechanism to obtain a feature fusion image; the feature fusion image is processed by a channel and spatial coupled attention mechanism to obtain a semantic complementary feature image; the operation of the channel and spatial coupled attention mechanism is: obtaining the pixel gradient feature map of the feature fusion image, after convolution processing, performing spatial attention processing and channel attention processing respectively, and multiplying element by element to obtain an attention mechanism feature image; extracting the context information of the attention mechanism feature image, and performing pyramid processing by a multi-modal dynamic mechanism to obtain a semantic complementary feature image; A cloud noise removal and reconstruction image generation module for performing several structure reconstruction processes on the semantic complementary feature image to obtain a structure reconstruction feature image; the structure reconstruction feature image is processed by convolution to obtain a cloud noise removal and reconstruction image.

[0051] This embodiment also provides an optical remote sensing cloud and cloud shadow removal device, including a processor and a memory. When the processor executes the computer program stored in the memory, the above-mentioned optical remote sensing cloud and cloud shadow removal method is implemented.

[0052] This embodiment also provides a computer-readable storage medium for storing a computer program. When the computer program is executed by a processor, the above-mentioned optical remote sensing cloud and cloud shadow removal method is implemented.

[0053] An optical remote sensing cloud and cloud shadow removal method provided in this embodiment. First, the optical image and the SAR image of the area to be processed are preprocessed respectively to solve the problems of high dynamic range and multiplicative speckle noise existing in the SAR image, and enhance the high-frequency information in the optical image. Then, the penetration ability of the SAR enhanced image and the spectral information of the optical enhanced image are subjected to feature fusion processing based on the mutual attention mechanism, solving the problems of image blurring and information loss in the optical remote sensing image under cloud occlusion, obtaining a feature fusion image, and performing channel and spatial coupled attention mechanism processing on the feature fusion image, enhancing the complementarity between the SAR image and the optical image, thereby accurately fusing multi-source information, effectively improving the structural consistency and clarity of the image after cloud removal, and obtaining a semantic complementary feature image. Finally, the semantic complementary feature image is subjected to several structure reconstruction processes, and during the structure reconstruction process, the semantic complementary convolution features reflecting the underlying texture edge information of the image, the semantic complementary attention features reflecting the key region information of the image channel and spatial dimensions, and the semantic complementary perception features reflecting the high-level semantic abstract information of the image are fused to strengthen the structural details, and the gradient information is used to provide a prior structure constraint for the reconstruction, which can retain both high-frequency details and low-frequency semantics during the reconstruction, improve the image reconstruction quality, and obtain a structure reconstruction feature image. The structure reconstruction feature image is subjected to convolution processing to obtain a cloud noise removal reconstruction image. When applied in the field of image cloud and cloud shadow removal, it has a good removal effect and a high image restoration quality.

Claims

1. An optical remote sensing cloud and cloud shadow removal method, characterized in that Including the following operations: S1. Obtain the optical image and SAR image of the area to be processed; the SAR image is subjected to dynamic logarithmic compression processing to obtain a dynamically compressed SAR image; the dynamically compressed SAR image is subjected to adaptive non-local mean filtering processing to obtain a filtered SAR image; the filtered SAR image is subjected to depth denoising and structure restoration processing to obtain a SAR enhanced image; The optical image is subjected to fast Fourier transform processing to map the optical image spatial domain to the frequency domain, and wavelet enhancement processing is performed in the frequency domain to obtain an optically textured edge enhanced image; after the optically textured edge enhanced image is subjected to inverse Fourier transform processing, it is added element by element to the optical image to obtain an optically enhanced image; S2. Perform feature fusion processing on the SAR enhanced image and the optically enhanced image based on a mutual attention mechanism to obtain a feature fusion image; The feature fusion image is processed by a channel and spatial coupled attention mechanism to obtain a semantically complementary feature image; The operation of the channel and spatial coupled attention mechanism is as follows: obtain the pixel gradient feature map of the feature fusion image, after convolution processing, perform spatial attention processing and channel attention processing respectively, and after element-by-element multiplication, obtain an attention mechanism feature image; extract the context information of the attention mechanism feature image, and after pyramid processing by a multi-modal dynamic mechanism, obtain a semantically complementary feature image; S3. The semantically complementary feature image is subjected to several structure reconstruction processes to obtain a structure reconstruction feature image; the structure reconstruction feature image is subjected to convolution processing to obtain a cloud noise removal and reconstruction image.

2. The optical remote sensing cloud and cloud shadow removal method according to claim 1, characterized in that In S1, the operation of depth denoising and structure restoration processing is as follows: The filtered SAR image is subjected to convolution processing to obtain a filtered SAR convolution image; the filtered SAR convolution image is subjected to several residual connection processes based on depthwise separable convolution and residual connection processes based on nonlinear processing to obtain a filtered SAR residual connection image; the filtered SAR residual connection image is subjected to convolution processing to obtain a SAR enhanced image.

3. The optical remote sensing cloud and cloud shadow removal method according to claim 1, characterized in that In S2, the operation of feature fusion processing based on a mutual attention mechanism is as follows: The SAR enhanced image is subjected to convolution processing and then multiplied element by element with the optically enhanced image, and then added element by element to the optically enhanced image to obtain an initial fusion image; the query feature of the initial fusion image, the value feature and key feature of the SAR enhanced image are respectively subjected to convolution processing and then subjected to attention mechanism processing to obtain an attention fusion image; the attention fusion image is subjected to global pooling, fully connected and feature reshaping, and then multiplied element by element with the attention fusion image to obtain an attention fusion enhanced image; the SAR enhanced image is subjected to global pooling, fully connected and feature reshaping, and then multiplied element by element with the SAR enhanced image to obtain a SAR enhanced feature image; the SAR enhanced feature image and the attention fusion enhanced image are added element by element to obtain a feature fusion image.

4. The optical remote sensing cloud and cloud shadow removal method according to claim 1, wherein In S2, the operation of pyramid processing by a multi-modal dynamic mechanism is as follows: The context information of the attention mechanism feature image is subjected to convolution processing to obtain context convolution features; the context convolution features are upsampled and element-wise added to the context convolution features to obtain initial semantic complementary features; the initial semantic complementary features are upsampled and element-wise added to the context convolution features to obtain a semantic complementary feature image.

5. The optical remote sensing cloud and cloud shadow removal method according to claim 1, characterized in that In S3, the operation of subjecting the semantic complementary feature image to structure reconstruction processing is as follows: The semantic complementary feature image is subjected to convolution, ReLU activation function, and convolution processing to obtain a semantic complementary convolution feature image; the semantic complementary feature image is subjected to a channel and spatial coupled attention mechanism processing to obtain a semantic complementary attention feature image; the semantic complementary feature image is subjected to multi-layer perceptron processing to obtain a semantic complementary perception feature image; The semantic complementary perception feature image and the semantic complementary feature image are element-wise multiplied to obtain a first semantic complementary fusion image; the semantic complementary convolution feature image and the semantic complementary attention feature image are element-wise added to obtain a second semantic complementary fusion image; the first semantic complementary fusion image and the second semantic complementary fusion image are element-wise added to obtain an initial structure reconstruction image; The structure gradient information of the initial structure reconstruction image is fused with the optical reference image, and after convolution processing, it is element-wise added to the initial structure reconstruction image to obtain a first structure reconstruction feature image for performing the operation of the second structure reconstruction processing.

6. The optical remote sensing cloud and cloud shadow removal method according to claim 1, characterized in that The dynamic compression of the SAR image in S1 is obtained based on the maximum brightness, standard deviation, and image information entropy of the SAR normalized image; the SAR normalized image is obtained by normalizing the SAR image.

7. The optical remote sensing cloud and cloud shadow removal method according to claim 1, characterized in that The smoothing parameter in the adaptive non-local mean filtering process in S1 is obtained based on the pixel information at each position in the dynamic compression SAR image.

8. An optical remote sensing cloud and cloud shadow removal system for implementing the optical remote sensing cloud and cloud shadow removal method according to claim 1, characterized in that, Including: An image preprocessing module for obtaining an optical image and an SAR image of the area to be processed; The SAR image is subjected to dynamic logarithmic compression processing to obtain a dynamic compression SAR image; the dynamic compression SAR image is subjected to adaptive non-local mean filtering processing to obtain a filtered SAR image; the filtered SAR image is subjected to depth denoising and structure restoration processing to obtain an SAR enhanced image; the optical image is subjected to fast Fourier transform processing to map the optical image from the spatial domain to the frequency domain, and wavelet enhancement processing is performed in the frequency domain to obtain an optically textured edge enhanced image; after the optically textured edge enhanced image is subjected to inverse Fourier transform processing, it is element-wise added to the optical image to obtain an optical enhanced image; A semantic complementary feature image generation module for performing feature fusion processing based on a mutual attention mechanism on the SAR enhanced image and the optical enhanced image to obtain a feature fusion image; The feature fusion image is subjected to a channel and spatial coupled attention mechanism processing to obtain a semantic complementary feature image; The operation of the channel and spatial coupled attention mechanism processing is as follows: obtaining the pixel gradient feature map of the feature fusion image, subjecting it to convolution processing, then performing spatial attention processing and channel attention processing respectively, and element-wise multiplying to obtain an attention mechanism feature image; extracting the context information of the attention mechanism feature image and subjecting it to pyramid processing of a multi-modal dynamic mechanism to obtain a semantic complementary feature image; The cloud noise removal and reconstruction image generation module is used to obtain a structure reconstruction feature image through several structural reconstruction processes on the semantic complementary feature image; the structure reconstruction feature image is processed through convolution to obtain a cloud noise removal and reconstruction image.

9. An optical remote sensing cloud and cloud shadow removal device, characterized in that, It includes a processor and a memory. Among them, when the processor executes the computer program stored in the memory, it implements the optical remote sensing cloud and cloud shadow removal method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, It is used to store a computer program. Among them, when the computer program is executed by a processor, it implements the optical remote sensing cloud and cloud shadow removal method according to any one of claims 1-7.

Citation Information

Patent Citations

  • SAR-assisted optical remote sensing image restoration method

    CN116309150A

  • Multi-modal remote sensing image cloud region and shadow removal method and system

    CN116977202A

  • Progressive repair frame cloud removal method for fusion of optical remote sensing image and SAR (Synthetic Aperture Radar) image

    CN117058059A

  • Image cloud removal method and system based on optical remote sensing image and SAR image

    CN117522738A

  • Progressive double-decoupling SAR (Synthetic Aperture Radar)-assisted remote sensing image thick cloud removal method

    CN117689579A

Cited By

  • Visible light image and infrared image fusion method and device, medium and equipment

    CN120782668A

  • Continuous shooting image super-resolution method based on discrete wavelet transform fusion

    CN121235906A