Optical remote sensing image cloud removal method based on iterative generative adversarial network

By constructing shared and private feature extraction modules through iterative generative adversarial networks (IGANs), and combining them with dual-attention feature fusion and reconstruction modules, the difficult problem of cloud and shadow removal in optical remote sensing images is solved, high-quality cloud-free images are generated, and the application effect of remote sensing images is improved.

CN120808192APending Publication Date: 2025-10-17NANJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510941522.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing methods are unable to effectively fuse the shared and private features of optical and synthetic aperture radar images, resulting in poor cloud and shadow removal in optical remote sensing images, especially insufficient reconstruction quality in areas covered by thick clouds.

Method used

A method based on iterative generative adversarial networks (IGAN) is adopted to optimize the feature extraction and fusion processes by constructing a shared feature extraction module, a progressive dual-attention feature fusion module, and a feature reconstruction module, combined with convolutional sparse coding and a deep iterative network, to generate high-quality cloud-free images.

Benefits of technology

It effectively alleviates the problem of optical images being covered by clouds in complex environments, generates high-quality cloud-free optical images, and improves the availability of remote sensing images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808192A_ABST
    Figure CN120808192A_ABST
Patent Text Reader

Abstract

The invention discloses an optical remote sensing image cloud removal method based on an iterative generative adversarial network, and belongs to the field of remote sensing image processing. Respectively inputting the optical image and the SAR image into a shared feature extraction module and a private feature extraction module to obtain shared features of the optical image and the SAR image, private features of the optical image and private features of the SAR image; sequentially fusing the shared features and private features of the optical image and the SAR image by using a progressive double-attention feature fusion module to perform feature reconstruction of the cloudless image; a feature reconstruction module is used for recovering a fusion result into a high-quality cloudless image; and the cloud-free image generated by the generator network and the real cloud-free image are superposed and input into the discriminator network, the quality of the generated cloud-free image is judged, and a discrimination result is fed back to the generator network. The problems that an existing deep learning method cannot fuse optical and SAR image shared features and private features at the same time, and the feature extraction process lacks interpretability are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of remote sensing image processing, and particularly relates to an optical remote sensing image cloud removal method based on an iterative generative adversarial network. BACKGROUND

[0002] With the continuous improvement of remote sensing detection technology, massive remote sensing data has increasingly become the main means of earth monitoring tasks. At present, various types of remote sensing images have appeared, and optical remote sensing images are indispensable in many applications such as geographic resource investigation, ecological environment monitoring, and meteorological disaster prediction due to their good visual effect and sufficient spectral information. However, data from Landsat ETM+ shows that clouds have obscured 66% of the earth's surface and covered about 35% of the land area. The imaging mode of optical remote sensing satellites inevitably leads to pollution of clouds and shadows, reducing the availability of these images, especially for earth observation tasks requiring continuous optical remote sensing time series images. Therefore, reconstructing high-quality cloud-free images covered by clouds and shadows in optical remote sensing has become a major challenge.

[0003] The challenge of optical image cloud and shadow removal methods based on optical and Synthetic Aperture Radar (SAR) image fusion is how to accurately reconstruct the spectral spatial information of cloud-covered areas using auxiliary information. Various methods have been developed to improve the performance of reconstructed images, which can be roughly classified into three categories according to the type of auxiliary information they use: single-image restoration-based methods, multi-temporal-based methods, and multi-sensor-based methods. Single-image restoration-based methods usually assume that cloud-polluted and cloud-free areas in the same image have similar geographical features. Based on this assumption, single-image restoration-based methods can use the spectral information of cloud-free areas to reconstruct cloud-covered areas. In addition, some single-image restoration-based methods use pixel correction techniques to handle places where surface information is not completely obscured in thin cloud and fog images, while removing thin clouds and fog. However, these methods often cannot effectively handle multi-cloud images where thick cloud coverage occupies a large proportion. To solve the problem of thick cloud coverage, some researchers have proposed multi-temporal methods that use cloud-free data from other time periods to replace the pixels of cloud-covered areas to reconstruct the spatio-temporal spectral information of thick cloud-covered areas. However, multi-temporal methods are usually limited by rapid topographic changes in a short period of time and long-term multi-cloud weather, and the computational overhead of this method is usually quite large. Multi-sensor-based methods aim to use complementary information from other satellite sensors to predict cloud-occluded areas. SAR is an active detection imaging system that captures detailed ground information by emitting electromagnetic pulses and receiving reflected signals. This function enables SAR to provide all-weather ground monitoring, even under continuous cloud cover. Therefore, SAR images can provide valuable information texture for cloud-covered areas. In the past few years, many cloud removal methods have been proposed that use SAR images as auxiliary data for optical remote sensing images. These methods mainly adopt two fusion strategies. The first strategy directly stacks the cloud and SAR images as network inputs to extract features, which ignores the special information of each modality. The second strategy extracts features from multi-cloud and SAR images separately, and then uses some specific techniques to fuse them, which ignores the shared information between different modalities. However, none of the existing methods considers both fusion strategies simultaneously, resulting in insufficient fusion of shared and private features affecting the performance of the method. In addition, their feature extraction process is usually implicit, lacking interpretability. SUMMARY

[0004] The application provides an optical remote sensing image cloud removal method based on an iterative generative adversarial network (IGAN), which is used for cloud and shadow removal in an optical remote sensing image, can improve the quality of cloud-free image reconstruction results, promotes wider application, and solves the problems that existing deep learning methods cannot simultaneously fuse shared features and private features of optical and SAR images and the feature extraction process lacks interpretability.

[0005] The application embodiment provides an optical remote sensing image cloud removal method based on an iterative generative adversarial network, which comprises the following steps:

[0006] Step 1, a shared feature extraction module and a private feature extraction module in a generator network are constructed, an optical image and a SAR image are respectively input into the shared feature extraction module and the private feature extraction module, shared features of the optical image and the SAR image, private features of the optical image and private features of the SAR image are obtained;

[0007] Step 2, a progressive double attention feature fusion module in the generator network is constructed, and the progressive double attention feature fusion module is used to sequentially fuse the shared features and the private features of the optical image and the SAR image to perform feature reconstruction of a cloud-free image;

[0008] Step 3, a feature reconstruction module in the generator network is constructed, and the feature reconstruction module is used to restore the fusion result to a high-quality cloud-free image;

[0009] Step 4, a discriminator network is constructed, the generated cloud-free image of the generator network is superimposed with a real cloud-free image and input into the discriminator network, the quality of the generated cloud-free image is judged, and the discrimination result is fed back to the generator network.

[0010] Optionally, in an embodiment of the application, step 1 specifically comprises:

[0011] An observation model based on convolutional sparse coding is established to extract shared features and private features from different modalities, and the observation model is expressed as follows:

[0012]

[0013] wherein, I o is a cloudy optical image, I s is a SAR image, is a multiplication operation, and are shared feature filters of I o and I s s i represents I o and I sshared feature, and I o and I s private feature filter of P i and V i I o and I s private feature of i, i is the number of filters;

[0014] According to the observation model, under the condition of knowing all filters, the cloud-free image F is obtained by optimizing equation (3):

[0015]

[0016] where P is the private feature of the optical image, V is the private feature of the SAR image, S is the shared feature, F norm, λ s , λ p , λ v are balance parameters, and R(.) is a regularization term.

[0017] Equation (3) is divided into three sub-problems, and each variable is updated alternately while fixing the other variables. The ISTA algorithm is used to solve the following three sub-problems:

[0018]

[0019] First, update P to solve the quadratic approximation of equation (4) to update the private feature P, which is expressed as follows:

[0020]

[0021] where P t-1 represents the result after the (t-1)th iteration, η p is the iteration step for updating P, and Δg(P t-1 ) is the derivative result of P.

[0022]

[0023] The optimal solution of equation (7) is defined by a proximal operator, as follows:

[0024] Δg(P t-1 )=X p *(X s *S t-1 +X p *P t-1 I o ) (9)

[0025]

[0026] where * is the transpose convolution operation, X s and X p denote the filters involved in the convolution operation, respectively and denote the filters involved in the convolution operation, respectively p is the filter, S t-1 is the (t-1)th updated S, P t-1 is the (t-1)th updated P, P t is the tth updated P, is the proximal operator depending on the regularization term;

[0027] The same method is used to update V, and a quadratic approximation of equation (5) is solved to update the private feature P, which is denoted as follows:

[0028]

[0029] where V t-1 denotes the result after (t-1) iterations, η v is the iteration step size for updating V, Δg(V t-1 ) is the derivative result of V;

[0030]

[0031] The optimal solution of equation (7) is defined by a proximal operator, which is denoted as follows:

[0032] Δg(V t-1 ) = Y v *(Y s *S t-1 + Y v *V t-1 -I s ) (13)

[0033]

[0034] where Y s and Y v denote the filters involved in the convolution operation, respectively and denote the filters involved in the convolution operation, respectively v is the filter, V t-1 is the (t-1)th updated V, V t is the tth updated V, is the proximal operator depending on the regularization term;

[0035] The same method is used to update S, and equation (6) is rewritten to obtain equation (15):

[0036]

[0037] where, when Further simplifying equation (15) gives equation (16):

[0038]

[0039] wherein, represents I s and I o superimposed, is a filter related to S, similarly, the same method as P and V is used to optimize the sub-problems, the optimization process is as follows:

[0040]

[0041] wherein, Δg(S t-1 ) is the derivative of S t-1 , η s is the iterative step length of updating S, L s is a filter, L s represents the filter involved in S, S t is the result of the tth update of S, is a proximal operator dependent on the regularization term;

[0042] Expanding equations (10), (14) and (19) into a deep iterative network, which contains three sub-networks to realize the prediction of P, V and S respectively, X s , X p , Y s , Y v , L s parameters are learned by convolutional layers, except that the channel number of L s is 128, the size of all convolution kernels is 3*3, and the channel number is 64, X p *, Y v *, L s * parameters are learned by transpose convolutional layers, and ResNet is used to implicitly learn prior information.

[0043] Optionally, in an embodiment of the present application, step 2 specifically comprises: the structure of the progressive double attention fusion module is composed of 2 cross attention blocks with the same structure and different parameters and 1 multi-head self-attention block, and the whole fusion process is represented by the following formula:

[0044] F=MSAB(CAB(CAB(P,S),V)) (20)

[0045] Wherein, CAB(.) and MSAB(.) represent cross attention block and multi-head self-attention block respectively, the target of cross attention block is to fuse the difference information from different branches, and the target of multi-head self-attention block is to improve the learning ability of global features to reconstruct the information of the area covered by thick clouds, the first cross attention block fuses the private feature P and the shared feature S of the optical image to obtain the intermediate feature F M , the private feature P is linearly reshaped into Q P , the shared feature S is linearly converted into K S , and the value V S , the matrix multiplication of Q P and K S transpose obtains the attention weight matrix, the feature similarity between Q P and K S is calculated, and the process is as follows:

[0046]

[0047] Wherein, d is a scaling factor;

[0048] The matrix decomposition method is used when calculating the attention weight to reduce the amount of calculation, then, the similar features between the private feature P and the shared feature S are removed, and the private feature P of the optical image is transferred to the intermediate feature F M , and the process is as follows:

[0049] F M = Linear(V S (1-Attention))+Q P (22)

[0050] F M =MLP(Norm(F M ))+F M (23)

[0051] Wherein, V S is the key value of S, Linear() is a linear transformation function, MLP(.) represents a multi-layer perceptron, Norm(.) represents a linear norm operation, and the fusion process of the intermediate feature F M and the private feature V in the second cross attention block is the same as that of the first cross attention block, and is defined as follows:

[0052] Q V ,K f ,V f =Linear(Reshape(V,F M ,F M ) (24)

[0053]

[0054] F = MLP(Norm(F)) + F (26)

[0055] where Q V is the query of V, K f is the key of the intermediate feature, V f is the value of the intermediate feature, Reshape() is a shape transformation operation, and F is the fusion result.

[0056] The self-attention mechanism captures the dependencies between pixels at any position and other positions through global perception. The multi-head self-attention block is integrated into the network, and the fusion result F of the cross-attention block is fed into the multi-head self-attention block to capture global information. The fusion result F is first put into a normalization layer to stabilize the distribution of the feature. Then, the attention map of the fusion result F is calculated through the multi-head self-attention mechanism. The multi-head self-attention mechanism calculates the attention result of each head respectively, and concentrates the attention weights of all heads. The process of the multi-head self-attention block is as follows:

[0057] Q, K, V = FW (27)

[0058] [Q1...Q i ], [K1...K i ], [V1...V i ] = Split(Q), Split(K), Split(V) (28)

[0059]

[0060] Multihead = concat(head1, head2...head i ) (30)

[0061] where Q, K, V are the three elements of the attention mechanism, Q i , K i and V i are the three elements of the attention mechanism of each head, W represents a deep convolution, and Split(.) represents a split function.

[0062] The gating mechanism G(.) is used to guide the attention result by a feedforward neural network, which consists of a kernel of 1*1 convolution layer and a GELU activation function. Then, the Hadamard product between G(F) and the attention result is performed to obtain the feature FG. Finally, the input F is added to the output using a skip connection to obtain the final result F fused :

[0063] F fused = MSAB(F) ⊙ G(F) + F (31).

[0064] Optionally, in an embodiment of the present application, the residual block is used as the backbone structure in the feature reconstruction module of the generator constructed in step 3, which is composed of five residual blocks and a convolutional layer, each residual block contains three convolutional layers and a skip connection, and the first two convolutional layers are accompanied by a rectified linear unit, the size of the convolution kernel in each residual block is 3*3, the number of input channels and output channels is 64, and the final convolutional layer has a convolution kernel size of 3*3, has 64 input channels and 13 output channels to match the spectral dimension of the generated cloud-free image.

[0065] Optionally, in an embodiment of the present application, in step 4, the structure of the discriminator is composed of five blocks with the same structure and a convolutional layer, each block contains a convolutional layer with a convolution kernel of 4*4, a normalization layer and an activation function Leaky ReLu, wherein the stride of the convolutional layer is 2, the convolution kernel size of the last convolutional layer is 3, and the stride is 1, and the last convolutional layer outputs one-dimensional data to represent the result of the discriminator.

[0066] The optical remote sensing image cloud removal method based on the iterative generative adversarial network has the following beneficial effects: effectively alleviating the problem of cloud coverage of optical images in complex environments, and generating high-quality cloud-free optical images.

[0067] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS

[0068] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the following description, taken in conjunction with the accompanying drawings, in which:

[0069] Figure 1 A flowchart of an optical remote sensing image cloud removal method based on an iterative generative adversarial network according to an embodiment of the present application is provided;

[0070] Figure 2 A generator structure diagram based on an iterative generative adversarial network is provided;

[0071] Figure 3 A discriminator structure diagram based on a generative adversarial network is provided;

[0072] Fig. 4(a) is a cloud-covered optical image of the SEN12MS-CR dataset;

[0073] Fig. 4(b) is a SAR image of the SEN12MS-CR dataset;

[0074] Fig. 4(c) is a cloud-covered optical image of the SMILE-CR dataset;

[0075] Fig. 4(d) is a SAR image of the SMILE-CR dataset;

[0076] Fig. 5(a) is a cloud-removed result of the SEN12MS-CR satellite dataset based on the Pix2pix method;

[0077] Fig. 5(b) is a cloud-removed result of the SEN12MS-CR satellite dataset based on the McGAN method;

[0078] Fig. 5(c) is a cloud-removed result of the SEN12MS-CR satellite dataset based on the SAR-opt-cGAN method;

[0079] Fig. 5(d) is a cloud-removed result of the SEN12MS-CR satellite dataset based on the Dsen2-CR method;

[0080] Fig. 5(e) is a cloud-removed result of the SEN12MS-CR satellite dataset based on the GLF-CR method;

[0081] Fig. 5(f) is a cloud-removed result of the SEN12MS-CR satellite dataset based on the Align-CR method;

[0082] Fig. 5(g) is a cloud-removed result of the SEN12MS-CR satellite dataset based on the IGAN method;

[0083] Fig. 6(a) is a cloud-removed result of the SMILE-CR satellite dataset based on the Pix2pix method;

[0084] Fig. 6(b) is a cloud-removed result of the SMILE-CR satellite dataset based on the McGAN method;

[0085] Fig. 6(c) is a cloud-removed result of the SMILE-CR satellite dataset based on the SAR-opt-cGAN method;

[0086] Fig. 6(d) is a cloud-removed result of the SMILE-CR satellite dataset based on the Dsen2-CR method;

[0087] Fig. 6(e) is a cloud-removed result of the SMILE-CR satellite dataset based on the GLF-CR method;

[0088] Fig. 6(f) is a cloud-removed result of the SMILE-CR satellite dataset based on the Align-CR method;

[0089] Fig. 6(g) is a cloud-removed result of the SMILE-CR satellite dataset based on the IGAN method. DETAILED DESCRIPTION

[0090] Embodiments of the present application are described below in detail with reference to the accompanying drawings, wherein the same or similar components are denoted by the same or similar reference numerals throughout. The embodiments described below are exemplary and are intended to explain the present application, and are not to be understood as limiting the present application.

[0091] Figure 1 A flowchart of an optical remote sensing image cloud removal method based on an iterative generative adversarial network according to an embodiment of the present application is provided.

[0092] As shown in the optical remote sensing image cloud removal method based on the iterative generative adversarial network, the method comprises the following steps: Figure 1

[0093] Step 1, constructing a shared feature extraction module and a private feature extraction module in a generator network, inputting an optical image I o and a SAR image I s into the shared feature extraction module and the private feature extraction module respectively to obtain shared features S of the optical image and the SAR image, private features P of the optical image, and private features of the SAR image;

[0094] Step 2, constructing a progressive dual attention feature fusion module in the generator network, and sequentially fusing the shared features and the private features of the optical image and the SAR image by using the progressive dual attention feature fusion module to perform feature reconstruction of a cloud-free image;

[0095] Step 3, constructing a feature reconstruction module in the generator network, and restoring the fusion result to a high-quality cloud-free image by using the feature reconstruction module;

[0096] Step 4, constructing a discriminator network, inputting the generated cloud-free image of the generator network and the real cloud-free image into the discriminator network, judging the quality of the generated cloud-free image, and feeding back the discrimination result to the generator network.

[0097] Optionally, in an embodiment of the present application, in step 1, an observation model based on convolutional sparse coding is first established to sufficiently extract shared and private features from different modalities. The observation model is expressed as follows:

[0098]

[0099] wherein, and are shared feature filters of I o and I s , and s i represents the shared features of I o and I s . ​and is I o and I s private feature filter, p i and v i is I o and I s private feature.

[0100] According to the above model, if all the filters are known, the cloud-free image F can be obtained by optimizing the following problem:

[0101]

[0102] It is difficult to solve each variable by jointly optimizing formula (3). Therefore, formula (3) is decomposed into three sub-problems, and each variable is updated alternately while fixing other variables. Specifically, the ISTA algorithm is used to solve the following three sub-problems:

[0103]

[0104] First, update P to solve the quadratic approximation of formula (4) to update the private feature P, which can be expressed as follows:

[0105]

[0106] where P t-1 represents the result after the (t-1)th iteration, η p is the iteration step size;

[0107]

[0108] The optimal solution of equation (7) can be defined by a proximal operator as follows:

[0109] Δg(P t-1 ) = X p *(X s *S t-1 + X p *P t-1 I o ) (9)

[0110]

[0111] where * is a transpose convolution operation that can be learned by a transpose convolution layer, X s and X p represent the filters involved in and respectively, is a proximal operator that depends on the regularization term.

[0112] Similar to the above method for updating V, a quadratic approximation to solve equation (5) is used to update the private feature P, which can be expressed as follows:

[0113]

[0114] where V t-1 denotes the result after (t-1) iterations, η v is the iteration step size;

[0115]

[0116] The optimal solution of equation (7) can be defined by a proximal operator as follows:

[0117] Δg(V t-1 ) = Y v *(Y s *S t-1 ) + Y v *V t-1 -I s (13)

[0118]

[0119] where * is a transposed convolution operation that can be learned by a transposed convolution layer, Y s and Y v denote the filters involved in and respectively, is a proximal operator that depends on the regularization term.

[0120] Finally, update S: rewrite equation (6) to get equation (15):

[0121]

[0122] where denotes the superposition of I s and I o . Similarly, use the same method as P and V to optimize this subproblem. The optimization process is shown as follows:

[0123]

[0124] where L s denotes the filter involved in is a proximal operator that depends on the regularization term.

[0125] Expand equations (10), (14), and (19) into a deep iterative network, which contains three subnetworks to realize the prediction of P, V, and S respectively. X s , X p , Ys 、Y v , L s The parameters are learned by the convolution layer. The size of all convolution kernels is 3*3 and the number of channels is 64 (except L s In addition, the number of channels is 128). p *、Y v *、L s The parameters of * are learned by the transposed convolutional layer. In addition, ResNet is used to implicitly learn prior information instead of using hand-crafted images. η is the inverse of the Lipschitz constant.

[0126] Optionally, in one embodiment of the present invention, in step 2, the main structure of the progressive dual-attention fusion module is composed of two Cross-Attention Blocks (CAB) with the same structure and different parameters and one Multi-head self-attention block (MSAB). The entire fusion process can be expressed as follows:

[0127] F=MSAB(CAB(CAB(P,S),V)) (20)

[0128] Among them, CAB(.) and MSAB(.) represent cross attention blocks and multi-head self-attention blocks respectively. The purpose of CAB is to fuse the difference information from different branches, while the goal of MSAB is to improve the learning ability of global features to reconstruct the information of areas covered by thick clouds. The first CAB fuses the private features P and shared features S of the optical image to obtain the intermediate features F M The private feature P is linearly reshaped into Q P , the shared feature S is linearly transformed into K S Sum V S .Q P Matrix multiplication and K S Transpose to get the attention weight matrix and calculate Q P and K S The process is as follows:

[0129]

[0130] Here, d is the scaling factor, which avoids the vanishing softmax gradient problem caused by the dot product becoming too large as d increases. Note that matrix factorization is used to reduce the amount of computation when calculating the attention weights. Afterwards, similar features between P and S can be removed, and the private features P of the optical image can be transferred to the intermediate features FM. This has the advantage of fusing the difference information between P and S, mitigating feature transfer in cloud-covered areas. This process is as follows:

[0131] F M = Linear(V S (1- Attention)) + Q P (22)

[0132] F M = MLP(Norm(F M )) + F M (23)

[0133] where MLP(.) denotes a multi-layer perceptron and Norm(.) denotes a linear norm operation. Subsequently, the fusion process of the intermediate feature F M and the private feature V in the second CAB is the same as the first CAB, which is defined as follows:

[0134] Q V ,K f ,V f = Linear(Reshape(V,F M ,F M ) (24)

[0135]

[0136] F = MLP(Norm(F)) + F (26)

[0137] The self-attention mechanism captures the dependency between pixels at any position and other positions through global receptive field. This ability enhances the utilization of non-local features. Therefore, the MSAB is integrated into the PDAFM. The fusion result F of the CAB is fed into the MSAB to better capture global information. The fusion result F is first put into the normalization layer to stabilize the distribution of the feature. Then, the attention map of F is calculated by the multi-head self-attention mechanism, which respectively calculates the attention result of each head and concentrates the attention weight of all heads. The process of MSAB is shown as follows:

[0138] Q,K,V = FW (27)

[0139] [Q1...Q i ],[K1...K i ],[V1...V i ] = Split(Q), Split(K), Split(V) (28)

[0140]

[0141] Multihead = concat(head1, head2... head i ) (30)

[0142] Where W represents the depthwise convolution and Split(.) represents the split function. Next, the gating mechanism G(.) is introduced, which is used in the feedforward neural network to guide the attention result. It consists of a convolutional layer with a kernel of 1*1 and a GELU activation function. Then, the Hadamard product between G(F) and the attention result is performed to obtain the feature FG. Finally, the input F is added to the output using a skip connection to obtain the final result F. fused :

[0143] F fused =MSAB(F)⊙G(F)+F (31).

[0144] Figure 2 This is the generator structure diagram based on iterative generative adversarial network. Figure 3 Diagram of the discriminator structure based on iterative generative adversarial network.

[0145] Figure 4(a) is a cloud-covered optical image of the SEN12MS-CR dataset, which shows a valley area. Figure 4(b) is a SAR image of the SEN12MS-CR dataset, which has the same scene as the optical image. Figure 4(c) is a cloud-covered optical image of the SMILE-CR dataset, which shows a plain area. Figure 4(d) is a SAR image of the SMILE-CR dataset, which has the same scene as the optical image.

[0146] Figure 5(a) shows the cloud removal result of the SEN12MS-CR satellite dataset based on the Pix2pix method; Figure 5(b) shows the cloud removal result of the SEN12MS-CR satellite dataset based on the McGAN method; Figure 5(c) shows the cloud removal result of the SEN12MS-CR satellite dataset based on the SAR-opt-cGAN method; Figure 5(d) shows the cloud removal result of the SEN12MS-CR satellite dataset based on the Dsen2-CR method; Figure 5(e) shows the cloud removal result of the SEN12MS-CR satellite dataset based on the GLF-CR method; Figure 5(f) shows the cloud removal result of the SEN12MS-CR satellite dataset based on the Align-CR method; Figure 5(g) shows the cloud removal result of the SEN12MS-CR satellite dataset based on the IGAN method.

[0147] Fig. 6(a) is a cloud removal result of the SMILE-CR satellite dataset based on the Pix2pix method; Fig. 6(b) is a cloud removal result of the SMILE-CR satellite dataset based on the McGAN method; Fig. 6(c) is a cloud removal result of the SMILE-CR satellite dataset based on the SAR-opt-cGAN method; Fig. 6(d) is a cloud removal result of the SMILE-CR satellite dataset based on the Dsen2-CR method; Fig. 6(e) is a cloud removal result of the SMILE-CR satellite dataset based on the GLF-CR method; Fig. 6(f) is a cloud removal result of the SMILE-CR satellite dataset based on the Align-CR method; and Fig. 6(g) is a cloud removal result of the SMILE-CR satellite dataset based on the IGAN method.

[0148] Table 1 is a quantitative evaluation result of the cloud removal result on the SEN12MS-CR and SMILE-CR datasets, indicators PSNR, SSIM, SAM, and MAE. The proposed A Dual Attention Blocks-Based Disentangled Iterative Generative Adversarial Network (IGAN) achieves the best result in all methods in all indicators.

[0149] Table 1 is a quantitative evaluation result of the cloud removal result on the SEN12MS-CR and SMILE-CR datasets, indicators PSNR, SSIM, SAM, and MAE. The proposed A Dual Attention Blocks-Based Disentangled Iterative Generative Adversarial Network (IGAN) achieves the best result in all methods in all indicators.

[0150]

[0151] According to the optical remote sensing image cloud removal method based on the iterative generative adversarial network proposed in the embodiments of the present application, the cloud removal method based on the iterative generative adversarial network can improve the quality of the cloud-free image reconstruction result and promote wider application.

[0152] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.

[0153] Furthermore, the terms "first", "second", etc. are used herein for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly pointing to the number of technical features indicated. Thus, features defined with "first", "second" can explicitly or implicitly include at least one of such features. In the description of the application, the meaning of "N" is at least two, for example two, three, etc., unless explicitly and specifically defined otherwise.

[0154] Any process or method descriptions or descriptions of the flow diagrams herein can be understood as indicating the possible implementation of the described processes or methods by process steps, program code, or other logic, and the present preferred embodiments contemplate a plurality of possible program code implementations of the processes or methods described herein, which can be executed by or to control the functioning of a computer or computer- based systems, and / or other functional entities. The program code can be written in any suitable form, including machine code, assembly code, interpreted code, object-oriented code, and / or high-level languages, and all equivalent techniques can be used to produce the desired results.

Claims

1. A cloud removal method for optical remote sensing images based on iterative generative adversarial networks, characterized in that: The following steps are involved: Step 1: construct a shared feature extraction module and a private feature extraction module in the generator network, input the optical image and the SAR image into the shared feature extraction module and the private feature extraction module respectively, and obtain the shared features of the optical image and the SAR image, the private features of the optical image, and the private features of the SAR image; Step 2: construct a progressive dual-attention feature fusion module in the generator network, and use the progressive dual-attention feature fusion module to sequentially fuse the shared features and private features of the optical image and the SAR image to reconstruct the features of the cloud-free image; Step 3: construct a feature reconstruction module in the generator network, and use the feature reconstruction module to restore the fusion result to a high-quality cloud-free image; Step 4: Construct a discriminator network, superimpose the cloud-free image generated by the generator network and the real cloud-free image and input them into the discriminator network, judge the quality of the generated cloud-free image and feed the identification result back to the generator network.

2. The method according to claim 1, characterized in that Step 1 specifically includes: An observation model based on convolutional sparse coding is established to extract shared features and private features from different modalities. The observation model is expressed as follows: Among them, I o is an optical image with clouds, I s is the SAR image, * is the transposed convolution operation, and For I o and I s The shared feature filter s i Indicates I o and I s The shared characteristics of and For I o and I s Private feature filter, p i and v i For I o and I s Private features of , i is the number of filters; According to the observation model, under the condition that all filters are known, the cloud-free image F is obtained by optimizing formula (3): Among them, P is the private feature of the optical image, V is the private feature of the SAR image, and S is the shared feature. is the F norm, λ s ,λ p ,λ v are all balance parameters, R(.) is the regularization term; Split equation (3) into three sub-problems, fix other variables while updating each variable alternately, and use the ISTA algorithm to solve the following three sub-problems: First, update P and solve the quadratic approximation of Equation (4) to update the private feature P, which is expressed as follows: Among them, P t-1 represents the result after (t-1) iterations, η p is the iterative step size for updating P, Δg(P t-1 ) is the derivative of P; The optimal solution of formula (7) is defined by a proximal operator as follows: Δg(P t-1 )=X p *(X s *S t-1 +X p *P t-1 I o ) (9) Among them, * is the transposed convolution operation, X s and X p Respectively and The filter involved, X p is the filter, S t-1 is the updated S, P for the t-1th time t-1 is the updated P for the t-1th time, P t is the t-th updated P, is the proximal operator that depends on the regularization term; Use the same method to update V and solve the quadratic approximation of Equation (5) to update the private feature P, which is expressed as follows: Among them, V t-1 represents the result after (t-1) iterations, η v is the iterative step size for updating V, Δg(V t-1 ) is the derivative of V; The optimal solution of formula (7) is defined by a proximal operator as follows: Δg(V t-1 )=Y v *(Y s *S t-1 +Y v *V t-1 -I s ) (13) Among them, Y s and Y v Respectively and The filter involved, Y v For the filter, V t-1 is the updated V for the t-1th time, V t is the t-th updated V, is a proximal operator that depends on the regularization term; Use the same method to update S, first rewrite formula (6) to obtain formula (15): Among them, when Further simplifying formula (15) yields formula (16): in, Indicates I s and I o The superposition of For the filter related to S, similarly, the same method as P and V is used to optimize the sub-problem. The optimization process is as follows: Among them, Δg(S t-1 ) is S t-1 The derivative of η s is the iterative step length for updating S, L s is the filter, L s represents the filter involved in S, S t is the result of S’s t-th update, is a proximal operator that depends on the regularization term; Expanding formulas (10), (14) and (19) transforms them into a deep iterative network, which contains three sub-networks to realize the prediction of P, V and S respectively, X s 、X p 、Y s 、Y v , L s The parameters are learned by the convolutional layer, except L s Except for the 128 channels, all convolution kernels are 3*3 in size and 64 in number of channels. p *、Y v *、L s The parameters of * are learned by the transposed convolutional layer, and ResNet is used to implicitly learn the prior information.

3. The method according to claim 1, characterized in that Step 2 specifically includes: the structure of the progressive dual-attention fusion module consists of two cross-attention blocks with the same structure and different parameters and one multi-head self-attention block. The entire fusion process is expressed by the following formula: F=MSAB(CAB(CAB(P,S),V)) (20) Among them, CAB(.) and MSAB(.) represent cross attention blocks and multi-head self-attention blocks respectively. The goal of the cross attention block is to fuse the difference information from different branches. The goal of the multi-head self-attention block is to improve the learning ability of global features to reconstruct the information of the area covered by thick clouds. The first cross attention block fuses the private feature P and the shared feature S of the optical image to obtain the intermediate feature F M , the private feature P is linearly reshaped into Q P , the shared feature S is linearly transformed into K S Sum V S , Q P Matrix multiplication and K S Transpose to get the attention weight matrix and calculate Q P and K S The feature similarity between them is as follows: Where d is the scaling factor; Matrix decomposition is used to reduce the amount of calculation when calculating the attention weight. After that, the similar features between the private feature P and the shared feature S are removed, and the private feature P of the optical image is transferred to the intermediate feature F. M , the process is as follows: F M =Linear(V S (1-Attention))+Q P (22) F M =MLP(Norm(F M ))+F M (23) Among them, V S is the key value of S, Linear() is the linear transformation function, MLP(.) represents the multi-layer perceptron, Norm(.) represents the linear norm operation, and the intermediate feature F in the second cross attention block M The fusion process of the private feature V is the same as the first cross-attention block and is defined as follows: Q V ,K f ,V f =Linear(Reshape(V,F M ,F M ) (24) F=MLP(Norm(F))+F (26) Among them, Q V For the query of V, K f The key for the middle feature, V f is the key value of the intermediate feature, Reshape() is the shape transformation operation, and F is the fusion result; The self-attention mechanism captures the dependency between pixels at any position and pixels at other positions through global perception. The multi-head self-attention block is integrated into the network, and the fusion result F of the cross-attention block is fed into the multi-head self-attention block to capture global information. The fusion result F is first put into the normalization layer to stabilize the distribution of features. Then, the attention map of the fusion result F is calculated by the multi-head self-attention mechanism. The multi-head self-attention mechanism calculates the attention result of each head separately and aggregates the attention weights of all heads. The process of the multi-head self-attention block is as follows: Q,K,V=FW (27) [Q1...Q i ],[K1...K i ],[V1...V i ]=Split(Q),Split(K),Split(V) (28) Multihead=concat(head1,head2...head i ) (30) Among them, Q, K, and V are the three elements of the attention mechanism. i , K i and V i The three elements of the attention mechanism for each head, W represents the depth of convolution, Split (.) represents the split function; The gating mechanism G(.) is used in the feedforward neural network to guide the attention result. It consists of a convolutional layer with a kernel of 1*1 and a GELU activation function. Then, the Hadamard product between G(F) and the attention result is performed to obtain the feature FG. Finally, the input F is added to the output using a skip connection to obtain the final result F. fused : F fused =MSAB(F)⊙G(F)+F (31)。 4. The method according to claim 1, wherein The feature reconstruction module of the generator constructed in step 3 uses the residual block as the backbone structure, which consists of five residual blocks and one convolution layer. Each residual block contains three convolution layers and a jump connection. The first two convolution layers are accompanied by a rectified linear unit. The size of the convolution kernel in each residual block is 3*3, and the number of input channels and output channels is 64. The final convolution layer has a convolution kernel size of 3*3, 64 input channels and 13 output channels to match the spectral dimension of the generated cloud-free image.

5. The method according to claim 1, wherein In step 4, the structure of the discriminator is composed of five blocks with the same structure and a convolution layer. Each block contains a convolution layer with a convolution kernel of 4*4, a normalization layer and an activation function Leaky ReLu. The stride of the convolution layer is 2, the convolution kernel size of the last convolution layer is 3, the stride is 1, and the last convolution layer outputs one-dimensional data to represent the result of the discriminator.

Citation Information

Cited By

  • Cross-modal sea surface target identification method and system for high-dynamic unmanned platform

    CN121544965A