UnCRtainTS cloud removal method fused with wavelet transform

The Tri-CloudNet model addresses the limitations of existing cloud removal methods by integrating wavelet transform and attention-based deep learning to enhance feature extraction and cloud separation, achieving superior cloud removal accuracy and detail preservation in remote sensing images.

CN120318113APending Publication Date: 2025-07-15YANGTZE DELTA REGION INST OF UNIV OF ELECTRONICS SCI & TECH OF CHINE (HUZHOU) +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510225635.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

When dealing with thick cloud areas, existing deep learning models have problems such as poor cloud removal effect, high training data collection and labeling costs, and limited model generalization capabilities, especially in complex thick cloud patterns, it is difficult to accurately restore obstructed surface information.

Method used

The UnCRtainTS decloud model Tri-CloudNet, which integrates wavelet transformation, isolates and processes low-frequency and high-frequency components through image classification, UnCRtainTS encoder, L-TAE module, Transformer multi-head self-attention mechanism and wavelet analysis, and combines time series features and spatial features to generate the final decloud image.

Benefits of technology

It significantly improves the accuracy and effect of thick cloud removal, reduces the root mean square error, improves the peak signal-to-noise ratio and structural similarity index, and enhances the model's robustness and detail resilience ability for different cloud types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318113A_ABST
    Figure CN120318113A_ABST
Patent Text Reader

Abstract

The invention discloses a UnCRtainTS cloud removal method fused with wavelet transform, belongs to the field of image processing, and particularly relates to a cloud removal method for aerial images. According to the method, the feature library of the cloud is constructed by using texture features extracted by wavelet transform. According to the step, the correlation between different wavebands is calculated to help to obtain the spectral difference between the cloud and the earth surface object, and the cloud recognition capability of the model is enhanced. And then, taking the high-frequency component after wavelet transformation and the extracted features as input, and inputting the high-frequency component and the extracted features into an improved UnCRtainTS cloud removal model. And after the model is output, combining a cloud removal result with the low-frequency component after wavelet transformation, and reconstructing a cloud-removed image by utilizing inverse wavelet transformation. Wavelet transform and an UnCRtainTS cloud removal model are fused to form a Tri-Cloud Net model, the advantages of a deep learning model for removing the cloud removal method are reserved, the deep learning model is fused with a traditional physical method to achieve innovation, and the defect that the cloud removal effect of pure deep learning is poor is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and particularly to a method for removing clouds from aerial images. Background Art

[0002] With the rapid development of artificial intelligence, deep learning for removing clouds in remote sensing images has many advantages over traditional physical methods, such as automatic feature learning, end-to-end mapping, strong generalization ability, high processing efficiency, good color restoration, more detail retention, strong adaptability, and data-driven. However, it also faces the challenge of poor cloud removal effect.

[0003] UnCRtainTS is a method for multi-temporal cloud removal. It combines a novel attention-based architecture and a formulation for multi-variable uncertainty prediction to handle multi-temporal cloud removal problems through an attention-based architecture. This architecture can handle various occlusion situations, from partially visible haze scenes to completely opaque cloud cover. For haze, thin clouds, and clouds, this model achieves the lowest root mean square error in the multi-temporal cloud removal model. However, the cloud removal effect is still not as good as that of haze and thin clouds.

[0004] Currently, for the problem of cloud removal in remote sensing images, researchers have proposed various processing methods, mainly including methods based on traditional image processing techniques and modern methods based on deep learning. Traditional methods mainly rely on techniques such as multi-temporal remote sensing image fusion, spectral analysis, image interpolation, and wavelet transform. By using multiple images taken at different time points or by extracting features of different spectral bands, thick cloud removal and surface information restoration are carried out. Such methods usually have some limitations. Traditional methods have high requirements for image quality and imaging conditions, are limited by the spatio-temporal resolution of multi-temporal images, and are easily affected by factors such as cloud dynamic changes and surface feature changes. Moreover, traditional methods have limited effects when dealing with complex thick cloud scenes. Especially for some thick cloud morphologies with diverse changes, traditional cloud removal methods often cannot accurately restore the occluded surface information.

[0005] With the rise of deep learning technology, more and more researchers have begun to attempt to use deep learning models to solve the problem of thick cloud removal. Deep learning models can automatically learn the complex mapping relationship between cloud layers and surface information through the training of massive data, and have powerful non-linear feature extraction and representation capabilities. Among various deep learning models, structures such as U-Net, GAN (Generative Adversarial Network), and Transformer have been widely applied to the task of thick cloud removal in remote sensing images due to their excellent performance in fields such as image restoration and image enhancement. Nevertheless, there are still some problems when existing deep learning models process thick cloud regions. For example, since thick cloud regions are usually accompanied by large brightness and contrast changes, traditional deep learning models may exhibit phenomena such as distortion and blurring when restoring occluded surface information. Training deep learning models requires a large amount of paired data (i.e., image pairs containing thick cloud and cloud-free scenes), and the collection and annotation costs of this kind of data are relatively high, and it is difficult to obtain in actual scenarios, which further limits the generalization ability and practicality of the models. Summary of the Invention

[0006] Aiming at the deficiencies in the background technology, the present invention proposes a Tri-CloudNet de-cloud model that combines wavelet transform with UnCRtainTS (Uncertainty Quantification for Cloud Removal in Optical Satellite TimeSeries) to solve the problem of poor cloud removal effect.

[0007] The technical solution of the present invention is a de-cloud method that combines wavelet transform with UnCRtainTS, and the method includes:

[0008] Step 1: Classify the image according to the characteristics of the image. Classify fog and thin clouds as category one, and classify thick clouds as category two; if the image is classified as category one, go to step 2; if the image is classified as category two, go to step 8;

[0009] Step 2: Input the image into the UnCRtainTS encoder to extract preliminary features;

[0010] Step 3: Input the preliminary features into the max-pooling layer to further reduce the resolution and enhance important features, and then strengthen the feature expression through the convolutional layer;

[0011] Step 4: The features then enter the L-TAE module to capture the global information of the time series. The L-TAE module represents the long-term dependence attention encoding module;

[0012] Step 5: After the L-TAE module, the data is input into the multi-head self-attention mechanism of the Transformer to enhance the correlation representation between different time steps;

[0013] Step 6: Perform upsampling layer to restore the image size, and then perform temporal aggregation on the upsampled image and the preliminary features output by the UnCRtainTS encoder to obtain optimized time-series features;

[0014] Step 7: Pass the optimized time-series features through the UnCRtainTS decoder to generate the cloud-removed image output;

[0015] Step 8: Decompose the image through wavelet analysis to separate the low-frequency component and the high-frequency component;

[0016] Step 9: Process the high-frequency component according to the methods in Steps 2 to 7 to output the high-frequency cloud-removed image;

[0017] Step 10: Combine the high-frequency cloud-removed image obtained in Step 9 with the low-frequency component obtained in Step 8, and then perform inverse wavelet transform to generate the final cloud-removed image.

[0018] Further, the specific method of Step 3 is:

[0019] Downsample the feature f at each time step t Assume that the downsampled feature is denoted as Its size is d m Denote the dimension of the downsampled feature, that is, the length of the feature vector, Denote the height of the downsampled feature, Denote the width of the downsampled feature;

[0020]

[0021] Further, the classification method of the image in Step 1 includes:

[0022] Method 1: Manual classification;

[0023] Method 2: Train a classification model with labeled images, use the convolutional layer of the classification model for feature extraction, and use the fully connected layer for classification.

[0024] Further, the implementation method of the Transformer multi-head self-attention mechanism in Step 5 is:

[0025] Assume the input sequence is X = [x1, x2, …, x n , each x i Represents an element in the sequence. To generate the query, key, and value vectors at each position, first use linear transformation to map X to the query Q, key K, and value V vector spaces respectively, W Q , WK , W V is the weight matrix for queries, keys, and values;

[0026] Q = XW Q , K = XW K , V = XW V

[0027] The attention weight matrix is determined by calculating the dot product of the query and the key. The specific formula is as follows:

[0028]

[0029] where d k is the dimensionality of the query and key vectors, QK T represents the product of Q and the transpose of K; d k is used to prevent the value of QK T from being too large, which would cause the output of the softmax function to be too concentrated; then, the softmax function is applied to each element of QK T to obtain a value between 0 and 1, which represents the weight of each value (Value) element. Finally, these weights are used to weight-sum the value (Value) matrix V to obtain the final output.

[0030] Therefore, softmax plays a role in normalization and weight assignment in the attention mechanism, enabling the model to focus more on important information while ignoring unimportant information.

[0031] Assume that h attention heads are used. Then, the formula for the i-th head head i is:

[0032]

[0033] Here, is the weight matrix independent for each head; the output of the multi-head attention is the concatenation of the results of all heads, that is:

[0034] MultiHead(Q, K, V) = Concat(head1,..., head h )W O

[0035] where Concat means concatenating the outputs of all heads in the feature dimension, and W O is the weight matrix of the output, which is used to map the concatenated representation back to the original dimension.

[0036] Furthermore, the method of time aggregation in step 6 includes three main gates: the forget gate f t , the input gate i t, output gate o t , candidate memory cell Memory cell update c t ;

[0037] The calculation method of temporal aggregation is as follows:

[0038]

[0039] h t = o t ⊙ tanh(c t )

[0040] where the output at time step t is denoted as h t , σ represents the sigmoid activation function, ⊙ represents element-wise multiplication, h t-1 is the aggregated feature of the previous time step, W C , W i , W o , W f represent the weights of the corresponding units, b i , b f , b o , b C are the biases of the corresponding units; tanh represents the hyperbolic tangent function, ensuring that the candidate memory value is between (-1, 1); h t is the aggregated feature of the current time step, U i represents the weight matrix connecting the input gate to the hidden state of the previous time step; U f represents the weight matrix connecting the forget gate to the hidden state of the previous time step; U o represents the weight matrix connecting the output gate to the hidden state of the previous time step; U c represents the weight matrix connecting the candidate memory cell to the hidden state of the previous time step;

[0041] These weight matrices are combined with the input features of the current time step through the weights W i , W f , W o , W C , and adjusted through the bias terms b i , b f , b o , b C to calculate the activation values of each gate and the value of the candidate memory cell; these activation values and the value of the candidate memory cell are then used to update the internal state and output of the LSTM cell.

[0042] Furthermore, the wavelet transform method in step 8 is as follows:

[0043] Decompose the image into a low-frequency component L and multi-level high-frequency components H using the Daubechies wavelet basis LH ,H HL ,H HH ,H LH ,H HL ,H HH which represent the high-frequency components in the horizontal, vertical, and diagonal directions respectively.

[0044] Furthermore, the forget gate f t controls the proportion of the memory cell information from the previous moment to be forgotten at the current moment. Let W f and b f be the weight matrix and bias of the forget gate respectively, h t-1 be the hidden state of the previous time step, x t be the input at the current time step, and σ denote the Sigmoid activation function; the calculation formula for the forget gate is:

[0045] f t = σ(W f · [h t-1 , x t + b f )

[0046] Based on the features of the current input and the previous hidden state, the forget gate determines the information to be discarded, thereby removing the noise features that cause inaccurate cloud removal.

[0047] Furthermore, the input gate i t controls the degree to which the new information at the current moment is updated in the memory cell, and the calculation formula is as follows:

[0048] i t = σ(W i · [h t-1 , x t + b i )

[0049] where W i and b i are the weight and bias of the input gate. By controlling the value of i t , the temporal aggregation method selectively introduces new image features according to the current input information.

[0050] Furthermore, the output gate o t controls the amount of information extracted from the current memory cell and updates the hidden state to be used as the input for the next time step; the calculation formula is:

[0051] o t = σ(W o · [h t-1 , x t+b o )

[0052] In the formula, W o and b o are the weights and biases of the output gate.

[0053] The present invention uses the open-source dataset SEN12MS-CR-TS to verify the experimental results. The experimental results in the eastern region of Asia show that the Tri-CloudNet model performs better in terms of accuracy. Specifically, the value of RMSE (root mean square error) is 0.042, the value of PSNR (peak signal-to-noise ratio) is 30.16, the value of SSIM (structural similarity index) is 0.91, and the value of SAM (spectral angle mapper) is 7.54. Compared with the original model, RMSE has decreased by 0.005, PSNR has increased by 1.26, SSIM has increased by 0.03, and SAM has decreased by 0.78. Description of the Drawings

[0054] Figure 1 is a schematic diagram of the temporal aggregation method.

[0055] Figure 2 is the model diagram of the cloud removal method of the present invention.

[0056] Figure 3 is the flowchart of the cloud removal method of the present invention.

[0057] Figure 4 is the comparison diagram of the implementation effects of the present invention. Detailed Embodiments

[0058] The dataset used by the UnCRtainTS deep learning network adopted by the present invention is SEN12MS-CR-TS, a multi-modal and multi-temporal dataset designed specifically for deep learning, with no intersection between adjacent samples, and is selected from the region of interest (ROI) in the eastern part of Asia. To enhance the robustness of the training data, each ROI has been independently measured 30 times repeatedly to ensure the representativeness and diversity of the data. The training set has collected 30 radar and optical satellite images for each region, and each image pair consists of 240 to 300 pairs of images, with an overall cloud coverage rate of about 50%, which conforms to the situation of real remote sensing images. These data cover various situations from clear images to semi-transparent haze, small cloud blocks, and even dense clouds. Each pair of images is composed of C-band Sentinel-1 radar data with dual-polarization function, and 1C-level top-of-atmosphere reflectance Sentinel-2 optical data that has been orthorectified and sub-pixel geometrically refined.

[0059] 1. Processing of thin clouds and foggy images:

[0060] (1) The image enters the UnCRtainTS encoder to extract a preliminary feature representation.

[0061] (2) The resolution is further reduced and important features are enhanced through the max pooling layer, and then the feature expression is strengthened through the convolutional layer.

[0062] (3) The features then enter the Long-Term Attention Encoding module (L-TAE) to capture the global information of the time series.

[0063] (4) After L-TAE, the data is input into the multi-head self-attention mechanism of the Transformer to enhance the correlation representation between different time steps.

[0064] (5) The image size is restored through the upsampling layer and input into the time aggregation module to further optimize the time series features.

[0065] (6) The de-clouded image output is generated through the UnCRtainTS decoder.

[0066] Let I cloudy represent the cloud-covered image, and F cloud-free represent the feature of the output de-clouded image. This process can be expressed by the following formula:

[0067] F cloud-free = UnCRtainTS decoder (Time aggregation module

[0068] (Upsample(Attention(L-TAE(Conv(MaxPool(UnCRtainTS encoder (I cloudy ))))))))

[0069] 2. Processing of thick cloud images:

[0070] (1) For the images identified as thick clouds, the images are first decomposed by wavelet analysis to separate the low-frequency and high-frequency components.

[0071] (2) The high-frequency components contain important texture information and detailed features, which can be used to remove the cloud features. The high-frequency components are used as the input and enter the UnCRtainTS encoder according to the same processing steps for thin cloud / foggy images, passing through modules such as max pooling, convolutional layer, LTAE, Transformer multi-head self-attention mechanism, upsampling, and time aggregation module.

[0072] (3) The output after being processed by the UnCRtainTS decoder is recombined with the previous wavelet low-frequency components, and the final de-clouded image is generated through the inverse wavelet transform.

[0073] The formula is expressed as follows:

[0074] F′ cloud-free = InverseWaveletTransform(L, UnCRtainTS decoder (Time aggregation module (Upsample(Attention(L - TAE(Conv(MaxPool(UnCRtainTS encoder (H)))))))))

[0075] Among them, L represents the low - frequency component, H represents the high - frequency component, and F′ cloud-free represents the complete image after cloud removal. Through this process design, the model can utilize the spatio - temporal features of multi - modal input and the multi - level attention mechanism to effectively process remote sensing images under different cloud conditions, providing better results for complex cloud removal tasks.

[0076] (1) Feature downsampling

[0077] Downsample the feature f at each time step t . Assume the downsampled feature is denoted as whose size is Here This can reduce the computational load while retaining important spatial information.

[0078]

[0079] (2) Temporal aggregation to process temporal features

[0080] Input the downsampled feature sequence into the time aggregation module. The time aggregation module will gradually process the features at each time step and finally output an aggregated temporal feature Let the hidden state of the time aggregation module be h t , and the output at time step t is denoted as h t . The update formula of the time aggregation module is:

[0081]

[0082] h t = o t ⊙tanh(c t )

[0083] Among them, i t , f t , o trepresent the input gate, forget gate, and output gate respectively, σ represents the sigmoid activation function, and ⊙ represents element-wise multiplication. After being processed by the temporal aggregation module, the final hidden state h T is the aggregated feature of the entire time series.

[0084] (3) Feature upsampling

[0085] The obtained aggregated feature has a size of To be consistent with the original spatial features, it is upsampled back to the original resolution through bilinear interpolation, i.e., [d m ×H×W]:

[0086]

[0087] After obtaining the aggregated feature it enters the decoding part. The decoder contains a series of MBConv blocks and a final point convolution for mapping the aggregated feature to the output image channel number C out . The formula for the final layer's point convolution is:

[0088]

[0089] where σ is the activation function. For the reconstructed image channels, the sigmoid function is used to compress the output values into the valid range; while for the uncertainty prediction channels, the softplus activation is used to ensure that the output values are positive. The advantage of the temporal aggregation module in the cloud removal task is that it can effectively extract important information in the time series and filter out noise features that may affect the cloud removal accuracy of the image, thereby improving the reliability of the model.

[0090] To verify the impact of adding the temporal aggregation module on the cloud removal effect, the cloud removal results of the Tri-CloudNet model with the temporal aggregation module removed and the Tri-CloudNet model were compared. Through various metrics in Table 1, including RMSE, PSNR, SSIM, and SAM, the performance of the model in cloud removal was analyzed and the effectiveness of the Tri-CloudNet algorithm was verified.

[0091] After removing the temporal aggregation module, the RMSE of the model increased to 0.049, the PSNR decreased to 28.22, the SSIM also decreased to 0.86, while the SAM increased to 8.45. The temporal aggregation module plays an important role in time series processing and can effectively capture the sequential relationship between images. Through the temporal information of multiple images, the temporal aggregation module can utilize the information of historical frames to enhance the cloud removal effect of the current frame. Therefore, after removing the temporal aggregation module, the model's dependence on historical images for cloud removal weakens, resulting in worse overall metrics. This indicates that the temporal aggregation module makes a significant contribution to cloud removal in time series modeling.

[0092]

[0093]

[0094] Table 1 Comparison of experimental results

[0095] Replace the self-attention mechanism in the UnCRtainTS model with the multi-head self-attention mechanism in the Transformer architecture. The multi-head self-attention mechanism can concurrently focus on the relationships between different time points, enhance the model's ability to capture cloud change information, and improve the cloud removal effect to capture cloud change information between different time points.

[0096] The multi-head self-attention mechanism of Transformer is a parallel mechanism for computing multiple attentions, used to capture dependencies between different positions in the input sequence. The input is first linearly mapped to multiple subspaces, each subspace called a "head", in order to focus on the information in the sequence from different perspectives. For each head, the input sequence is mapped to query, key, and value vector spaces. By calculating the similarity (inner product) between the query and the key, an attention weight matrix is obtained, and the matrix is used to weight the value vectors at each position. After concatenating the attention outputs of each head, the final self-attention output is generated through a linear transformation. This mechanism allows the model to learn the relationships between different positions at multiple scales and understand the features in the sequence more comprehensively.

[0097] During the implementation of the Transformer multi-head self-attention mechanism, assume the input sequence is X = [x1, x2, …, x n , and each x i represents an element in the sequence. To generate the query, key, and value vectors for each position, first use a linear transformation to map X to the query Q, key K, and value V vector spaces respectively, where W Q , W K , W V are the weight matrices for the query, key, and value.

[0098] Q = XWQ , K = XW K , V = XW V

[0099] The core of self-attention is to capture context information by calculating the correlation between each position in the sequence and other positions. The attention weight matrix is determined by calculating the dot product of the query and the key, is a scaling factor, usually taking the dimension size of the query and key vectors (d k ), which is used to stabilize the training process. The specific formula is as follows:

[0100]

[0101] A single attention head can only capture sequence relationships within one representation space. Multi-head self-attention enhances the model's expressive power by mapping the input to multiple subspaces and performing the above self-attention operations in different subspaces. Assuming h attention heads are used, the calculation formula for the i-th head is:

[0102]

[0103] Here, is the weight matrix independent for each head. The output of multi-head attention is the concatenation of the results of all heads, that is:

[0104] MultiHead(Q, K, V) = Concat(head1,..., head h )W O

[0105] where Concat means concatenating the outputs of all heads in the feature dimension, and W O is the weight matrix of the output, which is used to map the concatenated representation back to the original dimension.

[0106] The multi-head self-attention mechanism is incorporated into a standard layer structure in Transformer, including multi-head self-attention, residual connection, and layer normalization (Layer Normalization). The main steps are as follows:

[0107] (1) The input first passes through the multi-head self-attention mechanism to obtain a context-related representation.

[0108] (2) Add a residual connection to directly add the input to the output: X + MultiHead(Q, K, V).

[0109] (3) Stabilize the training through layer normalization.

[0110] This combination forms the basic computational unit of the Transformer, which is used to capture the dependencies at different positions in the sequence. In summary, the multi-head self-attention mechanism learns the dependencies between different positions from multiple subspaces by parallelizing multiple self-attention heads, which helps to enhance the model's performance in dealing with long-range dependencies.

[0111] After introducing the time aggregation module and the multi-head attention mechanism, in order to enhance the robustness of the Tri-CloudNet model in processing different cloud types (such as fog, thin clouds, thick clouds), weight parameters are introduced into the loss function, so that the model assigns different importance to different difficulty cloud removal tasks during training. According to experience, the processing errors of thin cloud or foggy images are given lower weights of 0.2 and 0.2, while the processing error of thick clouds is given a higher weight of 0.6.

[0112] To verify the impact of adding the Transformer multi-head attention mechanism on the cloud removal effect, the cloud removal results of the Tri-CloudNet model with the Transformer multi-head attention mechanism removed and the Tri-CloudNet model are compared. Through various indicators in Table 1, including RMSE, PSNR, SSIM, and SAM, the performance of the model in cloud removal is analyzed, and the effectiveness of the Tri-CloudNet algorithm is verified.

[0113] After removing the multi-head self-attention mechanism, the RMSE of the model rises to 0.048, the PSNR drops to 28.37, the SSIM is 0.87, and the SAM is 8.26. The multi-head self-attention mechanism has advantages in capturing long-range dependencies and context information of images. Removing this module will limit the model's ability to fuse spatial information in images, reduce the ability to restore details of cloud-removed images, and have a negative impact on PSNR and SSIM. The comparison results are shown in the following table;

[0114] Model ↓RMSE ↑PSNR ↑SSIM ↓SAM Wave-TSCloud Removes the Transformer Multi-Head Self-Attention Mechanism 0.048 28.37 0.87 8.26 Tri-CloudNet (3 Image Inputs) 0.042 30.16 0.91 7.54

[0115] Since wavelet transform can decompose an image into sub-bands of different scales, which helps to extract local features (such as texture) and global features (such as brightness distribution) of clouds, provides richer input information for the UnCRtainTS model, and increases the ability to extract multi-scale information. And since the UnCRtainTS model is good at capturing uncertainties in time series and space, but has limited ability to extract frequency domain information of details, while wavelet transform can just make up for this point, making the fused model have the advantages of both spatio-temporal features and frequency domain features. Therefore, the present invention proposes a remote sensing image processing method for thick cloud removal by combining wavelet transform and the UnCRtainTS model.

[0116] Combined with the wavelet transform of the traditional remote sensing cloud removal model and the improved UnCRtainTS deep learning model, first perform wavelet transform on the remote sensing image. We can choose the wavelet basis (Daubechies) to decompose the image into low-frequency and high-frequency components. In the high-frequency components, use the texture features (such as energy, entropy, etc.) extracted based on wavelet transform to establish a cloud feature library. By calculating the correlation between different bands, obtain the spectral differences between clouds and surface objects.

[0117] Take the high-frequency components after wavelet transform and the features extracted from them as inputs, and input them into the improved UnCRtainTS cloud removal model. Combine the output of the UnCRtainTS model with the low-frequency components after wavelet transform, and use inverse wavelet transform to reconstruct the cloud-removed image. Obtain the cloud-removed image by reconstructing the wavelet image. The thick cloud processing model based on the improved UnCRtainTS by wavelet transform is as follows:

[0118] The model performs wavelet decomposition on the input remote sensing image I, and uses the Daubechies wavelet basis to decompose the image into a low-frequency component L and multi-level high-frequency components H LH ,H HL ,H HH , representing the high-frequency details in the horizontal, vertical, and diagonal directions respectively. Let L represent the low-frequency component of the image, which contains the overall structural information of the image, and H LH ,H HL ,H HH High-frequency components, retaining the detail information in the horizontal, vertical, and diagonal directions respectively. The formula for wavelet decomposition is expressed as:

[0119]

[0120] In the high-frequency components H LH ,H HL ,H HH , construct a cloud feature library by calculating the texture features (energy E and entropy S) of each component for identifying thick cloud regions. The calculation formulas for these texture features are as follows:

[0121] Energy E:

[0122]

[0123] Entropy S:

[0124]

[0125] These features are used to construct a cloud feature library to assist the improved UnCRtainTS model in more accurately identifying cloud regions in the high-frequency components. Combine the extracted high-frequency features and the high-frequency components {H LH ,H HL ,HHH} is used as the input and fed into the improved UnCRtainTS model. The model further enhances the thick cloud removal ability by using a deep learning network architecture and outputs the high-frequency component H' after cloud removal. LH , H' HL , H' HH :

[0126] H' LH , H' HL , H' HH =UnCRtainTS({H LH , H HL , H HH}, Features)

[0127] The high-frequency component H' after cloud removal LH , H' HL , H' HH is combined with the low-frequency component L of the initial wavelet decomposition, and the image I' after cloud removal is reconstructed through inverse wavelet transform:

[0128] I' = InverseWaveletTransform(L, {H' LH , H' HL , H' HH})

[0129] The model combines the advantages of wavelet transform and UnCRtainTS model, retains the detailed features of the image, and effectively removes thick clouds, making the finally reconstructed image I' visually clearer.

[0130] Wavelet transform can separate the low-frequency and high-frequency information of an image, retain the details of the image, and at the same time provide texture features, which helps to enhance the cloud feature recognition ability. By combining the traditional remote sensing cloud removal model wavelet transform with the improved UnCRtainTS deep learning model, first perform wavelet transform on the remote sensing image. The wavelet basis (Daubechies) can be selected to decompose the image into low-frequency and high-frequency components. In the high-frequency components, use the texture features (such as energy, entropy, etc.) extracted based on wavelet transform to establish a cloud feature library. Take the high-frequency components after wavelet transform and the extracted energy and entropy features as the input and feed them into the improved UnCRtainTS cloud removal model. Combine the output of the UnCRtainTS model with the low-frequency components after wavelet transform, and use inverse wavelet transform to reconstruct the image after cloud removal, and obtain the image after cloud removal by reconstructing the wavelet image.

[0131] According to the characteristics of the images, first judge the cloud conditions of each image and select different processing paths accordingly. Compared with using visual judgment for manual classification of input cloudy pictures in wavelet transform, the cloud processing model of UnCRtainTS improved based on wavelet transform can achieve automatic classification of the input. In the training stage, the image dataset is clearly labeled with samples of thin clouds, thick clouds or foggy categories. Through supervised learning, the model learns the corresponding relationship between the characteristics of the input images and the categories, optimizes the cross-entropy loss function. After sufficient training, the model can automatically classify the cloud types of the input images. The convolutional layer is responsible for feature extraction, and the fully connected layer realizes classification. The specific classification algorithm process is as follows:

[0132] During the experiment, due to the use of wavelet transform, problems of multi-scale information alignment will occur. Wavelet transform introduces features of different scales, and it may be difficult for deep learning models to effectively capture the relationships between different scales when fusing these features. To solve this problem (not only for solving this problem, but the attention mechanism can just make up for this deficiency), an attention mechanism is introduced into the deep learning model to dynamically adjust the weight distribution of features of different scales and improve the model's sensitivity to multi-scale information. After effectively fusing multi-scale information, the model performs better in dealing with the continuity of the thickness change of clouds, especially in the cloud regions with blurred boundaries, where more accurate segmentation and reconstruction can be achieved.

[0133] This invention of removing clouds from remote sensing images uses four common evaluation metrics to comprehensively evaluate the performance of the cloud removal model, including RMSE (Root Mean Square Error), PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity Index) and SAM (Spectral Angle Mapper). These metrics can help quantify the differences and restoration effects between the cloud-removed images and the original images.

[0134] RMSE is used to measure the difference between the cloud-removed image and the original image. The smaller the RMSE, the better the cloud removal effect. Let x i and y i be the values of the cloud-removed image and the original image at the i-th pixel position respectively, and N be the total number of pixels in the image. The RMSE formula is as follows:

[0135]

[0136] PSNR is used to measure the restoration quality of the image. The higher the PSNR, the better the image quality. MAX I is the maximum value of the image pixels (usually 255 for 8-bit images), and MSE is the mean square error. The PSNR formula is:

[0137]

[0138] The SSIM is used to evaluate the similarity of the cloud-removed image in terms of brightness, contrast, and structure. Let μ x and μ y be the means of images x and y, and be the variances, σ x xy be the covariance, and C1 and C2 be constants used to stabilize the denominator. The closer the SSIM value is to 1, the more similar the images are and the better the cloud-removing effect. Its formula is:

[0139]

[0140] The SAM is used to evaluate the difference in spectral information of the cloud-removed image. The smaller the value, the closer the cloud-removed image is to the original image in terms of spectral information. Its formula is:

[0141]

[0142] where I original (i) and I recovered (i) are the spectral vectors of the original image and the cloud-removed image at pixel point i respectively, and ∥·∥ represents the norm of the vector. The smaller the SAM value, the more similar the cloud-removed image is to the original image in terms of spectral features.

[0143] Experiments and Results

[0144] To verify the effectiveness of the final model Tri-CloudNet proposed in this invention, a comparative experiment was designed. The comparative experiment was carried out in the same experimental environment and dataset to evaluate the contribution of the improved model to the overall cloud-removing effect.

[0145] Original UnCRtainTS: Using the UnCRtainTS model as the benchmark model, directly input three images at different times without using the time aggregation module, the Transformer multi-head self-attention mechanism, and wavelet transform to evaluate the basic cloud-removing effect of the input of three images without the improved module.

[0146] Introducing Tri-CloudNet: Based on the complete UnCRtainTS model, by introducing the long short-term memory network, the multi-head attention mechanism, and wavelet transform, the Tri-CloudNet model was formed to further verify the impact of multi-image input on improving the cloud-removing accuracy.

[0147] Through the results of the above comparative experiments, the effectiveness of the Tri-CloudNet model was verified, and the experimental results are shown in the following table:

[0148] Evaluation Metrics Original UnCRtainTS Tri-CloudNet ↓RMSE 0.047 0.042 ↑PSNR 28.90 30.16 ↑SSIM 0.88 0.91 ↓SAM 8.32 7.54

[0149] The RMSE of the original UnCRtainTS model is 0.047, the PSNR is 28.90, the SSIM is 0.88, and the SAM is 8.32. The RMSE of the Tri-CloudNet model is 0.042, the PSNR is 30.16, the SSIM is 0.91, and the SAM is 7.45. The Tri-CloudNet model combines the feature decomposition of wavelet transform, the context capture of the multi-head self-attention mechanism, and the time information fusion function of the long short-term memory network, and the de-clouding effect of the model has been significantly improved. Its lower RMSE shows that the prediction accuracy of the model has been further improved, the higher PSNR and SSIM indicate that the model has stronger capabilities in retaining image structure and texture details, and the lower SAM reflects that the model maintains excellent spectral consistency for the de-clouded image.

[0150] The present invention mainly uses the improved UnCRtainTS deep learning network model combined with wavelet transform for de-clouding, aiming to solve the problem of poor de-clouding effect of a single deep learning model. The results show that the de-clouding method using the improved UnCRtainTS deep learning network model combined with wavelet transform significantly reduces the appearance of clouds in terms of visual effect and performs better in terms of RMSE, PSNR, SSIM, and SAM during quantitative analysis, indicating that the improved model combined with the physical method has significant potential in dealing with cloud interference.

Claims

1. An UnCRtainTS cloud removal method integrated with wavelet transform, the method comprising: Step 1: Classify the image according to the characteristics of the image, classify the fog and thin clouds as category one, and classify the thick clouds as category two; if the image is classified as category one, go to step 2; if the image is classified as category two, go to step 8; Step 2: Input the image into the UnCRtainTS encoder to extract preliminary features; Step 3: Input the preliminary features into the max-pooling layer to further reduce the resolution and enhance important features, and then strengthen the feature expression through the convolutional layer; Step 4: The features then enter the L-TAE module to capture the global information of the time series, and the L-TAE module represents the long-term dependence attention encoding module; Step 5: After the L-TAE module, the data is input into the multi-head self-attention mechanism of the Transformer to enhance the correlation representation between different time steps; Step 6: Perform an upsampling layer to restore the image size, and then perform temporal aggregation on the upsampled image and the preliminary features output by the UnCRtainTS encoder to obtain optimized time series features; Step 7: Pass the optimized time series features through the UnCRtainTS decoder to generate the cloud-removed image output; Step 8: Decompose the image through wavelet analysis to separate the low-frequency component and the high-frequency component; Step 9: Process the high-frequency component according to the method from step 2 to step 7 to output the high-frequency cloud-removed image; Step 10: Combine the high-frequency cloud-removed image obtained in step 9 with the low-frequency component obtained in step 8, and then perform inverse wavelet transform to generate the final cloud-removed image.

2. The UnCRtainTS cloud removal method integrating wavelet transform according to claim 1, characterized in that, The specific method of step 3 is: Feature f for each time step t Perform downsampling. Assume the downsampled feature is represented as Its size is d m Indicates the dimension of the downsampled feature, i.e., the length of the feature vector, Indicates the height of the downsampled feature, Indicates the width of the downsampled feature; 3. The UnCRtainTS cloud removal method integrating wavelet transform according to claim 1, characterized in that, The image classification method in step 1 includes: Method 1: Manual classification; Method 2: Train a classification model with labeled images, use the convolutional layer of the classification model for feature extraction, and use the fully connected layer for classification.

4. The UnCRtainTS cloud removal method integrating wavelet transform according to claim 1, characterized in that The implementation method of the Transformer multi-head self-attention mechanism in step 5 is: Suppose the input sequence is X = [x1, x2, …, x n , where each x i represents an element in the sequence. To generate the query key and value vectors for each position, first use a linear transformation to map X to the query Q, key K, and value V vector spaces respectively. W Q , W K , and W V are the weight matrices for the query, key, and value; Q = XW Q , K = XW K , V = XW V The attention weight matrix is determined by calculating the dot product of the query and the key, and the specific formula is as follows: where d k is the dimension size of the query and key vectors, and QK T represents the product of the transposed matrices of Q and K; Suppose h attention heads are used, then the calculation formula for the i-th head head i is as follows: Here, is the weight matrix that is independent for each head; the output of multi-head attention is the concatenation of the results of all heads, i.e.: MultiHead(Q,K,V)=Concat(head1,…,head h )W O Among them, Concat means concatenating the outputs of all heads in the feature dimension, and W O is the weight matrix of the output, which is used to map the concatenated representation back to the original dimension.

5. The UnCRtainTS cloud removal method integrating wavelet transform according to claim 1, characterized in that, The method of time aggregation in step 6 includes three main gates: forget gate f t , input gate i t , output gate o t , candidate memory unit Memory unit update c t ; The calculation method of temporal aggregation is: h t = o t ⊙tanh(c t ) Among them, the output at time step t is denoted as h t , σ represents the sigmoid activation function, ⊙ represents element-wise multiplication, and h t-1 is the aggregated feature of the previous time step, and W C , W i , W o , W f represent the weights of the corresponding units, and b i , b f , b o , b C are the biases of the corresponding units; tanh represents the hyperbolic tangent function, ensuring that the candidate memory values are between (-1, 1); h t is the aggregated feature of the current time step, and U i represents the weight matrix connecting the input gate to the hidden state of the previous time step; U f represents the weight matrix connecting the forget gate to the hidden state of the previous time step; U o represents the weight matrix connecting the output gate to the hidden state of the previous time step; U c represents the weight matrix connecting the candidate memory unit to the hidden state of the previous time step; These weight matrices are combined with the input features at the current time step through weights W i , W f , W o , W C , and are adjusted through bias terms b i , b f , b o , b C to calculate the activation values of each gate and the values of the candidate memory units; these activation values and the values of the candidate memory units are then used to update the internal state and output of the LSTM unit.

6. The UnCRtainTS cloud removal method integrating wavelet transform according to claim 1, characterized in that The wavelet transform method in step 8 is: Decompose the image into a low-frequency component L and multi-level high-frequency components H using the Daubechies wavelet basis LH , H HL , H HH , H LH , H HL , H HH respectively represent the high-frequency components in the horizontal, vertical, and diagonal directions.

7. The UnCRtainTS cloud removal method integrating wavelet transform according to claim 1, characterized in that Forgotten gate f t Controls the forgetting ratio of the memory cell information at the previous moment at the current moment. Let W f and b f be the weight matrix and bias of the forgotten gate respectively, h t-1 is the hidden state of the previous time step, x t is the input of the current time step, and σ represents the Sigmoid activation function; the calculation formula of the forgotten gate is: f t = σ(W f · [h t-1 , x t + b f ) According to the features of the current input and the previous hidden state, the forgetting gate determines the information to be discarded, thereby removing the noise features that cause inaccurate cloud removal.

8. The UnCRtainTS cloud removal method integrating wavelet transform according to claim 1, characterized in that, Input gate i t Controls the degree to which the new information at the current moment is updated by the memory unit, and the calculation formula is as follows: i t = σ(W i · [h t-1 , x t + b i ) Among them, W i and b i are the weights and biases of the input gate. The input gate controls the value of i t so that the time aggregation method selectively introduces new image features according to the current input information.

9. The UnCRtainTS cloud removal method integrating wavelet transform according to claim 1, characterized in that Output gate o t Controls the amount of information retrieved from the current memory cell and updates the hidden state to be used as the input for the next time step; the calculation formula is: o t = σ(W o · [h t-1 , x t + b o ) W in the formula o and b o are the weights and biases of the output gate.