Processing method and device for cloud removal and image restoration of remote sensing image data

By adopting a neural network architecture combining attention mechanism and generative adversarial network in remote sensing image processing, combined with a multi-layer convolution module of MB Conv block and SwinTransformer block, the accuracy and detail loss of cloud removal and image recovery in remote sensing images are solved, and efficient and accurate cloud removal and high-quality image recovery are achieved.

CN120031729APending Publication Date: 2025-05-23ZHEJIANG UNIV OF TECH

Patent Information

Application Number
CN202510092915.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art has problems with accuracy and loss of detail in cloud removal and image recovery in remote sensing images, especially when dealing with complex surface types.

Method used

Adopting a neural network architecture based on attention mechanism and generative adversarial network, combining MB Conv blocks and SwinTransformer blocks, the potential association of multi-time sequence data is extracted through spatiotemporal encoding strategies to achieve cloud removal and image recovery.

Benefits of technology

Efficient and accurate cloud removal is achieved, and high-quality cloud-free images are generated, significantly improving image quality and detail fidelity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031729A_ABST
    Figure CN120031729A_ABST
Patent Text Reader

Abstract

The invention discloses a processing method and device for cloud removal and image recovery of remote sensing image data, and the method comprises the steps: 1, constructing a multi-time-sequence multi-mode remote sensing image data set, and carrying out the standardization preprocessing; 2, designing a spatio-temporal feature extraction network model based on a self-attention mechanism in combination with a generative adversarial network, including a multilayer hybrid convolution module used for spatial feature coding; the lightweight time self-attention encoder module is used for extracting multi-time sequence correlation characteristics of each pixel position; the time sequence feature aggregation module is used for generating spatial-temporal feature representation; the Markov discriminator module is used for alleviating the blurring problem during image recombination and improving the image generation quality; 3, training a network model by using the multi-time-sequence multi-modal data set; 4, realizing cloud removal and image reconstruction based on the training model, and generating a cloudless image; and 5, inlaying the reconstructed optical image to form a complete cloudless remote sensing image. According to the method, deep learning and multi-modal data analysis are combined, efficient spatial-temporal feature aggregation and cloud removal are realized through a self-attention mechanism, the image quality is remarkably improved, and the method is suitable for remote sensing application scenes with frequent clouds, such as surface monitoring and environment evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image processing, and in particular relates to a processing method and device for cloud removal and image restoration of remote sensing image data. Background Art

[0002] In remote sensing image data processing, the presence of clouds and their shadows often significantly interferes with image quality and subsequent data analysis, especially in applications such as multi-temporal image analysis and ground feature change monitoring. Cloud occlusion can lead to data loss, significantly reducing the accuracy and reliability of analysis. Therefore, how to effectively remove cloud interference in remote sensing images and accurately restore the image information obscured by clouds has become a key issue that needs to be urgently addressed in the field of remote sensing image data processing.

[0003] Traditional cloud removal methods usually rely on differences in spectral features for cloud detection, identifying and removing cloud areas by analyzing the differences in reflectance spectral characteristics between clouds and ground objects. Such methods are usually based on multispectral or hyperspectral image data, combined with color index and threshold segmentation techniques to achieve cloud detection. However, these methods are highly sensitive to lighting conditions and surface characteristics, and have poor adaptability to different surface types. In addition, when the cloud layer is thick or the occlusion range is large, it is difficult to accurately restore the information of the occluded area, especially when dealing with complex surface types (such as farmland and wetlands). Therefore, traditional cloud removal methods still have great limitations.

[0004] In order to improve the accuracy of cloud removal, cloud detection and removal methods based on machine learning have gradually emerged. These methods can achieve relatively accurate cloud segmentation by extracting the texture and spectral characteristics of clouds through model training. However, traditional machine learning methods are highly dependent on training data and have limited ability to recover information in cloud-occluded areas. In addition, super-resolution reconstruction and detail enhancement techniques have also been gradually introduced into the image restoration process to further improve the spatial detail performance of the image. However, the reconstruction accuracy of these technologies is restricted by the network structure design and data resolution, which may lead to detail loss or artifact generation.

[0005] In recent years, with the rapid development of deep learning technology, cloud removal and image restoration methods based on deep convolutional neural networks have shown great potential. These methods deeply mine image feature information through deep network architectures, and can accurately restore the texture and spatial details of the occluded area while removing cloud interference. For example, generative adversarial networks (GANs) can model the distribution of cloud and non-cloud images to generate restored images with high authenticity and detail integrity. However, such methods may still have problems such as detail loss and artifacts, and the reliability of images needs to be further improved. In addition, cloud removal methods based on deep learning still face many challenges in practical applications: for example, model training is highly dependent on high-quality labeled data, the complexity of the network structure significantly increases the demand for computing resources; and when processing high-resolution remote sensing images, the ability to restore details is still insufficient. Therefore, there is an urgent need for an efficient and accurate cloud removal and image restoration technology to meet the needs of complex application scenarios. Summary of the invention

[0006] The present invention aims to overcome the above-mentioned deficiencies in the prior art and provide a processing method and device for cloud removal and image restoration of remote sensing image data.

[0007] The present invention can efficiently and accurately remove cloud interference and achieve high-quality image reconstruction. Based on the neural network architecture of the attention mechanism combined with the generative adversarial network, the present invention fully explores the potential correlation of multi-time series data through the spatiotemporal coding strategy, without the need for manual annotation of data sets, thereby significantly simplifying the cumbersome process of sample annotation. In addition, by optimizing parameter settings and reducing computational complexity, the present invention effectively alleviates the challenges brought about by high computing resource requirements and provides a more feasible solution for practical applications.

[0008] The core process of the present invention includes the following steps: First, optical images and radar images are collected and standardized preprocessed to ensure the spatiotemporal consistency of multimodal data. Subsequently, a multi-layer convolution module combining MB Conv blocks and SwinTransformer blocks is used to extract multi-scale features on each frame of the time series, while maintaining the feature map of full resolution to retain the spatial detail information to the greatest extent. Next, the multi-time series feature maps are aggregated by the lightweight temporal self-attention (L-TAE) module to generate a unified comprehensive feature representation, thereby capturing the dynamic change relationship between clouds and ground objects. After that, the decoder is used to decode the aggregated feature representation layer by layer to generate high-quality cloudless images and restore the details and textures obscured by clouds. Finally, the VGG19 pre-trained model is used to calculate the perceptual loss between the generated image and the corresponding cloud-free image. The perceptual loss function is added to the loss function as a new loss function item to constrain the generator together. The discriminator uses a Markov discriminator, which can reduce the blurring problem during image reconstruction and improve the image generation quality.

[0009] The first aspect of the present invention provides a processing method for cloud removal and image restoration of remote sensing image data, and the specific implementation steps are as follows:

[0010] Step 1: Construct a multi-temporal and multi-modal remote sensing image dataset;

[0011] Select the study area, determine the geographical scope, time range and spatial resolution requirements; obtain the optical image data O (including 13 single bands such as red, green, blue and near infrared bands) and radar image data S of the study area. Divide the study area into several blocks, obtain and process remote sensing data in each block, and form a multi-time series and multi-modal image data set. The specific data processing process is as follows:

[0012] Step 1.1: Multispectral optical image data preprocessing;

[0013] Step 1.1.1: Data acquisition;

[0014] Acquire optical image data containing thirteen single bands, each single band image O i Represented as an m×n matrix, where element o i,j Represents the value of the image at the i-th row and j-th column.

[0015] Step 1.1.2: Data stacking;

[0016] The thirteen single-band images are stacked in the order of bands as slices of the third dimension to form a fused three-dimensional tensor P, which is defined as follows:

[0017] P ijk =O k(i,j), where k = 1, 2, ..., 13 (1)

[0018] Among them, P ijk Represents the pixel value at the i-th row and j-th column of the k-th band in the fused data P, which comes from the original single band O k The value of (i,j).

[0019] Step 1.1.3: Image output;

[0020] The fused image is cropped to a size of 256×256 suitable for the image input, and the resolution is resampled to 10m.

[0021] Step 1.2: Radar image data preprocessing;

[0022] Step 1.2.1: Data acquisition;

[0023] Get the original GRDH format image data from Copernicus DataHub, denoted as S raw .

[0024] Step 1.2.2: Orbit correction;

[0025] The original data is corrected using the precise orbit file to reduce the satellite position error. The corrected data S orbit It is expressed as:

[0026] S orbit =f orbit (S raw ,Precise Orbit) (2)

[0027] Among them, f orbit is the correction function, and Precise Orbit is the precise orbit file.

[0028] Step 1.2.3: Thermal noise removal;

[0029] By estimating the thermal noise value N thermal , remove the thermal noise in the data and get S thermal , which can be expressed as:

[0030] S thermal =S orbit -N thermal (3)

[0031] Among them, N thermal is the thermal noise value obtained by estimation.

[0032] Step 1.2.4: edge noise removal;

[0033] Remove the image edge noise and obtain the processed intermediate data S border , and its calculation formula is:

[0034] S border =f border (S thermal ) (4)

[0035] Among them, f border is a function used for edge noise removal.

[0036] Step 1.2.5: Calibration;

[0037] The radar reflection value is normalized to the physical measurement value through correction, and the calibrated data S correct It is expressed as:

[0038] S correct =K·S border (5)

[0039] where K is a calibration factor (depending on the characteristics of the satellite and radar system).

[0040] Step 1.2.6: Terrain correction;

[0041] The digital elevation model (DEM) is used to perform terrain geometry correction on radar images. terrain It is expressed as:

[0042] S terrain =f terrain (S filtered ,DEM) (6)

[0043] Among them, f terrain It is a geometric correction function based on terrain.

[0044] Step 1.2.7: Convert to logarithm;

[0045] Convert the corrected radar reflectance values ​​to logarithmic (dB) units for easier comparison and analysis. dB for:

[0046] S dB =10log 10 (S terrain ) (7)

[0047] Step 1.2.8: Image output;

[0048] The radar image was cropped to a size of 256 × 256 suitable for the image input and the resolution was resampled to 10 m.

[0049] Finally, the dataset obtained after processing in step 1 is: at each observation time point, each region of the dataset contains a tuple consisting of co-aligned two-band Sentinel-1 radar measurements and thirteen-band Sentinel-2 optical observations, and each region in the dataset has thirty repeated observations.

[0050] Step 2: Design a spatiotemporal feature extraction network model based on the self-attention mechanism;

[0051] The first part is the encoder, which is a convolutional module composed of multiple MB Conv blocks and Swin Transformer blocks, used to extract the basic features of the image; the second part is the attention-based temporal information aggregation module, which generates an attention mask for aggregating multi-time series features in the time dimension; the third part is the decoder, which is used to decode the aggregated feature map into the spatial resolution output of the target image; the fourth part is the discriminator, which is used to reduce the blur problem during image reconstruction and improve the quality of image generation.

[0052] Step 2.1: Construct a multi-layer mixed convolution module;

[0053] This module combines MBConv blocks and Swin Transformer blocks to extract local and global features of the input feature map.

[0054] MBConv block: This module is an efficient convolution operation that extracts and expresses spatial features through deep convolution and expansion factors. Specifically, this module receives the input feature map T∈R H*W*C , and generate rich spatial features through operations such as expansion and point-by-point convolution, and output feature F MBConv .

[0055] Swin Transformer block: This module uses sliding windows and self-attention mechanism to capture global features. Specifically, after receiving the input feature map, the module applies a multi-head self-attention mechanism to calculate the global correlation feature F of each pixel. swin The core process is:

[0056]

[0057] Among them, Q, K and V are query, key and value matrices respectively, and N is the number of pixels in the window.

[0058] Feature fusion convolution layer: After the MBConv and Swin Transformer blocks complete feature extraction, the output features F of the two are concatenated in the channel dimension through torch.cat() MBConv and F swin, and obtain a feature map that contains both spatial and global information. After that, the concatenated features are reduced in dimension through a 1×1 convolutional layer to achieve feature fusion. The calculation formula is:

[0059] F fusion =Conv 1×1 (concat(F MBConv , F swin )) (10)

[0060] Normalization and residual connection: Normalize the fused features to enhance the stability of training. Then, through the residual connection operation, the fused features F fusion Directly add it to the input feature X to realize the skip connection. The formula is:

[0061] X out =F fusion +X (11)

[0062] Among them, X is the input feature, X out is the output feature.

[0063] Step 2.2: Build a lightweight temporal self-attention (L-TAE) module;

[0064] Multi-head self-attention mechanism: The L-TAE module performs weighted processing on multi-frame features through a multi-head self-attention mechanism. For each pixel position (i, j), the attention weight is calculated using the query matrix Q', key matrix K' and value matrix V', and the formula is:

[0065]

[0066] Where W Q' ,W K' ,W V' is the linear transformation matrix, d K' is the dimension of the key vector. Attention weight Indicates the importance of time frame t at position (i, j).

[0067] Temporal aggregation: Based on the calculated attention weights, L-TAE performs weighted summation of all time frames to generate a temporal aggregation feature map The formula is:

[0068]

[0069] Feature map after aggregation Combining the temporal information of multiple frames can effectively capture the changes between clouds and ground objects.

[0070] Position encoding: In order to maintain the order information of the time series, L-TAE adds position encoding PE(t) to each frame feature, and its calculation formula is:

[0071]

[0072] Among them, PE(t) is the time step position encoding vector generated by trigonometric function, which is used to distinguish data from different time frames.

[0073] Step 2.3: Build the decoder module;

[0074] The decoder module aggregates feature maps through multiple decoding blocks The decoding is performed step by step. The decoding block is composed of a mixture of MBConv and Swin Transformer blocks to restore the spatial details of the image layer by layer. Assume that the output of the decoder layer l is g (l) , so the feature decoding process is:

[0075] g (l+1) =DecoderBlock(g (l) ) (16)

[0076] After completing layer-by-layer decoding, the decoder generates the final cloud-free image through a convolutional layer.

[0077]

[0078] Among them, σ is the Sigmoid activation function, which is used to limit the output to the range of [0,1].

[0079] Step 2.4: Build the identifier module;

[0080] The Markov Discriminator can be used to identify the authenticity of a local patch block with an N×N receptive field, without affecting the independence of other elements outside the patch block. Each patch only depends on the patch before it, and has nothing to do with the patch before it. This patch independence allows the Markov discriminator to reconstruct the image into a Markov random field, which can describe the features through the local area around each pixel. In terms of the size of the patch, the size can be smaller than the original image. Such a small patch size can produce fewer parameters than a large patch block, which can effectively reduce the resource usage of the network.

[0081] Step 3: Train the deep learning network model, including:

[0082] The multi-temporal and multi-modal remote sensing image dataset produced in step 1 is used as the input of the deep learning network model constructed in step 2. After setting the hyperparameters, the model is trained using the gradient descent algorithm until the loss function converges and a stable result is obtained.

[0083] Step 4: Reconstruct the cloud-removed image based on the trained model, including:

[0084] The multi-temporal and multi-modal data of the target area are taken as input, and the trained model is loaded to complete the image reconstruction of the target area. Specifically, the encoder is used to extract the spatial features of each frame of the image; the lightweight temporal self-attention module aggregates multi-frame information to capture the dynamic changes of clouds and ground objects; the decoder reconstructs the image based on the aggregated features to generate a clear image without clouds; the image quality is then optimized through the identifier; finally, the model outputs a high-quality cloud-free image V in a block form to support subsequent analysis and application.

[0085] Step 5: mosaicking and splicing the reconstructed optical images to form a complete cloud-free remote sensing image, including:

[0086] Step 5.1: Image preprocessing: Since the image data V is processed in blocks, direct splicing may affect the quality of the final image. Therefore, the image needs to be preprocessed, including cropping, adjusting brightness, contrast, and color balance.

[0087] Step 5.2: Image alignment: Use feature point extraction and matching to determine the correspondence between images and complete the image alignment operation.

[0088] Step 5.3: Image fusion: Use the fusion algorithm to merge the aligned images and adjust the color balance, brightness and contrast of the overall image to ensure the overall quality of the stitched image is good.

[0089] Finally, the mosaic results are output to obtain the cloud-free multispectral remote sensing image data of the entire study area.

[0090] The second aspect of the present invention relates to a device for cloud removal and image restoration of remote sensing image data, comprising a memory and one or more processors, wherein the memory stores executable code, and the one or more processors, when executing the executable code, are used to implement the processing method for cloud removal and image restoration of remote sensing image data of the present invention.

[0091] A third aspect of the present invention relates to a computer-readable storage medium having a program stored thereon, which, when executed by a processor, can implement the processing method of the present invention for cloud removal and image restoration of remote sensing image data.

[0092] The advantages of the present invention are:

[0093] 1) Based on a deep learning network framework, the present invention adopts an encoder-decoder structure combined with a generative adversarial network structure, which can efficiently remove cloud interference in satellite remote sensing images and significantly improve image quality.

[0094] 2) The present invention innovatively integrates the hybrid feature extraction structure of MBConv and Swin Transformer, which not only significantly reduces the computational complexity through the lightweight convolution operation of the MBConv module, but also uses the local window self-attention mechanism of Swin Transformer to enhance the capture of spatial context information, thus achieving a balance between feature extraction efficiency and accuracy.

[0095] 3) The present invention introduces observation data at different times as a reference to deeply mine the potential information of multi-time series data, greatly improving the reliability of the method in fields such as dynamic monitoring.

[0096] 4) The present invention fully combines optical remote sensing data and radar image data, and uses multimodal information to enhance the accuracy and detail fidelity of image reconstruction. The generated images have higher credibility and applicability and are suitable for complex application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0097] Attached Figure 1 is a flow chart of the method of the present invention.

[0098] FIG2 is an example diagram of the training data set of the present invention, from left to right are Figure 2a Radar images, Figure 2b Cloud cover optical data and Figure 2c Cloud removal optical prediction results.

[0099] Attached Figure 3 It is a flow chart of the network model of the de-clouding method in the present invention.

[0100] Attached Figure 4 It is a parameter index diagram in the present invention.

[0101] Attached Figure 5 It is the regional image to be processed in the preferred embodiment of the present invention.

[0102] Attached Figure 6 It is an image diagram of the reconstruction result of the region in the preferred embodiment of the present invention.

[0103] Attached Figure 7 It is a schematic diagram of the device of the present invention. DETAILED DESCRIPTION

[0104] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0105] Example 1

[0106] Figure 1 It shows a flow chart of a processing method for cloud removal and image restoration of remote sensing image data. The method of the present invention is further described in detail below with respect to each step in the embodiment process.

[0107] Step 1: Construct a multi-temporal and multi-modal remote sensing image dataset;

[0108] Determine the study area, select No. 2 data for optical remote sensing data, select No. 1 data for radar image data, select MGRS grid as the range of the area, select the time from January 2019 to December 2021, and select thirteen bands including B02, B03, B04, B08 and the No. 1 radar satellite data corresponding to almost the same time.

[0109] Step 1.1: Multispectral optical image data preprocessing;

[0110] Step 1.1.1: Data acquisition;

[0111] First, obtain thirteen single-band images, including bands B02, B03, B04, B08, etc. Store these image files in .tif format, each image represents a band, and the image size is 10980×10980. The file names can be stored as band_1.tif to band_13.tif, and all bands are stored in a List, recorded as bands, with a shape of (13, 10980, 10980).

[0112] Step 1.1.2: Data stacking;

[0113] Thirteen single-band images are stacked into a three-dimensional matrix (tensor). In this three-dimensional matrix, the first dimension represents the band, and the last two dimensions represent the rows and columns of the image. The multispectral image P can be obtained through the stack operation in numpy. P is a three-dimensional matrix of (13, 10980, 10980).

[0114] Step 1.1.3: Image output;

[0115] Finally, the multispectral image P is cropped and resampled to obtain several small images with 13 bands, 256×256 size, and 10m resolution as part of the feature input.

[0116] Step 1.2: Radar image data preprocessing;

[0117] Step 1.2.1: Data acquisition;

[0118] First, download the original GRDH format radar image data from Copernicus Data Hub. This data is recorded as S_raw, and you need to use rasterio or other related tools to read the GRDH format data.

[0119] Step 1.2.2: Orbit correction;

[0120] In this step, the original data is corrected using an accurate orbit file. This file is usually provided together with the image data. Then, the orbit file data is read and applied to the image for correction. Formula (2) is applied to obtain the corrected S_orbit.

[0121] Step 1.2.3: Thermal noise removal;

[0122] Thermal noise is a common noise in radar data and can usually be removed by estimating the thermal noise value. Here we can use the mean of S_orbit as an estimate of the noise and apply formula (3) to obtain S_thermal after thermal noise removal.

[0123] Step 1.2.4: edge noise removal;

[0124] Edge noise usually appears around the radar image and may be related to the generation of the image. Here, the noise can be removed by simple edge filtering, such as using Gaussian blur to smooth the image and reduce edge noise. Applying formula (4), we can obtain S_border after edge noise removal.

[0125] Step 1.2.5: Calibration;

[0126] A constant calibration coefficient is needed to convert the radar reflection value into a physical measurement value. The calibration coefficient is usually calculated based on the characteristics of the satellite system. The No. 1 SAR product usually chooses 10log10. Applying formula (5), we get S_correct.

[0127] Step 1.2.6: Terrain correction;

[0128] The data is geometrically corrected using the digital elevation model and formula (6) is applied to obtain S_terrain.

[0129] Step 1.2.7: Convert to logarithm;

[0130] Apply formula (7) to S_terrain obtained in the above steps to obtain S_db.

[0131] Step 1.2.8: Image output;

[0132] Finally, the radar image S_db is cropped and resampled to obtain several two-band, 256×256-size, 10m-resolution small images as part of the feature input.

[0133] Finally, the dataset obtained after step 1 is: each time point of observation is a tuple, consisting of co-registered 2-band Sentinel-1 radar measurements and 13-band Sentinel-2 optical observations. Each area in the dataset has 30 repeated measurements. The sample image of the constructed dataset is shown in Figure 2.

[0134] Step 2.1: Build the encoder module;

[0135] The main task of the encoder is to extract features at each time step of the input image sequence, retain spatial information and adjust the channel dimension. The shape of the input data is [T, Cin, H, W], specifically [3, 15, 256, 256], where T represents the input time series, Cin represents the number of input bands, H represents the height of the input image, and W represents the width of the input image. First, a 1×1 convolution is used to change the number of input channels from 15 to 128, thereby compressing information and improving computational efficiency. The output shape after point convolution is [3, 128, 256, 256]. Next, the data is mixed through multiple MBConv blocks and Swin Transformer modules to further encode the spatial information of each time step image, while retaining the input resolution. First, the number of channels is expanded from 128 to 256, and then reduced back to 128, and the input and output shapes are both [3, 128, 256, 256]. The final output feature map shape is [T, dm, H, W], that is, [3, 128, 256, 256].

[0136] Step 2.2: Build a lightweight temporal self-attention module;

[0137] The attention mechanism module (Temporal Aggregation) is used to integrate time series information and aggregate feature information at different time steps into a single feature map, thereby helping to remove clouds and improve the accuracy of remote sensing images. First, the module performs a maximum pooling operation on the feature map of each time step, reducing the resolution from 256×256 to 32×32 to reduce the amount of computation and aggregate similar features. Next, the number of channels of the downsampled feature map is expanded from 128 to 256 through a linear layer. Then, the L-TAE attention mechanism is applied to encode the temporal dimension on the downsampled feature map. Attention propagation is performed on the time series of each spatial position to focus on the effective information at a specific time step to remove clouds. On this basis, the attention encoding is upsampled back to the original resolution of 256×256 and applied to the high-resolution feature map of each time step. Finally, the temporal dimension is aggregated into a single feature map with an output shape of [128, 256, 256]. Through this process, the module effectively aggregates the time series information to generate a high-precision feature map after cloud removal.

[0138] Step 2.3: Build the decoder module;

[0139] The decoder module is responsible for reconstructing the aggregated feature maps into the final cloud-free image. The input feature map shape is [dm,H,W], i.e. [128,256,256].

[0140] First, multiple hybrid blocks (MBConv blocks and Swin Transformer modules) are used to further decode spatial details. The output resolution is kept unchanged, so the input and output shapes are both [128, 256, 256]. Next, a 1×1 point convolution is used to reduce the number of channels from 128 to 13. The shape of the feature map after dimensionality reduction is [13, 256, 256].

[0141] Subsequently, an activation function is applied to the reduced output to obtain the final reconstruction result. Using the sigmoid activation function, the output is normalized to the image pixel range of [0,1].

[0142] The final output is a reconstructed cloud-free image with a shape of [Cout, H, W], i.e. [13, 256, 256], where Cout represents the number of output bands. Through this process, the decoder module generates a cloud-free image, providing a higher guarantee for the quality and reliability of remote sensing images.

[0143] Step 2.4: Build the identifier module;

[0144] The Markov Discriminator can be used to identify the authenticity of a local patch block with an N×N receptive field, without affecting the independence of other elements outside the patch block. Each patch only depends on the patch before it, and has nothing to do with the patch before it. This patch independence allows the Markov discriminator to reconstruct the image into a Markov random field, which can describe the features through the local area around each pixel. In terms of the size of the patch, the size can be smaller than the original image. Such a small patch size can produce fewer parameters than a large patch block, which can effectively reduce the resource usage of the network.

[0145] Finally, construct Figure 3 The network model structure.

[0146] Step 3: Deep learning network model training;

[0147] Use the dataset produced in step 1 as the input of the network model built in step 2, and set the hyperparameters: learning rate is 0.001, total batch is 20, batch size is 4, momentum is 0.9, and weight decay is 0.0001.

[0148] The result is predicted through the set model framework, and then the loss value of the actual label and the predicted result is calculated according to the loss function. Then the reverse gradient propagation algorithm is used, and the model parameters are adjusted according to the learning rate. The deep network model is trained iteratively until the result obtained by the loss function tends to be stable. At this time, the network has converged and the accuracy of the model in the sample set is calculated. According to different accuracy requirements, the hyperparameters are adjusted using general methods to obtain the current optimal model. The model evaluation parameter indicators are as follows: Figure 4 shown.

[0149] Step 4: Reconstruct the cloud-removed image based on the trained model;

[0150] When applying the trained model to image reconstruction of the target area, we first need to prepare multi-time series data of the area to be reconstructed as input, such as Figure 5 As shown. Input these data into the trained model, the model will use its encoder to extract the spatial features of each frame of the image, and aggregate multiple frames of information through a lightweight temporal self-attention module to capture the dynamic changes of clouds and ground objects. Then, the decoder reconstructs the image based on the aggregated features to generate a clear image without clouds. Finally, the image quality is jointly controlled by the Markov discriminator and the perceptual loss generated by VGG19. This process will output a high-quality cloud-free image, which is helpful for subsequent analysis and application, as shown below. Figure 6 The reconstruction result image.

[0151] Step 5: Mosaic the time series data of multispectral images;

[0152] After the 10m resolution multispectral image dataset of each MGRS grid area is processed according to the above process, the images of multiple MGRS grid areas across Xinjiang are spliced ​​into seamless spatial continuous images according to the following image mosaic steps, and finally a 10m resolution multispectral cloud-free remote sensing image of the target area is obtained.

[0153] Step 5.1: Since the image data after model processing is processed in blocks, if it is directly spliced, it will affect the quality of the final image. Therefore, it is necessary to pre-process it first, including cropping, adjusting brightness, contrast, and color balance.

[0154] Step 5.2: Use feature point extraction and matching to determine the correspondence between images and perform image alignment.

[0155] Step 5.3: Merge the aligned images through the fusion algorithm, adjust the overall color balance, brightness and contrast to ensure the overall quality of the stitched image, and finally output the mosaic result to obtain the remote sensing cloud-free image data of the entire study area.

[0156] The present invention provides a processing method and device for cloud removal and image restoration of remote sensing image data, aiming to solve the influence of cloud interference on optical image quality. The method comprises: 1. constructing a multi-temporal multi-modal remote sensing image data set and performing standardized preprocessing; 2. designing a spatiotemporal feature extraction network model based on a self-attention mechanism combined with a generative adversarial network, including a multi-layer hybrid convolution module for spatial feature encoding; a lightweight temporal self-attention encoder module for extracting multi-temporal correlation features of each pixel position; and a temporal feature aggregation module for generating spatiotemporal feature representation; a Markov discriminator module for alleviating the blurring problem during image reorganization; 3. training the network model using a multi-temporal multi-modal data set; 4. realizing cloud removal and image reconstruction based on the training model to generate a cloud-free image; 5. mosaicking the reconstructed optical image to form a complete cloud-free remote sensing image. The present invention combines deep learning with multi-modal data analysis, realizes efficient spatiotemporal feature aggregation and cloud removal through a self-attention mechanism, significantly improves image quality, and is suitable for remote sensing application scenarios with frequent cloud layers such as surface monitoring and environmental assessment.

[0157] Example 2

[0158] Reference Figure 7, this embodiment relates to a device for cloud removal and image restoration of remote sensing image data, including: a memory for storing executable program code, and the program code implements the method described in Embodiment 1; one or more processors for executing the executable program code in the memory to implement cloud removal and image restoration of remote sensing image data. In addition, the device includes, but is not limited to, a single-machine deployed embedded device, a distributed computing system, or a cloud computing platform.

[0159] Embodiment 3

[0160] This embodiment relates to a computer-readable storage medium, on which a program is stored, and the program can implement the method described in Embodiment 1 when executed by a processor. The program supports local deployment, distributed computing, or running through a cloud computing platform to achieve cloud removal and image restoration of remote sensing image data.

[0161] The above is only a description of the embodiments of the present invention, but the protection scope of the present invention should not be regarded as limited to the specific forms stated in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art according to the concept of the present invention.

Claims

1. A method for cloud removal and image restoration of remote sensing image data, characterized in that: This is accomplished using the following steps: Step 1: Construct a multi-temporal and multi-modal remote sensing image dataset covering optical and radar images, and perform standardized preprocessing to improve data consistency; Step 2: Design a spatiotemporal feature extraction network model based on the self-attention mechanism combined with a generative adversarial network, the model includes: a multi-layer hybrid convolution module for extracting local and global spatial features of the image; a lightweight temporal self-attention encoder module for capturing spatiotemporal correlation features in multi-time series data; a time series feature aggregation module for generating a spatiotemporal feature representation that integrates multimodal and multi-time series features in a weighted manner; and a Markov discriminator for judging the quality of image generation. Step 3: Use the data set constructed in step 1 to train the deep learning network model and optimize the model parameters; Step 4: Remove clouds and reconstruct images based on the trained model to generate cloud-free images; Step 5: Mosaic the reconstructed optical images to form a complete cloud-free remote sensing image.

2. The method for cloud removal and image restoration of remote sensing image data according to claim 1, characterized in that: The step 1 of constructing a multi-temporal and multi-modal remote sensing image dataset covering optical radar images includes: Select the study area, determine the geographical scope, time range and spatial resolution requirements; obtain the optical image data O (including multi-band information) and radar image data S of the study area; standardize and crop the images to form a multi-temporal and multi-modal image data set that can be used for model input. The specific data processing process is as follows: Step 1.1: Multispectral optical image data preprocessing; Step 1.1.1: Data acquisition; Obtain optical image data including red, green, blue, near infrared bands, etc., a total of thirteen single-band bands. i Represented as an m×n matrix, where element o i,j Represents the value of the image at the i-th row and j-th column; Step 1.1.2: Data stacking; The thirteen single-band images are stacked in the order of the bands as slices of the third dimension to form a fused three-dimensional tensor P, which is defined as follows: P ijk =O k (i,j), where k = 1, 2, ..., 13 (1) Among them, P ijk Represents the pixel value at the i-th row and j-th column of the k-th band in the fused data P, which comes from the original single band O k The value of (i,j); Step 1.1.3: Image output; The fused image is cropped to a size of 256×256 suitable for the image input, and the resolution is resampled to 10m; Step 1.2: Radar image data preprocessing; Step 1.2.1: Data acquisition; Get the original GRDH format image data from Copernicus Data Hub, denoted as S raw ; Step 1.2.2: Orbit correction; The original data is corrected using the precise orbit file to reduce the satellite position error. The corrected data S orbit It is expressed as: S orbit =f orbit (S raw ,Precise Orbit) (2) Among them, f orbit is the correction function, Precise Orbit is the precise orbit file; Step 1.2.3: Thermal noise removal; By estimating the thermal noise value N thermal , remove the thermal noise in the data and get S thermal , which can be expressed as: S thermal =S orbit -N thermal (3) Among them, N thermal is the thermal noise value obtained by estimation; Step 1.2.4: edge noise removal; Remove the image edge noise and obtain the processed intermediate data S border , and its calculation formula is: S border =f border (S thermal ) (4) Among them, f border is a function used for edge noise removal; Step 1.2.5: Calibration; The radar reflection value is normalized to the physical measurement value through correction, and the calibrated data S correct It is expressed as: S correct =K·S border (5) Where K is the calibration factor (depending on the characteristics of the satellite and radar system); Step 1.2.6: Terrain correction; The digital elevation model (DEM) is used to perform terrain geometry correction on radar images. terrain It is expressed as: S terrain =f terrain (S filtered ,DEM) (6) Among them, f terrain It is a geometric correction function based on terrain; Step 1.2.7: Convert to logarithm; The corrected radar reflectance values ​​are converted to logarithmic (dB) units for easier comparison and analysis. The converted output S dB for: S dB =10log 10 (S terrain ) (7) Step 1.2.8: Image output; The radar image was cropped to a size of 256×256 suitable for the image input, and the resolution was resampled to 10m; Finally, the dataset obtained after processing in step 1 contains a tuple for each region of the dataset at each observation time point, consisting of co-registered two-band Sentinel-1 radar measurements and thirteen-band Sentinel-2 optical observations, and each region in the dataset has thirty repeated observations.

3. The method for cloud removal and image restoration of remote sensing image data according to claim 1, characterized in that: Step 2: Design a spatiotemporal feature extraction network model based on the self-attention mechanism combined with the generative adversarial network, including: The first part is the encoder, which uses a convolutional module composed of multiple MB Conv blocks and Swin Transformer blocks to extract the basic features of the image; the second part is the attention-based temporal information aggregation module, which generates an attention mask for aggregating multiple temporal sequence features in the temporal dimension; the third part is the decoder, which is used to decode the aggregated feature map into the spatial resolution output of the target image; the fourth part is the discriminator, which is used to reduce the blur problem during image reconstruction and improve the image generation quality; Step 2.1: Construct a multi-layer mixed convolution module; This module combines MBConv blocks and Swin Transformer blocks to extract local and global features of the input feature map; MBConv block: This module is an efficient convolution operation that extracts and expresses spatial features through deep convolution and expansion factors. Specifically, this module receives the input feature map T∈R H*W*C , and generate rich spatial features through operations such as expansion and point-by-point convolution, and output feature F MBConv ; Swin Transformer block: This module uses sliding windows and self-attention mechanisms to capture global features. Specifically, after receiving the input feature map, the module applies a multi-head self-attention mechanism to calculate the global correlation feature F of each pixel. swin , the core process is: Where Q, K and V are query, key and value matrices respectively, and N is the number of pixels in the window; Feature fusion convolution layer: After the MBConv and Swin Transformer blocks complete feature extraction, the output features F of the two are concatenated in the channel dimension through torch.cat() MBConv and F swin , and obtain a feature map that contains both spatial and global information; then, the concatenated features are reduced in dimension through a 1×1 convolutional layer to achieve feature fusion. The calculation formula is: F fusion =Conv 1×1 (concat(F MBConv ,F swin )) (10) Normalization and residual connection: The fused features are normalized to enhance the stability of training; then, the fused features F are connected through residual connection operation. fusion Directly add it to the input feature X to realize the skip connection. The formula is: X out =F fusion +X (11) Among them, X is the input feature, X out is the output feature; Step 2.2: Build a lightweight temporal self-attention (L-TAE) module; Multi-head self-attention mechanism: The L-TAE module performs weighted processing on multi-frame features through a multi-head self-attention mechanism. For each pixel position (i, j), the attention weight is calculated using the query matrix Q′, key matrix K′, and value matrix V′. The formula is: Among them, W Q′ ,W K′ ,W V′ is the linear transformation matrix, d K′ is the dimension of the key vector, the attention weight Indicates the importance of the time frame t at position (i, j); Temporal aggregation: Based on the calculated attention weights, L-TAE performs weighted summation of all time frames to generate a temporal aggregation feature map The formula is: Feature map after aggregation Combining multiple frames of temporal information, it can effectively capture the changes between clouds and ground objects; Position encoding: In order to maintain the order information of the time series, L-TAE adds position encoding PE(t) to each frame feature, and its calculation formula is: Among them, PE(t) is the time step position encoding vector generated by trigonometric function, which is used to distinguish data from different time frames; Step 2.3: Build the decoder module; The decoder module aggregates feature maps through multiple decoding blocks The decoding is performed step by step. The decoding block is composed of a mixture of MBConv and SwinTransformer blocks, which restores the spatial details of the image layer by layer. Assume that the output of the decoder layer l is g (l) , so the feature decoding process is: g (l+1) =DecoderBlock(g (l) ) (16) After completing layer-by-layer decoding, the decoder generates the final cloud-free image through a convolutional layer. Among them, σ is the Sigmoid activation function, which is used to limit the output to the range of [0,1]; Step 2.4: Build the identifier module; The Markov Discriminator can be used to identify the authenticity of a local patch block with an N×N receptive field, and will not affect the independence of other elements outside the patch block; each patch will only depend on the patch before it, and has nothing to do with the patch before it; this patch independence enables the Markov discriminator to reconstruct the image into a Markov random field, which can describe the features through the local area around each pixel; in terms of the size of the patch, the size can be smaller than the original image, so that a small patch size can produce fewer parameters than a large patch block, which can effectively reduce the resource usage of the network.

4. The method for cloud removal and image restoration of remote sensing image data according to claim 1, characterized in that: Step 3: Train the deep learning network model, including: The multi-temporal and multi-modal remote sensing image dataset produced in step 1 is used as the input of the deep learning network model constructed in step 2. After setting the hyperparameters, the model is trained using the gradient descent algorithm until the loss function converges and a stable result is obtained.

5. The method for cloud removal and image restoration of remote sensing image data according to claim 1, characterized in that: Step 4: Cloud removal and image reconstruction based on the trained model, including: The multi-temporal and multi-modal data of the target area are taken as input, and the trained model is loaded to complete the image reconstruction of the target area. Specifically, the encoder is used to extract the spatial features of each frame of the image. Multi-frame information is aggregated through a lightweight temporal self-attention module to capture the dynamic changes of clouds and ground objects. The aggregated features are input into the decoder for image reconstruction to generate a clear cloud-free image while retaining more detailed information about the objects. The image quality is then optimized through the identifier. Finally, the model outputs a high-quality cloud-free image V in a block form to support subsequent analysis and application.

6. The method for cloud removal and image restoration of remote sensing image data according to claim 1, characterized in that: The reconstructed optical images are mosaicked and spliced ​​as described in step 5 to form a complete cloud-free remote sensing image, including: Step 5.1: Image preprocessing: Since the image data V is processed in blocks, direct splicing may affect the quality of the final image. Therefore, the image needs to be preprocessed, including cropping, adjusting brightness, contrast, and color balance; Step 5.2: Image alignment: Use feature point extraction and matching to determine the correspondence between images and complete the image alignment operation; Step 5.3: Image fusion: Use the fusion algorithm to merge the aligned images and adjust the color balance, brightness and contrast of the overall image to ensure the overall quality of the stitched image is good; Finally, the mosaic results are output to obtain the cloud-free multispectral remote sensing image data of the entire study area.

7. A processing device for cloud removal and image restoration of remote sensing image data, characterized in that: include: A memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement a processing method for cloud removal and image restoration of remote sensing image data as described in any one of claims 1-6.

8. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, a processing method for cloud removal and image restoration of remote sensing image data as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Space-time fusion model method for remote sensing image cloud removal

    CN117408989A

  • Multispectral image cloud removal method based on improved conditional generative adversarial network

    CN117522737A

  • Image cloud removal method and system based on optical remote sensing image and SAR image

    CN117522738A

Cited By

  • Remote sensing image cloud removal method and system based on geographic coordinate embedding

    CN120495138A

  • Remote sensing image cloud removal method oriented to long time sequence cloud coverage

    CN120510055A

  • Remote sensing thick cloud image restoration method and system based on multi-scale feature fusion

    CN120782679A

  • Remote sensing image cloud classification method, system and device based on multi-modal data fusion and medium

    CN120913086A

  • Cloud detection method and device, electronic equipment and computer readable storage medium

    CN121033367A