Remote sensing image cloud removal reconstruction method and system

By constructing a cloud generation image and combining it with a multi-head self-attention and lateral connection upsampling structure, the feature extraction and fusion problems of existing remote sensing image declouding methods in complex cloud areas are solved, achieving efficient and accurate cloud removal and preservation of ground object details, and improving image quality.

CN120807309APending Publication Date: 2025-10-17SHANDONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510872938.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing remote sensing image declouding methods have difficulty in effectively extracting and fusing feature information at different scales when processing cloud areas with complex textures or blurred boundaries of objects, resulting in loss of detail information and blurred edges. In addition, they have high computational overhead and limited global dependency modeling capabilities, making it difficult to accurately restore long-distance relationships in occluded areas, resulting in texture distortion and object deformation in the reconstructed image.

Method used

By constructing cloud layer generation images, using multi-head self-attention processing and horizontally connected upsampling structure, combining local feature extraction and global feature extraction, fusing multi-scale feature maps, and adopting lightweight multi-head attention mechanism and shallow network structure, dynamic capture and detail restoration of clouds can be achieved.

Benefits of technology

The cloud removal effect of the cloud removal model is improved, texture distortion and ground object structure deformation in the reconstructed image are avoided, the authenticity and consistency of the image are improved, the ability to capture cloud structure and ground object boundary information is enhanced, and the computational overhead is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807309A_ABST
    Figure CN120807309A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of remote sensing image cloud removal processing, in particular to a remote sensing image cloud removal reconstruction method and system, and the method comprises the steps: screening three images with one cloudless image and two cloud images in a continuous time sequence for the same region, and forming a training set; performing down-sampling on each image in one sample in the training set to obtain a fourth feature map, reducing the number of channels of the fourth feature map to obtain an intermediate feature map, and performing transverse connection and nearest neighbor up-sampling processing to obtain a third aggregation feature map; and splicing all the third aggregated feature maps according to a channel direction to obtain a multi-scale feature map, performing feature extraction, fusing local and global feature maps to obtain a reconstructed image, and processing a to-be-processed cloud image by using the trained cloud removal model to obtain the reconstructed image. Multi-head self-attention processing is introduced in global feature extraction, the global relation between pixels can be accurately modeled on all scales, and the method is particularly suitable for processing remote sensing image scenes with irregular cloud layer distribution.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of remote sensing image cloud removal processing, in particular to a remote sensing image cloud removal reconstruction method and system. BACKGROUND

[0002] With the wide application of remote sensing images in the fields of agricultural monitoring, urban planning, disaster assessment, etc., obtaining high-quality, cloud-free remote sensing images has become a key prerequisite for improving data utilization efficiency. Due to meteorological conditions and orbital period limitations, a large number of remote sensing images are covered by clouds to varying degrees, affecting the accurate extraction of ground feature information.

[0003] In the existing method, the backbone network is mostly a stacked convolution structure, which has certain expansion capability, but when processing cloud areas with complex texture or blurred ground feature boundaries, it is difficult to effectively extract and fuse feature information at different scales, resulting in loss of detail information, edge blurring, etc., and slow inference speed and large computational overhead. In addition, the global dependency modeling capability is limited, such as the deep separable convolution used in the MAM module to simulate global attention, which reduces the computational load to a certain extent, but its modeling capability for cloud shape diversity and irregular spatial distribution is still weak, making it difficult to accurately restore long-distance relationships in the occluded area, resulting in texture distortion and ground feature deformation in the reconstructed image. SUMMARY

[0004] To solve the technical problems in the background art, the present application provides a.

[0005] The technical scheme of the present application is as follows:

[0006] A remote sensing image cloud removal reconstruction method comprises:

[0007] S1, obtaining multiple time-series remote sensing images of different regions, for the same region, selecting three images with one cloud-free image and two cloudy images in consecutive time series to form a sample, all samples forming a data set, generating a cloud layer generation image by setting cloud layer opacity, spatial aggregation, and shadow intensity from the cloud-free image in the data set, and updating it to the sample to divide a training set from the updated data set;

[0008] S2, down-sampling each image in a sample in the training set to obtain a first feature map, a second feature map, a third feature map, and a fourth feature map;

[0009] S3, reduce the fourth feature map channel number, obtain an intermediate feature map, the intermediate feature map is transversely connected with the third feature map after nearest neighbor up-sampling processing, to obtain a first aggregated feature map, the first aggregated feature map is transversely connected with the second feature map after nearest neighbor up-sampling processing, to obtain a second aggregated feature map, the second aggregated feature map is transversely connected with the first feature map after nearest neighbor up-sampling processing, to obtain a third aggregated feature map;

[0010] S4, according to the channel direction, the third aggregated feature map obtained by processing each image in the sample is spliced to obtain a multi-scale feature map, the value of the cloud mask is extracted, the local feature extraction and the global feature extraction are performed on the multi-scale feature map respectively, the local feature map and the global feature map are obtained, and the local feature map and the global feature map are fused based on the value of the cloud mask to obtain a primary fusion feature map; the fusion is refined to obtain a reconstructed image.

[0011] S5, all samples in the training set are executed in steps S2-S4 to obtain a trained cloud removal model, and the trained cloud removal model is used to process the cloud image to be processed to obtain a cloud-removed reconstructed image.

[0012] In S4, the value of the cloud mask is used to fuse the local feature map and the global feature map to obtain a primary fusion feature map, and the specific operation is as follows:

[0013] out = F l · α + F g ,

[0014] Wherein, out is the primary fusion feature map, alpha is the value of the cloud mask, F l is the local feature map, and F g is the global feature map.

[0015] In S3, the value of the cloud mask is extracted by cloud detection, and the specific operation is as follows: the spatial size of the third aggregated feature map is kept unchanged, the channel number of the third aggregated feature map is doubled, and the process of doubling the channel number is repeated to obtain an activated image. The cloud mask image is obtained by reducing the channel number of the activated image to 1 through convolution processing, and each pixel value is the value of the cloud mask.

[0016] In S4, the global feature extraction is performed, and the specific operation is as follows: after the multi-scale feature map is pooled and processed, the pooled feature map is obtained by element by point addition, and the enhanced global feature is obtained by multi-head self-attention processing, the spatial dimension is converted, and the global feature map is output.

[0017] The cloud layer generation image is generated by setting cloud layer opacity, spatial aggregation degree and shadow intensity, and the specific operation is as follows: the range of maximum and minimum cloud layer opacity is set to control the intensity of the cloud, wherein the maximum opacity is set to [0.7, 1.0], and the minimum opacity is set to [0.0, 0.05]; the spatial aggregation degree is set to [1, 10]; and the shadow intensity is set to [0.7, 0.9].

[0018] The refined fusion in S4 obtains the reconstructed image, and the specific operation is as follows: the local feature extraction and global feature extraction processing are performed on the primary fusion feature map, the fusion is performed according to the value of the cloud mask, and the channel number is adjusted to be the same as the image channel number in the sample to obtain the reconstructed image.

[0019] The specific operation of the local feature extraction in S4 is as follows:

[0020] The multi-scale feature map is subjected to depth convolution processing, and then subjected to normalization processing, point-by-point convolution processing is performed through an activation function, the channels are linearly combined and normalized to obtain the local feature map.

[0021] A remote sensing image cloud removal reconstruction system capable of realizing the above-mentioned remote sensing image cloud removal reconstruction method, comprising:

[0022] The training set updating module obtains a plurality of time series remote sensing images of different regions, selects three images including one cloud-free image and two cloudy images in the continuous time series for the same region to form a sample, and forms a data set from all samples. The cloud-free image in the data set is updated to the sample by setting cloud layer opacity, spatial aggregation degree and shadow intensity to generate a cloud layer generation image, and a training set is divided from the updated data set.

[0023] The downsampling module downsamples each image in a sample in the training set to obtain a first feature map, a second feature map, a third feature map and a fourth feature map.

[0024] The up-sampling module reduces the channel number of the fourth feature map to obtain an intermediate feature map, the intermediate feature map and the third feature map are transversely connected and subjected to nearest neighbor up-sampling processing to obtain a first aggregated feature map, the first aggregated feature map and the second feature map are transversely connected and subjected to nearest neighbor up-sampling processing to obtain a second aggregated feature map, and the second aggregated feature map and the first feature map are transversely connected and subjected to nearest neighbor up-sampling processing to obtain a third aggregated feature map.

[0025] The fusion module splices the third aggregated feature map obtained by processing each image in the sample in the channel direction to obtain a multi-scale feature map, extracts the value of the cloud mask, performs local feature extraction and global feature extraction on the multi-scale feature map respectively, obtains a local feature map and a global feature map, and obtains a primary fusion feature map based on the value of the cloud mask, the fused local feature map and the global feature map; and the refinement fusion obtains a reconstructed image.

[0026] The cloud removal module uses all samples in the training set to train the above modules to obtain a trained cloud removal model; and uses the trained cloud removal model to process the cloud image to be processed to obtain a reconstructed image.

[0027] A computer program product includes a memory and a processor, and a computer program is stored on the memory, and the program is executed by the processor to implement the steps in the remote sensing image cloud removal reconstruction described above.

[0028] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps in the remote sensing image cloud removal reconstruction described above when executing the program.

[0029] The beneficial effects of the present application are:

[0030] The present application generates a cloud layer generation image by setting the cloud layer opacity, spatial aggregation degree and shadow intensity, constructs a cloud layer in a cloud-free image, and trains a cloud removal model, thereby improving the cloud removal effect of the cloud removal model.

[0031] The present application introduces a multi-head self-attention process in the global feature extraction of the cloud removal model, dynamically captures spatial long-range dependencies of the obtained cloud removal model, effectively avoids the texture distortion and ground object structure deformation phenomenon in the reconstructed image, and greatly improves the authenticity and consistency of the cloud removal image.

[0032] The cloud removal model of the present application fully contains spatial details and semantic features of different scales through up-sampling processing of horizontal connection, primary fusion and detail fusion processing, can more effectively capture cloud layer structures and ground object boundary information under multi-scale compared with a traditional sequentially stacked convolution backbone network, and has obvious advantages in cloud layer area detail restoration and ground object contour preservation. BRIEF DESCRIPTION OF DRAWINGS

[0033] In the drawings:

[0034] Figure 1 is a cloud-free image;

[0035] Figure 2 is a cloud layer generation image;

[0036] Figure 3 is a reconstructed image. DETAILED DESCRIPTION

[0037] The technical solutions of the present invention are as follows:

[0038] A remote sensing image cloud removal and reconstruction method, comprising:

[0039] To address the problem of restoring cloud-affected farmland remote sensing images, a dataset of Sentinel-2 remote sensing images covering 50 areas in a region between 2019 and 2024 was collected using the Google Earth Engine (GEE) platform. Images in this dataset are acquired on average every 2–3 days, resulting in high temporal resolution. Each image is a 255×255 pixel tile containing four spectral bands: red, green, blue (RGB), and near-infrared (NIR). This comprehensive temporal and spatial coverage of the study area facilitates in-depth analysis of cloud cover and surface dynamics.

[0040] In practical applications, it's difficult to obtain precise, cloud-free image information for the areas corresponding to cloud-containing images. This makes it difficult to objectively and accurately compare and analyze the effectiveness of cloud removal. Therefore, an innovative cloud removal and validation method is proposed. First, a known cloud-free image is selected as a baseline sample, and a corresponding cloud-layer image is simulated. Then, cloud removal is performed on this simulated cloud-containing image. By comparing the images before and after cloud removal, the effectiveness of the cloud removal algorithm or model can be intuitively and effectively evaluated, providing a reliable basis for optimizing and improving cloud removal technology. The specific operational process is described in detail below.

[0041] S1. Obtain multiple time series of remote sensing images of different regions. For the same region, select three images with one cloud-free image and two cloud-covered images in the continuous time series. Figure 1 As shown, a sample is formed, and all samples form a data set. The cloud-free image is used to generate a cloud generation image by setting the cloud opacity, spatial aggregation, and shadow intensity. The cloud generation image is as follows Figure 2 As shown, update to the sample and divide the training set from the updated data set.

[0042] Each remote sensing image defaults to four-channel data, namely red, green, blue (R, G, B) and near-infrared (NIR) bands. Before use, you can first unify the size to (B, C = 4, H = 256, W = 256), the pixel value range is [0, 10000], and normalize it to the range of [0, 1] before entering the cloud removal model.

[0043] A known cloud-free image is selected as a reference sample, and a corresponding cloud layer generation image is simulated. The cloud layer generation process includes multiple key parameters to ensure the authenticity and diversity of the appearance of the cloud layer, and the cloud layer generation image is obtained. The specific operation is as follows: the range of maximum and minimum cloud layer opacity is set to control the intensity of the cloud, for example, the maximum opacity is set to [0.7, 1.0], and the minimum opacity is set to [0.0, 0.05], which ensures that part of the area remains semi-transparent, and other areas are completely covered by the cloud; the spatial aggregation degree is set to [1, 10]; a higher spatial aggregation degree, such as [8, 10], will generate a more compact and concentrated cloud layer structure, while a lower value, such as [1, 3], will generate a more dispersed cloud layer structure; in addition, in order to enhance the sense of reality, a shadow effect is introduced around the cloud layer, and the cloud shadow is simulated by adjusting the intensity variation. The shadow intensity is set to [0.7, 0.9] to simulate different degrees of shadow effect. This step considers the thickness and opacity of the cloud layer, making the depth and distribution of the shadow more natural.

[0044] To verify the cloud removal effect of the present application, the updated data set is divided into a training set, a validation set and a test set according to a ratio of 8:1:1.

[0045] S2, down-sampling each image in a sample in the training set to obtain a first feature map, down-sampling to obtain a second feature map, down-sampling to obtain a third feature map, and down-sampling to obtain a fourth feature map.

[0046] Each image in the sample input at this step is a four-channel image, i.e. red, green, blue (R, G, B) and near-infrared (NIR) band, with a uniform size of (B, C=4, H=256, W=256), and a pixel value range of 0-10000. After normalization to the range of [0, 1] before down-sampling, the final fourth feature map is an 8-channel image.

[0047] Image features are gradually extracted through multi-level down-sampling operations to obtain feature maps of different scales, providing effective feature representations for subsequent image processing tasks.

[0048] S3, reducing the channel number of the fourth feature map to obtain an intermediate feature map, horizontally connecting the intermediate feature map and the third feature map and performing nearest neighbor up-sampling processing to obtain a first aggregated feature map, horizontally connecting the first aggregated feature map and the second feature map and performing nearest neighbor up-sampling processing to obtain a second aggregated feature map, and horizontally connecting the second aggregated feature map and the first feature map and performing nearest neighbor up-sampling processing to obtain a third aggregated feature map.

[0049] The fourth feature map is reduced from 8 channels to 4 channels to facilitate subsequent horizontal connection, and the channel number of each of the two images horizontally connected is the same.

[0050] This step is realized by a shallow network structure, through transverse connection and layer-by-layer up-sampling operation, further, through nearest neighbor up-sampling instead of traditional bilinear interpolation, explicit extraction of spatial detail features at different resolution scales, enhancement of high resolution position information transmission ability, and fusion of deep context information. Cancel the traditional deep residual structure and additional smoothing convolution to reduce the parameter complexity, with good embedded deployment ability, faster inference speed, shallower network level, suitable for resource limited devices or real-time applications, can reduce the computational overhead. In the case of multi-channel input, a cross-channel feature fusion strategy is adopted to improve the overall prediction accuracy.

[0051] S4, splicing the third aggregated feature map obtained by processing each image in the sample according to the channel direction to obtain a multi-scale feature map, extracting the value of the cloud mask, performing local feature extraction and global feature extraction on the multi-scale feature map respectively to obtain a local feature map and a global feature map, and fusing the local feature map and the global feature map based on the value of the cloud mask to obtain a primary fusion feature map; refining the fusion to obtain a reconstructed image.

[0052] In the acquisition process of remote sensing images, clouds will interfere with the reflection and radiation signals of ground targets, causing the ground information in the images to be blurred, missing or distorted. By cloud detection to extract the cloud mask, the cloud-covered area can be accurately identified and separated from the image, avoiding interference of these invalid information on subsequent image processing and analysis, thereby improving the quality and usability of remote sensing data. Therefore, cloud detection is needed to extract the value of the cloud mask, and the specific operation is as follows: keep the spatial size of the third aggregated feature map unchanged, double the number of channels of the third aggregated feature map, repeat the process of doubling the number of channels to obtain an activated image, and through convolution processing, the number of channels of the activated image is reduced to 1 to obtain a cloud mask image, and each pixel value is the value of the cloud mask.

[0053] Further, the following steps are implemented:

[0054] The third aggregated feature map is input, 1x1 convolution is performed for channel number processing, the number of channels is increased from dim(48) to 2*dim(96), the spatial size is kept unchanged, then all are normalized and ReLU activated, and this step is repeated.

[0055] A 1x1 convolution is performed to reduce the number of channels to 1 to obtain a probability value map of each pixel belonging to a cloud, and a Sigmoid activation function is performed to limit the value to the interval [0, 1], and the module outputs a mask value with a cloud probability of [0, 1].

[0056] The local feature extraction generally describes the micro features such as the shape, texture, color and the like of a local region, and the features are generally robust to local geometric changes and illumination changes.

[0057] The local feature extraction is implemented by the following steps: inputting a multi-scale feature map, performing a deep convolution processing on the multi-scale feature map, performing a two-dimensional convolution on each input channel to obtain a new feature map, performing batch normalization on the new feature map after the deep convolution, and then performing an activation function ReLU to increase nonlinearity, so that the feature expression is sparse and stable, and then entering a point-by-point convolution layer, which is used for linear combination between channels, maps the number of channels, forms cross-channel information fusion, and outputs the local feature map after the point-by-point convolution.

[0058] The local feature extraction can capture subtle changes and detailed information in an image or data, is very effective for tasks that need to accurately describe local features, and has certain invariance to local geometric transformation and illumination change, so that the features extracted under different conditions still have comparability and stability.

[0059] The global feature extraction can reflect the overall semantic and structural information of an image or data, and is helpful for high-level classification and understanding of the image or data, and after the multi-scale feature map is pooled, the pooled feature map is obtained by element-by-element point addition, and then enhanced global features are obtained through multi-head self-attention processing, the spatial dimension is converted, and the global feature map is output.

[0060] Further, the specific operation is as follows: after the multi-scale feature map is pooled, the pooled feature map is obtained by element-by-element point addition, and then global attention mechanism enhancement processing is performed, the multi-scale feature map F i The two-dimensional adaptive average pooling function is used for pooling processing, and the calculation formula is

[0061]

[0062] is the pooled feature map, x and y are spatial coordinate indexes in the multi-scale feature map F i R h,w is a set of all pixel points (that is, pixels in the corresponding spatial region) corresponding to the output pooling unit (h, w) in the input feature map, and the coordinates of the pixel points are a set of (x, y);

[0063] All the pooled feature maps F are added element by element, to obtain the global feature map F globalThe size is unified to the spatial size of the smallest feature map, and D is the number of elements.

[0064] F global is re-encoded to obtain a two-dimensional sequence, linear mapping is performed to obtain the query Q, the key K, and the value vector V, the global attention relationship of all pixels is calculated to form an attention weight matrix, the value vector V is weighted and summed using the attention weight matrix A to obtain a feature representation:

[0065] A = Attention(Q, K, V)

[0066] X = AV

[0067] X represents the deep feature of each pixel after re-encoding combined with the relationship with the global pixels.

[0068] All feature representations are spliced into a two-dimensional sequence, the spatial dimension is converted, and a global feature map is output. The specific operation is as follows:

[0069] The fused feature X is rearranged and spliced with multiple heads. The number of heads is equal to the number of groups of X i , i = h = 8, restored to a two-dimensional sequence, and then fused by linear transformation and regularization operation of each subspace feature. Finally, the sequence is dimensionally transformed back to the original spatial dimension BxCxHxW, and the global feature is output.

[0070] The lightweight multi-head attention mechanism is used instead of the traditional self-attention approximation operation based on depthwise convolution, which adapts to irregular cloud distribution and heterogeneous scenes in remote sensing images, and strengthens the reconstruction ability of long-range dependence and boundary texture.

[0071] The global feature also needs to be converted in spatial dimension. Based on the two-dimensional sequence, each subspace feature is fused by linear transformation and regularization, and finally the sequence is dimensionally transformed back to the spatial dimension of BxCxHxW.

[0072] The local feature map and the global feature map are fused to obtain a primary fusion feature map. The specific operation is as follows:

[0073] out = F l · α + F g ,

[0074] Wherein, out is the primary fusion feature map, α is the value of the cloud mask, F l is the local feature map, and F g is the global feature map.

[0075] After primary fusion, it is also necessary to refine the fusion details to obtain the reconstructed image, and the specific operation is: local feature extraction and global feature extraction are performed on the primary fusion feature map, fusion is performed according to the value of the cloud mask, and the channel number is adjusted to be the same as the image channel number in the sample, to obtain the reconstructed image.

[0076] This step is actually a repetition of the primary fusion process, the difference being that the primary fusion has spliced the third aggregated feature map obtained by processing each image, and the refined fusion does not splice; for example, the primary fusion input is a 64*3 channel third aggregated feature map, and the output is a 64 channel primary fusion feature map; the refined fusion input is a 64 channel primary fusion feature map, and the output is also a 64 channel, and finally the channel number is adjusted to 4 channels through 1*1 convolution to adjust the channel number, and the 64 channel is adjusted to a 4 channel reconstructed image as the output.

[0077] By dividing the cloud layer removal task into multiple phased sub-tasks, the intermediate image generated in each stage is used as the input of the next stage, the modeling quality of the cloud layer edge and cloud shadow interference is gradually optimized, and the spatial consistency and detail restoration capability of the cloud removal model are improved.

[0078] After completing the image cloud removal reconstruction, the computer reverses the pixel value from [0, 1] to [0, 10000] range, and outputs the reconstructed image, as shown in Figure 3 , maintaining the spectral consistency and data format of the original image, facilitating subsequent storage and use.

[0079] S5, execute S2-S4 steps on all samples in the training set to obtain a trained cloud removal model, and use the trained cloud removal model to process the cloud image to obtain a cloud removal reconstructed image.

[0080] The above S2-S4 is the cloud removal model training process, and is also the specific processing process of the cloud removal model.

[0081] The cloud removal model trained by the training set is compared with the cloud removal model using the progressive multi-scale attention self-encoder architecture, the cloud removal model using the generative adversarial network, and the cloud removal model using the super-resolution reconstruction algorithm, and the PSNR and SSIM values of the images after cloud removal processing of the validation set compared with the original cloud-free images are shown in Table 1.

[0082] Table 1: PSNR and SSIM values of images after cloud removal processing of each model compared with original cloud-free images

[0083]

[0084] It can be seen that the application introduces global context modeling capability while maintaining high efficiency, significantly enhances the network's perception of long-distance dependence and complex spatial structure, thereby improving the overall feature expression capability and model performance, and the similarity between the cloud-removed image and the cloud-free image is higher, and the cloud-removal effect is better.

[0085] In the application, the traditional residual connection structure is replaced by a transverse connection structure. This structure realizes the collaborative expression of shallow spatial information and deep semantic information by fusing representations from different scale feature layers. Through transverse connection and layer-by-layer upsampling, the spatial details of the shallow layer and the semantic information of the deep layer are integrated, thereby preserving the spatial resolution while improving the model's expression capability for cloud layers and ground object structures and enhancing the cloud removal effect.

[0086] A remote sensing image cloud removal reconstruction system can implement the above-mentioned remote sensing image cloud removal reconstruction method, comprising:

[0087] A training set updating module obtains multiple time-series remote sensing images of different regions, selects three images with one cloud-free image and two cloudy images in the continuous time series for the same region to form a sample, and forms a data set from all samples. The cloud-free image is updated to the sample by setting the cloud layer opacity, spatial aggregation degree and shadow intensity to generate a cloud layer generation image, and a training set is divided from the updated data set.

[0088] A downsampling module downsamples each image in a sample in the training set to obtain a first feature map, a second feature map, a third feature map and a fourth feature map.

[0089] An upsampling module reduces the channel number of the fourth feature map to obtain an intermediate feature map. The intermediate feature map and the third feature map are transversely connected and processed by nearest neighbor upsampling to obtain a first aggregated feature map. The first aggregated feature map and the second feature map are transversely connected and processed by nearest neighbor upsampling to obtain a second aggregated feature map. The second aggregated feature map and the first feature map are transversely connected and processed by nearest neighbor upsampling to obtain a third aggregated feature map.

[0090] A fusion module splices the third aggregated feature map obtained by processing each image in the sample in the channel direction to obtain a multi-scale feature map. The value of the cloud mask is extracted, and the multi-scale feature map is subjected to local feature extraction and global feature extraction to obtain a local feature map and a global feature map. Based on the value of the cloud mask, the local feature map and the global feature map are fused to obtain a primary fusion feature map. Refinement fusion is performed to obtain a reconstructed image.

[0091] The de-clouding module uses all samples in the training set to train the above module to obtain a trained de-clouding model; the trained de-clouding model is used to process the clouded image to be processed to obtain a reconstructed image.

[0092] A computer program product includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, implements the steps in the above-mentioned remote sensing image cloud removal and reconstruction.

[0093] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned steps in cloud removal and reconstruction of remote sensing images when executing the program.

Claims

1. A remote sensing image cloud removal and reconstruction method, characterized in that: include: S1. Obtain multiple time series of remote sensing images from different regions. For the same region, select three images in the continuous time series that contain one cloud-free image and two cloud-covered images to form a sample. All samples form a dataset. The cloud-free images are then updated to generate cloud-covered images by setting cloud opacity, spatial aggregation, and shadow intensity. A training set is then generated from the updated dataset. S2. Downsample each image in a sample in the training set to obtain a first feature map, downsample to obtain a second feature map, downsample to obtain a third feature map, and downsample to obtain a fourth feature map; S3. Reduce the number of channels of the fourth feature map to obtain an intermediate feature map, horizontally connect the intermediate feature map with the third feature map and perform a nearest neighbor upsampling process to obtain a first aggregated feature map, horizontally connect the first aggregated feature map with the second feature map and perform a nearest neighbor upsampling process to obtain a second aggregated feature map, and horizontally connect the second aggregated feature map with the first feature map and perform a nearest neighbor upsampling process to obtain a third aggregated feature map; S4. Splice each image in the sample in the channel direction to obtain the third aggregated feature map to obtain a multi-scale feature map, extract the cloud mask value, perform local feature extraction and global feature extraction on the multi-scale feature map to obtain a local feature map and a global feature map, and fuse the local feature map and the global feature map based on the cloud mask value to obtain a primary fused feature map; refine the fusion to obtain a reconstructed image; S5. Execute steps S2-S4 on all samples in the training set to obtain a trained declouding model. Use the trained declouding model to process the clouded image to obtain a declouded reconstructed image.

2. The remote sensing image cloud removal and reconstruction method according to claim 1, characterized in that: Based on the cloud mask value described in S4, the local feature map and the global feature map are fused to obtain the primary fused feature map. The specific operation is: out=F l ·α+F g , Among them, out is the primary fusion feature map, α is the value of the cloud mask, F l is the local feature map, F g is the global feature map.

3. The remote sensing image cloud removal and reconstruction method according to claim 1, characterized in that: The cloud detection described in S3 extracts the cloud mask value. The specific operation is: keep the spatial size of the third aggregate feature map unchanged, double the number of channels of the third aggregate feature map, repeat the process of doubling the number of channels to obtain the activation image, and use convolution processing to reduce the number of channels of the activation image to a cloud mask image of 1. Each pixel value is the value of the cloud mask.

4. The remote sensing image cloud removal and reconstruction method according to claim 1, characterized in that: The global feature extraction described in S4 is specifically performed as follows: after pooling the multi-scale feature map, the pooled feature map is obtained by adding the elements point by point, and then the enhanced global features are obtained by multi-head self-attention processing, the spatial dimension is converted, and the global feature map is output.

5. The remote sensing image cloud removal and reconstruction method according to claim 1, characterized in that: As described in S1, cloud generation images are generated by setting cloud opacity, spatial aggregation, and shadow intensity. The specific operations are: setting the range of maximum and minimum cloud opacity to control the cloud intensity, where the maximum opacity is set to [0.7, 1.0] and the minimum opacity is set to [0.0, 0.05]; the spatial aggregation value range is set to [1, 10]; and the shadow intensity is set to [0.7, 0.9].

6. The remote sensing image cloud removal and reconstruction method according to claim 1, characterized in that: The refinement fusion described in S4 is used to obtain a reconstructed image. The specific operations are: performing local feature extraction and global feature extraction processing on the primary fusion feature map, fusing according to the value of the cloud mask, and adjusting the number of channels to be the same as the number of image channels in the sample to obtain a reconstructed image.

7. The remote sensing image cloud removal and reconstruction method according to claim 1, characterized in that: The specific operation of local feature extraction in S4 is: The multi-scale feature map is subjected to deep convolution processing, normalized, and point-by-point convolution processing through the activation function. The channels are linearly combined and normalized to obtain the local feature map.

8. A remote sensing image cloud removal and reconstruction system, characterized in that: Implementing a remote sensing image cloud removal and reconstruction method as claimed in any one of claims 1 to 7, comprising: The training set update module obtains remote sensing images from multiple time series of different regions. For the same region, three images with one cloud-free image and two cloud-covered images in the continuous time series are selected to form a sample. All samples form a data set. The cloud-free images are then updated to generate cloud-covered images by setting cloud opacity, spatial aggregation, and shadow intensity. The cloud-covered images are then updated to the sample, and the training set is divided from the updated data set. A downsampling module downsamples each image in a sample in the training set to obtain a first feature map, downsamples to obtain a second feature map, downsamples to obtain a third feature map, and downsamples to obtain a fourth feature map; An upsampling module reduces the number of channels of the fourth feature map to obtain an intermediate feature map, the intermediate feature map is horizontally connected to the third feature map and subjected to nearest neighbor upsampling processing to obtain a first aggregated feature map, the first aggregated feature map is horizontally connected to the second feature map and subjected to nearest neighbor upsampling processing to obtain a second aggregated feature map, and the second aggregated feature map is horizontally connected to the first feature map and subjected to nearest neighbor upsampling processing to obtain a third aggregated feature map; The fusion module splices the third aggregated feature map obtained by processing each image in the sample in the channel direction to obtain a multi-scale feature map, extracts the cloud mask value, performs local feature extraction and global feature extraction on the multi-scale feature map to obtain local feature maps and global feature maps, and obtains a primary fused feature map based on the cloud mask value and the fusion of the local feature map and the global feature map; refines the fusion to obtain the reconstructed image; The de-clouding module uses all samples in the training set to train the above module to obtain a trained de-clouding model; the trained de-clouding model is used to process the clouded image to be processed to obtain a reconstructed image.

9. A computer program product comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the program is executed by a processor, the steps in the remote sensing image cloud removal and reconstruction according to any one of claims 1 to 7 are implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps in the remote sensing image cloud removal and reconstruction according to any one of claims 1 to 7 are implemented.