Remote sensing cloud restoration method based on time-spectral domain fusion and temporal self-attention enhancement

By adopting a method based on time spectrum domain fusion and timing self-attention enhancement in remote sensing image cloud removal technology, the problem of cloud occlusion damages image details is solved, significantly improving the effect and consistency of image repair, and efficient cloud removal and detail fidelity are achieved.

CN119399082BActive Publication Date: 2025-05-06NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510013348.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

Existing remote sensing image cloud removal technology is difficult to effectively solve the damage to local details of the image by cloud occlusion, and the traditional attention mechanism does not pay enough attention to important information in edge areas, which limits the effect and consistency of image repair.

Method used

Using a remote sensing cloud layer repair method based on time spectrum domain fusion and timing self-attention enhancement, a generator, an adaptive image repair module and a multi-scale module discriminator are constructed, combined with the conditional GAN ​​loss function and the multi-head self-attention mechanism in the feature extractor, the feature expression ability and edge information capture ability are enhanced, and the progressive cloud removal from the outside to the inside is achieved.

Benefits of technology

It significantly improves the detail clarity and texture fidelity of the image, improves the cloud removal effect, enhances the model's ability to capture edge information, and improves the consistency and comprehensiveness of image repair.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399082B_ABST
    Figure CN119399082B_ABST
Patent Text Reader

Abstract

The present invention discloses a remote sensing cloud repair method based on time-spectrum domain fusion and temporal self-attention enhancement, and relates to the technical field of remote sensing cloud image processing. The present invention includes: receiving a remote sensing missing image as a source image; constructing a generator, and inputting the remote sensing missing image into the generator for preliminary processing, wherein the generator includes a feature extractor, a ConformerPlus module, and a cloud detection module, and a new multi-head self-attention mechanism is embedded in the ConformerPlus module. The present invention can extract and fuse the time-spectrum domain information of the data, thereby enhancing the expression ability of the features, and improving the model's ability to capture edge information, thereby guiding the model to achieve progressive cloud removal from the outside to the inside, while also improving detail performance and suppressing noise interference, effectively improving the detail clarity and texture fidelity of the image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing cloud image processing, and specifically to a remote sensing cloud repair method based on time-spectrum domain fusion and temporal self-attention enhancement. Background Art

[0002] The cloud occlusion problem in optical remote sensing images is one of the key challenges in remote sensing data processing. It leads to serious loss and quality degradation of surface information, which has a significant impact on the accuracy of core tasks such as object recognition, change detection and land cover classification, and may even cause the subsequent processing tasks to fail completely. Existing remote sensing image cloud removal technologies mainly include multi-temporal method, multi-spectral method and deep learning-based method: (1) Multi-temporal method: Use the complementary information of remote sensing images acquired at different times to repair the surface features of cloud-occluded areas. This method can effectively handle scenes with dynamic cloud cover, but its performance is easily affected by the difficulty of obtaining time series data and the dynamic changes of the surface; (2) Multi-spectral method: Repair cloud-occluded areas by correlating spectral bands, which has strong physical significance. However, this method relies heavily on high-spectral resolution data and has limited performance when the cloud layer is thick; (3) Deep learning-based method: Reconstruct images by building data-driven models, which has strong nonlinear modeling capabilities and can significantly improve the declouding effect. The cloud removal performance based on deep learning relies on large-scale labeled datasets and is easily limited by the quality and distribution of training data.

[0003] Existing image restoration methods usually rely on deep convolutional networks for feature extraction and information reconstruction. However, in complex remote sensing scenes, deep networks may introduce additional noise, resulting in a decrease in the quality of restored images. Existing methods generally lack targeted detail enhancement and restoration mechanisms, and it is difficult to effectively solve the damage to local image details caused by cloud occlusion, making it difficult to meet the requirements of high-precision image restoration. In addition, traditional attention mechanisms, such as the Conformer model, often pay insufficient attention to important information in edge areas when processing cloud-occluded images, thereby limiting the effect of image restoration. This problem limits the consistency and comprehensiveness of the restoration results, thereby affecting the overall performance. Summary of the invention

[0004] The purpose of the present invention is to provide a remote sensing cloud restoration method based on time-spectral domain fusion and temporal self-attention enhancement, which can extract and fuse time-spectral domain information of data, thereby enhancing the expression ability of features and improving the model's ability to capture edge information, thereby guiding the model to achieve progressive cloud removal from the outside to the inside, while also improving detail expression and suppressing noise interference, effectively improving the detail clarity and texture fidelity of the image.

[0005] To achieve the above object, the present invention provides the following technical solution: a remote sensing cloud repair method based on time-spectrum domain fusion and temporal self-attention enhancement, comprising the following steps:

[0006] receiving a remote sensing missing image as a source image;

[0007] Construct a generator and input the remote sensing missing image into the generator for preliminary processing. The generator includes a feature extractor, a ConformerPlus module, and a cloud detection module. The ConformerPlus module is embedded with a multi-head self-attention mechanism.

[0008] Construct an adaptive image restoration module, input the preliminarily processed remote sensing missing images into the adaptive image restoration module for further processing, and obtain the restored image data;

[0009] Construct a multi-scale module discriminator, connect the repaired image and the remote sensing missing image using concat and input them into the multi-scale module discriminator to determine whether the repaired image is true or false;

[0010] The conditional GAN ​​loss function, standard loss function, and cloud mask loss function are used to calculate the difference between the repaired image and the source image, and the gradient is updated through back propagation to gradually optimize the model parameters;

[0011] The peak signal-to-noise ratio and structural similarity index are used as quantitative evaluation indicators to comprehensively evaluate the performance of the overall network model consisting of the generator, adaptive image restoration module, and multi-scale module discriminator.

[0012] Save the trained overall network model and use the overall network model for image restoration.

[0013] Furthermore, a generator is constructed to input the remote sensing missing image into the generator for preliminary processing, as follows:

[0014] (21) The input of the generator is three remote sensing images with cloud occlusion, defined as , select one of the cloud images as the reference image, denoted as y, and on this basis, given , the model learns to generate a cloud-free image , the image Maintain high similarity with the reference image y in the non-missing area;

[0015] (22) Input image First, the feature extraction module is used to extract three features. , these feature representations model the local and global information of the input image in the temporal and spatial domains, laying the feature foundation for the subsequent cloud-free image generation, namely:

[0016]

[0017] Three feature representations All are fused together by the concat layer to obtain three parallel feature representations ,Right now:

[0018]

[0019] In the formula are passed to a single encoder pipeline, which performs Encode and fuse the feature representation output by the encoder through the Concat layer , decode the image and output the cloud-free image restoration result , where the encoder and decoder are convolutional layers with stride 2 respectively.

[0020] Furthermore, the feature extractor is used to extract and fuse the time-spectral domain features of the remote sensing missing image, and the specific model is as follows:

[0021] (31) A feature extractor based on the temporal self-attention mechanism is used. This extractor dynamically encodes the features at each position in the image and comprehensively integrates global semantic information and local detail features to cope with the complex distribution of objects and cloud occlusion in remote sensing scenes.

[0022] (32) Expanding the temporal and spectral features of the image data through a one-dimensional linear transformation layer to form a comprehensive feature vector that integrates the information of the two;

[0023] (33) These vectors are input into a Transformer model containing a single encoder, where residual connections and normalization operations are used to enhance the stability and transmission efficiency of features;

[0024] (34) The multi-head attention mechanism further captures the global and local dependencies between features, and the feedforward neural network refines the high-dimensional feature expression;

[0025] (35) Introducing the Dropout operation during the encoding process to improve the generalization ability of the model;

[0026] (36) The high-dimensional features are projected back into one-dimensional space through one-dimensional linear transformation to generate three representative time-domain-spectral feature vectors.

[0027] Furthermore, the ConformerPlus module is used to enhance multi-scale features to capture complex contextual relationships, and the ConformerPlus module integrates convolutional neural networks, new multi-head self-attention modules and Transformer architecture;

[0028] The new multi-head self-attention module is designed based on the Conformer module, and a new attention weight tensor is introduced in the new multi-head self-attention module;

[0029] The ConformerPlus module embeds the new multi-head self-attention module and the convolutional layer between two feedforward neural networks to improve the performance and robustness of the model. The feedforward neural network contains two linear transformations and a nonlinear activation function, and the global feature extraction network consists of ConformerPlus modules, for the Input of a ConformerPlus module , whose intermediate variables and output It can be obtained by the following formula:

[0030]

[0031] The input tensor Through the feedforward neural network, the feedforward neural network FFN contains two fully connected layers, Swish activation function and dropout regularization; It is a fusion value that combines the information of yi and the feature information extracted by the multi-head attention mechanism. is the feature mined by the convolution layer and The fusion value of the feature information; NMHSA is the new multi-head self-attention module, Conv is the convolution layer, and LayerNorm represents the normalization processing of the input data.

[0032] Furthermore, the cloud detection module is used to identify and mark the cloud distribution position in the input image, as follows:

[0033] (51) By analyzing the input image features, a corresponding cloud mask is generated to clearly identify the cloud-covered area and the potential information missing or occluded area;

[0034] (52) Based on the generated cloud mask, a selective contrast between the source and target images is performed, focusing only on the areas that are not obscured by clouds.

[0035] Furthermore, an adaptive image restoration module is constructed to input the preliminarily processed remote sensing missing images into the adaptive image restoration module for further processing. The specific model is as follows:

[0036] The adaptive image restoration module includes a random noise elimination module and a local contrast enhancement module, wherein the noise elimination module is constructed based on a deep convolutional network architecture, and a batch normalization layer and a residual block are introduced into the noise elimination module to enhance the perception and removal capabilities of noise;

[0037] The local contrast enhancement module is used to extract high-level features. During processing, the local contrast enhancement module extracts features through a convolutional network and uses the Tanh activation function to normalize the output.

[0038] Furthermore, a multi-scale module discriminator is constructed, and the repaired image and the remote sensing missing image are connected using concat and input into the multi-scale module discriminator to judge the authenticity of the repaired image, as follows:

[0039] (71) The image data is processed by a feature map of size 256×256×32, and the initial discrimination probability P1 is output;

[0040] (72) The image data is sequentially processed through feature maps of sizes 128×128×64 and 64×64×128 to generate the second-layer discriminant probability P2;

[0041] (73) The feature map is further compressed to the size of 32×32×256 and 16×16×512, and the third-layer discriminant probability P3 is output;

[0042] (74) In the feature extraction process, all convolution operations use a convolution kernel of size 4×4, with a step size of 2, and are combined with the LeakyReLU activation function to introduce nonlinear characteristics for efficient extraction of image features at different scales;

[0043] (75) After obtaining the three hierarchical discriminant probabilities of P1, P2 and P3, the data is further processed through the last layer of feature maps, which has a size of 1×1×1 and a step size of 1, and is combined with the Sigmoid activation function to generate the discriminant probability P4;

[0044] (76) The generated image is divided into multiple independent image blocks. Assuming that the pixels in each image block are independent of each other, the authenticity of each image block is judged separately. Then, the overall authenticity evaluation result of the entire image is obtained by taking the average of the judgment results of all image blocks.

[0045] Furthermore, the conditional GAN ​​loss function, standard loss function, and cloud mask loss function are used to calculate the difference between the repaired image and the original remote sensing missing image, and the gradient is updated through back propagation to gradually optimize the model parameters, as follows:

[0046] (81) The loss function used consists of three parts: conditional GAN ​​loss function, standard Loss function, cloud mask loss function, is defined as:

[0047]

[0048] Where L is the total loss function, represents the minimization operation on the parameters of the generator and adaptive image restoration module, Respectively represent the maximization operation of the discriminator parameters; L cGAN (GP,D) is the loss of the conditional generative adversarial network, is the gradient penalty coefficient, express The norm loss, L mask is the cloud mask loss, is the loss of cloud mask computation, where D represents the multi-scale module discriminator, and parameters G and P represent the generator and adaptive image inpainting module respectively;

[0049] (82) The first part is the conditional GAN ​​loss function, defined as:

[0050]

[0051] in, is the loss of the conditional generative adversarial network, and Denote the expectations of real data and generated data, respectively, D is the output of the discriminator for real data, It is the output of the generator and the adaptive image restoration module on the generated data;

[0052] (83) The second part is the standard Loss Function L 1 (GP), defined as:

[0053]

[0054] In the formula Represent the number of channels, width and height respectively, is the real data in coordinates Pixels at Indicates that the generated output image is at coordinates Pixels at

[0055] (84) The third part is the cloud mask loss function, which is defined as:

[0056]

[0057] Where L mask is the cloud mask loss, M and M′ are the real cloud mask and generated cloud mask respectively.

[0058] Furthermore, the peak signal-to-noise ratio and structural similarity index are used as quantitative evaluation indicators to comprehensively evaluate the performance of the overall network model consisting of the generator, the adaptive image restoration module, and the multi-scale module discriminator, as follows:

[0059] (91) Peak signal-to-noise ratio and structural similarity index were used as quantitative evaluation indicators. Peak signal-to-noise ratio is the ratio between the maximum signal and the background noise. It is a quality indicator for measuring image quality and is defined as:

[0060]

[0061] Where: n is the number of bits of each sampling value; MSE is the mean square error value; the larger the PSNR value of the quality indicator, the better the image declouding effect;

[0062] (92) The structural similarity index is an indicator used to measure the similarity between two images. The structural similarity is measured from three aspects: brightness, contrast, and structure. The mean is used as an estimate of brightness, the standard deviation is used as an estimate of contrast, and the covariance is used as a measure of structural similarity. It is defined as:

[0063]

[0064] in, is the structural similarity index, and They are The average value of and They are The standard deviation of for The covariance of are constants respectively to avoid calculation errors caused by the denominator being 0.

[0065] According to a second aspect of the present invention, the present invention provides a remote sensing cloud repair system based on time-spectrum domain fusion and temporal self-attention enhancement, which is used to implement the above-mentioned remote sensing cloud repair method based on time-spectrum domain fusion and temporal self-attention enhancement, comprising:

[0066] A data receiving unit, used for receiving a remote sensing missing image as a source image;

[0067] The generator construction unit is used to construct a generator and input the remote sensing missing image into the generator for preliminary processing, wherein the generator includes a feature extractor, a ConformerPlus module and a cloud detection module, and a multi-head self-attention mechanism is embedded in the ConformerPlus module;

[0068] An adaptive image restoration module construction unit is used to construct an adaptive image restoration module, input the preliminarily processed remote sensing missing image into the adaptive image restoration module for further processing, and obtain restored image data;

[0069] A determination unit is used to construct a multi-scale module discriminator, connect the repaired image and the remote sensing missing image using concat and input them into the multi-scale module discriminator to determine whether the repaired image is true or false;

[0070] The error calculation unit is used to calculate the difference between the repaired image and the source image using the conditional GAN ​​loss function, the standard loss function, and the cloud mask loss function, and to gradually optimize the model parameters by updating the gradient through back propagation;

[0071] An evaluation unit is used to comprehensively evaluate the performance of the overall network model consisting of the generator, the adaptive image restoration module, and the multi-scale module discriminator using the peak signal-to-noise ratio and the structural similarity index as quantitative evaluation indicators;

[0072] The restoration output unit is used to save the trained overall network model and use the overall network model to perform image restoration.

[0073] The present invention has at least the following beneficial effects:

[0074] The present invention adopts a feature extractor (FE) based on the temporal self-attention mechanism in the generative adversarial network (GAN) to effectively deal with the multi-heterogeneity of remote sensing data. The feature extractor can extract and fuse the temporal and spectral domain information of the data, thereby enhancing the expression ability of the features and improving the analysis and understanding ability of remote sensing data from different sources and types. In addition, the designed new multi-head self-attention mechanism (NMHSA) significantly improves the model's ability to capture edge information, thereby guiding the model to achieve progressive cloud removal from the outside to the inside, optimizing the overall image restoration effect;

[0075] After the remote sensing missing images were processed by GAN, an adaptive image enhancement module was designed to further optimize the image quality, improve detail performance and suppress noise interference, which effectively improved the image detail clarity and texture fidelity.

[0076] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 It is a schematic diagram of the process of the repair method of the present invention;

[0078] Figure 2 This is a schematic diagram of the principle framework of the repair method of the present invention;

[0079] Figure 3 Schematic diagram of the overall structure of the generator in an embodiment of the present invention;

[0080] Figure 4 Schematic diagram of the overall structure of a feature extractor in an embodiment of the present invention;

[0081] Figure 5 Schematic diagram of the overall structure of the ConformerPlus module in an embodiment of the present invention;

[0082] Figure 6 Schematic diagram of the overall structure of the adaptive image restoration module in an embodiment of the present invention;

[0083] Figure 7 Schematic diagram of the overall structure of a multi-scale module discriminator in an embodiment of the present invention;

[0084] Figure 8 This is the visualization comparison result of TGAN and TSGAN on the image declouding task on the Sen2_MTC cloud dataset according to an embodiment of the present invention. DETAILED DESCRIPTION

[0085] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0086] See also Figure 1 The present invention provides a technical solution: a remote sensing cloud repair method based on time-spectrum domain fusion and temporal self-attention enhancement, comprising:

[0087] S1. Construct a generator (G) and input the pre-received remote sensing missing image into the generator for preliminary processing, wherein the generator includes a feature extractor (FE), a ConformerPlus module (CPM) and a cloud detection module, and a multi-head self-attention mechanism is embedded in the ConformerPlus module;

[0088] It should be noted that the missing remote sensing images are collected by satellites or aircraft;

[0089] (S11) The overall structure of the generator in S1 is as follows Figure 3 As shown:

[0090] The input of the generator consists of three remote sensing images with cloud occlusion, defined as ,Since it is difficult to collect a real cloud-free image that completely corresponds to the missing image, it is assumed that the ground object remains relatively unchanged in the time scale, and one of the cloud images is selected as the reference image, denoted as y;

[0091] On this basis, given , the model learns to generate a cloud-free image , the image The model maintains a high similarity with the reference image y in the non-missing area and gradually optimizes the feature expression to achieve high-quality restoration by fusing the information of the non-missing part of the input image. Specifically:

[0092] Input Image First, the feature extraction module TempSA is used to extract three features: , these feature representations model the local and global information of the input image in the temporal and spatial domains, laying the feature foundation for the subsequent cloud-free image generation, namely:

[0093]

[0094] Three feature representations All are fused together by the concat layer to obtain three parallel feature representations ,Right now:

[0095]

[0096] in, are passed to a single encoder pipeline, which performs Encode and fuse the feature representation output by the encoder through the Concat layer , decode the image and output the cloud-free image restoration result , where the encoder and decoder are convolutional layers with stride 2 respectively;

[0097] (S12) FE in S1 adopts a feature extractor based on the temporal self-attention mechanism. The extractor dynamically encodes the features of each position in the image and comprehensively integrates global semantic information and local detail features to cope with the complex distribution of objects and cloud occlusion in remote sensing scenes. The design architecture is as follows Figure 4 As shown;

[0098] The feature extractor first expands the temporal and spectral dimension features of the image data through a one-dimensional linear transformation layer to form a comprehensive feature vector that integrates the information of the two. After dimensionality increase, these comprehensive feature vectors are input into the Transformer model containing a single encoder, where residual connections and normalization operations are used to enhance the stability and transmission efficiency of the features. The multi-head attention mechanism further captures the global and local dependencies between features, while the feedforward neural network refines the high-dimensional feature expression. In order to reduce the risk of overfitting, the Dropout operation is introduced in the position encoding process to improve the generalization ability of the model. Finally, the high-dimensional features are projected back to the one-dimensional space through a one-dimensional linear transformation to generate three representative time-domain-spectral feature vectors, providing high-quality data support for subsequent processing steps.

[0099] (S13) In S1, a new multi-head self-attention module (NMHSA) is designed on the basis of the Conformer module to construct the ConformerPlus module. The NMHSA module introduces a new attention weight tensor to improve the ability to capture edge information, thereby guiding the model to achieve progressive cloud removal from the outside to the inside and optimize the overall restoration effect; the ConformerPlus module integrates convolutional neural networks, NMHSA and Transformer architectures, thereby effectively enhancing the expression ability of key features. Its specific implementation is as follows: Figure 5 As shown;

[0100] The ConformerPlus module embeds the NMHSA module and the convolutional layer (Conv) between two feedforward neural networks (FFN). This dual FFN configuration is significantly better than using only a single FFN and can effectively improve the performance and robustness of the model. The feedforward network contains two linear transformations and a nonlinear activation function, and the global feature extraction network consists of ConformerPlus, for the Input of a ConformerPlus module , whose intermediate variables and output It can be obtained by the following formula:

[0101]

[0102] The input tensor Through the feedforward neural network (FFN), the network contains two fully connected layers, Swish activation function and dropout regularization to enhance the expressiveness of the model; at the same time, the NMHSA mechanism combines traditional multi-head attention with the attention self-adjustment matrix to effectively optimize the attention score; in addition, the convolution layer (Conv) integrates the gated linear unit (GLU), depthwise separable convolution and other activation and regularization techniques to efficiently process sequence data and extract features; finally, LayerNorm is used to normalize the input data to improve the training effect of the model; It is a fusion value that combines the information of yi and the feature information extracted by the multi-head attention mechanism. is the feature mined by the convolution layer and The fusion value of the feature information;

[0103] (S14) The cloud detection module in S1 is used to identify and mark the distribution of clouds in the input image. The module generates a corresponding cloud mask by analyzing image features to clearly identify cloud-covered areas and areas with potential information missing or occluded. Based on the generated cloud mask, a selective comparison between the source image (the original image containing clouds) and the target image (the image after clouds are removed) is further achieved, focusing only on areas not occluded by clouds. This differentiated comparison strategy can not only effectively guide the generator to optimize model training, but also improve the reconstruction accuracy in cloud-free areas, thereby enhancing the generalization ability of the generator in complex scenes.

[0104] S2. Construct an adaptive image inpainting module (AIIM), input the preliminarily processed remote sensing missing image into the adaptive image inpainting module for further processing, and obtain the inpainted image data. The adaptive image inpainting module can enhance the global consistency and detail fidelity of the inpainted image;

[0105] The adaptive image restoration module includes a random noise removal module (RNRM) and a local contrast enhancement module (LCEM), and its structure is as follows: Figure 6 As shown;

[0106] The overall structure of the adaptive image restoration module includes RNRM and LCEM. This separate design not only improves the adaptability and scalability of the system, but also can be optimized to meet the needs of diverse tasks according to different image characteristics and noise distribution. Specifically:

[0107] RNRM effectively enhances the ability to perceive and remove noise through a deep convolutional network architecture combined with batch normalization layers and residual blocks. The introduction of residual blocks alleviates the common gradient vanishing problem in deep networks and improves the training stability and convergence efficiency of the model. In addition, the gradually changing channel number strategy in the module design can capture multi-scale features, thereby achieving efficient removal of different types of noise.

[0108] LCEM focuses on extracting high-level features while retaining the original details of the image to avoid image distortion caused by over-enhancement. This module extracts features through a convolutional network and uses the Tanh activation function to normalize the output to ensure that the enhanced image is neither oversaturated nor overly dark.

[0109] S3. Construct a multiscale module discriminator (D), connect the repaired image and the remote sensing missing image using concat and input them into the multiscale module discriminator to determine whether the repaired image is true or false;

[0110] The multi-scale module discriminator architecture is as follows Figure 7 As shown in the figure, the multi-scale module discriminator extracts features from global to local and generates hierarchical discrimination probabilities by processing the input data through a series of feature maps in sequence. First, the data is processed by a feature map of size 256×256×32, and the initial discrimination probability P1 is output. Then, the data is processed by feature maps of size 128×128×64 and 64×64×128 in sequence to generate the second-layer discrimination probability P2. Then, the feature map is further compressed to the size of 32×32×256 and 16×16×512, and the third-layer discrimination probability P3 is output. This multi-scale feature extraction mechanism can fully capture the global information and local details of the input data, thereby improving the robustness and discrimination accuracy of the discriminator.

[0111] In the feature extraction process, all convolution operations use a 4×4 convolution kernel with a step size of 2, and are combined with the LeakyReLU activation function to introduce nonlinear characteristics; the dimension of the feature map is expressed in the form of H×W×C (image height × image width × number of channels); through this design, the discriminator can efficiently extract image features at different scales and comprehensively evaluate the authenticity of the input image, thereby providing strong guidance for the optimization of the generator;

[0112] After obtaining the three hierarchical discrimination probabilities of P1, P2 and P3, the data is further processed through the last layer of feature maps, which has a size of 1×1×1 and a step size of 1, and combined with the Sigmoid activation function to generate the discrimination probability P4; the core of the discriminator design is to divide the generated image into multiple independent image blocks, assuming that the pixels in each image block are independent of each other, and perform true or false judgment on each image block separately; finally, by taking the average of the judgment results of all image blocks, the overall true or false evaluation result of the entire image is formed;

[0113] S4. Using conditional GAN ​​loss function, standard The loss function and cloud mask loss function calculate the difference between the generated sample (the repaired image) and the real sample (the source image). By back-propagating the gradient, the model can gradually optimize the parameters.

[0114] (S41) The loss function used consists of three parts: conditional GAN ​​loss function, standard Loss function, cloud mask loss function, is defined as:

[0115]

[0116] Where L is the total loss function, represents the minimization operation on the parameters of the generator and adaptive image restoration module, Respectively represent the maximization operation of the discriminator parameters, L cGAN (GP,D) is the loss of the conditional generative adversarial network, is the gradient penalty coefficient, express The norm loss of is the loss of cloud mask computation; D represents the multi-scale module discriminator, and parameters G and P represent the generator and adaptive image restoration module respectively;

[0117] (S42) The first part is the conditional GAN ​​loss function, defined as:

[0118]

[0119] in, is the loss of cGAN, and Denote the expectations of real data and generated data, respectively, D is the output of the discriminator for real data, It is the output of the generator and the adaptive image restoration module on the generated data;

[0120] (S43) The second part is the standard The loss function is defined as:

[0121]

[0122] In the formula Represent the number of channels, width and height respectively, is the real data in coordinates Pixels at Indicates that the generated output image is at coordinates Pixels at

[0123] (S44) The third part is the cloud mask loss function, which is defined as:

[0124]

[0125] Where L mask is the cloud mask loss, M and M′ are the real cloud mask and the generated cloud mask respectively;

[0126] S5. Use peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) as quantitative evaluation indicators to comprehensively evaluate the performance of the proposed overall network model cloud removal method. The overall network model consists of a generator, an adaptive image restoration module, and a multi-scale module discriminator.

[0127] (S51) Peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) are used as quantitative evaluation indicators. PSNR calculates the ratio between the maximum signal and the background noise. It is a quality indicator for measuring image quality and is defined as:

[0128]

[0129] Where: n is the number of bits of each sampling value; MSE is the mean square error value; the larger the PSNR value, the better the image declouding effect;

[0130] (S52) SSIM is an indicator used to measure the similarity between two images. Structural similarity is measured from three aspects: brightness, contrast, and structure. The mean is used as an estimate of brightness, the standard deviation is used as an estimate of contrast, and the covariance is used as a measure of structural similarity. It is defined as:

[0131]

[0132] in, and They are The average value of and They are The standard deviation of for The covariance of are constants respectively, to avoid calculation errors caused by denominators being 0;

[0133] S6. Save the trained network model, test it, and then use the overall network model to perform image restoration;

[0134] S6 was conducted on an NVIDIA GeForce RTX 3090 GPU server equipped with 24GB of video memory, using the Adam optimizer with a momentum parameter of and The initial learning rate is set to 5e-4, and after 100 training epochs, the learning rate is gradually reduced to 1e-6 using a linear decay strategy; the batch size is 4.

[0135] Next, this embodiment is verified by combining specific experiments:

[0136] The experimental environment parameters are shown in Table 1:

[0137]

[0138] To verify the effectiveness of the cloud removal method of the TGAN model (i.e., the above-mentioned overall network model) in this embodiment, it is applied to the real Sen2-MTC cloud dataset, which contains 50 non-repetitive categories, each category includes 70 images, the pixel value range of the image is [0,10000], the image resolution is 256×256, and each image has four spectral channels: red (R), green (G), blue (B), and near infrared (NIR); these images represent different geographical areas and weather conditions, covering a variety of cloud types and changes, with strong representativeness and diversity, suitable for cloud removal tasks in remote sensing images.

[0139] To ensure the generalization ability of the model and the objectivity of the evaluation, the dataset is randomly divided into training set, validation set and test set in a ratio of 7:1:2. Specifically, 70% of the 70 images are used for training, 10% for validation, and the remaining 20% ​​for testing. To avoid category bias during data partitioning, images from the same category are assigned to the same subset to ensure the uniformity of the category distribution of images in each subset. This partitioning strategy helps to ensure that the model can be effectively trained and evaluated on different datasets, thereby improving its robustness and reliability in practical applications.

[0140] Experimental results and analysis:

[0141] In the experiment, in order to verify the performance of the proposed TGAN model in the remote sensing image data restoration task, the present invention conducted a comprehensive experiment based on the Sen2-MTC dataset and compared the experimental results with the results of TSGAN, STGAN, ST_net, AE and TGAN models. The results are shown in Tables 2 and 3.

[0142] Experimental results show that the TGAN model performs significantly better than other models in terms of peak signal-to-noise ratio (PSNR). Specifically, the current most advanced TSGAN model has SSIM of 0.662 and 0.609 on the validation set and test set, and PSNR of 21.259 and 18.308, respectively. The TGAN model of this embodiment achieves significant gains of 0.288 and 1.898 on the validation set and test set, respectively, in terms of PSNR. This shows that the TGAN model has higher robustness in restoring the overall image quality and reducing noise interference. Although the structural similarity index (SSIM) of TGAN has slightly decreased, which to some extent indicates that the similarity between the generated image and the reference image has decreased, the subjective visual evaluation results show that TGAN exhibits higher quality in terms of detail fidelity and global contrast of the generated image.

[0143]

[0144]

[0145] Corresponding experimental verification is carried out for the effect of the model, such as Figure 8 As shown in the figure, this figure shows a visual comparison of TGAN and TSGAN in the image declouding task. From top to bottom, there are three cloudy input images, cloud-free images generated by TSGAN, cloud-free images generated by TGAN, and corresponding cloud-free reference images. It can be observed from the figure that the images generated by TGAN are closer to the real images in terms of contrast, texture details, and overall visual perception, while the images generated by TSGAN are slightly lacking in some details and contrast. The above results show that the TGAN model not only performs better in quantitative indicators, but also shows stronger advantages in subjective perception quality, providing a reliable solution for high-quality remote sensing image restoration.

[0146] In summary, the present invention adopts a feature extractor (FE) based on the temporal self-attention mechanism in the generative adversarial network (GAN) to effectively deal with the multivariate heterogeneity of remote sensing data. The feature extractor can extract and fuse the temporal and spectral domain information of the data, thereby enhancing the expression ability of the features and improving the analysis and understanding ability of remote sensing data from different sources and types. In addition, the designed new multi-head self-attention mechanism (NMHSA) significantly improves the model's ability to capture edge information, thereby guiding the model to achieve progressive cloud removal from the outside to the inside, optimizing the overall image restoration effect;

[0147] After the remote sensing missing images were processed by GAN, an adaptive image enhancement module was designed to further optimize the image quality, improve detail performance and suppress noise interference, which effectively improved the image detail clarity and texture fidelity.

[0148] Embodiment 2:

[0149] like Figure 2 As shown in the figure, in the first stage, the model introduces a feature extractor based on the temporal self-attention mechanism and a new multi-head self-attention mechanism. The feature extractor captures the time domain and spectral domain characteristics of the data through a one-dimensional linear dimensionality increase layer, and uses a one-dimensional linear dimensionality reduction layer to replace the traditional maximum pooling operation, thereby enhancing the modeling ability of different time series positions; the new multi-head self-attention mechanism introduces a new weight allocation strategy to dynamically adjust the attention score to enhance the ability to capture edge information; the second stage adopts an adaptive image restoration module, including random noise elimination and local contrast enhancement sub-modules, to further optimize image quality, improve detail performance and suppress noise interference. In addition, the discriminator of the model achieves a balance between global consistency and local details through multi-scale module design.

[0150] In addition, this embodiment provides a remote sensing cloud repair system based on time-spectrum domain fusion and temporal self-attention enhancement, which is used to implement the remote sensing cloud repair method based on time-spectrum domain fusion and temporal self-attention enhancement described in Example 1, including:

[0151] A data receiving unit, used for receiving a remote sensing missing image as a source image;

[0152] The generator construction unit is used to construct a generator and input the remote sensing missing image into the generator for preliminary processing, wherein the generator includes a feature extractor, a ConformerPlus module and a cloud detection module, and a multi-head self-attention mechanism is embedded in the ConformerPlus module;

[0153] An adaptive image restoration module construction unit is used to construct an adaptive image restoration module, input the preliminarily processed remote sensing missing image into the adaptive image restoration module for further processing, and obtain restored image data;

[0154] A determination unit is used to construct a multi-scale module discriminator, connect the repaired image and the remote sensing missing image using concat and input them into the multi-scale module discriminator to determine whether the repaired image is true or false;

[0155] The error calculation unit is used to calculate the difference between the repaired image and the source image using the conditional GAN ​​loss function, the standard loss function, and the cloud mask loss function, and to gradually optimize the model parameters by updating the gradient through back propagation;

[0156] An evaluation unit is used to comprehensively evaluate the performance of the overall network model consisting of the generator, the adaptive image restoration module, and the multi-scale module discriminator using the peak signal-to-noise ratio and the structural similarity index as quantitative evaluation indicators;

[0157] The restoration output unit is used to save the trained overall network model and use the overall network model to perform image restoration.

[0158] Specifically, the above-mentioned data receiving unit, generator construction unit, adaptive image restoration unit construction unit, determination unit, error calculation unit, evaluation unit and restoration output unit can be embedded in a computer processing system. The computer calls the above-mentioned units to complete the task of repairing the remote sensing image according to the above-mentioned method for nighttime road surface recognition based on improved lighting conditions; the above-mentioned data receiving unit, generator construction unit, adaptive image restoration unit construction unit, determination unit, error calculation unit, evaluation unit and restoration output unit can perform operations according to the specific steps given in the above-mentioned method for nighttime road surface recognition based on improved lighting conditions.

[0159] It should be noted that it should be understood that the division of the various units of the above system is only the division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated, and these units can all be implemented in the form of software calling through processing elements; they can also be all implemented in the form of hardware; some units can also be implemented in the form of software calling through processing elements, and some units can be implemented in the form of hardware. For example, the data receiving unit can be a separately established processing element, or it can be integrated in a chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a processing element of the above device. The implementation of other units is similar. In addition, all or part of these units can be integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above units can be completed by the hardware integrated logic circuit in the processor element or the instructions in the form of software.

[0160] For example, the above units may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASICs), or one or more digital singnal processors (DSPs), or one or more field programmable gate arrays (FPGAs). For another example, when a certain unit above is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. For another example, these units may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0161] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0162] For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances. When an element is referred to as being "assembled on", "installed on", "fixed on" or "set on" another element, it can be directly on the other element or there can also be a centered element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be a centered element at the same time. The terms "vertical", "horizontal", "up", "down", "left", "right" and similar expressions used herein are only for illustrative purposes and are not intended to be the only implementation method.

[0163] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

[0164] In the description of this specification, the description with reference to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

Claims

1. A remote sensing cloud restoration method based on time-spectral domain fusion and temporal self-attention enhancement is characterized by: The following steps are involved: receiving a remote sensing missing image as a source image; Construct a generator and input the remote sensing missing image into the generator for preliminary processing. The generator includes a feature extractor, a ConformerPlus module, and a cloud detection module. The ConformerPlus module is embedded with a new multi-head self-attention mechanism. Construct an adaptive image restoration module, input the preliminarily processed remote sensing missing images into the adaptive image restoration module for further processing, and obtain the restored image data; Construct a multi-scale module discriminator, connect the repaired image and the remote sensing missing image using concat and input them into the multi-scale module discriminator to determine whether the repaired image is true or false; The conditional GAN ​​loss function, standard loss function, and cloud mask loss function are used to calculate the difference between the repaired image and the source image, and the gradient is updated through back propagation to gradually optimize the model parameters; The peak signal-to-noise ratio and structural similarity index are used as quantitative evaluation indicators to comprehensively evaluate the performance of the overall network model consisting of the generator, adaptive image restoration module, and multi-scale module discriminator. Save the trained overall network model and use the overall network model for image restoration.

2. The remote sensing cloud restoration method based on time-spectrum domain fusion and temporal self-attention enhancement according to claim 1 is characterized in that: Construct a generator and input the remote sensing missing image into the generator for preliminary processing, as follows: (21) The input of the generator is three remote sensing images with cloud occlusion, defined as , select one of the cloud images as the reference image, denoted as y, and on this basis, given , the model learns to generate a cloud-free image , the image Maintain high similarity with the reference image y in the non-missing area; (22) Input image First, the feature extraction module is used to extract three features. , these feature representations model the local and global information of the input image in the temporal and spatial domains, laying the feature foundation for the subsequent cloud-free image generation, namely: Three feature representations All are fused together by the concat layer to obtain three parallel feature representations ,Right now: In the formula are passed to a single encoder pipeline, which performs Encode and fuse the feature representation output by the encoder through the Concat layer , decode the image and output the cloud-free image restoration result , where the encoder and decoder are convolutional layers with stride 2 respectively.

3. The remote sensing cloud restoration method based on time-spectrum domain fusion and temporal self-attention enhancement according to claim 2 is characterized by: The feature extractor is used to extract and fuse the time-spectral domain features of the remote sensing missing image. The specific model is as follows: (31) A feature extractor based on the temporal self-attention mechanism is used. This extractor dynamically encodes the features at each position in the image and comprehensively integrates global semantic information and local detail features to cope with the complex distribution of objects and cloud occlusion in remote sensing scenes. (32) Expanding the temporal and spectral features of the image data through a one-dimensional linear transformation layer to form a comprehensive feature vector that integrates the information of the two; (33) These vectors are input into a Transformer model containing a single encoder, where residual connections and normalization operations are used to enhance the stability and transmission efficiency of features; (34) The multi-head attention mechanism further captures the global and local dependencies between features, and the feedforward neural network refines the high-dimensional feature expression; (35) Introducing the Dropout operation during the encoding process to improve the generalization ability of the model; (36) The high-dimensional features are projected back into one-dimensional space through one-dimensional linear transformation to generate three representative time-domain-spectral feature vectors.

4. The remote sensing cloud restoration method based on time-spectrum domain fusion and temporal self-attention enhancement according to claim 3 is characterized by: The ConformerPlus module is used to enhance multi-scale features to capture complex contextual relationships. The ConformerPlus module integrates convolutional neural networks, new multi-head self-attention modules and Transformer architecture; The new multi-head self-attention module is designed based on the Conformer module, and a new attention weight tensor is introduced in the new multi-head self-attention module; The ConformerPlus module embeds the new multi-head self-attention module and the convolutional layer between two feedforward neural networks to improve the performance and robustness of the model. The feedforward neural network contains two linear transformations and a nonlinear activation function, and the global feature extraction network consists of ConformerPlus modules, for the Input of a ConformerPlus module , whose intermediate variables and output It can be obtained by the following formula: The input tensor Through the feedforward neural network FFN, the feedforward neural network FFN contains two fully connected layers, Swish activation function and dropout regularization; It is a fusion value that combines the information of yi and the feature information extracted by the multi-head attention mechanism. is the feature mined by the convolution layer and The fusion value of the feature information; NMHSA is the new multi-head self-attention module, Conv is the convolution layer, and LayerNorm represents the normalization processing of the input data.

5. The remote sensing cloud restoration method based on time-spectrum domain fusion and temporal self-attention enhancement according to claim 4 is characterized by: The cloud detection module is used to identify and mark the cloud distribution position in the input image, as follows: (51) By analyzing the input image features, a corresponding cloud mask is generated to clearly identify the cloud-covered area and the potential information missing or blocked area; (52) Based on the generated cloud mask, a selective contrast between the source and target images is performed, focusing only on the areas that are not obscured by clouds.

6. The remote sensing cloud restoration method based on time-spectrum domain fusion and temporal self-attention enhancement according to claim 5 is characterized in that: Construct an adaptive image restoration module, and input the preliminarily processed remote sensing missing images into the adaptive image restoration module for further processing. The specific model is as follows: The adaptive image restoration module includes a random noise elimination module and a local contrast enhancement module, wherein the noise elimination module is constructed based on a deep convolutional network architecture, and a batch normalization layer and a residual block are introduced into the noise elimination module to enhance the perception and removal capabilities of noise; The local contrast enhancement module is used to extract high-level features. During processing, the local contrast enhancement module extracts features through a convolutional network and uses the Tanh activation function to normalize the output.

7. The remote sensing cloud restoration method based on time-spectrum domain fusion and temporal self-attention enhancement according to claim 6 is characterized in that: Construct a multi-scale module discriminator, connect the repaired image and the remote sensing missing image using concat and input them into the multi-scale module discriminator to determine whether the repaired image is true or false, as follows: (71) The image data is processed by a feature map of size 256×256×32, and the initial discrimination probability P1 is output; (72) The image data is sequentially processed through feature maps of sizes 128×128×64 and 64×64×128 to generate the second-layer discriminant probability P2; (73) The feature map is further compressed to the size of 32×32×256 and 16×16×512, and the third-layer discriminant probability P3 is output; (74) In the feature extraction process, all convolution operations use a convolution kernel of size 4×4, with a step size of 2, and are combined with the LeakyReLU activation function to introduce nonlinear characteristics for efficient extraction of image features at different scales; (75) After obtaining the three hierarchical discriminant probabilities of P1, P2 and P3, the data is further processed through the last layer of feature maps, which has a size of 1×1×1 and a step size of 1, and is combined with the Sigmoid activation function to generate the discriminant probability P4; (76) The generated image is divided into multiple independent image blocks. Assuming that the pixels in each image block are independent of each other, the authenticity of each image block is judged separately. Then, the overall authenticity evaluation result of the entire image is obtained by taking the average of the judgment results of all image blocks.

8. The remote sensing cloud restoration method based on time-spectrum domain fusion and temporal self-attention enhancement according to claim 7 is characterized in that: The conditional GAN ​​loss function, standard loss function, and cloud mask loss function are used to calculate the difference between the repaired image and the original remote sensing missing image, and the gradient is updated through back propagation to gradually optimize the model parameters as follows: (81) The loss function used consists of three parts: conditional GAN ​​loss function, standard Loss function, cloud mask loss function, is defined as: Where L is the total loss function, represents the minimization operation on the parameters of the generator and adaptive image restoration module, Respectively represent the maximization operation of the discriminator parameters; L cGAN (GP,D) is the loss of the conditional generative adversarial network, is the gradient penalty coefficient, express The norm loss, L mask is the cloud mask loss, is the loss of cloud mask computation; D represents the multi-scale module discriminator, and parameters G and P represent the generator and adaptive image restoration module respectively; (82) The first part is the conditional GAN ​​loss function, defined as: In the formula, is the loss of the conditional generative adversarial network, and Denote the expectations of real data and generated data, respectively, D is the output of the discriminator for real data, It is the output of the generator and the adaptive image restoration module on the generated data; (83) The second part is the standard The loss function is defined as: In the formula Represent the number of channels, width and height respectively, is the real data in coordinates Pixels at Indicates that the generated output image is at coordinates Pixels at (84) The third part is the cloud mask loss function, which is defined as: Where L mask is the cloud mask loss, M and M′ are the real cloud mask and generated cloud mask respectively.

9. The remote sensing cloud restoration method based on time-spectrum domain fusion and temporal self-attention enhancement according to claim 8 is characterized in that: The peak signal-to-noise ratio and structural similarity index are used as quantitative evaluation indicators to comprehensively evaluate the performance of the overall network model consisting of the generator, adaptive image restoration module, and multi-scale module discriminator. The details are as follows: (91) Peak signal-to-noise ratio and structural similarity index were used as quantitative evaluation indicators. Peak signal-to-noise ratio is the ratio between the maximum signal and the background noise. It is a quality indicator for measuring image quality and is defined as: Where: n is the number of bits of each sampling value; MSE is the mean square error value; the larger the PSNR value of the quality indicator, the better the image declouding effect; (92) The structural similarity index is an indicator used to measure the similarity between two images. The structural similarity is measured from three aspects: brightness, contrast, and structure. The mean is used as an estimate of brightness, the standard deviation is used as an estimate of contrast, and the covariance is used as a measure of structural similarity. It is defined as: in, is the structural similarity index, and They are The average value of and They are The standard deviation of for The covariance of are constants respectively to avoid calculation errors caused by the denominator being 0.

10. A remote sensing cloud repair system based on time-spectrum domain fusion and temporal self-attention enhancement, used to implement the remote sensing cloud repair method based on time-spectrum domain fusion and temporal self-attention enhancement as described in any one of claims 1 to 9, characterized in that: include: A data receiving unit, used for receiving a remote sensing missing image as a source image; The generator construction unit is used to construct a generator and input the remote sensing missing image into the generator for preliminary processing, wherein the generator includes a feature extractor, a ConformerPlus module and a cloud detection module, and a multi-head self-attention mechanism is embedded in the ConformerPlus module; An adaptive image restoration module construction unit is used to construct an adaptive image restoration module, input the preliminarily processed remote sensing missing image into the adaptive image restoration module for further processing, and obtain restored image data; A determination unit is used to construct a multi-scale module discriminator, connect the repaired image and the remote sensing missing image using concat and input them into the multi-scale module discriminator to determine whether the repaired image is true or false; The error calculation unit is used to calculate the difference between the repaired image and the source image using the conditional GAN ​​loss function, the standard loss function, and the cloud mask loss function, and to gradually optimize the model parameters by updating the gradient through back propagation; An evaluation unit is used to comprehensively evaluate the performance of the overall network model consisting of the generator, the adaptive image restoration module, and the multi-scale module discriminator using the peak signal-to-noise ratio and the structural similarity index as quantitative evaluation indicators; The restoration output unit is used to save the trained overall network model and use the overall network model to perform image restoration.

Citation Information

Patent Citations

  • Remote sensing image segmentation method of multi-path parallel network based on height perception

    CN113554032A

  • Face image restoration method based on multi-scale local self-attention generative adversarial network

    CN113962893A