Remote sensing image thin cloud removal method based on self-supervision domain self-adaption
By combining a thin cloud simulation model that takes into account both cloud distortion and spatial morphology with a self-supervised domain adaptive method, the problem of decreased accuracy caused by the inconsistency between the simulated thin cloud image and the real thin cloud image data distribution is solved, and efficient thin cloud removal from remote sensing images is achieved.
Patent Information
- Application Number
- CN202510831905.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-11-18
AI Technical Summary
Existing deep learning-based thin cloud removal methods suffer from inconsistent data distribution between simulated and real thin cloud images, leading to decreased accuracy when applied to real thin cloud images and making it difficult to effectively remove thin cloud interference from remote sensing images.
We employ a thin cloud simulation model (TCSCDSM) that takes into account both cloud distortion and spatial morphology. Combining self-supervised learning and domain adaptation methods, we train the declouding model using a simulated dataset. We design a mask attention mechanism and a multi-layer domain alignment strategy to enhance the model's feature alignment and generalization capabilities across different domains.
Without relying on or with minimal reliance on target domain labels, the accuracy and generalization ability of the thin cloud removal model are significantly improved, effectively removing thin clouds from remote sensing images and enhancing image utilization.
Smart Images

Figure CN120976056A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of thin cloud removal technology in remote sensing images, and specifically to a method for thin cloud removal in remote sensing images based on self-supervised domain adaptation. Background Technology
[0002] Optical remote sensing imagery, characterized by its large observation area, rich surface information, and strong timeliness, is widely used in numerous fields such as agricultural production, urban planning, and environmental protection. However, due to the limitations of the imaging mechanism of optical sensors, optical remote sensing imagery is highly susceptible to cloud cover, leading to distortion, blurring, or even complete loss of spectral and textural information (Ji et al., 2021). Furthermore, due to the absorption, reflection, and re-radiation of solar radiation by clouds, the contrast and sharpness of images obtained under cloud cover conditions are severely reduced (Tan et al., 2024). According to data from the International Satellite Cloud Climatology Project (ISCCP), the global average cloud cover is approximately 67% (Zhang et al., 2004). However, when the cloud cover of an image exceeds 10-15%, it is generally considered unsuitable for geoscientific research (Asner, 2001). Therefore, removing cloud interference from remote sensing imagery and restoring low-quality cloud-containing images to high-quality cloud-free images has significant practical application value and theoretical significance for improving the usability of remote sensing imagery under cloudy and foggy weather conditions.
[0003] Based on differences in optical thickness and spatial morphology, clouds in remote sensing images are generally classified into thick clouds and thin clouds. Thick clouds completely block reflected and radiated signals from the ground surface, and can usually only be processed through information compensation, i.e., filling in the target area covered by clouds with clear-sky images taken at adjacent time points (Zhai and Xue, 2024). Thin clouds, on the other hand, have a certain transmittance, allowing surface signals to penetrate them and reach the sensor. Surface information of the thin-clouded area can be recovered from the image itself, which is of great significance for improving image utilization (Skakun et al., 2022). Traditional thin cloud removal methods utilize the relationship between different spectral bands of thin cloud images to weaken the influence of thin clouds; however, their model construction has many prerequisites and relies excessively on expert knowledge, often only suitable for cloud removal work in specific scenarios (Xu et al., 2022).
[0004] With the development of artificial intelligence technology, the application of deep learning algorithms to remove thin clouds from remote sensing images has become increasingly popular, and its effectiveness and performance have significantly improved compared to traditional methods (Caraballo-Vega et al., 2023; Meraner et al., 2020; Tayebi et al., 2023). However, deep learning-based cloud removal models mainly employ supervised learning strategies, which means that training requires cloudless images of the same region corresponding to the thin cloud images. In reality, a region can only have one type of weather condition at a given time; clear weather and foggy weather cannot coexist. Currently, there are two main solutions to this problem. One is to find two images with and without clouds with a short time interval as training sample pairs. This method is costly and it is difficult to guarantee that the land cover remains unchanged. The other is to artificially add noise to cloudless images to simulate thin cloud images and pair them with cloudless images as training samples. Because finding adjacent "cloudy-cloudless" sample pairs at adjacent time intervals is costly and it's difficult to guarantee that the land cover remains unchanged, deep learning-based thin cloud removal methods typically use simulated thin cloud images to generate a large amount of labeled data for training. However, simulated thin cloud images and real thin cloud images have significant deviations in data distribution. Using simulated images for training may result in a large discrepancy between the mapping relationship between thin cloud and cloudless images learned by the model and the actual mapping, affecting the model's generalization ability when applied to real thin cloud images. Summary of the Invention
[0005] In view of this, the purpose of this invention is to provide a remote sensing image thin cloud removal method based on self-supervised domain adaptation, in order to solve the problem that the accuracy of existing cloud removal models decreases when applied to real thin cloud images due to the inconsistency between the data distribution of simulated thin cloud images and real thin cloud images.
[0006] To achieve the above objectives, this invention proposes a Thin Cloud Simulation Model (TCSCDSM) that considers both cloud distortion and spatial morphology. This model yields a simulated dataset of the source domain, and a domain-adaptive remote sensing image thin cloud removal model is constructed using self-supervised learning techniques. This model enables the removal of thin clouds from remote sensing images. Specifically, the method includes the following steps:
[0007] 1) Acquire remote sensing images of the thin clouds to be removed;
[0008] 2) Input the remote sensing image into the trained domain-adaptive remote sensing image thin cloud removal model to obtain the remote sensing image with thin clouds removed;
[0009] The training process of the domain-adaptive remote sensing image thin cloud removal model includes:
[0010] A thin cloud image dataset was created using a thin cloud simulation model that takes into account both cloud distortion and spatial morphology, which was used as a training sample for the thin cloud removal model. The thin cloud simulation model that takes into account both cloud distortion and spatial morphology consists of an atmospheric scattering model based on bidirectional transmission and a transmittance estimation method based on cirrus cloud bands, which are used to simulate the cloud distortion characteristics and spatial morphology characteristics of thin clouds, respectively.
[0011] The simulated dataset and the real dataset are regarded as the source domain and the target domain, respectively. A mask attention mechanism is designed in combination with land cover data. The cloud removal network uses the information of cloudless areas under the same land surface type to complete the information of thin cloud-covered areas by using the input mask and the attention mask, respectively.
[0012] A domain adaptation method is used to align the data distribution features of the source and target domains at the input, feature, and output layers, respectively.
[0013] The method of this invention has the following advantages: Based on the thin cloud simulation model proposed in this invention, which takes into account both cloud distortion and spatial morphology, and addressing the problem of decreased accuracy when the cloud removal model is applied to real thin cloud images due to the inconsistency in data distribution between simulated and real thin cloud images, this invention proposes a domain-adaptive remote sensing image thin cloud removal model. This thin cloud removal model belongs to the domain-adaptive method based on self-supervised learning. First, based on the concept of self-supervised learning, the simulated thin cloud image dataset created according to the thin cloud simulation model is used as the training sample for the cloud removal model, helping the model learn more robust feature representations while avoiding insufficient data labeling samples. Then, the simulated dataset and the real dataset are respectively regarded as the source and target domains. A mask attention mechanism is designed in conjunction with land cover data, encouraging the cloud removal network to use information from cloudless areas under the same land surface type to supplement the information of thin cloud-occluded areas through input masks and attention masks. Finally, a domain-adaptive method is used to align the data distribution features of the source and target domains from the input layer, feature layer, and output layer, enhancing the model's generalization ability and enabling it to achieve good cloud removal results even with little or no target domain labeling.
[0014] Furthermore, the atmospheric scattering model based on bidirectional transmission is as follows:
[0015] I(x) = t(x) 2 J(x)+(1-t(x))S
[0016] Where x represents the position of a point in the image, J and I represent the clear image to be restored and the observed hazy image, respectively, t represents the transmittance of the medium, t(x) = dt(x) = ut(x), S is solar radiation, and dt(x) and ut(x) are the downlink transmission and uplink transmission, respectively.
[0017] The method of this invention takes into account that the process of solar radiation reaching the ground and being reflected to the sensor mainly includes two processes: reflection (reflection from clouds and reflection from the ground) and transmission (transmission of solar radiation through clouds and transmission of ground reflection through clouds). Both processes are related to the thickness of the clouds. Therefore, it is necessary to consider not only the part of sunlight reflected from the ground objects that passes through clouds and fog to reach the sensor, but also the part of sunlight that is absorbed when it passes through clouds and fog to reach the ground objects (i.e., the downward transmission of sunlight through clouds and fog). Therefore, based on the characteristics of images acquired by remote sensing sensors, this invention considers the process of solar radiation passing through thin clouds to reach the ground and then being reflected by the ground to reach the sensor. It proposes an atmospheric scattering model based on bidirectional transmission to describe the relationship between solar radiation and thin clouds and to better simulate the distortion characteristics of thin clouds.
[0018] Furthermore, the transmittance estimation method based on the cirrus cloud band is as follows:
[0019]
[0020] Among them, t dcp The transmittance is estimated based on the dark channel prior algorithm, DN c It is the DN value of the cirrus cloud band.
[0021] This invention considers that the cirrus cloud band in remote sensing imagery is a water vapor absorption band, capable of reflecting cloud thickness, and is particularly sensitive to the thickness of thin clouds. Cloud thickness determines its transmittance; therefore, the cirrus cloud band can be used to reflect cloud transmittance. Based on the spectral characteristics of the cirrus cloud band, a transmittance estimation method based on the cirrus cloud band is designed to simulate the spatial morphological characteristics of thin clouds.
[0022] Furthermore, the domain adaptation method is used to align the data distribution features of the source and target domains at the input, feature, and output layers, respectively, including:
[0023] Three domain alignment strategies: input layer alignment, feature layer alignment, and output layer alignment;
[0024] The input layer alignment strategy includes: the image transformation generator of the input layer is used to convert the simulated image of the source domain into an image similar to the real image of the target domain, and then combined with the discriminator to enable the generator to convert the source domain image into a realistic target domain image through adversarial learning;
[0025] The feature layer alignment strategy includes: the discriminator of the feature layer is responsible for receiving the deep features of the source domain image and the target domain image output by the encoder, and reducing the difference in deep features to make the feature extraction process of the encoder for the source domain and the target domain similar;
[0026] The output layer alignment strategy includes: the output layer distinguishes whether the input comes from the source domain or the target domain, so that the output results of the source domain and target domain images after decoding are as close as possible.
[0027] The method of this invention proposes three domain alignment strategies: input layer alignment, feature layer alignment, and output layer alignment, to effectively reduce performance differences caused by domain offset. These three alignment strategies are seamlessly integrated into a unified model, thus allowing them to benefit from each other through an end-to-end training process.
[0028] Furthermore, when the source domain image is obtained based on the transmittance simulation of the target domain image, an input layer alignment strategy is executed.
[0029] Furthermore, the adversarial loss for input layer alignment is:
[0030]
[0031] in, For datasets from the source domain, For a dataset from the target domain, G i (x s )=x s →t This indicates that the generator transforms the source image into an image similar to the target domain, D i D represents i Discriminator.
[0032] Furthermore, feature layer alignment is achieved by adding a domain discriminator;
[0033] The domain discriminator takes the high-dimensional features of the source and target domains output by the encoder as input. If the domain discriminator can distinguish the high-dimensional features of the source and target domains, the adversarial gradient is backpropagated to the feature extractor. The adversarial loss for feature layer alignment is:
[0034]
[0035] Among them, D f D represents f Domain discriminator, where E represents the encoder.
[0036] Furthermore, by adding an output layer discriminator to distinguish whether the declouded image originates from x s→t The reconstruction is still based on the real target domain image x. t If the output layer discriminator obtains an accurate discrimination result, then the features extracted by the feature extractor include domain features. To keep the feature regions invariant, the following adversarial loss is used to supervise the feature extraction process:
[0037]
[0038] Among them, Do D represents o Output layer discriminator, where D represents the decoder.
[0039] Furthermore, during training, the encoder E utilizes the discriminator {D} after the input layer alignment process. f D o Collect and combat losses and To optimize, the decoder utilizes the data from the discriminator D. o Collect and fight against losses Optimize.
[0040] Furthermore, in each training iteration, all modules are updated sequentially in the following order:
[0041] First, the generator G is updated to obtain the transformed class-target domain image. Then, the discriminator D... i Updated to classify target image x s→t and the real target image x t Next, the encoder E is updated to use the data from x. s→t and x t Features are extracted from the discriminator D, and then the discriminator D is updated. f The decoder D is then updated to convert the extracted features into a cloudless image, based on its input features. Finally, the discriminator D... o The output features are updated to their source and target domains, thereby enhancing feature invariance.
[0042] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described in detail below with reference to the accompanying drawings. Attached Figure Description
[0043] Figure 1 This is a schematic diagram of the atmospheric scattering model based on bidirectional transmission used in this invention;
[0044] Figure 2 A graph showing the relationship between transmittance and DN values in cirrus cloud bands under different land cover types;
[0045] Figure 3 A schematic diagram showing the DN value and transmittance of cirrus clouds in the water region.
[0046] Figure 4 For image patch scale t dcp and DN c A schematic diagram of the results of a univariate linear regression;
[0047] Figure 5This is a schematic diagram of the overall architecture of the thin cloud removal model based on domain adaptation used in this invention;
[0048] Figure 6 This is a schematic diagram of the attention mask structure used in this invention. Detailed Implementation
[0049] The technical solution of the present invention will be clearly and completely described below with reference to specific embodiments. However, those skilled in the art should understand that the embodiments described below are only for illustrating the present invention and should not be regarded as limiting the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0050] Example of a self-supervised domain adaptive remote sensing image thin cloud removal method
[0051] This embodiment utilizes a Thin Cloud Simulation Model (TCSCDSM) that considers both cloud distortion and spatial morphology, along with a domain-adaptive remote sensing image thin cloud removal model, to achieve thin cloud removal from remote sensing images. Specifically, the method in this embodiment includes the following steps:
[0052] 1) Acquire remote sensing images of the thin clouds to be removed;
[0053] 2) Input the remote sensing image into the trained domain-adaptive remote sensing image thin cloud removal model to obtain the remote sensing image with thin clouds removed;
[0054] The training process of the domain-adaptive remote sensing image thin cloud removal model includes:
[0055] A thin cloud image dataset was created using a thin cloud simulation model that takes into account both cloud distortion and spatial morphology, which was used as a training sample for the thin cloud removal model. The thin cloud simulation model that takes into account both cloud distortion and spatial morphology consists of an atmospheric scattering model based on bidirectional transmission and a transmittance estimation method based on cirrus cloud bands, which are used to simulate the cloud distortion characteristics and spatial morphology characteristics of thin clouds, respectively.
[0056] The simulated dataset and the real dataset are regarded as the source domain and the target domain, respectively. A mask attention mechanism is designed in combination with land cover data. The cloud removal network uses the information of cloudless areas under the same land surface type to complete the information of thin cloud-covered areas by using the input mask and the attention mask, respectively.
[0057] A domain adaptation method is used to align the data distribution features of the source and target domains at the input, feature, and output layers, respectively.
[0058] This embodiment provides a detailed explanation of the thin cloud simulation model that takes into account both cloud distortion and spatial morphology, as well as the thin cloud removal model for remote sensing images based on domain adaptation.
[0059] I. Thin cloud simulation model that takes into account both cloud distortion and spatial morphology (TCSCDSM).
[0060] This embodiment considers that the spatial distribution characteristics of clouds can be summarized into two categories based on their morphological differences: one is uniformly distributed clouds, and the other is unevenly distributed clouds. Different cloud distribution characteristics have different effects on the quality and resolution of remote sensing images. In addition, the differences in the spatial morphology of clouds and fog are specifically manifested as differences in cloud and fog thickness, which in turn determines the degree of light distortion.
[0061] To address the above issues, this invention designs a thin cloud simulation model (TCSCDSM) that takes into account both cloud distortion and spatial morphology. Based on the characteristics of images acquired by remote sensing sensors, this model considers the process of solar radiation passing through thin clouds to the ground and then being reflected back to the sensor. It improves the bidirectional transmission-based atmospheric scattering model (ASTT) to describe the relationship between solar radiation and thin clouds, better simulating the distortion characteristics of thin clouds. Furthermore, considering the spectral characteristics of cirrus cloud bands, a transmittance estimation method based on cirrus cloud bands (Cirrus-TE) is designed to simulate the spatial morphology of thin clouds.
[0062] (1) Atmospheric scattering simulation taking into account the distortion characteristics of thin clouds.
[0063] Atmospheric scattering models are mathematical models used to describe the scattering effect of the atmosphere on electromagnetic waves. Commonly used, simple atmospheric scattering models are typically employed in image dehazing research to describe the process of sunlight passing through fog to reach the camera in natural scenes. However, remote sensing images are captured by sensors from space, covering a much larger area and being more significantly affected by atmospheric scattering effects. Simple atmospheric scattering models cannot effectively model this process. Therefore, this embodiment improves upon the simple atmospheric scattering model by adopting an Atmospheric Scattering Model Based on Two-Way Transmission (ASTT) to better simulate the distortion characteristics of thin clouds.
[0064] 1) A simple atmospheric scattering model.
[0065] In computer vision and computer graphics, the atmospheric scattering model widely used to describe the formation of hazy images is shown in the following equation, which is based on two scattering phenomena: atmospheric light and attenuation.
[0066] I(x)=t(x)J(x)+(1-t(x))A (1)
[0067] Where x represents the position of a point in the image, J and I represent the clear image to be restored and the observed hazy image, respectively, t represents the transmittance of the medium, and A is the global atmospheric light. This model is also widely used in remote sensing images to describe the distortion of thin clouds.
[0068] 2) Atmospheric scattering model based on bidirectional transmission (ASTT).
[0069] For formula (1), t(x)J(x) represents the radiation intensity of sunlight reflected from the ground object and reaching the sensor after transmission, and (1-t(x))A represents the sunlight radiation reflected by clouds and fog. Since this formula is mainly used to describe atmospheric scattering in natural scenes, it only considers the portion of sunlight reflected from the ground object that reaches the sensor through clouds and fog, but does not consider the portion of sunlight that is absorbed when it reaches the ground object through clouds and fog. Therefore, using this model to simulate thin cloud images is not optimal, and the downward transmission of sunlight through clouds and fog should also be taken into account.
[0070] like Figure 1 As shown, the process of solar radiation reaching the ground and being reflected back to the sensor mainly involves two processes: reflection (reflection from clouds and reflection from the ground) and transmission (transmission of solar radiation through clouds and transmission of ground reflection through clouds). Both processes are related to cloud thickness. This embodiment does not consider diffuse reflection and absorption; therefore, the sum of cloud reflection and transmission can be considered as 1, thereby improving the atmospheric scattering model:
[0071] I(x)=dt(x)·ut(x)·r(x)·S+(1-dt(x))·S (2)
[0072] Where S represents solar radiation, r(x) represents ground reflectivity, and dt(x) and ut(x) represent downlink and uplink transmission, respectively. r(x)·S can be considered as the image under cloudless conditions, i.e., J(x). Furthermore, this embodiment assumes that the downlink and uplink rays penetrate the same cloud thickness, so dt(x) equals ut(x). Therefore, the cloud distortion model considering bidirectional transmission of clouds can be written in the following form:
[0073] I(x) = t(x) 2 J(x)+(1-t(x))S (3)
[0074] In the formula, t(x) = dt(x) = ut(x), and S can be regarded as the global atmospheric light A in formula (1).
[0075] (2) Estimation of transmittance taking into account the spatial morphology of thin clouds.
[0076] The problem of thin cloud removal is estimating J, t, and A from a single input image I. This is difficult because vectors A, I(x), and J(x) are coplanar, and their endpoints are geometrically collinear in RGB space. If the solar radiation S and transmittance t are known, the true scene J can be recovered.
[0077]
[0078] In thin cloud simulation studies, cloudless images J are relatively easy to obtain. Therefore, only the solar radiation S and transmittance t are needed to simulate the thin cloud distorted image I using formula (3). Usually, for simplification, S is set as a constant, generally 1. Therefore, as long as the transmittance can be obtained, the thin cloud image can be simulated. For the transmittance image of a thin cloud image, the transmittance of a pixel reflects the thickness of the thin cloud (which determines the magnitude of the transmittance), while its horizontal expansion reflects the spatial morphology of the thin cloud. Therefore, accurately estimating the transmittance image from the thin cloud image is the key to simulating the spatial morphology of the thin cloud. The transmittance calculation methods can be divided into two categories: one is to create noise simulation transmittance through mathematical algorithms without relying on prior knowledge, and the other is to extract transmittance with the help of prior knowledge related to thin clouds. Based on the transmittance estimation method based on dark channel prior (DCP-TE), this invention proposes a new transmittance estimation method based on prior knowledge related to thin clouds, namely, transmittance estimation based on cirrus band (Cirrus-TE), to better simulate the spatial morphological characteristics of thin clouds.
[0079] 1) Transmittance estimation based on dark channel prior (DCP-TE).
[0080] Transmittance estimation based on dark channel priors extracts transmittance from thin cloud images to ensure high similarity in spatial morphology to real thin clouds. For cloud and fog images captured by a camera in general scenes, transmittance can be expressed by the following formula:
[0081] t(x)=e -kd(x) (5)
[0082] This formula describes the light that is not scattered and reaches the camera, where d(x) is the distance from a point in the image to the camera, and k is the atmospheric scattering coefficient. For images in typical scenes, since the field of view that the camera can capture is relatively small, k can be considered a constant. Therefore, the transmittance t can be estimated by estimating the scene depth, i.e., d(x).
[0083] However, for remote sensing imagery, since the sensor captures images from space towards the ground, d(x) can be considered a constant. Considering that the atmospheric scattering coefficient varies depending on the concentration of air particles and the wavelength in different regions, its transmittance can be expressed as follows:
[0084] t(x)=e -k(x,λ)d(x) (6)
[0085] Scattering is essentially a diffraction phenomenon that occurs when electromagnetic waves encounter atmospheric particles during propagation. The scattering intensity and particle size are related to the wavelength, as shown in the following formula:
[0086]
[0087] Where γ∈[0,4], it mainly depends on the size of suspended particles in the atmosphere. Thin clouds are composed of smoke, aerosols, and small water droplets, and their scattering phenomenon follows the Mie scattering law, showing that the scattering intensity is inversely proportional to the square of the wavelength, so γ≤2. For clear sky regions, the size of the constituent molecules is much smaller than the wavelength, γ=4. For thick cloud regions, the size of the constituent water droplets is 1-10μm, much larger than the wavelength, so γ≈0. Some studies have also shown that under cloud and fog conditions, γ varies from 0.5 to 1.
[0088] Since γ is an indefinite quantity, the calculation of t(x) requires the use of other features in the image. Currently, the most common method is to estimate it using the Dark Channel Prior (DCP) principle, which states that in the non-sky regions of a cloudless image, each pixel has at least one channel value close to 0. Based on this, the DCP of a color image J can be described as:
[0089]
[0090] Among them, J dc Let x and y represent the dark channel of image J, and let J represent the pixels of image J. c Ω is a color channel of image J. d (x) is a local region of size d×d at pixel x.
[0091] After minimizing equation (3) twice, it is transformed into the following form:
[0092]
[0093] Then, substituting (8) into (9), we can get:
[0094]
[0095] Finally, the parameter ω (0≤ω≤1) is introduced to represent the sense of depth. Finally, t(x) can be described as follows:
[0096]
[0097] While this method can extract t(x) from a cloudy image I, the result is inevitably biased due to the influence of I. Furthermore, when the transmittance extracted by this method is applied to another region for simulation, the simulated image will additionally exhibit the texture of the original image's features, indicating that transmittance extracted based on dark channel priors cannot be effectively applied to other regions. This makes it difficult to extract transmittance from thin cloud images and combine it with cloudless images from other regions to simulate pseudo-thin cloud images that conform to the spatial distribution and distortion characteristics of thin clouds, thereby creating high-quality "cloudless-thin cloud" sample pairs.
[0098] 2) Transmittance estimation based on cirrus band (Cirrus-TE).
[0099] The cirrus cloud band in remote sensing imagery is a water vapor absorption band, which can reflect cloud thickness, and is especially sensitive to the thickness of thin clouds. Cloud thickness determines its transmittance, so the cirrus cloud band can be used to reflect cloud transmittance.
[0100] Based on the above principles, ideally, the transmittance should be similar when the DN value is the same in the cirrus cloud band. The transmittance (t) is estimated a priori based on the dark channel. dcp The transmittance is calculated based on an atmospheric scattering model. Although it suffers from some interference when applied to other regions, it can still reflect the spatial distribution and variation of transmittance to a certain extent. Therefore, this invention uses transmittance based on prior estimation of the dark channel to verify the above conjecture (i.e., the DN value of the cirrus band (DN)). c (This can reflect transmittance) and proposes corresponding estimation methods.
[0101] ①t dcp and DN c Analysis of the interrelationships.
[0102] First, this invention calculates t under different land cover types. dcp and DN c The relationship, the visualization results are as follows Figure 2 As shown, when the dendritic density (DN) values in the cirrus cloud bands are similar, the transmittance of forests is significantly higher than that of built-up areas. This is mainly because forest areas have more shadows, and these shadows themselves have lower DN values, thus resulting in higher transmittance obtained based on the dark channel prior principle. Furthermore, built-up areas exhibit some discrete anomalies with lower transmittance in areas with lower DN values in the cirrus cloud bands. This is primarily because some features in built-up areas have high reflectivity across all bands, leading to lower transmittance in these areas under the same thickness of thin cloud cover.
[0103] In contrast, water images have weaker background texture, and the cloudy pixels in water areas are less affected by the water background. Although clouds may have different physical properties (e.g., aerosol type and water molecule density) from water to land areas, their mechanisms of reflecting atmospheric light are essentially the same. Therefore, transmittance extracted from water areas more accurately reflects the true transmittance of clouds in the atmosphere.
[0104] like Figure 3 As shown, the density (DN) value and transmittance of cirrus clouds in the water region exhibit a significant linear relationship, with few outliers, primarily caused by ships in the water. Therefore, clouds in the water region can be analyzed to quantify the relationship between the DN value and transmittance of cirrus clouds.
[0105] The above is about t dcp and DN c The relationship analysis is performed at the scale of the entire image, while deep learning training is based on the image patch scale. Therefore, this embodiment analyzes the relationship between the two at the image patch scale, and the results are as follows: Figure 4 As shown. At the image patch scale, t dcp and DN c They maintain this high degree of linear correlation. Furthermore, their slopes and intercepts exhibit clustered distribution characteristics, indicating a high degree of consistency between them.
[0106] ② Adaptive minimum transmittance estimation.
[0107] Based on the above analysis, DN c The first step in converting to transmittance is to estimate the linear relationship between the two. This problem can be transformed into estimating the maximum and minimum transmittance values in the image, as shown in the following equation:
[0108]
[0109] In the formula, t dcp The transmittance is estimated based on the dark channel prior algorithm, DN c It is the DN value of the cirrus band.
[0110] The above formulas use t respectively dcp The minimum and maximum values of DN are used as the minimum and maximum transmittance, respectively. However, while the estimation of the maximum transmittance in the image conforms to the dark channel prior theory, the estimation of the minimum transmittance is affected by the type of ground cover. For example, ground covers such as ice and snow have high reflectivity and large dark channel values, resulting in a large transmittance calculated from them, thus causing estimation errors. To address this, this embodiment proposes an adaptive minimum transmittance estimation method, namely, taking DN... c and t dcp All of them are located at the intersection of the smallest 1% of pixels in the image, and the minimum value is taken as the transmittance.
[0111] II. A domain-adaptive remote sensing image thin cloud removal model.
[0112] In traditional machine learning algorithms, it is typically assumed that the probability distributions of the training and test sets are the same. Then, corresponding models and discrimination criteria are designed to predict the output of the test samples. However, in reality, the probability distributions of the training and test sets differ in many current learning scenarios. For example, due to factors such as illumination, color, sensor characteristics, and imaging quality, there are distributional differences between the two domains. Therefore, domain adaptation studies how to overcome the inconsistency between the source and target domain probability distributions, enabling learning tasks on the target domain. This addresses the performance degradation problem of deep neural networks when deployed to target domain data with heterogeneous features or scarce labeled data.
[0113] Based on the differences in how label information is utilized, domain adaptation methods can be divided into four categories: fully supervised domain adaptation, which assumes that both the source and target domains have sufficient data and corresponding label information, but this is relatively rare in reality because labels for the target domain are often difficult to obtain; semi-supervised domain adaptation, which assumes that the source domain data is fully labeled, while the target domain data is only partially labeled, this is closer to real-world applications because in many cases, it is feasible to collect labels for a portion of the target domain data; unsupervised domain adaptation, which is the most common case, where the source domain data has labels, while the target domain data is completely unlabeled; and self-supervised domain adaptation, which is usually between semi-supervised and unsupervised, not directly relying on externally provided label information, but utilizing the structural information of the data itself by constructing pre-tasks. Among these four types of methods, domain adaptation methods based on self-supervised learning have become increasingly popular in the field of deep learning in recent years. This method can be performed simultaneously on source and target domain data, which helps to learn the feature representations shared by the two domains, thereby achieving domain adaptation. For example, by constructing some self-supervised tasks in the source and target domains, such as image reconstruction and image rotation, it can be ensured that the network can learn the structural information of the target domain.
[0114] The thin cloud removal model proposed in this invention belongs to the domain adaptation method based on self-supervised learning. First, based on the concept of self-supervised learning, a simulated thin cloud image dataset created according to TCSCDSM is used as the training sample for the cloud removal model. This helps the cloud removal model learn more robust feature representations while avoiding insufficient data labeling samples. Then, the simulated dataset and the real dataset are regarded as the source domain and target domain, respectively. A mask attention mechanism is designed in combination with land cover data. The input mask and attention mask encourage the cloud removal network to use information from cloudless areas under the same land surface type to complete the information of thin cloud-occluded areas. Finally, a domain adaptation method is used to align the data distribution features of the source domain and the target domain from the input layer, feature layer and output layer, respectively, to enhance the generalization ability of the model and enable it to achieve good cloud removal results even without using or with very little target domain labeling.
[0115] (1) Overall architecture.
[0116] The overall architecture of the Domain Adaptive Thin Cloud Removal (DACR) model proposed in this invention is as follows: Figure 5 As shown, DACR proposes three domain alignment strategies: input layer alignment, feature layer alignment, and output layer alignment, to effectively reduce performance differences caused by domain offset. These three alignment strategies are seamlessly integrated into a unified model, thus allowing them to benefit from each other through an end-to-end training process.
[0117] Specifically, the domain alignment function of DARPR consists of three parts: The input layer's image transformation generator converts the simulated image in the source domain into an image similar to the real image in the target domain. This generator, combined with a discriminator, uses adversarial learning to enable the generator to transform the source domain image into a realistic target domain-like image. The feature layer's discriminator receives deep features from the source and target domain images output by the encoder, minimizing their differences to make the encoder's feature extraction processes for the source and target domains as similar as possible. The output layer distinguishes whether the input comes from the source or target domain, ensuring that the decoder's outputs from the source and target domain images are as close as possible. Through this three-layer domain alignment strategy, the encoder and decoder are encouraged to produce declouding results in the target domain similar to those in the source domain.
[0118] (2) Domain alignment strategy.
[0119] Based on the location of the domain-adaptive alignment strategy proposed in this embodiment within the network, it can be categorized into three types: input layer, feature layer, and output layer alignment. These will be described in detail below:
[0120] 1) Input layer alignment.
[0121] Due to domain offset, images from different domains often exhibit distinct visual characteristics. Input layer alignment reduces the domain offset between the source and target domains by transforming the image appearance. Specifically, given a dataset from the source domain... And a dataset from the target domain. The goal of input layer alignment is to align the source image x s Convert to target image x t The resulting transformed image should appear as if it were extracted from the target domain, while the original content and its structural semantics remain unaffected.
[0122] This embodiment uses a simulated thin cloud dataset as the source domain and a real thin cloud dataset as the target domain. When the source domain image... Based on target domain image When simulating transmittance, an input layer alignment strategy is executed. This is achieved by constructing a generator G. i and Discriminator D i Generative adversarial networks (GANs) are used for pairwise image-to-image transformation. The generator aims to transform source images into images similar to the target domain, i.e., G. i (x s )=x s→t The discriminator competes with the generator to correctly distinguish spurious transformed images x. s→t and the real target image x s The two interact and optimize through adversarial learning. The adversarial loss for input layer alignment is:
[0123]
[0124] The discriminator attempts to maximize the target to distinguish G. i (x s )=x s→t and x t At the same time, the generator needs to minimize the objective to make x s Transform it into an image similar to the target domain.
[0125] 2) Feature layer alignment.
[0126] Through the aforementioned input layer alignment, the cloud removal network E♀D trained on the transformed target domain image can already achieve considerable performance on target domain data. However, when domain shift is severe, it is still insufficient to achieve the desired domain adaptation results. Therefore, this embodiment uses an additional domain discriminator D. f Perform feature layer alignment to further reduce the size of the transformed image x. s→t and the real target domain image x t Domain offset between them. Specifically, adversarial learning is applied directly to the feature space, making it impossible for the discriminator to distinguish which features come from which domain.
[0127] like Figure 5 As shown, D f The high-dimensional features of the source and target domains output by the encoder are used as input. These features should be consistent in the high-dimensional space after alignment by the input layer. Otherwise, adversarial gradients are backpropagated to the feature extractor E in order to minimize the gradient from x. s→t To x t The distance between the feature distributions. The adversarial loss for feature layer alignment is:
[0128]
[0129] 3) Output layer alignment.
[0130] For the output results after image transformation and E♀D, an output layer discriminator D is further added. o Construct an auxiliary task to distinguish whether the declouded image is from x s→t The reconstruction is still based on the real target domain image x. t Transformation. If discriminator D o Successfully classifying the domain of the generated image means that the extracted features still contain domain features. To ensure feature regions remain invariant, the following adversarial loss is used to supervise the feature extraction process:
[0131]
[0132] During domain alignment, the encoder and decoder are encouraged to extract domain-invariant features by connecting discriminators from three aspects. Adversarial learning within the feature space can effectively solve the problem of synthesizing thin cloud images x. s and real thin cloud image x t The gaps between domains.
[0133] (3) Cooperative learning encoder-decoder architecture.
[0134] In the collaborative learning framework proposed in this embodiment, seamless integration of input layer alignment and feature layer alignment is first achieved through an encoder with shared weights. Then, output layer alignment is fused through the encoder. Finally, a unified framework is trained in an end-to-end manner. Specifically, after the input layer alignment process, the encoder E utilizes the input layer alignment from the discriminator {D}. f D o Collect and combat losses and To optimize, the decoder utilizes the data from the discriminator D. o Collect and fight against losses Optimize.
[0135] The encoder-decoder architecture for DACR can adopt various network architectures from the field of image dehazing or thin cloud removal, such as the Pix2Pix model, which is the most common network in image reconstruction and has been proven to be effective and stable during training for image-to-image translation in different domains. After considering both training stability and memory consumption, the encoder-decoder structure of MSBDN was chosen. In each training iteration, all modules are updated sequentially in the following order: G→D i →E→D f →D→D o Specifically, the generator G is first updated to obtain the transformed class-target domain image. Then, the discriminator D... i Updated to classify target image x s→t and the real target image x t Next, the encoder E is updated to use the data from x. s→t and x t Features are extracted from the discriminator D, and then the discriminator D is updated. f The decoder D is then updated to convert the extracted features into a cloudless image, based on its input features. Finally, the discriminator D... o The output features are updated to their source and target domains, thereby enhancing feature invariance.
[0136] 1) Masked attention mechanism.
[0137] Furthermore, this embodiment leverages the advantages of a model architecture based on self-supervised learning to design a Mask Attention Mechanism (MAM). MAM artificially constructs a mask on the data input to the encoder, enabling the model to reconstruct data from thin cloud areas using data from cloudless regions, thus improving the thin cloud removal process in remote sensing images. MAM consists of two parts: an input mask and an attention mask.
[0138] The core idea of input masking is to utilize land cover data to identify key areas with minimal surface changes (such as water bodies) and combine this with thin cloud concentration data to generate accurate masks within these key areas. These masks help the model selectively focus on data from cloudless areas when performing cloud removal tasks, thereby using this data to reconstruct areas covered by thin clouds. The characteristic of input masking lies in its special handling capability for areas sensitive to surface changes. Through a sophisticated masking mechanism, the cloud removal model can more effectively utilize information from cloudless areas to reconstruct thin cloud regions. The advantage of this method is that it not only improves the accuracy of thin cloud removal but also enhances the model's adaptability to complex surface conditions while maintaining the continuity of surface features. This method randomly masks certain pixels of the input thin cloud image in areas with smooth image changes based on land cover data and encourages the network to complete the masked information during training. Input masking explicitly constructs a very challenging inpainting problem, forcing the network to reconstruct the masked content in the image using contextual information. Thin clouds themselves are equivalent to image occlusion. Therefore, the input mask can encourage de-clouding networks to use information from cloudless areas of the same land surface type to complete the information of the thin cloud-occluded areas.
[0139] While input masking has good results, it's not sufficient to build a cloud-free network solely based on input masking operations. During testing, an undamaged image is input to retain sufficient information. However, due to the inconsistency between training and testing, the network tends to increase the brightness of the output image. Attention mechanisms are commonly used in deep learning models to process spatial and spectral information; therefore, the training-test gap can be narrowed by performing the same masking operation within the attention mechanism. Attention masking is similar to input masking, but uses a different attention masking ratio. When some markers in the attention result are masked, the model adapts to the fact that the information from those markers is no longer reliable. Figure 6 The results of the attention mask are shown, which is added to the feature extractor of each layer in the model.
[0140] The method in this embodiment has significant advantages over other methods:
[0141] I. Experimental Dataset.
[0142] The dataset used in the experiment was created based on TCSCDSM, with the simulated dataset as the source domain, the training set from the real dataset as the target domain, and the test set used to evaluate the cloud removal effect. The simulated dataset expanded the real dataset by 8 times, that is, creating 8 different simulated thin cloud images on a cloudless image using different transmittances.
[0143] Furthermore, since the MAM proposed in this invention requires land cover information for each image block, this embodiment acquired MODIS MCD12Q1.061 land cover data. The classification standard used is LC_Prop2, and the data year is consistent with the acquisition year of the Landsat-8 image. The land cover types included are as follows:
[0144] Table 1. Surface types and their descriptions under the LC_Prop2 standard for MCD12Q1.061 data.
[0145]
[0146] II. Evaluation Indicators and Parameter Settings.
[0147] To quantitatively evaluate the effectiveness of thin cloud removal using the method of this invention, this invention introduces commonly used evaluation metrics in the field of image processing: Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), and Learned Perceptual Patch Similarity (LPIPS). PSNR is a metric for evaluating image quality, defined based on the root mean square error, and used to measure the image quality reference value between the maximum signal and background noise. Structural Similarity (SSIM) is a metric used to quantify the structural similarity between two images, mimicking the human visual system to measure perceptual sensitivity to local structural changes in an image. Learned Perceptual Patch Similarity (LPIPS), also known as "perceptual loss," is a metric that conforms to human perception and measures the differences between two images.
[0148] In model training, DARC was implemented in the PyTorch framework and trained on an NVIDIA GeForce RTX 3080Ti. During training, the batch size was set to 4, and the number of epochs was set to 50. Huber loss was chosen as the loss function, a combination of squared loss and absolute loss (L1 loss), exhibiting the smoothness of L2 loss when the error is small, while being less sensitive to outliers like L1 loss when the error is large. Furthermore, gradient descent was performed using the Adam optimizer with β1 = 0.5 and β2 = 0.999, and a "poly" learning rate descent strategy was employed to update the learning rate, with the initial learning rate and power set to 2e-4 and 0.9, respectively.
[0149] III. Ablation Experiment.
[0150] (1) Ablation experiment of MAM:
[0151] The ablation experiment results for MAM are shown in Table 2. First, without using the domain adaptation method, incorporating MAM into the model improves PSNR, SSIM, and LPIPS by 1.2%, 1.6%, and 1.2%, respectively, compared to using only the baseline network. Furthermore, by combining the domain alignment strategy, the performance is further improved by 1.1%, 0.5%, and 1.5%, respectively.
[0152] Table 2 Ablation Experiment Results of SS-MAM
[0153]
[0154] (2) Ablation experiment of domain alignment strategy.
[0155] This embodiment proposes three types of domain alignment strategies, each utilizing the labeled data differently. Input layer alignment and feature layer alignment do not require access to the target domain's labels, while output layer alignment does require access to the target domain's labels; therefore, they can be categorized into unsupervised domain adaptation and semi-supervised domain adaptation. Ablation experiments are conducted using each alignment strategy to analyze its effectiveness.
[0156] 1) Unsupervised domain-adaptive ablation experiments: This involves not using an output layer alignment strategy. This only requires accessing sample data from the target domain, without needing its label data. The results are shown in Table 4. Compared to the results in Table 2 without using a domain adaptation strategy, the cloud removal effect is significantly improved after using domain adaptation. Furthermore, input layer alignment performs better than feature layer alignment. Using both together improves PSNR, SSIM, and LPIPS by 6.3%, 3.8%, and 3.2%, respectively, compared to not using any domain alignment strategy.
[0157] 2) Ablation Experiments for Semi-Supervised Domain Adaptation: Using an output layer alignment strategy in the model can be considered semi-supervised domain adaptation because output layer alignment requires access to the target domain's label data. As shown in Table 3, output layer alignment performs best among the three alignment strategies, mainly because it uses the target domain's labels. Integrating the three alignment methods into the model improves PSNR, SSIM, and LPIPS by 10.3%, 5.7%, and 5.7%, respectively, compared to not using a domain alignment strategy.
[0158] Table 3 Ablation experimental results of different domain alignment strategies
[0159]
[0160] IV. Comparative Analysis.
[0161] This invention evaluates the performance of DARC by comparing it with six state-of-the-art methods, including Pix2Pix, SpAGan, McGan, DANet, CycleGan, and ColorMapGAN. Pix2Pix is a classic image dehazing model that also demonstrates good performance in thin cloud removal tasks. SpAGan, McGan, DANet, CycleGAN, and ColorMapGAN are designed for thin cloud removal, with the latter three based on domain adaptation algorithms. This embodiment evaluates the cloud removal effect of DARC by comparing it with other methods from both quantitative and qualitative perspectives.
[0162] Table 4 lists the quantitative results of all comparison methods. Under supervised learning training (training on the real dataset), DCR achieved the best cloud removal performance, showing improvements of 7.7%, 2.4%, and 5.7% in PSNR, SSIM, and LPIPS compared to McGan, the best-performing model among the others. Under self-supervised learning training, the accuracy improved significantly compared to supervised learning due to the greatly increased training samples. However, the proposed DCR still achieved the best cloud removal performance, showing improvements of 6%, 1.6%, and 6.7% in PSNR, SSIM, and LPIPS compared to McGan. Finally, DCR's domain adaptation capability was evaluated by comparing it with DANet, CycleGAN, and ColorMapGAN. Compared to ColorMapGAN, which has the strongest generalization ability among the three, DCR achieved improvements of 4.4%, 2.9%, and 3.5% in PSNR, SSIM, and LPIPS, respectively. In summary, the architecture and modules designed for thin cloud removal by DACR have shown excellent performance, and its domain adaptation capability is also stronger than that of existing models.
[0163] Table 4. Performance comparison of DARC and other thin cloud removal methods
[0164]
[0165] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.
Claims
1. A remote sensing image thin cloud removal method based on self-supervised domain adaptation, characterized in that, The method comprises the following steps: 1) obtaining a remote sensing image to be thinned cloud removed; 2) inputting the remote sensing image into a trained remote sensing image thin cloud removal model based on domain self-adaption to obtain a remote sensing image with thin clouds removed; The training process of the remote sensing image thin cloud removal model based on domain self-adaption comprises: A thin cloud simulation model considering cloud distortion and spatial form is adopted to create a simulated thin cloud image dataset as a sample for training the thin cloud removal model, the thin cloud simulation model considering cloud distortion and spatial form is composed of a bidirectional transmission-based atmospheric scattering model and a transmittance estimation method based on cirrus band, and is respectively used for simulating the cloud distortion characteristics and spatial form characteristics of thin clouds; The simulated dataset and the real dataset are respectively regarded as a source domain and a target domain, a mask attention mechanism is designed in combination with land cover data, and the cloud removal network is caused to utilize the information of cloud-free areas under the same ground surface type to complete the information of thin cloud blocked areas by inputting a mask and an attention mask respectively; The domain self-adaption method is adopted to align the data distribution characteristics of the source domain and the target domain from the input layer, the feature layer and the output layer respectively.
2. The method of claim 1, wherein, The bidirectional transmission-based atmospheric scattering model is: I(x) = t(x) 2 J(x) + (1 - t(x)) S Wherein, x represents the position of a certain point in the image, J and I respectively represent the clear image to be restored and the observed hazy image, t represents the medium transmittance, t(x)=dt(x)=ut(x), S is the solar radiation, and dt(x) and ut(x) are respectively the downward transmission and upward transmission. 3.The method of claim 2, wherein, The transmittance estimation method based on cirrus band is: where t dcp is the transmittance estimated based on the dark channel prior algorithm, DN c is the DN value of the cirrus band.
4. The method of claim 3, wherein, The domain self-adaption method adopted to align the data distribution characteristics of the source domain and the target domain from the input layer, the feature layer and the output layer respectively comprises: Three domain alignment strategies of input layer alignment, feature layer alignment and output layer alignment; The input layer alignment strategy comprises that an image transformer of the input layer is used to convert the simulated image of the source domain into an image similar to the real image of the target domain, and then the image transformer is combined with a discriminator to enable the generator to convert the source domain image into a realistic target domain-like image through adversarial learning; The feature layer alignment strategy comprises that a discriminator of the feature layer is responsible for receiving the deep features of the source domain image and the target domain image output by an encoder, and narrowing the difference between the deep features to make the feature extraction processes of the source domain and the target domain similar by the encoder; The output layer alignment strategy comprises that the output layer distinguishes whether the input is from the source domain or the target domain, so that the output results of the source domain and the target domain images through the decoder are as close as possible.
5. The method of claim 4, wherein, When the source domain image is simulated based on the transmittance of the target domain image, the input layer alignment strategy is performed.
6. The method of claim 5, wherein, The adversarial loss of the input layer alignment is: wherein, is a dataset from a source domain, is a dataset from a target domain, G i (x s ) = x s→t denotes a generator that converts source images to images similar to the target domain, D i denotes D i discriminator.
7. The method of claim 6, wherein, The feature layer alignment is performed by adding a domain discriminator; The domain discriminator takes the high-dimensional features of the source domain and the target domain output by the encoder as input, if the domain discriminator can distinguish the high-dimensional features of the source domain and the target domain, the adversarial gradient is back-propagated to the feature extractor, and the adversarial loss of the feature layer alignment is: where D f represents D f domain discriminator, E represents an encoder.
8. The method of claim 7, wherein, By adding an output layer discriminator to distinguish whether the de- clouded image is from x s→t reconstructed or transformed from real target domain image x t If the output layer discriminator gets accurate discrimination results, the features extracted by the feature extractor contain domain features. In order to make the feature region invariant, the following adversarial loss is used to supervise the feature extraction process: where D o represents D o Output layer discriminator, D represents decoder.
9. The method of claim 8, wherein, In training, the encoder E is optimized by collecting adversarial loss from discriminator {D f ,D o} after input layer alignment process and the decoder is optimized by collecting adversarial loss from discriminator D o . 10. The method of claim 9, wherein, In each training iteration, all modules are updated in the following order: The generator G is first updated to obtain a transformed class target domain image. Then, the discriminator D i is updated to distinguish between the class target image x s→t and the real target image x t . Next, the encoder E is updated for extracting features from x s→t and x t , and subsequently the discriminator D f is updated to input features to it, The decoder D is then updated to convert the extracted features into cloud-free images. Finally, the discriminator D o is updated to output features for both its source and target domains, thereby enhancing feature invariance.