Density-Aware Cloud Removal Method Based on Global-Local Fusion Transformer

By introducing global-local fusion transformer and density perception technology into the cloud removal method, using cloud coverage category tags to guide information flow, the problem of the existing method failing to reconstruct under different cloud densities is solved, and high-precision and low-cost cloud removal effect is achieved.

CN118967510BActive Publication Date: 2025-06-10UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410941792.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-15
Publication Date
2025-06-10
Estimated Expiration
2044-07-15

AI Technical Summary

Technical Problem

Existing cloud removal methods have the problem of reconstruction failure when dealing with different cloud densities, and rely on cloud occlusion maps or synthetic aperture radar images to increase labor costs.

Method used

The density-aware cloud removal method based on global-local fusion transformer is adopted to guide the information flow between different channels of the hyperspectral cloud coverage image through the cloud coverage category label, thereby improving the cloud removal accuracy.

Benefits of technology

Significantly improves the accuracy and effectiveness of cloud removal, reduces the introduction of noise and artifacts, and reduces labor costs, ensuring the integrity of the underlying surface details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118967510B_ABST
    Figure CN118967510B_ABST
Patent Text Reader

Abstract

The present invention discloses a density-aware cloud removal method based on a global-local fusion transformer, which improves the accuracy of cloud removal by using cloud cover class labels to guide the information flow between different channels of hyperspectral cloud-covered images. Specifically, the features of the cloud-covered image are first linearly transformed into queries, keys, and values, then the self-attention weights are calculated, and the attention map is optimized using cloud cover class labels. Finally, the relationship between all local window features in the optical cloud-covered image is guided by a global fusion method, significantly improving the accuracy and effect of cloud removal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and more specifically, relates to a density-aware cloud removal method based on a global-local fusion transformer. Background Art

[0002] In the field of remote sensing image processing, cloud cover is a common challenge affecting the extraction and analysis of ground features. Since the launch of the Terra and Aqua satellite platforms in 2000 and 2002 respectively, continuous observations have obtained various cloud property data of MODIS. The MODIS cloud mask data shows that the global cloud cover rate is about 67%, and the cloud cover rate in the land area is about 55%.

[0003] Given a cloud-covered image J, defined in the space where D and P represent the cloud-covered and cloud-free regions respectively, the goal of cloud removal is to restore the cloud-covered part J of the image D . Due to the significant information loss caused by cloud occlusion in optical satellite observations, this task is inherently challenging.

[0004] Traditional cloud removal strategies are similar to inpainting tasks, aiming to use the observable part J P to reconstruct the cloud-occluded region J D :

[0005] J D = G INP (J P ; T(J))

[0006] where G INP is an inpainting operator guided by the latent structure T(J) of the entire image. These structures may include image priors such as smoothness and non-local similarity, or features extracted from the data to facilitate the restoration process. However, for different cloud densities, these latent structures may be incomplete or missing, resulting in potential reconstruction failures.

[0007] However, although multi-temporal image cloud removal methods have improved the accuracy of the algorithm, they have extremely strict requirements for images and have significant limitations. The cloud removal task of a single remote sensing image is of great significance in supporting emergency response and disaster management. Summary of the Invention

[0008] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a density-aware cloud removal method based on a global-local fusion transformer, which uses cloud cover class labels to guide the information flow between different channels of hyperspectral cloud-covered images, thereby improving the cloud removal accuracy and solving the problems such as the increase in labor costs caused by relying on cloud occlusion maps or synthetic aperture radar images in existing methods.

[0009] To achieve the above-mentioned invention purpose, a density-aware cloud removal method based on a global-local fusion transformer according to the present invention is characterized by including the following steps:

[0010] (1) Extract the shallow features of the cloud-covered image;

[0011] Gmsi = SFEmsi(J)

[0012] where Gmsi represents the shallow features of the cloud-covered image, SFEmsi(·) is the shallow feature extraction function, and J represents the cloud-covered image;

[0013] (2) Extract the cloud coverage category;

[0014] Flatten the shallow features Gmsi, then input the flattened features into the cloud density classifier to obtain the cloud coverage category Glabel, and expand the cloud coverage category Glabel to the same spatial dimension as the shallow features Gmsi;

[0015] (3) Perform cloud removal operations on the cloud-covered image through the global-local fusion transformer;

[0016] (3.1) Assume that the global-local fusion transformer contains N layers, and each layer is composed of a density-aware global context interaction module and a spatial feature transformation module;

[0017] (3.2) Use the shallow features Gmsi and the cloud coverage category Glabel as the initial values of the input of the global-local fusion transformer, denoted as

[0018] (3.3) Calculate the input of the i-th layer of the global-local fusion transformer and

[0019]

[0020] where DGCI(·) represents the density-aware global context interaction module, and SFT(·) represents the spatial feature transformation module;

[0021] (3.4) Convert the shallow features into the query features of the cloud-containing image key features value features

[0022] Convert the cloud coverage category into the query features of the cloud coverage category key features value features

[0023] (3.5), Calculate the feature transformation matrix of the cloud-containing image and the feature transformation matrix of the cloud cover category

[0024]

[0025] where d is the query matrix or The corresponding number of rows, the superscript T represents the transpose, and B is a learnable position encoding matrix;

[0026] (3.6), Obtain the cloud cover-guided feature transformation matrix of the cloud-containing image through the gating function

[0027]

[0028] where Gate(·) represents the gating function that adjusts the attention weight, and ⊙ represents element-wise multiplication;

[0029] (3.7), Calculate the cloud-containing image features output by the density-aware global context interaction module in the i-th layer of the global-local fusion transformer and the cloud cover category features

[0030]

[0031] (3.8), Input the cloud-containing image features and cloud cover category features output by the density-aware global context interaction module in the i-th layer into the spatial feature transformation module to generate the mother parameter g of the scaling parameter γ and the offset parameter β;

[0032]

[0033] where ReLU(·) represents the activation function, and Conv(·) represents the convolution operation;

[0034] (3.9), Determine the scaling parameter γ and the offset parameter β according to the mother parameter g;

[0035]

[0036] γ = Sigmoid(γ)

[0037] where split(·) represents the splitting function, and Sigmoid(·) represents the activation function;

[0038] (3.10), Adjust the cloud-containing image features according to the scaling parameter γ and the offset parameter β

[0039]

[0040] (3.11) Repeat steps (3.3) to (3.10) to obtain the adjusted cloud image features of each layer of the global-local fusion transformer

[0041] (3.12) Reconstruct the cloud-free image J of high quality D ;

[0042]

[0043] Among them, IR(·) represents the cloud-free image reconstruction function.

[0044] The invention purpose of the present invention is realized as follows:

[0045] The density-aware cloud removal method based on the global-local fusion transformer of the present invention improves the accuracy of cloud removal by guiding the information flow between different channels of the hyperspectral cloud-covered image using the cloud coverage category label; specifically, the features of the cloud-covered image are first linearly transformed into queries, keys, and values, then the self-attention weights are calculated, and the attention map is optimized using the cloud coverage category label. Finally, the relationship between all local window features in the optical cloud-covered image is guided through the global fusion method, significantly improving the accuracy and effect of cloud removal.

[0046] Meanwhile, the density-aware cloud removal method based on the global-local fusion transformer of the present invention also has the following

[0047] Beneficial effects:

[0048] (1) The present invention introduces a cloud density classifier, which guides the cloud removal model to learn the spectral relationship of the multispectral image by initially estimating the cloud coverage situation, thereby improving the cloud removal performance.

[0049] (2) The present invention uses the global-local fusion transformer (DCR-GLFT) to combine the cloud density information with the cloud-ground image features. This method promotes more detailed and accurate cloud removal, ensures the integrity of the underlying surface details, and minimizes the introduction of noise and artifacts at the same time.

[0050] (3) The present invention demonstrates state-of-the-art performance on the well-known SEN12MS-CR cloud removal dataset. Even without using SAR data as guidance, it still achieves a PSNR of 28.93 and an SSIM of 0.838. This achievement highlights its significant progress in the field of single-image cloud removal. Brief Description of the Drawings

[0051] Figure 1 is the flow chart of the density-aware cloud removal method based on the global-local fusion transformer of the present invention;

[0052] Figure 2 It is a schematic diagram of flattening shallow features;

[0053] Figure 3 It is a schematic diagram of the density-aware global context interaction module;

[0054] Figure 4 It is a schematic diagram of the spatial feature transformation module;

[0055] Figure 5 It is a comparison diagram of cloud removal under different scenarios. Detailed implementation manners

[0056] The following describes the detailed implementation manners of the present invention with reference to the accompanying drawings, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.

[0057] Embodiment

[0058] Figure 1 It is a flowchart of the density-aware cloud removal method based on the global-local fusion transformer of the present invention.

[0059] In this embodiment, as Figure 1 shown, a density-aware cloud removal method based on the global-local fusion transformer of the present invention includes the following steps:

[0060] (1), Extract the shallow features of the cloud-covered image;

[0061] G msi = SFE msi (J)

[0062] wherein, G msi represents the shallow features of the cloud-covered image, SFE msi (·) is the shallow feature extraction function, and J represents the cloud-covered image;

[0063] (2), Extract the cloud coverage category;

[0064] Flatten the shallow features G msi , as Figure 2 shown, to convert it into a two-dimensional format suitable for classification; then input the flattened features into the cloud density classifier to obtain the cloud coverage category G label , and expand the cloud coverage category G label to the same spatial dimension as the shallow features G msi to match the spatial dimension of the input image and ensure the consistency of subsequent processing.

[0065] (3) Perform cloud removal operation on the cloud-covered image through the global-local fusion transformer;

[0066] (3.1) Assume that the global-local fusion transformer contains N layers, and each layer consists of a density-aware global context interaction module and a spatial feature transformation module;

[0067] (3.2) Take the shallow feature G msi and the cloud coverage category G label as the initial values of the input of the global-local fusion transformer, denoted as

[0068] (3.3) Calculate the input of the i-th layer of the global-local fusion transformer and

[0069]

[0070] where DGCI(·) represents the density-aware global context interaction module, and SFT(·) represents the spatial feature transformation module;

[0071] In this embodiment, the density-aware global context interaction module is as Figure 3 shown, integrating a vision transformer after each convolutional layer, splitting the input features into non-overlapping windows, and independently calculating the self-attention mechanism for each window; the spatial feature transformation module is as Figure 4 shown, used to calculate the adjustment parameters of the guiding features from the density labels to adjust the features of the input image;

[0072] (3.4) Convert the shallow feature into the query feature key feature value feature

[0073] Convert the cloud coverage category into the query feature key feature value feature

[0074]

[0075] (3.5) Calculate the feature transformation matrix of the cloud image and the feature transformation matrix

[0076]

[0077] where d is the query matrix or The corresponding row number, the superscript T represents transpose, and B is a learnable position encoding matrix;

[0078] (3.6), Obtain the cloud-covered guided cloud image feature transformation matrix through the gating function

[0079]

[0080] Among them, Gate(·) represents the gating function for adjusting the attention weight, and ⊙ represents element-wise multiplication;

[0081] (3.7), Calculate the cloud image features and cloud coverage category features output by the density-aware global context interaction module in the i-th layer of the global-local fusion transformer as well as the cloud coverage category features

[0082]

[0083] (3.8), Input the cloud image features and cloud coverage category features output by the density-aware global context interaction module in the i-th layer into the spatial feature transformation module to generate the parent parameter g of the scaling parameter γ and the offset parameter β;

[0084]

[0085] Among them, ReLU(·) represents the activation function, and Conv(·) represents the convolution operation;

[0086] (3.9), Determine the scaling parameter γ and the offset parameter β according to the parent parameter g;

[0087]

[0088] γ = Sigmoid(γ)

[0089] Among them, split(·) represents the splitting function, and Sigmoid(·) represents the activation function;

[0090] (3.10), Adjust the cloud image features according to the scaling parameter γ and the offset parameter β

[0091]

[0092] (3.11), Repeat steps (3.3) to (3.10) to obtain the adjusted cloud image features for each layer of the global-local fusion transformer

[0093] (3.12), Reconstruct the high-quality cloud-free image J D ;

[0094]

[0095] Among them, IR(·) represents the cloudless image reconstruction function.

[0096] In this embodiment, the SEN12MS-CR dataset is used for test experiments. This dataset is a multi-modal and single-temporal dataset, originating from the Copernicus mission satellites Sentinel-1 and Sentinel-2 of the European Space Agency, and is designed specifically for cloud removal in Earth observation. The dataset includes multi-spectral optical satellite observation data, covering both cloudy and cloudless conditions. The multi-spectral data includes multiple spectral channels, such as visible light, near-infrared, and short-wave infrared bands. In this embodiment, 175 regions of interest (ROIs) are randomly divided into 165 for training, 10 for validation, and 10 for testing. To prevent the overall performance from being biased towards any specific season, images from four different seasons are assigned to the validation set and the test set. Specifically, the training set, validation set, and test set contain 107,816, 7,655, and 6,747 samples respectively. The results of cloud removal are evaluated by normalizing the data based on peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), spectral angle mapper (SAM), and mean absolute error (MAE). The results are shown in the following table:

[0097] PSNR SSIM SAM MAE Experimental results 28.93 0.838 7.948 0.027

[0098] The experimental results are as Figure 5 shown, and are divided into two scenarios: (a) and (b). For each scenario, from left to right are the cloudy image, the cloud-removed image reconstructed by the method of the present invention, and the true cloudless image. Therefore, the cloud-removed image reconstructed by the present invention is similar to the true cloudless image and can quickly remove cloud cover.

[0099] Although the above-described illustrative specific embodiments of the present invention have been described to facilitate the understanding of the present invention by those skilled in the art, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.

Claims

1. A density-aware cloud removal method based on global-local fusion transformer, characterized in that: The following steps are involved: (1) Extract shallow features of cloud cover images; G msi =SFE msi (J) Among them, G msi Shallow features representing cloud cover images, SFE msi (·) is the shallow feature extraction function, J represents the cloud cover image; (2) Extract cloud cover categories; The shallow feature G msi Flatten the features and input them into the cloud density classifier to get the cloud cover category G label , and the cloud cover category G label Extended to the shallow feature G msi Same spatial dimensions; (3) Perform cloud removal on the cloud-covered image through a global-local fusion transformer; (3.1), suppose the global-local fusion transformer contains N layers, each layer consists of a density-aware global context interaction module and a spatial feature transformation module; (3.2) The shallow feature G msi and cloud cover category G label As the initial value of the global-local fusion transformer input, it is recorded as (3.3) Calculate the input of the i-th layer of the global-local fusion transformer and Among them, DGCI(·) represents the density-aware global context interaction module, and SFT(·) represents the spatial feature transformation module; (3.4) The shallow features Converted into query features with cloud images Key Features Value Features Cloud Cover Category Query features converted into cloud cover categories Key Features Value Features (3.5) Calculate the feature transformation matrix of the cloud image and cloud cover category feature transformation matrix Where d is the query matrix or The corresponding row number, the superscript T indicates transposition; (3.6) Obtain the cloud image feature transformation matrix guided by cloud coverage through the gating function Among them, Gate(·) represents the gating function for adjusting the attention weight, and ⊙ represents element-by-element multiplication; (3.7) Calculate the cloud image features output by the density-aware global context interaction module in the i-th layer of the global-local fusion transformer And cloud cover category characteristics (3.8) The cloud coverage category features output by the density-aware global context interaction module in the i-th layer Input to the spatial feature transformation module to generate the mother parameter g of the scaling parameter γ and the offset parameter β; Among them, ReLU(·) represents the activation function, and Conv(·) represents the convolution operation; (3.9), determine the scaling parameter γ and the offset parameter β according to the mother parameter g; γ=Sigmoid(γ) Among them, split(·) represents the segmentation function, Sigmoid(·) represents the activation function; (3.10) Adjust the cloud image features according to the scaling parameter γ and the offset parameter β (3.11) Repeat steps (3.3) to (3.10) to obtain the cloud image features after adjustment of each layer of the global-local fusion transformer (3.12), reconstruct high-quality cloud-free imagesJ D ; Where IR(·) represents the cloud-free image reconstruction function.

Citation Information

Patent Citations

  • Cloud image segmentation method based on fusion transformation network

    CN116091764A

  • SAR-assisted optical remote sensing image restoration method

    CN116309150A