A method and system for removing cloud areas and shadows from multimodal remote sensing images
Through the spectrum-feature-structure maintenance and combining deep learning and multi-level network structure, the problem of insufficient spatial structure information extraction in cloud and shadow removal of SAR image data is solved, and the cloud and shadow removal effect with high spectral fidelity and high structural clarity is achieved.
Patent Information
- Application Number
- CN202310733829.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-19
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-06-19
AI Technical Summary
When the prior art uses SAR image data and optical image data to remove cloud areas and shadows, the spatial structure information of SAR data cannot be fully extracted, resulting in blurred structures of cloud areas and shadow removal results.
The multimodal remote sensing image cloud area and shadow removal method based on the spectrum-feature-structure maintenance fusion network is adopted. Through deep learning methods combined with multi-level network structure design, features are extracted using variable convolution and variable multi-head self-attention mechanism, and combined with the first-order gradient operator to map the spectral domain to the structural domain, a three-level collaborative optimization loss function is designed for training.
The cloud area and shadow removal results that maintain high spectral fidelity and high structural clarity are achieved, which improves the quality of reconstructed cloudless optical remote sensing images, and solves the problem of insufficient extraction of spatial structure information in the existing methods.
Smart Images

Figure CN116977202B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - field of multi - modal remote sensing data and artificial intelligence, and relates to a method for removing cloud areas and shadows from multi - modal remote sensing images based on a spectral - feature - structure - preserving fusion network. Specifically, it includes a method for removing cloud areas and shadows by using synthetic aperture radar (SAR) data and optical remote sensing image data as inputs to collaboratively optimize the information reconstruction in the spectral domain, structural domain, and feature domain. Background Art
[0002] Continuous earth observation plays a crucial role in many research and monitoring tasks in the field of remote sensing, such as agricultural surveys, environmental protection, disaster prevention and mitigation, ocean development, and urbanization research. With the rapid development of remote sensing technology, using optical remote sensing images for surface monitoring has gradually become the mainstream method. However, optical remote sensing images are inevitably contaminated by cloud areas and cloud shadows, resulting in the inability to form continuous observations of the Earth's surface. According to the data analysis of the United States Geological Survey (USGS), the global annual average cloud cover rate reaches 66%. In addition, according to the statistics of Landsat ETM+ data, about 35% of the land is covered by clouds. It can be seen that the occlusion of clouds and shadows will greatly affect various earth - monitoring tasks, especially those that require consistent time series and must observe specific scenes at specific times, such as agricultural monitoring and natural disaster monitoring, thereby greatly reducing the usability of optical remote sensing images. To ensure seamless observation of the Earth's surface, removing cloud areas and shadows from optical remote sensing images has become an urgent problem to be solved.
[0003] Optical remote sensing image cloud area and shadow removal aims to reconstruct the information features lost due to cloud contamination by using complementary information. According to the different types of complementary information, cloud area and shadow removal methods can be divided into three categories: single-image data reconstruction method, multi-temporal data fusion method, and multi-modal data fusion method. The single-image data reconstruction method uses the original scene information from cloud-free areas or other spectra as a supplement to reconstruct the features of the missing cloud-containing areas. However, such methods rely on the existence of cloud-contaminated spatial areas or spectra, so they usually cannot reconstruct scenes covered by large-scale or opaque thick clouds. The multi-temporal data fusion method uses the information of the same ground scene collected at other times to restore the missing ground object information. The limitation of this method is that it assumes that there are almost no feature differences between the data obtained at different times. However, the temporal stability of ground object information cannot be guaranteed, which makes it difficult for the multi-temporal data fusion method to serve continuous monitoring methods or change detection methods in a short period. For the multi-modal data fusion method, cloud area and shadow removal requires the support of complementary other-modal remote sensing data. Different sensors have different imaging principles, and the multi-modal remote sensing images they capture also have obvious differences in the focus of describing the scene. Therefore, by fusing the complementary information in different-modal remote sensing images to reconstruct the cloud-contaminated area, the multi-modal data fusion method has good application potential. Among them, the fusion of SAR data and optical data has been a research hotspot in recent years. SAR is an all-weather sensor that can record radar backscattering intensity. SAR has strong penetration, and it can collect ground information without being affected by clouds or other atmospheric conditions, which determines that SAR image data can provide complementary content and structural information for the cloud-contaminated areas in optical remote sensing images. Based on the above advantages and characteristics, using SAR image data as a complementary information source has become the mainstream of multi-modal data fusion methods. Although cloud area and shadow removal methods based on SAR data have been continuously proposed and improved in recent years, most of the existing methods directly merge SAR and optical data in the channel dimension and perform fusion, which is not conducive to fully extracting the spatial structure information provided by SAR data. In addition, some cloud area and shadow removal methods based on SAR data focus on using pixel-level loss to repair the features in the spectral domain, resulting in blurred structures in the cloud area and shadow removal results. Therefore, multi-modal remote sensing image cloud area and shadow removal methods based on SAR data need further exploration and improvement. Summary of the Invention
[0004] The present invention mainly solves the problems that the existing technology does not fully extract the spatial structure information of SAR data and the cloud area and shadow removal results have blurred structures when jointly using SAR image data and optical image data for cloud area and shadow removal, and provides a multi-modal remote sensing image cloud area and shadow removal method based on a spectral-feature-structure-preserving fusion network, which can effectively generate cloud area and shadow removal results that simultaneously maintain high spectral fidelity and high structural clarity.
[0005] The method for removing cloud areas and shadows from multimodal remote sensing images based on a spectral-feature-structure preservation fusion network not only utilizes the powerful feature extraction and reasoning capabilities of deep learning methods but also combines a special multi-level network structure design. Data-driven deep learning methods can fully extract complementary features of multimodal remote sensing images and robustly handle various types of cloud coverages to reconstruct various ground object information; the multi-level structure can constrain the information in the spectral domain and the structural domain in the cloud removal results at different levels of the network, enhance the reliability of the results, can effectively extract the features of irregular ground objects based on deformable convolution, and capture the long-range dependencies between features based on the deformable multi-head self-attention mechanism to adaptively enhance the more representative spatial feature expressions.
[0006] The technical solution adopted by the present invention is: a method for removing cloud areas and shadows from multimodal remote sensing images, comprising the following steps:
[0007] Step 1, design a spectral-feature-structure preservation fusion network;
[0008] The spectral-feature-structure preservation fusion network includes a data fusion unit, a spectral reconstruction module, a structural reconstruction module, and a feature reconstruction module. The input multimodal remote sensing image data first undergoes feature-level fusion through the data fusion unit. The spectral reconstruction module includes multiple stacked residual groups, and the multi-level residual groups generate multi-level outputs in the spectral domain. In the structural reconstruction module, a first-order gradient operator is used to map the multi-level outputs in the spectral domain to the structural domain to generate multi-level structural domain outputs. Finally, in the feature reconstruction module, the spectral domain output generated by the last residual group and the label are jointly input into a pre-trained VGG network to obtain the cloud area and shadow removal results;
[0009] Step 2, design a loss function and train the spectral-feature-structure preservation fusion network;
[0010] Step 3, obtain the spectral domain output of the last residual group in the spectral-feature-structure preservation fusion network to obtain the cloud area and shadow removal results.
[0011] Furthermore, the input multi-modal remote sensing image data is first subjected to feature-level fusion by the data fusion unit and then input into the spectral reconstruction module. The data fusion unit consists of a concatenation layer Concat, a convolutional layer Conv, a rectified linear unit ReLU activation function, and a deformable multi-head self-attention mechanism DMSA. The input multi-modal remote sensing image data is first merged and mapped to the feature level in the data fusion unit, and then adaptively enhanced by DMSA. In DMSA, the input features are first passed through deformable convolutional layers to extract the query matrix Q, the key matrix K, and the value matrix V respectively. Subsequently, based on K, Q, and V, the long-range dependence relationship between features is captured to enhance the more representative spatial feature expression, as shown in (Equation 1):
[0012] DMSA = ReLU(Conv(Softmax(QK T ))V)) (Equation 1)
[0013] Among them, Softmax represents the classifier.
[0014] Furthermore, the feature after passing through the spectral reconstruction module is set as F N , and its expression is as follows:
[0015] F N = RG N (F N-1 ) = RG N (RG N-1 (…RG2(RG1(F0)))) (Equation 2)
[0016] Among them, N is the total number of residual groups, RG N is the Nth residual group, F0 is the feature output by the data fusion unit. The residual group contains a convolutional layer Conv, a residual block unit RBDMSA embedded with a deformable multi-head self-attention mechanism, and a short skip connection structure SSC. After passing through multiple stacked RBDMSAs in the residual group, the feature F is obtained through Conv. SSC is used to connect the input of the residual group and F to effectively transmit shallow features. After passing through Conv again, each level of the residual group will generate a corresponding level of output. In RBDMSA, it first passes through Conv, ReLU, and Conv in sequence, as shown in Equation 3, and then enters the deformable multi-head self-attention mechanism DMSA with the same structure as that in the data fusion unit;
[0017] F b+1 = F b + DMSA(Conv(ReLU(Conv(F b )))) (Equation 3)
[0018] Among them, F b is the input feature of the bth RBDMSA; Fb+1 It is the output feature of the b-th RBDMSA and also the input feature of the (b + 1)-th RBDMSA.
[0019] Furthermore, in the structure reconstruction module, a first-order gradient operator is used to map the multi-level output in the spectral domain to the structure domain to generate a multi-level structure domain output. The specific implementation method for obtaining the multi-level structure domain output is as follows;
[0020] I x (x, y) = I(x + 1, y) - I(x - 1, y) (Equation Four)
[0021] I y (x, y) = I(x, y + 1) - I(x, y - 1) (Equation Five)
[0022] I′(x, y) = (I x (x, y), I y (x, y)) (Equation Six)
[0023]
[0024] Among them, the gradient magnitudes in the x and y directions of the image I in the spectral domain at the position (x, y) are I x (x, y) and I y (x, y) respectively, and the gradient map obtained by mapping to the structure domain is The multi-level spectral and structure domain outputs are used to jointly constrain the reconstruction of spectral-level and structure-level features. Assuming that the original data of the remote sensing image is d, as shown in Equation Eight, the multi-level output generated for the given observation sample d is P(d), as shown in Equation Nine;
[0025] d = {d OPT , d SAR} (Equation Eight)
[0026]
[0027] Among them, d OPT is the optical image containing cloud cover in the modal remote sensing image sample and is the object for cloud area and shadow removal; d SAR is the SAR image in the multi-modal remote sensing image sample and is input as a complementary data source to achieve information reconstruction; p n (d) is the multi-level output in the spectral domain; is the multi-level output in the structure domain.
[0028] Furthermore, the specific implementation method of step 2 is as follows;
[0029] Using the SAR image d SAR and the optical image d OPTAs the input data of the spectral-feature-structure preservation fusion network, based on the three-level collaborative optimization loss function as shown in Equation X, the spectral-feature-structure preservation fusion network is optimized by the adaptive moment estimation through the backpropagation algorithm;
[0030]
[0031] Among them, the three-level collaborative optimization loss function consists of the spectral-level preservation loss the structure-level preservation loss and the perceptual loss λ T and λ P are two regularization constants respectively used to balance the three losses; λ n is a regularization constant used to balance the losses between multi-level outputs; t is an optical image without cloud cover in the same scene as the input data, serving as the label data of is the gradient map of t mapped to the structure domain, serving as the label data of N Assuming the original data of the remote sensing image is d, as shown in Equation XI, the label data for training the spectral-feature-structure preservation fusion network is T(d), and the spectral domain output p F l (p N ) and F l (t) respectively represent the feature representations of p N and t at the l-th layer in the VGG network, and L is the total number of network layers used for calculation.
[0032]
[0033] The present invention also provides a multi-modal remote sensing image cloud area and shadow removal system, including the following units:
[0034] The network design unit is used to design the spectral-feature-structure preservation fusion network;
[0035] The spectral-feature-structure-preserving fusion network includes a data fusion unit, a spectral reconstruction module, a structure reconstruction module, and a feature reconstruction module. The input multi-modal remote sensing image data first undergoes feature-level fusion through the data fusion unit. The spectral reconstruction module includes multiple stacked residual groups, and the multi-level residual groups generate multi-level outputs in the spectral domain. In the structure reconstruction module, a first-order gradient operator is used to map the multi-level outputs in the spectral domain to the structural domain to generate multi-level structural domain outputs. Finally, in the feature reconstruction module, the spectral domain output generated by the last residual group and the label are jointly input into a pre-trained VGG network to obtain the cloud area and shadow removal results;
[0036] A network training unit, which is used to design a loss function and train the spectral-feature-structure-preserving fusion network;
[0037] A shadow removal unit, which is used to obtain the spectral domain output of the last residual group in the spectral-feature-structure-preserving fusion network to obtain the cloud area and shadow removal results.
[0038] Furthermore, the input multi-modal remote sensing image data first undergoes feature-level fusion through the data fusion unit and then is input into the spectral reconstruction module. The data fusion unit consists of a concatenation layer Concat, a convolutional layer Conv, a rectified linear unit ReLU activation function, and a deformable multi-head self-attention mechanism DMSA. The input multi-modal remote sensing image data is first merged and mapped to the feature level in the data fusion unit, and then undergoes adaptive feature enhancement through DMSA. In DMSA, the input features first pass through a deformable convolutional layer to extract a query matrix Q, a key matrix K, and a value matrix V respectively. Subsequently, based on K, Q, and V, long-range dependencies between features are captured to enhance more representative spatial feature expressions, as shown in (Equation One):
[0039] DMSA = ReLU(Conv(Softmax(QK T )V)) (Equation One)
[0040] where Softmax represents a classifier.
[0041] Furthermore, let the feature after passing through the spectral reconstruction module be F N , and its expression is as follows:
[0042] F N = RG N (F N-1 ) = RG N (RG N-1 (…RG2(RG1(F0)))) (Equation Two)
[0043] where N is the total number of residual groups, and RG Nis the Nth residual group, F0 is the feature output by the data fusion unit. The residual group includes a convolutional layer Conv, a residual block unit RBDMSA with an embedded variable multi-head self-attention mechanism, and a short skip connection structure SSC. After passing through multiple stacked RBDMSAs in the residual group, the feature F is obtained through Conv. SSC is used to connect the input of the residual group and F to effectively transmit shallow features. After passing through Conv again, each level of the residual group generates a corresponding level of output. In RBDMSA, it first passes through Conv, ReLU, and Conv in sequence, as shown in Equation Three, and then enters the variable multi-head self-attention mechanism DMSA with the same structure as that in the data fusion unit;
[0044] F b+1 =F b +DMSA(Conv(ReLU(Conv(F b )))) (Equation Three)
[0045] where, F b is the input feature of the bth RBDMSA; F b+1 is the output feature of the bth RBDMSA and is also the input feature of the (b + 1)th RBDMSA.
[0046] Furthermore, in the structure reconstruction module, a first-order gradient operator is used to map the multi-level output in the spectral domain to the structural domain to generate multi-level structural domain outputs. The specific implementation method for obtaining the multi-level structural domain outputs is as follows;
[0047] I x (x,y) = I(x + 1,y) - I(x - 1,y) (Equation Four)
[0048] I y (x,y) = I(x,y + 1) - I(x,y - 1) (Equation Five)
[0049] I′(x,y) = (I x (x,y),I y (x,y)) (Equation Six)
[0050]
[0051] where, the gradient magnitudes in the x and y directions at the (x,y) position of the image I in the spectral domain are I x (x,y) and I y (x,y) respectively, and the gradient map obtained by mapping to the structural domain is The multi-level spectral and structural domain outputs are used to jointly constrain the reconstruction of spectral-level and structural-level features. Assuming the original data of the remote sensing image is d, as shown in Equation Eight, the multi-level output generated for the given observation sample d is P(d), as shown in Equation Nine;
[0052] d = {d OPT , d SAR}(Equation VIII)
[0053]
[0054] where d OPT is the optical image with cloud cover in the modal remote sensing image samples, which is the object for cloud area and shadow removal; d SAR is the SAR image in the multi-modal remote sensing image samples, which is input as a complementary data source to achieve information reconstruction; p n (d) is the multi-level output in the spectral domain; is the multi-level output in the structural domain.
[0055] Furthermore, the specific implementation method of the network training unit is as follows;
[0056] Using the SAR image d SAR and the optical image d with cloud cover OPT in the remote sensing image samples as the input data of the spectral-feature-structure preservation fusion network, based on the three-level collaborative optimization loss function, as shown in Equation X, the spectral-feature-structure preservation fusion network is optimized by the adaptive moment estimation through the backpropagation algorithm;
[0057]
[0058] where the three-level collaborative optimization loss function consists of the spectral-level preservation loss the structural-level preservation loss and the perceptual loss λ T , λ P are two regularization constants respectively used to balance the three losses; λ n is a regularization constant used to balance the losses between multi-level outputs; t is the optical image without cloud cover in the same scene as the input data, as the label data of; is the gradient map of t mapped to the structural domain, as the label data of; Assuming that the original remote sensing image data is d, as shown in Equation XI, the training spectral-feature-structure preservation fusion network label data is T(d), and the spectral domain output p N produced by the last residual group of the spectral-feature-structure preservation fusion network is jointly calculated with the spectral domain label t F l (p N ) and F l (t) respectively represent the feature representations of p N , t at the l-th layer in the VGG network, and L is used for Total number of calculated network layers
[0059]
[0060] Compared with the existing methods for removing cloud areas and shadows from multimodal remote sensing images, the present invention has the following advantages and positive effects:
[0061] (1) By constructing a spectral-feature-structure preservation fusion network embedded with a variable multi-head self-attention mechanism, the present invention successfully solves the problem of insufficient extraction of spatial structure information of SAR data, establishes long-range dependencies between features through the variable multi-head self-attention mechanism, and realizes adaptive enhancement of effective spatial features.
[0062] (2) The method of the present invention is based on a three-level collaborative optimization loss function, constrains the reconstruction of spectral domain and structural domain information at multiple levels of the network, and uses perceptual loss to constrain the reconstruction of information in the semantic feature domain, successfully overcoming the problem that the methods for removing cloud areas and shadows based on the fusion of SAR and optical data generally only focus on the reconstruction of spectral domain information, resulting in blurred object edges and textures in the cloud area and shadow removal results, and effectively improving the quality of cloud area and shadow removal results. Description of the Drawings
[0063] Figure 1 : is the overall flowchart of the embodiment of the present invention.
[0064] Figure 2 : is the structure diagram of the spectral-feature-structure preservation fusion network of the embodiment of the present invention.
[0065] Figure 3 : is the schematic diagram of the cloud area and shadow removal result of the embodiment of the present invention. Detailed Embodiment
[0066] To facilitate the understanding and implementation of the present invention by those of ordinary skill in the art, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0067] Please refer to Figure 1 , the method for removing cloud areas and shadows from multimodal remote sensing images based on a spectral-feature-structure preservation fusion network provided by the present invention includes the following steps:
[0068] Step 1: Design the spectral-feature-structure preservation fusion network structure according to the task specificity. Please refer to Figure 2, the Spectral-Feature-Structure Preservation Fusion Network is designed as a special deep multi-level network structure, which consists of four parts, namely the data fusion unit, the spectral reconstruction module, the structure reconstruction module, and the feature reconstruction module. The input multi-modal remote sensing image data first undergoes feature-level fusion in the data fusion unit. The data fusion unit is composed of a concatenation layer (Concat), a convolutional layer (Conv), a rectified linear unit (ReLU) activation function, and a deformable multi-head self-attention mechanism (DMSA). The input multi-modal remote sensing image data is first merged and mapped to the feature level in the data fusion unit, and then undergoes adaptive feature enhancement through DMSA. In DMSA, the input features first pass through a deformable convolutional layer (Deformable Conv) to extract the query matrix (Q), the key matrix (K), and the value matrix (V) respectively. Subsequently, based on K, Q, and V, the long-range dependence relationship between features is captured to enhance the more representative spatial feature expression (Equation One).
[0069] DMSA = ReLU(Conv(Softmax(QK T ))V)) (Equation One)
[0070] The feature F0 output by the data fusion unit then passes through the spectral reconstruction module, which is composed of stacked residual groups. As shown in Equation Two, the feature after passing through the spectral reconstruction module is F N .
[0071] F N = RG N (F N-1 ) = RG N (RG N-1 (…RG2(RG1(F0)))) (Equation Two)
[0072] where N is the total number of residual groups, RG Nis the Nth residual group. The residual group includes a Conv, a residual block unit with a deformable multi-head self-attention mechanism (RBDMSA), and a short skip connection structure (SSC). After passing through multiple stacked RBDMSAs in the residual group, the feature F is obtained through Conv. The SSC is used to connect the input of the residual group and F to effectively transmit shallow features. After passing through Conv again, each level of the residual group generates a corresponding level of output. In the RBDMSA, it first passes through Conv, ReLU, and Conv in sequence, as shown in Equation 3, and then enters the channel attention unit DMSA whose structure is consistent with that in the data fusion unit.
[0073] F b+1 = F b + DMSA(Conv(ReLU(Conv(F b )))) (Equation 3)
[0074] where F b is the input feature of the bth RBDMSA; F b+1 is the output feature of the bth RBDMSA and is also the input feature of the (b + 1)th RBDMSA. Through such an architecture, the RBDMSA can adaptively enhance effective spatial features and improve the representational ability of the network.
[0075] The multi-level residual groups in the spectral reconstruction module generate multi-level outputs in the spectral domain. To enable the deep learning network to focus on the reconstruction of geometric structure information, a first-order gradient operator is used in the structure reconstruction module to map the multi-level outputs in the spectral domain to the structural domain to generate multi-level structural domain outputs, as shown in Equation 4, Equation 5, Equation 6, and Equation 7.
[0076] I x (x, y) = I(x + 1, y) - I(x - 1, y) (Equation 4)
[0077] I y (x, y) = I(x, y + 1) - I(x, y - 1) (Equation 5)
[0078] I′(x, y) = (I x (x, y), I y (x, y)) (Equation 6)
[0079]
[0080] where the gradient magnitudes of the image I in the spectral domain at the (x, y) position in the x and y directions are I x (x, y) and Iy (x, y), the gradient map obtained by mapping to the domain is The multi-level spectrum and domain outputs are used to jointly constrain the reconstruction of spectral-level and structural-level features. Assuming the original data of the remote sensing image is d, as shown in Equation VIII, the multi-level output generated for the given observation sample d is P(d), as shown in Equation IX.
[0081] d = {d OPT , d SAR} (Equation VIII)
[0082]
[0083] Among them, d OPT is the optical image with cloud cover in the modal remote sensing image sample, which is the object to be removed for cloud areas and shadows; d SAR is the SAR image in the multi-modal remote sensing image sample, which is input as a complementary data source to achieve information reconstruction; p n (d) is the multi-level output in the spectral domain. The spectral domain output of each level is obtained by reducing the number of feature channels back to the same as that of the input d n through a convolutional layer; OPT The number of channels is the same; is the multi-level output in the domain. Each level of is obtained by mapping the corresponding p n (d) to the domain through a gradient operator (Equations IV, V, VI, and VII). Finally, in the feature reconstruction module, the spectral domain output generated by the last residual group and the label are jointly input into the pre-trained VGG-16 (or VGG-19) to make the cloud area and shadow removal results predicted by the network maintain a high similarity with the label in semantic features.
[0084] Step 2: Design a three-level collaborative optimization loss function to train the deep fusion network.
[0085] Taking the publicly available sample set as an example, the sample set includes triple Sentinel-1 SAR image data, Sentinel-2 optical image data with cloud cover, and Sentinel-2 optical image data without cloud cover. Taking the SAR image d SAR in the triple remote sensing image sample and the optical image d OPT with cloud cover in the multi-modal remote sensing image sample as the input, the multi-level output in the spectral domain and the multi-level output in the domain are obtained through the above steps.As the input data of the spectral-feature-structure preservation fusion network, based on the three-level collaborative optimization loss function (Equation 10), the spectral-feature-structure preservation fusion network is optimized by the backward propagation algorithm Nesterov-Adaptive Moment Estimation (Nadam). The three-level collaborative optimization loss function adds constraints on the structural reconstruction information and semantic feature alignment on the basis of the spectral-level constraint, which can guide the deep learning network to produce cloud area and shadow removal results with good appearance, clear ground object texture and boundaries.
[0086]
[0087] Among them, the three-level collaborative optimization loss function consists of the spectral-level preservation loss the structure-level preservation loss and the perceptual loss λ T and λ P are two regularization constants respectively used to balance the three losses; λ n is a regularization constant used to balance the losses between multi-level outputs; t is an optical image without cloud cover in the same scene as the input data, used as the label data of is the gradient map of t mapped to the structure domain, used as the label data of N Assume that the original data of the remote sensing image is d. As shown in Equation 11, the label data for training the spectral-feature-structure preservation fusion network is T(d). The spectral domain output p F l (p N ) and F l (t) respectively represent the feature representations of p N and t in the l-th layer of VGG-16, and L is the total number of network layers used for calculation.
[0088]
[0089] Step 3: Obtain the cloud area and shadow removal results. Use the trained spectral-feature-structure preservation fusion network to remove the cloud area and shadow from the optical remote sensing image, and obtain the spectral domain output of the last residual group in the spectral-feature-structure preservation fusion network to get the reconstructed cloud-free optical remote sensing image result as Figure 3 shown. Assume that the original data of the remote sensing image is d, and the cloud area and shadow removal result obtained for the given observation sample d is p N (d).
[0090] The present invention also provides a multi-modal remote sensing image cloud area and shadow removal system, including the following units:
[0091] A network design unit, configured to design a spectral-feature-structure-preserving fusion network;
[0092] The spectral-feature-structure-preserving fusion network includes a data fusion unit, a spectral reconstruction module, a structure reconstruction module, and a feature reconstruction module. The input multi-modal remote sensing image data first undergoes feature-level fusion through the data fusion unit. The spectral reconstruction module includes multiple stacked residual groups, and the multi-level residual groups generate multi-level outputs in the spectral domain. In the structure reconstruction module, a first-order gradient operator is used to map the multi-level outputs in the spectral domain to the structure domain to generate multi-level structure domain outputs. Finally, in the feature reconstruction module, the spectral domain output generated by the last residual group and the label are jointly input into a pre-trained VGG network to obtain the cloud area and shadow removal results;
[0093] A network training unit, configured to design a loss function and train the spectral-feature-structure-preserving fusion network;
[0094] A shadow removal unit, configured to obtain the spectral domain output of the last residual group in the spectral-feature-structure-preserving fusion network to obtain the cloud area and shadow removal results.
[0095] The specific implementation manners of each unit are the same as each step, and the present invention will not describe them.
[0096] The method of the present invention is based on the fusion of SAR and optical remote sensing image data. The RBDMSA module in the spectral-feature-structure-preserving fusion network is used to adaptively enhance effective spatial features. Different from the method of only merging SAR and optical data for fusion without feature enhancement processing, the method of the present invention is more conducive to extracting the spatial structure information provided by SAR image data. Training the spectral-feature-structure-preserving fusion network based on the three-level collaborative optimization loss can guide the network to simultaneously focus on the reconstruction of information in the spectral domain, the structure domain, and the semantic feature domain, and guide the network to generate cloud area and shadow removal results that not only have high spectral fidelity but also have clear ground object textures and contours. The deep multi-level structure in the spectral-feature-structure-preserving fusion network can gradually constrain the quality of the reconstruction results, ensuring the stability and robustness of the deep learning network model. Therefore, the method of the present invention is more suitable for large-scale scene applications including various cloud covers and ground object types. The present invention solves the problems commonly existing in the current research on multi-modal remote sensing image cloud area and shadow removal technology, such as insufficient extraction of the spatial structure information of SAR data and the fuzzy structure of cloud area and shadow removal results, and effectively improves the quality of the reconstructed cloud-free optical remote sensing images.
[0097] It should be understood that the parts not elaborated in this specification all belong to the prior art.
[0098] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art to which the present invention pertains may make various modifications or supplements to the described specific embodiments or use similar means for substitution, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A method for removing cloud areas and shadows from multimodal remote sensing images, characterized in that, It includes the following steps: Step 1, design a spectral-feature-structure-preserving fusion network; The spectral-feature-structure-preserving fusion network includes a data fusion unit, a spectral reconstruction module, a structure reconstruction module, and a feature reconstruction module. The input multi-modal remote sensing image data first undergoes feature-level fusion through the data fusion unit. The spectral reconstruction module includes multiple stacked residual groups, and the multi-level residual groups generate multi-level outputs in the spectral domain. In the structure reconstruction module, a first-order gradient operator is used to map the multi-level outputs in the spectral domain to the structural domain to generate multi-level structural domain outputs. Finally, in the feature reconstruction module, the spectral domain output generated by the last residual group and the label are jointly input into a pre-trained VGG network to obtain the cloud area and shadow removal results; Step 2, design a loss function and train the spectral-feature-structure-preserving fusion network; The specific implementation method of Step 2 is as follows; Taking the SAR image d in the remote sensing image samples SAR and the optical image d including cloud cover OPT as the input data of the spectral-feature-structure-preserving fusion network, based on the three-level collaborative optimization loss function, as shown in Equation 10, the spectral-feature-structure-preserving fusion network is optimized by the adaptive moment estimation of the backpropagation algorithm; Among them, the three-level collaborative optimization loss function consists of a spectral-level preservation loss a structure-level preservation loss and a perceptual loss . λ T , λ P are two regularization constants respectively used to balance the three losses; λ n is a regularization constant used to balance the losses between multi-level outputs; t is an optical image without cloud cover in the same scene as the input data, serving as the label data; is the gradient map of t mapped to the structure domain, serving as the label data; Assuming the original data of the remote sensing image is d, as shown in Equation 11, the label data for training the spectral-feature-structure preservation fusion network is T(d), and the spectral domain output p N produced by the last residual group of the spectral-feature-structure preservation fusion network is jointly calculated with the spectral domain label t F l (p N ) and F l (t) respectively represent the feature representations of p N , t at the l-th layer in the VGG network, and L is the total number of network layers used for calculation. Step 3, obtain the spectral domain output of the last residual group in the spectral-feature-structure-preserving fusion network to obtain the cloud area and shadow removal results.
2. A method for removing cloud areas and shadows from multimodal remote sensing images according to claim 1, characterized in that: The input multi-modal remote sensing image data first undergoes feature-level fusion through the data fusion unit and then is input into the spectral reconstruction module. The data fusion unit consists of a concatenation layer Concat, a convolutional layer Conv, a rectified linear unit ReLU activation function, and a deformable multi-head self-attention mechanism DMSA. The input multi-modal remote sensing image data is first merged and mapped to the feature level in the data fusion unit, and then undergoes adaptive feature enhancement through DMSA. In DMSA, the input features first pass through a deformable convolutional layer to extract the query matrix Q, the key matrix K, and the value matrix V respectively. Subsequently, based on K, Q, and V, the long-range dependence relationship between features is captured to enhance the more representative spatial feature expression, as shown in (Equation One): DMSA = ReLU(Conv(Softmax(QK T )(V)) (Equation 1) where Softmax represents the classifier.
3. A method for removing cloud areas and shadows from multimodal remote sensing images according to claim 1, characterized in that: Set the feature after passing through the spectral reconstruction module as F N , and its expression is as follows: F N = RG N (F N-1 ) = RG N (RG N-1 (…RG2(RG1(F0)))) (Equation Two) Where N is the total number of residual groups, RG N is the Nth residual group, F0 is the feature output by the data fusion unit. The residual group includes a convolutional layer Conv, a residual block unit RBDMSA with an embedded variable multi-head self-attention mechanism, and a short skip connection structure SSC. After passing through multiple stacked RBDMSAs in the residual group, the feature F is obtained through Conv. SSC is used to connect the input of the residual group and F to effectively transmit shallow features. After passing through Conv again, each level of the residual group generates a corresponding level of output; in RBDMSA, it first passes through Conv, ReLU, Conv in sequence, as shown in Equation III, and then enters the variable multi-head self-attention mechanism DMSA with the same structure as that in the data fusion unit; F b+1 = F b + DMSA(Conv(ReLU(Conv(F b )))) (Equation III) Among them, F b is the input feature of the b-th RBDMSA; F b+1 is the output feature of the b-th RBDMSA and also the input feature of the (b + 1)-th RBDMSA.
4. A method for removing cloud areas and shadows from multimodal remote sensing images according to claim 1, characterized in that: In the structure reconstruction module, a first-order gradient operator is used to map the multi-level outputs in the spectral domain to the structural domain to generate multi-level structural domain outputs. The specific implementation method for obtaining the multi-level structural domain outputs is as follows; I x (x,y) = I(x + 1,y) - I(x - 1,y) (Equation 4) I y (x,y) = I(x,y + 1) - I(x,y - 1) (Equation 5) I′(x,y) = (I x (x,y), I y (x,y)) (Equation 6) Among them, the gradient magnitudes in the x and y directions of the image I in the spectral domain at the (x, y) position are I x (x, y) and I y (x, y), and the gradient map obtained by mapping to the structural domain is The multi-level spectral and structural domain outputs are used to jointly constrain the reconstruction of spectral-level and structural-level features. Assuming that the original data of the remote sensing image is d, as shown in Equation VIII, the multi-level output generated for the given observed sample d is P(d), as shown in Equation IX; d = {d OPT , d SAR}(Formula VIII) Among them, d OPT is the optical image with cloud cover in the modal remote sensing image sample, which is the object for cloud area and shadow removal; d SAR is the SAR image in the multi-modal remote sensing image sample, which is input as a complementary data source to achieve information reconstruction; p n (d) is the multi-level output in the spectral domain; is the multi-level output in the structural domain.
5. A multi-modal remote sensing image cloud area and shadow removal system, characterized in that, It includes the following units: A network design unit for designing a spectral-feature-structure-preserving fusion network; The spectral-feature-structure-preserving fusion network includes a data fusion unit, a spectral reconstruction module, a structure reconstruction module, and a feature reconstruction module. The input multi-modal remote sensing image data first undergoes feature-level fusion through the data fusion unit. The spectral reconstruction module includes multiple stacked residual groups, and the multi-level residual groups generate multi-level outputs in the spectral domain. In the structure reconstruction module, a first-order gradient operator is used to map the multi-level outputs in the spectral domain to the structural domain to generate multi-level structural domain outputs. Finally, in the feature reconstruction module, the spectral domain output generated by the last residual group and the label are jointly input into a pre-trained VGG network to obtain the cloud area and shadow removal results; A network training unit for designing a loss function and training the spectral-feature-structure-preserving fusion network; The specific implementation method of the network training unit is as follows; Taking the SAR image d in the remote sensing image sample SAR and the optical image d including cloud cover OPT as the input data of the spectral-feature-structure-preserving fusion network, based on the three-level collaborative optimization loss function, as shown in Equation Ten, the spectral-feature-structure-preserving fusion network is optimized by the adaptive moment estimation of the backpropagation algorithm; Among them, the three-level collaborative optimization loss function consists of a spectral-level preservation loss a structure-level preservation loss and a perceptual loss λ T and λ P are two regularization constants respectively used to balance the three losses; λ n is a regularization constant used to balance the losses between multi-level outputs; t is an optical image without cloud cover in the same scene as the input data, serving as the label data; is the gradient map of t mapped to the structure domain, serving as the label data; Assuming the original data of the remote sensing image is d, as shown in Equation Eleven, the label data for training the spectral-feature-structure preservation fusion network is T(d), and the spectral domain output p N produced by the last residual group of the spectral-feature-structure preservation fusion network is jointly calculated with the spectral domain label t F l (p N ) and F l (t) respectively represent the feature representations of p N and t in the l-th layer of the VGG network, and L is the total number of network layers used for calculation, A shadow removal unit for obtaining the spectral domain output of the last residual group in the spectral-feature-structure-preserving fusion network to obtain the cloud area and shadow removal results.
6. A multi-modal remote sensing image cloud area and shadow removal system according to claim 5, characterized in that: The input multi-modal remote sensing image data is first fused at the feature level by the data fusion unit and then input into the spectral reconstruction module. The data fusion unit consists of a concatenation layer Concat, a convolutional layer Conv, a rectified linear unit ReLU activation function, and a deformable multi-head self-attention mechanism DMSA. The input multi-modal remote sensing image data is first merged and mapped to the feature level in the data fusion unit, and then adaptively enhanced by DMSA. In DMSA, the input features are first passed through a deformable convolutional layer to extract the query matrix Q, the key matrix K, and the value matrix V respectively. Subsequently, based on K, Q, and V, the long-range dependence relationship between features is captured to further enhance the more representative spatial feature expression, as shown in (Equation 1): DMSA = ReLU(Conv(Softmax(QK T )(V)) (Equation 1) Among them, Softmax represents the classifier.
7. A multi-modal remote sensing image cloud area and shadow removal system according to claim 5, characterized in that: Set the feature after passing through the spectral reconstruction module as F N , and its expression is as follows: F N = RG N (F N-1 ) = RG N (RG N-1 (…RG2(RG1(F0)))) (Equation II) Where N is the total number of residual groups, RG N is the Nth residual group, F0 is the feature output by the data fusion unit. The residual group includes a convolutional layer Conv, a residual block unit RBDMSA with an embedded variable multi-head self-attention mechanism, and a short skip connection structure SSC. After passing through multiple stacked RBDMSAs in the residual group, the feature F is obtained through Conv. The SSC is used to connect the input of the residual group and F to effectively transmit shallow features. After passing through Conv again, each level of the residual group generates a corresponding level of output. In the RBDMSA, it first passes through Conv, ReLU, and Conv in sequence, as shown in Equation III, and then enters the variable multi-head self-attention mechanism DMSA with the same structure as that in the data fusion unit; F b+1 = F b + DMSA(Conv(ReLU(Conv(F b )))) (Equation III) Among them, F b is the input feature of the b-th RBDMSA; F b+1 is the output feature of the b-th RBDMSA and also the input feature of the (b + 1)-th RBDMSA.
8. A multi-modal remote sensing image cloud area and shadow removal system according to claim 5, characterized in that: In the structure reconstruction module, a first-order gradient operator is used to map the multi-level output in the spectral domain to the structure domain to generate a multi-level structure domain output. The specific implementation method for obtaining the multi-level structure domain output is as follows; I x (x,y) = I(x + 1,y) - I(x - 1,y) (Equation Four) I y (x,y) = I(x,y + 1) - I(x,y - 1) (Equation 5) I′(x,y) = (I x (x,y), I y (x,y)) (Equation 6) Among them, the gradient magnitudes of the spectral-domain image I in the x and y directions at the (x, y) position are I x (x, y) and I y (x, y), and the gradient map obtained by mapping to the structural domain is The multi-level spectral and structural domain outputs are used to collaboratively constrain the reconstruction of spectral-level and structural-level features. Assuming that the original data of the remote sensing image is d, as shown in Equation VIII, the multi-level output generated for the given observed sample d is P(d), as shown in Equation IX; d = {d OPT , d SAR}(Formula VIII) Among them, d OPT is the optical image with cloud cover in the modal remote sensing image sample, which is the object for cloud area and shadow removal; d SAR is the SAR image in the multi-modal remote sensing image sample, which is input as a complementary data source to achieve information reconstruction; p n (d) is the multi-level output in the spectral domain; is the multi-level output in the structural domain.