Space-spectrum reconstruction method and system for cloud-containing optical remote sensing image fused with spaceborne SAR (Synthetic Aperture Radar)
Through a cloud-containing optical remote sensing image space spectrum reconstruction method that integrates satellite-borne SAR, multi-scale spatial feature extraction and space spectrum reconstruction technology are used to solve the problem of inaccurate details in cloud reconstruction, and achieve high-precision recovery of local objects in the cloud-covered area.
Patent Information
- Application Number
- CN202510140510.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-03
AI Technical Summary
The existing technology has problems such as rough texture details, obvious artifacts and local blur in the reconstruction of cloud areas, resulting in inaccurate recovery of local objects in the cloud-covered area, hindering the promotion of technology in high-precision application scenarios.
A cloud-containing optical remote sensing image null spectrum reconstruction method is adopted to integrate satellite-borne SAR. Through the synergy between shallow feature aggregation module, residual adaptive dynamic enhancement modulator and null spectrum reconstruction result output module, multi-scale spatial feature extraction and null spectrum reconstruction of optical images and SAR images are realized.
It significantly improves the reconstruction accuracy and visual quality of local objects in the cloud-covered area, fully integrates the complementary characteristics of optical images and SAR images, and provides better technical support for cloud-covered image reconstruction in complex scenarios.
Smart Images

Figure CN120088149A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and particularly to a method and system for spatial-spectral reconstruction of cloud-containing optical remote sensing images integrating spaceborne SAR. Background Art
[0002] Optical remote sensing images play an irreplaceable role in surface monitoring and earth observation due to their high spatial resolution, strong temporal continuity, and rich surface information coverage ability. Such images can quickly, macroscopically, and accurately reflect surface changes and have been widely used in fields such as environmental monitoring, disaster assessment, and agricultural management. However, limited by their passive imaging mechanism, optical remote sensing images are easily affected by cloud occlusion, resulting in the loss of some ground information. This data loss problem severely restricts the subsequent analysis and interpretation of images.
[0003] In contrast, Synthetic Aperture Radar (SAR) is an active remote sensing technology based on the microwave band. It detects surface targets by actively emitting and receiving signals and thus has the ability to work all day and all weather, effectively avoiding the influence of cloud occlusion. Due to this characteristic, using SAR images to assist in the cloud area reconstruction of optical images has become a research hotspot, providing a new technical path for solving the problem of missing information in optical images.
[0004] In terms of the technical implementation of cloud area reconstruction, according to the processing and utilization methods of SAR images, previous research methods are mainly divided into feature fusion-based methods and data generation-based methods. The former extracts key features from SAR images and integrates them into the missing areas of optical images to make up for information gaps. Such methods include, but are not limited to, wavelet transform fusion, multi-scale decomposition techniques, and multi-modal feature fusion based on Deep Convolutional Neural Networks (DCNN). The latter uses generative models such as Generative Adversarial Networks (GAN) to convert SAR images into pseudo-optical images to replace the areas of optical images occluded by clouds. The core idea of the GAN-based method is to train the Generator and the Discriminator to play against each other, enabling the Generator to learn how to generate pseudo-optical images similar to real optical images, thus generating high-quality reconstructed images in the cloud-occluded areas.
[0005] Both the feature fusion-based and data generation-based methods have improved the accuracy of spatial-spectral reconstruction of cloud-covered optical remote sensing imagery missing areas to a certain extent. However, due to the inherent differences between optical and SAR images in imaging geometry, ground object radiation characteristics, and texture expression, the current reconstruction results still have significant limitations, such as coarse texture details, obvious artifacts, and local blurring. These deficiencies pose a challenge to the accurate restoration of ground object details in cloud-covered areas and hinder the further promotion of the technology in high-precision application scenarios. At present, through deeper multimodal feature fusion strategies, more efficient generative adversarial model optimization methods, and adaptive processing mechanisms for optical-SAR differences, it is expected to further improve the accuracy and practicality of cloud area reconstruction, and provide more reliable technical support for seamless spatiotemporal monitoring of the surface. Summary of the invention
[0006] The present invention provides a method and system for spatial-spectral reconstruction of cloud-containing optical remote sensing images integrated with space-borne SAR, which are used to solve the defects in the prior art. The method aims to meet the urgent demand for spatiotemporal cloud-free remote sensing images for large-scale and long-time series ground environment monitoring, and combines the characteristics of SAR images that are not affected by bad weather such as clouds and fog, thereby providing valuable gain information for the restoration of cloud-containing optical images.
[0007] In a first aspect, the present invention provides a method for spatial-spectral reconstruction of cloud-containing optical remote sensing images fused with space-borne SAR, comprising: Acquire optical and SAR images; Inputting the optical image and the SAR image into a shallow feature aggregation module, and outputting aggregated low-level features of the optical image and the SAR image; Processing the aggregated low-level features through a residual adaptive dynamic enhancement modulator to obtain a multi-scale spatial feature map; The multi-scale spatial feature map is input into the spatial spectrum reconstruction result output module to perform spatial spectrum reconstruction to obtain a restored optical image with a preset clarity.
[0008] According to a method for reconstructing cloud-containing optical remote sensing images by integrating space-borne SAR provided by the present invention, the optical image and the SAR image are input into a shallow feature aggregation module, and the aggregated low-level features of the optical image and the SAR image are output, including: Determining that the shallow feature aggregation module includes an input tensor merging layer, a 3×3 convolutional layer with a stride of 1, and an attention layer; The input tensor merging layer merges the optical image and the SAR image to obtain a merged tensor :
[0009] In the formula, Indicates concatenating two tensors along the channel dimension, for optical images and SAR images The sizes of which are and respectively. The size of the merged tensor is ; Input the merged tensor into the 3×3 convolutional layer with a stride of 1 to obtain the convolutional output feature map :
[0010] In the formula, is the convolutional kernel, is the bias term, and the size of the convolutional output feature map is ; Input the convolutional output feature map into the attention layer to obtain the weighted output feature map :
[0011] In the formula, , , are the query matrix, key matrix, and value matrix respectively, obtained by performing a linear transformation on . is the dimension of the key vector, used to scale the dot product, is the normalization operation, used to ensure that the sum of the weights is 1, is the weighted output feature map, containing the weight adjustment at each spatial position; The output feature map after being weighted by the attention mechanism is , is the final attention-weighted output of the shallow feature aggregation module, containing the aggregated low-level features extracted from the optical image and the SAR image.
[0012] According to a method for spatial-spectral reconstruction of cloud-containing optical remote sensing images integrating spaceborne SAR provided by the present invention, process the aggregated low-level features through a residual adaptive dynamic enhancement modulator to obtain multi-scale spatial feature maps, including: Determine that the residual adaptive dynamic enhancement modulator includes multiple information gain blocks, and each information gain block includes a multi-scale spatial fine-grained coupling module, a sparse global statistical cooperation module, and a cross-domain dynamic interaction reconstruction module; Input the aggregated low-level features into the multi-scale spatial fine-grained coupling module to generate a cross-scale spatial attention map; Input the cross-scale spatial attention map into the sparse global statistical collaboration module to obtain a sparsified channel attention map; Input the sparsified channel attention map into the cross-domain dynamic interaction reconstruction module to obtain a multi-scale spatial feature map.
[0013] According to a spatial-spectral reconstruction method for cloud-containing optical remote sensing images integrated with spaceborne SAR provided by the present invention, input the aggregated low-level features into the multi-scale spatial fine-grained coupling module to generate a cross-scale spatial attention map, including: The multi-scale spatial fine-grained coupling module includes layer normalization, a downsampling module, neighborhood padding convolution, multi-head self-attention, and 1×1 convolution; Perform layer normalization to obtain layer-normalized features :
[0014] wherein, is the layer normalization operation; Perform downsampling on through an average pooling operation with a stride of to obtain a downsampled feature map :
[0015] wherein, is the average pooling operation; Perform linear projection on to adjust the channel dimension, and use a shortcut connection to retain the original features to obtain a feature map :
[0016] wherein, is the linear operation; Perform a 3×3 neighborhood padding convolution operation on the feature map to generate a query vector , a key vector , and a value vector :
[0017]
[0018]
[0019] wherein, , , is the local window size, is the 3×3 convolution operation; By introducing a cross-scale multiplication operation and adopting a multi-scale feature offset aggregation strategy, each local block of the query vector interacts with the relevant information of the key vector to generate a cross-scale spatial attention map :
[0020] Among them, is the relative position encoding, which is used to enhance the perception of spatial information, is the scaling factor, which is used to normalize the attention scores, is used to normalize the attention distribution. Each local block query vector can interact with the region corresponding to the corresponding region in the downsampled feature map. This operation alleviates the blur by using the clear block after downsampling as prior information.
[0021] According to a method for spectral-spatial reconstruction of cloud-containing optical remote sensing images integrating spaceborne SAR provided by the present invention, the cross-scale spatial attention map is input into the sparse global statistical collaboration module to obtain a sparsified channel attention map, including: Perform 1×1 convolution and 3×3 depth convolution operations on the cross-scale spatial attention map to generate channel context encoding and extract the potential feature information of the image; Based on the self-attention mechanism of the channel, calculate the similarity between the query and the key to calculate the attention weights:
[0022] Among them, and are the query and key respectively generated from the input features, is the scaling factor of the feature dimension, is the generated attention matrix; Adopt the Top- k selection operation to select the top in the attention matrix most relevant elements:
[0023] Among them, is the Top- k selection operator, is the th maximum attention score threshold between the query and key pairs; The softmax function is used to normalize the sparsified attention matrix to obtain a sparsity-optimized channel attention map:
[0024] where, is the Top- k selection operation, 、 、 are the query vector, key vector, and value vector respectively, and the final sparsified attention map ; The output of the multi-head self-attention is processed through a summation operation to obtain a sparsified channel attention map .
[0025] According to a method for spatio-spectral reconstruction of cloud-containing optical remote sensing images integrating spaceborne SAR provided by the present invention, the sparsified channel attention map is input into the cross-domain dynamic interaction reconstruction module to obtain multi-scale spatial feature maps, including: Use 1×1 convolution and 3×3 depth convolution operations to perform information aggregation on the input spatial feature and channel feature ; Based on the gating mechanism, use the GELU activation function to non-linearly activate the channel feature , and perform element-wise multiplication with the spatial feature to obtain an enhanced feature; Adjust the enhanced feature through linear projection and add it to the spatial feature to obtain a multi-scale spatial feature map .
[0026] According to a method for spatio-spectral reconstruction of cloud-containing optical remote sensing images integrating spaceborne SAR provided by the present invention, the multi-scale spatial feature map is input into the spatio-spectral reconstruction result output module for spatio-spectral reconstruction to obtain a restored optical image with a preset clarity, including: The spatio-spectral reconstruction result output module includes a 3×3 convolutional layer with a stride of 1, a polarization self-attention layer, and a 3×3 convolutional layer with a stride of 1 connected in sequence; Use the L1 loss and the structural similarity SSIM loss for spatio-spectral reconstruction, and output a restored optical image with a preset clarity.
[0027] According to a method for spatio-spectral reconstruction of cloud-containing optical remote sensing images integrating spaceborne SAR provided by the present invention, using the L1 loss and the structural similarity SSIM loss for spatio-spectral reconstruction, including: The total loss function combining the L1 loss and the SSIM loss is:
[0028] Among them, represents the output image, represents the corresponding target image, is a hyperparameter used to adjust the relative importance of the two loss functions; The L1 loss function is:
[0029] Among them, represents the total number of pixels in the image, represents the index of the image patch, represents the L1 norm, that is, the sum of absolute values; The SSIM loss function is: .
[0030] In a second aspect, the present invention also provides a spatial-spectral reconstruction system for cloud-containing optical remote sensing images integrating spaceborne SAR, including: An acquisition unit for acquiring optical images and SAR images; An aggregation unit for inputting the optical image and the SAR image into a shallow feature aggregation module and outputting aggregated low-level features of the optical image and the SAR image; A modulation unit for processing the aggregated low-level features through a residual adaptive dynamic enhancement modulator to obtain a multi-scale spatial feature map; An output unit for inputting the multi-scale spatial feature map into a spatial-spectral reconstruction result output module for spatial-spectral reconstruction to obtain a restored optical image with a preset clarity.
[0031] In a third aspect, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the spatial-spectral reconstruction method for cloud-containing optical remote sensing images integrating spaceborne SAR as described in any one of the above.
[0032] Compared with the prior art, the beneficial effects of the present invention are: Through the synergistic effect of the multi-scale spatial fine-grained coupling module and the sparse global statistical cooperation module, the present invention realizes the efficient capture of global semantic information and local spatial details. Among them, the multi-scale spatial fine-grained coupling module expands the receptive field range and enhances the interaction ability between cross-scale features through the cross-scale window attention mechanism, while retaining the tiny detail information in the SAR image. The sparse global statistical cooperation module, through Top- kThe channel selection strategy adaptively filters out the most relevant channel features, suppresses the interference of redundant information, and improves the quality of cross-domain feature fusion. The combination of the two can not only fully integrate the complementary characteristics of optical images and SAR images, but also significantly enhance the accuracy and robustness of feature expression, providing better technical support for the reconstruction of cloud-covered images in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0034] Figure 1 is a schematic flow chart of the method for spatio-spectral reconstruction of cloud-containing optical remote sensing images integrating spaceborne SAR provided by the present invention; Figure 2 is an overall framework diagram of the method for spatio-spectral reconstruction of cloud-containing optical remote sensing images integrating spaceborne SAR provided by the present invention; Figure 3 is a schematic diagram of the multi-scale spatial fine-grained coupling module provided by the present invention; Figure 4 is a schematic diagram of the sparse global statistical cooperation module provided by the present invention; Figure 5 is a schematic diagram of the cross-domain dynamic interaction reconstruction module provided by the present invention; Figure 6 is a sample diagram of the repaired cloud-containing optical remote sensing image provided by the present invention; Figure 7 is a schematic structural diagram of the spatio-spectral reconstruction system of cloud-containing optical remote sensing images integrating spaceborne SAR provided by the present invention; Figure 8 is a schematic structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0036] Aiming at the problems existing in the prior art, the present invention proposes a method and system for spatial-spectral reconstruction of cloud-containing optical remote sensing images integrating spaceborne SAR, constructs a system including a shallow feature aggregation module, a residual adaptive dynamic enhancement modulator and a spatial-spectral reconstruction result output module, captures global semantic information and local spatial details by designing a cross-scale window attention mechanism, and optimizes the global feature expression based on the Top- k channel selection strategy to reduce redundant information. In addition, it also uses channel information to guide the enhancement of spatial features, combines a gating mechanism to dynamically adjust the information flow, and fully integrates the complementary characteristics of optical images and SAR images. While overcoming the cloud cover problem, the system significantly improves the detail expression ability and visual quality of the reconstructed images, providing an efficient and accurate solution for remote sensing image reconstruction in complex environments.
[0037] Figure 1 is a schematic flow chart of the method for spatial-spectral reconstruction of cloud-containing optical remote sensing images integrating spaceborne SAR provided by an embodiment of the present invention. As Figure 1 shown, it includes: Step 100: Obtain an optical image and a SAR image; Step 200: Input the optical image and the SAR image into the shallow feature aggregation module, and output the aggregated low-level features of the optical image and the SAR image; Step 300: Process the aggregated low-level features through a residual adaptive dynamic enhancement modulator to obtain a multi-scale spatial feature map; Step 400: Input the multi-scale spatial feature map into the spatial-spectral reconstruction result output module for spatial-spectral reconstruction to obtain a restored optical image with a preset clarity.
[0038] Based on the above embodiment, as Figure 2 shown, it includes: Step 1, construct a shallow feature aggregation module. The shallow feature aggregation module is a double-branch series structure, including an input tensor merging layer, a 3×3 convolutional layer with a stride of 1, and an attention layer, which is used to extract and aggregate the low-level features of the input optical and SAR images.
[0039] Step 2, construct a residual adaptive dynamic enhancement modulator. The residual adaptive dynamic enhancement modulator is composed of multiple information gain blocks, and each information gain block includes a multi-scale spatial fine-grained coupling module, a sparse global statistical cooperation module, and a cross-domain dynamic interaction reconstruction module.
[0040] Step 3: Construct an empty-spectrum reconstruction result output module, which is structured as a 3×3 convolutional layer with a stride of 1 followed by a polarization self-attention layer and then another 3×3 convolutional layer with a stride of 1. It is used to perform empty-spectrum reconstruction on the feature map after shallow feature aggregation and enhancement processing to restore a clear optical image.
[0041] In one embodiment, Step 1 includes: Step 1.1: Perform an operation to merge the input optical image and SAR image. The formula is expressed as: (1) In the formula, denotes concatenating two tensors along the channel dimension. For the present invention, the cloud-containing optical remote sensing image and the SAR image have sizes of and respectively. Then, the size of the merged tensor is .
[0042] Step 1.2: Input the tensor merged in Step 1.1 into a 3×3 convolutional layer with a stride of 1. This convolutional layer is used to extract low-level features from the merged tensor. The formula for the convolution operation is expressed as: (2) In the formula, is the convolution kernel, is the bias term, and the output of the convolution operation has a size of .
[0043] Step 1.3: Input the feature map processed by the convolutional layer in Step 1.2 into the attention layer. The role of this layer is to further enhance the attention to important regions by learning the weighted coefficients at different positions. The present invention uses the self-attention mechanism to assign weights to each spatial position. The formula is expressed as: (3) In the formula, , , are the query, key, and value matrices respectively, obtained by performing a linear transformation on ; is the dimension of the key vector, used to scale the dot product; is the normalization operation, used to ensure that the sum of the weights is 1; is the output feature map after weighting, containing the weight adjustment for each spatial position. Finally, the output feature map after being weighted by the attention mechanism is expressed as 。 is the final attention-weighted output of the shallow feature aggregation module in step 1, with low-level feature information extracted and aggregated from optical and SAR images.
[0044] In one embodiment, step 2 includes: Step 2.1, the multi-scale spatial fine-grained coupling module adopts a cross-scale window attention mechanism, which includes key components such as layer normalization, downsampling module, neighborhood padding convolution, multi-head self-attention, and 1×1 convolution. First, the input features are processed by layer normalization, and then multi-scale downsampling is achieved by using stride average pooling and linear projection to extract cross-scale features. The neighborhood padding convolution is used to generate query vectors ( ), key vectors ( ), and value vectors ( ). Through the cross-scale feature aggregation strategy, the local region and global information are fully interacted to generate a cross-scale spatial attention map. With the support of the multi-head self-attention mechanism, the module fuses local and global features, and finally generates a multi-scale spatial feature map through 1×1 convolution, which not only expands the receptive field but also significantly enhances the expression ability of fine-grained spatial information. As Figure 3 shown, it specifically includes the following steps: Perform layer normalization (Layer Normalization) on the features obtained in step 1 to improve the stability of training and accelerate the convergence speed. It is expressed as: (4) Perform average pooling operation with a stride of on for downsampling to obtain the downsampled feature map , which is expressed as: (5) Perform linear projection on to adjust the channel dimension, and at the same time use a shortcut connection to retain the original features. The formula is: (6) By designing a neighborhood padding convolution, perform 3×3 convolution operation on the feature map to generate query vectors ( ), key vectors ( ), and value vectors ( ). The formula is: (7) (8) (9) Among them, , , is the local window size.
[0045] By introducing a cross-scale multiplication operation and adopting a multi-scale feature offset aggregation strategy, the query vector for each local block interacts with the relevant information of the key vector to generate a cross-scale spatial attention map, specifically as follows: (10) Among them, is the relative position encoding, used to enhance the perception of spatial information; is the scaling factor, used to normalize the attention scores; is used to normalize the attention distribution. Each local block query vector can interact with the region corresponding to the downsampled feature map in the corresponding region. This operation alleviates the blur by using the clear blocks of the downsampling as prior information.
[0046] Step 2.2, the sparse global statistical collaboration module realizes efficient sparse calculation in feature fusion by using the Top- k channel attention mechanism. This module first extracts the channel and spatial information of the input features through 1×1 convolution and 3×3 depth convolution, and then calculates the similarity between the query ( ) and the key ( ) to generate an attention matrix. Through the Top- k selection operation, only the most relevant channel features are retained, filtering out unnecessary redundant information. Finally, through softmax normalization processing, a sparse channel attention map is obtained to optimize the interaction and fusion of cross-domain features. This structure can effectively improve the computational efficiency of the model and avoid introducing irrelevant noise in the image reconstruction task. As Figure 4 shown, it specifically includes the following steps: Perform 1×1 convolution and 3×3 depth convolution operations on the input features to generate channel context encoding to extract the potential feature information of the image.
[0047] Calculate the attention weights by calculating the similarity between the query ( ) and the key ( ). The present invention designs a channel-based self-attention mechanism to avoid excessive calculation in the spatial dimension.
[0048] (11) Among them, and are the query and key respectively generated from the input features, is the scaling factor of the feature dimension, is the generated attention matrix.
[0049] By selecting the attention matrix among the top most relevant elements to achieve sparse attention. By setting a threshold , elements below this value are set to zero, thus only retaining the most relevant partial information. Specifically, the Top- k selection operation is implemented as follows: (12) where is the Top- k selection operator, is the maximum attention score threshold between the th query and key pair.
[0050] The softmax function is used to normalize the sparsified attention matrix, thereby obtaining a sparsity-optimized channel attention map. This process can be expressed as: (13) where is the Top- k selection operation, , , are the query, key, and value vectors respectively, and the final sparsified attention map is obtained after the softmax operation.
[0051] The outputs of the self-attention with multiple heads are processed through a summation operation to obtain a sparsity-optimized channel attention map , which captures the most important channel features.
[0052] Step 2.3, the cross-domain dynamic interaction reconstruction module aims to achieve effective feature interaction and fusion between the optical image and the SAR image by fusing the spatial and channel information between them. As Figure 5 shown, it specifically includes the following steps: Use 1×1 convolution and 3×3 depth convolution operations to perform information aggregation on the input spatial features and channel features . This step extracts spatial and channel information through convolution to prepare for subsequent feature enhancement.
[0053] To more effectively fuse information from different domains, the present invention designs and adopts a gating mechanism. By using the GELU activation function, the channel features are non-linearly activated and combined with the spatial features Perform element-wise multiplication. This gating mechanism can selectively enhance spatial features, thereby introducing global channel information into the spatial features.
[0054] Adjust the enhanced features through linear projection and combine them with the original spatial features to obtain the final reconstruction result .
[0055] The cross-domain dynamic interaction reconstruction module designed in the present invention realizes deep cooperation between different modality features through the guidance and enhancement of spatial features by channel information. This module makes full use of the advantages of SAR images in penetrating clouds and the strengths of optical images in expressing ground object details, forming a complementary and cooperative feature fusion method, greatly improving the reconstruction accuracy and visual effect of cloud-covered areas, and providing a robust technical guarantee for the high-quality reconstruction of remote sensing images.
[0056] In one embodiment, step 3 includes: Existing methods for repairing cloud-covered images usually use the L1 loss for information reconstruction. However, since the L1 loss mainly focuses on the pixel-level absolute difference and less considers the perceptual quality of the image, the structural information of the image is often ignored during the reconstruction process. To better preserve the structural details of the image, the present invention takes the L1 loss as the basic loss and simultaneously introduces the structural similarity (SSIM) loss to enhance the spatial detail representation ability of the reconstructed remote sensing image. Therefore, the mathematical expression of the loss function of the proposed method for spatio-spectral reconstruction of cloud-containing optical remote sensing images fused with spaceborne SAR is as follows: (14) where, represents the output image, represents the corresponding target image, is a hyperparameter used to adjust the relative importance of the two loss functions, usually set to 0.1.
[0057] The definition of the L1 loss is as follows: (15) where, represents the total number of pixels in the image, represents the index of the image patch, represents the L1 norm, that is, the sum of absolute values.
[0058] The definition of the loss is as follows: (16) By introducing the Structural Similarity (SSIM) loss into the total loss function, compared with the traditional L1 loss that relies on pixel-level absolute differences, the present invention can better preserve the structural details of images and improve the spatial detail expression ability of the reconstructed images in cloud-covered areas. In addition, based on the L1 loss, the present invention combines loss, and flexibly adjusts the weights of the two through the hyperparameter , achieving a balance between the pixel accuracy and perceptual quality of the reconstructed images and greatly improving the overall reconstruction effect of the images.
[0059] The following uses a specific embodiment to illustrate the technical solution of the present invention.
[0060] The experiment of this embodiment was carried out in the hardware environment of NVIDIA RTX A5000 and the software environment of Python.
[0061] By using the AdamW optimizer to optimize the constructed spatio-spectral reconstruction system of cloud-containing optical remote sensing images integrated with spaceborne SAR, the initial learning rate is set to 1×10 -5 , and the learning rate is halved every five iteration cycles, and a total of 50 iteration times are trained.
[0062] The dataset used in this embodiment is the SEN12MS-CR dataset released by the Technical University of Munich, Germany. The dataset has a total of 122,218 pairs of images. Each pair of images includes 1 Sentinel-1 SAR image, 1 Sentinel-2 cloud-containing optical image, and 1 Sentinel-2 cloud-free optical image at a nearby time. The training set, validation set, and test set are divided according to the ratio of 5:3:2, that is, the numbers of the training set, validation set, and test set are 61,109 pairs, 36,665 pairs, and 24,444 pairs respectively.
[0063] The restoration effect of the present invention is compared with four existing methods published in authoritative journals, namely: [Li C, Liu X, Li S. Transformer meets GAN: Cloud-free multispectral image reconstruction via multi-sensor data fusion in satellite images[J]. IEEE Transactions on Geoscience and Remote Sensing, 2023.] (comparison method 1), [Ma J, Chen Y, Pan J, et al. SCT-CR: A synergistic convolution-transformer modeling method using SAR-optical data fusion for cloud removal[J]. International Journal of Applied Earth Observation and Geoinformation, 2024, 130: 103909.] (comparison method 2), [Pan J, Xu J, Yu X, et al. HDRSA-Net: Hybrid dynamic residual self-attention network for SAR-assisted optical image cloud and shadow removal[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2024, 218: 258-275.] (comparison method 3), [Li C, Li S, Liu X. Breaking through clouds: A hierarchical fusion network empowered by dual-domain cross-modality interactive attention for cloud-free image reconstruction[J]. Information Fusion, 2025, 113: 102649.] (comparison method 4).
[0064] The corresponding restoration results are as Figure 6 shown Figure 6From top to bottom, each row in turn is: Sentinel-2 cloud-containing optical image to be repaired, Sentinel-1 SAR image, Sentinel-2 cloud-free optical image at a nearby time as the standard reference, result of comparison method 1, result of comparison method 2, result of comparison method 3, result of comparison method 4, result of the method of the present invention. The comparison results show that, compared with other methods, the cloud-containing optical remote sensing image spatio-spectral reconstruction system integrating spaceborne SAR designed by the present invention performs optimally in terms of detail reconstruction in cloud-covered areas and maintaining global semantic consistency. There are varying degrees of detail loss and texture blurring problems in the reconstruction process of other methods, while the present invention can significantly improve these deficiencies and effectively enhance the visual quality and structural integrity of the images.
[0065] The cloud-containing optical remote sensing image spatio-spectral reconstruction system integrating spaceborne SAR provided by the present invention will be described below. The cloud-containing optical remote sensing image spatio-spectral reconstruction system described below can be mutually referred to the cloud-containing optical remote sensing image spatio-spectral reconstruction method described above.
[0066] Figure 7 It is a schematic structural diagram of the cloud-containing optical remote sensing image spatio-spectral reconstruction system integrating spaceborne SAR provided by an embodiment of the present invention, as Figure 7 shown, including: an acquisition unit 71, an aggregation unit 72, a modulation unit 73 and an output unit 74, wherein: The acquisition unit 71 is used to acquire an optical image and a SAR image; the aggregation unit 72 is used to input the optical image and the SAR image into a shallow feature aggregation module and output the aggregated low-level features of the optical image and the SAR image; the modulation unit 73 processes the aggregated low-level features through a residual adaptive dynamic enhancement modulator to obtain a multi-scale spatial feature map; the output unit 74 is used to input the multi-scale spatial feature map into a spatio-spectral reconstruction result output module for spatio-spectral reconstruction to obtain a restored optical image with a preset clarity.
[0067] Figure 8 Illustrates a schematic structural diagram of an electronic device, as Figure 8As shown in the figure, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communications interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the method for spatial-spectral reconstruction of cloud-containing optical remote sensing images integrating spaceborne SAR. The method includes: acquiring an optical image and an SAR image; inputting the optical image and the SAR image into a shallow feature aggregation module to output the aggregated low-level features of the optical image and the SAR image; processing the aggregated low-level features through a residual adaptive dynamic enhancement modulator to obtain a multi-scale spatial feature map; inputting the multi-scale spatial feature map into a spatial-spectral reconstruction result output module for spatial-spectral reconstruction to obtain a restored optical image with a preset clarity.
[0068] In addition, when the logic instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0069] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0070] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A method for spatial-spectral reconstruction of cloud-containing optical remote sensing images fused with space-borne SAR, characterized in that: include: Acquire optical and SAR images; Inputting the optical image and the SAR image into a shallow feature aggregation module, and outputting aggregated low-level features of the optical image and the SAR image; Processing the aggregated low-level features through a residual adaptive dynamic enhancement modulator to obtain a multi-scale spatial feature map; The multi-scale spatial feature map is input into the spatial spectrum reconstruction result output module to perform spatial spectrum reconstruction to obtain a restored optical image with a preset clarity.
2. The method for reconstructing cloud-containing optical remote sensing images by integrating space-borne SAR according to claim 1, characterized in that: Inputting the optical image and the SAR image into a shallow feature aggregation module, and outputting aggregated low-level features of the optical image and the SAR image, including: Determining that the shallow feature aggregation module includes an input tensor merging layer, a 3×3 convolutional layer with a stride of 1, and an attention layer; The input tensor merging layer merges the optical image and the SAR image to obtain a merged tensor : In the formula, Indicates concatenation of two tensors according to the channel dimension, optical imaging and SAR images The sizes are and , merge the tensors The size is ; The tensors will be merged Input the 3×3 convolution layer with a step size of 1 to obtain the convolution output feature map : In the formula, is the convolution kernel, is the bias term, the convolution output feature map The size is ; The convolution output feature map Input the attention layer to get the weighted output feature map : In the formula, , , They are query matrix, key matrix and value matrix respectively. Performing linear transformation, we get is the dimension of the key vector, used to scale the dot product, is a normalization operation to ensure that the sum of the weights is 1. It is a weighted output feature map, which contains the weight adjustment of each spatial position; The output feature map after weighting by the attention mechanism is , It is the final attention-weighted output of the shallow feature aggregation module, including aggregated low-level features extracted from the optical image and the SAR image.
3. The method for reconstructing cloud-containing optical remote sensing images by integrating space-borne SAR according to claim 1, characterized in that: The aggregated low-level features are processed by a residual adaptive dynamic enhancement modulator to obtain a multi-scale spatial feature map, including: Determine that the residual adaptive dynamic enhancement modulator includes a plurality of information gain blocks, each of which includes a multi-scale spatial fine-grained coupling module, a sparse global statistical collaboration module, and a cross-domain dynamic interactive reconstruction module; Inputting the aggregated low-level features into the multi-scale spatial fine-grained coupling module to generate a cross-scale spatial attention map; Inputting the cross-scale spatial attention map into the sparse global statistical collaboration module to obtain a sparse channel attention map; The sparse channel attention map is input into the cross-domain dynamic interaction reconstruction module to obtain a multi-scale spatial feature map.
4. The method for reconstructing cloud-containing optical remote sensing images by integrating space-borne SAR according to claim 3, characterized in that: Inputting the aggregated low-level features into the multi-scale spatial fine-grained coupling module to generate a cross-scale spatial attention map, including: The multi-scale spatial fine-grained coupling module includes layer normalization, downsampling module, neighborhood padding convolution, multi-head self-attention and 1×1 convolution; Will Perform layer normalization to obtain layer normalized features : In the formula, is the layer normalization operation; By step length The average pooling operation of Downsample to obtain the downsampled feature map : In the formula, is the average pooling operation; right Perform linear projection to adjust the channel dimension, use shortcut connections to retain the original features, and obtain the feature map : In the formula, It is a linear operation; For feature maps Perform a 3×3 neighborhood padding convolution operation to generate a query vector , key vector Sum value vector : in, , , is the local window size, It is a 3×3 convolution operation; By introducing the cross-scale multiplication operation and adopting the multi-scale feature offset aggregation strategy, the query vector Each local block of Interact with relevant information to generate a cross-scale spatial attention map : in, It is a relative position encoding used to enhance spatial information perception. is a scaling factor used to normalize the attention scores, Used to normalize the attention distribution, each The local block query vector Can be compared with the corresponding area in the downsampled feature map This operation alleviates blur by using the downsampled clear blocks as prior information.
5. The method for reconstructing cloud-containing optical remote sensing images by integrating space-borne SAR according to claim 4, characterized in that: Inputting the cross-scale spatial attention map into the sparse global statistical collaboration module to obtain a sparse channel attention map, including: Performing 1×1 convolution and 3×3 deep convolution operations on the cross-scale spatial attention map to generate channel context encoding and extract potential feature information of the image; Channel-based self-attention mechanism to calculate queries and key The similarity between them is used to calculate the attention weight: in, and are the queries and keys generated from the input features, respectively. is the scaling factor of the feature dimension, is the generated attention matrix; Using Top- k Select the operation to select the attention matrix Center front The most relevant elements: in, It is Top- k Select the operator, It is The maximum attention score threshold between query and key pairs; The softmax function is used to normalize the sparse attention matrix to obtain the sparse optimized channel attention map: in, It is Top- k Select an action, , , They are query vector, key vector and value vector respectively. After softmax operation, the final sparse attention map is obtained ; The multi-head self-attention output is processed through the aggregation operation to obtain a sparse channel attention map .
6. The method for reconstructing cloud-containing optical remote sensing images by integrating space-borne SAR according to claim 5, characterized in that: The sparse channel attention map is input into the cross-domain dynamic interaction reconstruction module to obtain a multi-scale spatial feature map, including: The spatial features of the input are processed using 1×1 convolution and 3×3 depth convolution operations. and channel characteristics Aggregate information; Based on the gating mechanism, the GELU activation function is used to activate the channel features. Perform nonlinear activation and spatial features Perform element-by-element multiplication to obtain enhanced features; The enhanced features are adjusted by linear projection and compared with the spatial features. Add them together to get a multi-scale spatial feature map .
7. The method for reconstructing cloud-containing optical remote sensing images by integrating space-borne SAR according to claim 1, characterized in that: Inputting the multi-scale spatial feature map into the spatial spectrum reconstruction result output module for spatial spectrum reconstruction to obtain a restored optical image with a preset definition, including: The spatial spectrum reconstruction result output module includes a 3×3 convolution layer with a step size of 1, a polarization self-attention layer and a 3×3 convolution layer with a step size of 1 connected in sequence; L1 loss and structural similarity SSIM loss are used for spatial spectrum reconstruction, and the restored optical image of preset clarity is output.
8. The method for reconstructing cloud-containing optical remote sensing images by integrating space-borne SAR according to claim 7, characterized in that: L1 loss and structural similarity SSIM loss are used for spatial spectrum reconstruction, including: The total loss function that combines L1 loss and SSIM loss for: in, Represents the output image, represents the corresponding target image, is a hyperparameter used to adjust the relative importance of the two loss functions; The L1 loss function is: in, represents the total number of pixels in the image, represents the index of the image block, represents the L1 norm, that is, the sum of absolute values; The SSIM loss function is: 。 9. A cloud-containing optical remote sensing image spatial spectrum reconstruction system integrated with spaceborne SAR, characterized in that: include: An acquisition unit, used for acquiring optical images and SAR images; An aggregation unit, used for inputting the optical image and the SAR image into a shallow feature aggregation module, and outputting aggregated low-level features of the optical image and the SAR image; A modulation unit, processing the aggregated low-level features through a residual adaptive dynamic enhancement modulator to obtain a multi-scale spatial feature map; The output unit is used to input the multi-scale spatial feature map into the spatial spectrum reconstruction result output module for spatial spectrum reconstruction to obtain a restored optical image with a preset clarity.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for spatial-spectral reconstruction of cloud-containing optical remote sensing images fused with space-borne SAR as described in any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Image optimization method and device for iron tower climbing robot and storage medium
CN121121505A