Multimodal multi-temporal surface water probability gap filling method and device
By employing a multimodal, multi-temporal surface water probabilistic gap-filling method, deep neural networks and generative adversarial networks are used to process surface water probabilistic data, solving the problem of non-continuous pixel value data in surface water mapping and improving the accuracy and reliability of surface water mapping.
Patent Information
- Application Number
- CN202411790653.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2044-12-06
AI Technical Summary
Existing technologies struggle to handle the non-continuous pixel value data required for surface water mapping and lack sufficient high-quality training samples, resulting in insufficient generalization ability of the model under different environmental conditions and reducing the accuracy and reliability of surface water mapping.
A multimodal and multitemporal surface water probability gap filling method is adopted. By acquiring the target dataset for multimodal and multitemporal surface water probability gap filling in the target area, the gap filling is performed using a target deep neural network model, and a generative adversarial network is used to train the model to generate a surface water probability map without gaps, and target image enhancement processing is performed.
It effectively improves the accuracy and reliability of surface water mapping, solves the problem of processing non-continuous pixel value data, and enhances the model's generalization ability under different environmental conditions.
Smart Images

Figure CN119942360B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent interpretation technology of remote sensing images, and in particular to a method and apparatus for filling probabilistic gaps in multimodal and multitemporal surface water. Background Technology
[0002] In recent years, deep learning algorithms have been increasingly widely used in cloud removal tasks. These algorithms can be broadly categorized into single-temporal and multi-temporal methods. In single-temporal methods, researchers have effectively utilized Sentinel-1 information to remove clouds from Sentinel-2 images through residual learning. To address the domain differences between SAR (Synthetic Aperture Radar) and optical images, another study designed a two-stream network to hierarchically integrate SAR features into optical features. Furthermore, some researchers have introduced gated and dilated pyramid convolutions into coarse-to-fine structures, achieving multi-task reconstruction of remote sensing images. Compared to CNN (Convolutional Neural Network) structures, ViT (Vision Transformer) has demonstrated powerful capabilities in long-range dependency modeling in image processing, prompting its introduction into recent cloud removal research.
[0003] While single-temporal methods are relatively mature, they often overlook the multi-temporal information provided by satellite imagery, which is crucial for addressing the problem of thick cloud cover. In multi-temporal methods, researchers have proposed STGAN (Spatial-Temporal Generative Adversarial Network) methods that fuse spatiotemporal information, using multi-branch ResNet or U-Net as the generator and PatchGAN as the discriminator. Other researchers have embedded 3D convolutions into STGAN-like structures, using optical and SAR imagery as input, and achieved convincing results in both "sequence-to-point" and "sequence-to-sequence" generation modes. Furthermore, some studies have proposed a time-matching method within a GAN architecture to account for temporal differences in multi-temporal optical imagery. Although existing multi-temporal methods perform well, they primarily rely on convolutional layers for multimodal information extraction and fusion, which limits their sensitivity to long-range dependencies.
[0004] Cloud removal and probabilistic surface water gap-filling algorithms share similarities in input and output: the input is an optical image or its derivatives, containing pixels with continuous values; the algorithm may involve SAR imagery as auxiliary information; and the output is the gap-filled image. Since cloud interference is a significant challenge for high-frequency surface water mapping using remote sensing imagery, researchers have proposed various deep learning methods based on CNNs and Transformers for cloud removal from optical remote sensing imagery. However, these methods are rarely applied to surface water mapping because traditional water distribution maps do not have the same continuous pixel values as the original remote sensing imagery, making it difficult to handle the non-continuous pixel value data required for surface water mapping. Furthermore, the lack of sufficient high-quality training samples results in insufficient generalization ability of the model under different environmental conditions, reducing the accuracy and reliability of surface water mapping, a problem that urgently needs to be addressed. Summary of the Invention
[0005] This application is based on the inventor's understanding and insights into the following issues:
[0006] Large-scale, high-frequency, and high-resolution surface water mapping is of great significance in various applications such as water resource protection, government monitoring, and flood and drought control. Increasing the monitoring frequency is particularly important for supporting long-term time-series analysis and emergency response missions. Compared to SAR imagery, optical remote sensing imagery is widely used in surface water mapping due to its rich spectral information and high precision. However, optical imagery is significantly affected by clouds, cloud shadows, and topographical shading, leading to incomplete observations and spatial data gaps. Simultaneously, the frequency of satellite revisits also causes temporal gaps in surface water distribution maps. These data gaps hinder the continuity and reliability of surface water mapping.
[0007] In recent years, researchers have increasingly focused on gap filling in water body distribution maps. Some methods fill in missing pixels based on the spatial or temporal correlation of optical water body distribution maps. Due to the powerful cloud-penetrating capabilities of SAR, many existing methods use SAR-based water extraction results as supplementary information for optical surface water mapping. In fact, a large amount of research has focused on gap filling in different types of remote sensing data, such as surface temperature, soil moisture, and optical remote sensing imagery (in this case, generally referred to as cloud removal). Currently, deep learning methods based on CNNs and Transformers have received widespread attention in cloud removal tasks and have achieved reliable results. However, these advanced methods are rarely applied to gap filling in traditional water body distribution maps, mainly because water body distribution maps are usually discontinuous binary images, such as non-water body pixels having a value of 0 and water body pixels having a value of 1, which is not conducive to deep neural networks learning.
[0008] To more effectively preserve detailed information in high-resolution remote sensing imagery, numerous studies have explored fuzzy mapping techniques, where the value of each pixel represents the probability of that pixel belonging to a specific category. Compared to traditional water distribution maps, fuzzy water distribution maps, as probabilistic maps, contain richer information and can quantitatively reflect the uncertainties of water mapping methods. Against this backdrop, WP (Water Probability) data has been introduced into surface water mapping tasks, and its probability values can be converted into hard classification maps by adjusting thresholds. DW (Dynamic World), as a high-precision open-source land use and land cover dataset, provides WP data at a 10-meter resolution, making it highly valuable for applications. However, since WP data originates from Sentinel-2 (S2) optical observations, it still faces data integrity issues. Given that the pixel values in WP data are continuously distributed between 0 and 1, providing sufficient information for neural networks, the gap-filling problem in WP data can be considered a special case of cloud removal tasks and may potentially be solved using similar methods.
[0009] Traditional cloud removal methods utilize the spectral, temporal, and spatial information of optical remote sensing imagery to reconstruct areas obscured by clouds. For example, some researchers have found that cloud-covered images have a higher signal-to-noise ratio, thus developing a noise adjustment model to remove thin clouds. Another approach combines information from low-resolution imagery for cloud removal, improving the effect through multiple error correction steps. Furthermore, researchers have proposed a gap-filling method that does not rely on additional satellite data, using a harmonic model to fill in time series data. Although traditional methods do not require large amounts of training data, their effectiveness is limited due to their inability to accurately extract high-level features and fuse multimodal data. Since cloud interference is a major challenge in high-frequency surface water mapping using remote sensing imagery, researchers have proposed various deep learning methods based on CNNs and Transformers for cloud removal from optical remote sensing imagery. However, traditional water distribution maps do not have the same continuous pixel values as the original remote sensing imagery, and therefore are rarely used in surface water mapping, requiring further improvement.
[0010] This application provides a multimodal, multi-temporal surface water probabilistic gap filling method and apparatus to solve the problems in related technologies, such as the difficulty in processing non-continuous pixel value data required for surface water mapping and the lack of sufficient high-quality training samples, which leads to insufficient generalization ability of the model under different environmental conditions and reduces the accuracy and reliability of surface water mapping.
[0011] The first aspect of this application provides a method for filling probabilistic gaps in multimodal and multitemporal surface water probability, comprising the following steps: acquiring a target dataset for filling probabilistic gaps in multimodal and multitemporal surface water probability in a target region that meets preset conditions; performing surface water probability gap filling processing on the target dataset using a target deep neural network model to generate a surface water probability map without gaps; training the target deep neural network model using a target generative adversarial network to generate a trained model, and performing target image enhancement processing on the surface water probability map without gaps based on the trained model to generate a multimodal and multitemporal surface water probability filled image that meets preset enhancement conditions.
[0012] Optionally, in one embodiment of this application, the step of obtaining the target dataset for filling the multimodal and multitemporal surface water probability gaps in the target area that meets preset conditions includes: collecting synthetic aperture radar data and surface water probability data in the target area that meet preset time interval conditions and preset spatial resolution conditions; and processing the synthetic aperture radar data and the surface water probability data to generate the target dataset for filling the multimodal and multitemporal surface water probability gaps in the target area that meets the preset conditions.
[0013] Optionally, in one embodiment of this application, the step of using a target deep neural network model to perform surface water probability gap filling processing on the target dataset to generate a gap-free surface water probability map includes: performing preliminary surface water probability gap filling processing on the target dataset based on the multi-branch gated repair module in the target deep neural network model to generate a roughly filled surface water probability image; and refining and aligning the roughly filled surface water probability image based on the alignment and thinning module in the target deep neural network model, guided by target temporal synthetic aperture radar data, to generate the gap-free surface water probability map.
[0014] Optionally, in one embodiment of this application, training the target deep neural network model using a target generative adversarial network to generate a trained model includes: constructing a reconstruction loss function and an adversarial loss function in the target generative adversarial network; combining the reconstruction loss function and the adversarial loss function to obtain a joint loss function of the target deep neural network model; and training the target deep neural network model using the joint loss function to generate the trained model.
[0015] Optionally, in one embodiment of this application, the joint loss function is:
[0016]
[0017] Where λ1 and λ2 are hyperparameters set before training. To reconstruct the loss function, This is the adversarial loss function for the generator.
[0018] A second aspect of this application provides a multimodal, multitemporal surface water probability gap filling device, comprising: an acquisition module for acquiring a target dataset for multimodal, multitemporal surface water probability gap filling that meets preset conditions in a target area; a generation module for performing surface water probability gap filling processing on the target dataset using a target deep neural network model to generate a gap-free surface water probability map; and a processing module for training the target deep neural network model using a target generative adversarial network to generate a trained model, and performing target image enhancement processing on the gap-free surface water probability map based on the trained model to generate a multimodal, multitemporal surface water probability filling image that meets preset enhancement conditions.
[0019] Optionally, in one embodiment of this application, the acquisition module includes: a collection unit, used to collect synthetic aperture radar data and surface water probability data in the target area that meet preset time interval conditions and preset spatial resolution conditions; and a processing unit, used to process the synthetic aperture radar data and the surface water probability data to generate a target dataset for filling multimodal and multitemporal surface water probability gaps in the target area that meet the preset conditions.
[0020] Optionally, in one embodiment of this application, the generation module includes: a first generation unit, configured to perform preliminary surface water probability gap filling processing on the target dataset based on the multi-branch gated repair module in the target deep neural network model, to generate a roughly filled surface water probability image; and a second generation unit, configured to perform refinement and alignment processing on the roughly filled surface water probability image based on the alignment and refinement module in the target deep neural network model, guided by the target temporal synthetic aperture radar data, to generate the gap-free surface water probability map.
[0021] Optionally, in one embodiment of this application, the processing module includes: a construction unit for constructing a reconstruction loss function and an adversarial loss function in the target generative adversarial network; an acquisition unit for combining the reconstruction loss function and the adversarial loss function to obtain a joint loss function of the target deep neural network model; and a training unit for training the target deep neural network model using the joint loss function to generate the trained model.
[0022] Optionally, in one embodiment of this application, the joint loss function is:
[0023]
[0024] Where λ1 and λ2 are hyperparameters set before training. To reconstruct the loss function, This is the adversarial loss function for the generator.
[0025] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the multimodal, multitemporal surface water probabilistic gap filling method as described in the above embodiments.
[0026] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described multimodal, multitemporal surface water probabilistic gap-filling method.
[0027] A fifth aspect of this application provides a computer program product, including a computer program that, when executed, is used to implement the above-described multimodal, multitemporal surface water probabilistic gap filling method.
[0028] This application embodiment can acquire a target dataset for multimodal and multitemporal surface water probability gap filling that meets certain conditions in a target area. A target deep neural network model is used to perform surface water probability gap filling processing on the target dataset to generate a gap-free surface water probability map. A target generative adversarial network is then used to train the target deep neural network model to generate a trained model. This model is then used to perform target image enhancement processing on the gap-free surface water probability map to generate a multimodal and multitemporal surface water probability filled image that meets certain enhancement conditions, effectively improving the accuracy and reliability of surface water mapping. This solves the problems in related technologies, such as the difficulty in processing non-continuous pixel value data required for surface water mapping and the lack of sufficient high-quality training samples, which leads to insufficient generalization ability of the model under different environmental conditions and reduces the accuracy and reliability of surface water mapping.
[0029] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0030] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0031] Figure 1 This is a flowchart of a multimodal, multitemporal surface water probabilistic gap filling method provided according to an embodiment of this application;
[0032] Figure 2A sampling diagram illustrating a method for creating a multimodal, multi-temporal dataset according to a specific embodiment of this application;
[0033] Figure 3 This is a schematic diagram of training and testing sample pairs for a multimodal, multi-temporal dataset according to a specific embodiment of this application;
[0034] Figure 4 This is a schematic diagram of gated convolution in a multi-branch gated repair module according to a specific embodiment of this application;
[0035] Figure 5 This is a schematic diagram of the Restormer module in the SAR-guided alignment and refinement module of a specific embodiment of this application;
[0036] Figure 6 A neural network structure diagram of a multimodal, multitemporal surface water probabilistic gap-filling method according to a specific embodiment of this application;
[0037] Figure 7 This is a schematic diagram of a multi-branch SN-PatchGAN discriminator according to a specific embodiment of this application;
[0038] Figure 8 A comparison diagram showing the effects of a surface water probabilistic gap filling method and a frontier gap filling method in a specific embodiment of this application;
[0039] Figure 9 This is a schematic diagram of a multimodal, multitemporal surface water probabilistic gap-filling device provided according to an embodiment of this application;
[0040] Figure 10 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application. Detailed Implementation
[0041] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0042] The following description, with reference to the accompanying drawings, describes an embodiment of the multimodal and multitemporal surface water probability gap filling method and apparatus of this application. Addressing the problems mentioned in the background art regarding the difficulty in processing non-continuous pixel value data required for surface water mapping and the lack of sufficient high-quality training samples, which reduces the accuracy and reliability of surface water mapping, this application provides a multimodal and multitemporal surface water probability gap filling method. In this method, a target dataset for multimodal and multitemporal surface water probability gap filling that meets certain conditions in a target area can be obtained. A target deep neural network model is used to perform surface water probability gap filling processing on the target dataset to generate a gap-free surface water probability map. A target generative adversarial network is used to train the target deep neural network model to generate a trained model, thereby performing target image enhancement processing on the gap-free surface water probability map to generate a multimodal and multitemporal surface water probability filled image that meets certain enhancement conditions, effectively improving the accuracy and reliability of surface water mapping. This solves the problems in related technologies, such as the difficulty in processing non-continuous pixel value data required for surface water mapping and the lack of sufficient high-quality training samples, which reduces the accuracy and reliability of surface water mapping.
[0043] Specifically, Figure 1 This is a flowchart illustrating a multimodal, multitemporal surface water probabilistic gap filling method provided in an embodiment of this application.
[0044] like Figure 1 As shown, the multimodal, multi-temporal surface water probabilistic gap filling method includes the following steps:
[0045] In step S101, the target dataset for filling the multimodal and multitemporal surface water probability gaps in the target area that meets the preset conditions is obtained.
[0046] In this application embodiment, the target area is the area to fill the probability gap in surface water. This application will take the Chinese region as an example for illustration; the preset condition is to select a certain region of interest in the Chinese region, collect data according to the region of interest, and preprocess the data.
[0047] It is understood that the embodiments of this application can obtain a target dataset for multimodal and multitemporal surface water probabilistic gap filling that meets certain conditions in the target area. For example, the embodiments of this application can select 20 0.5°×0.5° regions of interest in China, covering more than 60,000 square kilometers, containing time series data of both optical and SAR modes, and collect data according to the regions of interest. The data is preprocessed, as will be described in the following steps, thereby obtaining a target dataset for multimodal and multitemporal surface water probabilistic gap filling, which effectively improves the feasibility of surface water probabilistic gap filling.
[0048] In one embodiment of this application, obtaining a target dataset for filling multimodal and multitemporal surface water probability gaps in a target area that meets preset conditions includes: collecting synthetic aperture radar data and surface water probability data in the target area that meet preset time interval and preset spatial resolution conditions; and processing the synthetic aperture radar data and surface water probability data to generate a target dataset for filling multimodal and multitemporal surface water probability gaps in the target area that meets preset conditions.
[0049] For example, in this embodiment, data can be collected according to regions of interest. The specific process is as follows: DynamicWorld (DW) is a near real-time global land use and land cover dataset, which can be obtained as an open-source resource on the Google Earth Engine (GEE) platform. Each probability map in DW is generated from a single Sentinel-2 image using a fully convolutional neural network, including nine land cover categories, and has high accuracy. In this embodiment, the "water body" category in the DW data is used as WP (Water Probability), while Sentinel-1 ground resolution data (hereinafter referred to as "S1") from the European Space Agency's Copernicus program is used as auxiliary data for reconstruction. Both types of data, namely synthetic aperture radar data and surface water probability data, have a spatial resolution of 10 meters and are aligned to the WGS84 coordinate system.
[0050] In this embodiment, all data collection processes are automatically completed on the GEE platform. S1 data requires preprocessing before processing, including thermal noise removal, radiometric calibration, and terrain correction. For each region of interest (ROI), data is collected three times a month, with an interval of approximately 10 days (e.g., January 1-10, 11-20, and 21-31). Within each time interval, to merge multiple WP or SAR data, the "median" function in the GEE platform is used to calculate the median of all values for each pixel within each ROI; the WP data range for DW is [0,1], while S1 data is restricted to [-25,0] and [-32.5,0] for VV and VH polarizations, respectively; for missing pixels, both WP and S1 data are assigned a value of 2 in all bands. After the above processing steps, data was collected from 36 time points and all 20 ROIs in 2023. Figure 2 This displays water probability data for a Region of Interest (ROI) at a specific time point and the VV band of Sentinel-1; yellow pixels indicate missing pixels.
[0051] Furthermore, the embodiments of this application can perform data preprocessing. The specific process is as follows: In order to adapt the original data for model input, this application implements multiple preprocessing steps. According to existing research, the effective pixels (i.e., pixel values not equal to 2) in the S1 data are rescaled to the range of [0,1]. The data of each region of interest is divided into 9680 non-overlapping 256×256 pixel patches, and each patch contains WP and S1 data at 36 time points.
[0052] The training and test datasets were generated by filtering paired data. First, it was ensured that the S1 data was complete. A time point in the WP data that was complete was selected as the target time point. At the same time, three other time points were selected as inputs. The selection of input times followed the following two criteria: (1) the gap coverage of the input data should be between 10% and 90%; (2) the input time should be as close as possible to the target time point. Taking "close" in criterion (2) as an example, if April 11 to 20 is selected as the target time point, then April 1 to 10 and April 21 to 30 are preferred as input times, followed by March 21 to 31 and May 1 to 10, and so on. Ultimately, a total of 4635 data tiles were obtained, which were randomly allocated to the training and test sets in a 7:3 ratio. The test set was further divided into three groups based on the missing coverage rate: 10%-30% (896 tiles), 30%-60% (690 tiles), and 60%-90% (1190 tiles). After the above processing, the three groups of data with different missing coverage rates are shown below. Figure 3 As shown, the left side is a water probability image with yellow missing pixels, and the right side is a SAR image; Figure 3 Only the VV band is shown, but both the VV and VH bands were used in training and testing. During the testing phase, image pairs were divided based on the proportion of missing pixels in the most recent input pair (input 1).
[0053] In step S102, the target dataset is filled with surface water probability gaps using a target deep neural network model to generate a surface water probability map without gaps.
[0054] It is understood that the embodiments of this application can utilize a target deep neural network model to perform surface water probability gap filling processing on the target dataset. For example, the embodiments of this application can combine deep neural networks of CNN and Transformer, and adopt a coarse-to-fine strategy to perform multimodal and multitemporal surface water probability filling processing, thereby generating a surface water probability map without gaps, effectively improving the accuracy of surface water probability gap filling.
[0055] In one embodiment of this application, a target deep neural network model is used to perform surface water probability gap filling processing on the target dataset to generate a gap-free surface water probability map. This includes: performing preliminary surface water probability gap filling processing on the target dataset based on the multi-branch gated repair module in the target deep neural network model to generate a roughly filled surface water probability image; and refining and aligning the roughly filled surface water probability image based on the alignment and thinning module in the target deep neural network model, guided by the target temporal synthetic aperture radar data, to generate a gap-free surface water probability map.
[0056] As one possible implementation, embodiments of this application can aggregate multi-temporal information through a multi-branch gated inpainting module. The specific process is as follows: Since gated convolution has been widely used in image inpainting research, unlike ordinary convolution, gated convolution takes the missing image and a mask representing the missing pixels as input, automatically generates a learnable soft mask from the input, and performs pixel-by-pixel and channel-by-channel gating operations on the output of ordinary convolution. Its formula is expressed as:
[0057] Feature=∑∑W f ·I
[0058] Mask = ∑∑W m ·I
[0059] O = φ(Feature) ⊙ σ(Mask)
[0060] Where I represents input, O represents output, and W represents input. m and W f There are two different convolution kernels, φ is an arbitrary activation function, σ refers to the sigmoid function that restricts the mask value range to (0,1), ⊙ is pixel-wise multiplication, Feature is the feature, and Mask is the mask.
[0061] While traditional gated convolution has proven effective, it is not suitable for image sequences such as multi-temporal remote sensing data due to differences in masks at different time points. To address this issue, this application proposes a Branch-Gated Inpainting Module. The module first performs branch feature extraction through multi-layer gated convolution, then aggregates features from different time points using 3D convolution, and finally generates a coarsely filled WP image through 2D convolution.
[0062] Among them, such as Figure 4The diagram shows the gated convolutional structure for each branch of BGIM. In BGIM's gated convolutional structure, multiple downsampling, dilated convolution, and upsampling operations are introduced to ensure a large receptive field and maintain consistent performance. Since each branch has a similar learning objective, weight sharing is used to reduce the number of parameters. 3D convolution is commonly used in video recognition tasks, capable of processing time-series inputs and extracting features from spatial, channel, and temporal dimensions. Considering the high computational intensity of 3D convolution, only a single, unpadded 3D convolutional layer in the temporal dimension is used to aggregate information from different time points.
[0063] Furthermore, this application embodiment can refine and align the coarse output under the guidance of target time SAR data: Although the missing pixels are processed by the BGIM module in the above steps, two problems still exist: (1) the time difference between the input time point and the target time point is not considered; (2) details are lost and blurring occurs in the output image. To solve these problems, this application proposes a SAR-guided SARM (Alignment and Refinement Module). In addition to the coarse output of BGIM, SARM also uses the target SAR image, which provides auxiliary information and promotes alignment with the target time point. In order to gradually integrate the coarse output with the auxiliary information and enhance the details in the missing pixels filled in the image, such as Figure 5 As shown, SARM introduces a transformer block from Restormer; compared to traditional ViT, Restormer uses a self-attention mechanism of DConvs (Depthwise Separable Convolutions) instead of fully connected layers, which is more efficient and easier to train. For input features... The self-attention mechanism of Restormer can be represented as:
[0064]
[0065] in, It is the output feature; It is a 3×3 depth-wise convolution; It is a 1×1 point-wise convolution; LN refers to layer normalization; These are variations of Q, K, and V, respectively; α is a learnable scaling parameter, and X is the input feature.
[0066] Furthermore, gated convolutions are introduced into the Restormer's feedforward network to focus on detailed information in the feature maps. For the input features... The feedforward network of Restormer can be represented as:
[0067]
[0068] Gating(X) is a gating operation.
[0069] like Figure 6 As shown, in SARM, the two data streams (coarse result and SAR) are passed layer by layer through a four-layer symmetric encoder-decoder structure. This process can be represented as:
[0070]
[0071] Among them, F SAR For features extracted from the input SAR image, F coarse The features are extracted from the coarse results. i represents the order of feature extraction, RB refers to the Restormer module, and Up refers to an upsampling layer.
[0072] Subsequently, subtraction is used to emphasize the temporal differences between the coarse result and the target SAR data. It should be noted that all upsampling or downsampling layers in the SARM are implemented through convolution and pixel-shuffle or pixel-unshuffle operations. Finally, a two-dimensional convolution is applied to generate the residual image. Final image with gaps filled Obtained through the following methods:
[0073] O = O coarse +R
[0074] Among them, O coarse The result is a rough estimate, and R is the residual image.
[0075] In step S103, a target deep neural network model is trained using a target generative adversarial network to generate a trained model. Based on the trained model, target image enhancement processing is performed on the surface water probability map without gaps to generate a multimodal, multitemporal surface water probability filling image that meets preset enhancement conditions.
[0076] In this embodiment of the application, the preset enhancement condition is the condition for achieving seamless surface water mapping.
[0077] It is understood that the embodiments of this application can utilize a target generative adversarial network to train a target deep neural network model. For example, by integrating a multi-branch SN-PatchGAN as a discriminator in the generative adversarial network training method, the target deep neural network model in the above steps can be trained to generate a trained model. Based on the trained model, target image enhancement processing can be performed on the surface water probability map without gaps to generate a multimodal, multitemporal surface water probability filling image that meets the enhancement conditions. This can effectively address cloud interference issues, improve the detail of the generated image, and significantly enhance the frequency, range, and resolution of surface water mapping, providing technical support and a new solution for high-frequency, large-range, and high-resolution surface water mapping.
[0078] In one embodiment of this application, a target deep neural network model is trained using a target generative adversarial network to generate a trained model, including: constructing a reconstruction loss function and an adversarial loss function in the target generative adversarial network; combining the reconstruction loss function and the adversarial loss function to obtain a joint loss function of the target deep neural network model; and training the target deep neural network model using the joint loss function to generate a trained model.
[0079] In some embodiments, such as Figure 7 As shown, the embodiments of this application can construct a reconstruction loss function in a target generative adversarial network. First, the reconstruction loss is introduced to evaluate the coarse output result and the final output result, using L1 loss, as shown in the following formula:
[0080]
[0081] Where ||·||1 is the L1 norm, The reconstruction loss function is T, where T is the target image (ground value).
[0082] Next, this embodiment constructs an adversarial loss function in the target generative adversarial network. While L1 loss is relatively stable, it often smooths the reconstructed image. Therefore, this embodiment introduces an adversarial loss to refine the output. Traditional adversarial losses suffer from instability, which SN-PatchGAN largely mitigates by using spectral normalization. To adapt to multi-temporal inputs, SN-PatchGAN can be modified to a branching form similar to BGIM in the above steps. Figure 7As shown, the SN-PatchGAN branch, acting as the discriminator, comprises five 2D convolutions (weight-shared), one 3D convolution, and one 2D convolution, all of which are spectral normalized. The Conditional Least Squares Generative Adversarial Network (LSGAN) is used as the objective function, which can be expressed as:
[0083]
[0084] Where G is the generator, D is the discriminator, x refers to the target ground truth, z refers to the coarse input, y refers to the mask at the three input time points, and p data Refers to the actual data distribution, p z Refers to the distribution of pseudo-data. In terms of mathematical expectation, For the generator's adversarial loss, The total loss of the discriminator, This is the adversarial loss of the discriminator.
[0085] Furthermore, embodiments of this application can combine the reconstruction loss function and the adversarial loss function to obtain a joint loss function, which is the loss function ultimately used by the target deep neural network model. The joint loss function is then used to train the target deep neural network model to generate the trained model, effectively improving the accuracy and reliability of surface water probabilistic gap filling.
[0086] In one embodiment of this application, the joint loss function is:
[0087]
[0088] Wherein, λ1 and λ2 are hyperparameters set before training, and can be set to 100 and 1 respectively; For reconstruction loss function; This is the adversarial loss function for the generator.
[0089] For example, the performance of the proposed model is compared with other cutting-edge gap-filling methods on the dataset constructed in this application. This application also introduces two simple methods for comparison: the surface water probability input with the fewest missing pixels (“Minimum Gap Method”) and the mosaicking result of all surface water probability inputs (“Mosaic Method”). The “mosaicing” process includes the following steps: for each pixel, if only one time point has a valid value, that value is taken as the result; if multiple time points have valid values, the average value is taken; if no valid values are found at any time point, a value of 0.5 is assigned to it to avoid extreme values in subsequent calculations. Missing pixels in the “Minimum Gap Method” are also set to 0.5. These two simple methods reflect the amount of useful information provided by the input data and emphasize the difficulty and importance of the gap-filling task.
[0090] like Figure 8 As shown, two sample blocks are selected in each gap coverage category to visualize and compare the performance of different methods. Figure 8 In the diagram, (1)-(6) represent different image patches in three categories of missing coverage; (a) minimum missing coverage method; (b) mosaic method; (c) DSen2-CR (deep Sentinel-2 correlation reconstruction); (d) GLF-CR (global-local fusion cloud removal); (e) HS2P (hyperspectral to panchromatic image fusion); (f) STGAN; (g) SEN12-MS-CR-TS (time series-based Sentinel-1 and Sentinel-2 multispectral cloud removal); (h) the method proposed in this application; (i) target ground truth image; the yellow pixels in (a) and (b) are missing pixels that cannot be filled; the red boxes mark areas with significant differences.
[0091] Except for sample block (1), the results of the "minimum gap method" and "mosaic method" differed significantly from the target, further emphasizing the necessity of the proposed method. In the 10%-30% category, all deep learning methods performed well, but (g.1) performed poorly due to overfitting. However, the reconstruction results of the model proposed in this application had the smallest deviation from the target in terms of detail. For sample blocks (3) and (4), only the model proposed in this application successfully recovered the details within the red box. For sample block (5), the results of DSen2-CR and GLF-CR showed serious distortion, while the results of STGAN and SEN12MS-CR-TS were also unsatisfactory. Although the "stitching" method showed that the river region in sample block (6) was almost completely polluted, the deep learning method was still able to effectively fill in the missing pixels, in which the reconstruction results of the model proposed in this application were almost identical to the target.
[0092] The accuracy of image reconstruction was evaluated using four commonly used metrics: MAE (Mean Absolute Error), RMSE (Root Mean Squared Error), PSNR (Peak Signal-to-Noise Ratio), and SSIM (Structural Similarity Index). The results are shown in Table 1, which compares the accuracy of this application with other gap-filling methods. The specific details of Table 1 are as follows:
[0093] Table 1
[0094]
[0095] Among the proposed methods, the simple "least gap method" and "mosaic method" performed poorly, especially under conditions of severe gaps. Compared to the previously best-performing method (SEN12MS-CR-TS), the model proposed in this application improved PSNR by +1.6384 and SSIM by +0.0193 in the category with the most severe gaps (60%-90%). The model proposed in this application exhibited the most consistent performance across different gap coverage rates; between the 10%-30% and 60%-90% categories, the differences in MAE, RMSE, PSNR, and SSIM were only 0.0005, 0.0002, 0.1739, and 0.0016, respectively. Furthermore, the model proposed in this application performed best in the 60%-90% category, demonstrating its effective utilization of multimodal and multitemporal information.
[0096] The multimodal and multitemporal surface water probability gap filling method proposed in this application can obtain a target dataset for multimodal and multitemporal surface water probability gap filling that meets certain conditions in a target area. A target deep neural network model is used to perform surface water probability gap filling processing on the target dataset to generate a gap-free surface water probability map. A target generative adversarial network is then used to train the target deep neural network model to generate the trained model. This model is then used to perform target image enhancement processing on the gap-free surface water probability map to generate a multimodal and multitemporal surface water probability filled image that meets certain enhancement conditions, effectively improving the accuracy and reliability of surface water mapping.
[0097] Next, referring to the accompanying drawings, a multimodal, multi-temporal surface water probabilistic gap-filling device according to an embodiment of this application is described.
[0098] Figure 9 This is a block diagram of a multimodal, multitemporal surface water probabilistic gap-filling device according to an embodiment of this application.
[0099] like Figure 9As shown, the multimodal and multitemporal surface water probabilistic gap filling device 10 includes: an acquisition module 100, a generation module 200, and a processing module 300.
[0100] Specifically, the acquisition module 100 is used to acquire the target dataset for filling the multimodal and multitemporal surface water probability gaps in the target area that meets preset conditions.
[0101] The generation module 200 is used to perform surface water probability gap filling processing on the target dataset using the target deep neural network model to generate a surface water probability map without gaps.
[0102] The processing module 300 is used to train a target deep neural network model using a target generative adversarial network to generate a trained model, and based on the trained model, to perform target image enhancement processing on a surface water probability map without gaps to generate a multimodal, multitemporal surface water probability filling image that meets preset enhancement conditions.
[0103] Optionally, in one embodiment of this application, the acquisition module 100 includes: a collection unit and a processing unit.
[0104] The acquisition unit is used to acquire synthetic aperture radar data and surface water probability data in the target area that meet the preset time interval and preset spatial resolution conditions.
[0105] The processing unit is used to process synthetic aperture radar data and surface water probability data to generate a target dataset for filling multimodal and multitemporal surface water probability gaps in the target area that meets preset conditions.
[0106] Optionally, in one embodiment of this application, the generation module 200 includes: a first generation unit and a second generation unit.
[0107] The first generation unit is used to perform preliminary surface water probability gap filling processing on the target dataset based on the multi-branch gated repair module in the target deep neural network model, so as to generate a surface water probability image with rough gap filling.
[0108] The second generation unit, based on the alignment and refinement module in the target deep neural network model, refines and aligns the roughly filled-in surface water probability image under the guidance of the target time synthetic aperture radar data, so as to generate a surface water probability map without gaps.
[0109] Optionally, in one embodiment of this application, the processing module 300 includes: a construction unit, an acquisition unit, and a training unit.
[0110] The building unit is used to construct the reconstruction loss function and the adversarial loss function in the target generative adversarial network.
[0111] The acquisition unit is used to combine the reconstruction loss function and the adversarial loss function to obtain the joint loss function of the target deep neural network model.
[0112] The training unit is used to train the target deep neural network model using a joint loss function to generate the trained model.
[0113] Optionally, in one embodiment of this application, the joint loss function is:
[0114]
[0115] Where λ1 and λ2 are hyperparameters set before training. To reconstruct the loss function, This is the adversarial loss function for the generator.
[0116] It should be noted that the foregoing explanation of the embodiment of the multimodal and multitemporal surface water probabilistic gap filling method also applies to the multimodal and multitemporal surface water probabilistic gap filling device of this embodiment, and will not be repeated here.
[0117] The multimodal and multitemporal surface water probability gap filling device proposed in this application can acquire a target dataset for multimodal and multitemporal surface water probability gap filling that meets certain conditions in a target area. It then uses a target deep neural network model to perform surface water probability gap filling processing on the target dataset to generate a gap-free surface water probability map. Furthermore, it uses a target generative adversarial network to train the target deep neural network model to generate a trained model, thereby performing target image enhancement processing on the gap-free surface water probability map to generate a multimodal and multitemporal surface water probability filling image that meets certain enhancement conditions. This effectively improves the accuracy and reliability of surface water mapping.
[0118] Figure 10 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include:
[0119] The memory 1001, the processor 1002, and the computer program stored on the memory 1001 and capable of running on the processor 1002.
[0120] When the processor 1002 executes the program, it implements the multimodal and multitemporal surface water probabilistic gap filling method provided in the above embodiments.
[0121] Furthermore, electronic devices also include:
[0122] Communication interface 1003 is used for communication between memory 1001 and processor 1002.
[0123] The memory 1001 is used to store computer programs that can run on the processor 1002.
[0124] The memory 1001 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0125] If the memory 1001, processor 1002, and communication interface 1003 are implemented independently, then the communication interface 1003, memory 1001, and processor 1002 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 10 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0126] Optionally, in a specific implementation, if the memory 1001, processor 1002, and communication interface 1003 are integrated on a single chip, then the memory 1001, processor 1002, and communication interface 1003 can communicate with each other through an internal interface.
[0127] The processor 1002 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.
[0128] This embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described multimodal and multitemporal surface water probabilistic gap filling method.
[0129] This embodiment also provides a computer program product, including a computer program, which, when executed, is used to implement the above-described multimodal, multitemporal surface water probabilistic gap filling method.
[0130] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0131] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0132] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0133] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0134] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0135] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it includes one or a combination of the steps of the method embodiments.
[0136] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0137] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A multimodal, multi-temporal surface water probabilistic gap filling method, characterized in that, Includes the following steps: Obtain the target dataset for filling in the probability gaps of multimodal and multitemporal surface water in the target region that meets the preset conditions; The target dataset is filled with surface water probability gaps using a target deep neural network model to generate a surface water probability map without gaps. The target deep neural network model is trained using a target generative adversarial network to generate a trained model. Based on the trained model, the target image enhancement processing is performed on the surface water probability map without gaps to generate a multimodal and multitemporal surface water probability filling image that meets preset enhancement conditions. The target dataset for obtaining multimodal and multitemporal surface water probability gap filling that meets preset conditions in the target area includes: Synthetic aperture radar data and surface water probability data that meet preset time interval and preset spatial resolution conditions are collected in the target area; The synthetic aperture radar data and the surface water probability data are processed to generate a target dataset for filling the multimodal and multitemporal surface water probability gaps in the target area that meets the preset conditions. The step of using a target deep neural network model to fill in the surface water probability gaps in the target dataset to generate a surface water probability map without gaps includes: Based on the multi-branch gated repair module in the target deep neural network model, the target dataset is subjected to preliminary surface water probability gap filling processing to generate a roughly filled surface water probability image. Based on the alignment and refinement module in the target deep neural network model, and guided by the target time synthetic aperture radar data, the roughly filled surface water probability image is refined and aligned to generate the missing surface water probability map.
2. The method according to claim 1, characterized in that, The step of training the target deep neural network model using a target generative adversarial network to generate the trained model includes: Construct the reconstruction loss function and adversarial loss function in the target generative adversarial network; By combining the reconstruction loss function and the adversarial loss function, the joint loss function of the target deep neural network model is obtained; The target deep neural network model is trained using the joint loss function to generate the trained model.
3. The method according to claim 2, characterized in that, The joint loss function is: in, and All of these are hyperparameters set before training. To reconstruct the loss function, This is the adversarial loss function for the generator.
4. A multimodal, multi-temporal surface water probabilistic gap-filling device, characterized in that, include: The acquisition module is used to acquire the target dataset for filling in the probability gaps of multimodal and multitemporal surface water in the target area that meets the preset conditions. The generation module is used to fill in the surface water probability gaps in the target dataset using a target deep neural network model to generate a surface water probability map without gaps. The processing module is used to train the target deep neural network model using a target generative adversarial network to generate a trained model, and based on the trained model, to perform target image enhancement processing on the surface water probability map without gaps to generate a multimodal multitemporal surface water probability filling image that meets preset enhancement conditions. The acquisition module includes: The acquisition unit is used to acquire synthetic aperture radar data and surface water probability data in the target area that meet the preset time interval conditions and preset spatial resolution conditions. The processing unit is used to process the synthetic aperture radar data and the surface water probability data to generate a target dataset for filling the multimodal and multitemporal surface water probability gaps in the target area that meets the preset conditions. The step of using a target deep neural network model to fill in the surface water probability gaps in the target dataset to generate a surface water probability map without gaps includes: Based on the multi-branch gated repair module in the target deep neural network model, the target dataset is subjected to preliminary surface water probability gap filling processing to generate a roughly filled surface water probability image. Based on the alignment and refinement module in the target deep neural network model, and guided by the target time synthetic aperture radar data, the roughly filled surface water probability image is refined and aligned to generate the missing surface water probability map.
5. An electronic device, characterized in that, include: The memory, the processor, and the computer program stored in the memory and executable on the processor, the processor executing the program to implement the multimodal multitemporal surface water probabilistic gap filling method as described in any one of claims 1-3.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the multimodal, multitemporal surface water probabilistic gap filling method as described in any one of claims 1-3.
7. A computer program product, comprising a computer program, characterized in that, The computer program is executed by a processor to implement the multimodal, multitemporal surface water probabilistic gap filling method as described in any one of claims 1-3.
Citation Information
Patent Citations
Remote sensing image cloud removal method fusing multi-temporal information and sub-channel dense convolution
CN114511786A
Water body remote sensing image expansion and eutrophication prediction method based on atmosphere-water quality multi-modal information
CN118735752A