Multi-modal multi-temporal surface water probability vacancy filling method and device
Through the multimodal multi-phase surface water probability gap filling method, deep neural networks and generative adversarial networks are used to process surface water probability data, the problem of insufficient non-continuous pixel value data processing and training samples in surface water mapping is solved, and higher mapping accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202411790653.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-06
AI Technical Summary
The prior art is difficult to effectively process the discontinuous pixel value data required for surface water mapping, and the lack of sufficient high-quality training samples, resulting in insufficient generalization capabilities of the model under different environmental conditions, reducing the accuracy and reliability of surface water mapping.
The multimodal multi-phase surface water probability vacancy filling method is adopted. By obtaining the multimodal multi-phase data set of the target area, the deep neural network model is used to fill the surface water probability vacancy, and the adversarial network training model is generated to generate the trained model. Finally, the target image enhancement processing is performed on the surface water probability map without vacancy.
It effectively improves the accuracy and reliability of surface water mapping, solves the problem of insufficient discontinuous pixel value data processing and training samples, and improves the generalization ability of the model under different environmental conditions.
Smart Images

Figure CN119942360A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of intelligent interpretation of remote sensing images, and in particular to a method and device for filling multi-modal and multi-temporal surface water probability gaps. Background Art
[0002] In recent years, deep learning algorithms have been increasingly used in cloud removal tasks. These algorithms can be roughly divided into single-phase and multi-phase methods. In the single-phase method, some researchers effectively use the information of Sentinel-1 to remove clouds from Sentinel-2 images through residual learning. In order to address the domain differences between SAR (Synthetic Aperture Radar) and optical images, other studies have designed a two-stream network to hierarchically integrate SAR features into optical features. In addition, some researchers introduced gated and dilated pyramid convolutions in the coarse-to-fine structure to achieve multi-task reconstruction of remote sensing images. Compared with the CNN (Convolutional Neural Network) structure, ViT (Vision Transformer) has shown strong capabilities in modeling long-distance dependencies in image processing, prompting it to be introduced into the latest cloud removal research.
[0003] Although single-temporal methods are relatively mature, they often overlook the multi-temporal information provided by satellite images, which is crucial for solving the problem of thick cloud occlusion. Among the multi-temporal methods, some researchers have proposed the STGAN (Spatial-Temporal Generative Adversarial Network) method that integrates spatiotemporal information, using multi-branch ResNet or U-Net as the generator and PatchGAN as the discriminator. Other researchers have embedded three-dimensional convolution in a structure similar to STGAN, using optical images and SAR images as input, and achieved convincing results in the "sequence to point" and "sequence to sequence" generation modes. In addition, a study has proposed a method for time matching in the GAN architecture to consider the temporal differences in multi-temporal optical images. Although the existing multi-temporal methods have good performance, they mainly rely on convolutional layers to extract and fuse multimodal information, which limits the sensitivity to long-distance dependencies.
[0004] The de-clouding and surface water probabilistic gap filling algorithms have similarities in input and output: the algorithm input is an optical image or its derivative, containing pixels with continuous values; the algorithm may involve SAR images as auxiliary information; the algorithm output is the gap-filled image. Since cloud interference is a major challenge in high-frequency surface water mapping using remote sensing images, researchers have proposed a variety of deep learning methods based on CNN and Transformer for optical remote sensing image de-clouding. However, these methods are rarely used in surface water mapping because traditional water distribution maps do not have continuous pixel values like the original remote sensing images, making it difficult to process the non-continuous pixel value data required for surface water mapping, and lack of sufficient high-quality training samples, resulting in insufficient generalization ability of the model under different environmental conditions, reducing the accuracy and reliability of surface water mapping, which needs to be solved urgently. Summary of the invention
[0005] This application is based on the following problems and understandings made by the inventor:
[0006] Large-scale, high-frequency, and high-resolution surface water mapping is of great significance in a variety of applications, including water resource protection, government supervision, and flood control and drought relief. Among them, increasing the monitoring frequency is particularly important to support long-term time series analysis and emergency response tasks. Compared with SAR images, optical remote sensing images are widely used in surface water mapping tasks due to their rich spectral information and high precision. However, optical images are easily affected by clouds, cloud shadows, and terrain shadows, resulting in incomplete observations, that is, spatial loss of data. At the same time, the revisit frequency of satellites will also cause temporal loss of surface water distribution maps. These data gaps hinder the continuity and reliability of surface water mapping.
[0007] In recent years, researchers have paid more and more attention to gap filling in water distribution maps. Some methods fill in missing pixels based on the spatial or temporal correlation of optical water distribution maps. Since SAR has a strong ability to penetrate clouds, many existing methods use SAR-based water extraction results as complementary information for optical surface water mapping. In fact, a large number of studies focus on the gap filling problem of different types of remote sensing data, such as surface temperature, soil moisture, and optical remote sensing images (in this case, it is generally called the declouding problem). At present, deep learning methods based on CNN and Transformer have received widespread attention in the declouding task and have achieved reliable results. However, these advanced methods are rarely used to fill gaps in traditional water distribution maps. The main reason is that water distribution maps are usually discontinuous binary images, such as non-water pixel values are 0 and water pixel values are 1, which is not conducive to the learning of deep neural networks.
[0008] In order to more effectively preserve the detailed information in high-resolution remote sensing images, many studies have explored fuzzy mapping techniques, in which the value of each pixel represents the probability that the pixel belongs to a specific category. Compared with traditional water distribution maps, fuzzy water distribution maps as probability maps contain richer information and can quantitatively reflect the uncertainty of water mapping methods. In this context, WP (Water Probability) data are introduced into the surface water mapping task, and its probability values can be converted into hard classification maps by adjusting the threshold. DW (Dynamic World), as a high-precision open source land use and land cover dataset, provides 10-meter resolution WP data with extremely high application value. However, since WP data comes from Sentinel-2 (S2) optical observations, it still faces data integrity issues. Given that the pixel values of WP data are continuously distributed between 0 and 1, which can provide sufficient information for neural networks, the vacancy filling problem of WP can be regarded as a special case of cloud removal tasks and may be solved by similar methods.
[0009] Traditional cloud removal methods use the spectral, temporal, and spatial information of optical remote sensing images to reconstruct areas obscured by clouds. For example, some researchers found that images with clouds have a higher signal-to-noise ratio, so they developed a noise adjustment model to remove thin clouds. Another method combines information from low-resolution images to remove clouds, and improves the cloud removal effect through multiple error correction steps. In addition, some researchers have proposed a gap filling method that does not rely on additional satellite data, using a harmonic model to fill the time series. Although traditional methods do not require a large amount of training data, their effectiveness is limited due to the inability to accurately extract high-level features and fuse multimodal data. Since cloud interference is a major challenge in using remote sensing images for high-frequency surface water mapping, researchers have proposed a variety of deep learning methods based on CNN and Transformer for optical remote sensing image declouding. However, traditional water distribution maps do not have the same continuous pixel values as the original remote sensing images, so they are rarely used in surface water mapping and need to be improved.
[0010] The present application provides a multimodal and multi-temporal surface water probability gap filling method and device to solve the problems in related technologies that it is difficult to process non-continuous pixel value data required for surface water mapping, and there is a lack of sufficient high-quality training samples, resulting in insufficient generalization ability of the model under different environmental conditions, reducing the accuracy and reliability of surface water mapping, and so on.
[0011] The first aspect of the present application provides a method for filling multimodal and multitemporal surface water probability gaps, comprising the following steps: obtaining a target data set for filling multimodal and multitemporal surface water probability gaps in a target area that meets preset conditions; performing surface water probability gap filling processing on the target data set using a target deep neural network model to generate a surface water probability map without gaps; training the target deep neural network model using a target generative adversarial network to generate a trained model, and based on the trained model, performing target image enhancement processing on the surface water probability map without gaps to generate a multimodal and multitemporal surface water probability filling image that meets preset enhancement conditions.
[0012] Optionally, in one embodiment of the present application, the obtaining of a target data set for filling multimodal and multi-temporal surface water probability gaps in a target area that meets preset conditions includes: collecting synthetic aperture radar data and surface water probability data in the target area that meet preset time interval conditions and preset spatial resolution conditions; processing the synthetic aperture radar data and the surface water probability data to generate a target data set for filling multimodal and multi-temporal surface water probability gaps in the target area that meets the preset conditions.
[0013] Optionally, in one embodiment of the present application, the target data set is processed for surface water probability gap filling using the target deep neural network model to generate a surface water probability map without gaps, including: based on the multi-branch gated repair module in the target deep neural network model, preliminary surface water probability gap filling processing is performed on the target data set to generate a surface water probability image with roughly filled gaps; based on the alignment and refinement module in the target deep neural network model, under the guidance of the target time synthetic aperture radar data, the surface water probability image with roughly filled gaps is refined and aligned to generate the surface water probability map with no gaps.
[0014] Optionally, in one embodiment of the present application, the target deep neural network model is trained using a target generative adversarial network to generate a trained model, including: constructing a reconstruction loss function and an adversarial loss function in the target generative adversarial network; combining the reconstruction loss function and the adversarial loss function to obtain a joint loss function of the target deep neural network model; and training the target deep neural network model using the joint loss function to generate the trained model.
[0015] Optionally, in one embodiment of the present application, the joint loss function is:
[0016]
[0017] Among them, λ1 and λ2 are hyperparameters set before training. To reconstruct the loss function, is the adversarial loss function of the generator.
[0018] The second aspect of the present application provides a multimodal and multi-temporal surface water probability gap filling device, including: an acquisition module, used to acquire a target data set for multimodal and multi-temporal surface water probability gap filling in a target area that meets preset conditions; a generation module, used to perform surface water probability gap filling processing on the target data set using a target deep neural network model to generate a surface water probability map without gaps; a processing module, used to train the target deep neural network model using a target generative adversarial network to generate a trained model, and based on the trained model, perform target image enhancement processing on the surface water probability map without gaps to generate a multimodal and multi-temporal surface water probability filling image that meets preset enhancement conditions.
[0019] Optionally, in one embodiment of the present application, the acquisition module includes: an acquisition unit, used to collect synthetic aperture radar data and surface water probability data in the target area that meet preset time interval conditions and preset spatial resolution conditions; a processing unit, used to process the synthetic aperture radar data and the surface water probability data to generate a target data set for filling multi-modal and multi-temporal surface water probability gaps in the target area that meets the preset conditions.
[0020] Optionally, in one embodiment of the present application, the generation module includes: a first generation unit, used to perform preliminary surface water probability gap filling processing on the target data set based on the multi-branch gated repair module in the target deep neural network model to generate a surface water probability image with roughly filled gaps; a second generation unit, used to refine and align the surface water probability image with roughly filled gaps based on the alignment and refinement module in the target deep neural network model and under the guidance of the target time synthetic aperture radar data to generate the surface water probability map without gaps.
[0021] Optionally, in one embodiment of the present application, the processing module includes: a construction unit, used to construct a reconstruction loss function and an adversarial loss function in the target generative adversarial network; an acquisition unit, used to combine the reconstruction loss function and the adversarial loss function to obtain a joint loss function of the target deep neural network model; and a training unit, used to train the target deep neural network model using the joint loss function to generate the trained model.
[0022] Optionally, in one embodiment of the present application, the joint loss function is:
[0023]
[0024] Among them, λ1 and λ2 are hyperparameters set before training. To reconstruct the loss function, is the adversarial loss function of the generator.
[0025] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the multimodal and multi-temporal surface water probability gap filling method as described in the above embodiment.
[0026] The fourth aspect of the present application provides a computer-readable storage medium, which stores a computer program. When the program is executed by a processor, it implements the above-mentioned multi-modal and multi-temporal surface water probability gap filling method.
[0027] The fifth aspect of the present application provides a computer program product, including a computer program, which, when executed, is used to implement the above-mentioned multi-modal and multi-temporal surface water probability gap filling method.
[0028] The embodiment of the present application can obtain a target data set that fills in multimodal and multi-temporal surface water probability gaps that meet certain conditions in the target area, use a target deep neural network model to perform surface water probability gap filling processing on the target data set to generate a surface water probability map without gaps, use a target generative adversarial network to train the target deep neural network model to generate a trained model, and then perform target image enhancement processing on the surface water probability map without gaps to generate a multimodal and multi-temporal surface water probability filling image that meets certain enhancement conditions, effectively improving the accuracy and reliability of surface water mapping. Thus, the problems in the related art that it is difficult to process the non-continuous pixel value data required for surface water mapping, and that there is a lack of sufficient high-quality training samples, resulting in insufficient generalization ability of the model under different environmental conditions, reducing the accuracy and reliability of surface water mapping, and the like are solved.
[0029] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0031] Figure 1 A flowchart of a multi-modal and multi-temporal surface water probability gap filling method provided according to an embodiment of the present application;
[0032] Figure 2A sampling schematic diagram of a method for producing a multi-modal and multi-temporal dataset according to a specific embodiment of the present application;
[0033] Figure 3 A schematic diagram of training and testing sample pairs of a multimodal multi-temporal dataset according to a specific embodiment of the present application;
[0034] Figure 4 A schematic diagram of a gated convolution in a multi-branch gated restoration module according to a specific embodiment of the present application;
[0035] Figure 5 A schematic diagram of a Restormer module in a SAR-guided alignment and refinement module according to a specific embodiment of the present application;
[0036] Figure 6 A neural network structure diagram of a multi-modal and multi-temporal surface water probability gap filling method according to a specific embodiment of the present application;
[0037] Figure 7 A schematic diagram of a multi-branch SN-PatchGAN discriminator according to a specific embodiment of the present application;
[0038] Figure 8 This is a comparison diagram of the effects of a surface water probability gap filling method and a frontier gap filling method according to a specific embodiment of the present application;
[0039] Fig. 9 It is a structural schematic diagram of a multi-modal and multi-temporal surface water probability gap filling device provided according to an embodiment of the present application;
[0040] Fig.10 It is a schematic diagram of the structure of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0041] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0042] The following describes the multimodal and multi-temporal surface water probability gap filling and device of the embodiment of the present application with reference to the accompanying drawings. In view of the problems that the related technologies mentioned in the background technology center are difficult to process the non-continuous pixel value data required for surface water mapping, and lack of sufficient high-quality training samples, which reduces the accuracy and reliability of surface water mapping, the present application provides a multimodal and multi-temporal surface water probability gap filling method, in which a target data set for multimodal and multi-temporal surface water probability gap filling that meets certain conditions in the target area can be obtained, and the target data set is processed for surface water probability gap filling using a target deep neural network model to generate a surface water probability map without gaps, and the target deep neural network model is trained using a target generative adversarial network to generate a trained model, thereby performing target image enhancement processing on the surface water probability map without gaps to generate a multimodal and multi-temporal surface water probability filling image that meets certain enhancement conditions, effectively improving the accuracy and reliability of surface water mapping. This solves the problems in related technologies such as difficulty in processing non-continuous pixel value data required for surface water mapping, lack of sufficient high-quality training samples, and reduced accuracy and reliability of surface water mapping.
[0043] Specifically, Figure 1 A schematic flow chart of a multi-modal and multi-temporal surface water probability gap filling method provided in an embodiment of the present application.
[0044] like Figure 1 As shown, the multi-modal and multi-temporal surface water probability gap filling method includes the following steps:
[0045] In step S101, a target dataset for multi-modal and multi-temporal surface water probability gap filling in a target area that meets preset conditions is obtained.
[0046] In the embodiment of the present application, the target area is the area where the surface water probability gap is filled, and the Chinese area will be used as an example for explanation in this application; the preset condition is to select a certain area of interest in the Chinese area, collect data according to the area of interest, and pre-process the data.
[0047] It can be understood that the embodiments of the present application can obtain a target data set for filling multimodal and multi-temporal surface water probability gaps that meets certain conditions in the target area. For example, the embodiments of the present application can select 20 0.5°×0.5° areas of interest in the Chinese region, covering more than 60,000 square kilometers, including time series data of optical and SAR modes, and collect data according to the areas of interest, and pre-process the data, which will be specifically explained in the following steps, so as to obtain a target data set for filling multimodal and multi-temporal surface water probability gaps, effectively improving the feasibility of filling surface water probability gaps.
[0048] Among them, in one embodiment of the present application, a target data set for filling multimodal and multi-temporal surface water probability gaps that meets preset conditions in a target area is obtained, including: collecting synthetic aperture radar data and surface water probability data that meet preset time interval conditions and preset spatial resolution conditions in the target area; processing the synthetic aperture radar data and the surface water probability data to generate a target data set for filling multimodal and multi-temporal surface water probability gaps that meets preset conditions in the target area.
[0049] For example, the embodiment of the present application can collect data according to the area of interest, and the specific process is as follows: DW (DynamicWorld) is a near real-time global land use and land cover dataset, which can be obtained as an open source resource on the GEE (Google Earth Engine) platform. Each probability map in DW is generated by a single Sentinel-2 image through a fully convolutional neural network, including 9 land cover categories, and has high accuracy. In the embodiment of the present application, the "water body" category in the DW data is used as WP (Water Probability, surface water probability), and the Sentinel-1 ground resolution data from the European Space Agency's Copernicus program (hereinafter referred to as "S1") is used as auxiliary data for reconstruction. These two types of data, namely synthetic aperture radar data and surface water probability data, have a spatial resolution of 10 meters and are aligned to the WGS84 coordinate system.
[0050] Among them, all data collection processes in the embodiments of the present application are automatically completed on the GEE platform, and the S1 data needs to be preprocessed before processing, including thermal noise removal, radiation calibration and terrain correction. For each area of interest, data is collected three times a month, with an interval of about 10 days (for example, January 1st to 10th, 11th to 20th, 21st to 31st). In each time interval, in order to merge multi-scene WP or SAR data, the "median" function in the GEE platform is used to calculate the median of all values of each pixel in each area of interest; the WP data value range of DW is [0,1], while the S1 data is limited to [-25,0] and [-32.5,0] in VV polarization and VH polarization, respectively; for vacant pixels, WP and S1 data are assigned a value of 2 in all bands. After the above processing steps, data were collected from 36 time points and all 20 areas of interest in 2023, among which, Figure 2 The water probability data of a ROI (Region of Interest) at a certain time point and the VV band of Sentinel-1 are shown; yellow pixels represent vacant pixels.
[0051] Furthermore, the embodiment of the present application can perform data preprocessing, and the specific process is as follows: In order to adapt the original data for model input, the present application implements multiple preprocessing steps. According to existing research, the valid pixels (i.e., pixel values not equal to 2) in the S1 data are rescaled to the [0,1] range. The data of each region of interest is divided into 9680 non-overlapping 256×256 pixel tiles, each of which contains WP and S1 data at 36 time points.
[0052] The training data set and the test data set are generated by screening paired data. First, we ensure that the S1 data has no gaps, select a time point with no gaps in the WP data as the target time point, and select three time points from other time points as input. The selection of input time follows the following two criteria: (1) the gap coverage of the input data should be between 10% and 90%; (2) the input time should be as close to the target time point as possible. Taking the "close" in criterion (2) as an example, if April 11 to 20 is selected as the target time point, April 1 to 10 and April 21 to 30 are preferred as input time points, followed by March 21 to 31 and May 1 to 10, and so on. Finally, a total of 4635 data tiles were obtained and randomly assigned to the training set and the test set in a ratio of 7:3. The test set was divided into three groups according to the vacancy coverage: 10%-30% (896 tiles), 30%-60% (690 tiles), and 60%-90% (1190 tiles). After the above processing, the three groups of data with different vacancy coverage are as follows: Figure 3 As shown, the left side is a water probability image with yellow vacant pixels, and the right side is a SAR image; Figure 3 Only the VV band is shown, but both VV and VH bands are used in training and testing. During the test phase, image pairs are divided according to the proportion of missing pixels in the most recent input pair (input 1).
[0053] In step S102, the target deep neural network model is used to fill in the surface water probability gaps in the target data set to generate a surface water probability map without gaps.
[0054] It can be understood that the embodiments of the present application can utilize the target deep neural network model to fill in the surface water probability gaps in the target data set. For example, the embodiments of the present application can combine the deep neural networks of CNN and Transformer and adopt a coarse-to-fine strategy to perform multi-modal and multi-temporal surface water probability filling processing, thereby generating a surface water probability map without gaps, effectively improving the accuracy of filling in surface water probability gaps.
[0055] Among them, in one embodiment of the present application, a target deep neural network model is used to perform surface water probability gap filling processing on a target data set to generate a surface water probability map without gaps, including: based on a multi-branch gated repair module in the target deep neural network model, preliminary surface water probability gap filling processing is performed on the target data set to generate a surface water probability image with roughly filled gaps; based on the alignment and refinement module in the target deep neural network model, under the guidance of the target time synthetic aperture radar data, the surface water probability image with roughly filled gaps is refined and aligned to generate a surface water probability map without gaps.
[0056] As a possible implementation method, the embodiment of the present application can aggregate multi-time series information through a multi-branch gated restoration module. The specific process is as follows: Since gated convolution has been widely used in image restoration research, unlike ordinary convolution, gated convolution takes a vacant image and a mask representing the vacant pixels as input, automatically generates a learnable soft mask from the input, and performs pixel-by-pixel and channel-by-channel gating operations on the output of ordinary convolution. The formula is expressed as follows:
[0057] Feature=∑∑W f I
[0058] Mask=∑∑W m I
[0059] O=φ(Feature)⊙σ(Mask)
[0060] Among them, I represents input, O represents output, and W m and W f are two different convolution kernels, φ is an arbitrary activation function, σ refers to the sigmoid function that limits the mask value range to (0,1), ⊙ is pixel-by-pixel multiplication, Feature is the feature, and Mask is the mask.
[0061] Although traditional gated convolution has proven its effectiveness, it is not suitable for image sequences such as multi-temporal remote sensing data due to the differences in masks at different time points. To solve this problem, this application proposes a (Branch-Gated Inpainting Module). The multi-branch gated inpainting module first performs branch feature extraction through multi-layer gated convolution, then aggregates features at different time points using three-dimensional convolution, and finally generates a roughly filled WP image through two-dimensional convolution.
[0062] Among them, Figure 4As shown in the figure, it is the gated convolution structure of each branch in BGIM. In the gated convolution structure of BGIM, multi-layer downsampling, hole convolution and upsampling operations are introduced to ensure a large receptive field and maintain consistent performance. Since each branch has similar learning objectives, weight sharing is used to reduce the number of parameters. Three-dimensional convolution is commonly used in video recognition tasks. It can process time series input and extract features from spatial, channel and time dimensions. Considering the high computational intensity of three-dimensional convolution, only one three-dimensional convolution layer without padding in the time dimension is used to aggregate information at different time points.
[0063] Furthermore, the embodiments of the present application can refine and align the coarse output under the guidance of the target time SAR data: Although the missing pixels are processed by the BGIM module in the above steps, there are still two problems: (1) The time difference between the input time point and the target time point is not taken into account; (2) Details are lost and blurring occurs in the output image. In order to solve these problems, the present application proposes a SAR-guided SARM (Alignment and Refinement Module). In addition to the coarse output of BGIM, SARM also uses the target SAR image, which provides auxiliary information and facilitates alignment with the target time point. In order to gradually merge the coarse output with the auxiliary information and enhance the details in the gap-filled image, such as Figure 5 As shown in Figure 1, SARM introduces the transformer block from Restormer; compared with the traditional ViT, Restormer uses the self-attention mechanism of DConvs (Depthwise Separable Convolutions) to replace the fully connected layer, which is more efficient and easier to train. Restormer's self-attention mechanism can be expressed as:
[0064]
[0065] in, is the output feature; It is a 3×3 depth-wise convolution; It is a 1×1 point-wise convolution; LN refers to layer normalization; are the deformations of Q, K, and V respectively; α is a learnable scaling parameter, and X is the input feature.
[0066] In addition, gated convolution is also introduced in the feedforward network of Restormer to focus on the detail information in the feature map. The feedforward network of Restormer can be expressed as:
[0067]
[0068] Among them, Gating(X) is a gating operation.
[0069] like Figure 6 As shown in Figure 2, in SARM, two data streams (rough result and SAR) are passed layer by layer through a four-layer symmetric encoder-decoder structure. This process can be expressed as:
[0070]
[0071] Among them, F SAR is the feature extracted from the input SAR image, F coarse are the features extracted from the rough result, i represents the order of feature extraction, RB refers to the Restormer module, and Up refers to an upsampling layer.
[0072] Subsequently, a subtraction operation is used to emphasize the phase difference between the rough result and the target SAR data. It should be noted that all upsampling or downsampling layers in SARM are implemented through convolution and pixel-shuffle or pixel-unshuffle operations. Finally, a two-dimensional convolution is applied to generate the residual image. Final filled-in image Obtained through:
[0073] O=O coarse +R
[0074] Among them, O coarse is a rough result, and R is the residual image.
[0075] In step S103, the target deep neural network model is trained using a target generative adversarial network to generate a trained model, and based on the trained model, the surface water probability map without gaps is subjected to target image enhancement processing to generate a multimodal and multi-temporal surface water probability filling image that meets preset enhancement conditions.
[0076] In the embodiment of the present application, the preset enhancement condition is a condition for achieving seamless surface water mapping.
[0077] It can be understood that the embodiments of the present application can utilize a target generative adversarial network to train a target deep neural network model. For example, by integrating a multi-branch SN-PatchGAN as a generative adversarial network training method of the discriminator, the target deep neural network model in the above steps is trained to generate a trained model, and based on the trained model, the target image enhancement processing is performed on the surface water probability map without vacancies to generate a multimodal and multi-phase surface water probability filling image that meets the enhancement conditions, which can effectively deal with cloud interference problems, improve the detail effect of the generated image, and significantly improve the frequency, range and resolution of surface water mapping, providing technical support and new solutions for high-frequency, large-range and high-resolution surface water mapping.
[0078] In one embodiment of the present application, a target deep neural network model is trained using a target generative adversarial network to generate a trained model, including: constructing a reconstruction loss function and an adversarial loss function in the target generative adversarial network; combining the reconstruction loss function and the adversarial loss function to obtain a joint loss function of the target deep neural network model; and training the target deep neural network model using the joint loss function to generate a trained model.
[0079] In some embodiments, Figure 7 As shown, the embodiment of the present application can construct a reconstruction loss function in the target generative adversarial network. First, the reconstruction loss is introduced to evaluate the rough output result and the final output result, and the L1 loss is adopted. The formula is as follows:
[0080]
[0081] Among them, ‖·‖1 is the L1 norm, is the reconstruction loss function, T is the target image (true value).
[0082] Next, the embodiment of the present application constructs an adversarial loss function in the target generative adversarial network. The L1 loss is relatively stable, but it often makes the reconstructed image smooth. Therefore, the embodiment of the present application introduces adversarial loss to refine the output. The traditional adversarial loss faces the problem of instability, and SN-PatchGAN alleviates this problem to a large extent by using spectral normalization. In order to adapt to multi-phase input, SN-PatchGAN can be modified to a branch form similar to BGIM in the above steps. Figure 7As shown in the figure, as a discriminator, the branch SN-PatchGAN includes 5 branch 2D convolutions (weight sharing), 1 3D convolution and 1 2D convolution, all of which are spectrally normalized. The conditional LSGAN (Conditional Least Squares Generative Adversarial Network) is used as the objective function, which can be expressed as:
[0083]
[0084] Among them, G is the generator, D is the discriminator, x refers to the target truth value, z refers to the rough input, y refers to the mask of the three input time points, and p data refers to the actual data distribution, p z refers to the pseudo data distribution, Refers to the mathematical expectation, is the adversarial loss of the generator, is the total loss of the discriminator, is the adversarial loss of the discriminator.
[0085] Furthermore, the embodiments of the present application can combine the reconstruction loss function and the adversarial loss function to obtain a joint loss function, that is, the loss function ultimately used by the target deep neural network model, thereby using the joint loss function to train the target deep neural network model to generate a trained model, thereby effectively improving the accuracy and reliability of filling surface water probability gaps.
[0086] In one embodiment of the present application, the joint loss function is:
[0087]
[0088] Among them, λ1 and λ2 are hyperparameters set before training and can be set to 100 and 1 respectively; is the reconstruction loss function; is the adversarial loss function of the generator.
[0089] For example, the model proposed in this application is compared with other cutting-edge gap filling methods on the data set constructed in this application. At the same time, this application introduces two simple methods for comparison: the surface water probability input with the least vacant pixels ("minimum gap method") and the mosaic result of all surface water probability inputs ("mosaic method"). The "stitching" process includes the following steps: for each pixel, if only one time point has a valid value, take the value as the result; if multiple time points have valid values, take the average value; if there are no valid values at all time points, assign a value of 0.5 to avoid extreme values in subsequent calculations. The vacant pixels in the "minimum gap method" are also set to 0.5. These two simple methods reflect the amount of useful information that the input data can provide and emphasize the difficulty and importance of the gap filling task.
[0090] like Figure 8 As shown, two sample blocks are selected in each vacancy coverage category to visually compare the performance of different methods, where Figure 8 (1)-(6) in the figure represent different image patches in three categories of vacancy coverage; (a) minimum vacancy method; (b) mosaic method; (c) DSen2-CR (deep Sentinel-2 correlation reconstruction); (d) GLF-CR (global-local fusion cloud removal); (e) HS2P (hyperspectral to panchromatic image fusion); (f) STGAN; (g) SEN12-MS-CR-TS (time series-based Sentinel-1 and Sentinel-2 multispectral cloud removal); (h) the method proposed in this application; (i) target true value image; the yellow pixels in (a) and (b) are vacant pixels that cannot be filled; the red boxes mark the areas where there are significant differences.
[0091] Except for sample block (1), the results of the "least gap method" and "mosaic method" methods are significantly different from the target, further emphasizing the necessity of the proposed method. In the 10%-30% category, all deep learning methods perform well, but (g.1) performs poorly due to overfitting. However, the reconstruction results of the model proposed in this application have the smallest deviation from the target in details. For sample blocks (3) and (4), only the model proposed in this application successfully restores the details within the red box. For sample block (5), the results of DSen2-CR and GLF-CR are severely distorted, while the results of STGAN and SEN12MS-CR-TS are also unsatisfactory. Although the "stitching" method shows that the river area in sample block (6) is almost completely polluted, the deep learning method can still effectively fill the missing pixels, and the reconstruction results of the model proposed in this application are almost the same as the target.
[0092] Four commonly used image reconstruction indicators, MAE (Mean Absolute Error), RMSE (Root Mean Squared Error), PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index), were used for accuracy evaluation. The results are shown in Table 1, which is a comparison table of the accuracy of this application and other gap filling methods. The specific Table 1 is as follows:
[0093] Table 1
[0094]
[0095] Among them, the simple methods "minimum vacancy method" and "mosaic method" performed poorly, especially under conditions of severe vacancies. Compared with the previous best performing method (SEN12MS-CR-TS), the model proposed in this application improved the PSNR by +1.6384 and the SSIM by +0.0193 in the category with the most severe vacancies (60%-90%). Under different vacancy coverage rates, the model proposed in this application showed the most consistent performance. The differences in MAE, RMSE, PSNR and SSIM between the 10%-30% and 60%-90% categories were only 0.0005, 0.0002, 0.1739 and 0.0016, respectively. In addition, the model proposed in this application performed best in the 60%-90% category, showing its effective use of multimodal and multi-temporal information.
[0096] According to the multimodal and multi-temporal surface water probability gap filling method proposed in the embodiment of the present application, a target data set for multimodal and multi-temporal surface water probability gap filling that meets certain conditions in the target area can be obtained, and the target data set can be processed for surface water probability gap filling using a target deep neural network model to generate a surface water probability map without gaps. The target deep neural network model is trained using a target generative adversarial network to generate a trained model, thereby performing target image enhancement processing on the surface water probability map without gaps to generate a multimodal and multi-temporal surface water probability filled image that meets certain enhancement conditions, thereby effectively improving the accuracy and reliability of surface water mapping.
[0097] Next, a multi-modal and multi-temporal surface water probability gap filling device proposed according to an embodiment of the present application will be described with reference to the accompanying drawings.
[0098] Fig. 9 It is a block diagram of a multi-modal and multi-temporal surface water probability gap filling device according to an embodiment of the present application.
[0099] like Fig. 9As shown, the multi-modal and multi-temporal surface water probability gap filling device 10 includes: an acquisition module 100, a generation module 200 and a processing module 300.
[0100] Specifically, the acquisition module 100 is used to acquire a target data set for filling multi-modal and multi-temporal surface water probability gaps in a target area that meets preset conditions.
[0101] The generation module 200 is used to use the target deep neural network model to fill the surface water probability gaps in the target data set to generate a surface water probability map without gaps.
[0102] The processing module 300 is used to train the target deep neural network model using the target generative adversarial network to generate a trained model, and based on the trained model, perform target image enhancement processing on the surface water probability map without vacancies to generate a multimodal and multi-temporal surface water probability filling image that meets the preset enhancement conditions.
[0103] Optionally, in one embodiment of the present application, the acquisition module 100 includes: a collection unit and a processing unit.
[0104] The acquisition unit is used to acquire synthetic aperture radar data and surface water probability data in the target area that meet preset time interval conditions and preset spatial resolution conditions.
[0105] The processing unit is used to process the synthetic aperture radar data and the surface water probability data to generate a target data set for filling multi-modal and multi-temporal surface water probability gaps in the target area that meets preset conditions.
[0106] Optionally, in one embodiment of the present application, the generation module 200 includes: a first generation unit and a second generation unit.
[0107] Among them, the first generation unit is used to perform preliminary surface water probability gap filling processing on the target data set based on the multi-branch gated repair module in the target deep neural network model to generate a surface water probability image with roughly filled gaps.
[0108] The second generation unit is used to refine and align the roughly filled-in surface water probability image based on the alignment and refinement module in the target deep neural network model and under the guidance of the target time synthetic aperture radar data to generate a surface water probability map without gaps.
[0109] Optionally, in one embodiment of the present application, the processing module 300 includes: a construction unit, an acquisition unit and a training unit.
[0110] Among them, the construction unit is used to construct the reconstruction loss function and the adversarial loss function in the target generative adversarial network.
[0111] The acquisition unit is used to combine the reconstruction loss function and the adversarial loss function to obtain the joint loss function of the target deep neural network model.
[0112] The training unit is used to train the target deep neural network model using the joint loss function to generate a trained model.
[0113] Optionally, in one embodiment of the present application, the joint loss function is:
[0114]
[0115] Among them, λ1 and λ2 are hyperparameters set before training. To reconstruct the loss function, is the adversarial loss function of the generator.
[0116] It should be noted that the aforementioned explanation of the embodiment of the multimodal and multi-temporal surface water probability gap filling method is also applicable to the multimodal and multi-temporal surface water probability gap filling device of this embodiment, and will not be repeated here.
[0117] According to the multimodal and multi-temporal surface water probability gap filling device proposed in the embodiment of the present application, a target data set for multimodal and multi-temporal surface water probability gap filling that meets certain conditions in the target area can be obtained, and the target data set can be processed for surface water probability gap filling using a target deep neural network model to generate a surface water probability map without gaps. The target deep neural network model is trained using a target generative adversarial network to generate a trained model, thereby performing target image enhancement processing on the surface water probability map without gaps to generate a multimodal and multi-temporal surface water probability filled image that meets certain enhancement conditions, thereby effectively improving the accuracy and reliability of surface water mapping.
[0118] Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0119] A memory 1001 , a processor 1002 , and a computer program stored in the memory 1001 and executable on the processor 1002 .
[0120] When the processor 1002 executes the program, the multi-modal and multi-temporal surface water probability gap filling method provided in the above embodiment is implemented.
[0121] Furthermore, the electronic device further comprises:
[0122] The communication interface 1003 is used for communication between the memory 1001 and the processor 1002 .
[0123] The memory 1001 is used to store computer programs that can be executed on the processor 1002 .
[0124] The memory 1001 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0125] If the memory 1001, the processor 1002 and the communication interface 1003 are implemented independently, the communication interface 1003, the memory 1001 and the processor 1002 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.10 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0126] Optionally, in a specific implementation, if the memory 1001, the processor 1002 and the communication interface 1003 are integrated on a chip, the memory 1001, the processor 1002 and the communication interface 1003 can communicate with each other through an internal interface.
[0127] The processor 1002 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0128] This embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned multi-modal and multi-temporal surface water probability gap filling method.
[0129] This embodiment also provides a computer program product, including a computer program, which, when executed, is used to implement the above multi-modal and multi-temporal surface water probability gap filling method.
[0130] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0131] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0132] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0133] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or N wirings (electronic devices), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways as necessary and then storing it in a computer memory.
[0134] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above embodiment, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one or a combination of the following technologies known in the art: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0135] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0136] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0137] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A multi-modal and multi-temporal surface water probability gap filling method, characterized in that: The following steps are involved: Obtain a target dataset for filling multi-modal and multi-temporal surface water probability gaps in the target area that meets preset conditions; Using a target deep neural network model to perform surface water probability gap filling processing on the target data set to generate a surface water probability map without gaps; The target deep neural network model is trained using a target generative adversarial network to generate a trained model, and based on the trained model, the surface water probability map without gaps is subjected to target image enhancement processing to generate a multimodal and multi-temporal surface water probability filling image that meets preset enhancement conditions.
2. The method according to claim 1, characterized in that: The method of obtaining a target dataset for filling multi-modal and multi-temporal surface water probability gaps in a target area that meets preset conditions includes: Collecting synthetic aperture radar data and surface water probability data in the target area that meet preset time interval conditions and preset spatial resolution conditions; The synthetic aperture radar data and the surface water probability data are processed to generate a target data set for filling multi-modal and multi-temporal surface water probability gaps in the target area that meets the preset conditions.
3. The method according to claim 1, characterized in that The method of using the target deep neural network model to perform surface water probability gap filling processing on the target data set to generate a surface water probability map without gaps includes: Based on the multi-branch gated restoration module in the target deep neural network model, the target data set is subjected to preliminary surface water probability gap filling processing to generate a roughly gap-filled surface water probability image; Based on the alignment and refinement module in the target deep neural network model, under the guidance of the target time synthetic aperture radar data, the roughly filled-in-gaps surface water probability image is refined and aligned to generate the gap-free surface water probability map.
4. The method according to claim 1, characterized in that The method of using a target generative adversarial network to train the target deep neural network model to generate a trained model includes: Constructing a reconstruction loss function and an adversarial loss function in the target generative adversarial network; Combining the reconstruction loss function and the adversarial loss function to obtain a joint loss function of the target deep neural network model; The target deep neural network model is trained using the joint loss function to generate the trained model.
5. The method according to claim 4, characterized in that The joint loss function is: Among them, λ1 and λ 21 These are all hyperparameters set before training. To reconstruct the loss function, is the adversarial loss function of the generator.
6. A multi-modal and multi-temporal surface water probability gap filling device, characterized in that: include: An acquisition module is used to acquire a target data set for filling multi-modal and multi-temporal surface water probability gaps in a target area that meets preset conditions; A generation module, used for performing surface water probability gap filling processing on the target data set using a target deep neural network model to generate a surface water probability map without gaps; A processing module is used to train the target deep neural network model using a target generative adversarial network to generate a trained model, and based on the trained model, perform target image enhancement processing on the surface water probability map without gaps to generate a multimodal and multi-temporal surface water probability filling image that meets preset enhancement conditions.
7. The device according to claim 6, characterized in that The acquisition module comprises: A collection unit, used for collecting synthetic aperture radar data and surface water probability data in the target area that meet preset time interval conditions and preset spatial resolution conditions; A processing unit is used to perform data processing on the synthetic aperture radar data and the surface water probability data to generate a target data set for filling multi-modal and multi-temporal surface water probability gaps in the target area that meets the preset conditions.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the multimodal and multi-temporal surface water probability gap filling method as described in any one of claims 1 to 5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the multi-modal and multi-temporal surface water probability gap filling method as described in any one of claims 1 to 5.
10. A computer program product, comprising a computer program, characterized in that The computer program is executed by a processor to implement the multi-modal and multi-temporal surface water probability gap filling method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
A method for construct a prediction and assessment model of time series surface water quality big data
CN109472321A
Remote sensing image cloud removal method fusing multi-temporal information and sub-channel dense convolution
CN114511786A
Water body remote sensing image expansion and eutrophication prediction method based on atmosphere-water quality multi-modal information
CN118735752A