A flood inundation area extraction method based on decision-level data fusion

By employing a decision-level data fusion method and utilizing deep learning and decision tree cascade models, remote sensing data from cloudless and cloudy areas are processed separately. This addresses the problem of insufficient extraction accuracy for flooded areas in existing technologies, achieving higher extraction accuracy and completeness.

CN115620162BActive Publication Date: 2025-11-18ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211321653.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-26
Publication Date
2025-11-18
Estimated Expiration
2042-10-26

AI Technical Summary

Technical Problem

Existing technologies for extracting data from flood-inundated areas lack consideration for regional differences in feature-level fusion methods of multispectral data and synthetic aperture radar data, resulting in the loss of some feature information and making it difficult to achieve high-precision extraction.

Method used

A decision-level data fusion approach was adopted, which uses deep learning neural networks to train multi-channel data to construct flood water body extraction models for cloudless and cloudy areas respectively. These models were then cascaded using decision trees and combined with pre- and post-disaster remote sensing images to extract flood-inundated areas.

Benefits of technology

It improved the extraction accuracy in flood-inundated areas, especially the integrity of water body extraction in cloud-covered areas, and significantly improved the average crossover ratio (mIoU). While maintaining the water body extraction accuracy in cloudless areas, it improved the overall extraction effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115620162B_ABST
    Figure CN115620162B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on decision level data fusion's flood submerged area extraction method.The method combines the image band of Sentinel-1, Sentinel-2 in training set, input in depth neural network, and multiple water body extraction models are trained;Then the water body extraction accuracy of each model in test set is calculated, and the model with the highest extraction accuracy for non-water body area without cloud and fog coverage, water body area without cloud and fog coverage and cloud and fog coverage area is selected respectively;Finally, the extraction results of the above three models in disaster area are fused using decision tree, to obtain the water body area before disaster and after disaster, and the flood submerged area of the study area is obtained by subtracting the water body range before disaster from the water body range after disaster.The advantage of the method is that the difference between images with and without cloud is fully considered, and different models are used to extract water bodies in cloudy and non-cloudy areas, and the extraction results are fused using decision tree, which ensures the water extraction accuracy in non-cloudy areas and makes the water extraction in cloudy areas more complete.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of remote sensing disaster detection, specifically involving a method for extracting flood-inundated areas based on decision-level data fusion. Background Technology

[0002] Floods occur rapidly, have a wide impact, and occur frequently, easily causing numerous casualties and severe property damage in affected areas. Efficiently initiating disaster response to flood events, especially accurately and quickly identifying flood-inundated areas, plays a crucial role in flood prevention and disaster reduction. Traditional flood-inundated area extraction is mainly achieved through sampling surveys, a time-consuming and labor-intensive process. In contrast, remote sensing data has a shorter acquisition cycle, is less restricted by ground conditions, and offers rich spectral information, providing the possibility for accurate and rapid extraction of flood-inundated areas. Most studies use multispectral data or synthetic aperture radar (SAR) data to extract flood-inundated areas. Multispectral data provides rich spectral information but is easily affected by clouds and fog; SAR data is less affected by clouds and fog but struggles to distinguish between water bodies and water-like surfaces. Using only multispectral or SAR data makes it difficult to extract high-precision flood-inundated areas. Currently, effectively integrating multispectral and SAR data to improve the utilization rate of feature information has become one of the main development directions for flood-inundated area extraction.

[0003] Currently, most multispectral and synthetic aperture radar (SAR) data fusion is achieved through feature-level data fusion methods. These methods can flexibly integrate feature variables from multispectral and SAR data, effectively improving the extraction accuracy of flood-inundated areas. However, most of these methods use a single model to extract the flood-inundated area of ​​the entire target region, lacking consideration for regional image differences and easily leading to the loss of some feature information. Therefore, this invention attempts to discuss cloudy and cloudless areas separately, using different models to extract the flood-inundated areas of these two regions, improving the utilization rate of feature information and thus improving the overall extraction accuracy of the target region. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method for extracting flood-inundated areas based on decision-level data fusion, so as to improve the utilization rate of feature information and improve the extraction accuracy of flood-inundated areas.

[0005] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows:

[0006] A method for extracting flood-inundated areas based on decision-level data fusion, comprising the following steps:

[0007] S1: Based on the disaster-stricken area to be detected and the time of the flood disaster, acquire pre-disaster and post-disaster remote sensing images of the area, and perform image preprocessing and data alignment.

[0008] S2: Extract different feature variables from cloudless sample sets with flood-inundated areas and combine them into different multi-channel data; use each set of multi-channel data as input to independently train a deep learning neural network, thereby obtaining the flood water body extraction model corresponding to each set of multi-channel data; in the last classification layer of the flood water body extraction model, water body areas and non-water body areas are divided by a fixed preset classification threshold.

[0009] S3: Using a pre-segmented sample set of non-water bodies with and without cloud cover, water bodies without cloud cover, and cloud-covered areas, each flood water body extraction model trained in S2 is tested, and the water body identification results of each flood water body extraction model for the three types of areas are obtained. Then, the following models are selected: the first model with the highest extraction accuracy for non-water bodies without cloud cover and an extraction accuracy of no less than the lower threshold for the other two types of areas; the second model with the highest extraction accuracy for water bodies without cloud cover and an extraction accuracy of no less than the lower threshold for the other two types of areas; and the third model with the highest extraction accuracy for cloud-covered areas and an extraction accuracy of no less than the lower threshold for the other two types of areas.

[0010] S4. A decision tree is used to cascade the first, second, and third models. In the cascaded model, the first model first extracts water body regions from the input image, designating regions with a probability greater than the first classification threshold as the first water body region, and the remaining regions as the first non-water body region. The second model then re-extracts water body regions from the first water body region, designating regions with a probability greater than or equal to the second classification threshold as the second water body region, and the remaining regions as the second non-water body region. Finally, the third model re-extracts water body regions from the second non-water body region, designating regions with a probability greater than or equal to the third classification threshold as the second water body region. The image is divided into three water bodies and three non-water bodies. The water body region is extracted from the input image by summing the water body region and the non-water body region. Based on the cloud sample set, the classification thresholds in the classification layers of the first, second, and third models are optimized in turn. First, the first classification threshold that makes the first model extract the highest accuracy for the first non-water body region is determined. Then, the second classification threshold that makes the second model extract the highest accuracy for the second water body region is determined. Finally, the third classification threshold that makes the third model extract the highest accuracy for the third water body region is determined.

[0011] S5: Using the pre-disaster and post-disaster remote sensing images obtained in S1 as input images, the cascaded model after parameter optimization extracts water bodies from the input images to obtain the water body area extraction results of the input images, thereby determining the pre-disaster and post-disaster water body areas; finally, the area where the post-disaster water body area increases compared to the pre-disaster water body area is identified as the final flood inundation area extraction result.

[0012] Preferably, in step S1, for the disaster-stricken area to be detected, the acquisition and preprocessing of pre-disaster and post-disaster remote sensing images of the target area are completed according to S11 to S13:

[0013] S11: Based on the disaster-stricken area to be detected and the time of the flood disaster, Sentinel-1 and Sentinel-2 images were acquired before and after the disaster, respectively; among them, the Sentinel-1 image is GRD product data in Sentinel-1A interferometric wide-swath mode, and the Sentinel-2 image is L1C level product data.

[0014] S12: Perform orbit file correction, radiometric calibration, and topographic correction on the acquired Sentinel-1 images, and then crop them according to the area of ​​the disaster-stricken area to be detected; perform atmospheric correction and resampling on the acquired Sentinel-2 images, and then crop them according to the area of ​​the disaster-stricken area to be detected.

[0015] S13: The two types of images processed in S12 are segmented separately, and each image is segmented into a series of image blocks.

[0016] Preferably, the Sentinel-2 image is resampled to a 10m resolution, and the image block size is 512×512 pixels.

[0017] Preferably, the pre-disaster Sentinel-1 and Sentinel-2 images and the post-disaster Sentinel-1 and Sentinel-2 images must ensure that the time interval does not exceed the interval threshold and the cloud and fog coverage does not exceed the coverage threshold.

[0018] Preferably, the cloudless sample set uses the Sen1Floods11 dataset.

[0019] Preferably, in step S2, for the disaster-stricken area to be detected, multiple flood water body extraction models are trained according to S21 to S23:

[0020] S21: The VV and VH polarization bands of Sentinel-1 imagery, Band 1 to Band 12 of Sentinel-2 imagery, and the Normalized Difference Water Index (NDWI) and Normalized Multiband Difference Water Index (NDMBWI) calculated based on Sentinel-2 imagery are combined as a set of feature variables; based on the cloudless sample set, the feature variables in the set of feature variables are combined in different ways to synthesize different multichannel data.

[0021] S22: Take each set of multi-channel data in S21 as input and train a UNet++ network independently to obtain the flood water body extraction model corresponding to each set of multi-channel data.

[0022] As a preferred option, when combining the feature variables in the feature variable set, the VV polarization band and VH polarization band of Sentinel-1 image are mandatory channels, and the Band1 to Band12 bands of Sentinel-2 image, as well as the Normalized Difference Water Index (NDWI) and Normalized Multiband Difference Water Index (NDMBWI) calculated based on Sentinel-2 image are optional channels.

[0023] Preferably, in step S3, the extraction precision is measured by recall rate, and the range of recall rate is [0,1].

[0024] Preferably, in step S3, the lower limit of the threshold is set to 0.5 to 0.7.

[0025] Preferably, in step S4, the parameter optimization range of the classification threshold in the classification layers of the first model, the second model, and the third model is [0,1].

[0026] Preferably, in step S4, the extraction accuracy is measured by recall rate.

[0027] Compared with the prior art, the present invention has the following specific advantages:

[0028] This invention fully considers the differences between remote sensing images under cloudy and cloudless conditions. It combines multispectral images and synthetic aperture radar images from the training set into multiple multi-channel data sets according to the needs of cloudy and cloudless conditions. These data are then input into a UNet++ network to train models with the highest accuracy in extracting cloud-covered areas, the highest accuracy in extracting cloudless areas, and the fewest misclassifications in cloudless areas. Next, a decision tree is used to fuse the extraction results of the three models. First, the models with the fewest misclassifications in cloudless areas and the model with the highest accuracy in extracting cloudless areas are used to extract the water body range in cloudless areas and exclude some non-water bodies. Then, the model with the highest accuracy in extracting cloud-covered areas is used to extract the water body range in the remaining areas. Finally, based on the extracted pre-disaster and post-disaster water body ranges, a high-precision flood-inundated area is calculated. Compared to other water body extraction models based on deep learning, this invention has a higher mean intersection over union (mIoU). While maintaining the accuracy of water body extraction in cloudless areas, it significantly improves the completeness of water body extraction in areas with high cloud coverage. This invention provides a reference for other disaster information extraction methods using decision-level data fusion. Attached Figure Description

[0029] Figure 1 Flowchart of a method for extracting flood-inundated areas based on decision-level data fusion;

[0030] Figure 2 A decision-level data fusion method;

[0031] Figure 3 To compare the results of extracting non-cloud-covered areas from the flooded Piura River in Peru on February 27, 2017, using a decision-level data fusion method and a method without decision-level data fusion;

[0032] Figure 4 To compare the results of extracting cloud-covered areas from the flooded Piura River in Peru on February 27, 2017, using a decision-level data fusion method and a method without decision-level data fusion;

[0033] Figure 5 The flood-inundated area of ​​the study area on February 27, 2017, was extracted using a decision-level data fusion method.

[0034] Figure 6 The flooded area of ​​the study area on February 27, 2017, provided by the Copernicus Emergency Response. Detailed Implementation

[0035] The present invention will be further described and illustrated below with reference to the accompanying drawings and specific embodiments.

[0036] like Figure 1The diagram shows a flowchart of a flood inundation area extraction method based on decision-level data fusion, provided in a preferred embodiment of the present invention. The main steps include four steps, S1 to S4:

[0037] S1: Based on the disaster-stricken area to be detected and the time of the flood disaster, acquire pre-disaster and post-disaster remote sensing images of the area, and perform image preprocessing and data alignment.

[0038] S2: Extract different feature variables from cloudless sample sets with flood-inundated areas and combine them into different multi-channel data; use each set of multi-channel data as input to independently train a deep learning neural network, thereby obtaining the flood water body extraction model corresponding to each set of multi-channel data; in the last classification layer of the flood water body extraction model, water body areas and non-water body areas are divided by a fixed preset classification threshold.

[0039] S3: Using a pre-segmented sample set of non-water bodies with and without cloud cover, water bodies without cloud cover, and cloud-covered areas, each flood water body extraction model trained in S2 is tested, and the water body identification results of each flood water body extraction model for the three types of areas are obtained. Then, the following models are selected: the first model with the highest extraction accuracy for non-water bodies without cloud cover and an extraction accuracy of no less than the lower threshold for the other two types of areas; the second model with the highest extraction accuracy for water bodies without cloud cover and an extraction accuracy of no less than the lower threshold for the other two types of areas; and the third model with the highest extraction accuracy for cloud-covered areas and an extraction accuracy of no less than the lower threshold for the other two types of areas.

[0040] S4. A decision tree is used to cascade the first, second, and third models. In the cascaded model, the first model first extracts water body regions from the input image, classifying regions with a probability of belonging to water less than a first classification threshold as first non-water regions, and the remaining regions as first water regions. The second model then re-extracts water body regions from the first water regions, classifying regions with a probability of belonging to water greater than or equal to a second classification threshold as second water regions, and the remaining regions as second non-water regions. Finally, the third model re-extracts water body regions from the second non-water regions, classifying regions with a probability of belonging to water greater than or equal to a third classification threshold as the second water regions. The three water bodies are designated as the third non-water body region. The sum of the second and third water body regions is used as the water body region extraction result in the input image, and the remaining regions are non-water body regions. Based on the cloud sample set, the classification thresholds in the classification layers of the first, second, and third models are optimized sequentially. First, a first classification threshold is determined to achieve the highest extraction accuracy of the first non-water body region by the first model. Then, a second classification threshold is determined to achieve the highest extraction accuracy of the second water body region by the second model. Finally, a third classification threshold is determined to achieve the highest extraction accuracy of the third water body region by the third model.

[0041] S5: Using the pre-disaster and post-disaster remote sensing images obtained in S1 as input images, the cascaded model after parameter optimization extracts water bodies from the input images to obtain the water body area extraction results of the input images, thereby determining the pre-disaster and post-disaster water body areas; finally, the area where the post-disaster water body area increases compared to the pre-disaster water body area is identified as the final flood inundation area extraction result.

[0042] The specific implementation methods of S1 to S5 in this embodiment and their effects are described in detail below.

[0043] In extracting flood-inundated areas from a target region, the type and quality of pre- and post-disaster imagery are crucial. First, to acquire more surface information, this method uses multispectral imagery, which provides rich spectral information, and synthetic aperture radar (SAR) data, which is unaffected by clouds and fog, to extract flood-inundated areas. Second, to achieve rapid and accurate extraction of flood-inundated areas, this method considers resolution, update cycle, and ease of acquisition, selecting Sentinel-1 and Sentinel-2 imagery to extract the pre- and post-disaster water body extent. However, directly downloaded Sentinel-1 and Sentinel-2 images cannot be used directly for water body extraction; preprocessing is required for both images. Furthermore, to facilitate feature extraction from Sentinel-1 and Sentinel-2 images, size segmentation is necessary after preprocessing.

[0044] Based on the above requirements, the image acquisition and preprocessing process of the target area in this embodiment is mainly achieved through step S1, and the specific method is described in detail below:

[0045] S11: Based on the disaster-stricken area to be detected and the time of the flood disaster, acquire pre-disaster and post-disaster Sentinel-1 and Sentinel-2 images. Among them, the Sentinel-1 image is GRD product data in Sentinel-1A interferometric wide-swath mode, and the Sentinel-2 image is L1C level product data.

[0046] In principle, Sentinel-1 and Sentinel-2 images should be acquired at exactly the same time, both before and after a disaster, to ensure complete alignment of the two types of data. However, since Sentinel-1 and Sentinel-2 satellites do not acquire images of the same area completely synchronously, in practical applications, Sentinel-1 and Sentinel-2 images with imaging times as close as possible should be acquired. Additionally, when there is significant cloud cover in the remote sensing imagery, it can easily lead to inaccurate water body extraction; therefore, the cloud-covered area selected for the Sentinel-2 imagery should not be too large. In practical applications, thresholds can be set to filter time and cloud cover conditions based on actual circumstances. Specifically, both pre-disaster and post-disaster Sentinel-1 and Sentinel-2 images must ensure that the time interval does not exceed the interval threshold, and the cloud cover rate does not exceed the coverage threshold. Specific thresholds can be optimized and adjusted.

[0047] S12: Perform orbital file correction, radiometric calibration, topographic correction, and cropping on the acquired Sentinel-1 imagery; perform atmospheric correction, resampling (all bands are resampled to 10m resolution), and cropping on the acquired Sentinel-2 imagery. It should be noted that the cropping of both Sentinel-1 and Sentinel-2 imagery must be performed according to the area of ​​the disaster-stricken region to be detected, and the cropped areas should be perfectly aligned.

[0048] S13: Select an appropriate segmentation size for segmenting the cropped remote sensing image from S12 to meet the input requirements of the deep learning neural network. To maintain consistency with the Sen1Floods11 training sample set, the method of this invention uses 512×512 pixels as the segmentation size for image patches and employs a conventional grid sampling method to segment the image of the disaster-stricken area to be detected. That is, each Sentinel-1 and Sentinel-2 image of the disaster-stricken area to be detected will be cropped into a series of 512×512 pixel image patches.

[0049] High-quality, labeled training datasets are crucial for target extraction using deep learning methods. This method selects the Sen1Floods11 dataset as the training sample set for the deep learning model. Sen1Floods11 contains 446 manually labeled flood data sets and corresponding Sentinel-1 and Sentinel-2 datasets, providing ample effective information for model training.

[0050] To further train a water extraction model with higher accuracy, it is necessary to select appropriate feature variables from the Sentinel-1 and Sentinel-2 data in the training set to form multi-channel data for input into the neural network.

[0051] First, the VV and VH polarization bands from the Sentinel-1 data were selected. Sentinel-1 can penetrate clouds and is unaffected by fog, providing rich texture information, which is beneficial for water extraction in foggy areas.

[0052] Secondly, 13 bands from Band 1 to Band 12 of the Sentinel-2 data were selected. Sentinel-2 data can provide rich surface information, and under cloudless conditions, Sentinel-2 data is the first choice for constructing water body extraction models.

[0053] Finally, to improve the efficiency and accuracy of model extraction, it is necessary to select appropriate water body indices as feature variables. The selection of water body indices considers both the surface conditions of the target area and the weather conditions during floods. The method of this invention selects the Normalized Difference Water Index (NDWI), which is sensitive to dried-up water bodies and vegetation, and the Normalized Multiband Water Index (NDMBWI), which can reduce the influence of clouds and shadows on water body extraction.

[0054] Based on the above theoretical analysis, the training process of the flood water body extraction model in this embodiment is mainly achieved through step S2, and its specific method is described in detail below:

[0055] S21: The VV and VH polarization bands of Sentinel-1 imagery, Band 1 to Band 12 of Sentinel-2 imagery, and the Normalized Difference Water Index (NDWI) and Normalized Multiband Water Index (NDMBWI) calculated based on Sentinel-2 imagery are collectively used as a set of feature variables. Based on the Sen1Floods11 training set, the feature variables in the set are combined in different ways to synthesize different multi-channel data. Specifically, for each set of feature variables, corresponding data can be extracted from each training sample in the Sen1Floods11 training set to construct multi-channel data.

[0056] S22: According to the pre-set network training parameters, multiple multi-channel data synthesized by different combination methods are input into the deep learning neural network for training. Each set of multi-channel data corresponds to a flood water body extraction model, thus obtaining multiple flood water body extraction models.

[0057] In this embodiment, UNet++ can be selected as the deep learning neural network for training the water extraction model. Based on the actual situation, 30 epochs are set, with 10 iterations per epoch, a fixed learning rate of 0.0005, and a batch size of 2 is fed into the model for training each time.

[0058] It should be noted that when constructing different combinations of feature variables based on the feature variable set, in principle, all different combinations should be exhausted as much as possible to obtain the globally optimal solution. However, considering practical efficiency, the feature variables in the feature variable set can be ranked by importance based on relevant previous research or existing technology reports, and then different combinations of feature variables can be constructed selectively. In this embodiment, based on relevant research, it is recommended that the following principle be adopted when combining the feature variables in the feature variable set: the VV polarization band and VH polarization band of Sentinel-1 image are mandatory channels, and the Band1 to Band12 bands of Sentinel-2 image, as well as the Normalized Difference Water Index (NDWI) and Normalized Multiband Difference Water Index (NDMBWI) calculated based on Sentinel-2 image are optional channels. In other words, subsequent multichannel data must include the VV and VH polarization bands of Sentinel-1 imagery, while the Band 1 to Band 12 bands of Sentinel-2 imagery, as well as the NDWI and NDMBWI indices, may or may not be included. Of course, this principle can also be adjusted as needed; for example, some or all bands of Sentinel-2 imagery can be included in the combination range.

[0059] After step S2 above, several flood water body extraction models have been built and trained. The next step is to select a suitable water body extraction model and complete the construction of a decision-level data fusion method through step S3.

[0060] In recent years, an increasing number of data fusion technologies have been applied to the extraction of water bodies from flood-inundated areas. However, most current data fusion methods often use a single model to extract water bodies during the process of extracting flood-inundated areas, with little consideration for the differences and complementarities between models. This invention attempts to use different models to extract water bodies, making sure that models suitable for cloudless areas extract areas of the target region less affected by clouds and fog, and models suitable for foggy areas extract areas of the target region more affected by clouds and fog. All extraction results are then combined using a decision tree approach. This invention ensures the accuracy of water body extraction in non-foggy areas while maximizing the completeness of water body extraction in foggy areas, thus improving the overall accuracy of water body extraction in the target region.

[0061] Based on the above analysis, in this embodiment, the optimal model selection for different types of regions is achieved through step S3. The specific method is described in detail below:

[0062] S31: Acquire sample data with partial cloud and fog coverage and annotate them to form a cloud-covered sample set. The image format and annotation of each sample in the cloud-covered sample set are the same as those in the cloudless sample set. However, in order to select the best model for different types of areas, it is necessary to pre-divide the images in each sample into three categories: non-water areas with and without cloud and fog coverage, water areas without cloud and fog coverage, and cloud-covered areas. Then, following the same procedure as step S21, combine them into multi-channel data corresponding to different flood water body extraction models.

[0063] S32: Input the multi-channel data obtained in S31 into the corresponding flood water body extraction model trained in S23 to obtain the water body extraction results of each flood water body extraction model. Note that the data types and order of each channel must be consistent with the data input during model training.

[0064] S33: Select suitable models by calculating the accuracy of each model. Specifically, to evaluate the extraction advantages of each model, this method compares the water body identification results of each flood water body extraction model for the three types of regions in the cloudless sample set, and then selects the first model with the highest extraction accuracy for non-water body regions without cloud or fog coverage and an extraction accuracy of no less than the lower threshold for the other two types of regions, the second model with the highest extraction accuracy for water body regions without cloud or fog coverage and an extraction accuracy of no less than the lower threshold for the other two types of regions, and the third model with the highest extraction accuracy for cloud-covered regions and an extraction accuracy of no less than the lower threshold for the other two types of regions.

[0065] It should be noted that the purpose of selecting the three models mentioned above is to integrate the extraction characteristics of different models for different regions, considering the differences and complementarities between models, so as to facilitate subsequent fusion through decision trees. The purpose of setting a lower threshold is that when selecting models, it is not only necessary to consider the extraction accuracy for a certain type of region, but also the balance of extraction accuracy for other regions. If the extraction accuracy for other regions is too low, they should be discarded. In this embodiment, in step S33 above, extraction accuracy can be used with recall as the indicator, with the recall value range being [0,1]. Furthermore, the lower threshold for recall can be set to 0.5 to 0.7. In practical applications, when determining the model with the highest extraction accuracy, the lower threshold should be used for initial screening, and the model with the highest extraction accuracy should be selected from the remaining models after screening.

[0066] After steps S31 to S33, three models with high extraction accuracy for different types of regions are obtained, and their results can be fused using a decision tree. In this embodiment, step S4 is used to fuse the optimal models for different types of regions. The specific method is described in detail below:

[0067] S41: A decision tree is used to cascade the first, second, and third models. In the cascaded model, the first model first extracts water body regions from the input image. Regions with a probability of belonging to water less than the first classification threshold k1 are designated as the first non-water body regions, and the remaining regions are designated as the first water body regions. The first non-water body regions may still contain misclassified water body regions, so the second model re-extracts water body regions from the first water body regions. Regions with a probability of belonging to water greater than or equal to the second classification threshold k2 are designated as the second water body regions, and the remaining regions are designated as the second non-water body regions. The second non-water body regions may still contain misclassified water body regions, so the third model re-extracts water body regions from the second non-water body regions. Regions with a probability of belonging to water greater than or equal to the third classification threshold k3 are designated as the third water body regions, and the remaining regions are designated as the third non-water body regions. Finally, the cascaded model uses the sum of the second and third water body regions as the water body region extraction result in the input image, and the remaining regions are considered non-water body regions. The water body classification process in the cascaded model is as follows: Figure 2 As shown.

[0068] S42: Based on the cloud sample set, the classification thresholds in the classification layers of the first model, the second model, and the third model are optimized sequentially. First, the first classification threshold k1 that makes the first model achieve the highest extraction accuracy for the first non-water body region is determined. Then, the second classification threshold k2 that makes the second model achieve the highest extraction accuracy for the second water body region is determined. Finally, the third classification threshold k3 that makes the third model achieve the highest extraction accuracy for the third water body region is determined.

[0069] In this embodiment, the parameter optimization range for the classification threshold in the classification layers of the first, second, and third models in step S42 is [0,1]. That is, different thresholds are continuously adjusted within this range to determine the classification threshold that achieves the highest extraction accuracy. In this embodiment, the extraction accuracy used in this step is also measured by recall rate.

[0070] Therefore, following the steps S1 to S4 described above, multiple water extraction models have been constructed using the UNet++ network. Decision trees have then been used to fuse the water extraction results of the model with the highest accuracy in extracting cloud-covered areas, the model with the highest accuracy in extracting cloudless areas, and the model with the fewest misclassifications in cloudless areas. The specific steps S5 for extracting flood-inundated areas in disaster-stricken regions are briefly described below:

[0071] S51: The pre-disaster remote sensing images and post-disaster remote sensing images obtained in S1 are used as input images. The cascaded model after parameter optimization extracts water bodies from the input images to obtain the water body area extraction results of the input images, thereby determining the pre-disaster water body area and the post-disaster water body area.

[0072] S42: Subtract the pre-disaster water body area from the post-disaster water body area of ​​the target area to obtain the flood-inundated area of ​​the target area. That is, the final determined flood-inundated area is the area that the post-disaster water body area increases relative to the pre-disaster water body area.

[0073] The following demonstrates the effect of applying the methods described in the above embodiments to specific examples. The specific process is as described above and will not be repeated here; the following mainly shows the specific parameter settings and the achieved results.

[0074] Example

[0075] The following describes the invention in detail using the 2017 floods in the Piura River in Peru as an example. The specific steps are as follows:

[0076] 1) Following step S1, acquire pre- and post-disaster images of the Piura River flood-affected area in Peru in February 2017 from the Google Earth platform. First, the pre-disaster images include pre-disaster Sentinel-1 and pre-disaster Sentinel-2 images. The pre-disaster Sentinel-1 image was captured on February 3, 2017, with a resolution of 10 m / pixel; the pre-disaster Sentinel-2 image was captured on February 16, 2017, with a resolution of 10 m / pixel. The post-disaster images include post-disaster Sentinel-1 and post-disaster Sentinel-2 images. The post-disaster Sentinel-1 image was captured on February 27, 2017, with a resolution of 10 m / pixel; the post-disaster Sentinel-2 image was captured on February 26, 2017, with a resolution of 10 m / pixel.

[0077] Based on image quality and the extent of the Piura River basin in Peru, the study area was selected, with latitude and longitude ranging from 80.312°W to 80.637°W and from 4.908°S to 5.085°S. Preprocessing of pre- and post-disaster images was completed according to step S1, and water body labeling was performed in the study area based on pre- and post-disaster Sentinel-2 images for subsequent analysis. Finally, 512×512 pixels were chosen as the sample segmentation size, and conventional grid sampling was used to segment the pre- and post-disaster images and label data of the study area, facilitating subsequent model extraction of water bodies.

[0078] 2) Based on the surface characteristics of the Piura River basin in Peru, the Normalized Difference Water Index (NDWI), which is sensitive to dried-up water bodies and vegetation, was selected as a feature variable. Considering the significant cloud cover during flood season, the Normalized Multiband Water Index (NDMBWI), which reduces the impact of cloud shadows, was selected as a feature variable. The original bands of NDWI, NDMBWI, and Sentinel-1 and Sentinel-2 were combined to create various multi-channel datasets. Following the aforementioned S2 step, the different multi-channel datasets were input into the UNet++ network for training. Each set of multi-channel data served as input, independently training a UNet++ network to obtain the flood water body extraction model corresponding to each set of multi-channel data. The final classification layer of the flood water body extraction model outputs the probability distribution of whether each pixel belongs to a water body. A fixed preset classification threshold can be used to divide water and non-water regions, effectively converting soft labels into hard labels. In this embodiment, the preset classification threshold in the classification layer of the UNet++ network is set to 0.5 by default.

[0079] In this embodiment, nine sets of feature variable combinations were constructed to form multi-channel data, namely:

[0080] Group 1: Only the VV polarization band and VH polarization band of Sentinel-1 imagery are used, denoted as group rS1.

[0081] The second group consists of the VV and VH polarization bands of Sentinel-1 imagery and the Band 6 band of Sentinel-2 imagery, denoted as S1+Band6 group.

[0082] The third group consists of the VV and VH polarization bands of Sentinel-1 imagery and the Band11 band of Sentinel-2 imagery, denoted as S1+Band11 group.

[0083] The fourth group consists of the VV and VH polarization bands of Sentinel-1 imagery and the Band 12 band of Sentinel-2 imagery, denoted as S1+Band12 group.

[0084] The fifth group consists of the VV and VH polarization bands of Sentinel-1 imagery and all 12 bands of Band 1 to Band 12 of Sentinel-2 imagery, denoted as group S1+S2.

[0085] The sixth group consists of the VV and VH polarization bands of Sentinel-1 images, all 12 bands of Band 1 to Band 12 of Sentinel-2 images, and the NDWI index calculated in Sentinel-2 images, denoted as S1+S2+NDWI group.

[0086] The seventh group consists of the VV and VH polarization bands of Sentinel-1 images, all 12 bands of Band 1 to Band 12 of Sentinel-2 images, and the NDMBWI index calculated in Sentinel-2 images, denoted as the S1+S2+NDMBWI group.

[0087] Group 8: The VV polarization band and VH polarization band of Sentinel-1 imagery, all 12 bands of Band 1 to Band 12 of Sentinel-2 imagery, and the NDWI and NDMBWI indices calculated in Sentinel-2 imagery are denoted as S1+S2+NDMBWI+NDWI group.

[0088] Group 9: All 12 bands from Band 1 to Band 12 of Sentinel-2 imagery are used alone and are designated as Group S2.

[0089] 3) Using the different flood water body extraction models trained in 2), tests were conducted on cloud-covered sample sets pre-segmented with non-water areas covered by clouds and fog (referred to as non-water areas without clouds and fog), water areas without cloud and fog (referred to as flood areas without clouds), and areas covered by clouds and fog (referred to as fog). The water body ranges of the study area before and after the disaster were extracted respectively, and the extraction results were compared with the pre-labeled data to calculate the extraction accuracy of each model. In this embodiment, following the steps in S3, the water body identification results of each flood water body extraction model for the three types of areas were calculated respectively, and then the following models were selected: the first model (model1) with the highest extraction accuracy for non-water areas without cloud and fog and an extraction accuracy of no less than 0.5 of the lower threshold for the other two types of areas; the second model (model2) with the highest extraction accuracy for water areas without cloud and fog and an extraction accuracy of no less than 0.5 of the lower threshold for the other two types of areas; and the third model (model3) with the highest extraction accuracy for fog-covered areas and an extraction accuracy of no less than 0.5 of the lower threshold for the other two types of areas.

[0090] The specific experimental results of some models are listed below, as shown in Table 1.

[0091] Table 1. Model recall rates in different regions

[0092]

[0093] By comparing the water extraction results of existing models, the recall rates of some models in cloudless non-water areas, cloudless flooded areas, and foggy areas were calculated to select the locally optimal model. The results show that the rS1 model performs best in foggy areas. In cloudless areas, the S1+S2+NDMBWI+NDWI model has the highest extraction accuracy, while the S1+Band6 and S1+Band12 models have relatively high accuracy in classifying non-water areas. Considering that the S1+Band6 model performs poorly in the other two areas (below the lower threshold of 0.5), this experiment ultimately selected the rS1 model, S1+Band12 model, and S1+S2+NDMBWI+NDWI model as the third, first, and second models, respectively, for decision-level data fusion.

[0094] 4) Following step S4 above, the first, second, and third models are cascaded using a decision tree. Then, the first, second, and third classification thresholds are determined sequentially by optimizing the classification thresholds in the classification layers of each model. In this embodiment, the rS1 model, S1+Band12 model, and S1+S2+NDMBWI+NDWI model are fused using a decision tree. During the fusion process, the probability thresholds k1 = 0.4, k2 = 0.95, and k3 = 0.5 are obtained in the decision tree.

[0095] 6) Finally, following step S5, the completed cascade model is used to extract the pre-disaster and post-disaster water body ranges of the target study area, and the flood inundation area of ​​the study area is calculated.

[0096] In this embodiment, the effectiveness of decision-level data fusion is analyzed by comparing the accuracy of water body extraction in the study area before and after the fusion. The experimental results are shown in Table 2. To further demonstrate that decision-level data fusion can achieve good extraction results in both cloudless and high-cloud-coverage areas, this experiment compares and analyzes the extraction results of the decision-level data fusion method and the three models involved in the fusion in cloudless and high-cloud-coverage areas. The specific results are as follows: Figure 3 and Figure 4 As shown.

[0097] Table 2 Results of decision-level data fusion

[0098]

[0099] After decision-level data fusion, the IoU of water extraction in the study area was 0.6911, significantly improving the accuracy compared to before fusion. The decision-level data fusion method successfully combined the extraction advantages of the rS1 model, the S1+Band12 model, and the S1+S2+NDMBWI+NDWI model. Specifically, in non-cloud-covered areas, decision-level data fusion effectively reduced the impact of bare soil and shadows on water extraction; in cloud-covered areas, the method yielded more complete water body extraction results.

[0100] Calculations show that the flooded area of ​​the Piura River in Peru on February 27, 2017, extracted using the decision-level data fusion method, was approximately 12,809,591 square meters. Furthermore, to demonstrate that decision-level data fusion is applicable to the extraction of flooded areas in the Piura River flood disaster in Peru, the results of this experimental example are compared with the flooded areas of the Peruvian floods on February 27, 2017, provided in the Copernicus Emergency Response. Using manually delineated areas as ground truth, the Intersection over Union (IoU) was calculated, and the results are as follows: Figure 5 and Figure 6 .

[0101] The results show that in the Piura River region of Peru, the IoU of the Copernicus emergency response is 0.4396, while the IoU of the decision-level data fusion method is 0.4985, indicating that the decision-level data fusion method has higher extraction accuracy.

[0102] From the perspective of detail extraction, the decision-level data fusion method proposed in this paper incorporates feature information from NDMBWI and other models, thus offering advantages in distinguishing between flooded water bodies and shadows. Furthermore, the decision-level data fusion approach combines the strengths of multiple models, demonstrating good extraction results for flood-prone areas and regions with fragmented flood distribution.

[0103] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention. Therefore, all technical solutions obtained through equivalent substitution or transformation fall within the protection scope of the present invention.

Claims

1. A method for extracting flood-inundated areas based on decision-level data fusion, characterized in that... The steps are as follows: S1: Based on the disaster-stricken area to be detected and the time of the flood disaster, acquire pre-disaster and post-disaster remote sensing images of the area, and perform image preprocessing and data alignment. S2: Extract different feature variables from cloudless sample sets with flood-inundated areas and combine them into different multi-channel data; use each set of multi-channel data as input to independently train a deep learning neural network, thereby obtaining the flood water body extraction model corresponding to each set of multi-channel data; in the last classification layer of the flood water body extraction model, water body areas and non-water body areas are divided by a fixed preset classification threshold. S3: Using a pre-segmented sample set of non-water bodies with and without cloud cover, water bodies without cloud cover, and cloud-covered areas, each flood water body extraction model trained in S2 is tested, and the water body identification results of each flood water body extraction model for the three types of areas are obtained. Then, the following models are selected: the first model with the highest extraction accuracy for non-water bodies without cloud cover and an extraction accuracy of no less than the lower threshold for the other two types of areas; the second model with the highest extraction accuracy for water bodies without cloud cover and an extraction accuracy of no less than the lower threshold for the other two types of areas; and the third model with the highest extraction accuracy for cloud-covered areas and an extraction accuracy of no less than the lower threshold for the other two types of areas. S4. Use decision trees to cascade the first model, the second model and the third model. In the cascaded model, the first model first extracts water areas from the input image. Areas with a probability of belonging to water less than the first classification threshold are designated as the first non-water areas, and the remaining areas are designated as the first water areas. The second model then re-extracts the water body region from the first water body region, and the regions with a probability of belonging to water greater than or equal to the second classification threshold are designated as the second water body region, while the remaining regions are designated as the second non-water body region. Then, the third model re-extracts water bodies from the second non-water body region. Regions with a probability of belonging to water greater than or equal to the third classification threshold are designated as the third water body region, and the remaining regions are designated as the third non-water body region. Finally, the sum of the second and third water body regions is used as the water body region extraction result in the input image, and the remaining regions are all non-water body regions. Based on the cloud sample set, the classification thresholds in the classification layers of the first model, the second model, and the third model are optimized sequentially. First, the first classification threshold that makes the first model achieve the highest extraction accuracy for the first non-water body region is determined. Then, the second classification threshold that makes the second model achieve the highest extraction accuracy for the second water body region is determined. Finally, the third classification threshold that makes the third model achieve the highest extraction accuracy for the third water body region is determined. S5: Using the pre-disaster and post-disaster remote sensing images obtained in S1 as input images, the cascaded model after parameter optimization extracts water bodies from the input images to obtain the water body area extraction results of the input images, thereby determining the pre-disaster and post-disaster water body areas; finally, the area where the post-disaster water body area increases compared to the pre-disaster water body area is identified as the final flood inundation area extraction result.

2. The method for extracting flood-inundated areas according to claim 1, characterized in that: In step S1, for the disaster-stricken area to be detected, the acquisition and preprocessing of pre-disaster and post-disaster remote sensing images of the target area are completed according to steps S11 to S13: S11: Based on the disaster-stricken area to be detected and the time of the flood disaster, acquire pre-disaster and post-disaster Sentinel-1 and Sentinel-2 images respectively; among them, the Sentinel-1 image is GRD product data in Sentinel-1A interferometric wide-swath mode, and the Sentinel-2 image is L1C level product data. S12: Perform orbit file correction, radiometric calibration, and topographic correction on the acquired Sentinel-1 images, and then crop them according to the area of ​​the disaster-stricken area to be detected; perform atmospheric correction and resampling on the acquired Sentinel-2 images, and then crop them according to the area of ​​the disaster-stricken area to be detected. S13: The two types of images processed in S12 are segmented separately, and each image is segmented into a series of image blocks.

3. The method for extracting flood-inundated areas as described in claim 2, characterized in that: The Sentinel-2 image was resampled to a 10m resolution, and the image block size was 512×512 pixels.

4. The method for extracting flood-inundated areas as described in claim 2, characterized in that: The pre-disaster Sentinel-1 and Sentinel-2 images, as well as the post-disaster Sentinel-1 and Sentinel-2 images, must ensure that the time interval does not exceed the interval threshold and the cloud and fog coverage does not exceed the coverage threshold.

5. The method for extracting flood-inundated areas according to claim 1, characterized in that: The cloudless sample set uses the Sen1Floods11 dataset.

6. The method for extracting flood-inundated areas according to claim 1, characterized in that: In step S2, for the disaster-stricken area to be detected, multiple flood water body extraction models are trained according to S21 to S23: S21: The VV and VH polarization bands of Sentinel-1 imagery, Band 1 to Band 12 of Sentinel-2 imagery, and the Normalized Difference Water Index (NDWI) and Normalized Multiband Difference Water Index (NDMBWI) calculated based on Sentinel-2 imagery are combined as a set of feature variables; based on the cloudless sample set, the feature variables in the set of feature variables are combined in different ways to synthesize different multichannel data. S22: Take each set of multi-channel data in S21 as input and train a UNet++ network independently to obtain the flood water body extraction model corresponding to each set of multi-channel data.

7. The method for extracting flood-inundated areas according to claim 6, characterized in that: When combining the feature variables in the feature variable set, the VV polarization band and VH polarization band of Sentinel-1 image are mandatory channels, while the Band1 to Band12 bands of Sentinel-2 image, as well as the Normalized Difference Water Index (NDWI) and Normalized Multiband Difference Water Index (NDMBWI) calculated based on Sentinel-2 image are optional channels.

8. The method for extracting flood-inundated areas according to claim 1, characterized in that: In step S3, the extraction precision is measured by recall rate, which has a range of [0,1].

9. The method for extracting flood-inundated areas according to claim 1, characterized in that: In step S3, the lower limit of the threshold is set to 0.5 to 0.

7.

10. The method for extracting flood-inundated areas according to claim 1, characterized in that: In step S4, the parameter optimization range of the classification threshold in the classification layers of the first model, the second model, and the third model is [0,1], and the extraction accuracy is measured by recall rate.

Citation Information

Patent Citations

  • Rapid extraction method for outburst flood inundation range

    CN113191292A

  • Closed Loop Machine Learning for a Ground-Based Air Defense System

    DE102021002194B3