A flood simulation verification method based on multi-source data fusion
By integrating multi-source data fusion methods that combine insurance claims, IoT monitoring, remote sensing imagery, and social media data, the problems of single data and poor timeliness in flood simulation verification are solved, achieving high-precision flood simulation and rapid response.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- EAST CHINA NORMAL UNIV
- Filing Date
- 2026-02-12
- Publication Date
- 2026-04-17
AI Technical Summary
Existing flood simulation and verification methods suffer from limited data, narrow spatial coverage, and poor timeliness, resulting in insufficient accuracy and reliability of simulation results, making it difficult to meet the needs of rapid post-disaster assessment and emergency response.
A multi-source data fusion approach is adopted, including insurance claims, IoT monitoring, remote sensing imagery and social media data. Through data preprocessing, flooding information extraction, basic verification index calculation and iterative optimization of model parameters, a comprehensive verification framework is constructed to improve simulation accuracy and practicality.
It enables multi-scale, comprehensive verification of flood inundation, significantly improving the accuracy and timeliness of simulation results and supporting rapid response in insurance claims and disaster emergency management.
Smart Images

Figure CN121683633B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of flood disaster simulation and risk assessment technology, and specifically relates to a flood simulation verification method that integrates multi-source data such as insurance claims, IoT monitoring, remote sensing images and social media, to improve the accuracy, reliability and practicality of flood inundation simulation. Background Technology
[0002] Floods are among the most frequent and economically devastating natural disasters globally. Rapid and accurate flood inundation simulation is a crucial foundation for disaster early warning, risk assessment, and emergency response. Currently, dynamic flood inundation simulations widely employ hydrodynamic models such as LISFLOOD, HEC-RAS, and FloodMap. However, the accuracy of these models heavily relies on high-precision topographic, rainfall, and underlying surface data, and is also affected by important parameters such as the Manning roughness coefficient, hydraulic conductivity, and viscosity coefficient. Therefore, validating simulation results remains a bottleneck restricting the application of these models.
[0003] Traditional verification methods suffer from the following technical shortcomings: 1) Insufficient data coverage. They often rely on sparsely distributed hydrological station observation data (only about 120,000 hydrological stations nationwide), mostly distributed along major river tributaries, failing to comprehensively reflect the water accumulation situation on complex underlying surfaces such as urban flooding areas and rural areas; 2) Poor timeliness. The compilation and release of station data are often delayed, making it difficult to meet the needs of post-disaster emergency response and rapid assessment; 3) Limited information dimensions. They can only provide water level or flow information, unable to directly verify the inundation range and depth. Therefore, due to the deficiencies of traditional verification, key parameters in existing models (such as the Manning roughness coefficient and hydraulic conductivity coefficient) largely rely on empirical values, resulting in significant uncertainty and directly affecting the reliability of simulation results.
[0004] With the development of big data information technology, multi-source data has provided new ideas for flood simulation verification, but different data sources have their own advantages and disadvantages. For example, remote sensing images can extract large-scale inundation areas, but it is difficult to obtain water depth information and is easily affected by factors such as recurrence cycles and weather conditions; social media data has real-time information, but there are problems such as information redundancy and insufficient authenticity. Insurance claims data, due to commercial verification, has advantages such as accurate location, reliable information, and wide coverage, making it a high-quality source of real disaster data. However, current technologies have not systematically and effectively integrated insurance claims data with other multi-source data to build a complete verification framework for model parameter optimization.
[0005] Therefore, there is an urgent need in this field for a flood simulation and verification method that integrates multi-source data such as insurance claims, IoT monitoring, remote sensing and social media, in order to improve the accuracy and practicality of the model. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a flood simulation verification method based on multi-source data fusion, so as to solve the problems of single data, narrow spatial coverage, poor timeliness and insufficient accuracy in traditional verification methods, and significantly improve the accuracy and practicality of flood numerical simulation.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] A flood simulation verification method based on multi-source data fusion includes the following steps:
[0009] S1: Collect multi-source verification data and basic geographic data, and perform preprocessing; the multi-source verification data includes insurance claims data, IoT monitoring data, remote sensing image data, and social media data;
[0010] S2: Extract and quantify real spatiotemporal inundation information from the preprocessed multi-source verification data to construct a comprehensive verification dataset; the spatiotemporal inundation information includes at least one of inundation range, inundation location and inundation depth;
[0011] S3: Based on the preprocessed basic geographic data, drive the flood inundation numerical model to simulate and obtain flood inundation simulation results;
[0012] S4: Perform multi-source data fusion verification by combining the flood inundation simulation results with the comprehensive verification dataset, specifically including:
[0013] S4-1: For each type of data source in the comprehensive verification dataset, calculate its corresponding basic verification index;
[0014] S4-2: Calculate the weight of each data source based on the number of effective spatial units and their inherent quality coefficients;
[0015] S4-3: Calculate a comprehensive score for each type of data source based on the aforementioned basic verification indicators;
[0016] S4-4: Based on the weights of each data source and the comprehensive score, calculate the overall comprehensive verification index through weighted fusion;
[0017] S4-5: Based on the overall comprehensive verification index, iteratively optimize the parameters of the flood inundation numerical model and re-simulate until the preset accuracy requirements are met;
[0018] S5: Outputs visual verification results, which can be applied to insurance claims assessment and disaster emergency management.
[0019] Furthermore, the inundation range described in step S2 is extracted using a semantic segmentation deep learning model or a thresholding method for water body extraction. When using the thresholding method, to optimize the pixel distribution characteristics of the dual-polarization SAR data, the following formula is used to transform the original dual-polarization data:
[0020]
[0021] In the formula, This refers to the channel data for receiving vertically polarized waves emitted by the satellite and for receiving vertically polarized backscattered signals. The data consists of vertically polarized waves emitted by the satellite, but receiving horizontally polarized backscattered signals. The Otsu algorithm is then used to determine the optimal segmentation threshold for water body extraction from the transformed data.
[0022] Furthermore, the basic verification metrics mentioned in step S4-1 include hit rate, fit statistic, deviation score, and root mean square error; wherein, the formula for calculating the hit rate is:
[0023]
[0024] In the formula, This indicates the simulated flood prediction, representing the inundated area or the number of grid sets. This indicates the actual observed flood inundation area or the number of flood location points.
[0025] The formula for calculating the fitted statistic is:
[0026]
[0027] In the formula, This represents the actual observed inundation area. The inundation area simulated by the model. for and The area of the overlapping portion;
[0028] The formula for calculating the deviation score is:
[0029]
[0030] In the formula, It is the total flooded area simulated by the model. This is the total flooded area actually observed;
[0031] The formula for calculating the root mean square error is:
[0032]
[0033] in, and These represent the predicted water depth and the observed water depth, respectively. It represents the total number of samples.
[0034] Further, in step S4-1, the corresponding basic verification index is calculated. When calculating the basic verification index using insurance claim data, a circular buffer zone with a radius of 50 m is established centered on the insurance claim location point. The maximum value of the simulated water depth within this buffer zone is extracted as the representative value of the simulated water depth at that point, which is used for comparison with the actual claim water depth. The calculation formula is:
[0035]
[0036] In the formula, It is the first Maximum raster value within a buffer, It is the first Each buffer zone Coordinates The simulated water depth raster pixel value at that location.
[0037] Further, step S4-2 involves calculating the weights of each data source, and their weights... The calculation formula is:
[0038]
[0039] In the formula, Total number of data sources For the first Number of valid spatial units for each data source For the first The inherent quality coefficient of each data source.
[0040] Furthermore, the inherent quality coefficient Specifically: IoT monitoring data The value ranges from 0.90 to 1.00; this is for insurance claims data. The value ranges from 0.85 to 0.95 for remote sensing image data. The value ranges from 0.80 to 0.90; social media data. The value range is 0.75-0.85.
[0041] Furthermore, step S4-3 involves calculating a comprehensive score, which is... The calculation formula is:
[0042]
[0043] In the formula, To ensure overall spatial consistency, To normalize the overall error, For balance coefficient, ;
[0044] The overall spatial consistency The calculation formula is:
[0045]
[0046] In the formula, and These represent the hit rate and fit statistic of the data source, respectively. As an indicator variable, when the data source is a point location. When the area is a planar region ;
[0047] The normalized comprehensive error The calculation formula is:
[0048]
[0049] In the formula, and The root mean square error and bias score of the data source. To determine the maximum observed water depth within the study area, As an indicator variable, when the data source provides measured water depth When only the flood range is provided It provides neither water depth nor range. .
[0050] Furthermore, step S4-4 involves calculating the overall comprehensive verification index, which is... The calculation formula is:
[0051]
[0052] In the formula, Total number of data sources For the first The weight of each data source, For the first A comprehensive score from multiple data sources.
[0053] Furthermore, the iterative optimization described in steps S4-5 specifically involves: adjusting the overall comprehensive verification index... If the result is not met, it is compared with a preset accuracy threshold. If the result is not met, the result is fed back to step S3, and the model's Manning roughness coefficient, permeability coefficient, capillary potential energy value, evapotranspiration value, and viscosity coefficient parameters are adjusted. After setting the parameters, the simulation is repeated, and the verification process in step S4 is repeated until... The preset accuracy threshold is reached or exceeded.
[0054] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0055] 1) Innovatively, insurance claims data is introduced as the core verification source. It comes directly from post-disaster on-site verification and has the characteristics of accurate positioning, authentic information and wide coverage. It can effectively make up for the lack of water depth information in traditional site data and remote sensing data, and realize more direct and accurate verification of simulated water depth.
[0056] 2) By integrating multi-source data such as insurance claims, IoT monitoring, remote sensing images, and social media, and assigning differentiated weights, the limitations of a single data source are overcome, a more reliable and comprehensive integrated verification framework is constructed, and multi-scale verification from point to surface is realized, which greatly improves the verification efficiency.
[0057] 3) By combining the verification results of multi-source data fusion, the model parameters are iteratively optimized, significantly improving simulation accuracy. This invention can not only be used for rapid verification and accuracy improvement of flood models, but also provide technical support for accurate insurance pricing and rapid claims settlement, as well as rapid response by emergency management departments, possessing significant practical value and application prospects. Attached Figure Description
[0058] Figure 1 This is a flowchart of the flood simulation verification method based on multi-source data fusion of the present invention;
[0059] Figure 2 This is a schematic diagram of the research area in Embodiment 1 of the present invention;
[0060] Figure 3 This is a spatiotemporal distribution map of rainfall in Embodiment 1 of the present invention. In the figure, (a), (b), (c), and (d) are the rainfall distribution time series interpolated at 5-hour intervals from several rain gauge stations, and (e) is the hourly rainfall time series from the three rain gauge stations with the largest process rainfall.
[0061] Figure 4 This is a distribution map of water depth data in insurance claim records according to Embodiment 1 of the present invention;
[0062] Figure 5 This is a distribution map of water depth monitoring points using an electronic water gauge, as described in Embodiment 1 of the present invention.
[0063] Figure 6 This is a schematic diagram of SAR satellite imagery used for verification of the present invention;
[0064] Figure 7 This is a distribution map of water accumulation points reported on social media in Embodiment 1 of the present invention;
[0065] Figure 8 This is a time series result diagram of flood inundation simulation in Embodiment 1 of the present invention;
[0066] Figure 9This is a schematic diagram illustrating the generation of a circular buffer zone with a radius of 50m at the insurance claim point according to the present invention;
[0067] Figure 10 This is a diagram showing the flood simulation verification results based on multi-source data fusion in Embodiment 1 of the present invention. Detailed Implementation
[0068] The present invention will be further described below with reference to the accompanying drawings, and embodiments of the present invention will be given.
[0069] See appendix Figure 1 This invention discloses a flood simulation verification method based on multi-source data fusion, belonging to the field of flood disaster simulation and risk assessment technology. The method includes the following steps: collecting at least two types of multi-source data from insurance claims, IoT monitoring, remote sensing imagery, and social media; extracting real inundation information after preprocessing to construct a comprehensive verification dataset; simulating the inundation range and water depth using a flood model, calculating the hit rate, fitting statistics, deviation score, and root mean square error index; assigning weights to each verification data source; and obtaining a comprehensive verification index through weighted fusion to quantitatively evaluate the simulation accuracy; iteratively optimizing model parameters based on the comprehensive verification index to improve simulation accuracy; and finally outputting visualized results for application in insurance claims and disaster emergency management. This invention solves the problems of single data sources, narrow coverage, and poor timeliness in traditional verification methods, achieving high-efficiency and comprehensive verification, and significantly improving the reliability and practical value of flood models.
[0070] This invention specifically includes:
[0071] S1: Multi-source data acquisition and preprocessing
[0072] Collect multi-source data for verification, including insurance claims, IoT monitoring, remote sensing imagery, and social media, as well as basic geographic data for simulation, such as digital elevation models (DEMs), land use data, rainfall observations, and drainage networks. Specific requirements for collecting multi-source verification data are as follows:
[0073] 1) Insurance claims data: obtained from the insurance company's database, including fields such as claim location (latitude and longitude coordinates), disaster time, water depth, compensation amount and loss type (vehicle flooding, house flooding, equipment damage, etc.);
[0074] 2) IoT monitoring data: This mainly refers to the real-time collection and analysis of data from different environments and systems using various internet-connected devices and sensors. In flood disaster application scenarios, data can be obtained through devices such as electronic water gauges (measurement accuracy ±1cm) and surveillance cameras deployed by local governments or enterprises, including device location, monitoring time, water depth data, and corresponding image and video data.
[0075] 3) Remote sensing image data: mainly obtained from the China Center for Resources Satellite Data and the European Space Agency's Copernicus Data Center, etc., synthetic aperture radar (SAR) satellite images during and after the disaster, with a resolution of no less than 10m, such as Sentinel-1 and Gaofen-3, to extract the flooding range.
[0076] 4) Social Media Data: Data was obtained from news reports, Baidu, WeChat official accounts, video accounts, Weibo, Douyin, Xiaohongshu, and other platforms using web crawling technology. This includes flood-related text, images, and videos with time and geographic location information. Key search terms included "heavy rain," "flooding," "waterlogging," "water accumulation," "submersion," and other limiting terms related to time and location.
[0077] The preprocessing of the collected multi-source data includes:
[0078] 1) Perform coordinate unification and spatial registration on multi-source verification data and basic geographic data to unify all data under the same geographic coordinate system;
[0079] 2) Perform preprocessing on insurance claims data, such as deleting duplicate records, cleaning data, and geocoding; for example, delete data that lacks detailed geographical locations.
[0080] 3) Perform preprocessing such as removing abnormal monitoring values and geocoding on IoT data from electronic water level gauges, surveillance cameras, etc.
[0081] 4) Perform preprocessing on the acquired multi-source SAR satellite image data, including orbit correction, thermal noise removal, filtering, radiometric calibration, Doppler topography correction, and decibel conversion (logarithmic transformation);
[0082] 5) Perform preprocessing on social media crowdsourced data, including outlier removal, data cleaning, and geocoding. For example, remove data reported before the flooding event, data containing advertisements, data without information on heavy rain and flooding, and data outside the study area. Simultaneously, compare the data with officially reported flood locations, manually check each one, and further remove duplicate, erroneous, outlier, and invalid data.
[0083] 6) Resampling, interpolation, and ASCII format conversion of basic geographic data such as topography and rainfall.
[0084] S2: Information Extraction and Quantification from Multi-Source Validation Data Overwhelming
[0085] From preprocessed multi-source verification data, including insurance claims, IoT monitoring, remote sensing imagery, and social media, we extracted and quantified real spatiotemporal inundation information, including inundation range, location, depth, and duration, to construct a comprehensive verification dataset. The geographic locations of the point dataset were vectorized using the Amap API coordinate point picking tool and ArcGIS.
[0086] 1) Extraction of flood information based on insurance claims data:
[0087] The pre-processed insurance claim points are used as verified real "disaster points". The actual inundation location and water depth information are extracted to form a high-confidence flood insurance claim record water accumulation point dataset.
[0088] 2) Extraction of flooding information based on IoT monitoring data:
[0089] Based on the precise location information of the pre-processed electronic water gauge or surveillance camera, a dataset of real continuous water depth monitoring points at the installation locations is extracted. For surveillance video, image recognition technology (such as the YOLOv5 object detection algorithm) is used to estimate the measured water depth using reference points.
[0090] 3) Extraction of flooding information based on SAR satellite imagery data:
[0091] Thresholding methods (such as Otsu's algorithm and adaptive algorithms) or semantic segmentation deep learning models (such as CNN and DeepLab-V3+) are used to extract water area from preprocessed multi-source, multi-temporal SAR satellite images. Accuracy evaluation is then used to form a dataset of the actual observed inundation extent. Among these, the Otsu's algorithm is a commonly used globally optimal thresholding method. This method searches for a threshold that maximizes the inter-class variance between foreground and background by traversing the image's grayscale values, thus achieving binarized thresholding segmentation. Let the number of pixels and grayscale values of the pre-image be... and grayscale range is The threshold for maximizing the inter-class variance of the foreground and background of the image is calculated as follows:
[0092]
[0093] In the formula, , Foreground and background pixel probabilities; , The average gray level of the foreground and background is used to make the inter-class variance... Maximum segmentation threshold That is, the threshold value we are looking for. The optimal segmentation threshold determined by the Otsu algorithm is:
[0094]
[0095] The Otsu algorithm is suitable for water body extraction in SAR images with abundant water bodies and obvious bimodal characteristics during or after disasters. For bipolar SAR data, to highlight the differences between water and non-water body classes, the natural exponential function bipolar data operation method is used to optimize the pixel distribution characteristics of the original SAR data, making the image pixel grayscale distribution more balanced and more conducive to the automatic threshold extraction of the Otsu algorithm. The calculation formula is as follows:
[0096]
[0097] In the formula, This refers to the channel data for receiving vertically polarized waves emitted by the satellite and for receiving vertically polarized backscattered signals. This refers to channel data that transmits vertically polarized waves from a satellite but receives horizontally polarized backscattered signals.
[0098] To assess the accuracy of water body extraction, sample points can be randomly collected within the study area. Visual interpretation using optical reference images is then used to determine the classification of the sample points, thus judging the accuracy of the extraction results. Accuracy assessment uses the Kappa coefficient and Overall Accuracy (OA) as indicators. OA reflects the proportion of correctly classified pixels in the remote sensing image; a higher OA value indicates better classification. The Kappa coefficient is also widely used for evaluating the accuracy of remote sensing classification. Kappa ranges from 0 to 1; a higher Kappa value indicates higher classification accuracy. Its calculation formula is as follows:
[0099]
[0100]
[0101]
[0102] In the formula, The total number of all samples, The total number of samples correctly classified for this category. The total number of samples that were incorrectly classified into this category. The total number of samples that were incorrectly classified into other categories. The total number of samples that were correctly classified into other categories.
[0103] 4) Extraction of information overwhelmed by social media crowdsourced data:
[0104] Natural language processing is performed on preprocessed social media text, image, and video data to determine the authenticity of disaster descriptions and extract the location points of water accumulation reports, forming a dataset of real flood and water accumulation report points.
[0105] S3: Numerical Simulation and Results Output of Flood Inundation
[0106] Two-dimensional hydrodynamic models (such as FloodMap-HydroInundation 2D) are used to simulate flood inundation.
[0107] 1) Model input: Input the preprocessed basic geographic data, including DEM data, rainfall data, land use data, etc., to drive the two-dimensional hydrodynamic model;
[0108] 2) Parameter settings: Based on empirical parameters, set initial Manning roughness coefficient, permeability coefficient, capillary potential energy value, evapotranspiration value and viscosity coefficient and other key parameter values;
[0109] 3) Simulation output: The simulation yields results such as the inundation range and inundation depth, and generates a flood inundation simulation raster map.
[0110] S4: Flood Simulation Verification Based on Multi-Source Data Fusion
[0111] (1) Calculation of basic verification indicators:
[0112] The simulated inundation results from the flood model in step S3 are spatially overlaid with the validation dataset constructed in step S2. Basic validation metrics are calculated, including Hit Rate (HR), Fit statistic (FAI), Bias score (BS), and Root Mean Square Error (RMSE), to quantitatively assess the consistency and error between the simulation results and actual observations. For different data characteristics, corresponding validation metrics are calculated. For point-based validation data (such as insurance claim points, IoT monitoring points, and social media reporting points), the main metrics calculated are Hit Rate and RMSE. For area-based validation data (such as inundation ranges retrieved from remote sensing images), due to the lack of water depth information, the main metrics calculated are Hit Rate, Fit statistic, and Bias score. The specific calculation methods for these metrics are as follows:
[0113] Hit Rate (HR) is a commonly used metric for evaluating the consistency between simulated flood inundation and actual inundation. HR represents the proportion of actual inundated area or inundated location points that fall within the simulated inundation zone, ranging from 1 to 0. The closer the value is to 1, the more accurate the model prediction. In the model simulation results, if the water depth of a grid cell is greater than 0.02m, that grid cell is defined as inundated. The calculation formula is as follows:
[0114]
[0115] in, This indicates the simulated flood prediction, representing the inundated area or the number of grid sets. This indicates the actual observed flood inundation area or the number of flood location points.
[0116] Fit Statistic (FAI) is commonly used to assess the degree of overlap between the simulated inundation extent and the actual observed inundation extent. Its value ranges from 1 to 0, with values closer to 1 indicating higher consistency. The calculation formula is as follows:
[0117]
[0118] in, This represents the actual observed inundation area. The inundation area simulated by the model. for and The area of the overlapping portion.
[0119] The bias score (BS) measures the presence and degree of bias in the model simulation. A BS of 0 indicates no model bias, while positive and negative scores indicate a tendency to overestimate and underestimate, respectively. The formula is as follows:
[0120]
[0121] in, It is the total flooded area simulated by the model. It is the total flooded area actually observed.
[0122] The root mean square error (RMSE) reflects the overall error level between the predicted water depth and the observed water depth. The smaller the value, the more accurate the model prediction. Its calculation formula is as follows:
[0123]
[0124] in, and These represent the predicted water depth and the observed water depth, respectively. It represents the total number of samples.
[0125] For insurance claim location data, considering that the locations of insurance claim companies or buildings are typically flat structures, a single insurance claim location cannot accurately reflect the actual flooded area. To more reasonably compare with the actual flooded area, an alternative method is proposed: establishing a circular buffer zone with a radius of 50 m centered on the insurance claim location. The simulated maximum water depth within the buffer zone is then extracted and compared with the actual flooded water depth. The specific steps are as follows: First, create a 50 m circular buffer zone for each insurance claim point; then, on the simulated water depth raster layer, extract the raster cell values of all raster cells whose center points fall within the buffer zone, forming the set of simulated water depth values corresponding to that buffer zone; finally, calculate the maximum value of the simulated water depth raster cells covered within each buffer zone. The calculation formula is as follows:
[0126]
[0127] in, It is the first Maximum raster value within a buffer, It is the first Each buffer zone Coordinates The simulated water depth raster cell value at the location, the maximum value function traverses the buffer All raster values within. This simulated water depth represents the local area where the claim settlement point is located and is used for subsequent accuracy and error calculations compared with the actual water depth.
[0128] (2) Calculation of weights for multi-source validation data:
[0129] To integrate different data sources, weights need to be assigned to them based on the quantity and inherent quality of their evidence. The weights of each data source are as follows: The calculation formula is as follows:
[0130]
[0131] in, Total number of data sources For the first The number of valid spatial units (e.g., points, pixels) of each data source. For the first The inherent quality coefficient of each data source is determined based on prior knowledge of its accuracy, reliability, and timeliness, and its value ranges from [0,1]. Typical data sources of various types... The reference values are as follows: IoT monitoring data is measured by professional sensors, with extremely high accuracy, but limited coverage; typically... Values range from 0.90 to 1.00; insurance claims data are accurate due to their commercial underwriting and claims processing, but may be subject to time and space lags. Value 0.85-0.95; Remote sensing image data is affected by clouds, all-weather, medium accuracy, typical. Values range from 0.80 to 0.90; social media data is timely but has high noise levels and low accuracy, typically... Value: 0.75-0.85. Actual The values are determined based on the quality of the data sources obtained from different flood-affected areas. Therefore, the final weights of each data source are... Based on the amount of evidence ( ) and quality coefficient ( The more evidence and the higher its quality, the better the decision will be made jointly. The larger the value, and the final weight of all data sources. The sum is 1.
[0132] (3) Calculation of comprehensive score for multi-source validation data:
[0133] Before weighted fusion, a comprehensive score is calculated for each data source by combining the comprehensive spatial consistency and the normalized comprehensive error. The specific steps are as follows:
[0134] 1) Overall spatial consistency ( ): Used to quantify the consistency of spatial distribution between model simulation and observation, its value combines point location hit rate and area overlap information. The calculation formula is:
[0135]
[0136] in, The hit rate of the data source. For fitting statistics; This is a data source type indicator variable; when the data source is a point location, When the data source is an area, . The value ranges from [0, 1], and a higher value indicates better spatial consistency.
[0137] 2) Normalized composite error ( This is used to comprehensively measure the simulation error of the model and normalize errors of different types for comparison. The calculation formula is:
[0138]
[0139] in, To determine the maximum observed water depth within the study area, The root mean square error of the data source. The deviation score; For data type indicator variables, when the data source provides measured water depth, When the data source only provides the flooding range, For data sources that provide neither water depth nor range, . The value range is [0, 1], and the smaller the value, the lower the error.
[0140] 3) Overall score By weighting the two indicators mentioned above, a comprehensive performance score for the model on a single data source is obtained. The calculation formula is as follows:
[0141]
[0142] in, To ensure overall spatial consistency, This is the normalized overall error; Balance coefficient This is used to adjust the relative importance of spatial consistency and error in the scoring, and is usually set to 0.5. The value range is [0,1], and the higher the value, the better the model's performance on that data source.
[0143] (4) Calculation of integrated verification indicators:
[0144] Based on the weights of the above data sources Overall score The overall comprehensive verification index is obtained by weighted average fusion. :
[0145]
[0146] in, Total number of data sources For the first The weights of each data source ( ), For the first A comprehensive score from multiple data sources. The metrics comprehensively reflect the overall performance of the model across all validation data sources; the closer the value is to 1, the higher the overall accuracy of the model simulation.
[0147] (5) Iterative optimization of model parameters
[0148] The calculated comprehensive verification index The result is compared with a preset accuracy threshold (e.g., 0.7% or 70%). If the target is not met, the result is fed back to the flood simulation step (S3), and key parameters (e.g., Manning roughness coefficient, hydraulic conductivity coefficient, etc.) are adjusted based on sensitivity analysis. The simulation and verification are then repeated, iterating until the accuracy requirements are met.
[0149] S5: Visualization and Application of Validation Results
[0150] The final, verified simulation results and accuracy assessment information will be visualized on a GIS platform, outputting a comparison chart of the simulated and actual flooding. The verified simulation results can be applied to practical scenarios such as accurate insurance pricing and rapid claims processing, flood disaster early warning, and emergency evacuation route optimization.
[0151] Example 1
[0152] Taking the flooding caused by Typhoon "Bamboo Grass" in a coastal city in the summer of 2025 as an example, the method of this invention is used to simulate and verify the flooding event.
[0153] S1: Multi-source data acquisition and preprocessing
[0154] Collect available multi-source data for verification and basic geographic data for simulation related to this flooding event, specifically including:
[0155] 1) Insurance claims verification data: 51 pieces of corporate property insurance claims data were obtained from an insurance company in the city within one week after the typhoon, including fields such as company name, specific address, water level, and data reporting time.
[0156] 2) IoT monitoring verification data: Real-time water level monitoring data of 558 electronic water gauge automatic water level piles installed by an insurance company in low-lying and flood-prone locations in the city during Typhoon "Bamboo Grass" were obtained, including detailed location information of the water level piles, status description (normal or awaiting maintenance), monitoring time and water depth value, etc.
[0157] 3) Remote sensing image verification data: In this embodiment, SAR satellite images of the city during the period of Typhoon "Bamboo Grass" were not obtained. To illustrate the application of remote sensing verification data, Sentinel-1 images of a portion of the city's central urban area during Typhoon "Fireworks" were used as a illustrative reference.
[0158] 4) Social media verification data: 602 pieces of data related to this flood event were collected from platforms such as news reports, WeChat official accounts, video accounts, Douyin, Weibo, and Xiaohongshu through web crawling technology. The data included text, images, and geolocation information.
[0159] 5) Basic Geographic Data: Topographic data uses the original 30m resolution digital elevation model (DEM) from the FABDEM website (https: / / data.bris.ac.uk); meteorological data uses real-time rainfall observation data from 446 meteorological monitoring stations with rainfall records during Typhoon "Bamboo Grass," provided by the Municipal Meteorological Bureau. For example... Figure 2 As shown, the terrain of this coastal city is generally high in the southwest and low in the northeast. The urban area has an elevation of approximately 4.0m–5.8m, while the suburbs have an elevation of approximately 3.6m–4.0m, with a dense river network. The region has a subtropical monsoon climate, with the rainy season concentrated in August and September, and frequent typhoons and torrential rains causing floods. Meteorological monitoring stations are mainly concentrated in the urban area, with relatively fewer stations in the suburbs.
[0160] Preprocess all collected data:
[0161] 1) Coordinate System 1: Convert all data to the GCS_WGS_1984 geographic coordinate system;
[0162] 2) Insurance claims data: After removing duplicate records and processing missing values, a total of 45 valid flood claims records were obtained, and the text addresses were converted into latitude and longitude coordinates through geocoding.
[0163] 3) IoT monitoring data: After removing abnormal monitoring values from the electronic water gauge data, 475 valid real-time water level monitoring data were finally obtained.
[0164] 4) Remote sensing image data: Perform preprocessing such as orbit correction, thermal noise removal, filtering, radiometric calibration, Doppler topographic correction and decibel conversion (logarithmic transformation) to obtain one scene of VV and VH dual-polarized Sentinel-1 image in interferometric wide swath imaging mode (IW) (resolution 10m).
[0165] 5) Social media data: Based on natural language processing methods, false, duplicate and flood-irrelevant content was identified and deleted, while valid flood report data with real geographical locations were retained, totaling 586 entries.
[0166] 6) Basic geographic data: The original 30m DEM raster data was resampled to 50m and cropped, outputting as an ASCII format file; hourly rainfall data from meteorological stations were spatially interpolated to 50m to generate a distributed rainfall ASCII format file. The spatiotemporal distribution of rainfall during the typhoon "Bamboo Grass" rainstorm in this city is shown in the attached figure. Figure 3 As shown in the figure, (a), (b), (c), and (d) are the time series of rainfall distribution at 5-hour intervals, interpolated from 446 rain gauge stations in the city, respectively, and (e) is the time series of hourly rainfall at the three rain gauge stations with the largest total rainfall in the city. It can be seen that a rainstorm occurred in the city before 05:00, with the rainfall intensity at its peak at 05:00, and nearly half of the area experiencing hourly rainfall exceeding 20 mm. The center of the rainstorm traversed the central part of the city from northwest to southeast, with the most concentrated area in the central city center. By 10:00, the rainfall gradually weakened, and the center of the rainstorm shifted northward to the northern part of the city. After 15:00, the hourly rainfall generally fell below 5 mm, and the rainfall process tended to end, consistent with the actual spatiotemporal distribution characteristics of rainfall. The 20-hour cumulative rainfall at the three rain gauge stations with the largest total rainfall in the city, G1, G2, and G3, was approximately 334 mm, 333 mm, and 303 mm, respectively, with the maximum hourly rainfall concentrated at 05:00.
[0167] S2: Information Extraction and Quantification from Multi-Source Validation Data Overwhelming
[0168] Extract actual inundation information (including inundation location, extent, and water depth) from the preprocessed multi-source data to construct a comprehensive validation dataset:
[0169] 1) Extraction of flood information based on insurance claims data:
[0170] Based on the detailed coordinates and water depth of the insurance claim points recorded by the insurance company, the actual flooding locations and flooding depths are extracted to form a high-confidence spatial point set. This embodiment extracts a total of 45 verified "true disaster points" (see attached). Figure 4All claim settlement points have a water depth of no less than 0.05m, with 28 claim settlement points having a water depth of ≥0.10m and 8 claim settlement points having a water depth of ≥0.50m. In terms of spatial distribution, claim settlement points are mainly concentrated in the central urban area of the city, with fewer in the suburbs.
[0171] 2) Extraction of flooding information based on IoT monitoring data:
[0172] In this embodiment, the IoT monitoring data mainly consists of electronic water gauge data. Utilizing the precise locations and continuous monitoring sequences recorded by the electronic water gauges, effective water depth monitoring points are extracted to form a spatial point dataset. This embodiment acquires a total of 475 effective monitoring points (see attached). Figure 5 Of these, 446 monitoring points had a water depth ≥ 0.05m, 363 monitoring points had a water depth ≥ 0.10m, and 42 monitoring points had a water depth ≥ 0.50m. The monitoring points were widely distributed but unevenly spaced, mainly concentrated in the city center (about 300 monitoring points), with relatively fewer in other areas, and no effective monitoring points in some suburbs.
[0173] 3) Extraction of flooding information based on SAR satellite imagery data:
[0174] In this embodiment, SAR satellite imagery of the city during the period of Typhoon "Bamboo Grass" impact was not obtained. To illustrate the application of the remote sensing verification data, see the attached... Figure 6 The image shows Sentinel-1 images of parts of the city's central urban area during Typhoon "Fireworks" ( ). Figure 6 a) and the extracted flooding range ( Figure 6 (b) For illustrative reference. This Sentinel-1 image is dual-polarized data. To highlight the differences between water and non-water bodies and optimize the pixel distribution characteristics of the original data, this method uses a natural exponential function to process the dual-polarized data. This processing makes the image grayscale distribution more balanced, thereby improving the robustness and accuracy of subsequent automatic thresholding segmentation using the Otsu algorithm. The calculation formula is:
[0175]
[0176] in, This refers to the channel data for receiving vertically polarized waves emitted by the satellite and for receiving vertically polarized backscattered signals. This refers to channel data that transmits vertically polarized waves from a satellite but receives horizontally polarized backscattered signals.
[0177] Using the data optimized by the aforementioned natural exponential function, the Otsu thresholding method is then employed for water body extraction. For example... Figure 6As shown in b, the extracted water body extent matches the actual water body area in the image with a high degree of consistency. Accuracy verification was performed through visual interpretation sampling, yielding an overall classification accuracy (OA) of 99.41% and a Kappa coefficient of 0.988. This indicates that the extraction results have high reliability and can therefore be used as a dataset of real-world observed inundation extent.
[0178] 4) Information extraction based on the flood of social media data:
[0179] By performing natural language processing on text from location-based social media platforms (WeChat, Weibo, Douyin, Xiaohongshu, WeChat Video Accounts, etc.), 586 valid reported water accumulation points were ultimately extracted (see attached). Figure 7 Its spatial distribution is relatively wide, mainly concentrated in the central urban area and northwest region of the city, with a small number also distributed in other areas, while some suburban areas have no effectively reported waterlogged points.
[0180] S3: Numerical Simulation and Results Output of Rainstorm Flooding
[0181] This embodiment uses the FloodMap-HydroInundation 2D two-dimensional hydrodynamic model to simulate the torrential rain and flooding process caused by Typhoon "Bamboo Grass" in this city in 2025. The entire simulation time is 20 hours to allow the flooding process to reach a stable state. For the torrential rain and flooding in the city center during Typhoon "Fireworks" in 2021, the simulation time is set to 24 hours. This model is simple to construct, has high computational efficiency, and combines hydrological processes such as evaporation, infiltration, and drainage with flooding. It can be used for torrential rain and flood simulation in urban environments and is now widely used in urban hydrological and hydrodynamic process research. Among them, the infiltration process is calculated using the widely used Green-Ampt equation, and the evaporation is estimated based on the empirical sine curve formula (approximately 3 mm / d) from previous studies. The simulation of surface flood evolution is based on the Saint-Venant equations to describe unsteady shallow water waves. It adopts a similar structure to the LISFLOOD-FP model, but uses a different method to calculate the time step. The Forward Courant-Freidrich-Levy (FCFL) method is used to dynamically select an appropriate time step to maintain model stability and minimize numerical diffusion, while simplifying the kinetic energy condition (convective acceleration term) of the water flow. The main governing equations of this model on a regular grid are expressed as follows:
[0182]
[0183] In the formula, In time Traffic, For time step, In time Traffic, It is the acceleration due to gravity. In time The water depth, This is the elevation of the bottom of the grid. This is the Manning coefficient.
[0184] In the model, the surface runoff loss of the urban stormwater drainage system is approximated by scaling the drainage capacity (mm / h) at each time step, assuming that the stormwater drainage system drains and / or pumps at maximum design capacity. Distributed drainage capacity can also be represented by the model at each basic unit. The design stormwater return period for the drainage facilities in the urban area of this city is 3 years. Since the extreme storm tides induced by Typhoon "Bamboo Grass" in 2025 and Typhoon "Fireworks" in 2021 approached or even exceeded the urban land elevation, the drainage system could not operate normally and discharge rainwater into surrounding rivers; therefore, the urban stormwater drainage capacity of this city was not considered in the simulations. Furthermore, the initial parameter settings for the model simulation in this city all adopted empirical coefficients calibrated in previous studies, namely a soil saturated hydraulic conductivity of 0.001 m / h and a Manning coefficient of 0.20.
[0185] As attached Figure 8 The flood simulation results during Typhoon "Pigsy" show that: at the initial stage of rainfall at 05:00, the overall flooding in the city was relatively light, with an average flood depth of about 0.05m and a total flooded area of 154.52km² (water depth ≥0.50m); by 10:00, as rainfall accumulated, the flooded area expanded significantly, with the area of water depth ≥0.02m reaching 4102.08km², the area of water depth ≥0.50m reaching approximately 469.08km², and the average flood depth rising to about 0.12m; by 15:00 and thereafter, the weakening rainfall caused the accumulated water to flow to lower elevations. The water level in low-lying areas rose further to an average depth of 0.15m, with some areas exceeding 2.0m. During this period, the area with a water depth ≥0.02m covered 3372.07 km², of which 561.95 km² were deep water areas ≥0.50m. By 20:00, the average water depth reached its peak at approximately 0.21m, with the area exceeding 0.02m covering 3151.33 km², including 595.87 km² of deep water areas ≥0.50m. Although infiltration reduced the affected area in some areas, the overall flooding remained severe. The flooding was widespread, with the central urban area and northwest regions being more severely affected, while the southern regions were relatively less flooded.
[0186] S4: Flood Simulation Verification Based on Multi-Source Data Fusion
[0187] (1) Calculation of basic verification indicators:
[0188] The simulated inundation results from the flood model in step S3 are spatially overlaid with the validation dataset constructed in step S2. Basic validation metrics are calculated, including Hit Rate (HR), Fit statistic (FAI), Bias score (BS), and Root Mean Square Error (RMSE), to quantitatively assess the consistency between the simulation results and actual observations. Different validation metrics are calculated for different data characteristics. For point-based validation data (such as insurance claim points, IoT monitoring points, and social media reporting points), the main metrics calculated are Hit Rate and RMSE. For area-based validation data (such as inundation areas retrieved from remote sensing images), due to the lack of water depth information, the main metrics calculated are Hit Rate, Fit statistic, and Bias score. The specific metric calculations are as follows:
[0189] Hit Rate (HR) is a commonly used metric for evaluating the consistency between simulated flood inundation and actual inundation. HR represents the proportion of actual inundated area or inundated location points that fall within the simulated inundation zone, ranging from 1 to 0. The closer the value is to 1, the more accurate the model prediction. In the model simulation results, if the water depth of a grid cell is greater than 0.02m, that grid cell is defined as inundated. The calculation formula is as follows:
[0190]
[0191] in, This indicates the simulated flood prediction, representing the inundated area or the number of grid sets. This indicates the actual observed flood inundation area or the number of flood location points.
[0192] Fit Statistic (FAI) is commonly used to assess the degree of overlap between the simulated inundation extent and the actual observed inundation extent. Its value ranges from 1 to 0, with values closer to 1 indicating higher consistency. The calculation formula is as follows:
[0193]
[0194] in, This represents the actual observed inundation area. The inundation area simulated by the model. for and The area of the overlapping portion.
[0195] The bias score (BS) measures the presence and degree of bias in the model simulation. A BS of 0 indicates no model bias, while positive and negative scores indicate a tendency to overestimate and underestimate, respectively. The formula is as follows:
[0196]
[0197] in, It is the total flooded area simulated by the model. It is the total flooded area actually observed.
[0198] The root mean square error (RMSE) reflects the overall error level between the predicted water depth and the observed water depth. The smaller the value, the more accurate the model prediction. Its calculation formula is as follows:
[0199]
[0200] in, and These represent the predicted water depth and the observed water depth, respectively. It represents the total number of samples.
[0201] Regarding insurance claim location data, considering that the locations of insurance claim companies or buildings are typically flat structures, a single insurance claim location cannot accurately reflect the actual flooded area. Therefore, to more reasonably compare with the actual flooded area, an alternative method is proposed: establishing a circular buffer zone with a 50m radius centered on the insurance claim location. (See appendix) Figure 9 The simulation of the maximum water depth within the buffer zone is compared with the actual flooding depth. The specific steps are as follows: First, a 50m circular buffer zone is created for each insurance claim point; then, on the simulated water depth raster layer, the raster cell values whose center points fall within the buffer zone are extracted to form the set of simulated water depth values corresponding to that buffer zone; finally, the maximum value of the simulated water depth raster cells covered within each buffer zone is calculated. The calculation formula is as follows:
[0202]
[0203] in, It is the first Maximum raster value within a buffer, It is the first Each buffer zone Coordinates The simulated water depth raster cell value at the location, the maximum value function traverses the buffer All raster values within. This simulated water depth represents the local area where the claim settlement point is located and is used for subsequent accuracy and error calculations compared with the actual water depth.
[0204] Based on the above method, the basic verification index is calculated for the submerged information features extracted from each data source: In Example 1, the hit rate of direct verification using a single insurance claim point. It is 77.78%. The accuracy is 0.36m; if a circular buffer zone with a radius of 50m is established around the claim point for verification, the hit rate will be... Increased to 95.56%, The value is 0.37m, therefore the result calculated using this method is used as the final basic verification index value for insurance claims data. The hit rate of verification using electronic water level gauge monitoring points is also considered. It was 88.56%. The accuracy is 0.36m. A portion of the remote sensing imagery was selected as a reference area for flood simulation verification, and the calculated hit rate was... It is 73.65%. It is 0.57. The hit rate was 0.021. This represents the accuracy of point-of-sale verification using social media reports. The RMSE is 89.25%, but since the data source does not contain water depth information, it is impossible to calculate the RMSE.
[0205] (2) Calculation of weights for multi-source validation data:
[0206] To integrate different data sources, weights need to be assigned to them based on the quantity and inherent quality of their evidence. The weights of each data source are as follows: The calculation formula is as follows:
[0207]
[0208] in, Total number of data sources For the first The number of valid spatial units (e.g., points, pixels) of each data source. For the first The inherent quality coefficient of each data source is determined based on prior knowledge of its accuracy, reliability, and timeliness, and its value ranges from [0,1]. Typical data sources of various types... The reference values are as follows: IoT monitoring data is measured by professional sensors, with extremely high accuracy, but limited coverage; typically... Values range from 0.90 to 1.00; insurance claims data are accurate due to their commercial underwriting and claims processing, but may be subject to time and space lags. Value 0.85-0.95; Remote sensing image data is affected by clouds, all-weather, medium accuracy, typical. Values range from 0.80 to 0.90; social media data is timely but has high noise levels and low accuracy, typically... The values are 0.75-0.85. Therefore, the final weights of each data source are... Based on the amount of evidence ( ) and quality coefficient ( The more evidence and the higher its quality, the better the decision will be made jointly. The larger the value, and the final weight of all data sources. The sum is 1.
[0209] In this embodiment, the parameters of each data source are as follows: Insurance claim points Number of IoT monitoring points Remote sensing image inversion of water body raster pixel count Social media points To comprehensively assess the robustness of the weight allocation, different data sources were used. Calculate using the lowest and highest values within the range. 1) Use Calculation of the lowest value within the range: Insurance claims data quality coefficient IoT monitoring data quality coefficient Remote sensing image data quality coefficient Social media data quality index The weights of the insurance claims data are calculated according to the formula as follows: IoT monitoring data weighting weighting of remote sensing image data Social media data weighting 2) Use Calculation of the highest value within the range: Insurance claims data quality coefficient IoT monitoring data quality coefficient Remote sensing image data quality coefficient Social media data quality index The weights of the insurance claims data are calculated according to the formula as follows: IoT monitoring data weighting weighting of remote sensing image data Social media data weighting .
[0210] (3) Calculation of comprehensive score for multi-source validation data:
[0211] Before weighted fusion, a comprehensive score is calculated for each data source by combining the comprehensive spatial consistency and the normalized comprehensive error. The specific steps are as follows:
[0212] 1) Overall spatial consistency ( ): Used to quantify the consistency of spatial distribution between model simulation and observation, its value combines point location hit rate and area overlap information. The calculation formula is:
[0213]
[0214] in, The hit rate of the data source. For fitting statistics; This is a data source type indicator variable; when the data source is a point location, When the data source is an area, . The value ranges from [0, 1], with higher values indicating better spatial consistency. In this embodiment, the calculation result is: the overall spatial consistency of insurance claims data. Spatial consistency of IoT monitoring data Spatial consistency of remote sensing image data Social media data overall spatial consistency .
[0215] 2) Normalized composite error ( This is used to comprehensively measure the simulation error of the model and normalize errors of different types for comparison. The calculation formula is:
[0216]
[0217] in, To determine the maximum observed water depth within the study area, The root mean square error of the data source. The deviation score; For data type indicator variables, when the data source provides measured water depth, When the data source only provides the flooding range, For data sources that provide neither water depth nor range, . The value ranges from [0, 1], with smaller values indicating lower errors. In this embodiment, the maximum water depth observed in the study area is... The calculation result is: Comprehensive error of insurance claims data. Overall error of IoT monitoring data Comprehensive error of remote sensing image data Social media data aggregation error .
[0218] 3) Overall score By weighting the two indicators mentioned above, a comprehensive performance score for the model on a single data source is obtained. The calculation formula is as follows:
[0219]
[0220] in, To ensure overall spatial consistency, This is the normalized overall error; Balance coefficient . The value range is [0,1], and a higher value indicates better model performance on that data source. In this embodiment, a balance coefficient is used. The calculation result is: Comprehensive score of insurance claims data. Comprehensive scoring of IoT monitoring data Comprehensive scoring of remote sensing image data Social media data comprehensive score .
[0221] (4) Calculation of integrated verification indicators:
[0222] Based on the weights of the above data sources Overall score The overall comprehensive verification index is obtained by weighted average fusion. :
[0223]
[0224] in, Total number of data sources For the first The weights of each data source ( ), For the first A comprehensive score from multiple data sources. The metric comprehensively reflects the model's overall performance across all validation data sources; the closer the value is to 1, the higher the overall accuracy of the model simulation. This embodiment uses various data sources... The calculation is performed using the lowest and highest values within the range of values. Values and data sources The values are used to calculate the final overall comprehensive verification index. The result is: 1) When using The weight is calculated based on the lowest value within the range of values. ;2) When using The weight is calculated based on the highest value within the range. Under two weighting values, the overall comprehensive verification index The calculation results show very little difference, indicating that the weight allocation has good robustness within this value range, and the overall accuracy of the model simulation is consistent and close to 87.67%. This value is significantly higher than the preset reliability threshold (70%), indicating that the overall credibility of this flood inundation simulation results is high, and the verification is successful. Spatial analysis further demonstrates the reliability of this comprehensive verification result; the severely inundated areas predicted by the model (such as the central urban area) are highly consistent with the spatial distribution of actual water accumulation points reflected by multi-source data (see attached). Figure 10 This demonstrates that the multi-source data fusion verification method employed in this invention yields more reliable and robust evaluation results compared to traditional methods that rely on a single data source.
[0225] (5) Iterative optimization of model parameters
[0226] The calculated comprehensive verification index The result is compared with a preset accuracy threshold (70%). If the threshold is not met, the result is fed back to flood simulation step S3, and key parameters (such as Manning roughness coefficient, hydraulic conductivity coefficient, etc.) are adjusted based on sensitivity analysis. The simulation and verification are then repeated iteratively until the accuracy requirements are met. In this embodiment, due to the simulation calculation... (87.67%) is already above the threshold, so there is no need to iterate again and the simulation result that has passed the verification can be output directly.
[0227] S5: Visualization and Application of Validation Results
[0228] The verified high-precision flood simulation inundation results (range and depth), multi-source verification comparison information, and comprehensive evaluation indicators will be used to generate thematic maps and comprehensive verification reports on a GIS platform for visualization and application to government departments and insurance companies. For example, flood simulation results can be imported into insurance claims systems to assist in loss assessment and rapid compensation; flood early warning information can be issued, and emergency evacuation routes can be planned.
[0229] This embodiment fully demonstrates the effectiveness and practicality of the method of the present invention. By integrating multi-source data such as insurance claims, IoT monitoring, remote sensing imagery, and social media, a comprehensive verification framework is constructed, and the model parameters are iteratively optimized based on weighted fusion comprehensive indicators, significantly improving the accuracy and reliability of flood simulation. This method is not only suitable for the verification and optimization of flood models, but can also be widely applied to fields such as insurance pricing, claims assessment, and disaster emergency management.
Claims
1. A flood simulation verification method based on multi-source data fusion, characterized in that, Includes the following steps: S1: Collect multi-source verification data and basic geographic data, and perform preprocessing; the multi-source verification data includes insurance claims data, IoT monitoring data, remote sensing image data, and social media data; S2: Extract and quantify real spatiotemporal inundation information from the preprocessed multi-source verification data to construct a comprehensive verification dataset; the spatiotemporal inundation information includes at least one of inundation range, inundation location and inundation depth; S3: Based on the preprocessed basic geographic data, drive the flood inundation numerical model to simulate and obtain the flood inundation simulation results; S4: Perform multi-source data fusion verification by combining the flood inundation simulation results with the comprehensive verification dataset, specifically including: S4-1: For each type of data source in the comprehensive verification dataset, calculate its corresponding basic verification index; S4-2: Calculate the weight of each data source based on the number of effective spatial units and their inherent quality coefficients; S4-3: Calculate a comprehensive score for each type of data source based on the aforementioned basic verification indicators; S4-4: Based on the weights of each data source and the comprehensive score, calculate the overall comprehensive verification index through weighted fusion; S4-5: Based on the overall comprehensive verification index, iteratively optimize the parameters of the flood inundation numerical model and re-simulate until the preset accuracy requirements are met; S5: Outputs visual verification results, which can be applied to insurance claims assessment and disaster emergency management.
2. The flood simulation verification method of claim 1, wherein, The inundation range described in step S2 is extracted using a semantic segmentation deep learning model or a thresholding method. When using the thresholding method, to optimize the pixel distribution characteristics of the dual-polarization SAR data, the following formula is used to transform the original dual-polarization data: In the formula, This refers to the channel data for receiving vertically polarized waves emitted by the satellite and for receiving vertically polarized backscattered signals. The data consists of vertically polarized waves emitted by the satellite, but receiving horizontally polarized backscattered signals. The Otsu algorithm is then used to determine the optimal segmentation threshold for water body extraction. 3.The flood simulation verification method of claim 1, wherein, The basic validation metrics mentioned in step S4-1 include hit rate, fit statistic, bias score, and root mean square error; wherein, the formula for calculating the hit rate is: wherein, represents the number of inundated areas or grid sets predicted as flood by simulation, represents the number of flood inundated areas or flood location point sets actually observed; The formula for calculating the fitted statistic is: In the formula, This represents the actual observed inundation area. The inundation area simulated by the model. for and The area of the overlapping portion; The formula for calculating the deviation score is: wherein, is the total inundated area modeled, is the total inundated area actually observed; The formula for calculating the root mean square error is: where, and represent the predicted and observed water depths, respectively; is the total number of samples.
4. The flood simulation verification method of claim 1, wherein, Step S4-1 involves calculating the corresponding basic verification index. When calculating the basic verification index using insurance claim data, a circular buffer zone with a radius of 50 m is established centered on the insurance claim location point. The maximum value of the simulated water depth within this buffer zone is extracted as the representative value of the simulated water depth at that point, used for comparison with the actual claim water depth. The calculation formula is as follows: where, is the maximum grid value in the buffer, is the buffer area in the buffer, is the simulated water depth grid cell value at the coordinates.
5. The flood simulation verification method of claim 1, wherein, The weight of each data source is calculated in step S4-2, and the weight of each data source is calculated according to the following formula: The calculation formula is as follows: In the formula, Total number of data sources For the first Number of valid spatial units for each data source For the first The inherent quality coefficient of each data source.
6. The flood simulation verification method of claim 5, wherein, The inherent mass coefficient Specifically: IoT monitoring data The value ranges from 0.90 to 1.00; this is for insurance claims data. The value ranges from 0.85 to 0.95 for remote sensing image data. The value ranges from 0.80 to 0.90; social media data. The value range is 0.75-0.
85.
7. The flood simulation verification method of claim 1, wherein, The calculation of the comprehensive score in step S4-3 is described as follows: the comprehensive score of the user's interest in the product is calculated according to the following formula: The calculation formula of the comprehensive score is as follows: wherein, is the comprehensive spatial consistency, is the normalized comprehensive error, is the balancing coefficient, ; The comprehensive spatial consistency degree The calculation formula is: In the formula, and These represent the hit rate and fit statistic of the data source, respectively. As an indicator variable, when the data source is a point location. When the area is a planar region ; The normalized overall error The formula for calculating the normalized overall error is: In the formula, and The root mean square error and bias score of the data source. To determine the maximum observed water depth within the study area, As an indicator variable, when the data source provides measured water depth When only the flood range is provided It provides neither water depth nor range. . 8.The flood simulation verification method of claim 1, wherein, Step S4-4 describes the calculation of the overall comprehensive verification index, which is the overall comprehensive verification index. The calculation formula is: In the formula, Total number of data sources For the first The weight of each data source, For the first A comprehensive score from multiple data sources. 9.The flood simulation verification method of claim 1, wherein, The iterative optimization described in steps S4-5 specifically involves: adjusting the overall comprehensive verification index... If the result is not met when compared with the preset accuracy threshold, the result is fed back to step S3. The model's Manning roughness coefficient, permeability coefficient, capillary potential energy value, evapotranspiration value, and viscosity coefficient parameters are adjusted, and the simulation is repeated. The verification process in step S4 is repeated until... The preset accuracy threshold is reached or exceeded.
Citation Information
Patent Citations
Basin rainstorm flood disaster-bearing body reset cost remote sensing simulation method
CN115907574A
Multimodal remote sensing-based flood maximum submerging water depth space simulation method
CN116305902A