A method for predicting urban waterlogging depth considering geographical similarity

By introducing geographical similarity theory and automatic machine learning framework, a urban waterlogging depth prediction model was constructed, which solved the problem of inaccurate simulation results caused by incomplete consideration of factors in the existing technology, and achieved more accurate and explainable waterlogging depth prediction, supporting urban waterlogging flood prevention decisions.

CN119272947BActive Publication Date: 2025-07-18NANJING NORMAL UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411795165.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2025-07-18
Estimated Expiration
2044-12-09

AI Technical Summary

Technical Problem

The existing technology cannot effectively consider various factors such as geographical location, climate, and terrain in urban flooding simulation, resulting in inaccurate simulation results and lack of interpretability, which cannot meet the requirements of calculation speed and real-time.

Method used

Using geographical similarity theory and combining automatic machine learning framework, we collect basic geographic information data, build a geographical similarity calculation unit, and use hydrological and hydrodynamic models and automatic machine learning model library to optimize the geographical similarity model and predict the depth of urban waterlogging.

Benefits of technology

It improves the accuracy and timeliness of urban flooding simulation results, enhances the interpretability of the model, and supports flood control and drainage decisions in different regions and granularities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119272947B_ABST
    Figure CN119272947B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for predicting urban waterlogging depth considering geographical similarity. The purpose is to introduce the geographical similarity theory into urban waterlogging modeling and analyze the geographical modeling mechanism contained in the urban waterlogging process under rainfall scenarios. The method includes the following steps: collecting basic geographical information data, calculating geographical environment factors, and constructing a geographical configuration data set; preprocessing the data and constructing a geographical similarity calculation unit; based on raster data, using a hydrological and hydrodynamic model to simulate the surface runoff generation and concentration process in the study area and constructing an urban waterlogging depth data set; based on the urban waterlogging data set and the geographical similarity calculation unit, using an automated machine learning model library to construct a geographical similarity model; optimizing hyperparameters to optimize the geographical similarity model; and calculating the predicted water depth. The beneficial effects of the present invention are to comprehensively consider geographical factors, introduce the geographical similarity theory into urban waterlogging modeling, and improve the accuracy and timeliness of urban waterlogging simulation results through automated machine learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of urban waterlogging in the field of hydrology and water resources, and in particular relates to a method for predicting water depth of urban waterlogging taking into account geographical similarities. Background Art

[0002] Against the backdrop of increasing global climate change, the continuous advancement of urbanization, and frequent extreme rainfall, the frequent occurrence of urban waterlogging has become a phenomenon that cannot be ignored. This extreme event not only poses a serious threat to public health, but also causes huge property losses and poses a serious challenge to social stability and economic development.

[0003] In response to the problem of urban waterlogging caused by heavy rains, some scholars have used physical modeling methods such as process mechanism expression and geographic calculation to reveal the spatial distribution, spatiotemporal process and evolution law of urban waterlogging geographical phenomena. However, the physical driving method cannot meet the requirements of urban waterlogging simulation for computing speed and real-time performance. Some scholars have achieved rapid simulation of urban waterlogging based on data-driven methods, but this method ignores relevant knowledge about waterlogging and the simulation results lack interpretability. Moreover, these methods fail to effectively take into account the complex geographical elements of the city and cannot accurately express the impact of geographical factors in the simulated area on the flood process.

[0004] Therefore, the present invention proposes a method for predicting urban waterlogging depth while taking into account geographical similarities. Under the condition of taking into account geographical similarities, it comprehensively considers multiple factors such as geographical location, climate, topography, vegetation coverage, etc., and uses an automatic machine learning framework for modeling to achieve a complete expression and prediction of the occurrence and development process of urban waterlogging. Summary of the invention

[0005] In view of the problems existing in the above-mentioned technologies, the present invention provides a method for predicting urban waterlogging depth taking into account geographical similarities, with the aim of introducing the geographical similarity theory into urban waterlogging modeling and analyzing the geographical modeling mechanism implied in the urban waterlogging process under rainfall scenarios, in order to achieve a more accurate and reliable prediction of urban waterlogging depth.

[0006] In order to achieve the above-mentioned purpose of the invention, the technical solution adopted by the present invention is: a method for predicting urban waterlogging depth taking into account geographical similarities, which comprises:

[0007] S1: Collect basic geographic information data, calculate geographic environmental factors, and construct geographic configuration data sets;

[0008] S2: data preprocessing, constructing geographic similarity calculation units;

[0009] S3: Based on raster data, the hydrological and hydrodynamic model is used to simulate the surface runoff generation and convergence process in the study area and construct an urban waterlogging depth dataset;

[0010] S4: Based on the urban waterlogging dataset and the geographical similarity calculation unit, construct a geographical similarity model using an automated machine learning model library;

[0011] S5: Hyperparameter optimization to optimize the geographical similarity model;

[0012] S6: Predict the water depth calculation.

[0013] Furthermore, the S1 includes the following steps:

[0014] S11: Select the Nanjing urban area division and obtain the basic geographical information data of this area from the National Catalogue Service For Geographic Information;

[0015] S12: Based on the digital elevation data, extract the geographical features of this area through the built-in calculation method of ArcMap, and construct a geographical configuration dataset based on 15 geographical factors including land use type, soil type, aspect, slope, plan curvature, profile curvature, roughness, rugosity, terrain moisture index, sediment transport index, water flow power index, terrain ruggedness index, terrain position index, normalized difference vegetation index, and terrain control index;

[0016] S13: Select potential explanatory variables through the Spearman correlation analysis method. The calculation formula is as follows:

[0017] (1),

[0018] Where, is the Spearman correlation coefficient value, is the total sample size, represents the rank difference between the

[0019] The collinearity between explanatory variables is tested using the variance inflation factor (VIF). Generally speaking, a VIF greater than 10 will be regarded as high collinearity, and the variable values with a VIF value exceeding 10 will be removed until finally all selected explanatory variables are below 10, so as to determine a geographical configuration dataset that fully expresses the geographical environment characteristics.

[0020] Furthermore, the S2 includes the following steps:

[0021] S21: Standardize and normalize each factor in the geographical configuration dataset, and perform feature expression on the processed data. The calculation formula is as follows:

[0022] (2),

[0023] (3),

[0024] (4),

[0025] Among them, is the processed result, is a certain piece of data, is the minimum value in the data, is the maximum value in the data, is the mean of the sample data, is the standard deviation of the sample data, is the explanatory variable at a certain position point value; is the set of explanatory variable values at all position points;

[0026] S22: Based on the hypothesis that there is a similarity relationship between the observation position and the unknown position , calculate the unknown position , and the similarity of the observation position , The formula for the similarity is as follows:

[0027] (5),

[0028] (6),

[0029] In the formula, is the initial similarity value, is the similarity based on the th covariate and positions, is the function that determines the similarity based on and by comparing the similarity of covariance levels of all covariances,

[0030] S23: Evaluate the similarity between the unknown position and the observation position by calculating the distance value. The calculation formula is as follows:

[0031] (7),

[0032] (8),

[0033] (9),

[0034] (10),

[0035] (11),

[0036] (12),

[0037] In the above formula, represents the Chebyshev distance value, represents the Manhattan distance value, represents the Euclidean distance value, represents the Minkowski distance value, represents the assumed total number of categories, represents the Canberra distance value, where is the true total number of categories, is the total sample size, are all unknown locations and the observed location of the explanatory variable is the square root of the mean deviation, is the set of observed locations is the set of unknown locations ;

[0038] The data that meets the similarity hypothesis is screened out through formulas (5)-(6). For the data that meets the similarity hypothesis, the similarity values between the observed location and the unknown location are calculated respectively by the similarity measurement methods of formulas (7)-(11). A weight vector matrix based on the average similarity is constructed through formula (12) as the similarity calculation unit.

[0039] Further, the S3 includes the following steps:

[0040] S31: Model the urban surface of the embodiment. To carry out the modeling work of the urban surface part of the embodiment area, the present invention obtains geospatial data such as land use types, remote sensing images, building and river network water system distributions, all of which adopt the WGS-84 coordinate system. When using a plane coordinate system, the map projection uniformly adopts the UTM zone 50N projection;

[0041] S32: Establish surface grid cells based on spatial discrete grids;

[0042] S33: Based on the grid data, use a hydrological and hydrodynamic model to model the urban surface runoff generation and concentration process, and input parameter information such as rainfall data, land use type data, elevation information (DEM), building data, infiltration rate, Manning coefficient grid, etc. into the grid cells;

[0043] Further, the S33 includes the following steps:

[0044] S331: Based on typical rainfall events in the region, using point-like time series data from the rainfall station, the spatial rainfall distribution in the embodiment region is estimated by spatial interpolation;

[0045] S332: Generate the Manning coefficient grid using land use type data and building data. Taking into account the research area and the runoff characteristics of the urban underlying surface, the present invention divides the urban surface grid into five types: grassland, water body, building, road, and impervious surface;

[0046] S333: Construct Manning coefficient grids according to different types as input data for hydrological and hydrodynamic models;

[0047] S34: adding the time step in the rainfall grid to the hydrological and hydrodynamic model input file to adjust the simulation time and total simulation time of the area to simulate waterlogging;

[0048] S35: Based on the simulation results of the study area, data exploratory analysis methods are used to find missing or incorrectly labeled data, correct and delete inappropriate data, and outlier deletion uses the MAD outlier identification method. The calculation formula is as follows:

[0049] (13),

[0050] In the above formula, is the mean of the data, is the standard deviation of the data. Through the above operations, the urban flooding depth dataset is constructed.

[0051] Further, the S4 comprises the following steps:

[0052] S41: Rainfall dataset, waterlogging depth dataset and similarity calculation unit are matched, i.e. the observed waterlogging depth value is used as a label for model learning;

[0053] Further, the S41 comprises the following steps:

[0054] S411: Time matching of rainfall sequence and water depth sequence. According to the existing rainfall sequence, the start time, time step, end time and other information of rainfall are obtained, and then the start time, time step and end time defined by the configured water depth result are matched (generally, this process cannot be matched directly). The matching is mainly to lengthen two inconsistent sequences. Due to the randomness of rainfall events, there is inconsistency in the time interval of rainfall, so it is necessary to add 0 value (default is non-rainfall) to the specified rainfall simulation time interval;

[0055] S412: Spatial matching of rainfall sequence and water depth sequence. For spatial matching, spatial resolution interpolation conversion is performed according to the spatial resolution data required by the experiment. The calculation formula is as follows:

[0056] (14),

[0057] In the above formula, is the observed value at the th discrete point, is the estimated value at grid point , is the number of samples participating in the calculation, is the distance between the interpolation point and the th station, is the power of the distance, takes a value of 1.6.

[0058] S42: High-dimensional space mapping. After completing the above similarity calculation and dataset matching, all similarity values are mapped to a high-dimensional space to convert it into a regression problem. The mapping is completed based on a kernel function. The calculation formula is as follows:

[0059] (15),

[0060] In the above formula, is the newly constructed kernel function, is the total number of samples, is a constant to be determined, is a combined kernel function, including the combination of a linear kernel function and a Gaussian kernel function, and represent different vectors in the input space, and are Lagrange operators. After completing the above steps, a spatio-temporal dataset of urban waterlogging hydrology is constructed.

[0061] S43: Perform data preprocessing on, and handle missing values, outliers, and extreme points;

[0062] Furthermore, S43 includes the following steps:

[0063] S431: In order to check the quality of the data in the spatio-temporal dataset of urban waterlogging hydrology, the interquartile range of the box plot is used to check for outliers. If a data point is above the upper quartile and exceeds 1.5 times the distance, or is below the lower quartile and less than 1.5 times the distance, these points are regarded as outliers;

[0064] S432: Since there are outliers in the water depth result data, the distribution of normal data is extremely compressed. The interquartile range detection algorithm is used to remove outliers, and the calculation formula is as follows:

[0065] (16),

[0066] represents the interquartile range value, represents the third quartile, represents the first quartile;

[0067] The DEM data is more in line with the true data distribution, so it is not processed. Since the median of rainfall data is generally 0, a part of 0 is removed to reduce the sample imbalance caused by too many 0 values;

[0068] S433: Remove data features with too high correlation to complete the dataset screening;

[0069] S434: Divide the dataset. Observe the distribution of the training set and the test set, and use the KDE distribution map to view and compare the distribution of feature variables in the training set and the test set. The training set and the test set of basically all variables generally satisfy the same distribution, meeting the requirements for dataset division.

[0070] S44: Use the AutoGluon automated machine learning framework, an automated machine learning library for training highly accurate machine learning models, and adopt multi-layer stacking integration and times of repetition fold bagging to optimize the training and validation process of the model;

[0071] Furthermore, the S44 includes the following steps:

[0072] S441: Feature construction. Feature construction is to create new features from the urban waterlogging hydrological spatio-temporal dataset. Generally, feature crossing (through some mathematical operations, such as arithmetic operations, , etc., to quickly derive new variables) and decomposing the original features are used to create new features, which are processed by AutoGluon;

[0073] S442: Divide the urban waterlogging hydrological spatio-temporal dataset. 70% of the data randomly selected from the urban waterlogging hydrological spatio-temporal dataset is used as the training set, and the remaining 30% is used as the test set for model input and validation;

[0074] S443: Establish a model library. AutoGluon first estimates the required training time. If it exceeds the remaining time of this layer, it jumps to the next stacked layer. After each new model is trained, it is immediately saved to disk for fault tolerance. This design makes the behavior of the framework highly predictable, ensuring that at least one model is trained within the specified time to generate predictions. Additionally, when AutoGluon checks the intermediate iteration process of sequentially trained models such as neural networks and tree models, model generation within a finite time can still be guaranteed. Furthermore, if a model fails during training, a jump to the next model operation is performed for this event. Generate multiple models;

[0075] S444: Model selection. Based on the candidate models, score them from the perspectives of model training time, computing resources, and whether the simulation results conform to hydrological rules, and select the top three scoring models as candidate models.

[0076] S45: Build a geographical similarity model. Due to the performance differences of each model under different geographical configurations, this application adapts locally from the top three models with comprehensive scores based on the accuracy of each model at the similarity calculation unit scale to build a similarity model that meets the specific geographical environment.

[0077] Further, the S5 includes the following steps:

[0078] S51: Set the hyperparameter search range to build a hyperparameter pool;

[0079] S52: Optimize the model by adjusting hyperparameters;

[0080] Further, the S52 includes the following steps:

[0081] S521: Conduct a search for hyperparameters through the Bayesian optimization algorithm;

[0082] S522: In this process, in order to limit the time consumption of model training, the parameter epoch_valid_improve is added to save model training time;

[0083] S523: Use the Gaussian process regression method and performance evaluation metrics in the Bayesian optimization algorithm for iterative search to obtain the optimal parameter combination,

[0084] S53: Based on the optimal parameter combination, complete the construction of the optimal geographical similarity model,

[0085] Further, the S6 includes the following steps:

[0086] S61: Design waterlogging scenarios according to different rainfall time resolution and spatial resolution parameters;

[0087] Further, S61 includes the following steps:

[0088] S611: Design rainfall scenarios with different return periods through the rainfall intensity formula. The rainfall intensity calculation formula for this experimental area is as follows:

[0089] (17),

[0090] In the above formula, is the designed rainfall intensity (mm / min), is the rainfall duration (min), is the design return period (years).

[0091] According to the rainfall intensity formula in the "Notice on Issuing the Rainfall Intensity Formula of Nanjing City" issued by the Nanjing Water Conservancy Bureau in 2024, design rainfall scenarios with a duration of 180 min, a comprehensive rain peak position coefficient of 0.393, and return periods of 1-year, 5-year, 10-year, 20-year, 50-year, and 100-year;

[0092] S612: Construct more scenarios through different time step settings, and use these multi-temporal and multi-spatial rainfall datasets as the input of the hydrological and hydrodynamic model to simulate the inundation depth under different waterlogging scenarios.

[0093] S62: Model pre-training with the same parameters. Use the optimal parameters in the fifth step as prior parameters for constructing the similarity model;

[0094] S63: Select different water depth thresholds, different time series rainfall data, and different computational units for similarity model training;

[0095] Further, S63 includes the following steps:

[0096] S631: Adopt spatial granular grid-based training, compare multiple rainfall scenarios, and keep the water depth thresholds in the top three in all selected experimental scenarios as the final results;

[0097] S632: The impact of different time series rainfall data on the final training results is one of the factors that need to be considered in the modeling. The selected historical time series rainfall data length is 1 - 30. Compare multiple rainfall scenarios, keep the water depth thresholds in the top three in all selected experimental scenarios as the final results, and compare the model performance;

[0098] S633: Compare two types of computational units, namely spatial granular grid type and neighborhood type;

[0099] S634: Basically, tree models and ensemble models have superior performance, and conduct training for different scenarios;

[0100] S64: Water depth prediction calculation. The geographical environment and rainfall at the position to be predicted are used as the inputs of the similarity model for water depth prediction calculation, and the predicted water depth result is output.

[0101] Specifically as follows:

[0102] S631: Adopt spatial granular grid-based training, select a water depth threshold of 1 mm - 10 mm, compare multiple rainstorm scenarios, and keep the water depth thresholds ranked in the top three in all selected experimental scenarios as the final result. Compare the root mean square error and mean absolute error under different water depth thresholds, and find that the smaller the water depth threshold, the smaller the root mean square error and mean absolute error. The water depth threshold with the largest correlation coefficient and goodness of fit is 5 mm, and finally select the best-performing water depth threshold of 5 mm;

[0103] S632: The impact of different time-series rainfall data on the final training result is one of the factors that need to be considered in modeling. The length of the historical time-series rainfall data selected is 1 - 30. Compare multiple rainstorm scenarios, and keep the water depth thresholds ranked in the top three in all selected experimental scenarios as the final result. The root mean square error and mean absolute error are kept between 20 mm - 52 mm and 13 mm - 38 mm respectively under all experimental scenarios and all rainfall time series at all times, and the correlation coefficient and goodness of fit are kept between 0.3 - 0.84 and 0.03 - 0.72 respectively. After comprehensive consideration, it is found that the rainfall time series of the first 26 - 29 moments performs well under multiple indicators;

[0104] S633: Compare two types of calculation units, namely spatial granular grid type and neighborhood type. The root mean square error, mean absolute error, correlation coefficient, and goodness of fit indicators of the neighborhood type calculation unit basically change in the same way. The overall root mean square error and mean absolute error are less than 30 mm and 17 mm respectively, and the correlation coefficient and goodness of fit are greater than 0.85 and 0.7 respectively, meeting the ideal results; the root mean square error and mean absolute error of the spatial granular grid type calculation unit are greater than 30 mm and 18 mm respectively, and the correlation coefficient and goodness of fit are both less than 0.85 and 0.7 respectively. It can be seen that the neighborhood type calculation unit has greatly improved the model performance;

[0105] S634: Finally, select a water depth threshold of 5 mm, the rainfall data of the first 26 historical moments, and the neighborhood type calculation unit for optimal configuration calculation. The tree model and the ensemble model are basically superior in performance, the error is basically stable within 52 mm, and the goodness of fit R2 is greater than 0.6010.

[0106] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0107] 1. Incorporate the geographical similarity theory into urban waterlogging modeling in combination with geographical laws, analyze the geographical modeling mechanisms underlying the urban rainfall process and urban waterlogging process under waterlogging scenarios, and improve the application of GIS theory in urban waterlogging modeling;

[0108] 2. Overcome the limitation of inaccurate flood simulation results caused by the simple elements considered in traditional models, incorporate geographical environment factors into model calculations, comprehensively utilize geographical factors and multi-temporal rainfall data, conduct urban waterlogging modeling from the perspectives of information geography and geographical similarity, characterize the spatio-temporal evolution characteristics of real urban rainstorm waterlogging events, and improve the accuracy and timeliness of rainstorm waterlogging simulation results;

[0109] 3. Aiming at the spatial granularity (spatial resolution) effect and parameter uncertainty of model applications, design different spatial granularity scenarios and different model configuration conditions to optimize parameters, which is conducive to enhancing the spatio-temporal simulation ability of urban waterlogging models at different spatial granularities, providing references for urban waterlogging modeling in different regions and at different granularities, and thus strongly supporting flood control and drainage decision-making in urban flood-prone areas;

[0110] 4. Based on the geographical similarity theory and spatial similarity calculation method, use geographical configuration factors to construct a geographical similarity calculation unit for the prior knowledge model of urban waterlogging depth prediction, making the urban waterlogging depth prediction method more interpretable;

[0111] 5. Use a high-precision urban waterlogging depth dataset to divide the training set and validation set. Through a structured automated machine learning framework, convert the water depth image data into structured tabular data and jointly train the model with geographical similarity units. Select the optimal geographical similarity model from multiple pre-evaluation model combinations to achieve the simulation and prediction of waterlogging depth based on rainfall and geographical configuration factors. BRIEF DESCRIPTION OF THE DRAWINGS

[0112] Figure 1 is the method framework diagram of the present invention;

[0113] Figure 2 is the data outlier inspection diagram of the present invention;

[0114] Figure 3 is the model hyperparameter training flow chart of the present invention;

[0115] Figure 4 is the model of the present invention 10-fold cross-validation schematic diagram;

[0116] Figure 5 is the schematic diagram of the water depth prediction result of the example area of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0117] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to embodiments and the accompanying drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention.

[0118] Taking a certain area as an example, the area is surrounded by the Yangtze River, the Qinhuai River, and the ring road, with no water volume exchange with the outside. The rainwater drainage network within the area is relatively independent, and there are various artificial ditches, lakes, and other water storage units. Therefore, it can be determined that the area of this embodiment belongs to a typical city with relatively clear boundary conditions. The overall terrain of the embodiment area is flat, with only a few mountains with a maximum altitude of 448.9 m, and the geographical features and hydrological types are relatively rich.

[0119] As Figure 1 shown, a method for predicting urban waterlogging depth considering geographical similarity disclosed in an embodiment of the present invention mainly includes the following steps:

[0120] S1: Collect basic geographic information data, calculate geographical environment factors, and construct a geographical configuration data set;

[0121] S2: Perform data preprocessing and construct a geographical similarity calculation unit;

[0122] S3: Based on raster data, use a hydrological and hydrodynamic model to simulate the surface runoff generation and concentration process in the study area, and construct an urban waterlogging depth data set;

[0123] S4: Based on the urban waterlogging data set and the geographical similarity calculation unit, use an automated machine learning model library to construct a geographical similarity model;

[0124] S5: Perform hyperparameter optimization to optimize the geographical similarity model;

[0125] S6: Calculate the predicted water depth.

[0126] Furthermore, the S1 includes the following steps:

[0127] S11: Select the Nanjing urban area and obtain the basic geographical information data of this area from the National Catalogue Service For Geographic Information;

[0128] S12: Based on the digital elevation data, extract the geographical features of this area through the built-in calculation method of ArcMap, and construct a geographical configuration data set based on 15 geographical factors including land use type, soil type, aspect, slope, plane curvature, profile curvature, roughness, rugosity, terrain humidity index, sediment transport index, water flow power index, terrain ruggedness index, terrain position index, normalized vegetation index, and terrain control index;

[0129] As shown in Table 1;

[0130] Table 1 Geographic Configuration Dataset

[0131]

[0132] S13: Select potential explanatory variables through the Spearman correlation analysis method. The calculation formula is as follows:

[0133] (1),

[0134] Wherein, is the Spearman correlation coefficient value, is the total sample size, represents the rank difference between the

[0135] The collinearity between explanatory variables is tested using the variance inflation factor (VIF). Generally speaking, a VIF greater than 10 is considered high collinearity, and the variable values with VIF values exceeding 10 will be removed until finally all selected explanatory variables are below 10, so as to determine the geographic configuration dataset that fully expresses the geographic environment characteristics.

[0136] Furthermore, the S2 includes the following steps:

[0137] S21: Standardize and normalize each factor in the geographic configuration dataset, and perform feature expression on the processed data. The calculation formula is as follows:

[0138] (2),

[0139] (3),

[0140] (4),

[0141] Wherein, is the processed result, is a certain piece of data, is the minimum value in the data, is the maximum value in the data, is the mean of the sample data, is the standard deviation of the sample data, is the explanatory variable at a certain location point value; is the set of explanatory variable values at all location points;

[0142] S22: Based on the observation location and the unknown location Assume a similarity relationship and calculate the unknown position , and the observed position , The formula for their similarity is as follows:

[0143] (5),

[0144] (6),

[0145] In the formula, is the initial similarity value, is based on the th covariate and the similarity of positions, is determined by comparing the similarity of covariance levels of all covariances to decide the similarity based on and function,

[0146] S23: Evaluate the similarity between the unknown position and the observed position by calculating the distance value. The calculation formula is as follows:

[0147] (7),

[0148] (8),

[0149] (9),

[0150] (10),

[0151] (11),

[0152] (12),

[0153] In the above formula, represents the Chebyshev distance value, represents the Manhattan distance value, represents the Euclidean distance value, represents the Minkowski distance value, represents the total number of assumed categories, represents the Canberra distance value, where is the total number of true categories, is the total sample size, is all unknown positions and the observed position the explanatory variable the square root of the mean deviation, is the observed position A collection of Unknown location A collection of;

[0154] The data that meets the similarity hypothesis are screened out through formulas (5)-(6). For the data that meets the similarity hypothesis, the similarity values of the observed position and the unknown position are calculated respectively through the similarity measurement method of formulas (7)-(11). The weight vector matrix based on the average similarity value is constructed through formula (12) as the similarity calculation unit.

[0155] Further, the S3 comprises the following steps:

[0156] S31: Modeling the urban surface of the embodiment. In order to carry out the modeling work of the urban surface part of the embodiment area, the present invention obtains the land use type, remote sensing images, buildings and river network distribution and other geographic spatial data, all of which use the WGS-84 coordinate system. When the plane coordinate system is used, the map projection uniformly uses the UTM zone 50N projection;

[0157] S32: Establishing surface grid cells based on spatial discrete grids;

[0158] S33: Based on the grid data, the hydrological and hydrodynamic model is used to model the urban surface runoff generation and convergence process, and the parameter information such as rainfall data, land use type data, elevation information (DEM), building data, infiltration rate, Manning coefficient grid, etc. are input into the grid unit;

[0159] Further, the S33 comprises the following steps:

[0160] S331: Based on typical rainfall events in the region, using point-like time series data from the rainfall station, the spatial rainfall distribution in the embodiment region is estimated by spatial interpolation;

[0161] S332: Generate the Manning coefficient grid using land use type data and building data. Taking into account the research area and the runoff characteristics of the urban underlying surface, the present invention divides the urban surface grid into five types: grassland, water body, building, road, and impervious surface;

[0162] S333: Construct the Manning coefficient grid according to different types, and the Manning roughness coefficient values are shown in Table 2 below:

[0163] Table 2 Manning roughness coefficient values for different surface types:

[0164]

[0165] S34: Add the time step in the rainfall grid to the input file of the hydrological and hydrodynamic model to adjust the simulation time and total simulation duration for the area, and conduct waterlogging simulation;

[0166] S35: Based on the simulation results of the study area, use data exploratory analysis methods to find missing or mislabeled data, correct and delete inappropriate data. The MAD outlier identification method is used for outlier deletion, and the calculation formula is as follows:

[0167] (13),

[0168] In the above formula, is the mean of the data, is the standard deviation of the data. Through the above operations, a waterlogging water depth dataset for the city is constructed.

[0169] Furthermore, the S4 includes the following steps:

[0170] S41: Match the rainfall dataset, waterlogging water depth dataset with the similarity calculation unit, that is, the observed waterlogging water depth value is used as a label for model learning;

[0171] Furthermore, the S41 includes the following steps:

[0172] S411: Match the time of the rainfall sequence and the water depth sequence. Obtain information such as the start time, time step, and end time of the rainfall based on the existing rainfall sequence, and then match according to the start time, time step, and end time defined by the configured water depth results (generally, this process cannot be directly matched). The matching is mainly to supplement the lengths of two inconsistent sequences. Due to the randomness of rainfall events and the inconsistency of the rainfall time intervals, it is necessary to supplement 0 values (default non-rainfall) to the specified rainfall simulation time interval;

[0173] S412: Match the space of the rainfall sequence and the water depth sequence. For space matching, spatial resolution interpolation conversion is performed according to the spatial resolution data required by the experiment, and the calculation formula is as follows:

[0174] (14),

[0175] In the above formula, is the observed value at the th discrete point, is the estimated value at the grid point , is the number of samples participating in the calculation, is the distance between the interpolation point and the th station, is the power of the distance, The value of is 1.6.

[0176] S42: High-dimensional space mapping. After completing the above similarity calculation and dataset matching, map all similarity values to a high-dimensional space to convert it into a regression problem. The mapping is completed based on a kernel function, and the calculation formula is as follows:

[0177] (15),

[0178] In the above formula, is a newly constructed kernel function, is the total number of samples, is a constant to be determined, is a combined kernel function, including the combination of a linear kernel function and a Gaussian kernel function, and represent different vectors in the input space, and are Lagrange operators. After completing the above steps, construct a spatio-temporal dataset of urban waterlogging hydrology;

[0179] S43: Perform data preprocessing on, and handle missing values, outliers, and extreme points;

[0180] Furthermore, the S43 includes the following steps:

[0181] S431: In order to check the quality of the data in the spatio-temporal dataset of urban waterlogging hydrology, use the interquartile range of the box plot to check for outliers. If a data point is above the upper quartile and exceeds 1.5 times of the distance, or below the lower quartile and less than 1.5 times of the distance, these points are regarded as outliers;

[0182] S432: As shown in the outlier check result Figure 2 , since there are outliers in the water depth result data, the distribution of normal data is extremely compressed. Use the interquartile range detection algorithm to remove outliers, and the calculation formula is as follows:

[0183] (16),

[0184] represents the interquartile range value, represents the third quartile, represents the first quartile;

[0185] The DEM data is more in line with the true data distribution, so it is not processed. Since the median of rainfall data is generally 0, removing a part of 0 can reduce the sample imbalance caused by too many 0 values;

[0186] S433: Remove data features with too high correlation to complete dataset screening;

[0187] S434: Divide the dataset. Observe the distributions of the training set and the test set, and use the KDE distribution plot to view and compare the distribution of feature variables in the training set and the test set. The training sets and test sets of basically all variables generally satisfy the same distribution, meeting the requirements for dataset division.

[0188] S44: Use AutoGluon Tabular (an automated machine learning library for training highly accurate machine learning models), an automated machine learning library for structured data, which adopts multi-layer stacking integration and times repetition bagging to optimize the model training and validation process;

[0189] Furthermore, the S44 includes the following steps:

[0190] S441: Feature construction. Feature construction creates new features from the urban waterlogging hydrological spatio-temporal dataset. Generally, feature crossing (deriving new variables quickly through some mathematical operations, such as arithmetic operations, , etc.) and decomposing the original features are used to create new features, and this part is processed by AutoGluon;

[0191] S442: Divide the urban waterlogging hydrological spatio-temporal dataset. Randomly select 70% of the data in the urban waterlogging hydrological spatio-temporal dataset as the training set, and the remaining 30% as the test set for model input and validation;

[0192] S443: Establish a model library. AutoGluon first estimates the required training time. If it exceeds the remaining time of this layer, it jumps to the next stacking layer. After each new model is trained, it is immediately saved to disk for fault tolerance. This design makes the behavior of the framework highly predictable, ensuring that at least one model can be trained within the specified time to generate predictions. In addition, when AutoGluon checks the intermediate iteration process of sequentially trained models such as neural networks and tree models, it can still ensure the generation of models within a limited time. Moreover, if a model fails during training, a jump to the next model operation will be performed for this event. Generate multiple models;

[0193] S444: Model selection. Score the candidate models from the perspectives of model time consumption, computing resources, and whether the simulation results conform to hydrological rules, and select the top three models with the highest scores as candidate models.

[0194] S45: Construct a geographical similarity model. Due to the performance differences of various models under different geographical configurations, this application adapts locally from the top three models in terms of comprehensive scores based on the accuracy of each model at the similarity calculation unit scale to construct a similarity model that meets the specific geographical environment.

[0195] Further, the said S5 includes the following steps:

[0196] S51: Set the hyperparameter search range to construct a hyperparameter pool. Taking the Tabular_nn and CatBoost models as examples, the hyperparameter search range is shown in Table 3;

[0197] Table 3 Hyperparameters and their ranges of the model

[0198]

[0199] S52: Optimize the model by adjusting the hyperparameters;

[0200] Further, the said S52 includes the following steps:

[0201] S521: Conduct a search for hyperparameters through the Bayesian optimization algorithm. Taking the CatBoost model as an example, the search for hyperparameters is shown in Table 4;

[0202] Table 4 Hyperparameters of the CatBoost model

[0203]

[0204] S522: In this process, in order to limit the time consumption of model training, the parameter of epoch_valid_improve is added to save the model training time;

[0205] S523: Use the Gaussian process regression method and performance evaluation indicators in the Bayesian optimization algorithm for iterative search to obtain the optimal parameter combination. The specific process is as Figure 3 shown.

[0206] S53: Based on the optimal parameter combination, complete the construction of the optimal geographical similarity model.

[0207] Further, the said S6 includes the following steps:

[0208] S61: Design waterlogging scenarios according to different rainfall time resolution and spatial resolution parameters;

[0209] Further, the said S61 includes the following steps:

[0210] S611: Design rainfall scenarios with different return periods through the rainstorm intensity formula. The rainstorm intensity calculation formula for this experimental area is as follows:

[0211] (17),

[0212] In the above formula, is the design rainfall intensity (mm / min), is the rainfall duration (min), is the design return period (years).

[0213] According to the rainfall intensity formula in the "Notice on Issuing the Rainfall Intensity Formula of Nanjing City" issued by the Nanjing Water Bureau in 2024, with a design duration of 180 min and a comprehensive rain peak position coefficient of 0.393, rainfall scenarios of once in 1 year, once in 5 years, once in 10 years, once in 20 years, once in 50 years, and once in 100 years;

[0214] S612: By setting different time steps, construct more scenarios, and use these multi - spatio - temporal rainfall datasets as inputs to the hydrological and hydrodynamic model to simulate the inundation depth under different waterlogging scenarios.

[0215] S62: Model pre - training based on the same parameters. Use the optimal parameters in the fifth step as prior parameters for constructing the similarity model;

[0216] S63: Select different water depth thresholds, different time - series rainfall data, and different calculation units for similarity model training;

[0217] Furthermore, S63 includes the following steps:

[0218] S631: Adopt spatial - granular grid - type training, select the water depth threshold from 1 mm to 10 mm, use the k - fold cross - validation method to determine the best similarity threshold. As Figure 4 shown, compare multiple rainfall scenarios, and keep the water depth thresholds in the top three in all selected experimental scenarios as the final result. Compare the root - mean - square error and mean absolute error under different water depth thresholds, and find that the smaller the water depth threshold, the smaller the root - mean - square error and mean absolute error, while the water depth threshold with the largest correlation coefficient and goodness of fit is 5 mm. Finally, select the best - performing water depth threshold of 5 mm;

[0219] S632: The impact of rainfall data at different time series on the final training results is one of the factors that need to be considered in the modeling. The selected length of historical time series rainfall data is 1 - 30. Multiple rainstorm scenarios are compared, and the water depth thresholds that rank in the top three in all selected experimental scenarios are taken as the final results. The root mean square error and mean absolute error in all experimental scenarios and at all moments of rainfall time series are kept between 20 mm - 52 mm and 13 mm - 38 mm respectively, and the correlation coefficient and goodness of fit are kept between 0.3 - 0.84 and 0.03 - 0.72 respectively. After comprehensive consideration, it is found that the rainfall time series at the first 26 - 29 moments performs well under multiple indicators;

[0220] S633: Compare two types of computational units, namely the spatial granular grid type and the neighborhood type. For the neighborhood type computational unit, the root mean square error, mean absolute error indicators, correlation coefficient, and goodness of fit indicators basically change in the same way. The overall root mean square error and mean absolute error are less than 30 mm and 17 mm respectively, and the correlation coefficient and goodness of fit are greater than 0.85 and 0.7 respectively, meeting the ideal results. For the spatial granular grid type computational unit, the root mean square error and mean absolute error are greater than 30 mm and 18 mm respectively, and the correlation coefficient and goodness of fit are both less than 0.85 and 0.7 respectively. It can be seen that the neighborhood type computational unit has greatly improved the model performance;

[0221] S634: Finally, select a water depth threshold of 5 mm, the rainfall data at the first 26 historical moments, and the neighborhood type computational unit for optimal configuration calculation. Basically, the tree model and the ensemble model have relatively superior performance, the error is basically stable within 52 mm, and the goodness of fit R2 is basically greater than 0.60. The training results for different scenarios are shown in Table 5:

[0222] Table 5 Model Training Results

[0223]

[0224] S64: Water depth prediction calculation. Use the geographical environment and rainfall at the location to be predicted as the input of the similarity model for future water depth prediction calculation. The accuracy of the calculation results is as Figure 5 shown, and its coefficient of determination is optimally 0.94. The curve trends of the water depth observed values and predicted values are consistent, and the predicted water depth results are output.

[0225] The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for predicting the water depth of urban waterlogging considering geographical similarity, characterized in that, It includes the following steps: S1: Collect basic geographic information data, calculate geographic environment factors, and construct a geographic configuration data set; S2: Data preprocessing, and construct a geographic similarity calculation unit; S3: Based on raster data, use a hydrological and hydrodynamic model to simulate the surface runoff generation and concentration process in the study area, and construct an urban waterlogging depth data set; S4: Based on the urban waterlogging depth data set and the geographic similarity calculation unit, use an automated machine learning model library to construct a geographic similarity model; S5: Hyperparameter optimization to optimize the geographic similarity model; S6: Predictive water depth calculation; Among them, S2 includes the following steps: S21: Standardize and normalize each factor in the geographic configuration data set, and perform feature expression on the processed data. The calculation formula is as follows: (2), (3), (4), Among them, is the processed result, is a certain piece of data, is the minimum value in the data, is the maximum value in the data, is the mean of the sample data, is the standard deviation of the sample data, , is the explanatory variable at a certain position point value; is the set of explanatory variable values at all position points; S22: Based on the assumption of a similarity relationship between the observation position and the unknown position , calculate the similarity of the unknown position , and the observation position , The formula is as follows: (5), (6), In the formula, is the initial similarity value, is based on the th covariate and position similarity, is determined by comparing the similarity of covariance levels of all covariances to be a function of the similarity based on and ​ S23: Evaluate the similarity between the unknown location and the observation location by calculating the distance value. The calculation formula is as follows: (7), (8), (9), (10), (11), (12), In the above formula, represents the Chebyshev distance value, represents the Manhattan distance value, represents the Euclidean distance value, represents the Minkowski distance value, represents the assumed total number of categories, represents the Canberra distance value, where is the actual total number of categories, is the total sample size, is all unknown positions and the explanatory variable of the observation position is the square root of the mean deviation, of for the observation position set, is the set of unknown positions ; Filter out the data that meets the similarity hypothesis through formulas (5)-(6). For the data that meets the similarity hypothesis, calculate the similarity values between the observation location and the unknown location respectively through the similarity measurement methods of formulas (7)-(11), and construct a weight vector matrix based on the similarity average value through formula (12) as the similarity calculation unit.

2. A method for predicting urban waterlogging depth considering geographic similarity according to claim 1, characterized in that S1 includes the following steps: S11: Obtain regional geographic information basic data from the National Geographic Information Resource Catalog Service System; S12: Based on digital elevation data, extract regional geographic features through the built-in calculation method of ArcMap, and construct a geographic configuration data set based on 15 geographic factors including land use type, soil type, slope aspect, slope gradient, plan curvature, profile curvature, roughness, rugosity, terrain wetness index, sediment transport index, stream power index, topographic ruggedness index, topographic position index, normalized difference vegetation index, and terrain control index 1; S13: Select potential explanatory variables through the Spearman correlation analysis method. The calculation formula is as follows: (1), Among them, is the Spearman correlation coefficient value, is the total sample size, represents the rank difference between the pairs of data.

3. A method for predicting the water depth of urban waterlogging considering geographical similarity according to claim 1, characterized in that, The construction of the urban waterlogging depth data set in S3 includes the following steps: S31: Model the urban surface to obtain geographic spatial data such as land use type, remote sensing image, building, and river network distribution. All use the WGS-84 coordinate system. When using a plane coordinate system, the map projection uniformly uses the UTM zone 50N projection; S32: Establish surface raster units based on spatial discrete grids; S33: Based on raster data, use a hydrological and hydrodynamic model to model the surface runoff generation and concentration process of the urban surface, and input rainfall data, land use type data, elevation information, building data, infiltration rate, and Manning coefficient grid parameter information into the raster units; S34: Add the time step in the rainfall grid to the input file of the hydrological and hydrodynamic model to adjust the regional simulation start time and total simulation duration for waterlogging simulation; S35: Based on the simulation results of the study area, use data exploratory analysis methods to find missing or mislabeled data, correct and delete inappropriate data. The MAD outlier identification method is used for outlier deletion, and the calculation formula is as follows: (13), In the above formula, is the mean value of the data, is the standard deviation of the data. Through the above operations, a dataset of urban waterlogging depths is constructed.

4. A method for predicting the water depth of urban waterlogging considering geographical similarity according to claim 1, characterized in that, S4 Construction of the urban waterlogging geographical similarity model, including the following steps: S41: Matching of the rainfall dataset, waterlogging depth dataset and similarity calculation unit, that is, the observed waterlogging depth value is used as a label for model learning, specifically as follows. S411: Temporal matching of the rainfall sequence and the water depth sequence. Obtain the start time, time step, and end time information of the rainfall according to the existing rainfall sequence, and then perform matching according to the start time, time step, and end time defined by the configured water depth results. S412: Spatial matching of the rainfall sequence and the water depth sequence. For spatial matching, spatial resolution interpolation conversion is performed according to the spatial resolution data required by the experiment. The calculation formula is as follows: (14), In the above formula, is the observed value at the th discrete point, is the estimated value at the grid point , is the number of samples participating in the calculation, is the distance between the interpolation point and the th station, is the power of the distance, takes the value of 1.

6. S42: High-dimensional space mapping. After completing the similarity calculation and dataset matching, map all similarity values to a high-dimensional space to convert it into a regression problem. The mapping is based on a kernel function, and the calculation formula is as follows: (15), In the above formula, is the newly constructed kernel function, is the total number of samples, is a constant to be determined, is the combined kernel function, including the combination of the linear kernel function and the Gaussian kernel function, and represent different vectors in the input space, and are Lagrange operators. After completing the above steps, a spatio-temporal dataset of urban waterlogging hydrology is constructed; S43: Perform data preprocessing, and process missing values, outliers, and anomalies. S44: An automated machine learning library that uses AutoGluon Tabular, an automated machine learning framework for structured data, to train highly accurate machine learning models, and adopts multi-layer stacked ensemble and times of repetition bagging to optimize the training and validation processes of the model; S45: Construct a geographical similarity model, and perform local adaptation from the top three models in terms of comprehensive scores to construct a similarity model that meets specific geographical environments.

5. A method for predicting the water depth of urban waterlogging considering geographical similarity according to claim 1, characterized in that, S5 Optimization of the urban waterlogging geographical similarity model, including the following steps: S51: Set the hyperparameter search range to construct a hyperparameter pool. S52: Select candidate models from the model library, randomly generate initial model parameters, encode them into strings, and generate new strings based on Bayesian prior probabilities to replace the old partial strings to achieve the creation of new parameters until the performance evaluation indicators are met to complete model optimization. S53: Use the Gaussian process regression method in the Bayesian optimization algorithm to iteratively adjust hyperparameters such as the learning rate parameter, tree depth parameter, regularization coefficient, hidden layer size, dropout rate, number of layers, absolute skewness threshold, number of iterations, and maximum number of epochs, search for the optimal parameter results of each candidate model, and then select the optimal model from multiple candidate models to complete the construction of the optimal geographical similarity model.

6. A method for predicting urban waterlogging depth considering geographical similarity according to claim 5, characterized in that The S6 includes the following steps: S61: Design waterlogging scenarios according to different rainfall time resolution and spatial resolution parameters. S62: Based on the pre-training of the model with the same parameters, use the optimal parameters as prior parameters for the construction of the similarity model. S63: Select different water depth thresholds, different temporal rainfall data, and different calculation units for similarity model training. S64: Water depth prediction calculation. Use the geographical environment and rainfall at the position to be predicted as the input of the similarity model for water depth prediction calculation, and output the predicted water depth result.

7. A method for predicting the water depth of urban waterlogging considering geographical similarity according to claim 5, characterized in that The data preprocessing of S43 includes the following steps: S431: Use the interquartile range of the box plot for outlier checking. If a data point is above the upper quartile and exceeds 1.5 times the distance, or is below the lower quartile and less than 1.5 times the distance, these points are considered outliers; S432: Use the interquartile range detection algorithm to remove outliers. The calculation formula is as follows: (16), represents the interquartile range value, represents the third quartile, represents the first quartile; S433: Remove data features with too high correlation to complete dataset screening. S434: Divide the dataset, observe the distributions of the training set and the test set, and use KDE distribution plots to view and compare the distributions of feature variables in the training set and the test set, meeting the requirements for dataset division.

8. A method for predicting the water depth of urban waterlogging considering geographical similarity according to claim 7, characterized in that, The S63 model training process includes the following steps: S631: Adopt spatial granular grid-based training, select a water depth threshold of 1 mm - 10 mm, compare multiple rainstorm scenarios, and keep the water depth thresholds ranked in the top three in all selected experimental scenarios as the final result. Compare the root mean square error and mean absolute error under different water depth thresholds. The water depth threshold with the largest correlation coefficient and goodness of fit is 5 mm. Finally, select the optimal performance water depth threshold of 5 mm. S632: Select the historical time series rainfall data length to be 1 - 30, compare multiple rainstorm scenarios, and keep the water depth thresholds ranked in the top three in all selected experimental scenarios as the final result. The root mean square error and mean absolute error are maintained between 20 mm - 52 mm and 13 mm - 38 mm respectively for all experimental scenarios and all rainfall time series moments, and the correlation coefficient and goodness of fit are maintained between 0.3 - 0.84 and 0.03 - 0.72 respectively. After comprehensive consideration, it is found that the rainfall time series of the first 26 - 29 moments performs well under multiple indicators. S633: Compare two types of computational units, namely spatial granular grid type and neighborhood type. For the neighborhood type computational unit, the root mean square error, mean absolute error, correlation coefficient, and goodness of fit indicators change in unison. The overall root mean square error and mean absolute error are less than 30 mm and 17 mm respectively, and the correlation coefficient and goodness of fit are greater than 0.85 and 0.7 respectively, meeting the ideal results. For the spatial granular grid type computational unit, the root mean square error and mean absolute error are greater than 30 mm and 18 mm respectively, and the correlation coefficient and goodness of fit are both less than 0.85 and 0.7 respectively. S634: Finally, select a water depth threshold of 5 mm, the rainfall data of the first 26 historical moments, and the neighborhood type computational unit for optimal configuration calculation. The tree model and the ensemble model are superior in performance, with the error stable within 52 mm and the goodness of fit R2 greater than 0.6010.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: When the processor executes the program, it implements a method for predicting urban waterlogging water depth considering geographical similarity as described in any one of claims 1 to 8 above.

Citation Information

Patent Citations

  • Multi-level urban inland inundation coupling simulation method considering climatic elements

    CN115859676A