Urban inland inundation factor identification method based on pipe network model and interpretable machine learning
By combining the MIKE Urban pipeline network model with interpretable machine learning algorithms, a model of urban flooding driving factors was established, which solved the problem that existing technologies could not quantify urban flooding factors. This enabled accurate identification of urban flooding and the formulation of renovation plans, thereby improving the city's flood control capacity and resilience.
Patent Information
- Application Number
- CN202511610526.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2026-02-06
AI Technical Summary
Existing urban flooding models cannot effectively quantify the flooding drivers in each catchment area, nor can they accurately identify the relative contributions of underground pipe networks and surface infiltration capacity, resulting in a lack of quantitative analysis methods for flood control renovation in old urban areas.
By combining the MIKE Urban pipe network model with interpretable machine learning algorithms and using SHAP value analysis, a model of urban flooding driving factors is established. By utilizing factors such as surface runoff, runoff coefficient, catchment area, pipe velocity, pipe flow rate, pipe length, and pipe diameter, an automatic machine learning model is constructed to achieve accurate identification and analysis of urban flooding factors.
It enables precise identification and quantification of urban flooding factors, allowing for the development of targeted renovation plans for different catchment areas, enhancing urban flood control capabilities and resilience, and optimizing the priority of underground drainage systems and surface sponge city renovation.
Smart Images

Figure CN121479986A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of urban stormwater management, specifically to a method for predicting and managing urban flooding by establishing a hybrid model using the MIKE Urban network model combined with machine learning algorithms. Background Technology
[0002] Global warming and accelerated urbanization have led to significant changes in urban climate and hydrological cycles, increasing the risk of flooding. In recent years, many cities in my country have experienced severe urban flooding due to torrential rains. Urban surfaces are becoming increasingly hardened due to the increase in impermeable areas from buildings and roads, making it difficult for rainwater to infiltrate and causing it to become surface runoff, exacerbating the risk of flooding. Therefore, promoting sponge city transformation (especially in old urban areas) to improve urban resilience and flood control capabilities to adapt to climate change and reduce urban flood risk is urgent. Currently, common urban stormwater and flooding models include SWMM and MIKE. SWMM's main functions cover the simulation of surface runoff, water flow changes in pipeline systems, the operation of low-impact development (LID) facilities, and rainwater storage and treatment processes within the study area. MIKE Urban is mainly used to simulate the operation of urban drainage systems and the development process of urban flooding. It can perform detailed modeling and analysis of urban drainage networks and waterways, and accurately simulate the movement and changes of water flow by solving complex hydraulic equations. However, these models cannot effectively quantify the driving factors of urban flooding in each catchment area. Current conventional methods for identifying urban flooding drivers include principal component analysis, Pearson correlation coefficient, multiple stepwise regression, and geographically weighted regression. However, these methods typically assume a linear relationship between the driving factors and urban flooding, ignoring the complex nonlinear relationships between them. This makes it difficult to accurately quantify the relative contribution of each factor to urban flooding. Furthermore, for different catchment areas, which factor—underground drainage capacity or surface infiltration capacity—plays a dominant role in flooding? From the perspective of the effectiveness of flood control, how should the priority be balanced between the renovation of underground pipe networks and the sponge city transformation of surface runoff coefficient? Currently, there is no quantitative analytical method that can accurately identify the primary factors causing urban flooding.
[0003] Against this backdrop, this invention proposes a method for analyzing urban flooding factors based on a combination of the MIKE Urban flooding network model and interpretable machine learning. Specifically, it employs the MIKE Urban hydrodynamic model combined with machine learning algorithms to construct a model of flooding influencing factors through big data fitting. This invention fully considers the complex topography and underground pipe network characteristics of different catchment areas. By coupling a high-precision stormwater model, automated machine learning, and SHAP value analysis, it extracts complex nonlinear relationships between multiple variables from a large amount of structured and unstructured data, achieving accurate identification of regional stormwater processes and dominant flooding factors, and formulating targeted renovation plans for each catchment area. Therefore, this invention not only promotes the organic integration of flooding control and urban master planning but also provides technical support for building an efficient, resilient, and sustainable urban stormwater management system, thereby significantly improving the city's ability to respond to stormwater disasters and ensuring the safety of residents' lives and property and the long-term stability of urban operations. Summary of the Invention
[0004] The present invention aims to provide a method for identifying urban flooding factors based on pipeline network models and interpretable machine learning.
[0005] To achieve the above objectives, the present invention provides the following technical solution: This invention provides a method for identifying urban flooding factors based on a pipeline network model and interpretable machine learning, comprising the following steps: (1) Obtain measured rainfall data or calculate rainfall data according to the local rainstorm intensity calculation formula, import it into the MIKE hydrodynamic model, divide several sub-catchment areas with the inspection well (node) as the center of the Thiessen polygon, and obtain the node water depth and waterlogging driving factor index through simulation. Among them, the waterlogging driving factors include: surface runoff, runoff coefficient, catchment area, pipe velocity, pipe flow rate, pipe length, pipe diameter, and pipe capacity per unit area of catchment area.
[0006] (2) Using the above-mentioned waterlogging driving factors as explanatory variables and the node water depth as the target variable, a regression model is established using a deep machine learning algorithm as a prediction model for node water depth.
[0007] (3) The water depth at the nodes of the inspection wells is predicted and the waterlogging factors in each catchment area are analyzed using the prediction model of the waterlogging driving factors.
[0008] Furthermore, in step (2), the regression model is established using the H2O package in R software.
[0009] Furthermore, in step (3), the methods for analyzing waterlogging factors in each catchment area include: Data on waterlogging drivers were obtained using surface runoff, runoff coefficient, catchment area, pipe velocity, pipe flow rate, pipe length, pipe diameter, and pipe capacity per unit area of the catchment as explanatory variables. These explanatory variables were then substituted into the waterlogging driver model, and the overall and individual catchment area waterlogging factors were analyzed using SHAP values.
[0010] The SHAP value analysis includes bee colony analysis for overall waterlogging factor analysis and single-catchment waterlogging factor analysis. Bee colony analysis assesses the global importance of waterlogging factors by calculating the average absolute SHAP value of all factors, and can explain the positive and negative impacts of waterlogging factors on prediction results under different rainfall conditions. Besides the global interpretation, SHAP can also provide interpretation for individual samples in the dataset. For a specific catchment area, pipe diameter is the feature with the greatest influence. By analyzing the SHAP values of different features, the factors with the greatest impact on node water depth can be identified. In practical stormwater management, this method can analyze the specific situation of each catchment area and select the most appropriate waterlogging solution. For example, in areas where pipe diameter is the main influencing factor, increasing pipe diameter is prioritized. In areas where runoff coefficient is the main influencing factor, measures such as improving the underlying surface permeability coefficient are prioritized.
[0011] The technical solution of the present invention has the following advantages: This invention provides a method for analyzing urban flooding factors based on the MIKE Urban flooding network model combined with interpretable machine learning. Rainfall data can be obtained from measured rainfall or calculated using storm rainfall formulas. This data is then imported into the MIKE hydrodynamic model to obtain flooding driving factor data, which is then used for automated machine learning to build a predictive model. Finally, SHAP interpretability analysis is used to achieve flooding factor analysis. Specifically, flooding driving factors, including surface runoff, runoff coefficient, catchment area, pipe velocity, pipe flow rate, pipe length, pipe diameter, and pipe capacity per unit area of the catchment area, are used as explanatory variables, while node water depth is used as a predictive variable. An automated machine learning algorithm is employed to build the model. Verification shows that the model built using the method provided by this invention has a strong fitting effect and can analyze flooding factors in several catchment areas of a city using only rainfall data. On the one hand, different flood management plans can be designed for different rainfall conditions and different rainfall recurrence periods; on the other hand, it can more accurately analyze the characteristics of rainfall and flooding in each catchment area and formulate the most effective flood control plan for that catchment area. This invention can effectively solve the problem that the primary factors causing flooding cannot be quantified during the flood control renovation of old urban areas. It can more comprehensively consider the complex factors affecting flooding and divide the perspective of flood control into two major systems based on spatial dimensions: above-ground and underground. The above-ground runoff coefficient plays a core dominant role, while the underground pipe diameter becomes a key limiting factor. The pipe diameter, pipe velocity, pipe flow rate, pipe capacity per unit area, and pipe length of the underground drainage pipe network system are coupled with each other to jointly construct a complex dynamic system of underground drainage. Attached Figure Description
[0012] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0013] Figure 1 This involves dividing the old city of Tianjin into several catchment areas; Figure 2 This describes the model training and testing results of each waterlogging driving factor on the node water depth in the embodiments of the present invention. Figure 3 This is a SHAP swarm diagram of various waterlogging drivers; Figure 4 This is a diagram analyzing the factors causing flooding in a single catchment area. Figure 5 This is a diagram analyzing the factors causing flooding in a single catchment area. Detailed Implementation
[0014] The following embodiments are provided to better understand the present invention and are not limited to the preferred embodiments described. They do not constitute a limitation on the content and scope of protection of the present invention. Any product that is the same as or similar to the present invention, derived by any person under the guidance of the present invention or by combining the features of the present invention with other prior art, falls within the protection scope of the present invention.
[0015] Where specific experimental steps or conditions are not specified in the embodiments, they can be performed according to the conventional experimental steps or conditions described in the literature in this field. All instruments used are commercially available conventional products, including but not limited to the instruments used in the embodiments of this application.
[0016] This embodiment provides a method for establishing an analysis model of urban flooding driving factors. The specific steps are as follows: (1) MIKE Urban flooding simulation Taking an old urban area in Tianjin as an example, the rainfall data for a 50-year return period with a duration of 60 minutes was calculated according to the rainfall formula (1), and imported into the MIKE Urban model. The old urban area was divided into 998 catchment areas according to the Thiessen polygon. Figure 1 The simulation obtains the node water depth and flooding situation, and exports the node water depth, surface runoff, runoff coefficient, catchment area, pipe velocity, pipe flow, pipe length, pipe diameter, and pipe capacity per unit area of the catchment area.
[0017] Formula (1) (2) Establishment of automated machine learning models This study utilizes the AutoML algorithm from the H2O package in R for research and analysis, employing the `h2o.automl` function to automatically select machine learning algorithms. To ensure no evaluation bias when selecting the model parameter range, all built-in models and their parameter combinations were searched. The original data was imported as a CSV file as a regular data frame, then converted to an H2O data frame format and input into the H2O AutoML model. The data was split using the built-in H2O function `ratio`. To ensure that all model results are free from randomness and data leakage, the 5x cross-validation built into the H2O package was used to evaluate the model's optimization results. H2O AutoML automatically selects the optimal algorithm; the Gradient Boosting Machine (GBM) algorithm was ultimately chosen and used to construct the initial regression model. The generalization ability of the optimal model was then evaluated based on its R² score on the test set.
[0018] This study used 998 samples, employing surface runoff, runoff coefficient, catchment area, pipe velocity, pipe flow rate, pipe length, pipe diameter, and pipe capacity per unit area of the catchment area as explanatory variables, and node water depth as the predictive variable. After comparing dataset splitting ratios (training set to test set sample size) of 0.6, 0.7, 0.8, and 0.9, a ratio of 0.77 was determined to provide the best predictive performance. A loop with 30 random seeds was used, running 30 times. Each deep learning model developed based on the training set was then used on the test set to verify the model's generalization ability.
[0019] Based on the fitting coefficients (R²) of the training set 2 And use a test set to evaluate the model's performance. R 2 < 0.3, 0.3≤R 2 < 0.4, 0.4≤R 2 < 0.6, 0.6≤R 2 ≤1.0 represents poor, weak, moderate, and strong fit between predicted and measured values, respectively. The deep learning algorithms in the H2O package are implemented by the "h2o" package and the "h2o.gbm" function in R software (v.4.3.3).
[0020] The fitting of predicted and observed values for node water depth based on the training set using a combination of eight waterlogging driving factors (explanatory variables) shows that the training set R... 2 The value is 0.82, and the test set R is... 2 The value is 0.61, indicating a high degree of agreement between the predicted and detected values, achieving a strong fitting effect. Figure 2 ).
[0021] Example 3: Analysis of Inland Flooding Factors The analysis of waterlogging factors using the prediction model (based on the training set) established in Example 1 is as follows: (1) Analysis of overall waterlogging factors by bee colony diagram The importance of factors influencing urban flooding to the overall situation is assessed by calculating the average absolute SAP value of all factors, which can also explain the positive and negative impacts of these factors on the prediction results. The average SAP value of each factor quantifies the global importance of factors influencing urban flooding and elucidates their positive and negative impacts on the prediction results. Figure 3A comprehensive analysis of the urban flooding swarm maps under all rainfall conditions revealed that the runoff coefficient is the primary factor influencing the water depth at nodes throughout the region. The runoff coefficient reflects the surface runoff generation capacity; a higher value indicates a higher proportion of rainfall converted into surface runoff. Impermeable surfaces such as urban roads and plazas have higher runoff coefficients because rainwater has difficulty infiltrating. Conversely, highly permeable surfaces such as grasslands and woodlands have lower runoff coefficients, allowing rainwater to infiltrate the soil in large quantities, reducing surface runoff. When the runoff coefficient is high, rainwater around roads and buildings accumulates rapidly. If drainage pipes cannot keep up with the rate of runoff generation, waterlogging can easily occur. Therefore, for this region, the focus should be on promoting sponge city transformation to improve the surface runoff coefficient, optimizing surface cover and land use structure, and effectively reducing the runoff coefficient through permeable paving, rainwater storage ponds, and other measures to achieve on-site absorption and utilization of rainwater, thereby alleviating pressure on the drainage system.
[0022] Furthermore, pipe diameter is the second most important factor, meaning that the drainage capacity of underground pipe networks is also a key factor affecting urban flooding. The size of the pipe diameter directly determines the cross-sectional area of the drainage pipe. A larger pipe diameter increases the cross-sectional area, allowing the pipe to carry more water, thus more effectively transporting rainwater and other accumulated water, accelerating drainage, and reducing the risk of road flooding. Conversely, if the pipe diameter is small, the drainage pipe can easily reach full capacity quickly during heavy rainfall and when large amounts of surface runoff are generated. Once the rainfall intensity exceeds the pipe's design drainage capacity, rainwater will accumulate inside the pipe, significantly slowing down the drainage speed, and may even backflow from weak points in the pipe or drain outlets onto the ground, leading to urban flooding. At the same time, pipe diameter also has a significant impact on the hydraulic gradient of the water flow. A larger pipe diameter increases the cross-sectional area of the water flow; at a given flow rate, the flow velocity is relatively lower, and the energy consumption of the water flow is reduced, resulting in a smaller hydraulic gradient. A smaller hydraulic gradient means that the water flows more smoothly in the pipe, the drainage process is smoother, and drainage efficiency is improved. If the pipe diameter is too small, the hydraulic gradient will be too large, especially in long-distance drainage pipelines, which may cause the water level at the end of the drainage point to be too high, hindering the timeliness of drainage. For example, in some flat urban areas, if the pipe diameter is not designed properly and the hydraulic gradient is too large, rainwater in some areas will not be able to drain in time, thus causing local flooding. Therefore, the rational design of pipe diameter is of great significance for ensuring the effective drainage function of underground pipe networks and mitigating the risk of urban flooding.
[0023] Meanwhile, pipe flow velocity is the third most important influencing factor. Increasing pipe flow velocity can improve drainage efficiency and reduce surface rainwater accumulation. For example, some old urban area renovation projects use underground pipe networks with large diameters and reasonable slopes, increasing pipe flow velocity and enabling rapid drainage during heavy rains, shortening the time of road flooding. The negative pressure effect generated by high flow velocity can enhance the water collection capacity of storm drains, accelerating the flow of surface rainwater into the underground pipe network and further reducing surface runoff accumulation. However, forcibly increasing pipe diameter or other methods to accelerate pipe flow velocity in low-lying areas or areas with poorly designed drainage networks can cause rainwater to rush into local drainage pipes in a short period, exceeding their drainage capacity. Because rainwater cannot be drained in time, local water accumulation occurs, exacerbating urban flooding. Furthermore, excessively fast flow velocities cause rainwater to flow rapidly downstream of the drainage network, putting enormous pressure on downstream drainage facilities such as pumping stations and waterways. If the downstream drainage system cannot handle the large influx of water in time, it may lead to poor drainage or even backflow, reducing the overall drainage efficiency of the drainage system and exacerbating urban flooding.
[0024] (2) Explanatory analysis of waterlogging factors in a single catchment area In addition to providing a global interpretation, SHAP can also provide an interpretation for individual samples within the dataset. Each observation can have its own SHAP value. Dividing the old city area into 998 samples, each corresponding to a small catchment area, SHAP value analysis of individual catchment areas reveals that when the surface runoff coefficient dominates, surface runoff reduction measures based on the sponge city concept should be prioritized, such as laying permeable pavements, modifying sunken green spaces, and adding rainwater storage tanks. Conversely, when pipe diameter becomes a critical factor, increasing pipe diameter to enhance drainage capacity is the priority for addressing urban flooding.
[0025] Taking catchment area No. 5 as an example, the SHAP value of the runoff coefficient is as high as +0.207, making it the dominant explanatory variable among the drivers of urban flooding, and it shows a significant positive correlation with the water depth at the predictor node. This indicates that an increase in the runoff coefficient directly exacerbates the severity of urban flooding. Meanwhile, the proportion of pipe diameter, another driver of urban flooding, is negligible. Therefore, this catchment area should prioritize the implementation of sponge city technologies, including permeable paving, installation of rainwater storage tanks, and construction of green roofs. Figure 4 Taking catchment area No. 6 as an example, the SHAP value of pipe diameter is -0.37, making it the largest driver of flooding among the eight explanatory variables. Pipe diameter is negatively correlated with node water depth; the larger the pipe diameter, the shallower the node water depth. Meanwhile, the SHAP value of runoff coefficient is +0.255, making it the second largest driver of flooding in this catchment area, but its influence is slightly less than that of pipe diameter. Therefore, for this catchment area, increasing pipe diameter should be prioritized, while reducing the runoff coefficient should be used as a supplementary measure to achieve more efficient flood prevention. Figure 5 ).
[0026] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. A method for identifying urban flooding factors based on a pipeline network model and interpretable machine learning, characterized in that... Includes the following steps: Step 1: Collect measured rainfall data in the city or calculate rainfall data for various scenarios using the city's rainfall formula, and import land use and underground pipe network data into the MIKE software; The city is divided into several catchment areas using the Thiessen polygon method. Rainfall flooding simulation is carried out, and the water depth of the inspection well is obtained as the node water depth. At the same time, the variable data affecting the node water depth are obtained, including: surface runoff, runoff coefficient, catchment area, pipe velocity, pipe flow rate, pipe length, pipe diameter, and pipe capacity per unit area of the catchment area. Step 2: Using node water depth as the target variable, and surface runoff, runoff coefficient, catchment area, pipe velocity, pipe flow rate, pipe length, pipe diameter, and pipe capacity per unit area of the catchment area as explanatory variables, an automatic machine learning method is used to establish a regression model as a predictive analysis model for node water depth. Step 3: Analyze the model using the SHAP interpretability machine learning method to identify the primary factors causing flooding in each catchment area.
2. The urban flooding factor identification method based on pipeline network model and interpretable machine learning according to claim 1, characterized in that, In step 2, the regression model is established using the H2O package in R software.
3. The urban flooding factor identification method based on pipeline network model and interpretable machine learning according to claim 2, characterized in that, In step 3, the analysis method for the urban flooding factors includes: performing SHAP analysis using a regression model established with the H2O package. It can identify the primary factors affecting overall urban flooding under the current rainfall conditions; It can also identify the primary factors causing flooding in each catchment area and propose the most effective stormwater management measures for each catchment area.
4. The urban flooding factor identification method based on pipeline network model and interpretable machine learning according to claim 3, characterized in that, In step 1, the water depth data of the 8 influencing nodes obtained by the MIKE hydrodynamic model can be divided into two major systems, above ground and underground, according to the spatial dimension. In the above-ground space, the runoff coefficient occupies a core position because it accurately quantifies the surface runoff generation capacity; In underground spaces, the diameter of drainage pipes directly determines the cross-sectional area of the drainage pipes, which is related to the upper limit of the drainage capacity of the pipeline system.
5. The urban flooding factor identification method based on pipeline network model and interpretable machine learning according to claim 4 further includes: The target variable and its corresponding explanatory variables are randomly divided into a training set and a test set. A prediction model is built using the training set, and the prediction ability of the prediction model is verified using the test set. 60% of the dataset is used as the training set, and 40% of the dataset is used as the test set.
6. The urban flooding factor identification method based on pipeline network model and interpretable machine learning according to claim 5, characterized in that, The predictive power of the prediction model is measured by the fitting coefficient R². When R² ≤ 0.3, the predicted values fit the observed values poorly, indicating poor predictive ability of the prediction model. When 0.3 < R² ≤ 0.4, the predicted value fits the observed value poorly, and the prediction ability of the prediction model is weak. When 0.4 < R² ≤ 0.6, the predicted value fits the observed value moderately, and the prediction ability of the prediction model is moderate. When 0.6 < R² ≤ 1.0, the predicted value fits the observed value well, and the prediction model has strong predictive ability.