Lightgbm and hydrological and hydrodynamic model based rapid flood prediction method for coastal cities

By combining LightGBM and hydrodynamic models, optimizing hyperparameters, and constructing a flood prediction model, the problems of long computation time and high data requirements in existing technologies are solved, enabling rapid and accurate prediction of urban floods.

CN116090625BActive Publication Date: 2026-05-05ZHENGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHENGZHOU UNIV
Filing Date
2023-01-01
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing hydrological and hydrodynamic models are computationally time-consuming and data-intensive in urban flood prediction, making it difficult to meet real-time prediction needs. Data-driven models are limited in application in areas lacking real-world data, and existing combined methods consume a lot of memory and have slow training speeds.

Method used

A method combining LightGBM and hydrodynamic models was adopted. A flood simulation model was constructed using PCSWMM, feature variables were extracted and LightGBM model was trained, and hyperparameters were optimized using K-fold cross-validation and grid search to achieve rapid prediction of urban floods.

Benefits of technology

It achieves accurate and rapid flood forecasting over large spatial areas or high-resolution regions, reduces memory consumption and improves training speed, and is suitable for real-time flood forecasting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116090625B_ABST
    Figure CN116090625B_ABST
Patent Text Reader

Abstract

The application discloses a coastal city flood rapid prediction method based on LightGBM and a hydrological and hydrodynamic model, and comprises the following steps: 1, obtaining sample data, including: designing multiple rainfall-tide level combination scenarios; constructing a flood simulation model of a target region based on a PCSWMM hydrological and hydrodynamic model; simulating the maximum water depth of each selected flood point under different rainfall-tide level combination scenarios by using the flood simulation model; extracting rainfall-related characteristic variables and tide-related characteristic variables, and taking the flood point position and the maximum water depth as sample data; 2, constructing and training a city flood rapid prediction model based on LightGBM; 3, predicting the maximum water depth of a submerged point in city flood by using the trained city flood prediction model. The application has excellent performance and calculation efficiency, and can realize accurate and rapid prediction of city flood. The application can also realize efficient and accurate prediction of a city region with a large space range or high spatial resolution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of urban flood early warning technology, and in particular relates to a rapid flood prediction method for coastal cities based on LightGBM and hydrodynamic models. Background Technology

[0002] Currently, urban flood prediction models can be divided into physical models and data-driven models. Physical models mainly refer to hydrological and hydrodynamic models, whose technology is relatively mature and computationally accurate. In recent years, with the advancement of numerical simulation technology, hydrological and hydrodynamic models have been widely used in the field of urban flood simulation. However, the complexity and long computation time of hydrological and hydrodynamic models limit their development, especially in urban areas with large spatial areas or high resolution requirements, where the computation time will be even longer, making it difficult to meet the needs of real-time flood prediction. In addition, hydrological and hydrodynamic models have high requirements for flood monitoring data, requiring calibration of a large number of parameters to accurately describe the flood process. The severe lack of flood monitoring data has affected their promotion and application.

[0003] Besides physical models, data-driven models are also widely used in urban flood forecasting. Data-driven models learn from large amounts of data to identify patterns between inputs and outputs, directly providing the model's output based on the inputs, thus enabling rapid flood prediction. However, data-driven models require a large amount of flood data to build effective flood prediction models. In many regions, the application of data-driven models is limited due to the scarcity of measured data.

[0004] One feasible solution is to combine physical models and data-driven models, using physical models to simulate large amounts of flood data to train data-driven models, thereby replacing physical models for rapid flood prediction. For example, Kabir... [1] Combining LISFLOOD-FP and convolutional neural networks (CNN) for rapid prediction of river flood depth. (Berkhahn) [2] A flood map database was generated using HYSTEM-EXTRAN 2D simulation, and a rapid urban flood prediction model based on an ensemble neural network was constructed. Lowe [3] A method combining MIKE and convolutional neural networks is proposed for predicting flood depth in urban rivers. However, the gradient boosting tools in the aforementioned studies use pre-sorting-based algorithms and a hierarchical growing tree strategy, which require one-hot encoding to represent class features, increasing memory consumption and reducing training speed.

[0005] The following references are used in this article:

[0006] [1] Kabir S, Patidar S, Xia X, et al. A deep convolutional neural network model for rapid prediction of fluvial flood inundation [J]. Journal of Hydrology, 2020, 590, 125481.

[0007] [2] Berkhahn S, Fuchs L, Neuweiler I. An ensemble neural networkmodel for real-time prediction of urban floods [J]. Journal of Hydrology, 2019, 575: 743-754.

[0008] [3] Löwe R, Böhm J, Jensen DG, et al. U-FLOOD – Topographic deeplearning for predicting urban pluvial flood water depth [J]. Journal ofHydrology, 2021, 603, 126898. Summary of the Invention

[0009] In view of this, this application proposes a rapid urban flood prediction method based on LightGBM and hydrodynamic models to further reduce memory consumption and improve training speed.

[0010] The technical solution of this application is implemented as follows:

[0011] This application proposes a rapid flood prediction method for coastal cities based on LightGBM and hydrodynamic models, including:

[0012] I. Obtaining sample data, including: designing various rainfall-tide combination scenarios; constructing a flood simulation model for the target area based on the hydrodynamic model of PCSWMM; simulating the maximum water depth at each selected flood point under different rainfall-tide combination scenarios using the flood simulation model; extracting feature variables related to rainfall and tide level, and using them together with the location of the flood point and the maximum water depth as sample data;

[0013] II. Construction and training of urban flood prediction models, including: construction of an urban flood prediction model based on LightGBM, and construction of an urban flood prediction model trained with sample data;

[0014] 3. Predict the maximum water depth at inundation points during urban flooding using a trained urban flood prediction model.

[0015] In some specific implementations, the rainfall-related characteristic variables include one or more of the following: cumulative rainfall, rainfall return period, peak rainfall, maximum 2-hour rainfall, maximum 3-hour rainfall, and cumulative rainfall before the peak; and the tide-related characteristic variables include one or more of the following: maximum tide level, tide return period, average tide level, and maximum 5-hour average tide level.

[0016] In some specific implementations, various rainfall-tide combination scenarios are designed, including:

[0017] A generalized extreme value function (GEM) is used to fit the frequency distributions of rainfall and tide. The parameters of the GEM, including the location, scale, and shape parameters of rainfall and tide distribution, are estimated using the maximum likelihood method. Various design rainfall and peak tide values ​​with return periods ranging from 5 to 100 years are obtained by calculating the GEM. Typical rainfall and tide processes are scaled using the same scaling method to obtain design rainfall and tide processes with different return periods. Multiple rainfall-tide combination scenarios are combined to design these scenarios, which serve as boundary conditions for the flood simulation model.

[0018] In some specific implementations, a flood simulation model for the target area is constructed based on the hydro-hydrodynamic model of PCSWMM, including:

[0019] First, the data collection includes digital elevation model, drainage network, one-dimensional manholes, sub-catchment areas, and obstacle counts. Then, a one-dimensional drainage model of manholes and pipes is constructed to obtain two-dimensional nodes and a two-dimensional network. Next, a two-dimensional surface inundation model is constructed based on the two-dimensional nodes and the two-dimensional grid. Finally, the one-dimensional and two-dimensional models are coupled by using bottom orifice connections to obtain the flood simulation model.

[0020] In some specific implementations, the method further includes: selecting actual measured rainfall and tide sequence data from the target area to calibrate the constructed flood simulation model.

[0021] In some specific implementations, the urban flood prediction model constructed using sample data includes:

[0022] Several flood-prone points were selected from the flood-prone areas, and several feature variables related to rainfall and tide level were extracted. The locations of the flood-prone points and the extracted feature variables were input into the flood simulation model to simulate the maximum water depth of each flood-prone point under different rainfall-tide combination scenarios. The LightGBM model was trained by using different rainfall-tide combination scenarios and flood-prone point locations as inputs and the maximum water depth of each flood-prone point as outputs.

[0023] In some specific implementations, the hyperparameters of the LightGBM model are optimized using K-fold cross-validation and grid search, specifically as follows:

[0024] Define the range and interval of hyperparameters for the LightGBM model. The range and interval of hyperparameters form a three-dimensional hyperparameter mesh. Use a mesh search method to traverse the hyperparameter mesh. Use K-fold cross-validation on each hyperparameter mesh to select the hyperparameter with the best generalization ability.

[0025] K-fold cross-validation involves: randomly and uniformly dividing the sample data into k subsets; using k-1 subsets for training the model each time it is trained and testing, and using the remaining subsets for testing; repeating this process k times, and averaging the results of the k tests as the model evaluation value.

[0026] In some specific implementations, the hyperparameters of the LightGBM model include the learning rate, the number of estimators, and the number of leaf nodes.

[0027] In some specific implementations, the contribution of the input feature variables of the LightGBM model to the prediction of urban flooding in the target area is also analyzed using the Gini index.

[0028] This application has the following beneficial effects on the prior art:

[0029] The method proposed in this application has excellent performance and computational efficiency, enabling accurate and rapid prediction of urban flooding. It can also achieve efficient and accurate prediction for urban areas with large spatial ranges or high spatial resolution. Attached Figure Description

[0030] Figure 1 A flowchart illustrating a specific implementation of this application;

[0031] Figure 2 This is a schematic diagram illustrating the principle of discretizing continuous feature values ​​in a histogram-based decision tree.

[0032] Figure 3 This is a schematic diagram illustrating the principle of 10-fold cross-validation.

[0033] Figure 4 This is a schematic diagram illustrating the principle of the grid search method.

[0034] Figure 5 The diagrams are for the growth strategy, where Figure (a) shows a schematic diagram of the hierarchical growth tree and Figure (b) shows a schematic diagram of the leaf-by-leaf growth strategy.

[0035] Figure 6 This is a map showing the geographical location of Haikou City.

[0036] Figure 7 Map of Haidian Island, Haikou City;

[0037] Figure 8 The data provided are DEM images of the study area in this example.

[0038] Figure 9 This is one-dimensional inspection well and pipeline network data for the study area in the embodiment;

[0039] Figure 10 This is a schematic diagram illustrating the division of the study area into several sub-catchment areas in the embodiment;

[0040] Figure 11 This is a schematic diagram of the water-blocking obstacles in the study area of ​​the embodiment;

[0041] Figure 12 The example shows the rainfall and tide sequence for the study area during Typhoon Rammasun;

[0042] Figure 13 These are design rainfall events with different return periods in the examples;

[0043] Figure 14 The examples illustrate design tidal levels with different return periods.

[0044] Figure 15 The locations of measured flood points in the study area are shown in the examples.

[0045] Figure 16 Figures (a)-(e) show the maximum water depth predicted by the LightGBM model under different return periods; where the return periods are 5 years, 20 years, 35 years, 75 years and 100 years, respectively.

[0046] Figure 17 The examples show the prediction results of LightGBM at flood points under different return periods.

[0047] Figure 18 The example shows the spatial distribution of PCSWMM simulated water depth and LightGBM predicted water depth in the study area when the return periods for both rainfall and tide are 5 years.

[0048] Figure 19The spatial distribution of PCSWMM simulated water depth and LightGBM predicted water depth in the study area is shown in the example where the return periods for both rainfall and tide level are 20 years.

[0049] Figure 20 The spatial distribution of PCSWMM simulated water depth and LightGBM predicted water depth in the study area is shown in the example where the return periods for both rainfall and tide level are 35 years.

[0050] Figure 21 The example shows the spatial distribution of PCSWMM simulated water depth and LightGBM predicted water depth in the study area when both rainfall and tide return periods are 50 years.

[0051] Figure 22 The spatial distribution of PCSWMM simulated water depth and LightGBM predicted water depth in the study area is shown in the example where the return periods for both rainfall and tide are 100 years.

[0052] Figure 23 The relative importance of urban flooding causative factors based on LightGBM in the embodiments;

[0053] Figure 24 The example shows the prediction results of RF, XGBoost, and KNN for 10 flood points under different return periods.

[0054] Figure 25 The scatter plots of the predicted and actual water depths for the LightGBM, RF, XGBoost, and KNN models on the test set are shown in the examples.

[0055] Figure 26 In this example, the simulated water depth of PCSWMM at the flooded location during Typhoon Rammasun is compared with the predicted water depth of LightGBM, RF, XGBoost, and KNN. Detailed Implementation

[0056] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0057] See Figure 1The flowchart shown illustrates a specific implementation of this application. First, various rainfall-tide combination scenarios are designed, and a one-dimensional coupled hydrodynamic model based on PCSWMM is constructed to simulate flood inundation distribution under different scenarios in the study area. Next, ten variables are extracted from the rainfall and tide sequences, and along with the location of flood points, are used as input variables for a machine learning model. A LightGBM model optimized using a grid search algorithm and K-fold cross-validation is employed to predict the maximum water depth at inundation points in urban flooding. Then, the constructed LightGBM model is evaluated and compared with other machine learning models on a test set based on evaluation metrics. Finally, the relative importance of the input variables for urban flood prediction is determined by calculating the Gini index.

[0058] The models involved in the embodiments of this application will be described in detail below.

[0059] 1. PCSWMM model

[0060] PCSWMM is a hydrodynamic model developed by the Computational Hydraulics Institute of Canada (CHI) based on the SWMM model. This application constructs a flood simulation model based on PCSWMM to provide data for machine learning models. The PCSWMM model includes surface runoff calculation, flow calculation, and one- and two-dimensional coupled calculations.

[0061] The surface runoff calculation uses the following Manning formula:

[0062] (1)

[0063] In equation (1), For outflow ( ); Width of the sub-catchment area (m); The Manning coefficient for the Earth's surface; Indicates the maximum depth of the depression (m); The average slope of the catchment area.

[0064] The surface runoff calculation using the nonlinear reservoir method involves solving the continuity equation and the Manning equation simultaneously. The continuity equation is as follows:

[0065] (2)

[0066] In equation (2), The total water volume of the catchment area ( ); For water depth ( ); For time ( ); For the area of ​​the sub-catchment area ( ); Net rainfall intensity ( ); Surface runoff ( ).

[0067] The flow rate calculation uses the dynamic wave method, and the continuity equation and momentum equation are as follows:

[0068] (3)

[0069] (4)

[0070] In equations (3)-(4): For traffic ( ); Distance ( ); The cross-sectional area of ​​the water passage ( ); For time ( ); For water depth ( ); The acceleration due to gravity ( ); This refers to the friction slope.

[0071] 2. LightGBM model

[0072] LightGBM is a machine learning framework based on Gradient Boosting Decision Trees (GBDT). It uses a histogram-based decision tree learning method, where continuous feature values ​​are discretized into k integers and stored in a histogram of width k. This method offers significant advantages in terms of training speed and memory consumption. See also... Figure 2 The figure shows the discretization of continuous feature values ​​in a histogram-based decision tree. Figure 3 The diagram illustrates 10-fold cross-validation, which randomly divides the data into k subsets, with k-1 subsets used for training and the remaining subset used as the test set. Figure 4 The grid search method is illustrated, which traverses all combinations of hyperparameters within a given search range to determine the optimal set of hyperparameters. This application further optimizes LightGBM by modifying the tree growth strategy, building upon the histogram method. Unlike most tree-based learning algorithms that use a hierarchical tree growth strategy, LightGBM employs a depth-constrained leaf-by-leaf tree growth strategy. The depth constraint refers to limiting the tree's growth depth by specifying the `max_depth` parameter of LightGBM.

[0073] See Figure 5 ,in Figure 5(a) shows a schematic diagram of a hierarchical growth tree. The hierarchical growth tree strategy can split leaf nodes at the same level as the decision tree grows, but the information gain of split leaf nodes at the same level is different. Processing leaves with low information gain has little effect on improving prediction accuracy and also increases memory consumption. Figure 5 (b) is a schematic diagram of the leaf-by-leaf growth strategy. The leaf-by-leaf growth tree only splits the leaf with the largest information gain, resulting in a lower model loss, but it may lead to overfitting. Therefore, the maximum depth of the tree should be limited.

[0074] In addition, LightGBM supports direct input of categorical features. Most other machine learning algorithms require one-hot encoding to represent categorical features, which is not optimal for tree-based learning algorithms, leading to sparse data and reduced computational efficiency.

[0075] In summary, LightGBM is a gradient boosting decision tree algorithm. The histogram-based approach reduces training time and memory usage. The leaf-by-leaf tree growth strategy improves model accuracy. LightGBM supports direct input of categorical features, simplifying data preprocessing.

[0076] 3. K-fold cross-validation and web search

[0077] For a given machine learning algorithm, its fit to the training set and its prediction accuracy on the test set are mainly determined by its hyperparameters. For example, the learning rate, number of estimators, and leaf node tree in LightGBM directly affect the reliability of the model. To maximize the performance of machine learning, this application's embodiments are based on the open-source Scikit-learn machine learning framework, using K-fold cross-validation and grid search algorithms to optimize the hyperparameters of LightGBM.

[0078] K-fold cross-validation is a non-exhaustive cross-validation method that can effectively evaluate the generalization ability of machine learning. For example... Figure 3 As shown, K-fold cross-validation randomly and uniformly divides the dataset into k subsets. During each training and testing phase, k-1 subsets are used for model training, with one subset reserved as the test set. This process is repeated k times. The final evaluation value is the average of the k test results. Typically, k is set to 10, meaning 10-fold cross-validation is used as the performance evaluation method during model training.

[0079] Grid search is an exhaustive algorithm widely used for hyperparameter optimization in machine learning models. (See...) Figure 4Grid search can traverse different combinations of hyperparameters within a given range of hyperparameters. The basic idea of ​​the grid search strategy is to search for the optimal set from the 1st to the nth dimension, keeping the values ​​of other dimensions unchanged when searching for the optimal solution in the i-th dimension. Then, the performance of the model under different sets of hyperparameters is compared using K-fold cross-validation to determine the model hyperparameters that are best suited for the current training data.

[0080] The following will use a specific application case to illustrate the specific implementation scheme and its technical effects in detail. Haikou City, as the capital of Hainan Province, is located at the northernmost tip of Hainan Island. It has a tropical monsoon climate with a relatively high average annual temperature. Frequent typhoons result in abundant rainfall. The majority of annual precipitation occurs from May to October, with a multi-year average rainfall of 1827 mm. The study area, Haidian Island, is located in the northern part of Haikou City, separated from the main urban area by two rivers. Figure 6-7 The island has an area of ​​approximately 14 km². 2 Haidian Island is the largest island in Haikou City. The city's underground drainage system and waterways together form Haidian Island's flood control and drainage system. However, due to Haidian Island's flat terrain and coastal location, the seawater has a backwater effect on the drainage outlets, leading to poor drainage and a tendency for seawater to backflow.

[0081] I. Obtaining Sample Data

[0082] This embodiment uses Haidian Island as the research object. First, a one-to-two-dimensional coupled flood simulation model based on PCSWMM is constructed. The required data includes a digital elevation model (DEM), drainage network, one-dimensional manholes, sub-catchment areas, and obstacle data. The DEM data comes from the Haikou Municipal Water Resources Bureau's elevation data for Haidian Island, see [link to relevant documentation] Figure 8 Data on drainage pipe networks and one-dimensional inspection wells were provided by the Haikou Municipal Water Resources Bureau. The data was preprocessed using ArcGIS software to obtain the one-dimensional inspection wells and pipe networks of the study area, as shown below. Figure 9 Based on the river system within the study area, each river section was generalized into different open channels. Considering the topographical features, drainage pipelines, and road layout of Haidian Island, the study area was divided into several sub-catchments, as shown in [reference needed]. Figure 10 The water-blocking obstacles are mainly buildings within the urban area. By identifying buildings in the urban satellite map, the water-blocking obstacles in the study area were identified, as shown in [see...]. Figure 11 During Typhoon Rammasun in July 2014, measured water depth data were obtained through field surveys and provided by the Haikou Municipal Water Resources Bureau. The measured water depth data was used to calibrate a one-dimensional coupled flood simulation model.

[0083] After constructing a flood simulation model using the above data, boundary conditions need to be added to the model to simulate the distribution of flood inundation data. The generalized extreme value (GEV) function is used to fit the distribution of rainfall and tide level, and the parameters of the GEV function are estimated using the maximum likelihood method. The generalized extreme value function is as follows:

[0084] (5)

[0085] In the formula: For position parameters, For scale parameters, For shape parameters.

[0086] For the rainfall distribution on Haidian Island, location parameters = 122.51, scale parameter = 46.78, shape parameter = 0.115. For tidal distribution, location parameter... = 2.349, scale parameter = 0.380, shape parameter = -0.048. Then, by calculating the GEV function, seven design rainfall amounts and peak tide levels with return periods ranging from 5 to 100 years were obtained.

[0087] Considering the worst-case scenario, the rainfall event on July 18, 2014, was selected as a typical rainfall and tidal event. (See...) Figure 12 Finally, because the proportional scaling method is simple to calculate and can preserve the shape of typical rainfall and tidal processes, it is used to scale typical rainfall and tidal processes using the same ratio, which is determined by the rainfall amount and peak tidal level. The ratio values ​​obtained using the proportional scaling method are listed in Table 1. Therefore, the temporal distributions of design rainfall and tidal levels with different return periods are obtained, see... Figure 13 and Figure 14 In this embodiment, both rainfall and tide durations are 24 hours, with a time resolution of 1 hour. The rainfall and tide data are combined to obtain 49 scenarios, which serve as boundary conditions for the flood simulation model.

[0088] To construct a coupled one-dimensional flood simulation model, a one-dimensional drainage model containing 2667 manholes and 2042 pipes was first built. Two-dimensional nodes and a two-dimensional mesh were obtained from the digital elevation model (DEM) using the two-dimensional modeling function of PCSWMM. Then, a two-dimensional surface inundation model was constructed based on the obtained two-dimensional nodes and the two-dimensional mesh. Finally, the one-dimensional and two-dimensional models were coupled by connecting the bottom orifices.

[0089] The model uses a hexagonal grid with a spatial resolution of 25m. In this embodiment, the flow calculation uses the fully dynamic wave method, and the infiltration model is the Horton model.

[0090] Table 1. Ratios obtained using the same multiple method

[0091]

[0092] Table 2 Comparison of observed and simulated water depths at Haidian Island

[0093]

[0094] To improve the accuracy of the constructed one-dimensional coupled model, the model was further calibrated and validated using actual measured rainfall and tide sequences from the time Typhoon Rammasun made landfall on Haidian Island in July 2014. The location distribution of the measured flood points is shown in [reference needed]. Figure 15 The comparison results between the simulated and measured water depths are shown in Table 2. As can be seen from Table 2, the simulated water depth values ​​are very close to the measured water depth values. 75% of the data have a relative error of less than 10%, and 87.5% have a relative error of less than or equal to 15%. The NSE value is 0.725, indicating that the model can effectively simulate the inundation situation in the study area.

[0095] II. Construction and Training of Urban Flood Prediction Models

[0096] Ten flood-prone locations were selected within the study area. These locations, provided by the Haikou Municipal Water Resources Bureau, are flood-prone areas evenly distributed across Haidian Island. Ten input variables were extracted from rainfall and tide data. The six rainfall-related variables were: cumulative rainfall, rainfall return period, peak rainfall, maximum 2-hour rainfall, maximum 3-hour rainfall, and cumulative rainfall before the peak. The four tide-related variables were: maximum tide level, tide return period, average tide level, and maximum 5-hour average tide level. The location of the flood-prone locations was also included as an input variable to represent different locations. The maximum water depth at each flood-prone location simulated under different rainfall-tide combination scenarios was used as the output variable of the LightGBM model.

[0097] Table 3 Inputs and Outputs of LightGBM and PCSWMM

[0098]

[0099] To increase data diversity and prevent model overfitting, Gaussian noise (0.01) was added to the data. After adding Gaussian noise, the dataset was Z-score normalized to eliminate the influence of different data units. In summary, PCSWMM was used to simulate 49 rainfall-tide combinations, resulting in a dataset containing 490 samples, each with 11 input variables and 1 output variable. Table 3 lists the inputs and outputs of LightGBM and PCSWMM.

[0100] Before building the LightGBM model, the dataset was divided into training, validation, and test sets. For the test set, five representative rainfall-tide combinations were selected, with the same return periods for rainfall and tides: 5 years, 20 years, 35 years, 75 years, and 100 years. The remaining data served as the training set. In other words, the training set contained 440 data points, and the test set contained 50.

[0101] The hyperparameters of a machine learning model affect its predictive performance, and optimizing these hyperparameters is beneficial for accurate predictions. This embodiment uses 10-fold cross-validation and a grid search algorithm to optimize the hyperparameters of the LightGBM model within a certain range.

[0102] The three hyperparameters of the LightGBM model were used as the optimization targets of the grid search algorithm. The learning rate (default value) is 0.1, which controls the gradient descent rate of the machine learning model. If the learning rate is too small, the training time will increase significantly; conversely, if the learning rate is too large, the model accuracy will decrease. The number of estimators (default value) is 100. Too many estimators can lead to overfitting, while too few estimators can result in insufficient model training. The number of leaves (default value) is 31. Too many leaves significantly increase the risk of overfitting, while too few leaves can lead to large model training errors. To optimize the hyperparameters of LightGBM, the range of hyperparameters for the model was first defined. These ranges and intervals of hyperparameters form a three-dimensional hyperparameter grid. The grid search algorithm was used to traverse each node in the hyperparameter grid. A 10-fold cross-validation method was used to evaluate the generalization ability of each hyperparameter combination based on the root mean square error (RMSE), and then the hyperparameter set with the best generalization ability was selected, as shown in Table 4.

[0103] Table 4. Optimization range and optimal values ​​of LightGBM hyperparameters based on grid search method.

[0104]

[0105] This embodiment uses Python 3.7 as the program and the web-based Google Colabnotebook as the platform. Based on the optimal hyperparameter set, a LightGBM model is constructed to make predictions on the test set, obtaining the maximum water depth at each inundation point. Figure 16 The maximum water depth predicted by the LightGBM model during the testing phase is shown. The prediction results are evenly distributed near the straight line, demonstrating excellent prediction performance. Figure 17The results show the predictions made by the LightGBM model for 10 inundation points on the test set, demonstrating strong agreement between the model predictions and the actual results. Furthermore, using rainfall and tide levels during Typhoon Rammasun as input to the LightGBM model, the predicted water depths at the 10 flooded points were obtained. The same rainfall and tide levels were then input into the calibrated PCSWMM model to obtain the simulated water depths.

[0106] To accurately evaluate the model's performance at each stage, the constructed LightGBM model was evaluated using root mean square error (RMSE), mean absolute percentage error (MAPE), sum of squared errors (SSE), and Nash efficiency coefficient (NSE). The evaluation results are shown in Table 5. Table 5 shows that the LightGBM model is feasible for water depth prediction, and the overfitting is acceptable. The distribution of PCSWMM simulated water depth and LightGBM model predicted water depth in the study area on the test set is shown in Table 5. Figure 18-22 As shown, the simulated water depth is very close to the predicted water depth.

[0107] Table 5 Predictive capabilities of LightGBM

[0108]

[0109] To test whether the LightGBM model can achieve real-time flood prediction, Table 5 shows the runtime required for LightGBM and PCSWMM to predict water depth under the same scenario. For the LightGBM model, it only takes 35 seconds (including model training time) to output the water depth at the flood point. Because they use different computing devices, a completely fair comparison of the computational speeds of the two models is difficult. However, the values ​​in Table 6 still demonstrate that the LightGBM model is computationally efficient and can replace the physical model.

[0110] Table 6 Runtime for LightGBM and PCSWMM to Predict Maximum Water Depth

[0111]

[0112] Furthermore, the importance of input variables for coastal urban flooding was analyzed by calculating the Gini index. The results are shown in […]. Figure 23The results showed that the maximum tide level was the most important flood-causing factor with a significance of 14.62%, followed by the average tide level (12.48%), the tide return period (11.61%), and the maximum 5-hour average tide level (11.36%), all with relative importance greater than 10%. The remaining seven flood-causing factors all had relative importance greater than 5%, namely: peak rainfall (8.38%), cumulative rainfall (8.09%), location of the flood point (7.93%), maximum 3-hour rainfall (7.23%), cumulative rainfall before the peak (7.02%), rainfall return period (6.20%), and maximum 2-hour rainfall (5.07%). The results indicate that the relative importance of flood-causing factors related to tide level is higher than that related to rainfall. This is because Haidian Island is surrounded by the sea and is significantly affected by tidal backwater, thus increasing the risk of urban flooding.

[0113] To verify the superiority of the LightGBM model, the prediction values ​​and performance of LightGBM were compared with those of RF, XGBoost, and KNN. Predictions were made for each model on the test set, and the predicted water depths at each inundation point were plotted. The prediction results of RF, XGBoost, and KNN are shown below. Figure 24 As shown in the figure. It also demonstrates good prediction performance for other inundation points. The scatter plots of predicted and actual water depths for LightGBM, RF, XGBoost, and KNN across the entire test set are shown in the figure. Figure 25 Intuitively, the scatter distribution concentration of LightGBM is similar to that of RF and XGBoost, while the data points of the KNN model are the most dispersed. Furthermore, this study trained LightGBM, RF, XGBoost, and KNN models on the same training set and tested the performance of each model at each stage. The calculated evaluation metrics are shown in Table 7. The results show that LightGBM outperforms the other models in all metrics during both the training and testing phases.

[0114] Table 7 Comparison of evaluation index values ​​for each model in the prediction phase

[0115]

[0116] To compare the predictive performance of each machine learning model during Typhoon Rammasun, rainfall and tide levels during the typhoon were used as input conditions to obtain the predicted water depth for each machine learning model, which was then plotted together with the simulated water depth from PCSWMM. Figure 26 The results show that the LightGBM model has the best prediction performance in this scenario.

[0117] To rapidly predict the maximum water depth at flood points in urban flooding, this application proposes a rapid urban flood prediction model based on a one-dimensional coupled hydrological and hydrodynamic model, combined with the LightGBM model and a grid search algorithm. Forty-nine rainfall-tide combination scenarios were simulated using PCSWMM to provide data for the LightGBM model. Eleven feature variables were extracted from rainfall, tide level, and flood point location as inputs to the LightGBM model, with the maximum flood depth at the flood point serving as the model's output. K-fold cross-validation and a grid search algorithm were used to optimize the hyperparameter set of the LightGBM model. The optimal values ​​for the learning rate, number of estimators, and number of leaf nodes were 0.11, 450, and 12, respectively. The LightGBM model achieved an NSE of 0.9896 on the test set, outperforming RF, XGBoost, and KNN. The results show that the LightGBM model can effectively predict water depth using a one-dimensional coupled hydrological and hydrodynamic model. For the LightGBM model, Gini index analysis revealed that variables related to tide level are the main variables for predicting flood depth. Due to its superior performance and computational efficiency, the urban flood prediction model constructed in this application embodiment can achieve efficient and accurate flood prediction.

[0118] The above are merely preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A rapid flood prediction method for coastal cities based on LightGBM and hydrodynamic models, characterized by: include: I. Obtaining sample data, including: designing various rainfall-tide combination scenarios; constructing a flood simulation model for the target area based on the hydrodynamic model of PCSWMM; simulating the maximum water depth at each selected flood point under different rainfall-tide combination scenarios using the flood simulation model; extracting feature variables related to rainfall and tide level, and using them together with the location of the flood point and the maximum water depth as sample data; II. Construction and training of urban flood prediction models, including: construction of an urban flood prediction model based on LightGBM, and construction of an urban flood prediction model trained with sample data; III. Predict the maximum water depth at inundation points during urban flooding using a trained urban flood prediction model; The hydro-hydrodynamic model based on PCSWMM is used to construct a flood simulation model for the target area, including: First, data on digital elevation models, drainage networks, one-dimensional manholes, sub-catchments, and obstacles are collected. Then, a one-dimensional drainage model of manholes and pipes is constructed to obtain two-dimensional nodes and a two-dimensional network. Next, a two-dimensional surface inundation model is constructed based on the two-dimensional nodes and the two-dimensional grid. Finally, the one-dimensional and two-dimensional models are coupled by using bottom orifice connections to obtain the flood simulation model.

2. The rapid flood prediction method for coastal cities based on LightGBM and hydrodynamic models as described in claim 1, characterized in that: The rainfall-related characteristic variables include one or more of the following: cumulative rainfall, rainfall return period, peak rainfall, maximum 2-hour rainfall, maximum 3-hour rainfall, and cumulative rainfall before the peak; and the tide-related characteristic variables include one or more of the following: maximum tide level, tide return period, average tide level, and maximum 5-hour average tide level.

3. The rapid flood prediction method for coastal cities based on LightGBM and hydrodynamic models as described in claim 1, characterized in that: The design includes various rainfall-tide combination scenarios, including: The frequency distributions of rainfall and tide levels are fitted using a generalized extreme value function (GEM). The parameters of the GEM, including location, scale, and shape parameters, are estimated using the maximum likelihood method. Various design rainfall and peak tide levels with return periods ranging from 5 to 100 years are obtained by calculating the GEM. Typical rainfall and tide events are scaled using the same scaling method to obtain design rainfall and tide events with different return periods. Various design rainfall-tide combination scenarios obtained by combining rainfall and tide data are used as boundary conditions for the flood simulation model.

4. The rapid flood prediction method for coastal cities based on LightGBM and hydrodynamic models as described in claim 1, characterized in that it also... include: The flood simulation model was calibrated using actual measured rainfall and tide sequences and inundation depths in the target area.

5. The rapid flood prediction method for coastal cities based on LightGBM and hydrodynamic models as described in claim 1, characterized in that: The urban flood prediction model constructed using sample data training includes: Several flood points were selected from flood-prone areas, and several feature variables related to rainfall and tide level were extracted. The locations of the flood points and the extracted feature variables were input into a flood simulation model to simulate the maximum water depth of each flood point under different rainfall-tide combination scenarios. The LightGBM model was trained using different rainfall-tide combination scenarios and flood point locations as inputs and the maximum water depth of each flood point as outputs.

6. The rapid flood prediction method for coastal cities based on LightGBM and hydrodynamic models as described in claim 1, characterized in that: It also includes optimizing the hyperparameters of the LightGBM model using K-fold cross-validation and grid search, specifically: Define the range and interval of hyperparameters for the LightGBM model. The range and interval of hyperparameters form a three-dimensional hyperparameter mesh. Use a mesh search method to traverse the hyperparameter mesh. Use K-fold cross-validation on each hyperparameter mesh to select the hyperparameter with the best generalization ability. K-fold cross-validation involves: randomly and uniformly dividing the sample data into k subsets; using k-1 subsets for training the model each time it is trained and testing, and using the remaining subsets for testing; repeating this process k times, and averaging the results of the k tests as the model evaluation value.

7. The rapid flood prediction method for coastal cities based on LightGBM and hydrodynamic models as described in claim 1, characterized in that: The hyperparameters of the LightGBM model include the learning rate, the number of estimators, and the number of leaf nodes.

8. The rapid flood prediction method for coastal cities based on LightGBM and hydrodynamic models as described in claim 1, characterized in that: It also includes using the Gini index to analyze the contribution of the input feature variables of the LightGBM model to urban flood prediction in the study area.