A method, apparatus and equipment for calculating the drainage and flood control capacity of low-lying urban roads.

By quantifying road feature data and rainfall data to form a scenario database, selecting key features and training a prediction model, the problem of complex and time-consuming drainage capacity calculation in existing technologies is solved, and rapid and accurate waterlogging risk assessment is achieved.

CN121684661BActive Publication Date: 2026-04-24BEIJING WATER SCI & TECH INST
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING WATER SCI & TECH INST
Filing Date
2026-02-11
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing methods for calculating drainage capacity are complex and time-consuming, making it difficult to quickly identify the risk of waterlogging and failing to meet the needs of large-scale road network assessments.

Method used

A parameter library is constructed by quantifying road feature data, a scenario library is formed by associating rainfall and water accumulation data, key features are selected and a prediction model is trained, features are selected using random forest regressors and cross-validation recursive feature elimination methods, and Bayesian optimization and LightGBM model are combined for model training to achieve fast and accurate water accumulation risk prediction.

Benefits of technology

It enables the calculation of the flooded area and water depth of one or more urban roads in a short time, provides reliable prediction of waterlogging risks, reduces reliance on complex data and professional software, and improves calculation efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121684661B_ABST
    Figure CN121684661B_ABST
Patent Text Reader

Abstract

This invention relates to the field of drainage and flood control technology, and discloses a method, apparatus, and equipment for calculating the drainage and flood control capacity of low-lying urban roads. The method includes: acquiring and quantifying road characteristics to construct a road characteristic parameter library; then associating rainfall and road waterlogging data with the parameter library to form a road waterlogging scenario parameter library; subsequently, selecting feature parameters from the scenario library and training a prediction model with the inundation area or waterlogging depth as the target; and finally, using the actual characteristic parameters of the target road, predicting its waterlogging index through the model. This invention solves the problems of traditional methods having a single data dimension and relying on experience. It can select prediction targets as needed to adapt to different scenarios, and combines feature selection and model training, taking into account factors such as terrain, rainfall, and drainage efficiency. This reduces the dependence on complex data and professional software, makes the calculation more efficient, and the results more realistic. It can quickly complete the calculation of road waterlogging indexes, providing a reliable basis for waterlogging risk prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of drainage and flood control technology, specifically to a method, apparatus, and equipment for calculating the drainage and flood control capacity of low-lying urban roads. Background Technology

[0002] In recent years, extreme rainfall events have become more frequent against the backdrop of global climate change, and urban flooding has become a prominent hidden danger threatening urban traffic operations and the safety of residents' lives and property. Flooding on urban roads during rainfall not only significantly reduces traffic efficiency and affects citizens' travel, but also easily leads to safety accidents; at the same time, long-term water accumulation accelerates the aging of road base layers and the corrosion of pipes and channels, increasing the maintenance costs of municipal facilities.

[0003] To accelerate urban flood control, various cities have implemented numerous flood control measures, including the construction and upgrading of stormwater pipe networks, the renovation of storm drain grates, and the construction of centralized water storage facilities. However, these measures have not yet been fully quantified and evaluated. Existing methods for analyzing urban road flooding risks and calculating drainage capacity are generally based on mathematical models of hydrological cycles. These models simulate pipe network drainage and surface runoff processes by constructing modules such as pipe network hydrodynamics, rainfall generation and runoff, and two-dimensional surface runoff. While these models have relatively fine parameter settings and can accurately reflect the actual operating status of the drainage system, their construction process is complex. They require a large amount of detailed data, including topography, permeability of different underlying surfaces, pipe network topology, and spatiotemporal distribution of rainfall. Data acquisition is difficult, costly, and highly dependent on specialized technical expertise. Furthermore, the calculation process is time-consuming, often taking several hours or even days for a single calculation. This makes it difficult to quickly identify flooding risks after a rapid response to rainfall forecasts and warnings, and it also fails to meet the practical application scenarios of large-scale road network drainage capacity assessment. Summary of the Invention

[0004] This invention provides a method, apparatus, and equipment for calculating the drainage and flood control capacity of low-lying urban roads, in order to solve the problems that existing drainage capacity calculation processes are complex, difficult to quickly identify waterlogging, and cannot meet the practical application needs of large-scale road network drainage capacity calculation and evaluation.

[0005] In a first aspect, the present invention provides a method for calculating the drainage and flood control capacity of low-lying urban roads, the method comprising:

[0006] Acquire road feature data, quantify the road feature data, and construct a road feature parameter library;

[0007] Based on road characteristic data, we obtain rainfall characteristic data and corresponding road waterlogging data, and combine them with the road characteristic parameter library to construct a road waterlogging scenario parameter library. The road waterlogging data includes the road flooded area and the road waterlogging depth.

[0008] Using the road flooding area or road water depth as the prediction target, feature parameters are selected from the road waterlogging scenario parameter database according to the prediction target, and the feature parameters are used as independent variables to train the model to obtain a trained prediction model.

[0009] Based on the prediction target, the actual characteristic parameters of the target road are obtained, and the prediction model is used to predict the flooded area or water depth of the target road.

[0010] This invention provides a method for calculating the drainage and flood control capacity of urban low-lying roads. It constructs a parameter library by quantifying road characteristics, establishes a scenario library by correlating rainfall and water accumulation data, and then selectively selects features and trains a model. This achieves accurate calculation of the drainage and flood control capacity of urban low-lying roads, solving the problems of traditional methods' single data dimension and reliance on experience. It allows for the selection of inundation area or water depth as the prediction target, adapting to the assessment needs of different scenarios. Simultaneously, the combination of feature selection and model training makes the calculation process more efficient and the results more realistic. While taking into account key factors such as topography, rainfall characteristics, and drainage system efficiency, it significantly reduces the reliance on complex data and specialized software. It can complete the calculation of inundation area, water depth, and other indicators for single or multiple urban roads in a short time, providing a reliable basis for predicting the water accumulation risk of target roads.

[0011] In one optional implementation, the road characteristic data includes: basic road information characterization parameters, source emission reduction capacity characterization parameters, process emission capacity characterization parameters, end-of-pipe storage capacity characterization parameters, and road flow capacity characterization parameters. The road characteristic data is quantified to construct a road characteristic parameter library, including:

[0012] Obtain the road area and the total area of ​​the road catchment area, calculate the ratio of the road area to the road catchment area, and obtain the quantitative result of the road area ratio as a parameter representing the basic information of the road.

[0013] The types of underlying surfaces, the area and runoff coefficients of each type of underlying surface are obtained within the road catchment area. The comprehensive runoff coefficient is calculated using the area weighting method. At the same time, the ratio of the permeable pavement area to the hardened underlying surface area within the road catchment area is calculated to obtain the permeable pavement rate. The comprehensive runoff coefficient and the permeable pavement rate are used as quantitative results of the source emission reduction capacity characterization parameters.

[0014] The average nearest neighbor distance of storm drain grates, storm drain network density, and storm drain network volume per unit catchment area are calculated as quantitative results of parameters characterizing process conveying and drainage capacity.

[0015] The ratio of cumulative storage capacity to total catchment area is used as the control ratio of centralized storage facilities. Based on pump performance parameters, the pump facility enhancement capacity data is calculated. The centralized storage facility control ratio and pump facility enhancement capacity data are used as the quantitative results of the end-point storage capacity characterization parameters.

[0016] The average road slope and the standard deviation of road slope are calculated based on road topographic raster data as quantitative results representing parameters of road traffic capacity.

[0017] In one optional implementation, the rainfall characteristics data for each rainfall event include: the cumulative rainfall amount of the rainfall event and the maximum rainfall per unit time; the road waterlogging data includes: the area of ​​road waterlogging or the maximum depth of waterlogging exceeding the depth threshold.

[0018] Based on road characteristic data, we obtain rainfall characteristic data for each event and corresponding road waterlogging data. Combined with a road characteristic parameter database, we construct a road waterlogging scenario parameter database, including:

[0019] By using historical waterlogging monitoring data and / or numerical model simulations, we can obtain rainfall characteristic data and corresponding road waterlogging data for multiple different roads and multiple rainfall events.

[0020] Based on different roads and different rainfall events, the road characteristic parameter database, rainfall event characteristic data, and corresponding road waterlogging data are classified and organized to construct a road waterlogging scenario parameter database.

[0021] The present invention provides a method for calculating the drainage and flood control capacity of low-lying urban roads. By quantifying the multi-dimensional characteristics of roads, the method makes the road drainage-related data more standardized and comprehensive, solving the problems of vague and single-dimensional descriptions of traditional road features. At the same time, the method classifies and associates road features with rainfall and water accumulation data to build a scenario library, achieving accurate data matching. This provides high-quality data support for subsequent screening of key features and training of prediction models, making the quantitative assessment of road water accumulation risk more accurate.

[0022] In one optional implementation, the predicted target is either the road flooding area or the road water depth. Feature parameters are selected from a road waterlogging scenario parameter database based on the predicted target, including:

[0023] The data in the road waterlogging scenario parameter database are standardized to obtain multiple sets of standard feature parameters, and these sets of standard feature parameters are divided into training set and validation set according to a preset ratio.

[0024] Using root mean square error as the feature selection metric, a random forest regressor and cross-validation recursive feature elimination are used to screen the standard feature parameters in the training set, and the key features that contribute the most to the prediction target are determined as the corresponding feature parameters.

[0025] In one optional implementation, the root mean square error is used as the feature selection metric, and a random forest regressor and cross-validation recursive feature elimination are used to select the standard feature parameters in the training set, including:

[0026] Configure a K-fold randomized cross-validation strategy to group the standard feature parameters and obtain multiple sets of standard feature data;

[0027] By embedding the random forest regressor into a cross-validation recursive feature elimination framework, a feature selection function is obtained.

[0028] Using root mean square error as the feature selection index, the feature selection function is used to iteratively select the standard feature data of each group in the training set. After removing the feature with the lowest contribution to the prediction target each time, the root mean square error of the feature selection function is calculated to determine the optimal feature and the optimal number of features corresponding to the minimum root mean square error.

[0029] Using the attribute extraction function of cross-validation recursive feature elimination, the importance of the optimal feature is calculated and extracted, and the optimal features are sorted from high to low according to their importance.

[0030] Based on the optimal features after sorting, different numbers of features are selected in sequence to construct performance evaluation models for performance evaluation, and the root mean square error values ​​corresponding to different numbers of features are recorded.

[0031] The present invention provides a method for calculating the drainage and flood control capacity of low-lying urban roads. By standardizing data dimensions, it avoids interference from differences in the dimensions of different features in the screening results. Using root mean square error as an indicator, it combines random forest and recursive feature elimination through cross-validation to accurately identify the key features that contribute the most to the prediction target. K-fold cross-validation ensures the stability of the screening and reduces the risk of overfitting. At the same time, by ranking the importance of features and testing the model performance of different numbers of features, it can obtain the optimal feature combination and provide a flexible selection basis for subsequent model optimization, thereby improving the efficiency and accuracy of road waterlogging prediction.

[0032] In one optional implementation, the feature parameters are used as independent variables for model training to obtain a trained prediction model, including:

[0033] The hyperparameters of the Light Gradient Boosting Machine (LightGBM) model based on decision trees were optimized using the Bayesian optimization method to obtain the hyperparameter-optimized LightGBM model. The Bayesian ridge regression model and the hyperparameter-optimized LightGBM model were then trained using the training set to determine the optimal number of iterations and cross-validation prediction results.

[0034] A stacked ensemble weight calculation system is constructed based on K-fold cross-validation to solve the optimal fusion weights of the Bayesian ridge regression model and the hyperparameter-optimized LightGBM model.

[0035] Using the optimal fusion weights, the cross-validation prediction results of the Bayesian Ridge Regression model and the hyperparameter-optimized LightGBM model are weighted and fused to obtain stacked prediction values. The training effect of the model is evaluated using the original data and standard data of each feature parameter.

[0036] A stacked ensemble model is trained based on all the data corresponding to the feature parameters to obtain a trained prediction model, and the training results are evaluated and verified using a validation set.

[0037] The present invention provides a method for calculating the drainage and flood control capacity of low-lying urban roads. By dividing the training set and validation set and retaining the original data and standard data, it takes into account both the stability of model training and the interpretability of results. Bayesian optimization is used to optimize the hyperparameters of LightGBM, and a stacked weight system is constructed by combining K-fold cross-validation. This integrates the advantages of linear models and LightGBM to improve prediction accuracy. Weighted fusion and dual data evaluation ensure the model performance. Finally, the stacked ensemble model trained with full data avoids overfitting and enhances generalization ability, which can more accurately and stably support road water accumulation prediction and improve the reliability of drainage and flood control capacity calculation.

[0038] Secondly, the present invention provides a device for calculating the drainage and flood control capacity of low-lying urban roads, the device comprising:

[0039] The road feature parameter library construction module is used to acquire road feature data, quantify the road feature data, and construct the road feature parameter library.

[0040] The road waterlogging scenario parameter library construction module is used to obtain rainfall characteristic data and corresponding road waterlogging data based on road characteristic data, and to construct a road waterlogging scenario parameter library in combination with the road characteristic parameter library. The road waterlogging data includes the road flooded area and the road waterlogging depth.

[0041] The parameter selection and model training module is used to select feature parameters from the road flooding area or road water depth as the prediction target, and use the feature parameters as independent variables to train the model to obtain a trained prediction model.

[0042] The actual prediction module is used to obtain the actual characteristic parameters of the target road based on the prediction target, and to use the prediction model to predict the flooded area or water depth of the target road.

[0043] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the method described in the first aspect or any corresponding embodiment thereof.

[0044] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to perform the method described in the first aspect or any corresponding embodiment thereof.

[0045] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to perform the method described in the first aspect or any corresponding embodiment thereof. Attached Figure Description

[0046] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0047] Figure 1 This is a schematic diagram of an application scenario according to an embodiment of the present invention;

[0048] Figure 2 This is a schematic diagram of the first process of the method for calculating the drainage and flood control capacity of urban low-lying roads according to an embodiment of the present invention.

[0049] Figure 3 This is a schematic diagram of the second process of the method for calculating the drainage and flood control capacity of urban low-lying roads according to an embodiment of the present invention;

[0050] Figure 4 This is a flowchart illustrating the prediction model construction process in the method for calculating the drainage and flood control capacity of urban low-lying roads according to an embodiment of the present invention.

[0051] Figure 5 This is a schematic diagram of the process for screening characteristic parameters in the method for calculating the drainage and flood control capacity of urban low-lying roads according to an embodiment of the present invention;

[0052] Figure 6 This is a schematic diagram of the stacked integrated model construction and training process in the urban low-lying road drainage and flood control capacity calculation method according to an embodiment of the present invention;

[0053] Figure 7This is a schematic diagram illustrating the verification results of predicting the flood-prone area of ​​roads in a specific embodiment of the method for calculating the drainage and flood control capacity of urban low-lying roads according to an embodiment of the present invention.

[0054] Figure 8 This is a schematic diagram of the verification results for predicting the maximum water accumulation depth of a road, in another specific embodiment of the method for calculating the drainage and flood control capacity of low-lying urban roads according to an embodiment of the present invention.

[0055] Figure 9 This is a structural block diagram of an urban low-lying road drainage and flood control capacity calculation device according to an embodiment of the present invention;

[0056] Figure 10 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0058] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.

[0059] In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0060] As an optional application scenario of this invention, such as Figure 1 As shown, the urban low-lying road drainage and flood control capacity calculation system may include at least one terminal device and at least one server. Figure 1 The system is illustrated in the example, which includes a computer 101, a mobile terminal 102, and a server 103, and the terminal devices such as the computer 101 and the mobile terminal 102 are connected to the server 103 through a network 110.

[0061] Specifically, the terminal device can be a smartphone, tablet, laptop, PDA, desktop computer, game console, smart TV, smart wearable device, in-vehicle terminal, VR (Virtual Reality) device, AR (Augmented Reality) device, etc. Server 103 can be a standalone physical server, a server cluster, a distributed system, or a cloud server providing cloud services. Network 110 can be a wired or wireless network, examples of which include, but are not limited to, the Internet, corporate intranet, local area network, wide area network, mobile communication network, and combinations thereof.

[0062] This invention provides a method for calculating the drainage and flood control capacity of low-lying urban roads. By quantifying data to construct a waterlogging scenario parameter database and selectively selecting features and training models, the method aims to improve the efficiency and adaptability of drainage capacity calculation and quickly identify waterlogging.

[0063] According to an embodiment of the present invention, an embodiment of a method for calculating the drainage and flood control capacity of low-lying urban roads is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0064] This embodiment provides a method for calculating the drainage and flood control capacity of low-lying urban roads, which can be used in the aforementioned computer system. Figure 2 This is a flowchart of a method for calculating the drainage and flood control capacity of urban low-lying roads according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps:

[0065] Step S201: Obtain road feature data, quantify the road feature data, and construct a road feature parameter library.

[0066] Specifically, basic characteristic data affecting road drainage capacity, such as land use type, road topographic slope, and distribution of drainage and water collection facilities, are collected within the urban road drainage zoning area. These basic characteristic data are then quantified to construct a road characteristic parameter database. The focus is on collecting, quantifying, and integrating basic characteristic data affecting road drainage capacity. The data collection scope covers basic road information and the entire drainage process from source to end, including road area, underlying surface type, detailed topography, and water collection and drainage facilities.

[0067] Step S202: Based on the road feature data, obtain the rainfall feature data of each event and the corresponding road waterlogging data, and combine it with the road feature parameter library to construct a road waterlogging scenario parameter library. The road waterlogging data includes the road flooded area and the road waterlogging depth.

[0068] Specifically, we collect road waterlogging scenario data and corresponding waterlogging scenario data from historical rainfall events or simulated rainfall events using digital models, including waterlogged area, waterlogged depth, and corresponding rainfall. Combined with a road feature parameter database, we construct a road waterlogging scenario database.

[0069] Step S203: Using the road flooding area or road water depth as the prediction target, select feature parameters from the road waterlogging scenario parameter library according to the prediction target, and use the feature parameters as independent variables to train the model to obtain the trained prediction model.

[0070] Specifically, the prediction model is trained by using the flooded area or water depth of roads as the prediction target (dependent variable) and numerical parameters such as topography, rainfall, and road attributes as input features (independent variables). The process involves "database preprocessing - feature parameter screening - model building and training - prediction effect verification" to ultimately achieve the simulation prediction of the flooded area or water depth of roads.

[0071] Step S204: Obtain the actual characteristic parameters of the target road based on the prediction target, and use the prediction model to predict the flooded area or water depth of the target road.

[0072] Specifically, by sorting out and integrating the main influencing factors and parameter data of urban road drainage, a standardized process was designed, core parameters were screened, and a rapid and simple evaluation model was built. A calculation model for urban road drainage and flood control capacity was constructed with the core parameters affecting road drainage capacity as input data. Finally, a rapid, simple, and universal calculation method was constructed with the road flood area or water depth as the prediction target.

[0073] The method for calculating the drainage and flood control capacity of low-lying urban roads provided in this embodiment constructs a parameter library by quantifying road features, forms a scenario library by associating rainfall and water accumulation data, and then selectively selects features and trains a model. This achieves accurate calculation of the drainage and flood control capacity of low-lying urban roads, solving the problems of single data dimensions and reliance on experience in traditional methods. It allows for the selection of inundation area or water depth as the prediction target as needed, adapting to the assessment needs of different scenarios. At the same time, the combination of feature selection and model training makes the calculation process more efficient and the results more realistic. While taking into account key factors such as topography, rainfall characteristics, and drainage system efficiency, it significantly reduces the reliance on complex data and professional software. It can complete the calculation of indicators such as water accumulation area and water depth of one or more urban roads in a short time, providing a reliable basis for predicting the water accumulation risk of target roads.

[0074] This embodiment provides a method for calculating the drainage and flood control capacity of low-lying urban roads, which can be used in the aforementioned computer system. Figure 3 This is a flowchart of a method for calculating the drainage and flood control capacity of urban low-lying roads according to an embodiment of the present invention, such as... Figure 3 As shown, the process includes the following steps:

[0075] Step S301: Obtain road feature data, quantify the road feature data, and construct a road feature parameter library.

[0076] Specifically, the road characteristic data includes: basic road information characterization parameters, source emission reduction capacity characterization parameters, process emission transport capacity characterization parameters, end-of-pipe storage capacity characterization parameters, and road flow discharge capacity characterization parameters. Step S301 includes:

[0077] Step S3011: Obtain the road area and the total area of ​​the road catchment zone, calculate the ratio of the road area to the road catchment zone area, and obtain the quantitative result of the road area ratio as a parameter representing the basic information of the road.

[0078] Specifically, the target road length and average width are recorded, along with the total area of ​​the road's catchment area. The road area ratio is calculated by dividing the road area by the road's catchment area, using the following formula:

[0079]

[0080] in, This indicates the percentage of the road area of ​​target road i. The length (km) of the target road i can be obtained through various methods, such as map software marking, on-site survey and measurement, and extraction of relevant appropriate layer files. The average width (km) of target road i can be obtained through various methods, such as map software marking, on-site survey and measurement, and extraction of relevant appropriate layer files. This represents the total catchment area (km²) of the target road i. 2 Data can be obtained through various methods, such as collecting planning documents and other materials, marking maps with software, conducting on-site surveys and measurements, and extracting relevant appropriate layer files.

[0081] Step S3012: Obtain the underlying surface type, the area and runoff coefficient of each underlying surface type within the road catchment area. Calculate the comprehensive runoff coefficient using the area weighting method. Simultaneously, calculate the ratio of permeable pavement area to hardened underlying surface area within the road catchment area to obtain the permeable pavement rate. The comprehensive runoff coefficient and permeable pavement rate serve as quantitative results for characterizing source emission reduction capacity.

[0082] Specifically, the emission reduction capacity at the source of roads includes two quantifiable parameters: 1) the runoff characteristics of the underlying surface are characterized by the comprehensive runoff coefficient of the catchment area where the road is located; 2) the infiltration, retention and storage capacity of the main source sponge facilities for rainwater in the process of sponge city construction within the catchment area is characterized by the permeable pavement rate.

[0083] 1) The underlying surface characteristics are quantitatively characterized using the comprehensive runoff coefficient to represent the underlying surface features of road catchment areas. Based on the distribution characteristics of the underlying surface in road catchment areas and combined with various underlying surface runoff coefficients, the comprehensive runoff coefficient is calculated using an area-weighted scoring method. The calculation formula is as follows:

[0084]

[0085] in, This represents the comprehensive runoff coefficient (dimensionless). This represents the runoff coefficient of the i-th type of underlying surface; the runoff coefficient values ​​for each type of underlying surface are taken with reference to the runoff coefficients in the relevant specifications; S represents the total catchment area of ​​the target road i (unit: km²). 2 This information can be obtained through various methods, including collecting planning documents and other data, marking maps with software, conducting on-site surveys and measurements, and extracting relevant appropriate layer files. This represents the area of ​​the nth type of underlying surface within the catchment area of ​​target road i (unit: km²). 2 The data was obtained through remote sensing image interpretation and land use type data collection. The underlying surface types within the catchment area where the road is located were classified and statistically analyzed into green space, hardened underlying surface, water area, and other permeable underlying surface. Among them, green space includes gardens, woodlands, grasslands, and park green spaces; hardened underlying surface includes commercial and service land, industrial and mining storage land, residential land, public management and public service land, transportation land, and special land; other permeable underlying surface includes cultivated land and other land; and water area includes water area and water conservancy facility land.

[0086] 2) The main parameters of the sponge city infrastructure are quantitatively characterized by the permeable pavement ratio. This is calculated and characterized by the ratio of the statistical results of the scale of the main source facilities (permeable pavement) within the road catchment area to the area of ​​the hardened underlying surface.

[0087]

[0088] in, Indicates the permeable pavement percentage (%) within the catchment area of ​​target road i; This represents the permeable pavement area within the catchment zone of target road i (unit: km²). 2 This information was obtained through on-site surveys and data collection. This represents the area of ​​the hardened underlying surface within the catchment zone of target road i (unit: km²). 2 ).

[0089] Step S3013: Calculate the average nearest distance of the storm drain grate, the density of the storm drain network, and the volume of the storm drain network per unit catchment area as quantitative results of the process flow and discharge capacity characterization parameters.

[0090] Specifically, the water conveyance capacity of a road includes two quantifiable parameters: 1) the distribution characteristics of water collection facilities are represented by the spacing of road storm drain grates; 2) the storm drain network capacity is represented by the density of the storm drain network and the volume of the storm drain network per unit catchment area.

[0091] 1) Distribution characteristics of water collection facilities: Conduct on-site surveys and record the number and location of rainwater grates distributed along the road. After vectorization of the survey data, use the "spatial statistics" function in geographic information system software (such as ArcGIS) to calculate the average nearest distance of each data point (unit: m).

[0092] 2) Drainage network distribution characteristics: Conduct on-site surveys and record the distribution characteristics of municipal stormwater drainage pipe network diameter, length, etc., within the road catchment area. Calculate the stormwater network density and stormwater network volume per unit catchment area. The formula for calculating stormwater network density is:

[0093]

[0094] in, Indicates the density (%) of municipal stormwater pipe network distribution within the catchment area of ​​target road i. The length of pipe segments of different diameters within the road catchment area (unit: m) is represented by the pipe length (obtained through on-site surveys and data collection); S represents the total catchment area of ​​the target road i (unit: km²). 2 ).

[0095] The formula for calculating the volume of stormwater pipes per unit catchment area is:

[0096]

[0097] in, This represents the volume of stormwater pipes per unit catchment area within the catchment zone of target road i (m³). 3 / km 2 ); The length of pipe sections of different diameters within the road catchment area (unit: m) is obtained through on-site surveys, data collection, and other methods. This indicates the cross-sectional area of ​​pipe sections of different diameters within the road catchment area (unit: m). 2The diameter of the stormwater drainage pipe network was calculated through on-site surveys and data collection; S represents the total catchment area of ​​the target road i (unit: km²). 2 ).

[0098] Step S3014: Calculate the ratio of cumulative storage volume to total catchment area as the control ratio of centralized storage facilities; calculate the pump facility enhancement capacity data based on pump performance parameters; and use the centralized storage facility control ratio and pump facility enhancement capacity data as the quantitative results of the end-point storage capacity characterization parameters.

[0099] Specifically, the stormwater storage capacity at the end of a road includes two quantifiable parameters: 1) The centralized stormwater storage facility control ratio characterizes the capacity of stormwater storage facilities or storage spaces within the catchment area. This is achieved through on-site surveys and statistics of stormwater storage facilities and water bodies with centralized stormwater storage functions within the road's catchment area, calculating the cumulative storage volume, and then using the ratio to the catchment area, expressed in meters. 3 / km 2 Recorded by unit.

[0100] 2) The performance parameters of the stormwater lift pumps characterize the centralized pumping capacity for stagnant water within the catchment area. This is achieved through on-site surveys and statistical analysis of the pump performance parameters at stormwater lift pump stations within the road catchment areas. The cumulative lifting capacity of the pump facilities is calculated, expressed in meters (m). 3 Records are kept in units of / h.

[0101] Step S3015: Calculate the average road slope and the standard deviation of road slope based on the road terrain raster data as quantitative results of road drainage capacity characterization parameters.

[0102] Specifically, road capacity includes two quantifiable parameters. First, the average road slope represents the overall trend of topographic change; second, the standard deviation of road slope represents the local topographic change characteristics.

[0103] The two feature parameters mentioned above are obtained by acquiring road terrain raster data and using the "Spatial Analysis" tool in ArcGIS software to calculate the slope, calculate the average road slope and slope standard deviation, and both feature parameters are represented as percentages (%).

[0104] After quantifying the 10 features of urban roads across the five aspects mentioned above, the parameters were categorized and organized to construct an urban road feature parameter library, as shown in Table 1. The library includes at least five urban roads.

[0105] Table 1 Road Feature Parameter Library

[0106]

[0107] Step S302: Based on the road feature data, obtain the rainfall feature data of each event and the corresponding road waterlogging data, and combine it with the road feature parameter library to construct a road waterlogging scenario parameter library. The road waterlogging data includes the road flooded area and the road waterlogging depth.

[0108] Specifically, step S302 includes:

[0109] Step S3021: Obtain rainfall characteristic data and corresponding road water accumulation data for multiple different roads and multiple rainfall events through historical water accumulation monitoring data and / or numerical model simulation.

[0110] Specifically, a road waterlogging scenario database is constructed, encompassing three categories: urban road characteristic parameters, rainfall event characteristic data, and road waterlogging conditions during corresponding rainfall events. Data sources include historical road waterlogging monitoring data, survey data, or results derived from numerical model simulations. Specifically, urban road characteristic parameters correspond to the parameter data of at least five roads obtained in step S301; rainfall event characteristic data must include both cumulative rainfall and maximum hourly rainfall data for each rainfall event, and should also include historical rainfall events corresponding to the waterlogging conditions on the target road and design rainfall data for different return periods in relevant planning and design standards; road waterlogging data is either "road waterlogged area exceeding 15cm" or "maximum road waterlogging depth."

[0111] The road waterlogging scenario database should contain data sets of no less than 5 different roads and no less than 45 rainfall events and corresponding waterlogging conditions. 80% to 90% of the data sets should be used for model building and training, and 10% to 20% of the data sets should be used to verify the accuracy of the model results.

[0112] Step S3022: According to different roads and different rainfall events, classify and organize the road feature parameter database, rainfall event feature data and corresponding road waterlogging data to construct a road waterlogging scenario parameter database.

[0113] Specifically, the three types of data were classified and organized according to different roads and different rainfall events, and a road waterlogging scenario parameter database as shown in Table 2 was constructed.

[0114] Table 2. Road Waterlogging Scenario Parameter Database

[0115]

[0116] The method for calculating the drainage and flood control capacity of low-lying urban roads provided in this embodiment quantifies the multi-dimensional characteristics of roads, making the road drainage-related data more standardized and comprehensive. This solves the problems of vague and single-dimensional descriptions of traditional road features. At the same time, it classifies and associates road features with rainfall and water accumulation data to build a scenario library, achieving accurate data matching. This provides high-quality data support for subsequent screening of key features and training of prediction models, making the quantitative assessment of road water accumulation risk more accurate.

[0117] Step S303: Using the road flooding area or road water depth as the prediction target, select feature parameters from the road waterlogging scenario parameter library according to the prediction target, and use the feature parameters as independent variables to train the model to obtain the trained prediction model.

[0118] Specifically, step S303 includes:

[0119] Step S3031: Standardize the data in the road waterlogging scenario parameter database to obtain multiple sets of standard feature parameters, and divide the multiple sets of standard feature parameters into training set and validation set according to a preset ratio.

[0120] Specifically, the predictive model is constructed using Python programming, such as... Figure 4 The diagram shows the flowchart of the prediction model construction process. The original data from the road flooding scenario parameter database are preprocessed. For each type of data in the database, Z-score outlier detection and Min-Max standardization (mapping the value range to [0,1]) are performed sequentially to eliminate scale differences caused by different units and numerical ranges, providing comparable and robust input features for subsequent modeling. The specific standardization process is a mature existing technology and will not be elaborated here.

[0121] The Z-score method transforms raw data into "standardized values ​​centered on the mean and measured in units of standard deviation". For example, by calculating the Z-score value of a single sample, "absolute Z-score < 3" is used as the criterion for outlier identification, thus identifying and reducing interference from extreme values.

[0122] Min-Max normalization maps the original data to the [0, 1] interval to eliminate dimensional and numerical range differences between different features, bringing all features to the same order of magnitude and thus providing a fair input basis for the training of subsequent machine learning models.

[0123] For multiple sets of data in the road waterlogging scenario parameter database, 80%–90% of the data sets are used for model building and training, and 10%–20% of the data sets are used for validating the accuracy of the model results. The feature matrices and target variables of the training set and the validation set are separated, while retaining both standardized and original formats of the target variables, thus balancing the stability of model training and the interpretability of the results.

[0124] Step S3032: Using root mean square error as the feature selection index, the standard feature parameters in the training set are selected by using a random forest regressor and cross-validation recursive feature elimination to determine the key features that contribute the most to the prediction target as the corresponding feature parameters.

[0125] Specifically, the combination of Random Forest Regressor (RFR) and Recursive Feature Elimination with Cross-Validation (RFECV) is the core technical approach. Its core objective is to select the subset of key features that contribute most to the prediction of the target "road flood area" or "maximum water depth" from the original 12 numerical features (including road feature parameters, rainfall feature data, etc.), while determining the optimal number of features to provide high-quality input for subsequent model construction and training, avoid redundant features, reduce data dimensionality, and avoid model overfitting.

[0126] like Figure 5 The diagram illustrates the feature parameter selection process. Feature indicators are selected based on the prediction target of road flooding area or maximum water depth. The root mean square error (RMSE) is calculated using cross-validation; a smaller RMSE value indicates higher model prediction accuracy. From the 12 input numerical feature parameters, core features that initially contribute significantly to the prediction objective function are selected. The contribution of each feature is quantified, and the RMSE is calculated using cross-validation to evaluate the suitability of different feature combinations.

[0127] In some optional implementations, the code for feature parameter filtering is as follows:

[0128] # Cross-validation settings

[0129] kf = KFold(n_splits=5, shuffle=True, random_state=42)

[0130] # Feature selection function

[0131] def top_features_analysis(X, y, tag, target_col_name):

[0132] # Initialize the random forest model

[0133] rf = RandomForestRegressor(n_estimators=50, n_jobs=-1, random_state=42)

[0134] # Feature selection using RFECV

[0135] rfecv = RFECV(estimator=rf, step=1, cv=kf, scoring='neg_mean_squared_error')

[0136] rfecv.fit(X, y)

[0137] print(f'\n===== {tag}Submerged Area Feature Selection Result=====')

[0138] print(f'RFECV recommended optimal number of features: {rfecv.n_features_}')

[0139] print(f'Feature ranking: {rfecv.ranking_}')

[0140] print(f'Length of feature importance array: {len(rfecv.estimator_.feature_importances_)}')

[0141] # Get the index of the best feature (feature with ranking=1)

[0142] optimal_feature_indices = [i for i in range(len(rfecv.ranking_)) ifrfecv.ranking_[i] == 1]

[0143] num_optimal_features = len(optimal_feature_indices)

[0144] print(f'\nRFECV selected the total number of optimal features: {num_optimal_features}')

[0145] # Create a mapping from the original feature indexes to the feature indexes preserved by RFECV

[0146] mask = rfecv.support_# A boolean array; True indicates that features are preserved.

[0147] original_to_selected = {}

[0148] selected_idx = 0

[0149] for original_idx in range(len(mask)):

[0150] if mask[original_idx]:

[0151] original_to_selected[original_idx] = selected_idx

[0152] selected_idx += 1

[0153] # Obtain the names and corresponding importance of the optimal features (using mapping to solve the indexing problem) optimal_feature_names = [feature_cols[i] for i in optimal_feature_indices]

[0154] optimal_feature_importances = []

[0155] for idx in optimal_feature_indices:

[0156] if idx in original_to_selected:

[0157] imp_idx = original_to_selected[idx]

[0158] optimal_feature_importances.append(rfecv.estimator_.feature_importances_[imp_idx])

[0159] else:

[0160] optimal_feature_importances.append(0.0)

[0161] print(f"Warning: Feature index {idx} is not in the retained feature set")

[0162] # Sort the best features by importance (in descending order)

[0163] sorted_pairs = sorted(zip(optimal_feature_names, optimal_feature_importances),

[0164] key=lambda x: x[1], reverse=True)

[0165] sorted_optimal_names = [pair[0] for pair in sorted_pairs]

[0166] sorted_optimal_importances = [pair[1] for pair in sorted_pairs]

[0167] # Output the best features sorted by importance

[0168] print('\nOptimal features sorted by importance:')

[0169] for i, (name, imp) in enumerate(zip(sorted_optimal_names, sorted_optimal_importances), 1):

[0170] print(f'Rank {i}: {name} (Importance: {imp:.4f})')

[0171] # Compare model performance with different numbers of optimal features

[0172] print(f'\n===== Performance Comparison of Models with Different Numbers of Optimal Features=====')

[0173] print(f'Different numbers of features will be selected from {num_optimal_features} optimal features for evaluation\n')

[0174] Based on the code for filtering using the aforementioned feature parameters, step S3032 includes:

[0175] Step a1: Configure the K-fold random shuffle cross-validation strategy to group the standard feature parameters and obtain multiple sets of standard feature data.

[0176] Specifically, a K-fold (1 to 10 folds) randomized cross-validation (KFold) strategy is configured to divide all standard feature parameters into multiple groups of standard feature data, ensuring the stability and randomness of the evaluation.

[0177] Step a2: Embed the random forest regressor into the cross-validation recursive feature elimination framework to obtain the feature selection function.

[0178] Specifically, a feature selection function is defined, using a random forest regressor (10~80 decision trees, running in full thread) as the base model, integrating RFECV to achieve iterative feature selection, and using mean squared error (RMSE) as the scoring metric to automatically determine the optimal number of features.

[0179] Step a3: Using root mean square error as the feature selection index, the feature selection function is used to iteratively select the standard feature data of each group in the training set. After removing the feature with the lowest contribution to the prediction target each time, the root mean square error of the feature selection function is calculated to determine the optimal feature and the optimal number of features corresponding to the minimum root mean square error.

[0180] Specifically, using the feature selection function, features are subtracted one by one from the grouped standard feature data. Each time, the feature with the smallest contribution to the prediction target is removed. Then, the root mean square error of the model is calculated. After multiple operations, the feature set corresponding to the smallest initial RMSE is selected, and the number of features in the feature set is determined.

[0181] Step a4: Utilize the attribute extraction function of cross-validation recursive feature elimination to calculate and extract the importance of the optimal feature, and sort the optimal features from high to low importance.

[0182] Specifically, using the attribute extraction function of RFECV, the names and corresponding importance of the core features are calculated and extracted, sorted in descending order of importance, and output.

[0183] Step a5: Based on the optimal features after sorting, select different numbers of features in sequence to construct performance evaluation models and perform performance evaluation, and record the root mean square error values ​​corresponding to different numbers of features.

[0184] Specifically, based on the selected optimal features, different numbers of core features (from 1 to the optimal number of features) were selected sequentially to construct models and evaluate their performance. RMSE and other metrics corresponding to different numbers of features were recorded to provide a basis for comparison in subsequent model optimization. Table 3 shows the contribution statistics of core feature metrics, and Table 4 shows the verification results of the impact of different combinations of metric numbers.

[0185] Table 3. Statistical Table of Contribution of Core Feature Indicators

[0186]

[0187] Table 4. Verification Results of the Influence of Indicator Quantity Combinations

[0188]

[0189] This approach integrates the feature importance assessment of random forests with the iterative selection capabilities of RFECV, balancing the accuracy and generalization of feature selection. Cross-validation is used throughout the feature selection process to effectively avoid overfitting and provide scientific feature selection support for subsequent model performance optimization.

[0190] The method for calculating the drainage and flood control capacity of low-lying urban roads provided in this embodiment standardizes data dimensions to avoid interference from differences in feature dimensions in the screening results. Using root mean square error as an indicator, it combines random forest and cross-validation recursive feature elimination to accurately identify the key features that contribute the most to the prediction target. K-fold cross-validation ensures the stability of the screening and reduces the risk of overfitting. At the same time, it ranks the importance of features and tests the model performance of different numbers of features, which not only obtains the optimal feature combination, but also provides a flexible selection basis for subsequent model optimization, thereby improving the efficiency and accuracy of road waterlogging prediction.

[0191] Step S3033: Use the Bayesian optimization method to optimize the hyperparameters of the LightGBM model to obtain the hyperparameter-optimized LightGBM model. Then, use the training set to train the Bayesian ridge regression model and the hyperparameter-optimized LightGBM (Light Gradient Boosting Machine) model to determine the optimal number of iterations and cross-validation prediction results.

[0192] Specifically, the Bayesian optimization method is used to optimize the hyperparameters of the LightGBM model, and the zero-value mask features in the training data are combined to improve the model's adaptability to special samples.

[0193] A stacked ensemble model of Bayesian ridge regression and LightGBM is constructed. The Bayesian ridge regression model and the optimized LightGBM model are trained separately. Overfitting is avoided by using an early stopping mechanism. The optimal number of iterations of LightGBM is recorded and the cross-validation prediction results are output. The model's fitting ability and generalization ability are balanced to achieve the prediction of the target factor.

[0194] LightGBM is a highly efficient Gradient Boosting Decision Tree (GBDT) optimization algorithm. Its core advantage lies in improving training efficiency and prediction accuracy through histogram optimization and leaf-wise growth, allowing the model to focus more on key feature combinations and enhance its ability to capture complex nonlinear relationships. LightGBM incorporates L1 / L2 regularization parameters, which can effectively suppress overfitting risks in small sample scenarios and enhance the model's generalization ability.

[0195] Bayesian optimization can focus on the optimal solution within a limited number of iterations, making it particularly suitable for hyperparameter tuning in small sample scenarios. It efficiently searches for the optimal combination of core parameters, significantly reducing the model's prediction error while shortening the time required for parameter tuning.

[0196] The code for training the prediction model using the training set is as follows:

[0197] # Prepare features and target variables

[0198] X_train = train_df[FEATURES].values

[0199] y_std_train = train_df[TARGET_FULL_STD].values ​​# Standardized target variable (for training)

[0200] y_raw_train = train_df[TARGET_FULL_RAW].values ​​# Original target variable

[0201] X_test = test_df[FEATURES].values

[0202] y_std_test = test_df[TARGET_FULL_STD].values ​​# Standardized target variable for validation set

[0203] y_raw_test = test_df[TARGET_FULL_RAW].values ​​# Original target variable for the validation set

[0204] # Bayesian optimization of LightGBM

[0205] best_lgb = optimize_lightgbm_for_zero(X_train, y_std_train, train_zero_mask)

[0206] # LOOCV Cross-Validation (with Early Stopping Mechanism)

[0207] loocv = LeaveOneOut()

[0208] oof_br_std = np.zeros(len(X_train)) # Bayesian ridge cross-validation prediction (standardized)

[0209] oof_lgb_std = np.zeros(len(X_train)) # LightGBM cross-validation predictions (normalized)

[0210] lgb_iterations = [] # Records the optimal number of iterations in each round

[0211] logger.info("\nStarting LOOCV cross-validation (with early stop mechanism)...")

[0212] for i, (tr_idx, val_idx) in enumerate(loocv.split(X_train)):

[0213] if i % 5 == 0:

[0214] logger.info(f"LOOCV iteration {i + 1} / {len(X_train)}")

[0215] X_tr, X_val = X_train[tr_idx], X_train[val_idx]

[0216] y_std_tr, y_std_val = y_std_train[tr_idx], y_std_train[val_idx]

[0217] # Training Bayesian Ridge Regression

[0218] br = BayesianRidge(alpha_1=1e-6, alpha_2=1e-6)

[0219] br.fit(X_tr, y_std_tr)

[0220] oof_br_std[val_idx] = br.predict(X_val)

[0221] # Training LightGBM (with early stop)

[0222] params = best_lgb.get_params()

[0223] if 'random_state' in params:

[0224] del params['random_state']

[0225] early_stopping = lgb.early_stopping(

[0226] stopping_rounds=EARLY_STOPPING_ROUNDS,

[0227] verbose=EARLY_STOPPING_VERBOSE )

[0229] if'verbose' in params:

[0230] del params['verbose']

[0231] lgb_model = lgb.LGBMRegressor(** params, random_state=42, verbose=-1)

[0232] lgb_model.fit(

[0233] X_tr, y_std_tr,

[0234] eval_set=[(X_val, y_std_val)],

[0235] callbacks=[early_stopping] )

[0237] lgb_iterations.append(lgb_model.best_iteration_)

[0238] oof_lgb_std[val_idx] = lgb_model.predict(X_val, num_iteration=lgb_model.best_iteration_)

[0239] # Calculate Stacking weights (K-fold cross-validation)

[0240] logger.info("\nStarting to calculate Stacking weights (based on K-fold cross-validation)...")

[0241] br_base = BayesianRidge(alpha_1=1e-6, alpha_2=1e-6)

[0242] # Preparing the LightGBM base model

[0243] lgb_base_params = best_lgb.get_params()

[0244] if 'random_state' in lgb_base_params:

[0245] del lgb_base_params['random_state']

[0246] if 'verbose' in lgb_base_params:

[0247] del lgb_base_params['verbose']

[0248] lgb_base = lgb.LGBMRegressor(** lgb_base_params, random_state=42,verbose=-1)

[0249] # Calculate weights

[0250] w_br, w_lgb = calculate_stacking_weights(

[0251] X_train, y_std_train, y_raw_train,

[0252] min_max_params, br_base, lgb_base )

[0254] #Training set stacking results

[0255] y_train_stack_std = w_br * oof_br_std + w_lgb * oof_lgb_std

[0256] y_train_stack_raw = y_train_stack_std * min_max_params['range'] +min_max_params['min']

[0257] # Calculate the training set stacking metric

[0258] tr_metrics = calculate_dual_metrics(

[0259] y_std_train, y_train_stack_std,

[0260] y_raw_train, y_train_stack_raw,

[0261] train_zero_mask, "Training set stack (LOOCV)" )

[0263] # Training the final model (with early stopping)

[0264] logger.info("Training the final model (with early stopping mechanism)...")

[0265] br_final = BayesianRidge(alpha_1=1e-6, alpha_2=1e-6).fit(X_train, y_std_train)

[0266] # Final LightGBM Model

[0267] final_early_stopping = lgb.early_stopping(

[0268] stopping_rounds=EARLY_STOPPING_ROUNDS,

[0269] verbose=EARLY_STOPPING_VERBOSE )

[0271] lgb_final = lgb.LGBMRegressor(**lgb_base_params, random_state=42,verbose=-1)

[0272] lgb_final.fit(

[0273] X_train, y_std_train,

[0274] eval_set=[(X_test, y_std_test)], # Use the validation set as the early stopping validation set

[0275] callbacks=[final_early_stopping]

[0276] Step S3034: Construct a stacked ensemble weight calculation system based on K-fold cross-validation to solve for the optimal fusion weights of the Bayesian Ridge Regression model and the hyperparameter-optimized LightGBM model.

[0277] Specifically, stacking is used to calculate the weights of Bayesian Ridge Regression and LightGBM based on K-fold cross-validation, couple the linear model (Bayesian Ridge Regression) and the tree model (LightGBM), and solve for the optimal fusion weights of Bayesian Ridge Regression and LightGBM to improve the overall prediction accuracy.

[0278] The model training results are quantitatively analyzed using three metrics: coefficient of determination (R²), root mean square error (RMSE), and mean absolute error (MAE). For example... Figure 6 The diagram shown illustrates the process of building and training a stacked ensemble model.

[0279] Step S3035: Using the optimal fusion weights, the cross-validation prediction results of the Bayesian Ridge Regression model and the hyperparameter-optimized LightGBM model are weighted and fused to obtain stacked prediction values. The model training effect is then evaluated using the original data and standard data of each feature parameter.

[0280] Specifically, the cross-validation prediction results of the two models are weighted and fused using the optimal weights to obtain the stacked prediction values ​​of the training set, and the training effect is evaluated through a dual-indicator system (standardization and original target variables).

[0281] Step S3036: Train a stacked ensemble model based on all the data corresponding to the feature parameters, and evaluate and verify the training results using the validation set to obtain a trained prediction model.

[0282] Specifically, the final ensemble model is trained based on the full training set data, and the result of K-fold cross-validation of the validation set is used as the LightGBM early stopping validation set to ensure the model's generalization ability.

[0283] After model training, the accuracy of the model results needs to be verified using at least 10% of the feature parameter data from the road waterlogging scenario parameter database. The feature parameters in the validation set are used as input to the prediction model to predict the target, yielding the model prediction results. These results are then compared with the corresponding road and rainfall waterlogging scenario data in the database. The model's prediction performance is verified using three indicators: coefficient of determination, root mean square error (RMSE), and mean absolute error (MAE). By calculating multiple indicators for both standardized and unstandardized target variables, a comprehensive and accurate assessment of the model's prediction performance is achieved. For standardized target variables, the core indicators—coefficient of determination, root mean square error, and MAE—are calculated to reflect the model's fitting accuracy in the standardized data space. Secondly, the above four types of indicator calculations are repeated for unstandardized target variables to ensure consistency between the evaluation results and actual application scenarios, highlighting the physical meaning and adaptability of the predicted values.

[0284] The method for calculating the drainage and flood control capacity of low-lying urban roads provided in this embodiment balances the stability of model training and the interpretability of results by dividing the training set and validation set and retaining the original data and standard data. It optimizes the hyperparameters of LightGBM using Bayesian optimization, constructs a stacked weight system by combining K-fold cross-validation, and integrates the advantages of linear models and LightGBM to improve prediction accuracy. The model performance is guaranteed by weighted fusion and dual data evaluation. Finally, the stacked ensemble model trained with full data avoids overfitting and enhances generalization ability, which can more accurately and stably support road water accumulation prediction and improve the reliability of drainage and flood control capacity calculation.

[0285] Step S304: Obtain the actual characteristic parameters of the target road based on the prediction target, and use the prediction model to predict the flooded area or water depth of the target road. For details, please refer to [link to relevant documentation]. Figure 2 Step S204 of the illustrated embodiment will not be described again here.

[0286] In one specific implementation, the goal was to predict the area of ​​urban roads submerged by more than 15 cm under different rainfall intensities. Five urban roads with a history of flooding were selected, and a road feature parameter database and a road flooding scenario database were constructed through data collection, field surveys, and numerical simulations. The construction results are shown in Tables 5 and 6. The road feature parameter database contains 50 sets of rainfall flooding scenario data for the 5 urban roads. The first 45 sets of data were used for feature parameter selection, model construction, and training; the last 5 sets of data were used to verify the prediction effect.

[0287] Table 5 Characteristic Indicators of Typical Road Drainage Zones

[0288]

[0289] Table 6 Typical Road Rainfall Scenario Database

[0290]

[0291] Through the process of "database preprocessing → feature parameter screening", nine key feature subsets that contribute most to the prediction target "road flooding area" were selected from the original 12 numerical features (including road feature parameters and rainfall feature data). Furthermore, using RMSE as a quantitative indicator, the impact of selecting different numbers of indicators ranked by contribution on the model prediction results was compared to determine the optimal number of features, which will serve as input parameters for subsequent model construction and training. The results are shown in Tables 7 and 8.

[0292] Table 7. Contribution of Key Characteristic Indicators ("Road Flooded Area")

[0293]

[0294] Table 8. Impact of Indicator Quantities on Verification Results ("Road Flooded Area")

[0295]

[0296] By comparing the effects of cross-validation of the number of feature indicators, and taking into account the model prediction effect, the number and dimension of input factors, and the ease of obtaining the indicators, seven parameters were finally selected as input factors for subsequent model construction: maximum hourly rainfall, rainfall, permeable pavement rate, average slope of waterlogged roads, average spacing of storm drain grates on waterlogged roads, storm drain density, and comprehensive runoff coefficient.

[0297] Using the seven selected core feature parameters as input data, the model was built and trained through the steps and methods of model construction and training, and prediction performance verification. The last five sets of data were used for prediction performance verification. The accuracy of the model training and testing results was verified by three indicators: coefficient of determination, root mean square error, and mean absolute error.

[0298] Table 9. Target Training and Verification Results for Road Flooding Area

[0299]

[0300] The error between the prediction results of the 5 sets of data in the test set and the corresponding waterlogging scenario data in the road waterlogging scenario database is as follows: Figure 7 As shown, the results indicate that the prediction error for the flooded area of ​​roads in the test set is within 50%.

[0301] In another specific embodiment, the objective is to predict the maximum water depth on urban roads under different rainfall intensities. Five urban roads that have historically experienced waterlogging were selected. Through data collection, field surveys, and numerical simulations, a road feature parameter database and a road waterlogging scenario database were constructed. The results are shown in Tables 5 and 6. The road feature parameter database contains 50 sets of rainfall waterlogging scenario data for the 5 urban roads. The first 45 sets of data were used for feature parameter selection, model construction, and training; the last 5 sets of data were used to verify the prediction results.

[0302] Through a process of "database preprocessing → feature parameter screening," nine key feature subsets that contribute most to the prediction target "maximum road water depth" were selected from the original 12 numerical features (including road feature parameters and rainfall feature data). Furthermore, using RMSE as a quantitative indicator, the impact of selecting different numbers of indicators ranked by contribution on the model's prediction results was compared to determine the optimal number of features, which will then serve as input parameters for subsequent model construction and training. The results are shown in Tables 10 and 11.

[0303] Table 10 Contribution of Key Characteristic Indicators (Maximum Road Flood Depth)

[0304]

[0305] Table 11 Impact of the Number of Indicators on Verification Results (Maximum Road Water Accumulation Depth)

[0306]

[0307] By comparing the effects of cross-validation of the number of feature indicators, and taking into account the model prediction effect, the number and dimension of input factors, and the ease of obtaining indicators, five parameters will be used as input factors: rainfall, maximum hourly rainfall, standard deviation of slope of waterlogged roads, volume of rainwater pipes per unit catchment area, and comprehensive runoff coefficient.

[0308] Using the five selected core feature parameters as input data, the model was built and trained through the steps and methods of model construction and training, and prediction performance verification. The last five sets of data were used for prediction performance verification. The accuracy of the model training and testing results was verified by three indicators: coefficient of determination, root mean square error, and mean absolute error.

[0309] Table 12 Target Training and Verification Results for Maximum Road Water Accumulation Depth

[0310]

[0311] The error between the prediction results of the 5 sets of data in the test set and the corresponding waterlogging scenario data in the road waterlogging scenario database is as follows: Figure 8As shown in the figure. The results indicate that the prediction error for the maximum water depth on roads in the test set is within 50%.

[0312] This embodiment also provides a device for calculating the drainage and flood control capacity of low-lying urban roads. This device is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0313] This embodiment provides a device for calculating the drainage and flood control capacity of low-lying urban roads, such as... Figure 9 As shown, it includes:

[0314] The road feature parameter library construction module 901 is used to acquire road feature data, quantify the road feature data, and construct a road feature parameter library.

[0315] The road waterlogging scenario parameter library construction module 902 is used to obtain the rainfall characteristic data and corresponding road waterlogging data based on the road characteristic data, and to construct the road waterlogging scenario parameter library in combination with the road characteristic parameter library. The road waterlogging data includes the road flooded area and the road waterlogging depth.

[0316] The parameter selection and model training module 903 is used to select feature parameters from the road flooding area or road water depth as the prediction target, and use the feature parameters as independent variables to train the model to obtain a trained prediction model.

[0317] The actual prediction module 904 is used to obtain the actual characteristic parameters of the target road based on the prediction target, and to use the prediction model to predict the flooded area or water depth of the target road.

[0318] In some optional implementations, the road feature parameter library construction module 901 includes:

[0319] The first data quantization unit is used to obtain the road area and the total area of ​​the road catchment area, calculate the ratio of the road area to the road catchment area, and obtain the quantification result of the road area ratio as a parameter representing the basic information of the road.

[0320] The second data quantification unit is used to obtain the underlying surface type, the area and runoff coefficient of each underlying surface type within the road catchment area. The comprehensive runoff coefficient is calculated using the area weighting method. At the same time, the ratio of the permeable pavement area to the hardened underlying surface area within the road catchment area is calculated to obtain the permeable pavement rate. The comprehensive runoff coefficient and the permeable pavement rate are used as the quantification results of the source emission reduction capacity characterization parameters.

[0321] The third data quantification unit is used to calculate the quantification results of the average nearest distance of the storm drain grate, the density of the storm drain network, and the volume of the storm drain network per unit catchment area as parameters characterizing the process flow and discharge capacity.

[0322] The fourth data quantification unit is used to calculate the ratio of the cumulative storage volume to the total area of ​​the catchment area as the control ratio of the centralized storage facility, and to calculate the pump facility improvement capacity data based on the pump performance parameters. The centralized storage facility control ratio and the pump facility improvement capacity data are used as the quantification results of the end-point storage capacity characterization parameters.

[0323] The fifth data quantization unit is used to calculate the quantification results of the average road slope and the standard deviation of road slope as parameters representing the road's drainage capacity based on road terrain raster data.

[0324] In some optional implementations, the road flooding scenario parameter library construction module 902 includes:

[0325] The raw data acquisition unit is used to acquire rainfall characteristic data and corresponding road water accumulation data for multiple different roads and multiple rainfall events through historical water accumulation monitoring data and / or numerical model simulation.

[0326] The data processing unit is used to classify and organize the road characteristic parameter database, rainfall characteristic data and corresponding road waterlogging data according to different roads and different rainfall events, and to build a road waterlogging scenario parameter database.

[0327] In some optional implementations, the parameter selection and model training module 903 includes:

[0328] The data standardization processing unit is used to standardize the data in the road waterlogging scenario parameter database to obtain multiple sets of standard feature parameters, and then divide the multiple sets of standard feature parameters into training set and validation set according to a preset ratio.

[0329] The feature parameter screening unit uses root mean square error as the feature screening index, and employs random forest regressors and cross-validation recursive feature elimination to screen each standard feature parameter in the training set, and determines the key feature that contributes the most to the prediction target as the corresponding feature parameter.

[0330] The cross-validation prediction result determination unit is used to optimize the hyperparameters of the LightGBM model using the Bayesian optimization method, obtain the hyperparameter-optimized LightGBM model, and train the Bayesian ridge regression model and the hyperparameter-optimized LightGBM model using the training set to determine the optimal number of iterations and the cross-validation prediction results.

[0331] The fusion weight determination unit is used to construct a stacked ensemble weight calculation system based on K-fold cross-validation to solve for the optimal fusion weights of the Bayesian ridge regression model and the hyperparameter-optimized LightGBM model.

[0332] The stacked fusion prediction unit is used to perform weighted fusion of the cross-validation prediction results of the Bayesian Ridge Regression model and the hyperparameter-optimized LightGBM model using the optimal fusion weights to obtain stacked prediction values, and to evaluate the model training effect using the original data and standard data of each feature parameter.

[0333] The stacked model training unit is used to train a stacked ensemble model based on all the data corresponding to the feature parameters, and to evaluate and verify the training results using a validation set to obtain a trained prediction model.

[0334] The urban low-lying road drainage and flood control capacity calculation device provided in this embodiment of the invention can execute the urban low-lying road drainage and flood control capacity calculation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method. Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0335] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.

[0336] The following is a detailed reference. Figure 10 This diagram illustrates a suitable structural schematic for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from memory 1008 into random access memory (RAM) 1003. The RAM 1003 also stores various programs and data required for the operation of the electronic device. The processor 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0337] Typically, the following devices can be connected to the I / O interface 1005: input devices 1006 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1007 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; memory devices 1008 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows electronic devices to exchange data via wireless or wired communication with other devices. Although Figure 10Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0338] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 1009, or installed from a memory 1008, or installed from a ROM 1002. When the computer program is executed by the processor 1001, it performs the functions defined in the method for calculating the drainage and flood control capacity of urban low-lying roads according to embodiments of the present invention.

[0339] Figure 10 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0340] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the method for calculating the drainage and flood control capacity of urban low-lying roads shown in the above embodiments is implemented.

[0341] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0342] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for calculating the drainage and flood control capacity of low-lying urban roads, characterized in that, The method includes: Acquire road characteristic data and quantify it to construct a road characteristic parameter library. The road characteristic data includes: basic road information characterization parameters, source emission reduction capacity characterization parameters, process emission capacity characterization parameters, end-of-pipe storage capacity characterization parameters, and road flow capacity characterization parameters. Quantifying the road characteristic data to construct the road characteristic parameter library includes: acquiring the road area and the total area of ​​the road catchment area; calculating the ratio of the road area to the road catchment area to obtain the road area ratio as the quantified result of the basic road information characterization parameters; acquiring the underlying surface type within the road catchment area, the area corresponding to each underlying surface type, and the runoff coefficient; calculating the comprehensive runoff coefficient using an area-weighted method; and simultaneously calculating the road catchment area... The permeable pavement ratio is obtained by calculating the ratio of the permeable pavement area to the hardened underlying surface area within the specified range. The comprehensive runoff coefficient and the permeable pavement ratio are used as quantitative results of the source emission reduction capacity characterization parameters. The average nearest distance of storm drain grates, storm drain network density, and storm drain network volume per unit catchment area are calculated as quantitative results of the process discharge capacity characterization parameters. The ratio of cumulative storage capacity to the total catchment area is calculated as the centralized storage facility control ratio. Based on pump performance parameters, the pump facility improvement capacity data is calculated. The centralized storage facility control ratio and pump facility improvement capacity data are used as quantitative results of the end-point storage capacity characterization parameters. Based on road topographic raster data, the average road slope and the standard deviation of road slope are calculated as quantitative results of the road drainage capacity characterization parameters. Based on the road feature data, rainfall feature data and corresponding road waterlogging data are obtained, and a road waterlogging scenario parameter library is constructed by combining the road feature parameter library. The road waterlogging data includes the road flooded area and the road waterlogging depth. Using the road flooding area or road water depth as the prediction target, feature parameters are selected from the road waterlogging scenario parameter database according to the prediction target, and the feature parameters are used as independent variables for model training to obtain a trained prediction model. This includes: standardizing the data in the road waterlogging scenario parameter database to obtain multiple sets of standard feature parameters, and dividing these sets of standard feature parameters into training and validation sets according to a preset ratio; using root mean square error as the feature selection index, a random forest regressor and cross-validation recursive feature elimination are used to select the standard feature parameters in the training set, determining the key features that contribute the most to the prediction target as the corresponding feature parameters; and optimizing the hyperparameters of the LightGBM model using a Bayesian optimization method to obtain the hyperparameter-optimized LightGBM model. The BM model is used to train a Bayesian ridge regression model and a hyperparameter-optimized LightGBM model using the training set, determining the optimal number of iterations and cross-validation prediction results. A stacked ensemble weight calculation system is constructed based on K-fold cross-validation to solve for the optimal fusion weights of the Bayesian ridge regression model and the hyperparameter-optimized LightGBM model. Using the optimal fusion weights, the cross-validation prediction results corresponding to the Bayesian ridge regression model and the hyperparameter-optimized LightGBM model are weighted and fused to obtain stacked prediction values. The model training effect is evaluated using the original data and standard data of each feature parameter. The stacked ensemble model is trained based on all data corresponding to the feature parameters to obtain the trained prediction model, and the training results are evaluated and validated using a validation set. Based on the prediction target, the actual characteristic parameters of the target road are obtained, and the prediction model is used to predict the flooded area or water depth of the target road.

2. The method according to claim 1, characterized in that, The rainfall characteristics data for each rainfall event include: the cumulative rainfall amount of the rainfall event and the maximum rainfall per unit time; the road waterlogging data includes: the area of ​​road waterlogging or the maximum depth of waterlogging exceeding the depth threshold. Based on the road feature data, rainfall characteristic data for each event and corresponding road waterlogging data are obtained. Combined with the road feature parameter library, a road waterlogging scenario parameter library is constructed, including: By using historical waterlogging monitoring data and / or numerical model simulations, we can obtain rainfall characteristic data and corresponding road waterlogging data for multiple different roads and multiple rainfall events. Based on different roads and different rainfall events, the road feature parameter database, the rainfall feature data of the aforementioned events, and the corresponding road waterlogging data are classified and organized to construct a road waterlogging scenario parameter database.

3. The method according to claim 1, characterized in that, Using root mean square error as the feature selection metric, a random forest regressor and cross-validation recursive feature elimination are used to select standard feature parameters in the training set, including: Configure a K-fold randomized cross-validation strategy to group the standard feature parameters and obtain multiple sets of standard feature data; By embedding the random forest regressor into a cross-validation recursive feature elimination framework, a feature selection function is obtained. Using the root mean square error as the feature selection index, the feature selection function is used to iteratively select each group of standard feature data in the training set. After removing the feature with the lowest contribution to the prediction target each time, the root mean square error of the feature selection function is calculated to determine the optimal feature and the number of optimal features when the root mean square error is minimized. The attribute extraction function of recursive feature elimination using cross-validation is used to calculate and extract the importance of the optimal feature, and the optimal features are sorted from high to low according to their importance. Based on the optimal features after sorting, different numbers of features are selected in sequence to construct performance evaluation models for performance evaluation, and the root mean square error values ​​corresponding to different numbers of features are recorded.

4. A device for calculating the drainage and flood control capacity of low-lying urban roads, characterized in that, The device includes: A road feature parameter library construction module is used to acquire road feature data, quantify the road feature data, and construct a road feature parameter library. The road feature data includes: basic road information characterization parameters, source emission reduction capacity characterization parameters, process emission capacity characterization parameters, end-of-pipe storage capacity characterization parameters, and road flow capacity characterization parameters. Quantifying the road feature data to construct the road feature parameter library includes: acquiring the road area and the total area of ​​the road catchment area; calculating the ratio of the road area to the road catchment area to obtain the road area ratio as the quantified result of the basic road information characterization parameters; acquiring the underlying surface type within the road catchment area, the area corresponding to each underlying surface type, and the runoff coefficient; calculating the comprehensive runoff coefficient using an area-weighted method; and simultaneously calculating... The permeable pavement ratio is obtained by calculating the ratio of the permeable pavement area to the hardened underlying surface area within the road catchment area. The comprehensive runoff coefficient and the permeable pavement ratio are used as quantitative results of the source emission reduction capacity characterization parameters. The average nearest distance of storm drain grates, storm drain network density, and storm drain network volume per unit catchment area are calculated as quantitative results of the process discharge capacity characterization parameters. The ratio of cumulative storage capacity to the total catchment area is calculated as the centralized storage facility control ratio. Based on pump performance parameters, the pump facility improvement capacity data is calculated. The centralized storage facility control ratio and the pump facility improvement capacity data are used as quantitative results of the end-point storage capacity characterization parameters. Based on road topographic raster data, the average road slope and the standard deviation of road slope are calculated as quantitative results of the road drainage capacity characterization parameters. The road waterlogging scenario parameter library construction module is used to obtain the rainfall characteristic data and the corresponding road waterlogging data based on the road characteristic data, and to construct the road waterlogging scenario parameter library in combination with the road characteristic parameter library. The road waterlogging data includes the road flooding area and the road waterlogging depth. The parameter selection and model training module is used to select feature parameters from a road flooding area or road water depth prediction target based on the prediction target. These feature parameters are then used as independent variables for model training to obtain a trained prediction model. The module includes: standardizing the data in the road waterlogging scenario parameter database to obtain multiple sets of standard feature parameters, and dividing these sets into training and validation sets according to a preset ratio; using root mean square error as the feature selection metric, employing a random forest regressor and cross-validation recursive feature elimination to select the standard feature parameters in the training set, identifying the key features that contribute most to the prediction target as the corresponding feature parameters; and optimizing the hyperparameters of the LightGBM model using a Bayesian optimization method to obtain the optimized hyperparameters. The LightGBM model is used, and a Bayesian ridge regression model and a hyperparameter-optimized LightGBM model are trained using the training set to determine the optimal number of iterations and cross-validation prediction results. A stacked ensemble weight calculation system is constructed based on K-fold cross-validation to solve for the optimal fusion weights of the Bayesian ridge regression model and the hyperparameter-optimized LightGBM model. Using the optimal fusion weights, the cross-validation prediction results corresponding to the Bayesian ridge regression model and the hyperparameter-optimized LightGBM model are weighted and fused to obtain stacked prediction values. The training effect of the model is evaluated using the original data and standard data of each feature parameter. The stacked ensemble model is trained based on all data corresponding to the feature parameters to obtain the trained prediction model, and the training results are evaluated and validated using the validation set. The actual prediction module is used to obtain the actual characteristic parameters of the target road based on the prediction target, and to use the prediction model to predict the flooded area or water depth of the target road.

5. An electronic device, characterized in that, include: A memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, the processor executing the computer instructions to perform the method of any one of claims 1 to 3.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the method of any one of claims 1 to 3.

7. A computer program product, characterized in that, Includes computer instructions for causing a computer to perform the method of any one of claims 1 to 3.

Citation Information

Patent Citations

  • Urban inland inundation simulation method based on topographic feature deep learning

    CN116933621A