Decision tree urban waterlogging meteorological model based on historical rainfall percentile and prediction method

Through the decision tree model based on historical precipitation percentile, the existing urban flooding meteorological warning model has been solved, and the early identification and forecasting of urban flooding has been realized, and the forecasting accuracy and real-time performance have been improved.

CN119940490AActive Publication Date: 2025-05-06NATIONAL METEOROLOGICAL CENTRE
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510083495.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-06
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The existing urban flooding meteorological warning model has a large amount of calculation, high basic data requirements, and is not suitable for real-time online applications. There are few machine learning methods, resulting in insufficient research on urban flooding forecast and early warning.

Method used

The decision tree urban flooding meteorological model based on historical precipitation percentage is used to determine the key impact factors of urban flooding by obtaining previous years' waterlogging data and no flooding data, and a decision regression tree model is constructed, and the model threshold is updated to consider the precipitation historical percentile.

Benefits of technology

It has realized early identification and forecasting of urban waterlogging, improved forecasting accuracy and real-time performance, and is suitable for scientific basis for urban flood control and flood prevention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940490A_ABST
    Figure CN119940490A_ABST
Patent Text Reader

Abstract

The invention discloses a decision-making tree urban waterlogging meteorological model based on historical rainfall percentile and a prediction method thereof, which are used for researching rainfall disaster-inducing factors of urban waterlogging on the basis of an urban waterlogging disaster situation, discovering and confirming that short-time heavy rainfall is a common characteristic of disaster occurrence, and predicting the urban waterlogging meteorological model based on the decision-making tree. According to the method, accumulated rainfall and heavy rainfall durations with the maximum different durations are selected as main disaster-causing meteorological factors, dichotomy data are established, 80% is used for training, 20% is used for testing, an optimal decision regression tree model with the maximum submerging depth is established through parameter adjustment by using a CART decision tree regression method, and on this basis, the optimal decision regression tree model is established through parameter adjustment. The historical percentile of the rainfall factor is introduced into the decision tree regression model, the decision tree regression model based on the historical percentile of rainfall considering regional features and site particularity is obtained, the model effect is better than that of a decision tree model not considering the historical percentile of rainfall, and the method can be used for early recognition and forecast of urban waterlogging.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of meteorological forecasting, and in particular relates to a decision tree urban waterlogging meteorological model and a forecasting method based on historical precipitation percentiles. Background Art

[0002] Urban waterlogging is caused by continuous rainfall or heavy rainfall that exceeds the drainage capacity of the urban drainage network system, making it impossible for rainwater to be discharged, resulting in waterlogging disasters on the ground. Meteorological warning for urban waterlogging is one of the effective means of urban disaster prevention and mitigation, and the research on warning models is the key to meteorological warning. Using quantitative precipitation forecasts to drive urban waterlogging models to predict the depth of waterlogging in the next few hours can increase the effective forecast period of urban waterlogging forecasts, improve forecast accuracy, and provide a scientific basis for urban flood control and waterlogging prevention.

[0003] However, the current research on urban waterlogging mainly focuses on waterlogging simulation and risk assessment, while there is less research on waterlogging forecasting and warning. In addition, the hydrological waterlogging model has high requirements for basic data in terms of business application, which is not suitable for widespread implementation. In addition, its large amount of calculation affects the real-time online application effect of the model, which affects the real-time implementation of urban waterlogging meteorological warning work. In recent years, machine learning methods have developed rapidly and achieved many results. However, due to the lack of urban waterlogging data, urban waterlogging machine learning methods are rarely used. Summary of the invention

[0004] Purpose of the invention: In view of the above-mentioned existing problems and shortcomings, the purpose of the present invention is to provide a decision tree urban waterlogging meteorological model and prediction method based on historical precipitation percentiles. The present invention establishes a regional decision tree regression model for regional urban waterlogging. On the basis of the analysis of disaster-causing factors, a decision tree regression model is established using the main meteorological factors that cause urban waterlogging. On this basis, the distribution characteristics of historical precipitation of different durations are considered to establish a decision tree urban waterlogging model, thereby realizing early identification of urban waterlogging.

[0005] Technical solution: To achieve the above invention purpose, the present invention adopts the following technical solution: a decision tree urban waterlogging meteorological model based on historical precipitation percentiles, comprising the following steps: S1, data acquisition: Obtain historical waterlogging data and historical waterlogging-free data of cities across the country as model samples, the historical waterlogging data include the time of urban waterlogging, maximum flooding depth, 24-hour precipitation from 1959 to 2023, and 1-hour precipitation from 2013 to 2023; if there is no waterlogging disaster within 3 days from the date of the occurrence of waterlogging disaster at the disaster site, extract the precipitation data of the third day after the occurrence of waterlogging disaster at the disaster site to establish historical waterlogging-free data; 80% of the model samples are used as training samples, and the remaining 20% ​​are used as test samples; S2, determine the key influencing factors of urban waterlogging, with the maximum accumulated precipitation in 1 hour, the maximum accumulated precipitation in 2 hours, the maximum accumulated precipitation in 3 hours, the maximum accumulated precipitation in 4 hours, the maximum accumulated precipitation in 5 hours, the maximum accumulated precipitation in 6 hours, the maximum accumulated precipitation in 12 hours, the maximum accumulated precipitation in 24 hours and the duration of heavy precipitation as the key influencing factors of the maximum flooding depth of the city; S3, establishment of a decision regression tree waterlogging model, building a decision tree waterlogging regression model with the maximum limit depth and the minimum training sample parameters, and testing the model with the training samples to obtain a tested decision regression tree waterlogging model; S4, in each node of the decision regression tree waterlogging model, the decision tree threshold is updated according to the value corresponding to the historical percentile of precipitation to complete the model update.

[0006] Furthermore, the model updating process in step S4 includes the following steps: (1) In each node of the decision regression tree waterlogging model, search for the sample closest to the decision tree threshold of the node among all the sample data of the node, and obtain the precipitation value of the closest sample; (2) Identify the station that is closest to the sample and sort the historical precipitation of the station; (3) Based on the precipitation closest to the sample, calculate the percentile of the historical precipitation ranking of the station to which it belongs; (4) Sort the historical precipitation of other stations in the node respectively, search the historical precipitation of each station according to the percentile determined in step (3), and obtain the precipitation value of the corresponding percentile in each station; (5) The corresponding percentile precipitation values ​​of each station obtained in step (4) are used as the thresholds of the node in turn to verify the accuracy of the model obtained by the sample evaluation test, and the one with the highest accuracy is selected as the threshold of the node; (6) The threshold values ​​of other nodes in the decision regression tree waterlogging model are updated according to steps (1) to (5).

[0007] Furthermore, in step S3, the value of the maximum limit depth is 2, and the value of the minimum training sample parameter is 7.

[0008] The present invention also provides a prediction method for a decision tree urban waterlogging meteorological model based on historical precipitation percentiles. Based on the above decision regression tree waterlogging model, by inputting quantitative precipitation estimation data and / or quantitative precipitation forecast data, a future urban waterlogging meteorological risk forecast is obtained.

[0009] Furthermore, the future urban waterlogging meteorological risk forecast is divided into blue, yellow, orange and red warning levels according to the severity of the waterlogging, and the maximum flooding depth ranges corresponding to the blue, yellow, orange and red warning levels are 0.05m~0.2m, 0.21m~0.35m, 0.36m~0.5m and greater than 0.5m, respectively.

[0010] Beneficial effects: Compared with the prior art, the present invention adopts a decision regression tree and uses the maximum cumulative precipitation of 3 hours, 6 hours, 1 hour, 12 hours and 24 hours as thresholds to predict the maximum water depth of urban waterlogging. On this basis, the decision tree regression model is introduced with the historical percentile of the precipitation factor to obtain a decision tree regression model based on the historical percentile of precipitation that considers both regional characteristics and site particularity. The model effect is better than the decision tree model that does not consider the historical percentile of precipitation, and can be used for early identification and prediction of urban waterlogging. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 It is a flow chart of the prediction method of the decision tree urban waterlogging meteorological model based on historical precipitation percentiles according to the present invention; Figure 2 ROC curve of the Logistic regression model test of the decision tree regression model parameters of the present invention; Figure 3 It is a structural diagram of the decision tree regression model of the present invention; Figure 4 It is a density function diagram of the maximum 1-12h maximum precipitation of urban waterlogging in the whole country from 2019 to 2022 in an embodiment of the present invention; Figure 5 It is a density function diagram of the maximum cumulative precipitation of urban flooding in 12~24h across the country from 2019 to 2022 in an embodiment of the present invention. DETAILED DESCRIPTION

[0012] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, various equivalent forms of modifications to the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.

[0013] The overall technical route of the present invention is mainly based on the historical disasters of urban waterlogging, and the maximum hourly precipitation of 1-24h and the duration of short-term precipitation are analyzed to find the key factors that cause urban waterlogging. Through the comparison of the effects of multiple regression methods, the best solution is found and confirmed, and then the transformation is carried out. The historical precipitation is introduced into the model for percentile test to obtain the best accuracy, thereby determining the model. The specific steps include: 1. Data Source The present invention collects urban waterlogging cases in China from 2019 to 2023 through the China Meteorological Administration's disaster reporting system and official media information, and matches them with the precipitation on the day. If there is no precipitation near the disaster site, the case is deleted. Through screening, 139 urban waterlogging cases were collected, including the maximum flooding depth at the time and location of urban waterlogging. The 24-hour precipitation from 1959 to 2023 and the historical precipitation data of the 1-hour observation station from 2013 to 2023 are from the National Meteorological Center.

[0014] At present, the disaster data has samples of urban waterlogging. The establishment of a decision tree regression model requires samples that have not occurred. The effect of the model is closely related to the selection of the value of 0. The present invention selects the value of 0 based on the situation that no precipitation disaster has occurred, no waterlogging has occurred before precipitation, etc., that is, when there is no waterlogging disaster within 3 days from the date of the occurrence of the waterlogging disaster at the disaster site, the precipitation data on the third day of the occurrence of the waterlogging disaster at the disaster site is extracted to establish the data of no waterlogging in previous years; after deduplication, the previous precipitation is extracted. After determining the modeling samples in this way, 80% of the samples are modeled and 20% are used for testing.

[0015] 2. Determination of urban waterlogging disaster factors The time of urban waterlogging in the present invention is accurate to the day. In order to analyze the characteristics of disaster-causing precipitation, the hourly precipitation on the day of urban waterlogging and the two days before it is extracted, and the 1-24-hour continuous cumulative precipitation is calculated by sliding. The maximum value is taken to analyze the main characteristics of the disaster-causing precipitation, and a density function diagram of the maximum 1-12h maximum precipitation of urban waterlogging in the country from 2019 to 2022 is prepared ( Figure 4 ) and the density function diagram of the maximum cumulative precipitation of 12~24h and the duration of short-term heavy precipitation in urban waterlogging across the country from 2019 to 2022 ( Figure 5 ).

[0016] Through the above analysis, it can be found that the main cause of urban waterlogging disasters is the maximum cumulative precipitation in 1-6 hours. The phase difference of each hourly density map is about 10-20mm, and the maximum cumulative precipitation density distribution in 12-24 hours is similar. Therefore, when modeling the present invention, a total of 9 factors, including the maximum cumulative precipitation in 1-6 hours, 12 hours, 24 hours, and the duration of heavy precipitation, are selected to conduct a decision tree modeling test on the maximum flooding depth of the city.

[0017] 3. Determine the model type In order to select the best method, the present invention uses four regression methods for testing, namely multiple regression, KNN, random forest and decision tree. In the KNN method, 5 nearest neighbors are considered and 30 leaf nodes are used. The random forest setting selects 3 leaf nodes for the latter three methods. The accuracy, absolute error and mean square error of the training data and the validation data are obtained by the five-fold cross validation method. The urban waterlogging meteorological regression model is established according to the four methods, and the accuracy and error of the four regression methods are compared, as shown in Table 1 below.

[0018] Table 1 Accuracy, mean absolute error and mean relative error of four regression methods

[0019] As can be seen from Table 1, for the training samples, random forest has the best effect, with an accuracy of 0.9, while KNN and decision tree methods are similar, at 0.61 and 0.63, respectively. The worst is multivariate regression, at only 0.53. However, after testing the validation samples, it was found that there was no significant difference in the effects of the four methods. The KNN method was relatively higher, while the random forest method was relatively lower, but the difference was too small to be representative. The difference in mean absolute error and mean square error was also small. The result of the validation sample is the key to the selection. Although random forest has the best effect on the test sample, it has no obvious advantage in accuracy for the validation sample.

[0020] Judging from the test results, there is no sufficient reason to choose a certain method. From the perspective of transparency and interpretability, machine learning methods include white box and black box models. The white box model emphasizes the high transparency and interpretability of the model, such as linear regression and decision tree methods, while the black box emphasizes predictability and ignores interpretability, such as random forests. In order to facilitate forecasters to better understand the internal structure and decision logic of the model, the present invention generates output based on input, and it is a better choice to establish an urban waterlogging meteorological forecast model through the decision tree method.

[0021] 4. Determination of decision tree regression model parameters In order to find better decision tree regression model parameters, two important parameters of growth and decision tree pruning, the maximum limit depth and the minimum training sample parameters, were tested and cross-validated. The specific test values ​​are shown in Table 2.

[0022] Table 2 Main test parameters

[0023] The test obtains the evaluation results of various parameters, such as Figure 2 As shown in the figure, it can be found that the best effect is achieved when the maximum limit depth is 2 and the minimum training sample parameter is 7. At this time, the AUC area of ​​the ROC curve is 0.82. Then, the parameters such as the number of samples for each branch are debugged, and the parameters whose test sample and experimental sample scores tend to be stable and consistent are selected for modeling.

[0024] 5. Building a decision tree regression model After testing, the decision regression tree waterlogging model is obtained as follows Figure 3, node in the figure is a node, X[i] represents i variables, indicating the maximum accumulated precipitation in a certain hour. From the decision regression tree, it can be found that the model uses the maximum accumulated precipitation in 3, 24, 1, and 12 hours and the duration of short-term heavy precipitation. The first layer is the maximum accumulated precipitation in 3 hours, and its threshold is 48.9mm. The mean square error of the model for samples without urban flooding is less than 0.001, while the root mean square error of the maximum waterlogging depth predicted for samples that do occur is 0.1~0.3, and the mean square error of the maximum flooding depth predicted to exceed 1 meter is 0.348. It can be seen that the effect is acceptable. At the same time, the accuracy of the test data after modeling is 0.63, and the validation data score is 0.43.

[0025] 6. Introducing the decision tree model into the historical percentile of precipitation In order to improve the accuracy of the test, the present invention introduces the historical percentile of precipitation into the decision tree based on the threshold of the decision tree. The main method is to test each percentile when 80% of the samples exceed the percentile of the station represented by the threshold, find the percentile where the precipitation is relatively exceeded in the historical ranking of the corresponding precipitation, and then extract the percentile of each meteorological station, update the threshold and then build a model to test its effect. The specific process is as follows: (1) In each node of the decision regression tree waterlogging model, search for the sample closest to the decision tree threshold of the node among all the sample data of the node, and obtain the precipitation value of the closest sample; (2) Identify the station that is closest to the sample and sort the historical precipitation of the station; (3) Using the precipitation closest to the sample as the standard, calculate the percentile of the historical precipitation ranking of the station to which it belongs; (4) Sort the historical precipitation of other stations in the node respectively, search the historical precipitation of each station according to the percentile determined in step (3), and obtain the precipitation value corresponding to the percentile of each station; (5) The corresponding percentile precipitation values ​​of each station obtained in step (4) are used as the thresholds of the node respectively, and the accuracy of the sample evaluation test model is tested, and the one with the highest accuracy is selected as the threshold of the node; (6) The thresholds of other nodes in the decision regression tree waterlogging model are updated according to steps (1) to (5). As shown in Table 3, the accuracy after the historical precipitation percentile is introduced.

[0026] Table 3 Accuracy before and after the introduction of historical precipitation percentiles

[0027] According to the results, it can be found that after introducing the historical percentile of precipitation, the overall accuracy of the test samples has been improved to 0.51, and the mean square error has been reduced from 0.13 to 0.11. The effect has been improved, especially for samples with a flooding depth greater than 0.5 meters, the mean square error has been reduced from 0.45 meters to 0.33 meters, and the improvement effect is obvious.

[0028] 7. Decision Tree Model Prediction Taking quantitative precipitation estimation and forecast as input data, the improved urban waterlogging decision tree model is used to predict the meteorological forecast of the maximum inundation depth of urban waterlogging.

[0029] Here, a decision tree regression method that introduces historical percentiles of precipitation is obtained based on regression experiments. Future urban waterlogging meteorological risk forecasts are made based on quantitative precipitation estimation (QPE) and quantitative precipitation forecast (QPF). The maximum inundation depth is forecasted and warned according to the blue, yellow, orange and red warning levels corresponding to the intervals of 0.05m~0.2m, 0.21m~0.35m, 0.36m~0.5m and greater than 0.5m. Example

[0030] The effect of the forecast model of the present invention is tested through the following examples: From 22:00 on July 10, 2024 to 7:00 on July 11, heavy rains to torrential rains occurred in Dianjiang County. Affected by the heavy rainfall, waterlogging occurred in Chengxi Town, Dianjiang County, and the water depth of some sections was about 2 meters, causing significant economic losses.

[0031] It can be seen that in addition to future precipitation forecasts, the model also captures historical precipitation characteristics. At the same time, it can also capture the characteristics of urban waterlogging for recent actual precipitation, which has a good indicative significance for the process.

[0032] In addition, 34 cases of urban waterlogging collected in 2024 were tested. Among the 34 cases, there were 7 red warnings, 6 orange warnings, 5 yellow warnings, and 4 blue warnings, with a hit rate of 65%, and 12 cases of missed reports, accounting for 35%, which is highly indicative.

[0033] In summary, based on the decision tree threshold, the present invention introduces the historical percentile of precipitation into the decision tree, and models it after updating the threshold. It is found that after introducing the historical percentile of precipitation, the overall accuracy of the test samples is improved to 0.51, and the mean square error is reduced from 0.13 to 0.11, and the effect is improved. In particular, for samples with a flooding depth greater than 0.5 meters, the mean square error is reduced from 0.45 meters to 0.33 meters, and the improvement effect is obvious.

[0034] The present invention is based on the urban waterlogging disaster, studies the precipitation disaster factors of urban waterlogging, finds and confirms that short-term heavy precipitation is a common feature of disaster occurrence, selects the maximum cumulative precipitation of different durations and the duration of heavy precipitation as the main meteorological factors causing disasters, establishes binary classification data, 80% of which is used for training and 20% for testing, and uses the CART decision tree regression method to establish the optimal decision regression tree model of the maximum flooding depth by adjusting parameters. The decision regression tree uses the maximum cumulative precipitation of 3, 6, 1 hours and 12, 24 hours as the threshold to predict the maximum water accumulation depth of urban waterlogging. On this basis, the decision tree regression model is introduced with the historical percentile of the precipitation factor, and a decision tree regression model based on the historical percentile of precipitation is obtained, which takes into account both regional characteristics and site specificity. The model effect is better than the decision tree model that does not consider the historical percentile of precipitation, and can be used for early identification and prediction of urban waterlogging.

[0035] The above descriptions are only some embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A decision tree urban waterlogging meteorological model based on historical precipitation percentiles, characterized by The following steps are involved: S1, data acquisition: Obtain historical waterlogging data and historical waterlogging-free data of cities across the country as model samples, the historical waterlogging data include the time of urban waterlogging, maximum flooding depth, 24-hour precipitation from 1959 to 2023, and 1-hour precipitation from 2013 to 2023; if there is no waterlogging disaster within 3 days from the date of the occurrence of waterlogging disaster at the disaster site, extract the precipitation data of the third day after the occurrence of waterlogging disaster at the disaster site to establish historical waterlogging-free data; 80% of the model samples are used as training samples, and the remaining 20% ​​are used as test samples; S2, determine the key influencing factors of urban waterlogging, with the maximum accumulated precipitation in 1 hour, the maximum accumulated precipitation in 2 hours, the maximum accumulated precipitation in 3 hours, the maximum accumulated precipitation in 4 hours, the maximum accumulated precipitation in 5 hours, the maximum accumulated precipitation in 6 hours, the maximum accumulated precipitation in 12 hours, the maximum accumulated precipitation in 24 hours and the duration of heavy precipitation as the key influencing factors of the maximum flooding depth of the city; S3, establishment of a decision regression tree waterlogging model, building a decision tree waterlogging regression model with the maximum limit depth and the minimum training sample parameters, and testing the model with the training samples to obtain a tested decision regression tree waterlogging model; S4, in each node of the decision regression tree waterlogging model, the decision tree threshold is updated according to the value corresponding to the historical percentile of precipitation to complete the model update.

2. The decision tree urban waterlogging meteorological model based on historical precipitation percentiles according to claim 1 is characterized by: The model updating process described in step S4 includes the following steps: (1) In each node of the decision regression tree waterlogging model, search for the sample closest to the decision tree threshold of the node among all the sample data of the node, and obtain the precipitation value of the closest sample; (2) Identify the station that is closest to the sample and sort the historical precipitation of the station; (3) Based on the precipitation closest to the sample, calculate the percentile of the historical precipitation ranking of the station to which it belongs; (4) Sort the historical precipitation of other stations in the node respectively, search the historical precipitation of each station according to the percentile determined in step (3), and obtain the precipitation value of the corresponding percentile in each station; (5) The corresponding percentile precipitation values ​​of each station obtained in step (4) are used as the thresholds of the node in turn to verify the accuracy of the model obtained by the sample evaluation test, and the one with the highest accuracy is selected as the threshold of the node; (6) The threshold values ​​of other nodes in the decision regression tree waterlogging model are updated according to steps (1) to (5).

3. The decision tree urban waterlogging meteorological model based on historical precipitation percentiles according to claim 1 is characterized by: In step S3, the value of the maximum limit depth is 2, and the value of the minimum training sample parameter is 7.

4. A prediction method for a decision tree urban waterlogging meteorological model based on historical precipitation percentiles, characterized by: Based on the decision regression tree waterlogging model obtained in any one of claims 1-3, by inputting quantitative precipitation estimation data and / or quantitative precipitation forecast data, a future urban waterlogging meteorological risk forecast is obtained.

5. The prediction method of the decision tree urban waterlogging meteorological model based on historical precipitation percentiles according to claim 4 is characterized by: The future urban waterlogging meteorological risk forecast is divided into blue, yellow, orange and red warning levels according to the severity of the waterlogging. The maximum flooding depths corresponding to the blue, yellow, orange and red warning levels are 0.05m~0.2m, 0.21m~0.35m, 0.36m~0.5m and greater than 0.5m respectively.

Citation Information

Patent Citations

  • Urban inland inundation rapid forecasting method based on multi-output machine learning algorithm

    CN114372625A

  • Urban inland inundation forecasting method based on secondary heterogeneous mode decomposition ESN model

    CN116596133A

  • Waterlogging model creation method and system applied to urban water treatment

    CN118627407A