A typhoon disaster risk assessment and dynamic prediction method

By constructing a machine learning model based on hazard, vulnerability, and exposure indicators, and combining it with real-time meteorological data, the problem of typhoon disaster risk assessment and dynamic forecasting at the district and county levels has been solved, achieving more accurate real-time forecasting and emergency decision-making.

CN116757305BActive Publication Date: 2026-07-31ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG UNIV
Filing Date
2023-03-30
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies are insufficient for dynamic assessment and real-time forecasting of typhoon disaster risks at the district and county level, and cannot support precise real-time emergency decision-making.

Method used

By selecting hazard, vulnerability, and exposure indicators, Pearson correlation analysis and principal component analysis are conducted to construct a machine learning model, which is then dynamically updated in conjunction with real-time meteorological data to achieve real-time forecasting of typhoon disaster risks.

Benefits of technology

It has achieved typhoon disaster risk assessment and dynamic forecasting at the district and county level. The machine learning model has achieved an accuracy of 76% cross-validation and 74% independent sample validation, supporting more refined real-time emergency decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116757305B_ABST
    Figure CN116757305B_ABST
Patent Text Reader

Abstract

This invention discloses a method for typhoon disaster risk assessment and dynamic forecasting, comprising: selecting forecast indicators, including at least hazard indicators, vulnerability indicators, and exposure indicators; extracting hazard indicators, extracting the maximum rainfall, process rainfall, and average maximum wind speed for several preset time periods within the administrative regions of each district / county; correlation analysis, performing Pearson correlation analysis between the selected forecast indicators and loss levels; principal component analysis, reducing the dimensionality of hazard indicators related to rainfall; sample set partitioning; constructing a machine learning model; training and testing the machine learning model; obtaining forecast results: dynamically updating the hazard indicators input to the machine learning model, thereby achieving real-time updated forecasts of typhoon disaster risk. The machine learning model of this invention updates the hazard indicators input to the machine learning model using measured and forecasted meteorological data, achieving hourly real-time updated forecasts for each district / county.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural disaster risk assessment technology, specifically to a method for typhoon disaster risk assessment and dynamic forecasting. Background Technology

[0002] Currently, there are three main methods for assessing natural disaster risks: one is the risk assessment method driven by physical models; the second is the risk assessment method based on indicator systems; and the third is the risk assessment method based on data.

[0003] Physical model-driven risk assessment, conducted by scholars both domestically and internationally, studies the evolution mechanism of disaster hazard, structural vulnerability, and disaster loss. Using physical equations and numerical simulations, it simulates various functional losses, economic losses, and social impacts that infrastructure systems may suffer under natural disasters. Simultaneously, simulation analysis platforms based on complex network theory utilize historical and existing disaster data, infrastructure data, and personnel data to perform static or dynamic simulations, enabling the analysis and assessment of natural disaster risks. The advantage of this method lies in its high precision, supporting refined emergency decision-making. However, it requires a high degree of data precision and completeness, as well as significant computational power. In situations where data is relatively scarce, this method cannot achieve efficient disaster risk assessment for administrative regions above the county level, and it cannot achieve large-scale, province-wide coverage of disaster risk assessment.

[0004] Risk assessment methods based on indicator systems typically construct a disaster risk indicator system based on various disaster risk factors, and then use weighting methods to determine the weights of the indicators. For example, the "National Technical Specification for Meteorological Disaster Risk Assessment (Typhoon)" uses this method for risk assessment. Although this method is easy to operate, it does not consider the interaction mechanism between disaster-causing factors and disaster-bearing bodies. Essentially, it is a relative qualitative risk assessment method, which is a static reflection of the long-term disaster risk level and patterns of disaster-bearing bodies, and cannot be applied to dynamic risk forecasting based on disaster events.

[0005] Many scholars have used historical disaster data combined with data-driven methods (such as mathematical statistics analysis and machine learning) to establish disaster loss prediction models and disaster risk assessment models. The advantage of data-driven models lies in their ability to extract characteristic parameters and key information from complex problems and massive amounts of data, avoiding extensive physical simulation analysis and thus saving computation time. Therefore, compared to other models, data-driven models are more conducive to achieving typhoon disaster risk assessments across all districts and counties in Zhejiang Province.

[0006] Current regional-scale typhoon disaster risk prediction research mainly focuses on the provincial level and above, based on the whole process of typhoon event disaster loss prediction, and has not yet achieved dynamic risk assessment at the county level during the real-time evolution of the disaster field, which cannot support more refined real-time emergency decision-making. Summary of the Invention

[0007] To overcome the shortcomings of the above technologies, this invention provides a method for typhoon disaster risk assessment and dynamic forecasting.

[0008] The technical solution adopted by this invention to overcome its technical problems is:

[0009] A method for typhoon disaster risk assessment and dynamic forecasting, comprising the following steps:

[0010] Step 1: Select forecast indicators: Forecast indicators should include at least hazard indicators, vulnerability indicators, and exposure indicators. Among them, hazard indicators should include at least the rainfall brought by strong winds and typhoons, and exposure indicators should include at least the exposure considering the number of disaster-bearing bodies and the disaster-prone environment considering topography and water systems.

[0011] Step 2: Extracting Hazard Indicators: During a single typhoon event, extract the maximum rainfall, total rainfall, and average maximum wind speed for several preset time periods within the administrative regions of each district and county. The specific time periods can be selected based on the actual situation, such as 1 hour, 3 hours, 6 hours, 12 hours, and 24 hours for the maximum rainfall, total rainfall, and average maximum wind speed.

[0012] Step 3: Correlation Analysis: Perform Pearson correlation analysis between the prediction indicators selected in Step 1 and the loss level. For different types of indicators, select indicators that are greater than the preset correlation threshold as input indicators.

[0013] Step 4: Principal Component Analysis: Principal component analysis is used to explore the degree of correlation between multiple potentially related variables and to find the direction of maximum or minimum correlation. Principal component analysis is used to reduce the dimensionality of risk indicators related to rainfall, thereby achieving the purpose of data compression or noise reduction and avoiding redundant variables that may interfere with subsequent model training.

[0014] Step 5: Sample set partitioning: Divide the samples into full samples and non-zero samples, and divide the variables into combinations of variables after dimensionality reduction and combinations of the original variables. Keep the data of one typhoon event as the validation set, and the data of the other typhoon events as the training set. The full samples include samples with disaster damage data of 0, and the non-zero samples do not include samples with disaster damage data of 0.

[0015] Step 6: Build a machine learning model;

[0016] Step 7: Machine Learning Model Training and Testing: Input the pre-defined training set data into the machine learning model for training. Use the XGBoost algorithm from the sklearn library to train the model, and combine it with the grid hyperparameter tuning algorithm to optimize the objective function and model hyperparameters. Use multiple evaluation metrics to evaluate the performance of the machine learning model. If all preset metrics reach the allowable error range, the training of the machine learning model is complete. Then, input the pre-defined test set into the trained machine learning model to test the actual generalization performance of the machine learning model.

[0017] Step 8: Obtain forecast results: By continuously updating the measured data and the latest forecast meteorological data, and inputting them into the machine learning model constructed in Step 6, the hazard indicators input into the machine learning model are dynamically updated, thereby realizing real-time updated forecasts of typhoon disaster risks.

[0018] Furthermore, in step one, both vulnerability indicators and exposure indicators can be obtained through yearbooks or publicly available data.

[0019] Vulnerability indicators are selected based on the characteristics of a region and the availability of data, such as the proportion of the primary industry, GDP per capita, per capita disposable income of urban residents, per capita disposable income of rural residents, year-end balance of urban and rural residents' savings deposits, mileage of highways within the territory, number of doctors per thousand people, and number of hospital beds per thousand people.

[0020] Exposure indicators are selected based on my country's definition of direct economic loss, which includes at least agricultural losses, infrastructure losses, and household property losses, and in combination with the availability of data in the region. Crop planting area, gross domestic product, agricultural output value, and year-end total population are selected as exposure indicators.

[0021] The disaster-prone environment should at least consider topographic factors and water system factors. Topographic factors include both elevation changes and topographic changes. Average elevation, average slope, average aspect, and river network density are selected as disaster-prone environmental indicators.

[0022] Furthermore, the meteorological stations that record meteorological data during a typhoon event are discretely distributed within the district / county administrative region. This means that each meteorological station has a rainfall and wind speed data. It is necessary to first average the hourly rainfall and hourly average wind speed within the district / county administrative region to obtain the average value of rainfall and wind speed for that hour, which is the representative value of rainfall and wind speed for that district / county administrative region.

[0023] In step two, let the duration of a typhoon event be a hours, and let the administrative area of ​​a certain district or county include b meteorological stations. The maximum rainfall during the N hours of the typhoon event is extracted by using a sliding time window method, and the specific calculation is as follows: (1) to (3):

[0024]

[0025]

[0026]

[0027] In the above formula, x represents the representative value of rainfall or maximum wind speed at time f in the administrative region of that district / county. fz x represents the rainfall or maximum wind speed recorded at the z-th weather station at time f. (N) Let N represent the maximum rainfall over N hours in the administrative region of the district / county, where N ≤ a, and w represent the average maximum wind speed in the administrative region of the district / county.

[0028] Furthermore, in step three, the formula for Pearson correlation analysis is as follows:

[0029]

[0030] In the above formula, X and Y represent two random variables, ρ X,Y The value of μ is between -1 and 1, and it is used to measure the correlation between variables X and Y. X μ Y Let σ represent the mean of two random variables X and Y, respectively. X σ Y These represent the standard deviations of two random variables X and Y, respectively. Generally, a correlation coefficient with an absolute value greater than or equal to 0.6 is considered a strong correlation. When selecting different types of indicators, indicators with a higher correlation within that type should be chosen.

[0031] Furthermore, in step four, after principal component analysis, variables with a cumulative variance contribution rate of 90% are selected as principal components after dimensionality reduction, forming different combinations of variables. These different combinations of variables are then input into the constructed machine learning model for comparison. Principal component analysis is implemented using the dimensionality reduction function in the software SPSS, so the principle will not be elaborated here.

[0032] Furthermore, in step six, constructing the machine learning model specifically includes:

[0033] 1) Construct the objective function

[0034] The objective function is divided into two parts: one part is the loss function and the other part is the regularization function. XGBoost, also known as Extreme Gradient Boosting Tree, is a type of machine learning algorithm. It belongs to the forward iterative machine learning model and contains multiple trees. Let the number of samples be n. For the t-th tree and the i-th sample, 1≤i≤n, the predicted value of the machine learning model is shown in the following formula (5):

[0035]

[0036] In the above formula, f represents the prediction result for sample i after the t-th iteration. k (x i () represents the prediction result for sample i after the k-th iteration. f represents the prediction result of the (t-1)th tree. t (x i () represents the prediction result of the t-th tree;

[0037] The original objective function is further obtained, as shown in equation (6):

[0038]

[0039] In the above formula, The loss function of a machine learning model. y represents the prediction value of the entire machine learning model for the i-th sample. i Let Ω(f) represent the true value of the i-th sample. j ) represents the complexity of the j-th tree, where is the regularization term in the original objective function;

[0040] The regularization term in equation (6) is split into equation (7):

[0041]

[0042] In the above formula, Obj (t) Let c represent the objective function of the t-th tree;

[0043] 2) Approximation of the second-order expansion of Taylor's formula

[0044]

[0045] In the above formula, g i This corresponds to the first derivative of the loss function, h. i This corresponds to the second derivative of the loss function;

[0046] 3) Tree parameterization

[0047] The complexity of a tree is calculated as follows (9):

[0048]

[0049] In the above formula, γ represents the penalty coefficient for the number of leaf nodes, T represents the number of leaf nodes in the current tree, and λ represents the penalty coefficient for the value of the leaf node. The L2 norm of the leaf node values;

[0050]

[0051]

[0052] In the above formula, G j H represents the sum of the first derivatives of the samples contained in leaf node j. j I represents the sum of the second derivatives of the samples contained in leaf node j. j This represents the set of samples contained in leaf node j;

[0053] Substituting equations (9) to (11) into equation (8) and simplifying, we get:

[0054]

[0055] in,

[0056] At this point, the machine learning model has been established.

[0057] Furthermore, in step seven, the performance of the machine learning model is evaluated using multiple evaluation metrics, including:

[0058]

[0059]

[0060]

[0061]

[0062]

[0063] In the above formula, Acc represents accuracy, CKS represents Cohen's Kappa Score, and F1 score is 0. l F1 score represents the damage level of a disaster of grade l. m This represents the macro average F1 score, F1 w c0 represents the weighted average F1 score, c0 represents the number of samples that correctly predicted the disaster level, n represents the sample size, and p represents the weighted average F1 score. e P represents the probability that the true value and the false value happen to coincide. 1l P represents the accuracy of level l disaster damage. 2l c represents the recall rate for level l disaster damage. 0l q represents the number of samples that correctly predicted level l disaster damage, and q represents the number of sample categories.

[0064] The beneficial effects of this invention are:

[0065] (1) By screening and analyzing indicators of risk, disaster-prone environment, exposure and vulnerability, the most suitable predictive indicators for the machine learning model are established to achieve more comprehensive and accurate prediction.

[0066] (2) Existing technologies have not narrowed the research scale of disaster loss prediction to the county / district level, but rather focus on the provincial level and above. This invention establishes a county / district-level machine learning model by collecting relevant data, enabling event-based forecasting of the direct economic loss level of typhoon disasters. The cross-validation accuracy of the machine learning model of this invention is 76%, and the independent samples test accuracy is 74%. This machine learning model, to a certain extent, fills the gap in county / district-level typhoon disaster loss prediction and has significant reference value.

[0067] (3) This invention explores the application of machine learning models in real-world scenarios. By using measured and forecasted meteorological data to update the hazard indicators input to the machine learning model, it achieves real-time hourly updates and forecasts for each district and county. Attached Figure Description

[0068] Figure 1 This is a flowchart illustrating the process of establishing a machine learning model according to an embodiment of the present invention.

[0069] Figure 2 The correlation coefficient between the hazard, disaster-prone environment, exposure, and vulnerability indicators and the loss level described in the embodiments of the present invention is given.

[0070] Figure 3 This is the confusion matrix of the test set samples described in the embodiments of the present invention.

[0071] Figure 4 This is a schematic diagram illustrating the distribution of direct economic losses caused by Typhoon Lekima in Zhejiang Province in 2019, as described in an embodiment of the present invention.

[0072] Figure 5 This is a prediction map of the level of direct economic loss in Zhejiang Province at four time points generated by the machine learning model described in this embodiment of the invention. Figure 5 (a) A map showing the predicted level of direct economic losses in Zhejiang Province at time 08:08:14. Figure 5 (b) A map showing the predicted level of direct economic losses in Zhejiang Province at time 08:09:12. Figure 5 (c) A map showing the predicted level of direct economic losses in Zhejiang Province at time 08:10:12. Figure 5 (d) A map showing the predicted level of direct economic losses in Zhejiang Province at time 08:11:12.

[0073] Figure 6 This is a schematic diagram illustrating the number of districts and counties predicted with different disaster levels at various times as described in this embodiment of the invention. Figure 6 (a) represents the number of districts and counties predicted to have a disaster level of 1 at various points in time over time. Figure 6(b) represents the number of districts and counties predicted to have a disaster level of 2 at various points in time. Figure 6 (c) represents the number of districts and counties with a disaster level of 3 predicted at each point in time over time. Figure 6 (d) represents the number of districts and counties with a disaster level of 4 predicted at each moment over time. Detailed Implementation

[0074] To facilitate a better understanding of the present invention by those skilled in the art, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The following are merely exemplary and do not limit the scope of protection of the present invention.

[0075] The typhoon disaster risk assessment and dynamic forecasting method described in this embodiment includes the following steps:

[0076] Step 1: Select forecast indicators: Forecast indicators should include at least hazard indicators, vulnerability indicators, and exposure indicators. Hazard indicators should include at least the rainfall brought by strong winds and typhoons, and exposure indicators should include at least the exposure considering the number of disaster-bearing bodies and the disaster-prone environment considering topography and water systems.

[0077] Step 2: Extract risk indicators: During a single typhoon event, extract the maximum rainfall, total rainfall, and average maximum wind speed for several preset time periods within the administrative regions of each district and county.

[0078] Step 3: Correlation Analysis: Perform Pearson correlation analysis between the prediction indicators selected in Step 1 and the loss level. For different types of indicators, select indicators that are greater than the preset correlation threshold as input indicators.

[0079] Step 4: Principal Component Analysis: Principal component analysis is used to reduce the dimensionality of hazard indicators related to rainfall.

[0080] Step 5: Sample set partitioning: Divide the samples into full samples and non-zero samples, and divide the variables into combinations of variables after dimensionality reduction and combinations of the original variables. Keep the data of one typhoon event as the validation set, and the data of the remaining typhoon events as the training set. The full samples include samples with disaster damage data of 0, and the non-zero samples do not include samples with disaster damage data of 0.

[0081] Step 6: Build a machine learning model.

[0082] Step 7: Machine Learning Model Training and Testing: Input the pre-defined training set data into the machine learning model for training. Use multiple evaluation metrics to evaluate the performance of the machine learning model. If all preset metrics reach the allowable error range, the training of the machine learning model is complete. Then, input the pre-defined test set into the trained machine learning model to test the actual generalization performance of the machine learning model.

[0083] Step 8: Obtain forecast results: By continuously updating the measured data and the latest forecast meteorological data, and inputting them into the machine learning model constructed in Step 6, the hazard indicators input into the machine learning model are dynamically updated, thereby realizing real-time updated forecasts of typhoon disaster risks.

[0084] This embodiment takes various districts and counties in Zhejiang Province as the research object, further utilizing more refined and accurate measured meteorological station data, and comprehensively considering hazard, disaster-prone environment, exposure, and vulnerability indicators to select predictive variables. Leveraging the advantages of the XGBoost machine learning algorithm in classification problems—high accuracy and speed—a typhoon disaster risk assessment model is constructed. Based on this, combined with measured and forecasted meteorological data, the hazard indicators input to the model are dynamically updated, achieving real-time updates and forecasts of typhoon disaster risks for all districts and counties in Zhejiang Province. The process of establishing the machine learning model is as follows: Figure 1 As shown.

[0085] Specifically, this embodiment collected county-level disaster data for 10 historical typhoon events that had a significant impact on Zhejiang Province from 2012 to 2019 (namely Haikui, Soulik, Trami, Fitow, Matmo, Chan-hom, Soudelor, Meranti, Maria, and Lekima), meteorological data from all weather stations in Zhejiang Province for 108 hours before and after the typhoons made landfall, and socio-economic data for each district and county in Zhejiang Province for the corresponding years. Detailed data descriptions are shown in Table 1. Simultaneously, to achieve real-time updates on typhoon disaster risk, hourly WRF meteorological forecast data for Typhoon Lekima in Zhejiang Province in 2019 were collected, with an accuracy of 1 km.

[0086] Table 1 Data Explanation

[0087]

[0088] Generally speaking, the risk level of typhoon disasters is a comprehensive reflection of multiple types of disasters. In order to simplify the model, this embodiment selects direct economic loss as the indicator to assess the severity of the disaster, and classifies the loss level as the risk level of typhoon disasters based on the disaster situation in Zhejiang Province over the years, as shown in Table 2.

[0089] Table 2. Typhoon Disaster Severity Classification Standards

[0090]

[0091] Typhoon disaster risk is the result of the combined effects of hazard (hazard indicators are the main disaster-causing factors), disaster-prone environment, exposure, and vulnerability indicators. Based on this and considering data availability, predictive indicators for machine learning models are selected. These mainly include the following four aspects:

[0092] (1) Hazard. Strong winds and rainfall brought by typhoons upon landfall constitute the hazard factors of typhoon disasters. In this embodiment, the typhoon impact process (uniformly selected 36 hours before landfall and 72 hours after landfall) is used as the statistical standard, and the county-level administrative division is used as the statistical unit. The maximum rainfall, process rainfall, and average maximum wind speed at 1 hour, 3 hours, 6 hours, 12 hours, and 24 hours are selected as hazard indicators. The average hourly rainfall and maximum wind speed of each meteorological station within the county are extracted as the representative hourly rainfall and maximum wind speed of that county to extract the hazard indicators for each county.

[0093] (2) Disaster-prone environment. The main factors considered are topography and water system. Topographic factors include elevation and topographic changes. Average elevation, average slope, average aspect and river network density are selected as disaster-prone environment indicators.

[0094] (3) Exposure indicators. In my country's statistics on direct economic losses, direct economic losses mainly consist of agricultural losses, infrastructure losses and household property losses. Based on the availability of data, the sown area of ​​crops, gross domestic product, agricultural output value and total population at the end of the year are selected as exposure indicators.

[0095] (4) Vulnerability Indicators. Based on the characteristics of each district and county in Zhejiang Province and the collection of basic data, the following were selected as vulnerability indicators: the proportion of primary industry, per capita GDP, per capita disposable income of urban residents, per capita disposable income of rural residents, year-end balance of urban and rural residents' savings deposits, mileage of highways within the territory, number of doctors per thousand people, and number of hospital beds per thousand people.

[0096] The initially selected indicators were correlated with the loss level, and categorized according to hazard, disaster-prone environment, exposure level, and vulnerability. The correlation coefficients are as follows: Figure 2 As shown in the table, darker colors indicate a stronger correlation between indicators, while lighter colors indicate a weaker correlation. To avoid indicator redundancy and improve the accuracy of machine learning model predictions, indicators with a high correlation to loss levels were selected as the final input factors. The final selected typhoon disaster risk assessment predictors are shown in Table 3.

[0097] Table 3 Predictive Variables for Typhoon Disaster Risk Assessment

[0098]

[0099] Considering the strong correlation among several hazard indicators related to rainfall, principal component analysis (PCA) was used to reduce the dimensionality of the six rainfall-related hazard indicators. The analysis revealed that the cumulative variance contribution rate of the first principal component reached 94.9% in the full sample and 93.7% in the non-zero sample. Therefore, the first principal component can be considered as a representative indicator related to rainfall. Thus, all predictive variables and the predictive variables after PCA were used as inputs to the machine learning model for comparative analysis, further analyzing the impact of the strong correlation among hazard indicators on the training of the machine learning model. Simultaneously, considering that zero samples (i.e., samples with zero disaster damage data) accounted for more than 50% of the sample, the sample was divided into a full sample (including samples with zero disaster damage data) and a non-zero sample (excluding samples with zero disaster damage data) to explore the impact of zero samples on the training effect of the machine learning model.

[0100] The training set was composed of samples from the previous nine typhoon events, while the 2019 typhoon Lekima was used as the test set. The machine learning model employed the GBTree algorithm, with the objective function set as the softmax function for multi-classification problems. Ten-fold cross-validation was used to train the model; the training set was randomly divided into ten parts, with nine parts used for fitting in each training iteration, and the last part used for validation. Grid-based parameter tuning was used to train the machine learning model, with Acc, CKS, and F1 scores applied. m and F1 w Four evaluation metrics are used to comprehensively evaluate the performance of the machine learning model, thereby determining the optimal hyperparameters and completing the training of the machine learning model. Finally, the machine learning model is cross-validated using Acc, CKS, and F1 scores. m and F1 w The values ​​were 0.76, 0.49, 0.48, and 0.74, respectively, representing the Acc, CKS, and F1 scores for the test set sample test. m and F1 w The confusion matrices for the test set samples are 0.74, 0.62, 0.68, and 0.73, respectively. Figure 3 As shown in the figure, the darker the color of the square, the more samples there are in that part; the lighter the color, the fewer samples there are.

[0101] Using Typhoon Lekima as a forecast case, the forecast time range is from 14:00 on August 8, 2019 to 12:00 on August 11, 2019 (hereinafter referred to as 080814~081112 in this embodiment), a total of 71 time points. Figure 4This is a schematic diagram showing the distribution of direct economic losses caused by Typhoon Lekima to Zhejiang Province in 2019. In a real typhoon event, meteorological stations throughout the province would monitor rainfall, wind speed, and other meteorological elements in real time. Simultaneously, meteorological forecast data would be continuously updated over time. The forecast data in this example is the WRF meteorological forecast data for Typhoon Lekima provided by the Zhejiang Provincial Meteorological Bureau, specifically including forecast fields (grid accuracy of 5 kilometers) for six time periods: 08:08:12, 08:09:00, 08:09:12, 08:10:00, 08:10:12, and 08:11:00. Each forecast field includes hourly forecasts of wind and rain data for the next 72 hours, and the forecast fields are updated every 12 hours. Therefore, at each current moment, by merging the measured meteorological data from the event start time (080814) to the present and the latest forecast meteorological data, and using this to extract hazard indicators as input to the machine learning model, a rolling forecast of the level of direct economic losses caused by Typhoon Lekima in Zhejiang Province can be made, helping decision-makers to intuitively understand the level of losses and risk changes caused by this typhoon.

[0102] The machine learning model generated 71 time-point prediction maps of direct economic loss levels in Zhejiang Province. Four of these time points were selected as representative time points, such as... Figure 5 As shown, specifically as follows Figure 5 (a) Figure 5 (b) Figure 5 (c) Figure 5 As shown in (d), the figure title indicates the current time, where the black solid line represents the historical path of the typhoon, and the black dots represent the location of the typhoon at the current time. Figure 6 This shows the number of districts and counties predicted with different levels of disaster at various points in time, specifically... Figure 6 (a) Figure 6 (b) Figure 6 (c) Figure 6 (d) represents the number of districts and counties at disaster levels 1, 2, 3 and 4 respectively. The solid line represents the number of districts and counties predicted by the machine learning model, and the dashed line represents the actual number of districts and counties.

[0103] from Figure 5(a) The results show that before the typhoon made landfall (Typhoon Lekima made landfall in Wenling City at 2:00 AM on August 10th), the meteorological forecasts issued by the meteorological bureau were too high, leading the machine learning model to overestimate the severity of the disaster in Zhejiang Province. However, as time progressed, the hazard indicators were continuously updated using data collected from actual meteorological stations, making the input hazard indicators increasingly closer to the actual situation. The accuracy rates of the direct economic loss level predictions at the four times shown in the figure (08:08:14, 08:09:12, 08:10:12, and 08:11:12) were 51.7%, 52.8%, 68.5%, and 74.2%, respectively, showing a gradual improvement and eventually stabilizing, ultimately equaling the prediction results when used as independent samples. Therefore, updating the machine learning model's forecasts using data from actual meteorological stations is beneficial for improving the prediction accuracy of the machine learning model in practical applications, and has practical and significant guiding value in disaster prevention decision-making.

[0104] from Figure 6 (a) Figure 6 (b) Figure 6 (c) Figure 6 (d) It can be seen that the number of counties at different levels fluctuates over time, and the curve fluctuates more significantly before the typhoon makes landfall. This is because predicting the typhoon trajectory before landfall is difficult, there is a gap between the forecasted meteorological data and the measured data, and the machine learning model has significant uncertainty in predicting the level of direct economic loss. After the typhoon makes landfall, the hazard indicators (rainfall and maximum wind speed) of the affected counties are basically fixed, and the shape of the curve gradually stabilizes.

[0105] The above description only outlines the basic principles and preferred embodiments of the present invention. Those skilled in the art can make many changes and modifications based on the above description, and these changes and modifications should fall within the protection scope of the present invention.

Claims

1. A typhoon disaster risk assessment and dynamic forecasting method, characterized in that, Including the following steps: Step 1: Select forecast indicators: Forecast indicators should include at least hazard indicators, vulnerability indicators, and exposure indicators. Among them, hazard indicators should include at least the rainfall brought by strong winds and typhoons, and exposure indicators should include at least the exposure considering the number of disaster-bearing bodies and the disaster-prone environment considering topography and water systems. Step 2: Extracting Hazard Indicators: During a single typhoon event, extract the maximum rainfall, total rainfall, and average maximum wind speed for several preset time periods within each district and county administrative region. In step two, let the process of a certain typhoon event take a total of Hours, assuming a certain district or county administrative region includes Each weather station extracts data during a typhoon event using a sliding time window method. The maximum hourly rainfall is calculated using the following formulas (1) to (3): ; ; ; In the above formula, This indicates the administrative region of the district / county. The representative value of rainfall or maximum wind speed at any given moment. Indicates the first Time of the first Rainfall or maximum wind speed recorded by a weather station Indicates the administrative region of the district / county Maximum hourly rainfall, , This indicates the average maximum wind speed over the administrative region of that district / county; Step 3: Correlation Analysis: Perform Pearson correlation analysis between the prediction indicators selected in Step 1 and the loss level. For different types of indicators, select indicators that are greater than the preset correlation threshold as input indicators. The types include hazard, disaster-prone environment, exposure, and vulnerability. Step 4: Principal Component Analysis: Principal component analysis is used to reduce the dimensionality of the risk indicators related to rainfall. After principal component analysis, the variables with a cumulative variance contribution rate of 90% are taken as the principal components after dimensionality reduction, forming different combinations of variables. Step 5: Sample set partitioning: Divide the samples into full samples and non-zero samples, and divide the variables into combinations of variables after dimensionality reduction and combinations of the original variables. Keep the data of one typhoon event as the validation set, and the data of the other typhoon events as the training set. The full samples include samples with disaster damage data of 0, and the non-zero samples do not include samples with disaster damage data of 0. Step 6: Build a machine learning model; Step 7: Machine Learning Model Training and Testing: Input the pre-defined training set data into the machine learning model for training. Use multiple evaluation metrics to assess the model's performance. Training is complete when all preset metrics reach the acceptable error range. Then, input the pre-defined test set into the trained model to test its actual generalization performance. The machine learning model uses the "gbtree" algorithm, with the objective function set as the softmax function for multi-classification problems. Ten-fold cross-validation is used to train the training set, and grid-based parameter tuning is employed. , , and Four evaluation metrics are used to comprehensively evaluate the performance of the machine learning model, thereby determining the optimal hyperparameters and completing the training of the machine learning model. Indicates accuracy rate. This represents Cohen's Kappa Score, which indicates the relative consistency of the model. Represents macro average Fraction, Indicates weighted average Fraction; Step 8: Obtain forecast results: By continuously updating the measured data and the latest forecast meteorological data, and inputting them into the machine learning model constructed in Step 6, the hazard indicators input into the machine learning model are dynamically updated, thereby realizing real-time updated forecasts of typhoon disaster risks.

2. The typhoon disaster risk assessment and dynamic forecasting method according to claim 1, characterized in that, In step one, both vulnerability indicators and exposure indicators can be obtained through yearbooks or publicly available data. Vulnerability indicators are selected based on the characteristics of a region and the availability of data, such as the proportion of the primary industry, GDP per capita, per capita disposable income of urban residents, per capita disposable income of rural residents, year-end balance of urban and rural residents' savings deposits, mileage of highways within the territory, number of doctors per thousand people, and number of hospital beds per thousand people. Exposure indicators should include at least agricultural losses, infrastructure losses, and household property losses. Combined with the availability of data in the region, crop planting area, gross domestic product, agricultural output value, and year-end total population should be selected as exposure indicators. The disaster-prone environment should at least consider topographic factors and water system factors. Topographic factors include both elevation changes and topographic changes. Average elevation, average slope, average aspect, and river network density are selected as disaster-prone environmental indicators.

3. The typhoon disaster risk assessment and dynamic forecasting method according to claim 1, characterized in that, In step three, the formula for Pearson correlation analysis is as follows: ; In the above formula, X and Y represent two random variables. The value of is between -1 and 1, and it is used to measure the correlation between variables X and Y. , Let X and Y represent the means of two random variables, respectively. Let X and Y represent the standard deviations of two random variables, respectively.

4. The typhoon disaster risk assessment and dynamic forecasting method according to claim 1, characterized in that, Step six, which involves building the machine learning model, specifically includes: 1) Construct the objective function The objective function consists of two parts: a loss function and a regularization function. XGBoost, also known as Extreme Gradient Boosting Tree, is a machine learning algorithm, belonging to the forward iterative machine learning model. It contains multiple trees. Assuming the number of samples is n, for the th... Tree, number samples, 1≤ For n ≤ n, the predicted value of the machine learning model is shown in equation (5): ; In the above formula, Indicates the first After the second iteration, the sample The prediction results Indicates the first After the second iteration, the sample The prediction results Indicates the first The predicted results for each tree, Indicates the first The predicted results for each tree; The original objective function is further obtained, as shown in equation (6): ; In the above formula, The loss function of a machine learning model. This indicates that the entire machine learning model is related to the first... The predicted value for each sample, Indicates the first The true value of each sample Indicates the first The complexity of the tree is represented here by the regularization term in the original objective function; The regularization term in equation (6) is split into equation (7): ; In the above formula, Indicates the first The objective function of the trees, Represents a constant; 2) Approximation of the second-order expansion of Taylor's formula ; In the above formula, This corresponds to the first derivative of the loss function. This corresponds to the second derivative of the loss function; 3) Tree parameterization The complexity of a tree is calculated as follows (9): ; In the above formula, This represents the penalty coefficient for the number of leaf nodes. This indicates the number of leaf nodes in the current tree. This represents the penalty coefficient for the leaf node value. Represents the value of the leaf node Norm; ; ; In the above formula, Represents leaf nodes The sum of the first derivatives of the included samples. Represents leaf nodes The sum of the second derivatives of the included samples. Represents leaf nodes The included sample set; Substituting equations (9) to (11) into equation (8) and simplifying, we get: ; in, .

5. The typhoon disaster risk assessment and dynamic forecasting method according to claim 1, characterized in that, Step seven involves evaluating the performance of the machine learning model using various metrics, including: ; ; ; ; ; In the above formula, express Level of disaster damage Fraction, This represents the number of samples where the disaster severity prediction was correct, where n represents the sample size. This represents the probability that the true value and the false value coincide by chance. express Accuracy of disaster level classification express Recall rate of level-based disasters express The number of samples that correctly predicted the level of disaster damage. Indicates the number of sample categories.