A method for predicting extreme events based on tail distribution of data
By constructing a distribution function model and homologous distribution mapping, the applicability and threshold selection of extreme disaster event prediction methods are solved, cross-domain data analysis and prediction are realized, and the accuracy and reliability of extreme event prediction are improved.
Patent Information
- Application Number
- CN202510631413.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The scope of application of existing extreme disaster event prediction methods is limited, and it is difficult to generalize to different disaster types or complex coupled scenarios. The selection of thresholds depends on manual experience, lacks dynamic adaptability, fails to effectively utilize cross-domain data, ignores the impact of dynamic factors on distribution rules, resulting in significant prediction errors.
By collecting situation-dependent and non-situation-dependent event data, building a distribution function model, filtering thresholds and splitting the data sets, judging rationality using goodness of fit coefficients, and predicting extreme events using homologous distribution mapping and data reconstruction.
Cross-domain data analysis has been realized, the objectivity and rationality of threshold selection has been improved, and the effective prediction of future extreme events has been effectively carried out, and the damage to society by extreme disasters has been reduced.
Smart Images

Figure CN120180043B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of disaster prediction, and in particular to a method for predicting heterogeneous extreme events based on tail distribution of data. Background Art
[0002] Existing extreme disaster prediction methods suffer from the following major flaws: First, traditional methods often rely on context-specific influencing factors (such as meteorological parameters and geological activity) to construct prediction models, which limits their applicability and makes them difficult to generalize to different disaster types or complex coupled scenarios. Second, the threshold selection process often relies on manual experience or static standards, lacks dynamic adaptability, and is prone to subjective bias. This is particularly true when analyzing tail data features, making it difficult to objectively identify key thresholds, which hinders the accurate capture of extreme events. Furthermore, existing technologies underutilize data that is scarce or non-context-dependent (such as simulation test data), failing to effectively enhance prediction robustness through cross-domain data fusion. Furthermore, they often fail to thoroughly explore tail data features and often overlook the influence of dynamic factors on distribution patterns, resulting in significant errors in model fitting and prediction of low-probability extreme events. These issues collectively limit the scientific nature and reliability of prediction results, necessitating the urgent need for a more universal, objective, and data-driven methodological system. Summary of the Invention
[0003] In view of the above-mentioned deficiencies in the prior art, the present invention provides a method for predicting heterogeneous extreme events based on tail distribution of data, which provides a more scientific and accurate method for forecasting and estimating various extreme events.
[0004] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:
[0005] A method for predicting extreme events of heterogeneous data tail distribution is provided, which includes the following steps:
[0006] S1: Collect historical data of scenario-dependent events as a scenario-dependent event dataset, clean the data in the scenario-dependent event dataset to generate a stable scenario-dependent event dataset; generate a non-scenario-dependent event dataset based on experimental simulation;
[0007] S2: Construct a distribution function set that fits the data distribution law, use the distribution function model to fit the distribution law of the scenario-dependent event dataset, screen the threshold of the tail data in the scenario-dependent event dataset, and split the tail data into several segmented datasets according to the dynamic factor of the tail data. Use the distribution function model to fit the optimal distribution function model of each segmented dataset as the prior data equation;
[0008] S3: Split the non-context-dependent event dataset into several continuous data sub-datasets, use the distribution function model in the distribution function set to fit the distribution law of the sub-datasets, and judge whether the split sub-datasets are reasonable based on the goodness of fit coefficient, and use the optimal distribution function model as the distribution equation of the sub-datasets;
[0009] S4: Compare the function types between the distribution equation and the prior data equation, and use the distribution equation of the same type with the smallest coefficient difference as the homologous distribution mapping equation for deriving scenario-dependent event data, and take the average value of the coefficient to construct the data reconstruction equation, and use the data reconstruction equation to predict the upcoming data of scenario-dependent events.
[0010] Furthermore, step S2 includes:
[0011] S21: Draw a data histogram using the data in the stable scenario-dependent event data set, select a tail threshold range according to the range of data distribution in the data histogram, and use the data within the tail threshold range as a trial calculation data set;
[0012] S22: constructing a distribution function set that fits the data distribution law, the distribution function set includes several distribution function models, inputting the data in the trial calculation data set into each distribution function model, and fitting the coefficients of each distribution function model;
[0013] S23: screening a candidate threshold set based on the steady-state of the coefficients fitted by each distribution function model within the tail threshold range, and determining the optimal distribution function model within the tail threshold range based on the goodness-of-fit coefficient corresponding to each distribution function model;
[0014] S24: Calculate the difference between the fitted value and the true value based on the optimal distribution function model, take the true value with the smallest difference as the super-threshold in the candidate threshold set, and use the super-threshold to filter the tail data in the scenario-dependent event dataset;
[0015] S25: Using the dynamic factor of the tail data as a trial calculation granularity, segmenting the tail data based on the trial calculation granularity to generate several segmented data sets; determining the prior data equation of each segmented data set based on the frequency of occurrence of the optimal distribution function model in the segmented data sets.
[0016] Furthermore, step S23 includes:
[0017] S231: Determine whether the tail threshold range is a candidate threshold set based on the steady-state of the coefficients fitted by each distribution function model within the tail threshold range; if it cannot be a candidate threshold set, return to step S21 and reselect the tail threshold range; if it can be a candidate threshold set, execute step S232;
[0018] The method for determining the steady-state condition of the coefficient is:
[0019] Calculate the stability coefficient of the coefficients fitted by each distribution function model using data within the tail threshold range ;
[0020] ;
[0021] in, i is the coefficient number, k is the number of coefficients, The first i coefficients, is the coefficient mean;
[0022] Set the threshold value of the stability coefficient ,like , then the fitted coefficients are determined to be stable, and the tail threshold range can be used as a candidate threshold set; otherwise, the fitted coefficients are unstable, and the tail threshold range cannot be used as a candidate threshold set;
[0023] S232: Get the coefficients of each distribution function model fitting Number of k , and use the residual square and mean square error between the true value of the data fitted by each distribution function model and the fitted value to calculate the fitting likelihood value , and then calculate the goodness of fit coefficient ;
[0024] ;
[0025] in, n is the amount of data in the candidate threshold set, AIC is the Akaike information criterion, and BIC is the Bayesian information criterion;
[0026] S233: Get the goodness of fit coefficient corresponding to each distribution function model , m is the type of distribution function model, For the m The goodness of fit coefficient of the distribution function model is selected, and the goodness of fit coefficient is selected. The minimum value in , minimum The corresponding distribution function model is taken as the optimal distribution function model within the tail threshold range.
[0027] Furthermore, step S25 includes:
[0028] S251: extracting the data change cycle pattern of the tail data, generating a dynamic factor of the tail data, and using the dynamic factor as the trial calculation granularity of the tail data. Based on the trial calculation granularity, the tail data is segmented to generate a plurality of segmented data sets.
[0029] S252: Execute steps S22-S23, input each segmented data set into the distribution function model, fit the optimal distribution function model corresponding to each segmented data set, and compare the differences between the optimal distribution function models corresponding to each segmented data set, and calculate the frequency of occurrence of each optimal distribution function model ;
[0030] ;
[0031] in, The first m The number of optimal distribution function models, U is the number of split data sets;
[0032] S253: Setting frequency threshold , if there exists an optimal distribution function model that satisfies , it is determined that the current trial calculation granularity meets the requirements, and the current trial calculation granularity is used as the best fitting granularity of the tail data, and step S255 is executed; otherwise, it is determined that the current trial calculation granularity does not meet the requirements, and step S254 is executed;
[0033] S254: Return to step S251, reselect the trial calculation granularity of the tail data, and execute steps S251-S253 until the best fitting granularity of the tail data is obtained;
[0034] S255: The optimal distribution function model corresponding to each segmented data set obtained by fitting the tail data under the optimal fitting granularity condition is used as a priori data equation.
[0035] Furthermore, step S3 includes:
[0036] S31: Split a sub-dataset containing continuous data from the non-context-dependent event data set, use the distribution function model in the distribution function set to fit the distribution law of the sub-dataset, and calculate the goodness of fit coefficient of each distribution function model to obtain the goodness of fit coefficient data set , For the m The goodness-of-fit coefficient of the distribution function model corresponding to the sub-data set;
[0037] S32: Set the threshold of the goodness-of-fit coefficient , if the goodness-of-fit coefficient data set Existence , then execute step S33; otherwise, return to step S31 and re-split a sub-dataset from the non-context-dependent event dataset;
[0038] S33: Extract the goodness of fit coefficient data set All satisfied The goodness of fit coefficient is formed into a goodness of fit coefficient data set , v Goodness coefficient data set The number of goodness-of-fit coefficients in , Goodness coefficient data set Middle v Goodness-of-fit coefficients;
[0039] S34: Fit goodness of fit coefficient data set The minimum value in As the optimal distribution function model corresponding to the sub-data set, and the optimal distribution function model is used as the distribution equation of the sub-data set;
[0040] S35: Delete the continuous data in the sub-dataset from the non-context-dependent event dataset, return to step S31, continue to split a sub-dataset containing continuous data from the remaining data in the non-context-dependent event dataset, and execute steps S31-S34;
[0041] S36: until the non-context-dependent event data set is split into several sub-data sets, and the distribution equation corresponding to each sub-data set is obtained.
[0042] Furthermore, step S4 includes:
[0043] S41: Compare the distribution equation corresponding to each sub-data set with the prior data equation corresponding to each split data set, compare the coefficient differences between the distribution equations and the prior data equations with the same function type, and use the distribution equation with the smallest coefficient difference as the homologous distribution mapping equation for deriving scenario-dependent event data;
[0044] S42: Calculate the average value of the coefficients between the homologous distribution mapping equation and the corresponding prior data equation, use the average value of the coefficients to replace the coefficients of the distribution equation to form a data reconstruction equation, and use the data reconstruction equation to predict the data of the scenario-dependent event that is about to occur.
[0045] The beneficial effects of the present invention are:
[0046] This method goes beyond traditional disaster prediction methods based on influencing factors. It uses historical context to predict disasters, thus offering broader applicability. By leveraging the experimental nature of non-contextually dependent events (such as coin tossing or the value of pi), this method can be used to analyze extreme events with contextual dependencies (such as earthquakes and typhoons). This enables cross-domain data analysis and prediction, effectively addressing the subjective threshold selection issues in data analysis of context-dependent extreme events.
[0047] This paper proposes a steady-state parameter-guided model-adaptive threshold selection method, which improves the objectivity and rationality of threshold selection. Furthermore, it innovatively employs a data derivation method based on differences in identically distributed coefficients, effectively predicting future extreme event data and evaluating historical data. For small amounts of non-context-dependent data, an adaptive fitting strategy is employed to fully exploit the characteristics of tail data.
[0048] This invention not only provides a scientific basis for disaster prevention and control, but also provides data support for relevant decision-making, which helps to reduce the damage caused by extreme disasters to society. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 Flowchart of the method for predicting heterogeneous extreme events based on the tail distribution of data.
[0050] Figure 2 This is a graph showing the proportion of earthquake occurrences in each segment to the total number of earthquakes.
[0051] Figure 3 This is a graph of the average excess of earthquake magnitude.
[0052] Figure 4 is the shape coefficient distribution diagram.
[0053] Figure 5 This is the distribution diagram of the modified scale coefficient. DETAILED DESCRIPTION
[0054] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.
[0055] like Figure 1 As shown in FIG, a method for predicting heterogeneous extreme events based on tail distribution of data includes the following steps:
[0056] S1: Collect historical data of scenario-dependent events as a scenario-dependent event dataset, clean the data in the scenario-dependent event dataset to generate a stable scenario-dependent event dataset; generate a non-scenario-dependent event dataset based on experimental simulation.
[0057] Scenario-dependent event data are data that are affected by multiple physical factors or vary with multiple factors (such as time, space, and the event location). Examples include earthquake characteristic data (magnitude, focal depth, peak ground acceleration, peak ground velocity, etc.), flood characteristic data (water level, flow, runoff, inundation depth, inundation range, etc.), and dam breach characteristic data (breach width, breach depth, breach flow, etc.). Non-scenario-dependent event data are data that are not affected by physical factors, such as coin tossing data and pi data.
[0058] This embodiment takes earthquake data and coin-tossing data as examples, and uses the distribution law of coin-tossing data to deduce earthquake data;
[0059] Coin tossing data: Based on a programming language, Monte Carlo simulations were used to simulate coin tossing 10 million, 50 million, 100 million, and 1 billion times, and the number of occurrences of consecutive numbers (00+11, 000+111, 0000+1111…) was obtained from them.
[0060] Earthquake data: Magnitude was used as the characteristic feature of earthquake data for analysis. The data spanned from October 11, 1800, to October 1, 2024, with a total of 4,271,800 data items. Data with a magnitude less than 2.5 were removed. Earthquakes with a magnitude less than 2.5 have minimal impact and are not considered when analyzing the tail patterns of extreme events.
[0061] S2: Construct a set of distribution functions that fit the data distribution law, use the distribution function model to fit the distribution law of the scenario-dependent event data set, screen the threshold of the tail data in the scenario-dependent event data set, and split the tail data according to the dynamic factor of the tail data to form several segmented data sets, and use the distribution function model to fit the optimal distribution function model of each segmented data set as the prior data equation.
[0062] Step S2 specifically includes:
[0063] S21: Draw a data histogram using the data in the stable scenario-dependent event data set, select a tail threshold range according to the range of data distribution in the data histogram, and use the data within the tail threshold range as a trial calculation data set;
[0064] S22: constructing a distribution function set that fits the data distribution law, the distribution function set includes several distribution function models, inputting the data in the trial calculation data set into each distribution function model, and fitting the coefficients of each distribution function model;
[0065] S23: screening a candidate threshold set based on the steady-state of the coefficients fitted by each distribution function model within the tail threshold range, and determining the optimal distribution function model within the tail threshold range based on the goodness-of-fit coefficient corresponding to each distribution function model;
[0066] Step S23 specifically includes:
[0067] S231: Determine whether the tail threshold range is a candidate threshold set based on the steady-state of the coefficients fitted by each distribution function model within the tail threshold range; if it cannot be a candidate threshold set, return to step S21 and reselect the tail threshold range; if it can be a candidate threshold set, execute step S232;
[0068] The method for determining the steady-state condition of the coefficient is:
[0069] Calculate the stability coefficient of the coefficients fitted by each distribution function model using data within the tail threshold range ;
[0070] ;
[0071] in, i is the coefficient number, k is the number of coefficients, The first i coefficients, is the average value of the coefficients. When fitting different distribution function models using data within the tail threshold range, due to the large amount of data within the tail threshold range, one distribution function model can fit multiple different coefficients. By calculating the differences between the coefficients, it can be determined whether the distribution function model is stable.
[0072] Set the threshold value of the stability coefficient ,like , then the fitted coefficients are determined to be stable, and the tail threshold range can be used as a candidate threshold set; otherwise, the fitted coefficients are unstable, and the tail threshold range cannot be used as a candidate threshold set;
[0073] S232: Get the coefficients of each distribution function model fitting Number of k , and use the residual square and mean square error between the true value of the data fitted by each distribution function model and the fitted value to calculate the fitting likelihood value , and then calculate the goodness of fit coefficient ;
[0074] ;
[0075] in, n is the amount of data in the candidate threshold set, AIC is the Akaike information criterion, and BIC is the Bayesian information criterion;
[0076] S233: Get the goodness of fit coefficient corresponding to each distribution function model , mis the type of distribution function model, For the m The goodness of fit coefficient of the distribution function model is selected, and the goodness of fit coefficient is selected. The minimum value in , minimum The corresponding distribution function model is taken as the optimal distribution function model within the tail threshold range.
[0077] S24: Calculate the difference between the fitted value and the true value based on the optimal distribution function model, use the true value with the smallest difference as the super-threshold value in the candidate threshold set, and use the super-threshold value to filter the tail data in the scenario-dependent event dataset. In this embodiment, the portion exceeding the super-threshold value is fitted as the tail data.
[0078] S25: Using the dynamic factor of the tail data as a trial calculation granularity, segmenting the tail data based on the trial calculation granularity to generate several segmented data sets; determining the prior data equation of each segmented data set based on the frequency of occurrence of the optimal distribution function model in the segmented data sets.
[0079] Step S25 specifically includes:
[0080] S251: Extract the data change cycle pattern of the tail data and generate the dynamic factor of the tail data. For example, for earthquake data, the magnitude of the earthquake doubles every year over time. In this case, time is the dynamic factor of the earthquake data. The dynamic factor is used as the trial calculation granularity of the tail data. The tail data is segmented based on the trial calculation granularity to generate several segmented data sets.
[0081] S252: Execute steps S22-S23, input each segmented data set into the distribution function model, fit the optimal distribution function model corresponding to each segmented data set, and compare the differences between the optimal distribution function models corresponding to each segmented data set, and calculate the frequency of occurrence of each optimal distribution function model ;
[0082] ;
[0083] in, The first m The number of optimal distribution function models, U is the number of split data sets;
[0084] S253: Setting frequency threshold , if there exists an optimal distribution function model that satisfies , it is determined that the current trial calculation granularity meets the requirements, and the current trial calculation granularity is used as the best fitting granularity of the tail data, and step S255 is executed; otherwise, it is determined that the current trial calculation granularity does not meet the requirements, and step S254 is executed;
[0085] S254: Return to step S251, reselect the trial calculation granularity of the tail data, and execute steps S251-S253 until the best fitting granularity of the tail data is obtained;
[0086] S255: The optimal distribution function model corresponding to each segmented data set obtained by fitting the tail data under the optimal fitting granularity condition is used as a priori data equation.
[0087] This embodiment uses 400,000 earthquake magnitude data as a granularity for data splitting. Based on the magnitude data histogram from March 17, 1800 to June 25, 1985, data with a frequency of around 0.4 is selected as the threshold. At this time, the threshold is 4.68, which is between 4 and 5.5 magnitudes and can be used as a threshold for earthquake data analysis.
[0088] The data simulation results are as follows:
[0089] Generalized Pareto Distribution (GPD): Goodness-of-Fit Coefficient N= 62573.2411 / 62599.7955
[0090] Lognormal Distribution: Goodness-of-Fit Coefficient N= 173310.9042 / 173328.6072
[0091] Cauchy distribution: Goodness-of-fit coefficient N= 105830.9287 / 105848.6317
[0092] Exponential distribution: Goodness-of-fit coefficient N= 63635.1960 / BIC = 63652.8990
[0093] Generalized Extreme Value Distribution (GEV): Goodness-of-Fit Coefficient N= 75518.1034 / 75544.6579
[0094] Power Law Distribution: Goodness-of-Fit Coefficient N= 69469.2683 / 69486.9713
[0095] Gumbel distribution: Goodness-of-fit coefficient N= 79309.7210 / 79327.4240
[0096] Weibull distribution: Goodness-of-fit coefficient N= 63728.1792 / 63754.7337
[0097] Fréchet distribution: Goodness-of-fit coefficient N=75518.1034 / 75544.6579
[0098] According to the simulation results of the above data, the optimal distribution function model is the generalized Pareto distribution (GPD), with a shape coefficient = -0.1520418596254437, a displacement coefficient = 0.010000097378072906, and a scale coefficient = 0.7852156395133192.
[0099] The magnitude fitting of the year split with a granularity of 400,000 earthquakes can be obtained, as shown in Table 1 below;
[0100] Table 1 Magnitude fitting with 400,000 magnitude as one granularity
[0101]
[0102] Note:
[0103] Time period granularity selection: 400,000 earthquake magnitude data are used as a time period unit for granularity separation, except for 2022.6-2024.11.1, (253884) and 2024.1.1-2024.10.1.
[0104] Processing of the number of earthquakes in a time period: After dividing the time periods, earthquakes with a magnitude greater than 2.5 are screened out, and earthquakes with a magnitude below 2.5 are considered mild.
[0105] Trial threshold: Select a threshold whose frequency is less than 0.4 and whose magnitude frequency does not change significantly after it exceeds the selected threshold.
[0106] Optimal distribution for fitting the excess amount: Since this study only focuses on extreme events and their tail characteristics, only the part that exceeds the trial calculation threshold is fitted. The fitting distributions are as follows: GPD fitting, lognormal fitting, Cauchy fitting, exponential distribution fitting, GEV fitting, power law distribution fitting, Weibull fitting, Gumbel fitting, and Fréchet fitting. These types of distributions are all distributions that express the tail characteristics of the data.
[0107] Heavy tails: Determine whether it is heavy tails based on the shape parameters of the optimal distribution being fitted.
[0108] Data segmentation is performed with 5-year magnitude data as the granularity: Since the five-year magnitude data is less than that in the previous section, the magnitude data with a frequency below 0.2 is taken as the threshold, and a total of 25 years of data from 1999 to 2024 are fitted with a granularity of five years. The fitting results are shown in Table 2 below.
[0109] Table 2 Magnitude fitting results with a granularity of five years
[0110]
[0111] The data is split using 10-year magnitude data as a granularity: fitting the data for a total of 40 years from 1984 to 2024, with 10 years as a granularity. The fitting results are shown in Table 3 below.
[0112] Table 3 Magnitude fitting results with a granularity of ten years
[0113]
[0114] The data is split using 20 years of magnitude data as a granularity: fitting the data from 1964 to 2024, a total of 60 years, with ten years as a granularity. The fitting results are shown in Table 4 below.
[0115] Table 4 Magnitude fitting results with a granularity of 20 years
[0116]
[0117] Data segmentation: According to the above fitting results, the threshold selection method with the number of earthquakes in a period as the standard and the magnitude data of every five years as one granularity does not show great fluctuations in magnitude. The threshold selection method with the heavy tail situation as the standard and the magnitude data of every five years as one granularity all shows heavy tail phenomenon. The threshold value fluctuates between 4.85 and 5.15 with little change when the trial threshold value is used as the standard. Therefore, the threshold value is subsequently selected with the magnitude data of five years as one granularity, and its fitting equation is obtained and deduced with the coin-tossing data.
[0118] S3: Split the non-scenario-dependent event dataset into several sub-datasets of continuous data, use the distribution function model in the distribution function set to fit the distribution law of the sub-dataset, and judge whether the split sub-datasets are reasonable based on the goodness of fit coefficient, and use the optimal distribution function model as the distribution equation of the sub-dataset.
[0119] Step S3 specifically includes:
[0120] S31: Split a sub-dataset containing continuous data from the non-context-dependent event data set, use the distribution function model in the distribution function set to fit the distribution law of the sub-dataset, and calculate the goodness of fit coefficient of each distribution function model to obtain the goodness of fit coefficient data set , For the m The goodness-of-fit coefficient of the distribution function model corresponding to the sub-data set;
[0121] S32: Set the threshold of the goodness-of-fit coefficient , if the goodness-of-fit coefficient data set Existence , then execute step S33; otherwise, return to step S31 and re-split a sub-dataset from the non-context-dependent event dataset;
[0122] S33: Extract the goodness of fit coefficient data set All satisfied The goodness of fit coefficient is formed into a goodness of fit coefficient data set , v Goodness coefficient data set The number of goodness-of-fit coefficients in , Goodness coefficient data set Middle v Goodness-of-fit coefficients;
[0123] S34: Fit goodness of fit coefficient data set The minimum value in As the optimal distribution function model corresponding to the sub-data set, and the optimal distribution function model is used as the distribution equation of the sub-data set;
[0124] S35: Delete the continuous data in the sub-dataset from the non-context-dependent event dataset, return to step S31, continue to split a sub-dataset containing continuous data from the remaining data in the non-context-dependent event dataset, and execute steps S31-S34;
[0125] S36: until the non-context-dependent event data set is split into several sub-data sets, and the distribution equation corresponding to each sub-data set is obtained.
[0126] S4: Compare the function types between the distribution equation and the prior data equation, and use the distribution equation of the same type with the smallest coefficient difference as the homologous distribution mapping equation for deriving scenario-dependent event data, and take the average value of the coefficient to construct the data reconstruction equation, and use the data reconstruction equation to predict the upcoming data of scenario-dependent events.
[0127] Step S4 specifically includes:
[0128] S41: Compare the distribution equation corresponding to each sub-data set with the prior data equation corresponding to each split data set, compare the coefficient differences between the distribution equations and the prior data equations with the same function type, and use the distribution equation with the smallest coefficient difference as the homologous distribution mapping equation for deriving scenario-dependent event data;
[0129] S42: Calculate the average value of the coefficients between the homologous distribution mapping equation and the corresponding prior data equation, use the average value of the coefficients to replace the coefficients of the distribution equation to form a data reconstruction equation, and use the data reconstruction equation to predict the data of the scenario-dependent event that is about to occur.
[0130] This example takes 2019-2024 as an example: if earthquake data is to be derived from coin tossing data, the tail characteristics of the coin tossing are generalized Pareto distribution, and according to the granularity trial, the AIC and BIC of the generalized Pareto are low. It is assumed that the part of the earthquake data exceeding the threshold conforms to the generalized Pareto distribution. Figure 4-Figure 5 As shown, according to the average excess of magnitude under different thresholds, the shape coefficient diagram of generalized Pareto distribution and the modified scale coefficient diagram, and according to Figure 3 , the threshold range is (5.1-5.5). Based on the fit between the super-threshold data and the distribution, the optimal threshold is 5.2. Refer to Table 5 below, which summarizes the fitting results of the generalized Pareto distribution at different thresholds.
[0131] Table 5 Generalized Pareto distribution fitting of super-threshold data at different thresholds
[0132]
[0133] According to Table 5, through the goodness of fit of the data under different thresholds, it is found that 5.2 is the optimal threshold. According to the coefficient graphs, it is reasonable to use 5.2 as the threshold, so 5.2 is used as the threshold from 2019 to 2024.
[0134] Earthquake data fitting: According to this method, the equation fitted by the earthquake data of two five-year periods, 2009-2014 and 2014-2019, and the coin-tossing data is the generalized Pareto equation, and its predicted historical and future situations are shown in Table 6 below.
[0135] Table 6 Coin tossing and earthquake deduction process
[0136]
[0137] Table 7 below compares the data obtained through the generalized Pareto equation with the data for the past five years and the next five years:
[0138] Table 7 Comparison of prediction results with earthquake data from the past five years and the next five years
[0139]
[0140] Figure 2 The data reconstruction equation obtained by this method is used to compare the earthquake occurrence ratio within the interval predicted by the data reconstruction equation with the earthquake occurrence ratio within the interval every five years and every ten years from 2004 to 2024. Therefore, the data reconstruction equation obtained by this method is quite representative of the data.
[0141] This method goes beyond traditional disaster prediction methods based on influencing factors. It uses historical context to predict disasters, thus offering broader applicability. By leveraging the experimental nature of non-contextually dependent events (such as coin tossing or the value of pi), this method can be used to analyze extreme events with contextual dependencies (such as earthquakes and typhoons). This enables cross-domain data analysis and prediction, effectively addressing the subjective threshold selection issues in data analysis of context-dependent extreme events.
[0142] This paper proposes a steady-state parameter-guided model-adaptive threshold selection method, which improves the objectivity and rationality of threshold selection. Furthermore, it innovatively employs a data derivation method based on differences in identically distributed coefficients, effectively predicting future extreme event data and evaluating historical data. For small amounts of non-context-dependent data, an adaptive fitting strategy is employed to fully exploit the characteristics of tail data.
[0143] This invention not only provides a scientific basis for disaster prevention and control, but also provides data support for relevant decision-making, which helps to reduce the damage caused by extreme disasters to society.
Claims
1. A method for predicting heterogeneous extreme events based on tail distribution of data, characterized in that: The following steps are involved: S1: Collect historical data on scenario-dependent events as a scenario-dependent event dataset. The scenario-dependent event data includes earthquake characteristic data, flood characteristic data, or dam break characteristic data. Clean the data in the scenario-dependent event dataset to generate a stable scenario-dependent event dataset. Generate a non-scenario-dependent event dataset based on experimental simulation. S2: Construct a distribution function set that fits the data distribution law, use the distribution function model to fit the distribution law of the scenario-dependent event dataset, screen the threshold of the tail data in the scenario-dependent event dataset, and split the tail data into several segmented datasets according to the dynamic factor of the tail data. Use the distribution function model to fit the optimal distribution function model of each segmented dataset as the prior data equation; S3: Split the non-context-dependent event dataset into several continuous data sub-datasets, use the distribution function model in the distribution function set to fit the distribution law of the sub-datasets, and judge whether the split sub-datasets are reasonable based on the goodness of fit coefficient, and use the optimal distribution function model as the distribution equation of the sub-datasets; S4: Compare the function types between the distribution equation and the prior data equation, and use the distribution equation of the same type with the smallest coefficient difference as the homologous distribution mapping equation for deriving scenario-dependent event data, and take the average value of the coefficient to construct the data reconstruction equation, and use the data reconstruction equation to predict the upcoming data of scenario-dependent events.
2. The method for predicting heterogeneous extreme events based on tail distribution of data according to claim 1, characterized in that: The step S2 comprises: S21: Draw a data histogram using the data in the stable scenario-dependent event data set, select a tail threshold range according to the range of data distribution in the data histogram, and use the data within the tail threshold range as a trial calculation data set; S22: constructing a distribution function set that fits the data distribution law, the distribution function set includes several distribution function models, inputting the data in the trial calculation data set into each distribution function model, and fitting the coefficients of each distribution function model; S23: screening a candidate threshold set based on the steady-state of the coefficients fitted by each distribution function model within the tail threshold range, and determining the optimal distribution function model within the tail threshold range based on the goodness-of-fit coefficient corresponding to each distribution function model; S24: Calculate the difference between the fitted value and the true value based on the optimal distribution function model, take the true value with the smallest difference as the super-threshold in the candidate threshold set, and use the super-threshold to filter the tail data in the scenario-dependent event dataset; S25: Using the dynamic factor of the tail data as a trial calculation granularity, segmenting the tail data based on the trial calculation granularity to generate several segmented data sets; determining the prior data equation of each segmented data set based on the frequency of occurrence of the optimal distribution function model in the segmented data sets.
3. The method for predicting heterogeneous extreme events based on tail distribution of data according to claim 2, characterized in that: The step S23 includes: S231: Determine whether the tail threshold range is a candidate threshold set based on the steady-state of the coefficients fitted by each distribution function model within the tail threshold range; if it cannot be a candidate threshold set, return to step S21 and reselect the tail threshold range; if it can be a candidate threshold set, execute step S232; The method for determining the steady-state condition of the coefficient is: Calculate the stability coefficient of the coefficients fitted by each distribution function model using data within the tail threshold range ; ; in, i is the coefficient number, k is the number of coefficients, The first i coefficients, is the coefficient mean; Set the threshold value of the stability coefficient ,like , then the fitted coefficients are determined to be stable, and the tail threshold range can be used as a candidate threshold set; otherwise, the fitted coefficients are unstable, and the tail threshold range cannot be used as a candidate threshold set; S232: Get the coefficients of each distribution function model fitting Number of k , and use the residual square and mean square error between the true value of the data fitted by each distribution function model and the fitted value to calculate the fitting likelihood value , and then calculate the goodness of fit coefficient ; ; in, n is the amount of data in the candidate threshold set, AIC is the Akaike information criterion, and BIC is the Bayesian information criterion; S233: Get the goodness of fit coefficient corresponding to each distribution function model , m is the type of distribution function model, For the m The goodness of fit coefficient of the distribution function model is selected, and the goodness of fit coefficient is selected. The minimum value in , minimum The corresponding distribution function model is taken as the optimal distribution function model within the tail threshold range.
4. The method for predicting heterogeneous extreme events based on tail distribution of data according to claim 3, characterized in that: The step S25 includes: S251: extracting the data change cycle pattern of the tail data, generating a dynamic factor of the tail data, and using the dynamic factor as the trial calculation granularity of the tail data. Based on the trial calculation granularity, the tail data is segmented to generate a plurality of segmented data sets. S252: Execute steps S22-S23, input each segmented data set into the distribution function model, fit the optimal distribution function model corresponding to each segmented data set, and compare the differences between the optimal distribution function models corresponding to each segmented data set, and calculate the frequency of occurrence of each optimal distribution function model ; ; in, The first m The number of optimal distribution function models, U is the number of split data sets; S253: Setting frequency threshold , if there exists an optimal distribution function model that satisfies , it is determined that the current trial calculation granularity meets the requirements, and the current trial calculation granularity is used as the best fitting granularity of the tail data, and step S255 is executed; otherwise, it is determined that the current trial calculation granularity does not meet the requirements, and step S254 is executed; S254: Return to step S251, reselect the trial calculation granularity of the tail data, and execute steps S251-S253 until the best fitting granularity of the tail data is obtained; S255: The optimal distribution function model corresponding to each segmented data set obtained by fitting the tail data under the optimal fitting granularity condition is used as a priori data equation.
5. The method for predicting heterogeneous extreme events based on tail distribution of data according to claim 4, characterized in that: The step S3 comprises: S31: Split a sub-dataset containing continuous data from the non-context-dependent event data set, use the distribution function model in the distribution function set to fit the distribution law of the sub-dataset, and calculate the goodness of fit coefficient of each distribution function model to obtain the goodness of fit coefficient data set , For the m The goodness-of-fit coefficient of the distribution function model corresponding to the sub-data set; S32: Set the threshold of the goodness of fit coefficient , if the goodness-of-fit coefficient data set Existence , then execute step S33; otherwise, return to step S31 and re-split a sub-dataset from the non-context-dependent event dataset; S33: Extract the goodness of fit coefficient data set All satisfied The goodness of fit coefficient is formed into a goodness of fit coefficient data set , v Goodness coefficient data set The number of goodness-of-fit coefficients in , Goodness coefficient data set Middle v Goodness-of-fit coefficients; S34: Fit goodness of fit coefficient data set The minimum value in As the optimal distribution function model corresponding to the sub-data set, and the optimal distribution function model is used as the distribution equation of the sub-data set; S35: Delete the continuous data in the sub-dataset from the non-context-dependent event dataset, return to step S31, continue to split a sub-dataset containing continuous data from the remaining data in the non-context-dependent event dataset, and execute steps S31-S34; S36: until the non-context-dependent event data set is split into several sub-data sets, and the distribution equation corresponding to each sub-data set is obtained.
6. The method for predicting heterogeneous extreme events based on tail distribution of data according to claim 5, characterized in that: The step S4 comprises: S41: Compare the distribution equation corresponding to each sub-data set with the prior data equation corresponding to each split data set, compare the coefficient differences between the distribution equations and the prior data equations with the same function type, and use the distribution equation with the smallest coefficient difference as the homologous distribution mapping equation for deriving scenario-dependent event data; S42: Calculate the average value of the coefficients between the homologous distribution mapping equation and the corresponding prior data equation, use the average value of the coefficients to replace the coefficients of the distribution equation to form a data reconstruction equation, and use the data reconstruction equation to predict the data of the scenario-dependent event that is about to occur.
Citation Information
Patent Citations
Coastal tide level extreme value prediction method based on extreme value mixed distribution
CN116227190A
Data shortage estuary composite flood disaster research method considering climatic change
CN118760899A