Heterogeneous extreme event prediction method for data tail distribution
By constructing a distribution function set to fit the distribution rules of extreme disaster event data, screening and splitting tail data, and constructing data reconstruction equations, the problem of limited application scope and subjectivity of threshold selection of extreme disaster event prediction methods in the existing technology is solved, and more scientific and reliable extreme event prediction is achieved.
Patent Information
- Application Number
- CN202510631413.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The scope of application of existing extreme disaster event prediction methods is limited, and it is difficult to generalize to different disaster types or complex coupled scenarios. The threshold selection process is subjective, and it is difficult to objectively identify key thresholds, which affects the accurate capture of extreme events.
By collecting historical data, the distribution function set is constructed to fit the data distribution rules, filter the tail data threshold, and split the tail data according to dynamic factors to form a segmented data set. The distribution function model is used to fit the optimal distribution function model of each segmented data set as a prior data equation, and then the data reconstruction equation is constructed to predict extreme events.
It improves the scientificity and reliability of extreme event prediction, enhances applicability and objectivity, effectively solves the subjectivity problem of threshold selection, and realizes more accurate capture and prediction of extreme event data.
Smart Images

Figure CN120180043A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of disaster prediction, and particularly to a method for predicting heterogeneous extreme events based on the data tail distribution. Background Art
[0002] The existing prediction methods for extreme disaster events mainly have the following defects: First, traditional methods mostly rely on influencing factors (such as meteorological parameters, geological activities, etc.) in specific scenarios to construct prediction models, resulting in limited applicability and difficulty in generalizing to different disaster types or complex coupling scenarios. Second, the threshold selection process often relies on manual experience or static standards, lacking dynamic self - adaptability and being prone to introducing subjective biases. Especially in the analysis of tail data characteristics, it is difficult to objectively identify key thresholds, affecting the accurate capture of extreme events. In addition, the existing technology makes insufficient use of scarce data or non - context - dependent data (such as simulation test data), and fails to effectively improve prediction robustness through cross - domain data fusion; at the same time, the mining of data tail characteristics is not deep enough, often ignoring the influence of dynamic factors on the distribution law, resulting in significant errors in the fitting and prediction of low - probability extreme events. These problems jointly restrict the scientific nature and reliability of prediction results, and there is an urgent need for a more general, objective and data - driven method system. Summary of the Invention
[0003] In view of the above - mentioned deficiencies of the prior art, the present invention provides a method for predicting heterogeneous extreme events based on the data tail distribution, providing a more scientific and accurate method for the prediction and estimation of various extreme events.
[0004] To achieve the above - mentioned invention purpose, the technical solution adopted by the present invention is as follows: Provide a method for predicting heterogeneous extreme events based on the data tail distribution, which includes the following steps: S1: Collect historical data of context - dependent events as a context - dependent event data set, and clean the data in the context - dependent event data set to generate a stable context - dependent event data set; generate a non - context - dependent event data set according to experimental simulations; S2: Construct a set of distribution functions for fitting the data distribution law, use the distribution function model to fit the distribution law of the context - dependent event data set, screen the thresholds of the tail data in the context - dependent event data set, and split the tail data according to the dynamic factors of the tail data to form several segmented data sets, and use the distribution function model to fit the optimal distribution function model of each segmented data set as the prior data equation; S3: Split the non - context - dependent event data set into several sub - data sets of continuous data, use the distribution function model in the set of distribution functions to fit the distribution law of the sub - data sets, and judge whether the split sub - data sets are reasonable according to the goodness - of - fit coefficient, and use the optimal distribution function model as the distribution equation of the sub - data sets; S4: Compare the types of functions between the distribution equation and the prior data equation. Take the distribution equation with the same type and the smallest coefficient difference as the homologous distribution mapping equation for deriving scenario-dependent event data. Take the average value of the coefficients to construct a data reconstruction equation, and use the data reconstruction equation to predict the data of the upcoming scenario-dependent event.
[0005] Further, step S2 includes: S21: Use the data in the stable scenario-dependent event dataset to draw a data histogram. Select the tail threshold range according to the data distribution range in the data histogram, and use the data within the tail threshold range as the trial calculation dataset. S22: Construct a set of distribution functions that fit the data distribution law. The set of distribution functions contains several distribution function models. Input the data in the trial calculation dataset into each distribution function model to fit the coefficients of each distribution function model. S23: According to the coefficient steady-state situation of each distribution function model fitted within the tail threshold range, screen the candidate threshold set, and determine the optimal distribution function model within the tail threshold range based on the goodness-of-fit coefficient corresponding to each distribution function model. S24: Calculate the difference between the fitted value and the true value based on the optimal distribution function model. Take the true value with the smallest difference as the super-threshold in the candidate threshold set, and use the super-threshold to screen the tail data in the scenario-dependent event dataset. S25: Take the dynamic factor of the tail data as the trial calculation granularity, and segment the tail data based on the trial calculation granularity to generate several segmented datasets; determine the prior data equation of each segmented dataset according to the frequency of the optimal distribution function model in the segmented datasets.
[0006] Further, step S23 includes: S231: According to the coefficient steady-state situation of each distribution function model fitted within the tail threshold range, determine whether the tail threshold range can be used as the candidate threshold set; if it cannot be used as the candidate threshold set, return to step S21 to re-select the tail threshold range; if it can be used as the candidate threshold set, execute step S232. The determination method of the coefficient steady-state situation is: Calculate the stability coefficient of the coefficients fitted by each distribution function model using the data within the tail threshold range ; ; where i is the coefficient number, k is the number of coefficients, is the i th coefficient fitted, is the coefficient average value; Set the threshold of the stability coefficient If , it is determined that the fitted coefficient is stable, and the tail threshold range can be used as the candidate threshold set; otherwise, the fitted coefficient is unstable, and the tail threshold range cannot be used as the candidate threshold set; S232: Obtain the coefficients fitted by each distribution function model The number of k , and calculate the fitting likelihood value using the mean square error of the sum of squared residuals between the true value and the fitted value of the data fitted by each distribution function model , and then calculate the coefficient of determination ; ; Among them, n is the data volume in the candidate threshold set, AIC is the Akaike information criterion, and BIC is the Bayesian information criterion; S233: Obtain the coefficient of determination corresponding to each distribution function model , m is the type of the distribution function model, is the coefficient of determination of the m th distribution function model, and screen out the minimum value in the coefficient of determination , and the distribution function model corresponding to the minimum value is used as the optimal distribution function model within the tail threshold range.
[0007] Furthermore, step S25 includes: S251: Extract the data change cycle law of the tail data, generate the dynamic factor of the tail data, and use the dynamic factor as the trial calculation granularity of the tail data. Based on the trial calculation granularity, divide the tail data to generate several divided data sets; S252: Execute steps S22 - S23, input each divided data set into the distribution function model, fit the optimal distribution function model corresponding to each divided data set, and compare the differences between the optimal distribution function models corresponding to each divided data set, and calculate the frequency of occurrence of each optimal distribution function model ; ; Among them, is the number of the m th optimal distribution function model fitted by several divided data sets, U is the number of divided data sets; S253: Set the frequency threshold , if there is an optimal distribution function model that satisfies , it is determined that the current trial calculation granularity meets the requirements. The current trial calculation granularity is used as the best fitting granularity of the tail data, and step S255 is executed; otherwise, it is determined that the current trial calculation granularity does not meet the requirements, and step S254 is executed; S254: Return to step S251, reselect the trial calculation granularity of the tail data, and execute steps S251 - S253; until the best fitting granularity of the tail data is obtained; S255: Use the optimal distribution function model corresponding to each split data set obtained by fitting the tail data under the best fitting granularity as the prior data equation.
[0008] Further, step S3 includes: S31: Split a sub - data set containing continuous data from the non - scenario - dependent event data set, fit the distribution law of the sub - data set using the distribution function models in the distribution function set, and calculate the goodness - of - fit coefficient of each distribution function model to obtain the goodness - of - fit coefficient data set , is the goodness - of - fit coefficient corresponding to the sub - data set fitted by the m th distribution function model; S32: Set the threshold of the goodness - of - fit coefficient . If there is in the goodness - of - fit coefficient data set , then step S33 is executed; otherwise, return to step S31 and re - split a sub - data set from the non - scenario - dependent event data set; S33: Extract all the goodness - of - fit coefficients in the goodness - of - fit coefficient data set that satisfy to form the goodness - of - fit coefficient data set , v is the number of goodness - of - fit coefficients in the goodness - of - fit coefficient data set , is the th goodness - of - fit coefficient in the goodness - of - fit coefficient data set v ; S34: Use the minimum value in the goodness - of - fit coefficient data set as the optimal distribution function model of the sub - data set, and use the optimal distribution function model as the distribution equation of the sub - data set; S35: Delete the continuous data in the sub - data set from the non - scenario - dependent event data set, return to step S31, continue to split a sub - data set containing continuous data from the remaining data in the non - scenario - dependent event data set, and execute steps S31 - S34; S36: Until the non - scenario - dependent event data set is split into several sub - data sets and the distribution equations corresponding to each sub - data set are obtained.
[0009] Further, step S4 includes: S41: Compare the distribution equation corresponding to each sub-dataset with the prior data equation corresponding to each segmented dataset, compare the coefficient differences between the distribution equations and the prior data equations with the same function type, and use the distribution equation with the smallest coefficient difference as the homologous distribution mapping equation for deriving scenario-dependent event data; S42: Calculate the average value of the coefficients between the homologous distribution mapping equation and the corresponding prior data equation, use the average value of the coefficients to replace the coefficients of the distribution equation to form a data reconstruction equation, and use the data reconstruction equation to predict the data of the upcoming scenario-dependent event.
[0010] The beneficial effects of the present invention are as follows: The present invention is not limited to the traditional method of predicting disaster events by the influencing factors of disaster events. Starting from the historical occurrence of disaster events for prediction, it has a wider range of applicability. By leveraging the testability of the occurrence of non-context-dependent events (such as coin-tossing, pi) to analyze extreme events with context-dependence (such as earthquakes, typhoons), cross-domain data analysis and prediction are achieved, effectively solving the subjectivity problem of threshold selection in the data analysis of context-dependent extreme events.
[0011] The present invention proposes a steady-state parameter-guided model adaptive threshold selection method, which improves the objectivity and rationality of threshold selection. At the same time, it creatively adopts a data derivation method based on the coefficient difference of the same distribution, which can effectively predict future extreme event data and evaluate historical data. For non-context-dependent data with a small amount of data, an adaptive fitting strategy is adopted to fully explore the tail data characteristics.
[0012] The present invention not only provides a scientific basis for disaster prevention and control, but also provides data support for relevant decisions, helping to reduce the damage caused by extreme disasters to society. Description of the Drawings
[0013] Figure 1 It is a flowchart of a prediction method for heterogeneous extreme events of the data tail distribution.
[0014] Figure 2 It is a diagram showing the proportion of the number of earthquakes occurring in each stage to the total number of earthquakes.
[0015] Figure 3 It is a diagram of the average excess of the magnitude.
[0016] Figure 4 It is a diagram of the shape coefficient distribution.
[0017] Figure 5 It is a diagram of the modified scale coefficient distribution. Detailed Embodiments
[0018] The specific embodiments of the present invention will be described below to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
[0019] As Figure 1 shown, a method for predicting heterogeneous extreme events in the data tail distribution includes the following steps: S1: Collect historical data of scenario-dependent events as a scenario-dependent event dataset, and clean the data in the scenario-dependent event dataset to generate a stable scenario-dependent event dataset; generate a non-scenario-dependent event dataset according to experimental simulations.
[0020] Scenario-dependent event data is data affected by multiple physical factors or varying with multiple factors (such as time, space, the situation of the event occurrence body, etc.), such as earthquake characteristic data (magnitude, focal depth, peak ground acceleration, peak ground velocity, etc.), flood characteristic data (water level, flow rate, runoff, inundation depth, inundation area, etc.), dam-break characteristic data (breach width, breach depth, breach discharge, etc.), etc. Non-scenario-dependent event data is data not affected by physical factors, such as coin-tossing data, pi data, etc.
[0021] In this embodiment, taking earthquake data and coin-tossing data as examples, the distribution law of coin-tossing data is used to deduce earthquake data; Coin-tossing data: Based on a programming language, 10 million times, 50 million times, 100 million times, and 1 billion times of coin-tossing data are simulated through Monte Carlo, and the occurrence times of consecutive numbers (00 + 11, 000 + 111, 0000 + 1111...) are obtained by traversing them respectively.
[0022] Earthquake data: Taking the magnitude as the earthquake data characteristic for analysis, the data time span is from October 11, 1800 to October 1, 2024, with a total of 4,271,800 pieces of data. The data with a magnitude less than 2.5 is removed. The magnitude less than 2.5 has extremely little impact and is not considered when analyzing the tail law of extreme events.
[0023] S2: Construct a set of distribution functions that fit the data distribution law, use the distribution function model to fit the distribution law of the scenario-dependent event dataset, screen the threshold of the tail data in the scenario-dependent event dataset, and split the tail data according to the dynamic factor of the tail data to form several segmentation datasets, and use the distribution function model to fit the optimal distribution function model of each segmentation dataset as the prior data equation.
[0024] Step S2 specifically includes: S21: Draw a data histogram using the data in the stable scenario-dependent event dataset, select the tail threshold range according to the range of data distribution in the data histogram, and use the data within the tail threshold range as the trial calculation dataset; S22: Construct a set of distribution functions that fit the data distribution law. The set of distribution functions contains several distribution function models. Input the data in the trial calculation dataset into each distribution function model to fit the coefficients of each distribution function model; S23: According to the coefficient steady-state situation of each distribution function model fitted within the tail threshold range, screen the candidate threshold set, and based on the goodness-of-fit coefficient corresponding to each distribution function model, determine the optimal distribution function model within the tail threshold range; Step S23 specifically includes: S231: According to the coefficient steady-state situation of each distribution function model fitted within the tail threshold range, determine whether the tail threshold range can be used as the candidate threshold set; if it cannot be used as the candidate threshold set, return to step S21 to re-select the tail threshold range; if it can be used as the candidate threshold set, execute step S232; The determination method of the coefficient steady-state situation is: Calculate the stability coefficient of the coefficients fitted by each distribution function model using the data within the tail threshold range ; ; Among them, i is the coefficient number, k is the number of coefficients, is the i th coefficient fitted, is the coefficient average value; when fitting different distribution function models using the data within the tail threshold range, due to the large amount of data within the tail threshold range, a distribution function model can fit multiple different coefficients. By calculating the difference between the coefficients, it can be determined whether the distribution function model is stable.
[0025] Set the threshold of the stability coefficient , if , it is determined that the fitted coefficients are stable, and the tail threshold range can be used as the candidate threshold set; otherwise, the fitted coefficients are unstable, and the tail threshold range cannot be used as the candidate threshold set; S232: Obtain the number of the coefficients k fitted by each distribution function model, and calculate the fitting likelihood value using the sum of squared residuals mean square error between the true value and the fitted value of the data fitted by each distribution function model, and then calculate the goodness-of-fit coefficient ; ; Among them, n is the data volume in the candidate threshold set, AIC is the Akaike information criterion, and BIC is the Bayesian information criterion; S233: Obtain the goodness-of-fit coefficient corresponding to each distribution function model , m is the type of the distribution function model, is the m th goodness-of-fit coefficient of the distribution function model, and screen out the minimum value in the goodness-of-fit coefficient The minimum value corresponding distribution function model is used as the optimal distribution function model within the tail threshold range.
[0026] S24: Calculate the difference between the fitted value and the true value fitted based on the optimal distribution function model, and use the true value with the smallest difference as the over-threshold value in the candidate threshold set, and use the over-threshold value to screen the tail data in the scenario-dependent event dataset; in this embodiment, the part exceeding the over-threshold value is fitted as the tail data; S25: Use the dynamic factor of the tail data as the trial calculation granularity, and based on the trial calculation granularity, divide the tail data to generate several divided datasets; according to the frequency of occurrence of the optimal distribution function model in the divided datasets, determine the prior data equation of each divided dataset.
[0027] Step S25 specifically includes: S251: Extract the data change cycle law of the tail data to generate the dynamic factor of the tail data. For example, for seismic data, the magnitude doubles annually over time. In this case, the time is the dynamic factor of the seismic data, and use the dynamic factor as the trial calculation granularity of the tail data, and based on the trial calculation granularity, divide the tail data to generate several divided datasets; S252: Execute steps S22 - S23, input each divided dataset into the distribution function model, fit the optimal distribution function model corresponding to each divided dataset, and compare the differences between the optimal distribution function models corresponding to each divided dataset, and calculate the frequency of occurrence of each optimal distribution function model ; ; Among them, is the number of the m th optimal distribution function model fitted from several divided datasets, U is the number of divided datasets; S253: Set the frequency threshold , if there is an optimal distribution function model that satisfies , it is determined that the current trial granularity meets the requirements. The current trial granularity is used as the best-fitting granularity of the tail data, and step S255 is executed; otherwise, it is determined that the current trial granularity does not meet the requirements, and step S254 is executed; S254: Return to step S251, reselect the trial granularity of the tail data, and execute steps S251 - S253; until the best-fitting granularity of the tail data is obtained; S255: Use the optimal distribution function model corresponding to each segmented data set fitted from the tail data under the best-fitting granularity condition as the prior data equation.
[0028] In this embodiment, the data is split with 400,000 magnitude data as one granularity. According to the magnitude data histogram during the period from March 17, 1800 to June 25, 1985, the data with a frequency near 0.4 is selected as the threshold. At this time, the threshold is 4.68, and the magnitude between 4 and 5.5 can be used as the threshold for earthquake data analysis.
[0029] The data simulation results are as follows: Generalized Pareto Distribution (GPD): Goodness-of-fit coefficient N= 62573.2411 / 62599.7955 Log-normal distribution: Goodness-of-fit coefficient N= 173310.9042 / 173328.6072 Cauchy distribution: Goodness-of-fit coefficient N= 105830.9287 / 105848.6317 Exponential distribution: Goodness-of-fit coefficient N= 63635.1960 / BIC = 63652.8990 Generalized Extreme Value Distribution (GEV): Goodness-of-fit coefficient N= 75518.1034 / 75544.6579 Power-law distribution: Goodness-of-fit coefficient N= 69469.2683 / 69486.9713 Gumbel distribution: Goodness-of-fit coefficient N= 79309.7210 / 79327.4240 Weibull distribution: Goodness-of-fit coefficient N= 63728.1792 / 63754.7337 Frechet distribution: Goodness-of-fit coefficient N= 75518.1034 / 75544.6579 According to the above data simulation results, the optimal distribution function model is the Generalized Pareto Distribution (GPD), with a shape coefficient of -0.1520418596254437, a location coefficient of 0.010000097378072906, and a scale coefficient of 0.7852156395133192.
[0030] The magnitude fitting situation with the year split by a granularity of 400,000 earthquake magnitudes can be obtained, as shown in Table 1 below; Table 1 Magnitude fitting situation with a granularity of 400,000 earthquake magnitudes
[0031] Note: Selection of time period granularity: Use 400,000 earthquake magnitude data as a time period unit for granularity separation, except for 2022.6 - 2024.11.1, (253,884) and 2024.1.1 - 2024.10.1.
[0032] Processing of the number of earthquakes in each time period: After dividing the time periods, select earthquake magnitudes greater than 2.5. Earthquake magnitudes below 2.5 are minor earthquake magnitudes.
[0033] Trial calculation threshold: Select a threshold where the frequency is less than 0.4 and the magnitude frequency does not change significantly after exceeding the selected threshold.
[0034] Optimal distribution fitting for the excess: Since this study only focuses on extreme events and their tail characteristics, only the part exceeding the trial calculation threshold is fitted. The several distributions for fitting are as follows: GPD fitting, lognormal fitting, Cauchy fitting, exponential distribution fitting, GEV fitting, power-law distribution fitting, Weibull fitting, Gumbel fitting, Fréchet fitting. These distributions are all used to express the tail characteristics of the data.
[0035] Whether it is heavy-tailed: Judge whether it is heavy-tailed according to the shape parameter of the optimal distribution obtained by fitting.
[0036] Data splitting with a granularity of 5-year earthquake magnitude data: Since the 5-year earthquake magnitude data is less than that in the previous section, earthquake magnitude data with a frequency below 0.2 is taken as the threshold, and a total of 25 years of data from 1999 to 2024 is fitted. With a granularity of five years, the fitting situation is shown in Table 2 below.
[0037] Table 2 Magnitude fitting situation with a granularity of five years
[0038] Data splitting with a granularity of 10-year earthquake magnitude data: Fit 40 years of data from 1984 to 2024. With a granularity of ten years, the fitting situation is shown in Table 3 below.
[0039] Table 3 Magnitude fitting situation with a granularity of ten years
[0040] The data is split with the magnitude data of 20 years as one granularity: Fit the data of 60 years from 1964 to 2024. Taking ten years as one granularity, the fitting situation is shown in Table 4 below.
[0041] Table 4 Magnitude fitting situation with 20 years as one granularity
[0042] Data segmentation situation: According to the above fitting results, taking the number of earthquakes occurring in a time period as the standard, for the threshold selection method with the magnitude data of every five years as one granularity, from the perspective of magnitude, there is no significant fluctuation. Taking the heavy-tailed situation as the standard, the threshold selection method with the magnitude data of every five years as one granularity all shows heavy-tailed phenomena. Taking the trial threshold as the standard, the threshold fluctuates between 4.85 and 5.15, and the change is not significant. Therefore, subsequently, the threshold is selected with the magnitude data of five years as one granularity, and its fitting equation is obtained for derivation with the coin-tossing data.
[0043] S3: Split the non-scenario-dependent event dataset into several sub-datasets of continuous data, use the distribution function models in the distribution function set to fit the distribution law of the sub-datasets, and judge whether the split sub-datasets are reasonable according to the goodness-of-fit coefficient, and use the optimal distribution function model as the distribution equation of the sub-datasets.
[0044] Step S3 specifically includes: S31: Split out a sub-dataset containing continuous data from the non-scenario-dependent event dataset, use the distribution function models in the distribution function set to fit the distribution law of the sub-dataset, and calculate the goodness-of-fit coefficient of each distribution function model to obtain the goodness-of-fit coefficient dataset , is the goodness-of-fit coefficient corresponding to the sub-dataset fitted by the m th distribution function model; S32: Set the threshold of the goodness-of-fit coefficient . If there exists in the goodness-of-fit coefficient dataset , then execute step S33; otherwise, return to step S31 to re-split a sub-dataset from the non-scenario-dependent event dataset; S33: Extract all the goodness-of-fit coefficients in the goodness-of-fit coefficient dataset that satisfy to form the goodness-of-fit coefficient dataset , v is the number of goodness-of-fit coefficients in the goodness-of-fit coefficient dataset , is the goodness-of-fit coefficient dataset in the v th goodness-of-fit coefficient; S34: Take the minimum value in the goodness-of-fit coefficient data set as the optimal distribution function model corresponding to the sub-data set, and use the optimal distribution function model as the distribution equation of the sub-data set; S35: Delete the continuous data in the sub-data set from the non-scenario-dependent event data set, return to step S31, continue to split a sub-data set containing continuous data from the remaining data in the non-scenario-dependent event data set, and execute steps S31 - S34; S36: Until the non-scenario-dependent event data set is split into several sub-data sets, and the distribution equation corresponding to each sub-data set is obtained.
[0045] S4: Compare the function types between the distribution equation and the prior data equation, take the distribution equation with the same type and the smallest coefficient difference as the homologous distribution mapping equation for deriving the scenario-dependent event data, and take the average value of the coefficients to construct a data reconstruction equation, and use the data reconstruction equation to predict the data about to occur for the scenario-dependent event.
[0046] Step S4 specifically includes: S41: Compare the distribution equation corresponding to each sub-data set with the prior data equation corresponding to each segmented data set, compare the coefficient differences between the distribution equations and the prior data equations with the same function type, and take the distribution equation with the smallest coefficient difference as the homologous distribution mapping equation for deriving the scenario-dependent event data; S42: And calculate the average value of the coefficients between the homologous distribution mapping equation and the corresponding prior data equation, use the average value of the coefficients to replace the coefficients of the distribution equation to form a data reconstruction equation, and use the data reconstruction equation to predict the data about to occur for the scenario-dependent event.
[0047] This embodiment takes 2019 - 2024 as an example: If we want to derive earthquake data from coin-tossing data, the tail feature of coin-tossing follows a generalized Pareto distribution, and according to the granularity trial calculation, the AIC and BIC of the generalized Pareto are relatively low. Assume that the part of the earthquake data exceeding the threshold conforms to the generalized Pareto distribution. As Figures 4 - 5 shown, according to the average excess amount diagram of magnitude under different thresholds, the shape coefficient diagram and the modified scale coefficient diagram of the generalized Pareto distribution, and according to Figure 3 , the threshold interval is obtained as (5.1 - 5.5). According to the goodness-of-fit between the data exceeding the threshold and the distribution, the best threshold is 5.2. Refer to Table 5 below, which is a summary of the fitting results of the generalized Pareto distribution for data exceeding the threshold under different thresholds.
[0048] Table 5 Fitting situation of generalized Pareto distribution for data exceeding the threshold under different thresholds
[0049] According to Table 5, by the goodness of fit of data at different thresholds, 5.2 is obtained as the optimal threshold. According to the coefficient diagrams, it is also reasonable to use 5.2 as the threshold. Therefore, 5.2 is used as the threshold from 2019 to 2024.
[0050] Fitting of seismic data: According to this method, for the seismic data from 2009 - 2014 and 2014 - 2019 in two five - year periods, the fitting equation with coin - tossing data is the generalized Pareto equation, and its prediction of historical and future situations is shown in Table 6 below.
[0051] Table 6 Deduction process of coin - tossing and earthquake
[0052] The following Table 7 shows the comparison between the data obtained through the generalized Pareto equation and the data of the past five years and the next five years: Table 7 Comparison of prediction results with seismic data of the past five years and the next five years
[0053] Figure 2 It is a comparison of the interval earthquake occurrence ratio predicted by the data reconstruction equation according to this method with the interval earthquake occurrence ratio every five years and every ten years from 2004 to 2024. Therefore, the data reconstruction equation obtained by this method has good representativeness for the data.
[0054] The present invention is not limited to the traditional method of predicting disaster events through influencing factors of disaster events. Starting from the historical situation of the occurrence of disaster events for disaster prediction, it has a wider applicability. By leveraging the testability of the occurrence of non - context - dependent events (such as coin - tossing, pi) to analyze extreme events with context - dependence (such as earthquakes, typhoons), cross - domain data analysis and prediction are realized, effectively solving the subjectivity problem of threshold selection in data analysis of context - dependent extreme events.
[0055] The present invention proposes a steady - state parameter - oriented model - adaptive threshold selection method, which improves the objectivity and rationality of threshold selection. At the same time, it creatively adopts a data derivation method based on the difference in the same - distribution coefficient, which can effectively predict future extreme event data and evaluate historical data. For non - context - dependent data with a small amount of data, an adaptive fitting strategy is adopted to fully explore the characteristics of tail data.
[0056] The present invention not only provides a scientific basis for disaster prevention and control, but also provides data support for relevant decisions, helping to reduce the damage caused by extreme disasters to society.
Claims
1. A method for predicting heterogeneous extreme events based on tail distribution of data, characterized in that: The following steps are involved: S1: Collect historical data of scenario-dependent events as a scenario-dependent event dataset, clean the data in the scenario-dependent event dataset to generate a stable scenario-dependent event dataset; generate a non-scenario-dependent event dataset based on experimental simulation; S2: Construct a distribution function set that fits the data distribution law, use the distribution function model to fit the distribution law of the scenario-dependent event data set, screen the threshold of the tail data in the scenario-dependent event data set, and split the tail data according to the dynamic factor of the tail data to form several segmented data sets, and use the distribution function model to fit the optimal distribution function model of each segmented data set as the prior data equation; S3: Split the non-scenario-dependent event data set into several sub-datasets of continuous data, use the distribution function model in the distribution function set to fit the distribution law of the sub-dataset, and judge whether the split sub-datasets are reasonable based on the goodness of fit coefficient, and use the optimal distribution function model as the distribution equation of the sub-dataset; S4: Compare the types of functions between the distribution equation and the prior data equation, and use the distribution equation of the same type with the smallest coefficient difference as the homologous distribution mapping equation for deriving scenario-dependent event data, and take the average value of the coefficient to construct the data reconstruction equation, and use the data reconstruction equation to predict the upcoming data of scenario-dependent events.
2. The method for predicting heterogeneous extreme events based on tail distribution of data according to claim 1, characterized in that: The step S2 comprises: S21: Draw a data histogram using the data in the stable scenario-dependent event data set, select a tail threshold range according to the range of data distribution in the data histogram, and use the data within the tail threshold range as a trial calculation data set; S22: constructing a distribution function set that fits the data distribution law, the distribution function set includes several distribution function models, inputting the data in the trial calculation data set into each distribution function model, and fitting the coefficients of each distribution function model; S23: screening the candidate threshold set according to the steady-state of the coefficients fitted by each distribution function model within the tail threshold range, and determining the optimal distribution function model within the tail threshold range based on the goodness of fit coefficient corresponding to each distribution function model; S24: Calculate the difference between the fitted value and the true value based on the optimal distribution function model, take the true value with the smallest difference as the super threshold in the candidate threshold set, and use the super threshold to filter the tail data in the scenario-dependent event data set; S25: Using the dynamic factor of the tail data as the trial calculation granularity, segmenting the tail data based on the trial calculation granularity to generate several segmented data sets; determining the prior data equation of each segmented data set according to the frequency of occurrence of the optimal distribution function model in the segmented data sets.
3. The method for predicting heterogeneous extreme events based on tail distribution of data according to claim 2, characterized in that: The step S23 comprises: S231: Determine whether the tail threshold range is used as a candidate threshold set according to the steady-state condition of the coefficients fitted by each distribution function model within the tail threshold range; if it cannot be used as a candidate threshold set, return to step S21 and reselect the tail threshold range; if it can be used as a candidate threshold set, execute step S232; The method for determining the steady-state condition of the coefficient is: Calculate the stability coefficient of the coefficients fitted by each distribution function model using data within the tail threshold range ; ; in, i is the coefficient number, k is the number of coefficients, For the fitted i coefficients, is the coefficient mean; Set the threshold value of the stability factor ,like , then the fitted coefficients are determined to be stable, and the tail threshold range can be used as a candidate threshold set; otherwise, the fitted coefficients are unstable, and the tail threshold range cannot be used as a candidate threshold set; S232: Get the coefficients of each distribution function model fitting Number of k , and use the residual square and mean square error between the true value of the data fitted by each distribution function model and the fitted value to calculate the fitting likelihood value , and then calculate the goodness of fit coefficient ; ; in, n is the amount of data in the candidate threshold set, AIC is the Akaike information criterion, and BIC is the Bayesian information criterion; S233: Get the goodness-of-fit coefficient corresponding to each distribution function model , m is the type of distribution function model, For the m The goodness-of-fit coefficient of the distribution function model is selected, and the goodness-of-fit coefficient is selected. The minimum value in , minimum The corresponding distribution function model is taken as the optimal distribution function model within the tail threshold range.
4. The method for predicting heterogeneous extreme events based on tail distribution of data according to claim 3, characterized in that: The step S25 comprises: S251: extracting the data change cycle law of the tail data, generating a dynamic factor of the tail data, and using the dynamic factor as the trial calculation granularity of the tail data, segmenting the tail data based on the trial calculation granularity to generate a number of segmented data sets; S252: Execute steps S22-S23, input each segmented data set into the distribution function model, fit the optimal distribution function model corresponding to each segmented data set, and compare the differences between the optimal distribution function models corresponding to each segmented data set, and calculate the frequency of occurrence of each optimal distribution function model ; ; in, The first m The number of optimal distribution function models, U is the number of split data sets; S253: Setting frequency threshold , if there exists an optimal distribution function model that satisfies , it is determined that the current trial calculation granularity meets the requirements, and the current trial calculation granularity is used as the best fitting granularity of the tail data, and step S255 is executed; otherwise, it is determined that the current trial calculation granularity does not meet the requirements, and step S254 is executed; S254: Return to step S251, reselect the trial calculation granularity of the tail data, and execute steps S251-S253 until the best fitting granularity of the tail data is obtained; S255: The optimal distribution function model corresponding to each segmented data set fitted by the tail data under the best fitting granularity condition is used as a priori data equation.
5. The method for predicting heterogeneous extreme events based on tail distribution of data according to claim 4, characterized in that: The step S3 comprises: S31: Split a sub-dataset containing continuous data from the non-scenario-dependent event data set, use the distribution function model in the distribution function set to fit the distribution law of the sub-dataset, and calculate the goodness-of-fit coefficient of each distribution function model to obtain the goodness-of-fit coefficient data set , For the m The goodness-of-fit coefficient corresponding to the distribution function model fitting sub-dataset; S32: Set the threshold of the goodness-of-fit coefficient , if the goodness-of-fit coefficient data set Existence , then execute step S33; otherwise, return to step S31 and re-split a sub-dataset from the non-context-dependent event data set; S33: Extract the goodness-of-fit coefficient data set All satisfied The goodness-of-fit coefficient of , v is the goodness coefficient data set The number of goodness-of-fit coefficients in , is the goodness coefficient data set Middle v goodness-of-fit coefficients; S34: Fit goodness of fit coefficient data set The minimum value in As the optimal distribution function model corresponding to the sub-data set, and the optimal distribution function model is used as the distribution equation of the sub-data set; S35: deleting the continuous data in the sub-dataset from the non-context-dependent event data set, returning to step S31, continuing to split a sub-dataset containing continuous data from the remaining data in the non-context-dependent event data set, and executing steps S31-S34; S36: until the non-context-dependent event data set is split into several sub-data sets, and the distribution equation corresponding to each sub-data set is obtained.
6. The method for predicting heterogeneous extreme events based on tail distribution of data according to claim 5, characterized in that: The step S4 comprises: S41: Compare the distribution equation corresponding to each sub-data set with the prior data equation corresponding to each segmented data set, compare the coefficient differences between the distribution equations and the prior data equations with the same function type, and use the distribution equation with the smallest coefficient difference as the homologous distribution mapping equation for deriving scenario-dependent event data; S42: Calculate the average value of coefficients between the homologous distribution mapping equation and the corresponding prior data equation, use the average value of coefficients to replace the coefficients of the distribution equation to form a data reconstruction equation, and use the data reconstruction equation to predict the data of the scenario-dependent event that is about to occur.
Citation Information
Patent Citations
Coastal tide level extreme value prediction method based on extreme value mixed distribution
CN116227190A
Data shortage estuary composite flood disaster research method considering climatic change
CN118760899A
Water quality index interval prediction method in non-consistency scene
CN119577698A
Frequency estimation of rare events by adaptive thresholding
US20100114526A1