An extreme natural disaster intensity probability prediction method and device based on minimum sample size calculation
Patent Information
- Application Number
- CN202610866759.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-09-25
AI Technical Summary
然而,上述方法大多基于纯数据驱动,尚未从统计推断的角度,给出关于现有样本规模是否足以支撑特定强度概率模型及尾部参数的可信估计的明确判据
[0018]与现有技术相比,本发明的有益效果包括:
Smart Images

Figure CN122817620A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of meteorological disaster risk management, and in particular to a method and device for predicting the probability of extreme natural disaster intensity based on minimum sample size calculation. Background Technology
[0002] In recent years, widespread power outages triggered by extreme natural disasters have occurred frequently in many places. Typhoons, torrential rains, and earthquakes, although seemingly low-probability, high-impact events, often cause widespread power outages.
[0003] Statistically, extreme natural disasters are characterized by data scarcity in both time and space. Statistical results show that tropical cyclones making landfall along the coast and reaching super typhoon strength, recorded torrential rain events, and earthquakes of magnitude 7 or higher are all extremely rare. The scarcity of samples for extreme events such as super typhoons, torrential rains, and strong earthquakes directly leads to significant errors in tail parameter estimation and insufficient extrapolation capabilities in models constructing intensity probability distributions based on historical data. Furthermore, due to significant differences in climate and geological conditions across regions, the southeastern coastal areas primarily face typhoon disasters, inland areas are more prone to torrential rain disasters, and areas located on active seismic zones have a higher earthquake risk. This regional difference in dominant disaster types further reduces the number of available local extreme disaster samples, making traditional probabilistic modeling and risk assessment relying on regional historical data face higher uncertainties.
[0004] To address the modeling difficulties caused by the scarcity of data on extreme natural disasters, scholars have proposed various methods, including Bayesian statistics, resampling techniques, and generative artificial intelligence. Bayesian methods, by setting a prior distribution, fuse empirical knowledge with observational data to progressively correct the disaster intensity distribution. Bootstrap resampling improves the stability of parameter estimation by repeatedly sampling with replacement from a finite sample. Bayes Bootstrap, building upon this, incorporates Bayesian ideas to further enhance the robustness of estimations in small sample situations. Meanwhile, data generation models such as generative adversarial networks have also been explored for disaster sample expansion. However, most of these methods are purely data-driven and lack clear criteria from a statistical inference perspective regarding whether the existing sample size is sufficient to support a reliable estimate of a specific intensity probability model and tail parameters. Especially in the context of tail modeling for extreme natural disaster intensity, a method for determining the minimum sample size is lacking.
[0005] Chinese patent application CN118644099A proposes a method for modeling and risk assessment of the probabilistic characteristics of typhoon intensity tails under small sample data. Addressing the problem that existing technologies struggle to effectively consider the thick tail characteristics of typhoon intensity, leading to an underestimation of potential risks to the power system, especially with small sample data, the patent proposes a mechanism-data fusion framework for modeling the probabilistic characteristics of typhoon intensity. Based on this framework, it elaborates on the determination of typhoon intensity probability distribution types and parameter estimation to achieve the assessment of power system tail risks. However, this patent primarily focuses on modeling the thick tail characteristics of extreme wind speeds at the typhoon tail. Like the aforementioned existing technologies, it still uses a method adapted to small sample data rather than determining the minimum required sample size for probabilistic characteristic modeling and risk assessment. When data is insufficient, it still forces modeling, which can easily lead to large parameter deviations and unreliable extrapolations.
[0006] In summary, what is needed now is a method for constructing probabilistic models of extreme natural disaster intensity that can know the gap between the minimum sample size required for modeling and the existing local sample size, accurately determine how many samples need to be migrated from the surrounding area or the global area, avoid blind migration or insufficient migration, and find the optimal balance between engineering feasibility, computational cost, and model credibility. Summary of the Invention
[0007] The purpose of this invention is to overcome the shortcomings of the existing technology by providing a method and device for predicting the probability of extreme natural disaster intensity based on minimum sample size calculation, thereby improving the statistical reliability of the load loss risk quantification results from the source and providing scientific support for disaster prevention and mitigation investment decisions and the improvement of power grid resilience.
[0008] The objective of this invention can be achieved through the following technical solutions: A method for predicting the intensity probability of extreme natural disasters based on minimum sample size calculation, the method comprising: Based on the nonlinear behavior and long-range correlation of extreme natural disasters, a basic mixed distribution model is established to characterize the probabilistic characteristics of the entire range of disasters; the basic mixed distribution model includes the Weibull distribution function and the truncated power-law distribution function; Based on the two-layer bootstrapping and Shapiro-Wilk normality test, the model parameters of the basic mixed distribution model are sampled and evaluated to obtain the minimum sample size required for parameter inference. Data samples of extreme natural disasters are collected based on the minimum sample size, and the model parameters of the basic mixed distribution model are calculated based on the maximum likelihood parameter estimation method of the mixed distribution to improve the basic mixed distribution model and obtain the intensity probability model of extreme natural disasters. The effectiveness of the extreme natural disaster intensity probability model is verified through a credibility test, and the intensity probability of extreme natural disasters is predicted based on the effective extreme natural disaster intensity probability model.
[0009] Furthermore, the Weibull distribution function characterizes the disaster intensity within a local intensity range, and the truncated power-law distribution function characterizes the extreme tails of the disaster intensity.
[0010] Furthermore, the truncated power-law distribution function is constructed based on the physical upper limit of extreme natural disasters.
[0011] Furthermore, the extreme natural disasters include typhoons, torrential rains, and earthquakes; the physical limits of the extreme natural disasters include the upper limit of typhoon wind speed, the upper limit of torrential rain intensity, and the upper limit of earthquake magnitude; wherein, The upper limit of typhoon wind speed is determined by air-sea thermal conditions, vertical wind shear, and multi-scale feedback processes. The upper limit of the intensity of the rainstorm can be determined by the water vapor content and the time and space of dynamic lifting. Regional stress reserves, fault geometry, and medium strength determine the upper limit of earthquake magnitude.
[0012] Furthermore, the expression for the basic mixed distribution model is: in, Disaster intensity The probability density function, i.e., the total mixture distribution, This represents the relative proportion of the Weibull distribution within the mixed distribution. Let Weibull distribution function be used. To truncate the power-law distribution function, As for the intensity of the disaster, Let the location parameters be those of the Weibull distribution. Let be the shape parameter of the Weibull distribution. Let be the scale parameter of the Weibull distribution. To truncate the shape parameter of the power-law distribution, The scaling parameter is used to truncate the power-law distribution.
[0013] Furthermore, the process of sampling and evaluating the model parameters of the basic mixed distribution model includes: Preliminary preparation: Based on the aforementioned basic mixed distribution model, determine the parameters to be evaluated and set the sampling parameters, which include the original observed sample, the candidate sample size set, the number of outer layer bootstrapping R, the number of inner layer bootstrapping B, and the significance threshold; Outer layer sampling: Based on the number of outer layer bootstrappings, perform R repeated samplings with replacement on the original observation sample to obtain R independent bootstrap sets; based on each candidate sample size j in the candidate sample size set, extract the first j samples from each bootstrap set to form R subsamples of size j. Inner layer sampling: Based on the number of inner layer bootstrappings, for each subsample of size j, perform B inner layer samplings with replacement to obtain B inner layer subsamples; for each inner layer subsample, use the mixture distribution maximum likelihood estimation algorithm to fit the basic mixture distribution model and calculate the estimated values of the model parameters; summarize the parameter estimation results of B inner layer samplings to form a parameter estimate sample under the candidate sample size j. Normality test and statistic calculation: Perform the Shapiro-Wilk normality test on the parameter estimate sample to obtain the p-value of the test; calculate the mean and standard deviation of the R p-values corresponding to the R outer set; construct the confidence interval of the p-value using a t-distribution with R-1 degrees of freedom; Sampling evaluation results: If the lower limit of the confidence interval is greater than the significance threshold, it is determined that the model parameter estimate meets the normal approximation requirement under the candidate sample size j, the sampling evaluation is passed, and the current candidate sample size j is output as the minimum sample size; otherwise, it is determined that the sampling is insufficient, a larger candidate sample size j is replaced, and the outer layer sampling, inner layer sampling, normality test and statistic calculation are repeated until the minimum sample size is found.
[0014] Furthermore, the mixed distribution maximum likelihood estimation algorithm is an EM iterative algorithm. The steps of the EM iterative algorithm include calculating the posterior probability and maximizing the log-likelihood function. The EM iterative algorithm is repeated until the parameter difference between two adjacent iterations is less than a preset allowable error.
[0015] Furthermore, the credibility test includes the KS test, which is used to verify the fitting effect of the mixed distribution model on the intensity of extreme natural disasters.
[0016] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method for predicting the probability of extreme natural disaster intensity based on minimum sample size calculation as described above.
[0017] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for predicting the probability of extreme natural disaster intensity based on a minimum sample size, as described above.
[0018] Compared with the prior art, the beneficial effects of the present invention include: 1. This invention addresses the problem of probabilistic modeling and quantitative determination of minimum sample size for extreme natural disaster intensity. By calculating the minimum sample size for probabilistic modeling of extreme disasters, it achieves a scientific determination of sample size, avoiding modeling failure due to insufficient samples. It can complete reliable modeling without relying on a large amount of historical data, controlling sample and parameter errors from the source of modeling, making quantitative results such as load loss risk and disaster probability more reliable, and avoiding risk misjudgment due to model bias. This invention also addresses the dual characteristics of local power-law tails and physical upper limit truncation of extreme disasters by constructing a hybrid probability model of Weibull distribution plus truncated power-law distribution, integrating the parameter characteristics of the two types of distributions to fully characterize the probabilistic law of the entire range of disaster intensity.
[0019] 2. This invention proposes a minimum sample size calculation method based on two-layer bootstrapping and Shapiro-Wilk normality test. By performing bootstrapping sampling and normality testing on the estimated model parameters, a confidence interval for the test p-value is constructed. Using the lower limit of this interval being greater than a preset significance level as a criterion, the minimum sample size that satisfies the normal approximation requirement for parameter inference can be determined. This allows for the clear identification of critical values for parameter estimation error, confidence interval width, and extrapolation failure risk when the sample size is insufficient during modeling. This ensures that the model results have a credible range, rather than being blindly output, preventing the use of unreliable models to guide engineering defenses. This improves the statistical reliability of the quantitative results of load shedding risk from the source, providing scientific support for disaster prevention and mitigation investment decisions and the improvement of power grid resilience.
[0020] 3. This invention adopts a two-layer resampling architecture with outer and inner bootstrap layers. The outer layer samples the original samples with replacement to generate multiple parent sets, and the inner layer resamples the subsamples of the parent set to generate parameter estimation samples. This solves the problem of large fluctuations in parameter estimation under small sample conditions and greatly improves the stability of parameter estimation under small sample conditions.
[0021] 4. This invention ensures that the parameters meet the asymptotic normality requirement through the SW normality test, making parameter inference and risk extrapolation more statistically based.
[0022] 5. This invention addresses the latent label characteristics of mixed distributions by employing the EM iterative algorithm. The E-step calculates the posterior probability of a sample belonging to one of the two distributions, and the M-step maximizes the log-likelihood function to update the parameters, thereby achieving accurate parameter fitting for the mixed distribution.
[0023] 6. This invention evaluates the model fitting effect through credibility verification methods such as the KS test, and directly applies the credible model to the assessment of the risk of load loss in urban power grids under extreme disasters, providing quantitative support for disaster prevention decision-making. Attached Figure Description
[0024] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a schematic diagram of the average p-value and confidence interval of the statistic based on the SW test in this invention; Figure 3 This is the probability distribution of earthquake intensity in a certain region according to an embodiment of the present invention. Detailed Implementation
[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0026] Example 1 This embodiment discloses a method for predicting the probability of extreme natural disaster intensity based on minimum sample size calculation. The method is as follows: Figure 1 The diagram includes steps S1-S4, each described in detail below: Step S1: Based on the nonlinear behavior and long-range correlation of extreme natural disasters, establish a basic mixed distribution model that characterizes the probabilistic characteristics of the entire disaster range.
[0027] The basic mixed distribution model includes the Weibull distribution function and the truncated power-law distribution function.
[0028] The Weibull distribution function characterizes the intensity of a disaster within a local intensity range, while the truncated power-law distribution function characterizes the extreme tails of the disaster intensity. The truncated power-law distribution function is constructed based on the physical upper limit of extreme natural disasters.
[0029] Extreme natural disasters include typhoons, torrential rains, and earthquakes; the physical limits for extreme natural disasters include the upper limit for typhoon wind speed, the upper limit for torrential rain intensity, and the upper limit for earthquake magnitude; among them, The upper limit of typhoon wind speed is determined by air-sea thermal conditions, vertical wind shear, and multi-scale feedback processes. The upper limit of rainfall intensity can be determined by the water vapor content and the time and space of dynamic lifting. Regional stress reserves, fault geometry, and medium strength determine the upper limit of earthquake magnitude.
[0030] Specifically, extreme natural disasters typically exhibit significant nonlinear behavior and long-range correlations. Within finite scales, their statistical characteristics often show an approximately scale-invariant pattern, and the tail of the event intensity distribution can be characterized by a power-law function. However, disaster intensity is inevitably limited by physical constraints such as energy supply, transmission, and release. Power-law behavior cannot extend infinitely across the entire intensity range. In regions close to physical extremes, the tail probability often exhibits accelerated decay relative to the power law. Therefore, the statistical form used for modeling extreme tails should retain power-law scaling characteristics within local intensity intervals while reflecting the tail truncation mechanism induced by the physical upper limit.
[0031] Taking typical disaster types as examples, the formation and development of typhoons are jointly constrained by air-sea thermal conditions, vertical wind shear, and multi-scale feedback processes, and their maximum wind speed has an upper limit determined by dynamic processes; rainstorm processes are limited by the available water vapor content and the temporal and spatial constraints of dynamic lifting, and their cumulative precipitation intensity cannot increase indefinitely; seismic activity originates from the long-term accumulation of fault stress and the release of energy by sudden slippage, and the magnitude is ultimately constrained by a combination of factors such as regional stress reserves, fault geometry, and medium strength. Therefore, in order to characterize the probabilistic characteristics of disasters across the entire range, a hybrid distribution model is used to describe disaster intensity. .
[0032] The expression for the basic mixed distribution model is: in, Disaster intensity The probability density function, i.e., the total mixture distribution, This represents the relative proportion of the Weibull distribution within the mixed distribution. Let Weibull distribution function be used. To truncate the power-law distribution function, As for the intensity of the disaster, Let the location parameters be those of the Weibull distribution. Let be the shape parameter of the Weibull distribution. Let be the scale parameter of the Weibull distribution. To truncate the shape parameter of the power-law distribution, The scaling parameter is used to truncate the power-law distribution.
[0033] Step S2: Based on the two-layer bootstrapping and Shapiro-Wilk normality test, the model parameters of the basic mixed distribution model are sampled and evaluated to obtain the minimum sample size required for parameter inference.
[0034] The process of sampling and evaluating the model parameters of the basic mixed distribution model includes: Preliminary preparation: Based on the basic mixed distribution model, determine the parameters to be evaluated and set the sampling parameters, including the original observed sample, the candidate sample size set, the number of outer bootstraps R, the number of inner bootstraps B, and the significance threshold; Outer layer sampling: Based on the number of outer layer bootstrappings, perform R repeated samplings with replacement on the original observation sample to obtain R independent bootstrap sets; based on the number of candidate samples j in the candidate sample size set, extract the first j samples from each bootstrap set to form R subsamples of size j. Inner layer sampling: Based on the number of inner layer bootstrappings, for each subsample of size j, perform B inner layer samplings with replacement to obtain B inner layer subsamples; for each inner layer subsample, use the mixture distribution maximum likelihood estimation algorithm to fit the basic mixture distribution model and calculate the estimated values of the model parameters; summarize the parameter estimation results of B inner layer samplings to form a parameter estimate sample under the candidate sample size j. Normality test and statistic calculation: Perform the Shapiro-Wilk normality test on the parameter estimate sample and obtain the p-value of the test; calculate the mean and standard deviation of the R p-values corresponding to the R outer set; construct the confidence interval of the p-value using a t-distribution with R-1 degrees of freedom; Sampling evaluation results: If the lower limit of the confidence interval is greater than the significance threshold, it is determined that the model parameter estimate meets the normal approximation requirement under the candidate sample size j, the sampling evaluation is passed, and the current candidate sample size j is output as the minimum sample size. Otherwise, it is determined that the sampling is insufficient, and a larger candidate sample size j is replaced. The outer layer sampling, inner layer sampling, normality test and statistic calculation are repeated until the minimum sample size is found.
[0035] Specifically, when the sample size is sufficiently large, the maximum likelihood estimator follows an asymptotically normal distribution under conventional regularization conditions. Therefore, as long as the parameter estimator approximates a normal distribution, the minimum sample size required for parameter inference can be determined. Thus, firstly, a candidate sample size set is defined by the number of bootstrapping iterations; secondly, for each candidate sample size, multiple parent sets are generated using outer bootstrapping, and then inner bootstrapping is performed on the parent sets to fit the model and obtain the estimator sample; next, the SW normality test is performed on the estimator sample to obtain a p-value sequence; then, the mean and standard deviation of the p-value sequence are calculated, and its confidence interval is constructed using a t-distribution; if the lower limit of the interval is greater than 0.05, the normal approximation is considered acceptable; at this point, the minimum candidate sample size that meets the conditions is the minimum sample size required for modeling. A more detailed description of the steps is as follows: Let the observed sample be The candidate sample size set is ,in outer layer bootstrap count R, inner layer bootstrap count B, and significance threshold. .
[0036] (1) To Sampling with replacement yields R bootstrap mother sets. For each For each r, take the first j samples from the parent set to form a subsample. .
[0037] (2) For fixed ,right B inner layer sampling was performed to obtain one by one Upper Estimated Parameters And calculate: .
[0038] (3) For sets Perform the Shapiro-Wilk normality test and obtain the p-value for this group, denoted as . Calculate the R p values for a fixed j: ; And construct using a t-distribution with R-1 degrees of freedom confidence interval .
[0039] (4) If the lower limit of the interval satisfies At that time, it is determined that a normal approximation is acceptable for a sample size of j, and we take: This is the minimum sample size required for modeling; if the sample size is insufficient, the sample size needs to be increased before estimating the parameters of the probability model.
[0040] Step S3: Collect data samples of extreme natural disasters based on the minimum sample size, and calculate the model parameters of the basic mixed distribution model based on the maximum likelihood parameter estimation method of mixed distribution, improve the basic mixed distribution model, and obtain the probability model of extreme natural disaster intensity.
[0041] The mixed distribution maximum likelihood estimation algorithm is the EM iterative algorithm. The steps of the EM iterative algorithm include calculating the posterior probability and maximizing the log-likelihood function. The EM iterative algorithm is repeated until the parameter difference between two adjacent iterations is less than the preset allowable error.
[0042] Specifically, the steps of the mixture distribution maximum likelihood estimation algorithm are as follows: Set hidden tags Representing samples respectively The probability of belonging to a Weibull distribution and a truncated power-law distribution, the log-likelihood function of the complete data is: .
[0043] E-Step: In Calculate the posterior probability: .
[0044] M-step: Maximizing the log-likelihood function: .
[0045] when When the iteration terminates, This is the allowable error.
[0046] Step S4: Verify the effectiveness of the extreme natural disaster intensity probability model through a credibility test, and make an extreme natural disaster intensity probability prediction based on the effective extreme natural disaster intensity probability model.
[0047] Credibility testing includes the KS test, which verifies the fitting effect of the mixed distribution model on the intensity of extreme natural disasters. Models that pass the test can be used to support disaster prevention investment decisions and resilience enhancement of urban power grids, reducing the risk of load loss from extreme disasters at the source.
[0048] Specifically, based on the effective extreme natural disaster intensity probability model, the probability of occurrence of disasters of different intensities such as typhoons, rainstorms, and earthquakes is calculated, the disaster resistance threshold of power grid equipment is matched, and the probability and scale of loss of substation outages, line trips, and regional load loss are quantitatively assessed, thereby improving the statistical reliability of risk calculation from the source.
[0049] Based on the probability distribution of disaster intensity, identify power grid lines, substations, and distribution facilities with insufficient disaster resistance and high failure probability, and clarify key protection targets.
[0050] For power grid reinforcement, renovation, and disaster preparedness solutions, models are used to simulate the operating conditions under different disaster intensities, quantitatively verify the disaster resistance effectiveness of the solutions, and support the selection of the best solution.
[0051] This embodiment also provides a practical application example based on the above method: Taking an earthquake disaster in a certain region as an example, the minimum sample size required for modeling is first calculated: To estimate the parameters of a probabilistic earthquake intensity model for a certain region, the minimum sample size is first calculated using the method described above. When the lower limit of the confidence interval for the p-value of the SW test meets the requirement of being greater than 0.05, the minimum sample size required for modeling is determined to be 165. Figure 2 As shown.
[0052] Step 2: Based on the minimum sample size, establish a probabilistic model for earthquake disasters.
[0053] For earthquake disasters, if a certain region only has 68 earthquake data samples, it is necessary to migrate data from surrounding areas to construct a system such as... Figure 3 The earthquake intensity probability distribution model shown has the following statistics: KS, and The values reached 0.1204, 0.7981, and 0.8086 respectively, all meeting the significance level requirements, indicating that the model can well characterize the probability distribution characteristics of earthquake intensity in a certain region.
[0054] Step 3: Probabilistic modeling of earthquake intensity in the surrounding area and the entire region is performed, with parameters shown in Table 1.
[0055] Table 1 Parameters of the Earthquake Intensity Probabilistic Model The differences in earthquake probability across regions are primarily determined by the distribution and activity of seismic zones. Surrounding area 2, located within a seismic zone, has a high probability of earthquake occurrence, significantly exceeding the average level. A certain region exhibits relatively low seismic activity. Surrounding area 1 is situated in a relatively stable block, far from major strong seismic zones, and has low geological tectonic activity.
[0056] Example 2 Based on Embodiment 1, this embodiment provides an electronic device, including: one or more processors and a memory, wherein the memory stores one or more programs, the one or more programs including instructions for executing the aforementioned method for predicting the probability of extreme natural disaster intensity based on minimum sample size calculation.
[0057] At the hardware level, the electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to implement the above-mentioned method for predicting the probability of extreme natural disaster intensity based on minimum sample size calculation. Of course, in addition to software implementation, this invention does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.
[0058] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0059] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0060] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for predicting the probability of extreme natural disaster intensity based on minimum sample size calculation, characterized in that, The method includes: Based on the nonlinear behavior and long-range correlation of extreme natural disasters, a basic mixed distribution model characterizing the probabilistic characteristics of the entire disaster range is established. This basic mixed distribution model includes a Weibull distribution function and a truncated power-law distribution function. Based on two-level bootstrapping and the Shapiro-Wilk normality test, the model parameters of the basic mixed distribution model are sampled and evaluated to obtain the minimum sample size required for parameter inference. Data samples of extreme natural disasters are collected based on this minimum sample size, and the model parameters of the basic mixed distribution model are calculated using the maximum likelihood parameter estimation method for mixed distributions. This refines the basic mixed distribution model, resulting in a probability model for the intensity of extreme natural disasters. The effectiveness of the probability model for the intensity of extreme natural disasters is verified through a credibility test, and probability predictions of the intensity of extreme natural disasters are made based on this effective probability model.
2. The method for predicting the probability of extreme natural disaster intensity based on minimum sample size calculation according to claim 1, characterized in that, The Weibull distribution function characterizes the disaster intensity within a local intensity range, and the truncated power-law distribution function characterizes the extreme tails of the disaster intensity.
3. The method for predicting the probability of extreme natural disaster intensity based on minimum sample size calculation according to claim 1, characterized in that, The truncated power-law distribution function is constructed based on the physical upper limit of extreme natural disasters.
4. The method for predicting the probability of extreme natural disaster intensity based on minimum sample size calculation according to claim 3, characterized in that, The extreme natural disasters mentioned include typhoons, torrential rains, and earthquakes; the physical limits for these extreme natural disasters include the upper limit for typhoon wind speed, the upper limit for torrential rain intensity, and the upper limit for earthquake magnitude; among which... The upper limit of typhoon wind speed is determined by air-sea thermal conditions, vertical wind shear, and multi-scale feedback processes. The upper limit of the intensity of the rainstorm can be determined by the water vapor content and the time and space of dynamic lifting. Regional stress reserves, fault geometry, and medium strength determine the upper limit of earthquake magnitude.
5. The method for predicting the probability of extreme natural disaster intensity based on minimum sample size calculation according to claim 1, characterized in that, The expression for the basic mixed distribution model is: in, Disaster intensity The probability density function, i.e., the total mixture distribution, This represents the relative proportion of the Weibull distribution within the mixed distribution. Let Weibull distribution function be used. To truncate the power-law distribution function, As for the intensity of the disaster, Let the location parameters be those of the Weibull distribution. Let be the shape parameter of the Weibull distribution. Let be the scale parameter of the Weibull distribution. To truncate the shape parameter of the power-law distribution, The scaling parameter is used to truncate the power-law distribution.
6. The method for predicting the probability of extreme natural disaster intensity based on minimum sample size calculation according to claim 1, characterized in that, The process of sampling and evaluating the model parameters of the basic mixed distribution model includes: Preliminary preparation: Based on the aforementioned basic mixed distribution model, determine the parameters to be evaluated and set the sampling parameters, which include the original observed sample, the candidate sample size set, the number of outer layer bootstrapping R, the number of inner layer bootstrapping B, and the significance threshold; Outer layer sampling: Based on the number of outer layer bootstrappings, perform R repeated samplings with replacement on the original observation sample to obtain R independent bootstrap sets; based on each candidate sample size j in the candidate sample size set, extract the first j samples from each bootstrap set to form R subsamples of size j. Inner layer sampling: Based on the number of inner layer bootstrappings, for each subsample of size j, perform B inner layer samplings with replacement to obtain B inner layer subsamples; for each inner layer subsample, use the mixture distribution maximum likelihood estimation algorithm to fit the basic mixture distribution model and calculate the estimated values of the model parameters; summarize the parameter estimation results of B inner layer samplings to form a parameter estimate sample under the candidate sample size j. Normality test and statistic calculation: Perform the Shapiro-Wilk normality test on the parameter estimate sample to obtain the p-value of the test; calculate the mean and standard deviation of the R p-values corresponding to the R outer set; construct the confidence interval of the p-value using a t-distribution with R-1 degrees of freedom; Sampling evaluation results: If the lower limit of the confidence interval is greater than the significance threshold, it is determined that the model parameter estimate meets the normal approximation requirement under the candidate sample size j, the sampling evaluation is passed, and the current candidate sample size j is output as the minimum sample size; otherwise, it is determined that the sampling is insufficient, a larger candidate sample size j is replaced, and the outer layer sampling, inner layer sampling, normality test and statistic calculation are repeated until the minimum sample size is found.
7. The method for predicting the probability of extreme natural disaster intensity based on minimum sample size calculation according to claim 6, characterized in that, The mixed distribution maximum likelihood estimation algorithm is an EM iterative algorithm. The steps of the EM iterative algorithm include calculating the posterior probability and maximizing the log-likelihood function. The EM iterative algorithm is repeated until the parameter difference between two adjacent iterations is less than the preset allowable error.
8. The method for predicting the probability of extreme natural disaster intensity based on minimum sample size calculation according to claim 1, characterized in that, The credibility test includes the KS test, which is used to verify the fitting effect of the mixed distribution model on the intensity of extreme natural disasters.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method for predicting the probability of extreme natural disaster intensity based on minimum sample size as described in any one of claims 1-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method for predicting the probability of extreme natural disaster intensity based on minimum sample size as described in any one of claims 1-8.
Citation Information
Patent Citations
Typhoon intensity tail probability characteristic modeling and risk assessment method under small sample data
CN118644099A