A regional flood frequency analysis method based on high-order linear moments

By using the higher-order linear moment method, the problem of insufficient estimation capability of linear moments in extreme flood events is solved. A regional flood frequency analysis framework is constructed, which improves the accuracy and reliability of design flood estimation and is applicable to regional flood risk assessment at sites with no or scarce data.

CN122434264APending Publication Date: 2026-07-21TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TAIYUAN UNIVERSITY OF TECHNOLOGY
Filing Date
2026-04-27
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing regional flood frequency analysis methods based on linear moments are less sensitive to extreme flood events, leading to significant deviations in the design results of floods with large return periods.

Method used

By employing the higher-order linear moment method, the ability to characterize the upper tail of flood distribution is enhanced by increasing the order of the linear moments. Combined with data screening, regional homogeneity testing, and distribution optimization, a regional frequency analysis framework is constructed to eliminate discordant sites, assess regional heterogeneity, determine the optimal distribution function and the order of the linear moments, and derive the flood design value.

Benefits of technology

It significantly improves the accuracy and reliability of design flood estimation for sites with no or scarce data, enhances the ability to estimate extreme flood events, and provides a more reliable basis for regional flood risk assessment and flood control engineering design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122434264A_ABST
    Figure CN122434264A_ABST
Patent Text Reader

Abstract

The present application relates to the field of regional flood frequency analysis, and particularly relates to a regional flood frequency analysis method based on high-order linear moment. The present application introduces high-order linear moment to construct a regional frequency analysis framework, effectively overcomes the defect that the traditional linear moment method is excessively sensitive to the lower tail of the distribution and insufficient to depict the upper tail, and significantly improves the estimation ability and stability of long return period design flood. The method can more reliably calculate the design flood value of the site without data or with scarce data, and provides a more reliable scientific basis for regional flood risk assessment and flood control engineering design. The present application constructs a complete technical process from data screening, regional homogeneity test, regional distribution optimization to cross-validation, which is suitable for various basins with a certain length of flow observation data and the need for design flood estimation of sites without data. The present application can also be extended to other regions with similar data basis, and has good practicability and wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of regional flood frequency analysis research, and specifically to a regional flood frequency analysis method based on higher-order linear moments. Background Technology

[0002] Regional flood frequency analysis extrapolates design floods for stations with no or scarce data by utilizing flood data from existing stations within hydrologically similar regions. Currently, regional frequency analysis methods based on linear moments (L-moments) are widely used due to their good stability in parameter estimation. However, linear moments have limitations in handling extreme flood events, mainly manifested in their weak sensitivity to the upper tail of the distribution (i.e., the portion with larger flows), leading to significant deviations in the design results for floods with long return periods.

[0003] Higher-order linear moments (LH-moments), as a higher-order extension of linear moments, can better characterize the upper tail features of flood distribution and reduce the impact of small sample events on extreme flood estimation. Although LH-moments have significant theoretical advantages, their application in regional flood frequency analysis is still in the exploratory stage, and a complete and systematic analytical framework has not yet been formed. Therefore, it is necessary to develop a systematic regional flood frequency analysis method based on LH-moments to improve the accuracy and reliability of design flood estimation for sites with no or scarce data. Summary of the Invention

[0004] The purpose of this invention is to overcome the problem that existing regional flood frequency analysis methods based on linear moments are insufficient in estimating extreme flood events (i.e., the upper tail of the distribution), and to provide a regional flood frequency analysis method based on higher-order linear moments. This method enhances the characterization of the upper tail of the flood distribution by increasing the order of the linear moments, effectively reducing the interference of low-flow data on the estimation of high-recurrence-period floods, thereby improving the accuracy and reliability of design flood estimation for sites with no or scarce data.

[0005] This invention is achieved using the following technical method: a regional flood frequency analysis method based on higher-order linear moments, comprising the following steps:

[0006] Step 1: Collect annual maximum flood peak flow data from multiple hydrological stations within the study area, and preprocess the data to form a candidate station set; the annual maximum flood peak flow data includes multi-year maximum flood peak flow data;

[0007] Step 2: Based on the annual maximum peak flow data of each station in the candidate station set, calculate the higher-order linear moments and higher-order linear moment coefficients of each station, and calculate the regional average higher-order linear moment coefficients.

[0008] Step 3: Based on the calculation results in Step 2, calculate the incoherence measure of each station based on higher-order linear moments. If the incoherence measure value of a certain order of a station is greater than the set threshold, the station is considered an incoherent station and is removed.

[0009] Step 4: Based on the calculation results in Step 3 and the remaining sites, the heterogeneity of the region is evaluated using a heterogeneity measure based on higher-order linear moments. If the value of the heterogeneity measure of a certain order is less than a set threshold, the study region is considered to meet the homogeneity requirements at that order. If the values ​​of the heterogeneity measures of all orders are greater than the set threshold, the study region needs to be divided into sub-regions and the heterogeneity of the sub-regions needs to be evaluated.

[0010] Step 5: Based on the calculation results of Step 4, determine the optimal regional distribution function and its optimal linear moment order for regional flood frequency analysis according to the goodness-of-fit test statistic; select different regional distribution functions and calculate the test statistics for multiple regional distribution functions respectively; when a certain test statistic is less than a set threshold, the corresponding regional distribution function is the optimal regional distribution function, and the corresponding order is the optimal linear moment order; if the test statistics for multiple regional distribution functions are all less than the set threshold, then take the regional distribution function with the smallest test statistic value as the optimal regional distribution function.

[0011] Step 6: Based on the optimal regional distribution function and the optimal linear moment order determined in Step 5, the flood design values ​​for stations with and without data in the region under different return periods are derived. For stations with data, the index flood is calculated using the method in Step 2, and then the flood design value is calculated. For stations without data, the index flood is calculated by the relationship between the higher-order linear moments of stations with data and the watershed characteristics in the region, and then the flood design values ​​for stations without data under different return periods are calculated.

[0012] Further, step 1 specifically involves: performing independence and stationarity tests on the collected annual maximum flood peak flow data using both serial autocorrelation and Mann-Kendall tests, respectively; eliminating stations that do not meet the independence and stationarity requirements; and forming a candidate station set, which includes... 1 site, of which .

[0013] Furthermore, step 2 specifically includes:

[0014] Step 2.1: For each station, sort its annual maximum peak flow data in ascending order. Calculate higher-order linear moments The calculation formula is:

[0015] (one)

[0016] (two)

[0017] (three)

[0018] (Four)

[0019] in The combination number is calculated as follows: (five)

[0020] In the formula, This refers to the sequence number of the annual maximum flood peak flow data for a given station, sorted in ascending order. For data length; when When =0, For ordinary linear moments, use Indicates; when When =1,2,3,4, For higher-order linear moments, respectively, use , , , express;

[0021] Step 2.2: Calculate the coefficient of variation of the higher-order linear moments of the samples at each site. skewness coefficient and kurtosis coefficient The calculation formulas are as follows:

[0022] , , (six)

[0023] contain The formulas for calculating the regional average higher-order linear moment coefficients of each station are as follows:

[0024] , , (seven)

[0025] In the formula, For the first Data length for each station; , , The first Each site The coefficient of variation of higher-order linear moments of the sample, the coefficient of skewness, and the coefficient of kurtosis. , and for Average higher-order linear moment coefficients in the first-order region;

[0026] Step 3: Based on the calculation results in Step 2, calculate the incoherence measure for each station based on higher-order linear moments. After removing discordant sites, a site set is obtained for subsequent calculations. This site set contains... One site, The specific process of this step is as follows: Let For the first Sites The higher-order linear moment coefficient vector of the first-order sample, the first-order sample... Disharmony measurement of individual sites The calculation formula is as follows:

[0027] (eight)

[0028] In the formula, =1,2,…, ; This is the region-averaged higher-order linear moment coefficient vector. ; for The inverse matrix, ;

[0029] If the calculated value of a certain station is of a certain order If the value exceeds the set threshold, the site is considered a discordant site; if all sites at all levels... If none of the values ​​exceed the set threshold, all sites are retained, and the number of sites is also marked as [value]. ;

[0030] Step 4: Calculation Regional heterogeneity measures for each site based on higher-order linear moments. Assess regional homogeneity; the specific method for this step is: calculate the order of... Regional heterogeneity measures for values ​​of 1, 2, 3, and 4. :

[0031] (Nine)

[0032] In the formula,

[0033] (ten)

[0034] In the formula, For the first Data length per site and The results were obtained using Monte Carlo simulation experiments. The mean and standard deviation; the specific process is as follows: assuming the region... Each of the sites conforms to a 4-parameter Kappa distribution, and its population parameters are the regional average higher-order linear moment coefficients, i.e. Monte Carlo simulation experiments were used to generate simulated data of the same length as the measured data from the stations. The higher-order linear moment coefficients of the simulated data generated for each station were obtained, and the simulated values ​​were calculated according to equations (vi) and (x). Repeated simulation This time received indivual Value, based on indivual Value calculation area average and standard deviation ;

[0035] In regional flood frequency analysis, the criteria for determining regional heterogeneity are as follows: <1, homogeneous regions are acceptable; 1≤ <2, possibly a homogeneous region; >2. Significantly heterogeneous regions; Select the order range that meets the above criteria based on the calculation results. If the calculation results are all heterogeneous regions, the study area needs to be divided into sub-regions and the heterogeneity of the sub-regions needs to be evaluated.

[0036] Step 5: Based on the calculation results of Step 4, use the goodness-of-fit test statistic. Determine the optimal distribution function and its optimal higher-order linear moment order for regional flood frequency analysis. The specific process of this step is as follows: select multiple candidate region distribution functions and calculate the test statistics for each region distribution function at different orders. :

[0037] (eleven)

[0038] In the formula,

[0039] (twelve)

[0040] In the formula, DIST is the distribution function of the candidate region. This is the kurtosis coefficient of the distribution function for this region; The kurtosis coefficient is the regional average, based on Each station is calculated using equation (vi); and These are simulations obtained from the Kappa distribution. The deviation and standard deviation are simulated in the same way as in step 4. For the number of simulations; For the first The regional average kurtosis coefficient of the second simulation is calculated by equations (vi) and (vii);

[0041] When the simulation statistics satisfy | | When a threshold is set, the distribution function of the specified region is accepted as the optimal distribution function, and the corresponding order is the optimal order. If multiple distribution functions satisfy the condition, the optimal order is usually chosen. The distribution function with the smallest value is taken as the optimal distribution function.

[0042] Furthermore, step 6 specifically involves:

[0043] Construct a regional growth curve (or regional frequency factor) based on the optimal regional distribution. That is, the quantile function (inverse function) of the regional distribution estimated based on the regional average higher-order linear moments, and the frequency is further derived using the index flood method. Time Station Flood design value The calculation formula is:

[0044] (Thirteen)

[0045] In the formula, For the site The flood index; For regional growth curves, When = 0, it is a normal linear moment; otherwise, it is a higher-order linear moment. , This refers to the return period. In actual calculations... This is the optimal order determined in step 5.

[0046] If the site For sites with available data, the indicator is flood. It can be calculated using equations (I) to (IV); if the station For sites without data, establish existing data sites within the region. (s=1,2,…,N2) estimated value and its catchment area The relationship between them is established in the logarithmic field using the least squares method, as shown below:

[0047] (fourteen)

[0048] In the formula, The coefficient is based on the catchment area of ​​stations without data. Formula (14) calculates index floods at data-free stations. .

[0049] Furthermore, it also includes step 7: performance evaluation of the regional frequency analysis method based on higher-order linear moments, using leave-one-out cross-validation to evaluate the robustness of the method in regional flood frequency analysis; this step specifically includes:

[0050] Step 7.1: Based on the region The data from each station is respectively derived from ordinary linear moments. and the optimal order of higher-order linear moments Fit the distribution of the region and calculate the given return period. or non-transcendental probability The following regional growth curve ;

[0051] Step 7.2: Sequentially assign each site... =1,2,…, Considered a site without data Remove the site and use the rest. -1 site recalculated regional growth curve ;

[0052] Step 7.3: For a given recurrence period Calculate the relative error of the region growth curve for each removed site using both the ordinary linear moment and higher-order linear moment methods:

[0053] (fifteen)

[0054] Therefore, for each recurrence period This yields two error vectors based on ordinary linear moments and higher-order linear moments: and ;

[0055] Step 7.4: By comparing the two error vectors and The robustness of a method is evaluated using statistical characteristics: comparing the mean of the absolute errors (Mean(| |) and standard deviation SD(| The smaller the value of |), the lower the average deviation of the method, the less sensitive the method is to removing different sites, and the better its robustness.

[0056] Furthermore, in step 3, The discrimination threshold and the number of stations in the region Related: When the number of sites is 5-10, The threshold values ​​were 1.333, 1.648, 1.917, 2.140, 2.329, and 2.491, respectively; when >10 o'clock, The threshold is 3. Dissonant sites are removed from the original site set, resulting in a set containing... The site set of each site is used for subsequent calculations, where .

[0057] Furthermore, in step 5, the candidate region distribution functions include the Generalized Extreme-value distribution (GEV), the Generalized Pareto distribution (GPA), the Generalized Logistic distribution (GLO), and the Pearson type III distribution (P-III).

[0058] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention provides a regional flood frequency analysis method based on higher-order linear moments. By introducing higher-order linear moments to construct a regional frequency analysis framework, it effectively overcomes the shortcomings of traditional linear moment methods, which are overly sensitive to the lower tail of the distribution (small events) and insufficiently characterize the upper tail (extreme floods), significantly improving the estimation ability and stability of long-return-period design floods. This method can more reliably derive design flood values ​​for stations with no or scarce data, providing a more reliable scientific basis for regional flood risk assessment and flood control engineering design.

[0059] The analytical method of this invention constructs a complete technical process from data screening, regional homogeneity testing, regional distribution optimization to cross-validation, exhibiting strong systematicity and good operability. This method is applicable to various watersheds with a certain length of flow observation data and requiring design flood estimation for data-free stations. In addition to the application results shown in this embodiment, this invention can also be extended to other regions with similar data foundations, demonstrating good practicality and broad applicability. Attached Figure Description

[0060] Figure 1 This is a flowchart of the regional flood frequency analysis method based on higher-order linear moments according to an embodiment of the present invention;

[0061] Figure 2 Quantile estimates for each station based on regional frequency analysis at different return periods;

[0062] Figure 3 A comparison chart of measured and regional simulation results of the annual maximum flood peak flow at Jiaokouhe Station in this embodiment of the invention;

[0063] Figure 4 Different methods in the embodiments of the present invention ( and The mean of the absolute error of the estimated regional growth factor ( Figure 4 (a) and standard deviation ( Figure 4 (b) Curve of variation with return period. Detailed Implementation

[0064] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0065] like Figure 1 As shown, the regional flood frequency analysis method based on higher-order linear moments of the present invention includes the following steps:

[0066] Step 1: Collect data from the study area The annual maximum flood peak discharge data of each hydrological station was collected and preprocessed. This data included maximum flood peak discharge data from multiple years (i.e., data from multiple years were collected for each station). Specifically, the following steps were performed: The collected annual maximum flood peak discharge data were tested for independence and stationarity using both serial autocorrelation and Mann-Kendall tests. Stations that did not meet the independence and stationarity requirements were removed, forming a candidate station set. This station set included… ( (Number of sites)

[0067] Step 2: Calculation The steps include: higher-order linear moments of each site, higher-order linear moment coefficients, and regional average higher-order linear moment coefficients; specifically, this step includes:

[0068] Step 2.1: For each station, sort its annual maximum peak flow data in ascending order. Calculate higher-order linear moments The calculation formula is:

[0069] (one)

[0070] (two)

[0071] (three)

[0072] (Four)

[0073] in The combination number is calculated as follows: (five)

[0074] In the formula, This refers to the sequence number of the annual maximum flood peak flow data for a given station, sorted in ascending order. For data length. When When =0, For ordinary linear moments, use Indicates; when When =1,2,3,4, For higher-order linear moments, respectively, use , , , express.

[0075] Step 2.2: Calculate the coefficient of variation of the higher-order linear moments of the samples at each site. (LH-Cv), skewness coefficient (LH-Cs) and kurtosis coefficient (LH-Ck), the calculation formula is:

[0076] , , (six)

[0077] contain The formula for calculating the regional average higher-order linear moment coefficients of each station is:

[0078] , , (seven)

[0079] In the formula, For the first Data length for each station; , , The first Each site The higher-order linear moments of variation (LH-Cv), skewness coefficient (LH-Cs), and kurtosis coefficient (LH-Ck) of the LH sample. , and for Average higher-order linear moment coefficients in the first-order region.

[0080] Step 3: Based on the calculation results in Step 2, calculate the incoherence measure for each station based on higher-order linear moments. After removing discordant sites, a site set is obtained for subsequent calculations. This site set contains... ( ( ) sites; the specific process of this step is: let For the first Sites The higher-order linear moment coefficient vector of the first-order sample, the first-order sample... Disharmony measurement of individual sites The calculation formula is as follows (there is a disharmony measure for each order). ):

[0081] (eight)

[0082] In the formula, =1,2,…, ; This is the region-averaged higher-order linear moment coefficient vector. ; for The inverse matrix, .

[0083] If the calculated value of a certain station is of a certain order If the value exceeds the set threshold, the site is considered a discordant site. The discrimination threshold is related to the number of stations in the region ( Related: When the number of sites is 5-10, The threshold values ​​were 1.333, 1.648, 1.917, 2.140, 2.329, and 2.491, respectively; when >10 o'clock, The threshold is 3. Dissonant sites are removed from the original site set, resulting in a set containing... ( A set of ) sites is used for subsequent calculations. If all sites at each level... If none of the values ​​exceed the set threshold, all sites are retained, and the number of sites is also marked as [value]. .

[0084] Step 4: Calculation Regional heterogeneity measures for each site based on higher-order linear moments. Assess regional homogeneity; the specific method for this step is: calculate the order of... Regional heterogeneity measures for values ​​of 1, 2, 3, and 4. :

[0085] (Nine)

[0086] In the formula,

[0087] (ten)

[0088] In the formula, For the first Data length per site and The results were obtained using Monte Carlo simulation experiments. The mean and standard deviation. The specific process is as follows: assuming the region... Each of the sites conforms to a 4-parameter Kappa distribution, and its overall parameters are the regional average higher-order linear moment coefficients ( The Monte Carlo simulation experiment was used to generate simulated data of the same length as the measured data from the stations. The higher-order linear moment coefficients of the simulated data generated for each station were obtained, and the simulated values ​​were calculated according to equations (vi) and (x). Repeated simulation This time received indivual Value, based on indivual Value calculation area average and standard deviation .

[0089] In regional flood frequency analysis, the criteria for determining regional heterogeneity are as follows: <1, homogeneous regions are acceptable; 1≤ <2, possibly a homogeneous region; >2. Significantly heterogeneous regions. Select the order range that meets the above criteria based on the calculation results. If all calculation results indicate heterogeneous regions, the study area needs to be divided into sub-regions and their heterogeneity assessed; for each sub-region, regional homogeneity should be assessed. If… If all values ​​are greater than the threshold, the stations in each sub-region need to be adjusted, and the regional homogeneity needs to be reassessed.

[0090] Step 5: Based on the calculation results of Step 4, use the goodness-of-fit test statistic. Determine the optimal distribution function and its optimal higher-order linear moment order for regional flood frequency analysis. The specific process of this step is as follows: Select the candidate region distribution functions: Generalized Extreme-value distribution (GEV), Generalized Pareto distribution (GPA), Generalized Logistic distribution (GLO), and Pearson type III distribution (P-III), and calculate the test statistics for each region distribution function at different orders. :

[0091] (eleven)

[0092] In the formula,

[0093] (twelve)

[0094] In the formula, DIST is the candidate distribution function. LH-Ck is the theoretical kurtosis coefficient of the distribution function in this region, which is the theoretical LH-Ck value calculated by determining the distribution parameters from the regional average higher-order linear moments. LH-Ck is the regional average (based on Each station is calculated using equation (vi). and These are simulations obtained from the Kappa distribution. The deviation and standard deviation (simulation process is the same as step 4); For the number of simulations; For the first The regional average LH-Ck of the sub-simulation is calculated by equations (VI) and (VII).

[0095] When the simulation statistics satisfy | | At a significance level of 5%, the specified distribution is accepted as the optimal distribution function, and the corresponding order is the optimal order. If multiple regional distribution functions satisfy the condition, it is usually chosen as the optimal distribution function. The distribution with the smallest value is taken as the optimal distribution function.

[0096] Step 6: Calculate the flood design values ​​for stations with and without data within the region under different return periods. This step specifically involves:

[0097] Construct a regional growth curve (or regional frequency factor) based on the optimal regional distribution. That is, the quantile function (inverse function) of the regional distribution estimated based on the regional average higher-order linear moments, and the frequency is further derived using the index flood method. Time Station Flood design value The calculation formula is:

[0098] (Thirteen)

[0099] In the formula, For the site The flood index; For regional growth curves, When = 0, it is a normal linear moment; otherwise, it is a higher-order linear moment. , This refers to the return period. In actual calculations... This is the optimal order determined in step 5.

[0100] If the site For sites with available data, the indicator is flood. It can be calculated using equations (I) to (IV); if the station For sites without data, establish existing data sites within the region. (s=1,2,…,N2) estimated value and its catchment area The relationship between them is established in the logarithmic field using the least squares method, as shown below:

[0101] (fourteen)

[0102] In the formula, The coefficient is based on the catchment area of ​​stations without data. Formula (14) calculates index floods at data-free stations. During the calculation, each station was analyzed separately. and By taking the logarithm and estimating the parameters c and d using the least squares method, we can obtain... and The relationship; if the catchment area of ​​the station without data is known to be... Substituting the values ​​into the formula, the station's... .

[0103] Step 7: Performance evaluation of the regional frequency analysis method based on higher-order linear moments. Leave-one-out cross-validation is used to evaluate the robustness of the method in regional flood frequency analysis; this step specifically includes:

[0104] Step 7.1: Based on the region The data for each station are respectively derived from ordinary linear moments (... ) and the optimal order of higher-order linear moments ( Fit the distribution of the region and calculate the given return period. (or non-transcendental probability) Regional growth curves under ) ;

[0105] Step 7.2: Sequentially assign each site... =(1,2,…, (This is considered a site without data) Remove the site and use the rest. -1 site recalculated regional growth curve ;

[0106] Step 7.3: For a given recurrence period Calculate the relative error of the region growth curve for each removed site using both the ordinary linear moment and higher-order linear moment methods:

[0107] (fifteen)

[0108] Therefore, for each recurrence period This yields two error vectors based on ordinary linear moments and higher-order linear moments: and .

[0109] Step 7.4: By comparing the two error vectors and The robustness of a method can be evaluated using statistical characteristics: comparing the mean of the absolute errors (Mean(| |)) and standard deviation (SD(|) The smaller the value of |), the lower the average deviation of the method, the less sensitive the method is to removing different sites, and the better its robustness.

[0110] Example:

[0111] This invention is based on a complete technical process for regional flood frequency analysis based on higher-order linear moments. It selects candidate stations from 11 hydrological stations in northern Shaanxi, performs homogeneity checks on the regions where these candidate stations are located, determines the optimal distribution function for the region based on fitting tests, and finally estimates the quantiles for unmeasured stations under a given return period. The specific implementation steps are as follows:

[0112] Step 1: Annual maximum flood peak flow data were collected from 11 hydrological stations in northern Shaanxi, including Shenmu, Jiaokouhe, Liujiahe, Suide, Huangling, Zaoyuan, Wuqi, and Xinghe. The independence and stationarity of the annual maximum flood peak flow data from these 11 stations were tested using both autocorrelation and Mann-Kendall tests. The test results are shown in Table 1. The Zhaoshiyao station failed the autocorrelation test. (Values ​​greater than the critical value), Zhao Shiyao, Zhangcunyi, and Zhidan stations failed the Mann-Kendal test, and their annual maximum flood peak flow sequences showed a significant trend ("+" indicates an upward trend, "-" indicates a downward trend, and "*" indicates a significant trend at the 5% significance level). After removing these three stations that did not meet the independence and stationarity tests, the remaining... =8 sites constitute the candidate site set.

[0113] .

[0114] Step 2: Calculate the higher-order linear moments and higher-order linear moment coefficients for each site in the candidate site set, and calculate the regional average higher-order linear moment coefficients.

[0115] Step 3: Based on the calculation results in Step 2, calculate the incoherence measure for each station based on higher-order linear moments. The minimum value of all calculated results for different orders is 0.035, and the maximum value is 1.925. Since there are 8 candidate sites in the set, then... The threshold is 2.140.

[0116] Therefore, the dissonance measure values ​​of each station at different orders did not exceed the threshold, meaning there were no dissonant stations. Subsequent calculations were then performed based on these eight stations. = =8.

[0117] Step 4: Based on the calculation results in Step 2, calculate the heterogeneity measure of the region based on higher-order linear moments according to Equation (IX). ( =1,2,3), to evaluate regional homogeneity. As shown in Table 2, when the order... When =0, ( =1,2) is greater than 1, when the order is greater than 1. When =1~4, ( =1,2,3) are all less than 1. In practical applications, This is often used to ultimately determine whether the homogeneity of a region is "acceptable." Therefore, under higher-order linear moments, the region is considered acceptablely homogeneous. Because at higher orders... =1~4 Commonly used values ​​are all less than 1 (which means homogeneity is acceptable), and further selection of the optimal order and optimal distribution function is needed based on the goodness-of-fit test.

[0118] .

[0119] Step 5: Based on the goodness-of-fit test statistic Determine the optimal distribution function and the optimal order of its higher-order linear moments for regional flood frequency analysis. The test statistic at a significance level of 5% is shown in Table 3.

[0120] .

[0121] The goodness-of-fit test statistics for the distribution functions of different orders in each region all satisfy | | 1.64, selected The distribution with the smallest GPA value is considered the optimal distribution, and its order is . =1.

[0122] Step 6: Calculate the flood design values ​​for stations with and without data within the region under different return periods (taking stations with data as an example). Given the return period... For years 50, 100, 200, and 500, the estimated peak flow quantiles for each station were calculated using the higher-order linear moment method based on regional distribution. The results are shown below. Figure 2 In addition, the empirical frequency of the measured peak flow at each station was calculated, and the estimated value of the peak flow at each station was obtained by using the higher-order linear moment method based on the regional distribution. The measured value and the estimated value under the same frequency were then compared. Figure 3 The measured and regional estimation results for Jiaokouhe Station are presented. As shown in the figure, the measured and regional estimation results are in good agreement. Therefore, using the higher-order linear moment method to calculate the regional flood frequency can yield highly accurate results.

[0123] Step 7: Construct a regional growth curve based on the determined regional distribution function, and use leave-one-out cross-validation to evaluate the robustness of the method in regional flood frequency analysis. Figure 4 Demonstrates the method of linear moments ( ) and higher-order linear moment method ( Mean(|) of the absolute error of the estimated regional growth factor |) and standard deviation SD(| |) Comparison. As shown in the graph, at a shorter return period ( <10 years), Mean(| |) and SD(| |) The estimation results are lower than those of the higher-order linear moment method, but when the return period is large ( (>10 years), the estimates of the higher-order linear moments method are lower, indicating that the higher-order linear moments method is more robust in extreme cases.

[0124] Table 4 shows the Mean(|) estimated by the two methods when the return period T = 100 years. |), SD(| |) and Max(| |) value. As shown in Table 4, the higher-order linear moment method outperforms the traditional linear moment method in all error statistics. Among them, Mean(| The SD(|) decreased from 3.4% to 3.0%, indicating that the estimation accuracy of higher-order linear moments is higher than that of the traditional linear moment method. The SD(|) of the higher-order linear moment method... A smaller |) indicates that its estimation results are less dependent on specific sites in the region, and are more robust. Even in the worst-case scenario (Max(|) The maximum error of the higher-order linear moment method is also lower than that of the traditional linear moment method. This result is consistent over long return periods, fully verifying the superiority of the higher-order linear moment method in flood frequency analysis in this study area.

[0125] .

Claims

1. A method for regional flood frequency analysis based on higher order linear moments, characterized in that, Includes the following steps: Step 1: Collect annual maximum flood peak flow data from multiple hydrological stations within the study area, and preprocess the data to form a candidate station set; the annual maximum flood peak flow data includes multi-year maximum flood peak flow data; Step 2: Based on the annual maximum peak flow data of each station in the candidate station set, calculate the higher-order linear moments and higher-order linear moment coefficients of each station, and calculate the regional average higher-order linear moment coefficients. Step 3: Based on the calculation results in Step 2, calculate the incoherence measure of each station based on higher-order linear moments. If the incoherence measure value of a certain order of a station is greater than the set threshold, the station is considered an incoherent station and is removed. Step 4: Based on the calculation results in Step 3 and the remaining sites, the heterogeneity of the region is evaluated using a heterogeneity measure based on higher-order linear moments. If the value of the heterogeneity measure of a certain order is less than a set threshold, the study region is considered to meet the homogeneity requirements at that order. If the values ​​of the heterogeneity measures of all orders are greater than the set threshold, the study region needs to be divided into sub-regions and the heterogeneity of the sub-regions needs to be evaluated. Step 5: Based on the calculation results of Step 4, determine the optimal regional distribution function and its optimal linear moment order for regional flood frequency analysis according to the goodness-of-fit test statistic; select different regional distribution functions and calculate the test statistics for multiple regional distribution functions respectively; when a certain test statistic is less than a set threshold, the corresponding regional distribution function is the optimal regional distribution function, and the corresponding order is the optimal linear moment order; if the test statistics for multiple regional distribution functions are all less than the set threshold, then take the regional distribution function with the smallest test statistic value as the optimal regional distribution function. Step 6: Based on the optimal regional distribution function and the optimal linear moment order determined in Step 5, the flood design values ​​for stations with and without data in the region under different return periods are derived. For stations with data, the index flood is calculated using the method in Step 2, and then the flood design value is calculated. For stations without data, the index flood is calculated by the relationship between the higher-order linear moments of stations with data and the watershed characteristics in the region, and then the flood design values ​​for stations without data under different return periods are calculated.

2. The regional flood frequency analysis method based on higher-order linear moments as described in claim 1, characterized in that, Step 1 specifically involves: performing independence and stationarity tests on the collected annual maximum flood peak flow data using both serial autocorrelation and Mann-Kendall tests, respectively; eliminating stations that do not meet the independence and stationarity requirements; and forming a candidate station set. This station set includes... 1 site, of which .

3. The regional flood frequency analysis method based on higher-order linear moments as described in claim 2, characterized in that, Step 2 specifically includes: Step 2.1: For each station, sort its annual maximum peak flow data in ascending order. Calculate higher-order linear moments The calculation formula is: (one) (two) (three) (Four) in The combination number is calculated as follows: (five) In the formula, This refers to the sequence number of the annual maximum flood peak flow data for a given station, sorted in ascending order. For data length; when When =0, For ordinary linear moments, use Indicates; when When =1,2,3,4, For higher-order linear moments, respectively, use , , , express; Step 2.2: Calculate the coefficient of variation of the higher-order linear moments of the samples at each site. skewness coefficient and kurtosis coefficient The calculation formulas are as follows: , , (six) contain The formulas for calculating the regional average higher-order linear moment coefficients of each station are as follows: , , (seven) In the formula, For the first Data length for each station; , , The first Each site The coefficient of variation of higher-order linear moments of the sample, the coefficient of skewness, and the coefficient of kurtosis. , and for Average higher-order linear moment coefficients in the first-order region; Step 3: Based on the calculation results in Step 2, calculate the incoherence measure for each station based on higher-order linear moments. After removing discordant sites, a site set is obtained for subsequent calculations. This site set contains... One site, The specific process of this step is as follows: Let For the first Sites The higher-order linear moment coefficient vector of the first-order sample, the first-order sample... Disharmony measurement of individual sites The calculation formula is as follows: (eight) In the formula, =1,2,…, ; This is the region-averaged higher-order linear moment coefficient vector. ; for The inverse matrix, ; If the calculated value of a certain station is a certain order If the value exceeds the set threshold, the site is considered a discordant site; if all sites at all levels... If none of the values ​​exceed the set threshold, all sites are retained, and the number of sites is also marked as [value]. ; Step 4: Calculation Regional heterogeneity measures for each site based on higher-order linear moments. Assess regional homogeneity; the specific method for this step is: calculate the order of... Regional heterogeneity measures for values ​​of 1, 2, 3, and 4. : (Nine) In the formula, (ten) In the formula, For the first Data length per site and The results were obtained using Monte Carlo simulation experiments. The mean and standard deviation; the specific process is as follows: assuming the region... Each of the sites conforms to a 4-parameter Kappa distribution, and its population parameters are the regional average higher-order linear moment coefficients, i.e. Monte Carlo simulation experiments were used to generate simulated data of the same length as the measured data from the stations. The higher-order linear moment coefficients of the simulated data generated for each station were obtained, and the simulated values ​​were calculated according to equations (vi) and (x). Repeated simulation This time received indivual Value, based on indivual Value calculation area average and standard deviation ; In regional flood frequency analysis, the criteria for determining regional heterogeneity are as follows: <1, homogeneous regions are acceptable; 1≤ <2, possibly a homogeneous region; >

2. Significantly heterogeneous regions; Select the order range that meets the above criteria based on the calculation results. If the calculation results are all heterogeneous regions, the study area needs to be divided into sub-regions and the heterogeneity of the sub-regions needs to be evaluated. Step 5: Based on the calculation results of Step 4, use the goodness-of-fit test statistic. Determine the optimal distribution function and its optimal higher-order linear moment order for regional flood frequency analysis. The specific process of this step is as follows: select multiple candidate region distribution functions and calculate the test statistics for each region distribution function at different orders. : (eleven) In the formula, (twelve) In the formula, DIST is the distribution function of the candidate region. This is the kurtosis coefficient of the distribution function for this region; The kurtosis coefficient is the regional average, based on Each station is calculated using equation (vi); and These are simulations obtained from the Kappa distribution. The deviation and standard deviation are simulated in the same way as in step 4. For the number of simulations; For the first The regional average kurtosis coefficient of the second simulation is calculated by equations (vi) and (vii); When the simulation statistics satisfy | | When setting a threshold, the specified regional distribution function is accepted as the optimal distribution function, and the corresponding order is the optimal order. If multiple regional distribution functions satisfy the condition, the optimal order is usually chosen. The distribution function of the region with the smallest value is taken as the optimal distribution function.

4. The regional flood frequency analysis method based on higher-order linear moments as described in claim 3, characterized in that, Step 6 specifically involves: Constructing a regional growth curve based on the regional optimal distribution function The frequency was estimated using the index flood method. Time Station Flood design value The calculation formula is: (Thirteen) In the formula, where For the site The flood index; The region growth curve, i.e., the frequency is Regional quantiles at time; , The return period; when calculating The optimal order determined in step 5; If the site For sites with available data, the indicator is flood. It can be calculated using equations (I) to (IV); if the station For sites without data, establish existing data sites within the region. (s=1,2,…,N2) estimated value and its catchment area The relationship between them is established in the logarithmic field using the least squares method, as shown below: (fourteen) In the formula, The coefficient is based on the catchment area of ​​stations without data. Formula (14) calculates index floods at data-free stations. .

5. The regional flood frequency analysis method based on higher-order linear moments as described in claim 4, characterized in that, The method also includes step 7: performance evaluation of the regional frequency analysis method based on higher-order linear moments, using leave-one-out cross-validation to evaluate the robustness of the method in regional flood frequency analysis; this step specifically includes: Step 7.1: Based on the region The data from each station is respectively derived from ordinary linear moments. and the optimal order of higher-order linear moments Fit the distribution of the region and calculate the given return period. or non-transcendental probability The following regional growth curve ; Step 7.2: Sequentially assign each site... =1,2,…, Considered a site without data Remove the site and use the rest. -1 site recalculated regional growth curve ; Step 7.3: For a given recurrence period Calculate the relative error of the region growth curve for each removed site using both the ordinary linear moment and higher-order linear moment methods: (fifteen) Therefore, for each recurrence period This yields two error vectors based on ordinary linear moments and higher-order linear moments: and ; Step 7.4: By comparing the two error vectors and The robustness of a method is evaluated using statistical characteristics: comparing the mean of the absolute errors (Mean(| |) and standard deviation SD(| The smaller the value of |), the lower the average deviation of the method, the less sensitive the method is to removing different sites, and the better its robustness.

6. The regional flood frequency analysis method based on higher-order linear moments as described in claim 3, characterized in that, In step 3, The discrimination threshold and the number of stations in the region Related: When the number of sites is 5-10, The threshold values ​​were 1.333, 1.648, 1.917, 2.140, 2.329, and 2.491, respectively; when >10 o'clock, The threshold is 3. Dissonant sites are removed from the original site set, resulting in a set containing... The site set of each site is used for subsequent calculations, where .

7. The regional flood frequency analysis method based on higher-order linear moments as described in claim 3, characterized in that, In step 5, the candidate region distribution functions include the Generalized Extreme-value distribution, the Generalized Pareto distribution, the Generalized Logistic distribution, and the Pearson type III distribution.