A runoff component model selection method based on benefit-risk balance criterion
Through the runoff component model selection method based on the benefit-risk balance criterion, the problem of balancing accuracy and stability in the runoff component model selection is solved, the scientificity and flexibility of the model selection are achieved, and a more reasonable runoff component model is recommended.
Patent Information
- Application Number
- CN202410240623.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-04
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-03-04
AI Technical Summary
Existing technologies fail to take into account both runoff component identification accuracy and model stability in the selection of runoff component models, and lack effective discrimination methods to provide a decision-making basis for the reasonable selection of runoff component model forms.
A runoff component model selection method based on the benefit-risk balance criterion is adopted. By constructing a set of runoff component models in different forms, calculating the fitting accuracy and stability indicators, and using the benefit-risk balance indicator BR=αB+(1-α)R to select the model, the benefits and risks of the model are balanced.
It provides a method that balances model accuracy and stability, and can flexibly recommend reasonable runoff component models based on decision makers' preferences, thereby improving the scientificity and reliability of model selection.
Smart Images

Figure CN118171970B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of hydrological runoff analysis, in particular to a runoff component model selection method based on a benefit-risk balance criterion. Background Art
[0002] Identifying runoff components is a crucial aspect of hydrological analysis and crucial for understanding the evolution of water resources in a river basin. Hydrology generally assumes that runoff sequences consist of deterministic components such as mutations, trends, and cyclical changes, along with the superposition of stochastic components separated from these deterministic components. In the process of runoff component identification, we first use time series variability diagnostic methods to diagnose evolutionary characteristics such as mutations, trends, and cycles. We then select appropriate runoff component models to quantitatively describe components with different variation characteristics, thereby fully identifying the various deterministic components of the runoff sequence and their patterns of variation.
[0003] Currently, a number of established methods exist for diagnosing and identifying trends, mutations, and periodicity, providing a theoretical foundation for comprehensive analysis of the evolving characteristics of runoff sequences. For the quantitative identification of runoff components, a linear superposition model is often used to sequentially extract and separate various components from a runoff sequence of a given length. The specific model form is then selected based on the goodness-of-fit of the components as a discriminant criterion. However, variations in sequence length can also lead to differences in runoff characteristics and component extraction results. This goodness-of-fit-based discriminant criterion essentially aims to maximize the accuracy of extracting deterministic components from a runoff sequence, but fails to consider the impact of varying runoff length on model accuracy and stability. Currently, a discriminant method that balances runoff component identification accuracy and model stability is lacking, providing a basis for rationally selecting a runoff component model form. Summary of the Invention
[0004] In response to the above-mentioned problems, the present invention provides a runoff component model selection method based on the benefit-risk balance criterion, which is used to balance the model's recognition accuracy of runoff components and its stability risk that changes with sequence length, providing a decision-making basis for the selection of runoff component models.
[0005] The present invention is achieved through the following technical solutions:
[0006] A runoff component model selection method based on the benefit-risk balance principle includes the following steps:
[0007] Step 1: Construct a runoff sequence sample set of varying lengths based on the measured runoff sequence;
[0008] Step 2: Establish a runoff component model consisting of the superposition of mutation component, trend component, periodic component, mean component and residual random component;
[0009] Step 3: Set the mutation component, trend component, periodic component, and mean component in step 2 to different combinations and separation orders to construct different forms of runoff component model sets, and record the model form as M(·);
[0010] Step 4: Select a model form M(·) from the runoff component model set in step 3, and extract the runoff deterministic components from the runoff sequence sample sets of varying lengths in step 1 in sequence until all model forms in the model set are traversed;
[0011] Step 5: Calculate the fitting accuracy of the runoff deterministic components identified by each model form M(·) in step 4 relative to the measured runoff series, and obtain a sample set of fitting accuracy for runoff series of varying lengths;
[0012] Step 6: Calculate the mean and standard deviation of the fitting accuracy sample set with a sample length of L under each model form in step 5, representing the benefit index and risk index of the model for runoff components;
[0013] Step 7: Determine the calculation formula of the benefit-risk balance index based on the benefit index and risk index;
[0014] Step 8: Based on the benefit and risk indicators of each model in step 6, use the calculation formula in step 7 to calculate the benefit-risk balance index of each model in turn, and use the minimum benefit-risk balance index criterion as the decision basis to select the best runoff component model from the runoff component model set constructed in step 3.
[0015] Furthermore, the step 1 specifically includes:
[0016] Suppose there are n samples of measured runoff time series X(t)=x1,x2,...,x n , split the measured sequence into two sequences of length m and nm samples according to the time sequence, and record X1=x1,x2,...,x m ; Take X1 as the initial sequence and gradually increase the length of the sequence, recorded as X i =x1,x2,...,x m+i-1 (i=1~n-m+1), until nm samples are increased to X i , the X i (i=1~n-m+1) is a runoff sequence sample set of varying lengths constructed based on the measured sequence.
[0017] Furthermore, the runoff component model in step 2 is described as follows:
[0018] X(t)=M(t)+T(t)+P(t)+μ+S(t) (1)
[0019] Where M(t), T(t), P(t), μ, and S(t) represent the mutation component, trend component, periodic component, mean component, and residual random component, respectively. The residual random component is the random residual sequence after removing the mean. In formula (1), each component is identified in sequence according to the given order, and the next component is identified after the residual sequence is obtained by removing the previous component.
[0020] Furthermore, the specific steps of extracting the mutation component, trend component, and periodic component in step 2 include:
[0021] Step 2.1: Use the mutation test method to perform mutation diagnosis on the runoff sequence. The original runoff sequence is segmented according to the mutation point. The first segment represents the natural runoff subsequence, and the subsequent segments are variant subsequences. The variant subsequences are restored to the same mean level as the natural runoff through mean transformation to eliminate the mean variation component. The subsequences of the original sequence segmented by K mutation points are X0, X1, ..., X K , the mutation component is expressed as follows:
[0022]
[0023] Where, is the mean of the natural subsequence, is the mean of the k-th variant sequence;
[0024] Step 2.2: Fit the runoff series using a univariate linear regression equation. The slope and intercept parameters are solved using the least squares regression method. The first-order term in the fitting equation is used as the trend component. The expression is as follows:
[0025]
[0026] In the formula, x is the dependent variable representing time, represents the fitted value of the trend component, and m is the slope parameter;
[0027] Step 2.3: Use the periodogram method to identify the significant periodic components of the runoff series. The significant periodic components are identified through the F test. If there are d significant harmonics, the periodic component expression is as follows:
[0028]
[0029] Where, It represents the fitting value of the periodic component, that is, the main periodic component of the harmonic accumulation, the main period T j =2π / w j .
[0030] Furthermore, the mutation point diagnosis in step 2.1 uses five methods, namely the MK test, sliding T test, Pettitt test, standard normal test, and Buishand test, to conduct preliminary mutation tests on the runoff series. Based on the "voting method", the years that are tested as significant mutation points by more than two methods are taken as preliminary identification results. The rank sum test method is further used to diagnose the final mutation point from the preliminary results.
[0031] Furthermore, the fitting accuracy in step 5 is calculated using the Pearson correlation coefficient, root mean square error, mean absolute error, mean absolute percentage error, and Nash coefficient.
[0032] Furthermore, in step 6, the mean and standard deviation of the fitting accuracy sample set with a sample length of L under each model form in step 5 are calculated respectively, specifically including:
[0033]
[0034] Where B is the average fitting accuracy of the model for the changing runoff sequence, representing the "benefit" of the model in identifying runoff components; R is the fluctuation of the model fitting accuracy under the changing runoff sequence, representing the stability of the model in responding to the changing runoff sequence, which is called "risk".
[0035] Furthermore, the calculation formula for the benefit-risk balance index in step 7 is:
[0036] BR=αB+(1-α)R (6)
[0037] In the formula, BR is the benefit-risk balance indicator, and α is the weight coefficient between 0 and 1;
[0038] If B is the average precision index based on the maximization target, B is converted into a minimization target consistent with the risk index R; different weight coefficients α are selected according to the decision maker's preference; BR indicators with different weights are calculated as the decision basis for model selection.
[0039] The present invention has the following advantages:
[0040] (1) In steps 1 and 3, the present invention considers the effects of the runoff sequence length and the separation order of different types of components on the runoff characteristics and component identification results. In step 7, a benefit-risk balance criterion is defined based on the deterministic component fitting accuracy of the runoff sequence with varying lengths, providing a decision-making basis for the reasonable selection of the runoff component model form by balancing the model accuracy and model stability.
[0041] (2) The proposed model selection method based on the benefit-risk balance criterion can balance the decision maker's preference for benefits and risks and flexibly give reasonable model recommendations. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 This is a flow chart of a runoff component model selection method based on the benefit-risk balance criterion proposed in an embodiment of the present invention;
[0043] Figure 2 This is a diagram showing the change process of the annual runoff sequence according to an embodiment of the present invention;
[0044] Figure 3 This is a comparison chart of the results of identifying sudden changes in runoff sequences of varying lengths under different component model forms according to an embodiment of the present invention;
[0045] Figure 4 A comparison chart of trend component identification results for runoff series of varying lengths under different component model forms according to an embodiment of the present invention;
[0046] Figure 5 This is a comparison chart of periodic component identification results for runoff sequences of varying lengths under different component model forms according to an embodiment of the present invention;
[0047] Figure 6 This is a comparison chart of the fitting accuracy of different component models for runoff sequences of varying lengths according to an embodiment of the present invention. DETAILED DESCRIPTION
[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0049] like Figure 1 As shown, the present embodiment provides a runoff component model selection method based on the benefit-risk balance criterion, comprising the following steps:
[0050] Step 1: Construct a runoff sequence sample set of varying lengths based on the measured runoff sequence. This example collects the annual runoff sequence of Pingshan Station from 1956 to 2010, such as Figure 2 As shown in Figure 2, the sequence is divided into segments 1 and 2 of length 30 and 25, respectively. The samples of segment 2 are gradually added to segment 1 to construct a sample set of runoff series with varying lengths (number of samples: 30 to 55) starting from 1956 and ending from 1986 to 2010.
[0051] Step 2: Establish a runoff component model consisting of the superposition of deterministic components such as mutation components, trend components, periodic components, mean components, and the remaining random components. The runoff component model is described as follows:
[0052] X(t)=M(t)+T(t)+P(t)+μ+S(t) (1)
[0053] Where M(t), T(t), P(t), μ, and S(t) represent the mutation component, trend component, periodic component, mean component (i.e., the mean of the residual sequence), and residual random component, respectively. The residual random component is the random residual sequence after removing the mean, i.e., the center.
[0054] The specific steps for extracting mutation components, trend components, and periodic components include:
[0055] Step 2.1: Five methods, including the MK test, sliding T test, Pettitt test, standard normality test (SNHT), and Buishand test, were used to test the runoff sequence for mutations. The years that were detected as significant mutation points by more than two methods were used as preliminary identification results based on the “voting method”. The rank sum test was then used to diagnose the final mutation point from the preliminary identification results. The original runoff sequence was segmented according to the final mutation point, and the subsequences of the original sequence segmented by K mutation points were recorded as X0, X1, …, X K , the mutation component is expressed as follows:
[0056]
[0057] Where, is the mean of the natural subsequence, is the mean of the k-th variant sequence;
[0058] Step 2.2: Fit the runoff series using a univariate linear regression equation. The slope and intercept parameters are solved using the least squares regression method. The first-order term in the fitting equation is used as the trend component. The expression is as follows:
[0059]
[0060] In the formula, x is the dependent variable representing time, represents the fitted value of the trend component, and m is the slope parameter;
[0061] Step 2.3: Use the periodogram method to identify the significant periodic components of the runoff series. The significant periodic components are identified through the F test. If there are d significant harmonics, the periodic component expression is as follows:
[0062]
[0063] Where, It represents the fitting value of the periodic component, that is, the main periodic component of the harmonic accumulation, the main period T j =2π / w j .
[0064] The mean component is the original sequence after removing the mutation component, trend component, and periodic component, and then calculating the mean of the remaining components. Further, after removing the mean from the remaining components, the remaining random component is obtained.
[0065] Step 3: Set the mutation component, trend component, periodic component, and mean component in step 2 to different combinations and separation orders to construct different forms of runoff component model sets, and denote the model form as M(·).
[0066] This example considers the separation order of sudden and trend components, as well as the impact of different combinations of sudden and trend components on periodic components. It designs runoff component identification schemes with different component combinations and separation orders, and constructs different runoff component models, as shown in Table 1. Each component model in Table 1 sequentially extracts and separates the current runoff component, and then uses the remaining sequence to identify the next component.
[0067] Table 1 Runoff component models constructed based on different component identification schemes
[0068]
[0069] The symbols in the brackets of the model name M represent different component categories. Mutation, trend, period, and mean represent mutation, trend, cycle, and mean components, respectively. Mut and Tre are simplified symbols for mutation and trend components.
[0070] Step 4: Select a model form M(·) from the runoff component model set in step 3, and perform the runoff sequence sample set X of varying lengths in step 1. i (i=1~n-m+1) Extract the deterministic components of runoff in turn until all 10 model forms in the model set are traversed. As shown in Table 1, in this embodiment, there are 2 schemes for identifying mutation components, 2 schemes for identifying trend components, and 5 schemes for identifying periodic components. The identification results of mutation, trend, and periodic components under different identification schemes are given respectively, as shown in Table 1. Figure 3 、 Figure 4 and Figure 5 shown.
[0071] Figures 3 to 5 The results show that the mutation, trend, and periodicity characteristics of runoff series are significantly affected by changes in runoff series length and the form of runoff component models (different combinations and separation orders). For example, removing the trend component weakens the rising mutation characteristics of the measured series; after removing the mutation component, the trend component of the remaining series tends to change more slowly after the mutation point (1998); removing the mutation component weakens the periodic fluctuation to a certain extent, and the periodicity is more affected by the mutation component than the trend component.
[0072] Step 5: Calculate the fitting accuracy of the deterministic components of runoff identified by each model form M(·) in Step 4 relative to the measured runoff series. In this embodiment, the Pearson Correlation Coefficient (PC) and Root Mean Square Error (RMSE) are used to calculate the fitting accuracy of the "deterministic components" identified by the model relative to the measured runoff series.
[0073] For the 10 models in this embodiment, the fitting accuracy sample sets of variable length runoff series can be obtained and recorded as V i (i=1~L), L=n-m+1, where the sample length is L=26. The fitting accuracy index PC and RMSE (m 3 / s) with the sample end year (1 9 The change process from 1985 to 2010 is as follows Figure 6 The results show that the fitting accuracy indicators of different models are different, and the PC and RMSE of different models fluctuate to varying degrees with the sequence length.
[0074] Step 6: Calculate the mean and standard deviation of the fitting accuracy sample set with a sample length of L under each model form in step 5, representing the benefit index and risk index of the model for runoff components. The calculation formula is as follows:
[0075]
[0076] Where B is the average fitting accuracy of the model for the changing runoff sequence, representing the "benefit" of the model in identifying runoff components; R is the fluctuation of the model fitting accuracy under the changing runoff sequence, representing the stability of the model in responding to the changing runoff sequence, which is called "risk".
[0077] Step 7: In order to balance the model's fitting accuracy (benefit) and stability (risk), the model with higher accuracy (smaller fitting error) and the smallest possible stability risk is selected. The benefit-risk balance index is defined as:
[0078] BR=αB+(1-α)R (6)
[0079] Where BR is the benefit-risk balance indicator, and α is the weight coefficient between 0 and 1.
[0080] In this embodiment, BR is calculated by setting five groups of weight coefficients α to 0, 0.2, 0.5, 0.8 and 1; the B index based on the precision index PC is minimized (the smaller the better) by taking a negative value and then calculating the BR index.
[0081] Step 8: Based on the benefit and risk indicators of each model in step 6, use formula (6) in step 7 to calculate the benefit-risk balance index of each model in turn. The comparison of BR indicators under the five weight coefficients is shown in Table 2. The bold indicators in Table 2 represent the optimal model under the current weight coefficient α.
[0082] Based on the "minimum benefit-risk balance index" criterion in step 7, the optimal runoff component model was selected from the 10 models constructed in step 2. The results showed that different precision index types and weight settings could lead to different model optimization results: when the risk index was minimized (α = 0), the PC recommended model was M(trend), and the RMSE recommended models were M(mutation) and M(mut_trend); when α = 0.2, the PC recommended models were M(mutation) and M(period), and the RMSE recommended model was M(mut_trend); when α = 0.5-1, the PC and RMSE recommended models were M(mut_period) and M(mut_tre_period), respectively.
[0083] Table 2 Comparison of model benefit-risk balance indicators based on different weight coefficients α
[0084]
[0085]
[0086] As can be seen, if the decision goal is to achieve a more complete fit, the model that sequentially separates the mutation and cyclical terms is preferred; if the decision goal is to achieve more stable accuracy over time, the model that only identifies mutations and / or trends is superior. This example demonstrates the superiority of the proposed model selection method based on the benefit-risk balance criterion. This method can balance the decision maker's preferences for benefits and risks and flexibly provide reasonable model recommendations.
[0087] The above description is merely a specific embodiment of a specific example of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A runoff component model selection method based on the benefit-risk balance criterion, characterized in that: The steps include: Step 1: Construct a runoff sequence sample set of varying lengths based on the measured runoff sequence; Step 2: Establish a runoff component model consisting of the superposition of mutation component, trend component, periodic component, mean component and residual random component; Step 3: Set the mutation component, trend component, periodic component, and mean component in step 2 to different combinations and separation orders, and construct different forms of runoff component model sets. The model form is recorded as ; Step 4: Select a model form from the runoff component model set in step 3 , extract the runoff deterministic components from the runoff sequence sample sets of varying lengths in step 1 in sequence until all the model forms in the model set are traversed; Step 5: Calculate the model forms in step 4 The fitting accuracy of the identified deterministic components of runoff relative to the measured runoff series is used to obtain a sample set of fitting accuracy for runoff series of varying lengths; Step 6: Calculate the mean and standard deviation of the fitting accuracy sample set with a sample length of L under each model form in step 5, representing the benefit index and risk index of the model for runoff components; Step 7: Determine the calculation formula of the benefit-risk balance index based on the benefit index and risk index; Step 8: Based on the benefit index and risk index of each model in step 6, the calculation formula in step 7 is used to calculate the benefit-risk balance index of each model in turn. The optimal runoff component model is selected from the runoff component model set constructed in step 3 based on the minimum benefit-risk balance index criterion; The runoff component model in step 2 is described as: (1); Where, They represent mutation component, trend component, cycle component, mean component, and residual random component respectively. The residual random component is the random residual sequence after removing the mean. In formula (1), each component is identified in the given order. After removing the previous component to obtain the residual sequence, the next component is identified. The specific steps for extracting mutation components, trend components, and periodic components in step 2 include: Step 2.1: Use the mutation test method to perform mutation diagnosis on the runoff sequence. The original runoff sequence is segmented according to the mutation point. The first segment represents the natural runoff subsequence, and the subsequent segments are variant subsequences. The variant subsequences are restored to the same mean level as the natural runoff through mean transformation to eliminate the mean variation component. The subsequences of the original sequence segmented by K mutation points are recorded as , the mutation component is expressed as follows: (2); Where, is the mean of the natural subsequence, is the mean of the k-th variant sequence; Step 2.2: Fit the runoff series using a univariate linear regression equation. The slope and intercept parameters are solved using the least squares regression method. The first-order term in the fitting equation is used as the trend component. The expression is as follows: (3); In the formula, x is the dependent variable representing time, represents the fitted value of the trend component, is the slope parameter; Step 2.3: Use the periodogram method to identify the significant periodic components of the runoff series. The significant periodic components are identified through the F test. If there are d significant harmonics, the periodic component expression is as follows: (4); Where, Represents the fitting value of the periodic component, that is, the main periodic component of the harmonic accumulation, the main period ; The calculation formula for the benefit-risk balance index in step 7 is: (6); Where B is the average fitting accuracy of the model for the changing runoff sequence, representing the "benefit" of the model in identifying runoff components; R is the fluctuation of the model fitting accuracy under the changing runoff sequence, representing the stability of the model in dealing with the changing runoff sequence, also known as "risk"; BR is the benefit-risk balance index. is a weight coefficient between 0 and 1; If B is the average accuracy index based on the maximization target, transform B into a minimization target consistent with the risk index R; select different weight coefficients according to the decision maker's preference ; Calculate BR indicators with different weights as the decision basis for model selection.
2. The runoff component model selection method based on the benefit-risk balance criterion according to claim 1, characterized in that: The step 1 specifically includes: Suppose there are n samples of measured runoff time series , divide the measured sequence into two sequences with lengths of m and nm samples according to the time sequence, record ;by is the initial sequence, and the length of the sequence is gradually increased, recorded as , until nm samples are increased to , is a sample set of runoff sequences of varying lengths constructed based on the measured sequences; .
3. The runoff component model selection method based on the benefit-risk balance criterion according to claim 1, characterized in that: In step 2.1, the mutation point diagnosis was performed using five methods: the MK test, the sliding T test, the Pettitt test, the standard normal test, and the Buishand test. Based on the "voting method," the years that were identified as significant mutation points by more than two methods were considered preliminary identification results. The rank sum test was then used to diagnose the final mutation point from the preliminary results.
4. The runoff component model selection method based on the benefit-risk balance criterion according to claim 1, characterized in that: The fitting accuracy in step 5 is calculated using the Pearson correlation coefficient, root mean square error, mean absolute error, mean absolute percentage error, and Nash coefficient.
5. The method for selecting a runoff component model based on the benefit-risk balance criterion according to claim 1, wherein: In step 6, the mean and standard deviation of the fitting accuracy sample set with a sample length of L under each model form in step 5 are calculated respectively, including: (5); Where B is the average fitting accuracy of the model for the changing runoff sequence, representing the "benefit" of the model in identifying runoff components; R is the fluctuation of the model fitting accuracy under the changing runoff sequence, representing the stability of the model in responding to the changing runoff sequence, also known as "risk".
Citation Information
Patent Citations
Runoff forecast sample set division method based on misjudgment risk minimum criterion
CN117131977A
Runoff forecasting method based on baseflow separation and artificial neural network model
WO2022110582A1