A medium- and long-term runoff time-varying probability prediction method based on the optimal combination of all factors of "input-structure-parameters"
By optimizing predictor factors, model structure, and parameters, employing copula entropy to screen driving factors, LSTM to generate single-value prediction schemes, combining information entropy and BMA to generate probabilistic prediction schemes, and utilizing GARCH models for error correction, the time-varying characteristics of medium- and long-term runoff prediction models were resolved, thereby improving the accuracy and reliability of predictions.
Patent Information
- Application Number
- CN202211206170.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-09-30
AI Technical Summary
Existing medium- and long-term runoff prediction models are unable to accurately reflect the time-varying characteristics of runoff processes, resulting in large prediction errors and high uncertainty, and failing to provide accurate and reliable information to support water resources planning and management.
The method of hierarchical combination optimization of all elements of "input-structure-parameter" is adopted. By optimizing the predictor, model structure and model parameters, the driving factors are screened by copula entropy, LSTM generates single-value prediction schemes, information entropy and BMA are combined to generate probabilistic prediction schemes, and GARCH model is used for error correction to generate runoff probabilistic prediction schemes that take error correction into account.
It improves the accuracy, reliability, and stability of runoff forecasting, providing effective decision-making information for water resource management.
Smart Images

Figure CN115496290B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to medium- and long-term hydrological forecasting technology in the field of water conservancy engineering, and in particular to a medium- and long-term runoff time-varying probability prediction method that optimizes the hierarchical combination of all elements of "input-structure-parameter". Background Technology
[0002] Medium- and long-term runoff forecasting refers to hydrological runoff forecasting with a forecast period exceeding months or seasons. It serves as the fundamental basis for water conservancy project scheduling plans and provides essential information support for optimal water resource allocation. Unlike short-duration flood forecasting, long-duration runoff has complex causes and is difficult to predict accurately, leading to large errors and high uncertainties in medium- and long-term runoff forecasts. Currently, single-value medium- and long-term runoff forecasting schemes still struggle to provide accurate and reliable information for water resource planning and management. Since water resource systems can buffer and tolerate runoff forecast errors through reservoir runoff regulation, under conditions where accurate single-value forecast results are difficult to obtain, identifying the variation characteristics of single-value forecasting scheme errors, integrating single-value forecasts and error simulations to construct probabilistic forecasting schemes, and conducting quantitative forecasts of runoff intervals and their probabilities can improve the supporting role of forecast information in scheduling.
[0003] The key indicators determining the quality of probabilistic prediction schemes mainly include three categories: accuracy, reliability, and stability. Accuracy primarily assesses the deviation between the single-value prediction scheme and the measured value; reliability primarily assesses the likelihood that the probabilistic prediction scheme will cover the measured value under normal conditions; and stability primarily assesses the potential deviation when the prediction interval cannot effectively cover the measured value under extraordinary (extreme) conditions. Pursuing accurate, reliable, and stable probabilistic prediction schemes is the ultimate goal of runoff prediction. The prediction model factors affecting these three indicators mainly include predictor factors (independent variables), prediction equations (model structure), and prediction errors: predictor factors and prediction equations reflect the deterministic physical processes of runoff formation; the former represents the key elements of the hydrological cycle affecting runoff formation, while the latter describes the response relationship between these elements and runoff through the construction of mathematical models. Prediction errors are a comprehensive reflection of the uncertainties in runoff prediction and the hydrological cycle, and are generally measured by the deviation between predicted runoff and actual runoff. Selecting different predictor factors and equations affects both the results of single-value predictions and the distribution characteristics of errors, ultimately producing different prediction effects. Therefore, the overall prediction effect can be changed by selecting predictor factors, changing the model structure, and optimizing model parameters.
[0004] Existing research has explored these three aspects: in terms of factor selection, meteorological and hydrological factors have been categorized and screened; in terms of model structure, a series of data-driven modeling methods, such as time series analysis, statistical regression, and neural networks, have been developed; and in terms of parameter setting, the stationarity of model parameters is usually assumed, and parameter calibration is performed by minimizing the sum of squared deviations. However, statically stationary prediction models are increasingly unable to objectively reflect runoff processes, especially the non-stationary and time-varying characteristics of the stochastic term. For example, the response of runoff processes is affected by the temporal and spatial continuity and complexity of meteorological and hydrological factors during runoff generation and confluence, exhibiting characteristics such as spatiotemporal correlation, peaks, and thick tails. Traditional stationary time-invariant prediction models face significant challenges. Therefore, how to construct accurate, reliable, and stable time-varying runoff prediction models, considering the time-varying characteristics of errors and uncertainties in medium- and long-term runoff forecasting, is an urgent scientific problem to be solved. Summary of the Invention
[0005] Purpose of the invention: The purpose of this invention is to provide a medium- to long-term runoff time-varying probability prediction method that optimizes the three elements of "input-structure-parameter" hierarchical combination. By optimizing the three elements of prediction factors, model structure, and model parameters respectively, the accuracy, reliability, and stability of runoff prediction are improved, providing effective and rich decision-making information for the efficient utilization of water resources.
[0006] Technical solution: The present invention provides a method for predicting the time-varying probability of medium- and long-term runoff through a hierarchical combination of all elements, including "input-structure-parameters," comprising the following steps:
[0007] S1. As the first layer of the hierarchical combination optimization of all elements of "input-structure-parameter", the model input is optimized: long-term meteorological and hydrological factor data including atmospheric circulation factors, precipitation, evaporation and runoff are collected in the study area to construct the initial factor set for runoff prediction. Based on copula entropy and runoff cause analysis, the combination of prediction factors is screened to obtain the driving factor set. The long short-term memory neural network LSTM model is used to generate a set of medium- and long-term runoff single-value prediction schemes.
[0008] S2. As the second layer of the hierarchical combination optimization of all elements of "input-structure-parameter", the model structure is optimized: information entropy is used as the indicator to evaluate the information value of the single-value prediction scheme set. Stepwise regression with inner and outer nested loop structure is used to optimize the combination of schemes in the single-value prediction scheme set. The optimal subset is generated by the Bayesian model average BMA method to generate a runoff set prediction scheme that simultaneously includes single-value process and probability interval.
[0009] S3, as the third layer of the hierarchical combination optimization of all elements of "input-structure-parameters", optimizes the model parameters: Based on the generalized autoregressive conditional heteroscedasticity model GARCH, the runoff prediction error is finely characterized; real-time error correction is performed using the dependency relationship of the error; Monte Carlo sampling is combined to generate a corrected error scenario sequence, which is then superimposed onto the single-valued process in step S2 to obtain a runoff time-varying probability prediction scheme considering error correction. The obtained probability prediction scheme is evaluated according to the multi-dimensional evaluation index system of "accuracy-reliability-stability" for runoff prediction. Further, step S1 includes the following steps:
[0010] S11, Generation of runoff prediction factor set;
[0011] Based on the meteorological conditions of the study area, the main weather elements controlling rainfall and runoff in the study area are identified, and the corresponding indicators are selected together with previous rainfall, evapotranspiration and runoff to form a runoff prediction factor set.
[0012] S12. Identification of driving factors based on copula entropy and runoff causation analysis;
[0013] Let N-dimensional random variables (X1, X2, ..., Xn) be defined. N ) has marginal distribution u i =F i (x i ), i = 1, 2, ..., N, where F is the cumulative probability density function, and the copula entropy is defined as:
[0014]
[0015] Where CE is the copula entropy and C is the copula function. Based on the copula entropy, all predictive factors related to runoff causes are ranked by correlation, and the top-ranked factors are selected as the driving factor set.
[0016] Generation of single-value runoff prediction schemes for S13 and LSTM models;
[0017] LSTM was used to predict monthly runoff into the lake. Different combinations of factors were used as dependent variables for learning, training, simulation and prediction.
[0018] Furthermore, the LSTM calculation method in step S13 is as follows:
[0019] (a) The Gate of Oblivion:
[0020] f t =σ(W f ·[h t-1 ,x t ]+b f )
[0021] Among them, f t Forget gate control of the current time period, it means that the historical output information of the control part is converted into the input of the current time period, reflecting the degree of loss of historical information in the current time period, W f Let b be the weight matrix of the forget gate. f σ is the bias term for the forget gate, and σ is the activation function;
[0022] (b) Input gate:
[0023] i t =σ(W i ·[h t-1 ,x t ]+b i )
[0024] Among them, i t This refers to input gate gating for the current time period, indicating that the input of the control part for the current time period is transformed into system state variables, reflecting the information input efficiency for the current time period. W i Let b be the weight matrix of the input gate. i This is the bias term for the input gate;
[0025] (c) The current long-term state c t ':
[0026] c t '=tanh(W c ·[h t-1 ,x t ]+b c )
[0027] Among them, W c Let b be the weight matrix of the current input long-term state. c For W c The bias term, tanh is the activation function;
[0028] (d) The long-term state at the current moment c t :
[0029] c t =f t ·c t-1 +i t ·c t '
[0030] (e) Output gate:
[0031] o t =σ(W o ·[h t-1 ,x t ]+b o )
[0032] Among them, ot This is the output gate for the current time period, indicating that the state variables of the control part of the system are transformed into the output for the current time period, reflecting the information output efficiency for the current time period. W o Let b be the weight matrix of the output gate. o For W o The bias term;
[0033] (f) Output value of LSTM:
[0034] h t =o t ·tanh(c t )
[0035] The calculation period is divided into a training period and a verification period. The set of driving factors W identified in step S12 during the training period is used as the input term xt in the above calculation steps, and the runoff is used as the output term ht. After training and calibration of the neural network, the parameters Wf, bf, Wi, bi, Wo, and bo are obtained. These parameters, along with the input term xt from the verification period, are substituted into the model to calculate the corresponding output term ht, resulting in a set of single-value prediction processes, which constitute a set of single-value prediction schemes.
[0036] Furthermore, step S2 includes the following steps:
[0037] S21. Optimal selection of solution set combinations based on information entropy and stepwise regression;
[0038] Information entropy is a statistic that measures the uncertainty of an information system, and its expression is:
[0039]
[0040] Where H is the information entropy, and X is a random variable with N discrete values, X = {x1, x2, ..., x...} N}, its probability distribution Furthermore, for a binary random variable (X,Y), its information entropy is also called joint entropy, defined as:
[0041]
[0042] Where p(x,y) is the joint probability density function of random variables X and Y; joint entropy represents the total amount of information contained in multiple random variables, and their equal amounts of information are defined as mutual information, i.e.:
[0043] T(X,Y)=H(X)+H(Y)-H(X,Y)
[0044] Therefore, joint entropy and mutual information in information entropy are used to describe the total amount of information and information content among candidate schemes in the single-value prediction scheme set, respectively, and this is used as the basis for the optimal combination of schemes.
[0045] S22. Generation of runoff single-value-probability ensemble prediction scheme based on BMA;
[0046] Let y be the predictor variable, y obs Let f be the sequence of measured samples, f = {f1, f2, ..., f...} k} represents the candidate model space, where K denotes the number of candidate scheme groups within the selected single-value prediction scheme combination in step S21; based on Bayesian theory, assuming that both the measured and predicted sequences follow a normal distribution, further refinement is made as follows:
[0047]
[0048] Where p(y|y obs ω represents the probability distribution of the predictor variable. i This represents the weight of the i-th model. The mean is f i The variance is The normal distribution;
[0049] Therefore, the BMA parameter ω is calibrated using the expectation-maximization algorithm. i and Monte Carlo sampling is used to give a probability prediction scenario sequence, and a runoff set prediction scheme containing single-valued processes and probability intervals is obtained.
[0050] Furthermore, step S22 uses Monte Carlo sampling to provide a probability prediction scenario sequence, the specific steps of which are as follows:
[0051] (a) The non-normally distributed runoff sequence is transformed into a normally distributed one using the BOX-COX transformation. The specific steps are as follows:
[0052]
[0053] Where λ is the transformation coefficient, the value of which is given by the maximum likelihood estimation; y' is the transformed runoff prediction sequence that follows a normal distribution;
[0054] (b) Calibrate BMA parameters ω using the EM algorithm i and
[0055] (c) Based on the BMA weights [ω1,ω2,...,ω] of each candidate scheme K Define the cumulative weight value.
[0056]
[0057] Generate a uniform random number u in the range [0,1]. This indicates that the i-th candidate solution f was selected in this sampling. i ;
[0058] (d) According to f i Probability distribution in time period j Randomly generate a runoff prediction sequence y' that follows a normal distribution;
[0059] (e) The generated runoff prediction sequence y', which follows a normal distribution, is restored based on the cumulative weight value formula. The calculation formula is as follows:
[0060]
[0061] (f) Repeat steps (c)-(e) M times to obtain a set of prediction scenarios with M scenarios, which is the runoff probability prediction scheme of BMA. The mean process is the runoff single-value prediction scheme of BMA. The two together constitute the runoff single-value-probability set prediction scheme.
[0062] Furthermore, step S3 includes the following steps:
[0063] S31. Fine characterization of runoff prediction error based on GARCH model;
[0064] Assume residual term ε t With conditional variance h t obey:
[0065]
[0066] Where c is a constant term, a i and b j The conditional variances h of lags i and j are respectively. t-i and residual term ε t-j The coefficients, m and n, are the lag orders of the GARCH conditional variance and residual terms, respectively, denoted as ε. t ~GARCH(m,n); maximum likelihood estimation is used for parameter a i b j Calibration was conducted;
[0067] S32. Runoff prediction error correction based on Monte Carlo sampling;
[0068] Based on the GARCH model and its parameters obtained in step S31, and combined with Monte Carlo sampling to generate a runoff prediction error simulation scenario sequence, for each time t, N sets of random numbers v following a standard normal distribution are first randomly generated using the Monte Carlo method. t Substituting into the formula in step S31, the conditional variance h is calculated. tThe simulated scenario sequence is obtained by taking the square root of the conditional variance term to obtain the error simulated scenario sequence, which is the correction for runoff prediction error;
[0069] S33. Generation of time-varying runoff probability prediction scheme considering error correction;
[0070] The relationship between prediction error and measured value is as follows:
[0071]
[0072] Among them, y t This is the predicted value for time period t. Let e be the measured true value for time period t. t This represents the prediction error for that period; if error correction is considered, i.e., when e t '=z(e t When ), we have:
[0073]
[0074]
[0075] Among them, e t ' represents the corrected runoff prediction error, and z is the error correction function, which reflects the statistical characteristics of historical prediction errors; This is the "theoretical value" of runoff, that is, a precise prediction of the measured runoff value after considering error correction; combined with the runoff prediction error simulation scenario sequence generated in step S32, i.e., e t ', and superimpose it on the runoff single-value prediction process of BMA in step S22 to obtain the runoff probability prediction scenario sequence, that is, the runoff probability time-varying prediction scheme considering error correction;
[0076] S34. Evaluation of runoff prediction accuracy;
[0077] A multi-dimensional evaluation index of "accuracy, reliability, and stability" for runoff prediction was constructed. The root mean square error (RMSE), Brier score (BS), and misalignment deviation (MD) were selected as indicators to evaluate the accuracy, reliability, and stability of runoff probability prediction, and to measure the merits of the prediction schemes. A comprehensive and systematic evaluation of the prediction schemes was conducted.
[0078] The present invention provides a medium- to long-term runoff time-varying probability prediction system based on a hierarchical combination of all elements including input, structure, and parameters, comprising:
[0079] The model input optimization module is used to collect long-term meteorological and hydrological factor data of the study area to construct the initial factor set for runoff prediction. Based on copula entropy and runoff causation analysis, the combination of prediction factors is screened to obtain the driving factor set. The LSTM based on deep learning is used to generate a set of medium- and long-term runoff single-value prediction schemes, which serves as the first layer of the hierarchical combination optimization of all elements of "input-structure-parameter".
[0080] The model structure optimization module uses information entropy as an indicator to evaluate the information value of the single-value prediction scheme set. It adopts stepwise regression with inner and outer nested loop structures to optimize the combination of schemes in the scheme set. The optimal subset is generated by BMA to generate a runoff set prediction scheme that simultaneously includes single-value processes and probability intervals, as the second layer of the hierarchical combination optimization of all elements of "input-structure-parameter".
[0081] The model parameter optimization module is used to construct a time-varying simulation model of runoff ensemble prediction error based on GARCH, perform real-time error correction using the dependency relationship of errors, generate a corrected error scenario sequence by combining Monte Carlo sampling, and superimpose it on the Bayesian model average single-value forecast result to obtain a runoff probability prediction scheme considering error correction. The obtained probability prediction scheme is evaluated according to the multi-dimensional evaluation index system of runoff prediction "accuracy-reliability-stability", which serves as the third layer of the hierarchical combination optimization of all elements of "input-structure-parameter".
[0082] An apparatus of the present invention includes a memory and a processor, wherein:
[0083] Memory is used to store computer programs that can run on a processor;
[0084] The processor is configured to, while running the computer program, execute the steps of the above-described method for predicting the time-varying probability of medium- and long-term runoff by a hierarchical combination of all elements of "input-structure-parameter" optimization.
[0085] The present invention provides a storage medium storing a computer program, which, when executed by at least one processor, implements the steps of the above-described method for the optimal selection of a hierarchical combination of all elements of "input-structure-parameter" for predicting the time-varying probability of medium- and long-term runoff.
[0086] Beneficial effects: Compared with the prior art, the significant technical effects of the present invention are as follows: The present invention optimizes and selects the three elements of the hydrological runoff prediction model, namely the prediction factors, model structure, and model parameters, respectively. By initially screening and selecting the combination of prediction driving factors and eliminating redundant prediction information, the invention achieves hierarchical combination optimization of the prediction model's "input-structure-parameters", thereby improving the accuracy, reliability, and stability of runoff prediction and providing effective and rich decision-making information for the efficient utilization of water resources. Attached Figure Description
[0087] Figure 1 This is a flowchart of the method of the present invention;
[0088] Figure 2 This is a flowchart of a stepwise regression process with an inner and outer double-loop nested structure when optimizing the combination of solutions;
[0089] Figure 3 This is a schematic diagram comparing the monthly runoff probability prediction using full-factor optimization and partial-factor optimization in the implementation case. Detailed Implementation
[0090] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. However, it will be apparent to those skilled in the art that the present invention can be practiced without one or more of these details. In other instances, some technical features well-known in the art have not been described to avoid confusion with the present invention. The order of the relevant steps in the present invention is not limiting; that is, those skilled in the art can adjust it. The order in the present invention is an exemplary writing style, not a limiting description.
[0091] This invention discloses a method for predicting the time-varying probability of medium- and long-term runoff using a hierarchical combination optimization of all elements, including "input-structure-parameters". The method involves collecting long-term series of medium- and long-term runoff, rainfall, evaporation, and atmospheric circulation factor data for the study area to form a runoff prediction factor set. Based on copula entropy and runoff causal analysis, the prediction factors are screened and freely combined to obtain a driving factor set. A set of single-valued runoff prediction schemes is generated using a long short-term memory (LSTM) neural network model. Information entropy is used as an indicator to evaluate the information value of the single-valued prediction scheme set. A stepwise regression with an inner and outer nested structure is used to optimize the combination of schemes in the scheme set. The optimal subset is then used to generate a runoff ensemble prediction scheme that simultaneously includes single-valued processes and probability intervals using a Bayesian model averaging (BMA). The runoff ensemble prediction error is finely characterized based on a generalized autoregressive conditional heteroscedasticity (GARCH) model. An error scenario sequence is generated by combining Monte Carlo sampling and superimposed onto the BMA single-valued process to obtain a runoff probability prediction scheme considering error correction. This invention addresses the three key elements of hydrological runoff prediction models: predictive factors, model structure, and model parameters. It employs a hierarchical combination optimization approach of "input-structure-parameter" to upgrade and regulate the prediction model, effectively improving prediction accuracy, reliability, and stability. This provides effective and abundant forecasting and decision-making information for water resources planning and management.
[0092] like Figure 1 As shown, the present invention provides a method for predicting the time-varying probability of medium- and long-term runoff through a hierarchical combination of all elements, including input, structure, and parameters, comprising the following steps:
[0093] S1. Input Optimization: Generation of single-value runoff prediction schemes.
[0094] As the first layer of the hierarchical combination optimization of all elements of "input-structure-parameter", the model input is optimized: long-term meteorological and hydrological factor data including atmospheric circulation factors, precipitation, evaporation and runoff are collected in the study area to construct the initial factor set for runoff prediction. Based on copula entropy and runoff cause analysis, the combination of prediction factors is screened to obtain the driving factor set. The LSTM model is used to generate a set of medium- and long-term runoff single-value prediction schemes.
[0095] Specifically, the following steps are included:
[0096] S11. Generation of the initial factor set for runoff prediction;
[0097] Based on the regional runoff generation and runoff characteristics, hydrological elements that have a direct or indirect impact on the current period are selected as forecasting factors. Physical or mathematical models are constructed to describe the complex nonlinear relationship between the forecasting factors (independent variables) and runoff (dependent variable), namely:
[0098] Q t =f(R) t-1 ,R t-2 ,…,R t-τ (1)
[0099] In the formula, Q t Let t be the runoff at time t, R be the predictor factor whose subscript indicates the corresponding time, f be the prediction equation, and τ be the lead time of the predictor factor.
[0100] The selection of predictive factors directly affects the accuracy of predictions. Generally, hydrological factors influencing the formation of medium- and long-term runoff include antecedent runoff, rainfall, and evapotranspiration, while meteorological factors include temperature, air pressure, and circulation indices. Meteorological factors provide qualitative trend predictions, while hydrological factors provide quantitative numerical predictions. Based on the meteorological conditions of the study area, the main weather factors controlling rainfall and runoff in the study area (such as monsoon, sea level pressure, subtropical high, etc.) are identified, and corresponding indicators (such as monsoon index, Southern Oscillation index, subtropical high area index, etc.) are selected. These, along with antecedent rainfall, evapotranspiration, and runoff volume, constitute the initial factor set for runoff prediction.
[0101] S12. Identification of driving factors based on copula entropy and runoff causation analysis;
[0102] Using all factors as model input not only significantly increases computation time but may also lead to model generalization. Therefore, initial screening of forecast factors is necessary, with the key being to find "driving factors" that are strongly correlated with runoff sequences and have appropriate causal explanations. Copula entropy is a nonlinear correlation measure with all orders and multiple variables, which can be used to describe the correlation between forecast factors and runoff sequences. Let N-dimensional random variables (X1, X2, ..., X...) be used. N ) has marginal distribution u i =F i (x i ), i = 1, 2, ..., N, where F is the cumulative probability density function, and the copula entropy can be defined as:
[0103]
[0104] In the formula, CE is the copula entropy, and C is the copula function. Therefore, this invention ranks all predictive factors related to runoff causes based on copula entropy, and selects the top r factors as the driving factor set W = {W1, W2, ... W}. r}
[0105] Generation of single-value runoff prediction schemes for S13 and LSTM models;
[0106] LSTM is used to predict runoff, employing different combinations of factors as dependent variables for learning, training, simulation, and prediction. LSTM is a variant of Recurrent Neural Networks (RNNs) and can solve the gradient explosion and vanishing problems of general RNN models in long-sequence regression. An LSTM consists of an input layer, one or more hidden layers, and an output layer. Each hidden layer contains two state variables, h and c, used to store short-term and long-term states, respectively. These are connected by gating units, which consist of a forget gate, an input gate, and an output gate. The forget gate determines how much of the input from the previous time step is retained in the current time step, while the input gate determines how much of the input from the current time step is stored in the long-term state c. The output gate determines how much of the long-term state c can be used as the output value of the LSTM. The LSTM calculation method is as follows:
[0107] (a) The Gate of Oblivion:
[0108] f t =σ(W f ·[h t-1 ,x t ]+b f (3)
[0109] In the formula, f tForget gate control of the current time period, it means that the historical output information of the control part is converted into the input of the current time period, reflecting the degree of loss of historical information in the current time period, W f Let b be the weight matrix of the forget gate. f σ is the bias term for the forget gate, and σ is the activation function, usually the Sigmoid function.
[0110] (b) Input gate:
[0111] i t =σ(W i ·[h t-1 ,x t ]+b i (4)
[0112] In the formula, i t This refers to input gate gating for the current time period, indicating that the input of the control part for the current time period is transformed into system state variables, reflecting the information input efficiency for the current time period. W i Let b be the weight matrix of the input gate. i This is the bias term for the input gate;
[0113] (c) The current long-term state c t ':
[0114] c t '=tanh(W c ·[h t-1 ,x t ]+b c (5)
[0115] In the formula, W c Let b be the weight matrix of the current input long-term state. c For W c The bias term, tanh is the activation function;
[0116] (d) The long-term state at the current moment c t :
[0117] c t =f t ·c t-1 +i t ·c t (6)
[0118] (e) Output gate:
[0119] o t =σ(W o ·[h t-1 ,x t ]+b o (7)
[0120] In the formula, ot This is the output gate for the current time period, indicating that the state variables of the control part of the system are transformed into the output for the current time period, reflecting the information output efficiency for the current time period. W o Let b be the weight matrix of the output gate. o For W o The bias term.
[0121] (f) The output value h of the LSTM t :
[0122] h t =o t ·tanh(c t ) (8)
[0123] The calculation period is divided into a training period and a validation period. The set of driving factors W identified in step S12 during the training period is used as the input term x in the above calculation steps. t Runoff volume is used as the output term h. t The parameters W are obtained after training and calibration of the neural network. f b f Wi, b i W o b o and the input item x during the verification period. t Substitute them together into the model to calculate the corresponding output term h. t This yields a set of single-value prediction processes, forming a single-value prediction scheme set, such as... Figure 3 As shown in (c), the prediction scheme that considers the selection of predictor factors can cover the measured true value as much as possible and has high reliability, but its interval width is too large and its accuracy needs to be improved.
[0124] S2, Structural Optimization: Generation of runoff ensemble prediction scheme based on BMA;
[0125] As the second layer of the "input-structure-parameter" all-element hierarchical combination optimization, the model structure is optimized: information entropy is used as the indicator to evaluate the information value of the single-value prediction scheme set, and stepwise regression with inner and outer nested loop structure is used to optimize the combination of schemes in the single-value prediction scheme set. The optimal subset is generated by using Bayesian model averaging (BMA) to generate a runoff set prediction scheme that simultaneously includes single-value processes and probability intervals.
[0126] Specifically, the following steps are included:
[0127] S21. Optimal selection of solution set combinations based on information entropy and stepwise regression;
[0128] Due to the "iso-parameter heterogeneity" of hydrological models and the consistency of driving factors, there is a certain degree of information redundancy among candidate schemes within the single-value prediction scheme set. This leads to an increase in computational scale and a decrease in computational efficiency. Excessive input may introduce a large amount of noise into the model, amplifying the uncertainty of the model output and reducing the model's stability. Therefore, it is necessary to remove redundant information and retain key information through combinatorial optimization. This process is a stepwise regression process with an inner and outer nested loop structure, such as... Figure 2 As shown, the inner loop includes: starting from the initial set of single-value prediction schemes, deleting a scheme in ascending order of its number, calculating the information content between the scheme and the remaining scheme subsets, selecting the scheme with the most information redundancy, updating its remaining scheme subset to the new set, and repeating the above operation until the number of elements in the remaining scheme subset is 1, at which point the loop terminates; the outer loop includes: for the remaining scheme subsets with different numbers of elements, calculating their total information content, comparing it with the total information content of the initial set, and when the ratio is less than a given threshold, taking the current set as the optimal combination, at which point the loop terminates. Specifically, the stepwise regression process with an inner and outer double-loop nested structure for scheme set combination optimization is as follows:
[0129] (1) Let M = F; m = |F|; where F is the initial set of single-value prediction schemes, M represents the updated set of single-value prediction schemes, and m represents the number of schemes in the updated set;
[0130] (2) Let i = 1, S i =M\{M i}, where M i Let S represent the i-th scheme. i This represents the remaining subset of single-valued prediction schemes after removing the i-th scheme;
[0131] (4) Calculate Obj i =C v (M i ,S i ), C v Represents the remainder subset S i With updated complete series M i The information is the same, and at this point, one inner loop is completed;
[0132] (5) For i, take values from 1 to m, that is, traverse and update each scheme in the entire set. Continue the inner loop and repeat the above steps to obtain the information redundancy of the remaining subset after deleting any scheme in the set. Select the scheme with the highest redundancy in the corresponding remaining subset to be deleted, and obtain the optimal remaining subset S. min The inner loop terminates;
[0133] (6) The optimal residual subset S minSet the new updated global set M, calculate the total information H between it and the initial global set F, and complete one outer loop.
[0134] (7) When the ratio of the total information of the updated set M to the total information of the initial set F (i.e., the relative total information) is greater than a given threshold β (i.e., H(M) / H(F)≤β does not hold), the outer loop continues to calculate the relative total information of the updated set M with different numbers of elements; when the relative total information is less than a given threshold or the number of elements in the updated set M decreases to 2 (i.e., m=2), the outer loop terminates, and the updated set M at this time is the determined optimal subset S. F .
[0135] This invention selects information entropy as an indicator to measure the information value of a prediction scheme. Information entropy is a statistic that measures the uncertainty of an information system. Starting from the probability distribution of random variables, it can characterize the independence, correlation, and redundancy between different distributions, and is therefore widely used in various scenarios such as prediction, optimization, and decision-making. The expression for information entropy is:
[0136]
[0137] In the formula, H is the information entropy, and X is a random variable with N discrete values, X = {x1, x2, ..., x...} N}, its probability distribution Furthermore, for a binary random variable (X,Y), its information entropy is also called joint entropy, defined as:
[0138]
[0139] In the formula, p(x,y) is the joint probability density function of random variables X and Y. Joint entropy represents the total amount of information contained in multiple random variables, and equal amounts of information can be defined as mutual information, i.e.:
[0140] T(X,Y)=H(X)+H(Y)-H(X,Y) (11)
[0141] Therefore, the joint entropy and mutual information in information entropy can be used to describe the total amount of information and the amount of information among the candidate schemes in the single-value prediction scheme set, respectively, and these can be used as the basis for the optimal combination of schemes.
[0142] S22. Generation of runoff single-value-interval ensemble prediction scheme based on BMA;
[0143] Bayesian model averaging (BMA) has wide applications in ensemble probability forecasting and hydrological uncertainty analysis. Based on Bayesian theory, this method analyzes the probability that each candidate model can become the optimal model and tends to select the model with the higher probability in the final model selection. Therefore, it can effectively eliminate the adverse effects of model uncertainty.
[0144] Let y be the forecast variable, y obs Let f be the sequence of measured samples, f = {f1, f2, ..., f...} k Let} be the candidate model space, where K represents the number of candidate scheme groups within the selected single-value prediction scheme combination in step S21. According to Bayesian theory, given sample y... obs The probability density function p(y|y) of the forecast variable y obs )for:
[0145]
[0146] In the formula, p(y|f i ,y obs ) is for a given sample y obs and model f i Under the given conditions, the probability density function of the predicted variable y, i.e., the "posterior distribution"; p(f i |y obs ) is for a given sample y obs The probability that the i-th prediction model is the optimal model under the given condition is the BMA weight ω.
[0147] Clearly, the forecast of variable y can be given in the form of deterministic single-valued results or probability density functions, thus enabling single-valued process forecasting and probability interval forecasting. Assuming that both the observed and forecast sequences follow a normal distribution, further refinement is possible:
[0148]
[0149] In the formula, p(y|y obs ω represents the probability distribution of the predictor variable. i This represents the weight of the i-th model. The mean is f i The variance is It follows a normal distribution.
[0150] Therefore, the BMA parameter ω is calibrated using the Expectation-Maximization (EM) algorithm. i and Calibration and Monte Carlo sampling can provide a probabilistic prediction scenario sequence. The specific steps are as follows:
[0151] (a) The non-normally distributed runoff sequence is transformed into a normally distributed one using the BOX-COX transformation. The specific steps are as follows:
[0152]
[0153] In the formula, λ is the transformation coefficient, the value of which is given by the maximum likelihood estimation; y' is the transformed runoff prediction sequence that follows a normal distribution.
[0154] (b) Calibrate BMA parameters ω using the EM algorithm i and Let the iteration number Iter = 0, and assume the initial weights are uniform weights, i.e. Calculate the initial variance Where T is the number of calculation periods. The initial likelihood function value is... When the iteration reaches the Iter-th iteration, there are hidden variables:
[0155]
[0156] At this point, the weights are updated, i.e. Recalculate variance With the likelihood function value l (Iter+1 The calculation formulas are as follows:
[0157]
[0158]
[0159] Repeat the above steps until l (Iter+1 )-l (Iter) If the result is ≤0.001, the iteration stops. Parameter calibration complete;
[0160] (c) Assuming that the selected single-value prediction scheme combination in step S21 contains K groups of candidate schemes, based on the BMA weights [ω1,ω2,...,ω] of each candidate scheme... K Define the cumulative weight value.
[0161]
[0162] Generate a uniform random number u in the range [0,1]. This indicates that the i-th candidate solution f was selected in this sampling. i ;
[0163] (d) According to f i Probability distribution in time period j Randomly generate a runoff prediction sequence y' that follows a normal distribution;
[0164] (e) The generated runoff prediction sequence y', which follows a normal distribution, is restored based on formula (13). The calculation formula is as follows:
[0165]
[0166] (f) Repeating steps (c)-(e) M times yields a set of prediction scenarios with M scenarios, which is the runoff probability prediction scheme for BMA. Its mean process is the runoff single-value prediction scheme for BMA. Together, they constitute a runoff ensemble prediction scheme containing single-value processes and probability intervals, such as... Figure 3 As shown in (b), the prediction scheme that considers the prediction factors and model structure optimization can reduce the interval width while covering the measured true value as much as possible, and has high accuracy. However, its prediction deviation is too large when the prediction is inaccurate, and its reliability needs to be improved.
[0167] S3. Parameter optimization: Generating a runoff probability prediction scheme that takes error correction into account;
[0168] As the third layer of the hierarchical combination optimization of all elements of "input-structure-parameter", the model parameters are optimized: a time-varying simulation model of runoff ensemble prediction error is constructed based on the generalized autoregressive conditional heteroscedasticity model (GARCH), real-time error correction is performed using the correlation of errors, and a corrected error scenario sequence is generated by combining Monte Carlo sampling. This sequence is then superimposed on the BMA single-value forecast results to obtain a runoff probability prediction scheme that takes error correction into account. The obtained probability prediction scheme is evaluated based on the multi-dimensional evaluation index system of "accuracy-reliability-stability" of runoff prediction.
[0169] Specifically, the following steps are included:
[0170] S31. Fine characterization of runoff prediction error based on GARCH model;
[0171] The Generalized Autoregressive Conditional Heteroskedasticity (GARCH) model is a regression model tailored for financial data. Besides sharing similarities with ordinary regression models (such as ARMA and ARCH), GARCH further models the variance of the error, making it particularly suitable for the analysis and forecasting of time series with volatile clustering. Such analysis can play a crucial guiding role in investor decision-making. The variance of runoff forecasts often exhibits volatile clustering and displays a "lumpy, heavy-tailed" statistical distribution, highly consistent with the patterns of financial series. Therefore, it can be used to describe the time-history variations of runoff forecast error series.
[0172] GARCH models are generally described by two basic equations: conditional mean and conditional variance. The conditional mean is given by the single-value prediction process in step S22, so only the variance (i.e., the error term) needs to be modeled to achieve a precise characterization of runoff prediction errors. Assume the residual term ε... t With conditional variance h t obey:
[0173]
[0174] In the formula, c is a constant term, and a i and b j The conditional variances h of lags i and j are respectively. t-i and residual term ε t-j The coefficients, m and n, are the lag orders of the GARCH conditional variance and residual terms, respectively, and can be denoted as ε. t ~GARCH(m,n). Maximum likelihood estimation is used for parameter a. i b j The calibration was carried out.
[0175] S32. Runoff prediction error correction based on Monte Carlo sampling;
[0176] Based on the GARCH model and its parameters obtained in step S31, a series of simulated scenarios for runoff prediction errors can be generated using Monte Carlo sampling. For each time t, N sets of random numbers v following a standard normal distribution are first randomly generated using the Monte Carlo method. t Substituting into formula (16), the conditional variance h is calculated. t The simulated scenario sequence can be used to obtain the error simulated scenario sequence by taking the square root of the conditional variance term. This sequence is the correction for runoff prediction error.
[0177] S33. Generation of time-varying runoff probability prediction scheme considering error correction;
[0178] Generally, the relationship between prediction error and measured value is as follows:
[0179]
[0180] In the formula, y t This is the predicted value for time period t. Let e be the measured true value for time period t. t This represents the prediction error for that time period. If error correction is considered, i.e., when e... t '=z(e t When ), we have:
[0181]
[0182]
[0183] In the formula, e t ' represents the corrected runoff prediction error, and z is the error correction function, which reflects the statistical characteristics of historical prediction errors; This is the "theoretical value" of runoff, that is, a precise prediction of the measured runoff value after considering error correction. Combined with the runoff prediction error simulation scenario sequence generated in step S32, i.e., e... tBy superimposing this onto the runoff single-value prediction process of BMA in step S22, the runoff probability prediction scenario sequence can be obtained, that is, the runoff time-varying probability prediction scheme considering error correction, such as... Figure 3 As shown in (a), the prediction scheme after considering all factors can further reduce the interval width while covering the measured true value as much as possible. At the same time, it has a smaller prediction deviation in the case of prediction inaccuracy and has higher stability.
[0184] S34. Evaluation of runoff prediction accuracy;
[0185] Generally, the evaluation of probabilistic prediction results should focus on both accuracy and reliability. Accuracy refers to the closeness of a certain quantile (median Q50 or mean) of the probabilistic prediction result distribution to the measured value. An accurate runoff prediction scheme satisfies the following: (a) unbiased prediction error; (b) the prediction error variance is as small as possible, which can be evaluated using the root mean square error (RMSE). Reliability refers to the ability of the probabilistic prediction result to cover the measured value within a certain confidence interval at a certain confidence level, usually evaluated using coverage rate (CR) and Brier score (BS). Furthermore, from the perspective of risk management and system security, a good prediction scheme should have strong risk resistance, i.e., stability. When the probabilistic prediction is inaccurate (i.e., the predicted interval does not cover the measured true value at a given confidence level), the closer its interval threshold is to the measured value, the more effective information the prediction scheme provides to decision-makers, and the stronger its risk resistance. Inaccuracy bias (MD) considers the minimum prediction bias under the condition of prediction inaccuracy and can be used to characterize the stability of runoff probabilistic prediction. Therefore, this invention constructs a multi-dimensional evaluation index for runoff prediction, encompassing accuracy, reliability, and stability. RMSE, BS, and MD are selected as indicators to evaluate the accuracy, reliability, and stability of runoff probability prediction, respectively, to measure the merits of the prediction scheme and provide a comprehensive and systematic evaluation. The specific calculation formula is as follows:
[0186]
[0187]
[0188]
[0189] In the formula, α is the confidence level, T is the number of time periods, p is the quantile, and prob is the confidence level. i With o i Let be the probability of the i-th prediction occurring and the actual situation (if the prediction result is within the allowable error range, then o). i =1, otherwise o i =0), and These represent the upper and lower limits of the permissible error for runoff prediction in the t-th time period, typically taken as ±30% of the relative error. and These represent the upper and lower limits of the runoff probability prediction interval for the t-th time period, respectively.
[0190] The present invention provides a medium- to long-term runoff time-varying probability prediction system based on a hierarchical combination of all elements including input, structure, and parameters, comprising:
[0191] The model input optimization module is used to collect long-term meteorological and hydrological factor data of the study area to construct the initial factor set for runoff prediction. Based on copula entropy and runoff causation analysis, the combination of prediction factors is screened to obtain the driving factor set. A long short-term memory neural network model based on deep learning is used to generate a set of medium- and long-term runoff single-value prediction schemes, which serves as the first layer of the hierarchical combination optimization of all elements of "input-structure-parameter".
[0192] The model structure optimization module is used to evaluate the information value of the single-value prediction scheme set by using information entropy as an indicator. It adopts stepwise regression with inner and outer nested loop structure to optimize the combination of schemes in the scheme set. The optimal subset is generated by averaging the Bayesian model to generate a runoff set prediction scheme that simultaneously includes single-value processes and probability intervals. This serves as the second layer of the hierarchical combination optimization of all elements of "input-structure-parameter".
[0193] The model parameter optimization module is used to construct a time-varying simulation model of runoff ensemble prediction error based on the generalized autoregressive conditional heteroscedasticity model. It uses the correlation of errors to perform real-time error correction, combines Monte Carlo sampling to generate a corrected error scenario sequence, and superimposes it onto the Bayesian model's average single-value forecast results to obtain a runoff probability prediction scheme that takes error correction into account. The obtained probability prediction scheme is evaluated according to the multi-dimensional evaluation index system of "accuracy-reliability-stability" of runoff prediction, serving as the third layer of the hierarchical combination optimization of all elements of "input-structure-parameter".
[0194] An apparatus of the present invention includes a memory and a processor, wherein:
[0195] Memory is used to store computer programs that can run on a processor;
[0196] The processor is configured to, when running the computer program, execute the steps of the above-described method for the optimal selection of the hierarchical combination of all elements of "input-structure-parameter" for predicting the time-varying probability of medium- and long-term runoff, and achieve the same technical effect as the above method.
[0197] The present invention provides a storage medium storing a computer program, which, when executed by at least one processor, implements the steps of the above-described method for the optimal selection of a hierarchical combination of all elements of "input-structure-parameter" for predicting the time-varying probability of medium- and long-term runoff, and achieves the same technical effect as the above-described method.
Claims
1. A method for predicting the time-varying probability of medium- and long-term runoff through hierarchical combination optimization of all elements including "input-structure-parameters", characterized in that, Includes the following steps: S1. As the first layer of the hierarchical combination optimization of all elements of "input-structure-parameter", the model input is optimized: collect long-term meteorological and hydrological factor data of the study area to construct the initial factor set for runoff prediction, screen the combination of prediction factors based on copula entropy and runoff cause analysis to obtain the driving factor set, and use the long short-term memory neural network LSTM model based on deep learning to generate a set of medium- and long-term runoff single-value prediction schemes. S2. As the second layer of the "input-structure-parameter" all-element hierarchical combination optimization, the model structure is optimized: information entropy is used as the indicator to evaluate the information value of the single-value prediction scheme set. Stepwise regression with inner and outer nested loop structure is used to optimize the combination of schemes in the single-value prediction scheme set. The optimal subset is generated by using Bayesian model average BMA to generate a runoff set prediction scheme that simultaneously includes single-value processes and probability intervals. include: S22. Generation of runoff single-value-probability ensemble prediction scheme based on BMA; Let y be the predictor variable, y obs Let f be the sequence of measured samples, f = {f1, f2, ..., f...} k Let} represent the candidate model space, where K represents the number of candidate scheme groups within the preferred single-value prediction scheme combination; based on Bayesian theory, it is assumed that both the measured and predicted sequences follow a normal distribution, further refined as follows: Where p(y|y obs ω represents the probability distribution of the predictor variable. i This represents the weight of the i-th model. The mean is f i The variance is The normal distribution; The BMA parameter ω is calibrated using the expectation-maximization algorithm. i and Monte Carlo sampling is used to give a probability prediction scenario sequence, and a runoff set prediction scheme containing single-valued processes and probability intervals is obtained. Monte Carlo sampling is used to generate a probability prediction scenario sequence. The specific steps are as follows: (a) The non-normally distributed runoff sequence is transformed into a normally distributed one using the BOX-COX transformation, specifically: Where λ is the transformation coefficient, the value of which is given by the maximum likelihood estimation; y' is the transformed runoff prediction sequence that follows a normal distribution; (b) Calibrate BMA parameters ω using the EM algorithm i and (c) Based on the BMA weights [ω1,ω2,...,ω] of each candidate scheme K Define the cumulative weight value. Generate a uniform random number u in the range [0,1]. This indicates that the i-th candidate solution f was selected in this sampling. i ; (d) According to f i Probability distribution in time period j Randomly generate a runoff prediction sequence y' that follows a normal distribution; (e) The generated runoff prediction sequence y', which follows a normal distribution, is restored based on the cumulative weight value formula. The calculation formula is as follows: (f) Repeat steps (c)-(e) M times to obtain a set of prediction scenarios with M scenarios, which is the runoff probability prediction scheme of BMA. The mean process is the runoff single value prediction scheme of BMA. The two together constitute the runoff single value-probability set prediction scheme. S3. As the third layer of the hierarchical combination optimization of all elements of "input-structure-parameter", the model parameters are optimized: the runoff prediction error is finely characterized based on the generalized autoregressive conditional heteroscedasticity model GARCH, the error is corrected in real time by using the correlation of errors, and the corrected error scenario sequence is generated by combining Monte Carlo sampling. This sequence is then superimposed on the single-valued process in step S2 to obtain a runoff time-varying probability prediction scheme that takes error correction into account. The obtained probability prediction scheme is evaluated according to the multi-dimensional evaluation index system of "accuracy-reliability-stability" of runoff prediction.
2. The method for predicting the time-varying probability of medium- and long-term runoff by hierarchical combination optimization of all elements (input-structure-parameters) according to claim 1, characterized in that, Step S1 includes the following steps: S11. Generation of the initial factor set for runoff prediction; Based on the meteorological conditions of the study area, the main weather elements controlling rainfall and runoff in the study area are identified, and the corresponding indicators are selected together with previous rainfall, evapotranspiration and runoff to form the initial factor set for runoff prediction. S12. Identification of driving factors based on copula entropy and runoff causation analysis; Let N-dimensional random variables (X1, X2, ..., Xn) be defined. N ) has marginal distribution u i =F i (x i ), i = 1, 2, ..., N, where F is the cumulative probability density function, and the copula entropy is defined as: Where CE is the copula entropy and C is the copula function. Based on the copula entropy, all predictive factors related to runoff causes are ranked by correlation, and the top-ranked factors are selected as the driving factor set. S13. Generation of a set of single-value runoff prediction schemes based on the LSTM model; LSTM is used to predict the inflow into the lake each month. Different combinations of factors are used as dependent variables for learning, training, simulation and prediction to obtain a set of single-value prediction schemes.
3. The method for predicting the time-varying probability of medium- and long-term runoff through a hierarchical combination of all elements (input-structure-parameters) as described in claim 2, is characterized in that... The LSTM calculation method in step S13 is as follows: (a) The Gate of Oblivion: f t =σ(W f ·[h t-1 ,x t ]+b f ) Among them, f t Forget gate control of the current time period, it means that the historical output information of the control part is converted into the input of the current time period, reflecting the degree of loss of historical information in the current time period, W f Let b be the weight matrix of the forget gate. f For W f The bias term, σ is the activation function; (b) Input gate: i t =σ(W i ·[h t-1 ,x t ]+b i ) Among them, i t This refers to input gate gating for the current time period, indicating that the input of the control part for the current time period is transformed into system state variables, reflecting the information input efficiency for the current time period. W i Let b be the weight matrix of the input gate. i For W i The bias term; (c) The current long-term state c t ': c t '=tanh(W c ·[h t-1 ,x t ]+b c ) Among them, W c Let b be the weight matrix of the current input long-term state. c For W c The bias term, tanh is the activation function; (d) The long-term state at the current moment c t : c t =f t ·c t-1 +i t ·c t ' (e) Output gate: the t =σ(W o ·[h t-1 ,x t ]+b o ) Among them, o t This is the output gate for the current time period, indicating that the state variables of the control part of the system are transformed into the output for the current time period, reflecting the information output efficiency for the current time period. W o Let b be the weight matrix of the output gate. o For W o The bias term; (f) The output value h of the LSTM t : h t =o t ·tanh(c t ) The calculation period is divided into a training period and a validation period. The set of driving factors W identified in step S12 during the training period is used as the input term x in the above calculation steps. t Runoff volume is used as the output term h. t The parameters W are obtained after training and calibration of the neural network. f b f Wi, b i W o b o and the input item x during the verification period. t Substitute them together into the model to calculate the corresponding output term h. t This yields a set of single-value prediction processes, forming a single-value prediction scheme set.
4. The method for predicting the time-varying probability of medium- and long-term runoff through a hierarchical combination of all elements (input-structure-parameters) as described in claim 1, characterized in that: The following steps are included before step S22: S21. Optimal selection of solution set combinations based on information entropy and stepwise regression; The expression for information entropy is: Where H is the information entropy, and X is a random variable with N discrete values, X = {x1, x2, ..., x...} N }, whose probability distribution is P={p(x1),p(x2),...,p(x... N )}, Furthermore, for a binary random variable (X,Y), its information entropy is also called joint entropy, defined as: Where p(x,y) is the joint probability density function of random variables X and Y; joint entropy represents the total amount of information contained in multiple random variables, and their equal amounts of information are defined as mutual information, i.e.: T(X,Y)=H(X)+H(Y)-H(X,Y) Therefore, joint entropy and mutual information in information entropy are used to describe the total amount of information and information content among candidate schemes in the single-value prediction scheme set, respectively, and these are used as the basis for the optimal combination of schemes.
5. The method for predicting the time-varying probability of medium- and long-term runoff through a hierarchical combination of all elements (input-structure-parameters) as described in claim 1, characterized in that... Step S3 includes the following steps: S31. Fine characterization of runoff prediction error based on GARCH model; The GARCH model is described by two fundamental equations: conditional mean and conditional variance. The conditional mean is given by the single-value prediction process in step S2. The conditional equations are calculated as follows: Assume residual term ε t With conditional variance h t obey: Where c is a constant term, a i and b j The conditional variances h of lags i and j are respectively. t-i and residual term ε t-j The coefficients, m and n, are the lag orders of the GARCH conditional variance and residual terms, respectively, denoted as ε. t ~GARCH(m,n); maximum likelihood estimation is used for parameter a i b j Perform calibration; S32. Runoff prediction error correction based on Monte Carlo sampling; Based on the GARCH model and its parameters obtained in step S31, and combined with Monte Carlo sampling to generate a runoff prediction error simulation scenario sequence, for each time t, N sets of random numbers v following a standard normal distribution are first randomly generated using the Monte Carlo method. t Substituting into the formula in step S31, the conditional variance h is calculated. t The simulated scenario sequence is obtained by taking the square root of the conditional variance term to obtain the error simulated scenario sequence, which is the correction for the runoff time-varying prediction error. S33. Generation of time-varying runoff probability prediction scheme considering error correction; The relationship between prediction error and measured value is as follows: Among them, y t This is the predicted value for time period t. Let e be the measured true value for time period t. t This represents the prediction error for that period; if error correction is considered, i.e., when e t '=z(e t When ), we have: Among them, e t ' represents the corrected runoff prediction error, and z is the error correction function, which reflects the statistical characteristics of historical prediction errors; This is the "theoretical value" of runoff, that is, a precise prediction of the measured runoff value after considering error correction; combined with the runoff prediction error simulation scenario sequence generated in step S32, i.e., e t ', and superimpose it on the runoff single-value prediction process of BMA in step S22 to obtain the runoff probability prediction scenario sequence, that is, the runoff probability prediction scheme considering error correction; S34. Evaluation of runoff prediction accuracy; A multi-dimensional evaluation index of "accuracy, reliability, and stability" for runoff prediction was constructed. The root mean square error (RMSE), Brier score (BS), and misalignment deviation (MD) were selected as indicators to evaluate the accuracy, reliability, and stability of runoff probability prediction, and to measure the merits of the prediction schemes. A comprehensive and systematic evaluation of the prediction schemes was conducted.
6. A medium- to long-term runoff time-varying probability prediction system with hierarchical combination optimization of all elements (input-structure-parameters), characterized in that, include: The model input optimization module is used to collect long-term meteorological and hydrological factor data of the study area to construct the initial factor set for runoff prediction. Based on copula entropy and runoff causation analysis, the combination of prediction factors is screened to obtain the driving factor set. A long short-term memory neural network model based on deep learning is used to generate a set of medium- and long-term runoff single-value prediction schemes, which serves as the first layer of the hierarchical combination optimization of all elements of "input-structure-parameter". The model structure optimization module uses information entropy as an indicator to evaluate the information value of the single-value prediction scheme set. It employs stepwise regression with inner and outer nested loop structures to optimize the combination of schemes in the scheme set. The optimal subset is then averaged using a Bayesian model to generate a runoff set prediction scheme that simultaneously includes single-value processes and probability intervals. This serves as the second layer of the hierarchical combination optimization of all elements, including "input-structure-parameters". It includes: a runoff single-value-probability set prediction scheme generation unit based on BMA. Let y be the predictor variable, y obs Let f be the sequence of measured samples, f = {f1, f2, ..., f...} k } represents the candidate model space, where K denotes the number of candidate scheme groups within the selected single-value prediction scheme combination in step S21; based on Bayesian theory, assuming that both the measured and predicted sequences follow a normal distribution, further refinement is made as follows: Where p(y|y obs ω represents the probability distribution of the predictor variable. i This represents the weight of the i-th model. The mean is f i The variance is The normal distribution; The BMA parameter ω is calibrated using the expectation-maximization algorithm. i and Monte Carlo sampling is used to give a probability prediction scenario sequence, and a runoff set prediction scheme containing single-valued processes and probability intervals is obtained. Monte Carlo sampling is used to provide a probability prediction scenario sequence, specifically including: (a) The non-normally distributed runoff sequence is transformed into a normally distributed one using the BOX-COX transformation, specifically: Where λ is the transformation coefficient, the value of which is given by the maximum likelihood estimation; y' is the transformed runoff prediction sequence that follows a normal distribution; (b) Calibrate BMA parameters ω using the EM algorithm i and (c) Based on the BMA weights [ω1,ω2,...,ω] of each candidate scheme K Define the cumulative weight value. Generate a uniform random number u in the range [0,1]. This indicates that the i-th candidate solution f was selected in this sampling. i ; (d) According to f i Probability distribution in time period j Randomly generate a runoff prediction sequence y' that follows a normal distribution; (e) The generated runoff prediction sequence y', which follows a normal distribution, is restored based on the cumulative weight value formula. The calculation formula is as follows: (f) Repeat steps (c)-(e) M times to obtain a set of prediction scenarios with M scenarios, which is the runoff probability prediction scheme of BMA. The mean process is the runoff single value prediction scheme of BMA. The two together constitute the runoff single value-probability set prediction scheme. The model parameter optimization module is used to construct a time-varying simulation model of runoff ensemble prediction error based on a generalized autoregressive conditional heteroscedasticity model. It uses the correlation of errors to perform real-time error correction, combines Monte Carlo sampling to generate a corrected error scenario sequence, and superimposes it onto the Bayesian model's average single-value forecast results to obtain a runoff probability prediction scheme that takes error correction into account. The obtained probability prediction scheme is evaluated according to the multi-dimensional evaluation index system of "accuracy-reliability-stability" of runoff prediction, serving as the third layer of the hierarchical combination optimization of all elements of "input-structure-parameter".
7. A device, characterized in that, Includes memory and processor, wherein: Memory is used to store computer programs that can run on a processor; A processor, configured to, while running the computer program, execute the steps of a method for predicting the time-varying probability of medium- and long-term runoff by a hierarchical combination of all elements of "input-structure-parameter" as described in any one of claims 1-5.
8. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by at least one processor, implements the steps of the medium- and long-term runoff time-varying probability prediction method as described in any one of claims 1-5, which is a hierarchical combination optimization of all elements of "input-structure-parameter".
Citation Information
Patent Citations
RBF neural network medium-and-long-term runoff prediction method based on runoff production mechanism
CN110188922A
Runoff probabilistic prediction method and system integrating deep learning and error correction
CN111275253A
Multi-station runoff medium-and-long-term rolling probability prediction method considering prediction uncertainty associated evolution characteristics
CN112116149A