A runoff prediction method based on variational Bayesian convolutional single-control memory network
By adopting a variational Bayesian convolution single-controlled memory network in the hydrological forecast model, combining the maximum translation Pearson correlation coefficient method and three-dimensional tensor input, the difficulties in parameter rate determination and calculation burden of the existing hydrological forecast model are solved, high-precision multi-foresight runoff prediction are achieved, and the model structure is simplified.
Patent Information
- Application Number
- CN202410336665.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-03-22
AI Technical Summary
The existing hydrological forecast models have difficulties in parameter rate determination, model input preparation and calculation burden, and although the prediction accuracy of the empirical hydrological model is high, the interpretability is poor, making it difficult to achieve high-precision runoff prediction in multiple foresight periods.
The runoff prediction method based on the variational Bayesian convolution single-control memory network is adopted. When analyzing the lag by the maximum translation Pearson correlation coefficient method, a three-dimensional tensor input model is constructed, and a convolutional single-control memory neural network and a variational Bayes framework is combined to achieve high-precision runoff probability prediction in multiple foreseeable periods.
It improves the accuracy and reliability of runoff prediction, simplifies the model structure, shortens the training time, and realizes the multi-foresee period prediction based on only historical data, which enhances the interpretability of prediction.
Smart Images

Figure CN118364941B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of runoff prediction, and specifically is a runoff prediction method based on a variational Bayesian convolutional single-controlled memory network. Background Art
[0002] Prior art background:
[0003] At present, the commonly used existing technologies in the industry are as follows:
[0004] Hydrological forecasting is a technique for predicting future runoff based on known hydrological and meteorological information. Obtaining high-precision, reliable, and multi-forecast probabilistic runoff forecasts is of great significance for reservoir operation and decision-making.
[0005] Hydrological prediction models can be divided into three types according to different principles: hydrological physical models, conceptual hydrological models and empirical hydrological models. Hydrological physical models simulate and predict hydrological processes by studying the mathematical equations of rainfall, evaporation, infiltration, runoff and confluence in the water cycle and solving the partial differential equations of water flow. Commonly used hydrological physical models include the Soil and Water Assessment Tool (SWAT), the European Hydrological System and the Hydrological Distributed Model. Conceptual hydrological models use approximate equations to describe the physical processes in the water cycle based on the mass conservation equation and the momentum conservation equation. Commonly used conceptual hydrological models include the Tank model, the Xin'anjiang (XAJ) model, the ARNO model and so on. Conceptual hydrological models usually need to calibrate the parameters of the approximate equations based on historical data to obtain a hydrological forecast model for a specific area. Empirical hydrological models seek the relationship between hydrological prediction variables and hydrological factors from historical hydrological data, and predict future runoff based on existing hydrological prediction factors. They are also called data-driven models. Commonly used empirical hydrological models include autoregressive moving average (ARMA), extreme learning machine (ELM), support vector regression (SVR) and neural network (NN).
[0006] Early empirical hydrological models are usually linear and difficult to capture the complex nonlinear characteristics of hydrological processes. With the rapid development of artificial intelligence and computer technology, machine learning methods represented by neural networks (NN) have effectively solved the problem of nonlinear sequence fitting. In recent years, due to the advancement of computing power of central processing units and graphics processing units, deep learning models have attracted the attention of many scholars and have gradually been applied in the field of hydrological forecasting, such as long short-term memory networks (LSTM), gated recurrent units (GRU) and convolutional neural networks (CNN), as well as convolutional long short-term memory neural networks (ConvLSTM) and convolutional gated recurrent unit neural networks (ConvGRU) that combine CNN with LSTM and GRU. From the perspective of applying deep learning to hydrological forecasting, proposing or improving deep learning models to improve the accuracy of hydrological forecasts or optimizing model structures to reduce training time are future research directions.
[0007] The problems existing in the prior art are:
[0008] Although the model parameters of hydrological physical models have strict physical meanings and can describe the water cycle process more intuitively, there are problems such as cumbersome parameter calibration process, complex model input preparation, and excessive calculation burden.
[0009] The conceptual hydrological model divides the evolution of upstream and tributary flow and interval flow prediction into two types of models, such as the Muskingum model for flow evolution and the Xin'anjiang model for interval flow prediction. However, flow evolution and interval flow prediction are not separated in the physical basin, but integrated. At the same time, the multi-forecast period prediction of the conceptual hydrological model depends on the multi-forecast period prediction of rainfall, and cannot be completed based on historical data alone.
[0010] The advantages of empirical hydrological models are simple input and output preparation, high computational efficiency, and high prediction accuracy. The disadvantage is that the model has poor interpretability.
[0011] Difficulty in solving the above technical problems:
[0012] The conceptual hydrological model takes into account hydrological processes such as the evolution of flow at upstream stations, the evolution of flow at tributary stations, and rainfall and evaporation in the river sections between upstream and downstream. One of the difficulties is how to transform the way these variables are considered in two categories in the conceptual hydrological model into an integrated way of considering these variables in the empirical hydrological model.
[0013] Currently, empirical hydrological models mostly use deep learning models for prediction. How to further improve the prediction accuracy and shorten the training time of the same type of deep learning is also one of the difficulties that the present invention intends to solve.
[0014] Precipitation forecasts for multiple future forecast periods are not always easy to obtain. At the same time, there are always errors in the forecasts and it is difficult to make a perfect forecast. How to complete multi-forecast period forecasts based only on historical rainfall, runoff and other information and quantify the uncertainty of the forecast is also one of the difficulties that the present invention intends to solve. Summary of the invention
[0015] In order to solve the current technical problems, the main purpose of the present invention is to provide a runoff prediction method based on a variational Bayesian convolutional single-controlled memory network. By adopting the present invention, reliable and high-precision runoff multi-forecast period probability prediction results can be obtained.
[0016] In order to achieve the above technical features, the purpose of the present invention is achieved as follows: a runoff prediction method based on a variational Bayesian convolutional single-control memory network, characterized in that the runoff prediction method analyzes the lag time t of the flow evolution from the upstream station to the downstream station according to the maximum translation Pearson correlation coefficient method lag , using the three-dimensional tensor based on site-variable-time as the input and output of the prediction model, a probabilistic prediction method based on the variational Bayesian framework and convolutional single-controlled memory neural network is proposed to obtain the multi-forecast period probabilistic prediction results of runoff.
[0017] The runoff prediction method specifically comprises the following steps:
[0018] S1, collects upstream site traffic Q U , tributary station flow Q T , Downstream site flow Q D , data on rainfall P and evaporation E in the river section between upstream and downstream stations;
[0019] S2, calculate the maximum translation Pearson correlation coefficient between the flow of upstream stations, tributary stations and downstream stations respectively. The translation period corresponding to the maximum Pearson correlation coefficient is the lag time t of the flow evolution of the station to the downstream station lag ;
[0020] S3, the upstream site traffic Q U , tributary station flow Q T , Downstream site flow Q D , the variables PE of the rainfall P minus the evaporation E of the section are arranged into a two-dimensional matrix, and each variable is considered to have lag The historical data of the time periods constitutes a three-dimensional input tensor with a shape of [T×M×N], where T is the time dimension, which is equal to t lag ; M and N are the number of rows and columns of the two-dimensional matrix in which the above variables are arranged; lag The downstream station flow data of each period is arranged into a three-dimensional output tensor with a shape of [T×1×1];
[0021] S4, prepare the data set according to the three-dimensional input and output format, and divide it into a training set and a test set;
[0022] S5, constructing a convolutional single-controlled memory neural network layer ConvSCM;
[0023] S6, based on the variational Bayesian framework and the ConvSCM model, a variational Bayesian convolutional single-controlled memory network model BConvSCM is constructed;
[0024] S7, train the parameters of the BConvSCM model on the training set, predict the flow of downstream sites on the test set, and obtain high-precision and reliable probabilistic prediction results of runoff for multiple forecast periods.
[0025] Preferably, in S2, the maximum translation Pearson correlation coefficient calculation formula is:
[0026]
[0027] Where, PCC t is the Pearson correlation coefficient between the flow of the upstream site after the t-step shift and the flow of the downstream site without the shift; X 1~(T-t+1) represents the upstream station flow samples with time period numbers from 1 to T-t+1, i.e., the independent variable samples; Y t~T represents the flow samples of downstream stations with time period numbers from t to T, i.e., the samples of the dependent variable; and are the means of the independent variable samples and the dependent variable samples respectively; T is the total number of samples;
[0028]
[0029] The translation period t changes from 0 to T-2, and the Pearson correlation coefficient corresponding to each translation period t is calculated. The normal period corresponding to the largest Pearson correlation coefficient is the lag time t of the flow evolution from the upstream site to the downstream site. lag .
[0030] Preferably, in S5, the forward propagation step and calculation formula of the t-th time period of the convolutional single-controlled memory neural network layer ConvSCM are:
[0031] net t =w h *h t-1 +w x *x t (3)
[0032] a t =tanh(net t )=tanh(w h *h t-1 +w x *xt ) (4)
[0033]
[0034] In the formula, w h and w x is the weight tensor; h t is the output of the ConvSCM layer in the tth period; net t and a t is the intermediate variable of the tth period; x t represents the input of the t-th period, and its shape is [M×N], where M and N are the number of rows and columns of the two-dimensional matrix after the correlation factor variables are rearranged; the complete input x = [x 1 ,x 2 ,…,x T ], its shape is [T×M×N], T is the number of selected historical periods, equal to t lag ; tanh(·) represents the tanh activation function; the symbols * and Represent the convolution operation and the Hadamard product respectively.
[0035] Preferably, in S6, the derivation process of the BConvSCM probability prediction model is as follows:
[0036] For any deterministic neural network model Y = F (X | w), where X and Y represent the forecast input and output respectively, and w represents the model weight parameter. In the deterministic model, the weight parameter w can be trained using the back propagation algorithm to obtain a specific value; however, in the probabilistic model, the weight parameter w is no longer a definite single value, but a probability distribution p (w), and the prediction result is no longer a definite single value but a probability density function p (Y | X):
[0037] p(Y|X)=∫p(Y|X,w)p(w)dw (6)
[0038] Where p(Y|X) represents the prediction probability density function corresponding to input X; p(w) is the model weight probability density function; p(Y|X,w) represents the conditional probability density under given input X and model weight w;
[0039] The core of probabilistic model training is to solve the posterior distribution of weight w according to the training set D = [X, Y]:
[0040]
[0041] Where p(w|D) is the posterior distribution of weight w given the training set D; p(D|w) and p(w) are the likelihood and prior respectively; p(D) is not affected by the weight w;
[0042] However, due to the lack of explicitness and multi-dimensional nesting of neural networks, the analytical solution of the posterior distribution p(w|D) is difficult to solve. The posterior distribution p(w|D) is approximated by using a variational distribution q(w|θ) whose function form is known but whose parameters need to be estimated; the optimal variational parameter θ is obtained by minimizing the KL divergence of q(w|θ) and p(w|D) * :
[0043]
[0044] In the formula, θ * represents the optimal parameters of the variational distribution after training; KL[q(w|θ)||p(w|D)] is the KL divergence of q(w|θ) and p(w|D); It is a function to solve the parameters corresponding to the minimum KL divergence;
[0045] KL[q(w|θ)||p(w|D)] cannot be solved directly, and its derivation is as follows:
[0046]
[0047] In the formula, E q(w|θ) (logp(D|w)) is called the log-likelihood expectation; logp(D) is not affected by the weight w;
[0048] Minimizing KL[q(w|θ)||p(w|D)] is equivalent to maximizing the evidence lower bound, which is calculated as follows:
[0049] ELBO=Ε q(w|θ) (logp(D|w))-KL[q(w|θ)||p(w)] (10)
[0050] The first term in ELBO is the expected log-likelihood Ε q(w|θ) (logp(D|w)) can be approximated by Monte Carlo simulation as follows:
[0051]
[0052] In the formula, [X n ,Y n ] is a sample sampled from the training set D; N is the total number of samples; for each sample, a set of weights w is randomly sampled from the variational distribution q(w|θ) n ; F(X n |w n ) is the model in sample X n and w n The predicted value under , F(·) represents the deterministic prediction model; weight Understood as a measure of the true value Y n and predicted values The loss function At this point the first term in ELBO can be calculated;
[0053] The second term in ELBO is the KL divergence of the variational distribution q(w|θ) and the prior p(w), using two mixed Gaussian distributions as the prior distribution p(w):
[0054]
[0055] Among them, λ∈[0,1],σ 1 and σ 2 is a model hyperparameter, which is set before training and does not need to be learned during training; N(w|0,σ 2 ) represents Gaussian distribution;
[0056] The variational distribution uses the following Gaussian distribution with weight parameter θ = (μ, σ):
[0057] q(w|θ)~N(μ,σ) (13)
[0058] w=μ+σ·ε ε~N(0,1) (14)
[0059] Among them, the weight w is not generated directly from θ through the reparameter operation, but is generated based on the reparameters (μ, σ) of θ, the random number ε and the Gaussian distribution; in BConvSCM, the training of the weight w becomes the training of the variational distribution parameters (μ, σ); since p(w) and q(w|θ) have specific function forms and parameters, the second term of ELBO is also computable;
[0060] From this derivation, ELBO can be solved. By maximizing ELBO through a gradient optimization algorithm, the optimal variational parameter θ can be solved. * .
[0061] Preferably, in S7, the prediction process of the BConvSCM probability prediction model is as follows:
[0062] Through model training, the form and parameters of the variational distribution have been determined, and the probability density function of the forecast results can be approximated by the following formula:
[0063] p(Y|X)=∫p(Y|F(X|w))·q(w|θ)dw(15)
[0064] Among them, p(Y|F(X|w)) and q(w|θ) can be calculated, but the analytical solution of p(Y|X) is difficult to obtain through integral derivation. Monte Carlo simulation is used to randomly generate multiple sets of weight variables [w 1 ,w 2 ,…,wn ], and then calculate multiple groups of prediction values
[0065]
[0066] The larger the number of simulations n is, the larger the sample The closer the estimated distribution is to the true distribution p(Y|X); a set of samples The probability density function of can be obtained by kernel density estimation method;
[0067] Each sample The shape of is [T×1×1]. The kernel density estimation method is used to estimate the prediction results of the downstream station flow in each forecast period, and the multi-forecast period probability prediction results of the downstream station runoff can be obtained.
[0068] The present invention has the following beneficial effects:
[0069] 1. This paper proposes a runoff prediction method based on a variational Bayesian convolutional single-control memory network, which uses the maximum shift Pearson correlation coefficient method to analyze the lag time t of the flow evolution from the upstream station to the downstream station. lag , the upstream site flow Q U , tributary station flow Q T , Downstream site flow Q D , the rainfall minus the evaporation variable PE of the interval river section is arranged into a two-dimensional matrix, and each variable is considered to be lag The historical data of each period is combined to form a three-dimensional input tensor; a probability prediction method based on a variational Bayesian framework and a convolutional single-controlled memory neural network is proposed to obtain multi-forecast period probability prediction results of runoff. The method proposed in this invention integrates the evolution of flow and the process of interval flow generation, which is more in line with the formation process of natural basin runoff and is conducive to improving the accuracy of runoff prediction.
[0070] 2. The present invention proposes a convolutional single-controlled memory neural network layer ConvSCM, which converts the four weight variables [w h ,w x ] is reduced to 1, which can shorten the training time; at the same time, the ConvSCM layer optimizes the model structure, which is beneficial to improve the runoff prediction accuracy.
[0071] 3. The time dimension of the three-dimensional tensor input into the model of the present invention takes the historical t lag The data of each period no longer rely on future rainfall information, and can achieve the prediction of future runoff for multiple forecast periods based on historical rainfall and runoff information. The present invention combines ConvSCM with the variational Bayesian framework to propose a new probability prediction model BConvSCM to obtain reliable runoff probability prediction results, which can provide richer information for reservoir scheduling and decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0073] Figure 1 It is a flow chart of a runoff prediction method based on a variational Bayesian convolutional single-controlled memory network provided in an embodiment of the present invention.
[0074] Figure 2 It is a schematic diagram of the ConvSCM layer network structure provided by an embodiment of the present invention.
[0075] Figure 3 It is a schematic diagram of the BConvSCM model provided by an embodiment of the present invention.
[0076] Figure 4 It is a schematic diagram of the study area from the Three Gorges Dam to the Luoshan Hydrological Station in the Yangtze River Basin provided by an embodiment of the present invention.
[0077] Figure 5 A schematic diagram of the model input and output structure provided by an embodiment of the present invention.
[0078] Figure 6 A single forecast period forecast result diagram provided by an embodiment of the present invention.
[0079] Figure 7 A multi-forecast period prediction result diagram provided by an embodiment of the present invention.
[0080] Figure 8 (a)(b)(c) Probability density function diagrams at typical time periods of 50, 500, and 1000 provided by an embodiment of the present invention.
[0081] Fig. 9 A reliability test diagram provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0082] The embodiments of the present invention are further described below in conjunction with the accompanying drawings.
[0083] See also Figure 1-7 The present invention proposes the maximum shift Pearson correlation coefficient method to analyze the lag time of the flow evolution from the upstream site to the downstream site, adopts the three-dimensional tensor based on site-variable-time as the input and output of the prediction model, and proposes a probability prediction method based on the variational Bayesian framework and convolutional single-controlled memory neural network to obtain the multi-forecast period probability prediction results of runoff.
[0084] Attached Figure 1 The flow chart of the runoff prediction method based on the variational Bayesian convolutional single-control memory network is shown, which specifically includes the following steps:
[0085] S1, collects upstream site traffic Q U, tributary station flow Q T , Downstream site flow Q D , data on rainfall P and evaporation E in the river section between upstream and downstream stations;
[0086] S2, calculate the maximum translation Pearson correlation coefficient between the flow of upstream stations, tributary stations and downstream stations respectively. The translation period corresponding to the maximum Pearson correlation coefficient is the lag time tlag of the flow evolution of the station to the downstream station;
[0087] The maximum translation Pearson correlation coefficient calculation formula is:
[0088]
[0089] Where PCC t is the Pearson correlation coefficient between the flow of the upstream site after the t-step shift and the flow of the downstream site without the shift; X 1~(T-t+1 ) represents the upstream station flow samples with time period numbers from 1 to T-t+1, i.e., the independent variable samples; Y t~T represents the flow samples of downstream stations with time period numbers from t to T, i.e., the samples of the dependent variable; and are the means of the independent variable samples and the dependent variable samples respectively; T is the total number of samples;
[0090]
[0091] The translation period t changes from 0 to T-2, and the Pearson correlation coefficient corresponding to each translation period t is calculated. The normal period corresponding to the largest Pearson correlation coefficient is the lag time t of the flow evolution from the upstream site to the downstream site. lag .
[0092] S3, the upstream site traffic Q U , tributary station flow Q T , Downstream site flow Q D , the variables PE of the rainfall P minus the evaporation E of the section are arranged into a two-dimensional matrix, and each variable is considered to have lag The historical data of the time periods constitutes a three-dimensional input tensor with a shape of [T×M×N], where T is the time dimension, which is equal to t lag ; M and N are the number of rows and columns of the two-dimensional matrix in which the above variables are arranged; lag The downstream station flow data of each period is arranged into a three-dimensional output tensor with a shape of [T×1×1];
[0093] S4, prepare the data set according to the three-dimensional input and output format, and divide it into a training set and a test set;
[0094] S5, constructing a convolutional single-controlled memory neural network layer ConvSCM;
[0095] Convolutional single-controlled memory neural network layer ConvSCM, its network structure is as shown in the attached Figure 2 As shown, the steps and calculation formula for the forward propagation of the tth period are:
[0096] net t =w h *h t-1 +w x *x t (3)
[0097] a t =tanh(net t )=tanh(w h *h t-1 +w x *x t ) (4)
[0098]
[0099] In the formula, w h and w x is the weight tensor; h t is the output of the ConvSCM layer in the tth period; net t and a t is the intermediate variable of the tth period; x t represents the input of the t-th period, and its shape is [M×N], where M and N are the number of rows and columns of the two-dimensional matrix after the correlation factor variables are rearranged; the complete input x = [x 1 ,x 2 ,…,x T ], its shape is [T×M×N], T is the number of selected historical periods, equal to t lag ; tanh(·) represents the tanh activation function; the symbols * and Represent the convolution operation and the Hadamard product respectively.
[0100] S6, based on the variational Bayesian framework and the ConvSCM model, a variational Bayesian convolutional single-controlled memory network model BConvSCM is constructed. The BConvSCM model structure is shown in the attached figure. Figure 3 As shown in the figure, it includes an input layer, three hidden layers (two ConvSCM layers and one variational Bayesian VBNN layer) and an output layer. The main parameters of the ConvSCM layer are as follows: kernel size is (2,2), strides is 1, filters is 2, and padding is "same".
[0101] The derivation process of the BConvSCM probability prediction model is as follows:
[0102] For any deterministic neural network model Y = F (X | w), where X and Y represent the forecast input and output respectively, and w represents the model weight parameter. In the deterministic model, the weight parameter w can be trained using the back propagation algorithm to obtain a specific value; however, in the probabilistic model, the weight parameter w is no longer a definite single value, but a probability distribution p (w), and the prediction result is no longer a definite single value but a probability density function p (Y | X):
[0103] p(Y|X)=∫p(Y|X,w)p(w)dw (6)
[0104] Where p(Y|X) represents the prediction probability density function corresponding to input X; p(w) is the model weight probability density function; p(Y|X,w) represents the conditional probability density under given input X and model weight w;
[0105] The core of probabilistic model training is to solve the posterior distribution of weight w according to the training set D = [X, Y]:
[0106]
[0107] Where p(w|D) is the posterior distribution of weight w given the training set D; p(D|w) and p(w) are the likelihood and prior respectively; p(D) is not affected by the weight w;
[0108] However, due to the lack of explicitness and multi-dimensional nesting of neural networks, the analytical solution of the posterior distribution p(w|D) is difficult to solve. The posterior distribution p(w|D) is approximated by using a variational distribution q(w|θ) whose function form is known but whose parameters need to be estimated; the optimal variational parameter θ is obtained by minimizing the KL divergence of q(w|θ) and p(w|D) * :
[0109]
[0110] In the formula, θ * represents the optimal parameters of the variational distribution after training; KL[q(w|θ)||p(w|D)] is the KL divergence of q(w|θ) and p(w|D); It is a function to solve the parameters corresponding to the minimum KL divergence;
[0111] KL[q(w|θ)||p(w|D)] cannot be solved directly, and its derivation is as follows:
[0112]
[0113] In the formula, E q(w|θ) (logp(D|w)) is called the log-likelihood expectation; logp(D) is not affected by the weight w;
[0114] Minimizing KL[q(w|θ)||p(w|D)] is equivalent to maximizing the evidence lower bound (ELBO), which is calculated as follows:
[0115] ELBO=Ε q(w|θ) (logp(D|w))-KL[q(w|θ)||p(w)] (10)
[0116] The first term in ELBO is the expected log-likelihood Ε q(w|θ) (logp(D|w)) can be approximated by Monte Carlo simulation as follows:
[0117]
[0118] In the formula, [X n ,Y n ] is a sample sampled from the training set D; N is the total number of samples; for each sample, a set of weights w is randomly sampled from the variational distribution q(w|θ) n ; F(X n |w n ) is the model in sample X n and w n The predicted value under , F(·) represents the deterministic prediction model; weight Understood as a measure of the true value Y n and predicted values The loss function At this point the first term in ELBO can be calculated;
[0119] The second term in ELBO is the KL divergence of the variational distribution q(w|θ) and the prior p(w), using two mixed Gaussian distributions as the prior distribution p(w):
[0120]
[0121] Among them, λ∈[0,1],σ 1 and σ 2 is a model hyperparameter, which is set before training and does not need to be learned during training; N(w|0,σ 2 ) represents Gaussian distribution;
[0122] The variational distribution uses the following Gaussian distribution with weight parameter θ = (μ, σ):
[0123] q(w|θ)~N(μ,σ) (13)
[0124] w=μ+σ·ε ε~N(0,1) (14)
[0125] Among them, through the re-parameter operation, the weight w is not directly generated according to θ, but is generated according to the re-parameters (μ, σ) of θ, the random number ε and the Gaussian distribution; in BConvSCM, the training of the weight w becomes the training of the variational distribution parameters (μ, σ); since p(w) and q(w|θ) have specific function forms and parameters, the second term of ELBO is also computable;
[0126] From this derivation, ELBO can be solved. By maximizing ELBO through a gradient optimization algorithm, the optimal variational parameter θ can be solved. * .
[0127] S7, train the parameters of the BConvSCM model on the training set, predict the flow of downstream sites on the test set, and obtain high-precision and reliable probabilistic prediction results of runoff for multiple forecast periods.
[0128] The prediction process of the BConvSCM probability prediction model is as follows:
[0129] Through model training, the form and parameters of the variational distribution have been determined, and the probability density function of the forecast results can be approximated by the following formula:
[0130] p(Y|X)=∫p(Y|F(X|w))·q(w|θ)dw (15)
[0131] Among them, p(Y|F(X|w)) and q(w|θ) can be calculated, but the analytical solution of p(Y|X) is difficult to obtain through integral derivation. Monte Carlo simulation is used to randomly generate multiple sets of weight variables [w 1 ,w 2 ,…,w n ], and then calculate multiple groups of prediction values
[0132]
[0133] The larger the number of simulations n is, the larger the sample The closer the estimated distribution is to the true distribution p(Y|X); a set of samples The probability density function of can be obtained by kernel density estimation method;
[0134] Each sample The shape of is [T×1×1]. The kernel density estimation method is used to estimate the prediction results of the downstream station flow in each forecast period, and the multi-forecast period probability prediction results of the downstream station runoff can be obtained.
[0135] Embodiment 2:
[0136] The application of the present invention is further described below in conjunction with specific experiments.
[0137] The present invention takes the area from the Three Gorges Dam to the Luoshan Hydrological Station in the middle and lower reaches of the Yangtze River as the object. The upstream flow in this area is the outflow of the Three Gorges Dam, the tributary flow includes the outflow of the Gaobazhou Dam on the Qingjiang River and the flow of Dongting Lake, and the downstream flow to be predicted is the flow of the Luoshan Hydrological Station. The data is from 2:00 on January 1, 2018 to 2:00 on January 1, 2021. The step length of each period is 6 hours, and the total number of samples is 4383. The data from 2018 to 2020 is used as the training set, and the data from 2021 is used as the test set.
[0138] (a) Lag analysis:
[0139] The outflow from the Three Gorges Dam, the outflow from the Gaobazhou Dam, the flow from Dongting Lake, the flow from Luoshan, the surface rainfall between the Three Gorges Dam and the Luoshan hydrological station, and the evaporation between the Three Gorges Dam and the Luoshan hydrological station are represented by the variables TGD_Q, GBZD_Q, DTL_Q, LS_Q, TGD-LS_P, and TGD-LS_E, respectively. The Pearson correlation coefficient PCC between the outflow from the Three Gorges Dam TGD_Q, the outflow from the Gaobazhou Dam GBZD_Q, the flow from Dongting Lake DTL_Q, the flow from Luoshan LS_Q and the untranslated flow from Luoshan LS_Q in each translation period is calculated, as shown in Table 1. The translation period corresponding to the largest PCC of the Three Gorges Dam outflow TGD_Q and the Gaobazhou Dam outflow GBZD_Q is 12, which means that it takes 12 periods (about 72 hours) for the flows of the Three Gorges Dam and Gaobazhou Dam to evolve to Luoshan Station. The translation period corresponding to the largest PCC of the Dongting Lake flow DTL_Q is 1, which means that it takes less than 1 period (less than 6 hours) for the flow of Dongting Lake to evolve to Luoshan Station.
[0140] Table 1 PCC values between TGD_Q, GBZD_Q, DTL_Q, LS_Q and the unshifted LS_Q in each shift period
[0141]
[0142] (b) Model input and output structure:
[0143] Based on the above time lag analysis, the Three Gorges Dam is the farthest from Luoshan Station, and the maximum time lag is 12. The input and output structure of the model of this embodiment is shown in Figure (5). The input shape is [12×2×3], and the five variables TGD-LS_P-TGD-LS_E, TGD_Q, GBZD_Q, DTL_Q, and LS_Q are rearranged into a [2×3] matrix (the last variable is vacant), and each variable considers the historical data of the past 12 time periods. The output shape is [12×1×1], where [1×1] is the flow of Luoshan Station, and the data of the next 12 time periods are output.
[0144] (c) Comparison of deterministic prediction and probabilistic prediction results:
[0145] In order to verify the prediction performance of the model proposed in this invention, the BConvSCM model is compared with the BConvLSTM, BConvGRU, BConv2D, MSK-XAJ and other models:
[0146] BConvSCM model: Its model structure and parameters are shown in Figure (3);
[0147] BConvLSTM model: replace ConvSCM in BConvSCM with ConvLSTM, and the rest is the same as BConvSCM;
[0148] BConvGRU model: replace ConvSCM in BConvSCM with ConvGRU, and the rest is the same as BConvSCM;
[0149] BConv2D model: replace ConvSCM in BConvSCM with Conv2D, and the rest is the same as BConvSCM;
[0150] MSK-XAJ model: This is a traditional conceptual hydrological model that uses the Muskingum model for river channel evolution and the Xin'anjiang model for runoff forecasting.
[0151] The coefficient of certainty (R 2 ) is used to evaluate the accuracy of deterministic prediction, and the continuous graded probability score (CRPS) is used to evaluate the comprehensive performance of probability prediction. The comparison index tables are shown in Table 2 and Table 3 respectively. 2 ) is closer to 1, the better, and the smaller the CRPS, the better. Whether in the deterministic prediction comparison index table or the probability prediction comparison index table, BConvSCM is the best among all models in most of the 12 forecast periods, verifying that the model BConvSCM proposed in the present invention has high deterministic prediction accuracy and good probability prediction comprehensive performance.
[0152] Table 2 Deterministic prediction indicators (R 2 )
[0153]
[0154] Table 3 Multi-forecast probability prediction index (CRPS) of each model
[0155]
[0156] (d) Training time comparison:
[0157] The training time of three similar models, BConvSCM, BConvLSTM and BConvGRU, is shown in Table 4. The training time of BConvSCM is 523s, which shortens the training time of BConvLSTM by 63%, indicating the performance of BConvSCM proposed by the present invention in shortening the training time.
[0158] Table 4. BConvSCM, BConvLSTM and BConvGRU training schedule
[0159]
[0160] (e) Single-forecast period and multi-forecast period forecast results:
[0161] The prediction results of each model in a single forecast period and a typical period of multiple forecast periods are plotted in Figures (6) and (7). By comparing the figures, whether in the single forecast period prediction or the multiple forecast period prediction, the prediction line of BConvSCM is closest to the observation value line, indicating that the prediction accuracy of the model proposed in the present invention is the highest.
[0162] (f) Probability density function graph:
[0163] The probability density function of the BConvSCM model in the typical time periods of 50, 500, and 1000 is shown in Figure (8). It can be seen from the figure that the probability density function is relatively appropriate, the predicted mean line is close to the observed value line, and the probability density function line is not too high, too low, too wide, or too narrow.
[0164] (g) Reliability verification diagram:
[0165] The reliability verification PIT diagram of the probability prediction results of the BConvSCM model is shown in Figure (9). It can be seen from the figure that the PIT values are all within the 5% Kolmogorov confidence band, indicating that the PIT values obey a uniform distribution, thereby verifying that the probability prediction results of the BConvSCM model are reliable.
Claims
1. A runoff prediction method based on variational Bayesian convolutional single-control memory network, characterized in that: The runoff prediction method analyzes the lag time of flow evolution from upstream stations to downstream stations based on the maximum translation Pearson correlation coefficient method. t lag , using the three-dimensional tensor based on site-variable-time as the input and output of the prediction model, a probabilistic prediction method based on the variational Bayesian framework and convolutional single-controlled memory neural network is proposed to obtain the multi-forecast period probabilistic prediction results of runoff; The runoff prediction method specifically comprises the following steps: S1, collects traffic from upstream sites Q U , tributary station flow Q T , Downstream site traffic Q D , rainfall in the river section between upstream and downstream stations P Evaporation E data; S2, calculate the maximum translation Pearson correlation coefficient between the flow of upstream stations, tributary stations and downstream stations respectively. The translation period corresponding to the maximum Pearson correlation coefficient is the lag time of the flow evolution of the station to the downstream station. t lag ; S3, which transfers traffic from upstream sites Q U , tributary station flow Q T , Downstream site traffic Q D , rainfall in the river section P Subtract evaporation E Variables P - E Arranged into a two-dimensional matrix, each variable is considered t lag The historical data of the time periods constitutes a three-dimensional input tensor with the shape of ,in T is the time dimension, equal to t lag ; M and N Arrange the above variables into the number of rows and columns of a two-dimensional matrix; t lag The downstream station flow data for each period is arranged into a three-dimensional output tensor with the shape of ; S4, prepare the data set according to the three-dimensional input and output format, and divide it into a training set and a test set; S5, constructing a convolutional single-controlled memory neural network layer ConvSCM; S6, based on the variational Bayesian framework and the ConvSCM model, a variational Bayesian convolutional single-controlled memory network model BConvSCM is constructed; S7, train the parameters of the BConvSCM model on the training set, predict the flow of downstream stations on the test set, and obtain high-precision and reliable probabilistic prediction results of runoff for multiple forecast periods; In S2, the maximum translation Pearson correlation coefficient calculation formula is: (1) In the formula, It is through t Pearson correlation coefficient between the flow of upstream stations after step period shift and the flow of downstream stations without shift; The representative time period is numbered from 1 to The upstream site flow sample is the independent variable sample; Represents the period number from t to T The downstream station flow sample is the dependent variable sample; and are the means of the independent variable samples and the dependent variable samples respectively; T is the total sample size; (2) Shift period t From 0 to T -2, calculate each translation period t The corresponding Pearson correlation coefficient. The largest Pearson correlation coefficient corresponds to the normal period, which is the lag time of the flow evolution from the upstream site to the downstream site. ; In S5, the convolutional single-controlled memory neural network layer ConvSCM t The steps and calculation formula for forward propagation in each period are as follows: (3) (4) (5) In the formula, and is the weight tensor; It is t The output of the ConvSCM layer in each period; and It is t The intermediate variables of the time period; Representative t The input of time periods has the shape of , M and N The number of rows and columns for the relevant factor variables to be rearranged into a two-dimensional matrix; the complete input , whose shape is , T is the number of historical periods selected, equal to t lag ; Represents the tanh activation function; symbol and Represent the convolution operation and the Hadamard product respectively.
2. According to claim 1, a runoff prediction method based on variational Bayesian convolutional single-controlled memory network is characterized in that: In S6, the derivation process of the BConvSCM probability prediction model is as follows: For any deterministic neural network model ,in and Represent the forecast input and output respectively, Represents the model weight parameter. In a deterministic model, the weight parameter The back propagation algorithm can be used to train the specific value; however, in the probability model, the weight parameter It is no longer a definite single value, but a probability distribution At the same time, the prediction result is no longer a definite single value but a probability density function : (6) In the formula, Represents input The corresponding forecast probability density function; is the model weight probability density function; Represents a given input and model weights The conditional probability density under ; The core of probabilistic model training is to Solving weights The posterior distribution of : (7) In the formula, Given a training set Lower weight The posterior distribution of and are likelihood and prior respectively; Not weighted The impact of However, due to the lack of explicitness of the form of neural networks and the existence of multi-dimensional nesting, the posterior distribution The analytical solution of is difficult to solve, and we can use a variational distribution whose function form is known but whose parameters need to be estimated. To approximate the posterior distribution By minimizing and The KL divergence is used to obtain the optimal variational parameters : (8) In the formula, Represents the optimal parameters of the variational distribution after training; yes and KL divergence of It is a function to solve the parameters corresponding to the minimum KL divergence; It cannot be solved directly, and its derivation is as follows: (9) In the formula, It is called the log-likelihood expectation; Not weighted Influence; minimize It is equivalent to maximizing the lower bound of evidence, and the calculation formula is as follows: (10) The first log-likelihood expectation in ELBO It can be approximated by Monte Carlo simulation as follows: (11) In the formula, From the training set Samples sampled in is the total number of samples; for each sample, from the variational distribution A set of weights is generated by random sampling ; The model is in the sample and The predicted value under represents a deterministic prediction model; weight Understood as a measure of true value and predicted values The loss function ; Now the first term in ELBO can be calculated; The second term in ELBO is the variational distribution and prior KL divergence, using two mixed Gaussian distributions as prior distributions : (12) in, , and It is a model hyperparameter, which is set before training and does not need to be learned during training; represents Gaussian distribution; The variational distribution adopts the following Gaussian distribution, with the weight parameter : (13) (14) Among them, the weights are manipulated by re-parameterization Not directly based on generated, but based on The weight parameter , random numbers and Gaussian distribution to generate; in BConvSCM, the weight The training becomes the variational distribution parameter training; and All have specific function forms and parameters, so the second term of ELBO is also computable; From this derivation, ELBO can be solved. By maximizing ELBO through a gradient optimization algorithm, the optimal variational parameters can be solved. .
3. According to claim 2, a runoff prediction method based on variational Bayesian convolutional single-controlled memory network is characterized in that: In S7, the prediction process of the BConvSCM probability prediction model is as follows: Through model training, the form and parameters of the variational distribution have been determined, and the probability density function of the forecast results can be approximated by the following formula: (15) in and It can be calculated, but The analytical solution of is difficult to obtain through integral derivation, so Monte Carlo simulation is used to obtain the variational distribution Randomly generate multiple sets of weight variables , and then calculate multiple groups of prediction values : (16) Number of simulations The larger the sample The closer the estimated distribution is to the true distribution A set of samples The probability density function of can be obtained by kernel density estimation method; Each sample The shapes are The kernel density estimation method is used to estimate the prediction results of the flow at the downstream site in each forecast period, and the multi-forecast period probability prediction results of the runoff at the downstream site can be obtained.