Monthly runoff sequence interval prediction method

By introducing the CNN-BiLSTM model and the crown porcupine optimization algorithm in the month runoff prediction model, combining the extreme difference segmentation method and non-parametric nuclear density estimation method, the problems of low accuracy of month runoff prediction and difficulty in characterizing uncertainty are solved, and high-precision month runoff interval prediction is achieved.

CN119990435APending Publication Date: 2025-05-13CHINA YANGTZE POWER
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510080027.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-19
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing monthly runoff prediction model has low prediction accuracy and is difficult to effectively characterize the uncertainty and fluctuation interval of runoff.

Method used

The CNN-BiLSTM model was used to combine the crown porcupine optimization algorithm (CPO) optimization parameters to establish a CPO-CNN-BiLSTM point prediction model, and the interval prediction model was constructed through extreme difference segmentation method and non-parametric kernel density estimation (NKDE) method to obtain the confidence interval of monthly runoff.

Benefits of technology

The accuracy and convergence rate of monthly runoff prediction can effectively characterize the uncertainty and fluctuation range of runoff, helping decision makers better understand and cope with data uncertainty and changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990435A_ABST
    Figure CN119990435A_ABST
Patent Text Reader

Abstract

The invention discloses a monthly runoff sequence interval prediction method which comprises the following steps: S1, monthly runoff data preparation: collecting monthly runoff data, preprocessing the data, and dividing the data into training data and test data; s2, establishing a CNN-BiLSTM point prediction model which comprises an input layer, a CNN layer, a BiLSTM layer and an output layer, and inputting the training data into the CNN-BiLSTM point prediction model to obtain a monthly runoff point prediction result; s3, a CPO-CNN-BiLSTM point prediction model is established; s4, a CPO-CNN-BiLSTM-NKDE interval prediction model is established, and the CPO-CNN-BiLSTM Aiming at the problem that hyper-parameters are difficult to select in the CNN-BiLSTM point prediction model, the CPO is used for optimizing the hidden layer node number, the initial learning rate and the regularization coefficient in the CNN-BiLSTM model, so that the optimal CPO-CNN-BiLSTM point prediction model is established, and the convergence rate and the accuracy of monthly runoff prediction are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of hydrological forecasting, in particular to a monthly runoff sequence interval prediction method. Background Art

[0002] Runoff prediction is not only related to flood prevention and disaster reduction, rational allocation of water resources, etc., but also has important value in meeting the needs of irrigation, power generation, etc. However, the complexity of runoff sequence in time and space has brought challenges to the accurate prediction of runoff. In order to improve the accuracy of runoff prediction, in recent years, scholars have proposed monthly runoff point prediction models based on machine learning methods such as BP neural network, artificial neural network, extreme learning machine, least squares support vector machine (LSSVM), etc., which effectively improved the accuracy of monthly runoff prediction. With the rapid development of data science and big data technology, deep learning methods such as convolutional neural network (CNN) and long short-term memory (LSTM) network have shown good capabilities in feature extraction and data processing. Some scholars have constructed monthly runoff point prediction models based on deep learning methods such as LSTM network and CNN. The experimental results show that the prediction accuracy of LSTM network and CNN is better than that of LSSVM and other models. Considering the strong spatial feature extraction capability of CNN and the strong temporal feature extraction capability of bidirectional long short-term memory (BiLSTM) network, some scholars have constructed a monthly runoff point prediction model based on CNN and BiLSTM network (CNN-BiLSTM point prediction model), whose prediction accuracy is better than that of single CNN model and BiLSTM network model.

[0003] However, the determined value obtained by point prediction is difficult to characterize the uncertainty of runoff, and cannot describe the probability of the predicted value and its possible fluctuation range. Interval prediction not only gives the runoff point prediction result, but also depicts the fluctuation range of runoff at a certain confidence level, which is more conducive to the rational arrangement of water resources allocation and the formulation of hydropower generation plans. The interval prediction model is generally implemented on the basis of the point prediction model combined with parametric or non-parametric density estimation methods. Among them, the non-parametric kernel density estimation (NKDE) does not need to assume the distribution of sample data in advance, and can estimate the probability density function only with the sample data itself, which has strong adaptability.

[0004] In addition, the prediction accuracy of the CNN-BiLSTM model depends on the selection of parameters such as the number of hidden layer units, learning rate and regularization coefficient of the BiLSTM network. The setting of the window width of the kernel function in NKDE affects the result of kernel density estimation. The Crested Porcupine Optimizer (CPO) is an intelligent optimization algorithm proposed by Abdel-Basset M et al. in 2024 to simulate the different defense strategies of crested porcupines when encountering predators. The CPO algorithm is suitable for solving various large-scale optimization problems. Summary of the invention

[0005] The purpose of the present invention is to overcome the above-mentioned shortcomings and provide a monthly runoff sequence interval prediction method, aiming to solve the problem of low prediction accuracy of the monthly runoff prediction model.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: a monthly runoff sequence interval prediction method, comprising the following steps:

[0007] S1: Monthly runoff data preparation: Collect monthly runoff data, pre-process the data, and divide the data into training data and test data;

[0008] S2: Establish a CNN-BiLSTM point prediction model: including: input layer, CNN layer, BiLSTM layer and output layer, input the training data into the CNN-BiLSTM point prediction model to obtain the monthly runoff point prediction results;

[0009] S3: Establish CPO-CNN-BiLSTM point prediction model: Use the crown porcupine optimization algorithm CPO to optimize the parameters of the CNN-BiLSTM point prediction model, bring the optimal parameters into the CNN-BiLSTM point prediction model for training, and obtain the trained monthly runoff prediction point model, namely the CPO-CNN-BiLSTM point prediction model; use the CPO-CNN-BiLSTM point prediction model to obtain the monthly runoff point prediction results;

[0010] S4: Establish a CPO-CNN-BiLSTM-NKDE interval prediction model: Use the extreme difference segmentation method to sort the point prediction results and divide them into low flow segment, medium flow segment and high flow segment, and then use the CPO optimized window width non-parametric kernel density estimation method NKDE to estimate the probability distribution of the prediction value error in each segment, and use the cubic spline interpolation method to perform curve fitting to obtain the quantile of each segment. Finally, superimpose the point prediction results and the quantile of the segment to which the point prediction results belong to obtain the interval prediction results of monthly runoff.

[0011] Preferably, the specific steps of S1 include:

[0012] S1.1: Obtain historical monthly runoff data from hydrological stations;

[0013] S1.2: Normalize the monthly runoff data. The normalization process is expressed as:

[0014]

[0015] In formula (1), x represents the measured value of monthly runoff, x * represents the normalized monthly runoff, x min ,x maxThey represent the minimum and maximum values ​​of the measured monthly runoff respectively;

[0016] S1.3: Determine the input dimension of the prediction model according to the delay time m of the partial autocorrelation coefficient of the monthly runoff; that is, the input corresponding to the monthly runoff output of the i-th month is the monthly runoff of the i-1th, i-2th, ..., imth months;

[0017] S1.4: Divide the normalized dataset into a training sample set and a test sample set.

[0018] More preferably, the specific steps of S2 include:

[0019] S2.1: Input the training sample set obtained in S1 into the CNN-BiLSTM point prediction model as the input of the model;

[0020] S2.2: Construct convolutional neural network layer CNN:

[0021] The CNN first extracts features from the monthly runoff data input by S2.1 through a one-dimensional convolution layer, with the number of convolution kernels being 32 and the step length being 1; then, after passing through two pooling layers and a smoothing layer, the data is filtered and flattened respectively, and feature information is output;

[0022] S2.3: Construct a bidirectional long short-term memory network layer BiLSTM:

[0023] The BiLSTM layer is composed of forward and backward LSTMs. The output feature information in S2.2 is input into the BiLSTM. After the feature information is enhanced and integrated by the Dropout layer and the fully connected layer, the prediction result is output.

[0024] S2.4: The predicted value u output by the CNN-BiLSTM point prediction model is denormalized to obtain the monthly runoff prediction value.

[0025] Preferably, the specific steps of S3 include:

[0026] S3.1: Initialize the parameters of CPO and CNN-BILSTM point prediction models, including the number of crested porcupine populations, the maximum number of iterations, the upper and lower boundaries of the solution space, and the dimension of the solution; the parameters optimized by CPO include the number of hidden layer nodes, the initial learning rate, and the regularization coefficient. The number of hidden layer nodes, the initial learning rate, and the regularization coefficient of the CNN-BILSTM point prediction model are used as the three-dimensional coordinates of the individual positions of the crested porcupine group, and the individual positions of the crested porcupines are randomly initialized;

[0027] S3.2: Calculate the fitness value of each crested porcupine individual;

[0028] S3.3: Update the fitness value of the crowned porcupine in the crowned porcupine group; calculate the information of the crowned porcupine with the best fitness from the first iteration to the current position, and save the information;

[0029] S3.4: Update the position vector of the crested porcupine with the best fitness, and calculate the new position vector of the crested porcupine group after the next iteration;

[0030] S3.5: Determine whether the maximum number of iterations has been reached. If so, end the iteration and return the optimal parameters; otherwise, return to S3.2;

[0031] S3.6: Bring the optimal parameters into the CNN-BiLSTM point prediction model for training, so as to obtain the trained monthly runoff prediction model CPO-CNN-BiLSTM point prediction model;

[0032] S3.7: Use the CPO-CNN-BiLSTM point prediction model to predict and obtain the monthly runoff point prediction result P pred .

[0033] Preferably, the specific steps of S4 include:

[0034] S4.1: Error distribution estimation of runoff sequence: After arranging the predicted values ​​of the monthly runoff training data from small to large, the extreme difference segmentation method is used to divide them into low flow segment, medium flow segment and high flow segment, and the relative error of each predicted value in the three segments is calculated; the window width h of NKDE is initialized, and the error probability density function and probability distribution function F(ε) of each segment are calculated by NKDE, where ε is the random variable of the prediction error; the probability distribution function of the three segments is curve fitted by cubic spline interpolation method, and the quantile corresponding to each segment is calculated; determine which segment the monthly runoff point prediction result in the training data belongs to, and according to the quantile corresponding to the segment, use formula (2) to calculate: When the significance level is α, the confidence interval with a confidence level of 1-α is:

[0035]

[0036] In formula (2): is the inverse function of F(ε), α1=α / 2, α2=1-α / 2;

[0037] S4.2: Use CPO to optimize the window width h and establish the CPO-CNN-BiLSTM-NKDE interval prediction model; recalculate the quantile corresponding to each segment according to the density distribution function after optimizing the window width h;

[0038] S4.3: Monthly runoff sequence interval prediction: Determine which segment each monthly runoff point prediction result obtained in S3 belongs to, and superimpose the point prediction result P according to the quantile corresponding to the segment obtained in S4.2.pred The confidence interval of monthly runoff is obtained by using the quantiles.

[0039] Beneficial effects of the present invention:

[0040] (1) Aiming at the problem of feature extraction of different information when historical monthly runoff data is input, the present invention adopts a CNN-BiLSTM model consisting of 1 convolutional layer, 2 pooling layers, 1 smoothing layer, 1 BiLSTM layer, 1 Dropout layer and a fully connected layer to extract the temporal feature information in the historical monthly runoff data, which effectively reduces the complexity of feature extraction;

[0041] (2) In view of the difficulty in selecting hyperparameters in the CNN-BiLSTM point prediction model, the present invention proposes to use CPO to optimize the number of hidden layer nodes, initial learning rate, and regularization coefficient in the CNN-BiLSTM model, thereby establishing an optimal CPO-CNN-BiLSTM point prediction model, which greatly improves the convergence rate and accuracy of monthly runoff prediction;

[0042] (3) Aiming at the problem that the determined value obtained by point prediction is difficult to characterize the uncertainty of monthly runoff, the present invention proposes to first use the extreme difference segmentation method to sort the point prediction results and divide them into low flow segment, medium flow segment and high flow segment, and then use the CPO optimization window width non-parametric kernel density estimation method to obtain the interval prediction results of monthly runoff. The interval prediction results can help decision makers better understand and deal with the uncertainty and changes of data, which is of great significance for assessing runoff risks and formulating flood control and drought relief strategies.

[0043] (4) The method of the present invention first preprocesses the monthly runoff data, establishes a CNN-BiLSTM point prediction model, and uses the CPO optimization model parameters to obtain a deterministic monthly runoff point prediction result. The NKDE with CPO optimized window width is further used to construct an interval prediction model to obtain the monthly runoff interval prediction result with a given confidence level, thereby solving the problem of low prediction accuracy of the monthly runoff prediction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a flowchart of a monthly runoff series interval prediction method;

[0045] Figure 2 It is the monthly runoff series curve;

[0046] Figure 3 This is the structure diagram of the CNN-BiLSTM model;

[0047] Figure 4 It is the monthly runoff sequence point prediction curve;

[0048] Figure 5 This is the prediction result of the monthly runoff series interval prediction model. DETAILED DESCRIPTION

[0049] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0050] Example 1: Figure 1 As shown, a monthly runoff series interval prediction method includes the following steps:

[0051] S1: Monthly runoff data preparation: Collect monthly runoff data, pre-process the data, and divide the data into training data and test data;

[0052] S2: Establish a CNN-BiLSTM point prediction model: including: input layer, CNN layer, BiLSTM layer and output layer, input the training data into the CNN-BiLSTM point prediction model to obtain the monthly runoff point prediction results;

[0053] S3: Establish CPO-CNN-BiLSTM point prediction model: Use the crown porcupine optimization algorithm CPO to optimize the parameters of the CNN-BiLSTM point prediction model, bring the optimal parameters into the CNN-BiLSTM point prediction model for training, and obtain the trained monthly runoff prediction point model, namely the CPO-CNN-BiLSTM point prediction model; use the CPO-CNN-BiLSTM point prediction model to obtain the monthly runoff point prediction results;

[0054] S4: Establish a CPO-CNN-BiLSTM-NKDE interval prediction model: Use the extreme difference segmentation method to sort the point prediction results and divide them into low flow segment, medium flow segment and high flow segment, and then use the CPO optimized window width non-parametric kernel density estimation method NKDE to estimate the probability distribution of the prediction value error in each segment, and use the cubic spline interpolation method to perform curve fitting to obtain the quantile of each segment. Finally, superimpose the point prediction results and the quantile of the segment to which the point prediction results belong to obtain the interval prediction results of monthly runoff.

[0055] Preferably, the specific steps of S1 include:

[0056] S1.1: Obtain historical monthly runoff data from hydrological stations;

[0057] S1.2: Normalize the monthly runoff data. The normalization process is expressed as:

[0058]

[0059] In formula (1), x represents the measured value of monthly runoff, x * represents the normalized monthly runoff, x min ,x max They represent the minimum and maximum values ​​of the measured monthly runoff respectively;

[0060] S1.3: Determine the input dimension of the prediction model according to the delay time m of the partial autocorrelation coefficient of the monthly runoff; that is, the input corresponding to the monthly runoff output of the i-th month is the monthly runoff of the i-1th, i-2th, ..., imth months;

[0061] S1.4: Divide the normalized dataset into a training sample set and a test sample set.

[0062] More preferably, the specific steps of S2 include:

[0063] S2.1: Input the training sample set obtained in S1 into the CNN-BiLSTM point prediction model as the input of the model;

[0064] S2.2: Construct convolutional neural network layer CNN:

[0065] The CNN first extracts features from the monthly runoff data input by S2.1 through a one-dimensional convolution layer, with the number of convolution kernels being 32 and the step length being 1; then, after passing through two pooling layers and a smoothing layer, the data is filtered and flattened respectively, and feature information is output;

[0066] S2.3: Construct a bidirectional long short-term memory network layer BiLSTM:

[0067] The BiLSTM layer is composed of forward and backward LSTMs. The output feature information in S2.2 is input into the BiLSTM. After the feature information is enhanced and integrated by the Dropout layer and the fully connected layer, the prediction result is output.

[0068] S2.4: The predicted value u output by the CNN-BiLSTM point prediction model is denormalized to obtain the monthly runoff prediction value.

[0069] Preferably, the specific steps of S3 include:

[0070] S3.1: Initialize the parameters of CPO and CNN-BILSTM point prediction models, including the number of crested porcupine populations, the maximum number of iterations, the upper and lower boundaries of the solution space, and the dimension of the solution; the parameters optimized by CPO include the number of hidden layer nodes, the initial learning rate, and the regularization coefficient. The number of hidden layer nodes, the initial learning rate, and the regularization coefficient of the CNN-BILSTM point prediction model are used as the three-dimensional coordinates of the individual positions of the crested porcupine group, and the individual positions of the crested porcupines are randomly initialized;

[0071] S3.2: Calculate the fitness value of each crested porcupine individual;

[0072] S3.3: Update the fitness value of the crowned porcupine in the crowned porcupine group; calculate the information of the crowned porcupine with the best fitness from the first iteration to the current position, and save the information;

[0073] S3.4: Update the position vector of the crested porcupine with the best fitness, and calculate the new position vector of the crested porcupine group after the next iteration;

[0074] S3.5: Determine whether the maximum number of iterations has been reached. If so, end the iteration and return the optimal parameters; otherwise, return to S3.2;

[0075] S3.6: Bring the optimal parameters into the CNN-BiLSTM point prediction model for training, so as to obtain the trained monthly runoff prediction model CPO-CNN-BiLSTM point prediction model;

[0076] S3.7: Use the CPO-CNN-BiLSTM point prediction model to predict and obtain the monthly runoff point prediction result P pred .

[0077] Preferably, the specific steps of S4 include:

[0078] S4.1: Error distribution estimation of runoff sequence: After arranging the predicted values ​​of the monthly runoff training data from small to large, the extreme difference segmentation method is used to divide them into low flow segment, medium flow segment and high flow segment, and the relative error of each predicted value in the three segments is calculated; the window width h of NKDE is initialized, and the error probability density function and probability distribution function F(ε) of each segment are calculated by NKDE, where ε is the random variable of the prediction error; the probability distribution function of the three segments is curve fitted by cubic spline interpolation method, and the quantile corresponding to each segment is calculated; determine which segment the monthly runoff point prediction result in the training data belongs to, and according to the quantile corresponding to the segment, use formula (2) to calculate: When the significance level is α, the confidence interval with a confidence level of 1-α is:

[0079]

[0080] In formula (2): is the inverse function of F(ε), α1=α / 2, α2=1-α / 2;

[0081] S4.2: Use CPO to optimize the window width h and establish the CPO-CNN-BiLSTM-NKDE interval prediction model; recalculate the quantile corresponding to each segment according to the density distribution function after optimizing the window width h;

[0082] S4.3: Monthly runoff sequence interval prediction: Determine which segment each monthly runoff point prediction result obtained in S3 belongs to, and superimpose the point prediction result P according to the quantile corresponding to the segment obtained in S4.2. pred The confidence interval of monthly runoff is obtained by using the quantiles.

[0083] Example 2: Figure 1 As shown, a monthly runoff series interval prediction method includes the following steps:

[0084] S1: Monthly runoff data preparation: Collect monthly runoff data, pre-process the data, and divide the data into training data and test data.

[0085] Specifically, the embodiment of the present invention uses the monthly runoff data of a river from 1959 to 2018 to illustrate the proposed method in detail. The Laida criterion is used to process the abnormalities of the monthly runoff data, and the interpolation method is used to supplement the missing values ​​to obtain 720 data. The change of monthly runoff over time is shown in Figure 2 shown.

[0086] The monthly runoff data from 1959 to 2008 (a total of 600 data, numbered 1 to 600) were selected as training data, and the monthly runoff data from 2009 to 2018 (a total of 120 data, numbered 601 to 720) were selected as test data. The delay time of the partial autocorrelation coefficient of the monthly runoff is 9, so the input dimension of the monthly runoff prediction model is 9, that is, the monthly runoff output of the i-th month (i = 601, 602, ..., 720) corresponds to the monthly runoff of the i-1, i-2, ..., i-9 months.

[0087] S2: Establish a CNN-BiLSTM point prediction model. It includes: input layer, CNN layer, BiLSTM layer and output layer. Input the training data into the CNN-BiLSTM point prediction model to obtain the actual monthly runoff point prediction results.

[0088] Specifically, based on the spatial feature extraction capability of CNN and the bidirectional temporal feature extraction capability of BiLSTM, CNN was used to extract the feature vector of monthly runoff, which was then input into the BiLSTM network for bidirectional loop training. The CNN-BiLSTM model was constructed. The model includes 1 convolutional layer, 2 pooling layers, 1 smoothing layer, 1 BiLSTM layer, 1 Dropout layer and a fully connected layer. Its structure is shown in the figure below. Figure 3 shown.

[0089] S3: Establish a CPO-CNN-BiLSTM point prediction model. Use the crown porcupine optimization algorithm (CPO) to optimize the parameters of the CNN-BiLSTM point prediction model, bring the optimal parameters into the CNN-BiLSTM point prediction model for training, and obtain a trained monthly runoff prediction point model (CPO-CNN-BiLSTM point prediction model). Use the CPO-CNN-BiLSTM point prediction model to obtain the monthly runoff point prediction results.

[0090] Specifically, the mean absolute percentage error (MAPE) is selected as the fitness value calculation function of CPO, that is,

[0091]

[0092] In the formula, Num is the total number of samples, t * and * are the measured and predicted values ​​of monthly runoff, respectively.

[0093] For the original runoff time series, CPO is used to optimize the number of hidden layer nodes, initial learning rate, and regularization coefficient of the prediction accuracy of the CNN-BiLSTM model. It is found that when the three parameters are 22, 0.0073, and 0.00010265, respectively, the corresponding fitness values ​​are optimal.

[0094] The CPO-CNN-BiLSTM point prediction model proposed by this method is compared with the CPO-LSSVM model, CPO-KELM model, CPO-LSTM model and CPO-BiLSTM model. In the CPO-LSSVM model, CPO is used to optimize the coefficients and penalty coefficients of the Gaussian kernel function in LSSVM; in the CPO-KELM model, CPO is used to optimize the kernel parameters and penalty coefficients of KELM; in the CPO-LSTM model and CPO-BiLSTM model, CPO is used to optimize the number of hidden layer nodes, initial learning rate and regularization coefficient of LSTM or BiLSTM. The point prediction curves of the five prediction models are shown in Figure 2. Figure 4 shown.

[0095] from Figure 4 It can be seen that although the prediction trends of the CPO-LSSVM model and the CPO-KELM model can well characterize the fluctuation of monthly runoff, the deviations in details are large, while the deviations between the prediction curves of the CPO-LSTM model, the CPO-BiLSTM model and the CPO-CNN-BiLSTM model and the measured value curve are very small. The error indicators RMSE, MRE and MAPE of the five prediction models are shown in Table 1.

[0096] Table 1 Point prediction error table

[0097]

[0098] A comprehensive comparison of the three error indicators of the five models shows that the three error indicators of the CPO-CNN-BiLSTM model are the best, with the highest prediction accuracy, followed by the CNN-BiLSTM model, the CPO-LSTM model, the CPO-KELM model and the CPO-LSSVM model. The prediction accuracy of the CPO-LSSVM model is the lowest. The RMSE, MRE and MAPE of the CPO-CNN-BiLSTM model are reduced by 22.49%, 26.13% and 25.45% respectively compared with the CPO-LSSVM model. This shows that the prediction results of the CPO-CNN-BiLSTM model have the highest fit with the measured values ​​and can more accurately follow the actual monthly runoff fluctuations.

[0099] S4: Establish a CPO-CNN-BiLSTM-NKDE interval prediction model. The point prediction results are sorted and divided into low flow segment, medium flow segment and high flow segment by using the extreme difference segmentation method. Then the non-parametric kernel density estimation method (NKDE) with CPO optimized window width is used to estimate the probability distribution of the prediction value error in each segment, and the cubic spline interpolation method is used for curve fitting to obtain the quantile of each segment. Finally, the point prediction results and the quantile of the segment to which the point prediction results belong are superimposed to obtain the interval prediction results of monthly runoff.

[0100] Specifically, according to the results of the CPO-CNN-BiLSTM point prediction model, the point prediction results are sorted and divided into low-flow segment, medium-flow segment and high-flow segment by using the extreme difference segmentation method, and the intervals are [5075, 6655], [6675, 8253] and [8256, 9836] respectively.

[0101] The comprehensive evaluation function (PICPAW) of prediction interval coverage (PICP) and prediction interval average bandwidth (PIAW) is selected as the fitness value calculation function of CPO optimization window width h. The mathematical expression of PICPAW is:

[0102] PICPAW=PIAW×[1+γ(PICP)×e -50(PICP-ξ) ] (4)

[0103] In formula (4), ξ represents the preset confidence level (0.9 in this method), γ takes 0 when PICP < ξ, otherwise it takes 1. CPO is used to optimize the window widths corresponding to the three segments, and it is found that the corresponding fitness values ​​are optimal when they are 0.0171403428, 0.0319370153 and 0.0170503989 respectively.

[0104] The CPO-CNN-LSTM-NKDE interval model proposed in this method is compared with the CPO-CNN-BiLSTM-NKDE model, CPO-LSTM-NKDE model, CPO-KELM-NKDE model and CPO-LSSVM-NKDE model. The structures of the four comparison models are the same as the CPO-CNN-BiLSTM-NKDE model, that is, based on the point prediction results, the monthly runoff prediction interval with a given confidence level is constructed in combination with NKDE. In the five models, CPO is used to optimize the window width of NKDE. The PICP and PICPPICPAW of the five prediction models are shown in Tables 2 and 3, respectively.

[0105] Table 2 Interval coverage of the prediction model

[0106]

[0107] Table 3 Interval bandwidth of the prediction model

[0108]

[0109] As can be seen from Table 2, at the 95% confidence level, the PICP of the CPO-CNN-BiLSTM-NKDE model is the largest; at the 90% confidence level, the interval PICPs of the five models are the same; at the 85% confidence level, the PICPs of the CPO-LSTM-NKDE model and the CPO-CNN-BiLSTM-NKDE model are the largest. That is, at these three confidence levels, the PICP of the CPO-CNN-BiLSTM-NKDE model can reach the maximum, and the PICPs at the three confidence levels are all higher than the corresponding confidence levels, indicating that the prediction interval of the model is the most reliable. As can be seen from Table 3, at the three confidence levels of 95%, 90% and 85%, the PIAW of the CPO-CNN-BiLSTM-NKDE model is the lowest, indicating that the prediction interval of the model is the most reliable. At the three confidence levels of 95%, 90% and 85%, the interval prediction results of the CPO-CNN-BiLSTM-NKDE model are as follows Figure 5 shown.

[0110] from Figure 5 It can be seen that the prediction interval of the CPO-CNN-BiLSTM-NKDE model can not only effectively cover the measured values, but also quantify the uncertainty of monthly runoff relatively accurately; the prediction interval can better represent the fluctuation of monthly runoff, especially when the monthly runoff fluctuates greatly, the model can also track the peak of monthly runoff relatively accurately, which proves the feasibility of the method.

[0111] The above embodiments are only preferred technical solutions of the present invention and should not be regarded as limiting the present invention. The protection scope of the present invention shall be the technical solutions recorded in the claims, including equivalent replacement solutions of the technical features in the technical solutions recorded in the claims. That is, equivalent replacement improvements within this scope are also within the protection scope of the present invention.

Claims

1. A monthly runoff series interval prediction method, characterized by: The steps include: S1: Monthly runoff data preparation: Collect monthly runoff data, pre-process the data, and divide the data into training data and test data; S2: Establish a CNN-BiLSTM point prediction model: including: input layer, CNN layer, BiLSTM layer and output layer, input the training data into the CNN-BiLSTM point prediction model to obtain the monthly runoff point prediction results; S3: Establish CPO-CNN-BiLSTM point prediction model: Use the crown porcupine optimization algorithm CPO to optimize the parameters of the CNN-BiLSTM point prediction model, bring the optimal parameters into the CNN-BiLSTM point prediction model for training, and obtain the trained monthly runoff prediction point model, namely the CPO-CNN-BiLSTM point prediction model; use the CPO-CNN-BiLSTM point prediction model to obtain the monthly runoff point prediction results; S4: Establish a CPO-CNN-BiLSTM-NKDE interval prediction model: Use the extreme difference segmentation method to sort the point prediction results and divide them into low flow segment, medium flow segment and high flow segment, and then use the CPO optimized window width non-parametric kernel density estimation method NKDE to estimate the probability distribution of the prediction value error in each segment, and use the cubic spline interpolation method to perform curve fitting to obtain the quantile of each segment. Finally, superimpose the point prediction results and the quantile of the segment to which the point prediction results belong to obtain the interval prediction results of monthly runoff.

2. A monthly runoff series interval prediction method according to claim 1, characterized in that: The specific steps of S1 include: S1.1: Obtain historical monthly runoff data from hydrological stations; S1.2: Normalize the monthly runoff data. The normalization process is expressed as: In formula (1), x represents the measured value of monthly runoff, x * represents the normalized monthly runoff, x min ,x max They represent the minimum and maximum values ​​of the measured monthly runoff respectively; S1.3: Determine the input dimension of the prediction model according to the delay time m of the partial autocorrelation coefficient of the monthly runoff; that is, the input corresponding to the monthly runoff output of the i-th month is the monthly runoff of the i-1th, i-2th, ..., imth months; S1.4: Divide the normalized dataset into a training sample set and a test sample set.

3. A monthly runoff series interval prediction method according to claim 2, characterized in that: The specific steps of S2 include: S2.1: Input the training sample set obtained in S1 into the CNN-BiLSTM point prediction model as the input of the model; S2.2: Construct convolutional neural network layer CNN: The CNN first extracts features from the monthly runoff data input by S2.1 through a one-dimensional convolution layer, with the number of convolution kernels being 32 and the step length being 1; then, after passing through two pooling layers and a smoothing layer, the data is filtered and flattened respectively, and feature information is output; S2.3: Construct a bidirectional long short-term memory network layer BiLSTM: The BiLSTM layer is composed of forward and backward LSTMs. The output feature information in S2.2 is input into the BiLSTM. After the feature information is enhanced and integrated by the Dropout layer and the fully connected layer, the prediction result is output. S2.4: The predicted value u output by the CNN-BiLSTM point prediction model is denormalized to obtain the monthly runoff prediction value.

4. A monthly runoff series interval prediction method according to claim 1, characterized in that: The specific steps of S3 include: S3.1: Initialize the parameters of CPO and CNN-BILSTM point prediction models, including the number of crested porcupine populations, the maximum number of iterations, the upper and lower boundaries of the solution space, and the dimension of the solution; the parameters optimized by CPO include the number of hidden layer nodes, the initial learning rate, and the regularization coefficient. The number of hidden layer nodes, the initial learning rate, and the regularization coefficient of the CNN-BILSTM point prediction model are used as the three-dimensional coordinates of the individual positions of the crested porcupine group, and the individual positions of the crested porcupines are randomly initialized; S3.2: Calculate the fitness value of each crested porcupine individual; S3.3: Update the fitness value of the crowned porcupine in the crowned porcupine group; calculate the information of the crowned porcupine with the best fitness from the first iteration to the current position, and save the information; S3.4: Update the position vector of the crested porcupine with the best fitness, and calculate the new position vector of the crested porcupine group after the next iteration; S3.5: Determine whether the maximum number of iterations has been reached. If so, end the iteration and return the optimal parameters; otherwise, return to S3.2; S3.6: Bring the optimal parameters into the CNN-BiLSTM point prediction model for training, so as to obtain the trained monthly runoff prediction model CPO-CNN-BiLSTM point prediction model; S3.7: Use the CPO-CNN-BiLSTM point prediction model to predict and obtain the monthly runoff point prediction result P pred .

5. A monthly runoff series interval prediction method according to claim 1, characterized in that: The specific steps of S4 include: S4.1: Error distribution estimation of runoff sequence: After arranging the predicted values ​​of the monthly runoff training data from small to large, the extreme difference segmentation method is used to divide them into low flow segment, medium flow segment and high flow segment, and the relative error of each predicted value in the three segments is calculated; the window width h of NKDE is initialized, and the error probability density function and probability distribution function F(ε) of each segment are calculated by NKDE, where ε is the random variable of the prediction error; the probability distribution function of the three segments is curve fitted by cubic spline interpolation method, and the quantile corresponding to each segment is calculated; determine which segment the monthly runoff point prediction result in the training data belongs to, and according to the quantile corresponding to the segment, use formula (2) to calculate: When the significance level is α, the confidence interval with a confidence level of 1-α is: In formula (2): is the inverse function of F(ε), α1=α / 2, α2=1-α / 2; S4.2: Use CPO to optimize the window width h and establish the CPO-CNN-BiLSTM-NKDE interval prediction model; recalculate the quantile corresponding to each segment according to the density distribution function after optimizing the window width h; S4.3: Monthly runoff sequence interval prediction: Determine which segment each monthly runoff point prediction result obtained in S3 belongs to, and superimpose the point prediction result P according to the quantile corresponding to the segment obtained in S4.

2. pred The confidence interval of monthly runoff is obtained by using the quantiles.

Citation Information

Cited By

  • Runoff interval prediction method and system based on multi-task learning framework

    CN121211985A

  • A runoff interval prediction method and system based on a multi-task learning framework

    CN121211985B