Method and system for forecasting medium-and-long-term runoff in watershed
By combining factor interaction detection technology and Markov chain correction method with the CNN-BiLSTM model, the problems of difficulty and low accuracy in medium- and long-term runoff forecasting for hydropower stations were solved, and high-precision medium- and long-term runoff forecasting for the basin was achieved, supporting water resources scheduling decisions.
Patent Information
- Application Number
- CN202510673635.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-05
AI Technical Summary
In the existing technology, the medium- and long-term runoff forecast of hydropower stations is difficult and has low accuracy. The traditional physical mechanism model has many parameters, which makes calibration difficult, and the data-driven model weakens the description of the hydrological cycle process.
The factor interaction detection technology is used to determine the optimal combination of forecast factors. The monthly runoff forecast model with the CNN-BiLSTM combined architecture is combined with the Markov chain correction method for rolling correction to obtain the target forecast factors that affect the monthly runoff changes in the basin.
It has greatly improved the accuracy of medium- and long-term runoff forecasts for hydropower stations, provided decision-making support for water resource scheduling, and improved the accuracy and reliability of forecasts by integrating remote correlation climate factor optimization and error correction strategies.
Smart Images

Figure CN120598262A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of water resource scheduling, and in particular to a method and system for forecasting medium- and long-term runoff in a river basin. Background Art
[0002] Medium- and long-term runoff forecasting is a crucial prerequisite for hydropower station operation management and the formulation of scheduling and operation plans. It is also crucial for optimizing water resource allocation, protecting the water environment, and combating floods and droughts. Climate change and the increasing inconsistency in runoff caused by human activities have significantly increased the difficulty and complexity of basin-wide runoff forecasting. Currently, basin-wide medium- and long-term runoff forecasting is primarily categorized into two main types: physical-based and data-driven models, depending on the modeling approach.
[0003] However, while traditional physical mechanism models can effectively depict the water cycle in a watershed, they suffer from numerous parameters that make calibration difficult and require a large number of data types. Compared to physical-driven models, data-driven models downplay the hydrological cycle and instead directly explore the relationship between predicted values and predictive factors from a data perspective. They are also more adaptable to watersheds lacking spatial information about the underlying surface. Summary of the Invention
[0004] To this end, the present invention provides a method and system for forecasting medium- and long-term runoff in a watershed, aiming to solve the technical problems of high difficulty and low accuracy in the existing medium- and long-term runoff forecasting for hydropower stations.
[0005] To achieve the above objectives, the present invention adopts the following technical solutions:
[0006] According to a first aspect of the present invention, the present invention provides a method for forecasting medium- and long-term runoff in a watershed, the method comprising:
[0007] Obtain target prediction factors that affect monthly runoff changes in the basin;
[0008] Determining the optimal combination of predictive factors among the target predictive factors using factor interaction detection technology;
[0009] The optimal prediction factor combination is divided into a model training set and a model test set, and a pre-built monthly runoff forecast model based on a CNN-BiLSTM combined architecture is trained and tested respectively to obtain a test prediction value;
[0010] The test prediction value is subjected to rolling correction using the Markov chain correction method to obtain a final prediction value.
[0011] Furthermore, the Markov chain correction method is used to correct the test prediction value to obtain the final prediction value, including:
[0012] Determine the relative error distribution of the model training set, and divide the relative error distribution of the model training set into a number of state intervals using a K-means clustering algorithm;
[0013] Calculating a state transition probability matrix corresponding to each state interval;
[0014] The target state interval of the test prediction value corresponding to the model test set is determined based on the state transition probability matrix, and the test prediction value is corrected using the Markov chain correction formula to obtain the final prediction value.
[0015] Furthermore, the acquisition of target prediction factors affecting the monthly runoff change in the watershed includes:
[0016] Collect multiple teleconnection climate factors that affect the monthly runoff changes in the basin as alternative forecast factors;
[0017] The candidate predictors were screened using the Spearman rank correlation analysis method to obtain target predictors, including:
[0018] The Spearman correlation coefficient of each candidate predictor was calculated using the following mathematical expression:
[0019]
[0020] Where ρ represents the Spearman correlation coefficient; represents the difference in the rank value of the i-th data pair; n represents the total number of observation samples;
[0021] The candidate predictor whose absolute value of the Spearman correlation coefficient is higher than a preset correlation threshold is screened as the target predictor.
[0022] Furthermore, the determining of the optimal prediction factor combination among the target prediction factors by using the factor interaction detection technology includes:
[0023] The factor detector and interaction detector in the geographic detector are used to screen the optimal predictor combination from the target predictor factors, including:
[0024] The factor detector is used to calculate the impact of the target prediction factor on the spatial heterogeneity of monthly runoff. The mathematical expression is as follows:
[0025]
[0026] S ST =Nσ 2
[0027] Among them, q represents the value of the influence degree of the target prediction factor on the spatial heterogeneity of monthly runoff; N represents the number of layers h in the region; N h represents the number of units in layer h; σ 2 represents the variance of the region; represents the variance of layer h; S ST represents the total variance of the entire region; S SW represents the sum of the variances within the layer;
[0028] Use the interactive detector to judge the interaction type of any two target prediction factors, and determine the optimal prediction factor combination in the target prediction factors according to the interaction type.
[0029] Further, the using the interactive detector to judge the interaction type of any two target prediction factors includes:
[0030] Compare the value of the influence degree of any two target prediction factors on the spatial heterogeneity of monthly runoff with the value of the influence degree of the two target prediction factors on the spatial heterogeneity of monthly runoff during interaction, and obtain a comparison result;
[0031] Judge the interaction type of the two target prediction factors according to the comparison result, including:
[0032] When the comparison result is q(X1∩X2) < Min(q(X1),(X2)), the corresponding interaction type is non-linear weakening; and / or,
[0033] When the comparison result is Min(q(X1),(X2)) < q(X1∩X2) < Max(q(X1),(X2)), the corresponding interaction type is single-factor non-linear weakening; and / or,
[0034] When the comparison result is q(X1∩X2) > Max(q(X1),(X2)), the corresponding interaction type is two-factor enhancement; and / or,
[0035] When the comparison result is q(X1∩X2) = q(X1) + q(X2), the corresponding interaction type is independent; and / or,
[0036] When the comparison result is q(X1∩X2) > q(X1) + q(X2), the corresponding interaction type is non-linear weakening.
[0037] Further, before dividing the optimal prediction factor combination into a model training set and a model test set, the method further includes:
[0038] Performing basin rainfall time lag analysis and previous runoff time lag analysis on the optimal forecast factor combination to construct a model data set for the monthly runoff forecast model, including:
[0039] The Pearson correlation coefficient is used to determine the lag time of the impact of basin rainfall on monthly runoff, including:
[0040]
[0041] Among them, x i and y i represents the i-th pair of observations in the optimal predictor combination; and represent the mean values of variables x and y respectively;
[0042] and / or,
[0043] The autocorrelation function and partial autocorrelation function are used to determine the lag time of the impact of previous runoff on current runoff, including:
[0044]
[0045] Among them, ρ k represents the value of the autocorrelation function at lag k; N represents the length of the time series; y t represents the value at time t; is the mean of the series; k represents the lag order; represents the value of the partial autocorrelation function at lag k; Represents the time series y t Linear estimation of .
[0046] Furthermore, the model training of the pre-built monthly runoff forecast model based on the CNN-BiLSTM combined architecture includes:
[0047] Based on the model training set, the model-related hyperparameters in the monthly runoff forecast model are optimized using the whale optimization algorithm; the mathematical expression of the whale optimization algorithm is as follows:
[0048]
[0049] Among them, X * (t) represents the position of the optimal solution, A and D represent the parameters that control the direction and distance of the whale's movement, respectively; D * represents the distance between the current individual and the optimal individual; b is a constant; l is a random number between [-1, 1]; the model-related hyperparameters include at least one of the number of convolutional layers, the number of output channels, the number of hidden layers, and the number of neurons.
[0050] Furthermore, determining the target state interval in which the test prediction value corresponding to the model test set is located based on the state transition probability matrix includes:
[0051] Calculate the sum of probabilities of the sample data in the model test set being in each state using the state transition probability matrix;
[0052] The state interval with the largest comprehensive probability is selected as the target state interval where the test prediction value is located.
[0053] Furthermore, the mathematical expression of the K-means clustering algorithm is as follows:
[0054]
[0055] Where J represents the sum of squared errors within the cluster; K represents the number of clusters; C i represents the point set of cluster i; x represents the point set belonging to C i Data points of u i represents the centroid of the i-th cluster; ‖x i -u i ‖ 2 Represents the data point x and the cluster centroid u i The square of the Euclidean distance between them;
[0056] and / or,
[0057] The mathematical expression of the Markov chain correction formula is:
[0058]
[0059] Wherein, F(x) represents the predicted menstrual flow value after correction, f(x) represents the predicted menstrual flow value before correction; ΔU represents the maximum value of the target state interval; ΔD represents the minimum value of the target state interval.
[0060] According to a second aspect of the present invention, the present invention provides a medium- to long-term runoff forecasting system for a watershed, characterized in that the system comprises:
[0061] Data acquisition module, used to obtain target prediction factors affecting the monthly runoff changes in the basin;
[0062] A data selection module is used to determine the optimal prediction factor combination among the target prediction factors by using factor interaction detection technology;
[0063] A model prediction module is used to divide the optimal prediction factor combination into a model training set and a model test set, and perform model training and model testing on a pre-built monthly runoff forecast model based on a CNN-BiLSTM combined architecture to obtain a test prediction value;
[0064] The error correction module is used to perform rolling correction on the test prediction value using the Markov chain correction method to obtain a final forecast.
[0065] The present invention adopts the above technical solution and has at least the following beneficial effects:
[0066] Through the scheme of the present invention, target forecast factors that affect the monthly runoff changes in the basin are obtained; the optimal forecast factor combination among the target forecast factors is determined using factor interaction detection technology; the optimal forecast factor combination is divided into a model training set and a model test set, and a pre-built monthly runoff forecast model based on the CNN-BiLSTM combined architecture is trained and tested respectively to obtain a test forecast value; the test forecast value is rolled corrected using the Markov chain correction method to obtain the final forecast value. Through the present invention, a medium- and long-term runoff forecast scheme for the basin is provided that integrates the optimization of teleconnection climate factors and error correction strategies, greatly improving the accuracy of medium- and long-term runoff forecasts for hydropower stations and providing decision support for water resource scheduling.
[0067] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0069] Figure 1 A schematic diagram showing a flow chart of a method for forecasting medium- and long-term runoff in a watershed provided by an embodiment of the present invention is shown;
[0070] Figure 2 A ranking diagram showing the degree of influence of target prediction factors on the spatial heterogeneity of monthly runoff provided by an embodiment of the present invention is shown;
[0071] Figure 3 shows a heat map of interactions between target prediction factors provided by an embodiment of the present invention;
[0072] Figure 4 A scatter plot showing relative error clustering of a model training set using a K-means clustering algorithm provided by an embodiment of the present invention is shown;
[0073] Figure 5 A comparison chart showing the effects of the monthly runoff forecast model after correction provided by an embodiment of the present invention and other deep learning forecast models in a test set is shown;
[0074] Figure 6 A scatter plot of the prediction results of the monthly runoff forecast model after correction provided by an embodiment of the present invention and other deep learning forecast models in the test set is shown;
[0075] Figure 7 A Taylor diagram showing the prediction results of the monthly runoff forecast model after correction provided by an embodiment of the present invention and other deep learning forecast models in the test set;
[0076] Figure 8 A schematic structural diagram of a basin medium- and long-term runoff forecasting system provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0077] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0078] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.
[0079] The embodiment of the present invention provides a method for long-term runoff forecasting in a watershed, such as Figure 1 As shown, it may at least include the following steps S101 to S104:
[0080] Step S101: Obtain target prediction factors that affect the monthly runoff change in the basin.
[0081] First, multiple teleconnection climate factors that affect the monthly runoff variation in the basin are collected as candidate predictors. In practical applications of the present invention, 130 large-scale teleconnection climate factors can be obtained from the National Climate Center of the China Meteorological Administration, including 88 atmospheric circulation indices, 26 sea surface temperature indices, and 16 other indices.
[0082] To ensure data integrity, embodiments of the present invention can also process the collected raw data set using a "elimination-interpolation" approach. For example, predictors with a missing rate greater than 10% in the data set are eliminated, and predictors with a missing rate less than 10% are interpolated by taking the mean value of that month over the past years.
[0083] Then, the Spearman rank correlation analysis method is used to screen the alternative prediction factors to obtain the target prediction factors. In other words, the 130 large-scale teleconnection climate factors collected are preliminarily screened based on the Spearman rank correlation analysis method, and the monthly runoff prediction factors with no significant correlation and redundancy are eliminated to obtain the target prediction factors. Spearman rank correlation analysis is a non-parametric statistical method, which is mainly used to determine the degree of correlation between two variables. It first measures the relationship between the two variables by calculating the ranks and then calculating their Spearman rank correlation coefficient. The value range of the Spearman correlation coefficient is between -1 and 1, where -1 indicates a complete negative correlation, 1 indicates a complete positive correlation, and 0 indicates no correlation.
[0084] Specifically, the Spearman correlation coefficient of each candidate predictor can be calculated using the following mathematical expression:
[0085]
[0086] Where ρ represents the Spearman correlation coefficient; represents the difference in rank value of the i-th data pair; n represents the total number of observed samples.
[0087] Alternative predictors whose absolute values of the Spearman correlation coefficients are higher than a preset correlation threshold are selected as target predictors. For example, the preset correlation threshold can be set to 0.6, that is, alternative climate factors whose absolute values of the Spearman correlation coefficients are greater than 0.6 are selected as target predictors.
[0088] The following uses the Pubugou Hydropower Station in the Dadu River Basin as an example to illustrate an embodiment of the present invention. First, the original data set consisting of 130 large-scale teleconnection climate factors is cleaned, and the Spearman correlation coefficient is calculated. The candidate prediction factors with an absolute value greater than 0.6 are selected as target prediction factors. To facilitate the display of the calculation results, 88 atmospheric circulation indices are numbered A1-A88, 26 sea temperature indices are numbered B1-B26, and 16 other indices are numbered C1-C16. The Spearman correlation coefficients corresponding to each target prediction factor are shown in Table 1 below:
[0089] Table 1 (Spearman correlation coefficient calculation results)
[0090]
[0091]
[0092]
[0093] It should be noted that, in practical applications, the preset relevant threshold value can be specifically set according to actual needs, and the present invention does not limit this.
[0094] Step S102: Determine the optimal prediction factor combination among the target prediction factors using factor interaction detection technology.
[0095] This embodiment of the present invention uses the factor detector and interaction detector within the geographic detector to determine the optimal combination of predictors from the screened target predictors. Geographic detectors are a set of statistical methods for detecting spatial heterogeneity and revealing the driving forces behind it. Their core concept is based on the assumption that if an independent variable has a significant impact on a dependent variable, then the spatial distributions of the independent and dependent variables should be similar.
[0096] Specifically, the factor detector is used to calculate the impact of the target prediction factor on the spatial heterogeneity of monthly runoff. The mathematical expression is as follows:
[0097]
[0098] S ST =Nσ 2
[0099] Where q represents the degree of influence of the target prediction factor on the spatial heterogeneity of monthly runoff, and its value range is between [0, 1]. The larger the q value, the greater the contribution of the target prediction factor to the spatial heterogeneity of monthly runoff and the stronger its explanatory power for monthly runoff. N represents the number of layers h in the region, h = 1, 2, 3...; N h represents the number of units in layer h; σ 2 represents the variance of the region; represents the variance of layer h; S ST represents the total variance of the entire region; S SW represents the sum of the within-stratum variances.
[0100] The interaction detector in the geographic detector primarily calculates the explanatory power of any two target predictors on monthly runoff when acting together. This evaluates the interaction between different target predictors. The spatial heterogeneity of monthly runoff, q, is compared between any two target predictors, X1 and X2, and the spatial heterogeneity of monthly runoff when the two predictors interact (X1∩X2). The comparison results are then used to determine the type of interaction between the two target predictors. The types of interactions between the two target predictors are shown in Table 2 below:
[0101] Table 2 (Types of Interactions between Target Predictors)
[0102]
[0103] Take the Pubugou Hydropower Station in the Dadu River Basin as an example. Figure 2 As shown in the figure, it is the q-value ranking diagram of the influence degree of each target prediction factor on the spatial heterogeneity of monthly runoff obtained by using the factor detector; Figure 3 As shown in Figure 2, it is a heat map of the interaction between any two target predictors obtained using the interaction detector. Figure 2 and Figure 3 Based on the calculation results, the East Asian trough intensity index, the North African subtropical high ridge position index, the Indian subtropical high ridge position index, the North Pacific subtropical high ridge position index, the Northern Hemisphere subtropical high ridge position index, the Tibet Plateau-2 index, the Western Pacific subtropical high ridge position index and the North African-North Atlantic-North American subtropical high ridge position index can be selected as the optimal forecast factor combination for the monthly runoff forecast model of the Pubugou section, with a time lag of 12 months.
[0104] In step S103, the optimal prediction factor combination is divided into a model training set and a model test set, and the pre-built monthly runoff forecast model based on the CNN-BiLSTM combined architecture is trained and tested respectively to obtain a test prediction value.
[0105] The embodiment of the present invention combines the selected optimal forecast factor combination with the rainfall and previous runoff data that have undergone time lag analysis to construct a data set for model training and testing.
[0106] Preferably, the Pearson correlation coefficient can be used to determine the lag time of the impact of basin rainfall on monthly runoff. The mathematical expression is as follows:
[0107]
[0108] Among them, x i and y i represents the i-th pair of observations in the optimal predictor combination; and represent the mean values of variables x and y respectively.
[0109] Preferably, the autocorrelation function (ACF) and partial autocorrelation function (PACF) can be used to determine the lag of the impact of previous runoff on current runoff. The mathematical expressions are as follows:
[0110]
[0111] Among them, ρ k represents the value of the autocorrelation function at lag k; N represents the length of the time series; y t express
[0112] The value at time t; is the mean of the series; k represents the lag order; represents the value of the partial autocorrelation function at lag k; Represents the time series y t Linear estimation of .
[0113] Therefore, basin rainfall time lag analysis and previous runoff time lag analysis are performed on the optimal forecast factor combination, a model dataset of the monthly runoff forecast model is constructed, and the model training set and model test set are divided. The pre-built monthly runoff forecast model based on the CNN-BiLSTM combined architecture is trained and tested.
[0114] The monthly runoff forecast model used in the embodiment of the present invention is based on a combined architecture of CNN-BiLSTM deep learning. The input step size of the monthly runoff forecast model is 12, and the output step size is 1, that is, the monthly runoff for the next month is predicted based on the historical data of the past year. It can be understood that after the CNN feature extraction and pooling operations, the model test set converts the extracted features into a format suitable for time series processing as the input of BiLSTM; then, through forward and backward propagation, the long-term and short-term dependencies in the time series are captured to form a CNN-BiLSTM combined architecture. At the same time, during the model training process, the whale optimization (WOA) algorithm can be used to optimize the four related hyperparameters of the number of convolutional layers, the number of output channels, the number of hidden layers, and the number of hidden units in the CNN-BiLSTM combined architecture to obtain the optimal hyperparameter combination. The whale optimization (WOA) algorithm includes three steps: surrounding prey, bubble net predation, and random prey search. The equal probability distribution strategy is used to dynamically adjust the whale's predation behavior to avoid falling into the local optimum. When P is less than 0.5, the shrinking and encircling mechanism is adopted. When P is greater than 0.5, the spiral updating mechanism is adopted. The mathematical expression is as follows:
[0115]
[0116] Among them, X *(t) represents the location of the optimal solution, A and D represent the parameters that control the direction and distance of the whale's movement, respectively. This mechanism ensures that the algorithm can focus on searching near the optimal solution; D * Represents the distance between the current individual and the optimal individual; b is a constant; l is a random number between [-1,1]. By moving in a spiral, the whale can avoid being limited to the current search path and expand the search space.
[0117] Taking the Pubugou Hydropower Station in the Dadu River Basin as an example, the first 80% of the model data set can be used as the model training set, and the last 20% as the model test set. In actual testing, the embodiment of the present invention uses root mean square error (RMSE), mean absolute error (MAE), mean percentage absolute error (MPAE), Kling-Gupta efficiency coefficient, and Nash efficiency coefficient (NSE) as evaluation indicators. At the same time, monthly runoff forecast models with CNN, LSTM, Transformer and other architectures are constructed for comparison. The indicator evaluation comparison results are shown in Table 3 below:
[0118] Table 3 (Comparison of evaluation indicators of different models)
[0119]
[0120] Based on the comparison of the above evaluation indicators, the monthly runoff forecast model based on the CNN-BiLSTM combined framework provided by the embodiment of the present invention has the highest prediction accuracy and good performance.
[0121] Step S104: Perform rolling correction on the test prediction value using the Markov chain correction method to obtain the final prediction value.
[0122] Specifically, the relative error distribution of the model training set can be determined, and the relative error distribution of the model training set can be divided into several state intervals using the K-means clustering algorithm; the state transition probability matrix corresponding to each state interval can be calculated; based on the state transition probability matrix, the target state interval where the test prediction value corresponding to the model test set is located can be determined, and the test prediction value can be corrected using the Markov chain correction formula to obtain the final prediction value.
[0123] Among them, the mathematical expression of the K-means clustering algorithm is as follows:
[0124]
[0125] Where J represents the sum of squared errors within the cluster; K represents the number of clusters; C i represents the point set of cluster i; x represents the point set belonging to C i Data points of u i represents the centroid of the i-th cluster; ‖x i -u i ‖ 2Represents the data point x and the cluster centroid u i The square of the Euclidean distance between them.
[0126] Taking the Pubugou Hydropower Station in the Dadu River Basin as an example, the K-means clustering algorithm is used to divide the relative error of the model training set into four error intervals. Figure 4 The following is a scatter plot of the classification results for the four error intervals. Based on the clustering results, the relative error of the training set can be divided into four Markov chain correction state intervals: [-0.3846, -0.0175), [-0.0175, 0.3137), [0.3137, 0.6449), and [0.6449, 1). The state transition probability matrix corresponding to each state interval is calculated as follows:
[0127]
[0128] Furthermore, the state transition probability matrix can be used to calculate the sum of the probabilities of the sample data in the model test set being in each state; the state interval with the largest probability integration is selected as the target state interval where the test prediction value is located; and the Markov chain correction formula is used to roll-correct the test set prediction value to obtain the final prediction value. The mathematical expression of the Markov chain correction formula is:
[0129]
[0130] Where F(x) represents the corrected menstrual flow prediction value, f(x) represents the monthly runoff prediction value before correction; ΔU represents the maximum value of the target state interval obtained by the state transition probability matrix for the model test set; ΔD represents the minimum value of the target state interval obtained by the state transition probability matrix for the model test set.
[0131] Taking the Pubugou Hydropower Station in the Dadu River Basin as an example, the state probability of each sample data in the model test set is calculated through the state transition probability matrix, and the probabilities of all samples in the same state are added together to obtain the sum of each state probability. The state with the largest sum of state probabilities is the relative error interval where the test prediction value error lies. Finally, the corrected monthly runoff forecast value, i.e. the final forecast value, is obtained through the Markov chain correction formula. Figure 5 、 Figure 6 As shown in the figure, there are the fitting comparison diagram and scatter plot of the prediction results of the CNN-BiLSTM monthly runoff forecast model after Markov chain error correction and other models on the test set. Figure 7 As shown, it is the Taylor diagram of the prediction results of each model in the test set.
[0132] comprehensive Figures 5-7Compared with the prediction results of the previous study, the monthly runoff forecast model based on the CNN-BiLSTM combined framework proposed in the embodiment of the present invention, after Markov chain correction (MC-CNN-BiLSTM), has better accuracy and reliability than other models, further improving the accuracy of monthly runoff forecast.
[0133] The embodiment of the present invention provides a method for forecasting medium- and long-term runoff in a watershed. Compared with the existing technology, it has at least the following advantages:
[0134] 1) The present invention uses the Spearman correlation analysis method for preliminary screening and the geographic detector selection method to determine the optimal combination of predictive factors affecting monthly runoff changes, and eliminates monthly runoff predictive factors with no significant correlation and redundancy;
[0135] 2) The present invention introduces a forecast model with a CNN-BiLSTM combined architecture, which overcomes the limitations of traditional data-driven models, such as their inability to fully capture the local spatial characteristics of runoff time series, and the low simulation accuracy of physical models in watersheds with complex underlying surfaces.
[0136] 3) The present invention introduces an error correction strategy and utilizes the “memoryless” feature of the Markov chain, that is, the state of the next stage depends only on the current state, which can further improve the accuracy of monthly runoff forecast.
[0137] Further, as Figure 1 The embodiment of the present invention provides a long-term runoff forecasting system for a watershed, such as Figure 8 As shown, the system may include: a data acquisition module 810, a data selection module 820, a model prediction module 830 and an error correction module 840.
[0138] The data acquisition module 810 can be used to obtain target prediction factors that affect the monthly runoff changes in the watershed;
[0139] The data selection module 820 can be used to determine the optimal prediction factor combination among the target prediction factors using factor interaction detection technology;
[0140] The model prediction module 830 can be used to divide the optimal prediction factor combination into a model training set and a model test set, and perform model training and model testing on the pre-built monthly runoff forecast model based on the CNN-BiLSTM combined architecture to obtain a test prediction value;
[0141] The error correction module 840 can be used to perform rolling correction on the test prediction value using the Markov chain correction method to obtain the final prediction value.
[0142] It should be noted that for other corresponding descriptions of the functional modules involved in the medium- and long-term runoff forecast system for a watershed provided by the embodiment of the present invention, please refer to Figure 1 The corresponding description of the method shown will not be repeated here.
[0143] Those skilled in the art will clearly understand that the specific working processes of the systems, devices, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and for the sake of brevity, they will not be further described here.
[0144] In addition, the functional units in various embodiments of the present invention may be physically independent of each other, or two or more functional units may be integrated together, or all functional units may be integrated into a single processing unit. The above-mentioned integrated functional units may be implemented in the form of hardware, software, or firmware.
[0145] Those skilled in the art will understand that if the integrated functional unit is implemented in the form of software and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention can essentially or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, which includes a number of instructions for enabling a computing device (such as a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention when running the instructions. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0146] Alternatively, all or part of the steps of implementing the aforementioned method embodiments may be accomplished by hardware associated with program instructions (such as a computing device such as a personal computer, a server, or a network device), and the program instructions may be stored in a computer-readable storage medium. When the program instructions are executed by a processor of a computing device, the computing device executes all or part of the steps of the method described in each embodiment of the present invention.
[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that within the spirit and principles of the present invention, they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not deviate from the scope of protection of the present invention.
Claims
1. A method for forecasting medium- to long-term runoff in a watershed, characterized in that: The method includes: Obtaining target prediction factors that affect the monthly runoff variation in the basin; Using factor interaction detection technology to determine the optimal combination of prediction factors among the target prediction factors; Dividing the optimal combination of prediction factors into a model training set and a model testing set, respectively training and testing a pre-constructed monthly runoff prediction model based on a CNN-BiLSTM combined architecture to obtain test prediction values; Using the Markov chain correction method to perform rolling correction on the test prediction values to obtain the final prediction values.
2. The method according to claim 1, characterized in that The step of using the Markov chain correction method to correct the test prediction values to obtain the final prediction values includes: Determining the relative error distribution of the model training set, and using the K-means clustering algorithm to divide the relative error distribution of the model training set into several state intervals; Calculating the state transition probability matrix corresponding to each state interval; Based on the state transition probability matrix, determining the target state interval where the test prediction value corresponding to the model testing set is located, and using the Markov chain correction formula to correct the test prediction value to obtain the final prediction value.
3. The method according to claim 1, characterized in that The step of obtaining target prediction factors that affect the monthly runoff variation in the basin includes: Collecting multiple teleconnection climate factors that affect the monthly runoff variation in the basin as alternative prediction factors; Using the Spearman rank correlation analysis method to screen the alternative prediction factors to obtain target prediction factors, including: Calculating the Spearman correlation coefficient of each alternative prediction factor using the following mathematical expression: Where ρ represents the Spearman correlation coefficient; represents the difference in the rank value of the i-th data pair; n represents the total number of observation samples; Screening the alternative prediction factors with the absolute value of the Spearman correlation coefficient higher than a preset correlation threshold as target prediction factors.
4. The method according to claim 1, wherein The step of using factor interaction detection technology to determine the optimal combination of prediction factors among the target prediction factors includes: Using the factor detector and interaction detector in the geographical detector to screen the optimal combination of prediction factors from the target prediction factors, including: Using the factor detector to calculate the value of the impact degree of the spatial heterogeneity of the target prediction factors on the monthly runoff, and the mathematical expression is as follows: S ST =Nσ 2 Where q represents the impact of the target forecast factor on the spatial heterogeneity of monthly runoff; N represents the number of layers h in the region; N h represents the number of units in layer h; σ 2 represents the variance of the region; represents the variance of layer h; S ST represents the total variance of the entire region; S SW represents the sum of variance within the layer; Using the interaction detector to judge the interaction type of any two target prediction factors, and determining the optimal combination of prediction factors among the target prediction factors according to the interaction type.
5. The method according to claim 4, characterized in that The step of using the interaction detector to judge the interaction type of any two target prediction factors includes: Comparing the value of the impact degree of the spatial heterogeneity of any two target prediction factors on the monthly runoff with the value of the impact degree of the spatial heterogeneity of the two target prediction factors on the monthly runoff during interaction to obtain a comparison result; Judging the interaction type of the two target prediction factors according to the comparison result, including: When the comparison result is q(X1∩X2)<Min(q(X1),(X2)), the corresponding interaction type is non-linear weakening; and / or, When the comparison result is Min(q(X1),(X2))<q(X1∩X2)<Max(q(X1),(X2)), the corresponding interaction type is single-factor non-linear weakening; and / or, When the comparison result is q(X1∩X2)>Max(q(X1),(X2)), the corresponding interaction type is two-factor enhancement; and / or, When the comparison result is q(X1∩X2)=q(X1)+q(X2), the corresponding interaction type is independent; and / or, When the comparison result is q(X1∩X2)>q(X1)+q(X2), the corresponding interaction type is nonlinear attenuation.
6. The method according to claim 1, characterized in that Before dividing the optimal prediction factor combination into a model training set and a model test set, the method further includes: Performing basin rainfall time lag analysis and previous runoff time lag analysis on the optimal forecast factor combination to construct a model data set for the monthly runoff forecast model, including: The Pearson correlation coefficient is used to determine the lag time of the impact of basin rainfall on monthly runoff, including: Among them, x i and y i represents the i-th pair of observations in the optimal predictor combination; and represent the mean values of variables x and y respectively; and / or, The autocorrelation function and partial autocorrelation function are used to determine the lag time of the impact of previous runoff on current runoff, including: Among them, ρ k represents the value of the autocorrelation function at lag k; N represents the length of the time series; y t represents the value at time t; is the mean of the series; k represents the lag order; represents the value of the partial autocorrelation function at lag k; Represents the time series y t Linear estimation of .
7. The method according to claim 1, characterized in that The model training of the pre-built monthly runoff forecast model based on the CNN-BiLSTM combined architecture includes: Based on the model training set, the model-related hyperparameters in the monthly runoff forecast model are optimized using the whale optimization algorithm; the mathematical expression of the whale optimization algorithm is as follows: Among them, X * (t) represents the position of the optimal solution, A and D represent the parameters that control the direction and distance of the whale's movement, respectively; D * represents the distance between the current individual and the optimal individual; b is a constant; l is a random number between [-1, 1]; the model-related hyperparameters include at least one of the number of convolutional layers, the number of output channels, the number of hidden layers, and the number of neurons.
8. The method according to claim 2, characterized in that The determining, based on the state transition probability matrix, a target state interval in which a test prediction value corresponding to the model test set is located includes: Calculate the sum of probabilities of the sample data in the model test set being in each state using the state transition probability matrix; The state interval with the largest comprehensive probability is selected as the target state interval where the test prediction value is located.
9. The method according to claim 8, characterized in that The mathematical expression of the K-means clustering algorithm is as follows: Where J represents the sum of squared errors within the cluster; K represents the number of clusters; C i represents the point set of cluster i; x represents the point set belonging to C i Data points of u i represents the centroid of the i-th cluster; ‖x i -u i ‖ 2 Represents the data point x and the cluster centroid u i The square of the Euclidean distance between them; and / or, The mathematical expression of the Markov chain correction formula is: Wherein, F(x) represents the predicted menstrual flow value after correction, f(x) represents the predicted menstrual flow value before correction; ΔU represents the maximum value of the target state interval; ΔD represents the minimum value of the target state interval.
10. A medium- to long-term runoff forecast system for a river basin, characterized in that: The system comprises: Data acquisition module, used to obtain target prediction factors affecting the monthly runoff changes in the basin; A data selection module is used to determine the optimal prediction factor combination among the target prediction factors by using factor interaction detection technology; A model prediction module is used to divide the optimal prediction factor combination into a model training set and a model test set, and perform model training and model testing on a pre-built monthly runoff forecast model based on a CNN-BiLSTM combined architecture to obtain a test prediction value; The error correction module is used to perform rolling correction on the test prediction value using the Markov chain correction method to obtain a final prediction value.