Long-term Inflow Runoff Prediction Method for Large Lakes in River Network Areas
By dividing the tributaries into lakes into three types: low, medium and high impact branches, using multiple models for differentiated predictions, and constructing conditional probability distribution and error correction, the complexity and accuracy of medium- and long-term lake runoff prediction in large lakes in the river network area are solved, and more accurate prediction and risk quantification are achieved.
Patent Information
- Application Number
- CN202510579559.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The prediction of medium- and long-term lake runoff in large lakes in the river network area has problems such as high model complexity, insufficient data, uncontrolled interval impact and low prediction accuracy. It is especially difficult to achieve satisfactory prediction results in watersheds with significant human activity interference.
The tributaries entering the lake are divided into three types: low, medium and high influence branches, and the SWAT model, physically constrained LSTM model and DBN-BP model are used for prediction, and combined with the conditional normalized flow model and the non-parametric residual model, runoff probability distribution and error correction are constructed to quantify uncertainty.
It improves the accuracy and reliability of medium- and long-term lake runoff prediction, provides more reliable water resource management information, quantifies the predicted risks, and optimizes the flood control scheduling strategy.
Smart Images

Figure CN120106399B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medium- and long-term hydrological forecasting, and particularly to a method for predicting medium- and long-term runoff into large lakes in a river network area. Background Art
[0002] As important fresh water resources, large lakes in river network areas play important roles in regional ecological balance, water resources dispatching, flood control and drought relief, etc. Accurate medium- and long-term runoff prediction into lakes can provide important data support and decision-making basis for overall planning and joint dispatching of water resources in river network areas. However, the hydro-hydraulic conditions in river network areas are complex and human activities are intense, resulting in characteristics such as non-linearity, time-variability and randomness of the runoff into lakes, making medium- and long-term scale prediction more complex and difficult.
[0003] Medium- and long-term runoff prediction methods mainly include two categories: process-driven models and data-driven models. Among them, process-driven models achieve runoff prediction by simulating the hydrological cycle process, which has strong interpretability but complex structures and has good fitting effects in basins with less human activities; data-driven models infer the development trend of future runoff by learning the complex linear or non-linear relationships between data features, with high calculation efficiency and strong adaptability, but highly dependent on the integrity and representativeness of training data and poor interpretability. The river network water systems in the catchment areas of large lakes are intricate, and it is difficult to obtain satisfactory prediction results with high modeling difficulty and large workload using a single model for prediction. Especially in basins with significant human activity interference, data-driven models often face challenges such as insufficient data. There are uncertainties in the input factors, model structures and parameters of runoff prediction models. Deterministic runoff prediction can only give a specific expected value and cannot quantify the degree of prediction uncertainty. At the same time, most of the existing technologies do not consider the influence of runoff changes in the uncontrolled section between the main stream gauging stations and the lake inlet, which to a certain extent limits the prediction accuracy of the runoff into lakes. Summary of the Invention
[0004] Object of the Invention: To provide a method for predicting medium- and long-term runoff of large lakes in a river network area to solve the above problems existing in the prior art.
[0005] Technical Solution: Provide a method for predicting medium- and long-term runoff of large lakes in a river network area, including the following steps:
[0006] S1. Collect data of the study area, determine the degree of influence of human activities on the tributaries flowing into the lake in the river network, and divide the tributaries into three types of influence branches: low, medium and high.
[0007] S2. Respectively use the SWAT model and the physically constrained LSTM model to predict the low- and medium-influence branches to obtain the runoff prediction results of the low- and medium-influence branches.
[0008] S3. Screen the runoff prediction results of low- and medium-impact branches that are significantly correlated with the high-impact branches as prediction factors, construct a DBN-BP model to establish the correlation between the runoff of high-impact branches and the prediction factors, and achieve the iterative prediction of the runoff of high-impact branches;
[0009] S4. Use the conditional normalizing flow model to construct the conditional probability distribution between the sum of the runoff of the three types of branches and the total inflow into the lake, obtain the deterministic and probabilistic prediction results of the total inflow into the lake considering the uncontrolled area, and construct a non-parametric residual model to achieve error correction.
[0010] By introducing the runoff probability prediction to generate the probability distribution function of the prediction results or the confidence interval at different levels, the uncertainty can be intuitively reflected. In addition, combining the residual model to achieve systematic error correction can further improve the accuracy and reliability of the inflow prediction into the lake, help decision-makers better evaluate risks, and optimize water resource management and flood control scheduling strategies.
[0011] According to one aspect of the present application, the step S1 is further as follows:
[0012] S11. Collect the data of the study area;
[0013] S12. Divide the catchment area of each tributary flowing into the lake in the river network, establish the set of human activity index factors within the catchment area, and screen out the important human activity indicators;
[0014] S13. Based on the important human activity indicators, construct the human activity evaluation matrix of the tributaries flowing into the lake in the river network, and use the modified fuzzy RANCOM method and the CRITIC method to determine the subjective and objective weights of each indicator respectively, and combine them to obtain the comprehensive weight of each indicator;
[0015] S14. According to the comprehensive weights of each indicator, use the TOPSIS method to rank the degree of influence of each tributary by human activities, and accordingly divide the tributaries into three types of impact branches: low, medium, and high.
[0016] According to one aspect of the present application, in step S13, the specific process of using the modified fuzzy RANCOM method to determine the subjective weight of each indicator is further as follows:
[0017] S13a. Experts rank each indicator according to the importance of the indicator;
[0018] S13b. Compare the ranking values of the indicators pairwise to generate a ranking comparison matrix;
[0019] S13c. Introduce the triangular fuzzy theory to convert the exact values in the permutation comparison matrix into triangular fuzzy intervals, generate a fuzzy ranking comparison matrix considering the uncertainty of expert judgment, and use the dynamic consensus strategy to identify and correct the evaluation results of low-consensus experts;
[0020] S13d. Sum the elements of each row of the modified fuzzy ranking comparison matrix according to the triangular fuzzy number addition rule to obtain the fuzzy comprehensive weight of each index, and normalize it to obtain the final fuzzy weight. Use the core value of the fuzzy weight as the subjective weight result of the index.
[0021] According to one aspect of the present application, step S2 is further as follows:
[0022] S21. Call the pre-built SWAT model to predict the low-impact branch runoff and obtain the deterministic prediction result of the low-impact branch runoff.
[0023] S22. Use the physically constrained LSTM model to predict the medium-impact branch runoff and obtain the deterministic prediction result of the medium-impact branch runoff.
[0024] According to one aspect of the present application, step S22 is further as follows:
[0025] S22a. Input the hydrometeorological data and human activity parameters, use the WEAP model to simulate the medium-impact branch runoff, and generate the runoff sequence and sub-process physical variables considering the impact of human activities.
[0026] S22b. Use the meteorological data and the output of the WEAP model as inputs to construct an LSTM model, and add physical condition constraints to the loss function of the LSTM model based on the output of the WEAP model to obtain the deterministic prediction result of the medium-impact branch runoff.
[0027] According to one aspect of the present application, step S3 is further as follows:
[0028] S31. Use the Laplace stepwise regression method to screen out a group of low- and medium-impact branches that are significantly correlated with the high-impact branch runoff.
[0029] S32. Calculate the predicted values of the screened low- and medium-impact branch runoff respectively, and construct an alternative prediction factor set for the high-impact branch based on this. Then use the MIC method to screen and obtain the optimal prediction factor set for the high-impact branch.
[0030] S33. Use the myxomycete algorithm-improved sooty tern optimization algorithm to optimize the hyperparameters of the DBN-BP model, use the optimal factor set of the high-impact branch as the model input, and obtain the deterministic prediction result of the high-impact branch runoff.
[0031] According to one aspect of the present application, in step S31, the screening process of the Laplace stepwise regression method is further as follows:
[0032] S31a. Calculate the Laplace scores between the runoff of each low- and medium-impact branch and the high-impact branch runoff, and arrange them in ascending order.
[0033] S31b. Construct a stepwise regression model. In each iteration, select the factor with the smallest Laplace score from the remaining factor set to add to the model, and eliminate the least significant factor obtained by the Bootstrap hypothesis test.
[0034] S31c. When the model performance is no longer significantly improved, the screening is stopped to obtain a set of low- and medium-impact branches that are significantly correlated with the runoff of high-impact branches.
[0035] According to one aspect of the present application, in step S33, the optimization process of the black tern hyperparameter optimization algorithm improved by the slime mold algorithm is further as follows:
[0036] S33a, use Circle chaos map to initialize individuals in the population;
[0037] S33b, adjust the individual position by moving variables to avoid collision, introduce slime mold foraging behavior to make the individual gradually converge to the optimal position, and update the individual position;
[0038] S33c, further updating individual positions by simulating the predatory behavior of the spiral flight of the sooty tern;
[0039] S33d. When the maximum number of iterations is reached, the optimization algorithm is terminated and the current solution set is returned as the hyperparameter optimization result.
[0040] According to one aspect of the present application, step S4 is further:
[0041] S41. Based on the sum of the historical measured values of the three types of branch runoff and their corresponding historical measured values of the runoff into the lake, a conditional normalized flow model is used to construct a conditional probability distribution between the two, and the sum of the predicted results of the three types of branch runoff is used as a conditional input to output the total runoff into the lake probability prediction result considering the uncontrolled interval;
[0042] S42, constructing a nonparametric residual model based on the error between the predicted value and the measured value of the total runoff into the lake, and generating residual samples;
[0043] S43. Combine the residual samples with the deterministic prediction values to generate residual-corrected runoff prediction values. Generate the probability distribution of runoff prediction values through repeated sampling, calculate the confidence interval or quantile, and quantify the uncertainty of the prediction results.
[0044] According to one aspect of the present application, in step S42, the construction process of the non-parametric residual model is further as follows:
[0045] S42a, calculating the original residual sequence based on the historical measured value and predicted value of the total runoff into the lake;
[0046] S42b. Fit the regression function and conditional volatility function of the residuals using a local linear estimator, and optimize the smoothness of the local linear regression using an adaptive bandwidth selector;
[0047] S42c. Use an autoregressive model to capture the temporal correlation of the standardized residuals, and fit the distribution of the random error terms using a mixture Gaussian distribution of positions. Generate residual samples according to the distribution.
[0048] Beneficial effects: The present application proposes a systematic solution to the key technical problems in the prediction of the runoff into large lakes in river network areas. Based on the quantitative assessment of the impact of human activities, the tributaries flowing into the lake are innovatively divided into three types of impact branches: low, medium, and high. Differentiated modeling is achieved according to the characteristics of the branches, and the iterative prediction of the runoff of the high-impact branches is realized using the prediction results of the low- and medium-impact branches; aiming at the problem of lack of measured data in the uncontrolled section, the deterministic and probabilistic joint prediction of the total runoff into the lake is realized by constructing a conditional probability distribution model, and the prediction results are further optimized using an error correction algorithm. This application can effectively improve the prediction accuracy of the medium- and long-term runoff into the lake and provide more reliable information for water resources planning and management. Description of the Drawings
[0049] Figure 1 It is a flowchart of the present invention.
[0050] Figure 2 It is a flowchart of step S1 of the present invention.
[0051] Figure 3 It is a flowchart of step S3 of the present invention.
[0052] Figure 4 It is a flowchart of step S4 of the present invention. Detailed Embodiments
[0053] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the implementation of the present invention does not depend on all the details described below. Some technical features may have been simplified or omitted to avoid interfering with the core content of the present invention. In addition, the order of the steps involved in the present invention is not fixed, and those skilled in the art can adjust it according to actual needs. The following described order is only an exemplary implementation scheme, aiming to help understand the technical idea of the present invention and does not constitute any limitation to the implementation manner of the present invention.
[0054] As Figures 1 to 4 shown, a method for predicting the medium- and long-term runoff into a large lake in a river network area is provided, including the following steps:
[0055] S1. Collect the data of the study area, determine the degree of influence of human activities on the tributaries flowing into the lake, and divide the tributaries into three types of impact branches: low, medium, and high;
[0056] S2. Respectively use the SWAT model and the physically constrained LSTM model to predict the low- and medium-impact branches, and obtain the runoff prediction results of the low- and medium-impact branches;
[0057] S3. Screen the runoff prediction results of the low- and medium-impact branches that are significantly correlated with the high-impact branch as prediction factors, construct a DBN-BP model to establish the correlation between the runoff of the high-impact branch and the prediction factors, and realize the iterative prediction of the runoff of the high-impact branch;
[0058] S4. Use the conditional normalizing flow model to construct the conditional probability distribution between the sum of the runoff of the three types of branches and the total inflow runoff into the lake, obtain the deterministic and probabilistic prediction results of the total inflow runoff into the lake considering the uncontrolled interval, and construct a non-parametric residual model to achieve error correction.
[0059] The applicable object of this embodiment is large lakes in river network areas, especially complex river networks with numerous tributaries flowing into the lake. Using a single model to directly predict the inflow runoff into the lake has high modeling difficulty and large workload, and it is difficult to obtain satisfactory prediction results. Generally speaking, the inflow runoff into a lake can be regarded as the convergence of runoff from different branches. Therefore, when predicting the inflow runoff into the lake, the tributaries can be divided into three categories according to the degree of human activity influence, and appropriate models and methods can be selected respectively to predict the runoff of the corresponding branches.
[0060] To divide the tributaries flowing into the lake into three types of impact branches: high, medium, and low, it is first necessary to quantify the degree of human activity influence on each tributary. In the present invention, a subjective and objective combined weighting method is used to evaluate the weights of important indicators of human activities. Among them, the subjective weighting method uses the modified fuzzy RANCOM method, considers the uncertainty of expert judgment by introducing triangular fuzzy numbers, and at the same time uses a dynamic consensus strategy to quantify the expert consensus degree and correct low-consensus judgments, so as to reduce the weight error caused by individual expert biases or misjudgments and weaken the influence of extreme opinions on the final result.
[0061] There is a certain interference of human activity influence on the runoff of the medium-impact branch. It is difficult to obtain satisfactory prediction results using a single physical hydrological model. The present invention combines the WEAP model and the LSTM model to predict the runoff of the medium-impact branch, uses the WEAP model to quantify the human activity influence, takes the output of the WEAP model and meteorological data as the input of the LSTM model, and based on the physical variable sub-processes such as the human activity correction amount output by the WEAP model and the natural runoff volume in the situation without human activities, adds three physical constraint terms of water balance, reservoir operation, and land use change to the loss function of the LSTM model, significantly enhancing the physical consistency of the model, thereby realizing the deep integration of physical mechanisms and data-driven, and significantly improving the prediction accuracy of the runoff of the branch affected by a certain degree of human activity, overcoming the defect that pure data-driven models are prone to produce unreasonable prediction results.
[0062] Since the runoff processes of the three types of influencing branches have certain similarities due to their similar climate and underlying surface conditions, the present invention creatively uses the runoff prediction results of the low- and medium-influencing branches to achieve the iterative prediction of the runoff of the high-influencing branch. Since there are numerous runoff prediction results for the low- and medium-influencing branches, screening effective and non-redundant factors as model inputs has a significant impact on the output of the model results. The present invention establishes a new factor screening method by integrating Laplacian score and stepwise regression model. First, the Laplacian score is used to preliminarily screen the features, and then the stepwise regression model is used to dynamically screen by adding the factors with the smallest Laplacian score into the model one by one. This method can consider the cross-correlation and autocorrelation between different feature subsets, significantly improve the factor screening efficiency and generalization ability, and is more suitable for the screening of high-dimensional data, laying a data foundation for the iterative prediction of the runoff of the high-influencing branch.
[0063] Taking the optimized factor set of the high-influencing branch obtained by screening as the input, the DBN-BP model is used to predict the runoff of the high-influencing branch. During the process of training the neural network model, it is necessary to optimize the hyperparameters. In the present invention, the slime mold algorithm is selected to improve the sooty tern optimization algorithm to optimize the hyperparameters in the DBN-BP model. By introducing the foraging behavior of the slime mold to update the position, the balance ability of the algorithm between global search and local exploitation is enhanced. This fusion mechanism breaks through the limitation of the traditional sooty tern optimization algorithm that only relies on the migration and predation behaviors of the sooty tern, guides individuals to jump out of the local optimum and move towards the high-quality area, thereby improving the convergence speed and making the algorithm more suitable for high-dimensional nonlinear optimization problems.
[0064] Excluding the already divided low-, medium-, and high-influencing branches, the runoff of a part of the uncontrolled interval needs to be considered. The runoff of the uncontrolled interval lacks actual observed data and is difficult to directly predict. In the present invention, through the conditional normalizing flow model, the conditional probability distribution between the sum of the runoff of the three branches and the actual inflow into the lake is established, and then the total inflow prediction result considering the uncontrolled interval is obtained based on the runoff prediction results of the three branches, breaking through the limitation of the traditional method that relies on the direct monitoring data of the uncontrolled interval and realizing the indirect quantification of the runoff contribution of the uncontrolled interval. On the basis of outputting the deterministic prediction and probability prediction results of the inflow into the lake considering the uncontrolled interval, due to a series of uncertain factors such as model parameters and structures, there must be an error between the output result and the actual observed value. To further improve the accuracy of the prediction result, the present invention uses a non-parametric residual model to generate a residual sample sequence to correct the error of the runoff prediction result. By directly inferring the statistical characteristics of the residuals from the data through local linear regression and kernel density estimation, the prior assumption of the functional form of the traditional parametric residual model is abandoned. This design avoids the risk of model misspecification, is suitable for the highly nonlinear and non-stationary scenarios of the hydrological system, significantly improves the reliability of the deterministic prediction and the risk quantification ability of the probability prediction, and provides key technical support for the refined decision-making of water resources management.
[0065] As Figure 2 shown, as a preferred embodiment, step S1 specifically includes the following steps:
[0066] S11. Collect data of the study area;
[0067] The data includes long-term mainstream, tributary and lake inflow runoff data provided by hydrological stations in the river network area, meteorological data such as precipitation, temperature, air pressure, humidity, wind speed, sunshine duration, etc. provided by meteorological stations in the river network area, and physical mechanism factor data such as water system distribution characteristics, soil data, digital elevation model, land use data, etc. obtained through geographical information databases, remote sensing images, geological surveys, etc. At the same time, necessary review and preprocessing are carried out on the collected data to ensure the integrity, accuracy and timeliness of the data, so as to understand the basic situation of the study area and provide a reliable data basis for subsequent analysis and modeling.
[0068] S12. Divide the catchment areas of each tributary flowing into the lake in the river network, establish a set of human activity index factors within the catchment areas, and screen out important human activity indicators;
[0069] A catchment area is the basic unit of the hydrological cycle, covering all areas that supply water to a certain river or lake. Based on the digital elevation model, each water system flowing into large lakes in the river network area is extracted and corrected, and the river channel flow direction is determined through hydrological analysis algorithms, and then the catchment areas of each water system flowing into the lake in the river network are divided;
[0070] To quantify the degree of influence of human activities on tributaries flowing into the lake in the river network, factors that can reflect the degree of influence of human activities on the catchment areas of each water system flowing into the lake are selected as index factors to establish a set of human activity index factors, including land use change, quantity and scale of water conservancy projects, human water withdrawal level, regional economic development level, and water ecosystem health status;
[0071] In a certain embodiment, the grey relational analysis is used to evaluate the relative importance of each human activity index factor on the lake inflow runoff, and then important human activity indicators are screened out. The specific steps are as follows: Given the sequences of each index factor in the lake catchment area and the sequence of lake inflow runoff, calculate the grey relational coefficients between the two to obtain the grey relational coefficient matrix, and then calculate the average value of each column of the grey relational coefficient matrix to obtain the grey relational degree of each index. Select the top 5 indicators with the largest grey relational degree as important human activity indicators.
[0072] S13. Based on the important human activity indicators, construct an evaluation matrix of human activities of tributaries flowing into the lake in the river network, and use the modified fuzzy RANCOM method and CRITIC method to determine the subjective and objective weights of each indicator respectively, and combine them to obtain the comprehensive weight of each indicator;
[0073] Based on the selected important indicators of human activities, an initial evaluation matrix R=(r pq ) P×Q is constructed for the tributaries flowing into the lake affected by human activities, where r pq represents the index value, p and q represent the tributary number and the index number respectively, and P and Q are the total numbers of tributaries and indicators. The initial evaluation matrix is normalized to obtain a standardized evaluation matrix to make the indicators comparable.
[0074] The modified fuzzy RANCOM method is used to determine the subjective weights of each important indicator of human activities: the modified fuzzy RANCOM method deals with the uncertainty in the subjective decision-making process by introducing triangular fuzzy numbers and adjusts the expert weights through dynamic consensus evaluation. The method steps are further as follows:
[0075] S13a. Experts rank the indicators according to the importance of the indicators; the smaller the ranking value, the more important the indicator. Different indicators are allowed to have the same ranking value to reflect the equal judgment of experts on the importance of some indicators;
[0076] S13b. Compare the ranking values of the indicators pairwise to generate a ranking comparison matrix; the determination rule of the matrix element A ij is as follows: if the indicator C i is more important than C j , then A ij =1; if the indicators C i and C j have the same importance, then A ij =0.5; if the indicator C i is less important than C j , then A ij =0;
[0077] S13c. Introduce the triangular fuzzy theory to convert the exact values in the permutation comparison matrix into triangular fuzzy intervals, generate a fuzzy ranking comparison matrix considering the uncertainty of expert judgments, and use the dynamic consensus strategy to identify and correct the evaluation results of low-consensus experts;
[0078] The specific conversion rule of the triangular fuzzy interval is as follows: if A ij =1, it is converted to [0.5, 1, 1]; if A ij =0.5, it is converted to [0, 0.5, 1]; if A ij =0, it is converted to [0, 0, 0.5];
[0079] For each expert, calculate the deviation CN k between his judgment and the average value. The formula is CN k =1 - {1 / [Q(Q - 1)]}×D(A ij k , A ijavg ) Definition, where D(A ij k , A ij avg ) is the triangular fuzzy distance between A ij k and A ij avg , and A ij avg is the weighted average of the sorting comparison results, and A ij k is the judgment result of the k-th expert, and Q is the total number of indicators. In a specific embodiment, if there is an expert who satisfies C k < 0.7, it is marked as a low-consensus expert, and the m indicators with the largest difference from the group in the fuzzy matrix of the low-consensus expert are re-evaluated, and then the corrected fuzzy sorting comparison matrix is obtained.
[0080] S13d. According to the triangular fuzzy number addition rule, sum each row of the corrected fuzzy sorting comparison matrix to obtain the fuzzy comprehensive weight of each indicator, and perform normalization processing to obtain the final fuzzy weight, and use the core value of the fuzzy weight as the subjective weight result of the indicator. The core value is the middle value of the triangular fuzzy interval.
[0081] The CRITIC method determines the objective weight of the indicator by calculating the variability and conflict of the indicator. The standard deviation is used to measure the degree of variability of the indicator. The larger the coefficient of variation of the standard deviation, the more information the indicator provides and the greater the weight. The conflict is calculated by the absolute value of the correlation coefficient between indicators. The larger the correlation coefficient, the higher the degree of information overlap between indicators and the smaller the weight. The product of the indicator variability and conflict is the information volume of the indicator. Normalize the information volume of each indicator to obtain the final objective weight.
[0082] The combined weight of each indicator is obtained by using the combined weighting method based on game theory. The core idea of this method is to integrate the subjective and objective weights of each indicator through a linear combination method to achieve the optimal coordination of different weighting methods and make the weighting result more scientific and reasonable. To minimize the deviation between the combined weight and the subjective and objective weights, the optimization objective is set as min∣∣l1·ω1 T + l2·ω2 T - ω c T ∣∣2, c = 1, 2; l1 and l2 are the linear combination coefficients of the subjective weight and the objective weight respectively, and ω1 and ω2 are the subjective weight and the objective weight respectively. A linear equation system is established through the optimal first derivative condition, and the optimal weight combination is solved, and the final combined weight W is obtained through normalization processing. From the formula W = l1 * · ω1 + l2 * · ω2; Definition, where l1* = l1 / (l1 + l2), l2 * = l2 / (l1 + l2).
[0083] S14. According to the comprehensive weights of each index, use the TOPSIS method to rank the degree of human activity influence on each tributary, and accordingly divide the tributaries into three categories of influence branches: low, medium, and high.
[0084] The TOPSIS method determines the degree of human activity influence on each tributary by calculating the relative closeness of each tributary to the tributary with the greatest degree of human activity influence. The specific steps are as follows:
[0085] Combine the index comprehensive weight W and the standardized evaluation matrix R = (r pq ) P×Q to construct the weighted decision matrix H = (h pq ) P×Q , where h pq = W q ·r pq ; calculate the positive ideal solution H + =(h1 + , h2 + ,...., h q + ,...., h Q + ) and the negative ideal solution H - =(h1 - , h2 - ,...., h q - ,...., h Q - ), where, h q + = max p (h pq ), h q - = min p (h pq ). Calculate the Euclidean distances d p + and d p - of each tributary from its positive and negative ideal solutions respectively. Finally, calculate the relative closeness δ p between tributary p and the tributary with the maximum human influence. It is defined by the formula δ p = d p - / (d p + + d p - ), δ p∈[0,1], the larger the value, the higher the degree of influence of the tributary by human activities.
[0086] According to the relative proximity δ p Sort all the tributaries flowing into the lake from high to low. In a certain embodiment, the degree of influence of the tributaries flowing into the lake by human activities is divided according to the following thresholds: δ p ≥0.7 is a high-impact branch, 0.4 ≤ δ p <0.7 is a medium-impact branch, δ p <0.4 is a low-impact branch. Accordingly, all the tributaries flowing into the lake in the river network area are divided into three categories: high-impact branches, medium-impact branches and low-impact branches.
[0087] As a preferred embodiment, step S2 specifically includes the following steps:
[0088] S21. Invoke the pre-constructed SWAT model to predict the runoff of the low-impact branches, and obtain the deterministic prediction results of the runoff of the low-impact branches;
[0089] The low-impact branches are mainly controlled by natural factors, are less disturbed by human activities, have strong regularity, and a physical model can be used to accurately predict the runoff of such branches. As a distributed hydrological model, the SWAT model can simulate different physical processes of the water cycle in the basin. According to the variability of the underlying surface and climate in the basin in time and space, the SWAT model divides the study area into several sub-basins, and divides several hydrological response units HRUs in each sub-basin. By calculating the runoff of each HRU to check the runoff at the outlet section, the accuracy of runoff simulation is enhanced.
[0090] The simulation of the basin hydrological process by the SWAT model can be divided into two parts: overland runoff generation and channel flow concentration. Among them, the overland runoff generation stage follows the principle of water balance:
[0091] SW t = SW0 + Σ t i=1 (P day,i - Q surf,i - E a,i - W seep,i - Q gw,i ).
[0092] In the formula, SW t is the soil water content on the t-th day, SW0 is the initial soil water content, P day,i is the precipitation on the i-th day, Q surf,i is the surface runoff depth on the i-th day, E a,i is the evapotranspiration on the i-th day, W seep,i is the infiltration and lateral flow of the soil profile stratum on the i-th day, Q gw,iis the groundwater content on the i-th day. The SCS curve method is used to calculate surface runoff, the Hargreaves method is used to calculate potential evapotranspiration, the infiltration and lateral flow of soil water are predicted through a kinematic storage model, and two groundwater aquifer systems are divided to simulate groundwater content.
[0093] During the river confluence stage, the Muskingum method is used to check the river flow, and the formula is: Q t = C0·I t + C1·I t-Δt +C2·Q t-Δt . In the formula, Q t is the outflow at time t of the river reach, I t is the inflow at time t of the river reach, Q t-Δt is the outflow at time (t - Δt) of the river reach, I t-Δt is the inflow at time (t - Δt) of the river reach, and C0, C1, and C2 are functions of the Muskingum parameters.
[0094] During the training and calibration process of the SWAT model, a large number of parameters with different physical meanings are involved. To improve the accuracy of water cycle simulation and the efficiency of parameter calibration, the present invention uses the global sensitivity analysis module in the SUFI-2 algorithm in SWAT-CUP to perform sensitivity analysis on the model parameters, screens out the parameters that have a more obvious impact on the runoff of the study basin for calibration, and calls the model to output the deterministic prediction results of the low-impact branch runoff.
[0095] S22. Use a physically constrained LSTM model to predict the medium-impact branch runoff and obtain the deterministic prediction results of the medium-impact branch runoff;
[0096] The medium-impact branch is moderately disturbed by human activities, and the runoff process still retains certain natural hydrological laws. A physically constrained LSTM model can be used to predict the medium-impact branch runoff. By integrating physical constraint terms into the deep learning model to achieve collaborative optimization inside the network, the accuracy, efficiency, and interpretability of the medium-impact branch runoff prediction are improved. Among them, the physical constraint terms ensure the physical consistency of the prediction results by considering water balance, reservoir operation, and land use changes, avoiding unreasonable negative runoff predictions, and the data-driven module effectively captures complex impacts such as human activity disturbances through the powerful non-linear fitting ability of machine learning.
[0097] According to one aspect of the present application, in the step S22, the process of using a physically constrained LSTM model to predict the medium-impact branch runoff is further as follows:
[0098] S22a. Input hydrometeorological data and human activity parameters, use the WEAP model to simulate the medium-impact branch runoff, and generate runoff sequences and sub-process physical variables considering the impact of human activities;
[0099] The Water Evaluation and Planning model (WEAP) is a semi - distributed hydrological model for integrated water resources management. Its core is to simulate the supply - demand relationship of water resources in the basin through the principle of water balance. For hydrological modeling, the WEAP model uses a soil moisture model, dividing the basin into two layers of soil. The root zone simulates precipitation infiltration, evapotranspiration, surface runoff, and shallow infiltration, while the deep soil simulates base flow and deep percolation. As a semi - distributed hydrological model, WEAP divides the basin into multiple sub - regions, each of which can be defined with different soil types and land uses. Each sub - region calculates the water balance independently, and then integrates the whole - basin simulation through river network confluence.
[0100] The input data of the WEAP model include hydrometeorological data such as precipitation, temperature, humidity, wind speed, and runoff. To simulate the impact of human activities on the branch runoff in the middle reach, human activity parameters such as water withdrawal in the basin, reservoir operation rules, and land - use changes still need to be input. The model output includes the runoff sequence considering the impact of human activities and physical sub - process variables. The physical sub - process variables include natural runoff volume under the scenario of no human activities, human activity correction volume, and reservoir regulation volume, etc.
[0101] S22b: Take the meteorological data and the output of the WEAP model as inputs, construct an LSTM model, and add physical condition constraints to the loss function of the LSTM model based on the output of the WEAP model to obtain the deterministic prediction result of the branch runoff in the middle reach.
[0102] The Long Short - Term Memory neural network (LSTM) is a recurrent neural network model. Its core structure includes three gating mechanisms: an input gate, a forget gate, and an output gate, which can effectively process time - series data. Input the meteorological data and the physical sub - process variables output by the WEAP model. By adding physical constraint terms to the loss function, embed hydrological theory into the LSTM model to make up for the deficiency of pure data - driven. The physical constraint terms include water balance constraints, reservoir operation constraints, and land - use change constraints.
[0103] To ensure that the predicted runoff of the model conforms to the water balance equation after considering human activities, construct a water balance constraint. The mean square error considering the water balance is: MSE wb = (1 / T)Σ t=1 T [Q LSTM,t - (Q natural,t - S t )] 2 ;
[0104] where Q LSTM,t is the runoff volume at time t predicted by the LSTM model, Q natural,t is the natural runoff volume at time t simulated by WEAP under the scenario of no human activities, and S tThe correction of human activities at time t for WEAP simulation, and T is the total time period.
[0105] To make the predicted runoff meet the reservoir operation rules, reservoir operation constraints are added. The upper bound constraint is that the predicted runoff does not exceed the flood control limit of the reservoir, which is defined by the formula Q LSTM,t ≤Q max The lower bound constraint is that the predicted runoff needs to be greater than the downstream ecological flow limit, which is defined by the formula Q LSTM,t ≥Q min Then the mean square error considering reservoir operation is: MSE re = (1 / T)Σ T t=1 [ReLU(Q min -Q LSTM,t ) + ReLU(Q LSTM,t - Q max )] 2 Among them, Q min is the minimum flow of downstream ecological or domestic water demand, and Q max is the maximum allowable flow for flood control safety. When the predicted flow violates any one of the constraints, a penalty is imposed through the ReLU function.
[0106] To reflect the impact of land use change on runoff, land use change constraints are added, which characterize the changes in runoff generation and concentration caused by urbanization and vegetation restoration in the affected branch catchments. The mean square error considering land use change is: MSE lu = (1 / T) Σ T t=1 (Q LSTM,t - Q lu,t ) 2 ; among them, Q lu,t is the runoff at time t under a specific land use scenario simulated by the WEAP model.
[0107] Combining the traditional mean square error and physical constraint terms, a composite loss function for LSTM is generated through weighted summation:
[0108] LOSS = λ data MSE data + λ wb MSE wb + λ re MSE re + λ lu MSE lu; where λ is a weight function that needs to be optimized through cross-validation. The output of the training model affects the deterministic prediction results of branch runoff.
[0109] like Figure 3 As shown, as a preferred embodiment, step S3 specifically includes the following steps:
[0110] S31. Use Laplace stepwise regression method to screen out a group of low- and medium-impact branches with significant correlation with the runoff of high-impact branches;
[0111] Since they are in the same river network area, high-impact branches have similar underlying surface and climate conditions as low- and medium-impact branches, and their runoff has certain spatiotemporal correlation characteristics. The runoff prediction results of low- and medium-impact branches can be used to realize the reciprocal prediction of the runoff value of high-impact branches, thereby making up for the shortcomings of insufficient data of high-impact branches, reducing monitoring costs and improving prediction efficiency.
[0112] First, low- and medium-impact branches that are significantly correlated with the high-impact branch runoff to be predicted are screened out. According to one aspect of the present application, in step S31, the Laplace stepwise regression method screening process is further as follows:
[0113] S31a, calculating the Laplace scores between each low-impact and medium-impact branch runoff and the high-impact branch runoff, and arranging them in ascending order;
[0114] The Laplace score is used to measure the correlation between low- and medium-impact branches and high-impact branches. The smaller the Laplace score, the more important the branch is. This is used to select branches with larger variance and that can retain the local manifold structure of the data as candidate predictors. The main steps are as follows:
[0115] The K nearest neighbor algorithm is used to determine the neighborhood relationship of high-impact branch runoff samples;
[0116] Calculate the weight F between samples ij , construct the weight matrix, if the feature m i is m j The K nearest neighbors of F ij = exp(-||m i -m j ||² / g), otherwise, F ij = 0, where ||m i -m j || is m i With m j The Euclidean distance between them, g is a constant, generally 1.
[0117] Runoff characteristics for low and medium impact branches r = [f r1 , f r2, ..., f rn ] is standardized to obtain f - r
[0118] Construct the Laplace matrix LS=EF, where E is the diagonal matrix and F is the weight matrix. Then the Laplace score expression of the rth influencing branch is LS r =(f - r T · LS·f - r / f - r T E f - r ); Arrange the branches in ascending order of Laplace scores. The smaller the score, the more important the branch.
[0119] S31b. Construct a stepwise regression model. In each iteration, select the factor with the smallest Laplace score from the remaining factor set to add to the model, and eliminate the least significant factor obtained by the Bootstrap hypothesis test.
[0120] The stepwise regression model introduces the most significant factors that affect the predicted object one by one, tests the factors introduced in the model one by one, eliminates the insignificant factors, and finally uses the selected related factors to establish a multivariate regression prediction equation. Starting from the empty model, factors are gradually added, which is divided into two processes: forward selection and backward elimination.
[0121] Among the remaining unselected factors, the factor with the smallest Laplace score is selected to be added to the model; for all factors in the currently selected factor set, Bootstrap resampling is used to generate M resampled data sets, and the regression coefficient distribution of each factor is calculated; if the factor passes the hypothesis test, it is retained, otherwise the model is eliminated.
[0122] S31c. When the model performance is no longer significantly improved, the screening is stopped to obtain a set of low- and medium-impact branches that are significantly correlated with the runoff of high-impact branches.
[0123] S32, respectively calculating the runoff prediction values of the selected low- and medium-impact branches, and constructing a set of candidate prediction factors for high-impact branches based on the values, and then using the MIC method to screen to obtain a set of preferred prediction factors for high-impact branches;
[0124] The predicted values of low- and medium-impact branch runoff obtained by screening are determined in step S2. In a specific embodiment, the set of candidate factors for high-impact branches includes the predicted runoff values of the low- and medium-impact branches for the current month and the previous 11 months. To improve the efficiency of predicting high-impact branch runoff, the MIC method is used to screen the set of optimal factors for high-impact branches; MIC is a two-dimensional variable correlation metric based on mutual information. Its main principle is to grid the sample space composed of variables, and the calculation formula is MIC(X;Y) = max αβ<B {MI(X;Y) / log2(min(α,β))};
[0125] where MI(X;Y) represents the mutual information between feature X and variable Y, which is defined by the formula MI(X;Y) = ∫ρ(X;Y)log(ρ(X,Y) / (ρ(X)ρ(Y)))dXdY. ρ(X,Y) is the joint probability density function of X and Y, and ρ(X) and ρ(Y) are the marginal probability distribution functions of X and Y respectively. α and β are parameters related to grid division, and B is the upper limit value of variable grid division, usually taking the 0.6th power of the data volume. The top n runoff predicted values with larger MIC values are selected as the set of optimal factors for high-impact branches.
[0126] S33. Use the improved Sooty Tern optimization algorithm with the slime mold algorithm to optimize the hyperparameters of the DBN-BP model, and use the set of optimal factors for high-impact branches as the model input to obtain the deterministic prediction result of high-impact branch runoff.
[0127] The DBN model consists of multiple restricted Boltzmann machines (RBMs). Through an unsupervised learning framework, the RBMs are trained layer by layer to extract features such as weight matrices and offsets. A backpropagation (BP) neural network is added to the last layer of the DBN to fine-tune the output of the DBN network model, and finally a supervised deep learning network for training is formed.
[0128] The training of the DBN-BP model is divided into two processes: forward training and backward fine-tuning. The contrast divergence algorithm is used to train the RBMs layer by layer, and the output of the hidden layer of the previous layer is used as the input of the next layer until the last layer of the RBM, so as to update the weight matrix and bias; the weight matrix and bias output by the last layer of the RBM are input into the BP neural network, and the network parameters are adjusted by backpropagating the error, and the fine-tuning of each layer of the RBM is completed.
[0129] Before training a neural network model, it is necessary to set hyperparameters first, that is, to select some configuration variables that have a significant impact on the model performance and training effect. In the present invention, the initial learning rate, the number of hidden layer nodes, the number of neural network training times, the initial weight, the training batch size, the regularization coefficient, and the dropout rate are selected as hyperparameters, and the slime mold algorithm is used to improve the sooty tern optimization algorithm for parameter optimization, so as to improve the overall reliability and generalization ability of the prediction model. The sooty tern optimization algorithm (STOA) is a bio-inspired optimization algorithm that optimizes the objective function by simulating the migration and foraging behaviors of sooty terns. By integrating the foraging mechanism of slime mold to improve STOA, the performance of the algorithm in terms of global search ability and local development accuracy is effectively improved, and it has a faster convergence speed and higher stability and robustness.
[0130] According to one aspect of the present application, in the step S33, the optimization process of improving the sooty tern optimization algorithm by the slime mold algorithm is further as follows:
[0131] S33a. Initialize the individuals in the population using the Circle chaotic map; the Circle chaotic map replaces random initialization to generate the initial population, solves the problem of uneven initial distribution of traditional algorithms, and enhances the diversity and coverage of the population;
[0132] S33b. Adjust the individual positions to avoid collisions through the movement variable, introduce the foraging behavior of slime mold to make the individuals gradually converge to the optimal position, and update the individual positions;
[0133] In the sooty tern optimization algorithm, a movement variable Mv is introduced to adjust the individual positions to avoid individual collisions, which is defined by P c =Mv×P s (iter), where P c is the position where no collision occurs between individuals, P s (iter) is the current individual position, Mv = Cv - {(iter / iter max ) × Cv}; Cv is a control variable, generally taken as 2, iter is the current iteration number, and iter max is the maximum iteration number;
[0134] Simulating the behavior of seabirds gradually approaching food, the individuals gradually converge to the optimal position. The original convergence formula is defined by Ms = r1 × (P best (iter) - P s (iter)), where Ms is the distance for an individual to move from the current position to the best position, P best (iter) is the current optimal solution, and r1 is a random number in the range [0, 0.5]; update the individual positions according to the results of collision avoidance and position convergence, and the update formula is P new =P c +Ms;
[0135] Introduce the foraging behavior of slime mold to update the position and enhance the global search ability. Define r2 as a random number uniformly distributed in [0,1], and po as the triggering probability threshold of the foraging behavior of slime mold. When r2 < po, P new (iter) = P best (iter) + vb×(V×P best (iter) - P s (iter)); otherwise P new (iter) = vc×P s (iter), where P new (iter) is the position update result at the current moment, vb and vc are random coefficients, controlling the direction and amplitude of position update, V is the slime mold weight factor, simulating the attraction intensity of slime mold to the food source, and is defined by the formula V = 1 + log{[(f(P best (iter)) - f(P s (iter))) / f(P best (iter))]} + 1}, where f is the fitness function.
[0136] S33c. Further update the individual position by simulating the predation behavior of sooty terns flying in a spiral;
[0137] Update the individual position by simulating the predation behavior of seabirds flying in a spiral in the air. During the spiral process, the position changes along the x-axis, y-axis, and z-axis are defined as: X’ = Rd×sin(θ); Y’ = Rd × cos(θ); Z’ = Rd ×θ;
[0138] where the radius is Rd = u× exp(ηv), and u, v, η are parameters defining the spiral shape, representing the initial radius, random number, and decay coefficient respectively. Therefore, considering the predation and migration behaviors, the individual position update is defined by P s (iter) = (P new (iter)×(X’ + Y’ + Z’))×P best (iter);
[0139] S33d. Terminate the optimization algorithm when the maximum number of iterations is reached, and return the current solution set as the hyperparameter optimization result.
[0140] As Figure 4 shown, as a preferred embodiment, step S4 specifically includes the following steps:
[0141] S41. Based on the sum of the historical measured values of the three types of branch runoff and their corresponding historical measured values of the inflow runoff into the lake, a conditional probability distribution between the two is constructed using a conditional normalizing flow model. The sum of the prediction results of the three types of branch runoff is used as the conditional input, and the probability prediction result of the total inflow runoff into the lake considering the uncontrolled interval is output;
[0142] Based on the prediction of the three types of branch runoff at low, medium, and high levels, it is still necessary to consider the inflow volume in the uncontrolled interval from the downstream of the hydrological monitoring section to the target control section (here referring to the lake inlet). Due to the limitations of actual hydrological monitoring conditions, it is usually impossible to directly set up a hydrological station at the connection between the river and the lake, resulting in a lack of direct observation data on the runoff changes (including direct runoff from precipitation, groundwater recharge, and the inflow of unmonitored tributaries, etc.) in this interval, making it one of the sources of uncertainty in runoff prediction. Based on the joint distribution characteristics of the sum of the three types of branch runoff and the total inflow runoff into the lake, this invention constructs a conditional probability distribution model, realizes the indirect estimation of the inflow volume in the uncontrolled interval, and finally obtains the probability prediction result of the total inflow runoff into the lake considering the uncontrolled interval.
[0143] Conditional normalizing flow is a generative model that maps a simple base distribution p z (z) to a complex conditional distribution p X (x|y) through a series of reversible neural network transformations. In a certain embodiment, the conditional input y is the sum of the historical measured values of the three types of branch runoff, the target output x is the historical measured value of the inflow runoff into the lake, and the conditional normalizing flow model is defined as x = f θ (z; y), where z ~ p z (z), f θ is a parameterized reversible transformation, and its inverse transformation is z = f θ -1 (x; y). According to the law of probability conservation, the conditional probability density formula is p X (x|y) = p Z (f θ -1 (x; y))|det(df θ -1 (x; y) / dx)|, where det(df θ -1 (x; y) / dx) is the Jacobian determinant of the transformation f θ .
[0144] The training objective of the conditional normalizing flow model is to maximize the conditional probability p X(x|y) likelihood function, with negative log-likelihood as the loss function, and the parameters of the conditional network and the flow model are gradually updated using the gradient descent method. In the prediction stage, the sum of the predicted values of the three types of branch runoff is used as the conditional input, and the trained conditional normalizing flow model is used to generate the probability prediction result of the total lake inflow considering the uncontrolled interval, and the conditional mean is taken as the deterministic prediction result of the total lake inflow.
[0145] S42. Construct a non-parametric residual model based on the error between the predicted value and the measured value of the total lake inflow to generate residual samples;
[0146] According to one aspect of the present application, the process of constructing a non-parametric residual model to generate residual samples in step S42 is further as follows:
[0147] S42a. According to the measured value y of the total lake inflow t and the predicted value y t * , calculate the original residual sequence e t =y t -y t * , and use it as the input of the non-parametric residual model;
[0148] S42b. Construct a non-parametric residual model, and estimate the regression function m of the residual n (R t ) and the conditional volatility function σ n (R t ), where R t =ln(y t * -1) is the logarithmic transformation of the deterministic prediction result, and the smoothness of the local linear regression is optimized using an adaptive bandwidth selector.
[0149] For each regression variable x, solve the local linear coefficients by minimizing the weighted squared error: (a * 1, b * 1) =argmin a,b {∑ n t=1 (e t -a-b(x-R t )) 2 ·K((x-R t ) / h(R t ))}, and the regression function estimate m n (x)=a * 1, a, b are the parameters to be optimized for local linear regression, representing the intercept and slope parameters respectively, a * 1 and b *1 is the optimized intercept and slope parameters, K(·) is the Gaussian kernel function, and h(R t ) is the adaptive bandwidth.
[0150] The conditional volatility function estimates by minimizing the objective function (a * 2, b * 2) = argmin a,b {∑ t=1 n (d * t - a - b(x - R t )) 2 ·K((x - R t ) / h(R t ))}, and the predicted value of the conditional volatility function σ n (x) = (a * 2) (1 / 2) , where d * t is the squared residual after detrending, represented by the formula d * t = (e t - m * n (R t )) 2 , σ n (x) is the conditional standard deviation of the residual, and a * 2 and b * 2 are the optimized intercept and slope parameters of the conditional volatility function, used to estimate the conditional variance.
[0151] S42c. Use the autoregressive model to capture the time correlation of the standardized residuals, and fit the distribution of the random error term with a location - mixture Gaussian distribution, and generate residual samples according to the distribution. The standardized residual is defined by the formula δ n,t = φ n ·δ n,t-1 + ε n,t , where φ n is the first - order autoregressive coefficient, and ε n,t is the random error term. Fit the distribution of the random error term with a location - mixture Gaussian distribution.
[0152] Use the Markov chain Monte Carlo method to generate samples of the random error term ε n,t r , generate the standardized residual δ n,t r through the AR(1) model, and generate the residual sample e t r = m * n (Rt ) +σ * n (R t )·δ n,t r 。
[0153] S43. Combine the residual samples with the deterministic prediction values to generate runoff prediction values after residual correction. Generate the probability distribution of the runoff prediction values through repeated sampling, calculate the confidence interval or quantile, and quantify the uncertainty of the prediction results.
[0154] Combine the residual samples with the deterministic prediction value y * t to generate the runoff prediction value y after residual correction t r* = y * t +e t r . Through multiple repeated samplings, generate the probability distribution of the runoff prediction values, and use kernel density estimation to calculate the confidence interval or quantile to complete the quantification of the prediction uncertainty.
[0155] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.
Claims
1. A long-term prediction method for the runoff into large lakes in a river network area, characterized in that, Including: S1. Collect data of the study area, determine the degree of human activity influence on the tributaries flowing into the lake in the river network, and divide the tributaries into three types of influence branches: low, medium, and high; S2. Use the SWAT model and the physically constrained LSTM model to predict the low and medium influence branches respectively, and obtain the runoff prediction results of the low and medium influence branches; S3. Screen the runoff prediction results of the low and medium influence branches that are significantly correlated with the high influence branch as prediction factors, construct a DBN-BP model to establish the correlation between the runoff of the high influence branch and the prediction factors, and realize the iterative prediction of the runoff of the high influence branch; S4. Use the conditional normalizing flow model to construct the conditional probability distribution between the sum of the runoff of the three types of branches and the total runoff into the lake, obtain the deterministic and probabilistic prediction results of the total runoff into the lake considering the uncontrolled interval, and construct a non-parametric residual model to realize error correction; Step S3 includes the following steps: S31. Use the Laplace stepwise regression method to screen out a group of low and medium influence branches that are significantly correlated with the runoff of the high influence branch; S32. Calculate the predicted values of the runoff of the screened low and medium influence branches respectively, and construct an alternative prediction factor set for the high influence branch accordingly. Then use the MIC method to screen and obtain the optimal prediction factor set for the high influence branch; S33. Use the improved sooty tern optimization algorithm with the slime mold algorithm to optimize the hyperparameters of the DBN-BP model, take the optimal factor set of the high influence branch as the model input, and obtain the deterministic prediction result of the runoff of the high influence branch.
2. The medium- and long-term prediction method for the incoming runoff of large lakes in a river network area according to claim 1, wherein Step S1 includes the following steps: S11. Collect data of the study area; S12. Divide the catchment area range of each tributary flowing into the lake in the river network, establish a set of human activity index factors within the catchment area, and screen out the important human activity indicators; S13. Based on the important human activity indicators, construct a human activity evaluation matrix for the tributaries flowing into the lake in the river network. Use the modified fuzzy RANCOM method and the CRITIC method to determine the subjective and objective weights of each indicator respectively, and combine them to obtain the comprehensive weight of each indicator; S14. According to the comprehensive weight of each indicator, use the TOPSIS method to rank the degree of human activity influence on each tributary, and divide the tributaries into three types of influence branches: low, medium, and high accordingly.
3. A long-term and medium-term prediction method for the inflow runoff into large lakes in a river network area according to claim 2, characterized in that The modified fuzzy RANCOM method in step S13 is further as follows: S13a. Experts rank each indicator according to the importance of the indicator; S13b. Compare the ranking values of the indicators pairwise to generate a ranking comparison matrix; S13c. Introduce the triangular fuzzy theory, transform the exact values in the ranking comparison matrix into triangular fuzzy intervals, generate a fuzzy ranking comparison matrix considering the uncertainty of expert judgment, and use the dynamic consensus strategy to identify and correct the evaluation results of low-consensus experts; S13d. According to the triangular fuzzy number addition rule, sum each row of the modified fuzzy ranking comparison matrix to obtain the fuzzy comprehensive weight of each indicator, and perform normalization processing to obtain the final fuzzy weight. Take the core value of the fuzzy weight as the subjective weight result of the indicator.
4. A long-term and medium-term inflow runoff prediction method for large lakes in a river network area according to claim 1, characterized in that, Step S2 includes the following steps: S21. Call the pre-constructed SWAT model to predict the runoff of the low influence branch, and obtain the deterministic prediction result of the runoff of the low influence branch; S22. Use the physically constrained LSTM model to predict the medium-influencing branch runoff and obtain the deterministic prediction result of the medium-influencing branch runoff.
5. A long-term and medium-term prediction method for the inflow runoff into large lakes in a river network area according to claim 4, characterized in that Step S22 is further as follows: S22a, input hydrological and meteorological data and human activity parameters, use WEAP model to simulate the impact branch runoff, and generate runoff sequence and sub-process physical variables considering the impact of human activities; S22b. Take meteorological data and WEAP model output as input, build LSTM model, and add physical condition constraints to the loss function of LSTM model based on the output of WEAP model to obtain the deterministic prediction result of middle-affected branch runoff.
6. A long-term and medium-term prediction method for the inflow runoff into large lakes in a river network area according to claim 1, characterized in that The Laplace stepwise regression screening process in step S31 is further as follows: S31a, calculating the Laplace scores between each low-impact and medium-impact branch runoff and the high-impact branch runoff, and arranging them in ascending order; S31b, construct a stepwise regression model, select the factor with the smallest Laplace score from the remaining factor set in each iteration to add to the model, and eliminate the least significant factor obtained by Bootstrap hypothesis test; S31c. When the model performance is no longer significantly improved, the screening is stopped to obtain a set of low- and medium-impact branches that are significantly correlated with the runoff of high-impact branches.
7. A long-term and medium-term prediction method for the inflow runoff of large lakes in a river network area according to claim 1, characterized in that The optimization process of the black tern optimization algorithm improved by the slime mold algorithm in step S33 is further as follows: S33a, use Circle chaos map to initialize individuals in the population; S33b, adjust the individual position by moving variables to avoid collision, introduce slime mold foraging behavior to make the individual gradually converge to the optimal position, and update the individual position; S33c, further updating individual positions by simulating the predatory behavior of the spiral flight of the sooty tern; S33d. When the maximum number of iterations is reached, the optimization algorithm is terminated and the current solution set is returned as the hyperparameter optimization result.
8. A long-term and medium-term prediction method for the inflow runoff into a large lake in a river network area according to claim 1, characterized in that Step S4 includes the following steps: S41. Based on the sum of the historical measured values of the three types of branch runoff and their corresponding historical measured values of the runoff into the lake, a conditional normalized flow model is used to construct a conditional probability distribution between the two, and the sum of the predicted results of the three types of branch runoff is used as a conditional input to output the total runoff into the lake probability prediction result considering the uncontrolled interval; S42, constructing a nonparametric residual model based on the error between the predicted value and the measured value of the total runoff into the lake, and generating residual samples; S43. Combine the residual samples with the deterministic prediction values to generate residual-corrected runoff prediction values. Generate the probability distribution of runoff prediction values through repeated sampling, calculate the confidence interval or quantile, and quantify the uncertainty of the prediction results.
9. A long-term and medium-term prediction method for the inflow runoff into large lakes in a river network area according to claim 8, characterized in that, The non-parametric residual model described in step S42 is further: S42a, calculating the original residual sequence based on the historical measured value and predicted value of the total runoff into the lake; S42b, fitting the regression function and conditional volatility function of the residuals through a local linear estimator, and optimizing the smoothness of the local linear regression using an adaptive bandwidth selector; S42c. An autoregressive model is used to capture the temporal correlation of the standardized residuals, and a location mixture Gaussian distribution is used to fit the distribution of the random error term, and residual samples are generated according to the distribution.
Citation Information
Patent Citations
Basin pollutant flux prediction method based on LSTM-BP space-time combination model
CN111639748A
Flood runoff forecasting method based on LSTM-SWAT coupling model
CN119474244A