A mid- and long-term runoff integrated probability prediction method based on deep learning

Through the combination of deep learning and an integrated framework, the uncertainty and interpretability problems in medium- and long-term runoff prediction are solved, and higher prediction accuracy and reliability are achieved, providing a scientific basis for water resource management.

CN119760557BActive Publication Date: 2025-05-06HOHAI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510249328.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-05-06
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

There are problems of insufficient uncertainty modeling, factor selection bias, model structure limitations and poor interpretability in medium and long-term runoff prediction, which affects forecast accuracy and reliability.

Method used

The medium- and long-term runoff integration probability prediction method is adopted based on deep learning, and a variety of factor screening methods are integrated through the improved Dempster-Shafer evidence theory, a deep learning basis model with multi-scale physically constrained adaptive loss function is constructed, and the multimodal adaptive Stacking ensemble framework and SHAP post-hoc interpretation method are used to improve the interpretability and prediction performance of the model.

Benefits of technology

It significantly improves the accuracy, reliability and practicality of medium- and long-term runoff prediction, and provides more scientific decision-making support for water resource management through full-factor optimization and uncertainty quantification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119760557B_ABST
    Figure CN119760557B_ABST
Patent Text Reader

Abstract

The present invention provides a medium- and long-term runoff integrated probability prediction method based on deep learning. By collecting and collating long-series runoff data in the basin, the improved time-lagged Pearson correlation coefficient method, the multi-scale maximum information coefficient method and the physical constraint-based random forest feature importance scoring method are used to screen the prediction factors, and the improved Dempster-Shafer evidence theory is used to fuse and construct the driving factor set. A deep learning base model is constructed based on a multi-scale physical constraint adaptive loss function, multi-modal data fusion and physical constraint enhancement mechanism are introduced, a multi-modal adaptive Stacking integration framework is established, the probability prediction interval is generated by the Bootstrap-Quantile method, and the key driving factors are identified by the SHAP method. The present invention proposes an "input-parameter-structure" full-factor optimization framework, which comprehensively considers multi-source uncertainties and improves prediction accuracy and reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to medium- and long-term hydrological forecasting in the field of water conservancy engineering technology, and in particular to a medium- and long-term integrated probability prediction method of multi-modal deep learning considering multi-source uncertainties and physical constraints. Background Art

[0002] Medium- and long-term runoff forecasting plays an important role in water resources planning and management. Accurate runoff forecasting can provide important guidance for optimal allocation of water resources and flood prevention and disaster reduction. However, due to the large temporal and spatial differences between climate, meteorology and underlying surface conditions, runoff presents complex characteristics such as high nonlinearity and non-stationarity in temporal and spatial scales, making medium- and long-term runoff forecasting extremely challenging. Therefore, it is imperative to develop robust and high-precision hydrological models and effectively apply them to actual scenarios to improve the scientificity and reliability of water resources management and scheduling.

[0003] In recent years, data-driven medium- and long-term runoff prediction methods have gradually become an important research direction in this field, especially the application of deep learning models has received widespread attention. Deep learning models have become a class of methods that cannot be ignored in basin runoff prediction due to their excellent modeling capabilities in dealing with complex nonlinear problems and large-scale data. Since the formation of runoff is affected by the coupling of multiple hydrological and meteorological elements, it can usually be regarded as a black box model with multiple inputs and single outputs. This structure is highly consistent with the architecture of deep learning neural networks, so deep learning has shown good application prospects in medium- and long-term runoff prediction. However, despite the success of these models, there are still some problems that need to be solved. On the one hand, the uncertainty caused by the selection of prediction factors, model structure and parameter settings often leads to insufficient forecast accuracy, and the current models usually lack systematic uncertainty assessment, which affects the reliability of their forecast results in practical applications. On the other hand, the deep learning model itself has a "black box" characteristic and lacks a physical explanation of the hydrological process, which not only affects the transparency of the model, but also limits its credibility in practical applications.

[0004] In order to improve the accuracy and reliability of the model, the research mainly focuses on two aspects: first, reducing the uncertainty of factor selection by screening suitable prediction factors, and second, optimizing the model structure and parameters to reduce the inherent uncertainty of the model. In terms of factor selection, there are large differences in the impact of hydrological, meteorological and climatic factors on basin runoff, and different factors have different degrees of influence on runoff at different time scales. Therefore, how to screen out the most predictive factors from massive meteorological, climatic and historical hydrological data is the key to improving the accuracy of runoff prediction. Traditional single factor screening methods usually focus on linear relationships or a specific data feature, but often ignore the interaction between multiple factors and the nonlinear characteristics of data, which affects the stability and reliability of the prediction results. At the same time, the selection of model structure and parameters is also an important factor affecting the medium- and long-term runoff prediction results. Deep learning models, especially multi-layer perceptron (MLP), long short-term memory network (LSTM) and Transformer models, although they show good prediction capabilities in some scenarios, when faced with complex, multi-dimensional and multi-time series data, single-structure deep learning models often find it difficult to maintain consistent prediction performance, resulting in greater uncertainty in prediction results. Therefore, by integrating multiple factor screening methods and multi-model fusion strategies, the deviation in input factor selection and the limitations of model structure can be effectively reduced, thereby improving the accuracy and reliability of runoff prediction.

[0005] In terms of model interpretability, the "black box" nature of deep learning models usually makes their prediction results difficult to explain and understand. In order to improve the transparency and credibility of the model, post-explanation methods, as an explanation tool independent of the model, have gradually been applied to fields such as hydrology and meteorology. These methods help researchers understand the decision-making process and key driving factors of the model by quantifying the contribution of input factors to the model prediction results, thereby enhancing the interpretability and transparency of the model. However, most of the current research focuses on the evaluation of static factor impacts, and has not fully explored the impact of multi-factor, long-time lag, and multi-category driving factors on runoff prediction at different time scales. Therefore, how to improve the interpretability of the model by introducing an adaptive mechanism and combining time series characteristics with nonlinear relationships between factors is still an important issue that needs to be solved.

[0006] In summary, although existing studies have made some progress in medium- and long-term runoff forecasting, existing methods still have problems such as insufficient uncertainty modeling, factor selection bias, model structure limitations, and poor interpretability. Therefore, it is of great theoretical and practical value to develop an integrated forecasting framework that can comprehensively consider multi-source uncertainty, optimize factor screening, improve model interpretability, and overcome existing limitations. Summary of the invention

[0007] The purpose of the invention is to provide a medium- and long-term runoff integrated probability prediction method based on deep learning to solve the above-mentioned problems existing in the prior art.

[0008] Technical solution, a medium- and long-term runoff integrated probability prediction method based on deep learning, including the following steps:

[0009] Runoff data were collected, and the Dempster-Shafer evidence theory based on conflict measurement and dynamic weight was used to integrate the results of multiple factor screening methods to obtain the set of driving factors for the runoff prediction model.

[0010] A deep learning base model based on a multi-scale physical constraint adaptive loss function is constructed, and the model parameters are optimized through the Bayesian optimization method to obtain a trained deep learning runoff prediction base model;

[0011] Construct a multimodal adaptive Stacking integration framework, use the deep learning runoff prediction base model as the first-layer base learner, and the decision tree regression model as the second-layer meta-learner to generate a single-value runoff prediction result of the integrated model;

[0012] An improved multi-level Bootstrap-Quantile fusion method is used to quantify the uncertainty of the single-value prediction results, and a medium- and long-term runoff probability prediction scheme is generated through a dynamic weighting mechanism;

[0013] The SHAP post-hoc explanation method was used to perform interpretability analysis and identify key drivers.

[0014] According to one aspect of the present application, the steps of integrating multiple factor screening methods based on conflict measurement and dynamic weight improved Dempster-Shafer evidence theory include:

[0015] Constructing identification frameworks and evidence sets, corresponding to the sets of predictors screened by multiple screening methods;

[0016] Define the basic trust function of each screening method, and normalize the evaluation index corresponding to the prediction factor screened by each method into a basic probability assignment;

[0017] Evidence fusion and dynamic weight adjustment are performed, and the joint basic probability assignment is calculated using the Dempster-Shafer combination rule. The weight is dynamically adjusted according to the Nash efficiency coefficient of each screening method on the validation set, and the evidence fusion result is optimized to obtain a comprehensive trust function.

[0018] The comprehensive trust function values ​​are arranged in descending order and the cumulative contribution rate is calculated. The predictive factors whose cumulative contribution rate reaches the preset threshold are selected to generate the final set of driving factors.

[0019] According to one aspect of the present application, the multi-scale physical constraint adaptive loss function comprises:

[0020] Multi-scale decomposition term L Decompose =∑ω k •MSE(IMF true , IMF pred );ω k is the weight of the kth mode; IMF true 、IMF pred are the modal components of the real runoff series and the predicted series;

[0021] Physical constraint L Physics =||P-ET-Q pred -ΔS|| 2 ; P is precipitation, ET is evapotranspiration, ΔS is the change in water storage;

[0022] Trend matching item L Trend =1-Cov(▽Q true ,▽Q pred ) / (σ▽Q true •σ▽Q pred );

[0023] Among them, ▽Q true and ▽Q pred Represent the real runoff sequence Q true and the predicted sequence Q pred The first-order difference of , i.e., trend change; Cov(▽Q true ,▽Q pred ) is the covariance between the two, which is used to measure the consistency of their trend changes; σ▽Q true and σ▽Q pred Represent the real sequence trend ▽Q true and predicting sequence trends ▽Q pred The standard deviation of

[0024] Dynamic weight parameters adaptively adjust the weight coefficients of the above three items according to the performance of the validation set.

[0025] According to one aspect of the present application, the steps of the Bayesian optimization method include:

[0026] Define the optimization objective function and dynamic weight adjustment mechanism, and set the mean of the multi-scale physical constraint adaptive loss function that minimizes the training samples as the optimization objective;

[0027] Set the search space for model hyperparameters and loss function weight parameters;

[0028] The Bayesian optimization algorithm is used for iterative optimization, and the global optimal solution is gradually approached through agent model construction, acquisition function selection and iterative optimization;

[0029] Based on the multi-indicator evaluation results on the validation set, the optimal hyperparameter configuration is selected to generate the base model prediction results of the deep learning model.

[0030] According to one aspect of the present application, the steps of the multimodal adaptive Stacking integration framework include:

[0031] The base learners are divided into three groups of expert models according to the three stages of flood season, normal water season and dry season, and the weights of the base learners are dynamically adjusted for special hydrological events.

[0032] The prediction results of the base learners are aggregated, and the physical constraints of water balance residuals and multimodal data are introduced to generate new training and test data sets.

[0033] Using new training data as input, the decision tree regression model is trained as a three-stage meta-learner to learn the combined relationship between the prediction results of the base learners; optimization is performed for the dry season, normal water season, and flood season respectively, and the prediction strategy is adjusted according to the historical hydrological characteristics of the same period;

[0034] Use the trained meta-learner combined with new test data to generate the final single-value prediction result.

[0035] According to one aspect of the present application, the steps of the improved multi-level Bootstrap-Quantile fusion method include:

[0036] The cross-validation method is used to divide the training set and the test set to obtain the prediction error sample set; stratified sampling is performed according to the runoff quantile, and oversampling is performed on historical flood peaks and extreme low water events;

[0037] Perform multiple Bootstrap resampling on the prediction results of each time period to generate multiple Bootstrap samples;

[0038] Combined with the Quantile method to calculate the probability distribution of errors, the weights are dynamically allocated based on the root mean square error of the prediction error; the hydrological extreme value sensitive distribution is constructed and the kurtosis adjustment parameter and skewness parameter are introduced;

[0039] The joint quantile function is calculated for the weighted prediction results to obtain the runoff prediction interval at a specific confidence level.

[0040] According to one aspect of the present application, the steps of the SHAP post-interpretation method include:

[0041] Calculate the SHAP value of each input feature to quantify its specific contribution in different time periods and different structural prediction models;

[0042] The SHAP visualization tool was used to draw Shapley value bee swarm diagrams and contribution ranking diagrams, analyze the key driving factors from the time dimension, and identify the input features that have an important impact on runoff prediction in different months.

[0043] According to one aspect of the present application, the evidence fusion process based on the conflict metric and the improved Dempster-Shafer evidence theory of dynamic weights includes:

[0044] The conflict metric K is used to calculate the conflict degree between the evidences of different screening methods. The K value is determined by the sum of the combinations of the intersection of the products of the basic probability assignments of each screening method being the empty set;

[0045] According to the conflict metric K, the calculation formula of the joint trust function is modified and the value of m(H j ) = ∑ B ∩ C = H j [w j •m 1 (B)•m 2 (C)•m 3 (C)] / (1-K) to correct for high conflict evidence; m 1 (B), m 2 (C), m 3 (C) is A 1 , A 2 , A 3 The basic trust function of , K is the conflict measure;

[0046] When the K value exceeds the preset threshold, the adaptive weight adjustment mechanism is triggered to increase the weight of the screening method to ensure the reliability of the evidence fusion results.

[0047] According to one aspect of the present application, the variational mode decomposition process of the multi-scale decomposition term includes:

[0048] The true runoff series Qtrue and the predicted series Qpred are decomposed into K intrinsic mode functions IMF by using variational mode decomposition algorithm;

[0049] A differentiated weight allocation strategy is implemented for the decomposed modal components, where low-frequency modal components are assigned a higher weight ωlow and high-frequency modal components are assigned a lower weight ωhigh, where ωlow+ωhigh=1;

[0050] For each modal component, the mean square error between the true value and the predicted value is calculated to obtain the weighted error LDecompose=∑ω k •MSE(IMF true , IMF pred); Ldecompose is the weighted sum of the errors of each modal component after variational mode decomposition; k is the total number of modal components decomposed by variational mode decomposition (VMD); ω k is the weight coefficient of the kth modal component; IMF true is the real runoff sequence component of the kth mode; IMF pred is the predicted runoff sequence component of the kth mode.

[0051] By dynamically adjusting the weight ratio of low-frequency and high-frequency modes, the model's ability to identify interannual changes and seasonal fluctuations in runoff is enhanced.

[0052] According to one aspect of the present application, the screening process of the multiple factor screening method includes:

[0053] Collect and preprocess the basin runoff data, remove factors with high missing rate, fill in missing values ​​and standardize them to obtain the preprocessed data set;

[0054] The improved time-lagged Pearson correlation coefficient method was used to calculate the correlation coefficient within the 1-12 month lag window, and the most significant factors were selected to form the solution set A1;

[0055] The multi-scale maximum information coefficient method is used to calculate the mutual information value at the monthly, quarterly and annual scales, and the most significant factors are selected to form the solution set A2;

[0056] The physical constraint random forest feature importance scoring method was used to calculate the variance increment score, and the most significant factors were selected to form the solution set A3.

[0057] According to one aspect of the present application, the deep learning base model is a hydrological time series adaptive hybrid neural network architecture, and its construction process includes:

[0058] The variational mode decomposition method is used to decompose the original runoff series into long-term trend component, seasonal component and random fluctuation component.

[0059] Construct a specialized neural network, use the multi-layer perceptron MLP to process the long-term trend component, use the long short-term memory network LSTM to process the seasonal component, and use the Transformer to process the random fluctuation component;

[0060] Design a dynamic attention fusion layer to calculate the attention weight coefficients α, β, and γ according to the hydrological characteristics of the target period, satisfying α+β+γ=1, where α, β, and γ are the output weights of the MLP, LSTM, and Transformer models, respectively;

[0061] The historical similar hydrological year characteristics are introduced as auxiliary input, and the k most similar hydrological year runoff data are selected by calculating the Euclidean distance to construct an enhanced feature matrix;

[0062] The weighted output results of the three sub-models are integrated to generate the final runoff prediction value Q_pred = α·Q_MLP +β·Q_LSTM + γ·Q_Transformer, where Q_MLP, Q_LSTM, and Q_Transformer are the prediction results of the three models respectively.

[0063] According to one aspect of the present application, the Bayesian optimization method is a Bayesian optimization method guided by hydrological parameter sensitivity, and the optimization process further includes:

[0064] The hydrological parameter sensitivity matrix S was constructed, and the sensitivity coefficient of each hyperparameter to the model performance was calculated by Morris screening method;

[0065] The hyperparameter prior distribution P is established based on the watershed characteristics including area, slope, and vegetation coverage;

[0066] Construct a multi-objective hydrological evaluation function H(θ) = w 1 NSE + w 2 (1-|RC_bias|) + w 3 (1-|WB_res|), where NSE is the Nash efficiency coefficient, RC_bias is the runoff coefficient bias, WB_res is the water balance residual, and w 1 、w 2 、w 3 is the weight coefficient and satisfies w 1 + w 2 + w 3 =1;

[0067] A phased optimization strategy is adopted to optimize the hyperparameter sets θ_wet and θ_dry for the wet season and the dry season, respectively;

[0068] Construct a Gaussian process surrogate model GP(x) and use the expected improvement acquisition function EI(x) to select the next set of hyperparameters;

[0069] According to the hydrological characteristics of different forecast months, the corresponding optimal hyperparameter set is switched to generate the final model prediction results.

[0070] According to one aspect of the present application, the specific output of the present invention includes:

[0071] The driving factor set of the runoff prediction model, the deep learning runoff prediction base model, the single value prediction results of the Stacking integrated model, the medium- and long-term runoff probability prediction scheme, and the Shapley value bee swarm diagram and contribution ranking diagram.

[0072] Beneficial effects: Compared with the prior art, the present invention proposes a medium- and long-term runoff probability prediction framework based on the integration idea of ​​"input-parameter-structure" full-factor optimization. First, the screening results of various factor screening methods are integrated through the improved Dempster-Shafer evidence theory, and a set of driving factors is constructed, which significantly improves the explanatory power and prediction accuracy of the input factors. The multi-scale physical constraint adaptive loss function MPALF is adopted, combined with variational mode decomposition (VMD), water balance equation residual and trend similarity measurement, to enhance the model's prediction ability for extreme values ​​and improve the model performance. The MLP, LSTM and Transformer multi-model structure is combined with the multi-modal adaptive Stacking integration framework to significantly improve the overall prediction accuracy and generalization ability. In terms of uncertainty quantification, the probability prediction interval is generated by the multi-level Bootstrap-Quantile fusion method, which effectively quantifies the uncertainty of the prediction results and provides support for risk decision-making. The characteristic Shapley value is calculated based on the SHAP method to identify key driving factors, enhance the transparency and credibility of the model, and provide a scientific basis for water resources management. The present invention significantly improves the accuracy, reliability and practicability of medium- and long-term runoff prediction through full-factor optimization and uncertainty quantification, and has important theoretical value and engineering application significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0074] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. However, it is obvious to those skilled in the art that the present invention can be implemented without one or more of these details. In other examples, in order to avoid confusion with the present invention, some technical features known in the art are not described. The order of the relevant steps in the present invention is not restrictive, that is, it can be adjusted by those skilled in the art. The order in the present invention is a case-based writing method, not a restrictive description.

[0075] The present invention comprehensively considers multi-source uncertainties and proposes a runoff probability prediction framework with full-factor optimization of "input-parameter-structure", which solves the problems of insufficient uncertainty modeling in medium- and long-term runoff prediction, factor selection bias, model structure limitations and poor interpretability, and significantly improves the accuracy, reliability and practicality of medium- and long-term runoff prediction, and has important theoretical value and engineering application significance.

[0076] like Figure 1 As shown, the following technical solution is proposed.

[0077] According to one aspect of the present application, a method for interpretable medium- and long-term integrated probability prediction based on deep learning is provided, characterized in that it includes the following steps:

[0078] Step S1, collecting and collating the long-term and medium-term runoff data of the basin where the station to be predicted is located, and screening different categories of runoff prediction factors by using the improved time-lagged Pearson correlation coefficient method, the multi-scale maximum information coefficient method, and the physical constraint-based random forest feature importance scoring method in turn, and using the Dempster-Shafer evidence theory based on conflict measurement and dynamic weight to integrate the prediction factor set generated by the single factor screening method, to obtain the driving factor set of the runoff prediction model;

[0079] Step S2, constructing a deep learning base model based on an improved multi-scale physical constraint adaptive loss function, wherein the loss function MPALF realizes multi-dimensional constraints through variational mode decomposition VMD, water balance equation residual and trend similarity measurement, and adopts Bayesian optimization method to optimize model parameters to obtain a trained deep learning runoff prediction base model;

[0080] Step S3, introducing multimodal data fusion and physical constraint enhancement mechanism, using the constructed deep learning runoff prediction base model as the first-layer base learner of the multimodal adaptive Stacking integration framework; using the pre-constructed decision tree regression model as the second-layer meta-learner of the framework, using the runoff prediction results output by the base learner, multimodal data and physical constraint items as input, constructing the final regression relationship, and generating the final Stacking integration model runoff single value prediction result;

[0081] Step S4, using an improved multi-level Bootstrap-Quantile fusion method to quantify the uncertainty of the single-value prediction results of the Stacking model, adjusting the model weights through a dynamic weighting mechanism, and generating a medium- and long-term runoff probability prediction scheme;

[0082] Step S5: Use the SHAP post-explanation method to perform interpretability analysis on each multi-factor driven deep learning runoff prediction model. By calculating and visualizing the Shapley value of each input feature, drawing a bee swarm diagram and a contribution ranking diagram, calculating the contribution of input features in different months to the prediction results, and identifying the key driving factors.

[0083] Medium- and long-term runoff prediction has important practical application value in water resources management, watershed planning and flood control. In order to improve the prediction accuracy, it usually depends on multiple factors, such as meteorological data, historical runoff data and soil moisture. However, due to the different characteristics of various factors, how to effectively screen and integrate different types of prediction factors has always been the core problem of improving the performance of prediction models. Existing factor screening methods mostly rely on a single algorithm, fail to fully consider the relationship between different types of factors and their changing characteristics at different time scales, and easily lead to instability of factor selection and insufficient model prediction performance. The present invention proposes a driving factor set construction method based on the integration of multiple factor screening methods, which can comprehensively screen factors from multiple dimensions and different time scales, reduce the uncertainty and deviation that may be caused by a single factor screening method, and improve the stability and reliability of factor selection. The improved Dempster-Shafer evidence theory is used for result fusion, which can not only effectively deal with the conflict between factors, but also dynamically adjust the weights of each screening method according to different data characteristics, so as to obtain a more accurate driving factor set. The implementation of this innovative method helps to improve the accuracy and stability of medium- and long-term runoff prediction models and provide more scientific decision support for water resources management.

[0084] According to one aspect of the present application, step S1 further comprises:

[0085] Step S11, collecting long-term and medium-term runoff data of the basin where the station to be predicted is located, including historical monthly runoff data, meteorological station observation data, "air-sea-land" teleconnection factor data, and remote sensing and GIS data, wherein the historical monthly runoff data is the runoff data of the 12 months before the prediction period, the meteorological station observation data covers 8 basic meteorological elements, and the teleconnection factor data includes 130 factors related to the atmosphere, the ocean and the land, which constitute the full set of initial prediction factors; preprocessing the initial prediction factor set data, if the data missing rate of a factor exceeds a preset threshold, it is removed, and the retained factor data is processed with an iterative filling method based on the MissForest algorithm to process the missing values, and the filled data is Z-score standardized, and finally a preprocessed data set containing m prediction factors is obtained, where m≥30 and is an integer;

[0086] In this embodiment, long-term mid- and long-term runoff data for the past 60 years in the basin where Hongze Lake is located were selected, and various data sources were collected, including historical monthly runoff data, meteorological station observation data, "atmosphere-ocean-land" teleconnection factor data, and remote sensing and GIS data. The data missing rates of each factor were checked, and it was found that the missing rates of 20 factors exceeded 10%, and they were excluded. The remaining 110 factors were filled iteratively using the MissForest algorithm to fill the missing values; the filled data were standardized by Z-score to obtain a preprocessed data set.

[0087] Step S12: Use the improved time-lagged Pearson correlation coefficient method to screen each prediction factor, calculate the Pearson correlation coefficient between runoff and each prediction factor within a 1-12 month lag window, and select the top p prediction factors with the largest absolute value of the correlation coefficient, where p < m, to form the prediction factor scheme set A 1 ;

[0088] PCC = ∑(X ij (t - k) - X j (t - k)) • (Y i (t) - Y(t) ) / (∑(X ij (t - k) - X j (t - k)) 2 • ∑ (Y i (t) - Y(t)) 2 ) 0.5

[0089] In the formula: PCC is the Pearson correlation coefficient; n is the length of the runoff sequence in the training period; m is the number of factors; Y i (t) is the runoff value in the t-th month of the i-th year; Y(t) is the average value of the runoff in the t-th month; X ij (t - k) is the j-th factor in the (t - k)-th month of the i-th year; X j (t - k) is the average value of the j-th factor in the (t - k)-th month;

[0090] Step S13: Use the multi-scale maximum information coefficient method to screen each prediction factor. On the monthly, quarterly, and annual time scales, calculate the mutual information value between the prediction factor and the target variable. The monthly scale is calculated according to the original sequence, and the quarterly and annual scales respectively take 3-month and 12-month moving average sequences. Select the top p prediction factors with the largest mutual information value MIC, where p < m, to form the prediction factor scheme set A 2 ;

[0091] MIC(X(t-k), Y(t)) = max(max(I) / ㏒ 2 (min(a,b)))

[0092] Where: I is the mutual information value of X(t-k) and Y(t), ab < B; a represents dividing the value range of X(t-k) into a segments; b represents dividing the value range of Y(t) into b segments; B is the upper limit value of variable grid division, generally taking the 0.6th power of the data volume.

[0093] Step S14: Use the random forest feature importance scoring method based on physical constraints to screen each predictor, and add the constructed physical constraint term L to the objective function RF , calculate the variance increment score VIM of each predictor, select the top p predictors with the largest variance increment, p < m, to form the predictor set A 3 ;

[0094] L RF = MSE + λ||▽Q - P + E|| 2

[0095] VIM(X(t-k), Y(t)) = [1 / n∑(MSE before - MSE after )] / S E

[0096] Where: λ = 0.1 is the constraint weight, MSE before is the mean square error before removing the predictor X(t-k), MSE after is the mean square error after removing the predictor X(t-k), S E is the standard error of T regression trees;

[0097] Step S15: Use the Dempster-Shafer evidence theory improved based on conflict measure and dynamic weight to integrate the factor sets A 1 , A 2 , A 3 by means of evidence fusion to construct the driving factor set of the final runoff prediction model.

[0098] According to one aspect of the present application, the step S15 is further:

[0099] Step S15a: Define the recognition framework Φ = {H 1 , H 2 ,…, H P}, where H i represents the event that the i-th factor is selected, and there is no intersection between the event elements; construct the evidence set D = {A 1 , A 2 , A 3}, corresponding to the predictor sets screened by three screening methods respectively;

[0100] Step S15b: Based on the Dempster-Shafer evidence theory, define the basic trust function m of each screening method i (H j ), represents the confidence of the i-th screening method on the j-th prediction factor, that is, the absolute value of the Pearson correlation coefficient, the maximum information coefficient value and the variance increment value corresponding to the first p prediction factors screened out by each of the three methods are normalized to the basic probability assignment BPA, the formula is as follows;

[0101] m i (H j )=s ij / ∑s ij

[0102] Step S15c: perform evidence fusion and dynamic weight adjustment, use the Dempster-Shafer combination rule to calculate the joint BPA, introduce a dynamic weight adjustment mechanism, and dynamically adjust the weight w according to the Nash efficiency coefficient NSE of the base learner of each screening method on the validation set. j , optimize the evidence fusion results and obtain the new comprehensive trust function m(H j );

[0103] w j =NSE j / ∑NSE j

[0104] NSE=1-∑(Q-Q*) 2 / ∑(Q-Q) 2

[0105] m(H j )=∑ B∩C=Hj [w j •m 1 (B)•m 2 (C)•m 3 (C)] / (1-K)

[0106] K=∑ B∩C=Φ [m 1 (B)•m 2 (C)•m 3 (C)]

[0107] In the formula, m 1 (B), m 2 (C), m 3 (C) is A 1 , A 2 , the basic trust function of A3, K is the conflict measure, which is used to correct the impact of high-conflict evidence.

[0108] Step S15d: The new comprehensive trust function value m(H j ) are arranged from high to low, the concept of cumulative contribution rate is introduced, the cumulative contribution rate of the comprehensive trust function is calculated, the prediction factors whose cumulative contribution rate reaches the preset threshold are selected, and the variance inflation factor VIF is used to finally generate the driving factor set X.

[0109] In this embodiment, the top 30 prediction factors with the largest absolute values ​​of correlation coefficients are selected to form a prediction factor solution set A. 1 , select the top 30 prediction factors with the largest mutual information value to form the prediction factor solution set A 2 , select the top 30 predictors with the largest variance increment to form the predictor solution set A 3 ; Define the recognition framework Φ = {H 1 ,H 2 ,…,H 30}, the absolute values ​​of the Pearson correlation coefficient, the maximum information coefficient and the variance increment corresponding to the top 30 predictors screened out by each of the three methods were normalized to the basic probability assignment BPA, and the joint BPA was calculated using the Dempster-Shafer combination rule. The dynamic weight adjustment mechanism was introduced to dynamically adjust the weight according to the Nash efficiency coefficient NSE of each screening method on the validation set to optimize the evidence fusion result. The new comprehensive trust function value m(H j ) are arranged from high to low, the cumulative contribution rate is calculated, and the prediction factors with a cumulative contribution rate of 90% are selected. The 15 most predictive driving factors are screened and integrated. These factors cover information from multiple aspects such as meteorology, teleconnection, remote sensing and GIS. The final set of driving factors X will be used to construct the runoff prediction model, which significantly improves the prediction accuracy and stability of the model. The specific effects are as follows: 1) Through multi-method screening and integration, the Nash efficiency coefficient NSE of the model is increased from 0.75 to 0.85; 2) Through dynamic weight adjustment and conflict measurement, the prediction stability of the model at different time scales is significantly improved; 3) Enhanced interpretability: Through physical constraints and feature importance scores, the interpretability of the model is enhanced, which helps to understand the impact mechanism of each factor on runoff.

[0110] According to one aspect of the present application, step S2 further comprises:

[0111] Step S21: introduce a multi-scale physical constraint adaptive loss function MPALF. The loss function MPALF includes a multi-scale decomposition term L Decompose , physical constraint L Physics 、Trend matching item L Trend , whose dynamic weight parameter λ 1 ,λ 2 ,λ 3After the initial value is set, the MPALF formula is adaptively adjusted according to the performance of the validation set as follows:

[0112] L MPALF =λ 1 •L Decompose +λ 2 •L Physics +λ 3 •L Physics

[0113] 1) Multi-scale decomposition term L Decompose :The real runoff sequence Q is transformed into true and the predicted sequence Q pred Decompose into K modal components IMF, calculate the weighted error of each mode, where ω k is the weight of the kth mode;

[0114] L Decompose =∑ω k •MSE(IMF true , IMF pred )

[0115] In this embodiment, the true runoff sequence Qtrue and the predicted sequence Qpred are decomposed into K=5 modal components (IMF1~IMF5), the mean square error (MSE) of each modal component is calculated, and weights are assigned (low-frequency mode ω=0.6, high-frequency mode ω=0.4).

[0116] 2) Physical constraint L Physics :By introducing the residual of the water balance equation as a constraint term, the predicted value is ensured to conform to the physical laws, where P is precipitation, ET is evapotranspiration, and ΔS is the change in water storage;

[0117] L Physics =||P-ET-Q pred -ΔS|| 2

[0118] 3) Trend matching item L Trend : Calculate the trend similarity between the predicted sequence and the true sequence;

[0119] L Trend =1-Cov(▽Q true ,▽Q pred ) / (σ▽Q true •σ▽Q pred )

[0120] Among them, ▽Q true and ▽Q pred Represent the real runoff sequence Q true and the predicted sequence Q predThe first-order difference of , i.e., trend change; Cov(▽Q true ,▽Q pred ) is the covariance between the two, which is used to measure the consistency of their trend changes; σ▽Q true and σ▽Q pred Represent the real sequence trend ▽Q true and predicting sequence trends ▽Q pred The standard deviation of , used to standardize the covariance values;

[0121] Step S22, respectively constructing hydrological time series adaptive hybrid MLP, LSTM, and Transformer deep learning base models that introduce an improved loss function MPALF;

[0122] 1) The original runoff series is decomposed into long-term trend component, seasonal component and random fluctuation component by using variational mode decomposition method;

[0123] 2) Construct a specialized neural network, use a multi-layer perceptron (MLP) to process the long-term trend component, use a long short-term memory (LSTM) to process the seasonal component, and use a Transformer to process the random fluctuation component;

[0124] 3) Design a dynamic attention fusion layer to calculate the attention weight coefficients α, β, and γ according to the hydrological characteristics of the target period, satisfying α+β+γ=1, where α, β, and γ are the output weights of the MLP, LSTM, and Transformer models respectively;

[0125] 4) Introducing the characteristics of similar historical hydrological years as auxiliary input, selecting the k most similar hydrological year runoff data by calculating the Euclidean distance, and constructing an enhanced feature matrix;

[0126] 5) Integrate the weighted output results of the three sub-models to generate the final runoff prediction value Q_pred = α·Q_MLP +β·Q_LSTM + γ·Q_Transformer, where Q_MLP, Q_LSTM, and Q_Transformer are the prediction results of the three models respectively.

[0127] In this embodiment, the runoff sequence is decomposed into three components: long-term trend (60% variance), seasonality (30%), and random fluctuation (10%) by variational mode decomposition (VMD). MLP (30 nodes in the input layer → 64 nodes in the hidden layer, ReLU), LSTM (50 units in both directions, Dropout=0.2), and Transformer (4-head self-attention + 2-layer feedforward network) are used for modeling, and the three similar annual runoff sequences with the smallest Euclidean distance in the past 10 years are introduced and concatenated with the current input after convolution feature extraction. The dynamic attention weights are adaptively adjusted according to the forecast period: the weights in the flood season (June-August) are α=0.2 (trend), β=0.6 (season), and γ=0.2 (random); the weights in the dry season (January-March) are α=0.5, β=0.3, and γ=0.2. The NSE of the fused model test set reached 0.92, an increase of 8.2% compared with the single LSTM model (NSE=0.85). The RMSE of the trend component dropped from 5.3 m³ / s to 3.2 m³ / s (a decrease of 40%), and the prediction interval coverage of extreme flood events (defined as runoff exceeding the 90th percentile) reached 93% (an increase of 35% compared with the unfused model), verifying the effectiveness of multi-component collaborative modeling and physical driving mechanism.

[0128] Step S23, input the driving factor set X of each monthly runoff prediction, and input the monthly runoff data as the target variable into the MLP, LSTM and Transformer deep learning base models respectively to obtain the monthly runoff prediction;

[0129] Step S24, using the Bayesian optimization method guided by hydrological parameter sensitivity to optimize the hyperparameters of the deep learning runoff prediction model MLP, LSTM and Transformer, and finally obtaining the prediction result of the deep learning base model with the best performance;

[0130] 1) Construct the hydrological parameter sensitivity matrix S and calculate the sensitivity coefficient of each hyperparameter to the model performance through the Morris screening method;

[0131] 2) Establish hyperparameter prior distribution P based on watershed characteristics including area, slope, and vegetation coverage;

[0132] 3) Construct a multi-objective hydrological evaluation function H(θ) = w 1 NSE + w 2 (1-|RC_bias|) + w 3 (1-|WB_res|), where NSE is the Nash efficiency coefficient, RC_bias is the runoff coefficient bias, WB_res is the water balance residual, and w 1 、w 2 、w 3 is the weight coefficient and satisfies w1 + w 2 + w 3 =1;

[0133] 4) A phased optimization strategy is adopted to optimize the hyperparameter sets θ_wet and θ_dry for the wet season and the dry season respectively;

[0134] 5) Construct a Gaussian process proxy model GP(x) and use the expected improvement acquisition function EI(x) to select the next set of hyperparameters;

[0135] 6) According to the hydrological characteristics of different forecast months, switch the corresponding optimal hyperparameter set to generate the final model prediction results.

[0136] In this embodiment, for the runoff prediction demand of the basin (area 1500 km², slope 12°, vegetation coverage 65%), the Bayesian optimization method guided by hydrological parameter sensitivity is used to optimize the hyperparameters of MLP, LSTM and Transformer models. First, the learning rate (sensitivity coefficient 0.45), the number of LSTM units (0.38) and the number of Transformer heads (0.25) are determined as key hyperparameters by Morris screening method, and based on the multi-objective evaluation function H(θ)=0.6·NSE+0.3·(1-|runoff coefficient deviation|)+0.1·(1-|water balance residual|), the optimization is carried out in stages: the optimal hyperparameter set θ_wet (learning rate 5e-4, LSTM unit 64) in the flood season and θ_dry (learning rate 1e-3, LSTM unit 32) in the dry season. After 50 rounds of iterative optimization with the Gaussian process surrogate model, the NSE of the model on the test set was improved from 0.78 to 0.88, the water balance residual was reduced by 60% (from 1.5 mm to 0.6 mm), and the RMSE in the flood season was reduced from 12.3 m³ / s to 8.5 m³ / s (a decrease of 30.9%), and the MAE in the dry season was reduced from 6.2 m³ / s to 4.1 m³ / s (a decrease of 33.9%), verifying the effectiveness of staged hyperparameter optimization and physical constraint driving.

[0137] According to one aspect of the present application, step S3 further comprises:

[0138] According to one aspect of the present application, the steps of the multimodal adaptive Stacking integration framework include:

[0139] 1) The base learners are divided into three groups of expert models according to the three stages of flood season, normal water season and dry season, and the weights of the base learners are dynamically adjusted according to special hydrological events;

[0140] 2) Summarize the prediction results of the base learners, introduce the physical constraints of water balance residuals and multimodal data, and generate new training and test data sets;

[0141] 3) Using new training data as input, the decision tree regression model is trained as a three-stage meta-learner to learn the combined relationship between the prediction results of the base learners; optimization is performed for the dry season, normal water season, and flood season respectively, and the prediction strategy is adjusted according to the historical hydrological characteristics of the same period;

[0142] 4) Use the trained meta-learner combined with new test data to generate the final single-value prediction result.

[0143] In this embodiment, based on the multimodal adaptive Stacking integrated framework, MLP, LSTM and Transformer are grouped into expert models according to the flood season (June-August), normal water season (April-May, September-October) and dry season (January-March), and the water balance residual (W0 = P−ET−Q−ΔS) and multimodal data (temperature, precipitation, monthly mean sea temperature in the Nino3.4 region) are introduced to construct a new data set. The decision tree regression meta-learner is trained for each hydrological period, and the weight of the base learner is dynamically adjusted (LSTM weight β=0.6 in the flood season, because of its advantage in temporal modeling of monsoon precipitation; MLP weight α=0.5 in the dry season, focusing on long-term trend prediction). The NSE of the optimized model on the test set reached 0.91, an increase of 7.1% compared with the single LSTM model (NSE=0.85). The water balance residual was reduced from 1.2 mm to 0.4 mm (a decrease of 66.7%), and the prediction interval coverage of extreme flood events (defined as runoff exceeding the historical 90% quantile) reached 93% (an increase of 30% compared with the non-staged model). At the same time, the RMSE in the flood season was reduced from 12.5 m³ / s to 8.8 m³ / s (a decrease of 29.6%), verifying the effectiveness of staged integration and physical constraint-driven, and providing a high-precision and high-reliability runoff prediction solution for reservoir operation.

[0144] According to one aspect of the present application, the steps of the improved multi-level Bootstrap-Quantile fusion method include:

[0145] The cross-validation method is used to divide the training set and the test set to obtain the prediction error sample set; stratified sampling is performed according to the runoff quantile, and oversampling is performed on historical flood peaks and extreme low water events;

[0146] Perform multiple Bootstrap resampling on the prediction results of each time period to generate multiple Bootstrap samples;

[0147] Combined with the Quantile method to calculate the probability distribution of errors, the weights are dynamically allocated based on the root mean square error of the prediction error; the hydrological extreme value sensitive distribution is constructed and the kurtosis adjustment parameter and skewness parameter are introduced;

[0148] The joint quantile function is calculated for the weighted prediction results to obtain the runoff prediction interval at a specific confidence level.

[0149] In this embodiment, an improved multi-level Bootstrap-Quantile fusion method is used. First, the 60-year runoff data of the basin is divided into a training set and a test set through 5-fold cross validation, and stratified sampling is performed according to the runoff quantile, and historical flood peaks (such as the 1998 flood) and extreme low water events (such as the 2001 drought) are oversampled. The prediction results of each time period are resampled 1000 times by Bootstrap to generate a subset of 80% of the sample size, and the error distribution is calculated in combination with the Quantile method, and the weight (η=0.1) is dynamically assigned based on RMSE. The kurtosis adjustment parameter (k=2.1) and the skewness parameter (s=-0.8) are introduced to construct the hydrological extreme value sensitive distribution, and finally a runoff prediction interval with a 90% confidence interval is generated. The optimized model's prediction interval coverage for extreme floods on the test set increased from 80% to 95%, and the dry season prediction error decreased by 35% (RMSE dropped from 6.5 m³ / s to 4.2 m³ / s), verifying that the multi-level Bootstrap-Quantile method can significantly improve its ability to capture extreme events.

[0150] According to one aspect of the present application, step S5 is further:

[0151] Step S51, using the SHAP post-explanation method to perform interpretability analysis on the deep learning monthly runoff prediction models MLP, LSTM, and Transformer driven by multi-category and long-lag factors respectively;

[0152] Step S52, calculating the SHAP value of each input feature, quantifying its specific contribution in the runoff prediction model in different time periods and different structures, and exploring the importance and significance of the features of different months to the prediction results;

[0153] Step S53, use the SHAP visualization tool to draw the Shapley value bee swarm diagram and contribution ranking diagram, analyze the key driving factors from the time dimension, and identify the input features that have an important impact on runoff prediction in different months; first use the SHAP value to draw a bee swarm diagram to show the contribution distribution of each input feature to the prediction result; then, based on the SHAP contribution ranking diagram, identify the input features that have an important impact on the prediction result in each time period, especially the changes in key factors between multiple months.

[0154] In this embodiment, the SHAP post-explanation method is used to perform interpretability analysis on the MLP, LSTM and Transformer models, and the SHAP value of each input feature is calculated to quantify its contribution to runoff prediction. The analysis found that the SHAP value of the precipitation factor in the flood season (June-August) is significantly higher than that in other months (the SHAP value in June is 0.35, and the SHAP value in December is 0.08), while the influence of the sea temperature factor in the dry season (January-March) is more prominent (the SHAP value in January is 0.22, and the SHAP value in July is 0.05). Through the SHAP bee colony diagram and contribution ranking diagram, precipitation, evapotranspiration and ENSO index are identified as key driving factors, and their importance changes dynamically with the month. This analysis not only reveals the physical mechanism of model prediction, but also provides a scientific basis for optimizing factor selection and improving prediction accuracy.

Claims

1. A medium- and long-term runoff integrated probability prediction method based on deep learning, characterized in that: The following steps are involved: Runoff data were collected, and the Dempster-Shafer evidence theory based on conflict measurement and dynamic weight was used to integrate the results of multiple factor screening methods to obtain the set of driving factors for the runoff prediction model. A deep learning base model based on a multi-scale physical constraint adaptive loss function is constructed, and the model parameters are selected through the Bayesian optimization method to obtain a trained deep learning runoff prediction base model. Construct a multimodal adaptive Stacking integration framework, use the deep learning runoff prediction base model as the first-layer base learner, and the decision tree regression model as the second-layer meta-learner to generate a single-value runoff prediction result of the integrated model; An improved multi-level Bootstrap-Quantile fusion method is used to quantify the uncertainty of the single-value prediction results, and a medium- and long-term runoff probability prediction scheme is generated through a dynamic weighting mechanism; The SHAP post-hoc explanation method was used to conduct interpretability analysis and identify key driving factors; The steps of integrating multiple factor screening methods based on conflict measurement and dynamic weight-improved Dempster-Shafer evidence theory include: Constructing identification frameworks and evidence sets, corresponding to the sets of predictors screened by multiple screening methods; Define the basic trust function of each screening method, and normalize the evaluation index corresponding to the prediction factor screened by each method into a basic probability assignment; Evidence fusion and dynamic weight adjustment are performed, and the joint basic probability assignment is calculated using the Dempster-Shafer combination rule. The weight is dynamically adjusted according to the Nash efficiency coefficient of each screening method on the validation set, and the evidence fusion result is optimized to obtain a comprehensive trust function. Arrange the comprehensive trust function values ​​in descending order and calculate the cumulative contribution rate, select the predictive factors whose cumulative contribution rate reaches the preset threshold, and generate the final set of driving factors; The steps of the improved multi-level Bootstrap-Quantile fusion method include: The cross-validation method is used to divide the training set and the test set to obtain the prediction error sample set; stratified sampling is performed according to the runoff quantile, and oversampling is performed on historical flood peaks and extreme low water events; Perform multiple Bootstrap resampling on the prediction results of each time period to generate multiple Bootstrap samples; Combined with the Quantile method to calculate the probability distribution of errors, the weights are dynamically allocated based on the root mean square error of the prediction error; the hydrological extreme value sensitive distribution is constructed and the kurtosis adjustment parameter and skewness parameter are introduced; The joint quantile function is calculated for the weighted prediction results to obtain the runoff prediction interval at the confidence level.

2. The medium- and long-term runoff integrated probability prediction method based on deep learning according to claim 1 is characterized in that: The multi-scale physical constraint adaptive loss function includes: Multi-scale decomposition term L Decompose =∑ ω k • MSE ( IMF true , IMF pred ); ω k For the k The weight of each mode; IMF true 、 IMF pred are the modal components of the real runoff series and the predicted series; Physical Constraints L Physics =|| P-ET-Q pred - Δ S || 2 ; P is precipitation, ET is evapotranspiration, and ΔS is the change in water storage; Trend Matches L Trend =1 -Cov (▽ Q true ,▽ Q pred ) / (σ▽ Q true •σ▽ Q pred ); Among them, Q true and Q pred Represent the real runoff series Q true and the predicted sequence Q pred The first-order difference of , i.e., trend change; Cov (▽ Q true ,▽ Q pred ) is the covariance between the two, which is used to measure the consistency of their trend changes; σ▽ Q true and σ▽ Q pred Represent the real sequence trend ▽ Q true and predicting sequence trends▽ Q pred The standard deviation of Dynamic weight parameters adaptively adjust the weight coefficients of the above three items according to the performance of the validation set.

3. The medium- and long-term runoff integrated probability prediction method based on deep learning according to claim 2 is characterized in that: The steps of the Bayesian optimization method include: Define the optimization objective function and dynamic weight adjustment mechanism, and set the mean of the multi-scale physical constraint adaptive loss function that minimizes the training samples as the optimization objective; Set the search space for model hyperparameters and loss function weight parameters; The Bayesian optimization algorithm is used for iterative optimization, and the global optimal solution is gradually approached through agent model construction, acquisition function selection and iterative optimization; Based on the multi-indicator evaluation results on the validation set, the optimal hyperparameter configuration is selected to generate the base model prediction results of the deep learning model.

4. The medium- and long-term runoff integrated probability prediction method based on deep learning according to claim 1 is characterized in that: The steps of the multi-modal adaptive Stacking integration framework include: The base learners are divided into three groups of expert models according to the three stages of flood season, normal water season and dry season, and the weights of the base learners are dynamically adjusted according to hydrological events. The prediction results of the base learners are aggregated, and the physical constraints of water balance residuals and multimodal data are introduced to generate new training and test data sets. Using new training data as input, the decision tree regression model is trained as a three-stage meta-learner to learn the combined relationship between the prediction results of the base learners; optimization is performed for the dry season, normal water season, and flood season respectively, and the prediction strategy is adjusted according to the historical hydrological characteristics of the same period; Use the trained meta-learner combined with new test data to generate the final single-value prediction result.

5. The medium- and long-term runoff integrated probability prediction method based on deep learning according to claim 1 is characterized in that: The steps of the SHAP post hoc interpretation method include: Calculate the SHAP value of each input feature to quantify its specific contribution in different time periods and different structural prediction models; The SHAP visualization tool was used to draw Shapley value bee swarm diagrams and contribution ranking diagrams, analyze the key driving factors from the time dimension, and identify the input features that have an important impact on runoff prediction in different months.

6. The medium- and long-term runoff integrated probability prediction method based on deep learning according to claim 2 is characterized in that: The evidence fusion process of the Dempster-Shafer evidence theory based on conflict measurement and dynamic weight improvement includes: The conflict metric K is used to calculate the conflict degree between the evidences of different screening methods. The K value is determined by the sum of the combinations of the intersection of the products of the basic probability assignments of each screening method being the empty set; According to the conflict metric K, the calculation formula of the joint trust function is modified and the m(H j ) = ∑ B ∩ C = H j [w j •m1(B)•m2(C)•m3(C)] / (1-K) to correct for high conflict evidence; m 1(B), m 2(C) m 3(C) is the basic trust function of A1, A2, and A3, K for conflict measurement; When the K value exceeds the preset threshold, the adaptive weight adjustment mechanism is triggered to increase the weight of the screening method to ensure the reliability of the evidence fusion results.

7. The medium- and long-term runoff integrated probability prediction method based on deep learning according to claim 2 is characterized in that: The variational mode decomposition process of the multi-scale decomposition term includes: The true runoff sequence Qtrue and the predicted sequence Qpred are decomposed into K intrinsic mode functions IMF by using variational mode decomposition algorithm; A differentiated weight allocation strategy is implemented for the decomposed modal components, where low-frequency modal components are assigned a higher weight ωlow and high-frequency modal components are assigned a lower weight ωhigh, where ωlow+ωhigh=1; For each modal component, the mean square error between the true value and the predicted value is calculated to obtain the weighted error LDecompose=∑ω k •MSE( IMF true , IMF pred ); Ldecompose is the weighted sum of the errors of each modal component after variational mode decomposition; k is the total number of modal components decomposed by variational mode decomposition (VMD); ω k is the weight coefficient of the kth modal component; IMF true is the real runoff sequence component of the kth mode; IMF pred is the predicted runoff sequence component of the kth mode; By dynamically adjusting the weight ratio of low-frequency and high-frequency modes, the model's ability to identify interannual changes and seasonal fluctuations in runoff is enhanced.

8. The medium- and long-term runoff integrated probability prediction method based on deep learning according to claim 1 is characterized in that: The screening process of the multiple factor screening method includes: Collect and preprocess the basin runoff data, remove factors with high missing rate, fill in missing values ​​and standardize them to obtain the preprocessed data set; The improved time-lagged Pearson correlation coefficient method was used to calculate the correlation coefficient within the 1-12 month lag window, and the most significant factors were selected to form the solution set A1; The multi-scale maximum information coefficient method is used to calculate the mutual information value at the monthly, quarterly and annual scales, and the most significant factors are selected to form the solution set A2; The physical constraint random forest feature importance scoring method was used to calculate the variance increment score, and the most significant factors were selected to form the solution set A3.

9. The medium- and long-term runoff integrated probability prediction method based on deep learning according to claim 1 is characterized in that: The deep learning base model is a hydrological time series adaptive hybrid neural network architecture, and its construction process includes: The variational mode decomposition method is used to decompose the original runoff series into long-term trend component, seasonal component and random fluctuation component. Construct a specialized neural network, use the multi-layer perceptron MLP to process the long-term trend component, use the long short-term memory network LSTM to process the seasonal component, and use the Transformer to process the random fluctuation component; Design a dynamic attention fusion layer to calculate the attention weight coefficients α, β, and γ according to the hydrological characteristics of the target period, satisfying α+β+γ=1, where α, β, and γ are the output weights of the MLP, LSTM, and Transformer models, respectively; The historical similar hydrological year characteristics are introduced as auxiliary input, and the k most similar hydrological year runoff data are selected by calculating the Euclidean distance to construct an enhanced feature matrix; The weighted output results of the three sub-models are integrated to generate the final runoff prediction value Q_pred = α·Q_MLP + β·Q_LSTM + γ·Q_Transformer, where Q_MLP, Q_LSTM, and Q_Transformer are the prediction results of the three models respectively.

10. The medium- and long-term runoff integrated probability prediction method based on deep learning according to claim 4 is characterized in that: The Bayesian optimization method is a Bayesian optimization method guided by hydrological parameter sensitivity, and the optimization process also includes: The hydrological parameter sensitivity matrix S was constructed, and the sensitivity coefficient of each hyperparameter to the model performance was calculated by Morris screening method; The hyperparameter prior distribution P is established based on the watershed characteristics including area, slope, and vegetation coverage; Construct a multi-objective hydrological evaluation function H(θ) = w1·NSE + w2·(1-|RC_bias|) + w3·(1-|WB_res|), where NSE is the Nash efficiency coefficient, RC_bias is the runoff coefficient deviation, WB_res is the water balance residual, w1, w2, and w3 are weight coefficients and satisfy w1+w2+w3=1; A staged optimization strategy is adopted to select the hyperparameter sets θ_wet and θ_dry for the wet and dry seasons respectively; Construct a Gaussian process surrogate model GP(x) and use the expected improvement acquisition function EI(x) to select the next set of hyperparameters; According to the hydrological characteristics of different forecast months, the corresponding optimal hyperparameter set is switched to generate the final model prediction results.

Citation Information

Patent Citations

  • SE20095C1

  • Customer waiting time and response time confidence interval prediction method for G / G / 1 queuing system

    CN112163686A

  • Soil risk assessment method for solid waste stockpiling place

    CN115841248A