Long-term discharge probability prediction method for dual-driven reservoirs
Through the dual-drive reservoir long-term outflow probability prediction method, combined with the causal consistency probability inflow prediction and the mechanism-constrained outflow prediction model, the problem of inconsistent probability prediction results in the reservoir long-term outflow probability prediction and the scheduling scheme violates the physical mechanism, and more accurate and consistent prediction results are achieved.
Patent Information
- Application Number
- CN202510909221.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-07-02
AI Technical Summary
The existing long-term probability prediction method for reservoirs has the problem of crossing quantiles in the quantile results of probability prediction and the final scheduling scheme violates the physical mechanism.
A dual-drive method is adopted, combining historical hydrological meteorological observation data, atmospheric circulation index and reservoir historical operation data, a causal consistency probability inflow prediction model is constructed, a multi-dimensional probability inflow information set is generated, and a mechanism-constrained outflow prediction model is used to generate a long-term outflow probability prediction scheme for reservoirs that conform to physical laws.
Generating a reservoir outbound prediction plan that combines probability self-consistentness and physical reliability improves the accuracy and consistency of predictions, and solves the problems of inconsistent probability prediction and decision-making results in the existing technology do not conform to the physical mechanism.
Smart Images

Figure CN120410281B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a prediction method, in particular to a method for predicting the probability of long-term discharge from a reservoir. Background Art
[0002] Reservoir operation decisions are essentially a sequential decision-making process based on future inflow forecasts. Long-term (e.g., several months to a year into the future) forecasts of discharge probabilities are crucial for determining the trade-off between the benefits and risks of reservoirs. Against the backdrop of intensifying global climate change and the profound impact of human activities, the nonlinearity and uncertainty of river basin hydrometeorological processes have significantly increased, leading to increasingly severe randomness and volatility in reservoir inflow sequences. Therefore, how to abandon traditional empirical operation models and develop advanced forecasting methods that can accurately characterize future inflow uncertainties and make reliable operation decisions, thereby promoting the transformation of reservoir management to a refined, intelligent, and risk-informed model, is a major issue that urgently needs to be addressed in the field of water resources engineering.
[0003] Currently, research on long-term reservoir runoff forecasting and operational decision-making has made significant progress, forming three major technical approaches: rainfall-runoff physical models, traditional time series statistical models, and data-driven machine learning models. Physically process-based distributed hydrological models, such as SWAT and VIC, can simulate the complete evolutionary chain from rainfall to runoff by meticulously characterizing the underlying surface characteristics and the physical processes of the water cycle, offering good mechanistic interpretability. Traditional time series statistical models, such as the autoregressive moving average (ARIMA) model and its variants, construct forecasting models by analyzing statistical characteristics such as autocorrelation and trend in historical runoff series. In recent years, with the advancement of artificial intelligence (AI), machine learning models, such as artificial neural networks (ANNs), support vector machines (SVMs), and long-short-term memory networks (LSTMs), have also gained widespread application in hydrological forecasting due to their powerful nonlinear mapping capabilities. These models are capable of learning complex response relationships from high-dimensional historical hydrometeorological data. Furthermore, to address forecasting uncertainties, researchers have begun focusing on probabilistic forecasting. Commonly employed methods include Bayesian model averaging, integrated forecasting systems, and quantile regression techniques. These aim to provide a probability distribution or confidence interval encompassing multiple possibilities for future runoff, rather than a single, deterministic prediction value. Furthermore, utilizing large-scale climate factors (such as the El Niño-Southern Oscillation (ENSO) and the North Atlantic Oscillation (NAO)) as predictors and exploring their teleconnected influences on regional runoff has also become an important approach to improving forecast accuracy.
[0004] For the above-mentioned runoff prediction problem of the coupled meteorological-hydrological-engineering multi-process, the single-driven model is difficult to achieve accurate modeling of the entire chain due to its inherent limitations. Although the data-driven model can effectively explore the nonlinear statistical laws between meteorological-hydrological factors and runoff, its black box characteristics make it difficult to express the mechanism of the engineering control process. Although the single-mechanism driven model can describe the engineering control process through water balance equations and scheduling rules, due to the incomplete understanding and description of the dynamically changing hydrological process, it cannot cover the complex response relationship between meteorological-hydrological processes and it is difficult to capture the cross-scale nonlinear correlation between atmospheric circulation factors and runoff. Summary of the Invention
[0005] The purpose of the invention is to solve the problem that the probability prediction results of the existing methods have quantile crossing and the final scheduling plan violates the physical mechanism.
[0006] The technical solution is a dual-driven reservoir long-term discharge probability prediction method, including:
[0007] Based on the basic data formed by historical hydrological and meteorological observation data, atmospheric circulation index data and historical reservoir operation data, an initial factor set for inflow prediction with different time lag effects is constructed;
[0008] The causal consistency probabilistic inflow prediction model is used to process the initial inflow prediction factor set and generate a multi-dimensional probabilistic inflow information set.
[0009] The multi-scenario outflow prediction factor set is constructed by integrating the multi-dimensional probabilistic inflow information set and the dispatch factors extracted from the basic data;
[0010] The mechanism-constrained outflow prediction model is used to process the multi-scenario outflow prediction factor set to generate the final long-term reservoir outflow probability prediction scheme.
[0011] Preferably, the step of applying the causal consistency probabilistic inflow prediction model to process the initial inflow prediction factor set to generate a multi-dimensional probabilistic inflow information set includes:
[0012] The causal consistency probabilistic inflow prediction model is used to make predictions, and a physically causal consistent probabilistic inflow prediction sequence without quantile crossing is obtained.
[0013] During the prediction process, the hydrological process context encoding vector is extracted from the internal hidden layer state of the model;
[0014] Outputting a predicted distribution shape parameter vector representing the runoff probability distribution shape from a specific module of the model;
[0015] The probabilistic inflow prediction sequence, hydrological process context encoding vector and prediction distribution morphology parameter vector are integrated into a multidimensional probabilistic inflow information set.
[0016] Preferably, the network architecture of the causal consistency probabilistic inflow prediction model includes:
[0017] The upstream feature extraction backbone network receives the initial factor set for inflow prediction as input, performs spatiotemporal feature extraction, and outputs a high-dimensional time series feature code from which the hydrological process context code vector is extracted;
[0018] The parallel morphology controller module selects physical driving factors from the initial inflow prediction factor set as input and outputs the predicted distribution morphology parameter vector after learning;
[0019] The double-anchor symmetric chain output layer receives high-dimensional time series feature encoding and uses the predicted distribution morphology parameter vector as conditional input to generate a probabilistic inflow prediction sequence that is consistent with physical causality.
[0020] Preferably, the step of generating a physically causally consistent probabilistic inflow prediction sequence at the output layer of the double-anchor symmetric chain includes:
[0021] Based on the high-dimensional time series feature encoding and the predicted distribution morphology parameter vector, two anchor points of the probability distribution are predicted in parallel: the lowest quantile anchor point and the highest quantile anchor point;
[0022] Predict a set of normalized monotonic internal position parameter sequences;
[0023] By performing convex combination operations on the lowest quantile anchor point, the highest quantile anchor point and the monotonic internal position parameter sequence, all intermediate quantiles are reconstructed and generated to form the final physically causally consistent probabilistic inflow prediction sequence.
[0024] Preferably, the step of outputting the predicted distribution morphology parameter vector by the morphology controller module comprises:
[0025] From the initial set of inflow prediction factors, factors that have physical driving effects on runoff distribution are screened out to form a subset of physical driving factors.
[0026] Feed a subset of physical drivers into a lightweight feedforward neural network for learning;
[0027] A lightweight feedforward neural network is used to output the predicted distribution morphological parameter vector, which contains the scale and skewness information of the predicted distribution and is used to quantitatively describe the key morphology of the runoff probability distribution at the prediction moment.
[0028] Preferably, the causal consistency probability inflow prediction model is trained using a gradient heterogeneous modulation paradigm based on internal attention deconstruction, and the training steps include:
[0029] An auxiliary supervision branch is introduced after the upstream feature extraction backbone network to generate auxiliary prediction outputs;
[0030] Perform model forward propagation and calculate the final Pinball loss of the main task based on the physically causally consistent probabilistic inflow prediction sequence;
[0031] The auxiliary loss for the auxiliary task is calculated based on the auxiliary prediction output.
[0032] Preferably, the training step further includes generating a heterogeneous gradient modulation matrix, which includes:
[0033] During the forward propagation of the model, the internal attention weight tensor is parsed and extracted from its internal attention mechanism;
[0034] According to the final Pinball loss of the main task, the base modulation strength in scalar form is calculated through a dynamic gating function;
[0035] The base modulation strength is broadcasted and element-wise multiplied with the internal attention weight tensor to generate a heterogeneous gradient modulation matrix.
[0036] Preferably, the training step further comprises performing modulated backpropagation, the step comprising:
[0037] During the gradient backpropagation process, the gradient flow from the auxiliary loss back to the upstream feature extraction backbone network is multiplied element-by-element with the heterogeneous gradient modulation matrix;
[0038] The modulated gradient flow obtained after multiplication is applied to update the weights of the upstream feature extraction backbone network, thereby realizing non-uniform intelligent training of the front-end feature extraction module.
[0039] Preferably, the step of applying the mechanism-constrained outflow prediction model to process the multi-scenario outflow prediction factor set to generate a final reservoir long-term discharge probability prediction scheme includes:
[0040] Construct a mechanism-constrained outflow prediction model based on gradient boosting decision tree;
[0041] The model is trained using a mechanism-constrained composite loss function as the optimization objective, where the composite loss function forcibly couples the physical mechanisms of water balance and outflow boundary.
[0042] Using the trained model, the corresponding outflow prediction value is generated for each scenario in the multi-scenario outflow prediction factor set to form a multi-scenario outflow prediction sequence, which is then integrated into a long-term reservoir outflow probability prediction scheme.
[0043] Preferably, the mechanism-constrained composite loss function combines at least the following three parts in a weighted summation manner:
[0044] The root mean square error term, which is the basic prediction accuracy term, is used to measure the deviation between the predicted outflow value and the true value;
[0045] The water balance penalty term constructed based on the water balance equation is used to penalize the prediction results that violate the law of water conservation;
[0046] The outflow boundary penalty term constructed according to the reservoir operation regulations is used to penalize the prediction results that exceed the upper limit of discharge capacity or fall below the lower limit of ecological flow.
[0047] Preferably, the water balance penalty term is obtained by calculating the square of the difference between the net flow formed by the scenario inflow and the predicted outflow in a unit time step and the storage change during the period;
[0048] The outbound flow boundary penalty term is obtained through a one-way penalty function, which only generates a penalty based on the value of the predicted outbound flow when the value exceeds the preset outbound flow upper limit or falls below the outbound flow lower limit.
[0049] Preferably, the steps of training the model using the mechanism-constrained composite loss function as the optimization objective specifically include:
[0050] For the mechanism-constrained composite loss function, the calculation formulas of its partial derivative (gradient) with respect to the model prediction value and the second-order partial derivative (Hessian) are derived;
[0051] Write the calculation formulas of gradient and Hessian into executable custom objective functions that meet the requirements of the gradient boosting decision tree framework API;
[0052] During training, an executable custom objective function is provided to the model training interface, so that in each iterative construction of the model, its node splitting decision and the determination of leaf node values are directly guided by the physical mechanism penalty term in the composite loss function.
[0053] The beneficial effect is that it can solve the technical problems existing in the existing technology and generate a reservoir discharge prediction scheme with both probabilistic self-consistency and physical reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a flow chart of the present invention.
[0055] Figure 2 This is a flow chart of the present invention for generating a multi-dimensional probability inflow information set.
[0056] Figure 3 This is a flow chart of the double-anchor symmetric chain output layer generating a physically causally consistent probabilistic inflow prediction sequence.
[0057] Figure 4 It is a flow chart of the output of the predicted distribution morphological parameter vector by the morphological controller module of the present invention. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that the specific embodiments herein are only used to explain the present invention and are not intended to limit the present invention in any form.
[0059] First, the following embodiment is provided. According to one aspect of the present application, a method for predicting the long-term discharge probability of a dual-driven reservoir is provided, comprising the following steps:
[0060] According to one aspect of the present application, step S1 further comprises:
[0061] Step S11, generating the initial factor set for reservoir inflow prediction: There are many factors that affect runoff, and the impact of different factors has multi-order time lag differences. Meteorological factors (such as atmospheric circulation, temperature, air pressure, etc.) usually have long time lag effects, while hydrological factors (such as previous runoff, rainfall, evapotranspiration, soil moisture, etc.) show short time lag responses. According to the meteorological conditions and runoff characteristics of the study area, the main weather factors that control the rainfall runoff in the study area are determined, and the hydrological and meteorological factors that have a direct or indirect impact on the current period are selected as forecast factors, including previous rainfall, evapotranspiration, temperature, soil moisture, atmospheric circulation index and reservoir inflow, and considering the previous time lag effect, together form the initial factor set for runoff prediction X={x t ,x t-1 ,x t-2 ,…,x t-τ};
[0062] Step S12, screening of the preliminary driving factor set based on DS evidence theory: The reliability of predictive factor screening depends on the coordinated evaluation of multi-source statistical indicators. Traditional methods based on a single indicator are usually difficult to fully capture the complex linear, nonlinear and conditional dependencies between runoff and meteorological factors. The Pearson correlation coefficient and the Spearman rank correlation coefficient can capture linear and monotonic nonlinear relationships respectively, while partial mutual information can be used to identify local conditional dependencies. Therefore, the DS evidence theory can be used to integrate the correlation calculation results of the three complementary indicators of Pearson correlation coefficient, Spearman rank correlation coefficient and partial mutual information to obtain the final correlation ranking, and select several factors with the highest ranking as the preliminary driving factor set X'={x' t ,x' t-1 , x' t-2 ,…, x' t-τ}, the factors selected in this step ensure physical rationality and statistical relevance between them and the target predictor;
[0063] Step S13, screening of key driving factor sets based on SHAP interpretability technology: The coarse screening method based on DS evidence theory screens out a preliminary driving factor set through multi-evidence fusion, but its essence is still driven by static statistical indicators, and does not consider the nonlinear interaction effects and dynamic contribution differences between factors after model training. To overcome this limitation, the SHAP interpretability technology is used to quantify the marginal contribution of each factor to the model prediction, so as to achieve fine screening optimization based on mechanism interpretability. Compared with traditional feature selection methods, SHAP can not only identify the importance of global features, but also reveal the direction and degree of influence of factors on prediction results in different scenarios, which helps to improve the stability and generalization ability of the model. The gradient interpreter is selected as the interpretation method. Based on the gradient calculation of the loss function by the neural network, the specific contribution of each factor in the preliminary driving factor set to the reservoir inflow prediction results output by the QRCNN-BiLSTM model is measured, and the factors are ranked according to the contribution. Several factors with the highest ranking are selected as the final key driving factor set X”={x” t , x” t-1 , x” t-2 ,…, x” t-τ}.
[0064] In this example, a factor screening strategy combining coarse and fine screening was adopted. First, based on the DS evidence theory, a coarse screening was performed through multi-evidence fusion to screen out a preliminary set of driving factors, thereby ensuring the physical rationality and statistical correlation between the factors. Next, a fine screening was performed using the SHAP interpretability technique to further optimize the selection of factors, quantify the marginal contribution of each factor to the model prediction results, and reveal the direction and degree of its influence in the specific model. This coarse and fine screening factor screening method can fully capture the complex nonlinear relationships and dynamic contribution differences between factors, effectively improving the stability, generalization ability, and accuracy of the prediction model.
[0065] According to one aspect of the present application, step S2 is further:
[0066] Step S21: Couple a multi-scale convolutional neural network and a bidirectional long short-term memory network to construct a multi-model nested structure, wherein quantile regression is introduced into this structure to construct a QR-MC-CNN-BiLSTM model as an inflow prediction equation;
[0067] Step S22: Divide the dataset into a training set and a test set. The samples in the training set are used for hyperparameter optimization and model training. The model hyperparameters are adaptively optimized using the Tornado optimization algorithm and K-fold cross-validation is performed to obtain the optimal hyperparameter combination. The QR-MC-CNN-BiLSTM model is trained on the training set using this combination to obtain model parameters such as weights and biases.
[0068] Step S23: Use the test set samples to make predictions, generate reservoir inflow prediction results at different quantiles, take the reservoir inflow prediction process at the low quantile as the lower bound of the interval, and the reservoir inflow prediction process at the high quantile as the upper bound of the interval, generate an inflow probability prediction scheme containing a probability interval, and obtain the inflow single value prediction result by calculating the expected value of the inflow prediction results at different quantiles:
[0069] E( I t )=∑ a j=1 I ρ(j) t •P(j);
[0070] Where: I t is the inflow prediction result vector corresponding to different quantiles in time period t; a is the number of discrete quantiles; j is the quantile sequence number, and the corresponding quantiles are arranged from low to high; ρ(j) is the jth quantile α j The corresponding cumulative probability, ρ(j)=α j ; I ρ(j) t is the inflow prediction result under the jth quantile; P(j) is the relative probability within the adjacent quantile interval, that is, the probability density; E[•] represents the expected value operation, which is achieved by weighted summation of the quantile probability density.
[0071] According to one aspect of the present application, step S3 is further:
[0072] Step S31: Using the inflow prediction results generated at different quantiles of the reservoir inflow prediction model as multi-scenario input, set the discrete quantile set α = {α1, α2, ..., α k}, each quantile scenario corresponds to a complete inflow prediction sequence {I α t} T t=1 , where T represents the length of the time series;
[0073] In this example, in order to overcome the problem that traditional outflow prediction relies on a single inflow scenario and lacks uncertainty characterization, multi-quantile inflow prediction is introduced to construct multi-scenario input, which improves the adaptability and robustness of the prediction model to different hydrological scenarios.
[0074] Step S32: In order to enable the LightGBM model to identify the quantile corresponding to the current prediction scenario, a corresponding one-hot encoding vector e is constructed for each quantile. α ∈R k, where the i-th bit is 1 and the rest are 0. The encoding vector is embedded into the input features at each time step t to form an input variable containing a semantic quantile index. This guides the outflow prediction model to distinguish between scheduling responses at different quantiles and improves the model's ability to classify and learn regulatory response patterns. At each time step t, the same encoding is replicated for all α scenarios to ensure consistency with the sequence length.
[0075] In this embodiment, in order to overcome the problem that LightGBM cannot identify different quantile scenarios, the quantile one-hot encoding embedding method is adopted to guide the model to explicitly identify the current confidence scenario, enhance its classification recognition and differentiated modeling capabilities, and improve the prediction accuracy in boundary scenarios.
[0076] Step S33: In the prediction process of each quantile based on the QR-MC-CNN-BiLSTM model, corresponding to each time step t, the hidden layer output vector h is extracted from the BiLSTM structure. α t ∈R 2d , as a semantic encoding representation of historical meteorological and hydrological dynamics. This vector is obtained by sliding window input and only depends on the time step t and its historical window sequence to avoid future information leakage. The hidden state, as a compressed expression of the temporal context, is embedded as an important input feature of the outflow prediction model to improve the prediction performance and context adaptability in multi-quantile scenarios.
[0077] In this embodiment, in order to solve the problem of missing contextual information in the inflow prediction results, the BiLSTM hidden layer output is extracted as the historical window semantic vector and used as the background input of the scheduling behavior, so that the scheduling prediction has long-term lag characteristics, factor interaction relationships and seasonal pattern perception capabilities, which effectively improves the expressiveness and context adaptability of the outflow prediction in multiple scenarios.
[0078] Step S34: Integrate the predicted inflows, one-hot encoding vectors, and BiLSTM hidden layer outputs under different quantile scenarios obtained in steps S31 to S33 to construct an inflow factor scenario set;
[0079] Step S35: Based on the actual operation and dispatching rules of the reservoir and the engineering control mechanism, factors that have a significant impact on the reservoir outflow are extracted from the historical dispatching data to construct a stable key reservoir storage factor set. This factor set is consistent under all quantile scenarios, including the initial storage capacity V of the reservoir. t-1 , Reservoir storage capacity at the end of the period V t and reservoir autocorrelation lagged outflow {O t-1 ,O t-2 ,…, O t-τ};
[0080] In this embodiment, the system introduces reservoir storage characteristics as structural input, which effectively enhances the model's ability to identify engineering regulation responses and improves scheduling consistency and physical interpretability.
[0081] Step S36: Combine the several groups of inflow scenarios formed in steps S31 to S34 with the key reservoir regulation factors extracted in step S35 to construct a multi-scenario factor set for reservoir outflow prediction. The scenario set contains the same number of outflow prediction input scenarios as the number of inflow prediction scenario groups. Each scenario group consists of a set of inflow-related factors at a specific quantile and a fixed set of key reservoir regulation factors:
[0082] z α t ={I α t ,O t-1 ,O t-2 ,…,O t-τ, V t-1 ,V t ,e α ,h α t};
[0083] The final outflow prediction multi-scenario factor set structure is z∈R T˙k˙n , where k is the number of quantile scenarios and n is the feature dimension of each sample after splicing.
[0084] According to one aspect of the present application, step S4 is further:
[0085] Step S41: construct an outflow prediction model based on the LightGBM model, taking into account the physical mechanism of reservoir operation as its penalty term, including water balance constraints, reservoir outflow boundary constraints and non-negative constraints, guiding the model to conform to the prior knowledge of reservoir operation, and constructing a LightGBM model based on mechanism-driven simulation engineering regulation;
[0086] Step S42: Use the reservoir outflow prediction multi-scenario factor set constructed in step S3 as the independent variable factor of the LightGBM model, and the measured reservoir outflow as the dependent variable factor. Divide the above data samples into a training set and a test set, input the independent variable factors and dependent variable factors of the training period into the LightGBM model based on mechanism-driven simulation engineering regulation for training, obtain the mapping relationship between the independent variable factors and the dependent variable factors, and input the outflow prediction factor sets under different scenarios in the test period into the trained model to calculate multiple groups of outflow prediction scenarios;
[0087] The multiple groups of outflow prediction scenarios obtained in step S43 and step S42 constitute the conditional distribution samples of outflow. Based on the kernel density estimation method, the outflow results are probabilistically modeled and fitted with their probability density functions to obtain the reservoir outflow probability prediction scheme. The expected values under the multiple groups of outflow prediction scenarios are used as point prediction values to obtain the single value prediction results of the reservoir outflow:
[0088] E( O t )=∑ a j=1 O ρ(j) t •P(j);
[0089] Where: O t is the outflow prediction result vector corresponding to different quantiles in time period t; O ρ(j) t is the outflow prediction result under the j-th quantile.
[0090] In this example, water balance and boundary constraints are embedded in the LightGBM loss function, enabling explicit modeling of the engineering control mechanism. This effectively improves the physical consistency and boundary rationality of outflow forecasts, enhancing the model's usability and credibility in actual scheduling scenarios. Furthermore, the kernel density estimation method is used to further refine the probabilistic outflow forecasting scheme, providing multi-level prediction results for decision-making.
[0091] During use, the following problems were found:
[0092] First, the hydrological system is a typical non-stationary system, and its driving mechanism undergoes structural changes with sudden changes in climate modes (such as the occurrence of El Niño / La Niña events and the transition from warm and dry to cold and wet). Factors that are highly correlated with runoff under normal modes (such as previous rainfall) may see their correlation drop sharply under extreme drought modes, while another factor (such as deep soil moisture and upper-air circulation index) may jump to the top in importance. Existing factor screening methods cannot capture the dynamic non-stationarity of this correlation structure. They only provide the average optimal solution under all modes. This inevitably leads to suboptimal or even erroneous model input information at the moment of mode change and during the maintenance of the new mode.
[0093] Second, quantile regression uses the Pinball loss function to independently model different quantiles, resulting in the model being unable to guarantee the physical and causal consistency of the prediction results. For example, for inflow predictions at the same moment, there may be a quantile crossover / folding phenomenon where the 10% quantile prediction value is greater than the 90% quantile prediction value. On a deeper level, different quantiles are related in physical causes. Extreme rainfall events that lead to high quantiles (major floods) will inevitably raise the low quantiles (flow lower limit) at that moment. They share the same physical drive, but the response amplitude is different. The existing model severs the physical causal relationship and regards each quantile as an independent prediction task, resulting in the predicted probability distribution being fragmented and contrary to common sense. To this end, the following embodiments are further provided.
[0094] like Figures 1 to 4 Another set of embodiments of the present application is described as follows:
[0095] Example 1: This example provides a complete process for a dual-driven reservoir long-term discharge probability prediction method that couples meteorological, hydrological, and engineering processes. The steps are as follows:
[0096] Step S100: Based on basic data formed by historical hydrological and meteorological observation data, atmospheric circulation index data and historical reservoir operation data, an initial factor set for inflow prediction with different time lag effects is constructed.
[0097] In this embodiment, this step is the data preparation stage. First, the collected raw time series data are preprocessed, including cleaning outliers, interpolating missing values (for example, using linear interpolation or spline interpolation), and normalizing (for example, using Z-score standardization or Min-Max normalization) to eliminate the influence of data of different dimensions and obtain cleaned data. Subsequently, based on hydrological physical mechanisms and statistical correlation analysis (such as Pearson correlation coefficient and mutual information method), lag terms of different time scales are constructed for each variable in the cleaned data, and preliminary screening is performed in combination with domain knowledge, ultimately forming a high-dimensional initial factor set for inflow prediction that contains rich prediction information. This factor set is the data basis for all subsequent model training and prediction.
[0098] Step S200: Applying the causal consistency probabilistic inflow prediction model to process the initial inflow prediction factor set to generate a multi-dimensional probabilistic inflow information set.
[0099] In this embodiment, this step is used to generate an information-rich and internally consistent probabilistic inflow forecast. Unlike traditional methods, the multidimensional probabilistic inflow information set generated in this step not only includes the probabilistic forecast itself but also provides in-depth information on the forecast process and the resulting form. This provides downstream shipping decisions with a multi-dimensional basis that goes far beyond a single forecast value. The generation of this multidimensional information set lays a solid data foundation for subsequent refined and highly reliable shipping decisions.
[0100] Step S300: A multi-scenario outflow prediction factor set is constructed by fusing the multi-dimensional probabilistic inflow information set and the scheduling factors extracted from the basic data.
[0101] In this embodiment, this step integrates the complex probabilistic information about future inflows generated upstream (step S200) with scheduling factors that represent the current project state and historical operating experience (e.g., initial storage, historical outflows over the same period, downstream water demand, etc.). Specifically, for each inflow scenario represented by a quantile in the multidimensional probabilistic inflow information set, the scheduling factors are horizontally concatenated to construct a unique, high-dimensional feature vector for each scenario. The feature vectors for all scenarios together form a multi-scenario outflow prediction factor set, which describes the current scheduling decision context facing the reservoir under different future inflow possibilities.
[0102] Step S400: Apply the mechanism-constrained outflow prediction model to process the multi-scenario outflow prediction factor set to generate a final long-term reservoir outflow probability prediction scheme.
[0103] In this embodiment, this step is used to generate a final discharge plan that conforms to both data patterns and strictly adheres to the laws of physics. It receives as input the multi-scenario outflow prediction factor set constructed in the previous step and processes it through a model whose internal learning mechanism is deeply transformed by physical rules. This model can generate a discharge decision recommendation that satisfies water balance and boundary constraints for each input scheduling scenario. Ultimately, the discharge decision recommendations under all scenarios are integrated (for example, by fitting their probability distributions through kernel density estimation and calculating the expected value), forming a complete long-term reservoir discharge probability prediction plan that includes a multi-scenario prediction set, a probability density function, and a single-valued expected prediction.
[0104] This embodiment solves the two major pain points of the existing technology, namely, inconsistent probability predictions and decision results that do not conform to physical mechanisms, through two core driving links.
[0105] Example 2: This example describes in detail the construction, training (AD-HGM paradigm) and application process of the causal consistency probabilistic inflow prediction model (PDP model).
[0106] The network architecture of the causal consistency probabilistic inflow prediction model includes:
[0107] The upstream feature extraction backbone network receives the initial set of inflow prediction factors as input, performs spatiotemporal feature extraction, and outputs a high-dimensional time series feature code. The high-dimensional time series feature code is a compressed, information-dense vector sequence extracted from the original set of factors that characterizes the complex, nonlinear relationships of hydrological processes in both time and space.
[0108] In this embodiment, the backbone network adopts a heterogeneous nested structure of a multi-scale convolutional neural network (MC-CNN) and a bidirectional long short-term memory network (BiLSTM). The MC-CNN part sets three parallel 1D convolution branches, and the convolution kernel sizes can be set to 3, 6, and 12 respectively to capture local hydrological patterns at different time scales such as monthly, quarterly, and semi-annual. The output of each branch is max-pooled and then spliced in the feature dimension. The spliced features are fed into a two-layer BiLSTM network, where the number of hidden units in each BiLSTM layer can be set to 128. The BiLSTM network is used to analyze long-range temporal dependencies in hydrological sequences. Ultimately, the hidden state sequence of the last layer of BiLSTM is the high-dimensional time series feature encoding.
[0109] The combination of MC-CNN and BiLSTM can fully utilize the advantages of CNN in capturing local and multi-scale patterns and the ability of LSTM in handling long sequence dependency problems, thereby achieving a more comprehensive and in-depth feature representation of hydrological processes.
[0110] In some embodiments, the backbone network may also adopt other structures with time series feature extraction capabilities, for example, using a Transformer encoder based on a self-attention mechanism, or using a gated recurrent unit (GRU) instead of BiLSTM.
[0111] The parallel morphology controller module is used to select physical driving factors from the initial set of inflow prediction factors as input and, after learning, outputs a predicted distribution morphology parameter vector Θ. The predicted distribution morphology parameter vector Θ is a low-dimensional vector whose each element directly corresponds to a key morphology parameter of the runoff probability distribution at the prediction time, such as location, scale (indicating fatness), and skewness (indicating symmetry).
[0112] In some processes, strong physical signal factors that have a decisive influence on runoff distribution morphology are first selected from the initial set of inflow prediction factors based on prior domain knowledge, such as the ENSO index, regional cumulative rainfall, and extreme temperature indices, to form a subset of physical driving factors. This subset is then input into an independent two-layer fully connected neural network (i.e., a feedforward network). The hidden layer dimension of this network can be set to 16, and the number of neurons in the output layer is equal to the number of required morphological parameters (for example, if a split normal distribution is used, the location, scale, and skewness parameters are required). To ensure the physical meaning of the parameters, the output layer uses a specific activation function, such as the Soft Plus function to ensure that the scale parameter is positive and the Tanh function to ensure that the skewness parameter is within the range [-1, 1].
[0113] By establishing a direct mapping path from large-scale physical causes to specific predicted distribution forms, the model's predictions are no longer purely black-box fitting, but are guided and constrained by physical laws, and the generated probability distribution is more consistent with the actual physical characteristics of the hydrological process.
[0114] The double anchor symmetric chain (DASC) output layer is used to receive high-dimensional time series feature encoding and use the predicted distribution morphology parameter vector Θ as conditional input to ultimately generate a probabilistic inflow prediction sequence that is consistent with physical causality.
[0115] This output layer eliminates quantile crossings by:
[0116] a. Concatenate the high-dimensional time series feature encoding with the predicted distribution morphology parameter vector Θ to form an enhanced conditional feature vector.
[0117] b. Based on the enhanced feature vector, two independent linear output neurons are used to predict the two anchor points of the probability distribution: the lowest quantile anchor point q min and the highest quantile anchor point q max .
[0118] c. At the same time, another set of independent output neurons (whose number is equal to the number of required intermediate quantiles) predicts a set of initial internal position offsets. These offsets are activated to ensure non-negativity, and then through cumulative summation operations, a monotonic internal position parameter sequence {β i}.
[0119] d. Finally, through the convex combination formula q i =(1-β i )q min +β i q max, using the anchor point and location parameters obtained in steps b and c, calculate the values of all intermediate quantiles to form a non-crossover physically causally consistent probabilistic inflow forecast sequence.
[0120] Through the DASC layer, instead of directly predicting each quantile, we predict the upper and lower bounds of the entire distribution and the relative position within it. Due to the mathematical properties of convex combinations, as long as {β i}sequence is monotonic, the output {q i The sequence is also necessarily monotonic. This structure eliminates the possibility of quantile crossing and ensures the theoretical consistency and usability of the probability prediction results.
[0121] Through the collaborative work of the above three modules, the PDP model can receive the initial inflow prediction factor set and ultimately generate a multidimensional probabilistic inflow information set containing three parts of information:
[0122] Physically causally consistent probabilistic inflow prediction sequence from the DASC output layer.
[0123] The hydrological process context encoding vector is extracted from the hidden state of the BiLSTM module of the backbone network, that is, the high-dimensional temporal feature encoding.
[0124] The predicted distribution morphology parameter vector Θ from the morphology controller module.
[0125] In order to solve the possible gradient learning efficiency mismatch problem of heterogeneous networks (CNN and LSTM) in the PDP model, the AD-HGM paradigm is adopted for training.
[0126] Constructing dual tasks and calculating losses involves the following steps:
[0127] In the upstream feature extraction backbone of the PDP model, specifically after the MC-CNN module and before the BiLSTM module, an auxiliary supervision branch is introduced. This branch is typically a fully connected layer with an output dimension identical to the prediction target, used to generate auxiliary prediction outputs. During each forward propagation, a batch of initial factors for inflow predictions are fed into the model, and the following are calculated:
[0128] Final Pinball loss L final : Calculated based on the final output of the main task (i.e., the physically causally consistent probabilistic inflow prediction sequence output by the DASC layer) and the true value. Pinball loss is a standard loss function for quantile regression.
[0129] Auxiliary loss L aux : The auxiliary prediction output based on the auxiliary supervision branch output is calculated with the true value, usually using the mean square error (MSE) loss.
[0130] Generating a heterogeneous gradient modulation matrix includes the following steps:
[0131] a. During the forward propagation process, the internal attention weight tensor At for different features when processing each time step is parsed and extracted from the BiLSTM module of the backbone network.
[0132] b. According to the final Pinball loss L of the main task final The current value of the dynamic gating function, such as Gt=Softplus(w•(L final -L threshold )+b), calculate the basic modulation intensity Gt in scalar form. threshold is a hyperparameter. This gating function generates a larger modulation intensity when the primary task loss is large, thereby strengthening the training of the front-end network. Larger, smaller, etc., all indicate values greater than or less than a preset threshold.
[0133] c. Broadcast the scalar basis modulation intensity Gt and the internal attention weight tensor At and multiply them element-wise to obtain the final heterogeneous gradient modulation matrix Mt=GtAt, the Hadamard product of the two matrices.
[0134] Each element of the Mt matrix incorporates dual information from the global learning state (represented by Gt) and the local feature preference (represented by At). Therefore, it is no longer a uniform, indiscriminate gradient scaling factor, but a highly differentiated, sophisticated gradient router.
[0135] Perform modulated backpropagation, including:
[0136] During the gradient backpropagation process, the gradient hook function of the deep learning framework (such as PyTorch's register_backward_hook) is used to intervene in the gradient flow.
[0137] Specifically, from the auxiliary loss L aux The gradient flow returned to the MC-CNN module must be multiplied by the heterogeneous gradient modulation matrix Mt element-wise before applying the gradient to update the weights of MC-CNN. The final Pinball loss L of the main task final The gradient flow propagates normally and is not affected.
[0138] The above operation intelligently distributes the auxiliary supervision gradients. The auxiliary gradients are strengthened for important features of the primary task (i.e., those with higher weights in At); otherwise, they are weakened. Furthermore, when the primary task has been well learned (i.e., Gt is small), the overall impact of the auxiliary gradients is reduced, avoiding unnecessary interference with the already converged primary task. Experimental data shows that training using the AD-HGM paradigm accelerates model convergence by approximately 15% compared to traditional auxiliary learning methods, and improves final prediction accuracy (measured by CRPS) by approximately 5-10%.
[0139] Through the above training process, a PDP model that can efficiently generate multi-dimensional inflow information is finally obtained.
[0140] Example 3: Detailed description of the construction and application process of the mechanism-constrained outflow prediction model. Specifically, the following steps are included:
[0141] Design a mechanism-constrained composite loss function. Specifically, to embed the physical rules into the model, a mechanism-constrained composite loss function L=w1•L is constructed, which consists of a weighted sum of three parts: RMSE +w2•L balance +w3•L boundary The weights w1, w2, and w3 are hyperparameters that can be determined through methods such as cross-validation. For example, they can be set to w1=1.0, w2=0.8, and w3=1.2.
[0142] Basic prediction accuracy term L RMSE , the standard root mean square error is used to ensure the basic fitting ability of the model.
[0143] Water balance penalty term L balance , which is used to measure the degree to which the model prediction results violate the physical laws of water balance.
[0144] Based on the water balance equation (inflow - outflow) × Δt = storage change, a penalty term in the form of mean square error is constructed: L balance =((I sce -O pred )•Δt-(V end -V start )) 2 Among them, I sce is the scenario inflow, O pred is the outflow predicted by the model, Δt is the time step, V start and V end These variables can be obtained from the multi-scenario outflow prediction factor set.
[0145] Outbound traffic boundary penalty term L boundary, which is used to measure the degree to which the model prediction results violate the physical or ecological boundaries defined in the reservoir operation regulations.
[0146] The rectified linear unit (ReLU) function is used to construct a unidirectional penalty. Specifically, L boundary =ReLU( pred- O max )+ReLU(O min- O pred ). Among them, O max and O min are the upper limit (such as the maximum discharge capacity) and lower limit (such as the minimum ecological flow) of the outflow defined in the dispatching regulations. The ReLU function ensures that only when the predicted value O pred This only produces a penalty greater than zero when the boundary is crossed.
[0147] The training mechanism constrains the LightGBM model, and LightGBM is selected as the basic model framework.
[0148] Each iteration of GBDT (i.e., building a new tree) requires the first- and second-order derivatives of the loss function with respect to the previous prediction (i.e., the gradient and Hessian). Therefore, existing GBDT models cannot directly use this composite loss function. To address this, the following solution is provided:
[0149] a. For the mechanism constraint composite loss function L, find its relation to the model prediction value O pred Analytical expressions for the first-order partial derivative (gradient) and second-order partial derivative (Hessian) of .
[0150] b. Write a Python function that complies with the LightGBM API. This function receives the predicted value and the true value (and other physical parameters) as input and returns a tuple (grad, hess) containing the gradient and Hessian values for each sample according to the formula derived in step a.
[0151] c. Call the LightGBM training API and pass the executable custom objective function implemented in step b as the fobj parameter. Use the multi-scenario outflow prediction factor set as input and the composite loss as the optimization objective to train the model.
[0152] Under this setting, when LightGBM constructs each decision tree, its splitting decisions and leaf node values are directly guided by the physical mechanism penalty (reflected by the gradient and Hessian). This allows the model to learn to obey the laws of physics while simultaneously learning the patterns in the data, rather than passively correcting itself after prediction. After training is complete, the trained mechanism-constrained outflow model is obtained.
[0153] The multi-scenario outflow prediction factor set for the test period is input into the trained mechanism-constrained outflow model. For each input quantile scenario, a corresponding outflow prediction value that satisfies the physical constraints is generated, forming a multi-scenario outflow prediction sequence. This sequence can be fitted with a probability density curve (PDF) using methods such as kernel density estimation (KDE), and its expected value is calculated as a single-value prediction result, ultimately resulting in a complete long-term reservoir discharge probability prediction scheme. Tests have shown that the discharge scheme generated using this method reduces the proportion of water balance constraint violations by over 99% compared to the unconstrained model, and all predictions fall within the scheduling boundaries, demonstrating extremely high reliability and practicality.
[0154] Example 4 describes the detailed workflow of the Dual Anchor Symmetric Chain (DASC) output layer. This module solves the quantile crossing problem that is prevalent in the prior art when generating probability prediction results.
[0155] The workflow of the output layer includes the following steps:
[0156] Step S310: Receive and integrate conditional information. This step receives two key information streams from upstream modules: the high-dimensional temporal feature encoding from the upstream feature extraction backbone network ( ), and the predicted distributed morphological parameter vector Θ from the morphological controller.
[0157] Specifically, these two vectors are concatenated along the feature dimension to form a richer, enhanced conditional feature vector that incorporates temporal dynamics and physical distribution priors. This vector serves as the unified input for all subsequent predictions, ensuring that the forecast results are guided by both historical data patterns and physical causes.
[0158] Step S320: Parallel prediction of distribution anchor points and internal position parameters. Based on the enhanced conditional feature vector, this step generates three different intermediate results through three parallel and independent prediction heads (PredictionHead).
[0159] The prediction head generally refers to several layers at the end of a neural network for generating specific outputs. In this embodiment, each head may be one or more fully connected layers.
[0160] Head 1: used to predict the lowest quantile anchor point q min It is a linear layer with an output unit of 1. Preferably, it can be followed by a Soft plus activation function to ensure that the predicted flow anchor point is non-negative, which is consistent with physical reality.
[0161] Head 2: used to predict the highest quantile anchor point q max Its structure is the same as that of the first head, and it can also be connected to a Soft plus activation function. min and qmax Together they define the support interval of the predicted probability distribution.
[0162] Head 3: Used to predict a set of unconstrained raw values of internal location parameters. The number of its output units is equal to the number of middle quantiles to be predicted (for example, if 9 quantiles are predicted, 9 raw values are output).
[0163] Step S330: Implement the monotonicity constraint of the internal position parameters. This is a key technical improvement to ensure that the quantiles do not cross. A series of mathematical transformations are performed on the original value sequence of the internal position parameters output by the head three in step S320 to generate a strictly monotonic monotonic internal position parameter sequence {β i}.
[0164] First, the original value sequence is passed through the Soft plus activation function to ensure that all elements are non-negative.
[0165] Secondly, a cumulative summation operation is performed on the sequence after the soft plus process. After this step, each element in the sequence is equal to the sum of itself and all previous elements, thus ensuring the strict monotony of the sequence.
[0166] Finally, the cumulative summed sequence is divided by the value of its last element (i.e., the sum of the sequence), and normalized to the interval [0,1] to obtain the final monotone internal position parameter sequence {β i}. This sequence satisfies 0≤β1≤β2≤…≤β N ≤1.
[0167] Step S340: Reconstruct the final quantile sequence through convex combination. This step uses the anchor points and monotonic position parameters generated previously to calculate the predicted values of all quantiles.
[0168] Specifically, applying the convex combination formula q i =(1-β i )•q min +β i •q max , is a monotone internal position parameter sequence {β i Each β in i , calculate the corresponding quantile prediction value q i . Due to q min ≤q max And {β i}sequence is monotonic, the calculated {q i}The sequence must also be monotonic.
[0169] As an example, for a certain moment, the model predicts q min =100m 3 / s,qmax =500m 3 / s. For the 25%, 50%, and 75% quantiles, the monotonic internal position parameter sequence generated in step S330 is {β1=0.2,β2=0.5,β3=0.75}. Then, the corresponding quantile prediction values are:
[0170] q 25% =(1-0.2)•100+0.2•500=80+100=180m 3 / s.
[0171] q 50% =(1-0.5)•100+0.5•500=50+250=300m 3 / s.
[0172] q 75% =(1-0.75)•100+0.75•500=25+375=400m 3 / s.
[0173] It can be seen that the output sequence (180, 300, 400) is strictly monotonic and has no crossover.
[0174] Through the above steps, the DASC output layer transforms the quantile prediction problem into a structured, constrained parameter prediction problem. Its characteristic is that it ensures the physical consistency of the output through mathematical construction rather than relying on model training itself, thereby 100% avoiding the quantile crossing phenomenon.
[0175] Example 5. This example details the complete workflow of the internal attention deconstruction-based gradient heterogeneous modulation (AD-HGM) paradigm for training the PDP model in a single training iteration. The main feature is the creation and application of a gradient modulation matrix, which enables differentiated and intelligent training of different modules in a heterogeneous network.
[0176] In a complete training iteration workflow, the steps are as follows:
[0177] Step S410: Forward propagation and multi-task loss calculation: Input a batch of inflow prediction initial factor sets and their corresponding true inflow labels into the constructed PDP model.
[0178] After the data flows through the MC-CNN module, a part of the features enters the auxiliary supervision branch and is processed to obtain the auxiliary prediction output. This output is compared with the true label to calculate the auxiliary loss L aux (For example, using MSE loss).
[0179] The main stream of data continues to flow through the BiLSTM module and the DASC output layer to obtain the final physical causal consistent probabilistic inflow prediction sequence.
[0180] During the BiLSTM module processing, the internal attention weight tensor At is parsed and temporarily stored.
[0181] Compare the predicted sequence output by DASC with the true label and calculate the final Pinball loss L for the main task final .
[0182] Step S420: Dynamically generate heterogeneous gradient modulation matrices. After the loss is calculated in the forward propagation and before the backpropagation begins, dynamically generate the key matrix for modulating the gradient.
[0183] The scalar loss value L calculated in step S410 is final , input into the dynamic gating function, such as Gt=1 / (1+exp(-k•(L final -L0))), where k and L0 are hyperparameters. This function calculates the scalar base modulation strength Gt. L final The larger it is, the closer Gt is to 1, and vice versa.
[0184] The scalar Gt is expanded to the same dimension as the internal attention weight tensor At temporarily stored in step S410 through a broadcast operation, and then element-wise multiplication is performed to obtain the final heterogeneous gradient modulation matrix Mt=Gt*At (Hadamard product).
[0185] Step S430: Modulated back propagation and weight update. This is the core step of the AD-HGM paradigm.
[0186] Start the automatic differentiation engine and begin gradient backpropagation.
[0187] From the final Pinball loss L final The returned gradient flow propagates normally and updates the weights of all related modules such as BiLSTM and DASC layers.
[0188] From the auxiliary loss L aux The returned gradient flow is intercepted by the preset gradient hook before reaching the MC-CNN module.
[0189] The intercepted gradient flow is element-wise multiplied with the heterogeneous gradient modulation matrix Mt generated in step S420.
[0190] Use the modulated gradient flow obtained after multiplication to update the weights of the MC-CNN module.
[0191] Through this complete iteration, the auxiliary task's training intensity on the front-end feature extractor (MC-CNN) is finely and dynamically regulated by the main task's learning state (Gt) and feature preference (At). This improvement avoids gradient conflicts and learning efficiency mismatches that can arise in simple multi-task learning, ensuring harmonious and efficient end-to-end training of the entire heterogeneous network.
[0192] Example 6: Detailed explanation of how to implement and apply a custom loss function containing physical mechanisms in LightGBM.
[0193] Step S510: Express the physical rules to be implemented using a clear mathematical formula and integrate them into a unified loss function. For example, the mechanism-constrained composite loss function defined in Example 3 is used: L(y,y*)=w1•(yy*) 2 +w2•((I sce -y*)•Δt-ΔV)2+w3•(ReLU(y*-O max )+ReLU(O min -y*)).
[0194] Among them, y* is the model prediction value O pred , y is the actual outflow (only for the first item), other physical parameters such as I sce ,ΔV,O max ,O min etc. are input as known conditions.
[0195] Step S520: For the composite loss function L(y,y*), calculate its first-order partial derivative (gradient grad) and second-order partial derivative with respect to the model prediction value y*.
[0196] grad=dL / dy =2w1(y*-y)-2w2((I sce -y*)•Δt-ΔV)•Δt+w3(I(y*>O max )-I(y* <O min ));
[0197] hess=d 2 L / dy* 2 =2w1+2w2Δt2; It should be noted that d is the partial derivative operator.
[0198] Note: I(•) is the indicator function. The treatment of the derivative of ReLU at non-differentiable points can be done by using methods such as subgradient. This is a simplified example.
[0199] Step S530: In Python, write a function whose signature must meet the requirements of LightGBM, usually defcustom_objective(y_true,y_pred).
[0200] Inside this function, y true It not only contains the real label, but also needs to be designed to pass additional physical parameters (such as I sce ,ΔV, etc.) of the data structure (e.g., dictionary or object). The function body uses the formula of step S520 to pred and from y true The various parameters parsed out in Calculate the grad and hess values corresponding to each sample.
[0201] The function finally returns a tuple (grad, hess) containing the gradient values and Hessian values of all samples.
[0202] When calling the lightgbm.train() function, pass this custom objective function as the fobj parameter.
[0203] No constraint checking or post-processing is performed outside the model. Instead, the penalty of physical rules is injected directly into the GBDT model's genes, which are constructed layer by layer and node by node, by providing the underlying gradient and Hessian information. This makes the final trained model naturally have the instinct to obey physical mechanisms.
[0204] In a specific embodiment of the present invention, atmospheric circulation indices refer to quantitative indicators used to describe the state of atmospheric motion on a global or regional scale. These indices typically reflect climate anomalies that occur over a long period (months to years) and are caused by the interaction between the ocean and the atmosphere. The technical principle behind this is that these large-scale climate anomalies, through atmospheric teleconnection mechanisms, can have a significant and predictable impact on key meteorological and hydrological factors such as long-term rainfall and temperature in a specific region. Therefore, introducing these indices as predictive factors into the model helps the model capture low-frequency, long-term hydrological variations, thereby improving the accuracy and reliability of long-term forecasts.
[0205] Specifically, when constructing the initial set of factors for inflow prediction, a range of atmospheric circulation indices can be obtained from publicly available meteorological data service providers. Examples of these indices mentioned in the proposal include the ENSO (El Niño-Southern Oscillation) and the NAO (North Atlantic Oscillation). In practical applications, the Nino3.4 sea temperature anomaly index can be used to quantify the state of the ENSO, or a time series of the NAO index can be used. Before being used as model input, these indices undergo preprocessing, similar to other meteorological and hydrological data, including data cleaning, interpolation, and normalization.
[0206] In a specific embodiment of the present invention, a dispatch factor refers to a key variable extracted from historical reservoir operating data that directly reflects the current state of the reservoir project and historical dispatch and operation experience. While future inflows are certainly important in reservoir dispatch decisions, the current reservoir state (such as water storage) and existing operating rules or experience are also core factors in determining the next step. The purpose of introducing a dispatch factor is to provide the downstream mechanism-constrained outflow prediction model with the real-time engineering context necessary for decision-making, so that its predicted outflow flow not only responds to future water inflows but also conforms to current engineering realities.
[0207] Specifically, in the process of constructing a multi-scenario outflow prediction factor set, the scheduling factors are integrated with the multidimensional probabilistic inflow information set output by the upstream model. Examples of examples provided in the scheme include the beginning-of-period storage, the end-of-period storage target, and the historical outflow for the same period. For example, when forecasting for the next month, the actual water storage at the beginning of the current month, the storage corresponding to the target water storage level at the end of the month set according to the scheduling plan, and the average outflow for that month over the past ten years are used as scheduling factors and spliced with the forecast information for each inflow scenario.
[0208] In a specific embodiment of the present invention, physics-driven parameterized distributional morphology (PDP) is the name of the overall design concept and architecture of a deep neural network model designed for causally consistent probabilistic inflow forecasting. Its core technical principle lies in abandoning the traditional model approach of black-box fitting by mixing all factors into an input. Instead, it innovatively designs parallel modules (i.e., morphology controllers) that utilize specific factors strongly related to the physical causes of runoff to directly learn and control (i.e., parameterize) the probabilistic distributional morphology of the final forecast results, thereby achieving a structured, physically interpretable reconstruction of the predicted distribution.
[0209] Specifically, the network architecture, built on the PDP principle, consists of three core modules working together. First, there is the upstream feature extraction backbone network, which extracts deep spatiotemporal features. For example, it uses a combination of MC-CNN and BiLSTM. Second, there is the parallel morphology controller module, which receives physical driving factors such as the atmospheric circulation index and outputs a prediction distribution morphology parameter vector Θ that describes information such as the scale and skewness of the prediction distribution. Finally, there is the dual-anchor symmetric chain (DASC) output layer, which receives the feature encoding from the backbone network and, guided by the morphology parameter vector Θ, generates the final probabilistic inflow prediction sequence in a cross-over-free manner.
[0210] In a specific embodiment of the present invention, the Dual Anchor Symmetric Chain (DASC) refers to an output layer structure applied to the PDP network architecture. By changing the prediction target and utilizing mathematical construction, this solves the quantile crossing problem caused by independent prediction of each quantile in the existing technology. This ensures that the probability prediction sequence output by the model is strictly monotonic and inherently theoretically self-consistent, making it possible to construct and apply subsequent probability density functions.
[0211] Specifically, the workflow of the DASC output layer is not to directly predict the quantile q i It first predicts two structural anchor points of the entire probability distribution in parallel: the lowest quantile anchor point q min and the highest quantile anchor point q max。 At the same time, it also predicts a set of normalized, monotonic internal position parameters {β i}. Finally, all the middle quantiles q i is obtained by the convex combination formula q i =(1-β i )•q min +β i •q max Calculate and reconstruct the generated. min ,q max And the monotonic sequence {β i The mathematical properties of} ensure that the final calculated quantile sequence {qi} must be monotonic, thus eliminating the possibility of quantile crossing at the structural level.
[0212] In a specific embodiment of the present invention, internal attention-deconstructed gradient heterogeneous modulation (AD-HGM) is an advanced training paradigm for training heterogeneous deep neural networks such as PDP. This paradigm aims to address the problem that in complex models containing different types of sub-networks (such as CNNs and RNNs), due to differences in network structure and depth, the gradient signal returned from the final task may not be able to fully and effectively train all modules, especially the front-end feature extraction module. AD-HGM addresses this gradient learning efficiency mismatch by introducing auxiliary tasks that are intelligently controlled by the main task.
[0213] Specifically, the AD-HGM paradigm dynamically generates a heterogeneous gradient modulation matrix Mt. This matrix is obtained by multiplying two parts: one part is based on the final loss L of the main task. final The calculated global scalar basis modulation strength Gt reflects the overall learning state of the model. The other part is the internal attention weight tensor At parsed from the model (such as the BiLSTM module), which reflects the model's local preference for different features when processing sequences. During backpropagation, the auxiliary task loss L auxThe returned gradient stream is element-wise multiplied by the modulation matrix Mt before updating the front-end network (such as MC-CNN). This operation makes the distribution of auxiliary gradients highly intelligent, enabling efficient and conflict-free training of the front-end feature extraction module.
[0214] In a specific embodiment of the present invention, a mechanism-constrained composite loss function directly encodes the physical laws (such as water balance) and engineering boundaries (such as discharge limits) that must be adhered to during reservoir operation into a differentiable mathematical penalty term. This penalty term is then combined with a traditional loss term used to measure prediction accuracy. By optimizing this composite loss function, the model is forced to learn and adhere to these physical laws while simultaneously learning from the data.
[0215] Specifically, the composite loss function is constructed by weighted summation of at least three sub-terms:
[0216] Basic prediction accuracy items, such as the root mean square error, are used to ensure the basic fitting accuracy of model predictions.
[0217] The water balance penalty term is a penalty term constructed based on the water balance equation (i.e., (inflow - outflow) × Δt = storage change) and is used to penalize prediction results that violate this physical law.
[0218] The outflow boundary penalty term is constructed based on physical boundaries such as the upper limit of discharge capacity and the lower limit of ecological flow defined in the reservoir operation regulations. In this way, the physical mechanism is integrated into the model learning process as an endogenous constraint rather than an external post-processing method.
[0219] In a specific embodiment of the present invention, the mechanism constraint model generates a corresponding outflow forecast value that satisfies physical constraints for multiple input scenarios (each scenario corresponds to an inflow quantile), thereby forming a multi-scenario outflow forecast sequence containing multiple possible outflow values. This sequence can be viewed as a set of discrete samples drawn from the conditional probability distribution of future outflows. To obtain a continuous probability distribution description, this solution uses the KDE method to process these sequence samples and fit a probability density curve for the final outflow. Based on this curve, its expected value, variance, specific confidence interval, and other parameters can be further calculated to form a probabilistic forecast solution.
[0220] The preferred embodiments of the present invention are described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the scope of protection of the present invention.
Claims
1. A method for predicting the long-term discharge probability of a dual-drive reservoir, characterized in that: include: Based on the basic data formed by historical hydrological and meteorological observation data, atmospheric circulation index data and historical reservoir operation data, an initial factor set for inflow prediction with different time lag effects is constructed; The causal consistency probabilistic inflow prediction model is used to process the initial inflow prediction factor set and generate a multi-dimensional probabilistic inflow information set. The multi-scenario outflow prediction factor set is constructed by integrating the multi-dimensional probabilistic inflow information set and the dispatch factors extracted from the basic data; The mechanism-constrained outflow prediction model is used to process the multi-scenario outflow prediction factor set to generate the final long-term reservoir outflow probability prediction scheme; The causal consistency probabilistic inflow prediction model is applied to process the initial inflow prediction factor set to generate a multi-dimensional probabilistic inflow information set, including: The causal consistency probabilistic inflow prediction model is used to make predictions, and a physically causal consistent probabilistic inflow prediction sequence without quantile crossing is obtained. During the prediction process, the hydrological process context encoding vector is extracted from the internal hidden layer state of the model; Outputting a predicted distribution shape parameter vector representing the runoff probability distribution shape from a specific module of the model; The probabilistic inflow prediction sequence, hydrological process context encoding vector and prediction distribution morphology parameter vector are integrated into a multi-dimensional probabilistic inflow information set; The network architecture of the causal consistency probabilistic inflow prediction model includes: The upstream feature extraction backbone network receives the initial factor set for inflow prediction as input, performs spatiotemporal feature extraction, and outputs a high-dimensional time series feature code from which the hydrological process context code vector is extracted; The parallel morphology controller module selects physical driving factors from the initial inflow prediction factor set as input and outputs the predicted distribution morphology parameter vector after learning; The double-anchor symmetric chain output layer receives high-dimensional time series feature encoding and uses the predicted distribution morphology parameter vector as conditional input to generate a probabilistic inflow prediction sequence that is consistent with physical causality.
2. The method according to claim 1, characterized in that The steps of generating a physically causally consistent probabilistic inflow prediction sequence at the output layer of the double-anchor symmetric chain include: Based on the high-dimensional time series feature encoding and the predicted distribution morphology parameter vector, two anchor points of the probability distribution are predicted in parallel: the lowest quantile anchor point and the highest quantile anchor point; Predict a set of normalized monotonic internal position parameter sequences; By performing convex combination operations on the lowest quantile anchor point, the highest quantile anchor point and the monotonic internal position parameter sequence, all intermediate quantiles are reconstructed and generated to form the final physically causally consistent probabilistic inflow prediction sequence.
3. The method according to claim 2, characterized in that The steps of outputting the predicted distribution morphology parameter vector by the morphology controller module include: From the initial set of inflow prediction factors, factors that have physical driving effects on runoff distribution are screened out to form a subset of physical driving factors. Feed a subset of physical drivers into a lightweight feedforward neural network for learning; A lightweight feedforward neural network is used to output the predicted distribution morphological parameter vector, which contains the scale and skewness information of the predicted distribution and is used to quantitatively describe the key morphology of the runoff probability distribution at the prediction moment.
4. The method according to claim 1, wherein The causal consistency probability inflow prediction model is trained using a gradient heterogeneous modulation paradigm based on internal attention deconstruction. The training steps include: An auxiliary supervision branch is introduced after the upstream feature extraction backbone network to generate auxiliary prediction outputs; Perform model forward propagation and calculate the final Pinball loss of the main task based on the physically causally consistent probabilistic inflow prediction sequence; The auxiliary loss for the auxiliary task is calculated based on the auxiliary prediction output.
5. The method according to claim 4, characterized in that The training step also includes generating a heterogeneous gradient modulation matrix, which includes: During the forward propagation of the model, the internal attention weight tensor is parsed and extracted from its internal attention mechanism; According to the final Pinball loss of the main task, the base modulation strength in scalar form is calculated through a dynamic gating function; The base modulation strength is broadcasted and element-wise multiplied with the internal attention weight tensor to generate a heterogeneous gradient modulation matrix.
6. The method according to claim 5, characterized in that The training step also includes performing modulated backpropagation, which includes: During the gradient backpropagation process, the gradient flow from the auxiliary loss back to the upstream feature extraction backbone network is multiplied element-by-element with the heterogeneous gradient modulation matrix; The modulated gradient flow obtained after multiplication is applied to update the weights of the upstream feature extraction backbone network, thereby realizing non-uniform intelligent training of the front-end feature extraction module.
7. The method according to claim 1, characterized in that The steps of applying the mechanism-constrained outflow prediction model to process the multi-scenario outflow prediction factor set and generate the final reservoir long-term outflow probability prediction scheme include: Construct a mechanism-constrained outflow prediction model based on gradient boosting decision tree; The model is trained using a mechanism-constrained composite loss function as the optimization objective, where the composite loss function forcibly couples the physical mechanisms of water balance and outflow boundary. Using the trained model, the corresponding outflow prediction value is generated for each scenario in the multi-scenario outflow prediction factor set to form a multi-scenario outflow prediction sequence, which is then integrated into a long-term reservoir outflow probability prediction scheme.
8. The method according to claim 7, characterized in that The mechanism-constrained composite loss function combines at least the following three parts through weighted summation: The root mean square error term, which is the basic prediction accuracy term, is used to measure the deviation between the predicted outflow value and the true value; The water balance penalty term constructed based on the water balance equation is used to penalize the prediction results that violate the law of water conservation; The outflow boundary penalty term constructed according to the reservoir operation regulations is used to penalize the prediction results that exceed the upper limit of discharge capacity or fall below the lower limit of ecological flow.
Citation Information
Patent Citations
Meteorological forecast-based small and medium-sized reservoir capacity prediction method and system
CN116911178A
Reservoir irrigation area water resource allocation method based on big data analysis
CN119250476A