Hydrological Forecasting Method Driven by Both Physical Processes and Explainable Deep Learning

By combining physical processes and interpretable deep learning in hydrological forecasting, model parameters and hyperparameters are optimized, and model interpretability is improved through Shapley additive interpretation method, the problem of incomplete understanding of complex hydrological processes and insufficient interpretability of deep learning models in existing hydrological forecasting technologies is solved, and hydrological forecasting with high precision and low resource consumption is achieved.

CN119204355BActive Publication Date: 2025-06-03HOHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411707422.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-06-03
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

The existing hydrological forecasting technology has problems such as incomplete understanding of complex hydrological processes, data-driven models ignore physical mechanisms, and insufficient interpretability of deep learning models, resulting in deviations and uncertainties in forecast results.

Method used

The dual-driven hydrological forecasting method based on physical processes and interpretable deep learning is adopted. By constructing a basin hydrological model and deep learning model, combining the aurora optimization algorithm to optimize model parameters and hyperparameters, a dual-drive correction model is generated, and the interpretability of the model is improved through the Shapley additive interpretation method.

Benefits of technology

It significantly improves the accuracy and efficiency of real-time hydrological forecasting, reduces computing resource consumption, enhances the interpretability and reliability of the model, and ensures the accuracy and practicality of the forecast results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119204355B_ABST
    Figure CN119204355B_ABST
Patent Text Reader

Abstract

The present invention discloses a hydrological forecasting method driven by both physical processes and interpretable deep learning, which relates to the technical field of hydrological forecasting and is used to address the problems that process-driven models are difficult to fully generalize dynamic hydrological processes and data-driven models ignore physical processes. By integrating traditional hydrological models and interpretable deep learning, the advantages of both are fully utilized, thus significantly improving the accuracy and efficiency of real-time hydrological forecasting. The parameters of the hybrid model are optimized using a specifically improved aurora optimization algorithm for the head, which not only improves the prediction accuracy, ensures the reliability and stability of the model, but also reduces the consumption of computing resources. Moreover, by using the Shapley additive explanation method to explain the deep learning model, the interpretability of the model is further enhanced, enabling users to clearly understand the impact of input variables on the prediction results. This feature provides important support for the practical application of the model, especially in the field of hydrological forecasting where decision-making transparency is required.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hydrological forecasting, and more specifically, to a hydrological forecasting method driven by both physical processes and interpretable deep learning. Background Art

[0002] Hydrological forecasting is an extremely important non-engineering measure and plays a key role in watershed water resources management and flood disaster mitigation. The accuracy and efficiency of hydrological forecasting are crucial in areas such as flood warning, flood risk analysis, and reservoir operation. However, due to the complexity of hydrological processes and the limitations of human understanding, hydrological forecasting itself contains various uncertainties, such as data observation, model structure, and model parameters. These uncertainties lead to biases in the forecasting results and pose new challenges to the theory and methods of hydrological forecasting. Therefore, comprehensively analyzing uncertainty information and migration transformation principles is of profound significance for improving existing forecasting models.

[0003] Hydrological forecasting technologies are mainly divided into process-driven models and data-driven models. Process-driven models usually face challenges in generalizing the complex nonlinear relationship between environmental variables and runoff. This limitation stems from the incomplete understanding and representation of dynamic hydrological processes, thus bringing errors and uncertainties to flow prediction. On the other hand, the heavy dependence of data-driven models on data often leads to the neglect of the basic physical mechanisms governing hydrological processes. Therefore, developing a hybrid model that combines process-driven models and data-driven models is a promising and resource-saving strategy.

[0004] In recent years, deep learning, as a specialized subset of machine learning, has received extensive attention in hydrological research. Compared with traditional machine learning techniques, deep neural networks are distinguished by their complex architectures and larger numbers of neurons. However, interpretability remains an obvious limitation of deep learning, and the internal mechanisms of the model are usually difficult to directly clarify during the forecasting process. This opacity leads to insufficient credibility of the forecasting results, thus limiting the practical application of such models, especially in the field of real-time, hourly-scale hydrological forecasting. In addition, there is still an obvious trade-off between model accuracy and interpretability in deep learning models. Therefore, how to optimize the utilization of the unique advantages of process-driven models and data-driven models when constructing hybrid models, quantitatively evaluate the uncertainty of model outputs, and confirm the credibility of these hybrid models need to be comprehensively addressed within a unified theoretical framework.

[0005] In view of the above problems, the present invention proposes a solution. Summary of the Invention

[0006] To overcome the above-mentioned deficiencies of the prior art, embodiments of the present invention provide a hydrological forecasting method driven by both physical processes and interpretable deep learning to solve the problems raised in the above background art.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] A hydrological forecasting method driven by both physical processes and interpretable deep learning includes the following steps:

[0009] Construct a basin hydrological model based on a physical model, analyze basin characteristics and select a suitable rainfall-runoff model, calibrate and verify the model parameters through a head-specific improved aurora optimization algorithm, and generate an initial forecast result;

[0010] Construct a deep learning model with the forecast residuals of the basin hydrological model as the training target, and optimize the hyperparameters of the deep learning model through a head-specific improved aurora optimization algorithm;

[0011] Based on the initial forecast results of the basin hydrological model and the measured data, construct a multi-dimensional multi-step forecast residual matrix and a multi-dimensional multi-step hydrological element matrix, where the matrix represents the difference between the actual physical process and the simulated physical process of the basin;

[0012] Synchronously correct the forecast residuals predicted by the deep learning post-processing model and the original forecast results of the basin hydrological model for multiple lead times to generate a dual-driven correction model;

[0013] Generate real-time hydrological forecast results through the dual-driven correction model, and obtain multi-step corrected forecast results within different lead times;

[0014] Based on the corrected forecast results, use the Shapley additive explanation method to analyze the importance of the input variables of the deep learning post-processing model, quantify the contribution of each input variable to the correction result, and optimize the design of the model input variables and the correction mechanism.

[0015] In a preferred embodiment, the construction of the multi-dimensional multi-step forecast residual matrix includes:

[0016] Strip the forecast values from the result set of the basin hydrological model, and analyze the temporal correlation of the forecast differences in combination with the measured values;

[0017] Based on the temporal lag factor and the lead time setting, generate a multi-dimensional multi-step forecast residual matrix:

[0018] ;

[0019] In the formula, represents the forecast residual value at time t with a lead time of m and a previous lag period of s, represents From the moment to time t and from time t to the measured values at time are expressed as From the moment to time t and from time t to the predicted values at time

[0020] In a preferred embodiment, multi-horizon synchronous calibration is performed on the initial prediction results of the watershed hydrological model, specifically including:

[0021] The runoff prediction results of the dual-drive calibration model at time t and within the prediction horizon m are expressed as follows:

[0022]

[0023] where is the prediction residual at time t + m obtained through the deep learning post-processing model, is the predicted flow at time t + m obtained through the rainfall-runoff model, and represent the calculation processes of the deep learning post-processing model and the rainfall-runoff model respectively, is the prediction difference of the rainfall-runoff model at time t, are the hydrological elements required for the calculation of the rainfall-runoff model at time t + m, is the initial calculation condition of the rainfall-runoff model at time t + m.

[0024] In a preferred embodiment, the construction of the physical model includes the following steps:

[0025] The soil layer is divided into three layers to calculate the evapotranspiration respectively;

[0026] Based on the rainfall-runoff conditions, runoff generation is calculated, and the runoff is divided into surface runoff, subsurface runoff and baseflow;

[0027] Runoff routing calculation is performed by the Muskingum method to generate the hourly flow prediction results of the watershed.

[0028] In a preferred embodiment, the construction of the multi-dimensional multi-step hydrological element matrix includes:

[0029] Select the key hydrological elements of the watershed, including rainfall, rainfall-runoff, soil moisture and evapotranspiration;

[0030] Based on the temporal cross-correlation and lag characteristics, a multi-dimensional multi-step hydrological element matrix is generated, and the matrix is expressed as:

[0031] ;

[0032] where Denote the matrix value of hydrological elements at time t with a prediction period of m and a front lag period of s2. Denote From time to time t and from time t to The hydrological element values at time.

[0033] In a preferred embodiment, the specific steps for optimizing the model parameters using the head-specific improved aurora optimization algorithm include:

[0034] Generate an initial population based on the chaotic mapping method, where the chaotic mapping parameters are determined by analyzing the variation characteristics of the model parameters;

[0035] During the update process of the population individuals, introduce the gyroscopic motion, aurora ellipse trail, and particle collision mechanism to optimize the objective function;

[0036] Continuously iterate to update the population position until the objective function reaches the optimum or the maximum number of iterations is reached.

[0037] In a preferred embodiment, use Bayesian inference and Markov Chain Monte Carlo algorithm (MCMC) to quantify the uncertainty of the dual-drive correction model, specifically including:

[0038] Transform the output results of the dual-drive correction model into a normal distribution and generate an initial prior distribution;

[0039] Combined with the observed data, iteratively update the posterior distribution based on the Bayesian inference method to quantify the uncertainty of the model prediction results;

[0040] Generate the probability prediction interval for each moment through posterior distribution sampling analysis.

[0041] In a preferred embodiment, perform interpretability analysis on the deep learning post-processing model through the Shapley additive explanation method, including the following steps:

[0042] Calculate the Shapley values of the forecast residual matrix and the hydrological element matrix on the model output based on the Shapley additive explanation method;

[0043] Quantify the contribution of each input variable to the model output;

[0044] Combined with the runoff generation mechanism and error sources of the watershed hydrological model, explain the influence of input variables on the overall prediction results of the dual-drive model.

[0045] In a preferred embodiment, the dual-drive fusion mechanism of the deep learning post-processing model and the watershed hydrological model is applicable to different types of watershed hydrological models, including but not limited to the Xin'anjiang model, the SCS-CN model, and the TOPMODEL model.

[0046] Technical effects and advantages of the hydrological forecasting method driven by both physical processes and interpretable deep learning of the present invention:

[0047] By integrating traditional hydrological models with interpretable deep learning, the present invention makes full use of the advantages of both, thereby significantly improving the accuracy and efficiency of real-time hydrological forecasting. The parameters of the hybrid model are optimized by using a head-specific improved aurora optimization algorithm, which can not only improve the prediction accuracy, ensure the reliability and stability of the model, but also reduce the consumption of computing resources.

[0048] By using the Shapley additive explanation method to explain the deep learning model, the present invention further enhances the interpretability of the model, enabling users to clearly understand the impact of input variables on the prediction results. This feature provides important support for the practical application of the model, especially in the field of hydrological forecasting where decision-making transparency is required.

[0049] The present invention feeds the predicted forecast residuals back into the original forecast results to obtain corrected forecast results, and conducts comprehensive evaluation and comparison through the established evaluation index system, thereby ensuring the accuracy and practicality of the forecast results. Description of the Drawings

[0050] Figure 1 It is the flow chart of the dual-driven hydrological forecasting of the present invention;

[0051] Figure 2 It is the schematic diagram of the flow of the head-specific improved aurora optimization algorithm of the present invention;

[0052] Figure 3 It is the schematic diagram of the calculation flow of the Markov chain Monte Carlo algorithm based on the non-reversible sampler of the present invention. Detailed Embodiments

[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0054] The present invention combines traditional hydrological models with emerging interpretable deep learning, fully leveraging the advantages of both, and using the insights of interpretable deep learning to improve the accuracy and efficiency of real-time hydrological forecasting. The present invention calibrates the parameters of the Xin'anjiang model through a head-specific improved aurora optimization algorithm, outputs the original forecasting results and forecasting residuals; then, uses the head-specific improved aurora optimization algorithm to optimize the hyperparameters of the deep learning model to predict the forecasting residuals, and uses the Shapley additive explanation method to interpret the deep learning model. Finally, the predicted forecasting residuals are fed back into the original forecasting results to obtain the corrected forecasting results, and the Bayesian inference method is used to obtain the forecasting interval. Through the established evaluation index system for comprehensive evaluation and comparison, the accuracy and practicality of the forecasting results are ensured.

[0055] Embodiment, the present invention provides a hydrological forecasting method driven by both physical processes and interpretable deep learning. As Figure 1 shown, taking the Chitan Reservoir Basin as a research case, it includes the following steps:

[0056] S1: Systematically collect the hydrological data of the target basin, perform data cleaning and preprocessing operations on the data set, obtain the sub-flood data set of the target basin hourly, and divide the data set into three segments: training set, validation set, and test set;

[0057] S2: Construct a rainfall-runoff model that meets the conditions according to the basin characteristics, including lumped rainfall-runoff models, distributed rainfall-runoff models, and semi-distributed rainfall-runoff models. On the basis of S1, use the aurora optimization algorithm improved specifically for the head, take the range constraint conditions of the model parameter group as the search space, generate the initial population through the head-specific improved initialization of the objective function, where the objective function is used to describe the optimization objective of the hydrological forecasting system, the constraint conditions are used to describe the constraints of the physical basin on the rainfall-runoff model, and the individuals in the population are used to describe the different parameter expressions of the rainfall-runoff model for the physical basin. Perform fine calibration and strict verification of the parameters of the constructed rainfall-runoff model, record in detail the runoff forecasting results of each model run, and obtain the runoff forecasting results of the rainfall-runoff model on the test set based on the calibrated parameters;

[0058] S3: Separate the forecast values from the result set alone, analyze the temporal autocorrelation of the forecast differences, and construct a multi-dimensional multi-step forecasting residual matrix based on the measured values and the physical process model. The matrix characterizes the differences between the actual hydrological process and the simulated hydrological process, including differences in different time steps, differences in different residual variables, and differences in temporal sequence; analyze the average temporal cross-correlation of each hydrological element, and reorganize the S1 data set into a multi-dimensional multi-step hydrological element matrix according to the node with the largest average cross-correlation. The matrix characterizes the actual physical process of the basin, and perform normalization processing on it and the forecasting residual matrix respectively;

[0059] S4: Input the hydrological element matrix and the forecast residual matrix into the deep learning prediction model and the deep learning post-processing model respectively. Use the aurora optimization algorithm with specific improvements at the head to carefully optimize and strictly verify the hyperparameters of the models. Among them, the individuals in the former population are used to describe the non-linear expression of the prediction model for basin elements, and the individuals in the latter population are used to describe the non-linear expression of the post-processing model for forecast differences. Record in detail the runoff forecast results and residual prediction results of each model run, and obtain the prediction results of the two types of deep learning models on the test set based on the optimized hyperparameters respectively.

[0060] S5: Integrate the well-trained deep learning post-processing model in S4 and the calibrated rainfall-runoff model in S2 to obtain a dual-driven correction model, that is, use the rainfall-runoff model to drive the deep learning post-processing model and respond to the rainfall-runoff model in the form of feedback. Use the deep learning post-processing model to predict the forecast residuals within different lead times, screen the original forecast results of the rainfall-runoff model respectively to match the moments of the forecast residuals within different lead times, and then fuse the forecast residuals and the original forecast results at the same moment in the form of feedback for each lead time to obtain the runoff forecast results of the dual-driven correction model within different lead times.

[0061] S6: On the basis of deterministic forecasting, split the output result matrix into one-dimensional arrays according to the lead time, and perform normal distribution transformation that conforms to hydrological data respectively. Obtain the parameter initial value from the artificially assumed initial prior distribution, and then further establish a probabilistic forecasting model to quantify the forecasting uncertainty of the rainfall-runoff model and the dual-driven correction model in S5. Iteratively apply the Bayesian inference method in multiple one-dimensional arrays, and combine the Markov chain Monte Carlo algorithm based on the non-reversible sampler to perform posterior probability distribution sampling analysis on normal distributions with different parameters. Average a large number of sampling results to obtain the parameters of each normal distribution. Furthermore, use the deterministic forecast as the mean to construct the probabilistic forecast distribution at each moment. According to the specific requirements of hydrological forecast applications, calculate and determine the forecast interval at each moment under the specified confidence level to provide more comprehensive and reliable forecast information.

[0062] S7: Due to the complexity and black-box nature of the internal operation mechanism of deep learning, conduct a detailed interpretability analysis on the deep learning post-processing model in the dual-driven correction model. Re-enter the constructed forecast residual matrix into the deep learning post-processing model, record the Shapley values of the input variables at each moment within different lead times, and combine the runoff generation mechanism and possible error sources of the rainfall-runoff model to explain the impact of the errors generated by the rainfall-runoff model at different moments on the overall prediction of the dual-driven correction model, and further understand the actual runoff generation process of the rainfall-runoff model under the calibrated parameters, so as to improve the transparency and credibility of the model.

[0063] S8: Use an index system consisting of 7 quantitative evaluation indicators to evaluate the deterministic forecast results and probability forecast results respectively; conduct an all-round comparative analysis and evaluation of the proposed dual-driven hydrological forecasting method from multiple dimensions such as the bandwidth of the forecast interval, the coverage of the forecast interval, the comprehensive score of the probability forecast, the reliability of the probability forecast, the forecast accuracy, and the forecast error, so as to improve the comprehensiveness and reliability of the forecast results.

[0064] Furthermore, the specific steps of step S1 are as follows:

[0065] S11: According to the actual situations such as the geographical distribution, operation time, and control area of hydrological stations in the target basin, systematically collect hydrological station data with as long a time series as possible, as uniform or representative a spatial distribution as possible; standardize and time-segment the collected hydrological data according to hydrological specifications to obtain a standard hourly hydrological data set;

[0066] S12: Based on the recession characteristics of surface runoff, subsurface runoff, and groundwater runoff, as well as the mutual relationships among hydrological data, smoothly segment the hydrological data set according to the rising and recession conditions of a flood to obtain an hourly sub-flood data set;

[0067] S13: Considering the requirements of hydrological forecasting for hydrological data, divide the sub-flood data set into a training set (calibration set), a validation set, and a test set, with a ratio of 7:2:2. In the research case, use the data of the first 14 floods in the Chitan Reservoir Basin as the training set, the data of the middle 4 floods as the validation set, and the data of the last 4 floods as the test set.

[0068] Furthermore, the specific steps of step S2 are as follows:

[0069] S21: Analyze the characteristics of runoff generation and concentration in the basin, and select and construct a rainfall-runoff model that conforms to the basin's hydrological conditions: The lumped rainfall-runoff model does not consider the spatial heterogeneity of hydrological elements, the distributed rainfall-runoff model fully considers the spatial heterogeneity of hydrological elements, and the semi-distributed rainfall-runoff model is between the two. In the research case, the Chitan Reservoir Basin belongs to humid and semi-humid regions. According to hydrological conditions such as annual precipitation, the Xin'anjiang model is suitable for selecting the rainfall-runoff model;

[0070] S22: Use the range constraint conditions of the model parameter group as the search space, generate an initial population by initializing the objective function through specific improvements at the head, where the objective function is used to describe the optimization objective of the hydrological forecasting system, the constraint conditions are used to describe the constraints of the physical basin on the rainfall-runoff model, and the individuals in the population are used to describe the different parameter expressions of the rainfall-runoff model for the physical basin;

[0071] S23: Specifically, the objective function aims to minimize the hydrological forecast error, and the constraint conditions are composed of the underlying surface conditions, runoff generation conditions, and confluence conditions of the basin; for N floods, the calculation expression of the objective function is as follows:

[0072]

[0073] In the formula, t 0 is the initial time series of hydrological data, T is the total time series length of hydrological data, is the measured value of the nth flood at time t, is the forecast value of the rainfall-runoff model for the nth flood and at time t. For the multi-dimensional parameter group, is the upper limit of the model parameter group, is the lower limit of the model parameter group, and the constraint conditions are expressed as follows:

[0074]

[0075] In the case study, the value ranges of the rainfall-runoff model parameter group are shown in Table 1:

[0076] Table 1 Value ranges of the rainfall-runoff model parameter group

[0077]

[0078] Different from random generation, the population is initialized through specific improvements at the head to make individuals more conform to the nature of hydrological data. The calculation expression is as follows:

[0079]

[0080] In the formula, D is the number of multi-dimensional parameter groups, is the chaotic mapping parameter between 0 and 1. The parameter change coefficient can be determined as the chaotic mapping parameter by analyzing the relationship between model parameters; in the research case, the Logistic method is used to determine the chaotic mapping parameter, and the formula is , r takes 0.8, and the initial uses a random number.

[0081] S24: Based on the individual fitness values in the initial population, that is, the objective function values, the positions of individuals in the population are continuously iteratively updated through the aurora optimization algorithm with specific improvements at the head (the specific improvements at the head are reflected in S23). The calculation process is as Figure 2 shown, that is, the different parameter expressions of the rainfall-runoff model for the physical basin are updated, and the population after updating the individual positions is obtained until the forecast requirements are met or the maximum number of iterations is reached; specifically, the individual positions are updated through rotational motion, aurora elliptical trails, and particle collisions. The individual positions and The calculation expression is as follows:

[0082]

[0083]

[0084] In the formula, are four independent random numbers between 0 and 1, and are constants that control the degree of individual position update, determined by the number of iterations and the fitness value, is the coefficient that affects , is the coefficient that affects , is any individual in the population. In the research case, the population size of the head-specific improved aurora optimization algorithm is 50, and the maximum number of iterations is 10.

[0085] S25: Select the parameter combination that makes the objective function optimal during the iteration process, that is, enable the rainfall-runoff model to best describe the physical basin, and conduct hydrological forecasting experiments in the validation set on the condition of being close to the physical phenomena actually occurring in the basin; record the runoff forecasting results of each model run in detail, and further obtain the runoff forecasting results of the model on the test set after the experiment.

[0086] Furthermore, the specific steps of step S3 are as follows:

[0087] S31: The result set of the rainfall-runoff model includes forecast values, measured values, and model details. Separate the forecast values from the result set, analyze the temporal correlation of the forecast differences, and combine the measured values and the physical process model to determine the time lag of the forecast differences, and set the lead time according to the hydrological forecasting requirements, so as to construct a multi-dimensional multi-step forecast residual matrix. The calculation expression is as follows:

[0088]

[0089] In the formula, represents the forecast residual value at time t with a lead time of m and a previous lag period of s, represents the measured values from time to time t and from time t to represents the forecast values from time to time t and from time t to

[0090] S32: Before performing deep learning prediction, it is necessary to first calculate the cross-correlation of multi-dimensional hydrological element independent variables with single-dimensional or multi-dimensional dependent variables, analyze the average time-series cross-correlation, and then determine the time lag of hydrological elements based on the node with the maximum average cross-correlation. The lead time setting is the same as in S31. In this way, the S1 dataset is reorganized into a multi-dimensional and multi-step hydrological element matrix, and the calculation expression is as follows:

[0091]

[0092] In the formula, represents the value of the hydrological element matrix at time t with a lead time of m and a previous lag period of s2, represents from time to time t and from time t to the hydrological element values at time. The multi-dimensional and multi-step hydrological element matrix characterizes the actual physical process of the basin; in the research case, the cross-correlation coefficient is used to determine the lag term s2 of hydrological elements, and s2 = 10 is taken, and the lead time m = 4 is selected.

[0093] S33: Normalize the S31 forecast residual matrix and the S32 hydrological element matrix according to the following expression to avoid the influence of magnitude:

[0094]

[0095] In the formula, x t is the matrix data at time t, min(·) and max(·) represent the minimum and maximum functions respectively, and x represents the entire matrix data.

[0096] Furthermore, the specific steps of step S4 are as follows:

[0097] S41: Construct a deep learning prediction model and a deep learning post-processing model respectively. Using the range constraints of the two model hyperparameter groups as the search space, generate an initial population through the head-specific improved initialization objective function. Among them, the individuals in the former population are used to describe the non-linear expression of the prediction model for basin elements, and the individuals in the latter population are used to describe the non-linear expression of the post-processing model for forecast differences; in the research case, the long short-term memory network is selected for the deep learning model, and the range of the deep learning model hyperparameter group is shown in Table 2.

[0098] Table 2 Search Space of Deep Learning Model Hyperparameters

[0099]

[0100] S42: Specifically, the objective function of the deep learning prediction model aims to optimize the approximation of the actual runoff generation and concentration in the basin, and the objective function of the deep learning post - processing model aims to optimize the approximation of the prediction difference variation law of the physical process model. The calculation formulas of the two objective functions can be expressed as follows:

[0101]

[0102] In the formula, is expressed as the forecast residual at time t or the true value of the hydrological element, is expressed as the forecast residual at time t or the predicted value of the hydrological element;

[0103] S43: Input the hydrological element matrix and the forecast residual matrix into the constructed deep learning prediction model and deep learning post - processing model respectively. Continuously update the positions of individuals in the population through the head - specific improved aurora optimization algorithm (the detailed steps are the same as S2), that is, update the different parameter expressions of the deep learning model for the runoff generation and concentration process in the basin or the prediction difference variation law, and obtain the population after updating the individual positions until the forecast requirements are met or the maximum number of iterations is reached. In the research case, the population size of the head - improved aurora optimization algorithm is taken as 50, and the maximum number of iterations is taken as 10.

[0104] S44: Select the parameter combinations that make the objective function optimal during the iteration process of the two models respectively, that is, make the models best describe the basin elements or prediction differences. Conduct pre - test experiments in the validation set on the condition of approximating the basin elements or prediction difference variation law; Record in detail the runoff forecast results and residual prediction results of each run of the two models, and further obtain the prediction results of the two models on the test set after the experiment.

[0105] Furthermore, the specific steps of step S5 are as follows:

[0106] S51: Drive the deep learning post - processing model trained in S4 with the rainfall - runoff model calibrated in S2, and respond to the rainfall - runoff model in the form of feedback to form a dual - driven hydrological forecasting model based on physical processes and deep learning;

[0107] S52: Specifically, use the deep learning post - processing model to predict the forecast residuals with a time step of within different lead times. According to the specific moments of the forecast residuals within each lead time, respectively screen the original forecast results of the rainfall - runoff model to match the moments of the forecast residuals within different lead times, and then fuse the forecast residuals and the original forecast results at the same moment in the form of feedback for each lead time to obtain the runoff forecast results of the dual - driven correction model within different lead times. The runoff forecast result of the dual - driven correction model at time t and within lead time m is expressed as follows:

[0108]

[0109] In the formula, is the forecast residual at time t + m obtained through the deep learning post - processing model, is the forecast flow at time t + m obtained through the rainfall - runoff model, and respectively represent the calculation processes of the deep learning post - processing model and the rainfall - runoff model, is the forecast difference of the rainfall - runoff model at time t, are the hydrological elements required for the calculation of the rainfall - runoff model at time t + m, is the initial calculation condition of the rainfall - runoff model at time t + m.

[0110] Furthermore, the specific steps of step S6 are as follows:

[0111] S61: Split the runoff forecast result matrix output by the rainfall - runoff model and the dual - drive correction model according to the lead time to obtain multiple one - dimensional arrays, and each array represents a lead time; select a suitable normal distribution according to the characteristics of the hydrological data, perform normal distribution transformation on the multiple one - dimensional arrays respectively, obtain the initial value of the parameter from the artificially assumed initial prior distribution, and then further establish a probabilistic forecast model. In the research case, the initial prior distribution is assumed to be a semi - Cauchy distribution.

[0112] S62: Specifically, iteratively apply the Bayesian inference method in the multiple one - dimensional arrays, that is, the expression is , where is the posterior probability density of the parameter given the hydrological element sample information DS, is the likelihood function of the hydrological element sample information DS given the parameter , is the prior probability distribution of the parameter ;

[0113] S63: On the basis of S62, combined with the Markov chain Monte Carlo algorithm based on the no - U - turn sampler, the algorithm calculation process is as Figure 3 shown, perform posterior probability distribution sampling analysis on the normal distributions with different parameters, and the calculation expression is as follows:

[0114]

[0115] In the formula, is the logical function, which judges the true - false situation of the expression and assigns values. When the judgment is True, it assigns a value of 1, and when the judgment is False, it assigns a value of 0, and They are the direction vectors for left and right parameter sampling respectively. and are the left and right candidate parameters generated according to the current state respectively , is the algorithm identifier before update, which is used to judge whether the current sampling is over;

[0116] S64: Repeat S63, and take the average of a large number of sampling results as the reasonable parameters of each normal distribution, and further take the deterministic forecast as the mean value , construct the probability forecast normal distribution at each moment t, and each lead time has a unique normal distribution; according to the specific requirements of hydrological forecast applications, calculate and determine the forecast interval at each moment t at the specified confidence level, and obtain the probability forecast results of the rainfall-runoff model and the dual-driving calibration model at different lead times. In the research case, the maximum number of samplings is set to 5000 times, the number of samplings for adjusting the benchmark is set to 2500 times, and 4 are run in parallel.

[0117] Furthermore, the specific steps of step S7 are as follows:

[0118] S71: Reorganize the measured data into a multi-dimensional multi-step measured matrix with a lead time of m and a lag time of s (the same as step S32), and input it together with the S31 forecast residual matrix into the deep learning post-processing model, and record the Shapley value of each input variable at each moment within different lead times. The calculation expression of the Shapley value is as follows:

[0119]

[0120] In the formula, is the contribution of the i-th dimensional forecast residual matrix to the output of the dual-driving calibration model; k is the dimension of the forecast residual matrix; K is the forecast residual matrix; S is the forecast residual matrix before the i-th dimension in the sequence, which is a subset of K; is the value generated by the cooperation of the forecast residual matrices included in the subset S;

[0121] S72: Considering the differences in hydrological lead times, and combining the runoff generation mechanism and possible error sources of the rainfall-runoff model (including model parameter errors, model structure errors, measurement data errors, etc.), explain the impact of the errors generated by the rainfall-runoff model at different moments on the overall prediction of the dual-driving calibration model, and further understand the actual runoff generation process of the rainfall-runoff model under the calibrated parameters. In the research case, special attention should be paid to the forecast errors of the rainfall-runoff model at the current moment and the nearest three moments, and the constructed forecast errors of the rainfall-runoff model have strong transitivity in a short time.

[0122] Furthermore, the specific steps of step S8 are as follows:

[0123] S81: Construct an index system consisting of 7 quantitative evaluation indicators. The 3 deterministic evaluation indicators include Nash-Sutcliffe Efficiency (NSE), Root Mean Square Error (RMSE), and Mean Absolute Percentage Error (MAPE), and their expressions are as follows:

[0124]

[0125]

[0126]

[0127] In the formula, is the measured value of the hydrological element, is the predicted value of the hydrological model;

[0128] The 4 probabilistic evaluation indicators include Coverage Ratio (CR), Bandwidth (B), Continuous Ranked Probability Score (CRPS), and Probability Integral Transform (PIT), and their expressions are as follows:

[0129]

[0130]

[0131]

[0132]

[0133] In the formula, is a conditional function that takes 1 when the expression in the parentheses holds, and 0 otherwise, and are the lower and upper bounds of the prediction interval at time t, respectively, and are the probability density function and cumulative distribution function of the predicted value respectively;

[0134] The closer the NSE value is to 1, the higher the prediction accuracy of the model. The smaller the RMSE and MAPE values, the smaller the prediction error of the model. CR and B are used to evaluate the prediction interval of probabilistic prediction. The larger the CR value, the more real information the prediction result contains. The larger the B value, the larger the bandwidth of the interval, that is, the greater the uncertainty of the prediction. The confidence level of interval prediction is taken as 95%. CRPS is used to evaluate the overall performance of probabilistic hydrological prediction. The smaller the CRPS value, the better the comprehensive performance of probabilistic prediction. PIT is used to evaluate the reliability of probabilistic prediction and to test whether the normal distribution assumption of the model is reasonable. If the PIT value follows U(0, 1), it means that the probabilistic prediction result is reliable.

[0135] S82: According to the predicted and measured values of the dual-drive calibration model, a comprehensive comparative analysis and evaluation of the proposed dual-drive hydrological forecasting method are carried out from multiple dimensions, such as the bandwidth of the prediction interval, the coverage of the prediction interval, the comprehensive score of probability prediction, the reliability of probability prediction, the prediction accuracy, and the prediction error. In the research case, the evaluation results are shown in Table 3. The prediction effect of the dual-drive calibration model is better than that of the single rainfall-runoff model at different lead times, and the uncertainty brought is much smaller than that of the rainfall-runoff model, and the prediction results are more reliable.

[0136] Table 3 Evaluation index results of the dual-drive calibration model and the rainfall-runoff model in the test set floods

[0137]

[0138] According to the above examples, aiming at the problems of the complexity and uncertainty of the hydrological process in traditional hydrological forecasting methods, as well as the neglect of physical mechanisms by data-driven models, the present invention proposes a hydrological forecasting method based on the dual drive of physical processes and interpretable deep learning. The present invention uses the head-specific optimized aurora optimization algorithm to optimize the parameters of the dual-drive calibration model, which can not only improve the prediction accuracy, ensure the reliability and stability of the model, but also reduce the consumption of computing resources. The present invention constructs a probability prediction interval for each moment with deterministic prediction as the mean, supplements the uncertainty information of the model prediction, and increases the reliability of the dual-drive calibration model. The present invention combines the Shapley value and hydrological forecasting knowledge to conduct an interpretability analysis of the deep learning model, so that the importance of input variables to the prediction results can be clearly presented. This feature provides better decision-making support for users, especially in the field of hydrological forecasting that requires transparency. The present invention combines the established evaluation index system for comprehensive evaluation and comparison, ensuring the accuracy and practicability of the prediction results, thereby providing effective technical support for flood forecasting and flood risk analysis.

[0139] The above formulas are all dimensionless and take their numerical calculations. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula that is closest to the real situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0140] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product.

[0141] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application of the technical solution and the inventive constraints. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0142] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0143] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0144] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A hydrological forecasting method based on physical processes and interpretable deep learning, characterized by: The steps include: Based on the physical model, a basin hydrological model is constructed, basin characteristics are analyzed and an appropriate rainfall-runoff model is selected. The model parameters are calibrated and verified through the head-specific improved Aurora optimization algorithm to generate initial forecast results. A deep learning model was constructed, with the forecast residuals of the basin hydrological model as the training target, and the hyperparameters of the deep learning model were tuned through the head-specific improved Aurora optimization algorithm; Based on the initial forecast results of the basin hydrological model and the measured data, a multi-dimensional multi-step forecast residual matrix and a multi-dimensional multi-step hydrological element matrix are constructed. The matrix represents the difference between the actual physical process and the simulated physical process of the basin. The forecast residuals predicted by the deep learning post-processing model are synchronously corrected with the original forecast results of the basin hydrological model for multiple forecast periods to generate a dual-drive correction model; Generate real-time hydrological forecast results through the dual-drive correction model, and obtain multi-step correction forecast results in different forecast periods; Based on the correction and forecast results, the Shapley additive interpretation method is used to perform importance analysis on the input variables of the deep learning post-processing model, quantify the contribution of each input variable to the correction results, and optimize the design of the model input variables and the correction mechanism; The initial forecast results of the basin hydrological model are calibrated synchronously for multiple forecast periods, including: Runoff forecast results of the dual-drive correction model at time t and within the forecast period m The expression is as follows: In the formula, is the forecast residual at time t + m obtained by the deep learning post-processing model, is the forecast flow at time t + m obtained by the rainfall-runoff model, and They represent the calculation process of the deep learning post-processing model and the rainfall-runoff model respectively. is the forecast difference of rainfall-runoff model at time t, The hydrological elements required for the rainfall-runoff model calculation at time t + m, is the initial calculation condition of the rainfall-runoff model at time t + m.

2. The hydrological forecasting method based on physical process and interpretable deep learning dual drive according to claim 1 is characterized in that: The construction of the multi-dimensional multi-step prediction residual matrix includes: The forecast values ​​are separated from the result set of the basin hydrological model, and the time series correlation of the forecast differences is analyzed in combination with the measured values; Based on the time series lag factor and the forecast period setting, a multi-dimensional multi-step forecast residual matrix is ​​generated: ; In the formula, It represents the forecast residual value at time t with a forecast period of m and a forward lag period of s. express time to time t and time t to The measured value at time, express time to time t and time t to The forecast value at time.

3. The hydrological forecasting method based on physical process and interpretable deep learning dual drive according to claim 1 is characterized in that: The construction of the physical model includes the following steps: The soil layer was divided into three layers and the evapotranspiration was calculated separately; Calculate runoff based on rainfall runoff conditions and divide runoff into surface runoff, subsoil runoff and underground runoff; The Muskingum method is used to calculate runoff and generate hourly flow forecast results for the basin.

4. The hydrological forecasting method based on physical process and interpretable deep learning dual drive according to claim 3 is characterized in that: The construction of the multi-dimensional and multi-step hydrological element matrix includes: Select key hydrological elements of the watershed, including rainfall, rainfall runoff, soil moisture, and evapotranspiration; Based on the time series mutual correlation and hysteresis characteristics, a multi-dimensional and multi-step hydrological element matrix is ​​generated, and the matrix is ​​expressed as: ; In the formula, It represents the hydrological element matrix value with a forecast period of m at time t and a front lag period of s2. express time to time t and time t to The value of the hydrological element at the time.

5. The hydrological forecasting method based on physical process and interpretable deep learning dual drive according to claim 4 is characterized in that: The specific steps to tune the hyperparameters of deep learning models using the head-specific improved Aurora optimization algorithm include: An initial population is generated based on a chaos mapping method, wherein the chaos mapping parameters are determined by analyzing the variation characteristics of the model parameters; In the process of population individual renewal, the rotational motion, auroral elliptical trail and particle collision mechanism are introduced to optimize the objective function; Continue to iterate and update the population position until the objective function reaches the optimal value or the maximum number of iterations is reached.

6. The hydrological forecasting method based on physical process and interpretable deep learning dual drive according to claim 1 is characterized in that: The uncertainty of the dual-drive correction model is quantified using Bayesian inference and Markov Monte Carlo algorithms, including: The output results of the dual-drive correction model are transformed into a normal distribution and an initial prior distribution is generated; Combined with observed data, the posterior distribution is iteratively updated based on the Bayesian inference method to quantify the uncertainty of the model prediction results; The probability forecast interval for each moment is generated through posterior distribution sampling analysis.

7. The hydrological forecasting method based on physical process and interpretable deep learning dual drive according to claim 1 is characterized in that: The interpretability analysis of the deep learning post-processing model is performed using the Shapley additive interpretation method, which includes the following steps: The Shapley values ​​of the forecast residual matrix and the hydrological element matrix for the model output are calculated based on the Shapley additive interpretation method; Quantify the contribution of each input variable to the model output; Combined with the flow generation mechanism and error sources of the basin hydrological model, the impact of input variables on the overall prediction results of the dual-drive model is explained.

8. The hydrological forecasting method based on physical process and interpretable deep learning dual drive according to any one of claims 1 to 7, characterized in that: The dual-driven fusion mechanism of deep learning post-processing model and watershed hydrological model is applicable to different types of watershed hydrological models, including but not limited to Xinanjiang model, SCS-CN model and TOPMODEL model.

Citation Information

Patent Citations

  • Multi-step daily runoff forecasting method based on meteorological information and deep learning algorithm

    CN113255986A

  • Drainage basin water, wind and light resource integrated forecasting method based on physical-data dual drive

    CN116108956A