Basin runoff simulation method fusing random Xinanjiang model and machine learning
By introducing random disturbance terms and wavelet transform processing into the Xin'anjiang model with three water sources, and combining it with a machine learning model, the problem that existing models cannot capture random disturbances under extreme weather conditions is solved, thus improving the accuracy of runoff simulation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-26
- Publication Date
- 2026-03-27
AI Technical Summary
Existing runoff simulation models cannot capture random disturbances within the system under extreme weather conditions. Traditional physical models lack uncertainty quantification, while data-driven models lack physical constraints, and hybrid models fail to effectively integrate random information.
In the Xin'anjiang River model with three water sources, a perturbation term with random distribution characteristics is introduced and transformed into a stochastic differential equation. Statistical features are extracted by combining Monte Carlo simulation and discrete wavelet transform, and then input into a machine learning model to improve simulation accuracy.
By introducing random perturbation terms and wavelet transform processing, the accuracy of low and high flow rate simulations of watershed runoff was improved, and the predictive ability of the model under extreme conditions was enhanced.
Smart Images

Figure CN121744947A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a watershed runoff simulation method that integrates the stochastic Xin'anjiang model with machine learning, belonging to the field of watershed runoff simulation technology. Background Technology
[0002] Runoff simulation is a core issue in hydrological research, and it is of great significance for water resource management, flood control and disaster reduction, and ecological protection. Currently, rainfall-runoff simulation methods are mainly divided into two categories: physical mechanism models and data-driven models. Physical mechanism models, such as the conceptual hydrological model represented by the Xin'anjiang Three-Source Model, have been widely used in humid and semi-humid regions of China. These models are based on water balance equations and runoff generation and confluence mechanisms, possessing clear physical meaning and interpretability. However, the traditional Xin'anjiang model is usually a deterministic model, with both its inputs and outputs being deterministic values, making it difficult to quantify the inherent uncertainties and random disturbances in parameter estimation and observational data. Data-driven models, such as Gated Cyclic Unit (GRU) and Extreme Gradient Boosting (XGBoost) models, are widely used in runoff simulation. These models have strong nonlinear fitting capabilities, but they lack physical constraints and are highly dependent on training data. To combine the advantages of these two types of models, researchers have proposed hybrid modeling techniques, which have improved simulation accuracy to some extent and have become a research hotspot in the field of runoff simulation in recent years.
[0003] However, most existing hybrid models are based on deterministic physical models, which ignore the random fluctuations in rainfall-runoff processes. Especially under extreme weather conditions, deterministic models cannot capture the random disturbances within the system. Although the academic community has developed stochastic hydrological models based on stochastic differential equations, few studies have effectively integrated the rich statistical information generated by stochastic hydrological models into the framework of machine learning. Summary of the Invention
[0004] The purpose of this invention is to overcome the technical defects of the existing technology and solve the above-mentioned technical problems. It proposes a watershed runoff simulation method based on the coupling of a stochastic three-source Xin'anjiang model and machine learning. The method introduces a perturbation term with random distribution characteristics into the state variable soil moisture content of the deterministic three-source Xin'anjiang model, transforming it into a stochastic differential equation and outputting random trajectory information. After information purification by discrete wavelet transform, the information is input into the machine learning model, enabling the machine learning model to more accurately capture multi-scale hydrological signals. This improves the model's runoff simulation effect by increasing the accuracy of low and high flow simulations in the watershed.
[0005] The present invention specifically adopts the following technical solution: a watershed runoff simulation method that integrates the stochastic Xin'anjiang model and machine learning, comprising the following steps: Step SS1: Collect hydrological and meteorological data within the study basin, including measured runoff Q, rainfall P, and potential evapotranspiration E, and divide them into periodic and validation periods according to a 7:3 ratio; Step SS2: Construct the Xin'anjiang River model with three water sources, retain the original runoff generation and confluence mechanism of the Xin'anjiang River model with three water sources, and introduce a perturbation term with random distribution characteristics into the state update equation of soil water storage to transform it into a stochastic differential equation. The perturbation term includes, but is not limited to, Lévy stable distribution noise, Gaussian white noise and Poisson noise, so as to transform the deterministic physical model into a stochastic model that can output probability distribution. Step SS3: Use Monte Carlo simulation to transform the single output of the stochastic model into a set of trajectories containing probability distribution information, and perform statistical analysis on the generated trajectory set at each time step to extract key statistical features, including mean trajectory, quantile trajectory and range trajectory. Step SS4: Perform multi-scale analysis on the statistical feature sequence based on discrete wavelet transform, decompose the mean, quantile and range time series, decompose the one-dimensional time series into low-frequency approximation coefficients and high-frequency detail coefficients, flatten and splice the wavelet coefficients of all levels obtained by decomposition, and construct a high-dimensional feature vector containing multi-scale information. Step SS5: Construct two machine learning models with different architectures, XGBoost and Gated Recurrent Unit (GRU), and design two input methods: Input method 1 uses only the mean trajectory and wavelet coefficients of the random model; Input method 2 uses the mean, range, quantile trajectory and all wavelet coefficients; then use different information combinations to simulate watershed runoff.
[0006] In a preferred embodiment, step SS2 specifically includes: Taking Lévy stable distribution noise as an example, specifically, Lévy stable distribution noise is introduced into the state variable soil water storage, and the discrete stochastic state update equation is constructed as follows: ; In the formula, W t and W t+1 The watershed soil water storage at time t and time t+1 are respectively; WM is the average tension water storage capacity of the watershed; P t Let t be the rainfall at time t; E t For the first t Potential evaporation at any moment; R t η is the production flow rate at time t; σ is the noise intensity coefficient; η is the noise level coefficient. t It is random noise that follows a Lévy stable distribution.
[0007] In a preferred embodiment, step SS2 further includes: Lévy stable distribution noise η t ~S(α,β,γ,δ) is controlled by the characteristic exponent α and the skewness parameter β, and its characteristic function ψ(t) is defined as: ; In the formula, 0 < α ≤ 2 controls the tail thickness of the distribution; -1 ≤ β ≤ 1 controls the asymmetry of the distribution; γ > 0 is the scale parameter; δ is the position parameter; i is the imaginary unit; and sgn(t) is the sign function.
[0008] In a preferred embodiment, step SS3 specifically includes: Given a total number of trajectories N in the Monte Carlo simulation, at the current time t, obtain a set of N simulated flow rates. Calculate the mean locus, quantile locus, and range locus: (1) Mean locus for: ; (2) Quantile locus for: Given a confidence level τ, calculate the set value at that quantile: ; Obtain the lower boundary trajectory and upper boundary trajectory ; (3) Range trajectory for: .
[0009] In a preferred embodiment, step SS4 specifically includes: Using specific wavelet basis functions ψ(t) and scaling functions φ(t), an arbitrary statistical sequence x(t) is decomposed into J-levels, and the approximation coefficients cA of the j-th level are obtained. j and detail coefficient cD j The calculation formula is as follows: ; ; In the formula, x [n] represents the nth sample value of the statistical sequence; [k] represents the position index of the wavelet coefficients, indicating the coefficients at different positions in the j-th level decomposition; h[n] represents the low-pass filter coefficients; g[n] represents the high-pass filter coefficients; and the constructed high-dimensional feature vector V feature For the concatenation of all decomposition coefficients: ; in, .
[0010] In a preferred embodiment, step SS5 specifically includes: The extreme gradient boosting XGBoost model employs a residual learning strategy, using measured runoff at time t. With the mean trajectory of the random model The difference is taken as the target variable: ; The final prediction result is: ; In the formula, To study the measured runoff of the watershed at time t; Let be the value of the mean runoff trajectory at time t; The residual between the measured runoff and the mean runoff trajectory at time t; Let be the residual prediction value of the machine learning model at time t; For extreme gradient boosting The final simulated runoff results obtained from the model.
[0011] In a preferred embodiment, step SS5 further includes: The gated cyclic unit (GRU) model employs a direct mapping strategy, using the measured runoff at time t. As the target variable, the final predicted runoff result is directly output by the model. .
[0012] In a preferred embodiment, step SS5 further includes: Input method 1 is: ; The input vector X1 contains only the mean trajectory of the output of the stochastic model and its corresponding wavelet coefficients at all levels.
[0013] In a preferred embodiment, step SS5 further includes: Input method 2 is as follows: The input vector X2 contains the mean trajectory, lower boundary trajectory, upper boundary trajectory, range trajectory, and wavelet coefficients corresponding to all the above trajectories output by the random model. ; In the formula, WT(Q) mean ), WT(Q min ), WT(Q max ) and WT(Q range ) represents the set of coefficients after wavelet transform.
[0014] The beneficial effects achieved by this invention are as follows: This invention addresses how to solve the problem of traditional physical mechanism models, such as the Xin'anjiang Three-Source Watershed model. Although these models have clear physical meaning and interpretability, they are usually deterministic models, with both inputs and outputs being deterministic values. It is difficult to quantify the inherent uncertainties and random disturbances in parameter estimation and observation data. Although data-driven models have strong nonlinear fitting capabilities, they lack physical constraints and are highly dependent on training data. Most existing hybrid models are based on deterministic physical models, which ignore the random fluctuations in rainfall-runoff processes. Especially under extreme weather conditions, deterministic models cannot capture the technical requirements of random disturbances within the system. By introducing a disturbance term with random distribution characteristics into the state variables of the Xin'anjiang Three-Source Watershed model, the model is randomized. Monte Carlo simulation is used to generate random trajectories to extract statistical features. Discrete wavelet transform is used to remove redundant noise and purify information. The resulting high-quality information is input into a machine learning model, enabling it to more accurately capture multi-scale hydrological signals. This improves the model's runoff simulation effect by increasing the accuracy of low and high flow simulations in the watershed. Attached Figure Description
[0015] Figure 1 This is a flowchart of the watershed runoff simulation method that integrates the stochastic Xin'anjiang model and machine learning according to the present invention. Detailed Implementation
[0016] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.
[0017] This invention randomizes the deterministic model by introducing a perturbation term with random distribution characteristics into the state variable soil moisture content of the three-source Xin'anjiang model. Discrete wavelet processing is then applied to the statistical characteristics output by the randomized three-source Xin'anjiang model to improve data quality, which is then used as input to two machine learning models with different architectures. The resulting stochastic hybrid model provides the final runoff simulation results for the studied watershed, such as... Figure 1 As shown, the technical solution of the present invention specifically includes the following steps: Step 1: Collect hydrological and meteorological data within the study basin, including measured runoff Q, rainfall P, and potential evapotranspiration E, and divide them into periodic and validation periods according to a 7:3 ratio.
[0018] (1) Based on the hydrological and meteorological data in the research basin, a deterministic three-source Xin'anjiang model is established, which includes four modules: evapotranspiration, runoff generation, water source distribution and runoff confluence.
[0019] (2) The hydrological and meteorological data in the study basin were divided into calibration period and validation period according to 7:3. The differential evolution algorithm was used as the parameter calibration method to select the optimal parameters. The performance of the constructed deterministic three-source Xin'anjiang model was evaluated by Nash Efficiency Coefficient (NSE), Kling-Gupta Efficiency (KGE) and Percent Bias (PBIAS).
[0020] Step 2: Construct a stochastic three-source Xin'anjiang model, retaining its original runoff generation and confluence mechanism. Introduce a perturbation term with random distribution characteristics into the state update equation of soil water storage, transforming it into a stochastic differential equation form. This stochastic perturbation term includes, but is not limited to, Lévy stable distribution noise, Gaussian white noise, and Poisson noise. This invention takes Lévy stable distribution noise as an example to realize the transformation of a deterministic physical model into a stochastic model that can output a probability distribution.
[0021] (1) Based on the Xin'anjiang model architecture with three water sources, Lévy stable distribution noise is introduced into the state variable soil water storage, and the discrete stochastic state update equation is constructed as follows: ; In the formula, W t and W t+1 The watershed soil water storage at time t and time t+1 are respectively; WM is the average tension water storage capacity of the watershed; P t Let t be the rainfall at time t; E t For the first t Potential evaporation at any moment; R t η is the production flow rate at time t; σ is the noise intensity coefficient; η is the noise level coefficient. t It is random noise that follows a Lévy stable distribution.
[0022] Lévy stable distribution noise η t ~S(α,β,γ,δ) is controlled by the characteristic exponent α and the skewness parameter β, and its characteristic function ψ(t) is defined as: ; In the formula, 0 < α ≤ 2 controls the tail thickness of the distribution; -1 ≤ β ≤ 1 controls the asymmetry of the distribution; γ > 0 is the scale parameter; δ is the position parameter; i is the imaginary unit; and sgn(t) is the sign function.
[0023] (2) The traditional parameters, Lévy stable distribution noise characteristic index α, skewness parameter β and noise intensity σ in the Xin'anjiang model are used as parameters to be calibrated. The differential evolution algorithm is used to optimize the parameters. Multiple random trajectories are selected and the average value is taken in each calculation. The optimal Kling-Gupta efficiency coefficient (KGE) is used as the target, so that the probability distribution interval of the random model can effectively capture the random fluctuations of the measured runoff, realize the quantitative analysis of uncertainty, and ensure that the model generates high-quality statistical features.
[0024] Step 3: Use Monte Carlo simulation to transform the single output of the stochastic model into a set of trajectories containing probability distribution information, and perform statistical analysis on the generated trajectory set at each time step to extract key statistical features, such as mean trajectory, quantile trajectory and range trajectory.
[0025] (1) Using the optimal parameters calibrated in step 2, input the watershed hydrological and meteorological data, perform Monte Carlo simulation, and generate multiple sets of independent random flow trajectories. The method is as follows: Given a total number of trajectories N in the Monte Carlo simulation, at the current time t, obtain a set of N simulated flow rates. ; (2) Calculate the statistical characteristics based on the set of random flow trajectories obtained from the Monte Carlo simulation. The specific calculation method is as follows: Mean locus : ; quantile trajectory : Set a confidence level τ (2.5% and 97.5% are used in this invention), and calculate the value of the set at that quantile: ; Obtain the lower boundary trajectory and upper boundary trajectory .
[0026] Range Trajectory : .
[0027] Step 4: Perform multi-scale analysis on the statistical feature sequence based on discrete wavelet transform, decompose the time series such as mean, quantile and range, decompose the one-dimensional time series into low-frequency approximation coefficients and high-frequency detail coefficients, flatten and splice the wavelet coefficients of all levels obtained by decomposition, and construct a high-dimensional feature vector containing multi-scale information.
[0028] (1) Since the optimal wavelet is different for machine learning models with different architectures, grid search is used to select the optimal wavelet basis and hyperparameters suitable for machine learning models with different architectures.
[0029] (2) Based on the selected optimal wavelet basis, the statistical feature sequence is analyzed by multiple scales, decomposed into approximation coefficients and detail coefficients, and a high-dimensional feature vector containing multi-scale information is constructed.
[0030] The specific implementation process of discrete wavelet transform is as follows: Using specific wavelet basis functions ψ(t) and scaling functions φ(t), an arbitrary statistical sequence x(t) is decomposed into J-levels, where the j-th level... Approximation coefficient cA j and detail coefficient cD j The calculation formula is as follows: ; ; In the formula, x [n] represents the nth sample value of the statistical sequence; [k] represents the position index of the wavelet coefficients, indicating the coefficients at different positions in the j-th level decomposition; h[n] represents the low-pass filter coefficients; g[n] represents the high-pass filter coefficients; and the constructed high-dimensional feature vector V feature For the concatenation of all decomposition coefficients: ; Step 5: Construct two machine learning models with different architectures: Extreme Gradient Boosting (XGBoost) and Gated Recurrent Unit (GRU), and design two input methods: 1. Using only the mean trajectory and its wavelet coefficients of the stochastic model; 2. Using the mean, range, quantile trajectories and all their wavelet coefficients. Simulate watershed runoff using different information combinations.
[0031] (1) Construct an extreme gradient boosting XGBoost model using a residual learning strategy; construct a gated recurrent unit (GRU) model using a direct mapping strategy. The specific approach is as follows: For the extreme gradient boosting XGBoost model, a residual learning strategy is adopted, using the measured runoff Q at time t. obs(t) With the mean trajectory Q of the random model mean(t) The difference is taken as the target variable: ; The final prediction result is: ; In the formula, To study the measured runoff of the watershed at time t; Let be the value of the mean runoff trajectory at time t; The residual between the measured runoff and the mean runoff trajectory at time t; Let be the residual prediction value of the machine learning model at time t; For extreme gradient boosting The final simulated runoff results obtained from the model.
[0032] For the gated cyclic unit (GRU) model, a direct mapping strategy is adopted, using the measured runoff at time t. As the target variable, a gated recurrent unit is used to capture the nonlinear temporal relationship between the input feature vector and runoff, and the final prediction result is directly output by the model. .
[0033] (2) Two input methods were used: one using only the mean trajectory and its wavelet coefficients of the stochastic model, and the other using the mean, range, quantile trajectory and all its wavelet coefficients. Two different machine learning models with different architectures were input to obtain the final stochastic hybrid model runoff simulation results. Model evaluation indices for low flow, high flow, and overall performance were calculated. The model evaluation indices used were the Nash Efficiency Coefficient (NSE), Kling-Gupta Efficiency (KGE), and the coefficient of determination (R²). 2 The input method and the root mean square error (RMSE) are discussed. The specific design methods for the two input methods are as follows: Input Method 1: The input vector X1 contains only the mean trajectory of the random model output and its corresponding wavelet coefficients at all levels. ; Input Method 2: The input vector X2 contains the mean trajectory, lower boundary trajectory, upper boundary trajectory, range trajectory, and wavelet coefficients corresponding to all the above trajectories output by the stochastic model. ; In the formula WT(Q mean ), WT(Q min ), WT(Q max ) and WT(Q range ) represents the set of coefficients after wavelet transform.
[0034] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A watershed runoff simulation method integrating the stochastic Xin'anjiang model and machine learning, characterized in that, Includes the following steps: Step SS1: Collect hydrological and meteorological data within the study basin, including measured runoff Q, rainfall P, and potential evapotranspiration E, and divide them into periodic and validation periods according to a 7:3 ratio; Step SS2: Construct the Xin'anjiang River model with three water sources, retain the original runoff generation and confluence mechanism of the Xin'anjiang River model with three water sources, and introduce a perturbation term with random distribution characteristics into the state update equation of soil water storage to transform it into a stochastic differential equation. The perturbation term includes, but is not limited to, Lévy stable distribution noise, Gaussian white noise and Poisson noise, so as to transform the deterministic physical model into a stochastic model that can output probability distribution. Step SS3: Use Monte Carlo simulation to transform the single output of the stochastic model into a set of trajectories containing probability distribution information, and perform statistical analysis on the generated trajectory set at each time step to extract key statistical features, including mean trajectory, quantile trajectory and range trajectory. Step SS4: Perform multi-scale analysis on the statistical feature sequence based on discrete wavelet transform, decompose the mean, quantile and range time series, decompose the one-dimensional time series into low-frequency approximation coefficients and high-frequency detail coefficients, flatten and splice the wavelet coefficients of all levels obtained by decomposition, and construct a high-dimensional feature vector containing multi-scale information. Step SS5: Construct two machine learning models with different architectures, XGBoost and Gated Recurrent Unit (GRU), and design two input methods: Input method 1 uses only the mean trajectory and wavelet coefficients of the random model; Input method 2 uses the mean, range, quantile trajectory and all wavelet coefficients; then use different information combinations to simulate watershed runoff.
2. The watershed runoff simulation method integrating the stochastic Xin'anjiang model and machine learning as described in claim 1, characterized in that, Step SS2 specifically includes: Taking Lévy stable distribution noise as an example, specifically, Lévy stable distribution noise is introduced into the state variable soil water storage, and the discrete stochastic state update equation is constructed as follows: ; In the formula, W t and W t+1 The watershed soil water storage at time t and time t+1 are respectively; WM is the average tension water storage capacity of the watershed; P t Let t be the rainfall at time t; E t For the first t Potential evaporation at any moment; R t η is the production flow rate at time t; σ is the noise intensity coefficient; η is the noise level coefficient. t It is random noise that follows a Lévy stable distribution.
3. The watershed runoff simulation method integrating the stochastic Xin'anjiang model and machine learning as described in claim 2, characterized in that, Step SS2 specifically also includes: Lévy stable distribution noise η t ~S(α,β,γ,δ) is controlled by the characteristic exponent α and the skewness parameter β, and its characteristic function ψ(t) is defined as: ; In the formula, 0 < α ≤ 2 controls the tail thickness of the distribution; -1 ≤ β ≤ 1 controls the asymmetry of the distribution; γ > 0 is the scale parameter; δ is the position parameter; i is the imaginary unit; and sgn(t) is the sign function.
4. The watershed runoff simulation method integrating the stochastic Xin'anjiang model and machine learning as described in claim 1, characterized in that, Step SS3 specifically includes: Given a total number of trajectories N in the Monte Carlo simulation, at the current time t, obtain a set of N simulated flow rates. Calculate the mean locus, quantile locus, and range locus: (1) Mean locus for: ; (2) Quantile locus for: Given a confidence level τ, calculate the set value at that quantile: ; Obtain the lower boundary trajectory and upper boundary trajectory ; (3) Range trajectory for: 。 5. The watershed runoff simulation method integrating the stochastic Xin'anjiang model and machine learning according to claim 1, characterized in that, Step SS4 specifically includes: Using specific wavelet basis functions ψ(t) and scaling functions φ(t), an arbitrary statistical sequence x(t) is decomposed into J-levels, and the approximation coefficients cA of the j-th level are obtained. j and detail coefficient cD j The calculation formula is as follows: ; ; In the formula, x [n] represents the nth sample value of the statistical sequence; [k] represents the position index of the wavelet coefficients, indicating the coefficients at different positions in the j-th level decomposition; h[n] represents the low-pass filter coefficients; g[n] represents the high-pass filter coefficients; and the constructed high-dimensional feature vector V feature For the concatenation of all decomposition coefficients: ; in, .
6. The watershed runoff simulation method integrating the stochastic Xin'anjiang model and machine learning according to claim 1, characterized in that, Step SS5 specifically includes: The extreme gradient boosting XGBoost model employs a residual learning strategy, using measured runoff at time t. With the mean trajectory of the random model The difference is taken as the target variable: ; The final prediction result is: ; In the formula, To study the measured runoff of the watershed at time t; Let be the value of the mean runoff trajectory at time t; The residual between the measured runoff and the mean runoff trajectory at time t; Let be the residual prediction value of the machine learning model at time t; For extreme gradient boosting The final simulated runoff results obtained from the model.
7. The watershed runoff simulation method integrating the stochastic Xin'anjiang model and machine learning according to claim 6, characterized in that, Step SS5 specifically also includes: The gated cyclic unit (GRU) model employs a direct mapping strategy, using the measured runoff at time t. As the target variable, the final predicted runoff result is directly output by the model. .
8. The watershed runoff simulation method integrating the stochastic Xin'anjiang model and machine learning as described in claim 6, characterized in that, Step SS5 specifically also includes: Input method 1 is: ; The input vector X1 contains only the mean trajectory of the output of the stochastic model and its corresponding wavelet coefficients at all levels.
9. The watershed runoff simulation method integrating the stochastic Xin'anjiang model and machine learning according to claim 6, characterized in that, Step SS5 specifically also includes: Input method 2 is as follows: The input vector X2 contains the mean trajectory, lower boundary trajectory, upper boundary trajectory, range trajectory, and wavelet coefficients corresponding to all the above trajectories output by the random model. ; In the formula, WT(Q) mean ), WT(Q min ), WT(Q max ) and WT(Q range ) represents the set of coefficients after wavelet transform.
Citation Information
Patent Citations
Xin'an River model parameter optimization method based on Monte-Carlo algorithm
CN106446388A
Basin runoff forecasting method and device and storage medium
CN116384538A
Hydrological ensemble forecasting method and system based on space-time adaptive generative adversarial-random differential disturbance and medium
CN119989278A
Xinanjiang model correction method based on CLDAS soil humidity assimilation
CN121522775A
Big data-based hydrologic forecasting method
WO2022032872A1