Hydrological model parameter attribution method based on deep learning interpretability technology

Through the deep learning interpretability technology, the challenge of parameter setting and data acquisition of deep learning hydrological model is solved, the interpretability and accuracy of the model is improved, and the parameter basis for physical practical significance is provided to assist water resource management and environmental protection.

CN120494304APending Publication Date: 2025-08-15NANJING HYDRAULIC RES INST

Patent Information

Application Number
CN202510984185.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing deep learning hydrological model has challenges in parameter setting and data acquisition, and the lack of interpretability technology research has led to unclear physical mechanisms of the model, affecting its application effect in the fields of water resource management and environmental protection.

Method used

The parameter attribution method of hydrological model based on deep learning interpretability technology is adopted. By constructing a distributed hydrological model, the parameter rate determination is performed using an ant colony algorithm, the mapping relationship between meteorological factors and lower surface features is learned in combination with the deep residual network, and the parameter attribution analysis is performed using the feature importance method to clarify the degree of influence of each factor.

Benefits of technology

It improves the interpretability and accuracy of the hydrological model, weakens the uncertainty of the black box model, provides parameter basis for physical practical significance, and provides more accurate prediction tools for water resource management and environmental protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494304A_ABST
    Figure CN120494304A_ABST
Patent Text Reader

Abstract

The invention discloses a hydrological model parameter attribution method based on a deep learning interpretability technology. The hydrological model parameter attribution method comprises the following steps: collecting data; constructing a distributed hydrological model, performing optimization parameter calibration by adopting an ant colony algorithm, and selecting three indexes to construct a threshold constraint function so as to perform model precision screening; using a deep residual network model to learn a correlation between the parameters of the region with good precision and meteorological factors and underlying surface feature factors, and complementing parameter distribution on the region lacking data or the region with unqualified hydrological simulation precision; a machine learning feature importance interpretability technology is applied, feature factors are selected to calculate arrangement importance, and finally an attribution analysis result is output. The intervention of the interpretability technology is beneficial to weakening the uncertainty influence of original machine learning'black box ', and well explains the decision process of the hydrological model, so that model parameters have more physical practical significance, and a more accurate prediction tool is provided for the fields of water resource management, environmental protection and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of hydrological models, and specifically relates to a hydrological model parameter attribution method based on deep learning interpretability technology. Background Art

[0002] Hydrological models, as crucial tools for understanding and predicting hydrological processes, are widely used in water resources management, climate change research, and natural disaster warning. The rise of deep learning technology has enabled significant breakthroughs in parameter setting and data acquisition challenges in traditional hydrological models, offering new approaches to addressing these challenges.

[0003] The deep learning process includes model understanding and model interpretation. Model interpretation effectively improves model credibility and provides transparency in examining prediction results. Understanding why a machine learning model makes a certain decision and which features play the most important role in the decision also allows us to determine whether the model conforms to common sense. Deep learning networks are often considered black box models, with unclear physical mechanisms and poor interpretability.

[0004] As deep learning is increasingly used in various fields, researchers are not only satisfied with the ideal effects of the model, but also have more in-depth thoughts on the fundamental reasons for the model's effects. More and more research focuses on the interpretability analysis of deep learning models. The main methods of interpretability technology currently include the locally interpretable agnostic model LIME, the SHAP model, the feature importance method, etc.

[0005] Among them, the feature importance method is a method for studying the identification of deep learning input factors. It is a commonly used deep learning model interpretability technique. It checks deep learning network models by permuting feature importance. It has fast calculation speed and can maintain the consistency of feature importance and attributes. For example, the paper: ALTMANNA, TOLOŞIL, SANDERO et al., Permutation importance: a corrected feature importance measure [J]. Bioinformatics, 2010, 26(10): 1340-1347, by changing the arrangement of a column of data in a data table while keeping the other features unchanged, to see how much impact it has on prediction accuracy.

[0006] At present, most of the research on hydrological models in the field of deep learning focuses on improving prediction results, such as improving prediction accuracy by optimizing parameters or improving the model structure framework, but there is less research on deep learning interpretability technology. Paper: Tian Ye et al. Performance and interpretability of LSTM variant model in runoff prediction [J]. Water Resources Protection (3): 188-194. It describes the use of permutation importance method and integral gradient method to explore the interpretability of LSTM model for basin runoff prediction. Huang Qiang et al. Combined forecast of runoff in the source area of the Yellow River based on interpretable machine learning [J]. People's Yellow River (9): 50-59. Based on the SHAP deep learning interpretability analysis framework, the contribution of Yellow River source prediction factors to runoff forecast results was identified. However, the research on interpretability methods of deep residual network ResNet is still blank and needs further research. Summary of the Invention

[0007] The purpose of this invention is to provide a hydrological model parameter attribution method based on deep learning interpretability technology, which can accurately analyze the impact of meteorological factors and underlying surface characteristics on hydrological model parameters, clarify the importance and action mechanism of each factor, and provide a reliable basis for the optimization and practical application of hydrological models.

[0008] In order to solve the above technical problems, the present invention adopts the following solution: The hydrological model parameter attribution method based on deep learning interpretability technology includes the following steps: Step S1: Data collection and preprocessing: Construct a grid-scale hydrological dataset for the study area, which includes meteorological data, runoff depth data, and underlying surface watershed characteristic values; for areas with sparse watershed observation stations, additional stations are added to ensure the spatiotemporal representativeness of the hydrological sequence.

[0009] Step S2: Distributed hydrological model construction and parameter calibration: Select a distributed hydrological model suitable for the study area and divide the model into a warm-up period, a calibration period, and a verification period. Use the ant colony algorithm to optimize the parameters of the distributed hydrological model, construct a parameter calibration model with the Nash benefit coefficient (NSE) as the objective function, and dynamically adjust the parameter search path through pheromone concentration and heuristic functions. Construct an accuracy constraint function based on the percentage deviation (PBIAS), the Nash benefit coefficient (NSE), and the correlation coefficient (R²) to screen out high-precision simulation areas.

[0010] Step S3: Parameter completion based on deep residual network: Using the parameters of the high-precision area in step S2 as labels, the deep residual network model is trained to learn the mapping relationship between parameters and meteorological factors and underlying surface characteristics; the mean square error is used as the loss function to optimize the weights of the deep residual network model to achieve parameter spatial distribution completion in areas with missing data or low-precision simulation areas.

[0011] Step S4: Parameter attribution analysis based on characteristic importance: Use the permutation importance method to calculate the characteristic factor weight coefficient, clarify the degree of influence of different characteristic factors on the hydrological model parameters, and distinguish the positive and negative correlations, so as to achieve attribution analysis of the hydrological model parameters and provide a basis for in-depth understanding of the hydrological process, optimization model and related decision-making.

[0012] Further optimization, in step S1, the grid-scale hydrological dataset must meet temporal continuity and spatial consistency; wherein, meteorological data include precipitation, temperature, wind speed, potential evapotranspiration, runoff and atmospheric pressure. Runoff depth data comes from a high-resolution grid runoff depth dataset based on long-term observations, collected from the Global Runoff DataBase, and interpolated to the grid scale using the Thiessen polygon method. This database, supported by the World Meteorological Organization, has been collected over 30 years of hydrological observation data. It includes runoff data from 10,374 observation stations in 159 countries and regions around the world, with an average record length of 45 years. All runoff observation data from all stations are publicly available.

[0013] The underlying surface watershed characteristic values include digital elevation model (DEM), land use type, vegetation index, soil properties, composite terrain index, hydraulic conductivity and leaf area index.

[0014] Further optimization, the step S2 includes the following steps: Step S2.1: Select a distributed hydrological model and divide it into a warm-up period, a rate period, and a verification period. The warm-up period is set to 1-2 years for model initialization. The rate period and verification period have a ratio of 2:1.

[0015] Step S2.2: Calibrate the parameters of the distributed hydrological model using the ant colony algorithm.

[0016] Step S2.3: Construct the constraint function and perform accuracy screening. Specifically, after the parameter calibration is completed, select the percentage deviation PBIAS, Nash efficiency coefficient NSE and correlation coefficient R 2 Three indicators to construct the accuracy constraint function Y , evaluate the accuracy of the hydrological model after calibrating the parameters; if and only if the correlation coefficient R 2 When the NSE is greater than 0.5, the Nash efficiency coefficient NSE is greater than 0.6, and the percentage deviation PBIAS is less than 20%, the hydrological model passes the accuracy test and the corresponding area is a good accuracy area.

[0017] The precision constraint function is expressed by the following formula: ; In the above formula, is an indicator function, if the condition is met, the value is 1, otherwise it is 0; Y = 1 indicates that the distributed hydrological model passes the accuracy test; Y = 0 indicates that the distributed hydrological model fails the accuracy test.

[0018] Further optimization, step S2.2 uses the ant colony algorithm to calibrate the parameters of the distributed hydrological model, specifically including: Step S2.2.1: Set up the distributed hydrological model to include n Parameters to be calibrated , forming a parameter set , expressed as: ; i ∈[1, n ].

[0019] Step S2.2.2, algorithm implementation, specifically includes: Step S2.2.2.1: Initialization: Set the number of ants M and the maximum number of iterations T max , initial pheromone concentration , C is a constant.

[0020] Step S2.2.2.2, path selection: Each ant constructs a complete parameter set path; set the k When Ant builds the parameter set path, select i Parameter No. j The probability of a discrete value is: ; In the above formula, M is the total number of ants, k ∈[1,M]; w is the loop variable for enumerating all traversable nodes, N is the number of ants k The optional parameter space of is the pheromone concentration at the jth discrete value of the ith parameter at the tth iteration; α is the pheromone heuristic factor, β is the heuristic function factor; n ij is the heuristic function value of the jth discrete value of the i-th parameter. The heuristic function is defined as the inverse of the model performance index caused by the historical value of the parameter. The formula is expressed as: ; In the above formula, is the target function, and its variable is the Nash efficiency coefficient NSE; Representative i The first parameter j A discrete value.

[0021] Step S2.2.2.3, pheromone update: After each iteration, the pheromone is updated according to the ant search results. The formula is: ; In the above formula, ρ is the pheromone volatility coefficient, and its value range is [0,1]; Indicates that at the (t+1)th iteration, i Parameter No. j The pheromone concentration on the path corresponding to the discrete value; For the k Only ants on the path ( i , j ) is positively correlated with the model performance, that is, the better the model performance, the more pheromone the ants leave on the path they pass through.

[0022] In each iteration, the performance index of the hydrological model corresponding to the parameter set constructed by each ant is calculated, and the pheromone concentration is updated according to the better solution found by the ant; the path selection and pheromone update process is repeated until the maximum number of iterations T is reached. max , and obtain the optimal or approximately optimal parameter set.

[0023] Further optimization, the step S3 specifically includes the following steps: Step S3.1, construct a deep residual network model: the model includes a convolutional layer, a pooling layer, and a fully connected layer connected in sequence; the convolutional layer is used to extract the local spatial features of the input matrix, which is a three-dimensional matrix integrating meteorological data and underlying surface eigenvalues; the pooling layer is used to reduce the dimension of the features, and the fully connected layer is used to integrate the features and output parameter prediction values; the deep residual network adopts a residual module structure, which is expressed as follows: ; In the above formula, Represents the input features of the lth layer; represents the residual mapping function of the lth layer; is the set of trainable weights for layer l.

[0024] The fully connected layer outputs the predicted value of the hydrological model parameters: ; In the above formula, For the g The prediction parameters of the input samples, is the set of trainable weights in the deep residual network; Represents the three-dimensional matrix of characteristic factors of the g-th input sample; f (⋅) represents the mapping function of the deep residual network model.

[0025] Step 3.2: Training and validation of the deep residual network model: Divide the data in the area with good accuracy into training set, test set, and validation set in a ratio of 7:2:1; The mean square error is used as the loss function, which is conducive to training the deep residual network with the training set. The hydrological model parameters of the high-precision area are used as training labels. The weights of the deep residual network model are adjusted through the back propagation algorithm to minimize the loss function until the model converges. The loss function expression is: ; In the above formula, Indicates the g In the sample i The predicted value of the parameter; Indicates the first sample in the gth sample i The actual calibration value of the parameter; G represents the number of samples; The trained deep residual network model is used to predict the corresponding hydrological model parameters in the test set and compared with the corresponding actual calibration parameters to evaluate the rationality of the results. If the comparison results are reasonable, it means that the deep residual network model is effective. If not, the deep residual network model needs to be further optimized until the results meet expectations.

[0026] Step S3.3, parameter completion: For areas with missing data or areas that failed the accuracy test in step S2, an input matrix is constructed based on their meteorological data and underlying surface data, and the trained deep residual network is input. The hydrological model prediction parameters of the corresponding area are output to form a complete parameter spatial distribution covering the study area.

[0027] Further optimization, the step S4 specifically includes: Step S4.1: Calculate the characteristic factor weight coefficient based on the permutation importance method: Step S4.1.1: Use the Nash benefit coefficient or loss function value of the deep residual network model on the validation set as the original reference score S to measure the prediction accuracy of the deep residual network model.

[0028] Step S4.1.2, conduct a single feature shuffling experiment: First, fix the trained deep residual network model, keep other features unchanged, randomly shuffle the order of a certain feature, and generate a perturbed dataset; use the perturbed dataset to predict parameters and calculate the reference score after perturbation.

[0029] Step S4.1.3: Perform the above operations on all features in turn, repeat the shuffling of each feature Q times, and take the average change value as the importance measure.

[0030] Step S4.2, weight coefficient calculation: The importance of each feature is defined as the feature factor weight coefficient, which is calculated using the following formula: ; In the above formula, S is the original reference score of the model; Q refers to the number of repetitions of the input factor; The weight coefficient of the feature factor represents the degree of influence of different inputs on the output. The weight coefficient of the feature factor ranges from -1 to 1. A value greater than 0 indicates a positive correlation, and a value less than 0 indicates a negative correlation.

[0031] Step S4.3: Outputting the attribution results, specifically including: First, the top 3–5 key characteristic factors, such as composite topographic index, hydraulic conductivity, and leaf area index, were extracted by sorting by absolute value of weights, and their physical mechanisms were analyzed, such as the composite topographic index reflecting the effect of topography on water retention.

[0032] Then, the differences in factor weights were statistically analyzed by region (such as climate zones and topographic regions) to reveal region-specific driving mechanisms. For example, the weight of precipitation P in arid areas is significantly higher than that in humid areas.

[0033] Finally, visualization results, such as feature importance histograms and correlation matrices, are output to provide a basis for distributed hydrological model parameter optimization and physical process improvement.

[0034] Compared with the prior art, the present invention has the following beneficial effects: The method described in the present invention incorporates interpretability technology on the basis of a deep residual network and is applied to the field of hydrological models to conduct effective attribution analysis of parameters. This invention helps to reduce the uncertainty impact of "black box" models, helps to detect model logical errors, adjust the model to improve its accuracy, and helps to correctly understand the meaning of the model structure itself. The intervention of interpretability technology can well explain the decision-making process of hydrological models, make model parameters more physically realistic, and also facilitate the coupling and improvement of subsequent model modules, providing more accurate prediction tools for fields such as water resources management and environmental protection. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 Flowchart of the hydrological model parameter attribution method based on deep learning interpretability technology; Figure 2 is the weight of the influencing factors of global basin water storage capacity parameters. DETAILED DESCRIPTION

[0036] The technical solution of the present invention is described in detail below with reference to the embodiments, but the protection scope of the present invention is not limited to the embodiments.

[0037] This example uses the global-scale watershed storage capacity parameters as an example.

[0038] like Figure 1As shown in the figure, the hydrological model parameter attribution method based on deep learning interpretability technology includes the following steps: Step S1: Data Collection and Collation: All 10,374 observation sites in the Global Runoff Database (GRDC) were screened. Hydrological model calibration requires long-term, continuous, daily runoff observations. Furthermore, vector files of basin boundaries are required to extract basin characteristics. Basin selection criteria required basins with basin boundary vector files in the dataset, observations lasting at least 10 years, and a data missing rate of no more than 10%.

[0039] Over 7,000 basins in the GRDC have provided drainage boundary files, but 1,676 observation sites in the GRDC had data missing rates below 1%, 2,291 sites had data missing rates below 5%, and 2,625 sites had data missing rates below 10%. Therefore, 2,625 basin observation sites that met the requirements were ultimately selected from the global GRDC runoff data.

[0040] Step S2: Distributed hydrological model construction and parameter calibration: Using a distributed hydrological model, the warm-up period, calibration period and verification period are divided, and the ant colony algorithm is used to calibrate the hydraulic model parameters; the percentage deviation PBIAS, Nash benefit coefficient NSE and correlation coefficient R are selected. 2 As an accuracy indicator, a threshold constraint function is established to screen the model accuracy, and high-precision simulation areas are screened for subsequent deep learning attribution analysis. Specifically, it includes: Step S2.1: In this embodiment, the data collection sequence time is 1970-2016, wherein the warm-up period is 2 years, the model rate period is 30 years, and the model validation period is 15 years.

[0041] The input of the hydrological model includes precipitation, evaporation, etc. Some data are taken for display, such as the meteorological data of the Black River Basin (No. 4119430) in Wisconsin from January 1 to January 15, 1971, including rainfall and evaporation data, as shown in Table 1.

[0042] Table 1 Meteorological data of some basins date Precipitation.mm Runoff.mm Evaporation capacity.mm Atmospheric pressure.pa Radiation.Wm2 Temperature.K VPD.kpa Wind speed.ms 1971 / 1 / 1 2.12 0.30 0.27 96710 -14.39 269.92 0.11 2.92 1971 / 1 / 2 0.10 0.30 0.34 97191 -17.28 265.79 0.10 3.72 1971 / 1 / 3 5.66 0.26 0.13 98065 -26.52 262.12 0.09 2.06 1971 / 1 / 4 15.93 0.25 0.64 96432 -0.06 266.96 0.08 6.41 1971 / 1 / 5 0.12 0.25 0.39 97385 -19.88 254.99 0.07 4.63 1971 / 1 / 6 0.06 0.26 0.21 98369 -25.52 250.85 0.05 3.16 1971 / 1 / 7 0.07 0.28 0.19 98597 -22.49 251.45 0.05 2.59 1971 / 1 / 8 0.04 0.30 0.20 98203 -20.51 254.56 0.06 2.49 1971 / 1 / 9 2.25 0.31 0.51 96663 -21.38 264.54 0.11 4.88 1971 / 1 / 10 0.07 0.31 0.42 96812 -10.82 263.72 0.09 3.67 1971 / 1 / 11 2.22 0.31 0.33 97658 -6.66 259.99 0.09 1.91 1971 / 1 / 12 0.30 0.28 0.30 98682 -17.40 257.01 0.08 2.79 1971 / 1 / 13 2.06 0.28 0.53 98149 6.86 261.53 0.08 3.58 1971 / 1 / 14 0.41 0.30 0.43 97254 1.39 263.16 0.08 3.49 1971 / 1 / 15 0.27 0.32 0.24 98806 -22.34 253.20 0.06 2.70 The output of the hydrological model is a simulated runoff sequence, and the parameters include the maximum watershed storage capacity, soil moisture distribution curve index, and the proportion of impervious area.

[0043] Step S2.2: Calibrate the parameters of the distributed hydrological model using the ant colony algorithm Search for the optimal parameter set in the parameter space and construct the optimization equation to calibrate the model parameters: ; In the above formula, T is the number of time steps, To observe the flow, To simulate traffic.

[0044] Ant colony algorithm is a bionic heuristic optimization method that simulates the behavior of ants in finding the optimal path through pheromone transmission during foraging. For the parameter calibration problem, the number of ants M is set, and the parameter search space is defined as the upper and lower limits of each parameter. Each ant represents the search path of a parameter set, initializes the pheromone matrix, sets the initial pheromone concentration as a constant, and each ant selects the parameter value according to the current pheromone concentration and the heuristic function; then inputs the current parameter set into the hydrological model to calculate the simulated flow At the same time, the loss function is calculated according to the above formula, and then the pheromone concentration is updated according to the ant optimization effect. The optimal parameter set and corresponding loss value in the current iteration are retained until the iteration ends, the optimal parameter set is output, and the rate determination is completed.

[0045] Step S2.3: Construct the constraint function and perform accuracy screening. Specifically, after the parameter calibration is completed, select the percentage deviation PBIAS, Nash efficiency coefficient NSE and correlation coefficient R 2 Three indicators to construct the accuracy constraint function Y , evaluate the accuracy of the hydrological model after calibrating the parameters; if and only if the correlation coefficient R 2 When the NSE is greater than 0.5, the Nash efficiency coefficient NSE is greater than 0.6, and the percentage deviation PBIAS is less than 20%, the hydrological model passes the accuracy test and the corresponding area is a good accuracy area.

[0046] The precision constraint function is expressed by the following formula: ; Where, is an indicator function, if the condition is met, the value is 1, otherwise it is 0; Y=1 means the model passes the accuracy test; Y=0 means the model fails the accuracy test.

[0047] In this embodiment, the simulation results of the watershed storage capacity of the hydrological model after parameter calibration and the results of the evaluation model using the Nash efficiency coefficient NSE are shown in Table 2.

[0048] Table 2. Simulated water storage capacity station results based on the global watershed dataset Basin number latitude longitude NSE_Rate Regular NSE_Inspection Period Basin water storage capacity / mm 1 26.91 99.95 74.65 77.02 305.34 2 26.59 101.72 79.82 80.53 265.80 3 26.91 102.91 87.52 84.37 257.32 4 28.84 104.36 79.33 75.11 742.21 5 28.46 101.88 82.42 86.75 327.77 6 26.77 101.84 84.75 86.66 424.36 7 28.80 104.42 75.49 74.79 576.47 8 29.85 106.42 60.14 65.95 414.19 9 30.44 110.34 43.46 51.66 484.27 10 30.49 111.17 31.64 41.83 718.05 11 30.41 111.44 23.09 28.99 556.38 12 32.84 110.08 69.63 65.39 302.51 13 32.36 110.48 62.87 70.30 283.80 14 32.03 112.15 48.58 44.99 613.77 15 31.19 112.56 52.48 52.50 529.75

[0049] Step S3: Parameter completion based on deep residual network: Using the parameters of the high-precision area in step S2 as labels, the deep residual network model is trained to learn the mapping relationship between parameters and meteorological factors and underlying surface features; the mean square error is used as the loss function to optimize the weights of the deep residual network model to complete the parameter spatial distribution in areas with missing data or low-precision simulation areas. The specific steps include the following: Step S3.1, construct a deep residual network model: In the deep residual network model, the input three-dimensional matrix is composed of meteorological factors and underlying surface eigenvalues. The local spatial features are first extracted through the convolution layer, and then the dimensionality reduction and compression are achieved through the pooling layer. Finally, the feature integration and mapping are completed in the fully connected layer, and the output is a one-dimensional hydrological model parameter prediction value, which realizes the accurate regression prediction of the basin grid parameters.

[0050] Step 3.2, training and validation of the deep residual network model: divide the data in the area with good accuracy into training set, test set and validation set in a ratio of 7:2:1.

[0051] The mean square error is used as the loss function, which is conducive to training the deep residual network with the training set. The hydrological model parameters of the high-precision area are used as training labels. The weights of the deep residual network model are adjusted through the back propagation algorithm to minimize the loss function until the model converges.

[0052] For the test set, the parameters predicted by the deep residual network model are compared point by point with the corresponding calibration parameters. The determination coefficient R² and root mean square error RMSE are used as performance evaluation criteria to verify the accuracy and stability of the model.

[0053] Step S3.3, parameter completion: For areas with missing data or areas that failed the accuracy test in step S2, the input matrix is reconstructed based on the existing meteorological and underlying surface information, and reasoning is performed through the trained deep residual network model. The hydrological model prediction parameters of the corresponding area are output to form a complete parameter space distribution covering the study area. This method effectively fills the gaps in the original data and makes the parameter space distribution more coherent and reasonable.

[0054] Experiments have shown that this deep learning method can flexibly adapt to the reconstruction of hydrological parameters under different natural conditions, especially in watershed environments with complex terrain and diverse climates, showing wide applicability and excellent completion effects.

[0055] Step S4: Parameter attribution analysis based on characteristic importance: Use the permutation importance method to calculate the characteristic factor weight coefficient, clarify the degree of influence of different characteristic factors on the hydrological model parameters, and distinguish the positive and negative correlations, so as to achieve attribution analysis of the hydrological model parameters.

[0056] The input factors used in the deep residual network model include global meteorological data, soil and vegetation data, topographic data, and runoff characteristics, all of which may affect the water storage capacity of a watershed. Discretely distributed factors such as soil texture, soil type, and land use, which are categorized by type, are expressed as "type" and represented as discrete integers in the attribution analysis. Meteorological data include variables such as precipitation and near-surface temperature, which have direct or indirect effects on processes such as evaporation, transpiration, and runoff. Soil data include soil thickness, root zone depth, soil type, and land use type, all of which directly or indirectly affect soil water storage.

[0057] In this example, 13 characteristic importance factors were finally selected for attribution analysis of watershed water storage capacity on a global scale. Since different underlying surface conditions have a significant impact on watershed water storage capacity, and the underlying surfaces within the same zone in the global climate zoning have locational similarities, the equatorial zone, arid zone, and humid zone were divided for attribution analysis.

[0058] In this example, due to space limitations, some characteristic importance factors and basin water storage capacity are shown, as shown in Table 3.

[0059] Table 3 Global river basin water storage capacity and some important characteristic factors Serial number Basin water storage capacity Composite Terrain Index Elevation / m Hydraulic conductivity Land use Rainfall / mm Root zone depth / mm Soil thickness / mm Soil texture Soil thickness slope 1 394.1617 0.4633 2.88 0.1149 10 1731.1 160.45 50 3 3 0.0065 2 296.7248 0.4345 1.11 0.1113 14 1453 212.4 50 0 0 0.0051 3 355.0561 0.7097 1353.0601 0.0592 6 1869 285.85 1 9 9 10.8906 4 408.9005 0.6464 1170.95 0.0507 4 1953.1 383.05 1 9 7 7.5663 5 262.3802 0.7649 0.45 0.1004 8 1490.2 336.8 50 3 3 0.0028 6 392.8598 0.5126 28.38 0.121 11 1643.1 293.07 50 9 7 0.1212 7 455.0065 0.5695 26.2 0.1091 11 1852.7 284.36 50 6 6 0.0908 8 480.6681 0.5375 22.43 0.1103 10 1938.6 298.15 50 9 6 0.0462 9 484.6116 0.5702 25.94 0.0958 0 2201.1001 289.18 50 6 6 0.1731 10 624.9331 0.6122 133.45 0.0732 7 2590.3 336.72 1 6 6 1.3555 11 405.2319 0.7301 1390.89 0.0466 1 1683.5 266.74 1 12 9 7.4527 12 410.148 0.7588 1355.5601 0.0452 4 1913.7 368.43 1 12 9 8.9387 13 435.652 0.5907 257.11 0.0529 6 1938.2 458.49 6 9 7 1.2524 14 248.3777 0.8484 1.46 0.1015 9 1284.8 0 50 6 6 0.0216 15 479.4843 0.5243 19.25 0.1325 11 1403.4 351.88 50 9 6 0.1498 .

[0060] like Figure 2 Figure 2 shows the estimated importance of various characteristic factors. Among them, the composite topographic index, hydraulic conductivity, and leaf area index (LAI) are the main factors influencing the spatial evolution of watershed water storage capacity parameters, explaining over 60% of the spatial distribution. At the global scale, the study found that watershed water storage capacity is significantly positively correlated with six characteristics: root zone depth (RD), soil thickness (SD), leaf area index (LAI), standard vegetation index (NDVI), annual mean temperature (T), and annual mean precipitation (P). The factor weight coefficients are 0.29, 0.17, 0.36, 0.30, 0.12, and 0.10, respectively. In contrast, five characteristics: land use (LUCC), hydraulic conductivity (K), elevation (H), composite topographic index (CTI), and slope (S) are significantly negatively correlated with watershed water storage capacity. The factor weight coefficients are -0.29, -0.41, -0.10, -0.51, and -0.26, respectively. The other two characteristics, soil texture and soil type, showed a weak negative correlation with the watershed water storage capacity, with factor characteristic weight coefficients of -0.09 and -0.05, respectively.

[0061] The following three feature importance factors are selected for additional explanation from the perspective of physical mechanism: The Leaf Area Index (LAI) refers to the total leaf area per unit area and is commonly used to describe the density and richness of vegetation cover. An increase in LAI increases the density of plant roots, increasing the water storage space in the soil and altering the permeability of water, thereby affecting the water storage capacity of a watershed. LAI also affects runoff. Higher vegetation cover increases the vegetation's ability to intercept rainfall, reducing runoff and correspondingly increasing the water storage capacity of a watershed.

[0062] Hydraulic conductivity refers to the ability of water within soil or rock to migrate downward. Different soils and rocks have different hydraulic conductivities, which affect the movement and distribution of water within a watershed. If the hydraulic conductivity within a watershed is high, water within the soil or rock can move downward more quickly, reducing the water content within the soil or rock and diminishing its water storage capacity. Conversely, if the hydraulic conductivity within a watershed is low, water within the soil or rock encounters greater resistance, slowing its movement. This results in a relatively high water content within the soil or rock and a corresponding increase in the water storage capacity of the watershed.

[0063] The Composite Terrain Index (CTI) describes the potential for runoff generation within a watershed. It is closely related to the basin's topography and is the natural logarithm mean of the ratio of the upstream contribution area to the length of the flow path from each point within the basin to the basin outlet. The CTI effectively reflects the potential for runoff generation within a watershed, specifically, it is closely related to factors such as soil moisture storage, groundwater levels, and surface runoff. A higher CTI value indicates more complex and undulating topography within the basin, leading to more tortuous flow paths. This leads to increased water retention and infiltration as it flows through such terrain, thereby increasing soil moisture storage and groundwater levels and reducing surface runoff. Conversely, a lower CTI value indicates flatter topography and simpler flow paths within the basin, leading to faster water collection into the river channel and increased surface runoff. Therefore, basins with high CTI values have relatively more water storage capacity, and more water will be retained in underground storage, while basins with low CTI values have relatively weaker water storage capacity, and most of the water will quickly flow into the river in the form of surface runoff.

[0064] from Figure 2As can be seen in the data, in the equatorial and humid zones, watershed storage capacity has a significant positive correlation with six characteristics: root zone depth (RD), soil thickness (SD), leaf area index (LAI), standard vegetation index (NDVI), annual mean temperature (T), and annual mean precipitation (P). In contrast, four characteristics, land use (LUCC), hydraulic conductivity (K), composite topographic index (CTI), and slope (S), have a significant negative correlation with watershed storage capacity. Elevation (H) shows a weak positive correlation in humid regions, while the remaining two characteristics, soil texture and soil type, show a weak negative correlation with watershed storage capacity. It can be seen that the influence of soil thickness and root zone depth becomes greater in temperate climate zones.

[0065] Soil thickness significantly impacts a watershed's water storage capacity. Thicker soil layers hold more soil water because they have larger pore spaces and more soil organisms. These pores and organisms can store more water and accelerate its infiltration and transport. Therefore, the thicker the soil layer, the greater its water storage capacity. On the other hand, thinner soil layers limit their water storage capacity. This is because thin soil layers have less pore space and less water storage capacity, and the soil water flows more rapidly, resulting in much of the water not being effectively stored. Consequently, soil water easily flows out of the watershed and goes unutilized.

[0066] The root zone depth refers to the depth at which plant roots are distributed in the soil. The root zone depth affects the soil pore structure. Root growth will destroy the soil structure, forming different soil pores. If the root system is shallow, the pores formed in the soil are mainly distributed in the surface layer, making the surface soil prone to water saturation and leakage, affecting water storage capacity. The root zone depth affects the vegetation coverage and vegetation type. The distribution and growth status of plants will affect the runoff process in the basin, thereby affecting the water storage capacity of the basin. If the root system is shallow, the distribution of plants in the soil surface will also be shallow, and the evaporation of soil moisture and the utilization of water storage capacity will also be weakened.

[0067] In inland and arid zones, watershed storage capacity shows a significant positive correlation with six characteristics: root zone depth (RD), soil thickness (SD), leaf area index (LAI), standard vegetation index (NDVI), annual mean temperature (T), and annual mean precipitation (P). In contrast, four characteristics—land use (LUCC), hydraulic conductivity (K), composite topographic index (CTI), and slope (S)—are significantly negatively correlated with watershed storage capacity. Elevation (H) shows a weak negative correlation in arid regions, while the remaining two characteristics—soil texture and soil type—show weak negative correlations with watershed storage capacity. It can be seen that the influence of land use type and annual mean precipitation becomes greater in arid climate zones.

[0068] Increased rainfall can have complex effects on a basin's water storage capacity, depending on the interactions between various hydrological processes within the basin. Generally speaking, increased rainfall typically leads to an increase in a basin's water storage capacity. As rainwater infiltrates, the water content in the soil increases, which improves the soil's water storage capacity and, in turn, increases the basin's water storage capacity. However, in some specific cases, it can also lead to a decrease in water storage capacity. Increased rainfall also promotes vegetation growth, and transpiration from vegetation consumes some rainwater, but also converts some rainwater back into atmospheric water vapor, reducing the basin's water storage capacity.

[0069] In polar regions, watershed water storage capacity showed a significant positive correlation with root zone depth (RD), leaf area index (LAI), and average annual precipitation (P). In contrast, soil texture, land use (LUCC), hydraulic conductivity (K), elevation (H), and composite topographic index (CTI) showed significant negative correlations with watershed water storage capacity. Soil thickness (SD), slope (S), standard vegetation index (NDVI), and soil type showed a weak negative correlation with watershed water storage capacity, while average annual temperature (T) showed a weak positive correlation.

[0070] As above, although the present invention has been shown and described with reference to specific preferred embodiments, it should not be construed as limiting the present invention itself. Various changes may be made to it in form and detail without departing from the spirit and scope of the present invention as defined in the appended claims.

Claims

1. A hydrological model parameter attribution method based on deep learning interpretability technology, characterized by: The following steps are involved: Step S1: Data collection and preprocessing: Construct a grid-scale hydrological dataset for the study area, which includes meteorological data, runoff depth data, and underlying surface watershed characteristics. For areas with sparse observation stations in the watershed, additional stations are added to ensure the spatiotemporal representativeness of the hydrological series. Step S2: Distributed hydrological model construction and parameter calibration: A distributed hydrological model suitable for the study area is selected and divided into a warm-up period, a calibration period, and a validation period. An ant colony algorithm is used to optimize the parameters of the distributed hydrological model. A parameter calibration model with the Nash benefit coefficient (NSE) as the objective function is constructed. The parameter search path is dynamically adjusted using pheromone concentration and a heuristic function. An accuracy constraint function is constructed based on the percentage deviation (PBIAS), the Nash benefit coefficient (NSE), and the correlation coefficient (R²) to screen out high-precision simulation areas. Step S3: Parameter completion based on deep residual network: Using the parameters of the high-precision area in step S2 as labels, the deep residual network model is trained to learn the mapping relationship between parameters and meteorological factors and underlying surface characteristics; the mean square error is used as the loss function to optimize the weights of the deep residual network model to complete the parameter spatial distribution in areas with missing data or low-precision simulation areas; Step S4: Parameter attribution analysis based on characteristic importance: Use the permutation importance method to calculate the characteristic factor weight coefficient, clarify the degree of influence of different characteristic factors on the hydrological model parameters, and distinguish the positive and negative correlations, so as to achieve attribution analysis of the hydrological model parameters.

2. The hydrological model parameter attribution method based on deep learning interpretability technology according to claim 1, characterized in that: In step S1, the grid-scale hydrological dataset must satisfy temporal continuity and spatial consistency; Among them, meteorological data include precipitation, temperature, wind speed, potential evapotranspiration, runoff and atmospheric pressure; Runoff depth data are derived from a high-resolution gridded runoff depth dataset based on long-term observations, collected from the Global Runoff DataBase, and interpolated to the grid scale using the Thiessen polygon method. The underlying surface watershed characteristic values include digital elevation model (DEM), land use type, vegetation index, soil properties, composite terrain index, hydraulic conductivity and leaf area index.

3. The hydrological model parameter attribution method based on deep learning interpretability technology according to claim 2, characterized in that: The step S2 comprises the following steps: Step S2.1: Select a distributed hydrological model and divide it into a warm-up period, a rate-based period, and a validation period. The warm-up period is set to 1-2 years for model initialization; the rate-based period and the validation period have a ratio of 2:

1. Step S2.2: Calibrate the parameters of the distributed hydrological model using the ant colony algorithm; Step S2.3: Construct the constraint function and perform accuracy screening. Specifically, after the parameter calibration is completed, select the percentage deviation PBIAS, Nash efficiency coefficient NSE and correlation coefficient R 2 Three indicators to construct the accuracy constraint function Y , evaluate the accuracy of the distributed hydrological model after calibrating the parameters; if and only if the correlation coefficient R 2 When the NSE is greater than 0.5, the Nash efficiency coefficient NSE is greater than 0.6, and the percentage deviation PBIAS is less than 20%, the hydrological model passes the accuracy test and the corresponding area is a good accuracy area. The accuracy constraint function is expressed by the following formula: ; In the above formula, is an indicator function, if the condition is met, the value is 1, otherwise it is 0; Y = 1 indicates that the distributed hydrological model passes the accuracy test; Y = 0 indicates that the distributed hydrological model fails the accuracy test.

4. The hydrological model parameter attribution method based on deep learning interpretability technology according to claim 3 is characterized by: The step S2.2 uses the ant colony algorithm to calibrate the parameters of the distributed hydrological model, specifically including: Step S2.2.1: Set up the distributed hydrological model to include n Parameters to be calibrated , forming a parameter set , expressed as: ; i ∈[1, n ]; Step S2.2.2, algorithm implementation, specifically includes: Step S2.2.2.1: Initialization: Set the number of ants M and the maximum number of iterations T max , initial pheromone concentration , C is a constant; Step S2.2.2.2, path selection: Each ant constructs a complete parameter set path; set the k When Ant builds the parameter set path, select i Parameter No. j The probability of a discrete value is: ; In the above formula, M is the total number of ants, k ∈[1,M]; w is the loop variable for enumerating all traversable nodes, N is the number of ants k The optional parameter space of is the pheromone concentration at the jth discrete value of the ith parameter at the tth iteration; α is the pheromone heuristic factor, β is the heuristic function factor; n ij is the heuristic function value of the jth discrete value of the i-th parameter. The heuristic function is defined as the inverse of the model performance index caused by the historical value of the parameter. The formula is expressed as: ; In the above formula, is the target function, and its variable is the Nash efficiency coefficient NSE; Representative i The first parameter j discrete values; Step S2.2.2.3, pheromone update: After each iteration, the pheromone is updated according to the ant search results. The formula is: ; In the above formula, ρ is the pheromone volatility coefficient, and its value range is [0,1]; Indicates that at the (t+1)th iteration, i Parameter No. j The pheromone concentration on the path corresponding to the discrete value; For the k Only ants on the path ( i , j The amount of pheromone left on the path is positively correlated with the model performance, that is, the better the model performance, the more pheromone the ants leave on the path they pass through; In each iteration, the performance index of the hydrological model corresponding to the parameter set constructed by each ant is calculated, and the pheromone concentration is updated according to the better solution found by the ant; the path selection and pheromone update process is repeated until the maximum number of iterations T is reached. max , and obtain the optimal or approximately optimal parameter set.

5. The hydrological model parameter attribution method based on deep learning interpretability technology according to claim 4 is characterized by: The step S3 specifically includes the following steps: Step S3.1, construct a deep residual network model: the model includes a convolutional layer, a pooling layer, and a fully connected layer connected in sequence; the convolutional layer is used to extract the local spatial features of the input matrix, which is a three-dimensional matrix integrating meteorological data and underlying surface eigenvalues; the pooling layer is used to reduce the dimension of the features, and the fully connected layer is used to integrate the features and output parameter prediction values; the deep residual network adopts a residual module structure, which is expressed as follows: ; In the above formula, Indicates the l The input features of the layer; Indicates the l The residual mapping function of the layer; For the l The set of trainable weights of the layer; The fully connected layer outputs the predicted value of the hydrological model parameters: ; In the above formula, For the g The prediction parameters of the input samples, is the set of trainable weights in the deep residual network; Represents the three-dimensional matrix of characteristic factors of the g-th input sample; f (⋅) represents the mapping function of the deep residual network model; Step 3.2: Training and validation of the deep residual network model: Divide the data in the area with good accuracy into training set, test set, and validation set in a ratio of 7:2:1; The mean square error is used as the loss function, which is conducive to training the deep residual network with the training set. The hydrological model parameters of the high-precision area are used as training labels. The weights of the deep residual network model are adjusted through the back propagation algorithm to minimize the loss function until the model converges. The loss function The expression is: ; In the above formula, Indicates the g In the sample i The predicted value of the parameter; Indicates the first sample in the gth sample i The actual calibration value of the parameter; G represents the number of samples; The trained deep residual network model is used to predict the corresponding hydrological model parameters in the test set and compared with the corresponding actual calibration parameters to evaluate the rationality of the results. If the comparison results are reasonable, it means that the deep residual network model is effective. If not, the deep residual network model needs to be further optimized until the results meet the expectations. Step S3.3, parameter completion: For areas with missing data or areas that failed the accuracy test in step S2, a three-dimensional matrix is constructed based on its meteorological data and underlying surface data, and the trained deep residual network is input to output the hydrological model prediction parameters of the corresponding area, forming a complete parameter spatial distribution covering the study area.

6. The hydrological model parameter attribution method based on deep learning interpretability technology according to claim 5, characterized in that: The step S4 specifically includes: Step S4.1: Calculate the characteristic factor weight coefficient based on the permutation importance method: Step S4.1.1, using the Nash benefit coefficient or loss function value of the deep residual network model on the validation set as the original reference score S, to measure the prediction accuracy of the deep residual network model; Step S4.1.2: Conduct a single-feature shuffling experiment: First, fix the trained deep residual network model, keep other features unchanged, and randomly shuffle the order of a feature to generate a perturbed dataset; use this perturbed dataset to predict parameters and calculate the reference score after perturbation; Step S4.1.3: Perform the above operation on all features in turn, repeat the shuffle Q times for each feature, and take the average change value as the importance measure; Step S4.2, weight coefficient calculation: The importance of each feature is defined as the feature factor weight coefficient, which is calculated using the following formula: ; In the above formula, S is the original reference score of the model; Q refers to the number of repetitions of the input factor; represents the reference score obtained after the qth perturbation of the pth feature; The characteristic factor weight coefficient ranges from -1 to 1, where greater than 0 indicates positive correlation and less than 0 indicates negative correlation; Step S4.3: Outputting the attribution results, specifically including: First, sort by absolute value of weight, extract the top 3-5 key characteristic factors and analyze their physical mechanisms; Then, the differences in characteristic factor weights were statistically analyzed by region to reveal the region-specific driving mechanism; Finally, the visualization results are output to provide a basis for distributed hydrological model parameter optimization and physical process improvement.

Citation Information

Patent Citations

  • Regional land water reserve change attribution analysis method based on interpretable deep learning

    CN117132023A

  • Sare data region hydrological model parameter reconstruction method based on image deep learning

    CN117763970A

  • Method for an explainable autoencoder and an explainable generative adversarial network

    WO2022101515A1

Cited By

  • Method and system for constructing arid irrigation area basic model

    CN121388608A

  • Evaporation complementary relation parameter determination method based on interpretable machine learning

    CN121936306A