Hydrological multi-model assembly optimization method based on environmental factor response and machine learning

Through the multi-model assembly optimization method based on environmental factor response and machine learning, the high-precision simulation problem of hydrological models in different watersheds and climatic conditions is solved, the flexible adaptation and efficient optimization of the model are achieved, and the accuracy and applicability of hydrological simulation are improved.

CN120493680APending Publication Date: 2025-08-15HOHAI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510389688.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

When existing hydrological models face the environmental characteristics and variable climatic conditions of different river basins, it is difficult to achieve high-precision runoff simulation, lack a flexible model assembly and optimization mechanism, and cannot fully consider the changes in the relationship between environmental factors and model response.

Method used

Based on the method of environmental factor response and machine learning, the optimal model method for different confluent partitions is identified through multi-model assembly and optimization, combined with hydrological statistics and numerical simulation, a random forest is used to construct alternative models, and parameter sensitivity is analyzed through SHAP and Sobol methods, and the model combination is optimized by the SCE-UA algorithm.

Benefits of technology

It significantly improves the accuracy of flood simulation and model flexibility, can maintain high simulation performance under limited data or complex watershed environment, reduce waste of computing resources, adapt to the impact of climate change and human activities, and provide scientific decision-making support for water resource management and flood warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493680A_ABST
    Figure CN120493680A_ABST
Patent Text Reader

Abstract

The invention discloses a hydrological multi-model assembly optimization method based on environmental factor response and interpretable machine learning. The method comprises the following steps: carrying out runoff partition of grids based on environmental factor response; environmental indexes TWI, SPI and CN of the grids are calculated, and runoff production partitions are divided according to a comprehensive score calculation result; a model method with the best simulation effect of different runoff and confluence partitions is effectively identified by means of a machine learning algorithm, and a random forest algorithm is used to construct a substitution model of runoff and confluence modules; performing model interpretation of the substitution model by applying the SHAP, and analyzing sensitive input of the model; the reliability of the machine learning substitution model is verified by calculating and comparing the parameter sensitivity; the method adopts the SCE-UA algorithm to optimize the runoff production model and the confluence model, has wide applicability, can adapt to various environmental conditions under the influence of climate changes and human activities, improves the availability of the models, and provides scientific decision support for water resource management and flood early warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of hydrological model optimization and application, and specifically relates to a hydrological multi-model combination optimization method based on environmental factor response and machine learning. Background Art

[0002] Current hydrological models typically use a single model or fixed model parameters to simulate runoff and runoff, making it difficult to fully adapt to the environmental characteristics and changing climatic conditions of different basins. This is especially true in basins with complex underlying surfaces or limited data, where simulation results can exhibit significant errors. Existing technologies lack a flexible model combination optimization mechanism and are unable to fully account for changes in the relationship between environmental factors and model responses. Therefore, there is an urgent need for a multi-model intelligent optimization solution based on environmental trigger responses. By deeply exploring the response relationship between environmental conditions and model modules, an adaptive model combination system can be established to optimize model combinations and achieve more accurate runoff simulation. Summary of the Invention

[0003] In response to the problems existing in the prior art, the present invention provides a hydrological multi-model combination optimization method based on environmental factor response and machine learning. The grid runoff zoning is performed based on the environmental factor response, and the model method with the best simulation effect of different runoff zoning is effectively identified with the help of machine learning algorithms. The reliability of the machine learning alternative model is verified by calculating and comparing parameter sensitivity, giving full play to the advantages of traditional hydrological physical models and machine learning algorithms, and can provide data support for the input of hydrological models.

[0004] To solve the above technical problems, the present invention provides the following technical solution: a hydrological multi-model combination optimization method based on environmental factor response and machine learning, comprising the following steps:

[0005] S1. Obtain data on various environmental factors in the target area, including topography, soil type, soil moisture, climate zones, vegetation types, and humidity. Preprocess the data and divide the sub-basin into sub-basins based on the underlying surface conditions.

[0006] S2. Calculate the terrain wetness index (TWI), standardized precipitation index (SPI), and runoff curve number index (CN) of the grid, perform normalization and weight distribution calculations, and divide the runoff generation zones according to the comprehensive score calculation results;

[0007] S3. Based on the setup of multiple model modules, combined with hydrological statistics and numerical simulation, the distributed Xin'anjiang model, the terrain hydrological model, the soil and water assessment tool model, the variable infiltration capacity model, the Sacramento soil water balance model, and the Weather Research and Forecasting hydrological coupling model were split. Multiple runoff generation and confluence models were formed by combining multiple confluence methods such as Muskingum, diffusion wave, kinematic wave, and Congee algorithm. Random forests were used to construct alternative models for each runoff generation and confluence module.

[0008] S4. Use the model interpretation method SHAP to interpret the alternative model and analyze the model's sensitive inputs;

[0009] S5. Use Morris screening method to screen out sensitive environmental parameters, and then use Sobol method to further quantitatively analyze the sensitivity of the parameters;

[0010] S6. Compare the results of the sensitivity analysis of environmental parameters and surrogate model inputs. The sensitivity ranking of environmental parameters is consistent with the sensitivity ranking of the input parameters of the surrogate model, verifying the reliability of the surrogate model.

[0011] S7. Use the SCE-UA algorithm optimization model to determine the most suitable runoff model for each runoff zone and the confluence model for each sub-basin, and implement a hydrological multi-model combination optimization method based on environmental factor response and machine learning.

[0012] Furthermore, in the aforementioned step S2, calculating the standardized precipitation index SPI includes the following sub-steps: A-2.1, calculating the probability density function of the precipitation amount x on a preset time scale that follows a gamma distribution:

[0013]

[0014] Where β and γ are the scale and shape parameters of the gamma distribution, respectively, β>0, γ>0; β and γ are obtained using the maximum likelihood estimation method, as follows:

[0015]

[0016] Where x i is the precipitation sample, is the average precipitation amount;

[0017] A-2.2. Using the monthly precipitation x0, calculate the probability of the event x < x0:

[0018]

[0019] A-2.3. Calculate the cumulative probability of the preset time scale as follows:

[0020]

[0021] If the monthly precipitation is 0, the probability estimation method is as follows:

[0022]

[0023] In the formula, m represents the number of samples without precipitation, and n represents the total number of samples;

[0024] A-2.4. Normalize the cumulative probability. Use the standard normal distribution function to infer the exponential value of the obtained probability to obtain the SPI, as shown in the following formula:

[0025] SPI=φ(P)

[0026] Where φ is the standard normal distribution function, P is the cumulative probability value, and the larger the SPI, the wetter it is.

[0027] Furthermore, the terrain wetness index TWI is calculated as follows:

[0028]

[0029] Where a is the upstream contribution area per unit profile length, and β is the slope.

[0030] Furthermore, the aforementioned runoff curve number CN ranges from 0 to 100 and is comprehensively assessed by factors affecting previous rainfall, topography, surface land type, and soil properties.

[0031] Furthermore, in the aforementioned step S2, the runoff zones are divided according to the comprehensive score calculation results, specifically including the following sub-steps:

[0032] B-2.1. Normalize the SPI, CN, and TWI of each grid to obtain SPI', CN', and TWI';

[0033] B-2.2, Construct a new indicator X XS , represents the wetness of the grid, S x Bigger means more moist:

[0034] X XS =SPI'-0.5×CN'+0.5×TWI'

[0035] B-2.3, based on the calculated X XS The value divides each grid into 5 runoff zones X A 、X B 、X C 、X D and X E , mark it.

[0036] Furthermore, the model interpretation method SHAP in step 4 above explains the output of the machine learning model. The SHAP value is used to measure the contribution of each input feature to the predicted output. For a given model f and input sample x, the SHAP value of feature i is The calculation is as follows:

[0037]

[0038] Where M is the input set, S is any subset that does not contain input i, |S| is the size of set S, |M| is the number of features, and f x (S∪{i})-f x (S) Model prediction results with and without feature i.

[0039] Furthermore, the aforementioned step S5 includes the following sub-steps:

[0040] S5.1. Determine the range of model parameters, the number of levels p and the disturbance factor Δ = 1 / (p-1);

[0041] S5.2, for the parameter x i The values of are normalized and discretized:

[0042] x i ∈{0,1 / (p-1),2 / (p-1),…,1-1 / (p-1)}

[0043] S5.3, generate Morris's sample matrix B * :

[0044] B * ={(J1x * +Δ / 2)[(2B-J m )D * +J m ]}P *

[0045] Where B * is a (m+1)×m parameter sample matrix, each row represents a parameter sample, m is the number of model parameters; J1 is a (m+1)×1 matrix, J m is a (m+1)×m matrix, and all matrix elements are 1; B is a matrix (b ij ) (m+1)×m , when j≤i,i=1,2,…,m, b ij =1, otherwise 0; D * is an m×m random confusion matrix; the generated parameter sample matrix B * It has the following characteristics: the model parameter samples in two adjacent rows have only one parameter value that is different, and the values of the other m-1 model parameters are exactly the same;

[0046] S5.4, calculate the basis effect EE of each model parameter i and its mean And the standard deviation:

[0047]

[0048] S5.5, the variance decomposition formula of the Sobol method calculation model is:

[0049]

[0050] Where, f=f(x1,x2,…,x m ) is the structure of the model, V(f) is the total variance of the model output, V i is the variance term generated by the i-th parameter, V ij is the variance term generated by the interaction of the i-th and j-th parameters, V 1,2,…,m is the variance generated by all parameters;

[0051] S5.6, define the first-order sensitivity S i , second-order sensitivity S ij With the total sensitivity S Ti index:

[0052]

[0053] Among them, V -i represents the variance term without the effect of the i-th parameter, S i With S ij Respectively represent the impact of a single model parameter and the combination of two model parameters on the model output, S Ti represents the impact of all parameter combinations including the i-th model parameter on the model output, S Ti -S i Used to analyze the interaction effect between the i-th model parameter and other model parameters;

[0054] S5.7, based on the characteristics and requirements of hydrological model calculations, the Nash-Sutcliffe efficiency coefficient (NSE) is selected to characterize the model output results. The calculation formula for NSE is as follows:

[0055]

[0056] Among them, Q si is the simulated runoff, Q oi is the measured runoff, is the average measured runoff, and n is the sequence length.

[0057] Furthermore, the aforementioned step S7 includes the following sub-steps:

[0058] S7.1. Initialize the parameters of the SCE-UA algorithm: set the number of complex shapes h ≥ 1, the number of complex internal points m ≥ n + 1, and the sample size s = h × m;

[0059] S7.2. Generate sample points and randomly sample s sample points x in the feasible parameter space. i And calculate the function value f i =f(xi ), generate samples using a uniform probability distribution;

[0060] S7.3. Sort the s sample points (x i , f i ) in ascending order of function values, and denote it as buffer D = {(x i , f i ), i = 1, 2,..., s};

[0061] S7.4. Divide the s sample points obtained by sampling into p composite shapes A1,..., A p :

[0062]

[0063] S7.5. According to the competitive composite shape evolution CCE algorithm, perform evolution on each composite shape;

[0064] S7.6. Combine the vertices of all evolved composite shapes into a new sample group and sort them by function value;

[0065] S7.7. Check the reduced number of composite shapes. If the required minimum number of composite shapes p in the group min < p, remove the composite shape with the lowest ranked point or randomly remove a composite shape.

[0066] Compared with the prior art, the beneficial technical effects of the present invention adopting the above technical solutions are as follows: A hydrological multi-model combination optimization method based on environmental factor response and machine learning provided by the present invention can accurately identify the response relationship between environmental factors and model modules by systematically collecting and analyzing environmental factor data of multiple typical basins, thereby significantly improving the accuracy of flood simulation. In addition, this method enhances the flexibility of the model, enabling it to make intelligent selections and adjustments according to different environmental conditions, and still maintaining high simulation performance even in the case of limited data or complex basin environments. At the same time, by inventing new indicators considering rainfall, terrain, and soil for runoff generation zoning, resource utilization is optimized, unnecessary waste of computing resources is reduced, and the calculation efficiency of the optimization model is improved. The reliability of the machine learning model is verified by comparing parameter sensitivity and analyzing the sensitivity of alternative model inputs. In addition, the method of the present invention has wide applicability, can adapt to various environmental conditions under the influence of climate change and human activities, improve the usability of the model, and provide scientific decision-making support for water resource management and flood warning. Description of the Drawings

[0067] Figure 1 is the flow chart of the method of the present invention. Detailed Embodiment

[0068] In order to better understand the technical content of the present invention, specific embodiments are given and described below with reference to the accompanying drawings.

[0069] Various aspects of the present invention are described herein with reference to the accompanying drawings, which show a number of illustrative embodiments. The embodiments of the present invention are not limited to those described in the accompanying drawings. It should be understood that the present invention can be implemented by any of the various concepts and embodiments described above, as well as the concepts and implementations described in detail below, because the concepts and embodiments disclosed herein are not limited to any particular implementation. In addition, some aspects disclosed herein may be used alone or in any appropriate combination with other aspects disclosed herein.

[0070] refer to Figure 1 The present invention provides a hydrological multi-model combination optimization method based on environmental factor response and machine learning, comprising the following steps:

[0071] S1. Collect environmental factor data for multiple typical watersheds, including topography (digital elevation model (DEM), river network density, slope, etc.), soil type (such as sand, clay, loam, etc.), soil moisture, climatic conditions (precipitation, temperature, air humidity, etc.), vegetation type (such as NDVI, LAI), land use type, and observation data from hydrological stations and meteorological stations; preprocess the data and use normalization methods to convert data of different dimensions to the same scale range, fill in missing data, and remove outliers.

[0072] S2. Calculate the Terrain Wetness Index (TWI), Standardized Precipitation Index (SPI), and Runoff Curve Number (CN) of the grid, perform normalization and weight distribution, and divide runoff zones based on the comprehensive score calculation results. Specifically, the steps are as follows: As a preferred embodiment of the present invention, calculating the Standardized Precipitation Index (SPI) includes the following substeps: A-2.1. Calculate the probability density function of the gamma distribution of rainfall x at a preset time scale:

[0073]

[0074] Where β and γ are the scale and shape parameters of the gamma distribution, respectively, β>0, γ>0; β and γ are obtained using the maximum likelihood estimation method, as follows:

[0075]

[0076] Where x i is the precipitation sample, is the average precipitation amount;

[0077] A-2.2. Using the monthly precipitation x0, calculate the probability of the event x < x0:

[0078]

[0079] A-2.3. Calculate the cumulative probability of the preset time scale as follows:

[0080]

[0081] If the monthly precipitation is 0, the probability estimation method is as follows:

[0082]

[0083] In the formula, m represents the number of samples without precipitation, and n represents the total number of samples;

[0084] A-2.4. Normalize the cumulative probability. Use the standard normal distribution function to infer the exponential value of the obtained probability to obtain the SPI, as shown in the following formula:

[0085] SPI=φ(P)

[0086] Where φ is the standard normal distribution function, P is the cumulative probability value, and the larger the SPI, the more humid it is.

[0087] As a preferred embodiment of the present invention, the terrain wetness index TWI is calculated as follows:

[0088]

[0089] Where a is the upstream contribution area per unit profile length, and β is the slope (expressed in radians).

[0090] As a preferred embodiment of the present invention, the runoff curve number CN can be used to represent land use type and soil characteristics, and can also reflect the difficulty of directly generating runoff. The value range of CN is 0 to 100, and is comprehensively evaluated by factors such as previous rainfall, topography, surface land type and soil properties. The smaller the CN value, the wetter the area. The index is shown in Table 1: CN Index Value Table, as follows:

[0091] Table 1

[0092]

[0093] As a preferred embodiment of the present invention, in step S2, the runoff zones are divided according to the comprehensive score calculation results, which specifically includes the following sub-steps:

[0094] B-2.1. Normalize the SPI, CN, and TWI of each grid to obtain SPI', CN', and TWI';

[0095] B-2.2, Construct a new indicator X XS , represents the wetness of the grid, S x Bigger means more moist:

[0096] X XS =SPI'-0.5×CN'+0.5×TWI'

[0097] B-2.3, based on the calculated X XS The value divides each grid into 5 runoff zones X A 、X B 、X C 、X D and X E , mark it.

[0098] S3. Based on the setting of multiple model modules, combined with hydrological statistics and numerical simulation, the distributed Xin'anjiang Model, Topography-based Hydrological Model (TOPMODEL), Soil and Water Assessment Tool (SWAT), Variable Infiltration Capacity Model (VIC), Sacramento Soil Moisture Accounting Model (SAC), and Weather Research and Forecasting Hydrological Model (WRF-Hydro) were split. Multiple runoff generation and confluence models were formed by combining multiple runoff methods such as Muskingum, diffusion wave, kinematic wave, and Congee algorithm. Random forest was used to construct alternative models for each runoff generation and confluence module.

[0099] S4. Use the model interpretation method SHAP to interpret the alternative model and analyze the model’s sensitive inputs.

[0100] The model interpretation method SHAP can explain the output of the machine learning model. The SHAP value can be used to measure the contribution of each input feature to the predicted output. For a given model f and input sample x, the SHAP value of feature i is The calculation method is as follows:

[0101]

[0102] Where M is the input set, S is any subset that does not contain input i, |S| is the size of set S, |M| is the number of features, and f x (S∪{i})-f x (S) Model prediction results with and without feature i.

[0103] S5. Use the Morris screening method to screen out sensitive environmental parameters, and then use the Sobol method to further quantitatively analyze the sensitivity of the parameters; specifically, it includes the following sub-steps:

[0104] S5.1. Determine the range of model parameters, the number of levels p and the disturbance factor Δ = 1 / (p-1);

[0105] S5.2, for the parameter x i The values of are normalized and discretized:

[0106] x i ∈{0,1 / (p-1),2 / (p-1),…,1-1 / (p-1)}

[0107] S5.3, generate Morris's sample matrix B * :

[0108] B * ={(J1x * +Δ / 2)[(2B-J m )D * +J m ]}P *

[0109] Where B * is a (m+1)×m parameter sample matrix, each row represents a parameter sample, m is the number of model parameters; J1 is a (m+1)×1 matrix, J m is a (m+1)×m matrix, and all matrix elements are 1; B is a matrix (b ij ) (m+1)×m , when j≤i,i=1,2,…,m, b ij =1, otherwise 0; D * is an m×m random confusion matrix; the generated parameter sample matrix B * It has the following characteristics: the model parameter samples in two adjacent rows have only one parameter value that is different, and the values of the other m-1 model parameters are exactly the same;

[0110] S5.4, calculate the basis effect EE of each model parameter i and its mean And the standard deviation:

[0111]

[0112] S5.5, the variance decomposition formula of the Sobol method calculation model is:

[0113]

[0114] Where, f=f(x1,x2,…,x m) is the structure of the model, V(f) is the total variance of the model output, V i is the variance term generated by the i-th parameter, V ij is the variance term generated by the interaction of the i-th and j-th parameters, V 1,2,…,m is the variance generated by all parameters;

[0115] S5.6, define the first-order sensitivity S i , second-order sensitivity S ij With the total sensitivity S Ti index:

[0116]

[0117] Among them, V -i represents the variance term without the effect of the i-th parameter, S i With S ij Respectively represent the impact of a single model parameter and the combination of two model parameters on the model output, S Ti represents the impact of all parameter combinations including the i-th model parameter on the model output, S Ti -S i Used to analyze the interaction effect between the i-th model parameter and other model parameters;

[0118] S5.7, based on the characteristics and requirements of hydrological model calculations, the Nash-Sutcliffe efficiency coefficient (NSE) is selected to characterize the model output results. The calculation formula for NSE is as follows:

[0119]

[0120] Among them, Q si is the simulated runoff (m 3 / s), Q oi is the measured runoff (m 3 / s), is the average measured runoff (m 3 / s), n is the sequence length.

[0121] S6. Compare the results of the sensitivity analysis of environmental parameters and surrogate model inputs. The sensitivity ranking of environmental parameters is consistent with the sensitivity ranking of the input parameters of the surrogate model, verifying the reliability of the surrogate model.

[0122] S7. Use the SCE-UA algorithm to optimize the model to determine the most suitable runoff model for each runoff zone and the most suitable confluence model for each sub-basin, and implement a hydrological multi-model combination optimization method based on environmental factor response and machine learning, which specifically includes the following sub-steps:

[0123] S7.1. Initialize the parameters of the SCE-UA algorithm: Set the number of complexes \(h\geq1\), the number of interior points in the complex \(m\geq n + 1\), and the sample size \(s=h\times m\).

[0124] S7.2. Generate sample points. Randomly sample \(s\) sample points \(x\) within the feasible parameter space i and calculate the function values \(f\) i = \(f(x\) i ). Generate samples using a uniform probability distribution.

[0125] S7.3. Sort the \(s\) sample points \((x\) i , \(f\) i ) in ascending order of function values, and denote it as the buffer \(D=\{(x\) i , \(f\) i ), \(i = 1,2,\cdots,s\}\).

[0126] S7.4. Divide the \(s\) sample points obtained by sampling into \(p\) complexes \(A_1,\cdots,A\) p :

[0127]

[0128] S7.5. According to the Competitive Complex Evolution (CCE) algorithm, perform evolution on each complex.

[0129] S7.6. Combine the vertices of all evolved complexes into a new sample group and sort them by function values.

[0130] S7.7. Check the reduced number of complexes. If the required minimum number of complexes \(p\) min in the population is \(< p\), remove the complex with the lowest ranked point or randomly remove a complex.

[0131] The experimental results show that the present invention provides a systematic and effective multi-model combination optimization method, which can optimize the combined use of hydrological models by deeply exploring the response relationship between environmental factors and model modules. Through a comprehensive analysis of multiple typical basins, this method not only improves the flexibility and adaptability of the model, but also ensures that high-precision hydrological simulation results can still be obtained under limited data and changing environmental conditions. In the future, this method is expected to play an important role in flood forecasting, watershed management and ecological protection, etc., providing a scientific basis and technical support for the sustainable utilization of water resources.

[0132] Although the present invention has been described above with preferred embodiments, it is not intended to limit the present invention. Those with ordinary knowledge in the technical field to which the present invention pertains can make various modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be determined by what is defined in the claims.

Claims

1. A hydrological multi-model combination optimization method based on environmental factor response and machine learning, characterized in that: The following steps are involved: S1. Obtain data on various environmental factors in the target area, including topography, soil type, soil moisture, climate zones, vegetation types, and humidity. Preprocess the data and divide the sub-basin into sub-basins based on the underlying surface conditions. S2. Calculate the terrain wetness index (TWI), standardized precipitation index (SPI), and runoff curve number index (CN) of the grid, perform normalization and weight distribution calculations, and divide the runoff generation zones according to the comprehensive score calculation results; S3. Based on the setup of multiple model modules, combined with hydrological statistics and numerical simulation, the distributed Xin'anjiang model, the terrain hydrological model, the soil and water assessment tool model, the variable infiltration capacity model, the Sacramento soil water balance model, and the Weather Research and Forecasting hydrological coupling model were split. Multiple runoff generation and confluence models were formed by combining multiple confluence methods such as Muskingum, diffusion wave, kinematic wave, and Congee algorithm. Random forests were used to construct alternative models for each runoff generation and confluence module. S4. Use the model interpretation method SHAP to interpret the alternative model and analyze the model's sensitive inputs; S5. Use Morris screening method to screen out sensitive environmental parameters, and then use Sobol method to further quantitatively analyze the sensitivity of the parameters; S6. Compare the results of the sensitivity analysis of environmental parameters and surrogate model inputs. The sensitivity ranking of environmental parameters is consistent with the sensitivity ranking of the input parameters of the surrogate model, verifying the reliability of the surrogate model. S7. Use the SCE-UA algorithm optimization model to determine the most suitable runoff model for each runoff zone and the confluence model for each sub-basin, and implement a hydrological multi-model combination optimization method based on environmental factor response and machine learning.

2. The method for optimizing the combination of multiple hydrological models based on environmental factor response and machine learning according to claim 1, wherein in step S2, calculating the standardized precipitation index (SPI) comprises the following sub-steps: A-2.

1. Calculate the probability density function of the gamma distribution of the rainfall x at the preset time scale: Where β and γ are the scale and shape parameters of the gamma distribution, respectively, β>0, γ>0; β and γ are obtained using the maximum likelihood estimation method, as follows: Where x i is the precipitation sample, is the average precipitation amount; A-2.

2. Using the monthly precipitation x0, calculate the probability of the event x < x0: A-2.

3. Calculate the cumulative probability of the preset time scale as follows: If the monthly precipitation is 0, the probability estimation method is as follows: In the formula, m represents the number of samples without precipitation, and n represents the total number of samples; A-2.

4. Normalize the cumulative probability. Use the standard normal distribution function to infer the exponential value of the obtained probability to obtain the SPI, as shown in the following formula: SPI=φ(P) Where φ is the standard normal distribution function, P is the cumulative probability value, and the larger the SPI, the wetter it is.

3. The method for optimizing the combination of multiple hydrological models based on environmental factor response and machine learning according to claim 1, characterized in that: The terrain wetness index TWI is calculated as follows: Where a is the upstream contribution area per unit profile length, and β is the slope.

4. The method for optimizing the combination of multiple hydrological models based on environmental factor response and machine learning according to claim 1, characterized in that: The runoff curve number CN ranges from 0 to 100 and is a comprehensive assessment of the preceding influencing rainfall, topography, surface land type, and soil properties.

5. The method for optimizing the combination of multiple hydrological models based on environmental factor response and machine learning according to claim 1, characterized in that: In step S2, the runoff zones are divided according to the comprehensive score calculation results, which specifically includes the following sub-steps: B-2.

1. Normalize the SPI, CN, and TWI of each grid to obtain SPI', CN', and TWI'; B-2.2, Construct a new indicator X XS , represents the wetness of the grid, S x Bigger means more moist: X XS =SPI’-0.5×CN’+0.5×TWI′ B-2.3, based on the calculated X XS The value divides each grid into 5 runoff zones X A 、X B 、X C 、X D and X E , mark it.

6. The method for optimizing the combination of multiple hydrological models based on environmental factor response and machine learning according to claim 1, characterized in that: In step 4, the model interpretation method SHAP explains the output of the machine learning model. The SHAP value is used to measure the contribution of each input feature to the predicted output. For a given model f and input sample x, the SHAP value of feature i The calculation is as follows: Where M is the input set, S is any subset that does not contain input i, |S| is the size of set S, |M| is the number of features, and f x (S∪{i})-f x (S) Model prediction results with and without feature i.

7. The method for optimizing the combination of multiple hydrological models based on environmental factor response and machine learning according to claim 1, characterized in that: Step S5 includes the following sub-steps: S5.

1. Determine the range of model parameters, the number of levels p and the disturbance factor Δ = 1 / (p-1); S5.2, for the parameter x i The values of are normalized and discretized: x i ∈{0,1 / (p-1),2 / (p-1),…,1-1 / (p-1)} S5.3, generate Morris's sample matrix B * : B * ={(J1x * +Δ / 2)[(2B-J m )D * +J m ]}P * Where B * is a (m+1)×m parameter sample matrix, each row represents a parameter sample, m is the number of model parameters; J1 is a (m+1)×1 matrix, J m is a (m+1)×m matrix, and all matrix elements are 1; B is a matrix (b ij ) (m+1)×m , when j≤i,i=1,2,…,m, b ij =1, otherwise 0; D * is an m×m random confusion matrix; the generated parameter sample matrix B * It has the following characteristics: the model parameter samples in two adjacent rows have only one parameter value that is different, and the values of the other m-1 model parameters are exactly the same; S5.4, calculate the basis effect EE of each model parameter i and its mean And the standard deviation: S5.5, the variance decomposition formula of the Sobol method calculation model is: Where, f=f(x1,x2,…,x m ) is the structure of the model, V(f) is the total variance of the model output, V i is the variance term generated by the i-th parameter, V ij is the variance term generated by the interaction of the i-th and j-th parameters, V 1,2,…,m is the variance generated by all parameters; S5.6, define the first-order sensitivity S i , second-order sensitivity S ij With the total sensitivity S Ti index: Among them, V -i represents the variance term without the effect of the i-th parameter, S i With S ij Respectively represent the impact of a single model parameter and the combination of two model parameters on the model output, S Ti represents the impact of all parameter combinations including the i-th model parameter on the model output, S Ti -S i Used to analyze the interaction effect between the i-th model parameter and other model parameters; S5.7, based on the characteristics and requirements of hydrological model calculations, the Nash-Sutcliffe efficiency coefficient (NSE) is selected to characterize the model output results. The calculation formula for NSE is as follows: Among them, Q si is the simulated runoff, Q oi is the measured runoff, is the average measured runoff, and n is the sequence length.

8. The method for optimizing the combination of multiple hydrological models based on environmental factor response and machine learning according to claim 1, characterized in that: Step S7 includes the following sub-steps: S7.

1. Initialize the parameters of the SCE-UA algorithm: set the number of complex shapes h ≥ 1, the number of complex internal points m ≥ n + 1, and the sample size s = h × m; S7.

2. Generate sample points and randomly sample s sample points x in the feasible parameter space. i And calculate the function value f i =f(x i ), use uniform probability distribution to generate samples; S7.3, for s sample points (x i ,f i ) are sorted in ascending order according to the function value, and are recorded as buffer D = {(x i ,f i ),i=1,2,…,s}; S7.

4. Divide the s sample points obtained by sampling into p composite shapes A1,…,A p : S7.5, performing evolution on each complex according to the competitive complex evolution CCE algorithm; S7.

6. After evolution, all vertices of the composite shape are combined into a new sample group and sorted by function value; S7.

7. Check the reduced number of complexes. If the minimum number of complexes p required in the population min < is less than p, remove the complex with the lowest ranking points or randomly remove a complex.