Hydrological parameter calibration and regionalization method based on proxy model and local observation guidance

By combining local observation data and diversified selection strategies within a unified proxy model framework, the computational efficiency of hydrological model parameter calibration and the local consistency and physical rationality of regionalization results are improved. This solves the problems of low computational efficiency and unstable regionalization results in existing methods and is applicable to hydrological model parameter calibration and regionalization under multi-basin conditions.

CN121092986APending Publication Date: 2025-12-09WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511084052.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing surrogate model methods suffer from low computational efficiency, weak spatial generalization ability, and poor regionalization rationality in hydrological model parameter calibration. They fail to effectively utilize local observation information, resulting in insufficient physical rationality and stability of regionalization results.

Method used

Within the framework of a unified agent model, diverse selection strategies are designed by combining local observation data. The machine learning agent model is trained using large sample data from multiple watersheds, and the parameters are optimized using evolutionary algorithms. The parameters are then regionalized in conjunction with local observation guidance.

Benefits of technology

It significantly improves the efficiency of parameter calibration and the local consistency and physical rationality of regionalized results, and is applicable to watersheds with no or sparse observations, thus expanding the applicability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121092986A_ABST
    Figure CN121092986A_ABST
Patent Text Reader

Abstract

The invention discloses a hydrological parameter calibration and regionalization method based on an agent model and local observation guidance, and relates to the field of large-scale watershed hydrological modeling. According to the method, parameter samples are generated in a plurality of observation drainage basins, a unified proxy model is trained in combination with geographic attributes, a candidate parameter set is generated by utilizing an evolutionary algorithm, and an optimal parameter combination is screened out from the candidate set by adopting a plurality of instructive selection strategies in combination with local observation data, so that regionalized inference of hydrological parameters is realized. Compared with an existing method, the method can utilize hydrological information of various observation types to assist parameter regionalization, remarkably improves the reasonability, stability and physical consistency of a regionalization result, is suitable for probabilistic and physical process hydrological models, and has the advantages of high efficiency, robustness, universality and expandability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of hydrological simulation and surface process modeling technology, specifically involving a hydrological parameter calibration and regionalization method based on surrogate models and local observation guidance. It is applicable to the efficient parameter calibration and reasonable regionalization inference of hydrological models under large-scale and multi-basin conditions, and is particularly suitable for watersheds with incomplete or scarce observation data. Background Technology

[0002] With the intensification of global climate change and the increasing demand for water resource management, hydrological models have become important tools in fields such as water security assessment, flood forecasting, and climate change impact analysis. Hydrological models can simulate key processes of the watershed hydrological cycle, have a clear physical basis and interpretability, and are widely used in scientific research and engineering practice.

[0003] However, hydrological models typically contain a large number of internal parameters that are difficult to observe or determine directly (such as soil hydraulic properties, vegetation cover influencing factors, and snow ablation coefficients), which have a significant impact on simulation accuracy. Due to the complexity of the model structure and the heavy computational burden, traditional single-basin independent calibration methods are difficult to extend to regionalized and large-scale conditions.

[0004] Existing methods mainly face the following problems: 1. Low computational efficiency: Running a highly complex physical model for parameter search in a single watershed results in a large amount of computation and a long time. 2. Weak spatial generalization ability: It fails to effectively utilize the common patterns of multi-basin data, and the calibration results are difficult to transfer to unobserved basins; 3. Poor regionalization rationality: The constraints of local observation information were not considered, and the regionalization results were insufficient in terms of physical rationality and local consistency. 4. Lack of a unified framework: The methods rely on empirical parameter tuning and local calibration, making it difficult to form a standardized and scalable modeling process.

[0005] In recent years, machine learning surrogate models have been introduced into the field of hydrological modeling to replace complex physical models for parameter sensitivity analysis and optimization searches, effectively improving computational efficiency. For example, some studies have attempted to use machine learning methods such as Gaussian process regression to construct surrogate models within a single watershed, accelerating parameter calibration. However, these methods typically model independently within each watershed, failing to fully utilize the hydrological similarities between multiple watersheds and making it difficult to generalize to other watersheds. A few studies have proposed training a unified surrogate model based on large-sample watershed data to achieve synchronous parameter calibration across multiple watersheds, the so-called "large-sample simulator" method. However, these methods still primarily focus on improving computational efficiency and calibration accuracy, with insufficient consideration given to the physical rationality, local consistency, and stability of regionalized results.

[0006] Especially against the backdrop of the rapid development of hydrogeographic big data, multi-source observation data (such as satellite remote sensing, snow water equivalent, soil moisture, etc.) are widely available. However, existing surrogate model-based methods have failed to effectively utilize these local observation information to guide regional inference, resulting in low reliability of regionalization results and limiting their ability to be promoted in practical applications. Summary of the Invention

[0007] To address the shortcomings of existing surrogate model methods (including the large-sample simulator LSE), which focus solely on improving computational efficiency and fail to fully utilize the constraints of local observation data, resulting in insufficient physical rationality and stability of regionalized results, this invention provides a hydrological parameter calibration and regionalization method based on surrogate models and local observation guidance. This method, within a unified surrogate model framework, combines local observation data and designs diverse selection strategies, effectively improving parameter calibration efficiency and the local consistency and physical rationality of regionalized results.

[0008] According to one aspect of the present invention, a method for calibrating and regionalizing hydrological parameters based on a surrogate model and local observations is provided, comprising: Step 1: Obtain multiple representative hydrological basins with runoff observation data, collect the static geographic attributes of each basin, and determine the parameters of the hydrological model to be calibrated and their value ranges. Step 2: In each watershed, multiple parameter combinations are generated in the parameter space using a random sampling method to construct an initial parameter sample set; Step 3: Input the initial parameter samples into the hydrological model for simulation, calculate the objective function value between the simulation results and the measured runoff data, and form a large sample training dataset of the correspondence between parameter combinations and performance indicators. Step 4: Construct a unified machine learning proxy model based on large sample training data. The proxy model takes parameter combinations and static geographic attributes as input and performance indicators as output to capture the pattern of how parameters affect simulation performance under different geographic backgrounds. Step 5: Based on the trained surrogate model, use an optimization algorithm to generate a candidate parameter set, and input the candidate parameter set into the hydrological model for simulation to obtain new parameter combinations and performance indicators, and supplement them to the large sample training dataset. Step 6: Repeat steps 4 and 5 to iteratively optimize the proxy model until the preset stopping criterion is reached; Step 7: Classify the uncalibrated target watersheds. For different categories of target watersheds, collect their static geographic attributes and determine the parameter range. Step 8: Implement different local observation-guided regional selection strategies for different types of target watersheds.

[0009] As a further technical solution, in step 7, the uncalibrated target watersheds are classified into categories, resulting in: Category 1 target watersheds: Watersheds with complete observation data but not included in the training watershed; The second type of target watershed: watersheds with partial observation data; The third type of target watershed: watersheds without observational data.

[0010] As a further technical solution, in step 8, different local observation-guided regional selection strategies are implemented for different types of target watersheds, including: For the first and second type of target watersheds, a candidate parameter set is generated using a surrogate model, and a fitting index is calculated based on local observation data. The optimal or suboptimal parameter combination that meets the local constraints is selected through a variety of guiding selection strategies.

[0011] As a further technical solution, guiding selection strategies for the first and second type of target watersheds include model evaluation indicators based on hydrological observation data, including but not limited to Nash efficiency coefficient, Kling-Gupta coefficient, root mean square error, absolute error, or a combination strategy; the hydrological observation data used includes but is not limited to multi-timescale flow data, climatological data, snow water equivalent, soil moisture, and total water storage.

[0012] As a further technical solution, in step 8, different local observation-guided regional selection strategies are implemented for different types of target watersheds, including: For the third type of target watershed, a candidate parameter set is generated using a surrogate model, and the optimal parameter combination is selected based on the objective function value predicted by the surrogate model.

[0013] As a further technical solution, in step 1, the static geographic attributes of each watershed are collected, including: We collect topographic indicators, climate statistics, land cover and soil types, and multi-year daily runoff observation data for each watershed, excluding statistical characteristics derived from measured runoff.

[0014] As a further technical solution, a unified machine learning agent model is constructed, including: Choose a machine learning proxy model from the random forest proxy model, artificial neural network proxy model, gradient boosting tree proxy model, or a combination of the above models.

[0015] As a further technical solution, the method also includes: New samples are added during each round of surrogate model training, and cross-validation is used to evaluate the performance of the surrogate model.

[0016] According to one aspect of the present invention, a hydrological parameter calibration and regionalization system based on a surrogate model and local observation guidance is provided for implementing the method, the system comprising: The first main module is used to acquire multiple representative hydrological basins with runoff observation data, collect the static geographic attributes of each basin, and determine the parameters of the hydrological model to be calibrated and their value ranges. The second main module is used to generate multiple parameter combinations in the parameter space using a random sampling method in each watershed, and to construct an initial parameter sample set; The third main module is used to input the initial parameter samples into the hydrological model for simulation, calculate the objective function value between the simulation results and the measured runoff data, and form a large sample training dataset of the correspondence between parameter combinations and performance indicators. The fourth main module is used to build a unified machine learning proxy model based on large sample training data. The proxy model takes parameter combinations and static geographic attributes as input and performance indicators as output to capture the pattern of how parameters affect simulation performance under different geographic backgrounds. The fifth main module is used to generate a candidate parameter set based on the trained surrogate model using an optimization algorithm, and input the candidate parameter set into the hydrological model for simulation to obtain new parameter combinations and performance indicators and supplement them to the large sample training dataset. The sixth main module is used to repeat main modules four and five, iteratively optimizing the proxy model until the preset stopping criteria are met; The seventh main module is used to classify uncalibrated target watersheds. For different categories of target watersheds, static geographic attributes are collected and parameter ranges are determined. The eighth main module is used to implement different local observation-guided regional selection strategies for different types of target watersheds.

[0017] According to one aspect of the present invention, a non-transitory computer-readable storage medium is provided, the non-transitory computer-readable storage medium storing computer instructions that cause the computer to execute the described hydrological parameter calibration and regionalization method based on surrogate models and local observation guidance.

[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention, by combining a unified proxy model and a local observation guidance strategy, significantly enhances the local consistency, physical rationality, and stability of regionalized results while improving computational efficiency. It effectively overcomes the shortcomings of existing methods based on LSE and other methods that fail to fully utilize local observation information, and has significant practical value and broad prospects for promotion. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This invention provides an overall technical flowchart for a hydrological parameter calibration and regionalization method based on a surrogate model and local observation guidance, illustrating the complete process of large sample construction, unified surrogate model training, evolutionary algorithm optimization, iterative updates, and inference for different watersheds.

[0021] Figure 2 This is a schematic diagram of the spatial distribution of 531 sample watersheds used in the embodiments of the present invention, which shows the geographical location and distribution characteristics of the multiple watersheds used to construct the unified surrogate model and perform five-fold cross-validation.

[0022] Figure 3 The graph shows the trend of the HBV hydrological model parameter calibration accuracy (NSE and KGE) increasing with the number of iterations of the surrogate model in Embodiment 1 of the present invention, which verifies the effectiveness of the iterative optimization framework.

[0023] Figure 4 This is a regional accuracy comparison chart of the Pd, Pm, and Py strategies based on runoff time series observation data used in Embodiment 1 of the present invention, which demonstrates that the local observation guidance strategy is superior to random selection (Prand).

[0024] Figure 5 The regional accuracy comparison chart for parameter selection guided by runoff climatological statistics (P50, Pavg, Pcomb) in Embodiment 2 of the present invention verifies the advantages of the guidance strategy based on climatological observation data compared with the lack of observation guidance (Prand, Pemu).

[0025] Figure 6 This is a regional accuracy comparison chart of the parameter selection guided by snow water equivalent (SWE) observation data in Embodiment 3 of the present invention, which shows that the simulation effect of the SWE guidance strategy in the uncalibrated watershed is better than the result without observation guidance. Detailed Implementation

[0026] To address the shortcomings of existing surrogate model methods (including the large-sample simulator LSE), which focus solely on improving computational efficiency and fail to fully utilize local observation data constraints, resulting in insufficient physical plausibility and stability of regionalized results, there is an urgent need to propose a method that, within a unified surrogate model framework, incorporates local observation data and designs diverse selection strategies to guide parameter regionalization inference. This would improve the local consistency, physical plausibility, and stability of regionalized results, and expand the model's applicability in watersheds with no or sparse observations. Therefore, this invention proposes a hydrological parameter calibration and regionalization method based on a surrogate model and local observation guidance. By combining a unified surrogate model with diverse local observation guidance strategies, it effectively improves parameter calibration efficiency and the local consistency and physical plausibility of regionalized results.

[0027] This invention proposes a unified technical process that integrates large-sample data from multiple watersheds, machine learning proxy modeling, evolutionary algorithm optimization, and local observation-guided selection strategies, such as... Figure 1 As shown, this includes global calibration, parameter selection guided by local observations, and unguided parameter regionalization. The figure illustrates the complete process of large-sample construction, unified surrogate model training, evolutionary algorithm optimization, iterative updates, and inference across different watersheds. Specifically, it includes the following steps: 1. Data Preparation and Parameter Definition: Several representative hydrological watersheds were selected, and static geographic attribute characteristics (such as topography, climate statistics, land cover, and soil type) and observational data (such as runoff, snowmelt equivalent, and soil moisture) were collected for each watershed for performance evaluation. The set of parameters to be calibrated and their value ranges were determined based on the model structure and sensitivity analysis. Static geographic attributes did not include statistical characteristics derived from measured runoff to ensure that the surrogate model could still be used for regional parameter inference in watersheds without observations.

[0028] 2. Initial Sample Construction: Multiple parameter combinations are generated in the parameter space using random sampling methods within each watershed. These combinations are then input into the hydrological model for simulation, and performance indicators (such as Nash efficiency coefficient NSE, Kling-Gupta coefficient KGE, etc.) are calculated to form a large-sample training dataset containing parameter values, geographical attributes, and performance indicators.

[0029] 3. Unified surrogate model training: Based on large-sample training data, a unified machine learning surrogate model is constructed to approximate the parameter-performance relationship. The surrogate model takes parameter combinations and watershed static attributes as input and performance indicators as output, capturing the patterns of parameter influence on simulation performance under different geographical backgrounds. The surrogate model is preferably a machine learning regression model, including random forest, gradient boosting tree, artificial neural network, support vector regression, and combinations thereof. The machine learning surrogate model is selected from random forest, artificial neural network, gradient boosting tree, or combinations thereof.

[0030] 4. Parameter calibration based on surrogate model: On the trained surrogate model, an evolutionary algorithm is used to search for global parameters. The optimization algorithm is selected from single-objective genetic algorithm, multi-objective evolutionary algorithm, non-dominated sorting genetic algorithm (NSGA-II) or other evolutionary algorithms. Candidate parameter sets are generated through different objective function configurations and initial conditions.

[0031] 5. Iterative Update of the Surrogate Model: The candidate parameter set is applied to the hydrological model for simulation. New performance samples are obtained and added to the training dataset. The surrogate model is then retrained, forming an iterative optimization mechanism until the preset convergence condition is met or the maximum number of iterations is reached. New samples are added in each round of surrogate model training, and cross-validation is used to evaluate the performance of the surrogate model to improve its generalization ability.

[0032] 6. Target Watershed Delineation and Regional Inference: The target watersheds are divided into three categories: a) watersheds with complete observation data but not included in the training set; b) watersheds with partial observation data (such as short-term runoff, water level, snow water equivalent, soil moisture, etc.); c) watersheds with no observation data. Geographic attributes are collected for each category, and parameter ranges are determined.

[0033] 7. Regionalized Selection Strategies Guided by Local Observations: For Class I and Class II watersheds, various guiding selection strategies (such as those based on daily-scale Pd, monthly-scale Pm, annual-scale Py, median P50, mean Pavg, and combined strategies Pcomb) are employed, combining candidate parameter sets output by the surrogate model with local observation data, to screen for optimal or suboptimal parameter combinations that satisfy local constraints, thereby improving the rationality and consistency of regionalization results. For watersheds without observations, the optimal parameters are selected by minimizing the objective function value predicted by the surrogate model. The guiding selection strategies for Class I and Class II target watersheds include model evaluation indicators based on observation data, including but not limited to Nash efficiency coefficient, Kling-Gupta coefficient, root mean square error, absolute error, or combined strategies; the hydrological observation data used include, but are not limited to, multi-timescale flow data, climatological data, snow water equivalent, soil moisture, and total water storage.

[0034] The objective function is used to measure the fitting accuracy between simulated runoff and measured runoff, and the objective function includes the Nash efficiency coefficient or other equivalent hydrological evaluation indicators.

[0035] 8. Regionalization Result Validation: The generalization ability of the regionalization results is evaluated through methods such as spatial cross-validation to verify the applicability and reliability of the method of the present invention in unobserved watersheds.

[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. In addition, the technical features of the various embodiments or individual embodiments provided by the present invention can be arbitrarily combined to form new technical solutions. Such combinations are not bound by the order of steps and / or structural composition patterns, but must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0037] Example 1 This embodiment uses 531 watersheds from the US CAMELS dataset as experimental areas to demonstrate the complete process of global calibration of hydrological model parameters based on a random forest surrogate model and local observation-guided regionalization methods. These datasets are public datasets used to verify the effectiveness of the method of this invention, and this invention is not limited to specific data sources. The specific steps are as follows: 1. Data Preparation 531 watersheds with runoff observation data were selected from the CAMELS dataset as training watersheds, and static geographic attribute features of each watershed were collected, including but not limited to: (1) Topographic indicators (such as elevation, slope, and topographic relief); (2) Climate statistics (such as multi-year average precipitation, evapotranspiration, and temperature); (3) Land cover type and soil type.

[0038] Simultaneously, multi-year diurnal runoff observation data for each watershed were collected for hydrological model performance evaluation. The hydrological model used was the HBV (Hydrologiska Byråns Vattenbalansavdelning) model. Based on the HBV structure and sensitivity analysis results, 12 key parameters requiring calibration and their value ranges were determined. This method is not dependent on a specific hydrological model and can also be applied to other similar models. The watersheds selected in this study are as follows: Figure 2 As shown.

[0039] 2. Initial Sample Construction Within each training watershed, 200 sets of parameter combinations were randomly generated in the parameter space using the Latin hypercube sampling method. Each combination was then input into the HBV model to drive the simulation. The Nash efficiency coefficient (NSE) between the simulated runoff and the measured runoff was calculated, forming a large-sample training dataset containing parameter combinations, geographical attributes, and performance indicators.

[0040] 3. Proxy Model Training Based on large-sample training data, a unified random forest surrogate model is trained to approximate the parameter-performance mapping relationship of the HBV model. The hyperparameters of the surrogate model are set as follows: Number of trees (n_estimators): 100; Maximum tree depth (max_depth): None; Minimum number of samples per leaf node (min_samples_leaf): 1; Feature selection mode (max_features): "auto".

[0041] The proxy model takes parameter combinations and static geographic attributes as input features and NSE index as output.

[0042] 4. Global parameter calibration and iterative update of surrogate model Based on the trained surrogate model, a genetic algorithm (GA) is used for global parameter calibration. In each iteration, 50 new candidate parameter combinations are generated, input into the HBV model for simulation, and then the new samples and corresponding NSEs are added to the training dataset to retrain the surrogate model. This iterative process continues until the surrogate model's performance converges or the maximum number of iterations is reached. Figure 3 As shown, the median values ​​of NSE and KGE gradually increase during the iteration process, eventually converging to 0.74 and 0.78, respectively.

[0043] 5. Guidance on Target Watershed Delineation and Local Observation Five-fold cross-validation was used to validate the uncalibrated watershed within the target area. The specific steps are as follows: (1) Divide the entire watershed into 5 equal sections. Each time, select 1 section (about 20%) as the target watershed for uncalibrated purposes and the remaining 4 sections as the training watershed. (2) After constructing a surrogate model based on the training watershed and performing global calibration iteration, 20 sets of candidate parameter combinations predicted by the surrogate model are used to simulate in each uncalibrated watershed to obtain 20 sets of simulated flow time series; (3) For each uncalibrated watershed, the Nash efficiency coefficient (NSE) is calculated as an accuracy index using limited observed flow data (time series). For each candidate parameter combination, the difference in NSE index between its simulated time series and observed time series at different time scales is calculated, and the optimal parameter combination is selected accordingly.

[0044] Specifically, the following three local observation guidance strategies are adopted: Pd (Daily Strategy): On a daily scale, calculate the NSE index of the simulated daily flow time series and the observed daily flow time series for each candidate parameter combination, and select the parameter combination with the highest NSE as the optimal solution.

[0045] Pm (Monthly Strategy): The observed and simulated flow sequences are aggregated into monthly-scale sequences, the NSE is calculated, and the parameter combination with the highest NSE is selected as the optimal solution.

[0046] Py (Yearly Strategy): The observed and simulated flow sequences are aggregated into annual-scale sequences, the NSE is calculated, and the parameter combination with the highest NSE is selected as the optimal solution.

[0047] In addition, a baseline strategy is set: Prand (random strategy): Randomly selects one of 20 candidate parameter combinations as the regionalization result, without using any local observation data for guidance, representing the traditional information-free regionalization method.

[0048] By comparing the regionalization effects of Pd, Pm, Py, and Prand in uncalibrated watersheds, such as... Figure 4 As shown in the figure. The results indicate that all three local observation-guided strategies are significantly superior to Prand, verifying that using local observation information for parameter selection can effectively improve the accuracy and rationality of regional inference.

[0049] 6. Technical Effects This embodiment verifies that the method of the present invention, which combines random forest proxy modeling with local observation guidance strategy, can significantly improve the efficiency of hydrological model parameter calibration and regionalization, without the need to run a large number of real HBV hydrological models, and shows good applicability and promotion capability in watersheds with no or sparse observations.

[0050] Example 2 This embodiment is based on the method framework described in Embodiment 1, the difference being that the local observation guidance stage uses hydrological statistics (P50, Pavg, Pcomb) as guidance strategies instead of time-series observation guidance strategies. This embodiment also uses 531 watersheds from the US CAMELS dataset as the experimental area. It should be noted that the method of this invention does not depend on a specific watershed cluster and can also be applied to similar watershed clusters in other regions. The specific steps are as follows: 1. Data Preparation We selected 531 watersheds with runoff observation data from the US CAMELS dataset as experimental subjects, collected static geographic attribute characteristics and multi-year diurnal runoff observation data for each watershed, and used them for model performance evaluation. The hydrological model used was the HBV model, and 12 key calibration parameters and their value ranges were determined.

[0051] 2. Initial Sample Construction In each training watershed, 200 sets of parameter combinations were randomly generated using the Latin hypercube sampling method. These combinations were then input into the HBV model for driving simulation, and the Nash efficiency coefficient (NSE) was calculated to form a large-sample training dataset.

[0052] 3. Proxy Model Training A unified random forest proxy model was trained based on the training dataset. The hyperparameter settings were the same as in Example 1: number of trees 100, maximum depth None, minimum number of leaf node samples 1, and feature selection mode "auto".

[0053] 4. Global parameter calibration and iterative update of surrogate model A genetic algorithm (GA) is used for global parameter calibration based on the surrogate model. Each iteration generates 50 new candidate parameter combinations, which are input into the HBV model for simulation and added to the training dataset to retrain the surrogate model. Through multiple iterations, the performance of the surrogate model gradually improves and converges.

[0054] 5. Target watershed verification and local observation guidance A five-fold cross-validation method was employed, with each fold designating approximately 20% of the watershed as uncalibrated and the remaining 80% as training watersheds. For each uncalibrated watershed, the optimal solution was selected from 20 candidate parameter combinations predicted by the trained surrogate model. The specific method is as follows: For each candidate parameter combination, it is input into the HBV model for driving simulation to obtain the simulated runoff sequence of the uncalibrated watershed, and the accuracy index is calculated based on the limited observation data of the uncalibrated watershed.

[0055] The following local observation-guided strategies were used for parameter selection: P50 Strategy: Calculate the annual median Q50 of the simulated flow time series. sim Compared with the observed median value Q50 obs The absolute difference is used to select the candidate parameter with the smallest difference as the optimal solution.

[0056] Pavg strategy: Calculate and simulate multi-year average flow Qavg sim Compared with the observed multi-year average flow rate Qavg obs The absolute difference is used to select the candidate parameter with the smallest difference as the optimal solution.

[0057] Pcomb Strategy: Construct a combined index based on multiple statistics (such as median, mean, and coefficient of variation). In this embodiment, the mean absolute error of Pavg and P50d is used, and the candidate parameter with the smallest overall difference is selected as the optimal solution.

[0058] The following non-local observation guidance strategy is also set as a benchmark for comparison: Prand strategy: Randomly select one from a set of 20 candidate parameters as the regionalization result.

[0059] Pemu strategy: directly select the parameter combination with the lowest predicted objective function value from the surrogate model as the regionalization result, without using local observation data constraints.

[0060] In addition, the same Pd strategy as in Example 1 was adopted to achieve a direct comparison of the accuracy of each strategy with that of Example 1.

[0061] The parameter combinations selected using the above strategies are input into the HBV model for simulation. Performance indicators such as NSE and KGE are calculated to evaluate the regionalization effect of each strategy. The accuracy comparison results of each strategy in the uncalibrated watershed are as follows: Figure 5 As shown in the figure, P50, Pavg, and Pcomb are all significantly better than Prand and Pemu, which lack local observation guidance, and Pcomb is better than P50 and Pavg, indicating that combining multiple statistics can further improve regionalization accuracy. Meanwhile, since P50, Pavg, and Pcomb only utilize climatological statistics data, their accuracy is slightly lower than Pd, which is based on daily series observations, verifying the impact of different local observation information on regionalization effectiveness.

[0062] 6. Technical Effects This embodiment demonstrates that by combining random forest surrogate modeling with a local observation guidance strategy based on hydrological statistics, the method of this invention can effectively improve the accuracy and rationality of parameter regionalization. Compared with a single statistical guidance strategy, the combination of multiple statistics (Pcomb) further improves the stability and accuracy of the regionalization results. Simultaneously, it verifies the differences in the regionalization inference effects of different types of local observation data, proving the flexibility and applicability of the method of this invention under different observation information availability conditions.

[0063] Example 3 This embodiment is based on the method framework described in Embodiment 1, with the difference being that the selected hydrological model is the process-based complex hydrological model SUMMA (Structure for Unifying Multiple Modeling Alternatives), and the local observation guidance step uses snow cover water equivalent (SWE) observation data, which is validated only on a subset of watersheds with SWE observation data. The SUMMA model has stronger physical process simulation capabilities and can simulate multiple key hydrological variables such as snow cover and hydrothermal energy, but its computational efficiency is lower than that of HBV.

[0064] 1. Data Preparation A subset of watersheds with SWE observation data from the US CAMELS dataset was selected as the experimental subjects. Static geographic attribute characteristics and multi-year SWE observation time series data for each watershed were collected for model performance evaluation. The hydrological model used was SUMMA. Based on the model structure and sensitivity analysis results, 13 key calibration parameters and their value ranges were determined.

[0065] 2. Initial Sample Construction In the SWE watershed subset, 200 sets of parameter combinations were randomly generated using the Latin hypercube sampling method, input into the SUMMA model for driving simulation, and the objective function value (standardized Nash efficiency coefficient NSE) was calculated to form a large sample training dataset.

[0066] 3. Proxy Model Training A unified random forest proxy model was trained based on the training dataset. The hyperparameter settings were similar to those in Example 1: number of trees 50, maximum depth None, minimum number of leaf node samples 1, and feature selection mode "auto".

[0067] 4. Global parameter calibration and iterative update of surrogate model A genetic algorithm (GA) is used for global parameter calibration based on the surrogate model. Each iteration generates 50 new candidate parameter combinations, which are input into the SUMMA model for simulation and added to the training dataset to retrain the surrogate model. Through multiple iterations, the performance of the surrogate model gradually improves and converges.

[0068] 5. Local parameter selection guided by SWE observations A five-fold cross-validation method was employed, with each fold selecting approximately 20% of the watersheds in the SWE watershed subset as uncalibrated watersheds and the remainder as training watersheds. For each uncalibrated watershed, the optimal solution was selected from a set of 20 candidate parameters predicted by the trained surrogate model. The specific method is as follows: For each candidate parameter combination, it is input into the SUMMA model for driving simulation, resulting in the simulated SWE time series of the uncalibrated watershed.

[0069] The Nash efficiency coefficient (NSE) is calculated based on the observed SWE time series, and the candidate parameter with the largest NSE is selected as the optimal solution.

[0070] The parameter combinations selected using the SWE-guided strategy are input into the SUMMA model for simulation. Performance metrics such as NSE of the SWE time series are calculated to evaluate the regionalization effect. Results are as follows: Figure 6 As shown, the SWE-guided strategy significantly outperforms the Prand and Pemu strategies, which do not employ local information guidance, in uncalibrated watersheds, verifying the effectiveness and rationality of using SWE local observation information for parameter selection.

[0071] 6. Technical Effects This embodiment demonstrates that by combining random forest proxy modeling and SWE local observation guidance strategy, the method of the present invention can effectively improve the accuracy and rationality of parameter regionalization under the SUMMA hydrological model based on complex processes. In particular, under conditions of multivariate simulation capability and low observation, it further verifies the applicability and generalizability of the method of the present invention.

[0072] The implementation of the various embodiments of the present invention is based on programmed processing through a device with processor functionality. Therefore, in practical engineering, the technical solutions and functions of the various embodiments of the present invention are encapsulated into various modules. Based on this reality, and building upon the above embodiments, the embodiments of the present invention provide a hydrological parameter calibration and regionalization system guided by a surrogate model and local observations. This system is used to execute a hydrological parameter calibration and regionalization method guided by a surrogate model and local observations as described in the above method embodiments.

[0073] The system comprises: a first main module for acquiring multiple representative hydrological basins with runoff observation data, collecting static geographic attributes of each basin, and determining the parameters of the hydrological model to be calibrated and their value ranges; a second main module for generating multiple parameter combinations in the parameter space using a random sampling method within each basin, constructing an initial parameter sample set; a third main module for inputting the initial parameter samples into the hydrological model for simulation, calculating the objective function value between the simulation results and measured runoff data, forming a large-sample training dataset corresponding to the parameter combinations and performance indicators; and a fourth main module for constructing a unified machine learning proxy model based on the large-sample training data, wherein the proxy model takes parameter combinations and static geographic attributes as input. The system uses performance indicators as outputs to capture the patterns of parameter influence on simulation performance under different geographical backgrounds; the fifth main module is used to generate candidate parameter sets based on the trained surrogate model using optimization algorithms, and input the candidate parameter sets into the hydrological model for simulation to obtain new parameter combinations and performance indicators, which are then added to the large-sample training dataset; the sixth main module is used to repeat the fourth and fifth main modules to iteratively optimize the surrogate model until a preset stopping criterion is reached; the seventh main module is used to classify uncalibrated target watersheds, and for different categories of target watersheds, their static geographical attributes are collected and parameter ranges are determined; the eighth main module is used to implement different local observation-guided regional selection strategies for different categories of target watersheds.

[0074] This invention provides a hydrological parameter calibration and regionalization system based on a surrogate model and local observation guidance. Addressing the problems of existing surrogate model methods (including the large-sample simulator LSE) which only focus on improving computational efficiency, fail to fully utilize local observation data constraints, and result in insufficient physical rationality and stability of regionalization results, this invention employs several modules within a unified surrogate model framework. By combining local observation data and designing diverse selection strategies, it effectively improves parameter calibration efficiency and the local consistency and physical rationality of regionalization results.

[0075] It should be noted that the system embodiments provided by the present invention are used not only to implement the methods in the above method embodiments, but also to implement the methods in other method embodiments provided by the present invention. The only difference is that corresponding functional modules are set. The principle is basically the same as that of the above system embodiments provided by the present invention. As long as those skilled in the art can improve the modules in the above system embodiments by referring to the specific technical solutions in other method embodiments and combining technical features to obtain corresponding technical means and technical solutions composed of these technical means, on the basis of the above system embodiments, and on the premise of ensuring the practicality of the technical solutions, they can obtain corresponding system-like embodiments for implementing the methods in other method-like embodiments.

[0076] Based on the same inventive concept as the foregoing embodiments, this embodiment of the invention also provides a non-transitory computer-readable storage medium storing computer instructions that cause the computer to execute the hydrological parameter calibration and regionalization method based on the surrogate model and local observation guidance.

[0077] In summary, this invention generates parameter samples across multiple observation basins, trains a unified surrogate model using geographic attributes, generates a candidate parameter set using an evolutionary algorithm, and employs various guided selection strategies based on local observation data to select the optimal parameter combination from the candidate set, thereby achieving regional inference of hydrological parameters. Compared with existing methods, this invention can utilize hydrological information from various observation types to assist in parameter regionalization, significantly improving the rationality, stability, and physical consistency of the regionalization results. It is applicable to probabilistic and physical process hydrological models and has the advantages of high efficiency, robustness, versatility, and scalability.

[0078] The terms “comprising” and “having”, and any variations thereof, in the specification, claims, and accompanying drawings of this invention are intended to cover a non-exclusive inclusion, such as a process, method, system, product, or apparatus that includes a series of steps or units, not necessarily limited to those explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.

Claims

1. A method for calibrating and regionalizing hydrological parameters based on surrogate models and local observations, characterized in that, include: Step 1: Obtain multiple representative hydrological basins with runoff observation data, collect the static geographic attributes of each basin, and determine the parameters of the hydrological model to be calibrated and their value ranges. Step 2: In each watershed, multiple parameter combinations are generated in the parameter space using a random sampling method to construct an initial parameter sample set; Step 3: Input the initial parameter samples into the hydrological model for simulation, calculate the objective function value between the simulation results and the measured runoff data, and form a large sample training dataset of the correspondence between parameter combinations and performance indicators. Step 4: Construct a unified machine learning proxy model based on large sample training data. The proxy model takes parameter combinations and static geographic attributes as input and performance indicators as output to capture the pattern of how parameters affect simulation performance under different geographic backgrounds. Step 5: Based on the trained surrogate model, use an optimization algorithm to generate a candidate parameter set, and input the candidate parameter set into the hydrological model for simulation to obtain new parameter combinations and performance indicators, and supplement them to the large sample training dataset. Step 6: Repeat steps 4 and 5 to iteratively optimize the proxy model until the preset stopping criterion is reached; Step 7: Classify the uncalibrated target watersheds. For different categories of target watersheds, collect their static geographic attributes and determine the parameter range. Step 8: Implement different local observation-guided regional selection strategies for different types of target watersheds.

2. The hydrological parameter calibration and regionalization method based on a surrogate model and local observation guidance as described in claim 1, characterized in that, In step 7, the uncalibrated target watersheds are classified into categories, resulting in: Category 1 target watersheds: Watersheds with complete observation data but not included in the training watershed; The second type of target watershed: watersheds with partial observation data; The third type of target watershed: watersheds without observational data.

3. The hydrological parameter calibration and regionalization method based on a surrogate model and local observation guidance as described in claim 2, characterized in that, In step 8, different local observation-guided regional selection strategies are implemented for different categories of target watersheds, including: For the first and second type of target watersheds, a candidate parameter set is generated using a surrogate model, and a fitting index is calculated based on local observation data. The optimal or suboptimal parameter combination that meets the local constraints is selected through a variety of guiding selection strategies.

4. The hydrological parameter calibration and regionalization method based on a surrogate model and local observation guidance as described in claim 3, characterized in that, Guiding selection strategies for the first and second category of target watersheds include model evaluation indicators based on hydrological observation data, including but not limited to Nash efficiency coefficient, Kling-Gupta coefficient, root mean square error, absolute error, or a combination of these strategies; the hydrological observation data used include but are not limited to multi-timescale flow data, climatological data, snow water equivalent, soil moisture, and total water storage.

5. The hydrological parameter calibration and regionalization method based on a surrogate model and local observation guidance as described in claim 2, characterized in that, In step 8, different local observation-guided regional selection strategies are implemented for different categories of target watersheds, including: For the third type of target watershed, a candidate parameter set is generated using a surrogate model, and the optimal parameter combination is selected based on the objective function value predicted by the surrogate model.

6. The hydrological parameter calibration and regionalization method based on a surrogate model and local observation guidance as described in claim 1, characterized in that, In step 1, static geographic attributes of each watershed are collected, including: We collect topographic indicators, climate statistics, land cover and soil types, and multi-year daily runoff observation data for each watershed, excluding statistical characteristics derived from measured runoff.

7. The hydrological parameter calibration and regionalization method based on a surrogate model and local observation guidance as described in claim 1, characterized in that, Construct a unified machine learning agent model, including: Choose a machine learning proxy model from the random forest proxy model, artificial neural network proxy model, gradient boosting tree proxy model, or a combination of the above models.

8. The hydrological parameter calibration and regionalization method based on a surrogate model and local observation guidance as described in claim 1, characterized in that, The method further includes: New samples are added during each round of surrogate model training, and cross-validation is used to evaluate the performance of the surrogate model.

9. A hydrological parameter calibration and regionalization system based on a surrogate model and local observation guidance, used to implement the method described in any one of claims 1 to 8, characterized in that, The system includes: The first main module is used to acquire multiple representative hydrological basins with runoff observation data, collect the static geographic attributes of each basin, and determine the parameters of the hydrological model to be calibrated and their value ranges. The second main module is used to generate multiple parameter combinations in the parameter space using a random sampling method in each watershed, and to construct an initial parameter sample set; The third main module is used to input the initial parameter samples into the hydrological model for simulation, calculate the objective function value between the simulation results and the measured runoff data, and form a large sample training dataset of the correspondence between parameter combinations and performance indicators. The fourth main module is used to build a unified machine learning proxy model based on large sample training data. The proxy model takes parameter combinations and static geographic attributes as input and performance indicators as output to capture the pattern of how parameters affect simulation performance under different geographic backgrounds. The fifth main module is used to generate a candidate parameter set based on the trained surrogate model using an optimization algorithm, and input the candidate parameter set into the hydrological model for simulation to obtain new parameter combinations and performance indicators and supplement them to the large sample training dataset. The sixth main module is used to repeat main modules four and five, iteratively optimizing the proxy model until the preset stopping criteria are met; The seventh main module is used to classify uncalibrated target watersheds. For different categories of target watersheds, static geographic attributes are collected and parameter ranges are determined. The eighth main module is used to implement different local observation-guided regional selection strategies for different types of target watersheds.

10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions that cause the computer to execute the hydrological parameter calibration and regionalization method based on a surrogate model and local observation guidance as described in any one of claims 1 to 8.