Partitioned cooperative reinforcement learning parameter calibration method for semi-distributed hydrological model

CN122839804APending Publication Date: 2026-09-29BEIJING JINSHUI INFORMATION TECH DEV CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610919224.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]为此,本发明所要解决的技术问题在于克服现有技术中半分布式水文模型参数率定时参数维度高、空间异质性表达不足、上下游模拟不协调以及产流参数与路由参数割裂的问题

Benefits of technology

本发明所述的面向半分布式水文模型的分区协同强化学习参数率定方法,本发明通过构建站点上下游有向拓扑图并确定计算顺序,划分产流分区并建立站点与分区的对应关系,构建包含全局共享参数、分区独立参数和河道路由参数的分层参数空间,有效降低了半分布式水文模型高维参数率定的难度并保留了流域空间异质性;在此基础上,按照所述有向拓扑图所确定的顺序运行模型,将各站点的本地产流过程与上游站点经河道路由后的来水过程进行耦合,显著增强了上下游站点模拟的一致性和物理合理性;进而基于出口站、关键控制站和分区代表站计算多控制站协同奖励,并利用强化学习智能体根据所述协同奖励更新策略对三类参数进行协同率定,避免了产流参数与路由参数的割裂优化,实现了流域整体模拟精度和关键控制站点模拟效果的协同提升。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122839804A_ABST
    Figure CN122839804A_ABST
Patent Text Reader

Abstract

The application discloses a partitioned cooperative reinforcement learning parameter calibration method for a semi-distributed hydrological model. The method comprises the following steps: constructing a directed topological graph of upstream and downstream stations and determining a calculation sequence; dividing a basin into multiple runoff generation partitions, establishing a correspondence between stations and partitions, and constructing a hierarchical parameter space containing global shared parameters, partition independent parameters and river routing parameters; running the model according to the topological sequence, coupling the local runoff of each station with the upstream inflow after river routing to obtain a simulated flow process; calculating the multi-control station cooperative reward based on the outlet station, the key control station and the representative station of the partition, and updating the strategy of the three types of parameters according to the reward by using the reinforcement learning agent for cooperative calibration. The application can reduce the difficulty of high-dimensional parameter calibration of the semi-distributed model, enhance the consistency of upstream and downstream simulation, avoid the separation of runoff and routing parameters, and improve the overall simulation accuracy of the basin and the simulation effect of the key control station.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hydrological model parameter calibration technology, and in particular to a method, apparatus, equipment, and computer storage medium for calibrating parameters of semi-distributed hydrological models using partitioned collaborative reinforcement learning. Background Technology

[0002] Hydrological model parameter calibration is a crucial technical step in hydrological forecasting, flood control scheduling, and the construction of digital twin watersheds, and its accuracy directly impacts the scientific validity of watershed water resources management decisions. Among related technologies, a parameter calibration system ranging from single-site to multi-site calibration has been constructed using manual trial-and-error methods, heuristic optimization algorithms (such as genetic algorithms and SCE-UA), and intelligent optimization methods based on reinforcement learning. Specifically, existing technologies cover the entire process from constructing the parameter search space to evaluating model simulation results, including key aspects such as setting upper and lower limits for parameters, optimizing the objective function, and the interaction between the agent and the environment.

[0003] However, existing parameter calibration methods often involve simply evaluating multiple stations side-by-side or searching all parameters as ordinary continuous vectors. This approach does not fully consider the spatial structure of the semi-distributed hydrological model, which may result in high parameter dimensionality, a large search space, or incoordination between upstream and downstream station simulations. Consequently, it may affect the synergistic optimization of runoff generation parameters and routing parameters, and ultimately limit the improvement of the overall watershed simulation accuracy. Summary of the Invention

[0004] Therefore, the technical problem to be solved by the present invention is to overcome the problems of high parameter dimensionality, insufficient expression of spatial heterogeneity, lack of coordination between upstream and downstream simulation, and separation of runoff generation parameters and routing parameters in the existing semi-distributed hydrological model parameters.

[0005] To address the aforementioned technical problems, this invention provides a partitioned collaborative reinforcement learning parameter calibration method for semi-distributed hydrological models, comprising:

[0006] Obtain station information, river network topology, and hydrological and meteorological time series of the target watershed; construct a directed topology graph of the upstream and downstream stations; and determine the calculation order of the semi-distributed model from upstream to downstream based on the directed topology graph. Based on the watershed's river system structure and station distribution, the target watershed is divided into multiple runoff-producing zones. The correspondence between each hydrological station and the runoff-producing zone is established, and a hierarchical parameter space containing globally shared parameters, zone-independent parameters, and river routing parameters is constructed. According to the calculation order determined by the upstream and downstream directed topology graph of the station, the semi-distributed hydrological model is run sequentially. For the current station, its local runoff process is coupled with the inflow process after the river routing calculation of all directly upstream stations through the corresponding river routing parameters to obtain the simulated flow process of the current station. Based on the simulated and measured flow processes of exit stations, key control stations, and representative stations in different zones, a multi-control station collaborative reward is calculated to comprehensively evaluate the simulation effect of each control station. A reinforcement learning agent is then used to update its policy based on the collaborative reward, and the global shared parameters, zone-independent parameters, and river routing parameters are collaboratively calibrated.

[0007] Preferably, the construction of a directed topology graph of the upstream and downstream sites, and the determination of the semi-distributed model computation order from upstream to downstream based on the directed topology graph, specifically includes: Obtain the station number, control area, sub-basin attributes, rainfall time series, evaporation time series, and measured flow time series of multiple hydrological stations within the target watershed, as well as the connection relationships between upstream stations, downstream stations, and river segments; Based on the river segment connection relationship, a directed topology graph of upstream and downstream stations is constructed with hydrological stations as nodes and river channel connections between stations as directed edges. The stations in the basin are then topologically sorted according to the directed topology graph to obtain the hydrological simulation order from upstream to downstream.

[0008] Preferably, the step of dividing the target watershed into multiple runoff-producing zones and establishing a correspondence between each hydrological station and the runoff-producing zones specifically includes: Based on the watershed's river system structure, station distribution, sub-watershed attributes, control section locations, and hydrological response characteristics, the target watershed is divided into multiple runoff-generating zones, and a station-zone correspondence table between hydrological stations and runoff-generating zones is established. Specifically, for multiple sites within the same flow generation zone, some zone parameters are shared; for different flow generation zones, independent zone parameters are set.

[0009] Preferably, the hierarchical parameter space includes: globally shared parameters for characterizing the overall hydrological response characteristics of the target watershed; zone-independent parameters for characterizing the differences in water storage capacity and runoff response between different runoff-producing zones; and river routing parameters for characterizing the characteristics of flow propagation, lag, peak reduction, and confluence in different river sections. The globally shared parameters include one or more of evapotranspiration conversion parameters, runoff curve parameters, runoff distribution parameters, groundwater recession parameters, and interflow recession parameters. The zone-independent parameters include one or more of tension water capacity, upper layer water storage capacity, lower layer water storage capacity, free water capacity, and soil water storage parameters. The river routing parameters include one or more of water storage time constants, weighting coefficients, river section distance-related parameters, wave velocity-related parameters, and other parameters for characterizing the flow propagation characteristics of river sections.

[0010] Preferably, the semi-distributed hydrological model is run sequentially according to the calculation order determined by the upstream and downstream directed topology graph of the station. For the current station, its local runoff process is coupled with the inflow process calculated by the corresponding river routing parameters of all directly upstream stations to obtain the simulated flow process of the current station, specifically including: For the current site, its runoff generation zone is determined according to the site-zone correspondence table, and the globally shared parameters and the zone-independent parameters corresponding to the zone are extracted. The local runoff generation process of the current site is calculated based on the rainfall time series, evaporation time series, watershed area and model parameters. Obtain the set of direct upstream stations of the current station. For any upstream station, obtain its simulated outflow process and calculate the river routing based on the river routing parameters corresponding to the connected river segment to obtain the upstream inflow process after routing. The local runoff generation process is superimposed and coupled with the upstream water inflow processes after all routing to obtain the simulated flow process of the current site; The river routing calculation adopts the Muskingan method, which determines the calculation coefficients through the water storage time constant, weight coefficients and calculation step size, so as to characterize the functional relationship between the current inflow, the previous inflow and the previous outflow and the current outflow.

[0011] Preferably, the calculation of the multi-control station collaborative reward, which comprehensively evaluates the simulation effect of each control station, based on the simulated flow process and the measured flow process of the exit station, key control station, and regional representative station, specifically includes: For each control station, the Nash efficiency coefficient, water volume error, flood peak error, peak occurrence time error, and upstream-downstream consistency error are calculated based on its simulated flow sequence and measured flow sequence. Each error index is then converted into a positive reward through an exponential or linear function. Based on the preset weights of each evaluation indicator and the weight of each site, the converted positive rewards are weighted and summed to obtain the collaborative rewards for multiple control stations. Among the site weights, the weight of the exit station is greater than that of the key control station, and the weight of the key control station is greater than that of the regional representative station.

[0012] Preferably, the method further includes the following steps: Before co-calibrating the global shared parameters, partition-independent parameters, and river routing parameters, an initialization search phase is performed. Several initial parameter combinations are generated within the parameter feasible region through random sampling, Latin hypercube sampling, historical parameter hot start, heuristic algorithms, or empirical parameter range screening. The parameter region with better performance is selected as the initial parameter space for formal reinforcement learning training. Furthermore, the process of co-calibrating the global shared parameters, partition-independent parameters, and river routing parameters is divided into a phased co-calibration phase. The phased co-calibration phase includes: a first phase, calibrating the global shared parameters; a second phase, calibrating the partition-independent parameters; a third phase, calibrating the river routing parameters; and a fourth phase, performing joint fine-tuning of all parameters. In each phase, the reinforcement learning agent only outputs actions to the parameter subspace corresponding to the current phase, while the remaining parameters remain fixed or are only allowed to be fine-tuned within a small range. After each phase, the optimal parameters of that phase are used as the initial parameters for the next phase.

[0013] The present invention also includes a partitioned collaborative reinforcement learning parameter calibration device for semi-distributed hydrological models, comprising: The data acquisition and topology construction module is used to acquire station information, river network topology, and hydrological and meteorological time series of the target watershed, construct a directed topology graph of the upstream and downstream of the station, and determine the semi-distributed model calculation order from upstream to downstream based on the directed topology graph; The partitioning and parameter space construction module is used to divide the target watershed into multiple runoff-producing zones according to the watershed system structure and station distribution, establish the correspondence between each hydrological station and the runoff-producing zone, and construct a hierarchical parameter space containing globally shared parameters, zone-independent parameters, and river routing parameters. The model running module is used to run the semi-distributed hydrological model sequentially according to the calculation order determined by the upstream and downstream directed topology graph of the station. For the current station, its local runoff process is coupled with the inflow process after the river routing calculation of all directly upstream stations through the corresponding river routing parameters to obtain the simulated flow process of the current station. The collaborative reward calculation module is used to calculate the collaborative reward of multiple control stations based on the simulated and measured flow processes of the exit station, key control station, and regional representative station, which can comprehensively evaluate the simulation effect of each control station. The reinforcement learning module is used to update the policy of the reinforcement learning agent based on the cooperative reward and to perform cooperative calibration on the global shared parameters, partition-independent parameters and river routing parameters.

[0014] This invention also provides a partitioned collaborative reinforcement learning parameter calibration device for semi-distributed hydrological models, comprising: Memory, used to store computer programs; The processor is used to implement the steps of the above-described method for calibrating parameters of a partitioned collaborative reinforcement learning model for a semi-distributed hydrological model when executing the computer program.

[0015] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described method for calibrating parameters of a partitioned collaborative reinforcement learning model for a semi-distributed hydrological model.

[0016] The technical solution of the present invention has the following advantages compared with the prior art: The present invention describes a partitioned collaborative reinforcement learning parameter calibration method for semi-distributed hydrological models. This method constructs a directed topology graph of upstream and downstream stations and determines the calculation order. It divides runoff-generating zones and establishes a correspondence between stations and zones, constructing a hierarchical parameter space containing globally shared parameters, zone-specific parameters, and river routing parameters. This effectively reduces the difficulty of high-dimensional parameter calibration for semi-distributed hydrological models while preserving the spatial heterogeneity of the watershed. Based on this, the model is run according to the order determined by the directed topology graph, coupling the local runoff-generating process of each station with the inflow process of upstream stations via river routing. This significantly enhances the consistency and physical rationality of the simulation of upstream and downstream stations. Furthermore, it calculates multi-control station collaborative rewards based on outlet stations, key control stations, and representative stations of each zone. A reinforcement learning agent is then used to collaboratively calibrate the three types of parameters according to the collaborative reward update strategy, avoiding the fragmented optimization of runoff-generating parameters and routing parameters. This achieves a synergistic improvement in the overall watershed simulation accuracy and the simulation effect of key control stations. Attached Figure Description

[0017] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings, wherein: Figure 1 This is a flowchart illustrating the implementation of a partitioned collaborative reinforcement learning parameter calibration method for semi-distributed hydrological models provided by this invention. Figure 2 This is a flowchart of a partitioned collaborative reinforcement learning parameter calibration method for a semi-distributed hydrological model according to an embodiment of the present invention; Figure 3This is a schematic diagram of the topological relationship between stations, zones, and river segments; Figure 4 A schematic diagram of the hierarchical parameter space of a semi-distributed hydrological model; Figure 5 A schematic diagram illustrating the interaction between a reinforcement learning agent and a semi-distributed hydrological model; Figure 6 This is a structural block diagram of a partitioned collaborative reinforcement learning parameter calibration device for a semi-distributed hydrological model, provided in an embodiment of the present invention. Detailed Implementation

[0018] The core of this invention is to provide a method, apparatus, device, and computer storage medium for calibrating parameters of a semi-distributed hydrological model through regional collaborative reinforcement learning, so as to achieve collaborative calibration of globally shared parameters, regionally independent parameters, and river routing parameters.

[0019] To enable those skilled in the art to better understand the present invention, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Please refer to Figure 1 and Figure 2 , Figure 1 This is a flowchart illustrating the implementation of a partitioned collaborative reinforcement learning parameter calibration method for semi-distributed hydrological models provided by the present invention. Figure 2 This is a flowchart of a partitioned collaborative reinforcement learning parameter calibration method for a semi-distributed hydrological model according to an embodiment of the present invention; the specific operation steps are as follows: S101: Obtain station information, river network topology, and hydrological and meteorological time series of the target watershed; construct a directed topology graph of the upstream and downstream stations; and determine the semi-distributed model calculation order from upstream to downstream based on the directed topology graph. Obtaining basic data for the target watershed is a prerequisite for parameter calibration, aiming to provide a spatial structural framework and time-series driving force for subsequent model construction and calculation. This step first requires collecting station information for each hydrological station within the target watershed. This information includes at least a unique station number and basic attributes such as the control area or local catchment area corresponding to each station. Simultaneously, the river network connectivity of the watershed needs to be obtained. This relationship defines the upstream and downstream connections and river segment affiliations between stations, forming the basis for constructing the watershed's spatial topology. Furthermore, the hydrological and meteorological time series required to drive the semi-distributed hydrological model need to be obtained. This series includes at least rainfall time series, evaporation or potential evapotranspiration time series, and measured flow time series for subsequent evaluation. Based on the obtained river network connectivity, a directed upstream and downstream topology graph representing the direction of water flow propagation between stations is constructed. This directed topology graph uses hydrological stations as nodes and river connections between stations as directed edges. If a station is upstream of another station and there is a river connection between them, a directed edge is constructed from the upstream station to the downstream station. Based on this, a topological sort of all stations within the watershed is performed according to the directed topological graph, thereby determining a hydrological simulation calculation order from upstream to downstream that conforms to the physical propagation laws of water flow. As one implementation method, this directed topological graph can be represented as follows: ,in This represents a collection of hydrological stations. This represents the set of river segments connected, if the stations Located at the station If the upstream of the two rivers is connected by a river channel, then a directed edge is constructed. Based on this, the watershed stations were topologically sorted to obtain the hydrological simulation order from upstream to downstream.

[0021] By constructing a directed topology graph of upstream and downstream stations and determining the calculation order from upstream to downstream, a rigorous spatial calculation logic is provided for the semi-distributed hydrological model. This ensures that the outflow from upstream stations can be transmitted to downstream stations through river channels according to physical laws, thus laying a structural foundation for the subsequent coupled calculation of runoff generation and routing, as well as ensuring the consistency of upstream and downstream simulations.

[0022] S102: Based on the watershed system structure and station distribution, the target watershed is divided into multiple runoff-producing zones, the correspondence between each hydrological station and the runoff-producing zones is established, and a hierarchical parameter space containing globally shared parameters, zone-independent parameters, and river routing parameters is constructed. Based on the watershed's river system structure, station distribution, and hydrological response characteristics, the target watershed is divided into multiple runoff-generating zones. Each runoff-generating zone contains one or more hydrological stations, and a correspondence between hydrological stations and runoff-generating zones is established. This correspondence is used in subsequent model calculations to allow each station to access specific parameters of its respective zone. Building upon this, a hierarchical parameter space for the semi-distributed hydrological model is constructed. This space is organized according to the physical scope of the parameters and includes three categories: globally shared parameters, zone-independent parameters, and river routing parameters. Globally shared parameters characterize the overall hydrological response characteristics of the target watershed and are used uniformly across the entire watershed. Zone-independent parameters characterize the spatial differences in water storage capacity and runoff response among different runoff-generating zones; stations within the same zone share some zone parameters, while parameters between different zones are independent. River routing parameters characterize the flow propagation, lag, and peak-shaving characteristics of different river sections. Through the construction of this hierarchical parameter space, the high-dimensional parameters of the semi-distributed hydrological model are structurally organized according to their physical meaning and scope of application, thereby reducing the difficulty of parameter calibration while preserving the spatial heterogeneity within the watershed. As one implementation method, a site-zone correspondence table can be predefined, and several upstream tributary sites can be configured as the first runoff-producing zone, while the sites near the confluence area of ​​the main stream and the outlet can be configured as another runoff-producing zone. A globally shared parameter set, a zone-independent parameter set for each runoff-producing zone, and a river routing parameter set for each river segment can be constructed respectively.

[0023] This step establishes a mapping relationship between runoff-producing zones and stations, and constructs a hierarchical parameter space. This enables a structured organization of parameters for the semi-distributed hydrological model, reduces the difficulty of searching high-dimensional parameter spaces, and preserves the ability to express spatial differences between different regions, thus providing a foundation for subsequent zone-based coordinated calibration.

[0024] S103: According to the calculation order determined by the upstream and downstream directed topology graph of the station, the semi-distributed hydrological model is run sequentially. For the current station, its local runoff process is coupled with the inflow process after the river routing calculation of all directly upstream stations through the corresponding river routing parameters to obtain the simulated flow process of the current station. Following the directed topological order of upstream and downstream stations, a semi-distributed hydrological model is run sequentially to couple the local runoff generation process of each station with the inflow process of all directly upstream stations after corresponding river routing, thus obtaining the simulated flow process of each station. Specifically, this method simulates each hydrological station in the basin sequentially according to the calculation order from upstream to downstream determined by the pre-constructed directed topological graph of the upstream and downstream stations. For any current station, firstly, based on its runoff generation zone, the corresponding zone-independent parameters and globally shared parameters are called, and combined with the station's local catchment area, rainfall time series, and evaporation time series, the local runoff generation process of the station is calculated through the hydrological model. At the same time, the simulated outflow processes of all directly upstream stations of the station are obtained, and for each directly upstream station, the river routing process of the outflow process of the upstream station is calculated according to the river routing parameters corresponding to the river segment connecting the upstream station and the current station, in order to simulate the propagation, lag, and peak-shaving effects of water flow in the river channel, thereby obtaining the upstream inflow process after routing. Finally, the local runoff process at the current station is superimposed and coupled with the inflow processes of all directly upstream stations after river routing, thus obtaining the complete simulated flow process at the current station. As one implementation method, when the Muskingan method is used for river routing, the route calculation can be expressed as... ,in Indicates the current moment of entry into the stream. Indicates the outflow at the current moment. The calculation coefficients are determined by the water storage time constant, weighting coefficients, and calculation step size. By performing the above-mentioned coupled calculations of runoff generation and routing station by station in topological order, the simulated flow process of each station in the entire watershed can be obtained.

[0025] This technical step, by explicitly performing coupled calculations of runoff generation and routing according to the directional topological order of upstream and downstream stations, ensures that the outflow process of upstream stations can accurately participate in the flow simulation of downstream stations after being routed by river roads. This effectively solves the problem of incoordination in the simulation process caused by ignoring the hydraulic connection between upstream and downstream, and significantly improves the consistency and physical rationality of the simulation results of the semi-distributed hydrological model at the watershed scale.

[0026] S104: Based on the simulated and measured flow processes of the exit station, key control station, and representative station in the region, calculate the multi-control station collaborative reward that can comprehensively evaluate the simulation effect of each control station, and use the reinforcement learning agent to update its policy according to the collaborative reward to perform collaborative calibration on the global shared parameters, regional independent parameters, and river routing parameters.

[0027] Based on the simulated and measured flow processes of exit stations, key control stations, and representative stations in different regions, a multi-control station collaborative reward is calculated. A reinforcement learning agent is then used to collaboratively calibrate globally shared parameters, region-specific parameters, and river routing parameters according to the collaborative reward update strategy. Specifically, this method first constructs a multi-control station collaborative reward function. This function comprehensively evaluates the simulation performance of each control station, which includes at least exit stations, key control stations, and representative stations in different regions. By comparing the simulated and measured flow processes of each station, evaluation indicators reflecting simulation accuracy are calculated and converted into positive reward values. These indicators are then assigned corresponding weights based on the importance of each control station, and finally, a weighted sum is obtained to obtain the global collaborative reward. This reward function not only focuses on the simulation performance of exit stations but also considers the simulation accuracy of key control stations and representative stations in different regions, thereby guiding the reinforcement learning agent to simultaneously consider the simulation consistency of multiple key locations within the watershed during the optimization process. The reinforcement learning agent outputs parameters and actions based on the current state, runs a semi-distributed hydrological model to obtain the next state and the cooperative reward, and uses this reward to update its policy network and value network. Through repeated iterations, the agent gradually learns the parameter combination that optimizes the overall simulation effect of multiple control stations. As one implementation method, the evaluation indicators may include one or more of the following: Nash efficiency coefficient, water volume error, flood peak error, peak occurrence time error, and upstream-downstream consistency error. Each indicator is converted into a positive reward through an exponential or linear function, and a minimum reward cutoff value is set to avoid the adverse effects of negative rewards on training. The weights of each control station can be set such that the weight of the outlet station is greater than that of the key control station, and the weight of the key control station is greater than that of the representative station in the region, to reflect the differences in importance of different stations in watershed hydrological control.

[0028] Through the aforementioned multi-control station collaborative reward mechanism, this invention can effectively avoid the problem of simulation distortion of intermediate control stations caused by optimizing only the exit station. It enables the reinforcement learning agent to simultaneously consider the simulation accuracy of multiple key locations within the watershed during the calibration process, thereby improving the simulation consistency and reliability of the semi-distributed hydrological model across the entire watershed and achieving collaborative optimization of globally shared parameters, zone-independent parameters, and river routing parameters.

[0029] Based on the above embodiments, in some embodiments, the construction of a directed topology graph of upstream and downstream sites, and the determination of the semi-distributed model computation order from upstream to downstream based on the directed topology graph, specifically includes: Obtain the station number, control area, sub-basin attributes, rainfall time series, evaporation time series, and measured flow time series of multiple hydrological stations within the target watershed, as well as the connection relationships between upstream stations, downstream stations, and river segments; Based on the river segment connection relationship, a directed topology graph of upstream and downstream stations is constructed with hydrological stations as nodes and river channel connections between stations as directed edges. The stations in the basin are then topologically sorted according to the directed topology graph to obtain the hydrological simulation order from upstream to downstream.

[0030] In this embodiment, step S101 specifically includes the following process. First, station information of multiple hydrological stations within the target watershed is obtained. This station information includes station number, control area or local catchment area, and sub-watershed attributes. Simultaneously, hydrological and meteorological time series are obtained, specifically including rainfall time series, evaporation or potential evapotranspiration time series, and measured flow time series. Furthermore, river network topology is obtained, including upstream stations, downstream stations, and river segment connections. Based on the obtained river segment connections, a directed topology graph of upstream and downstream stations is constructed. Specifically, if station i is upstream of station j, and there is a river channel connection between them, a directed edge e_ij is constructed, thus forming a directed topology graph G=(V,E), where V represents the set of hydrological stations and E represents the set of river segment connections. Then, all hydrological stations within the watershed are topologically sorted according to this directed topology graph G to obtain the hydrological simulation order from upstream to downstream. The topological sorting results are used for subsequent calculations of the semi-distributed hydrological model to ensure that the simulated outflow from the upstream station can be transmitted to the downstream station via the river route, thus providing a calculation basis for running the semi-distributed hydrological model in the subsequent steps according to the directional topological order of the upstream and downstream stations.

[0031] Through the above methods, this embodiment can accurately construct the upstream and downstream directed topology graphs of the stations and determine the calculation order based on the acquired station information, river network topology, and hydrological and meteorological time series. This provides an ordered topological foundation for the subsequent runoff generation-routing coupling calculation of the semi-distributed hydrological model, thereby ensuring the consistency and accuracy of the upstream and downstream station simulation process.

[0032] Based on the above embodiments, in some embodiments, the step of dividing the target watershed into multiple runoff-producing zones and establishing the correspondence between each hydrological station and the runoff-producing zones specifically includes: Based on the watershed's river system structure, station distribution, sub-watershed attributes, control section locations, and hydrological response characteristics, the target watershed is divided into multiple runoff-generating zones, and a station-zone correspondence table between hydrological stations and runoff-generating zones is established. Specifically, for multiple sites within the same flow generation zone, some zone parameters are shared; for different flow generation zones, independent zone parameters are set.

[0033] In this embodiment, the process of dividing runoff-generating zones and establishing station-zone mapping relationships in step S102 is as follows: First, based on the watershed's river system structure, station distribution, sub-watershed attributes, control section locations, and hydrological response characteristics, the target watershed is divided into multiple runoff-generating zones, and the set of runoff-generating zones is represented as follows: ,in Indicates the number of runoff zones. Indicates the first Each runoff-producing zone contains one or more hydrological stations. Then, the correspondence between the hydrological stations and the runoff-producing zones is established. Let the set of all hydrological stations be denoted as . The set of runoff zones is Define the site-partition mapping table This is used to indicate the runoff zone to which each hydrological station belongs; when When, it indicates the first Hydrological stations Belongs to the Individual production flow zones The station-zone correspondence table can be stored through configuration files, parameter tables, dictionary structures, or database tables. For example, several upstream tributary stations can be configured as the first runoff-generating zone, and stations near the confluence of the main stream and its outlet can be configured as another runoff-generating zone. For multiple stations within the same runoff-generating zone, some zone parameters are shared to express the similar underlying surface conditions and runoff response characteristics within that zone; for different runoff-generating zones, independent zone parameters are set to express the spatial differences between different areas. During subsequent model calculations, stations call the independent zone parameters of their respective zones according to this correspondence table. Through the above method, the specific processing steps from inputting basic watershed data to runoff-generating zone division, establishing station-zone mapping relationships, and sharing and independently setting zone parameters are realized. Finally, the station-zone correspondence table and zone parameter configuration rules are output, providing a basis for the subsequent construction of a hierarchical parameter space and the operation of a semi-distributed hydrological model.

[0034] This specific implementation effectively reduces the difficulty of high-dimensional parameter calibration in semi-distributed models by establishing a site-partition mapping table and setting up a mechanism for sharing and separating partition parameters. Simultaneously, it preserves the ability to express spatial heterogeneity between different flow-generating partitions, improving the interpretability and physical consistency of the parameter calibration results. For example... Figure 3 As shown, the topological relationships between runoff zones, zone exit stations, main stations, river segment routing parameters, and exit control stations are illustrated to explain the coupling process between local runoff and upstream water flow after it travels along the river route.

[0035] Based on the above embodiments, in some embodiments, the hierarchical parameter space includes: globally shared parameters for characterizing the overall hydrological response characteristics of the target watershed; zone-independent parameters for characterizing the differences in water storage capacity and runoff response of different runoff-producing zones; and river routing parameters for characterizing the characteristics of flow propagation, lag, peak reduction, and confluence of different river sections. The globally shared parameters include one or more of evapotranspiration conversion parameters, runoff curve parameters, runoff distribution parameters, groundwater recession parameters, and interflow recession parameters. The zone-independent parameters include one or more of tension water capacity, upper layer water storage capacity, lower layer water storage capacity, free water capacity, and soil water storage parameters. The river routing parameters include one or more of water storage time constants, weighting coefficients, river section distance-related parameters, wave velocity-related parameters, and other parameters for characterizing the characteristics of flow propagation in river sections.

[0036] In this embodiment, the hierarchical parameter space constructed in step S102 specifically includes three types of parameters. The first type is globally shared parameters, used to characterize the overall hydrological response characteristics of the target watershed. This set of globally shared parameters is denoted as... ,in The first category represents the number of globally shared parameters. As an example, globally shared parameters may include one or more of the following: evapotranspiration conversion parameters, runoff generation curve parameters, runoff distribution parameters, groundwater recession parameters, and interflow recession parameters. These parameters are used uniformly across the entire watershed to reflect the overall evapotranspiration capacity, runoff generation mechanism, and baseflow recession pattern of the watershed. The second category consists of zone-independent parameters, used to characterize the differences in water storage capacity and runoff response between different runoff generation zones. This set of zone-independent parameters is denoted as […]. ,in For the number of production flow zones, the first The parameters for each partition are , This indicates the number of parameters for each zone. As an example, zone-independent parameters may include one or more of the following: tensional water capacity, upper reservoir capacity, lower reservoir capacity, free water capacity, and soil water storage parameters. These parameters are independent of each other across different runoff-producing zones to express the differences in underlying surface conditions and runoff characteristics across different regions. The third category consists of river routing parameters, used to characterize the propagation, lag, peak reduction, and confluence characteristics of different river segments. This set of river routing parameters is denoted as... ,in For the number of river sections, the first The parameters of the river section are , This represents the number of parameters for each river segment. As an example, river routing parameters may include one or more of the following: Muskingen storage time constant, Muskingen weighting coefficient, channel bottom width, Manning roughness, channel longitudinal slope, and wave velocity-related parameters. Ultimately, the complete parameter set of the semi-distributed hydrological model is represented as follows: After constructing this hierarchical parameter space, subsequent steps S103 and S104 will be based on this space for model running and reinforcement learning calibration.

[0037] By dividing the semi-distributed hydrological model parameters into globally shared parameters, zone-independent parameters, and river routing parameters according to their physical scope, this embodiment effectively reduces the search difficulty of high-dimensional parameter space, avoids parameter redundancy, and preserves the spatial difference representation ability of different zones and the water flow propagation characteristics of river sections, thereby improving the interpretability of parameter calibration results and the overall accuracy of watershed simulation. Figure 4 As shown, the hierarchical organization of globally shared parameters, partition-independent parameters, and river routing parameters is illustrated, forming a unified parameter vector.

[0038] Based on the above embodiments, in some embodiments, the semi-distributed hydrological model is run sequentially according to the calculation order determined by the upstream and downstream directed topology graph of the station. For the current station, its local runoff process is coupled with the inflow process after river routing calculation of all directly upstream stations through corresponding river routing parameters to obtain the simulated flow process of the current station, specifically including: For the current site, its runoff generation zone is determined according to the site-zone correspondence table, and the globally shared parameters and the zone-independent parameters corresponding to the zone are extracted. The local runoff generation process of the current site is calculated based on the rainfall time series, evaporation time series, watershed area and model parameters. Obtain the set of direct upstream stations of the current station. For any upstream station, obtain its simulated outflow process and calculate the river routing based on the river routing parameters corresponding to the connected river segment to obtain the upstream inflow process after routing. The local runoff generation process is superimposed and coupled with the upstream water inflow processes after all routing to obtain the simulated flow process of the current site; The river routing calculation adopts the Muskingan method, which determines the calculation coefficients through the water storage time constant, weight coefficients and calculation step size, so as to characterize the functional relationship between the current inflow, the previous inflow and the previous outflow and the current outflow.

[0039] In this embodiment, step S103 specifically includes the following sub-steps. First, for the hydrological station to be simulated, the runoff-generating zone to which the station belongs is determined according to the pre-established station-zone mapping relationship, and globally shared parameters and zone-independent parameters corresponding to the runoff-generating zone are extracted from the hierarchical parameter space. Then, using the local catchment area of ​​the station, the rainfall time series and evaporation time series of the zone to which it belongs as inputs, combined with the extracted globally shared parameters and zone-independent parameters, the runoff calculation module in the hydrological model is run to calculate the local runoff process of the interval or zone corresponding to the station. This local runoff process reflects the direct contribution of rainfall-runoff transformation within the control area of ​​the current station. Next, the set of directly upstream stations of the current station is obtained. For each upstream station in the set, its simulated outflow process is obtained, and the river routing parameters corresponding to the river segment connecting the upstream station and the current station are used to calculate the river routing of the outflow process of the upstream station to obtain the routed upstream inflow process. In one possible implementation, the river routing calculation adopts the Muskingan method, and its routing calculation can be expressed as follows: ,in This represents the inflow at the current moment, i.e., the outflow process from the upstream station. This indicates the outflow at the current moment, that is, the upstream water flow process after routing. , , For the water storage time constant Weighting coefficients and calculate step size The calculation coefficients were determined. Finally, the calculated local runoff process was superimposed and coupled with the inflow processes of all directly upstream stations after passing through river routes, thus obtaining the complete simulated flow process of the current station. This simulated flow process serves as the input data for subsequent calculations of the coordinated reward of multiple control stations.

[0040] Through the above specific implementation methods, this method achieves accurate calculation of the simulated flow process at stations in a semi-distributed hydrological model. Its beneficial effects are: it explicitly couples the process of upstream water flowing through the river with the local confluence process, ensuring the physical consistency of water volume and flood peak propagation between upstream and downstream stations, avoiding simulation distortion caused by ignoring topological relationships, and providing an accurate and coordinated simulation data foundation for subsequent parameter calibration based on multi-control station collaborative rewards.

[0041] Based on the above embodiments, in some embodiments, the calculation of the multi-control station collaborative reward, which can comprehensively evaluate the simulation effect of each control station, based on the simulated traffic process and the measured traffic process of the exit station, key control station, and regional representative station, specifically includes: For each control station, the Nash efficiency coefficient, water volume error, flood peak error, peak occurrence time error, and upstream-downstream consistency error are calculated based on its simulated flow sequence and measured flow sequence. Each error index is then converted into a positive reward through an exponential or linear function. Based on the preset weights of each evaluation indicator and the weight of each site, the converted positive rewards are weighted and summed to obtain the collaborative rewards for multiple control stations. Among the site weights, the weight of the exit station is greater than that of the key control station, and the weight of the key control station is greater than that of the regional representative station.

[0042] In this embodiment, the process of calculating the multi-control station collaborative reward and updating the policy using the reinforcement learning agent in step S104 is as follows. First, for each control station i in the control station set C, based on the simulated traffic sequence obtained by that station in step S103... and the pre-acquired measured flow sequence Calculate multiple evaluation metrics. Specifically, the Nash efficiency coefficient. The calculation formula is: ,in This represents the average measured flow rate. Water volume error. The calculation formula is: ,in To prevent extremely small positive numbers with a denominator of zero. Peak flow error. The calculation formula is: Peak occurrence time error The calculation formula is: ,in and The simulated and measured flood peak times are respectively used. In addition, an upstream-downstream consistency error is constructed. This is used to constrain the consistency between the simulated traffic process at the current site and the traffic process derived from the topology. Subsequently, the aforementioned error metrics are converted into positive rewards, where... , , , , , The attenuation coefficient is... This is the minimum reward cutoff value. The reward for a single control station i. It is obtained by weighted summation of the rewards from each evaluation indicator, i.e. ,in to Weights were assigned to each evaluation indicator. Finally, site weights were set according to the importance of the control sites. The weight of the exit station is greater than that of the key control station, and the weight of the key control station is greater than that of the representative station of the region. The collaborative reward of multiple control stations is obtained by weighted summation. The reinforcement learning agent constructs interactive experience based on the current state, actions, cooperative rewards, and the next state, and uses the proximal policy optimization algorithm to update the policy network and value network, thereby driving the agent to perform cooperative calibration on globally shared parameters, partition-independent parameters, and river routing parameters.

[0043] This implementation method constructs a multi-index reward function that includes Nash efficiency coefficient, water volume error, flood peak error, peak occurrence time error, and upstream-downstream consistency error, and sets differentiated weights based on the importance of each station. This allows the reinforcement learning agent to consider the simulation effects of outlet stations, key control stations, and representative stations in different areas, significantly improving the simulation accuracy of the semi-distributed hydrological model across the entire watershed and the consistency of the upstream-downstream propagation process. Figure 5 As shown, this demonstrates the closed-loop process by which a reinforcement learning agent outputs action vectors as parameters, and a semi-distributed hydrological model environment completes hierarchical parameter space mapping, topology routing calculation, local runoff calculation, and multi-control station evaluation, and then feeds back the status and collaborative rewards to the agent.

[0044] In this embodiment, in step S104, the reinforcement learning agent uses the Proximal Policy Optimization (PPO) algorithm to update the policy, thereby improving search stability in a continuous high-dimensional parameter space. Specifically, in each iteration, the reinforcement learning agent updates the policy based on the current state space. Output parameter action vector The action vector comprises globally shared parameter action sub-vectors, partition-independent parameter action sub-vectors, and river routing parameter action sub-vectors. The agent inputs the action vectors into a semi-distributed hydrological model environment. The environment runs the model sequentially according to the directed topological order of upstream and downstream stations, calculating the simulated flow process for each station. Based on the simulated and measured flow processes of the outlet station, key control station, and representative partition station, the collaborative reward for multiple control stations is calculated. The collaborative reward is obtained by summing the weighted rewards of each control station, where each station's reward incorporates the exponentially converted values ​​of the Nash efficiency coefficient, water volume error, flood peak error, peak occurrence time error, and upstream-downstream consistency error. The agent receives this collaborative reward and the next state after executing the action. Building interactive experiences During the policy update phase, the agent utilizes accumulated interaction experience batches to update the policy network and value network using the pruning objective function of the proximal policy optimization algorithm. This objective function is expressed as: ,in This represents the probability ratio between the old and new strategies. This is the estimated value of the dominance function. This is a pruning hyperparameter. Through this pruning mechanism, the algorithm limits the magnitude of each policy update, avoiding drastic fluctuations in policy performance due to excessively large single-step updates. This enables stable and efficient search within a continuous high-dimensional action space composed of globally shared parameters, partition-independent parameters, and river routing parameters. The agent continuously repeats the above parameter action output, model execution, reward calculation, and policy update process until the training termination condition is met, ultimately outputting a hierarchical parameter set after co-calibration.

[0045] This specific implementation effectively limits the update range of reinforcement learning policies in continuous high-dimensional parameter spaces by introducing a proximal policy optimization algorithm and its pruning objective function. This avoids training oscillations or divergences caused by excessive parameter adjustments, and significantly improves the convergence stability and search efficiency of the parameter calibration process of the semi-distributed hydrological model.

[0046] Based on the above embodiments, in some embodiments, this method further includes the following steps: Before co-calibrating the global shared parameters, partition-independent parameters, and river routing parameters, an initialization search phase is performed. Several initial parameter combinations are generated within the parameter feasible region through random sampling, Latin hypercube sampling, historical parameter hot start, heuristic algorithms, or empirical parameter range screening. The parameter region with better performance is selected as the initial parameter space for formal reinforcement learning training. Furthermore, the process of co-calibrating the global shared parameters, partition-independent parameters, and river routing parameters is divided into a phased co-calibration phase. The phased co-calibration phase includes: a first phase, calibrating the global shared parameters; a second phase, calibrating the partition-independent parameters; a third phase, calibrating the river routing parameters; and a fourth phase, performing joint fine-tuning of all parameters. In each phase, the reinforcement learning agent only outputs actions to the parameter subspace corresponding to the current phase, while the remaining parameters remain fixed or are only allowed to be fine-tuned within a small range. After each phase, the optimal parameters of that phase are used as the initial parameters for the next phase.

[0047] In this embodiment, before the formal execution of the reinforcement learning calibration process, an initialization search phase is first performed to provide a high-performance initial parameter space for subsequent reinforcement learning training. Specifically, this initialization search phase generates several initial parameter combinations within the parameter feasible region, which is composed of globally shared parameters, partition-independent parameters, and river routing parameters, through various sampling or search methods. The parameter feasible region is defined by the lower and upper limits of each parameter to be calibrated. One implementation method is to use random sampling, uniformly or randomly selecting multiple sets of parameter vectors within the parameter feasible region, each set containing the normalized or actual values ​​of all parameters to be calibrated. Another implementation method is to use Latin hypercube sampling, dividing the value range of each parameter into several intervals and randomly selecting a sample point within each interval, thereby ensuring that the generated initial parameter combinations have better spatial filling and representativeness in the parameter space. As another implementation method, a historical parameter hot-start method can be used to extract several sets of well-performing parameter combinations from historical calibration records, existing model configurations, or empirical parameter libraries of the watershed or similar watersheds, as initial parameter combinations. As yet another implementation method, heuristic algorithms, such as genetic algorithms or particle swarm optimization, can be used to perform a short-time search within the parameter feasible region to generate several candidate parameter combinations. As yet another implementation method, an empirical parameter range screening method can be used to select several representative parameter combinations from the parameter feasible region based on the physical meaning of the hydrological model parameters and watershed characteristics. For each initial parameter combination generated through any of the above methods, it is assigned to a semi-distributed hydrological model, and the model is run sequentially according to the directed topological order of upstream and downstream stations to calculate the simulated flow process at each station. Then, based on the simulated and measured flow processes at the outlet station, key control station, and representative sub-station, the multi-control station collaborative reward is calculated. The calculation method for this multi-control station collaborative reward is consistent with the reward function used in the subsequent reinforcement learning training phase, thus ensuring the consistency of the evaluation criteria. Then, from all initial parameter combinations, several parameter combinations with higher multi-control station collaborative reward values ​​are selected, or parameter combinations with reward values ​​exceeding a preset threshold are selected. The parameter regions corresponding to these parameter combinations are used as the initial parameter space for formal reinforcement learning training. This initial parameter space can be a narrowed parameter search range or a specific set of initial parameter vectors, used to initialize the policy network or value network of the reinforcement learning agent, or as the starting point for parameter exploration during reinforcement learning training. Through this initial search phase, the reinforcement learning agent can be placed in a parameter region with better performance in the early stages of training, avoiding a large amount of blind exploration in an ineffective or inefficient parameter space, thereby effectively improving the convergence speed of reinforcement learning training and the stability of the final calibration results.

[0048] This specific implementation introduces an initialization search phase before reinforcement learning calibration, using various sampling or search strategies to generate initial parameter combinations and select optimal parameter regions. This provides a high-quality initial parameter space for the reinforcement learning agent, significantly reducing the initial exploration difficulty of training in a high-dimensional continuous parameter space, and effectively improving the convergence efficiency of reinforcement learning training and the stability and reliability of the final calibration results.

[0049] In this embodiment, to further improve the training stability and convergence efficiency of reinforcement learning in a high-dimensional continuous parameter space, the method also includes a phased collaborative calibration step. Specifically, before the reinforcement learning agent interacts and trains with the semi-distributed hydrological model environment, an initialization search phase is first executed. This phase generates several initial parameter combinations within the parameter feasible region through random sampling, Latin hypercube sampling, or historical parameter warm-start, and runs the semi-distributed hydrological model on each initial parameter combination to calculate the collaborative reward for multiple control stations. The parameter region with the best performance is selected as the initial parameter space for the formal training of reinforcement learning. Subsequently, the reinforcement learning calibration process is divided into four progressively advancing phases. In the first phase, the reinforcement learning agent only outputs actions to the globally shared parameter subspace. At this time, the independent parameters of the partitions and the river routing parameters remain fixed or are only allowed to be fine-tuned within a small range. This phase mainly adjusts the common hydrological response parameters of the entire basin to enable the model to obtain basic water balance and overall runoff response capability. After the first phase, the optimal globally shared parameters obtained from this phase are used as the initial parameters for the next phase. In the second stage, the reinforcement learning agent, while keeping the globally shared parameters fixed, outputs actions only to the subspace of independent parameters for each region. River routing parameters remain fixed or only allowed for minor adjustments. This stage primarily adjusts the water storage capacity and runoff response parameters of different runoff-producing regions to enhance the spatial heterogeneity of the model's representation. After the second stage, the optimal subspace-independent parameters obtained from this training are used as the initial parameters for the next stage. In the third stage, while keeping both the globally shared and subspace-independent parameters fixed, the reinforcement learning agent outputs actions only to the subspace of river routing parameters. This stage primarily adjusts the routing parameters for each river segment to make the peak propagation time, peak reduction process, and downstream flow process more reasonable. After the third stage, the optimal river routing parameters obtained from this training are used as the initial parameters for the next stage. In the fourth stage, the reinforcement learning agent outputs joint actions to the complete parameter space comprised of the globally shared parameters, subspace-independent parameters, and river routing parameters, performing joint fine-tuning of all parameters to achieve optimal results for the outlet stations, key control stations, and representative subspace stations. By using the aforementioned phased collaborative calibration method, the reinforcement learning agent only needs to explore the parameter subspace corresponding to the current stage at each stage, reducing the dimensionality of a single search and thus improving the training stability and convergence efficiency in the continuous parameter space.

[0050] Based on the above embodiments, in some embodiments, when the training termination condition is met, the reinforcement learning calibration process ends, and the following parameters are output: globally shared parameters, partition-independent parameters of each runoff-producing partition, river routing parameters of each river segment, station-partition correspondence table, upstream and downstream topological relationships of stations, evaluation indicators for the control station calibration period and validation period, semi-distributed hydrological model parameter configuration file, model running results, and training process records.

[0051] In this embodiment, when the reinforcement learning calibration process meets a preset training termination condition, the entire calibration process ends, and a series of data and files characterizing the calibration results of the semi-distributed hydrological model parameters are output. Specifically, the training termination condition can be one or more of various scenarios. For example, when the number of training rounds for the reinforcement learning agent reaches a preset maximum number of training rounds, or when the cumulative number of runs of the semi-distributed hydrological model reaches a preset maximum number of model runs, the termination condition is determined to be met. Furthermore, when the collaborative reward of multiple control stations does not show a significant increase in several consecutive iterations, i.e., the change in reward value is less than a preset reward increase threshold, it indicates that the model performance has stabilized, and this can also be used as a termination condition. Alternatively, the termination of the calibration process can be triggered when the simulation accuracy of the exit station or key control station, such as the Nash efficiency coefficient, reaches a preset accuracy threshold, or when the change in all parameters to be calibrated is less than a preset parameter change threshold in several consecutive iterations.

[0052] After any of the above training termination conditions are met, the system performs an output operation. First, it outputs the complete parameter set of the semi-distributed hydrological model determined after collaborative calibration, specifically including globally shared parameters, zone-specific parameters for each runoff-producing zone, and river routing parameters for each river segment. Simultaneously, it outputs the station-zone correspondence table established in step S102 and the upstream-downstream topological relationships of the stations constructed in step S101. This information clarifies the spatial assignment of parameters and the computational order of the model. Furthermore, it outputs the evaluation indicators for each control station during the calibration and validation periods. These evaluation indicators are based on various sub-indicators involved in the multi-control station collaborative reward, such as Nash efficiency coefficient, water volume error, flood peak error, peak occurrence time error, and upstream-downstream consistency error. Finally, it outputs the parameter configuration file for the semi-distributed hydrological model. This file contains all calibrated parameter values ​​and their corresponding station and zone information, as well as the model's running results and a record of the entire training process, including the status, actions, rewards, and parameter change trajectories for each iteration. These output results collectively constitute the complete calibration outcome, providing a directly usable parameter set and performance evaluation basis for subsequent hydrological forecasting and model applications.

[0053] Through the above specific implementation methods, the present invention can automatically determine the convergence state of the calibration process and output a complete and detailed calibration result when the conditions are met. It not only provides the optimal parameter combination, but also includes spatial mapping relationships and performance evaluation, thereby significantly improving the automation level and interpretability of the semi-distributed hydrological model parameter calibration process.

[0054] Based on the above embodiments, this embodiment will describe in detail the complete implementation of the partitioned collaborative reinforcement learning parameter calibration method for semi-distributed hydrological models, including specific technical aspects such as data acquisition and preprocessing, runoff generation partitioning, hierarchical parameter space construction, reinforcement learning state and action space encoding, model topology sequential operation, multi-control station collaborative reward calculation, agent policy update, and initialization search and phased collaborative calibration.

[0055] First, basic data for the target watershed is acquired. Specifically, station information for multiple hydrological stations within the target watershed is obtained, including station numbers, controlled areas, or local catchment areas; river network connectivity is acquired, including upstream and downstream station connections and river segment connections; sub-watershed attributes are acquired; and hydrological and meteorological time series are acquired, including rainfall time series, evaporation or potential evapotranspiration time series, and measured flow time series. Based on the river network connectivity, a directed topology graph of upstream and downstream stations is constructed, represented as follows: ,in This represents a collection of hydrological stations. This represents a set of river segments connected. If the station... Located at the station If the upstream of the two rivers is connected by a river channel, then a directed edge is constructed. Based on the directed topological graph The watershed stations are topologically sorted to obtain the hydrological simulation order from upstream to downstream. This topological order is used for subsequent semi-distributed hydrological model calculations, enabling the simulated outflow from upstream stations to be transmitted to downstream stations via river routes.

[0056] Then, based on the watershed's river system structure, station distribution, sub-watershed attributes, control section locations, and hydrological response characteristics, the target watershed is divided into multiple runoff-generating zones, denoted as... ,in Indicates the number of runoff zones. Indicates the first There are several runoff-producing zones, each containing one or more hydrological stations. Establish the correspondence between hydrological stations and runoff-producing zones, assuming the set of all hydrological stations is denoted as . The set of runoff zones is Define the site-partition mapping table This is used to indicate the runoff zone to which each hydrological station belongs; when When, it indicates the first Hydrological stations Belongs to the Individual production flow zones The site-zone correspondence table can be stored through configuration files, parameter tables, dictionary structures, or database tables. For example, several upstream tributary sites can be configured as the first runoff-generating zone, and sites near the main stream confluence area and outlet can be configured as another runoff-generating zone. For multiple sites within the same runoff-generating zone, some zone parameters are shared to express the similar underlying surface conditions and runoff response characteristics within that zone; for different runoff-generating zones, independent zone parameters are set to express the spatial differences between different areas.

[0057] Next, a hierarchical parameter space for the semi-distributed hydrological model is constructed, comprising three categories: globally shared parameters, zone-independent parameters, and river routing parameters. The globally shared parameters are denoted as... ,in This indicates the number of globally shared parameters. Globally shared parameters are used uniformly across the entire watershed to characterize the overall hydrological response of the target watershed. For example, in the Xin'anjiang model, these include one or more of the following: evapotranspiration conversion parameters, runoff generation curve parameters, runoff distribution parameters, groundwater recession parameters, interflow recession parameters, and watershed-wide unified response parameters. Zonal independent parameters are denoted as... , of which The partition-independent parameters for each flow generation partition are: , This represents the number of partition-independent parameters for each runoff-producing zone. Partition-independent parameters characterize the differences in water storage capacity and runoff response among different runoff-producing zones. For example, in the Xin'anjiang model, these include one or more of the following: tensional water capacity, upper layer water storage capacity, lower layer water storage capacity, free water capacity, soil water storage parameters, and partition-specific runoff response parameters. River routing parameters are denoted as... , of which The river channel of this river section is determined by the following parameters: , Indicates the number of river sections. This represents the number of routing parameters for each river segment. These parameters characterize the flow propagation, lag, peak reduction, and confluence characteristics of different river segments. They include one or more of the following: Muskingan storage time constant, Muskingan weighting coefficient, channel bottom width, Manning roughness, riverbed longitudinal slope, and wave velocity-related parameters. Ultimately, the complete parameter set of the semi-distributed hydrological model is represented as follows: .

[0058] For each parameter to be calibrated Set parameter lower limit and parameter upper limit Normalize the actual parameter values ​​to ,in The global parameter subvector, partition parameter subvector, and routing parameter subvector are concatenated sequentially to obtain a normalized continuous parameter vector $X = [X_g, X_z, X_r]$, where Represents a global parameter subvector. Represents a subvector of partitioning parameters. Represents a subvector of routing parameters. Each partition's independent parameters are encoded according to the production flow partitioning order. The routing parameters for each river segment are encoded according to the topological order of the river segments. The state space of the reinforcement learning agent is defined as follows: ,in Indicates the first The state space at the next iteration. This represents the current normalized parameter vector. This represents the change between the current parameter vector and the parameter vector of the previous round or the historical best parameter vector. This represents the set of simulation error features for each control station. The set of simulation error features includes one or more of the following: exit station evaluation indicators, key control station evaluation indicators, representative zone station evaluation indicators, upstream and downstream propagation consistency indicators, Nash efficiency coefficient, water volume error, flood peak error, and peak occurrence time error. The action space of a reinforcement learning agent is defined as... ,in Represents a globally shared parameter action vector. Represents a partition-independent parameter action subvector. The action space represents the action vector of the river path parameters. The action space can represent the normalized values ​​of the parameters or the adjustment amount of the parameters relative to the previous round.

[0059] The reinforcement learning agent adjusts its behavior according to the current state. Output parameter action vector When the action vector represents the normalized values ​​of the parameters, the actual parameters are obtained through mapping using the following formula: Where $a_i^t$ represents the first... The agent outputs the first iteration. A normalized action value. When the action vector represents the parameter adjustment amount, the parameter is updated using the following formula: ,in Indicates the first The parameter adjustment amount in the next iteration, Clip represents the boundary clipping function. Further physical and numerical stability constraints are imposed on the river routing parameters; for example, for Muskingan routing parameters, the water storage time constant is required. The weight coefficients are greater than the lower bound of stability corresponding to the model's computational step size. Within the preset feasible range; for the Muskingen-Congee routing parameters, the riverbed width, Manning roughness, and longitudinal slope are required to be within a physically reasonable range. After parameter boundary mapping and constraint processing, the global shared parameters for the current iteration, the partition-independent parameters for each runoff-producing zone, and the river routing parameters for each river segment are obtained.

[0060] Following the established station topology, the semi-distributed hydrological model is run sequentially from upstream to downstream. For the current station... First, based on the site-zone mapping relationship, determine the corresponding flow generation zone. And extract globally shared parameters Partition-independent parameters corresponding to this partition Based on rainfall time series, evaporation time series, watershed area, and model parameters, the hydrological model is run to calculate the local runoff process for the corresponding interval or zone at the current station. Then, obtain the set of direct upstream sites of the current site. For any upstream station Obtain its simulated outflow process And according to the connecting river sections The corresponding river road is determined by parameters. The river route calculation is performed to obtain the upstream water flow process after the route is taken. Finally, the local runoff generation process is coupled with all upstream water routing processes to obtain the simulated flow process at the current site. When the river route is routed using the Muskingan method, the route calculation can be expressed as follows: ,in Indicates the current moment of entry into the stream. Indicates the outflow at the current moment. , , For the water storage time constant Weighting coefficients and calculate step size Determined calculation coefficients. When the river routing method adopts the Muskingen-Kangi method, the water storage time constant and weight coefficients can be calculated from parameters such as river length, wave velocity, river width, riverbed longitudinal slope, and Manning roughness, and can also be used as routing parameters in reinforcement learning calibration.

[0061] Let the set of control stations be ,in Indicates the collection point at the exit station. This represents the set of key control stations. This represents the set of representative stations for a given zone. (For control stations...) According to the simulated flow sequence and measured flow sequence Calculate the Nash efficiency coefficient Water quantity error Flood peak error Peak time error ,in This represents the average measured flow rate. To prevent extremely small positive numbers with a denominator of zero, Indicates the time of the simulated flood peak. This indicates the measured time of the flood peak. To constrain the consistency of upstream and downstream propagation, an upstream-downstream consistency error is further constructed. This is used to constrain the consistency between the simulated traffic process at the current site and the traffic process derived from the topology. The above metrics are then converted into positive rewards: , , , , ,in The attenuation coefficient is... This is the minimum reward cutoff value. The reward for a single control site is... ,in Assign weights to each evaluation indicator. Set site weights based on the importance of the control sites. The weight of the exit station is greater than that of the key control station, and the weight of the key control station is greater than that of the regional representative station. Ultimately, the collaborative reward for multiple control stations is... .

[0062] Reinforcement learning agents construct interactive experiences based on the current state, actions, rewards, and the next state. ,in This indicates whether the current training round has terminated. The agent updates the policy network and value network using interactive experience, employing a proximal policy optimization algorithm. By limiting the update magnitude between the old and new policies, the stability of the continuous parameter space search process is improved. The policy update objective can be represented as... ,in This represents the ratio of the probability of the new strategy to the probability of the old strategy. This represents the estimated value of the advantage function. This indicates the clipping range hyperparameter.

[0063] To improve the efficiency and stability of reinforcement learning training, this embodiment can adopt a combination of initial search and phased collaborative calibration. In the initial search phase, several initial parameter combinations are generated within the feasible parameter domain through methods such as random sampling, Latin hypercube sampling, historical parameter warm-start, heuristic algorithms, or empirical parameter range screening. A semi-distributed hydrological model is run on each initial parameter combination to calculate the collaborative reward of multiple control stations, and the parameter region with better performance is selected as the initial parameter space for formal reinforcement learning training. In the phased collaborative calibration phase, the reinforcement learning calibration process is divided into the following stages: The first stage calibrates the globally shared parameters, mainly adjusting the common hydrological response parameters across the entire basin to achieve basic water balance and overall runoff response capability. The second stage calibrates the regionally independent parameters, adjusting the water storage capacity and runoff response parameters of different runoff-producing zones based on the already optimized globally shared parameters, allowing the model to express spatial heterogeneity. The third stage calibrates the river routing parameters, mainly adjusting the routing parameters of each river segment to make the peak propagation time, peak reduction process, and downstream flow process more reasonable. The fourth stage performs joint fine-tuning of all parameters, jointly optimizing the globally shared parameters, regionally independent parameters, and river routing parameters to achieve optimal results for the outlet stations, key control stations, and representative stations of each zone. In each stage, the reinforcement learning agent only outputs actions to the parameter subspace corresponding to the current stage; other parameters remain fixed or are only allowed to be fine-tuned within a small range. After each stage, the optimal parameters of that stage are used as the initial parameters for the next stage.

[0064] The reinforcement learning calibration process ends when the training termination conditions are met. These conditions include reaching the maximum number of training rounds, reaching the maximum number of model runs, no significant improvement in multi-control station collaborative rewards for several consecutive rounds, the simulation accuracy of the outlet station reaching a preset threshold, the simulation accuracy of the key control station reaching a preset threshold, or the parameter change being less than a preset threshold—one or more of these conditions. The final output includes: globally shared parameters, zone-specific parameters for each runoff-producing zone, river routing parameters for each river segment, a station-zone correspondence table, upstream and downstream topological relationships of stations, evaluation indicators for the control station calibration and validation periods, a semi-distributed hydrological model parameter configuration file, model running results, and training process records.

[0065] This embodiment constructs a hierarchical parameter space comprising globally shared parameters, partition-independent parameters, and river routing parameters. It performs runoff-routing coupling calculations according to the upstream and downstream topological order of the stations. Combined with a multi-control station collaborative reward function and a phased collaborative calibration strategy, it achieves efficient, stable, and collaborative calibration of semi-distributed hydrological model parameters. This effectively solves the technical problems in existing methods, such as high parameter dimensionality, insufficient expression of spatial heterogeneity, lack of coordination between upstream and downstream simulations, and the separation of runoff and routing parameters.

[0066] Please refer to Figure 6 , Figure 6 This invention provides a structural block diagram of a partitioned collaborative reinforcement learning parameter calibration device for semi-distributed hydrological models; the specific device may include: The data acquisition and topology construction module 100 is used to acquire station information, river network topology and hydrological and meteorological time series of the target watershed, construct a directed topology graph of upstream and downstream stations, and determine the semi-distributed model calculation order from upstream to downstream based on the directed topology graph; The partitioning and parameter space construction module 200 is used to divide the target watershed into multiple runoff-producing zones according to the watershed system structure and station distribution, establish the correspondence between each hydrological station and the runoff-producing zones, and construct a hierarchical parameter space containing globally shared parameters, zone-independent parameters, and river routing parameters. The model running module 300 is used to run the semi-distributed hydrological model sequentially according to the calculation order determined by the upstream and downstream directed topology graph of the station. For the current station, its local runoff process is coupled with the inflow process after the river routing calculation of all directly upstream stations through the corresponding river routing parameters to obtain the simulated flow process of the current station. The collaborative reward calculation module 400 is used to calculate the multi-control station collaborative reward based on the simulated and measured flow processes of the exit station, key control station, and regional representative station, which can comprehensively evaluate the simulation effect of each control station. The reinforcement learning module 500 is used to update the policy of the reinforcement learning agent based on the cooperative reward and to perform cooperative calibration on the global shared parameters, partition-independent parameters and river routing parameters.

[0067] The partitioned collaborative reinforcement learning parameter calibration device for semi-distributed hydrological models in this embodiment is used to implement the aforementioned partitioned collaborative reinforcement learning parameter calibration method for semi-distributed hydrological models. Therefore, the specific implementation of the partitioned collaborative reinforcement learning parameter calibration device for semi-distributed hydrological models can be found in the previous embodiment section of the partitioned collaborative reinforcement learning parameter calibration method for semi-distributed hydrological models. For example, the data acquisition and topology construction module 100, the partition and parameter space construction module 200, the model running module 300, the collaborative reward calculation module 400, and the reinforcement learning module 500 are respectively used to implement steps S101-S104 in the aforementioned partitioned collaborative reinforcement learning parameter calibration method for semi-distributed hydrological models. Therefore, its specific implementation can be referred to the description of the corresponding embodiments, which will not be repeated here.

[0068] A specific embodiment of the present invention also provides a partitioned collaborative reinforcement learning parameter calibration device for a semi-distributed hydrological model, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of the partitioned collaborative reinforcement learning parameter calibration method for a semi-distributed hydrological model described above.

[0069] A specific embodiment of the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described method for calibrating partitioned collaborative reinforcement learning parameters for a semi-distributed hydrological model.

[0070] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0071] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0072] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0073] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0074] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A method for calibrating parameters of a partitioned collaborative reinforcement learning model for a semi-distributed hydrological model, characterized in that, include: Obtain station information, river network topology, and hydrological and meteorological time series of the target watershed; construct a directed topology graph of the upstream and downstream stations; and determine the calculation order of the semi-distributed model from upstream to downstream based on the directed topology graph. Based on the watershed's river system structure and station distribution, the target watershed is divided into multiple runoff-producing zones. The correspondence between each hydrological station and the runoff-producing zones is established, and a hierarchical parameter space containing globally shared parameters, zone-independent parameters, and river routing parameters is constructed. According to the calculation order determined by the upstream and downstream directed topology graph of the station, the semi-distributed hydrological model is run sequentially. For the current station, its local runoff process is coupled with the inflow process after the river routing calculation of all directly upstream stations through the corresponding river routing parameters to obtain the simulated flow process of the current station. Based on the simulated and measured flow processes of exit stations, key control stations, and representative stations in different zones, a multi-control station collaborative reward is calculated to comprehensively evaluate the simulation effect of each control station. A reinforcement learning agent is then used to update its policy based on the collaborative reward, and the global shared parameters, zone-independent parameters, and river routing parameters are collaboratively calibrated.

2. The method according to claim 1, characterized in that, The construction of a directed topology graph of the upstream and downstream sites, and the determination of the semi-distributed model computation order from upstream to downstream based on the directed topology graph, specifically includes: Obtain the station number, control area, sub-basin attributes, rainfall time series, evaporation time series, and measured flow time series of multiple hydrological stations within the target watershed, as well as the connection relationships between upstream stations, downstream stations, and river segments; Based on the river segment connection relationship, a directed topology graph of upstream and downstream stations is constructed with hydrological stations as nodes and river channel connections between stations as directed edges. The stations in the basin are then topologically sorted according to the directed topology graph to obtain the hydrological simulation order from upstream to downstream.

3. The method according to claim 1, characterized in that, The process of dividing the target watershed into multiple runoff-producing zones and establishing a correspondence between each hydrological station and the runoff-producing zones specifically includes: Based on the watershed's river system structure, station distribution, sub-watershed attributes, control section locations, and hydrological response characteristics, the target watershed is divided into multiple runoff-generating zones, and a station-zone correspondence table between hydrological stations and runoff-generating zones is established. Specifically, for multiple sites within the same flow generation zone, some zone parameters are shared; for different flow generation zones, independent zone parameters are set.

4. The method according to claim 1, characterized in that, The hierarchical parameter space includes: globally shared parameters characterizing the overall hydrological response characteristics of the target watershed; zone-independent parameters characterizing the differences in water storage capacity and runoff response between different runoff-producing zones; and river routing parameters characterizing the characteristics of flow propagation, lag, peak reduction, and confluence in different river sections. The globally shared parameters include one or more of the following: evapotranspiration conversion parameters, runoff curve parameters, runoff distribution parameters, groundwater recession parameters, and interflow recession parameters. The zone-independent parameters include one or more of the following: tension water capacity, upper layer water storage capacity, lower layer water storage capacity, free water capacity, and soil water storage parameters. The river routing parameters include one or more of the following: water storage time constant, weighting coefficient, river section distance-related parameters, wave velocity-related parameters, and other parameters characterizing the flow propagation characteristics of river sections.

5. The method according to claim 1, characterized in that, The semi-distributed hydrological model is run sequentially according to the calculation order determined by the upstream and downstream directed topology graph of the station. For the current station, its local runoff process is coupled with the inflow process calculated by the corresponding river routing parameters of all directly upstream stations to obtain the simulated flow process of the current station, specifically including: For the current site, its runoff generation zone is determined according to the site-zone correspondence table, and the globally shared parameters and the zone-independent parameters corresponding to the zone are extracted. The local runoff generation process of the current site is calculated based on the rainfall time series, evaporation time series, watershed area and model parameters. Obtain the set of direct upstream stations of the current station. For any upstream station, obtain its simulated outflow process and calculate the river routing based on the river routing parameters corresponding to the connected river segment to obtain the upstream inflow process after routing. The local runoff generation process is superimposed and coupled with the upstream water inflow processes after all routing to obtain the simulated flow process of the current site; The river routing calculation adopts the Muskingan method, which determines the calculation coefficients through the water storage time constant, weight coefficients and calculation step size, so as to characterize the functional relationship between the current inflow, the previous inflow and the previous outflow and the current outflow.

6. The method according to claim 1, characterized in that, The simulated and measured flow processes based on exit stations, key control stations, and representative regional stations are used to calculate the multi-control station collaborative reward, which comprehensively evaluates the simulation effect of each control station. Specifically, this includes: For each control station, the Nash efficiency coefficient, water volume error, flood peak error, peak occurrence time error, and upstream-downstream consistency error are calculated based on its simulated flow sequence and measured flow sequence. Each error index is then converted into a positive reward through an exponential or linear function. Based on the preset weights of each evaluation indicator and the weight of each site, the converted positive rewards are weighted and summed to obtain the collaborative rewards for multiple control stations. Among the site weights, the weight of the exit station is greater than that of the key control station, and the weight of the key control station is greater than that of the regional representative station.

7. The method according to claim 1, characterized in that, It also includes the following steps: Before co-calibrating the global shared parameters, partition-independent parameters, and river routing parameters, an initialization search phase is performed. Several initial parameter combinations are generated within the parameter feasible region through random sampling, Latin hypercube sampling, historical parameter hot start, heuristic algorithms, or empirical parameter range screening. The parameter region with better performance is selected as the initial parameter space for formal reinforcement learning training. Furthermore, the process of co-calibrating the global shared parameters, partition-independent parameters, and river routing parameters is divided into a phased co-calibration phase. The phased co-calibration phase includes: a first phase, calibrating the global shared parameters; a second phase, calibrating the partition-independent parameters; a third phase, calibrating the river routing parameters; and a fourth phase, performing joint fine-tuning of all parameters. In each phase, the reinforcement learning agent only outputs actions to the parameter subspace corresponding to the current phase, while the remaining parameters remain fixed or are only allowed to be fine-tuned within a small range. After each phase, the optimal parameters of that phase are used as the initial parameters for the next phase.

8. A partitioned collaborative reinforcement learning parameter calibration device for semi-distributed hydrological models, characterized in that, include: The data acquisition and topology construction module is used to acquire station information, river network topology, and hydrological and meteorological time series of the target watershed, construct a directed topology graph of the upstream and downstream of the station, and determine the semi-distributed model calculation order from upstream to downstream based on the directed topology graph; The partitioning and parameter space construction module is used to divide the target watershed into multiple runoff-producing zones according to the watershed system structure and station distribution, establish the correspondence between each hydrological station and the runoff-producing zone, and construct a hierarchical parameter space containing globally shared parameters, zone-independent parameters, and river routing parameters. The model running module is used to run the semi-distributed hydrological model sequentially according to the calculation order determined by the upstream and downstream directed topology graph of the station. For the current station, its local runoff process is coupled with the inflow process after the river routing calculation of all directly upstream stations through the corresponding river routing parameters to obtain the simulated flow process of the current station. The collaborative reward calculation module is used to calculate the collaborative reward of multiple control stations based on the simulated and measured flow processes of the exit station, key control station, and regional representative station, which can comprehensively evaluate the simulation effect of each control station. The reinforcement learning module is used to update the policy of the reinforcement learning agent based on the cooperative reward and to perform cooperative calibration on the global shared parameters, partition-independent parameters and river routing parameters.

9. A partitioned collaborative reinforcement learning parameter calibration device for semi-distributed hydrological models, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of a partitioned collaborative reinforcement learning parameter calibration method for a semi-distributed hydrological model as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the partitioned collaborative reinforcement learning parameter calibration method for a semi-distributed hydrological model as described in any one of claims 1 to 7.