A training method for reservoir operation model and reservoir operation system
By combining hydrological models and reinforcement learning technology, the parameters of the reservoir operation model are adjusted, which solves the problem of inaccurate operation decisions of the reservoir operation model in a dynamic environment and ensures the safety and rationality of reservoir operation.
Patent Information
- Application Number
- CN202510860392.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Existing reservoir scheduling models trained based on reinforcement learning technology have difficulty in accurately determining scheduling decisions under rapidly changing meteorological conditions, and fail to effectively consider the physical laws of reservoir water storage.
By combining hydrological models and reinforcement learning technology, the parameters of the reservoir operation model are adjusted, including the value of the reservoir operation model's own output, strategy loss and reward value constraints. Considering that the change of reservoir water storage should conform to physical laws, a multi-layer neural network is used for training.
The accuracy of reservoir operation decision-making in a dynamic environment is improved, the safety and rationality of reservoir operation are ensured, and system instability caused by frequent or large-scale actions is avoided.
Smart Images

Figure CN120373808B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent water conservancy technology, and in particular to a training method for a reservoir scheduling model and a reservoir scheduling system. Background Art
[0002] In the water conservancy industry, reservoir scheduling decisions require consideration of multiple factors, such as reservoir water level, inflow, and weather. Determining scheduling decisions based on manual experience or optimization algorithms based on mathematical models is difficult to adapt to rapidly changing weather conditions and complex scheduling objectives. However, reinforcement learning, as a dynamic decision-making optimization method, can automatically find optimal strategies in dynamically changing environments. Therefore, using reservoir scheduling models derived from reinforcement learning technology to determine scheduling decisions has become a mainstream approach.
[0003] Therefore, how to improve the accuracy of scheduling decisions output by the reservoir scheduling model trained based on reinforcement learning technology is an important issue.
[0004] Based on this, this application specification provides a training method for a reservoir scheduling model and a reservoir scheduling system. Summary of the Invention
[0005] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a training method for a reservoir scheduling model and a reservoir scheduling system. When training the reservoir scheduling model, the method obtains an updated reservoir state and a reward value corresponding to the predicted gate opening through a hydrological model based on the gate opening predicted by the reservoir scheduling model. The parameters of the reservoir scheduling model are adjusted based on the actual water storage of the reservoir when the gate opening is the predicted gate opening, and based on the actual water storage, combined with the updated state and reward value output by the hydrological model. The adjustment direction not only includes the loss corresponding to the value output by the reservoir scheduling model itself, the loss corresponding to the strategy, and the reward value constraints, but also takes into account physical constraints, namely, that the change in reservoir water storage should conform to physical laws, thereby improving the prediction accuracy of the trained reservoir scheduling model.
[0006] In order to achieve the above objectives, the main technical solutions adopted by the present invention include:
[0007] In a first aspect, an embodiment of the present invention provides a training method for a reservoir operation model, wherein the reservoir operation model includes a first policy network, a second policy network, a first value network, and a second value network; the method includes:
[0008] Acquire a first state of a reservoir; wherein the state of the reservoir includes at least a water level, an inflow, and a rainfall, and the reservoir is a simulated reservoir in a simulation environment provided based on a preset hydrological model;
[0009] Inputting the first state into the first strategy network to obtain a first gate opening of the reservoir; and inputting the first state and the first gate opening into the first value network to obtain a first value;
[0010] Determining a second state and a reward value of the reservoir using the hydrological model according to the first gate opening and the first state;
[0011] Determine, based on the second state, the actual water storage capacity of the reservoir when the gate opening of the reservoir is the first gate opening; input the second state into the second strategy network to obtain a second gate opening of the reservoir; input the second gate opening and the second state into the second value network to obtain a second value;
[0012] The parameters of the reservoir scheduling model are adjusted according to the first gate opening, the actual water storage capacity, the reward value, the first value, and the second value.
[0013] Optionally, adjusting the parameters of the reservoir operation model specifically includes:
[0014] determining the outflow of the reservoir based on the first gate opening; and determining the time duration for the state of the reservoir to be updated from the first state to the second state under a simulation environment provided by the hydrological model;
[0015] Determine the actual water storage capacity change rate based on the actual water storage capacity and the duration; and determine the predicted water storage capacity change rate based on the outflow flow, the first state and the duration;
[0016] determining a first loss according to the actual water storage capacity change rate and the predicted water storage capacity change rate;
[0017] Adjusting parameters of the reservoir operation model according to the first loss.
[0018] Optionally, adjusting the parameters of the reservoir operation model specifically includes:
[0019] determining a second loss based on the first value and the second value; and adjusting parameters of the first value network based on the second loss, the first loss, and the reward value;
[0020] A third loss is determined based on a third value corresponding to the first gate opening output by the first value network after parameter adjustment; and parameters of the first strategy network are adjusted based on the third loss and the first loss.
[0021] Optionally, adjusting the parameters of the reservoir operation model specifically includes:
[0022] The parameters of the second strategy network and the second value network are adjusted by soft updating.
[0023] Optionally, during the training process, for each round of training, the order of adjusting the parameters of the first value network is before adjusting the parameters of the first policy network, and the order of adjusting the parameters of the second value network and the second policy network is after adjusting the parameters of the first policy network.
[0024] Optionally, the reward value is determined by the hydrological model based on a preset reward function, and the reward function at least includes: a water level constraint function and a flow constraint function.
[0025] Optionally, obtaining the first state of the reservoir specifically includes:
[0026] In response to a user operation, determining a state space of the reservoir selected by the user, the state space including at least water level, inflow, and rainfall;
[0027] Based on the state space of the reservoir selected by the user, the state of the reservoir is initialized, and the initialized state of the reservoir is used as the first state of the reservoir.
[0028] Optionally, the method further includes:
[0029] Obtaining the status of a target reservoir, wherein the status of the target reservoir includes water level, inflow, and rainfall;
[0030] Inputting the state of the target reservoir into a trained reservoir scheduling model to obtain a target gate opening output by the trained reservoir scheduling model;
[0031] Based on the target gate opening, the gate opening of the target reservoir is controlled.
[0032] In a second aspect, an embodiment of the present invention provides a reservoir scheduling system, the system being used to execute any of the methods described in the first aspect, the system comprising: a model scheduling interface, an environment simulation unit, and an intelligent agent unit;
[0033] The model scheduling interface is used to determine the state space, action space and reward function selected by the user in response to the user's operation; the state space includes at least: water level, inflow and rainfall, and the action space includes at least reservoir gate opening;
[0034] The environment simulation unit is configured to load a preset hydrological model based on the state space, action space, and reward function selected by the user, and initialize a simulation environment to simulate dynamic changes in the state of the reservoir;
[0035] The intelligent agent unit is used to initialize the parameters of the reservoir operation model based on the state space and action space selected by the user, and determine the target action and target value corresponding to the current state through the reservoir operation model; and send the current state and the target action to the environment simulation unit;
[0036] The environmental simulation unit is configured to simulate, in the simulation environment, a process of updating the current state of the reservoir based on the target action using the loaded hydrological model, thereby obtaining an updated state corresponding to the current state; determine a reward value corresponding to the target action based on the reward function selected by the user; and send the reward value and the updated state to the agent unit;
[0037] The intelligent agent unit is used to determine at least the actual water storage capacity of the reservoir based on the updated state; and adjust the parameters of the reservoir scheduling model based on the reward value, the actual water storage capacity, the target action and the target value.
[0038] Optionally, the system further comprises: a reservoir operation model scheduling engine and a decision execution unit;
[0039] The reservoir operation model scheduling engine is used to generate scheduling decisions based on the trained reservoir operation model;
[0040] The decision execution unit is used to receive the scheduling decision; and control the gate opening of the reservoir according to the scheduling decision.
[0041] The beneficial effect of the present invention is that when adjusting the parameters of the reservoir scheduling model, the adjustment direction not only includes the loss corresponding to the value output by the reservoir scheduling model itself, the loss corresponding to the strategy, and the constraint of the reward value, but also takes into account the physical constraint, that is, the change in the reservoir water storage capacity should conform to the physical laws, thereby improving the accuracy of the prediction of the trained reservoir scheduling model. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 A schematic diagram of a reservoir dispatching system provided in this specification;
[0043] Figure 2 A schematic diagram of a reservoir dispatching system provided in this specification;
[0044] Figure 3 A flow chart of a training method for a reservoir operation model provided in this specification;
[0045] Figure 4 This is a structural diagram of a reservoir scheduling model provided in this manual. DETAILED DESCRIPTION
[0046] In order to better explain the present invention and facilitate understanding, the present invention is described in detail below through specific implementation methods in conjunction with the accompanying drawings.
[0047] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided to enable a clearer and more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0048] The core components of reinforcement learning (RL) technology are the agent and the environment. The environment refers to the space or situation in which the agent is located. The agent can perceive its environment and influence it by taking actions. The environment can receive actions from the agent and return a new state and corresponding rewards. In this way, the agent can learn how to make decisions or what actions to take to maximize cumulative rewards. In this specification, the reservoir operation model can be trained based on reinforcement learning technology. The reservoir operation model is the agent. The purpose is to use reinforcement learning technology to enable the reservoir operation model to learn how to make optimal decisions to manage reservoir resources based on different states.
[0049] like Figure 1 As shown, Figure 1 This is a schematic diagram of a reservoir operation system provided in this specification. In one or more embodiments of this specification, the reservoir operation system can provide a simulation environment for training a reservoir operation model, and the system includes at least: a model operation interface, an environment simulation unit, and an intelligent agent unit.
[0050] The model scheduling interface can determine the state space, action space, and reward function selected by the user in response to the user's operation. In one or more embodiments of this specification, the state space includes but is not limited to information such as the reservoir water level, inflow, and rainfall, and the action space includes at least the reservoir gate opening.
[0051] It should be noted that this reward function can be determined based on the actual mission objectives (e.g., power generation efficiency, flood control effectiveness). In one or more embodiments of this specification, this reward function includes at least a water level reward function and a flow rate reward function. The water level reward function ensures that the reservoir water level is maintained within a safe range, while the flow rate reward function ensures that the amount of water passing through the reservoir per unit time is maintained within a safe range.
[0052] In one or more embodiments of this specification, both the state space and action space are selectable. Users can select different state spaces and action spaces based on different scenarios and scheduling objectives. In this specification, the action space is the reservoir gate opening. In other words, the reservoir scheduling model in this specification is the model used to determine the reservoir gate opening. When the user selects a different action space, the reservoir scheduling model can be trained to execute other decisions.
[0053] The environment simulation unit can load a preset hydrological model based on the state space, action space, and reward function selected by the user, and initialize the simulation environment to simulate the dynamic changes in the state of the reservoir. The hydrological model can be used to simulate the actual operation of the reservoir. It should be understood that when initializing the simulation environment, not only the hydrological model may be used, but other tools or components, such as random event generators, may also be used, and this manual does not impose specific restrictions. The simulation environment is intended to provide a virtual test scenario that interacts with the reservoir operation model.
[0054] The environment simulation unit can initialize the reservoir's water level, inflow, and rainfall based on a hydrological model, as well as the reservoir's gate opening, and evaluate the reward value of the current action based on a user-selected reward function. It can also receive actions from the agent unit to update the reservoir's status.
[0055] The agent unit can initialize the parameters of the reservoir operation model, such as the weights of the neural network, based on the state space and action space selected by the user. The agent unit can use the reservoir operation model to determine the target action and target value corresponding to the current state and return the target action and target value to the environment simulation unit. As previously mentioned, since the action space includes the reservoir gate opening, the reservoir operation model can determine the target gate opening and the value corresponding to the target gate opening based on the current state.
[0056] After receiving the target action sent by the agent, the environment simulation unit simulates the process of updating the reservoir's current state based on the target action within the simulation environment using the loaded hydrological model. This process then generates an updated state corresponding to the current state. The reward value corresponding to the target action is determined based on the user-selected reward function. The reward value and updated state are then returned to the agent unit.
[0057] The intelligent unit can at least determine the actual water storage capacity of the reservoir based on the updated state, and can adjust the parameters of the reservoir scheduling model based on the reward value, the actual water storage capacity, the target action and the target value. Based on the above, it can be seen that the present application aims to constrain the adjustment of the parameters of the reservoir scheduling model based on the physical laws that are followed during the dynamic change of the reservoir, so as to improve the accuracy of the reservoir scheduling model prediction. Therefore, when the action space at least includes the reservoir gate opening, the physical law can be the reservoir water storage capacity change rate. That is to say, based on the gate opening predicted by the reservoir scheduling model, the actual water storage capacity of the reservoir can be simulated and the time length for the reservoir to change from the initial state to the updated state can be determined, so that the actual water storage capacity change rate can be calculated. At the same time, the predicted water storage capacity change rate can also be calculated based on the predicted gate opening and time length. Therefore, according to the actual water storage capacity change rate and the predicted water storage capacity change rate, the parameters of the reservoir scheduling model can be adjusted.
[0058] Repeat the above process. When the specified conditions are met, such as the number of iterative training reaches the specified number, the training can be terminated to obtain a trained reservoir scheduling model. In practical applications, the reservoir scheduling model can be used to determine the gate opening of the reservoir.
[0059] like Figure 2 As shown, Figure 2 This is a schematic diagram of a reservoir scheduling system provided in the present application specification. In one or more embodiments of this specification, the system may also include a reservoir scheduling model scheduling engine and a decision execution unit, and the system is used to apply the reservoir scheduling model.
[0060] After the reservoir scheduling model is trained, the reservoir scheduling model scheduling engine can generate scheduling decisions in real time according to the trained reservoir scheduling model and send them to the decision execution unit, so that the decision execution unit can control the opening of the reservoir gate based on the received scheduling decision.
[0061] based on Figure 1 as well as Figure 2 The reservoir operation system shown here constrains parameter adjustments based on the actual physical laws that the reservoir operation model conforms to, improving the predictive capabilities of the trained reservoir operation model and thus enhancing safety in practical applications. It also supports configurations for different state spaces, action spaces, and reward functions to meet diverse user needs.
[0062] This application specification provides a training method for a reservoir operation model. Figure 3 As shown, Figure 3 This is a flow chart of a training method for a reservoir operation model provided in this specification, which specifically includes the following steps:
[0063] S300: Obtaining a first state of a reservoir; wherein the state of the reservoir includes at least a water level, an inflow flow, and a rainfall amount, and the reservoir is a simulated reservoir in a simulation environment provided based on a preset hydrological model.
[0064] The execution subject of the training method for the reservoir operation model can be any of the aforementioned reservoir operation systems, or any computing device with computing capabilities, such as a server, terminal, etc. For the convenience of description, the following description uses the reservoir operation system as the execution subject.
[0065] The reservoir scheduling system can obtain a first state of the reservoir. This first state can be an initialized state of the reservoir, where the state space includes at least water level, inflow, and rainfall. As previously described, a user can select a state space, and the reservoir scheduling system can determine the state space of the user-selected reservoir in response to the user's operation. Based on the state space of the user-selected data, the system can initialize the state of the reservoir and use the initialized state of the reservoir as the first state of the reservoir. That is, in this specification, the state of the reservoir can include the water level, inflow, and rainfall of the reservoir.
[0066] In one or more embodiments of the present specification, during the training process, the reservoir is a simulated reservoir in a simulation environment provided based on a preset hydrological model.
[0067] S302: Input the first state into the first strategy network to obtain a first gate opening of the reservoir; and input the first state and the first gate opening into the first value network to obtain a first value.
[0068] like Figure 4 As shown in the figure, a schematic diagram of the structure of a reservoir operation model provided in the present application specification is shown. It can be seen that the reservoir operation model includes a first strategy network, a second strategy network, a first value network, and a second value network. The reservoir operation system can input a first state into the first strategy network to obtain a first gate opening output by the first strategy network. The first strategy network is used to determine the gate opening based on the current state. Then, the first gate opening and the first state can be input into the first value network to obtain a first value. The first value refers to the evaluation value of the first gate opening output by the second strategy network for the first state by the first value network. In other words, the first value network is used to evaluate the value of the gate opening determined by the first strategy network based on the current state.
[0069] S304: Determine a second state of the reservoir using the hydrological model according to the first gate opening and the first state; and determine a reward value.
[0070] As previously described in this specification, the reservoir operation system can initialize a simulation environment based on a preset hydrological model to simulate the dynamic changes of the reservoir. Based on the first gate opening and the first state, the reservoir operation system can determine, using the hydrological model, the second state of the reservoir after the gate opening of the reservoir is updated to the first gate opening, and determine a reward value.
[0071] In one or more embodiments of the present specification, the reward value is determined by the reservoir scheduling system based on a preset reward function, and the reward function at least includes: a water level reward function and a flow reward function.
[0072] Specifically, the water level reward function is used to ensure that the water level of the reservoir is maintained within the water level safety range and as close to the target water level as possible to achieve safe operation and reasonable water storage of the reservoir. The water level reward function is specifically shown in the following formula:
[0073] .
[0074] in, is the target water level, is the upper limit of the water level safety range, It is the lower limit of the water level safety range, which can be predetermined based on actual needs. is the water level of the reservoir in the second state. is the water level reward value.
[0075] That is, according to the water level reward function, when it is determined that the water level in the second state is within the preset water level safety range, the first formula is used , determine the water level reward value, when it is determined that the water level in the second state is not less than the upper limit of the preset water level safety range, use the second formula , determine the water level reward value.
[0076] The flow reward function is used to ensure that the flow of the reservoir is maintained within the flow safety range, as shown in the following formula:
[0077] .
[0078] in, is the inflow flow in the second state, It is the maximum inflow supported by the reservoir, which can be predetermined based on actual demand. The traffic reward value.
[0079] That is, according to the flow reward function, when it is determined that the inflow in the second state is not greater than the maximum inflow supported by the reservoir, the flow reward value is zero. When it is determined that the inflow in the second state is greater than the maximum inflow supported by the reservoir, the third formula is used. , determine the traffic reward value.
[0080] Therefore, the aforementioned reward value can be obtained according to the obtained flow reward value and water level reward value.
[0081] By determining the reward value in the above way, the safe operation and reasonable water storage of the reservoir can be achieved, and the flood risk downstream of the reservoir can be avoided.
[0082] Of course, other reward functions can also be designed and other reward values can be determined, such as an action execution reward function, which can be used to encourage the reservoir scheduling model to generate reasonable and stable actions and avoid system instability due to frequent or excessive action adjustments. This manual does not impose specific restrictions and relevant reward functions can be designed according to actual needs.
[0083] S306: Determine the actual water storage capacity of the reservoir when the gate opening of the reservoir is the first gate opening according to the second state; and input the second state into the second strategy network to obtain the second gate opening of the reservoir, and input the second gate opening and the second state into the second value network to obtain the second value.
[0084] When training a reservoir operation model based on reinforcement learning technology, if only the strategy and value loss corresponding to the reservoir operation model are considered, while ignoring the physical laws that the reservoir state conforms to during the change process, the strategy output by the reservoir operation model will be inaccurate. Therefore, in this specification, the physical laws that the reservoir state changes in accordance with are considered, that is, when the state of the reservoir changes from one state to another, the actual rate of change of the reservoir's water storage capacity should be consistent with the theoretical rate of change of the water storage capacity calculated based on the predicted gate opening. Therefore, in this specification, the reservoir operation system can determine the actual water storage capacity of the reservoir when the gate opening of the reservoir is the first gate opening based on the second state, and simulate the process of the reservoir changing from the first state to the second state through the hydrological model, and obtain the actual water storage capacity of the reservoir in the second state. The reservoir scheduling system can also determine the outflow of the reservoir based on the opening of the first gate, and determine the predicted water storage capacity of the reservoir based on the outflow and the first state. Specifically, the predicted water storage capacity of the reservoir can be calculated based on the inflow when the reservoir changes from the first state to the second state, the water storage capacity in the first state and the outflow.
[0085] It should be noted that when determining the actual water storage capacity of the reservoir, in a simulation environment, the hydrological model can simulate the gate opening being switched to the first gate opening in the first state, and the state of the reservoir after the simulated state switching can be obtained. The reservoir dispatching system can monitor the actual water storage capacity of the reservoir after the state switching, or based on the water level of the reservoir after the simulated switching and the shape and size of the reservoir, use the capacity curve method or the stratified volume method, etc. to determine the actual water storage capacity of the reservoir. When determining the outflow flow of the reservoir according to the first gate opening, it can be determined according to the hydraulic formula. There are many calculation methods for determining the actual water storage capacity and the outflow flow. This manual does not specifically limit the above-mentioned methods for determining the actual water storage capacity and the outflow flow.
[0086] Furthermore, the reservoir operation model can input the second state into a second policy network to obtain the second gate opening of the reservoir output by the second policy network. The second policy network is used to determine the action based on the state obtained after executing the action output by the first policy network. In other words, the second policy network is used to determine the gate opening of the reservoir in the second state. The second gate opening and the second state can then be input into a second value network to obtain a second value. The second value network is used to evaluate the value of the action output by the second policy network. It can be seen that the input data for the second value network and the second policy network when making predictions depends on the outputs of the first policy network and the first value network. The first value network and the first policy network are used to predict the action and evaluate the value of the current state, while the second value network and the second policy network are used to predict the action and evaluate the value of the next state corresponding to the current state.
[0087] In general, the second value network and the second policy network are used to assist in correcting the parameter adjustments of the first policy network and the first value network, preventing the first policy network and the first value network from becoming unstable due to frequent and significant changes during the learning process. Therefore, in this specification, the reservoir operation system can adjust the parameters of the reservoir operation model based on the first value and the second value.
[0088] S308: Adjusting the parameters of the reservoir scheduling model according to the first gate opening, the actual water storage capacity, the reward value, the first value, and the second value.
[0089] As mentioned above, the reservoir scheduling system can determine the outflow rate of the reservoir based on the opening of the first gate, and further determine the predicted water storage capacity of the reservoir. Therefore, when adjusting the parameters of the reservoir scheduling model according to the opening of the first gate and the actual water storage capacity, the parameters of the reservoir scheduling model can be adjusted based on the predicted water storage capacity and the actual water storage capacity. In one or more embodiments of the present specification, the reservoir scheduling model can determine the duration for the state of the reservoir to be updated from the first state to the second state. Specifically, the process of the reservoir changing from the first state to the second state can be simulated by a hydrological model, thereby determining the duration for the state of the reservoir to be updated from the first state to the second state. Furthermore, based on this duration and the actual water storage capacity, the time inverse of the actual water storage capacity can be calculated by differential or difference calculation to obtain the actual water storage capacity change rate. And, based on this duration and the predicted water storage capacity, the predicted water storage capacity change rate is determined.
[0090] Then, based on the obtained predicted water storage capacity change rate and the actual water storage capacity change rate, a first loss can be determined, and the parameters of the reservoir operation model can be adjusted according to the first loss. In one or more embodiments of this specification, the first loss can be the mean square error loss of the residual of the water storage capacity change rate.
[0091] In this specification, the adjustment of the parameters of the reservoir operation model is divided into three parts: adjustment of the parameters of the first value network, adjustment of the parameters of the first policy network, and adjustment of the parameters of the second value network and the second policy network. Furthermore, in one or more embodiments of this specification, during the training of the reservoir operation model, for each round of training, the order of adjusting the parameters of the first value network is before adjusting the parameters of the first policy network, and the order of adjusting the parameters of the second value network and the second policy network is after adjusting the parameters of the first policy network.
[0092] When adjusting the parameters of the first value network, the first loss can be first obtained based on the above method. When the reservoir operation system adjusts the parameters of the reservoir operation model based on the reward value, the first value, and the second value, it can determine the second loss based on the first and second values, and adjust the parameters of the first value network based on the second loss, the first loss, and the reward value. The second loss can be the mean squared error between the first and second values.
[0093] In one or more embodiments of the present specification, the sum of the second loss, the first loss, and the reward value can be used as the total loss. Based on the total loss, a gradient that minimizes the total loss is determined, and the parameters of the first value network are adjusted according to the direction of gradient descent. Alternatively, the second loss, the first loss, and the reward value can be weighted based on a first preset weight to obtain the weighted losses. Based on the sum of the weighted losses, a gradient that minimizes the sum of the weighted losses is determined, and the parameters of the first value network are adjusted according to the direction of gradient descent.
[0094] After adjusting the parameters of the first value network, the parameters of the first policy network can be adjusted. Specifically, the first gate opening and the first state can be re-input into the first value network to obtain a third value output by the adjusted first value network. In other words, the first gate opening predicted by the first policy network is re-evaluated based on the adjusted first value network. Based on this third value, a third loss can be determined. Since the value used by the first value network to evaluate the first gate opening is also used to evaluate the score or quality of the action output by the first policy network, the third loss can be determined with the goal of maximizing the third value. The parameters of the first policy network can then be adjusted based on the third loss and the first loss.
[0095] Similarly, when adjusting the parameters of the first policy network based on the third loss and the first loss, the parameters of the first policy network can be adjusted based on the sum of the first loss and the third loss, and the gradient that minimizes the sum of the first loss and the third loss can be determined, and the parameters of the first policy network can be adjusted according to the direction of gradient descent. Alternatively, the first loss and the third loss can be weighted based on a second preset weight to obtain the weighted first loss and third loss, and then the parameters of the first policy network can be adjusted based on the sum of the weighted first loss and the third loss, and the gradient that minimizes the sum of the weighted first loss and the third loss can be determined, and the parameters of the first policy network can be adjusted according to the direction of gradient descent.
[0096] After adjusting the parameters of the first value network and the first network, the parameters of the second strategy network and the second value network can be adjusted using a soft update based on the adjustments to the parameters of the first value network and the first strategy network. Soft updates are currently available and are not described in detail in this specification.
[0097] According to the above method, iterative training of the reservoir scheduling model is performed. It should be noted that this specification does not impose any specific restrictions on when the training of the reservoir scheduling model is determined to be completed. For example, the training of the reservoir scheduling model is determined to be completed when the number of training iterations reaches a preset threshold, or when the determined loss is less than a preset value.
[0098] based on Figure 3 The training method of the reservoir scheduling model shown in the figure adjusts the parameters of the reservoir scheduling model. The adjustment direction not only includes the loss corresponding to the value output by the reservoir scheduling model itself, the loss corresponding to the strategy, and the constraint of the reward value, but also considers the physical constraint, that is, the change of the reservoir water storage capacity should conform to the physical laws, thereby improving the prediction accuracy of the trained reservoir scheduling model.
[0099] In one or more embodiments of the present specification, the aforementioned preset hydrological model may be a basin hydrological model based on pygrdhm, such as the Guiren hydrological model pygrdhm, and the environment based on the hydrological model is a basin hydrological model environment based on pygrdhm. The Guiren hydrological model pygrdhm simulates the hydrological processes of the basin, including hydrological processes such as rainfall runoff, reservoir storage and discharge, and river confluence. The GrdhmEnv environment module can be developed in advance to provide a platform for the reservoir scheduling model to interact with the basin hydrological system. The aforementioned preset hydrological model may also be the Xin'anjiang model, which is a distributed basin hydrological model that divides the basin into multiple calculation units and considers processes such as rainfall, soil water storage, and groundwater flow to calculate the runoff output of the basin. This specification does not limit what specific model the hydrological model is.
[0100] Furthermore, as previously mentioned, the second value network is used to approximate the value of the state-action pair at the next moment from the current moment, while the action at the next moment is approximated by the second policy network. In one or more embodiments of this specification, after obtaining the value output by the second network, physical losses can be determined in conjunction with the expert system's scheduling plan to train the reservoir operation model. The expert system's scheduling plan involves constraints such as discharge capacity curves, sluice and weir facility constraints, control rule constraints, and gate opening constraints.
[0101] The discharge capacity curve constraint specifies that, given a specified reservoir discharge capacity curve, the outflow cannot exceed the maximum discharge capacity of the reservoir at each water level. The sluice and weir facility constraint specifies that, given a specified reservoir sluice and weir facility, the outflow cannot exceed the maximum flow capacity of the reservoir sluice and weir. The control rule constraint specifies that, given a specified reservoir control rule, the outflow cannot exceed the outflow limit specified by the control rule. The gate opening constraint specifies that the gate opening cannot exceed the maximum gate opening.
[0102] Among them, the discharge capacity curve constraint is expressed as , the sluice and weir facility constraints are expressed as , the control rule constraint is expressed as , the gate opening constraint is expressed as . Then, the constraints corresponding to the scheduling scheme of the comprehensive expert system are:
[0103] ,
[0104] ,
[0105] .
[0106] in: Generates outflow from a reservoir for an Actor network. It is the maximum discharge capacity of the reservoir at water level H, determined by the discharge capacity curve. It is the maximum flow capacity of the sluice weir when the reservoir is at water level H. This is the upper outflow limit specified by the reservoir control rule. It is the comprehensive maximum outflow of the reservoir when the water level is H. : The opening of the gate generated by the Actor network. is the maximum opening of the gate. is the outflow of the expert system reservoir. is the opening of the expert system gate. is the scheduling plan update rate. _real is the actual reservoir outflow implemented by the environment. _real is the gate opening actually executed by the environment.
[0107] Therefore, the value output by the second policy network is: .
[0108] In one or more embodiments of this specification, the reservoir operation system may also use a Gaussian noise function to process the action input into the reservoir operation model. The value output by the first value network is: .
[0109] Then, the parameters of the first value network are updated by minimizing the loss value (mean square error loss). The loss function when the first value network is updated is:
[0110] ,
[0111] .
[0112] in, , is the Gaussian noise function. The aforementioned physical loss, that is, the loss determined based on gate opening constraints or outflow flow constraints, Losses determined based on other data items.
[0113] When adjusting the parameters of the reservoir operation model, the total loss can be minimized. To update the parameters of the reservoir operation model ,in, is a hyperparameter that balances the two losses.
[0114] That is, in step S308, when adjusting the parameters of the first value network based on the second loss, the first loss, and the reward value, the parameters of the first value network may be adjusted based on the physical loss, the second loss, the first loss, and the reward value. When adjusting the parameters of the first policy network based on the third loss and the first loss, the parameters of the first policy network may be adjusted based on the physical loss, the third loss, and the first loss.
[0115] The physical loss training of the reservoir operation model based on the expert system such as the aforementioned gate opening constraints and outflow flow constraints improves the prediction performance of the trained reservoir operation model and avoids the reservoir operation model predicting data that does not conform to the actual physical laws of the reservoir.
[0116] Furthermore, in one or more embodiments of the present specification, an experience replay mechanism may be introduced, that is, the data generated by the interaction between the first strategy network and the second strategy network and the simulation environment are stored as experience data samples in a preset experience pool, so that when the reservoir scheduling model is trained, batches of these experience data samples are extracted for training. In other words, the data samples used to train the reservoir scheduling model may come from the data generated by the interaction between the first strategy network and the second strategy network and the simulation environment, thereby removing the correlation and dependency of the data samples and making the reservoir scheduling model easier to converge.
[0117] Furthermore, in one or more embodiments of the present specification, the first policy network and the first value network each include a feature extraction layer and a fully connected layer. The feature extraction network is used to extract features based on input data, and the fully connected layer is used to output results based on the features extracted by the feature extraction network. The network layer structure of the second policy network can be the same as that of the first policy network, and the structure of the second value network can be the same as that of the first value network.
[0118] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0119] In the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," "connect," "fixed," etc. should be understood broadly. For example, they may refer to fixed connection, detachable connection, or integration; mechanical connection or electrical connection; direct connection or indirect connection through an intermediate medium; and internal communication between two components or interaction between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0120] In the present invention, unless otherwise expressly specified or limited, when a first feature is "above" or "below" a second feature, it may mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. Furthermore, when a first feature is "above," "above," or "above" a second feature, it may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is at a higher level than the second feature. When a first feature is "below," "below," or "below" a second feature, it may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is at a lower level than the second feature.
[0121] In the description of this specification, the terms "one embodiment", "some embodiments", "embodiments", "examples", "specific examples" or "some examples" refer to the specific features, structures, materials or characteristics described in conjunction with the embodiment or example and included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and features of different embodiments or examples, unless they are mutually inconsistent.
[0122] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may alter, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A training method for a reservoir operation model, characterized in that: The reservoir operation model includes a first strategy network, a second strategy network, a first value network and a second value network; the method includes: Acquire a first state of a reservoir; wherein the state of the reservoir includes at least a water level, an inflow, and a rainfall, and the reservoir is a simulated reservoir in a simulation environment provided based on a preset hydrological model; Inputting the first state into the first strategy network to obtain a first gate opening of the reservoir; and inputting the first state and the first gate opening into the first value network to obtain a first value; Determining a second state of the reservoir using the hydrological model according to the first gate opening and the first state; and determining a reward value; Determine, based on the second state, the actual water storage capacity of the reservoir when the gate opening of the reservoir is the first gate opening; input the second state into the second strategy network to obtain a second gate opening of the reservoir; input the second gate opening and the second state into the second value network to obtain a second value; adjusting the parameters of the reservoir operation model according to the first gate opening, the actual water storage, the reward value, the first value, and the second value; Adjusting the parameters of the reservoir operation model specifically includes: determining the outflow of the reservoir based on the first gate opening; and determining the time duration for the state of the reservoir to be updated from the first state to the second state under a simulation environment provided by the hydrological model; Determine the actual water storage capacity change rate based on the actual water storage capacity and the duration; and determine the predicted water storage capacity change rate based on the outflow flow, the first state and the duration; determining a first loss according to the actual water storage capacity change rate and the predicted water storage capacity change rate; adjusting parameters of the reservoir operation model according to the first loss; Adjusting the parameters of the reservoir operation model specifically includes: determining a second loss based on the first value and the second value; and adjusting parameters of the first value network based on the second loss, the first loss, and the reward value; determining a third loss based on a third value corresponding to the first gate opening output by the first value network after parameter adjustment; and adjusting parameters of the first strategy network based on the third loss and the first loss; Adjusting the parameters of the reservoir operation model specifically includes: The parameters of the second strategy network and the second value network are adjusted by soft updating.
2. The method according to claim 1, wherein During the training process, for each round of training, the order of adjusting the parameters of the first value network is before adjusting the parameters of the first policy network, and the order of adjusting the parameters of the second value network and the second policy network is after adjusting the parameters of the first policy network.
3. The method according to claim 1, wherein The reward value is determined by the hydrological model based on a preset reward function, and the reward function at least includes: a water level constraint function and a flow constraint function.
4. The method according to claim 1, wherein Get the first state of the reservoir, including: In response to a user operation, determining a state space of the reservoir selected by the user, the state space including at least water level, inflow, and rainfall; Based on the state space of the reservoir selected by the user, the state of the reservoir is initialized, and the initialized state of the reservoir is used as the first state of the reservoir.
5. The method according to claim 1, wherein The method further comprises: Obtaining the status of a target reservoir, wherein the status of the target reservoir includes water level, inflow, and rainfall; Inputting the state of the target reservoir into a trained reservoir scheduling model to obtain a target gate opening output by the trained reservoir scheduling model; Based on the target gate opening, the gate opening of the target reservoir is controlled.
6. A reservoir dispatching system, characterized in that: The system is used to execute the method according to any one of claims 1 to 5 above, and the system comprises: a model scheduling interface, an environment simulation unit, and an agent unit; The model scheduling interface is used to determine the state space, action space and reward function selected by the user in response to the user's operation; the state space includes at least: water level, inflow and rainfall, and the action space includes at least reservoir gate opening; The environment simulation unit is configured to load a preset hydrological model based on the state space, action space, and reward function selected by the user, and initialize a simulation environment to simulate dynamic changes in the state of the reservoir; The intelligent agent unit is used to initialize the parameters of the reservoir operation model based on the state space and action space selected by the user, and determine the target action and target value corresponding to the current state through the reservoir operation model; and send the current state and the target action to the environment simulation unit; The environmental simulation unit is configured to simulate, in the simulation environment, a process of updating the current state of the reservoir based on the target action using the loaded hydrological model, thereby obtaining an updated state corresponding to the current state; determine a reward value corresponding to the target action based on the reward function selected by the user; and send the reward value and the updated state to the agent unit; The intelligent agent unit is used to determine at least the actual water storage capacity of the reservoir based on the updated state; and adjust the parameters of the reservoir scheduling model based on the reward value, the actual water storage capacity, the target action and the target value.
7. The system according to claim 6, wherein: The system further comprises: a reservoir operation model scheduling engine and a decision execution unit; The reservoir operation model scheduling engine is used to generate scheduling decisions based on the trained reservoir operation model; The decision execution unit is used to receive the scheduling decision; and control the gate opening of the reservoir according to the scheduling decision.
Citation Information
Patent Citations
Reservoir gate group multi-target flood control optimization scheduling method and system
CN115719041A
Brake pump group joint optimization scheduling method based on Multi-Agent PPO reinforcement learning
CN116738874A