Reservoir gate dispatching method and system
Through the construction and training of deep reinforcement learning models, the problems of gate opening number and flow allocation errors in traditional reservoir flood control scheduling are solved, and more efficient and accurate reservoir flood control scheduling is achieved.
Patent Information
- Application Number
- CN202510004872.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-02
AI Technical Summary
In traditional reservoir flood control scheduling, there are errors when distributing outflows to the gate, and the number of openings and flow distribution of the gates cannot be effectively optimized.
By constructing a deep reinforcement learning model, multiple constraints are determined, reward functions are established, and the probability distribution of the number of gate openings is trained to output the model, and the gate is opened according to this distribution.
The problem of errors in the allocation of outbound flow to the gate is solved, more accurate and efficient gate scheduling is achieved, and the optimization effect of reservoir flood control scheduling is improved.
Smart Images

Figure CN120069374A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of gate scheduling, and particularly relates to a reservoir gate scheduling method and system. Background Art
[0002] Among all natural disasters, flood disasters have the most serious impacts. At the same time, the proportion of the global population facing flood risks is increasing year by year. In the face of the increasingly severe flood risks, scientific researchers around the world have carried out a large number of research works aimed at deeply understanding the flood formation mechanism, prediction methods, and prevention and control technologies, providing a scientific basis for flood risk management and response.
[0003] Flood disasters have always been an important restricting factor for the social and economic development of our country. According to statistics, the economic losses caused by floods in our country rank first among various disasters, seriously affecting the national economic stability and social sustainable development. The frequent occurrence of flood disasters not only has a huge impact on agricultural production, urban construction, and infrastructure, but also directly threatens the lives and property safety of the people.
[0004] As one of the widely adopted measures for flood control in our country, reservoirs play an important role in realizing the safe and efficient utilization of flood resources. The flood control scheduling of reservoirs realizes the regulation of floods by adjusting the outflow discharge of the reservoirs, aiming to reduce flood losses, protect the safety of protected objects, and improve the comprehensive benefits of the reservoirs.
[0005] However, in traditional reservoir flood control scheduling, the decision variable is the outflow discharge, and the flow is not distributed to the gates. Due to the operation rules such as the opening and closing sequence and symmetric opening and closing of the spillway gates of the reservoir, there will be errors when the outflow discharge is distributed to the gates, and there is room for further optimization in the traditional scheduling scheme. Summary of the Invention
[0006] To solve the above problems, the present disclosure provides a reservoir gate scheduling method and system. By constructing a deep reinforcement learning model, determining various constraint conditions, establishing a reward function using the constraint conditions and the objective function, training the model to output the probability distribution of the number of opened gates, and opening the gates according to this distribution, the problem of errors in distributing the outflow discharge to the gates is solved.
[0007] The following is the technology of the present invention:
[0008] A reservoir gate scheduling method, characterized by comprising:
[0009] Obtaining the basic information of the reservoir;
[0010] Construct the objective function and constraint conditions for the multi-objective optimal scheduling of the reservoir spillway gate based on the basic information of the reservoir; the objective of the objective function is to minimize the sum of the maximum reservoir outflow and the average adjustment times of the reservoir spillway gate; the constraint conditions include water balance constraint, reservoir water level constraint, storage capacity curve constraint, maximum discharge capacity constraint, variable range of outflow constraint, and spillway discharge constraint;
[0011] Construct a deep reinforcement learning framework; the state variables of the deep reinforcement learning framework are the reservoir water level at the beginning of time step t, the inflow of the reservoir at time step t, the outflow of the reservoir at time step t - 1, and the gate combination situation; the action variable is the gate combination situation at time step t; the reward function of the deep reinforcement learning framework is constructed based on the objective function and constraint conditions;
[0012] The construction of the reward function includes:
[0013] Determine the basic reward function according to the objective function, and the basic reward function is:
[0014] Provide an immediate reward in each time step, and incorporate the reservoir water level constraint as a penalty term into the reward function to obtain the final reward function as:
[0015]
[0016] In the formula, α' and β' are weight coefficients; c 1 and c 2 are penalty coefficients;
[0017] Use the proximal policy optimization algorithm to train the deep reinforcement learning framework to obtain a trained deep reinforcement learning model; use the trained deep reinforcement learning model for reservoir gate scheduling.
[0018] Furthermore,
[0019] The constraint conditions for the reservoir spillway gate scheduling are:
[0020] Water balance constraint:
[0021] V t+1 = V t +(I t - Q t )·Δt
[0022]
[0023] In the formula, V t is the reservoir storage volume at the beginning of time step t, m 3 ; I t is the reservoir inflow at time step t, m3 / s; Q t is the reservoir outflow at time step t, m3 / s; Δt is the duration of time period t, h; M is the number of gate types; J m is the number of gates of the m-th type; q j,m,t is the discharge of the j-th gate of the m-th type in time period t, m 3 / s;
[0024] Reservoir water level constraint:
[0025] Z min ≤Z t ≤Z max
[0026] In the formula, Z t is the reservoir water level at the beginning of time period t, m; Z max and Z min are the upper and lower limits of the reservoir water level respectively, m;
[0027] Storage curve constraint:
[0028] Z t =f zv (V t )
[0029] In the formula, f zv (·) is the functional relationship between the reservoir water level and the water storage volume;
[0030] Maximum discharge capacity constraint:
[0031] Q t ≤Q t,max
[0032] In the formula, Q t,max is the maximum discharge capacity of the reservoir in time period t, m 3 / s;
[0033] Outflow variation constraint:
[0034] |Q t+1 -Q t |≤ΔQ
[0035] In the formula, ΔQ is the allowable variation of the reservoir outflow between adjacent time periods, m 3 / s;
[0036] Spillway discharge constraint:
[0037] q j,m,t =f m,qz (O j,m,t ,Z t )
[0038] In the formula, O j,m,t is the opening of the j-th gate of the m-th type in time period t; fm,qz (·) is the functional relationship between the discharge of the m-th type of gate, the gate opening, and the upstream water level;
[0039] Furthermore,
[0040] The objective function of the multi-objective optimal scheduling of the reservoir spillway gate is:
[0041]
[0042] In the formula, T is the total scheduling duration, h; t is the current scheduling period, h; G is the average number of gate adjustments; α and β are weight coefficients.
[0043] Furthermore,
[0044] In the proximal policy optimization algorithm, the objective function of the policy network is:
[0045]
[0046] The value network expression is:
[0047]
[0048] In the formula: π θ (a t ∣s t ) is the policy of the agent; is the estimated value of the advantage function, calculated using the generalized advantage estimation method; ∈ is a hyperparameter used to limit the magnitude of policy updates; c H is a hyperparameter used to control the strength of entropy regularization; V φ (s t ) is the estimated value of the value network for the state s t ; R t is the actual return starting from the state s t ;
[0049] Furthermore,
[0050] The
[0051] calculated using the generalized advantage estimation method expression is:
[0052]
[0053] In the formula: γ is the discount factor; λ is a hyperparameter used to control the bias and variance.
[0054] Furthermore,
[0055] Training the deep reinforcement learning framework using the proximal policy optimization algorithm to obtain a trained deep reinforcement learning model, including:
[0056] Initialize the parameters of the policy network and the value network;
[0057] Obtain the initial environmental state s of the reservoir 0 , that is, the inflow and reservoir water level at the initial moment, as well as the outflow and gate combination conditions at the previous moment;
[0058] The policy network outputs the probability distribution of different gate combination conditions according to the initial state, and performs the action a of the reservoir spillway gate according to the probability distribution of different gate combination conditions t ;
[0059] Execute the action a t After that, obtain the reward value r t , and at this time the environmental state becomes s t+1 ;
[0060] Store the experience (s t , a t , r t , s t+1 ) into the buffer. When the number of experiences reaches the buffer capacity, maximize to update the parameters of the policy network, and minimize to update the parameters of the value network;
[0061] Every time the parameters of the policy network and the value network are updated, evaluate the current policy π θ (a t |s t ) until the obtained cumulative reward value tends to be stable, and obtain a trained deep reinforcement learning model.
[0062] A reservoir spillway gate scheduling system, characterized by including:
[0063] A data acquisition module for acquiring basic reservoir information;
[0064] A framework construction module for constructing the objective function and constraint conditions for the multi-objective optimal scheduling of the reservoir spillway gate based on the basic reservoir information; the objective of the objective function is to minimize the sum of the maximum reservoir outflow and the average adjustment times of the reservoir spillway gate; the constraint conditions include water balance constraint, reservoir water level constraint, storage capacity curve constraint, maximum discharge capacity constraint, outflow variation constraint, and spillway discharge constraint;
[0065] A model construction module for constructing a deep reinforcement learning framework; the state variables of the deep reinforcement learning framework are the reservoir water level at the beginning of period t, the inflow rate in period t, the outflow rate in period t-1, and the gate combination situation; the action variable is the gate combination situation in period t; the reward function of the deep reinforcement learning framework is constructed based on the objective function and the constraint conditions;
[0066] The construction of the reward function includes:
[0067] Determine the basic reward function according to the objective function, and the basic reward function is:
[0068] Provide an immediate reward in each time step, and incorporate the reservoir water level constraint as a penalty term into the reward function to obtain the final reward function as:
[0069]
[0070] In the formula, α' and β' are weight coefficients; c 1 and c 2 are penalty coefficients;
[0071] A model training module for training the deep reinforcement learning framework using the proximal policy optimization algorithm to obtain a trained deep reinforcement learning model;
[0072] A control module for using the trained deep reinforcement learning model to perform reservoir gate scheduling.
[0073] Compared with the prior art, the present disclosure has the following advantages:
[0074] The present invention provides a data basis for the follow-up by obtaining the basic information of the reservoir, including the reservoir inflow rate in a certain determined period, the water level-capacity curve, the discharge capacity curve of the water discharge facility, etc.; constructing the objective function while considering the maximum outflow rate and the number of gate openings and closings, paying attention to both flood control safety and the operation efficiency and service life of the gates; and then using the deep reinforcement learning model to automatically make decisions on the opening of the gates.
[0075] In the training of the deep reinforcement learning model, the proximal policy optimization algorithm is used, including a policy network and a value network. The policy network outputs the probability distribution of the number of gate openings based on the reservoir state, and the value network outputs the current state value. Determine the constraint conditions including various actual situations and establish a reward function. Use the upper and lower limits of the water level as penalty terms and provide intermediate rewards to solve the problem of sparse rewards. Train the model to output the probability distribution by inputting relevant period information, and finally open the gates according to the probability distribution, so as to realize the reservoir spillway gate scheduling under the constraint conditions and solve the problems of the traditional scheduling scheme.
[0076] Other features and advantages of the present disclosure will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present disclosure. The objectives and other advantages of the present disclosure may be realized and attained by the structure particularly pointed out in the specification, claims as well as the drawings. Description of the Drawings
[0077] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or in the prior art, the following briefly introduces the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings may be obtained based on these drawings.
[0078] Figure 1 Shows a schematic diagram of the method of the present invention;
[0079] Figure 2 Shows a flowchart of the model training process;
[0080] Figure 3 Shows a flowchart of the model testing process;
[0081] Figure 4 Shows a cumulative reward curve graph;
[0082] Figure 5 Shows a graph of the outbound flow process;
[0083] Figure 6 Shows a graph of the reservoir water level process;
[0084] Figure 7 Shows a schematic diagram of the number of deep holes opened in each period. Detailed Description of the Embodiments
[0085] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following clearly and completely describes the technical solutions in the embodiments of the present disclosure with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.
[0086] As Figure 1 is a schematic diagram of the method of the present invention, the specific implementation details of the present invention include:
[0087] 1. Step 1: Obtain the basic information of the reservoir.
[0088] The basic information includes: the reservoir inflow, the water level-storage curve, and the discharge capacity curve of the water discharge facilities during a certain period.
[0089] 2. Step 2: Determine the objective function.
[0090] Consider both the maximum outflow rate during the scheduling period and the number of gate openings and closings. Adopt the common maximum peak shaving criterion in reservoir flood control scheduling and ignore the inflow from the intermediate section. Take the maximum reservoir outflow rate as one objective function.
[0091] In addition, frequent opening and closing of the gates during the flood season will shorten the service life of the gates and increase the operating cost. Therefore, take the average number of adjustments of the reservoir spillway gates during the flood season as another objective function.
[0092] Meanwhile, use the weighting method to integrate the two objectives into a total objective function in the following form:
[0093]
[0094] In the formula, T is the total scheduling duration, h; t is the current scheduling period, h; G is the average number of gate adjustments; α and β are weighting coefficients.
[0095] 3. Step 3: Determine the constraint conditions.
[0096] The constraint conditions include:
[0097] 1) Water balance constraint:
[0098] V t+1 =V t +(I t -Q t )·Δt
[0099]
[0100] In the formula, V t is the reservoir water storage at the beginning of the t period, m 3 ; I t is the reservoir inflow rate during the t period, m 3 / s; Q t is the reservoir outflow rate during the t period, m 3 / s; Δt is the duration of the t period, h; M is the number of gate types; J m is the number of the m-th type of gates; q j,m,t is the discharge of the j-th gate of the m-th type during the t period, m 3 / s;
[0101] 2) Reservoir water level constraint
[0102] Z min ≤Z t ≤Z max
[0103] In the formula: Z t is the reservoir water level at the beginning of period t, in m; Z max and Z min are the upper and lower limits of the reservoir water level, in m respectively.
[0104] 3) Storage curve constraint
[0105] Z t = f zv (V t )
[0106] In the formula: f zv (·) is the functional relationship between the reservoir water level and the water storage volume.
[0107] 4) Maximum discharge capacity constraint
[0108] Q t ≤ Q t,max
[0109] In the formula: Q t,max is the maximum outflow discharge of the reservoir in period t, in m 3 / s.
[0110] 5) Constraint on the variation range of outflow discharge between adjacent periods
[0111] |Q t+1 - Q t | ≤ ΔQ
[0112] In the formula: ΔQ is the allowable variation range of the reservoir outflow discharge between adjacent periods, in m 3 / s.
[0113] 6) Spillway discharge constraint
[0114] q j,m,t = f m,qz (O j,m,t , Z t )
[0115] In the formula, O j,m,t is the opening of the j-th gate of the m-th type at time t; f m,qz (·) is the functional relationship between the discharge of the m-th type of gate and the gate opening and the upstream water level;
[0116] 4. Step Four: Determine the basic elements of the reinforcement learning model, including action variables, state variables, and reward functions.
[0117] Specifically, select the gate combination situation as the decision variable and define it as the action variable a t :
[0118] a t = {N t}
[0119] Select the reservoir water level at the beginning of period t, the inflow during period t, the outflow during period t-1, and the combination of reservoir spillway gates during period t-1 as the state variables of the environment. State s t is defined as:
[0120] s t ={I t ,Z t ,Q t-1 ,N t-1}
[0121] Reward is the most important element in reinforcement learning. A good reward function can guide the agent to train in the
[0122] "correct direction" so that it gradually learns the optimal policy. On the contrary, an inappropriate reward function will hinder the agent from learning a reasonable policy. Usually, the reward function is directly related to the objective function. Since the agent tends to learn the policy that maximizes the cumulative reward, the basic reward in the present invention
[0123] function is expressed as follows:
[0124]
[0125] During the reinforcement learning training process, the agent usually does not receive rewards. Only when a specific goal is achieved, the environment will provide a certain reward to the agent. In the present invention, according to the above reward function, the environment only provides rewards to the agent at the end of the scheduling period, and this reward reflects the maximum outflow of the reservoir and the number of gate adjustments during the entire scheduling period. This form of reward function makes it difficult for the agent to learn a reasonable policy and there is a problem of sparse rewards.
[0126] A common method to alleviate this problem is reward shaping. The environment provides intermediate rewards to the agent at each time step to guide the agent to learn a suitable policy. These intermediate rewards are usually much smaller than the final reward to highlight the final reward value. In addition, the actions selected by the agent must satisfy the reservoir spillway gate constraint conditions. In the present invention, the upper and lower limits of the reservoir water level are incorporated into the reward function as penalty terms. To sum up, the calculation formula of the reward function in the present invention is as follows:
[0127]
[0128] where: α' and β' are weight coefficients; c 1 and c 2 are penalty coefficients.
[0129] 5. Step Five: Determine the deep reinforcement learning algorithm and optimize a specific inflow process according to the elements of reinforcement learning.
[0130] The present invention selects the Proximal Policy Optimization (PPO) algorithm for solution.
[0131] The Proximal Policy Optimization algorithm is an improvement of the policy gradient method, mainly solving the stability and efficiency problems of the policy gradient method during policy update. The Proximal Policy Optimization algorithm improves the stability of training by restricting the amplitude of policy update. By introducing a "clip" mechanism, it prevents the policy from being updated too quickly, thus avoiding policy collapse. The core idea of the Proximal Policy Optimization algorithm is to use the idea of Trust Region Methods to control the size of each policy update, so that the difference between the new and old policies is not too large.
[0132] There are two neural networks in the Proximal Policy Optimization algorithm, namely the policy network and the value network. The policy network takes the state as input and outputs the probability of each possible action. The value network takes the state as input and outputs the value of that state, that is, the expected cumulative reward of the agent in that state.
[0133] The update objective of the policy network is to maximize the clip function, and the update objective of the value network is to minimize the mean squared error loss. The corresponding functional forms are as follows:
[0134]
[0135] In the formula: π θ (a t ∣s t ) is the policy of the agent; is the estimated value of the advantage function; ∈ is a very small value (such as 0.2) used to limit the amplitude of policy update; V φ (s t ) is the estimated value of the value network for state s t ; R t is the actual return starting from state s t .
[0136] The generalized advantage estimation method is used to calculate The expression is:
[0137]
[0138] In the formula: γ is the discount factor; λ is a hyperparameter used to control bias and variance.
[0139] 6. Step Six: Determine the hyperparameters of the deep reinforcement learning algorithm, including the total number of training times, learning rate, batch size, hidden layer size, etc.
[0140] 7. Step 7: Train the model. The agent and the environment interact continuously, and the neural network parameters are updated according to the loss function. Each time the parameters are updated, evaluate the current policy and output the cumulative reward curve graph.
[0141] Specifically, train according to the basic idea of the Proximal Policy Optimization algorithm. The specific training process is as follows:
[0142]
[0143] During the training phase, the agent calculates the probability distribution of actions according to the policy network and selects actions based on this probability distribution. Considering the constraints, when the agent selects an action, it no longer selects from the overall action space, but selects actions from the corresponding available action space according to the current received environmental state s t ,.
[0144] Such as Figure 2 is the flowchart of the model training process of the present invention.
[0145] 8. Step 8: According to the cumulative reward curve graph in Step 7, select the neural network parameters with the largest cumulative reward.
[0146] 9. Step 9: Use the network parameters in Step 8 as the network parameters during scheduling to determine the scheduling process.
[0147] The evaluation of the current policy in Step 7 and the determination of the scheduling process in Step 9. The interaction process between the agent and the environment is as Figure 3 shown. The agent calculates the probabilities of all actions in the corresponding state according to the policy network and selects the action with the largest probability in the corresponding available action space.
[0148] The following is a specific demonstration example of the present invention:
[0149] Taking a certain reservoir as an example, it specifically includes the following steps:
[0150] Step 1, input the basic information of the reservoir. Select 0:00 on August 11, 2020 to 24:00 on August 25, 2020 as the start and end times of scheduling, with a time resolution of 2 hours. Import the inflow sequence, water level-storage capacity curve, and discharge capacity curve of the spillway facilities of a certain reservoir. The form of the discharge capacity curve of the spillway facilities is shown in Table 1.
[0151] Table 1 Discharge capacity of spillway facilities
[0152]
[0153] Step 2, determine the objective function and select appropriate values of α and β.
[0154] Step 3, determine the constraint conditions. Select appropriate upper and lower limits of the reservoir water level, upper limit of the discharge flow rate, maximum variation range of the discharge flow rate in adjacent periods, initial reservoir water level, discharge flow rate at the initial moment, and gate combination.
[0155] Step 4, determine the basic elements of the reinforcement learning model. Select appropriate values of α', β', c 1 , c 2 .
[0156] Step 5, determine the deep reinforcement learning algorithm. Select the Proximal Policy Optimization (PPO) algorithm for solution.
[0157] Step 6, determine the hyperparameters of the deep reinforcement learning. The total number of training times is 2000, the learning rate is 0.0003, the batch size is 1024, and the hidden layer size is 128.
[0158] Step 7, train the model. The agent and the environment interact continuously. Each time the parameters are updated, the current policy is evaluated, and the cumulative reward curve graph is output, in the form as Figure 4 shown.
[0159] Step 8, according to the cumulative reward values that can be obtained by each policy in Step 4, select the network parameters with the largest cumulative reward value.
[0160] Step 9, use the network parameters selected in Step 8 to output the scheduling process, such as Figure 5 , 6 , as shown in 7.
[0161] Table 2 Scheduling Results
[0162]
[0163] During the flood process from August 11th to August 25th, 2020, compared with the actual scheduling, the maximum discharge flow rate of the optimized scheduling decreased by 4.34%. The number of deep holes opened by the optimized scheduling plan changed from 2 at the start of the scheduling to 4, 6, 8 in sequence, and then remained unchanged. Compared with the actual scheduling process, the optimized scheduling can effectively control the maximum discharge flow rate and ensure flood control safety. Compared with the actual scheduling process, the optimized scheduling plan has a larger number of gates put into use in the initial stage, a larger discharge flow rate, and keeps the water level at a lower level. Except for the first few scheduling periods, the number of opened gates of the optimized scheduling plan is almost unchanged, and the discharge flow rate changes slightly within a small range following the change of the water level. When the flood peak recedes, the discharge flow rate is greater than the inflow rate, and the water level gradually decreases. Through comprehensive analysis, it can be seen that the number of gate openings and closings of the optimized scheduling plan is significantly reduced compared with the actual scheduling, ensuring the safe operation of the gates during the flood season.
[0164] In summary, based on the engineering problem of large errors in the traditional gate optimization scheduling scheme, the present invention utilizes deep reinforcement learning technology to optimize the operation of the gates during the scheduling period. On the basis of the traditional reservoir flood control optimization scheduling, considering the gate constraint conditions, the objective function takes into account both the flood control objective and the number of gate openings and closings, uses the gate combination situation as the decision variable, and optimizes the gate operation scheme during the scheduling period for the known inflow process, ensuring flood control safety while making the gate operation safe and efficient.
[0165] Based on the method of the present invention, the embodiments of the present disclosure also provide a system corresponding to the above method, which includes:
[0166] Obtain the basic information of the reservoir;
[0167] Construct the objective function and constraint conditions for the multi-objective optimization scheduling of the reservoir spillway gates based on the basic information of the reservoir; the objective of the objective function is to minimize the sum of the maximum reservoir outflow and the average adjustment times of the reservoir spillway gates; the constraint conditions include water balance constraint, reservoir water level constraint, storage capacity curve constraint, maximum discharge capacity constraint, outflow variation constraint, and spillway discharge constraint;
[0168] Construct a deep reinforcement learning framework; the state variables of the deep reinforcement learning framework are the reservoir water level at the beginning of time step t, the inflow at time step t, the outflow at time step t-1, and the gate combination situation; the action variable is the gate combination situation at time step t; the reward function of the deep reinforcement learning framework is constructed based on the objective function and constraint conditions;
[0169] The construction of the reward function includes:
[0170] Determine the basic reward function according to the objective function, and the basic reward function is:
[0171] Provide an immediate reward in each time step, and incorporate the reservoir water level constraint as a penalty term into the reward function to obtain the final reward function as:
[0172]
[0173] In the formula, α' and β' are weight coefficients; c 1 and c 2 are penalty coefficients;
[0174] Use the proximal policy optimization algorithm to train the deep reinforcement learning framework to obtain a trained deep reinforcement learning model; use the trained deep reinforcement learning model for reservoir gate scheduling.
[0175] Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure.
Claims
1. A reservoir gate dispatching method, characterized in that: include: Obtain basic information about the reservoir; Based on the basic information of the reservoir, the objective function and constraint conditions of the multi-objective optimization scheduling of the reservoir spillway gate are constructed; the goal of the objective function is to minimize the sum of the maximum outflow of the reservoir and the average adjustment times of the reservoir spillway gate; the constraint conditions include water balance constraint, reservoir water level constraint, storage capacity curve constraint, maximum discharge capacity constraint, outflow amplitude constraint, and spillway discharge constraint; Construct a deep reinforcement learning framework; the state variables of the deep reinforcement learning framework are the reservoir water level at the beginning of period t, the inflow in period t, the outflow in period t-1, and the gate combination; the action variable is the gate combination in period t; The reward function of the deep reinforcement learning framework is constructed based on the objective function and the constraints; The construction of the reward function includes: Determine the basic reward function according to the objective function. The basic reward function is: Providing an immediate reward at each time step and incorporating the reservoir water level constraint as a penalty term into the reward function, the final reward function is: In the formula, α' and β' are weight coefficients; c1 and c2 are penalty coefficients; The proximal strategy optimization algorithm is used to train the deep reinforcement learning framework to obtain a trained deep reinforcement learning model; the trained deep reinforcement learning model is used to perform reservoir gate scheduling.
2. A reservoir gate dispatching method according to claim 1, characterized in that: The constraints for the reservoir spillway gate scheduling are: Water balance constraints: V t+1 =V t +(I t -Q t )·Δt Where V t is the reservoir water storage at the beginning of period t, m 3 ; I t is the reservoir inflow during period t, m 3 / s;Q t is the outflow of the reservoir during period t, m 3 / s; Δt is the duration of time period t, h; M is the number of gate types; J m is the number of gates of the mth type; q j,m,t is the discharge of the jth gate of the mth type in period t, m 3 / s; Reservoir water level constraints: WITH min ≤Z t ≤Z max In the formula, Z t is the reservoir water level at the beginning of period t, m; Z max and Z min are the upper and lower limits of the reservoir water level, m; Storage capacity curve constraints: With t =f zv (In t ) In the formula, f zv (·) is the functional relationship between reservoir water level and water storage capacity; Maximum discharge capacity constraint: Q t ≤Q t,max In the formula, Q t,max is the maximum discharge capacity of the reservoir during period t, m 3 / s; Constraints on outbound traffic fluctuation: |Q t+1 -Q t |≤ΔQ Where ΔQ is the allowable variation of the reservoir outflow in adjacent time periods, m 3 / s; Spillway discharge constraints: q j,m,t =f m,qz (O j,m,t ,Z t ) In the formula, O j,m,t is the opening of the jth gate of the mth type in time period t; f m,qz (·) is the functional relationship between the discharge of the mth type of gate and the gate opening and upstream water level.
3. A reservoir gate dispatching method according to claim 1, characterized in that: The objective function of the multi-objective optimization scheduling of the reservoir spillway gate is: Where, T is the total scheduling time, h; t is the current scheduling period, h; G is the average number of gate adjustments; α and β are weight coefficients.
4. A reservoir gate dispatching method according to claim 1, characterized in that: In the proximal policy optimization algorithm, the objective function of the policy network is: The value network expression is: Where: π θ (a t ∣s t ) is the agent’s strategy; is the estimate of the advantage function, calculated using the generalized advantage estimation method; ∈ is a hyperparameter used to limit the magnitude of the policy update; c H is a hyperparameter used to control the strength of entropy regularization; V φ (s t ) is the value network for state s t The estimated value of R t From the state s t The actual return of the start.
5. A reservoir gate dispatching method according to claim 4, characterized in that: The generalized advantage estimation method is used to calculate The expression is: Where: γ is the discount factor; λ is a hyperparameter used to control bias and variance.
6. A reservoir gate dispatching method according to claim 5, characterized in that: The method uses the proximal strategy optimization algorithm to train the deep reinforcement learning framework to obtain a trained deep reinforcement learning model; including: Initialize the parameters of the policy network and the value network; Obtain the initial environmental state s0 of the reservoir environment, that is, the inflow and reservoir water level at the initial moment, as well as the outflow and gate combination at the previous moment; The strategy network outputs the probability distribution of different gate combinations according to the initial state, and executes the action of the reservoir spillway gate according to the probability distribution of different gate combinations. t ; Execute action a t After that, the reward value r is obtained t , then the environment state becomes s t+1 ; The experience (s t ,a t ,r t ,s t+1 ) is stored in the buffer. When the amount of experience reaches the buffer capacity, it is maximized. To update the parameters of the policy network, minimize To update the parameters of the value network; Each time the parameters of the policy network and the value network are updated, the current policy π θ (a t ∣s t ) is evaluated until the cumulative reward value obtained tends to be stable, and a trained deep reinforcement learning model is obtained.
7. A reservoir spillway gate dispatching system, characterized in that: include: Data acquisition module, used to obtain basic information of reservoirs; A framework construction module is used to construct the objective function and constraint conditions of multi-objective optimization scheduling of the reservoir spillway gate based on the basic information of the reservoir; the goal of the objective function is to minimize the sum of the maximum outflow of the reservoir and the average adjustment times of the reservoir spillway gate; the constraint conditions include water balance constraint, reservoir water level constraint, storage capacity curve constraint, maximum discharge capacity constraint, outflow amplitude constraint, and spillway discharge constraint; A model building module is used to build a deep reinforcement learning framework; the state variables of the deep reinforcement learning framework are the reservoir water level at the beginning of period t, the inflow during period t, the outflow during period t-1, and the gate combination; the action variable is the gate combination during period t; The reward function of the deep reinforcement learning framework is constructed based on the objective function and the constraints; The construction of the reward function includes: Determine the basic reward function according to the objective function. The basic reward function is: Providing an immediate reward at each time step and incorporating the reservoir water level constraint as a penalty term into the reward function, the final reward function is: In the formula, α' and β' are weight coefficients; c1 and c2 are penalty coefficients; The model training module is used to train the deep reinforcement learning framework using the proximal policy optimization algorithm to obtain a trained deep reinforcement learning model; A control module is used to dispatch reservoir gates using the trained deep reinforcement learning model.
Citation Information
Patent Citations
Reinforcement learning model FQI-based reservoir flood control optimal scheduling method
CN112966445A
Reservoir gate group multi-target flood control optimization scheduling method and system
CN115719041A
Method and system for extracting scheduling rules of regulation and storage engineering in water network system
CN118014325A
Reservoir flood control scheduling method and system
CN118536766A
Water conservancy project flood discharge device
CN221167690U
Cited By
Reservoir dispatching model training method and reservoir dispatching system
CN120373808A
Plain river gatekeeper group intelligent regulation and control method based on deep reinforcement learning
CN120764853A
Reservoir scheduling optimization method and system based on aerospace big data
CN121766733A
Porous gate group scheduling method and system based on probability distribution
CN121961160A
Intelligent reconstruction method and device for gravity dam
CN122088158A