Cascade Hydropower Scheduling Method, System, Equipment and Storage Medium Based on TD3 Algorithm
By applying the scheduling method based on the TD3 algorithm in cascade hydropower stations, the problem that existing scheduling methods are difficult to achieve refined scheduling is solved, and the power generation efficiency and operational economics are improved.
Patent Information
- Application Number
- CN202411199335.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-08-29
AI Technical Summary
The existing hydropower station scheduling methods are difficult to achieve refined scheduling, which has affected the economics of the operation of cascade hydropower stations.
The cascaded hydropower scheduling method based on the TD3 algorithm is adopted, and the scheduling problem is transformed into the Markov decision-making process by constructing a long-term optimization scheduling model, and the dual delay-determination strategy gradient algorithm is used for solving it to output the long-term scheduling decision results.
It improves the power generation efficiency of cascade hydropower stations, enhances the refinement of scheduling strategies, and improves the economical operation.
Smart Images

Figure CN119026867B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hydropower station scheduling, and particularly relates to a cascade hydropower scheduling method, system, device and storage medium based on the TD3 algorithm. Background Technique
[0002] The scheduling of hydropower stations is an important guarantee for making full use of water resources and improving power generation efficiency. In recent years, with the cascade rolling development and centralized commissioning of hydropower bases in the southwestern region of China, the scale of hydropower scheduling has expanded rapidly, the grid's refined requirements for hydropower have increased day by day, and the timeliness and result availability of hydropower scheduling are facing great challenges.
[0003] The optimal scheduling of cascade hydropower stations in a basin is a typical discrete, multi-dimensional, non-linear, large-scale, multi-time and space decision-making optimization problem, and its modeling and solution are very difficult. When establishing a mathematical model with existing methods, it is difficult to consider all influencing factors, and the adaptability to changing environments is poor. In terms of optimization solution methods, they mainly include traditional mathematical programming methods and traditional machine learning methods. The mathematical programming method has problems such as time-consuming calculation, excessive memory occupation, and low calculation efficiency, while traditional machine learning algorithms have problems such as being easily limited to local optima and poor reproducibility of results.
[0004] Currently, with the rapid development of big data and artificial intelligence algorithms, reinforcement learning and even deep reinforcement learning algorithms have been initially applied in the optimal scheduling of cascade hydropower stations. However, most of the current deep reinforcement learning algorithms adopt the DQN (Deep Q Network) algorithm. Since the DQN network outputs discrete action values, it affects the refinement degree of the scheduling strategy, and further affects the economy of the operation of cascade hydropower stations. In this context, in order to further refine the intelligent scheduling decision of cascade hydropower stations and improve the economy of the operation of cascade hydropower stations, it is necessary to study new deep reinforcement learning methods.
[0005] The information disclosed in this background technical section is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art. Summary of the Invention
[0006] The present invention provides a cascade hydropower scheduling method, system, device and storage medium based on the TD3 algorithm, thereby effectively solving the problems in the background technology.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is: a cascade hydropower scheduling method based on the TD3 algorithm, including the following steps:
[0008] Construct a long-term optimal scheduling model based on the basic data and operation conditions of cascade hydropower stations;
[0009] Convert the scheduling problem in the long-term optimal scheduling model into a Markov decision process;
[0010] Use the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3) to solve the Markov decision process, and obtain the long-term scheduling decision plans for each power station in the cascade hydropower station;
[0011] Based on the actual cascade hydropower station, output the long-term scheduling decision results.
[0012] Furthermore, in the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3):
[0013] Use two value networks to evaluate the action-value;
[0014] Update the value function more frequently than the policy function;
[0015] Soft update the parameters of the target policy network and the target value network;
[0016] Clip the gradient used for updating the policy network parameters to a set range;
[0017] Use exploration noise and policy noise to smooth the policy expectation during exploration.
[0018] Furthermore, the process of using the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3) to solve the Markov decision process includes the following steps:
[0019] Step 1: Initialize the predicted value network and The network parameters are w1 and w2 respectively;
[0020] Step 2: Initialize the target value network and The network parameters are w1′ and w2′ respectively;
[0021] Step 3: Initialize the predicted policy network μ θ (s) and the target policy network μ θ′ (s), and the network parameters are θ and θ′ respectively;
[0022] Step 4: Make the parameters w1′ and w2′ of the target value network consistent with the parameters w1 and w2 of the predicted value network;
[0023] Step 5: Initialize the experience pool and set its capacity D;
[0024] Step 6: Initialize the environment module, generate the corresponding cascade hydropower station scheduling model, and determine the output of the initial state s0;
[0025] Step 7: Determine the values of the hyperparameters, including the total number of iterations M and the discount factor γ;
[0026] Step 8: Iteratively calculate the following steps from 1 to the final M;
[0027] Step 9: Initialize the environment and determine the output state s0 at the initial moment;
[0028] Step 10: Input the state s0 into the prediction policy network a = μ θ (s), and superimpose the exploration noise N to generate the corresponding action a0 = μ θ (s0) + N;
[0029] Step 11: Input the generated execution action a0 into the cascade hydropower station environment to obtain the reward r0 and the data of the next moment state s1, and store the data (s0, a0, r0, s1) in the experience pool in the form of a unit;
[0030] Step 12: Repeat Step 10 and Step 11 to generate a series of data groups and save them in the experience pool. When the time step reaches the maximum value of the environment, return to Step 9 and continue the operation;
[0031] Step 13: When the number of cyclically stored data groups reaches the set quantity, randomly sample n experience transfer samples (s, a, r, s′) from the experience pool as training data for training calculation:
[0032] Step 13.1: Calculate the perturbed target policy network action: It adds a truncated noise to the target action generated by the target policy network;
[0033] Step 13.2: Calculate the updated target: Calculate the final target value by substituting the perturbed target policy action and the next moment state into the two target value networks;
[0034] Step 14: Update the parameters w1 and w2 of the two prediction value networks according to the minimized loss function using the calculated target value:
[0035] Step 15: Update the parameters θ of the prediction policy network according to the maximized objective function using the updated prediction value network Q w1 (s, a); Among them, the parameters of the prediction value network are updated twice and then the parameters of the prediction policy network are updated once;
[0036] Step 16: Soft update the parameters of the target policy network and the target value network: Among them, τ is a hyperparameter much smaller than 1;
[0037] Step 17: Continuously update the network parameters using Steps 13 to 16, and save the network parameter values under the maximum reward until the final M iterations are reached.
[0038] The present invention further includes a cascade hydropower scheduling system based on the TD3 algorithm, which uses the method as described above and includes:
[0039] A model construction unit for constructing a long-term optimal scheduling model based on the basic data and operation conditions of cascade hydropower stations;
[0040] A conversion unit for converting the scheduling problem in the long-term optimal scheduling model into a Markov decision process;
[0041] A solving unit for solving the Markov decision process using the Twin Delayed - Deep Deterministic Policy Gradient algorithm TD3 to obtain the long-term scheduling decision schemes for each power station in the cascade hydropower stations;
[0042] An application unit for outputting the long-term scheduling decision results based on the actual cascade hydropower stations.
[0043] The present invention further includes a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method as described above is implemented.
[0044] The present invention further includes a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method as described above is implemented.
[0045] The beneficial effects of the present invention are as follows: Compared with the existing deep reinforcement learning algorithm based on DQN, the present invention can output the continuous action space value in this state according to the state quantity; compared with the existing deep reinforcement learning algorithm based on Deep Deterministic Policy Gradient (DDPG), the TD3 algorithm combines the deep deterministic policy gradient algorithm and double Q-learning, reduces the problem of network overestimation, and uses delayed policy updates and adds noise to smooth the target policy to solve the problem of error accumulation, which can effectively improve the power generation efficiency of cascade hydropower stations. Description of the Drawings
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 It is a flowchart of the method of the present invention;
[0048] Figure 2 It is a schematic structural diagram of the system of the present invention;
[0049] Figure 3 It is a flowchart of a long-term scheduling intelligent decision-making method and system for cascade hydropower stations based on the twin-delayed deterministic policy gradient algorithm;
[0050] Figure 4 It is the storage capacity - water level curve of Jinping-I Hydropower Station and Ertan Hydropower Station;
[0051] Figure 5 It is the TD3 algorithm process;
[0052] Figure 6 It is the water level of Jinping-I Hydropower Station after intelligent scheduling;
[0053] Figure 7 It is the power generation flow rate of Jinping-I Hydropower Station after intelligent scheduling;
[0054] Figure 8 It is the power generation power of Jinping-I Hydropower Station after intelligent scheduling;
[0055] Figure 9 It is the water level of Ertan Hydropower Station after intelligent scheduling;
[0056] Figure 10 It is the power generation flow rate of Ertan Hydropower Station after intelligent scheduling;
[0057] Figure 11 It is the power generation power of Ertan Hydropower Station after intelligent scheduling;
[0058] Figure 12 It is a schematic structural diagram of a computer device. Specific embodiments
[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0060] Embodiment 1:
[0061] As Figure 1 shown: A cascade hydropower scheduling method based on the TD3 algorithm includes the following steps:
[0062] Construct a long-term optimal scheduling model based on the basic data and operation conditions of cascade hydropower stations;
[0063] Convert the scheduling problem in the long-term optimal scheduling model into a Markov decision process;
[0064] Use the twin-delayed deterministic policy gradient algorithm TD3 to solve the Markov decision process, and obtain the long-term scheduling decision schemes of each power station in the cascade hydropower stations;
[0065] Output long-term scheduling decision results based on actual cascade hydropower stations.
[0066] Compared with the existing deep reinforcement learning algorithm based on DQN, it can output the continuous action space value in this state according to the state variables; compared with the existing deep reinforcement learning algorithm based on DDPG (Deep Deterministic Policy Gradient), the TD3 (Twin Delayed Deep Deterministic policygradient) algorithm combines the deep deterministic policy gradient algorithm and double Q-learning, reduces the network overestimation problem, and uses delayed policy updates and adding noise to smooth the target policy to solve the error accumulation problem, which can effectively improve the power generation efficiency of cascade hydropower stations.
[0067] In the twin-delayed deterministic policy gradient algorithm TD3:
[0068] Use two value networks to evaluate the action-value;
[0069] Update the value function more frequently than the policy function;
[0070] Soft-update the parameters of the target policy network and the target value network;
[0071] Clip the gradient used for policy network parameter update to a set range;
[0072] Use exploration noise and policy noise to smooth the policy expectation during exploration.
[0073] Among them, using the twin-delayed deterministic policy gradient algorithm TD3 to solve the Markov decision process includes the following steps:
[0074] Step 1: Initialize the predictive value networks Q w1 (s,a) and Q w2 (s,a), and the network parameters are w1 and w2 respectively;
[0075] Step 2: Initialize the target value networks Q w1′ (s,a) and Q w′2 (s,a), and the network parameters are w1′ and w2′ respectively;
[0076] Step 3: Initialize the predictive policy network μ θ (s) and the target policy network μ θ′ (s), and the network parameters are θ and θ′ respectively;
[0077] Step 4: Let the target value network parameters w1′ and w2′ be consistent with the predicted value network parameters w1 and w2;
[0078] Step 5: Initialize the experience pool and set its capacity D;
[0079] Step 6: Initialize the environment module, generate the corresponding cascade hydropower station scheduling model, and determine the output of the initial state s0;
[0080] Step 7: Determine the hyperparameter values, including the total number of iterations M and the discount factor γ;
[0081] Step 8: Perform iterative calculations for the following steps from 1 to the final M;
[0082] Step 9: Initialize the environment and determine the output of the initial state s0;
[0083] Step 10: Input the state s0 into the predicted policy network a = μ θ (s), and superimpose exploration noise N to generate the corresponding action a0 = μ θ (s0) + N;
[0084] Step 11: Input the generated execution action a0 into the cascade hydropower station environment to obtain the reward r0 and the data of the next state s1, and store the data (s0, a0, r0, s1) in the experience pool in the form of a unit;
[0085] Step 12: Repeat Step 10 and Step 11 to generate a series of data groups and save them in the experience pool. When the time step reaches the maximum value of the environment, return to Step 9 and continue the operation;
[0086] Step 13: When the number of loop - stored data groups reaches the set quantity, randomly sample n experience transition samples (s, a, r, s′) from the experience pool as training data for training calculations:
[0087] Step 13.1: Calculate the perturbed target policy network action: It adds a truncated noise to the target action generated by the target policy network;
[0088] Step 13.2: Calculate the updated target: Calculate the final target value by substituting the perturbed target policy action and the next state into the two target value networks;
[0089] Step 14: Update the two predicted value network parameters w1 and w2 using the calculated target value according to minimizing the loss function:
[0090] Step 15: According to maximizing the objective function, use the updated predicted value network Update the parameters θ of the prediction policy network: Among them, the parameters of the prediction policy network are updated once after the parameters of the prediction value network are updated twice;
[0091] Step 16: Soft-update the parameters of the target policy network and the target value network: where τ is a hyperparameter much smaller than 1;
[0092] Step 17: Continuously update the network parameters using Steps 13 to 16, and save the network parameter values under the maximum reward until the final M iterations are reached.
[0093] As Figure 2 shown, this embodiment also includes a cascade hydropower scheduling system based on the TD3 algorithm. Using the method as described above, it includes:
[0094] A model construction unit for constructing a long-term optimal scheduling model based on the basic data and operation conditions of cascade hydropower stations;
[0095] A conversion unit for converting the scheduling problem in the long-term optimal scheduling model into a Markov decision process;
[0096] A solving unit for using the Twin Delayed - Deterministic Policy Gradient algorithm TD3 to solve the Markov decision process and obtain the long-term scheduling decision schemes for each power station in the cascade hydropower station;
[0097] An application unit for outputting the long-term scheduling decision results based on the actual cascade hydropower station.
[0098] Embodiment 2:
[0099] As Figure 3 shown, the long-term scheduling intelligent decision-making method and system for cascade hydropower stations based on the Twin Delayed - Deterministic Policy Gradient algorithm include the following steps:
[0100] S1, Investigate the basic data and operation conditions of the cascade hydropower station, and construct a long-term optimal scheduling model of the cascade hydropower station;
[0101] The research object is the cascade hydropower stations in the lower reaches of the Yalong River Basin, mainly including 5 hydropower stations: Jinping-I, Jinping-II, Guandi, Ertan, and Tongzilin. Among them, Jinping-I Hydropower Station is annually scheduled, Ertan Reservoir is seasonally regulated, and Jinping-II Reservoir, Guandi Reservoir, and Tongzilin Reservoir are all daily regulated. Therefore, during the long-term scheduling process, long-term optimal scheduling is carried out for Jinping-I Hydropower Station and Ertan Hydropower Station. The optimal scheduling time is the whole year from June 1, 2020 to May 31, 2021, and the time scale is set to one day.
[0102] The basic data of Jinping-I and Ertan hydropower stations include the highest and lowest water levels, the maximum and minimum reservoir capacities, the maximum power generation flow rate, and the maximum output. The specific data are shown in Table 1. The reservoir capacity-water level curves of the two hydropower stations are shown in Figure 4 .
[0103] Table 1 Basic Data of Jinping-I and Ertan Hydropower Stations
[0104] Power station Jinping-I Hydropower Station Ertan Hydropower Station Highest water level (m) 1880 1200 Lowest water level (m) 1800 1155 <![CDATA[Maximum storage capacity (100 million m 3 )]]> 77.6 57.9 <![CDATA[Minimum reservoir capacity (100 million m 3 )]]> 28.5 24.2 <![CDATA[Maximum power generation flow rate (m 3 / s)]]> 2024 2226 Maximum output (MW) 3600 3300
[0105] To establish a long-term optimal operation model for cascade hydropower stations, the objective function is as follows:
[0106]
[0107] In the formula, E is the total power generation of the cascade hydropower stations during the operation period, in kWh; N is the number of hydropower stations in the cascade basin; T is the number of time periods in an operation period, here it is 365 days in a year; Δt is the number of hours in a time scale, in h; i is the hydropower station number; t is the time period number; P i,t is the average power generation of hydropower station i at time t, in kW; k i is the output coefficient of hydropower station i, here it takes the value of 8.5; is the average water diversion flow for power generation of hydropower station i at time t, in m 3 / s; H i,t is the average net head for power generation of hydropower station i at time t, in m.
[0108] The constraint conditions include:
[0109] (1) Water balance constraint:
[0110]
[0111] In the formula, V i,t and V i,t+1 are the initial and final reservoir capacities of hydropower station i at time t, in m 3 ; and are the average inflow and average outflow of hydropower station i at time t, in m 3 / s.
[0112] The water balance constraint between the two hydropower stations is:
[0113]
[0114] In the formula, is the outflow of the upper-level hydropower station at time t, in m 3 / s.
[0115] (2) Reservoir capacity constraint
[0116]
[0117] Wherein, and are respectively the minimum and maximum values of the reservoir storage of hydropower station i at time t, m 3 .
[0118] (3) Discharge flow constraint
[0119]
[0120] Wherein, and are respectively the minimum and maximum values of the discharge flow of hydropower station i at time t, m 3 / s.
[0121] (4) Power generation flow constraint
[0122]
[0123] Wherein, and are respectively the minimum and maximum values of the diverted flow for power generation of hydropower station i at time t, m 3 / s.
[0124] (5) Unit output constraint
[0125]
[0126] Wherein, is the maximum allowable output of hydropower station i at time t, kW.
[0127] (6) Water level limit constraint
[0128]
[0129] Wherein, and are respectively the minimum and maximum values of the net head for power generation of hydropower station i at time t, m.
[0130] (7) Reservoir storage - water level relationship constraint
[0131] H i,t = f(V i,t ) (9)
[0132] The reservoir storage - water level curve reflects the functional correspondence between the reservoir storage and the water head of hydropower station i at time t.
[0133] (8) Initial water level constraint
[0134]
[0135] In the formula, is the initial head of hydropower station i, in m.
[0136] S2 is used to transform the long-term optimal scheduling problem of cascade hydropower stations into a Markov decision process;
[0137] According to the characteristics of the long-term scheduling problem of cascade hydropower stations, the Markov decision process defines the state, action, state transition function, and reward function in reinforcement learning, and constructs a learning process. Among them,
[0138] State s t : It is used to describe the state information of cascade hydropower stations at the current moment, and provides reference guidance for the scheduling decision of the intelligent agent through the state information of cascade hydropower stations. It is defined as:
[0139]
[0140] In the formula, s t is the state quantity at time t, which mainly includes time t, and the inflow, storage capacity, and water level height of hydropower station i at this moment.
[0141] Action a t : The long-term scheduling of cascade hydropower stations realizes the long-term optimal scheduling decision of cascade hydropower stations by controlling the power generation flow of each hydropower station. Therefore, action a t is the power generation diversion flow of each hydropower station, expressed as:
[0142]
[0143] State transition function: When the upstream inflow at this moment is determined, the storage capacity and water level at the next moment are determined by controlling the power generation flow at this moment. Therefore, the operating environment of cascade hydropower stations is also determined, and the state transition probability of the Markov process is 1, that is:
[0144] p(s′|s,a) = 1 (13)
[0145] In the formula, after the intelligent agent executes action a in state s, a definite environmental state s′ at the next moment can be obtained, and this state transition process satisfies various set constraints.
[0146] Reward function r t (s t ,a t ):In the long-term optimal scheduling of cascade hydropower stations, the reward value is used to reflect the objective function value in a certain time period, and further the sum of all reward values within the scheduling period is used to represent the objective function of cascade hydropower stations within the scheduling period.
[0147]
[0148] In the formula, during the scheduling process of the cascade hydropower stations, the reward value for each time period is the sum of the power generation of each hydropower station in this time period. When the environmental state H of the cascade hydropower stations is determined i,t and the action value are obtained, the corresponding environmental reward value can be generated for the learning of optimal scheduling.
[0149] S3. Use the TD3 deep reinforcement learning algorithm to solve the Markov decision process to obtain the long-term scheduling decision-making scheme of each power station in the cascade hydropower stations;
[0150] The specific steps are shown in Figure 5 and specifically are:
[0151] Step 1: Initialize the predicted value network and The network parameters are w1 and w2 respectively;
[0152] Step 2: Initialize the target value network and The network parameters are w1′ and w2′ respectively;
[0153] Step 3: Initialize the predicted policy network μ θ (s) and the target policy network μ θ′ (s), the network parameters are θ and θ′ respectively;
[0154] Step 4: Make the target value network parameters w1′ and w2′ consistent with the predicted value network parameters w1 and w2;
[0155] Step 5: Initialize the experience pool and set its capacity D;
[0156] Step 6: Initialize the environment module, generate the corresponding cascade hydropower station scheduling model, and determine the output initial moment state s0;
[0157] Step 7: Determine the hyperparameter values, including the total number of iterations M, the discount factor γ, etc.;
[0158] Step 8: Perform iterative calculations for the following steps from 1 to the final M;
[0159] Step 9: Initialize the environment and determine the output initial moment state s0;
[0160] Step 10: Input the state s0 into the predicted policy network a = μ θ (s), and superimpose the exploration noise N to generate the corresponding action a0 = μ θ (s0) + N;
[0161] Step 11: Input the generated execution action a0 into the cascade hydropower station environment to obtain the reward r0 and the data of the next moment state s1, and store the data (s0, a0, r0, s1) in the experience pool in the form of a unit;
[0162] Step 12: Repeat Step 10 and Step 11 to generate a series of data groups and save them in the experience pool. When the time step reaches the environmental maximum value, return to Step 9 and continue the operation;
[0163] Step 13: When the number of cyclically stored data groups reaches the set quantity, randomly sample n experience transition samples (s, a, r, s′) from the experience pool as training data for training calculation:
[0164] Step 13.1: Calculate the perturbed target policy network action: It adds a truncated noise to the target action generated by the target policy network, so as to achieve the purpose of smoothing the target policy;
[0165] Step 13.2: Calculate the updated target: By substituting the perturbed target policy action and the next moment state into the two target value networks for the final target value calculation;
[0166] Step 14: According to minimizing the loss function, update the parameters w1 and w2 of the two predictive value networks using the calculated target value:
[0167] Step 15: According to maximizing the target function, use the updated predictive value network to update the parameters θ of the predictive policy network: Among them, the parameters of the predictive value network are updated twice and then the parameters of the predictive policy network are updated once, so as to achieve the purpose of delayed policy update;
[0168] Step 16: Soft update the parameters of the target policy network and the target value network: where τ is a hyperparameter much smaller than 1;
[0169] Step 17: Continuously update the network parameters using Steps 13 to 16, and save the network parameter values under the maximum reward until the final M iterations are reached.
[0170] S4. Use TD3 for long-term scheduling intelligent decision-making to conduct long-term scheduling intelligent decision-making on Jinping-I and Ertan hydropower stations on the Yalong River from June 1, 2020 to May 31, 2021 for a whole year, and the scheduling time scale is set to one day. The hyperparameter settings of TD3 are shown in Table 2.
[0171] Table 2 Main hyperparameters of the TD3 algorithm
[0172] Sequence Hyperparameter Value Specific description 1 episode num 300 Total number of time loops 2 learning rate <![CDATA[3×10 -9 > Learning rate 3 optimizer Adam Selected optimizer 4 <![CDATA max_buffer > 40000 <![CDATA[Experience Pool size > 5 batch size 64 Number of samples for batch sampling 6 targetupdaterate 0.005 Target network update rate 7 disdcount factor 0.996 Discount factor 8 exploration noise N(0,0.1) Exploration noise 9 policy noise 0.2 Target policy noise 10 noiseclip 0.5 Range of target policy noise limit 11 policy delay 2 Delayed policy update value
[0173] The dispatching decision results of Jinping-I Hydropower Station and Ertan Hydropower Station on the Yalong River for the whole year from June 1, 2020 to May 31, 2021 can be seen respectively in Figure 6 and Figure 11 .
[0174] It can be seen from Figures 6 to 8 that at the initial stage of dispatching, the generated power flow under TD3 dispatching is larger than the actual generated power flow, and the water level rises more slowly; at the middle stage of dispatching, the generated power flow under TD3 dispatching is less than the actual generated power flow, and at the same time the water level remains higher than the actual water level; at the end stage of dispatching, the TD3 dispatching method quickly increases the generated power flow to release the stored water energy, causing the water level to drop rapidly. From the perspective of generating power, at the initial stage of dispatching, the generating power under TD3 dispatching is more stably maintained at a high level compared with the actual generating power, which is closely related to the maintained high value of the generated power flow; at the middle stage of dispatching, the generating power under TD3 dispatching begins to decline, and the decline range is greater than the actual generating power, so as to ensure that the water level can be maintained at a high level; at the end stage of dispatching, the generating power under TD3 dispatching begins to increase rapidly, quickly converting the stored water energy into generated electricity.
[0175] It can be seen from Figures 9 to 11 that at the initial stage of dispatching, the generated power flow of Ertan Hydropower Station under TD3 dispatching remains at a high value, so the rising rate of the TD3 dispatching water level is slower than the actual reservoir water level; at the middle stage of dispatching, Ertan Hydropower Station has been operating at the highest water level under TD3 dispatching, and at the same time its generated power flow is smaller than the actual generated power flow; while at the end stage of dispatching, the TD3 generated power flow increases rapidly and is much higher than the actual generated power flow, and the water level drops rapidly. From the perspective of generating power, at the initial stage of dispatching, the generating power of Ertan Hydropower Station under TD3 dispatching is greater than the actual generating power, and reaches the maximum generating power after a period of time; at the middle stage of dispatching, the generating powers of Ertan Hydropower Station under TD3 dispatching and the actual generating power both begin to decline, and the generating power under TD3 dispatching is maintained at a low level, so as to ensure that the water level is at a high level; at the end stage of dispatching, the generating power of TD3 dispatching begins to rise and exceeds the actual generating power, so as to increase the generated electricity as much as possible before the end of the dispatching cycle.
[0176] Furthermore, the DQN method and the DDPG method are further used for the long-term dispatching intelligent decision of the cascade hydropower stations on the Yalong River. The generated electricity of the cascade hydropower stations under each method is shown in Table 3.
[0177] Table 3 Generated Electricity of Cascade Hydropower Stations under Actual Dispatching, DQN and DDPG Algorithms (Unit: 100 million kWh)
[0178] Power station Actual scheduling DQN scheduling DDPG scheduling TD3 scheduling Jinping-I 188.85 190.05 195.23 195.51 Ertan 158.20 192.08 193.62 194.06 Cascade total 347.05 382.13 388.85 389.57
[0179] As can be seen from Table 3, in the intelligent scheduling decision based on the TD3 deep reinforcement learning algorithm, the total annual power generation of the cascade hydropower stations is 3,895,700 MWh respectively, which is 4,252 million MWh higher than the actual historical scheduling decision annual power generation, 744 million MWh higher than the DQN scheduling decision annual power generation, and 72 million MWh higher than the DDPG scheduling decision annual power generation. Compared with the actual scheduling decision, the method of the present invention can significantly improve the power generation benefit of the cascade hydropower stations; at the same time, compared with the existing DQN and DDPG deep reinforcement learning scheduling decisions, since the present invention adopts continuous action values and combines the deep deterministic policy gradient algorithm and double Q learning, the scheduling strategy is more refined and has better power generation benefits.
[0180] Please refer to Figure 12 the structural schematic diagram of the computer device provided by the embodiment of the present application shown. A computer device 400 provided by an embodiment of the present application includes: a processor 410 and a memory 420. The memory 420 stores a computer program executable by the processor 410. When the computer program is executed by the processor 410, the above method is executed.
[0181] An embodiment of the present application also provides a storage medium 430. A computer program is stored on the storage medium 430. When the computer program is run by the processor 410, the above method is executed.
[0182] Among them, the storage medium 430 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (abbreviation: SRAM), electrically erasable programmable read-only memory (abbreviation: EEPROM), erasable programmable read-only memory (abbreviation: EPROM), programmable read-only memory (abbreviation: PROM), read-only memory (abbreviation: ROM), magnetic memory, flash memory, magnetic disk or optical disc.
[0183] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The meaning of "a plurality" is two or more, unless otherwise specifically defined.
[0184] In the present invention, unless otherwise clearly defined and limited, terms such as "install", "connect", "link", "fix", etc. shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal connection of two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0185] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not have to be directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0186] Any process or method description shown in the flowchart or described in other ways herein can be understood to represent a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present invention includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0187] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0188] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0189] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0190] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A cascade hydropower dispatching method based on TD3 algorithm, characterized in that: The steps include: Construct a long-term optimization dispatch model based on the basic data and operation status of cascade hydropower stations; Converting the scheduling problem in the long-term optimal scheduling model into a Markov decision process; The Markov decision process is solved by using the double-delay-deterministic policy gradient algorithm TD3 to obtain a long-term scheduling decision plan for each power station in the cascade hydropower station; Based on actual cascade hydropower stations, output long-term scheduling decision results; In the dual-delay-deterministic policy gradient algorithm TD3: Use two valuation networks to evaluate action-values; The frequency of updating the valuation function is greater than the policy function; Soft update of target policy network and target value network parameters; Clip the gradient used to update the policy network parameters to a set range; Use exploration noise and policy noise to smooth policy expectations during exploration; The method of solving the Markov decision process using the double-delayed-deterministic policy gradient algorithm TD3 comprises the following steps: Step 1: Initialize the prediction value network and The network parameters are w1 and w2 respectively; Step 2: Initialize the target value network and The network parameters are w1′ and w2′; Step 3: Initialize the prediction strategy network μ θ (s) and the target policy network μ θ′ (s), the network parameters are θ and θ′ respectively; Step 4: Make the target value network parameters w1′ and w2′ consistent with the predicted value network parameters w1 and w2; Step 5: Initialize the experience pool and set its capacity D; Step 6: Initialize the environment module, generate the corresponding cascade hydropower station dispatching model, and determine the output initial state s0; Step 7: Determine the hyperparameter values, including the total number of iterations M and the discount factor γ; Step 8: Perform iterative calculations of the following steps starting from 1 to the final M; Step 9: Initialize the environment and determine the output initial state s0; Step 10: Input state s0 into the prediction policy network a = μ θ (s), and superimpose the exploration noise N to generate the corresponding action a0 = μ θ (s0)+N; Step 11: Input the generated execution action a0 into the cascade hydropower station environment, obtain the reward r0 and the next state s1 data, and store the (s0, a0, r0, s1) data in the experience pool in the form of a unit; Step 12: Repeat steps 10 and 11 to generate a series of data groups and save them in the experience pool. When the time step reaches the maximum value of the environment, return to step 9 and continue the operation; Step 13: When the number of data groups stored in the loop reaches the set number, n experience transfer samples (s, a, r, s′) are randomly sampled from the experience pool as training data for training calculation: Step 13.1: Compute the perturbed target policy network action: It adds a truncation noise to the target action generated by the target policy network; Step 13.2: Calculate the updated target: The final target value is calculated by substituting the target strategy action after disturbance and the state at the next moment into two target value networks; Step 14: According to the minimized loss function, the two prediction value network parameters w1 and w2 are updated using the calculated target value: Step 15: Use the updated prediction value network according to the maximized objective function Update the prediction strategy network parameters θ: Among them, the prediction value network parameters are updated twice and then the prediction strategy network parameters are updated once; Step 16: Soft update target policy network and target value network parameters: Where τ is a hyperparameter much smaller than 1; Step 17: Continue to update the network parameters using steps 13 to 16, and save the network parameter values under the maximum reward until the final M iterations are reached.
2. A cascade hydropower dispatching system based on TD3 algorithm, characterized in that: Using the method as claimed in claim 1, comprising: Model building unit, used to build a long-term optimization dispatching model based on the basic data and operation status of cascade hydropower stations; A conversion unit, used for converting the scheduling problem in the long-term optimization scheduling model into a Markov decision process; A solution unit, used to solve the Markov decision process using a double-delayed-deterministic policy gradient algorithm TD3 to obtain a long-term dispatch decision plan for each power station in the cascade hydropower station; The application unit is used to output long-term scheduling decision results based on actual cascade hydropower stations.
3. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to claim 1 is implemented.
4. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to claim 1 is implemented.
Citation Information
Patent Citations
Cascade reservoir random optimization scheduling method based on deep Q learning
CN110930016A
Intelligent scheduling method for wind-light-water complementary system based on deep reinforcement learning
CN116345450A