Short-term optimization scheduling method for electric power system containing pumped storage and electrochemical energy storage and related device
By applying a double-layer agent with deep reinforcement learning in the new power system, the Markov decision-making model is optimized to optimize the scheduling strategies of pumped storage and electrochemical energy storage, the problems of high and low efficiency of distributed energy scheduling in the new power system are solved, and the system operation cost is reduced and the power supply reliability is improved.
Patent Information
- Application Number
- CN202510013466.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-27
AI Technical Summary
In the new power system, the dispatch of distributed energy is problematic of high cost and low efficiency, especially when dealing with complex and high-dimensional problems and dealing with uncertainty and dynamic changes.
A two-layer agent based on deep reinforcement learning combined with Markov decision-making model is adopted to construct a new power system optimization scheduling model, and the scheduling strategies of pumped storage and electrochemical energy storage are optimized through deep Q networks and deep deterministic strategy gradient algorithms.
The new power system has been optimized for scheduling within the next 24 hours to one week, minimizing system operating costs and improving power supply reliability and economic benefits.
Smart Images

Figure BDA0005229126000000021 
Figure BDA0005229126000000022 
Figure BDA0005229126000000034
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of power grid operation and maintenance. Technical Background
[0002] The new power system includes various distributed energy sources, which are divided into dispatchable energy sources (such as pumped-storage power stations and electrochemical energy storage) and non-dispatchable energy sources (such as photovoltaic power generation and wind power generation). How to reasonably dispatch these energy sources to minimize the system operation cost has become an issue of widespread concern. Since new energy sources such as photovoltaic and wind energy are greatly affected by the environment, and at the same time, the change of user load is also likely to cause power balance problems, these uncertainties have brought great challenges to the optimal dispatch of the new power system.
[0003] At present, researchers have carried out extensive research in aspects such as energy prediction and management, energy storage systems and battery technologies. In terms of energy prediction, the focus is on improving the prediction accuracy of photovoltaic and wind power generation, and using advanced prediction models and algorithms to achieve accurate prediction; in terms of energy storage technology, researchers focus on the application of energy storage systems to improve the stability and economy of the system. However, the existing optimization methods have many limitations. For example, model-based methods are limited by model accuracy when dealing with complex high-dimensional problems, genetic algorithms are slow and require a large amount of computing resources, and particle swarm algorithms are prone to falling into local optima and have a slow convergence speed. Summary of the Invention
[0004] The present invention aims to provide a short-term optimal dispatch method for a power system with pumped-storage and electrochemical energy storage based on deep reinforcement learning, which is applied to a new power system with pumped-storage and electrochemical energy storage, and provides a feasible method for the optimal dispatch of the power system within the next 24 hours to one week. This method maximizes the reduction of the system operation cost and improves the economic benefits by optimizing the dispatch strategies of various dispatchable energy sources.
[0005] The present invention builds a new power system model and uses the deep reinforcement learning method to find the optimal dispatch strategy to reduce the energy cost and improve the power supply reliability.
[0006] The technical solution adopted by the present invention is as follows: A short-term optimal dispatch method for a power system with pumped-storage and electrochemical energy storage, including:
[0007] Construct a new power system optimal dispatch model considering the coordination among power sources (WE, PV, PS), local power grids, large power grids, loads, and energy storage batteries, establish an objective function for minimizing the system operation cost, and determine the dispatchable energy source constraint conditions;
[0008] Convert the new power system optimal dispatch model into a two-layer agent joint Markov decision model, and train and optimize the decision model;
[0009] Apply the trained decision-making model to the actual power system, output the optimal decision according to the current state, and achieve the optimal scheduling of the new power system.
[0010] The objective function is as follows:
[0011]
[0012] In the above formula, C ic,t is the interaction cost between the new power system and the large power grid, C bat,t is the depreciation cost of the electrochemical energy storage, C oe,t is the equipment operation and maintenance cost, C ps,t is the pumping cost; among them,
[0013]
[0014] In the formula, p buy,t is the electricity purchase price of the new power system from the large power grid at time t, P buy,t is the electricity purchase power of the new power system from the large power grid at time t, p sell,t is the electricity selling price of the new power system to the large power grid at time t, P sell,t is the electricity selling power of the new power system to the large power grid at time t, ρ bat is the charge-discharge depreciation cost coefficient of the electrochemical energy storage, P bat,t is the active power of the electrochemical energy storage at time t, P bat,t >0 indicates that the electrochemical energy storage is discharging, P bat,t <0 indicates that the electrochemical energy storage is charging, k oe_i is the coefficient of the operation cost of various distributed energy sources in the new power system at time t, P i,t is the active power of each energy source in the new power system at time t, P ps,t is the active power of the pumped-storage power station at time t, C s is the pumping cost coefficient of the pumped-storage power station, ε is a judgment parameter, P ps,t When >0, the pumped-storage power station generates electricity, and the value is 1, P ps,t When <0, the pumped-storage power station stores energy, and the value is 0.
[0015] The adjustable energy constraint conditions include:
[0016] Power balance constraint: P ps,t +P bat,t +P we,t +P pv,t +P buy,t =P ld,t +P sell,t
[0017] In the formula, P ps,t is the active power of the pumped-storage power station at time t, Pbat,t The active power of the electrochemical energy storage device at time t, P we,t Is the active power output by the wind power generation system at time t, P pv,t Is the active power output by the photovoltaic system at time t, P buy,t Is the power purchase power from the local area network to the wide area network at time t, P sell,t Is the power selling power from the local area network to the wide area network at time t, P ld,t Is the load power at time t;
[0018] Pumped storage power station operation constraints:
[0019] In the formula, And Represent the upper and lower limits of the rated output of the pumped storage power station;
[0020] Pumped storage power station reservoir capacity constraint: E min ≤E t ≤E max
[0021] In the formula, E min And E max Represent the lower and upper limits of the reservoir capacity of the pumped storage power station; E t The calculation formula of is as follows:
[0022]
[0023] In the formula, E t+1 Is the capacity of the reservoir in the next time period, E t Is the capacity of the reservoir in time period t, ΔE t Is the change in reservoir capacity in time period t, P ps,t Is the active power output of the pumped storage power station in time period t, η g Is the power generation efficiency of the pumped storage power station, η f Is the pumping efficiency of the pumped storage power station, P ps,t >0 indicates that the pumped storage power station is generating electricity, P ps,t <0 indicates that the pumped storage power station is pumping energy storage;
[0024] Electrochemical energy storage device operation constraints:
[0025] In the formula, Are the lower and upper limits of the active power output of the electrochemical energy storage device;
[0026] In order to avoid damage to the electrochemical energy storage caused by deep charge and discharge, the state of charge of the electrochemical energy storage needs to satisfy: SOC min ≤SOC t ≤SOC max ;
[0027] Wherein, SOC min and SOC max are respectively the lower limit and the upper limit allowed by SOC; specifically, the calculation formula of SOC t is as follows:
[0028] SOC bat,t+1 = SOC bat,t - η bat · P bat,t · Δt / E bat
[0029]
[0030] Wherein, SOC bat,t+1 is the state of charge of the electrochemical energy storage within the next time period, SOC bat,t is the state of charge of the electrochemical energy storage within the current time period, E bat is the capacity of the electrochemical energy storage, η c is the charging efficiency of the electrochemical energy storage, η d is the discharging efficiency of the electrochemical energy storage, P bat,t > 0 indicates that the electrochemical energy storage discharges, and P bat,t < 0 indicates that the electrochemical energy storage charges;
[0031] Interaction constraint of external energy:
[0032] Wherein, P gso,export_max is the maximum power supplied by the new power system to the large power grid, and P gso,import_max is the maximum power supplied by the large power grid to the new power system.
[0033] The specific double-layer agent joint Markov decision model is as follows: The upper-layer agent outputs the charge and discharge decisions of the electrochemical energy storage and the pumped-storage power station, and returns the decision actions to the environment and transmits them to the lower-layer agent. This charge and discharge decision is a discrete decision, which is solved using the deep Q-network (DQN) algorithm; the lower-layer agent receives the charge and discharge decisions transmitted by the upper-layer agent, solves the optimal action, that is, the dispatching strategy of the dispatchable energy. This decision is a continuous variable decision, which is solved using the deep deterministic policy gradient (DDPG) algorithm; both the upper and lower layers of agents can perceive the current state of the environment, and the action decision spaces of the upper-layer agent and the lower-layer agent jointly act on the environment, thereby realizing the mixed decision of continuous actions and discrete actions.
[0034] Converting the new power system optimal dispatching model into a double-layer agent joint Markov decision model includes upper-layer agent design, lower-layer agent design, environment construction, and agent construction;
[0035] The upper-layer agent design includes:
[0036] State - space design: In this new - type power - system optimal scheduling model, the environmental state refers to the wind - power generation power \(P\) we,t , the photovoltaic - power generation power \(P\) pv,t , the load power \(P\) ld,t , the selling - electricity price \(p\) sell,t , the purchasing - electricity price \(p\) buy,t , the reservoir capacity \(E\) t , the state of charge (SOC) of the electrochemical energy storage bat,t and the scheduling period \(t\) in which it is located; thus, for this optimal - scheduling problem, its state space can be expressed as:
[0037] \(s\) t =\([P\) we,t , \(P\) pv,t , \(P\) ld,t , \(p\) sell,t , \(p\) buy,t , \(E\) t , \(SOC\) bat,t , \(t]\)
[0038] Action - space design: To simplify the model, the charge - discharge decisions of the electrochemical energy storage and the pumped - storage power station are defined as 3 states:
[0039] \(a\) upper_t =\(\{-1,0, + 1\}\)
[0040] When the action value is - 1, the electrochemical energy storage and the pumped - storage power station are in the discharge / generation state; when the action value is 0, no operation is performed on them; when the action value is + 1, the electrochemical energy storage and the pumped - storage power station are in the charge / energy - storage state;
[0041] Reward - function design: The objective function is transformed into a reward function, and the reward is set to be negative;
[0042] \(r\) upper_t =\(-\mu(C\) ic,t +C\) bat,t +C\) oe,t +C\) ps,t )
[0043]
[0044] \(C\) ic,t is the interaction cost between the local power grid and the external power grid, \(C\) bat,t is the depreciation cost of the electrochemical energy storage, \(C\) oe,t is the equipment operation and maintenance cost, \(C\) ps,tLet \(C_w\) be the pumping cost, and \(\mu\) is a flag for the charge and discharge of the electrochemical energy storage and the pumped-storage power station. When the output power of the non-dispatchable energy can meet the load demand, \(\mu = 1\), and a positive reward is given to the agent; when the output power of the non-dispatchable energy cannot meet the load demand, \(\mu=-1\), and a negative reward is given to the agent.
[0045] The design of the lower-layer agent includes:
[0046] State space design: The state space of the lower-layer agent is the same as that of the upper-layer agent. In this new power system optimal scheduling model, the environmental state refers to the wind power generation power \(P_w\) we,t , the photovoltaic power generation power \(P_p\) pv,t , the load power \(P_l\) ld,t , the selling electricity price \(p_s\) sell,t , the purchasing electricity price \(p_b\) buy,t , the reservoir capacity \(E\) t , the state of charge (SOC) of the electrochemical energy storage bat,t and the scheduling period \(t\) in which it is located. Therefore, for this optimal scheduling problem, its state space can be expressed as:
[0047] \(s\) t =\([P_w\) we,t , \(P_p\) pv,t , \(P_l\) ld,t , \(p_s\) sell,t , \(p_b\) buy,t , \(E\) t , \(SOC\) bat,t , \(t]\)
[0048] Action space design: In this new power system optimal scheduling model, the goal is to determine the energy storage or power generation power \(P_p\) of the pumped-storage power station ps,t , the charge and discharge power \(P_e\) of the electrochemical energy storage bat,t , the power purchase power \(P_b\) of the new power system buy,t and the power selling power \(P_s\) of the new power system to the large power grid sell,t ; for this optimal scheduling problem, its action space can be expressed as:
[0049] \(a\) lower_t =\(\{P_p\) ps,t , \(P_e\) bat,t , \(P_b\) buy,t , \(P_s\) sell,t}\)
[0050] Reward function design: After the lower-layer agent executes the action \(a\) lower_t , it obtains an immediate reward \(r\) lower_t , and the reward function is set as:
[0051] \(r\) lower_t =-\(\eta(C_w\) ic,t +\(C_e\) bat,t +\(C_b\)oe,t +C ps,t ),
[0052] wherein, C ic,t is the interaction cost between the new power system and the large power grid, C bat,t is the depreciation cost of the electrochemical energy storage, C oe,t is the equipment operation and maintenance cost, C ps,t is the pumping cost, and η is the normalization coefficient;
[0053] The environment structure includes:
[0054] Define the state space: Collect the current wind power generation power, photovoltaic power generation power, and load power and standardize them; Collect the current selling electricity price and purchasing electricity price and standardize them; Collect the current state of charge of the electrochemical energy storage and the reservoir capacity of the pumped-storage power station and normalize them; Obtain the current time period and normalize it; Concatenate the above normalized values into a state vector, and this state vector constitutes the state space;
[0055] Define the reward function: Include defining the reward functions of the upper-level agent and the lower-level agent;
[0056] Define the dispatchable energy constraint conditions: Include power balance constraints, pumped-storage power station operation constraints, electrochemical energy storage operation constraints, and interaction constraints between the new power system and the large power grid;
[0057] The agent structure includes:
[0058] Define the internal operations of the agent: Include defining the interaction process between the agent and the environment, the process of storing and extracting experience data from the memory, etc.;
[0059] Define the modeling and optimization process of the neural network: Include constructing the neural network, defining the optimizer, setting the loss function, etc.
[0060] The steps to train and optimize the model are as follows:
[0061] Training of the upper-level agent:
[0062] Step 1: Initialize the environment and the upper-level agent; Initialize the environment, including initializing the state space and the reward function; Initialize the upper-level agent, including initializing the main value Q network and the target Q network, and copying the main value Q network parameters to the target Q network. The network parameters of the main value Q network and the target Q network are μ and μ' respectively; Both networks are convolutional networks, and their network structures are the same. The input layer contains 7 neurons, the output layer contains 3 neurons, three convolutional layers and two fully connected layers are set, and the hidden layer uses the rectified linear unit ReLU activation function;
[0063] Step 2: Interaction between the upper-layer agent and the environment; the upper-layer agent obtains the current state s from the environment t , and based on the ε-greedy strategy, selects an action a using the main value Q-network upper_t and returns this action information to the environment. The environment receives the action information of the upper-layer agent and calculates the reward r upper_t and generates the next state s t+1 which is fed back to the upper-layer agent. The upper-layer agent stores the quadruple data [s t , a upper_t , r upper_t , s t+1 in the experience replay pool;
[0064]
[0065] Step 3: Update of network parameters in the upper-layer agent; extract the quadruple data, input [s t , a upper_t into the main value Q-network to obtain Q(s t , a upper_t ; μ), and input [s t+1 into the target Q-network to obtain Q(s t+1 , a upper_t+1 ; μ') and find the maximum Q value among them Input it into the optimizer, calculate the target value y and use this target value y as the label and the loss function:
[0066]
[0067] Use the gradient descent method to update the parameters of the main value Q-network. The expression for updating the parameters of the main value Q-network is as follows, where 0 ≤ α ≤ 1:
[0068]
[0069] After the parameters of the main value Q-network are updated N times, copy the parameters of the main value Q-network to the target Q-network;
[0070] Repeat step 3, the update of network parameters in the upper-layer agent, continuously extract quadruple data, calculate the target value, and update the network parameters until the change in the parameters of the main value Q-network is less than the set threshold, and the training terminates; the upper-layer agent outputs charge and discharge decisions to the environment and the lower-layer agent;
[0071] Training of the lower-layer agent:
[0072] Step 4: Initialize the environment and the lower-level agent; Initialize the environment, including initializing the state space, reward function, etc.; Initialize the main policy network and the main value network, and copy the parameters of the main policy network and the main value network to the target policy network and the target value network. The parameters of the main policy network and the target policy network are w and w' respectively, and the parameters of the main value network and the target value network are θ and θ' respectively. The input layer of the policy network contains 7 neurons, and the output layer contains 6 neurons. It consists of two fully connected layers, where the hidden layer uses the ReLU activation function and the output layer uses the tanh activation function. The input layer of the value network contains 13 neurons, and the output layer contains 1 neuron. It consists of two fully connected layers, where the hidden layer uses the ReLU activation function;
[0073] Step 5: The lower-level agent interacts with the environment; The lower-level agent obtains the current state s from the environment t , and based on the ε-greedy strategy, selects an action a using the main policy network lower_t and returns this action information to the environment. The environment receives the action information of the lower-level agent, calculates the reward r lower_t and generates the next state s t+1 and feeds it back to the lower-level agent. The lower-level agent stores the quadruple [s t , a lower_t , r lower_t , s t+1 into the experience replay pool;
[0074]
[0075] Step 6: Update the network parameters in the lower-level agent; Extract the quadruple data, input [s t , a lower_t into the main value network to obtain Q(s t , a lower_t ; θ), input [s t+1 into the target policy network to obtain a lower_t+1 , and input [s t+1 , a lower_t+1 into the target value network to obtain Q(s t+1 , a lower_t+1 ; θ'), calculate the target value y t , and use this target value y t as the label:
[0076] y t = r lower_t + λQ(s t+1 , a lower_t+1 ; θ')
[0077] Minimizing the loss function makes the Q-value output by the main value network close to the target Q-value, and updates the parameters of the main value network:
[0078]
[0079] The main policy network uses the Q(s t ,a lower_t ; θ) generated by the main value network to update the gradient of the action a lower_t and the gradient of the action a lower_t with respect to the network parameter w; the policy gradient during the update of the main policy network is as follows, and the parameters w of the main policy network are updated using gradient ascent, where 0 ≤ β ≤ 1:
[0080]
[0081] After the parameters of the main policy network and the main value network are updated N times, the main network parameters are copied to the target network; the updates of the target policy network and the target value network adopt soft updates. A learning rate τ is introduced, 0 < τ < 1. The old target network parameters and the new corresponding network parameters are weighted averaged and then assigned to the target network, that is, the network parameters are updated slowly:
[0082]
[0083] Repeat step 6, continuously extract quadruple data, calculate the target, and update the network parameters until the parameter changes of the main policy network and the main value network are less than the set threshold, and the training terminates; the lower-layer agent outputs the scheduling policy to the environment;
[0084] Step 7: Repeat iteration; repeat steps 2, 3, 5, and 6. The upper and lower layer agents continuously interact with the environment, generate new experience data, and update the network parameters in the agents until the maximum number of training times is reached, and the training terminates.
[0085] Applying the trained model to the actual power system, outputting the optimal decision according to the current state, and realizing the optimal scheduling of the new power system, specifically: after the double-layer agent joint Markov decision model is trained, the agent outputs the corresponding action according to the state information input at the current time period, that is, the decision variable to be solved, to realize the direct mapping of the system from the current state to the scheduling decision.
[0086] The present invention also provides a short-term optimal scheduling device for a power system including pumped storage and electrochemical energy storage, including:
[0087] A construction module for constructing an optimal scheduling model for a new power system considering the coordinated cooperation among power sources (WE, PV, PS), the local network, the large power grid, the load, and the energy storage battery, establishing an objective function for minimizing the system operation cost, and determining the schedulable energy constraint conditions;
[0088] A training module, which is used to convert the new power system optimal scheduling model into a two-layer agent joint Markov decision model and train and optimize the decision model;
[0089] A scheduling module, which is used to apply the trained decision model to the actual power system, output the optimal decision according to the current state, and realize the optimal scheduling of the new power system.
[0090] The objective function is:
[0091]
[0092] In the above formula, C ic,t is the interaction cost between the new power system and the large power grid, C bat,t is the depreciation cost of the electrochemical energy storage, C oe,t is the equipment operation and maintenance cost, C ps,t is the pumping cost; where
[0093]
[0094] In the formula, p buy,t is the electricity purchase price of the new power system from the large power grid at time t, P buy,t is the electricity purchase power of the new power system from the large power grid at time t, p sell,t is the electricity selling price of the new power system to the large power grid at time t, P sell,t is the electricity selling power of the new power system to the large power grid at time t, ρ bat is the charge and discharge depreciation cost coefficient of the electrochemical energy storage, P bat,t is the active power of the electrochemical energy storage at time t, P bat,t >0 indicates that the electrochemical energy storage discharges, P bat,t <0 indicates that the electrochemical energy storage charges, k oe_i is the coefficient of the operating cost of various distributed energy sources in the new power system at time t, P i,t is the active power of each energy source in the new power system at time t, P ps,t is the active power of the pumped-storage power station at time t, C s is the pumping cost coefficient of the pumped-storage power station, ε is a judgment parameter, P ps,t >0 means the pumped-storage power station generates electricity, and the value is 1, P ps,t <0 means the pumped-storage power station stores energy, and the value is 0.
[0095] The adjustable energy constraint conditions include:
[0096] Power balance constraint: P ps,t +P bat,t +P we,t +P pv,t +Pbuy,t = P ld,t + P sell,t
[0097] Wherein, P ps,t is the active power of the pumped - storage power station at time t, P bat,t is the active power of the electrochemical energy storage device at time t, P we,t is the active power output by the wind power generation system at time t, P pv,t is the active power output by the photovoltaic system at time t, P buy,t is the power purchase from the local network to the wide - area network at time t, P sell,t is the power sale from the local network to the wide - area network at time t, P ld,t is the load power at time t;
[0098] Operating constraints of the pumped - storage power station:
[0099] Wherein, and represent the upper and lower limits of the rated output of the pumped - storage power station;
[0100] Storage capacity constraint of the pumped - storage power station: E min ≤ E t ≤ E max
[0101] Wherein, E min and E max represent the lower and upper limits of the reservoir storage capacity of the pumped - storage power station; The calculation formula of E t is as follows:
[0102]
[0103] Wherein, E t+1 is the capacity of the reservoir in the next time period, E t is the capacity of the reservoir in time period t, ΔE t is the change in the storage capacity in time period t, P ps,t is the active power output of the pumped - storage power station in time period t, η g is the power generation efficiency of the pumped - storage power station, η f is the pumping efficiency of the pumped - storage power station, P ps,t > 0 indicates that the pumped - storage power station is generating electricity, P ps,t < 0 indicates that the pumped - storage power station is pumping and storing energy;
[0104] Operating constraints of the electrochemical energy storage device:
[0105] Wherein, are the lower and upper limits of the active power output of the electrochemical energy storage device;
[0106] To avoid damage to electrochemical energy storage caused by deep charge and discharge, the state of charge of electrochemical energy storage needs to satisfy: SOC min ≤SOC t ≤SOC max ;
[0107] where SOC min , SOC max are the lower and upper limits allowed for SOC respectively; specifically, the calculation formula of SOC t is as follows:
[0108] SOC bat,t+1 = SOC bat,t -η bat ·P bat,t ·Δt / E bat
[0109]
[0110] where SOC bat,t+1 is the state of charge of the electrochemical energy storage in the next time period, SOC bat,t is the state of charge of the electrochemical energy storage in the current time period, E bat is the capacity of the electrochemical energy storage, η c is the charging efficiency of the electrochemical energy storage, η d is the discharging efficiency of the electrochemical energy storage, P bat,t >0 indicates that the electrochemical energy storage discharges, P bat,t <0 indicates that the electrochemical energy storage charges;
[0111] Interaction constraints of external energy sources:
[0112] where P gso,export_max is the maximum power supplied by the new power system to the large power grid, and P gso,import_max is the maximum power supplied by the large power grid to the new power system.
[0113] The specific double-layer intelligent agent combined Markov decision model is as follows: The upper-layer intelligent agent outputs the charge and discharge decisions of the electrochemical energy storage and the pumped-storage power station, and returns the decision actions to the environment and transmits them to the lower-layer intelligent agent. This charge and discharge decision is a discrete decision, which is solved using the deep Q-network algorithm; the lower-layer intelligent agent receives the charge and discharge decisions transmitted by the upper-layer intelligent agent, solves the optimal action, that is, the dispatching strategy of the dispatchable energy. This decision is a continuous variable decision, which is solved using the deep deterministic policy gradient algorithm; both the upper and lower layers of intelligent agents can perceive the current state of the environment. The action decision space of the upper-layer intelligent agent and the action decision space of the lower-layer intelligent agent jointly act on the environment, thereby realizing the mixed decision of continuous actions and discrete actions.
[0114] The training module is specifically used for training the upper-layer agent and the lower-layer agent, and the steps are as follows:
[0115] Step 1: Initialize the environment and the upper-layer agent; Initialize the environment, including initializing the state space and the reward function; Initialize the upper-layer agent, including initializing the main value Q-network and the target Q-network, and copy the parameters of the main value Q-network to the target Q-network. The network parameters of the main value Q-network and the target Q-network are μ and μ' respectively; Both networks are convolutional networks, and they have the same network structure. The input layer contains 7 neurons, the output layer contains 3 neurons, three convolutional layers and two fully connected layers are set, and the ReLU activation function is used in the hidden layer;
[0116] Step 2: The upper-layer agent interacts with the environment; The upper-layer agent obtains the current state s from the environment t , and based on the ε-greedy greedy strategy, selects an action a based on the main value Q-network upper_t and returns this action information to the environment. The environment receives the action information of the upper-layer agent, calculates the reward r upper_t and generates the next state s t+1 and feeds it back to the upper-layer agent. The upper-layer agent stores the quadruple data [s t , a upper_t , r upper_t , s t+1 into the experience replay pool;
[0117]
[0118] Step 3: Update the network parameters in the upper-layer agent; Extract the quadruple data, input [s t , a upper_t into the main value Q-network to get Q(s t , a upper_t ; μ), [s t+1 into the target Q-network to get Q(s t+1 , a upper_t+1 ; μ') and find the maximum Q value among them Input it into the optimizer, calculate the target value y and use this target value y as the label and the loss function:
[0119]
[0120] Use the gradient descent method to update the parameters of the main value Q-network. The expression for updating the parameters of the main value Q-network is as follows, where 0 ≤ α ≤ 1:
[0121]
[0122] After the main value Q-network parameters are updated N times, copy the main value Q-network parameters to the target Q-network;
[0123] Repeat the network parameter update in the upper-layer intelligent agent in step 3, continuously extract quadruple data, calculate the target value, and update the network parameters until the parameter change of the main value Q-network is less than the set threshold, and the training terminates; the upper-layer intelligent agent outputs charge and discharge decisions to the environment and the lower-layer intelligent agent;
[0124] Step 4: Initialize the environment and the lower-layer intelligent agent; Initialize the environment, including initializing the state space, reward function, etc.; Initialize the main policy network and the main value network, and copy the parameters of the main policy network and the main value network to the target policy network and the target value network. The parameters of the main policy network and the target policy network are w and w' respectively, and the parameters of the main value network and the target value network are θ and θ' respectively; The input layer of the policy network contains 7 neurons, the output layer contains 6 neurons, and it consists of two fully connected layers. The ReLU activation function is used in the hidden layer, and the tanh activation function is used in the output layer; The input layer of the value network contains 13 neurons, the output layer contains 1 neuron, and it consists of two fully connected layers. The ReLU activation function is used in the hidden layer;
[0125] Step 5: The lower-layer intelligent agent interacts with the environment; The lower-layer intelligent agent obtains the current state s from the environment t , and based on the ε-greedy greedy strategy, selects an action a based on the main policy network lower_t and returns this action information to the environment. The environment receives the action information of the lower-layer intelligent agent, calculates the reward r lower_t and generates the next state s t+1 and feeds it back to the lower-layer intelligent agent. The lower-layer intelligent agent stores the quadruple [s t , a lower_t , r lower_t , s t+1 into the experience replay pool;
[0126]
[0127] Step 6: Update the network parameters in the lower-layer intelligent agent; Extract the quadruple data, input [s t , a lower_t into the main value network to get Q(s t , a lower_t ; θ), input [s t+1 into the target policy network to get a lower_t+1 , and input [s t+1 , a lower_t+1 into the target value network to get Q(s t+1 , a lower_t+1 ; θ'), and calculate the target value yt , take this target value y t as the label:
[0128]
[0129] Minimize the loss function to make the Q-value output by the main value network close to the target Q-value, and update the parameters of the main value network:
[0130]
[0131] The main policy network uses the Q(s t , a lower_t ; θ) to update the gradient of the action a lower_t and the gradient of the action a lower_t with respect to the network parameter w; the policy gradient of the main policy network during update is as follows. Use gradient ascent to update the main policy network parameter w, where 0 ≤ β ≤ 1:
[0132]
[0133] After the main policy network and the main value network parameters are updated N times, copy the main network parameters to the target network; the update of the target policy network and the target value network adopts soft update. Introduce a learning rate τ, 0 < τ < 1, and perform weighted averaging on the old target network parameters and the new corresponding network parameters, and then assign the result to the target network, that is, the network parameters are updated slowly:
[0134]
[0135] Repeat step 6, continuously extract quadruple data, calculate the target, and update the network parameters until the parameter changes of the main policy network and the main value network are less than the set threshold, and the training terminates; the lower-level agent outputs the scheduling policy to the environment;
[0136] Step 7: Repeat iteration; repeat steps 2, 3, 5, and 6. The upper and lower-level agents continuously interact with the environment, generate new experience data, and update the network parameters in the agents until the maximum number of training times is reached, and the training terminates.
[0137] The scheduling module is specifically used for: after the joint Markov decision model training of the double-layer agent is completed, the agent outputs the corresponding action, that is, the decision variable to be solved, according to the state information input at the current time period, realizing the direct mapping from the current state of the system to the scheduling decision.
[0138] On the other hand, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the short-term optimal scheduling method for a power system with pumped storage and electrochemical energy storage described in the first aspect.
[0139] On the other hand, the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the short-term optimal scheduling method of the power system with pumped-storage and electrochemical energy storage described in the first aspect.
[0140] Compared with the prior art, the beneficial effects of the present invention are as follows: Compared with traditional optimal scheduling technologies, the deep reinforcement learning adopted by the present invention has played an obvious advantage in the optimal scheduling of power systems. Especially in dealing with high-dimensional complex problems, coping with uncertainties and dynamic changes, achieving multi-objective optimization, and maximizing long-term benefits, the powerful learning ability of deep reinforcement learning can dynamically adapt to the complex environment of microgrids, provide more flexible and intelligent scheduling strategies, and promote the intelligent development of new power systems; On the basis of Markov decision-making, the present invention conducts hierarchical optimization, reduces complexity and improves computational efficiency by decomposing complex problems and processing sub-problems in parallel, while enhancing the adaptability and robustness of the system, enabling it to better cope with the volatility of renewable energy and demand changes. BRIEF DESCRIPTION OF THE DRAWINGS
[0141] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0142] Figure 1 It is a simplified structure diagram of a new power system.
[0143] Figure 2 It is a schematic structural diagram of a two-layer agent joint Markov decision model.
[0144] Figure 3 It is a flow chart for training and optimizing a two-layer agent joint Markov decision model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0145] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0146] The first aspect of the present invention provides a short-term optimal scheduling method for a power system including pumped-storage energy and electrochemical energy storage, comprising the following steps: First, an optimal scheduling model of the new power system is created. Constraint conditions for dispatchable energy, power balance constraint conditions, and constraint conditions for interaction with the large power grid are established. The objective function is designed to minimize the operating costs of the new power system, which include the interaction costs between the local new power system and the large power grid, the depreciation costs of the electrochemical energy storage, the equipment operation and maintenance costs, and the pumping costs.
[0147] The objective function of the local area network model is set as follows:
[0148]
[0149] In the above formula, C ic,t is the interaction cost between the new power system and the large power grid, C bat,t is the depreciation cost of the electrochemical energy storage, C oe,t is the equipment operation and maintenance cost, C ps,t is the pumping cost.
[0150] The constraint condition formula of the new power system model is as follows:
[0151] P ps,t + P bat,t + P we,t + P pv,t + P buy,t = P ld,t + P sell,t (Formula 1)
[0152]
[0153]
[0154] In the above formula, Formula (1) is the power balance constraint, Formula (2) is the dispatchable energy constraint, and Formula (3) is the interaction constraint with the large power grid.
[0155] Then, the above new power system model is converted into a Markov decision model. The problem of minimizing the operating cost of the new power system is converted into the problem of maximizing the reward in the Markov decision process. A two-layer intelligent agent decision framework is designed for the optimal scheduling of the new power system, where the upper-layer intelligent agent decides the charge and discharge processes of the pumped-storage power station and the electrochemical energy storage, and the lower-layer intelligent agent generates the scheduling strategies of each dispatchable energy at the current moment.
[0156] Refer to Figure 1, consider an independent new power system and establish a mathematical model for it. The simplified model of the system includes a pumped-storage power station (PS), an electrochemical energy storage (BAT), a load (LD), a wind power generation (WE), a photovoltaic system (PV), and a large power grid (GSO). This new power system model takes into account the coordinated cooperation among power sources (WE, PV, PS), the local grid and the large power grid, the load (LD), and the energy storage battery (source, grid, load, storage). WE and PV are non-dispatchable energy sources, and PS is a dispatchable energy source. WE, PV, and PS jointly provide energy for the load. When the power generation of these energy sources exceeds the load demand, the pumped-storage power station and the electrochemical energy storage device charge and store energy. When the power generation cannot meet the load demand, the pumped-storage power station and the electrochemical energy storage discharge.
[0157] The purpose of optimizing the dispatching of the new power system is to minimize the operating cost of the system. The operating cost specifically includes the interaction cost between the new power system and the large power grid, the depreciation cost of the electrochemical energy storage, the equipment operation and maintenance cost, and the pumping cost. Accordingly, the objective function of the new power system is established:
[0158]
[0159] In the formula, C ic,t is the interaction cost between the new power system and the large power grid, C bat,t is the depreciation cost of the electrochemical energy storage, C oe,t is the equipment operation and maintenance cost, C ps,t is the pumping cost of the pumped-storage power station. Among them,
[0160]
[0161] In the formula, p buy,t is the electricity purchase price of the new power system from the large power grid at time t, P buy,t is the electricity purchase power of the new power system from the large power grid at time t, p sell,t is the electricity selling price of the new power system to the large power grid at time t, P sell,t is the electricity selling power of the new power system to the large power grid at time t, ρ bat is the charge and discharge depreciation cost coefficient of the electrochemical energy storage, P bat,t is the active power of the electrochemical energy storage at time t, P bat,t >0 indicates that the electrochemical energy storage discharges, P bat,t <0 indicates that the electrochemical energy storage charges, k oe_i is the coefficient of the operating cost of various distributed energy sources in the new power system at time t, P i,t is the active power of each energy source in the new power system at time t, P ps,t is the active power of the pumped-storage power station at time t, C sis the pumping cost coefficient of the pumped - storage power station, ε is the judgment parameter, P ps,t When P ps,t > 0, the pumped - storage power station generates electricity and takes the value of 1. When P
[0162] The constraint conditions formula of the new power system is as follows:
[0163] (1) Power balance constraint: P ps,t + P bat,t + P we,t + P pv,t + P buy,t = P ld,t + P sell,t
[0164] In the formula, P ps,t is the active power of the pumped - storage power station at time t, P bat,t is the active power of the electrochemical energy storage device at time t, P we,t is the active power output by the wind power generation system at time t, P pv,t is the active power output by the photovoltaic system at time t, P buy,t is the power purchase from the local network to the wide - area network at time t, P sell,t is the power sale from the local network to the wide - area network at time t, P ld,t is the load power at time t.
[0165] (2) Operation constraints of the pumped - storage power station:
[0166] In the formula, and represent the upper and lower limits of the rated output of the pumped - storage power station.
[0167] For a pumped - storage system, there is also a reservoir capacity constraint: E min ≤ E t ≤ E max
[0168] In the formula, E min and E max represent the lower and upper limits of the reservoir capacity of the pumped - storage power station. Specifically, the calculation formula of E t is as follows:
[0169] E t+1 = E t - ΔE t
[0170]
[0171] In the formula, E t+1 is the capacity of the reservoir in the next time period, E tis the reservoir capacity during period t, ΔE t is the change in reservoir capacity during period t, P ps,t is the active power output of the pumped-storage power station during period t, η g is the power generation efficiency of the pumped-storage power station, η f is the pumping efficiency of the pumped-storage power station, P ps,t > 0 indicates that the pumped-storage power station is generating electricity, P ps,t < 0 indicates that the pumped-storage power station is pumping and storing energy.
[0172] (3) Operating constraints of electrochemical energy storage equipment:
[0173] In the formula, are the lower and upper limits of the active power output of the electrochemical energy storage equipment.
[0174] To avoid damage to the electrochemical energy storage caused by deep charge and discharge, the state of charge of the electrochemical energy storage needs to be within a certain range: SOC min ≤SOC t ≤SOC max
[0175] In the formula, SOC min , SOC max are the lower and upper limits allowed for SOC, respectively. Specifically, the calculation formula for SOC t is as follows:
[0176] SOC bat,t+1 = SOC bat,t -η bat ·P bat,t ·Δt / E bat
[0177]
[0178] In the formula, SOC bat,t+1 is the state of charge of the electrochemical energy storage in the next period, SOC bat,t is the state of charge of the electrochemical energy storage in the current period, E bat is the capacity of the electrochemical energy storage, η c is the charging efficiency of the electrochemical energy storage, η d is the discharging efficiency of the electrochemical energy storage, P bat,t > 0 indicates that the electrochemical energy storage is discharging, P bat,t < 0 indicates that the electrochemical energy storage is charging.
[0179] (4) Interaction constraints of external energy:
[0180] In the formula, P gso,export_max is the maximum power supplied by the new power system to the large power grid, Pgso,import_max The maximum power supplied by the large power grid to the new power system.
[0181] Next, the above new power system model is transformed into a Markov decision model, and the problem of minimizing the operating cost of the new power system is converted into the problem of solving the maximum Markov reward.
[0182] In the optimization scheduling decision-making process of the new power system, the agent outputs the charge and discharge power decisions of the electrochemical energy storage and the pumped-storage power station and the scheduling strategy of the dispatchable energy to meet the actual power consumption needs of the load. However, the charge and discharge decisions of the electrochemical energy storage and the pumped-storage power station are discrete variable decisions, and the scheduling strategy of the dispatchable energy is a continuous variable decision. Therefore, the energy scheduling problem of the new power system is a mixed decision problem with discrete and continuous action variables with complex constraints.
[0183] To solve this problem, in the present invention, a two-layer agent structure is designed. The upper-layer agent outputs the charge and discharge decisions of the electrochemical energy storage and the pumped-storage power station, and returns the decision actions to the environment and transmits them to the lower-layer agent. This decision is a discrete decision and is solved using the deep Q-network DQN algorithm; the lower-layer agent receives the charge and discharge decisions transmitted by the upper-layer agent, and on this basis, uses the AC network to solve the optimal action using the deep deterministic policy gradient algorithm DDPG algorithm, that is, the scheduling strategy of the dispatchable energy. This decision is a continuous variable decision and is solved using the DDPG algorithm. The agents in both the upper and lower layers can perceive the current state of the environment, and the action decision spaces of the upper-layer agent and the lower-layer agent act on the environment together. Thus, a mixed decision of continuous actions and discrete actions is realized.
[0184] Next, the Markov decision framework designs of the upper-layer agent and the lower-layer agent in the present invention are introduced separately.
[0185] (1) Design of the upper-layer agent:
[0186] (a) Design of the state space: In this new power system optimization scheduling model, the environmental state refers to the wind power generation power P we,t , the photovoltaic power generation power P pv,t , the load power P ld,t , the selling electricity price p sell,t , the purchasing electricity price p buy,t , the reservoir capacity E t , the state of charge SOC of the electrochemical energy storage bat,t and the scheduling period t in which it is located. Thus, for this optimization scheduling problem, its state space can be expressed as:
[0187] s t =[P we,t ,P pv,t ,Pld,t , p sell,t , p buy,t , E t , SOC bat,t , t]
[0188] (b) Action space design: To simplify the model, the charge and discharge decisions of the electrochemical energy storage and pumped-storage power stations are defined as three states:
[0189] a upper_t = {-1, 0, +1}
[0190] When the action value is -1, the electrochemical energy storage and pumped-storage power stations are in the discharge / generation state; when the action value is 0, no operation is performed on them; when the action value is +1, the electrochemical energy storage and pumped-storage power stations are in the charge (energy storage) state.
[0191] (c) Reward function design: The reward is set to be negative because the deep reinforcement learning algorithm is an algorithm for solving the maximum cumulative reward, while the optimization of the new power system is a minimization problem.
[0192] r upper_t = -μ(C ic,t + C bat,t + C oe,t + C ps,t )
[0193]
[0194] C ic,t is the interaction cost between the local network and the external power grid, C bat,t is the depreciation cost of the electrochemical energy storage, C oe,t is the equipment operation and maintenance cost, C ps,t is the pumping cost, and μ is the flag for the charge and discharge of the electrochemical energy storage and pumped-storage power stations (when the output power of the non-dispatchable energy can meet the load demand, μ takes 1, giving a positive reward to the agent; when the output power of the non-dispatchable energy cannot meet the load demand, μ takes -1, giving a negative reward to the agent).
[0195] (2) Lower-layer agent design:
[0196] (a) State space design (the state space of the lower-layer agent is the same as that of the upper-layer agent): In this new power system optimal scheduling model, the environmental state refers to the wind power generation power P we,t , the photovoltaic power generation power P pv,t , the load power P ld,t , the selling electricity price p sell,t , the purchasing electricity price p buy,t , the reservoir capacity E t , the state of charge SOC of the electrochemical energy storagebat,t and the scheduling period t it is in. Thus, for this optimal scheduling problem, its state space can be expressed as:
[0197] s t =[P we,t , P pv,t , P ld,t , p sell,t , p buy,t , E t , SOC bat,t , t]
[0198] (b) Action space design: In this new power system optimal scheduling model, the goal is to determine the energy storage or generation power P ps,t of the pumped-storage power station, the charge and discharge power P bat,t of the electrochemical energy storage, the power purchase P buy,t of the new power system, and the power sale P sell,t of the new power system to the large power grid. Thus, for this optimal scheduling problem, its action space can be expressed as:
[0199] a lower_t ={P ps,t , P bat,t , P buy,t , P sell,t}
[0200] (c) Reward function design: After the lower-level agent executes the action a lower_t , it obtains the immediate reward r lower_t . Therefore, the reward function is set as:
[0201] r lower_t =-η(C ic,t +C bat,t +C oe,t +C ps,t ),
[0202] where C ic,t is the interaction cost between the new power system and the large power grid, C bat,t is the depreciation cost of the electrochemical energy storage, C oe,t is the equipment operation and maintenance cost, C ps,t is the pumping cost, and η is the normalization coefficient.
[0203] To achieve the optimal scheduling decision of the new power system, the construction of the environment and the agent will be described separately below:
[0204] (1) Construction of the environment:
[0205] (a) Define the state space: Collect the current wind power generation, photovoltaic power generation, and load power and standardize them; collect the current selling electricity price and purchasing electricity price and standardize them; collect the current state of charge of the electrochemical energy storage and the reservoir capacity of the pumped-storage power station and normalize them; obtain the current time period and normalize it. Concatenate the above normalized values into a state vector, and this state vector constitutes the state space.
[0206] (b) Define the reward function: including defining the reward functions for the upper-layer agent and the lower-layer agent.
[0207] (c) Define the dispatchable energy constraint conditions: including power balance constraints, pumped-storage power station operation constraints, electrochemical energy storage operation constraints, and interaction constraints between the new power system and the large power grid.
[0208] (2) Construction of the agent:
[0209] (a) Define the internal operations of the agent: including defining the interaction process between the agent and the environment, the process of storing and extracting experience data from the memory, etc.
[0210] (b) Define the modeling and optimization process of the neural network: including constructing the neural network, defining the optimizer, setting the loss function, etc.
[0211] The following specifically describes the training steps of the upper-layer agent and the lower-layer agent.
[0212] (1) Training steps of the upper-layer agent:
[0213] Step 1: Initialize the environment and the upper-layer agent. Initialize the environment, including initializing the state space, reward function, etc. Initialize the upper-layer agent, including initializing the main value Q network and the target Q network, and copy the parameters of the main value Q network to the target Q network. The network parameters of the main value Q network and the target Q network are μ and μ' respectively (both networks are convolutional networks, they have the same network structure, the input layer contains 7 neurons, the output layer contains 3 neurons, set three convolutional layers, two fully connected layers, and the hidden layer uses the ReLU (Rectified Linear Unit) activation function).
[0214] Step 2: The upper-layer agent interacts with the environment. The upper-layer agent obtains the current state s t from the environment, and uses the ε-greedy strategy (greedy strategy) to select an action a upper_t based on the main value Q network and returns this action information to the environment. The environment receives the action information of the upper-layer agent, calculates the reward r upper_t and generates the next state s t+1 and feeds it back to the upper-layer agent. The upper-layer agent stores the quadruple data (experience data) [s t , aupper_t , r upper_t , s t+1 Store it in the experience replay pool (this process lasts for a certain number of time cycles).
[0215]
[0216] Step 3: Update the network parameters in the upper-level agent. Extract the quadruple data, and input the [s t , a upper_t into the main value Q-network to obtain Q(s t , a upper_t ; μ), and input the [s t+1 into the target Q-network to obtain Q(s t+1 , a upper_t+1 ; μ') and find the maximum Q value among them Input it into the optimizer to calculate the target value y (using this target value y as a label) and the loss function:
[0217]
[0218] Use the gradient descent method to update the parameters of the main value Q-network. The expression for updating the parameters of the main value Q-network is as follows, where 0 ≤ α ≤ 1:
[0219]
[0220] After the parameters of the main value Q-network are updated 500 times, copy the parameters of the main value Q-network to the target Q-network.
[0221] Repeat Step 3, continuously extract quadruple data, calculate the target value, and update the network parameters until the change in the parameters of the main value Q-network is less than 0.0001, and the training terminates. The upper-level agent outputs the charge and discharge decisions to the environment and the lower-level agent.
[0222] (2) Training steps for the lower-level agent:
[0223] Step 4: Initialize the environment and the lower-level agent. Initialize the environment, including initializing the state space, reward function, etc. Initialize the main policy network and the main value network, and copy the parameters of the main policy network and the main value network to the target policy network and the target value network. The parameters of the main policy network and the target policy network are w and w' respectively, and the parameters of the main value network and the target value network are θ and θ' respectively (the input layer of the policy network contains 7 neurons, the output layer contains 6 neurons, and it consists of two fully connected layers, where the hidden layer uses the ReLU activation function and the output layer uses the tanh (hyperbolic tangent function) activation function; the input layer of the value network contains 13 neurons, the output layer contains 1 neuron, and it consists of two fully connected layers, where the hidden layer uses the ReLU activation function).
[0224] Step 5: The lower-level agent interacts with the environment. The lower-level agent obtains the current state s from the environment t , and based on the ε-greedy policy, selects an action a using the main policy network lower_t and returns this action information to the environment. The environment receives the action information of the lower-level agent, calculates the reward r lower_t and generates the next state s t+1 and feeds it back to the lower-level agent. The lower-level agent stores the quadruple [s t , a lower_t , r lower_t , s t+1 in the experience replay pool (this process continues for a certain number of time steps).
[0225]
[0226] Step 6: Update the network parameters in the lower-level agent. Extract the quadruple data, input [s t , a lower_t into the main value network to obtain Q(s t , a lower_t ; θ), input [s t+1 into the target policy network to obtain a lower_t+1 , and input [s t+1 , a lower_t+1 into the target value network to obtain Q(s t+1 , a lower_t+1 ; θ'), calculate the target value y t (using this target value y t as the label):
[0227] y t = r lower_t + λQ(s t+1 , a lower_t+1 ; θ')
[0228] Minimize the loss function to make the Q value output by the main value network close to the target Q value, and update the main value network parameters:
[0229]
[0230] The main policy network uses the gradient of Q(s t , a lower_t ; θ) generated by the main value network with respect to the action a lower_t and the gradient of the action a lower_t with respect to the network parameter w to update. The policy gradient of the main policy network during update is as follows. Use gradient ascent to update the main policy network parameter w, where 0 ≤ β ≤ 1:
[0231]
[0232] After the main policy network and the main value network parameters are updated 500 times, the main network parameters are copied to the target network. The update of the target policy network and the target value network adopts soft update (introducing a learning rate τ (0 < τ < 1), taking the weighted average of the old target network parameters and the new corresponding network parameters, and then assigning them to the target network, that is, the network parameters are updated slowly):
[0233]
[0234] Repeat step 6, continuously extract quadruple data, calculate the target, and update the network parameters until the parameter changes of the main policy network and the main value network are less than 0.0001, and the training terminates. The lower-level agent outputs a scheduling policy to the environment.
[0235] Step 7: Repeat iteration. Repeat steps 2, 3, 5, and 6. The upper and lower-level agents continuously interact with the environment, generate new experience data, and update the network parameters in the agents until the maximum number of training times, 500 times, is reached, and the training terminates.
[0236] After the agent decision-making model is trained, the agent outputs corresponding actions according to the state information input in the current period, that is, the decision variables to be solved. These decision variables specifically include: the charge and discharge strategies of energy storage devices, including when to charge, when to discharge, and the discharge power, etc.; the allocation and utilization strategies of energy resources, the output power of each dispatchable energy at a specific time point; the energy interaction strategy between the new power system and the large power grid, deciding when to purchase electric energy from the large power grid and the purchase power or when to sell electric energy to the large power grid and the selling power.
[0237] The agent decision-making model proposed by the present invention makes dynamic scheduling decisions using the state information observed in each period, realizing the direct mapping from the current state of the system to the scheduling decision.
[0238] The embodiment of the present invention also provides a short-term optimal scheduling device for a power system including pumped storage and electrochemical energy storage, including:
[0239] A construction module for constructing an optimal scheduling model of a new power system considering the coordinated cooperation among power sources (WE, PV, PS), the local power grid and the large power grid, load, and energy storage batteries, establishing an objective function for minimizing the system operation cost, and determining the dispatchable energy constraint conditions;
[0240] A training module for converting the optimal scheduling model of the new power system into a two-layer agent joint Markov decision model and training and optimizing the decision model;
[0241] The scheduling module is used to apply the trained decision-making model to the actual power system, output the optimal decision according to the current state, and achieve the optimal scheduling of the new power system.
[0242] The objective function is as follows:
[0243]
[0244] In the above formula, C ic,t is the interaction cost between the new power system and the large power grid, C bat,t is the depreciation cost of the electrochemical energy storage, C oe,t is the equipment operation and maintenance cost, C ps,t is the pumping cost; among them,
[0245]
[0246] In the formula, p buy,t is the electricity purchase price of the new power system from the large power grid at time t, P buy,t is the electricity purchase power of the new power system from the large power grid at time t, p sell,t is the electricity selling price of the new power system to the large power grid at time t, P sell,t is the electricity selling power of the new power system to the large power grid at time t, ρ bat is the charge and discharge depreciation cost coefficient of the electrochemical energy storage, P bat,t is the active power of the electrochemical energy storage at time t, P bat,t >0 indicates that the electrochemical energy storage discharges, P bat,t <0 indicates that the electrochemical energy storage charges, k oe_i is the coefficient of the operating cost of various distributed energy sources in the new power system at time t, P i,t is the active power of each energy source in the new power system at time t, P ps,t is the active power of the pumped-storage power station at time t, C s is the pumping cost coefficient of the pumped-storage power station, ε is a judgment parameter, P ps,t >0 means the pumped-storage power station generates electricity, and the value is 1, P ps,t <0 means the pumped-storage power station stores energy, and the value is 0.
[0247] The adjustable energy constraint conditions include:
[0248] Power balance constraint: P ps,t +P bat,t +P we,t +P pv,t +P buy,t =P ld,t +P sell,t
[0249] In the formula, P ps,tThe active power of the pumped - storage power station at time t, P bat,t The active power of the electrochemical energy storage device at time t, P we,t The active power output by the wind power generation system at time t, P pv,t The active power output by the photovoltaic system at time t, P buy,t The power purchase from the local network to the wide - area network at time t, P sell,t The power sale from the local network to the wide - area network at time t, P ld,t The load power at time t;
[0250] Operating constraints of the pumped - storage power station:
[0251] In the formula, and represent the upper and lower limits of the rated output of the pumped - storage power station;
[0252] Storage capacity constraint of the pumped - storage power station: E min ≤E t ≤E max
[0253] In the formula, E min and E max represent the lower and upper limits of the reservoir storage capacity of the pumped - storage power station; E t The calculation formula of is as follows:
[0254]
[0255] In the formula, E t+1 is the reservoir capacity in the next time period, E t is the reservoir capacity in time period t, ΔE t is the change in storage capacity in time period t, P ps,t is the active power output of the pumped - storage power station in time period t, η g is the power generation efficiency of the pumped - storage power station, η f is the pumping efficiency of the pumped - storage power station, P ps,t >0 indicates that the pumped - storage power station is generating electricity, P ps,t <0 indicates that the pumped - storage power station is pumping and storing energy;
[0256] Operating constraints of the electrochemical energy storage device:
[0257] In the formula, are the lower and upper limits of the active power output of the electrochemical energy storage device;
[0258] To avoid damage to the electrochemical energy storage caused by deep charge - discharge, the state of charge of the electrochemical energy storage needs to satisfy: SOC min ≤SOC t ≤SOC max ;
[0259] Wherein, SOC min , SOC max are respectively the lower limit and the upper limit allowed by SOC; specifically, the calculation formula of SOC t is as follows:
[0260] SOC bat,t+1 = SOC bat,t -η bat ·P bat,t ·Δt / E bat
[0261]
[0262] Wherein, SOC bat,t+1 is the state of charge of the electrochemical energy storage in the next time period, SOC bat,t is the state of charge of the electrochemical energy storage in the current time period, E bat is the capacity of the electrochemical energy storage, η c is the charging efficiency of the electrochemical energy storage, η d is the discharging efficiency of the electrochemical energy storage, P bat,t > 0 indicates that the electrochemical energy storage discharges, P bat,t < 0 indicates that the electrochemical energy storage charges;
[0263] Interaction constraint of external energy sources:
[0264] Wherein, P gso,export_max is the maximum power supplied by the new power system to the large power grid, P gso,import_max is the maximum power supplied by the large power grid to the new power system.
[0265] The specific double-layer agent joint Markov decision model is as follows: The upper-layer agent outputs the charge and discharge decisions of the electrochemical energy storage and the pumped-storage power station, and returns the decision actions to the environment and transmits them to the lower-layer agent. This charge and discharge decision is a discrete decision, which is solved by using the deep Q-network algorithm; the lower-layer agent receives the charge and discharge decisions transmitted by the upper-layer agent, solves the optimal action, that is, the dispatching strategy of the dispatchable energy. This decision is a continuous variable decision, which is solved by using the deep deterministic policy gradient algorithm; both the upper and lower layers of agents can perceive the current state of the environment, and the action decision spaces of the upper-layer agent and the lower-layer agent jointly act on the environment, thereby realizing the hybrid decision of continuous actions and discrete actions.
[0266] The training module is specifically used for the training of the upper-layer agent and the training of the lower-layer agent, and the steps are as follows:
[0267] Step 1: Initialize the environment and the upper-level agent; Initialize the environment, including initializing the state space and the reward function; Initialize the upper-level agent, including initializing the main value Q-network and the target Q-network, and copy the parameters of the main value Q-network to the target Q-network. The network parameters of the main value Q-network and the target Q-network are μ and μ' respectively; Both networks are convolutional networks with the same network structure. The input layer contains 7 neurons, the output layer contains 3 neurons, and three convolutional layers and two fully connected layers are set. The hidden layer uses the rectified linear unit ReLU activation function;
[0268] Step 2: The upper-level agent interacts with the environment; The upper-level agent obtains the current state s from the environment t , and based on the ε-greedy greedy strategy, selects an action a based on the main value Q-network upper_t and returns this action information to the environment. The environment receives the action information of the upper-level agent, calculates the reward r upper_t and generates the next state s t+1 and feeds it back to the upper-level agent. The upper-level agent stores the quadruple data [s t , a upper_t , r upper_t , s t+1 into the experience replay pool;
[0269]
[0270] Step 3: Update the network parameters in the upper-level agent; Extract the quadruple data, input [s t , a upper_t into the main value Q-network to get Q(s t , a upper_t ; μ), [s t+1 into the target Q-network to get Q(s t+1 , a upper_t+1 ; μ') and find the maximum Q value among them Input it into the optimizer, calculate the target value y and use this target value y as the label and the loss function:
[0271]
[0272] Use the gradient descent method to update the parameters of the main value Q-network. The expression for updating the parameters of the main value Q-network is as follows, where 0 ≤ α ≤ 1:
[0273]
[0274] After the parameters of the main value Q-network are updated N times, copy the parameters of the main value Q-network to the target Q-network;
[0275] Repeat the network parameter update in the upper-layer agent in Step 3, continuously extract quadruple data, calculate the target value, and update the network parameters until the parameter change of the main value Q network is less than the set threshold, and the training terminates; the upper-layer agent outputs charge and discharge decisions to the environment and the lower-layer agent.
[0276] Step 4: Initialize the environment and the lower-layer agent; Initialize the environment, including initializing the state space, reward function, etc.; Initialize the main policy network and the main value network, and copy the parameters of the main policy network and the main value network to the target policy network and the target value network. The parameters of the main policy network and the target policy network are w and w′ respectively, and the parameters of the main value network and the target value network are θ and θ′ respectively; The input layer of the policy network contains 7 neurons, and the output layer contains 6 neurons, which consists of two fully connected layers. The ReLU activation function is used in the hidden layer, and the hyperbolic tangent function tanh activation function is used in the output layer; The input layer of the value network contains 13 neurons, and the output layer contains 1 neuron, which consists of two fully connected layers. The ReLU activation function is used in the hidden layer.
[0277] Step 5: The lower-layer agent interacts with the environment; The lower-layer agent obtains the current state s from the environment t , and based on the ε-greedy greedy strategy, selects an action a based on the main policy network lower_t and returns this action information to the environment. The environment receives the action information of the lower-layer agent, calculates the reward r lower_t and generates the next state s t+1 and feeds it back to the lower-layer agent. The lower-layer agent stores the quadruple [s t , a lower_t , r lower_t , s t+1 into the experience replay pool.
[0278]
[0279] Step 6: Update the network parameters in the lower-layer agent; Extract the quadruple data, input the [s t , a lower t in the quadruple into the main value network to get Q(s t , a lower_t ; θ), input [s t+1 into the target policy network to get a lower_t+1 , and input [s t+1 , a lower_t+1 into the target value network to get Q(s t+1 , a lower_t+1 ; θ'), calculate the target value y t , and use this target value y t as the label:
[0280] y t = r lower_t + λQ(s t+1 , a lower_t+1 ; θ')
[0281] Minimize the loss function to make the Q - value output by the main value network close to the target Q - value, and update the parameters of the main value network:
[0282]
[0283] The main policy network uses the Q(s t , a lower_t ; θ) generated by the main value network to update the gradient of the action a lower_t with respect to the network parameters w and the gradient of the action a lower_t with respect to the network parameters w; the policy gradient of the main policy network during update is as follows. Use gradient ascent to update the parameters w of the main policy network, where 0 ≤ β ≤ 1:
[0284]
[0285] After the parameters of the main policy network and the main value network are updated N times, copy the main network parameters to the target network; the update of the target policy network and the target value network adopts soft update. Introduce a learning rate τ, 0 < τ < 1, and take the weighted average of the old target network parameters and the new corresponding network parameters, and then assign them to the target network, that is, the network parameters are updated slowly:
[0286]
[0287] Repeat step 6, continuously extract quadruple data, calculate the target, and update the network parameters until the parameter changes of the main policy network and the main value network are less than the set threshold, and the training terminates; the lower - layer agent outputs a scheduling policy to the environment;
[0288] Step 7: Repeat iteration; repeat steps 2, 3, 5, and 6. The upper - layer and lower - layer agents continuously interact with the environment, generate new experience data, and update the network parameters in the agents until the maximum number of training times is reached, and the training terminates.
[0289] The scheduling module is specifically used for: after the training of the two - layer agent joint Markov decision model is completed, the agent outputs the corresponding action according to the state information input in the current period, that is, the decision variable to be solved, to realize the direct mapping from the current state of the system to the scheduling decision.
[0290] On the other hand, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the short - term optimal scheduling method of the power system with pumped - storage energy storage and electrochemical energy storage described in the first aspect.
[0291] On the other hand, the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the short-term optimal scheduling method of the power system with pumped storage and electrochemical energy storage described in the first aspect.
[0292] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0293] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1 a device for the functions specified in one block or multiple blocks.
[0294] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements in the process Figure 1 one process or multiple processes and / or blocks Figure 1 a device for the functions specified in one block or multiple blocks.
[0295] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process or multiple processes and / or blocks Figure 1 one process or multiple processes and / or blocks Figure 1 a device for the functions specified in one block or multiple blocks.
[0296] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent substitutions can still be made to the specific implementation manners of the present invention, and any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.
Claims
1. A short-term optimization scheduling method for a power system including pumped storage and electrochemical energy storage, characterized in that: include: Construct a new power system optimization dispatch model that takes into account the coordination between power sources, local networks and large power grids, loads, and energy storage batteries, establish an objective function that minimizes system operating costs, and determine dispatchable energy constraints; Converting the novel power system optimization dispatching model into a two-layer agent joint Markov decision model, and training and optimizing the decision model; The trained decision-making model is applied to the actual power system, and the optimal decision is output according to the current state to achieve optimal scheduling of the new power system.
2. The short-term optimal dispatching method for a power system including pumped storage and electrochemical energy storage according to claim 1 is characterized in that: The objective function is: In the above formula, C ic,t is the interaction cost between the new power system and the large power grid, C bat,t is the depreciation cost of electrochemical energy storage, C oe,t is the equipment operation and maintenance cost, C ps,t is the pumping cost; In the formula, p buy,t is the electricity price of the new power system purchased from the large power grid at time t, P buy,t is the power purchased by the new power system from the large power grid at time t, p sell,t is the electricity price sold by the new power system to the large power grid at time t, P sell,t is the power sold by the new power system to the large power grid at time t, ρ bat The charge and discharge depreciation cost coefficient of electrochemical energy storage, P bat,t is the active power of electrochemical energy storage at time t, P bat,t >0 indicates electrochemical energy storage discharge, P bat,t <0 means electrochemical energy storage charging, k oe_i is the coefficient of the operating cost of various distributed energy sources in the new power system at time t, P i,t is the active power of each energy source in the new power system at time t, P ps,t is the active power of the pumped storage power station at time t, C s is the pumping cost coefficient of the pumped storage power station, ε is the judgment parameter, P ps,t >0, the pumped storage power station generates electricity, the value is 1, P ps,t When <0, the pumped storage power station stores energy and the value is 0.
3. The short-term optimal dispatching method for a power system including pumped storage and electrochemical energy storage according to claim 2 is characterized in that: The dispatchable energy constraints include: Power balance constraint: P ps,t +P bat,t +P we,t +P pv,t +P buy,t =P ld,t +P sell,t Where P ps,t is the active power of the pumped storage power station at time t, P bat,t The active power of the electrochemical energy storage device at time t, P we,t is the active power output of the wind power generation system at time t, P pv,t is the active power output of the photovoltaic system at time t, P buy,t is the power purchased by the local network from the wide area network at time t, P sell,t is the power sold from the local network to the wide area network at time t, P ld,t is the load power at time t; Pumped storage power station operation constraints: In the formula, and Indicates the upper and lower limits of the rated output of the pumped storage power station; Pumped storage power station storage capacity constraint: E min ≤E t ≤E max In the formula, E min and E max Indicates the lower and upper limits of the reservoir capacity of the pumped storage power station; E t The calculation formula is as follows: E t+1 =E t -ΔE t In the formula, E t+1 is the capacity of the reservoir in the next period, E t is the capacity of the reservoir in period t, ΔE t is the change in storage capacity during period t, P ps,t is the active power output of the pumped storage power station in period t, η g is the power generation efficiency of the pumped storage power station, η f is the pumping efficiency of the pumped storage power station, P ps,t >0 indicates that the pumped storage power station generates electricity, P ps,t <0 indicates pumped storage power station; Electrochemical energy storage equipment operation constraints: In the formula, The lower and upper limits of the active power output of the electrochemical energy storage device; In order to avoid damage to electrochemical energy storage caused by deep charging and discharging, the state of charge of electrochemical energy storage needs to meet the following requirements: SOC min ≤SOC t ≤SOC max ; In the formula, SOC min , SOC max are the lower and upper limits of SOC respectively; specifically, SOC t The calculation formula is as follows: SOCIETY bat,t+1 =SOC bat,t -η bat ·P bat,t ·Δt / E bat In the formula, SOC bat,t+1 SOC is the state of charge of electrochemical energy storage in the next period of time. bat,t is the state of charge of the electrochemical energy storage in the current period, E bat is the capacity of electrochemical energy storage, η c is the charging efficiency of electrochemical energy storage, η d is the discharge efficiency of electrochemical energy storage, P bat,t >0 indicates electrochemical energy storage discharge, P bat,t <0 indicates electrochemical energy storage charging; Interaction constraints of external energy sources: Where P gso,export_max The maximum power supplied by the new power system to the large power grid, P gso,import_max The maximum power that can be supplied to the new power system from the large power grid.
4. The short-term optimal dispatching method for a power system including pumped storage and electrochemical energy storage according to claim 1 is characterized in that: The two-layer intelligent agent joint Markov decision model is specifically as follows: the upper-layer intelligent agent outputs the charging and discharging decisions of electrochemical energy storage and pumped-storage power stations, and returns the decision actions to the environment and transmits them to the lower-layer intelligent agent. This charging and discharging decision is a discrete decision, which is solved using a deep Q-network algorithm; the lower-layer intelligent agent receives the charging and discharging decision transmitted by the upper-layer intelligent agent, and solves the optimal action, that is, the scheduling strategy for scheduling energy. This decision is a continuous variable decision, which is solved using a deep deterministic policy gradient algorithm; the upper and lower-layer intelligent agents can both perceive the current state of the environment, and the action decision space of the upper-layer intelligent agent and the action decision space of the lower-layer intelligent agent act together on the environment, thereby realizing a mixed decision of continuous and discrete actions.
5. The short-term optimal dispatching method for a power system including pumped storage and electrochemical energy storage according to claim 4 is characterized in that: The converting of the novel power system optimization dispatching model into a two-layer intelligent agent joint Markov decision model includes upper-layer intelligent agent design, lower-layer intelligent agent design, environment construction, and intelligent agent construction; The upper-level agent design includes: State space design: In this new power system optimization dispatch model, the environmental state refers to the wind power generation power P we,t , Photovoltaic power generation power P pv,t , load power P ld,t 、Electricity sales price sell,t 、Purchase electricity price buy,t 、Storage capacity E t , Electrochemical Energy Storage State of Charge SOC bat,t And the scheduling period t; therefore, for this optimization scheduling problem, its state space can be expressed as: s t =[P we,t ,P pv,t ,P ld,t ,p sell,t ,p buy,t ,E t ,SOC bat,t ,t] Action space design: In order to simplify the model, the charging and discharging decisions of electrochemical energy storage and pumped storage power stations are defined as three states: a upper_t ={-1,0,+1} When the action value is -1, the electrochemical energy storage and pumped storage power station are in the discharge / generation state; when the action value is 0, no operation is performed on them; when the action value is +1, the electrochemical energy storage and pumped storage power station are in the charging / storage state; Reward function design: The objective function is converted into a reward function, and the reward is set to a negative value; r upper_t =-μ(C ic,t +C bat,t +C oe,t +C ps,t ) C ic,t is the interaction cost between the local network and the external power grid, C bat,t is the depreciation cost of electrochemical energy storage, C oe,t is the equipment operation and maintenance cost, C ps,t is the pumping cost, μ is the symbol of the charging and discharging of electrochemical energy storage and pumped storage power station. When the output power of the non-dispatchable energy can meet the load demand, μ takes 1, giving the agent a positive reward; when the output power of the non-dispatchable energy cannot meet the load demand, μ takes -1, giving the agent a negative reward; The lower-level agent design includes: State space design: The state space of the lower-level agent is consistent with that of the upper-level agent: In this new power system optimization dispatch model, the environmental state refers to the wind power generation power P we,t , Photovoltaic power generation power P pv,t , load power P ld,t 、Electricity sales price sell,t 、Purchase electricity price buy,t 、Storage capacity E t , Electrochemical Energy Storage State of Charge SOC bat,t And the scheduling period t; therefore, for this optimization scheduling problem, its state space can be expressed as: s t =[P we,t ,P pv,t ,P ld,t ,p sell,t ,p buy,t ,E t ,SOC bat,t ,t] Action space design: In this new power system optimization dispatch model, the goal is to determine the storage or generation power P of the pumped storage power station. ps,t 、Charge and discharge power P of electrochemical energy storage bat,t 、Purchasing power of new power system buy,t And the power sold by the new power system to the large power grid P sell,t ; For this optimization scheduling problem, the action space can be expressed as: a lower_t ={P ps,t ,P bat,t ,P buy,t ,P sell,t } Reward function design: The lower-level agent performs action a lower_t After that, you will get an immediate reward. lower_t , set the reward function to: r lower_t =-η(C ic,t +C bat,t +C oe,t +C ps,t ), In the formula, C ic,t is the interaction cost between the new power system and the large power grid, C bat,t is the depreciation cost of electrochemical energy storage, C oe,t is the equipment operation and maintenance cost, C ps,t is the pumping cost, η is the normalization coefficient; The environment structure includes: Define the state space: collect the current wind power generation power, photovoltaic power generation power, and load power and standardize them; collect the current electricity sales price and electricity purchase price and standardize them; collect the current state of charge of the electrochemical energy storage and the storage capacity of the pumped storage power station and normalize them; obtain the current time period and normalize it; splice the above normalized values into a state vector, which constitutes the state space; Define reward function: including defining reward functions for upper-layer agents and lower-layer agents; Define dispatchable energy constraints: including power balance constraints, pumped storage power station operation constraints, electrochemical energy storage operation constraints, and interaction constraints between new power systems and large power grids; The intelligent agent structure includes: Define the internal operations of the agent: including defining the interaction process between the agent and the environment, the process of storing and extracting experience data from the memory, etc. Define the modeling and optimization process of the neural network: including building a neural network, defining an optimizer, setting a loss function, etc.
6. The short-term optimal dispatching method for a power system including pumped storage and electrochemical energy storage according to claim 1 is characterized in that: The steps for training and optimizing the model are as follows: Training of upper-level agents: Step 1: Initialize the environment and the upper-layer agent; initialize the environment, including initializing the state space and reward function; initialize the upper-layer agent, including initializing the principal value Q network and the target Q network, and copy the principal value Q network parameters to the target Q network. The network parameters of the principal value Q network and the target Q network are μ and μ' respectively; both networks are convolutional networks with the same network structure. The input layer contains 7 neurons and the output layer contains 3 neurons. Three convolutional layers and two fully connected layers are set. The hidden layer uses the rectified linear unit ReLU activation function; Step 2: The upper-layer agent interacts with the environment; the upper-layer agent obtains the current state s from the environment t , using the ε-greedy greedy strategy, select action a based on the principal value Q network upper_t And return this action information to the environment, the environment receives the action information of the upper-level agent and calculates the reward r upper_t And generate the next moment state s t+1 Feedback to the upper-level agent, the upper-level agent will send the four-tuple data [s t ,a upper_t ,r upper_t ,s t+1 ]Stored in the experience replay pool; Step 3: Update the network parameters in the upper agent; extract the four-tuple data and convert [s t ,a upper_t ] is input into the main value Q network to obtain Q(s t ,a upper_t ; μ), [s t+1 ] is input into the target Q network to obtain Q(s t+1 ,a upper_t+1 ; μ') and find the maximum Q value among them This is fed into the optimizer, which calculates the target value y and uses this target value y as the label and loss function: The gradient descent method is used to update the parameters of the main value Q network. The expression for updating the parameters of the main value Q network is as follows, where 0≤α≤1: After the main value Q network parameters are updated N times, the main value Q network parameters are copied to the target Q network; Repeat step 3 to update the network parameters in the upper agent, continuously extracting quadruple data, calculating target values, and updating network parameters until the parameter change of the main value Q network is less than the set threshold, and the training is terminated; the upper agent outputs the charging and discharging decision to the environment and the lower agent; Training of lower-level agents: Step 4: Initialize the environment and lower-layer agents; initialize the environment, including initializing the state space, reward function, etc.; initialize the main policy network and the main value network, and copy the parameters of the main policy network and the main value network to the target policy network and the target value network. The parameters of the main policy network and the target policy network are w and w′, respectively, and the parameters of the main value network and the target value network are θ and θ′, respectively; the policy network input layer contains 7 neurons, the output layer contains 6 neurons, and is composed of two fully connected layers, in which the hidden layer uses the ReLU activation function, and the output layer uses the hyperbolic tangent function tanh activation function; the value network input layer contains 13 neurons, the output layer contains 1 neuron, and is composed of two fully connected layers, in which the hidden layer uses the ReLU activation function; Step 5: The lower-level agent interacts with the environment; the lower-level agent obtains the current state s from the environment t , using the ε-greedy greedy strategy, select action a based on the main strategy network lower_t And return this action information to the environment, the environment receives the action information of the lower-level agent and calculates the reward r lower_t And generate the next moment state s t+1 Feedback to the lower-level agent, the lower-level agent will convert the four-tuple [s t ,a lower_t ,r lower_t ,s t+1 ]Stored in the experience replay pool; Step 6: Update the network parameters in the lower-level agent; extract the four-tuple data and convert [s t ,a lower_t ] is input into the main value network to obtain Q(s t ,a lower_t ;θ), [s t+1 ] is input into the target policy network to obtain a lower_t+1 , and [s t+1 ,a lower_t+1 ] Input the target value network to get Q(s t+1 ,a lower_t+1 ; θ'), calculate the target value y t , this target value y t As a tag: y t =r lower_t +λQ(s t+1 ,a lower_t+1 (i') Minimize the loss function to make the Q value output by the main value network close to the target Q value, and update the parameters of the main value network: The main policy network uses Q(s) generated by the main value network t ,a lower_t ;θ) for action a lower_t The gradient and action a lower_t Update the gradient of the network parameter w; the policy gradient of the main policy network during update is as follows, using gradient ascent to update the main policy network parameter w, where 0≤β≤1: After the main policy network and the main value network parameters are updated N times, the main network parameters are copied to the target network; the target policy network and the target value network are updated using soft updates, introducing a learning rate τ, 0<τ<1, and taking the weighted average of the old target network parameters and the new corresponding network parameters, and then assigning them to the target network, that is, the network parameters are updated slowly: Repeat step 6, continuously extracting quadruple data, calculating targets, and updating network parameters until the parameter changes of the main strategy network and the main value network are less than the set threshold, and the training is terminated; the lower-level agent outputs the scheduling strategy to the environment; Step 7: Repeat the iteration; Repeat steps 2, 3, 5, and 6. The upper and lower layers of agents continuously interact with the environment, generate new experience data, and update the network parameters in the agents until the maximum number of training times is reached and the training is terminated.
7. The short-term optimal dispatching method for a power system including pumped storage and electrochemical energy storage according to claim 1 is characterized in that: The trained model is applied to the actual power system, and the optimal decision is output according to the current state to realize the optimal scheduling of the new power system. Specifically, after the training of the two-layer intelligent agent combined with the Markov decision model is completed, the intelligent agent outputs the corresponding action according to the state information input in the current period, that is, the decision variable that needs to be solved, so as to realize the direct mapping of the system from the current state to the scheduling decision.
8. A short-term optimization dispatching device for a power system including pumped storage and electrochemical energy storage, characterized in that: include: The construction module is used to construct a new power system optimization dispatching model that takes into account the coordination between power sources, local networks and large power grids, loads, and energy storage batteries, establish an objective function that minimizes system operating costs, and determine dispatchable energy constraints; A training module, used to convert the novel power system optimization dispatching model into a two-layer intelligent agent joint Markov decision model, and to train and optimize the decision model; The scheduling module is used to apply the trained decision model to the actual power system, output the optimal decision according to the current state, and realize the optimized scheduling of the new power system.
9. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, a short-term optimization dispatching method for a power system containing pumped storage and electrochemical energy storage is implemented as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the short-term optimal scheduling method for a power system containing pumped storage and electrochemical energy storage as described in any one of claims 1 to 7.
Citation Information
Cited By
Intelligent hydraulic engineering data monitoring method and system
CN120355181A