Reinforcement learning-based grid-side energy storage optimal dispatching method and apparatus, terminal device, and computer-readable storage medium
By constructing a state space and reward function based on reinforcement learning for grid-side energy storage optimization scheduling, and training an energy storage optimization scheduling model, the problem that traditional algorithms cannot adapt to the uncertainty of new energy power generation is solved, and the optimization scheduling of energy storage systems is realized, thereby improving the stability and reliability of the power grid.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- GUANGDONG POWER GRID CO LTD
- Filing Date
- 2025-02-26
- Publication Date
- 2026-05-21
AI Technical Summary
Traditional optimization algorithms cannot effectively adapt to the uncertainties and dynamic changes in new energy power generation, resulting in the energy storage system failing to achieve optimal utilization.
A grid-side energy storage optimization scheduling method based on reinforcement learning is adopted. By constructing a state space, action space and reward function, the grid-side energy storage optimization scheduling model is trained to optimize the charging and discharging of the energy storage system and the generator output, so as to maximize the total value of the energy storage system.
It improves the stability and reliability of the power grid, adapts to the uncertain and dynamically changing energy environment, and enables the optimized scheduling of energy storage systems.
Smart Images

Figure CN2025079406_21052026_PF_FP_ABST
Abstract
Description
A reinforcement learning-based method, device, terminal equipment, and computer-readable storage medium for optimized scheduling of grid-side energy storage. Technical Field
[0001] This invention relates to the field of energy storage optimization scheduling technology, and in particular to a grid-side energy storage optimization scheduling method, apparatus, terminal equipment, and computer-readable storage medium based on reinforcement learning. Background Technology
[0002] With the transformation of the energy structure, the installed capacity of new energy power generation, especially wind and solar power, has grown rapidly. However, wind and solar power generation is characterized by significant randomness and intermittency, making it difficult to accurately predict their output. This leads to significant uncertainties and volatility challenges in grid dispatch and operation. To ensure the safety and stability of the power system, energy storage technology, as an important means of regulating grid balance, has gradually attracted attention. Energy storage systems can not only store energy when there is a surplus of new energy generation and release energy during peak load periods, but also enhance the flexibility and reliability of the grid through frequency regulation and backup power. Despite these advantages, how to effectively plan and manage energy storage systems, especially considering the overall value of the system, has become a key issue facing the current power system.
[0003] Existing grid-side energy storage system planning typically relies on traditional optimization algorithms. These algorithms are usually based on deterministic models and cannot effectively adapt to uncertainties and dynamically changing energy environments. Consequently, they struggle to cope with the randomness and volatility of new energy power generation, resulting in the energy storage system failing to achieve optimal utilization. Summary of the Invention
[0004] This invention provides a grid-side energy storage optimization scheduling method, device, terminal equipment, and computer-readable storage medium based on reinforcement learning. It solves the problem that traditional optimization algorithms cannot effectively adapt to uncertain and dynamically changing energy environments, thereby better adapting to grid demands and improving grid stability and reliability.
[0005] One embodiment of the present invention provides a grid-side energy storage optimization scheduling method based on reinforcement learning, comprising:
[0006] Obtain the current state space of the energy storage system; the state space includes: remaining capacity, load, renewable energy output, and electricity price;
[0007] The state space of the current energy storage system is input into a preset grid-side energy storage optimization scheduling model, so that the grid-side energy storage optimization scheduling model outputs the corresponding action space based on the current state space of the energy storage system. The action space includes: energy storage charging, energy storage discharging, energy storage scheduling capacity, and generator output. When training the grid-side energy storage optimization scheduling model, a reward function is constructed based on the value of the grid-side energy storage system, and then the loss function of the grid-side energy storage optimization scheduling model is updated based on the reward function until the loss function converges.
[0008] Based on the corresponding action space, the grid-side energy storage is optimized and scheduled.
[0009] Furthermore, the value creation of grid-side energy storage systems includes:
[0010] The value of grid-side energy storage systems is determined by the following factors: reducing grid losses, delaying new grid construction investment, improving grid operation reliability, smoothing renewable energy generation output, and reducing thermal power spinning reserve capacity.
[0011] Furthermore, to reduce grid loss value, the system is constructed based on the operating parameters of the power system and electricity prices. The operating parameters of the power system include: distribution network voltage, power factor, load power factor, line resistance under grid load conditions considering energy storage systems, and charging and discharging power of energy storage systems.
[0012] Delaying the investment value of new power grid construction, and constructing based on the investment cost of power grid infrastructure upgrades, the peak load reduction of energy storage systems, and the expected peak demand of energy storage systems;
[0013] To enhance the value of power grid operation reliability, the following parameters are constructed: power grid operation reliability parameters include: outage frequency, loss assessment rate of important users, remaining energy of energy storage systems, power supply reliability, total annual operating hours of important users, and power required for power supply to key users.
[0014] The value of smoothing renewable energy power generation output is constructed based on renewable energy power generation parameters, including: the feed-in tariff of renewable energy, the upper limit power of the node, the lower limit power of the node, the value of the energy storage system exceeding the upper limit power after charging and discharging, the value of the energy storage system exceeding the lower limit power after charging and discharging, the upper and lower fluctuation limits of wind farm output during grid dispatch, the actual active power output of the wind farm, and the charging and discharging power of the energy storage system participating in smoothing the fluctuations in wind farm output.
[0015] To reduce the value of thermal power spinning reserve capacity, the thermal power spinning reserve parameters are constructed based on the parameters of thermal power spinning reserve. These parameters include: the cost of thermal power plant spinning reserve in the absence of energy storage in the power system, the total thermal power replenishment time within one year, and the time period for replenishing power due to wind power output prediction errors.
[0016] Furthermore, the construction of the reward function includes:
[0017] A reward function is constructed based on the value of the grid-side energy storage system, the present value of the total investment cost of the grid-side energy storage system, and the present value of the total operating cost of the grid-side energy storage system.
[0018] Furthermore, the present value of the total investment cost of the grid-side energy storage system is constructed based on the investment cost, discount rate, unit cost of energy storage capacity, configured energy storage capacity, unit cost of energy storage converter, and rated power of energy storage converter;
[0019] The present value of the total operating cost of the grid-side energy storage system is constructed based on the discount rate, generator active power output, generator operating time, and generator economic operating parameters.
[0020] Furthermore, the pre-defined grid-side energy storage optimization scheduling model is trained in the following way:
[0021] For the grid-side energy storage optimization scheduling model to be trained, the state space, action space and reward function of the model are defined;
[0022] Construct the policy network based on the state space, action space, and current policy network parameters;
[0023] Obtain the initialized state space and initialized action space;
[0024] The initialized state space and initialized action space are input into the grid-side energy storage optimization scheduling model to be trained for iterative updates, so that the grid-side energy storage optimization scheduling model to be trained updates the policy network through the DDPG algorithm until the loss function converges, thus obtaining the preset grid-side energy storage optimization scheduling model.
[0025] Based on the above method embodiments, the present invention provides corresponding device embodiments, including: an energy storage status acquisition module, an action strategy output module, and an optimization scheduling module;
[0026] The energy storage status acquisition module is used to acquire the current state space of the energy storage system; the state space includes: remaining capacity, load, renewable energy output, and electricity price;
[0027] The action strategy output module is used to input the current state space of the energy storage system into the preset grid-side energy storage optimization scheduling model, so that the grid-side energy storage optimization scheduling model outputs the corresponding action space according to the current state space of the energy storage system. The action space includes: energy storage charging, energy storage discharging, energy storage scheduling capacity and generator output. Among them, when training the grid-side energy storage optimization scheduling model, a reward function is constructed based on the value of the grid-side energy storage system, and then the loss function of the grid-side energy storage optimization scheduling model is updated according to the reward function until the loss function converges.
[0028] The optimization scheduling module is used to optimize the scheduling of grid-side energy storage based on the corresponding action space.
[0029] Furthermore, the action strategy output module includes: a sub-module for constructing the value of grid-side energy storage systems;
[0030] The grid-side energy storage system value construction submodule is used to construct the value of the grid-side energy storage system based on the value of reducing grid losses, delaying new grid construction investment, improving grid operation reliability, smoothing renewable energy generation output, and reducing thermal power spinning reserve capacity.
[0031] Based on the above method embodiments, the present invention provides a corresponding terminal device embodiment, including: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the steps of the reinforcement learning-based grid-side energy storage optimization scheduling method as described in the present invention.
[0032] Based on the above method embodiments, the present invention provides a corresponding computer-readable storage medium embodiment, including: a stored computer program, which, when the computer program is running, controls the device where the computer-readable storage medium is located to execute the steps of the reinforcement learning-based grid-side energy storage optimization scheduling method described in the present invention.
[0033] Compared with the prior art, the beneficial effects of this embodiment are as follows:
[0034] This invention constructs a state space for a current energy storage system by acquiring its remaining capacity, load, renewable energy output, and electricity price. This state space is then input into a pre-defined grid-side energy storage optimization scheduling model. Based on this state space, the model outputs a corresponding action space, yielding the capacity for energy storage charging, discharging, and scheduling, as well as the generator output. During training, a reward function is constructed based on the value of the grid-side energy storage system. This reward function is then used to update the model's loss function until convergence. The grid-side energy storage optimization scheduling model, based on a reinforcement learning framework, learns through continuous interaction with the environment, enabling it to adapt to uncertainties and dynamic changes in the energy environment. Finally, based on the action space output by the model, optimized scheduling of the grid-side energy storage is performed to achieve optimal scheduling of the grid-side energy storage system.
[0035] In summary, this invention maximizes the total value of grid-side energy storage systems by constructing a reinforcement learning-based grid-side energy storage optimization scheduling model. It optimizes the scheduling of grid-side energy storage through the corresponding action space, solving the problem that traditional optimization algorithms cannot effectively adapt to uncertain and dynamically changing energy environments. This allows for better adaptation to grid demands and improves grid stability and reliability. Attached Figure Description
[0036] Figure 1 is a flowchart illustrating a grid-side energy storage optimization scheduling method based on reinforcement learning according to an embodiment of the present invention.
[0037] Figure 2 is a flowchart illustrating the training process of a grid-side energy storage optimization scheduling model provided in an embodiment of the present invention;
[0038] Figure 3 is a schematic diagram of the structure of a grid-side energy storage optimization scheduling device based on reinforcement learning provided in an embodiment of the present invention. Detailed Implementation
[0039] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0040] As shown in Figure 1, an embodiment of the present invention provides a grid-side energy storage optimization scheduling method based on reinforcement learning, which includes at least the following steps:
[0041] Step S1: Obtain the current state space of the energy storage system; the state space includes: remaining capacity, load, renewable energy output, and electricity price;
[0042] For step S1, obtain the remaining capacity C of the current energy storage system. t Load P L Renewable energy output P w and electricity prices t By acquiring the above information and combining it into a state vector s t ={C t ,P L ,P w ,e t This vector represents the current state space of the energy storage system.
[0043] Step S2: Input the state space of the current energy storage system into the preset grid-side energy storage optimization scheduling model, so that the grid-side energy storage optimization scheduling model outputs the corresponding action space according to the state space of the current energy storage system; the action space includes: energy storage charging, energy storage discharging, energy storage scheduling capacity and generator output; wherein, when training the grid-side energy storage optimization scheduling model, a reward function is constructed according to the value of the grid-side energy storage system, and then the loss function of the grid-side energy storage optimization scheduling model is updated according to the reward function until the loss function converges;
[0044] In a preferred embodiment, the value creation of a grid-side energy storage system includes:
[0045] The value of grid-side energy storage systems is determined by the following factors: reducing grid losses, delaying new grid construction investment, improving grid operation reliability, smoothing renewable energy generation output, and reducing thermal power spinning reserve capacity.
[0046] In a preferred embodiment, the value of grid loss is reduced by constructing a system based on the operating parameters of the power system and the electricity price. The operating parameters of the power system include: distribution network voltage, power factor, load power factor, line resistance under grid load conditions considering the energy storage system, and the charging and discharging power of the energy storage system.
[0047] Delaying the investment value of new power grid construction, and constructing based on the investment cost of power grid infrastructure upgrades, the peak load reduction of energy storage systems, and the expected peak demand of energy storage systems;
[0048] To enhance the value of power grid operation reliability, the following parameters are constructed: power grid operation reliability parameters include: outage frequency, loss assessment rate of important users, remaining energy of energy storage systems, power supply reliability, total annual operating hours of important users, and power required for power supply to key users.
[0049] The value of smoothing renewable energy power generation output is constructed based on renewable energy power generation parameters, including: the feed-in tariff of renewable energy, the upper limit power of the node, the lower limit power of the node, the value of the energy storage system exceeding the upper limit power after charging and discharging, the value of the energy storage system exceeding the lower limit power after charging and discharging, the upper and lower fluctuation limits of wind farm output during grid dispatch, the actual active power output of the wind farm, and the charging and discharging power of the energy storage system participating in smoothing the fluctuations in wind farm output.
[0050] To reduce the value of thermal power spinning reserve capacity, the thermal power spinning reserve parameters are constructed based on the parameters of thermal power spinning reserve. These parameters include: the cost of thermal power plant spinning reserve in the absence of energy storage in the power system, the total thermal power replenishment time within one year, and the time period for replenishing power due to wind power output prediction errors.
[0051] In a preferred embodiment, the construction of the reward function includes:
[0052] A reward function is constructed based on the value of the grid-side energy storage system, the present value of the total investment cost of the grid-side energy storage system, and the present value of the total operating cost of the grid-side energy storage system.
[0053] In a preferred embodiment, the present value of the total investment cost of the grid-side energy storage system is constructed based on the investment cost, discount rate, unit cost of energy storage capacity, configured energy storage capacity, unit cost of energy storage converter, and rated power of energy storage converter;
[0054] The present value of the total operating cost of the grid-side energy storage system is constructed based on the discount rate, generator active power output, generator operating time, and generator economic operating parameters.
[0055] For step S2, the current state vector s of the energy storage system obtained in step S1 is... t ={C t ,P L ,P w ,e t The input is provided to the pre-defined grid-side energy storage optimization scheduling model. This model is based on reinforcement learning, receiving the current state space of the energy storage system as input and generating the corresponding action space through its internal policy network and reward function. This operational space includes energy storage and charging. Energy storage and discharge Energy storage dispatch capacity and generator output P Gi,yIn this context, the policy network refers to the network used to generate actions in the reinforcement learning model, predicting the optimal action selection based on the current state; while the reward function is the function used in reinforcement learning to evaluate the quality of actions, assessing the effectiveness of the grid-side energy storage optimization scheduling model after taking a certain action in the environment. In this embodiment, the reward function is constructed based on the value of the grid-side energy storage system.
[0056] Specifically, in this embodiment, the reward function of the grid-side energy storage optimization scheduling model is constructed as follows:
[0057] First, determine the value of reducing grid losses: K1 = K′1 - K″1
[0058] Where K1 represents the value of reducing grid losses, K'1 represents the line loss cost without an energy storage system, K"1 represents the line loss cost with an energy storage system, U represents voltage, a represents power factor, and cos a represents load power factor. e represents the line resistance at node i within time period t, considering both the energy storage system and the grid load. t This represents the electricity price within the time period t. This represents the charging power of the energy storage system at node i during time period t. T represents the discharge power of the energy storage system at node i during time period t. H T represents the set of periods with high user demand. L This represents the set of periods with low user demand, and n represents the number of nodes in the grid-side energy storage system.
[0059] Second, determine the investment value of delaying new power grid construction:
[0060] Where K2 represents the value of delaying new power grid construction, C up,i E represents the investment cost of upgrading the power grid infrastructure at node i. reduction,i E represents the peak load reduction provided by the energy storage system at node i. peak,i This indicates the expected peak demand at node i that the energy storage system helps alleviate.
[0061] Third, determine the value of improving power grid operation reliability: K3 = K′3 - K″3
[0062] Where K3 represents the value of improving grid operation reliability, K'3 represents the cost of power shortage before configuring the energy storage system due to insufficient power supply, and K"3 represents the cost of power shortage after configuring the energy storage system due to insufficient power supply. s Indicates the frequency of power outages (times / year). The loss assessment rate for critical users at node i refers to the economic loss caused by each unit of unmet energy demand by the critical user. This represents the remaining energy of the energy storage system at node i during time period t. This indicates that the expected demand from key users was not met due to power grid failures. Indicates when Hours This indicates the reliability of the power supply at node i. This represents the total annual running hours of important users at node i. This represents the power required to ensure the power supply to critical users at node i.
[0063] Fourth, determine the value of smoothing renewable energy generation output:
[0064] Where K4 represents the value of smoothed renewable energy generation output, e f This indicates the feed-in tariff for renewable energy. This represents the power exceeding the upper limit at node i within time period t. This represents the power at node i that exceeds the lower limit during time period t. This represents the power exceeding the upper limit at node i within time period t after the energy storage system has been charged and discharged. This represents the power at node i exceeding the lower limit after the energy storage system has been charged and discharged within time period t. and These represent the upper and lower fluctuation limits of the wind farm output at node i, as determined by the power grid dispatching system within time period t. and These represent the charging / discharging power of the energy storage system at node i during time period t, which participates in smoothing the output fluctuations of the wind farm.
[0065] Fifth, determine the value of reducing the spinning reserve capacity of thermal power plants: K5 = K′5 - K″5
[0066] Where K5 represents the reduced value of thermal power spinning reserve capacity, K'5 represents the cost of thermal power plant spinning reserve without energy storage system, K"5 represents the cost of thermal power plant spinning reserve with energy storage system, and C CR This represents the cost of a thermal power plant's spinning reserve in the absence of energy storage in the power system. h represents the total thermal power replenishment time within a year, and t1 and t2 represent the time periods when additional power is needed due to wind power output forecast errors.
[0067] Next, based on the value of reducing grid losses, delaying new grid construction investment, improving grid operation reliability, smoothing renewable energy generation output, and reducing thermal power spinning reserve capacity, the value of grid-side energy storage systems is constructed as follows: ∑K=K1+K2+K3+K4+K5
[0068] Among them, K ES The values of the grid-side energy storage system are represented by ρ, T, K1, K2, K3, K4, K5, and K6 respectively. ρ represents the discount rate, T represents the scheduling period of the grid-side energy storage system, K1 represents the value of reducing grid losses, K2 represents the value of delaying new grid construction investment, K3 represents the value of improving grid operation reliability, K4 represents the value of smoothing renewable energy generation output, and K5 represents the value of reducing thermal power spinning reserve capacity.
[0069] Finally, based on the value of the grid-side energy storage system, the present value of the total investment cost of the grid-side energy storage system, and the present value of the total operating cost of the grid-side energy storage system, the reward function for the grid-side energy storage optimal scheduling model is constructed: R = K ES -C inv -C ope
[0070] Where R represents the reward function of the grid-side energy storage optimization scheduling model, and C inv C represents the present value of the total investment cost of the grid-side energy storage system. ope C' represents the present value of the total operating cost of the grid-side energy storage system. y Let ρ represent the investment cost at the beginning of year y, ρ represent the discount rate, and C represent the investment cost at the beginning of year y. B The unit cost representing energy storage capacity. C represents the energy storage capacity configured at node i. P This indicates the unit cost of the energy storage converter. P represents the rated power of the energy storage converter at node i. Gi,y h represents the typical active power output of the generator at node i in year y. i,y Let a represent the operating time of the generator at node i in year y. i b i and c i All represent the economic operating parameters of the generator at node i.
[0071] Next, the training process of the grid-side energy storage optimization scheduling model will be explained in detail:
[0072] As shown in Figure 2, the training process of the grid-side energy storage optimization scheduling model includes the following steps:
[0073] Step S201: For the grid-side energy storage optimization scheduling model to be trained, define the state space, action space and reward function of the model;
[0074] For step S201, for the grid-side energy storage optimization scheduling model to be trained, it is first necessary to define the state space, action space, and reward function. The state space is a set used to describe the current state of the system. In this embodiment, the state space may include remaining capacity, load, renewable energy output, and electricity price s. t ={C t ,P L ,P w ,e t This state information can be obtained through sensors or other monitoring devices and used as input to the model. The action space is a set of actions that the model can select. In this embodiment, the action space may include energy storage charging, energy storage discharging, energy storage scheduling capacity, and generator output. The model needs to select an action from the action space based on the current state to make a scheduling decision. The reward function is used to evaluate the quality of the action selected by the model in each state. In this embodiment, the reward function R of the grid-side energy storage optimization scheduling model is constructed based on the value of the grid-side energy storage system, the present value of the total investment cost of the grid-side energy storage system, and the present value of the total operating cost of the grid-side energy storage system.
[0075] Compared to existing energy storage system planning, which typically focuses on reducing single investment costs, this invention comprehensively considers various factors such as reducing grid losses, delaying new grid construction investment, improving grid operation reliability, smoothing renewable energy output, and reducing thermal power spinning reserve capacity. By combining the role of energy storage systems with the overall benefits of the power system, this invention conducts a system value analysis that more comprehensively considers the economic and stability issues of the power system in long-term operation, thereby maximizing the economic efficiency and optimizing the stability of the power system in long-term operation and improving the overall operating efficiency and reliability of the power system.
[0076] Step S202: Construct the policy network based on the state space, action space, and current policy network parameters;
[0077] For step S202, the state space s defined in step S201 is used. t Action space a t and the current policy network parameters θ Q Construct the following policy network: Q(s) t ,a t |θ Q )
[0078] Where Q(s) t ,a t |θ QThe symbol () represents the policy network, also known as a Q-network. A Q-network is a state-action value function network based on deep reinforcement learning. It learns and optimizes to predict the long-term rewards obtained by taking different actions in a given state. The input to the Q-function is the state and the action, and the output is the corresponding Q-value, representing the long-term reward obtained by taking that action in that state.
[0079] Step S203: Obtain the initialized state space and the initialized action space;
[0080] For step S203, before training the reinforcement learning model, it is necessary to obtain the initialized state space and the initialized action space.
[0081] Step S204: Input the initialized state space and initialized action space into the network-side energy storage optimization scheduling model to be trained for iterative updates, so that the network-side energy storage optimization scheduling model to be trained updates the policy network through the DDPG algorithm until the loss function converges, and obtain the preset network-side energy storage optimization scheduling model.
[0082] For step S204, the initialized state space and initialized action space obtained in step S203 are input into the network-side energy storage optimization scheduling model to be trained. A policy network and a value network are initialized using the Deep Deterministic Policy Gradient (DDPG) reinforcement learning algorithm. The policy network is used to select actions, while the value network is used to evaluate the value of actions. In each training iteration, an initial state is selected from the state space, and the policy network is used to predict the next state s. t+1 The value of an action is evaluated using a value network. The Bellman equation represents the expected reward of performing an action in the current state as the immediate reward plus a discounted reward from a future state. Specifically, the Bellman equation can be expressed as: y t =r t +γQ(s t+1 ,π(s t+1 |θ π )|θ Q )
[0083] Among them, y t Indicates the current state s t Next, execute action a t The expected reward obtained, r t This represents the current immediate reward, based on the environmental state s. t Q(s) is calculated using the reward function R. t+1 ,π(s t+1 |θ π )|θ Q) represents the Q-value function of the state-action sequence at the next moment, γ represents the discount factor used to indicate the importance of future rewards, and s t+1 Indicates the execution of action a t The next state after that.
[0084] The parameters of the Q-network are updated by minimizing the loss value (squared error loss). The loss function for updating the Q-network is:
[0085] Here, E represents the expected value operation, which is used to estimate the long-term expected return of performing a series of actions under a certain strategy.
[0086] The Q-network is updated using policy gradient updates:
[0087] in, Represents the policy gradient. This represents the gradient of the Q-network. This represents the gradient of the policy network. Here, the policy gradient update is based on the value of the Q network to adjust the policy network so that the action selected by the policy can maximize the Q value.
[0088] '').
[0089] Finally, the target Q-network Q'(s,a|θ) is updated to obtain the desired network. Q ) and the target policy network π'(s|θ) π ).
[0090] Step S3: Optimize the scheduling of grid-side energy storage based on the corresponding action space.
[0091] For step S3, based on the action space output by the grid-side energy storage optimization scheduling model in step S2, including the capacity of energy storage charging, energy storage discharging, energy storage scheduling, and generator output, the charging, discharging, and scheduling behavior of the energy storage system can be controlled, thereby achieving optimized scheduling of grid-side energy storage.
[0092] This invention optimizes the capacity planning strategy of energy storage systems through reinforcement learning algorithms, balances the economy and lifespan of energy storage systems, and ensures that energy storage systems continuously improve the operational stability and economic benefits of the power grid in dynamic electricity markets and new energy fluctuation environments.
[0093] As shown in Figure 3, based on the above method embodiments, corresponding device embodiments are provided;
[0094] An embodiment of the present invention provides a grid-side energy storage optimization scheduling device based on reinforcement learning, comprising: an energy storage status acquisition module, an action strategy output module, and an optimization scheduling module;
[0095] The energy storage status acquisition module is used to acquire the current state space of the energy storage system; the state space includes: remaining capacity, load, renewable energy output, and electricity price;
[0096] The action strategy output module is used to input the current state space of the energy storage system into the preset grid-side energy storage optimization scheduling model, so that the grid-side energy storage optimization scheduling model outputs the corresponding action space according to the current state space of the energy storage system. The action space includes: energy storage charging, energy storage discharging, energy storage scheduling capacity and generator output. Among them, when training the grid-side energy storage optimization scheduling model, a reward function is constructed based on the value of the grid-side energy storage system, and then the loss function of the grid-side energy storage optimization scheduling model is updated according to the reward function until the loss function converges.
[0097] The optimization scheduling module is used to optimize the scheduling of grid-side energy storage based on the corresponding action space.
[0098] In a preferred embodiment, the action strategy output module includes: a grid-side energy storage system value construction submodule;
[0099] The grid-side energy storage system value construction submodule is used to construct the value of the grid-side energy storage system based on the value of reducing grid losses, delaying new grid construction investment, improving grid operation reliability, smoothing renewable energy generation output, and reducing thermal power spinning reserve capacity.
[0100] It is understood that the above-described device embodiments correspond to the method embodiments of the present invention, and can implement the reinforcement learning-based grid-side energy storage optimization scheduling method provided by any of the above-described method embodiments of the present invention.
[0101] It should be noted that the device embodiments described above are merely illustrative, and some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0102] Based on the above embodiments of the grid-side energy storage optimization scheduling method based on reinforcement learning, another embodiment of the present invention provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the grid-side energy storage optimization scheduling method based on reinforcement learning of any embodiment of the present invention.
[0103] For example, in this embodiment, the computer program can be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the terminal device.
[0104] The terminal device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.
[0105] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device via various interfaces and lines.
[0106] Based on the above-described method embodiments, another embodiment is provided: another embodiment of the present invention provides a computer-readable storage medium including a stored computer program, wherein, when the computer program is running, it controls the device where the computer-readable storage medium is located to execute the reinforcement learning-based grid-side energy storage optimization scheduling method described in any of the above-described method embodiments of the present invention.
[0107] The modules / units integrated into the reinforcement learning-based network-side energy storage optimization scheduling device / terminal equipment, if implemented as software functional units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.
[0108] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A grid-side energy storage optimization scheduling method based on reinforcement learning, characterized in that, include: Obtain the current state space of the energy storage system; The state space includes: remaining capacity, load, renewable energy output, and electricity price; The state space of the current energy storage system is input into a preset grid-side energy storage optimization scheduling model, so that the grid-side energy storage optimization scheduling model outputs a corresponding action space based on the current state space of the energy storage system. The action space includes: energy storage charging, energy storage discharging, energy storage scheduling capacity, and generator output. During the training of the grid-side energy storage optimization scheduling model, a reward function is constructed based on the value of the grid-side energy storage system, and then the loss function of the grid-side energy storage optimization scheduling model is updated based on the reward function until the loss function converges. Based on the corresponding action space, the grid-side energy storage is optimized and scheduled.
2. The method of claim 1, wherein, The value creation of the grid-side energy storage system includes: The value of the grid-side energy storage system is constructed based on the value of reducing grid losses, delaying new grid construction investment, improving grid operation reliability, smoothing renewable energy generation output, and reducing thermal power spinning reserve capacity.
3. The method of claim 2, wherein, The value of reducing grid losses is constructed based on the operating parameters of the power system and electricity prices; the operating parameters of the power system include: distribution network voltage, power factor, load power factor, line resistance under grid load conditions considering energy storage system, and charging and discharging power of energy storage system; The proposed value of delaying new grid construction investment is constructed based on the investment cost of upgrading grid infrastructure, the peak load reduction of energy storage systems, and the expected peak demand of energy storage systems. The value of improving power grid operation reliability is constructed based on power grid operation reliability parameters; these parameters include: power outage frequency, loss assessment rate of important users, remaining energy of energy storage systems, power supply reliability, total annual operating hours of important users, and power required for power supply to key users. The value of smoothing renewable energy power generation output is constructed based on renewable energy power generation parameters. These renewable energy power generation parameters include: the on-grid price of renewable energy, the upper limit power of the node, the lower limit power of the node, the value of the energy storage system exceeding the upper limit power after charging and discharging, the value of the energy storage system exceeding the lower limit power after charging and discharging, the upper and lower fluctuation limits of wind farm output during grid dispatch, the actual active power output of the wind farm, and the charging and discharging power of the energy storage system participating in smoothing the fluctuations in wind farm output. The reduction in the value of thermal power spinning reserve capacity is constructed based on thermal power spinning reserve parameters; these parameters include: the cost of thermal power plant spinning reserve in the absence of energy storage in the power system, the total thermal power replenishment time within one year, and the time period for replenishing power due to wind power output prediction errors.
4. The method of claim 3, wherein, The construction of the reward function includes: The reward function is constructed based on the value of the grid-side energy storage system, the present value of the total investment cost of the grid-side energy storage system, and the present value of the total operating cost of the grid-side energy storage system.
5. The method of claim 4, wherein, The present value of the total investment cost of the grid-side energy storage system is constructed based on the investment cost, discount rate, unit cost of energy storage capacity, configured energy storage capacity, unit cost of energy storage converter, and rated power of energy storage converter; The present value of the total operating cost of the grid-side energy storage system is constructed based on the discount rate, generator active power output, generator operating time, and generator economic operating parameters.
6. The method of claim 5, wherein, The preset grid-side energy storage optimization scheduling model is trained in the following way: For the grid-side energy storage optimization scheduling model to be trained, the state space, action space and reward function of the model are defined; Construct the policy network based on the state space, action space, and current policy network parameters; Obtain the initialized state space and initialized action space; The initialized state space and initialized action space are input into the grid-side energy storage optimization scheduling model to be trained for iterative updates, so that the grid-side energy storage optimization scheduling model to be trained updates the policy network through the DDPG algorithm until the loss function converges, thus obtaining the preset grid-side energy storage optimization scheduling model.
7. A grid-side energy storage optimal scheduling device based on reinforcement learning, characterized in that, include: Energy storage status acquisition module, action strategy output module, and optimization scheduling module; The energy storage status acquisition module is used to acquire the current state space of the energy storage system; The state space includes: remaining capacity, load, renewable energy output, and electricity price; The action strategy output module is used to input the state space of the current energy storage system into a preset grid-side energy storage optimization scheduling model, so that the grid-side energy storage optimization scheduling model outputs a corresponding action space based on the current state space of the energy storage system. The action space includes: energy storage charging, energy storage discharging, energy storage scheduling capacity, and generator output. During the training of the grid-side energy storage optimization scheduling model, a reward function is constructed based on the value of the grid-side energy storage system, and then the loss function of the grid-side energy storage optimization scheduling model is updated based on the reward function until the loss function converges. The optimization scheduling module is used to optimize the scheduling of grid-side energy storage according to the corresponding action space.
8. The grid-side energy storage optimal scheduling apparatus based on reinforcement learning according to claim 7, characterized in that, The action strategy output module includes: a grid-side energy storage system value construction sub-module; The grid-side energy storage system value construction submodule is used to construct the value of the grid-side energy storage system based on the value of reducing grid losses, delaying grid construction investment, improving grid operation reliability, smoothing renewable energy power generation output, and reducing thermal power spinning reserve capacity.
9. A terminal device, comprising: include: The processor, the memory, and the computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements the reinforcement learning-based grid-side energy storage optimization scheduling method as described in any one of claims 1-6.
10. A computer-readable storage medium, characterized in that, include: A stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the grid-side energy storage optimization scheduling method based on reinforcement learning as described in any one of claims 1-6.