Power system source-grid-load-storage joint regulation method and device based on deep reinforcement learning

CN116247742BActive Publication Date: 2026-08-28GUANGZHOU POWER SUPPLY BUREAU GUANGDONG POWER GRID CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310233509.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-13
Publication Date
2026-08-28
Estimated Expiration
2043-03-13

AI Technical Summary

Technical Problem

但强化学习模型训练困难,面对电力系统高维状态与动作空间,极容易训练失败,无法给出有效的决策

Benefits of technology

[0054] This invention provides a power system source-grid-load-storage scheduling decision-making method by leveraging the superior decision-making capabilities of traditional reinforcement learning in uncertain scenarios. To improve the training efficiency of the reinforcement learning agent, a preliminary source-grid-storage scheduling scheme is given based on predicted information, serving as the foundation for the agent's regulation. A reinforcement learning architecture is designed based on this scheme, and the agent is trained through deep reinforcement learning. The agent modifies the source-grid-storage scheduling scheme based on observable environmental values ​​to eliminate inaccurate predictions and power imbalances caused by unexpected accidents during actual operation. Furthermore, the reward function in the reinforcement learning architecture, by defining line power flow margin and energy storage charging/discharging rewards, guides the agent to reserve sufficient margin and reserves for uncertain scenarios, thereby improving the safety of system operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116247742B_ABST
    Figure CN116247742B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on deep reinforcement learning's power system source network load storage joint regulation method, device, computer equipment and storage medium, the method includes: according to power system prediction information, with economy as target, preliminary generation source load storage scheduling scheme, the source load storage scheduling scheme as the basis of intelligent agent regulation;Design source network load storage joint scheduling reinforcement learning architecture, through deep reinforcement learning training intelligent agent, with security as target, learn to correct source load storage scheduling scheme, realize source network load storage joint scheduling;Wherein, the reward function in the reinforcement learning architecture, by defining line flow margin reward and energy storage reward, guide intelligent agent to reserve enough margin and sufficient backup for uncertainty scenario.The application is trained intelligent agent by designing reinforcement learning architecture and using deep reinforcement learning, learns to correct source load storage scheduling scheme, to eliminate inaccurate prediction and power imbalance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power dispatching technology, and in particular to a method, apparatus, computer equipment, and storage medium for the joint regulation of power system sources, grid, load, and storage based on deep reinforcement learning. Background Technology

[0002] With the development of new power systems, the proportion of renewable energy on the source side is increasing. However, the randomness, intermittency, and anti-peak-shaving characteristics of renewable energy make power supply reliability lower compared to conventional power sources. Meanwhile, peak output surges on the load side, and unexpected power fluctuations leading to power imbalances place higher demands on the system's regulation capabilities. Energy storage, as a flexible load with bidirectional regulation capabilities, provides the power grid with more dispatchable resources. How to jointly dispatch power generation, grid, load, and storage, reserving sufficient margins and reserves for uncertain scenarios to improve system economy and security, is a current hot topic.

[0003] Currently, the joint scheduling of power generation, grid, load and storage mainly relies on establishing mathematical models and then using traditional optimization algorithms to find the optimal solution. When faced with the randomness of power generation and load, stochastic programming or robust optimization are often used. It is difficult to achieve a proper balance between economy and robustness, and real-time decision-making faces the problem of low solution efficiency.

[0004] Deep reinforcement learning algorithms combine the excellent representational capabilities of deep learning with the excellent decision-making capabilities of reinforcement learning. They can effectively address sequential decision-making problems in complex, nonlinear, and uncertain scenarios, naturally meeting the regulation requirements of new power systems. However, reinforcement learning models are difficult to train, and are prone to training failure when faced with the high-dimensional state and action space of power systems, failing to provide effective decisions. Summary of the Invention

[0005] To address the shortcomings of the existing technologies, this invention provides a method, device, computer equipment, and storage medium for the joint regulation of power system sources, grid, load, and energy storage based on deep reinforcement learning. The method initially proposes a source-load-energy storage scheduling scheme based on predicted information, serving as the foundation for intelligent agent regulation. A reinforcement learning architecture is designed based on this scheme. The intelligent agent learns to correct unit output, wind and solar curtailment, and energy storage charging and discharging based on observable environmental values, thereby eliminating inaccurate predictions and power imbalances caused by unexpected accidents during actual operation.

[0006] The first objective of this invention is to provide a method for joint regulation of power system sources, grid, load and storage based on deep reinforcement learning.

[0007] The second objective of this invention is to provide a power system source-grid-load-storage joint control device based on deep reinforcement learning.

[0008] A third objective of this invention is to provide a computer device.

[0009] A fourth objective of this invention is to provide a storage medium.

[0010] The first objective of this invention can be achieved by adopting the following technical solution:

[0011] A power system source-grid-load-storage joint control method based on deep reinforcement learning, the method comprising:

[0012] Based on power system forecast information, and with economic efficiency as the objective, a preliminary source-load-storage scheduling scheme is generated, which serves as the basis for intelligent agent regulation.

[0013] The design employs a reinforcement learning architecture for the joint scheduling of grid, load, and storage. This architecture trains an agent using deep reinforcement learning, aiming to improve the source-load-storage scheduling scheme and achieve joint scheduling of these components. The reward function within this architecture defines line power flow margin rewards and energy storage rewards, guiding the agent to reserve sufficient margin and backup capacity for uncertain scenarios.

[0014] Furthermore, the power flow margin reward for the aforementioned line is:

[0015]

[0016] Where, n L P represents the number of line connections. l For line l power flow, P l,max This represents the maximum power flow value of line l.

[0017] Furthermore, the energy storage reward is:

[0018]

[0019] Where SOC is the current charged capacity of the energy storage. min SOC max These represent the minimum and maximum energy storage capacities, respectively; P b The energy storage charging and discharging power after the intelligent agent makes a decision.

[0020] Furthermore, the reinforcement learning architecture includes:

[0021] Environmental state variables: These are abstract representations of the current state of the power system environment and represent information that an intelligent agent can acquire and needs.

[0022] Operational space: Specifically for adjustable generating units, new energy generating units, and energy storage equipment;

[0023] State transition: The next environment state is determined solely by the current environment state and the action performed.

[0024] Reward function: Used to guide the training of the agent, making it make decisions in the direction of maximizing cumulative reward.

[0025] Furthermore, the environmental state quantity is the power grid state s at time t. t Described as:

[0026] s t =(P disp ,P g ,P w,fc ,P s,fc ,P l,fc ,ΔP up,max ,ΔP down,max ,S l ,SOC,ρ,ρ cur )

[0027] P g (t)=P G (t)+P disp (t)

[0028] ΔP up,max =min(P Gmax -P G (t+1),P G,rampup )

[0029] ΔP down,max =min(P G (t+1)-P Gmin ,P G,rampdown )

[0030]

[0031] Among them, P disp (t) represents the rescheduled value at time t, and is the cumulative rescheduled action of the agent unit; P G (t) represents the output of the adjustable generator at time t in the day-ahead source-load-storage dispatch scheme; P g (t) represents the unit output at time t in the final scheme of superimposed source-load-storage scheduling and agent actions; P w,fc ,P s,fc ,P l,fc These are the predicted values ​​for wind power, solar power, and load, respectively; ΔP up,max ,ΔP down,max These represent the maximum upward and downward adjustment ranges of the unit, respectively; S l The line status is represented by SOC, which is the current charge capacity of the energy storage. P represents the line status. Gmax P Gmin These represent the upper and lower limits of the unit's output, respectively; P G,rampup P G,rampdown These represent the upper and lower limits of the unit's ramp rate, respectively; ρ is the line load rate; PL For line power flow; P L,max ρ represents the maximum power flow value of the line. cur This represents the reduction ratio of new energy sources and the accumulated value of intelligent agent actions.

[0032] Furthermore, for some states and actions of the agent, the following state transition relationships exist:

[0033]

[0034]

[0035] P W =ρ cur *P W,max

[0036] P S =ρ cur *P S,max

[0037] P b =P B +ΔP B

[0038] Among them, P W P S P b The actual output of wind power, photovoltaic power, and energy storage after the decision was made.

[0039] Furthermore, the source-load-storage scheduling scheme is solved day-ahead. During intraday real-time scheduling, the trained agent quickly outputs action values ​​based on the environmental state quantities provided at each time step, including the unit output correction, wind power and photovoltaic reduction ratios, and energy storage charging and discharging power, thereby realizing the joint regulation of source, grid, load, and storage.

[0040] Furthermore, the environment is processed as follows during agent training:

[0041] If the line load rate is outside the set range, it is considered a soft overload and will disconnect after m time steps; where m is the set value.

[0042] If the line load rate exceeds the maximum value of the set range, it is considered a hard overload and the line is immediately disconnected.

[0043] The line will automatically reconnect after n time steps after a disconnection; where n is a set value and n>m;

[0044] Set the probability of a line disconnection occurring at each time step;

[0045] The time step mentioned here differs from the time step in the source-load-storage scheduling scheme: the time interval / granularity is smaller. The second objective of this invention can be achieved by adopting the following technical solution:

[0046] A power system source-grid-load-storage joint control device based on deep reinforcement learning, the device comprising:

[0047] The scheduling scheme generation module is used to generate a preliminary source-load-storage scheduling scheme based on power system forecast information and with economic efficiency as the objective. The source-load-storage scheduling scheme serves as the basis for intelligent agent regulation.

[0048] The scheduling scheme correction module is used to design the reinforcement learning architecture for the source-grid-load-storage joint scheduling. It trains the agent through deep reinforcement learning, with security as the goal, and learns to correct the source-grid-load-storage scheduling scheme to achieve joint scheduling of source, grid, load, and storage. The reward function in the reinforcement learning architecture defines the line power flow margin reward and energy storage reward to guide the agent to reserve sufficient margin and backup for uncertain scenarios.

[0049] The third objective of this invention can be achieved by adopting the following technical solution:

[0050] A computer device includes a processor and a memory for storing processor-executable programs, wherein when the processor executes the program stored in the memory, it implements the above-described power system source-grid-load-storage joint control method.

[0051] The fourth objective of this invention can be achieved by adopting the following technical solution:

[0052] A storage medium storing a program, which, when executed by a processor, implements the aforementioned power system source-grid-load-storage joint control method.

[0053] The present invention has the following advantages over the prior art:

[0054] This invention provides a power system source-grid-load-storage scheduling decision-making method by leveraging the superior decision-making capabilities of traditional reinforcement learning in uncertain scenarios. To improve the training efficiency of the reinforcement learning agent, a preliminary source-grid-storage scheduling scheme is given based on predicted information, serving as the foundation for the agent's regulation. A reinforcement learning architecture is designed based on this scheme, and the agent is trained through deep reinforcement learning. The agent modifies the source-grid-storage scheduling scheme based on observable environmental values ​​to eliminate inaccurate predictions and power imbalances caused by unexpected accidents during actual operation. Furthermore, the reward function in the reinforcement learning architecture, by defining line power flow margin and energy storage charging / discharging rewards, guides the agent to reserve sufficient margin and reserves for uncertain scenarios, thereby improving the safety of system operation. Attached Figure Description

[0055] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0056] Figure 1 This is a flowchart of the power system source-grid-load-storage joint control method based on deep reinforcement learning, according to Embodiment 1 of the present invention.

[0057] Figure 2 This is a schematic diagram of the power system source-grid-load-storage joint control method based on deep reinforcement learning according to Embodiment 1 of the present invention.

[0058] Figure 3 This is a schematic diagram of the source-load-storage scheduling scheme of Embodiment 1 of the present invention.

[0059] Figure 4 This is a schematic diagram of the intelligent agent control scheme of Embodiment 1 of the present invention.

[0060] Figure 5 This is a schematic diagram of the energy storage charging and discharging scheme of Embodiment 1 of the present invention.

[0061] Figure 6 This is a structural block diagram of the power system source-grid-load-storage joint control device based on deep reinforcement learning according to Embodiment 2 of the present invention.

[0062] Figure 7 This is a structural block diagram of the computer device according to Embodiment 3 of the present invention. Detailed Implementation

[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. It should be understood that the specific embodiments described are merely used to explain this application and are not intended to limit this application.

[0064] Example 1:

[0065] like Figure 1 , 2 As shown in the figure, the power system source-grid-load-storage joint control method based on deep reinforcement learning provided in this embodiment includes the following steps:

[0066] S101. Based on the forecast information, a preliminary source-load-storage scheduling scheme is generated as the basic scheme.

[0067] This embodiment generates a preliminary source-load-storage scheduling scheme based on predicted information through mathematical modeling. The source-load-storage scheduling scheme serves as the basis for subsequent revisions.

[0068] Based on grid information, new energy and load forecast information, and the unit start-up and shutdown schedule, the power system arranges unit output and energy storage charging and discharging schemes with economic efficiency as the main objective and at 1-hour intervals.

[0069] (1) Control objectives.

[0070] The control objective is:

[0071]

[0072] Among them, C G (t) represents the total cost of the adjustable generator set at time t; C B (t) represents the total scheduling cost of energy storage at time t.

[0073] (1-1) Total cost of adjustable generator:

[0074] C G (t)=aP G (t) 2 +bP G (t)+c

[0075] Among them, P G (t) represents the output of the adjustable generator at time t, and a, b, and c are the generator's comprehensive cost coefficients, respectively.

[0076] (1-2) Total cost of energy storage dispatch.

[0077] This embodiment uses a battery-based energy storage system as an example for cost analysis. The total cost includes maintenance and depreciation expenses, and its overall cost has a quadratic relationship with the charging and discharging power.

[0078] C B (t)=ω B P B (t) 2

[0079] Where: P B (t) represents the energy storage charging and discharging power. When its value is positive, the energy storage is in a charging state and can be considered as a load; otherwise, it is in a discharging state and can be considered as a power generation device. B Cost coefficient for energy storage output.

[0080] (2) Constraints.

[0081] (2-1) Active power balance constraint:

[0082] P w (t)+P s (t)+P G (t)=P l (t)+P B (t)

[0083] Where: P w (t), P s (t), P l (t) is divided into wind power, photovoltaic power and load forecast power.

[0084] (2-2) Adjustable unit output and ramp-up constraints:

[0085] P Gmin <P G (t)<P Gmax

[0086] P G,rampdown <P G (t)-P G (t-1)<P G,rampup

[0087] Where: P Gmax P Gmin These represent the upper and lower limits of the unit's output, P G,rampup P G,rampdown These are the upper and lower limits of the unit's ramp rate, respectively.

[0088] (2-3) Energy storage charge and discharge limits:

[0089] P B,pro <P B (t)<P B,ab

[0090] Among them, P B,pro For the maximum discharge power of energy storage, P B,ab This represents the maximum charging power for energy storage.

[0091] (2-4) Constraints on energy storage operation:

[0092]

[0093]

[0094] Where: SOC(t) is the charge capacity of the energy storage, E(t) is the energy value stored at time t, and E max This represents the maximum energy value of the energy storage. The state value of the energy storage charge capacity is related to the energy storage output at the previous moment; η ab η proThese represent the charging and discharging efficiencies of the energy storage. During energy storage operation, due to capacity limitations, the State of Charge (SOC) is subject to upper and lower limits as shown in the following equations. Furthermore, to ensure continuity in subsequent optimizations, the energy storage needs to maintain its initial value at the end of the optimization cycle, which imposes the following constraints:

[0095] SOC min ≤SOC(t)≤SOC max

[0096] Soc(0)=Soc(T)

[0097] Among them, SOC min SOC max These represent the minimum and maximum energy storage capacities, respectively; T represents the optimized operating cycle, which in this embodiment refers to the 24th hour of the day.

[0098] (2-5) Rotate spare constraint.

[0099] To ensure the system has sufficient reserve capacity during the day to cope with uncertainties, the spinning reserve constraint must be met during operation:

[0100]

[0101] Where ε is the spinning reserve rate.

[0102] This embodiment aims to achieve power balance with economic efficiency as the goal, and uses mathematical modeling to generate a scheduling scheme with a 1-hour time step as the base value for each scheduling unit. The time interval in this embodiment is not limited to 1 hour.

[0103] In this embodiment, the modeling method is a basic scheme for unit output and energy storage charging and discharging, employing a deterministic modeling approach to provide a benchmark only for agent training. Using other algorithms, such as stochastic programming or robust optimization, will not affect the design and training of the reinforcement learning framework.

[0104] S102. Design a reinforcement learning architecture for joint scheduling of source, grid, load and storage.

[0105] By using reinforcement learning algorithms, the agent is trained to correct the scheduling scheme generated by mathematical modeling, resulting in a joint source-grid-load-storage scheduling decision scheme with a 5-minute time interval. The scheduling scheme generated in step S101 will be referred to as the basic scheme below.

[0106] Based on the requirement of the five-tuple (S, A, P, R, γ) structure of Markov decision processes in reinforcement learning, power system environment modeling focuses on the setting of state S, action A, reward R, and discount factor γ. In this embodiment, continuous decision-making at 5-minute time intervals within a day is considered, i.e., 288 steps per scenario, and a suitable γ = 0.99 is selected.

[0107] (1) Environmental state quantity.

[0108] Environmental state variables are abstract representations of the current state of the power system environment; they represent information that an intelligent agent can acquire and requires. The power grid state s at time t... t It can be described as:

[0109] s t =(P disp ,P g ,P w,fc ,P s,fc ,P l,fc ,ΔP up,max ,ΔP down,max ,S l ,SOC,ρ,ρ cur )

[0110] P g (t)=P G (t)+P disp (t)

[0111] ΔP up,max =min(P Gmax -P G (t+1),P G,rampup )

[0112] ΔP down,max =min(P G (t+1)-P Gmin ,P G,rampdown )

[0113]

[0114] Among them, P disp It is P disp (t) is an abbreviation for rescheduling at time t, representing the cumulative rescheduling actions of the agent unit, P. g It is P g (t) is an abbreviation representing the unit's output at time t. (Note: the subscript G is the basic scheduling scheme, and the lowercase g is the final scheme that superimposes the basic scheme and the agent's actions); P w,fc ,P s,fc ,P l,fc These are the predicted values ​​for wind power, solar power, and load, respectively, and ΔP. up,max ,ΔP down,max These represent the maximum upward and downward adjustment ranges of the unit, respectively; S l The line status is represented by a 0-1 variable, where 0 indicates a disconnected line and 1 indicates a connected line. SOC represents the current charged capacity of the energy storage, ρ is the line load factor, and P... L For line power flow, P L,maxρ represents the maximum power flow value of the line. cur This represents the reduction ratio of new energy sources and the accumulated value of intelligent agent actions.

[0115] (2) Action space.

[0116] Actions are control variables for the current time step. In this embodiment, the action space is defined in three categories, which are respectively for adjustable generator units, new energy generator units, and energy storage devices.

[0117] a t =[ΔP disp ,Δρ cur ,ΔP B ]

[0118] Where: ΔP disp For the rescheduled power of the unit, Δρ cur For renewable energy units, including wind and solar power, the reduction ratio, ΔP B It contributes to the redistribution of energy storage.

[0119] (3) State transition.

[0120] As defined by the Markov decision process, the next environmental state is determined solely by the current environmental state and the action performed:

[0121]

[0122] Let P be the state transition function of the environment, representing the state transition in state s. t =s take action a t =a After the state transitions to the next state s t+1 = s′ probability. Due to the various uncertainties and high nonlinearity in the power system, the environment in which the reinforcement learning agent interacts with the power system, i.e., the state transition process, is composed of a power flow simulator. The power flow simulator calculates the power flow of the grid based on the environmental state and the adjustment amount given by the agent, and outputs line power, line current, unit output, etc., while simultaneously providing feedback rewards.

[0123] For some states and actions of an agent, the following state transition relationships exist:

[0124]

[0125]

[0126] P W =ρ cur *P W,max

[0127] P S =ρ cur *P S,max

[0128] P b =P B +ΔP B

[0129] Among them, P W P S P b The actual output of wind power, photovoltaic power, and energy storage after the decision was made.

[0130] (4) Reward function.

[0131] Reward r t This is used to guide the training of the agent, causing it to make decisions in a direction that maximizes cumulative rewards. The reward function is designed to meet the goal of intraday safety.

[0132] (4-1) For basic power supply and power consumption balance, set a positive reward of +1 for each time step of survival;

[0133] (4-2) Considering the system line margin (considering only the connecting lines), set a line load rate bonus with a value range of [0,1]. The smaller the line load rate, the larger the bonus.

[0134]

[0135] Where, n L P represents the number of connecting lines. l For line l power flow, P l,max This represents the maximum power flow value of line l.

[0136] (4-3) To encourage energy storage to have sufficient reserves, i.e., to discharge when the charge capacity is large and charge when the charge capacity is small, an energy storage reward is set:

[0137]

[0138] The reward obtained in a single time step is the sum of the three rewards mentioned above.

[0139] In this embodiment, the reward function defines the line power flow margin and the energy storage charging and discharging reward to guide the agent to reserve sufficient margin and backup for uncertain scenarios, thereby improving the safety of system operation.

[0140] S103, Basic scheme for regulation using intraday intelligent agents.

[0141] (1) Agent training.

[0142] Intelligent agents need to make effective decisions in various uncertain environments. To ensure that the strategies generated by the agent are robust to uncertain environments and thus guarantee the safety of power system operation, the environment is processed as follows in this specific embodiment:

[0143] (1-1) When the line load rate is greater than 1 and less than 1.35, it is considered a soft overload and can be disconnected after 4 time steps (i.e. 20 minutes).

[0144] (1-2) When the line load rate is greater than 1.35, it is considered a hard overload and the line should be disconnected immediately.

[0145] (1-3) The line will automatically reconnect after 16 time steps after the line is disconnected;

[0146] (1-4) There is a 1% probability of line disconnection at each time step;

[0147] Each time step is 5 minutes long, for a total of 288 steps, with one day constituting one training round. Following the reinforcement learning algorithm and training process, multiple training rounds are sufficient to complete the training of the agent.

[0148] (2) Agent-assisted regulation process.

[0149] The basic solution can be solved day-ahead. During real-time scheduling within the day, the trained agent can quickly output action values ​​based on the environmental state quantities provided at each time step, including the unit output correction, wind power / solar power reduction ratio, and energy storage charging and discharging power, thereby realizing the joint regulation of source, grid, load and storage.

[0150] This embodiment trains an agent using deep reinforcement learning to correct scheduling decisions with safety as the goal.

[0151] In this embodiment, the test system is an improved IEEE 14-node system, including renewable energy units such as wind power and photovoltaic power generation, and an energy storage system with a maximum charging and discharging power of 15MW. The basic scheduling scheme is solved using the Gurobi solver, and the SAC algorithm is selected for training the agent in reinforcement learning. One scenario is selected, such as... Figure 2 As shown, a control scheme with a 1-hour resolution is obtained according to step S101, and a trained agent is obtained based on the framework designed in step S102. This agent can continuously correct the unit output and energy storage charging / discharging based on the basic scheme. Figure 3 , 4 As shown, it is feasible to use reinforcement learning methods for active power correction in the joint dispatching of power system sources, grid, load and storage, and it can quickly correct the basic scheme in real time at 5-minute time step intervals.

[0152] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware, and the corresponding program can be stored in a computer-readable storage medium.

[0153] It should be noted that although the method operations of the above embodiments are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the order of execution of the described steps may be changed. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0154] Example 2:

[0155] like Figure 6 As shown, this embodiment provides a power system source-grid-load-storage joint control device based on deep reinforcement learning. The device includes a scheduling scheme generation module 601 and a scheduling scheme correction module 602, wherein:

[0156] The scheduling scheme generation module 601 is used to generate a preliminary source-load-storage scheduling scheme based on power system forecast information and with economic efficiency as the objective. The source-load-storage scheduling scheme serves as the basis for intelligent agent regulation.

[0157] The scheduling scheme correction module 602 is used to design the reinforcement learning architecture for the source-grid-load-storage joint scheduling. It trains the agent through deep reinforcement learning, with security as the goal, and learns to correct the source-grid-load-storage scheduling scheme to achieve joint scheduling of source, grid, load, and storage. The reward function in the reinforcement learning architecture defines the line power flow margin reward and energy storage reward to guide the agent to reserve sufficient margin and backup for uncertain scenarios.

[0158] The specific implementation of each module in this embodiment can be found in Embodiment 1 above, and will not be repeated here. It should be noted that the device provided in this embodiment is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure can be divided into different functional modules to complete all or part of the functions described above.

[0159] Example 3:

[0160] This embodiment provides a computer device, which can be a computer, such as... Figure 7As shown, the system is connected via a system bus 701 to a processor 702, a memory, an input device 703, a display 704, and a network interface 705. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium 706 and internal memory 707. The non-volatile storage medium 706 stores the operating system, computer programs, and a database. The internal memory 707 provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. When the processor 702 executes the computer programs stored in the memory, it implements the power system source-grid-load-storage joint control method of Embodiment 1, as follows:

[0161] Based on power system forecast information, and with economic efficiency as the objective, a preliminary source-load-storage scheduling scheme is generated, which serves as the basis for intelligent agent regulation.

[0162] The design employs a reinforcement learning architecture for the joint scheduling of grid, load, and storage. This architecture trains an agent using deep reinforcement learning, aiming to improve the source-load-storage scheduling scheme and achieve joint scheduling of these components. The reward function within this architecture defines line power flow margin rewards and energy storage rewards, guiding the agent to reserve sufficient margin and backup capacity for uncertain scenarios.

[0163] Example 4:

[0164] This embodiment provides a storage medium, which is a computer-readable storage medium, storing a computer program. When the computer program is executed by a processor, it implements the power system source-grid-load-storage joint control method of Embodiment 1 above, as follows:

[0165] Based on power system forecast information, and with economic efficiency as the objective, a preliminary source-load-storage scheduling scheme is generated, which serves as the basis for intelligent agent regulation.

[0166] The design employs a reinforcement learning architecture for the joint scheduling of grid, load, and storage. This architecture trains an agent using deep reinforcement learning, aiming to improve the source-load-storage scheduling scheme and achieve joint scheduling of these components. The reward function within this architecture defines line power flow margin rewards and energy storage rewards, guiding the agent to reserve sufficient margin and backup capacity for uncertain scenarios.

[0167] It should be noted that the computer-readable storage medium in this embodiment can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0168] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope disclosed in the present invention, based on the technical solution and inventive concept of the present invention, shall fall within the scope of protection of the present invention.

Claims

1. A power system source-grid-load-storage joint control method based on deep reinforcement learning, characterized in that, The method includes: Based on power system forecast information, and with economic efficiency as the objective, a preliminary integrated dispatch scheme for power generation, grid, load, and storage is generated; this integrated dispatch scheme serves as the basis for intelligent agent regulation. A reinforcement learning architecture for the joint scheduling of power generation, grid, load, and energy storage is designed. The agent is trained through deep reinforcement learning. With safety as the goal, the agent learns to modify the joint scheduling scheme of power generation, grid, load, and energy storage to achieve joint scheduling of power generation, grid, load, and energy storage. The reward function in the reinforcement learning architecture defines the line power flow margin reward and the energy storage reward to guide the agent to reserve sufficient margin and backup for uncertain scenarios. The reinforcement learning architecture includes: Environmental state variables: These are abstract representations of the current state of the power system environment and represent information that an intelligent agent can acquire and needs. Operational space: Specifically for adjustable generating units, new energy generating units, and energy storage equipment; State transition: The next environment state is determined solely by the current environment state and the action performed. Reward function: Used to guide the training of the agent, making it make decisions in the direction of maximizing cumulative reward; Wherein, the environmental state quantity is the power grid state at time t. Described as: ; In the formula, yes abbreviation, Let be the rescheduled value at time t, and be the cumulative rescheduled actions of the agent unit; The output of the adjustable generator at time t in the current source-grid-load-storage joint dispatch scheme; yes abbreviation, The output of the generating units at time t in the final scheme of superimposing the source-grid-load-storage joint scheduling scheme and the intelligent agent action; These are the predicted values ​​for wind power, solar power, and load, respectively. These refer to the maximum upward and downward adjustment range of the generator unit, respectively. Line status; This represents the current charge capacity of the energy storage. , These are the upper and lower limits of the unit's output, respectively. , These are the upper and lower limits of the unit's ramp rate, respectively. Line load rate; For line power flow; This represents the maximum power flow value of the line. The reduction ratio of new energy sources; the accumulated value of intelligent agent actions; For the agent's partial states and actions, the following state transition relationships exist: ; In the formula, The energy storage charging and discharging power after the intelligent agent makes a decision. For the rescheduled power of the unit, For new energy units, It contributes to the redistribution of energy storage.

2. The power system source-grid-load-storage joint control method according to claim 1, characterized in that, The power flow margin bonus for the aforementioned line is: ; in, This refers to the number of connecting wires in the circuit.

3. The power system source-grid-load-storage joint control method according to claim 1, characterized in that, The energy storage reward is: ; in, For the current charged capacity of energy storage, , These represent the minimum and maximum energy storage capacities, respectively.

4. The power system source-grid-load-storage joint control method according to claim 1, characterized in that, The proposed source-grid-load-storage joint scheduling scheme is solved day-ahead. During intraday real-time scheduling, the trained agent quickly outputs action values ​​based on the environmental state quantities provided at each time step, including the unit output correction, wind and solar power reduction ratios, and energy storage charging and discharging power, thereby achieving joint regulation of source-grid-load-storage.

5. The power system source-grid-load-storage joint control method according to any one of claims 1 to 4, characterized in that, The environment is processed as follows during agent training: If the line load rate is outside the set range, it is considered a soft overload and will disconnect after m time steps; where m is the set value. If the line load rate exceeds the maximum value of the set range, it is considered a hard overload and the line is immediately disconnected. The line will automatically reconnect after n time steps after a disconnection; where n is a set value and n>m; Set the probability of a line disconnection occurring at each time step; The time step is smaller than the time step in the source-grid-load-storage joint scheduling scheme.

6. A power system source-grid-load-storage joint control device based on deep reinforcement learning, characterized in that, The device includes: The dispatch scheme generation module is used to generate a preliminary source-grid-load-storage joint dispatch scheme based on power system forecast information and with economic efficiency as the objective. The source-grid-load-storage joint dispatch scheme serves as the basis for intelligent agent regulation. The scheduling scheme correction module is used to design a reinforcement learning architecture for the joint scheduling of source, grid, load and storage. It trains the agent through deep reinforcement learning and learns to correct the joint scheduling scheme of source, grid, load and storage with safety as the goal, so as to realize the joint scheduling of source, grid, load and storage. The reward function in the reinforcement learning architecture defines the line power flow margin reward and energy storage reward to guide the agent to reserve sufficient margin and sufficient backup for uncertain scenarios. The reinforcement learning architecture includes: Environmental state variables: These are abstract representations of the current state of the power system environment and represent information that an intelligent agent can acquire and needs. Operational space: Specifically for adjustable generating units, new energy generating units, and energy storage equipment; State transition: The next environment state is determined solely by the current environment state and the action performed. Reward function: Used to guide the training of the agent, making it make decisions in the direction of maximizing cumulative reward; Wherein, the environmental state quantity is the power grid state at time t. Described as: ; In the formula, Let be the rescheduled value at time t, and be the cumulative rescheduled actions of the agent unit; The output of the adjustable generator at time t in the current source-grid-load-storage joint dispatch scheme; The output of the generating units at time t in the final scheme of superimposing the source-grid-load-storage joint scheduling scheme and the intelligent agent action; These are the predicted values ​​for wind power, solar power, and load, respectively. These refer to the maximum upward and downward adjustment range of the generator unit, respectively. Line status; This represents the current charge capacity of the energy storage. , These are the upper and lower limits of the unit's output, respectively. , These are the upper and lower limits of the unit's ramp rate, respectively. Line load rate; For line power flow; This represents the maximum power flow value of the line. The reduction ratio of new energy sources; the accumulated value of intelligent agent actions; For the agent's partial states and actions, the following state transition relationships exist: ; In the formula, The energy storage charging and discharging power after the intelligent agent makes a decision. For the rescheduled power of the unit, For new energy units, It contributes to the redistribution of energy storage.

7. A storage medium storing a program, characterized in that, When the program is executed by the processor, it implements the power system source-grid-load-storage joint control method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Power grid real-time scheduling optimization method and system, computer equipment and storage medium

    CN115241885A