Energy storage and other power supply cooperative AGC control method and device based on deep reinforcement learning, electronic equipment and storage medium
Through the AGC control method of energy storage and other power supplies based on deep reinforcement learning, the problem of power fluctuations in the power grid frequency and connection line power after the new energy in the wind and light base is connected to the grid, the optimal normal setting of energy storage SOC and the stability of the grid frequency are achieved, and the control performance standards of the power system are met.
Patent Information
- Application Number
- CN202510147467.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-13
AI Technical Summary
The grid connection of new energy in large-scale scenery and light bases has led to insufficient traditional controllable adjustment resources, the power fluctuations in new energy affect the stability of the power grid frequency, and the current flow of the gate connection line is affected, increasing the difficulty of control and assessment pressure.
The AGC control method for energy storage and other power supplies is adopted based on deep reinforcement learning. By building a load frequency control model of the two-region interconnection system, the deep reinforcement learning algorithm is used to determine the optimal normal setting value of energy storage SOCs in each period of the day, and a model reward function is designed to formulate an adaptive energy storage control strategy.
Real-time adjustment of the regional control deviation allocation ratio of energy storage and other power supplies is achieved, the unit output is optimized, the CPS indicators of the power system are met, the frequency and connection line power deviations are reduced, and the service life of energy storage is extended.
Smart Images

Figure CN119995045A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power system operation control technology, and in particular to a method, device, electronic device and storage medium for coordinated AGC control of energy storage and other power sources based on deep reinforcement learning. Background Art
[0002] Large-scale wind and solar bases are basically far away from the load center and need to transmit electricity through ultra-high voltage transmission channels. Due to the intermittent and unstable nature of wind power and photovoltaics, when large-scale new energy is connected to the grid, traditional controllable regulation resources are seriously insufficient. At the same time, there are problems such as sharp rise and fall of new energy power, frequent injection of impact power, and serious lack of rapid regulation capacity of the power grid, which seriously affect the transmission and absorption of new energy in the base. Frequency is an important stability indicator in the power system. Its deviation from the normal range will cause instability or even collapse of the power system. On the other hand, when the active power-frequency deviation control mode is adopted in the gateway interconnection line, the fluctuation of new energy power will also seriously affect the flow of the interconnection line, which may cause overload problems of the interconnection line between the interconnected areas, increase the difficulty of gateway control, and make it face severe gateway assessment pressure.
[0003] Electrochemical energy storage with bidirectional rapid regulation performance can quickly respond to frequency changes and track power fluctuations, which can greatly alleviate the situation of insufficient thermal power regulation capacity and effectively deal with the impact of rapid changes in frequency characteristics caused by sudden rise and fall in power. It can play an important role in solving the difficulties in power balance caused by insufficient rotational inertia of high-proportion new energy power systems and volatility and intermittency of new energy output. It is of great significance to alleviate the pressure of checkpoint assessment and improve frequency stability. It can effectively control the power of regional power grid interconnection lines and the frequency of the power grid, and reduce the failures and wear of conventional frequency regulation units.
[0004] In the process of joint frequency regulation, there are still problems such as poor energy storage charge state, short life and unstable frequency regulation performance due to incomplete frequency regulation strategy. Conventional AGC strategy does not consider its impact on power grid flow safety constraints. At this time, it will be difficult to maintain system frequency and tie line power within the operating range by relying solely on conventional AGC hysteresis control. In this context, a deep reinforcement learning-based energy storage and other power sources collaborative AGC control method is needed to solve the above problems. Summary of the invention
[0005] The purpose of the present invention is to propose a method, device, electronic device and storage medium for coordinated AGC control of energy storage and other power sources based on deep reinforcement learning, so as to formulate an AGC control strategy for coordinated energy storage and other power sources that meets CPS assessment, cross-section limit control and automatic recovery of energy storage SOC. The method comprises the following steps:
[0006] Build a load frequency control model for a two-region interconnected system with energy storage;
[0007] Based on the load frequency control model of the two-region interconnected system and the actual historical data of wind power fluctuations, a deep reinforcement learning algorithm is used to determine the optimal normal setting value of the energy storage SOC at all times of the day;
[0008] According to the optimal normal setting value of the energy storage SOC in each period of the day, the absolute value of the regional control deviation is distinguished, and the model reward function containing control performance standards, section over-limit control and energy storage SOC recovery to the optimal normal setting value of the corresponding period is designed based on the absolute value. The energy storage adaptive control strategy is formulated according to the model reward function.
[0009] The load frequency control model of the two-region interconnected system includes:
[0010] Energy storage system model
[0011]
[0012] SOC of energy storage system
[0013]
[0014] Among them, T b is the response time constant of the energy storage converter; Δu b It is the control input signal of the energy storage module; ΔP b is the output power of the energy storage system; η d is the discharge efficiency of energy storage, η c is the charging efficiency of energy storage, P bess Δt is the charging and discharging power of the energy storage within the time Δt, and C is the rated capacity of the energy storage system;
[0015] Speed Regulator Model in Thermal Power Unit Module
[0016]
[0017] Among them, ΔP a is the speed regulator power deviation, T a is the speed regulator time constant; Δu a is the speed regulator control input signal, R is the droop coefficient, Δf1 is the frequency deviation in the first area;
[0018] Steam turbine model in the thermal power unit module
[0019]
[0020] Among them, ΔP h is the turbine output power, T t is the turbine time constant;
[0021] Generator-Load Model
[0022]
[0023] Where M is the inertia time constant of the regional unit, D is the load damping coefficient, ΔP h is the turbine output power, ΔP 12 is the tie line power deviation, ΔP L is the wind power fluctuation;
[0024] Tie line power deviation model between areas 1 and 2
[0025]
[0026] Area Control Deviation ACE
[0027] ACE=ΔP 12 +BΔf
[0028] Among them, T 12 is the tie line synchronization coefficient, Δf1 and Δf2 are the frequency deviations of the first area and the second area respectively; B is the frequency deviation factor.
[0029] Based on the load frequency control model of the two-region interconnected system and the actual wind power fluctuation historical data, the deep reinforcement learning algorithm is used to determine the optimal normal setting value of the energy storage SOC at each time period throughout the day. The specific steps include:
[0030] Set the energy storage initial SOC to different values and initialize the parameters of the DQN algorithm;
[0031] The actual wind power fluctuation historical data is imported into the environmental constraints of the two-region load frequency control model with energy storage, the allocation ratio of energy storage to other power sources is initialized to 0, and the DQN algorithm is pre-trained.
[0032] The trained agent is used for online control to calculate the energy storage SOC in real time. If the energy storage SOC exceeds the limit, the ACE allocation ratio of the energy storage is set to 0 to meet the energy storage SOC constraint. The real-time tie line power deviation ΔP during the online control process is recorded. 12 The frequency deviation Δf1 value of the value and area 1 is adjusted until all the set time periods are completed online;
[0033] Statistics of the tie line power deviation ΔP in each period obtained through online control 12 The absolute value of the frequency deviation Δf1 in area 1 is used to determine the optimal normal setting value of the energy storage SOC in each time period.
[0034] The model reward function is defined as follows:
[0035] |ACE|<x:
[0036] R=(SOC-SOC in ) 2
[0037] |ACE|>x:
[0038]
[0039] Where: R is the model reward function, P limit is the cross-sectional limit, x is the dividing point of the absolute value of ACE, SOC in Indicates the optimal normal setting value of energy storage SOC at each time period throughout the day.
[0040] Energy storage and other power sources coordinated AGC control device based on deep reinforcement learning, including:
[0041] Model building module, used to build a load frequency control model of a two-region interconnected system with energy storage;
[0042] The value determination module is used to determine the optimal normal setting value of the energy storage SOC at each time period throughout the day using a deep reinforcement learning algorithm based on the load frequency control model of the two-region interconnected system and the actual wind power fluctuation historical data;
[0043] The function design module is used to distinguish the absolute value of the regional control deviation according to the optimal normal setting value of the energy storage SOC in each period of the day, and to design the model reward function containing the control performance standard, section over-limit control and energy storage SOC recovery to the optimal normal setting value of the corresponding period based on the absolute value, and formulate the energy storage adaptive control strategy according to the model reward function.
[0044] The load frequency control model of the two-region interconnected system in the model building module includes:
[0045] Energy storage system model
[0046]
[0047] SOC of energy storage system
[0048]
[0049] Among them, T b is the response time constant of the energy storage converter; Δu b It is the control input signal of the energy storage module; ΔP b is the output power of the energy storage system; η d is the discharge efficiency of energy storage, η c is the charging efficiency of energy storage, P bess Δt is the charging and discharging power of the energy storage within the time Δt, and C is the rated capacity of the energy storage system;
[0050] Speed Regulator Model in Thermal Power Unit Module
[0051]
[0052] Among them, ΔP a is the speed regulator power deviation, T a is the speed regulator time constant; Δu a is the speed regulator control input signal, R is the droop coefficient, Δf1 is the frequency deviation in the first area;
[0053] Steam turbine model in the thermal power unit module
[0054]
[0055] Among them, ΔP h is the turbine output power, T t is the turbine time constant;
[0056] Generator-Load Model
[0057]
[0058] Where M is the inertia time constant of the regional unit, D is the load damping coefficient, ΔP h is the turbine output power, ΔP 12 is the tie line power deviation, ΔP L is the wind power fluctuation;
[0059] Tie line power deviation model between areas 1 and 2
[0060]
[0061] Area Control Deviation ACE
[0062] ACE=ΔP 12 +BΔf
[0063] Among them, T 12 is the tie line synchronization coefficient, Δf1 and Δf2 are the frequency deviations of the first area and the second area respectively; B is the frequency deviation factor.
[0064] In the value determination module, based on the load frequency control model of the two-region interconnected system and the actual wind power fluctuation historical data, the deep reinforcement learning algorithm is used to determine the optimal normal setting value of the energy storage SOC at each time period throughout the day. The specific steps include:
[0065] Set the energy storage initial SOC to different values and initialize the parameters of the DQN algorithm;
[0066] The actual wind power fluctuation historical data is imported into the environmental constraints of the two-region load frequency control model with energy storage, the allocation ratio of energy storage to other power sources is initialized to 0, and the DQN algorithm is pre-trained.
[0067] The trained agent is used for online control to calculate the energy storage SOC in real time. If the energy storage SOC exceeds the limit, the ACE allocation ratio of the energy storage is set to 0 to meet the energy storage SOC constraint. The real-time tie line power deviation ΔP during the online control process is recorded. 12 The frequency deviation Δf1 value of the value and area 1 is adjusted until all the set time periods are completed online;
[0068] Statistics of the tie line power deviation ΔP in each period obtained through online control 12 The absolute value of the frequency deviation Δf1 in area 1 is used to determine the optimal normal setting value of the energy storage SOC in each time period.
[0069] The model reward function in the function design module is defined as follows:
[0070] |ACE|<x:
[0071] R=(SOC-SOC in ) 2
[0072] |ACE|>x:
[0073]
[0074] Where: R is the model reward function, P limit is the cross-sectional limit, x is the dividing point of the absolute value of ACE, SOC in Indicates the optimal normal setting value of energy storage SOC at each time period throughout the day.
[0075] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, each step of a method for coordinated AGC control of energy storage and other power sources based on deep reinforcement learning is implemented.
[0076] A storage medium stores a computer program, which, when executed by a processor, implements various steps in a method for coordinated AGC control of energy storage and other power sources based on deep reinforcement learning.
[0077] The beneficial effects of the present invention are:
[0078] 1. The method of the present invention can adjust the regional control deviation ACE allocation ratio of energy storage and other power sources in real time, thereby realizing the output adjustment of the unit.
[0079] 2. The model reward function set in the present invention enables the algorithm to meet the CPS index of the power system and continue to participate in AGC control. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] Figure 1 It is a flow chart of the AGC control method of energy storage and other power sources in collaboration based on deep reinforcement learning of the present invention;
[0081] Figure 2 The present invention provides a load frequency control model of a two-area interconnected system based on a deep reinforcement learning-based energy storage and other power supply collaborative AGC control strategy;
[0082] Figure 3 It is a structural block diagram of an AGC control strategy for energy storage and other power sources in coordination based on deep reinforcement learning provided by the present invention;
[0083] Figure 4 is an actual wind power fluctuation curve used in pre-training of the DQN algorithm in the embodiment provided by the present invention;
[0084] Figure 5 is an actual wind power fluctuation curve of a certain day in an embodiment provided by the present invention;
[0085] Figure 6 : is a CPS1 curve of the control effect of different controllers of the embodiment provided by the present invention;
[0086] Figure 7 is a tie line power deviation curve of different controller control effects of the embodiment provided by the present invention;
[0087] Figure 8 is a region-load deviation curve of control effects of different controllers of the embodiment provided by the present invention;
[0088] Fig. 9 is a regional control deviation curve of the control effect of different controllers of the embodiment provided by the present invention;
[0089] Fig.10 It is the energy storage SOC variation curve of different controller control effects of the embodiments provided by the present invention. DETAILED DESCRIPTION
[0090] The present invention proposes a method, device, electronic device and storage medium for coordinated AGC control of energy storage and other power sources based on deep reinforcement learning. The present invention is further described below in conjunction with the accompanying drawings and specific embodiments.
[0091] The English abbreviations in this embodiment are as follows:
[0092] CPS: Control Performance Standard.
[0093] DQN: Deep Q network algorithm, English Deep-Q-Network.
[0094] Figure 1 This is a flow chart of the AGC control method of energy storage and other power sources in collaboration based on deep reinforcement learning of the present invention, which specifically includes:
[0095] S101: Build a load frequency control model for a two-region interconnected system with energy storage, such as Figure 2 As shown, the specific contents include:
[0096] The energy storage system is represented by a first-order inertia link, and its dynamic physical model is:
[0097]
[0098] Among them, T b is the response time constant of the energy storage converter; Δu b It is the control input signal of the energy storage module; ΔP b Output power to the energy storage system.
[0099] The calculation of the SOC of the energy storage system satisfies the following equation:
[0100]
[0101] Among them, η d is the discharge efficiency of energy storage, η c is the charging efficiency of energy storage, P bess Δt is the charging and discharging power of the energy storage within Δt time, and C is the rated capacity of the energy storage system.
[0102] The thermal power unit module includes a governor model and a steam turbine model. The governor power deviation ΔP a The dynamic physical model is:
[0103]
[0104] Among them, T a is the speed regulator time constant; Δu a is the speed regulator control input signal, R is the droop coefficient, and Δf1 is the frequency deviation in the first area.
[0105] Steam turbine output power ΔP h The dynamic physical model is:
[0106]
[0107] Among them, Tt is the turbine time constant.
[0108] The generator-load model Δf is:
[0109]
[0110] Where M is the inertia time constant of the regional unit, D is the load damping coefficient, ΔP h is the turbine output power, ΔP 12 is the tie line power deviation, ΔP L is the wind power fluctuation.
[0111] The power deviation model of the tie line between areas 1 and 2 is:
[0112]
[0113] Among them, T 12 is the tie line synchronization coefficient, Δf1 and Δf2 are the frequency deviations of the first and second areas respectively.
[0114] The interconnected system adopts the tie line power frequency deviation control TBC mode, which requires simultaneous detection of tie line power and system frequency deviation. The calculation formula of regional control deviation ACE is:
[0115] ACE=ΔP 12 +BΔf (7)
[0116] Wherein, B is the frequency deviation factor.
[0117] The DQN algorithm is used as a controller to solve the problem of power allocation between energy storage and other power sources in the interconnected power system. In each AGC control cycle, the ACE allocation ratio of energy storage and other power sources is adjusted in real time according to the ACE signal input in the control area, thereby realizing the output adjustment of the unit. Figure 3 This is a structural block diagram of an AGC control strategy for energy storage and other power sources in collaboration based on deep reinforcement learning in an embodiment of the present invention.
[0118] This method consists of two parts: offline training and online application. The offline training process iteratively updates the agent parameters θ i In each AGC control cycle, the agent generates different control instructions to interact with the load frequency control model of the interconnected system. Through exploration, the agent's parameter θ i The agent parameters θ will be updated based on the regional control deviation and the LFC reward function. i After the update is completed, it enters the online application stage. The intelligent body will be updated based on the observed system status, regional control deviation and the network parameters θ completed by the intelligent body. iGenerate the optimal allocation ratio and send it to energy storage and other power sources.
[0119] According to the above definition, the learning steps of the DQN algorithm are as follows:
[0120] (1) Initialize neural network parameters and experience memory library;
[0121] (2) Observe the system status s t ;
[0122] (3) Calculate action instruction a t ;
[0123] (4) Observe the state s at the next moment t+1 And calculate the reward value r t ;
[0124] (5) Storage t ,a t ,r t ,s t+1 ) into the experience memory bank;
[0125] (6) Uniformly sample n (s) from the experience base t ,a t ,r t ,s t+1 ), update the neural network parameters;
[0126] (7) Let t = t + 1, return to step (2), and loop until the algorithm converges or all training samples are traversed.
[0127] S102: In order to give full play to the flexible adjustment ability of energy storage, based on the actual historical data of wind power fluctuation, the energy storage SOC normal setting value with the best effect of smoothing the interconnection line power fluctuation and wind power fluctuation in each period of the day is obtained. The reward function R1 is designed to minimize the absolute value of the regional control deviation, which is defined as:
[0128] R1=-|ACE|
[0129] Among them, ACE is the regional control deviation.
[0130] The discrete-time Markov decision process can be used to solve the problem of dynamic optimization of power allocation of AGC. A four-dimensional array is used to describe the state space S = (s1, s2, s3, s4) in the Markov decision process. Among them, s1 is the regional control deviation ACE, s2 is the wind power fluctuation value, s3 is the interconnection line power fluctuation value, and s4 is the frequency fluctuation value. The action space A is the energy storage system allocation ratio a1, a1∈[0,1], and the above allocation ratio is discretized into 11 control instructions.
[0131] By setting different energy storage SOC values for operation simulation, the optimal normal setting value of energy storage SOC in different periods is determined based on the DQN algorithm. The specific steps are as follows:
[0132] (1): Set the initial SOC of the energy storage to different values and initialize the parameters of the DQN algorithm;
[0133] (2): The actual wind power fluctuation data of a certain power grid is imported into the load frequency control model of the interconnected system, the allocation ratio of energy storage to other power sources is initialized to 0, and the DQN algorithm is pre-trained;
[0134] (3): The trained agent is used for online control, and the energy storage SOC is calculated in real time according to formula (2). If the energy storage SOC exceeds the limit, the ACE ratio allocated to the energy storage is set to 0 to meet the energy storage SOC constraint. The real-time tie line power deviation ΔP is recorded during the online control process. 12 Value and frequency deviation Δf1 value of region one.
[0135] (4): Determine whether all the set time periods have been controlled online. If so, proceed to step (5). If not, proceed to step (3).
[0136] (5): Statistics of the tie line power deviation ΔP in each period obtained through online control 12 The absolute value of the frequency deviation Δf1 in area 1 is used to determine the optimal normal setting value of the energy storage SOC in each time period.
[0137] S103: Based on the optimal normal setting value of the energy storage SOC in each time period determined by the simulation operation, when the regional control deviation ACE is small, the energy storage SOC is controlled to be restored to the optimal normal value of the corresponding time period, so that the energy storage can better participate in the AGC control.
[0138] For the load frequency control model of the interconnected system between the two regions, the size of ACE is distinguished. When the absolute value of ACE is large, the goal is to meet the CPS assessment index; when the absolute value of ACE is small, the energy storage SOC is restored. The model reward function is defined as:
[0139] |ACE|>x:
[0140]
[0141] |ACE|<x:
[0142] R=(SOC-SOC in ) 2
[0143] Where: P limit is the cross-sectional limit, x is the dividing point of the absolute value of ACE, SOC inIndicates the optimal normal setting value of energy storage SOC in each time period.
[0144] On the basis of step S102, the value of CPS1 and the historical state information of the model at the previous moment are added to the state space, so that the state space becomes a ten-dimensional array S = (s1, s2, s3, s4, s5, s6, s7, s8, s9, s 10 ), where s5 is the CPS1 value of region 1. s6 is the regional control deviation ACE at the previous moment, s7 is the wind power fluctuation value at the previous moment, s8 is the tie line power fluctuation value at the previous moment, s9 is the frequency fluctuation value at the previous moment, and s 10 is the CPS1 value of region 1 at the previous moment. The action space A is still the ACE ratio a1 allocated to the energy storage system, a1∈[-1,1], and the above allocation ratio is also discretized into 11 control instructions.
[0145] Assume that the total installed capacity of the regional power grid is 1000MW, and the benchmark power is 1000MW. The charging and discharging efficiency of the energy storage is η d , η c are all set to 0.9, the upper and lower limits of the energy storage SOC are set to 0.1 and 0.9 respectively, and the rated capacity C of the energy storage system is 15MW.h. limit Set to 0.1 (pu), and the absolute value of ACE is 0.05 (pu). Figure 1 The interconnected system load frequency control model is shown in Figure 1, and the model-related parameters are shown in Table 1. The DQN algorithm applied to the controller is written in Python. The estimated network learning rate, experience memory size and sampling time of the DQN algorithm are set to 0.001, 3000 and 1.5 minutes, respectively.
[0146] Table 1 Nominal parameters of interconnected power systems
[0147]
[0148] In order to determine the optimal normal setting value of the energy storage SOC in each period of the day mentioned in this embodiment, the actual wind power fluctuation data of the Eastern Menggu Power Grid in 2023 is used for pre-training, and the effectiveness of the proposed control method is verified.
[0149] In both area 1 and area 2, the Figure 4The actual wind power fluctuation curve of the Mengdong Power Grid shown in the figure simulates the actual disturbance of the system. The model is pre-trained 8000 times, and the initial SOC values of the energy storage are set to 0.3, 0.4, 0.5 and 0.6 respectively. After the pre-training, the wind power fluctuation data of the Mengdong Power Grid for 15 days in January 2023 is used for online control. The whole day is divided into 24 time periods, with 1.5 minutes as an AGC control cycle. 40 values are sampled in each time period every day, and the interconnection line power deviation ΔP of a total of 600 sampling points in the same time period within 15 days is counted. 12 The probability that the absolute value of is greater than 0.04 (pu) and the absolute value of the frequency deviation Δf1 in region 1 is greater than 0.15 Hz.
[0150] According to the statistical results, the energy storage SOC value with the lowest probability of interconnection line power fluctuation in each period is selected first. If the probability of interconnection line power fluctuation in each period is equal, the energy storage SOC value with the lowest probability of frequency fluctuation in each period is selected to determine the normal setting value of energy storage SOC with the best effect of smoothing wind power fluctuation and interconnection line power fluctuation in each period of the day. The results are shown in Table 2.
[0151] Table 2 Optimal normal setting values of energy storage SOC at different time periods throughout the day
[0152]
[0153]
[0154] To verify the effectiveness of the proposed control method, the Mengdong Power Grid was used in 2023. Figure 4 The actual wind power fluctuation data shown in Figure 1 is used for pre-learning to update the intelligent agent parameters. After the pre-learning is completed, the actual wind power fluctuation curve of the Mengdong power grid on a certain day used in both area 1 and area 2 is used as the actual wind power disturbance of the system to simulate the system. The fluctuation curve is as follows: Figure 5 shown.
[0155] In order to compare and analyze the effectiveness of the control method proposed in this embodiment, the PI control method is used and the frequency regulation demand is divided into the high-frequency part and the low-frequency part, and the energy storage system is formulated to participate in the AGC control strategy (hereinafter referred to as the "reference method") to compare the results, as shown in Table 3.
[0156] Table 3 Simulation results under disturbance
[0157]
[0158] For the proposed DQN algorithm, according to Figure 6It can be seen that the CPS1 index of area 1 of PI control is less than 100%, indicating that the power allocation control is out of range. The CPS1 index of the literature method is greater than or equal to 100%, but CPS1 directly meets the CPS assessment standard, which is 5.21% lower than that of the DQN method. The control index of the DQN algorithm is the most ideal. Figure 7 It can be obtained that the power deviation value of the interconnection line ΔP of the control method of the present invention is 12 The section limit of 0.1 (pu) is not exceeded, and the section safety constraint requirements are met. Compared with the PI control method and the literature method, the average values are reduced by 13.51% and 23.81% respectively. Figure 8 It can be seen that the average value of the frequency deviation fluctuation Δf1 of the control method of the present invention is reduced by 23.81% and 90.8% respectively compared with the reference method and the PI control method, which shows that the control method of the present invention can effectively restore the dynamic indicators of the system to 0 faster.
[0159] Compared with the literature method and PI control, the DQN algorithm is used as the controller. Fig. 9 It can be seen that the fluctuation of regional control deviation ACE has been significantly improved; Fig.10 It can be seen that when the regional control deviation ACE is small, the SOC of the energy storage system recovers to the optimal normal setting value of the corresponding period, so that the energy storage can continue to participate in AGC control. However, both the literature method and PI control have the situation that the energy storage cannot participate in AGC control during certain periods of the day.
[0160] In order to solve the problem of insufficient grid-connected active power regulation of large-scale wind and solar bases and heavy assessment pressure at the tie-line checkpoint, the embodiment of the present invention proposes an AGC control strategy for energy storage and other power sources that takes into account the recovery of SOC to the optimal normal setting value at all times of the day. By designing a controller based on the DQN algorithm to meet the CPS assessment standard while taking into account the recovery of energy storage SOC, the following conclusions are drawn through theoretical analysis and simulation verification:
[0161] 1) Based on the actual wind power fluctuation curve, the DQN algorithm is used to determine the energy storage SOC normal setting value that is best for smoothing the interconnection line power fluctuation and wind power fluctuation at different times of the day. The simulation results show that the optimal normal value of the energy storage SOC is different at different times of the day.
[0162] 2) Considering the difference in the absolute value of ACE, when the absolute value of ACE is large, simulation is carried out starting from the control performance standard, proving that when the ACE fluctuates greatly, the proposed DQN algorithm can meet the CPS index of the power system. When the absolute value of ACE is small, the energy storage SOC is restored to the optimal setting value for each period of the day, and the energy storage of the method proposed in the present invention can continue to participate in AGC control.
[0163] This embodiment also includes an AGC control device for energy storage and other power sources in coordination based on deep reinforcement learning, including:
[0164] Model building module, used to build a load frequency control model of a two-region interconnected system with energy storage;
[0165] The value determination module is used to determine the optimal normal setting value of the energy storage SOC at each time period throughout the day using a deep reinforcement learning algorithm based on the load frequency control model of the two-region interconnected system and the actual wind power fluctuation historical data;
[0166] The function design module is used to distinguish the absolute value of the regional control deviation according to the optimal normal setting value of the energy storage SOC in each period of the day, and to design the model reward function containing the control performance standard, section over-limit control and energy storage SOC recovery to the optimal normal setting value of the corresponding period based on the absolute value, and formulate the energy storage adaptive control strategy according to the model reward function.
[0167] This embodiment also includes an electronic device and a storage medium. An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, each step in the AGC control method for energy storage and other power sources based on deep reinforcement learning is implemented. A storage medium stores a computer program, and when the computer program is executed by the processor, each step in the AGC control method for energy storage and other power sources based on deep reinforcement learning is implemented.
[0168] In summary, the method of the present invention can adjust the regional control deviation ACE allocation ratio of energy storage and other power sources in real time, thereby realizing the output adjustment of the unit. The model reward function set by it enables the algorithm to meet the CPS index of the power system and continue to participate in AGC control.
[0169] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiments of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.
[0170] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0171] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0172] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0173] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0174] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A method for coordinating AGC control of energy storage and other power sources based on deep reinforcement learning, characterized in that: The following steps are involved: Build a load frequency control model for a two-region interconnected system with energy storage; Based on the load frequency control model of the two-region interconnected system and the actual historical data of wind power fluctuations, a deep reinforcement learning algorithm is used to determine the optimal normal setting value of the energy storage SOC at all times of the day; According to the optimal normal setting value of the energy storage SOC in each period of the day, the absolute value of the regional control deviation is distinguished, and the model reward function containing control performance standards, section over-limit control and energy storage SOC recovery to the optimal normal setting value of the corresponding period is designed based on the absolute value. The energy storage adaptive control strategy is formulated according to the model reward function.
2. According to the AGC control method of energy storage and other power sources based on deep reinforcement learning in claim 1, it is characterized in that: The load frequency control model of the two-region interconnected system includes: Energy storage system model SOC of energy storage system Among them, T b is the response time constant of the energy storage converter; Δu b It is the control input signal of the energy storage module; ΔP b is the output power of the energy storage system; η d is the discharge efficiency of energy storage, η c is the charging efficiency of energy storage, P bess Δt is the charging and discharging power of the energy storage within the time Δt, and C is the rated capacity of the energy storage system; Speed Regulator Model in Thermal Power Unit Module Among them, ΔP a is the speed regulator power deviation, T a is the speed regulator time constant; Δu a is the speed regulator control input signal, R is the droop coefficient, Δf1 is the frequency deviation in the first area; Steam turbine model in the thermal power unit module Among them, ΔP h is the turbine output power, T t is the turbine time constant; Generator-Load Model Where M is the inertia time constant of the regional unit, D is the load damping coefficient, ΔP h is the turbine output power, ΔP 12 is the tie line power deviation, ΔP L is the wind power fluctuation; Tie line power deviation model between areas 1 and 2 Area Control Deviation ACE ACE=ΔP 12 +BΔf Among them, T 12 is the tie line synchronization coefficient, Δf1 and Δf2 are the frequency deviations of the first area and the second area respectively; B is the frequency deviation factor.
3. According to claim 1, the AGC control method for energy storage and other power sources based on deep reinforcement learning is characterized in that: Based on the load frequency control model of the two-region interconnected system and the actual wind power fluctuation historical data, the deep reinforcement learning algorithm is used to determine the optimal normal setting value of the energy storage SOC at each time period throughout the day. The specific steps include: Set the energy storage initial SOC to different values and initialize the parameters of the DQN algorithm; The actual wind power fluctuation historical data is imported into the environmental constraints of the two-region load frequency control model with energy storage, the allocation ratio of energy storage to other power sources is initialized to 0, and the DQN algorithm is pre-trained. The trained agent is used for online control to calculate the energy storage SOC in real time. If the energy storage SOC exceeds the limit, the ACE allocation ratio of the energy storage is set to 0 to meet the energy storage SOC constraint. The real-time tie line power deviation ΔP during the online control process is recorded. 12 The frequency deviation Δf1 value of the value and area 1 is adjusted until all the set time periods are completed online; Statistics of the tie line power deviation ΔP in each period obtained through online control 12 The absolute value of the frequency deviation Δf1 in area 1 is used to determine the optimal normal setting value of the energy storage SOC in each time period.
4. According to claim 1, the AGC control method for energy storage and other power sources based on deep reinforcement learning is characterized in that: The model reward function is defined as follows: |ACE|<x: R=(SOC-SOC in ) 2 |ACE|>x: Where: R is the model reward function, P limit is the cross-sectional limit, x is the dividing point of the absolute value of ACE, SOC in Indicates the optimal normal setting value of energy storage SOC at each time period throughout the day.
5. AGC control device for energy storage and other power sources based on deep reinforcement learning, characterized in that: include: Model building module, used to build a load frequency control model of a two-region interconnected system with energy storage; The value determination module is used to determine the optimal normal setting value of the energy storage SOC at each time period throughout the day using a deep reinforcement learning algorithm based on the load frequency control model of the two-region interconnected system and the actual wind power fluctuation historical data; The function design module is used to distinguish the absolute value of the regional control deviation according to the optimal normal setting value of the energy storage SOC in each period of the day, and to design the model reward function containing the control performance standard, section over-limit control and energy storage SOC recovery to the optimal normal setting value of the corresponding period based on the absolute value, and formulate the energy storage adaptive control strategy according to the model reward function.
6. According to claim 5, the energy storage and other power supply coordinated AGC control device based on deep reinforcement learning is characterized in that: The load frequency control model of the two-region interconnected system in the model building module includes: Energy storage system model SOC of energy storage system Among them, T b is the response time constant of the energy storage converter; Δu b It is the control input signal of the energy storage module; ΔP b is the output power of the energy storage system; η d is the discharge efficiency of energy storage, η c is the charging efficiency of energy storage, P bess Δt is the charging and discharging power of the energy storage within the time Δt, and C is the rated capacity of the energy storage system; Speed Regulator Model in Thermal Power Unit Module Among them, ΔP a is the speed regulator power deviation, T a is the speed regulator time constant; Δu a is the speed regulator control input signal, R is the droop coefficient, Δf1 is the frequency deviation in the first area; Steam turbine model in the thermal power unit module Among them, ΔP h is the turbine output power, T t is the turbine time constant; Generator-Load Model Where M is the inertia time constant of the regional unit, D is the load damping coefficient, ΔP h is the turbine output power, ΔP 12 is the tie line power deviation, ΔP L is the wind power fluctuation; Tie line power deviation model between areas 1 and 2 Area Control Deviation ACE ACE=ΔP 12 +BΔf Among them, T 12 is the tie line synchronization coefficient, Δf1 and Δf2 are the frequency deviations of the first area and the second area respectively; B is the frequency deviation factor.
7. According to claim 5, the energy storage and other power supply coordinated AGC control device based on deep reinforcement learning is characterized in that: The numerical determination module uses a deep reinforcement learning algorithm to determine the optimal normal setting value of the energy storage SOC for each period of the day based on the load frequency control model of the two-region interconnected system and the actual wind power fluctuation historical data, and specifically includes the following steps: Set the energy storage initial SOC to different values and initialize the parameters of the DQN algorithm; The actual wind power fluctuation historical data is imported into the environmental constraints of the two-region load frequency control model with energy storage, the allocation ratio of energy storage to other power sources is initialized to 0, and the DQN algorithm is pre-trained. The trained agent is used for online control to calculate the energy storage SOC in real time. If the energy storage SOC exceeds the limit, the ACE allocation ratio of the energy storage is set to 0 to meet the energy storage SOC constraint. The real-time tie line power deviation ΔP during the online control process is recorded. 12 The frequency deviation Δf1 value of the value and area 1 is adjusted until all the set time periods are completed online; Statistics of the tie line power deviation ΔP in each period obtained through online control 12 The absolute value of the frequency deviation Δf1 in area 1 is used to determine the optimal normal setting value of the energy storage SOC in each time period.
8. According to claim 5, the energy storage and other power supply coordinated AGC control device based on deep reinforcement learning is characterized in that: The model reward function in the function design module is defined as follows: |ACE|<x: R=(SOC-SOC in ) 2 |ACE|>x: Where: R is the model reward function, P limit is the cross-sectional limit, x is the dividing point of the absolute value of ACE, SOC in Indicates the optimal normal setting value of energy storage SOC at each time period throughout the day.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, each step of the AGC control method for coordinating energy storage with other power sources based on deep reinforcement learning is implemented as described in any one of claims 1 to 4.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, each step of the AGC control method for coordinating energy storage with other power sources based on deep reinforcement learning is implemented as described in any one of claims 1 to 4.
Citation Information
Cited By
Fire storage system agc frequency control method, system, device and medium based on deep recurrent q network
CN122659954A