An automatic power generation control method, system, medium, and processor for large-scale grid connection of new energy sources.
By establishing a multi-regional interconnected power grid model and a deep reinforcement learning controller, the problem of insufficient adaptive capability of traditional control strategies in the grid connection of new energy sources is solved, and distributed collaborative control of grid frequency stability and new energy consumption is realized.
Patent Information
- Application Number
- CN202411537587.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2044-10-31
AI Technical Summary
Traditional centralized automatic generation control modes are difficult to meet the grid frequency regulation requirements of large-scale new energy grid connection, especially in terms of power distribution and utilization. Moreover, existing control strategies lack adaptive capabilities and are unable to respond quickly to changes in the power system.
A multi-regional interconnected power grid model is established, and a deep reinforcement learning controller is used to calculate regional control deviations. By iteratively updating the total control power output, distributed power allocation and frequency stability are achieved by coordinating thermal power units, hydropower units, and electric vehicle clusters.
It has achieved frequency stability in multi-regional power grids and the absorption of new energy sources in new power systems, and improved the coordinated control efficiency of generator units and the control performance of the power grid.
Smart Images

Figure CN119448430B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automatic power generation control technology, and in particular to an automatic power generation control method, system, medium, and processor for large-scale grid connection of new energy sources. Background Technology
[0002] With the increasing scarcity of fossil fuels and the growing severity of environmental problems such as the greenhouse effect, the transition to renewable energy is an inevitable trend in China's and even the world's energy development. To achieve this transition, it is necessary to build a clean, low-carbon, safe, and efficient energy system, implement renewable energy substitution initiatives, deepen power system reform, and construct a new power system dominated by new energy sources. Against this backdrop of building a new power system, the proportion of new energy sources integrated into the power system is constantly increasing. However, the intermittent, random, and unpredictable nature of these new energy sources poses numerous challenges to the power grid. In particular, when the volatility of new energy sources is factored into short timescales of seconds and minutes, the power grid urgently needs frequency regulation resources.
[0003] With the rapid development of smart grids, the installed capacity is constantly expanding, and new energy and distributed energy are constantly being connected. The traditional centralized automatic generation control (AGC) mode can hardly meet the development and operation conditions of the power grid, and the distributed AGC control mode is imperative.
[0004] From the perspective of distributed energy utilization, due to limitations in grid structure, power plant capacity, and unit regulation rates, different types of power plants exhibit significant differences in power allocation and utilization rates. Power plants controlled by centralized AGC (Automatic Generation Control) can only be allocated through provincial dispatch centers, making coordinated control between distributed energy plants and new energy plants difficult. These small-capacity power plants without fixed power generation targets struggle to achieve high utilization rates under centralized AGC. Clearly, centralized AGC lacks economic viability and scientific rigor, making it difficult to meet the development needs of the power grid. This underscores the crucial importance of researching distributed AGC.
[0005] Currently, most research primarily uses traditional control methods (such as Q-learning) as frequency regulation control strategies. However, for new power systems dominated by renewable energy sources, discrete Q-learning and its derived control algorithms exhibit poor adaptability, failing to quickly achieve system control of individual generator units when the power system changes, and also suffer from poor control accuracy. Therefore, there is an urgent need in this field to research a reinforcement learning control strategy or method for distributed AGC (Automatic Generation Control). Summary of the Invention
[0006] To address the problems in existing technologies, this invention provides an automatic power generation control method, system, medium, and processor for large-scale grid connection of new energy sources. The specific technical solution is as follows:
[0007] An automatic power generation control method for large-scale grid connection of new energy sources includes the following steps:
[0008] A multi-regional interconnected power grid model is constructed, and the model parameters are initialized. The multi-regional interconnected power grid model is then integrated into the power grid. The multi-regional interconnected power grid model includes a deep reinforcement learning controller for each region, wind turbines, photovoltaic power plants, thermal power units, hydropower units, electric vehicle clusters, and loads. The deep reinforcement learning controller is connected to the thermal power units, hydropower units, and electric vehicle clusters, respectively.
[0009] Calculate the power allocation ratio of thermal power units, hydropower units, and electric vehicle clusters based on the capacity of thermal power units, hydropower units, and electric vehicle clusters in each region.
[0010] Based on the output power of wind turbines, photovoltaic power plants, thermal power units, hydropower units, the power of electric vehicle clusters participating in frequency regulation, and the load in each region, the real-time frequency deviation of the power grid in each region and the exchange power between adjacent regions are calculated, and the regional control deviation of each region is calculated.
[0011] The regional control deviation of each region is input into the deep reinforcement learning controller of the corresponding region. The controller is iteratively updated through its internal control method and outputs the total control power of the corresponding region.
[0012] Based on the total control power of the corresponding region and the power allocation ratio of thermal power units, hydropower units, and electric vehicle clusters, the allocated power of thermal power units, hydropower units, and electric vehicle clusters is obtained, and then the power command of thermal power units, hydropower units, and electric vehicle clusters in the corresponding region is obtained.
[0013] Thermal power units, hydropower units, and electric vehicle clusters execute corresponding power commands and interact with the power grid.
[0014] Preferably, the specific calculation method for the power allocation ratio of thermal power units, hydropower units, and electric vehicle clusters is as follows:
[0015]
[0016] In the formula: a f a h a ev These represent the power allocation ratios for thermal power units, hydropower units, and electric vehicle clusters, respectively. f +a h +a ev =1 and a f a h a ev ∈[0,1];E f Eh These are the rated capacities of thermal power units and hydropower units, respectively, E ev,all The available capacity for electric vehicle clusters.
[0017] Preferably, the calculation method for the regional control deviation of each region is as follows:
[0018] Δf k =L -1 {(P g,k +P pv,k +P h,k +P f,k +P ev,k -P L,k -P tie,kj )×K p / (1+sT p )};
[0019] P tie,kj =L -1 {(Δf k -Δf j )×(2πT kj / s)};
[0020] ACE k =P tie,kj +β k Δf k ;
[0021] Where, Δf k Let L be the real-time frequency deviation of the power grid in region k. -1 For the inverse Laplace transform, P g,k Let P be the output power of the wind turbine in region k. pv,k Let P be the output power of the photovoltaic power station in region k. h,k Let P be the output power of the hydropower unit in region k. f,k Let P be the output power of the thermal power unit in region k. ev,k P represents the power of the electric vehicle cluster in region k participating in frequency modulation. L,k For the load of region k, P tie,kj The power exchanged between the tie lines of regions k and j, where regions k and j are adjacent, K p T represents the coefficients of the frequency response function. p Let s be the time constant of the frequency response function, s be the Laplace operator, and Δf be the time constant. j Let T be the real-time frequency deviation of the power grid in region j. kj Let be the time constant of the connection line between regions k and j; ACE k For the regional control deviation of region k, β k is the frequency deviation coefficient for region k.
[0022] Preferably, the power P of the electric vehicle cluster in region k participating in frequency modulation ev,k Specifically as follows:
[0023] A controllable dynamic change model for the number of electric vehicles is established, as follows:
[0024] Let the starting time be t0, the time when electric vehicles in region k connect to the grid be t1, and the charging time of electric vehicles be T2. Therefore, the time t3 when electric vehicles enter operation frequency regulation mode is:
[0025] t3 = t1 + T2;
[0026] Electric vehicles determine their eligibility to participate in frequency regulation services based on their own state of charge (SBC). The calculation of the electric vehicle's SBC is as follows:
[0027]
[0028] Where: SOC0 is the current state of charge of a single electric vehicle; E ev For the existing capacity of a single electric vehicle; E ev,max This refers to the maximum capacity of a single electric vehicle.
[0029] The charging time T2 for electric vehicles is calculated using the following formula:
[0030]
[0031] in, The charging power for a single electric vehicle; η ev For electric vehicle charging efficiency; SOC min The minimum state of charge required for electric vehicles to participate in frequency regulation;
[0032] The Monte Carlo method was used to randomly sample the controllable entry / exit timetables of electric vehicles and perform statistical analysis to obtain the number N of electric vehicles in the controllable state at each time point in the region. in '(t i The number of electric vehicles N in a controllable state out '(t i The cumulative number N of electric vehicles entering a controllable state at time t can be calculated using the following formula. in (t) and the cumulative number of electric vehicles N that have exited the controllable state out (t):
[0033]
[0034]
[0035] Let N0 be the initial number of controllable electric vehicles at time t0, and N be the number of controllable electric vehicles at time t. c(t) is shown in the following equation: N c (t)=N0+N in (t)-N out (t);
[0036] Therefore, the available capacity of the electric vehicle cluster at time t for:
[0037]
[0038] When participating in frequency regulation, the total usable capacity of electric vehicles needs to be limited. If the existing capacity is less than the minimum capacity limit, the usable capacity of electric vehicles is zero.
[0039] Lower limit of the capacity of the electric vehicle cluster at time t for:
[0040]
[0041] Capacity limit of the electric vehicle cluster at time t for:
[0042]
[0043] Available capacity E of the electric vehicle cluster at time t ev,all The calculation formula is as follows:
[0044]
[0045] Calculate a based on the existing capacity. ev The power P of the electric vehicle cluster in region k participating in frequency modulation ev,k The calculation method is as follows:
[0046] P ev,k =P t,k ·a ev ;
[0047] Among them, P t,k This represents the total power command output by the deep reinforcement learning controller at time t in region k.
[0048] Preferably, the deep reinforcement learning controller includes two policy networks. and 2 evaluation networks
[0049] The policy network consists of an input layer, a first hidden layer, a second hidden layer, a third hidden layer, a fourth hidden layer, and an output layer. The policy network is used to output action 'a' in state S, where state S represents the area control deviation of the corresponding region, and action 'a' represents the total control power deviation of the corresponding region. The mathematical expression for the policy network is as follows:
[0050]
[0051] In the formula: y u The input to the policy network is y1 and y2, which are the outputs of the first and second hidden layers, respectively; μ and logσ are the outputs of the third and fourth hidden layers, respectively; W1 μ , W μ W logσ These are the weight matrices for the first hidden layer, the second hidden layer, the third hidden layer, and the fourth hidden layer, respectively. b μ b logσ Let be the bias vectors of the first, second, third, and fourth hidden layers; the output action 'a' follows a normal distribution N(μ,σ), where the mean μ and standard deviation σ are the mean and standard deviation of the normal distribution followed by the output action 'a', respectively; tanh(x) = (e^(μ,σ)) / (σ) x -e -x ) / (e x +e -x Let A be the activation function and A be the maximum action value.
[0052] The evaluation network consists of an input layer, a first hidden layer, a second hidden layer, and an output layer. The mathematical expression for evaluating the network is as follows: This network is used to determine the value of action a in state S and output a value Q.
[0053]
[0054] Where, x u To evaluate the network input, x1 and x2 are the outputs of the first and second hidden layers, respectively; Q is the output of the output layer; W1 Q , These are the weight matrices for the first hidden layer, the second hidden layer, and the output layer, respectively. represents the bias vectors for the first hidden layer, the second hidden layer, and the output layer; relu(x) = max(0,x) is the activation function, ensuring that the output value is not negative.
[0055] Preferably, the regional control deviation of each region is input into the deep reinforcement learning controller of the corresponding region, and the total control power of the corresponding region is iteratively updated and output through the internal control method of the deep reinforcement learning controller. Specifically, this includes the following steps:
[0056] Initialize two evaluation networks and 2 policy networks Set the maximum action value, learning rate α, discount factor γ, entropy regularization coefficient τ1, environmental experience pool B1, and expert experience pool B2 and their upper limit B. 2max Random action N, optimal action proportion coefficient ρ, expert experience proportion ξ;
[0057] Input the area control deviation (ACE) value of the region at time t and use it as the environmental state S at time t. t ;
[0058] Determine whether the amount of information stored in expert experience pool B2 has reached its upper limit. 2max If the amount of information stored in expert experience pool B2 reaches its upper limit B... 2max Then, through the policy network at time t Output the environmental state S at time t t The following action a t If the amount of information stored in expert experience pool B2 has not reached the upper limit B... 2max Then, based on expert experience, the environmental state S at time t is selected. t The following action
[0059] Calculate real-time reward r t ;
[0060] Environmental information is stored in environmental experience buffer pool B1, and expert information is stored in expert experience buffer pool B2. Environmental information includes the environmental state S from the previous time step. t-1 Corresponding real-time strategies action a t-1 The reward r from the previous moment t-1 The environmental state S at this moment t Expert information includes the environmental state S at the previous moment. t-1 Actions selected based on corresponding expert experience Rewards from the previous moment The environmental state S at this moment t ;
[0061] Based on the proportion of expert experience ξ, samples are randomly drawn from the expert experience buffer pool B2 with equal probability. Each data point is then drawn from the environmental experience buffer pool B1 with the same probability. The two sets of data together constitute the training data, where G is the number of training samples set for mini-batch descent.
[0062] The parameters of the evaluation network and policy network are updated using the training data, and the proportion of expert experience ξ is updated accordingly.
[0063] Output power action a, calculate the total power command for the corresponding region at time t.
[0064] Preferably, the calculation of the total power command for the corresponding region at time t is as follows:
[0065] If the information stored in expert experience buffer pool B2 has not exceeded the storage limit:
[0066]
[0067] If the information stored in expert experience buffer pool B2 exceeds the storage limit:
[0068] P t =P t-1 +a t ;
[0069] Among them, P t For the total power command at time t in the corresponding region, P t-1 This represents the total power command for the corresponding region at time t-1.
[0070] An automatic power generation control system for large-scale grid connection of new energy sources, applied to the method described, includes:
[0071] The model building module is used to build a multi-regional interconnected power grid model, initialize model parameters, and integrate the multi-regional interconnected power grid model into the power grid. The multi-regional interconnected power grid model includes deep reinforcement learning controllers, wind turbines, photovoltaic power plants, thermal power units, hydropower units, electric vehicle clusters, and loads in each region. The deep reinforcement learning controllers are connected to the thermal power units, hydropower units, and electric vehicle clusters, respectively.
[0072] The power allocation ratio calculation module is used to calculate the power allocation ratio of thermal power units, hydropower units, and electric vehicle clusters based on the capacity of thermal power units, hydropower units, and electric vehicle clusters in each region.
[0073] The regional control deviation calculation module is used to calculate the real-time frequency deviation of the power grid and the exchange power between adjacent regions based on the output power of wind turbines, photovoltaic power plants, thermal power plants, hydropower plants, the power of electric vehicle clusters participating in frequency regulation, and the load in each region, and to calculate the regional control deviation of each region. The deep reinforcement learning controller is used to input the regional control deviation of each region into the deep reinforcement learning controller of the corresponding region, and to iteratively update and output the total control power of the corresponding region through the internal control method of the deep reinforcement learning controller. The total power command calculation module is used to obtain the allocated power of thermal power plants, hydropower plants, and electric vehicle clusters based on the total control power of the corresponding region and the power allocation ratio of thermal power plants, hydropower plants, and electric vehicle clusters, and then obtain the power command of thermal power plants, hydropower plants, and electric vehicle clusters in the corresponding region.
[0074] The power grid interaction module is used by thermal power units, hydropower units, and electric vehicle clusters to execute corresponding power commands and interact with the power grid.
[0075] A computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to perform the automatic power generation control method for large-scale grid connection of new energy sources.
[0076] A processor for running a program, wherein the program executes the automatic power generation control method for large-scale grid connection of new energy sources.
[0077] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0078] This invention establishes a distributed power grid control mechanism. It integrates a multi-regional interconnected power grid model into the grid and calculates the power allocation ratios of thermal power units, hydropower units, and electric vehicle clusters. It also calculates the regional control deviations for each region. A distributed deep reinforcement learning controller takes the corresponding region's control deviation as input, iteratively updates the controller through its internal control method, and outputs the total control power for the corresponding region. This yields the power commands for the corresponding thermal power units, hydropower units, and electric vehicle clusters. The thermal power units, hydropower units, and electric vehicle clusters execute the corresponding power commands to achieve system closed-loop control and interact with the power grid. The control method of this invention, through the designed controller, can achieve CPS index compliance and frequency stability in multi-regional power grids, the absorption of large-scale new energy sources in new power systems, and the coordinated control of generator units. Attached Figure Description
[0079] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.
[0080] Figure 1 This is a schematic flowchart of the method of the present invention;
[0081] Figure 2 This is a multi-regional interconnected power grid model built in an embodiment of the present invention;
[0082] Figure 3 This is a principle model for wind-solar-thermal-hydro-electric vehicles in a multi-regional interconnected power grid model.
[0083] Figure 4 This is a schematic diagram of the wind and solar power curves in the embodiment;
[0084] Figure 5 A graph showing the change in capacity of electric vehicles over a day;
[0085] Figure 6 This is a schematic diagram of the policy network structure;
[0086] Figure 7 A schematic diagram of the network structure for evaluation;
[0087] Figure 8 Flowchart of the internal control method for a deep reinforcement learning controller;
[0088] Figure 9 This is a schematic diagram illustrating the updates of the policy network and the evaluation network.
[0089] Figure 10 This is a comparison curve of the power grid control performance indicators of two regions under random disturbances according to the present invention, compared with different methods. Figure 10 (a) Waveform diagram of random disturbance. Figure 10 (b) is the frequency curve under random disturbance; Figure 10 (c) is a graph of the 10-minute ACE average value under random perturbation; Figure 10 (d) is the 10-minute CPS1 average curve under random perturbation;
[0090] Figure 11 This is a system schematic diagram of the present invention. Detailed Implementation
[0091] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0092] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0093] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0094] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0095] Example 1:
[0096] like Figure 1 As shown, this embodiment provides an automatic power generation control method for large-scale grid connection of new energy sources, specifically including the following steps:
[0097] Step S1 involves building a multi-regional interconnected power grid model, initializing model parameters, and integrating the multi-regional interconnected power grid model into the power grid. The multi-regional interconnected power grid model includes deep reinforcement learning controllers for each region, wind turbines, photovoltaic power plants, thermal power units, hydropower units, electric vehicle clusters, and loads. The deep reinforcement learning controllers are connected to the thermal power units, hydropower units, and electric vehicle clusters, respectively. In this embodiment, the invention is based on the real-world context of large-scale grid connection of new energy sources (wind power, photovoltaic power, and electric vehicle clusters), such as... Figure 2 As shown, a realistic two-region interconnected power grid model was constructed to verify the cooperative control strategy based on deep reinforcement learning designed in this invention. The regional power tie lines are AC power tie lines between different regions. When there is an imbalance between power generation and consumption in one region, power from other regions is mobilized through these tie lines to maintain the stability of the entire power grid. To verify the controller's control capability, the load is a randomly generated load signal using MATLAB / Simulink.
[0098] Step S2: Calculate the power allocation ratio of thermal power units, hydropower units, and electric vehicle clusters based on the capacity of each region. The calculation method is the same for each region, and the specific calculation is as follows:
[0099]
[0100] In the formula: a f a h a ev These represent the power allocation ratios for thermal power units, hydropower units, and electric vehicle clusters, respectively. f +a h +a ev =1 and a f a h a ev ∈[0,1];E f E h These are the rated capacities of thermal power units and hydropower units, respectively; they are two fixed values. ev,all The available capacity of the electric vehicle cluster is a value that changes in real time.
[0101] Step S3: Based on the output power of wind turbines, photovoltaic power plants, thermal power units, hydropower units, the power of electric vehicle clusters participating in frequency regulation, and the load in each region, calculate the real-time frequency deviation of the power grid and the exchange power between adjacent regions for each region, and calculate the regional control deviation for each region. The calculation method for the regional control deviation of each region is as follows:
[0102] Δf k =L -1 {(P g,k +P pv,k +P h,k +P f,k +P ev,k -P L,k -P tie,kj )×K p / (1+sT p (2)
[0103] P tie,kj =L -1 {(Δf k -Δf j )×(2πT kj / s)};(3)
[0104] ACE k =P tie,kj +β k Δf k (4)
[0105] Where, Δf k Let L be the real-time frequency deviation of the power grid in region k. -1 For the inverse Laplace transform, P g,k Let P be the output power of the wind turbine in region k. pv,k Let P be the output power of the photovoltaic power station in region k. h,k Let P be the output power of the hydropower unit in region k. f,k Let P be the output power of the thermal power unit in region k. ev,k Let P be the power of the electric vehicle cluster in region k participating in frequency modulation. L,k For the load of region k, P tie,kj The power exchanged between the tie lines of regions k and j, where regions k and j are adjacent, K p T represents the coefficients of the frequency response function. p Let s be the time constant of the frequency response function, s be the Laplace operator, and Δf be the time constant. j Let T be the real-time frequency deviation of the power grid in region j. kj Let be the time constant of the connection line between regions k and j; ACE k For the regional control deviation of region k, β kis the frequency deviation coefficient for region k.
[0106] like Figure 3 As shown, the generator sets are thermal power steam turbine units and hydroelectric power units considering generator output rate control (GRC). Thermal power generator sets include a governor, GRC, steam turbine, and power limiting module; hydroelectric generator sets include a governor, GRC, turbine, and power limiting module. These modules are used to simulate the actual operating conditions of real thermal and hydroelectric power units. Combined with... Figure 2 and Figure 3 The formulas for calculating the output power of thermal power units and hydropower units can be obtained:
[0107] The output power P of the thermal power unit in region k f,k as follows:
[0108] P f,k =L -1 {(P t,k ·a f -Δf / R1) / (1+sT g (1+sT) f (5)
[0109] Among them, P t,k Let a be the total power command output by the deep reinforcement learning controller at time t in region k. f R1 is the power distribution ratio of the thermal power unit, and T is the droop coefficient of the thermal power unit. g T is the time delay constant of the governor of the thermal power unit. f This represents the time constant of the thermal power unit.
[0110] The output power P of the hydropower unit in region k h,k as follows:
[0111] P h,k =L -1 {(P t,k ·a h -Δf / R2)·(sT rs +1)(-sT w1s +1) / (1+sT gh (1+sT) rh )(1+sT rh (6)
[0112] Among them, a h R2 is the power distribution ratio of the hydropower unit, and T is the droop coefficient of the hydropower unit. rs T w1s T rh T is the time constant of the hydropower unit; ghThis is the time delay constant of the hydropower unit speed governor.
[0113] The wind turbine output power P in region k g,k The calculation is as follows:
[0114]
[0115] In the formula: V is the rated power of the fan; w This refers to the actual wind speed; Rated wind speed; To cut in wind speed; To cut off the wind speed.
[0116] The output power P of the photovoltaic power station in region k pv,k as follows:
[0117]
[0118] In the formula: c is the rated power generation capacity of the photovoltaic power station; pv T represents the temperature conversion power factor of photovoltaics; T represents the current air temperature; T ref This is a reference temperature value; s pv The light intensity at the current moment.
[0119] Wind power output is highly uncertain, with peak output in the evening; solar power output occurs during the day, peaking around midday. Specifically... Figure 4 As shown, the wind and solar power outputs described above can effectively simulate actual wind and solar power outputs.
[0120] Electric vehicles participate in power system frequency regulation through vehicle-to-grid (V2G) technology, achieving rapid response to grid frequency through centralized control. They discharge when grid load is high and charge when load is low, aiding in peak shaving and frequency regulation. The power P of the electric vehicle cluster in region k participating in frequency regulation is... ev,k Specifically as follows:
[0121] A controllable dynamic change model for the number of electric vehicles is established, as follows:
[0122] Let the starting time be t0, the time when electric vehicles in region k connect to the grid be t1, and the charging time of electric vehicles be T2. Therefore, the time t3 when electric vehicles enter operation frequency regulation mode is:
[0123] t3 = t1 + T2; (9)
[0124] Electric vehicles need to determine whether they can participate in frequency regulation services based on their own state of charge (SOC0). The calculation of the electric vehicle's SOC0 is as follows:
[0125]
[0126] Where: SOC0 is the current state of charge of a single electric vehicle; E ev For the existing capacity of a single electric vehicle; E ev,max This refers to the maximum capacity of a single electric vehicle.
[0127] The charging time T2 for electric vehicles is calculated using the following formula:
[0128]
[0129] in, The charging power for a single electric vehicle; η ev For electric vehicle charging efficiency; SOC min The minimum state of charge required for electric vehicles to participate in frequency regulation;
[0130] The Monte Carlo method was used to randomly sample the controllable entry / exit timetables of electric vehicles and perform statistical analysis to obtain the number N of electric vehicles in the controllable state at each time point in the region. in '(t i The number of electric vehicles N in a controllable state out '(t i The cumulative number N of electric vehicles entering a controllable state at time t can be calculated using the following formula. in (t) and the cumulative number of electric vehicles N that have exited the controllable state out (t):
[0131]
[0132]
[0133] Let N0 be the initial number of controllable electric vehicles at time t0, and N be the number of controllable electric vehicles at time t. c (t) is shown in the following formula:
[0134] N c (t)=N0+N in (t)-N out (t); (14)
[0135] Therefore, the available capacity of the electric vehicle cluster at time t for:
[0136]
[0137] When participating in frequency regulation, the total usable capacity of electric vehicles needs to be limited. If the existing capacity is less than the minimum capacity limit, the usable capacity of electric vehicles is zero.
[0138] Lower capacity limit of the electric vehicle cluster at time t for:
[0139]
[0140] Capacity limit of the electric vehicle cluster at time t for:
[0141]
[0142] Available capacity E of the electric vehicle cluster at time t ev,all The calculation formula is as follows:
[0143]
[0144] Figure 5 The figure shows the change in capacity of an electric vehicle over a day, assuming a State of Charge (SOC) of the electric vehicle. min The value is 0.6. The electric private car connects to the power grid at times t1 between 8:00 and 9:00 (assuming t1 follows a uniform distribution, i.e., t1~U(8,9)), 17:30 and 19:00, or 21:00 and 22:30 (the car owner goes to leisure and entertainment immediately after get off work and does not drive home immediately). The electric private car disconnects from the power grid (i.e., the time t4 when it leaves the controllable state) between 7:30 and 8:30 (assuming t4 follows a uniform distribution, i.e., t4~U(7.5,8.5)), and 17:00 and 18:30.
[0145] The capacity calculated according to formula (18) and a calculated in step S2 are obtained. ev Thus, the power P of the electric vehicle cluster in region k participating in frequency modulation can be obtained. ev,k :
[0146] P ev,k =P t,k ·a ev (19)
[0147] Among them, P t,k This represents the total power command output by the deep reinforcement learning controller at time t in region k.
[0148] This invention incorporates electric vehicle clusters to simulate the diversity of frequency regulation resources in a real power grid and further verifies the collaborative capabilities of EL-GAC deep reinforcement learning. Furthermore, to effectively address the integration problem of wind and solar power, which are highly stochastic, wind and solar power are directly integrated into the power grid. The adaptive nature of the EL-GAC deep reinforcement learning controller and its collaborative capabilities with different generator sets are used to achieve the integration of new energy sources.
[0149] Step S4: Input the regional control deviation of each region into the deep reinforcement learning controller of the corresponding region, and iteratively update it through the internal control method of the deep reinforcement learning controller and output the total control power of the corresponding region.
[0150] Figure 2 The EL-GAC deep reinforcement learning controller is an important component of this invention. It can monitor the frequency and ACE of the power grid, and then output the specific power generation of the unit to eliminate the impact of large-scale energy grid connection on the power system, maintain the frequency stability of the power grid, improve the control performance standard (CPS) of the AGC system, and realize the problem of large-scale new energy consumption.
[0151] Deep reinforcement learning controllers are a new generation of artificial intelligence control methods based on deep reinforcement learning. Their primary advantage is the ability to quickly learn optimal cooperative strategies by rationally utilizing expert data, and they possess efficient self-updating and adaptive capabilities. They interact with the power system using ACE (Active Power Response) as input and the current system state, with the generator's total power command as the action output. The resulting control data is then stored in a database, and finally, proportionally used for learning and updating the cooperative control strategy. This invention employs a continuous controller, replacing the action set of a discrete controller with an action range. The neural network output is mapped to [-1, 1] using a tanh activation function and then expanded to an appropriate range. This invention limits the action range to [-25, 25] MW. ACE is a comprehensive indicator of power deviation and frequency deviation; as long as ACE is maintained within the acceptable range, the power grid indicators will tend to stabilize.
[0152] The deep reinforcement learning controller includes two policy networks. and 2 evaluation networks like Figure 6 As shown, the policy network includes an input layer, a first hidden layer, a second hidden layer, a third hidden layer, a fourth hidden layer, and an output layer. The policy network is used to output action 'a' in state S, where state S represents the area control deviation of the corresponding region, and action 'a' represents the total control power deviation of the corresponding region. The mathematical expression for the policy network is as follows:
[0153]
[0154] In the formula: y u The input to the policy network is y1 and y2, which are the outputs of the first and second hidden layers, respectively; μ and logσ are the outputs of the third and fourth hidden layers, respectively; W1 μ , W μ W logσThese are the weight matrices for the first hidden layer, the second hidden layer, the third hidden layer, and the fourth hidden layer, respectively. b μ b logσ Let be the bias vectors of the first, second, third, and fourth hidden layers; the output action 'a' follows a normal distribution N(μ,σ), where the mean μ and standard deviation σ are the mean and standard deviation of the normal distribution followed by the output action 'a', respectively; tanh(x) = (e^(μ,σ)) / (σ) x -e -x ) / (e x +e -x ) is the activation function, and A is the maximum action value; the policy network samples the mean and standard deviation of the output into a normal distribution to obtain the action output.
[0155] like Figure 7 As shown, the evaluation network includes an input layer, a first hidden layer, a second hidden layer, and an output layer. The mathematical expression for evaluating the network is as follows: This network is used to determine the value of action a in state S and output a value Q.
[0156]
[0157] Where, x u To evaluate the network input, x1 and x2 are the outputs of the first and second hidden layers, respectively; Q is the output of the output layer; W1 Q , These are the weight matrices for the first hidden layer, the second hidden layer, and the output layer, respectively. represents the bias vectors for the first hidden layer, the second hidden layer, and the output layer; relu(x) = max(0,x) is the activation function, ensuring that the output value is not negative.
[0158] The evaluation network has multiple neurons in each layer, and the connections of each neuron in each layer are multiplied by a coefficient, such as... Figure 3 The input layer is connected to the first hidden layer, and the matrix composed of these coefficients is the weight matrix. After multiplying the weight matrix, the corresponding bias is added. These biases can form a bias vector, which is then passed through an activation function to obtain the output. It is connected to the next layer in the same way until the final output value Q is obtained.
[0159] like Figure 8 As shown, taking the regional control deviation as input, the total control power of the region is iteratively updated and output through the internal control method of the deep reinforcement learning controller, specifically including the following steps:
[0160] Step S41, initialize two evaluation networks and 2 policy networks To ensure stable operation of the algorithm, the maximum action value is set to A = 25; the learning rate α = 0.001; the discount factor γ = 0.99; and the entropy regularization coefficient τ1 = 0.2. The size of the experience pool is a crucial parameter in deep reinforcement learning, significantly impacting the stability of the learning process, data efficiency, and final performance. Therefore, environment experience pool B1 and expert experience pool B2 are set, along with their upper limits B1 and B2. 2max B 2max =200; random actions N=30; optimal action proportion coefficient ρ=0.1; expert experience proportion ξ.
[0161] Step S42: Input the regional control deviation (ACE) value of the region at time t and use it as the environmental state S at time t. t .
[0162] Step S43: Determine whether the amount of information stored in the expert experience pool B2 has reached the upper limit B. 2max If the amount of information stored in expert experience pool B2 reaches its upper limit B... 2max Then, through the policy network at time t Output the environmental state S at time t t The following action a t If the amount of information stored in expert experience pool B2 has not reached the upper limit B... 2max Then, based on expert experience, the environmental state S at time t is selected. t The following action
[0163] Step S44, calculate the real-time reward r t This invention considers the reliability and power quality of novel power systems, setting the controller's objective as minimizing ACE and Δf. The reward function for the controller is generated by linearly weighting ACE and Δf, as detailed below:
[0164]
[0165] in, Let t be the ratio of the frequency deviation Δf at time t to 0.2, and k1 and k2 be weighting factors, both of which are set to 0.5 in this invention.
[0166] Step S45: Store environmental information in environmental experience buffer pool B1 and expert information in expert experience buffer pool B2. Environmental information includes the environmental state S from the previous time step. t-1 Corresponding real-time strategies action a t-1 The reward r from the previous moment t-1 The environmental state S at this moment t Expert information includes the environmental state S at the previous moment. t-1Actions selected based on corresponding expert experience Rewards from the previous moment The environmental state S at this moment t .
[0167] Step S46: Randomly draw from the expert experience buffer pool B2 with the same probability according to the proportion of expert experience ξ. Each data point is then drawn from the environmental experience buffer pool B1 with the same probability. The two sets of data together constitute the training data, where G is the amount of data required for mini-batch descent, and in this embodiment, G = 64.
[0168] Step S47: Update the parameters of the evaluation network and policy network using the training data, and update the expert experience ratio ξ. A larger ξ indicates a higher proportion of expert information participating in the update, which can accelerate the convergence speed of the algorithm; a smaller ξ indicates a higher proportion of environmental information participating in the update, which can increase the probability of learning the most accurate power command. The expert experience ratio ξ changes with the number of iterations, and the calculation formula is:
[0169]
[0170] In the formula: i represents the iteration rate of ξ, which is set to 200. The algorithm decreases ξ by 0.1 every i iterations, and stops updating ξ when ξ is 0.1.
[0171] Among them, evaluation network Network parameters The update is as follows:
[0172]
[0173] θ2←τ2·θ1-(1-τ2)·θ2; (25)
[0174] In the formula: α is the learning rate; Let γ be the gradient of the target value with respect to θ1 (i.e., the partial derivative with respect to θ1); γ is the discount factor used to calculate the target value; τ2 is a constant, specifically a decimal close to 0, used to control the update rate. This method can smoothly update the target network, avoiding training instability caused by the target network changing too quickly.
[0175] Policy Network The structure is the same as shown in formula (6). It takes the state (ACE) as input, calculates the mean μ and variance σ through three linear layers, then passes the normal distribution N(μ,σ) and the activation function tanh and multiplies it by the maximum action A, and finally outputs the action (power command) corresponding to the state. parameter Updates require approval According to state st-1 Sampling action, by Calculate the Q-value and sort and filter to obtain state s. t-1 The optimal set of actions
[0176]
[0177] In the formula: s t-1 Let N be the state at time t-1, and N be the sample network at time s. t-1 The total number of actions generated below; ρ is the ratio of optimal actions; a1 to a ρN For the most valuable One action. Among them This indicates rounding down to the nearest integer.
[0178] Then use renew Parameters w1 and Parameter w2. During the update process. Instead of using entropy to avoid policy collapse, To encourage exploration, a higher level of entropy is added. The updated formula is as follows:
[0179]
[0180] In the formula: α is the learning rate; and They are respectively and In state s t-1 The probability of the output action (power deviation); τ1 is the regularization parameter. The larger τ1 is, the higher the entropy and the greater the randomness. The cross-entropy is calculated using the following formula:
[0181]
[0182] Step S48, output power action a, calculate the total power command for the corresponding region at time t, if the information stored in B2 has not exceeded the storage limit:
[0183]
[0184] If the information stored in B2 exceeds the storage limit:
[0185] P t =P t-1 +a t (31)
[0186] Among them, P t For the total power command at time t in the corresponding region, P t-1 This is the total power command for the corresponding region at time t-1.
[0187] If the total power command P output by the deep reinforcement learning controller at time t in the computation region k is... t,k If the information stored in B2 does not exceed the storage limit:
[0188]
[0189] When the information stored in B2 exceeds the storage limit:
[0190] P t,k =P t-1,k +a t,k (33)
[0191] Among them, P t-1,k Let a be the total power command output by the deep reinforcement learning controller in region k at time t-1. t,k For deep reinforcement learning controller policy networks in region k The output is the environmental state S at time t. t The following action For a deep reinforcement learning controller in region k, the environmental state S at time t is selected based on expert experience. t The following action.
[0192] If the policy network output results in a large reward value without significantly increasing it, the policy network training is complete, and the algorithm iteration stops. The algorithm will also stop if the simulation system stops running the corresponding algorithm.
[0193] In the face of new power systems dominated by new energy sources, discrete Q-learning and its derived control algorithms have poor adaptive capabilities, fail to properly store and utilize expert data, and cannot quickly achieve system control of individual generator units when the power system changes, resulting in poor control accuracy. This invention studies continuous deep reinforcement learning, enabling the reasonable storage and utilization of expert data. It can maintain frequency stability in various regions of the distributed power grid under highly stochastic load conditions; improve the distributed power grid's ability to absorb intermittent energy sources such as wind and solar power; and enhance inter-regional coordinated control capabilities.
[0194] Step S5: Based on the total control power of the corresponding area in Step S2 and the power allocation ratio of thermal power units, hydropower units, and electric vehicle clusters, the allocated power of thermal power units, hydropower units, and electric vehicle clusters is obtained, and then the power command of thermal power units, hydropower units, and electric vehicle clusters in the corresponding area is obtained.
[0195] The total power commands for thermal power units, hydropower units, and electric vehicle clusters are obtained according to the following formulas (5), (6), and (19), respectively.
[0196] In step S6, the thermal power unit, hydropower unit, and electric vehicle cluster execute the corresponding power commands and interact with the power grid, then return to step S2 to begin the next cycle.
[0197] The following combination Figure 10 Compared with two different methods, the superior control performance of the proposed controller is demonstrated:
[0198] In power systems, Area Control Error (ACE) is a crucial indicator that measures the degree of imbalance between generation and load within a controlled area. The ACE value is the difference between actual power and planned power. Actual power is the difference between the total generation capacity and total load capacity of the area, while planned power is the power value determined based on day-ahead or real-time market trading plans. ACE is vital for the stable operation of the power system because it directly affects the grid's frequency and voltage levels. A high ACE value in an area indicates an imbalance between generation and load, which may lead to fluctuations in grid frequency and thus affect the stability of the power system. Therefore, power system operators use Automatic Generation Control (AGC) systems to adjust generator output to reduce ACE and maintain grid stability.
[0199] The evaluation of regional power grids relies on the CPS (Circuit Power Response) index. The controller should strive to maintain stable CPS and frequency (f / Hz), with specific indicators as follows:
[0200] 1) If CPS1≥200% and CPS2 is any value, the CPS index is qualified.
[0201] 2) If 100% ≤ CPS1 < 200% and CPS2 ≥ 90%, the CPS index is qualified.
[0202] 3) If CPS1 < 100%, then the CPS index is unqualified.
[0203] 4) If f∈(49.80,50.20)Hz, the frequency index is qualified.
[0204] Please see Figure 10 (a) In order to simulate the load changes of irregular random disturbances in the actual power grid, the present invention introduces irregular random disturbances with strong disturbances to simulate the real power grid environment and performs 24-hour real-time simulation to verify the actual engineering application effect of EL-GAC.
[0205] Please see Figure 10(b) Under random disturbances, the controller proposed based on EL-GAC can consistently maintain the average absolute value of the regional frequency deviation within 0.0003Hz, compared to 0.001Hz for the SAC controller and 0.009Hz for the DQN controller. In actual power grids, the frequency needs to be stable at 50Hz, and the frequency deviation is the difference between the actual frequency and 50Hz; the smaller this value, the better the frequency performance.
[0206] Please see Figure 10 (c) Under random disturbances, the controller proposed based on EL-GAC can always maintain the average absolute value of ACE within 1.06MW, while the SAC controller achieves 4.4MW and the DQN controller achieves 3.4MW. In actual power grids, the smaller the ACE value, the better the performance.
[0207] Please see Figure 10 (d) Under random disturbances, the controller proposed based on EL-GAC can consistently maintain the 10-minute CPS average above 199.939%, the SAC controller above 199.984%, and the DQN controller above 199.977%. In actual power grids, the closer the CPS is to 200%, the better the performance.
[0208] In summary, this invention uses a deep reinforcement learning controller to ensure that the CPS index of a multi-regional power grid meets the requirements and the frequency is stable. Compared with existing reinforcement learning controllers (DQN, SAC), the deep reinforcement learning controller has a simple structure, self-learning capability, and can quickly learn new cooperative control strategies using expert data when the power system changes, thus maintaining good control capabilities.
[0209] The control method of this invention can achieve qualified CPS indicators and frequency stability in multi-regional power grids, large-scale absorption of new energy sources in new power systems, and coordinated control of generator sets through the designed controller.
[0210] Example 2:
[0211] like Figure 11 As shown, based on the same inventive concept as Embodiment 1, this embodiment provides an automatic power generation control system for large-scale grid connection of new energy sources, applied to the method described, including:
[0212] The model building module is used to build a multi-regional interconnected power grid model, initialize model parameters, and integrate the multi-regional interconnected power grid model into the power grid. The multi-regional interconnected power grid model includes deep reinforcement learning controllers, wind turbines, photovoltaic power plants, thermal power units, hydropower units, electric vehicle clusters, and loads in each region. The deep reinforcement learning controllers are connected to the thermal power units, hydropower units, and electric vehicle clusters, respectively.
[0213] The power allocation ratio calculation module is used to calculate the power allocation ratio of thermal power units, hydropower units, and electric vehicle clusters based on the capacity of each region.
[0214] The regional control deviation calculation module is used to calculate the real-time frequency deviation of the power grid and the exchange power between adjacent regions based on the output power of wind turbines, photovoltaic power plants, thermal power plants, hydropower plants, the power of electric vehicle clusters participating in frequency regulation, and the load in each region, and to calculate the regional control deviation of each region.
[0215] The deep reinforcement learning controller is used to input the regional control deviation of each region into the deep reinforcement learning controller of the corresponding region, and iteratively update it through the internal control method of the deep reinforcement learning controller and output the total control power of the corresponding region.
[0216] The total power command calculation module is used to obtain the allocated power of the thermal power units, hydropower units, and electric vehicle clusters based on the total control power of the corresponding area and the power allocation ratio of thermal power units, hydropower units, and electric vehicle clusters, and then obtain the power command of the thermal power units, hydropower units, and electric vehicle clusters in the corresponding area.
[0217] The power grid interaction module is used by thermal power units, hydropower units, and electric vehicle clusters to execute corresponding power commands and interact with the power grid.
[0218] The specific working principle is described in the above method description and will not be repeated here.
[0219] Example 3:
[0220] Based on the same inventive concept as Embodiment 1, this embodiment provides a computer-readable storage medium: the computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to execute the automatic power generation control method for large-scale new energy grid connection.
[0221] Example 4:
[0222] Based on the same inventive concept as Embodiment 1, this embodiment provides a processor for running a program, wherein the program executes the automatic power generation control method for large-scale new energy grid connection.
[0223] Those skilled in the art will recognize that the modules of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components of the examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of the invention.
[0224] In the embodiments provided by this invention, it should be understood that the division of modules is only a logical functional division. In actual implementation, there may be other division methods, such as multiple modules can be combined into one module, one module can be split into multiple modules, or some features can be ignored.
[0225] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0226] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks. Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.
Claims
1. An automatic power generation control method for large-scale grid connection of new energy sources, characterized in that, Includes the following steps: A multi-regional interconnected power grid model is constructed, and the model parameters are initialized. The multi-regional interconnected power grid model is then integrated into the power grid. The multi-regional interconnected power grid model includes a deep reinforcement learning controller for each region, wind turbines, photovoltaic power plants, thermal power units, hydropower units, electric vehicle clusters, and loads. The deep reinforcement learning controller is connected to the thermal power units, hydropower units, and electric vehicle clusters, respectively. Calculate the power allocation ratio of thermal power units, hydropower units, and electric vehicle clusters based on the capacity of thermal power units, hydropower units, and electric vehicle clusters in each region. Based on the output power of wind turbines, photovoltaic power plants, thermal power units, hydropower units, the power of electric vehicle clusters participating in frequency regulation, and the load in each region, the real-time frequency deviation of the power grid in each region and the exchange power between adjacent regions are calculated, and the regional control deviation of each region is calculated. The regional control deviation of each region is input into the deep reinforcement learning controller of the corresponding region. The controller is iteratively updated through its internal control method and outputs the total control power of the corresponding region. Based on the total control power of the corresponding region and the power allocation ratio of thermal power units, hydropower units, and electric vehicle clusters, the allocated power of thermal power units, hydropower units, and electric vehicle clusters is obtained, and then the power command of thermal power units, hydropower units, and electric vehicle clusters in the corresponding region is obtained. Thermal power units, hydropower units, and electric vehicle clusters execute corresponding power commands and interact with the power grid; The calculation methods for the regional control deviation of each area are as follows: Δf k =L -1 {(P g,k +P pv,k +P h,k +P f,k +P ev,k -P L,k -P tie,kj )×K p / (1+sT p )}; P tie,kj =L -1 {(Δf k -Δf j )×(2πT kj / s)}; ACE k =P tie,kj +β k Δf k ; Where, Δf k Let L be the real-time frequency deviation of the power grid in region k. -1 For the inverse Laplace transform, P g,k Let P be the output power of the wind turbine in region k. pv,k Let P be the output power of the photovoltaic power station in region k. h,k Let P be the output power of the hydropower unit in region k. f,k Let P be the output power of the thermal power unit in region k. ev,k P represents the power of the electric vehicle cluster in region k participating in frequency modulation. L,k For the load of region k, P tie,kj The power exchanged between the tie lines of regions k and j, where regions k and j are adjacent, K p T represents the coefficients of the frequency response function. p Let s be the time constant of the frequency response function, s be the Laplace operator, and Δf be the time constant. j Let T be the real-time frequency deviation of the power grid in region j. kj Let be the time constant of the connection line between regions k and j; ACE k For the regional control deviation of region k, β k is the frequency deviation coefficient for region k.
2. The automatic power generation control method for large-scale new energy grid connection according to claim 1, characterized in that, The specific calculation method for the power allocation ratio of thermal power units, hydropower units, and electric vehicle clusters is as follows: In the formula: a f a h a ev These represent the power allocation ratios for thermal power units, hydropower units, and electric vehicle clusters, respectively. f +a h +a ev =1 and a f a h a ev ∈[0,1];E f E h These are the rated capacities of thermal power units and hydropower units, respectively, E ev,all The available capacity for electric vehicle clusters.
3. The automatic power generation control method for large-scale new energy grid connection according to claim 1, characterized in that, The power P of the electric vehicle cluster in region k participating in frequency modulation ev,k Specifically as follows: A controllable dynamic change model for the number of electric vehicles is established, as follows: Let the starting time be t0, the time when electric vehicles in region k connect to the grid be t1, and the charging time of electric vehicles be T2. Therefore, the time t3 when electric vehicles enter operation frequency regulation mode is: t3 = t1 + T2; Electric vehicles determine their eligibility to participate in frequency regulation services based on their own state of charge (SBC). The calculation of the electric vehicle's SBC is as follows: Where: SOC0 is the current state of charge of a single electric vehicle; E ev For the existing capacity of a single electric vehicle; E ev,max This refers to the maximum capacity of a single electric vehicle. The charging time T2 for electric vehicles is calculated using the following formula: in, The charging power for a single electric vehicle; η ev For electric vehicle charging efficiency; SOC min The minimum state of charge required for electric vehicles to participate in frequency regulation; The Monte Carlo method was used to randomly sample the controllable entry / exit timetables of electric vehicles and perform statistical analysis to obtain the number N of electric vehicles in the controllable state at each time point in the region. in '(t i The number of electric vehicles N in a controllable state out '(t i The cumulative number N of electric vehicles entering a controllable state at time t can be calculated using the following formula. in (t) and the cumulative number of electric vehicles N that have exited the controllable state out (t): Let N0 be the initial number of controllable electric vehicles at time t0, and N be the number of controllable electric vehicles at time t. c (t) is shown in the following equation: N c (t)=N0+N in (t)-N out (t); Therefore, the available capacity of the electric vehicle cluster at time t for: When participating in frequency regulation, the total usable capacity of electric vehicles needs to be limited. If the existing capacity is less than the minimum capacity limit, the usable capacity of electric vehicles is zero. Lower limit of the capacity of the electric vehicle cluster at time t for: Capacity limit of the electric vehicle cluster at time t for: Available capacity E of the electric vehicle cluster at time t ev,all The calculation formula is as follows: Calculate a based on the existing capacity. ev The power P of the electric vehicle cluster in region k participating in frequency modulation ev,k The calculation method is as follows: P ev,k =P t,k ·a ev ; Among them, P t,k a is the total power command output by the deep reinforcement learning controller in region k at time t; ev The power allocation ratio for electric vehicle clusters.
4. The automatic power generation control method for large-scale new energy grid connection according to claim 1, characterized in that, The deep reinforcement learning controller includes two policy networks. and 2 evaluation networks The policy network consists of an input layer, a first hidden layer, a second hidden layer, a third hidden layer, a fourth hidden layer, and an output layer. The policy network is used to output action 'a' in state S, where state S represents the area control deviation of the corresponding region, and action 'a' represents the total control power deviation of the corresponding region. The mathematical expression for the policy network is as follows: In the formula: y u y1 and y2 are the inputs to the policy network; y1 and y2 are the outputs of the first and second hidden layers, respectively; μ and logσ are the outputs of the third and fourth hidden layers, respectively. W μ W logσ These are the weight matrices for the first hidden layer, the second hidden layer, the third hidden layer, and the fourth hidden layer, respectively. b μ b logσ Let be the bias vectors of the first, second, third, and fourth hidden layers; the output action 'a' follows a normal distribution N(μ,σ), where the mean μ and standard deviation σ are the mean and standard deviation of the normal distribution followed by the output action 'a', respectively; tanh(x) = (e^(μ,σ)) / (σ) x -e -x ) / (e x +e -x Let A be the activation function and A be the maximum action value. The evaluation network consists of an input layer, a first hidden layer, a second hidden layer, and an output layer. The mathematical expression for evaluating the network is as follows: This is used to determine the value of action a in state S and output a value Q. Where, x u To evaluate the network input, x1 and x2 are the outputs of the first and second hidden layers, respectively; Q is the output of the output layer. These are the weight matrices for the first hidden layer, the second hidden layer, and the output layer, respectively. represents the bias vectors for the first hidden layer, the second hidden layer, and the output layer; relu(x) = max(0,x) is the activation function, ensuring that the output value is not negative.
5. The automatic power generation control method for large-scale new energy grid connection according to claim 1, characterized in that, The control deviation of each region is input into the deep reinforcement learning controller of the corresponding region. The controller is then iteratively updated using its internal control method, and the total control power of the corresponding region is output. The specific steps include: Initialize two evaluation networks and 2 policy networks Set the maximum action value, learning rate α, discount factor γ, entropy regularization coefficient τ1, environmental experience pool B1, and expert experience pool B2 and their upper limit B. 2max Random action N, optimal action proportion coefficient ρ, expert experience proportion ξ; Input the regional control deviation (ACE) value of the corresponding region at time t and use it as the environmental state S at time t. t ; Determine whether the amount of information stored in expert experience pool B2 has reached its upper limit. 2max If the amount of information stored in expert experience pool B2 reaches its upper limit B... 2max Then, through the policy network at time t Output the environmental state S at time t t The following action a t If the amount of information stored in expert experience pool B2 has not reached the upper limit B... 2max Then, based on expert experience, the environmental state S at time t is selected. t The following action Calculate real-time reward r t ; Environmental information is stored in environmental experience buffer pool B1, and expert information is stored in expert experience buffer pool B2. Environmental information includes the environmental state S from the previous time step. t-1 Corresponding real-time strategies action a t-1 The reward r from the previous moment t-1 The environmental state S at this moment t Expert information includes the environmental state S at the previous moment. t-1 Actions selected based on corresponding expert experience Rewards from the previous moment The environmental state S at this moment t ; Based on the proportion of expert experience ξ, samples are randomly drawn from the expert experience buffer pool B2 with equal probability. Each data point is then drawn from the environmental experience buffer pool B1 with the same probability. The two sets of data together constitute the training data, where G is the number of training samples set for mini-batch descent. The parameters of the evaluation network and policy network are updated using the training data, and the proportion of expert experience ξ is updated accordingly. Output power action a, calculate the total power command for the corresponding region at time t.
6. The automatic power generation control method for large-scale new energy grid connection according to claim 5, characterized in that, The specific instructions for calculating the total power in the corresponding region at time t are as follows: If the information stored in expert experience buffer pool B2 has not exceeded the storage limit: If the information stored in expert experience buffer pool B2 exceeds the storage limit: P t =P t-1 +a t ; Among them, P t For the total power command at time t in the corresponding region, P t-1 This is the total power command for the corresponding region at time t-1.
7. An automatic power generation control system for large-scale grid connection of new energy sources, characterized in that, The method applied to any one of claims 1 to 6 includes: The model building module is used to build a multi-regional interconnected power grid model, initialize model parameters, and integrate the multi-regional interconnected power grid model into the power grid. The multi-regional interconnected power grid model includes deep reinforcement learning controllers, wind turbines, photovoltaic power plants, thermal power units, hydropower units, electric vehicle clusters, and loads in each region. The deep reinforcement learning controllers are connected to the thermal power units, hydropower units, and electric vehicle clusters, respectively. The power allocation ratio calculation module is used to calculate the power allocation ratio of thermal power units, hydropower units, and electric vehicle clusters based on the capacity of thermal power units, hydropower units, and electric vehicle clusters in each region. The regional control deviation calculation module is used to calculate the real-time frequency deviation of the power grid and the exchange power between adjacent regions based on the output power of wind turbines, photovoltaic power plants, thermal power units, hydropower units, the power of electric vehicle clusters participating in frequency regulation, and the load in each region, and to calculate the regional control deviation of each region. The deep reinforcement learning controller is used to input the regional control deviation of each region into the deep reinforcement learning controller of the corresponding region, and iteratively update the control power of the corresponding region through the internal control method of the deep reinforcement learning controller. The total power command calculation module is used to obtain the allocated power of the thermal power unit, hydropower unit, and electric vehicle cluster based on the total control power of the corresponding area and the power allocation ratio of the thermal power unit, hydropower unit, and electric vehicle cluster, and then obtain the power command of the thermal power unit, hydropower unit, and electric vehicle cluster in the corresponding area. The power grid interaction module is used by thermal power units, hydropower units, and electric vehicle clusters to execute corresponding power commands and interact with the power grid.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device containing the computer-readable storage medium to perform the automatic power generation control method for large-scale new energy grid connection as described in any one of claims 1 to 6.
9. A processor, characterized in that, The processor is used to run a program, wherein the program executes the automatic power generation control method for large-scale new energy grid connection as described in any one of claims 1 to 6.
Citation Information
Patent Citations
AGC unit dynamic optimization method based on deep reinforcement learning
CN112186811A
KR20220151330A