A multi-regional coordinated control method, system, medium, and processor applicable to large-scale renewable energy grid connection.
By constructing a multi-regional interconnected power grid model and using a reinforcement learning controller, the problem of traditional AGC control methods being unable to coordinate the control of heterogeneous frequency regulation resources in new power systems was solved, thereby improving power grid frequency stability and load balance, as well as enhancing adaptability and response speed.
Patent Information
- Application Number
- CN202411537481.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2044-10-31
AI Technical Summary
Traditional centralized AGC control methods are difficult to effectively coordinate and control heterogeneous frequency regulation resources in new power systems, resulting in low utilization of frequency regulation resources, difficulty in coping with the intermittency and volatility of new energy power generation, and impact on grid frequency stability and load balance.
A multi-regional interconnected power grid model is constructed using a deep reinforcement learning algorithm. The reinforcement learning controller calculates the regional control deviation and outputs the total control power command to coordinate the power allocation of thermal power units, hydropower units, and biomass power units, thereby achieving frequency stability and coordinated control of the distributed power grid.
It improves grid frequency stability and load balance, enhances adaptability and response speed to new energy power generation, and realizes multi-regional collaborative control of distributed grids, which is superior to the coordination and adaptability of PI controllers.
Smart Images

Figure CN119482717B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power generation control technology, and in particular to a multi-regional coordinated control method, system, medium, and processor suitable for large-scale grid connection of new energy sources. Background Technology
[0002] Under the new power system, the power structure will shift from being dominated by coal-fired power generation capacity with controllable and continuous output to being dominated by new energy power generation capacity with strong uncertainty and weak controllability. However, due to the inherent characteristics of new energy power generation, such as intermittency and volatility, it brings unprecedented challenges to the optimization and control of the new power system.
[0003] Automatic generation control (AGC) plays a crucial role in the control and operation of modern interconnected power systems, enabling a balance between connected loads plus losses and total power generation. Traditional centralized AGC suffers from poor dynamic performance, large frequency deviation, and long transient times. In particular, it has a low degree of coordinated control between different regions and struggles to handle the strong random disturbances brought about by the integration of new energy sources. Therefore, distributed AGC is imperative.
[0004] Meanwhile, the development of new power systems will inevitably be accompanied by a significant increase in new energy generating units, which will reduce the rotational inertia of the system. This necessitates exploring various inverter-based resources to support frequency regulation services. Heterogeneous frequency regulation resources exhibit different system models, capacities, and response rates. Furthermore, the shift from traditional large-capacity generators to a series of small-capacity distributed generators (including photovoltaic arrays, wind turbines, energy storage systems, and gas turbines) has significantly increased the number of distributed generators in the power system. Centralized AGC-controlled power plant power allocation can only be executed through provincial dispatch centers, making it difficult to achieve coordinated control of various heterogeneous frequency regulation resources, resulting in low utilization of regulation resources. This further highlights the importance of researching distributed AGC.
[0005] Currently, frequency control strategies are dominated by PI control and PID control, while research on continuous reinforcement learning algorithms as frequency control strategies is limited. Therefore, there is a need for a multi-region coordinated control method, system, medium, and processor that applies reinforcement learning algorithms and is suitable for large-scale grid connection of new energy sources. Summary of the Invention
[0006] To address the shortcomings of traditional AGC control methods in existing technologies, this invention provides a multi-regional coordinated control method, system, medium, and processor suitable for large-scale grid connection of new energy sources, possessing advantages such as strong exploratory capabilities, high adjustment accuracy, and fast response speed. The specific technical solution is as follows:
[0007] A multi-regional coordinated control method applicable to large-scale renewable energy grid connection includes the following steps:
[0008] A multi-regional interconnected power grid model is constructed, and the model parameters are initialized. The multi-regional interconnected power grid model is then integrated into the power grid. The multi-regional interconnected power grid model includes reinforcement learning controllers for each region, wind power plants, photovoltaic power plants, thermal power units, hydropower units, bio-generator units, supercapacitor energy storage units, and loads. The deep reinforcement learning controllers are connected to the thermal power units, hydropower units, and bio-generator units, respectively.
[0009] Based on the output power of wind power plants, photovoltaic power plants, thermal power units, hydropower units, biomass generator units, supercapacitor energy storage units, and load in each region, calculate the real-time frequency deviation of the power grid in each region and the exchange power between adjacent regions, and calculate the regional control deviation of each region.
[0010] The regional control deviation of each region is input into the reinforcement learning controller of the corresponding region. The internal control method of the reinforcement learning controller is used to iteratively update the control power of the corresponding region and output the total control power of the corresponding region.
[0011] Based on the total control power of the corresponding area and the power allocation ratio of thermal power units, hydropower units, and bio-power units, the allocated power of thermal power units, hydropower units, and bio-power units is obtained, and then the power command of thermal power units, hydropower units, and bio-power units in the corresponding area is obtained.
[0012] Thermal power units, hydropower units, and biomass power units execute corresponding power commands and interact with the power grid.
[0013] Preferably, the calculation method for the regional control deviation of each region is as follows:
[0014] Δf i =L -1 {(P f,i +P s,i +P B,i -P C,i -P solar,i -P wind,i -P L,i -P tie,ij )·K p / (1+sT p )};
[0015] P tie,ij =L -1 {(Δf i -Δf j )·(2πT ij / s)};
[0016] ACE i =P tie,ij +B i Δf i ;
[0017] Where, Δf i Let L be the real-time frequency deviation of the power grid in region i. -1 For the inverse Laplace transform, P f,i Let P be the output power of the thermal power unit in region i. s,i P represents the output power of the hydropower unit in region i. B,i Let P be the output power of the bio-generator in region i. C,i P represents the output power of the supercapacitor energy storage unit in region i. solar,i Let P be the output power of the photovoltaic power plant in region i. wind,i Let P be the output power of the wind farm in region i. L,i Let P be the load power of region i. tie,ij The switching power of the tie lines between adjacent regions i and j, K p T represents the coefficients of the frequency response function. p Let s be the time constant of the frequency response function, s be the Laplace operator, and Δf be the time constant. j Let T be the real-time frequency deviation of the power grid in region j. ij Let be the connection line time constant between regions i and j; ACE i β represents the regional control deviation for region i. i Let be the frequency deviation coefficient for region i.
[0018] Preferably, the calculation methods for the output power of thermal power units, hydropower units, and biomass generator units in each region are as follows:
[0019]
[0020] Among them, P f,i P represents the output power of the thermal power units in region i. s,i P represents the output power of the hydropower unit in region i. B,i L represents the output power of the bio-generator in region i; -1 Let s be the inverse Laplace transform, s be the Laplace operator, and P be the inverse Laplace transform. i The reinforcement learning controller for region i outputs the total control power for region i; a f a s a B These are the power allocation ratio coefficients for thermal power units, hydropower units, and biomass power units, respectively; Δf i R represents the real-time frequency deviation of the power grid in region i; f R S R BThese are the droop coefficients for thermal power units, hydropower units, and biomass power units, respectively.
[0021] T g T gh T gb The governor time delay constants for thermal power units, hydropower units, and biomass generator units are respectively; T t T is the time constant of the thermal power unit; rs T rh T w1s These are the time constants of the hydroelectric generator units; T wb is the time constant of the bio-generator.
[0022] Preferably, the calculation method for the supercapacitor energy storage power in each region is as follows:
[0023]
[0024] Among them, P C,i L represents the supercapacitor energy storage power in region i; -1 Let s be the inverse Laplace transform, s be the Laplace operator, and Δf be the inverse Laplace transform. i K represents the real-time frequency deviation of the power grid in region i; C For supercapacitor energy storage units, T1, T2, T3, T4, T c This is the time constant of the supercapacitor energy storage unit.
[0025] Preferably, the reinforcement learning controller includes two policy networks μ(.;Φ1) and μ(.;Φ2) and two evaluation networks Q(.;θ1) and Q(.;θ2). The policy networks include an input layer, a first hidden layer, a second hidden layer, a third hidden layer, and an output layer. The policy networks μ(.;Φ1) and μ(.;Φ2) are used to output action a in state S, where state S is the regional control deviation of the corresponding region, and action a is the output of the total control power deviation of the corresponding region. The mathematical expression of the policy network is as follows:
[0026]
[0027] Where, x u x1 is the input to the policy network; x2 and x3 are the outputs of the first hidden layer, the second hidden layer, and the third hidden layer, respectively. These are the weight matrices for the first hidden layer, the second hidden layer, and the third hidden layer, respectively. Here are the bias vectors for the first, second, and third hidden layers; in the output action a, tanh(x) = (e x -e -x ) / (e x +e -x) is the activation function, which maps the output to [-1, 1]; A is the maximum action value;
[0028] The evaluation network consists of an input layer, a first hidden layer, a second hidden layer, and an output layer. The evaluation networks Q(.;θ1) and Q(.;θ2) are used to evaluate the output action a in state S and output the value Q. The mathematical expression for the evaluation network is as follows:
[0029]
[0030] In the formula, y u To evaluate the network input, x1 and x2 are the outputs of the first and second hidden layers, respectively; Q is the output of the output layer. These are the weight matrices for the first hidden layer, the second hidden layer, and the output layer, respectively. y is the bias vector for the first hidden layer, the second hidden layer, and the output layer; relu(y) = max(0,y) is the activation function, which ensures that the output value is positive.
[0031] Preferably, the process of inputting the regional control deviation of each region into the reinforcement learning controller of the corresponding region, iteratively updating the control power of the corresponding region through the internal control method of the reinforcement learning controller, and outputting the total control power of the corresponding region specifically includes the following steps:
[0032] Initialize two evaluation networks Q(.;θ1) and Q(.;θ2) and two policy networks μ(.;Φ1) and μ(.;Φ2), and set the learning rate α, discount factor γ, number of candidate actions N, number of sub-experience pools m, number of clusters k, maximum action value, and time period T. p Sampling batch b c ;
[0033] Input the area control deviation (ACE) value of the region at time t and use it as the environmental state S at time t. t ;
[0034] The action a at time t is output based on the policy network μ(.;Φ1). t ;
[0035] If time step t is greater than the time step t at the start of training and update start After updating the policy network and evaluation network, the policy network μ(.;Φ1) outputs the action a at time t. t ;
[0036] Time step t is less than or equal to the time step t from the start of training and update. start Then, the action a at time t is directly output from the policy network μ(.;Φ1). t ;
[0037] Based on the action a at time tt The total control power P at time t is calculated. t :
[0038] P t =P t-1 +a t ;
[0039] Among them, P t For the total power command at time t in the corresponding region, P t-1 This represents the total power command for the corresponding region at time t-1.
[0040] Preferably, the specific steps for updating the policy network and the evaluation network are as follows:
[0041] Calculate the reward r at time t t ;
[0042] The environmental state S at time t-1 t-1 Action a at time t-1 t-1 The reward r at time t-1 t-1 The environmental state S at time t t Constitute a sample (s) t-1 ,a t-1 ,r t-1 ,s t Store in the c-th sub-experience pool B c ,in if Then sample b c One sample, and for sub-experience pool B c The samples are clustered into k clusters using k-means clustering, which are then used for policy networks and evaluation network updates.
[0043] if Sample b c One sample is used for updating the policy network and evaluation network;
[0044] After sampling is completed, the policy network μ(s) is used. t ;Φ1) Generate a set of candidate actions and from the candidate action set Select the best action This optimal action is used to calculate the target value of the evaluation network and does not output any interaction with the environment;
[0045] Update the evaluation network based on its target value;
[0046] The updated evaluation network is used to update the strategy network.
[0047] A multi-regional coordinated control system applicable to large-scale renewable energy grid connection, applied to the method described, includes:
[0048] The model building module is used to build a multi-regional interconnected power grid model, initialize model parameters, and integrate the multi-regional interconnected power grid model into the power grid. The multi-regional interconnected power grid model includes reinforcement learning controllers for each region, wind power plants, photovoltaic power plants, thermal power units, hydropower units, bio-generator units, supercapacitor energy storage units, and loads. The deep reinforcement learning controllers are connected to the thermal power units, hydropower units, and bio-generator units, respectively.
[0049] The regional control deviation calculation module is used to calculate the real-time frequency deviation of the power grid and the exchange power between adjacent regions based on the output power of wind power plants, photovoltaic power plants, thermal power units, hydropower units, biomass generator units, supercapacitor energy storage units, and load in each region, and to calculate the regional control deviation of each region.
[0050] The reinforcement learning controller module is used to input the regional control deviation of each region into the reinforcement learning controller of the corresponding region, and iteratively update the control power of the corresponding region through the internal control method of the reinforcement learning controller.
[0051] The power command calculation module is used to obtain the allocated power of the thermal power unit, hydropower unit, and bio-power unit based on the total control power of the corresponding area and the power allocation ratio of the thermal power unit, hydropower unit, and bio-power unit, and then obtain the power command of the thermal power unit, hydropower unit, and bio-power unit in the corresponding area.
[0052] The interaction module is used by thermal power units, hydropower units, and biomass generator units to execute corresponding power commands and interact with the power grid.
[0053] A computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to perform the multi-regional coordinated control method applicable to large-scale new energy grid connection.
[0054] A processor for running a program, wherein the program executes the multi-regional coordinated control method applicable to large-scale renewable energy grid connection.
[0055] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0056] The method of this invention includes the following steps: building a multi-regional interconnected power grid model and initializing the model parameters; integrating the multi-regional interconnected power grid model into the power grid; calculating the regional control deviation of each region; outputting the total control power of the corresponding region through a reinforcement learning controller; obtaining the power commands for the thermal power units, hydropower units, and biomass power units in the corresponding region; and executing the corresponding power commands and interacting with the power grid. The method of this invention is beneficial for maintaining the frequency stability and CPS (Continuous Power Scaling) performance of distributed multi-regional power grids. Compared with PI controllers, it has a simpler structure, stronger adaptability, and higher coordination. In complex operating conditions containing large-scale renewable energy sources, it can respond quickly and accurately to the load, realizing multi-regional coordinated control of the distributed power grid. Attached Figure Description
[0057] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.
[0058] Figure 1 This is a flowchart of the method of the present invention;
[0059] Figure 2 This is a schematic diagram of the model constructed in this embodiment of the invention;
[0060] Figure 3 This is a schematic diagram of the generator unit model;
[0061] Figure 4 Power output curve for the landscape;
[0062] Figure 5 This is a schematic diagram of the structure of a policy network.
[0063] Figure 6 To evaluate the structural principle diagram of the network;
[0064] Figure 7 This is a flowchart of the controller's internal algorithm.
[0065] Figure 8 The diagram shows the power grid control performance indicators for two regions under sinusoidal disturbances. Figure 8 (a) is a graph of frequency deviation under sinusoidal perturbation. Figure 8 (b) is a graph of the 10-minute average ACE under sinusoidal perturbation. Figure 8 (c) is the 10-minute average CPS1 curve under sinusoidal perturbation;
[0066] Figure 9 This is a system schematic diagram of the present invention. Detailed Implementation
[0067] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0068] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0069] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0070] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0071] Example 1:
[0072] like Figure 1 As shown, this embodiment provides a multi-regional coordinated control method suitable for large-scale new energy grid connection, characterized by the following steps:
[0073] Step S1: Build a multi-regional interconnected power grid model, initialize the model parameters, and integrate the multi-regional interconnected power grid model into the power grid; the multi-regional interconnected power grid model includes reinforcement learning controllers, wind power plants, photovoltaic power plants, thermal power units, hydropower units, bio-generator units, supercapacitor energy storage units, and loads in each region; the deep reinforcement learning controllers are connected to the thermal power units, hydropower units, and bio-generator units respectively.
[0074] Please see Figure 2 This invention, based on the background of a new type of power system with large-scale grid integration of new energy sources (wind power, solar power, biomass energy, and supercapacitor energy storage), constructs a distributed two-region interconnected power grid model to verify the reinforcement learning multi-region cooperative strategy designed in this invention. This strategy is mainly applicable to distributed multi-region power grids. The reinforcement learning controller is the CER-AC-TD3 reinforcement learning controller.
[0075] Step S2: Calculate the real-time frequency deviation of the power grid and the exchange power between adjacent areas for each region based on the output power of wind power plants, photovoltaic power plants, thermal power units, hydropower units, biomass generator units, supercapacitor energy storage units, and loads in each region, and calculate the regional control deviation for each region.
[0076] The specific calculation methods for the regional control deviations of each area are as follows:
[0077] Δf i =L -1 {(P f,i +P s,i +P B,i -P C,i -P solar,i -P wind,i -P L,i -P tie,ij )·K p / (1+sT p (1)
[0078] P tie,ij =L -1 {(Δf i -Δf j )·(2πT ij / s)}; (2)
[0079] ACE i =P tie,ij +B i Δf i (3)
[0080] Where, Δf i Let L be the real-time frequency deviation of the power grid in region i. -1 For the inverse Laplace transform, P f,i Let P be the output power of the thermal power unit in region i. s,i P represents the output power of the hydropower unit in region i. B,i Let P be the output power of the bio-generator in region i. C,i P represents the output power of the supercapacitor energy storage unit in region i. solar,i Let P be the output power of the photovoltaic power plant in region i. wind,i Let P be the output power of the wind farm in region i. L,i Let P be the load power of region i. tie,ij The switching power of the tie lines between adjacent regions i and j, K p T represents the coefficients of the frequency response function. p Let s be the time constant of the frequency response function, s be the Laplace operator, and Δf be the time constant. j Let T be the real-time frequency deviation of the power grid in region j.ij Let be the connection line time constant between regions i and j; ACE i β represents the regional control deviation for region i. i Let be the frequency deviation coefficient for region i.
[0081] like Figure 3 As shown, Figure 3 for Figure 2 All unit models included in the document. Generator units are thermal power turbine units, hydropower turbine units, biomass turbine units, and supercapacitor energy storage units, all considering generator output rate control (GRC). The thermal power turbine units, hydropower turbine units, and biomass turbine units include speed governors, GRC, turbines, and power limiting modules. These four modules simulate the actual operation of a reheat turbine unit. The regional power tie lines are AC power tie lines between different regions. When there is an imbalance between power generation and consumption in one region, power from other regions is called upon through these tie lines to maintain the stability of the entire power grid. Supercapacitor energy storage units have high power density, fast response speed, rapid charging and discharging capabilities, and long service life, reducing frequent unit operations and offering significant advantages in participating in grid frequency regulation. To verify the controller's control capabilities, the load is a randomly generated load signal using MATLAB / Simulink.
[0082] The calculation methods for the output power of thermal power units, hydropower units, and biomass generator units in each region are as follows:
[0083]
[0084] Among them, P f,i P represents the output power of the thermal power units in region i. s,i P represents the output power of the hydropower unit in region i. B,i L represents the output power of the bio-generator in region i; -1 Let s be the inverse Laplace transform, s be the Laplace operator, and P be the inverse Laplace transform. i The reinforcement learning controller for region i outputs the total control power for region i; a f a s a B These are the power allocation ratio coefficients for thermal power units, hydropower units, and biomass power units, respectively; Δf i R represents the real-time frequency deviation of the power grid in region i; f R S R B These are the droop coefficients for thermal power units, hydropower units, and biomass power units, respectively.
[0085] T g Tgh T gb The governor time delay constants for thermal power units, hydropower units, and biomass generator units are respectively; T t T is the time constant of the thermal power unit; rs T rh T w1s These are the time constants of the hydroelectric generator units; T wb is the time constant of the bio-generator.
[0086] The calculation methods for the supercapacitor energy storage capacity in each region are as follows:
[0087]
[0088] Among them, P C,i L represents the supercapacitor energy storage power in region i; -1 Let f be the inverse Laplace transform, s be the Laplace operator, and f be the inverse Laplace transform. i K represents the real-time frequency deviation of the power grid in region i; C For supercapacitor energy storage units, T1, T2, T3, T4, T c This is the time constant of the supercapacitor energy storage unit.
[0089] Please see Figure 4 Because wind and solar power are highly volatile and intermittent, the model is simplified and treated as a stochastic load. The wind and solar power output curves are as follows: Figure 4 As shown.
[0090] The power output formula for wind power is shown below:
[0091]
[0092] In the formula: P wind,i , These represent the rated output power of wind farm i in region i; V wind,i , These are the actual wind speed and the rated wind speed, respectively. These are the cut-in wind speed and the cut-out wind speed, respectively.
[0093] The power output formula for a photovoltaic power plant is shown below:
[0094]
[0095] In the formula: c is the rated power generation capacity of the photovoltaic power plant; solar,i T represents the temperature conversion power factor of photovoltaics; T represents the current air temperature; T ref This is a reference temperature value; s solar,i The light intensity at the current moment.
[0096] Step S3: Input the regional control deviation of each region into the reinforcement learning controller of the corresponding region, iteratively update it through the internal control method of the reinforcement learning controller, and output the total control power of the corresponding region.
[0097] The reinforcement learning controller consists of two policy networks μ(.;Φ1) and μ(.;Φ2) and two evaluation networks Q(.;θ1) and Q(.;θ2). The policy networks are responsible for making action decisions, while the evaluation networks evaluate these actions and provide feedback to help the policy networks learn and improve their policies so that they output better actions.
[0098] Please see Figure 5 The policy network includes an input layer, a first hidden layer, a second hidden layer, a third hidden layer, and an output layer. Policy network members μ(.;Φ1) and μ(.;Φ2) are used to output action a in state S, where state S represents the region control deviation, and action a represents the total control power deviation of the corresponding region. The mathematical expression of the policy network is as follows:
[0099]
[0100] Where, x u x1 is the input to the policy network; x2 and x3 are the outputs of the first hidden layer, the second hidden layer, and the third hidden layer, respectively. These are the weight matrices for the first hidden layer, the second hidden layer, and the third hidden layer, respectively. Here are the bias vectors for the first, second, and third hidden layers; in the output action a, tanh(x) = (e x -e -x ) / (e x +e -x ) is the activation function that maps the output to [-1,1]; A is the maximum action value.
[0101] Please see Figure 6 The evaluation network consists of an input layer, a first hidden layer, a second hidden layer, and an output layer. The evaluation networks Q(.;θ1) and Q(·;θ2) are used to evaluate the value of the action a in state S, outputting a value Q. Their inputs are state S and action a, and their output is the value of the state-action pair, used to assess the value of the state-action pair. The mathematical expression for the evaluation network is as follows:
[0102]
[0103] In the formula, y u To evaluate the network input, x1 and x2 are the outputs of the first and second hidden layers, respectively; Q is the output of the output layer. These are the weight matrices for the first hidden layer, the second hidden layer, and the output layer, respectively. y is the bias vector for the first hidden layer, the second hidden layer, and the output layer; relu(y) = max(0,y) is the activation function, which ensures that the output value is positive.
[0104] In the evaluation network, each layer contains multiple neurons, and the connections between layers are achieved by multiplying them by coefficients; the matrix formed by these coefficients is the weight matrix. The input is then processed by the weight matrix and a corresponding bias is added; these biases can be represented as a bias vector. Next, the output is obtained through an activation function and passed to the next layer in the same way, until the final Q-value is obtained.
[0105] Please see Figure 7 The process involves inputting the regional control deviations of each region into the corresponding reinforcement learning controller, iteratively updating the controller using its internal control methods, and outputting the total control power for the corresponding region. Specifically, this includes the following steps:
[0106] Step S31: Initialize two evaluation networks Q(.;θ1) and Q(.;θ2) and two policy networks μ(.;Φ1) and μ(.;Φ2), setting the learning rate α = 0.0003, discount factor γ = 0.99, number of candidate actions N = 32, number of sub-experience pools m = 10, number of clusters k = 2, maximum action value A = 25, and time interval length T. p =10000, sampling batch Where, n min In this embodiment, n represents the minimum number of samples required for each sampling. min =128, n max In this embodiment, n represents the maximum number of samples that can be sampled in each sampling session. max =256.
[0107] Initialize two policy networks μ(·;Φ1), μ(·;Φ2), with the current state s as input and a deterministic action a as output. These networks are used to select actions based on the state. Two evaluation networks Q(·;θ1), Q(·;θ2), with state s and action a as input and the value of the state-action pair as output, are used to evaluate the value of the state-action pair. For each network, initialize its target network θ′1,θ′2,φ′1,φ′2, which is used to delay updates to improve stability. Initialize experience replay buffers B1,B2,…,B m-1 B m It is used to store state-action transition data for interaction with the environment (current state, action, response, next state).
[0108] Step S32: Input the regional control deviation (ACE) value of the region at time t and use it as the environmental state S at time t. t If the actual load is greater than the predetermined load, ACE is positive, indicating that the load needs to be reduced or power generation needs to be increased; if the actual load is less than the predetermined load, ACE is negative, indicating that the load needs to be increased or power generation needs to be reduced.
[0109] Step S33: Output the action a at time t according to the policy network μ(·;Φ1). t .in:
[0110]
[0111] Where μ(;Φ1) represents the output action a of the policy network, and ε represents noise. This indicates that ε follows a Gaussian distribution with a mean of 0 and a standard deviation of δ.
[0112] Step S34, if time step t is greater than the time step t at the start of training and update start After updating the policy network and evaluation network, the policy network μ(.;Φ1) then outputs the action a at time t according to the above formula. t ;
[0113] Step S35, if time step t is less than or equal to the time step t at the start of training update start Then directly output the action a at time t output by the policy network μ(.;Φ1) in step S33. t In this embodiment, t start =256, obtained from numerous experiments, setting it to 256 yields better results;
[0114] Step S36, based on the action a at time t t The total control power P at time t is calculated. t :
[0115] P t =P t-1 +a t (12)
[0116] Among them, P t For the total power command at time t in the corresponding region, P t-1 This represents the total power command for the corresponding region at time t-1.
[0117] The specific steps for updating the policy network and evaluation network are as follows:
[0118] Calculate the reward r at time t t ;
[0119] The environmental state S at time t-1 t-1 Action a at time t-1t-1 The reward r at time t-1 t-1 The environmental state S at time t t Constitute a sample (s) t-1 ,a t-1 ,r t-1 ,s t Store in the c-th sub-experience pool B c ,in if Then sample b c One sample, and for sub-experience pool B c The samples are clustered into k clusters using k-means clustering. The above process is repeated at each time step for updating the policy network and evaluation network.
[0120] if Sample b c One sample is used for updating the policy network and evaluation network;
[0121] The sample sampling probability is as follows: during playback at time step t The sample probability in the equation is defined as follows:
[0122]
[0123] In the formula, j represents the current time period. for Rounding up, f(c) = 2c + 8, f(c) is a function that satisfies The function is defined by the parameters set in the experiment. Simultaneously, the sampling batch size b for time period j is dynamically set. j This ensures that the number of replays for samples explored in any two sub-experience pools is the same during training, i.e., for If 1≤x≤y≤n, b j It can be guaranteed that:
[0124]
[0125] In the formula, For time step jT P Replay The probability of the middle sample. For time step jT P Replay The probability of the middle sample, b j denoted as the sampling batch size for time period j.
[0126] Replaying samples at time step t (s) t-1 ,a t-1 ,r t-1 ,s t ) i The probability is defined as:
[0127]
[0128] In the formula, (s i-1 ,a i-1 ,r i-1 ,s i ) i For the sample corresponding to time step i, The sub-experience pool corresponding to the samples generated at time step i is added. To replay at time step t The probability of the middle sample. The sub-experience pool represents the time step t. The number of samples in Sub-experience pool Cluster tags and (s) i-1 ,a i-1 ,r i-1 ,s i ) i The consistent number of samples, where k represents the number of clusters.
[0129] Sample each sub-experience pool according to the above probabilities;
[0130] After sampling is completed, the policy network μ(s) is used. t ;Φ1) Generate a set of candidate actions and from the candidate action set Select the best action This optimal action is used to calculate the target value of the evaluation network and does not output any interaction with the environment; candidate action set. a i N candidate actions are obtained by adding noise to the output action 'a' of the policy network, and the optimal action is selected according to the following formula.
[0131]
[0132] The objective value of the evaluation network is calculated as follows:
[0133]
[0134] In the formula, γ is the discount factor, and θ1′ and θ2′ are the values calculated by formula (20) in the previous time step, respectively;
[0135] The evaluation network is updated based on its target value; specifically, it is updated according to the following formula:
[0136]
[0137] The updated evaluation network is used to update the policy network. Specifically, every d steps, the policy network is updated according to the following formula.
[0138]
[0139] The policy network is updated to ensure it outputs the optimal action. In the formula, b c This indicates the number of samples taken at time step t. For value Q with respect to θ i The gradient (i.e., with respect to θ) i (Find the partial derivative) For strategy μ to Φ i The gradient (i.e., for Φ) i (Calculate partial derivatives). Finally, update the target network:
[0140] θ′ i ←τθ i +(1-τ)θ′ i (20)
[0141] φ′ i ←τφ i +(1-τ)φ′ i ;(twenty one)
[0142] In the formula, τ is the soft update coefficient, and θ i ,φ i These are the current values of the policy network and the evaluation network, respectively.
[0143] The above process uses an action candidate mechanism to improve the accuracy of Q-value function estimation and improves stability through policy delay updates and updating the target network. The Q-value is used to evaluate the quality of the agent's actions, i.e., whether the output power deviation is accurate and whether a smaller ACE can be obtained. Improving the accuracy of Q-value function estimation enables the agent to output more accurate power commands. The termination condition is a manually set time; the training automatically terminates after the set time is reached.
[0144] The CER-AC-TD3 reinforcement learning controller is used to provide specific power generation values for the generating units to maintain grid frequency stability and control performance standard (CPS) indicators. The establishment of new power systems will inevitably lead to a more diversified and complex grid structure. Discrete reinforcement learning algorithms, when facing systems with high-dimensional and continuous state and action spaces, cannot meet high-precision power control requirements at different load sections. Furthermore, the agent struggles to perceive more characteristic state information from the environment, resulting in less than ideal convergence speed and control accuracy for algorithms based on discrete action sets. The controller described is a new generation of artificial intelligence control method based on continuous reinforcement learning algorithms. It can adapt to the development requirements of new power systems and has stronger self-updating and adaptive capabilities. It takes Area Control Error (ACE) as input and outputs action 'a' to calculate the total power command of the generator set. The target Q-value is updated using real-time feedback data ACE, Δf, and dynamic sampling of experience pool data. The controller selects the optimal action to approximate the best action value through an action candidate mechanism and dynamically sets the sampling batch, thereby reducing underestimation of the error and improving the learning speed. Simultaneously, clustering algorithms and experience replay mechanisms are placed in a divide-and-conquer framework to effectively mine experience samples hidden in all exploration transitions during training and fully replay various types of samples. The experience pool is a collection of historical data, a two-dimensional matrix of system parameters (current state, action, reward, next state). Dynamic sampling of historical data can effectively mine experience samples hidden in all exploration transitions during training, improving the accuracy of Q-value updates and obtaining more accurate and effective control strategies. ACE is a comprehensive performance index of power deviation and frequency deviation. It is a direct reflection of regional control indicators and the quality of real-time operation. As long as ACE is within the acceptable range, the power grid indicators will tend to be stable.
[0145] Step S4: Based on the total control power of the corresponding area and the power allocation ratio of the thermal power unit, hydropower unit, and bio-power unit, the allocated power of the thermal power unit, hydropower unit, and bio-power unit is obtained, thereby obtaining the power command of the thermal power unit, hydropower unit, and bio-power unit in the corresponding area. The output power of the thermal power unit, hydropower unit, and bio-power unit is calculated according to formulas (4)-(6). In this embodiment, a f a S a B The values are 0.8, 0.1, and 0.1 respectively.
[0146] In step S5, the thermal power unit, hydropower unit, and biomass generator unit execute the corresponding power commands and interact with the power grid.
[0147] Step S6: Return to step S2 above and calculate the regional control deviation of each area in real time according to formulas (1)-(3).
[0148] Please see Figure 8 The following is combined Figure 8 Compared with two different methods, the superior control performance of the proposed controller is demonstrated:
[0149] Please see Figure 8 (a) Under sinusoidal disturbance, the average absolute value of the regional frequency deviation of the controller proposed by CER-AC-TD3 is within 0.0011Hz, the AC-TD3 controller is 0.0023Hz, and the DQN controller is 0.0029Hz. The smaller the regional frequency deviation, the better. The proposed algorithm reduces the frequency deviation by 52.17% to 62.07% compared with other algorithms.
[0150] Please see Figure 8 (b) Under sinusoidal disturbance, the controller proposed based on CER-AC-TD3 can always keep the average absolute value of the ACE deviation at 0.166MW, while the AC-TD3 controller is 2.721MW and the DQN controller is 21.195MW. The smaller the area control deviation, the better. The proposed algorithm reduces the deviation by 93.90% to 99.22% compared with other algorithms.
[0151] Please see Figure 8 (c) Under sinusoidal disturbance, the average CPS of the controller proposed by CER-AC-TD3 in 10 minutes is 199.99%, the ACTD3 controller is 199.91%, and the DQN controller is 198.77%. Compared with other algorithms, it is closest to the evaluation standard of 200%.
[0152] In summary, it can be seen that the DQN algorithm, limited by the predefined action and state space of discrete algorithms, makes it difficult for the agent to perceive more feature state information from the environment, resulting in poor control performance. The controller based on the CER-AC-TD3 algorithm of this invention exhibits significantly improved control performance compared to other controllers, possessing higher frequency stability, higher control accuracy, and the ability to achieve multi-regional coordination in distributed power grids, thus better adapting to the development needs of new power systems. The designed CER-AC-TD3 controller can achieve frequency stability and coordinated control of distributed multi-regional power grids.
[0153] The CER-AC-TD3 controller developed in this invention can maintain frequency stability and CPS index compliance in distributed multi-regional power grids. Compared with PI controllers, it has a simpler structure, stronger adaptability, and higher coordination. In complex operating conditions with large-scale new energy sources, it can respond quickly and accurately to loads and realize multi-regional collaborative control of distributed power grids.
[0154] Example 2:
[0155] like Figure 9 As shown, based on the same inventive concept as Embodiment 1, this embodiment provides a multi-regional coordinated control system suitable for large-scale new energy grid connection, applied to the method described, including:
[0156] The model building module is used to build a multi-regional interconnected power grid model, initialize model parameters, and integrate the multi-regional interconnected power grid model into the power grid. The multi-regional interconnected power grid model includes reinforcement learning controllers for each region, wind power plants, photovoltaic power plants, thermal power units, hydropower units, bio-generator units, supercapacitor energy storage units, and loads. The deep reinforcement learning controllers are connected to the thermal power units, hydropower units, and bio-generator units, respectively.
[0157] The regional control deviation calculation module is used to calculate the real-time frequency deviation of the power grid and the exchange power between adjacent regions based on the output power of wind power plants, photovoltaic power plants, thermal power units, hydropower units, biomass generator units, supercapacitor energy storage units, and load in each region, and to calculate the regional control deviation of each region.
[0158] The reinforcement learning controller module is used to input the regional control deviation of each region into the reinforcement learning controller of the corresponding region, and iteratively update the control power of the corresponding region through the internal control method of the reinforcement learning controller.
[0159] The power command calculation module is used to obtain the allocated power of the thermal power unit, hydropower unit, and bio-power unit based on the total control power of the corresponding area and the power allocation ratio of the thermal power unit, hydropower unit, and bio-power unit, and then obtain the power command of the thermal power unit, hydropower unit, and bio-power unit in the corresponding area.
[0160] The interaction module is used by thermal power units, hydropower units, and biomass generator units to execute corresponding power commands and interact with the power grid.
[0161] The specific working principle is described in the method description above and will not be repeated here.
[0162] Example 3:
[0163] Based on the same inventive concept as Embodiment 1, this embodiment provides a computer-readable storage medium, which includes a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to execute the multi-regional coordinated control method applicable to large-scale new energy grid connection.
[0164] Example 4:
[0165] Based on the same inventive concept as Embodiment 1, this embodiment provides a processor for running a program, wherein the program executes the multi-regional coordinated control method applicable to large-scale new energy grid connection.
[0166] Those skilled in the art will recognize that the modules of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components of the examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of the invention.
[0167] In the embodiments provided by this invention, it should be understood that the division of modules is only a logical functional division. In actual implementation, there may be other division methods, such as multiple modules can be combined into one module, one module can be split into multiple modules, or some features can be ignored.
[0168] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0169] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.
Claims
1. A multi-regional coordinated control method applicable to large-scale renewable energy grid connection, characterized in that, Includes the following steps: A multi-regional interconnected power grid model is constructed, and the model parameters are initialized. The multi-regional interconnected power grid model is then integrated into the power grid. The multi-regional interconnected power grid model includes reinforcement learning controllers for each region, wind power plants, photovoltaic power plants, thermal power units, hydropower units, bio-generator units, supercapacitor energy storage units, and loads. The reinforcement learning controllers are connected to the thermal power units, hydropower units, and bio-generator units, respectively. Based on the output power of wind power plants, photovoltaic power plants, thermal power units, hydropower units, biomass generator units, supercapacitor energy storage units, and load in each region, calculate the real-time frequency deviation of the power grid in each region and the exchange power between adjacent regions, and calculate the regional control deviation of each region. The regional control deviation of each region is input into the reinforcement learning controller of the corresponding region. The internal control method of the reinforcement learning controller is used to iteratively update the control power of the corresponding region and output the total control power of the corresponding region. Based on the total control power of the corresponding area and the power allocation ratio of thermal power units, hydropower units, and bio-power units, the allocated power of thermal power units, hydropower units, and bio-power units is obtained, and then the power command of thermal power units, hydropower units, and bio-power units in the corresponding area is obtained. Thermal power units, hydropower units, and biomass power units execute corresponding power commands and interact with the power grid; The specific calculation methods for the regional control deviations of each area are as follows: Δf i =L -1 {(P f,i +P s,i +P B,i -P C,i -P solar,i -P wind,i -P L,i -P tie,ij )·K p / (1+sT p )}; P tie,ij =L -1 {(Δf i -Δf j )·(2πT ij / s)}; ACE i =P tie,ij +B i Δf i ; Where, Δf i Let L be the real-time frequency deviation of the power grid in region i. -1 For the inverse Laplace transform, P f,i Let P be the output power of the thermal power unit in region i. s,i P represents the output power of the hydropower unit in region i. B,i Let P be the output power of the bio-generator in region i. C,i P represents the output power of the supercapacitor energy storage unit in region i. solar,i Let P be the output power of the photovoltaic power plant in region i. wind,i Let P be the output power of the wind farm in region i. L,i Let P be the load power of region i. tie,ij The switching power of the tie lines between adjacent regions i and j, K p T represents the coefficients of the frequency response function. p Let s be the time constant of the frequency response function, s be the Laplace operator, and Δf be the time constant. j Let T be the real-time frequency deviation of the power grid in region j. ij Let be the time constant of the connection line between regions i and j; ACE i B represents the regional control deviation for region i. i Let be the frequency deviation coefficient for region i.
2. The multi-regional coordinated control method applicable to large-scale new energy grid connection according to claim 1, characterized in that, The calculation methods for the output power of thermal power units, hydropower units, and biomass generator units in each region are as follows: Among them, P f,i P represents the output power of the thermal power units in region i. s,i P represents the output power of the hydropower unit in region i. B,i L represents the output power of the bio-generator in region i; -1 Let s be the inverse Laplace transform, s be the Laplace operator, and P be the inverse Laplace transform. i The reinforcement learning controller for region i outputs the total control power for region i; a f a s a B These are the power allocation ratio coefficients for thermal power units, hydropower units, and biomass power units, respectively; Δf i R represents the real-time frequency deviation of the power grid in region i; f R S R B These are the droop coefficients for thermal power units, hydropower units, and biomass power units, respectively. T g T gh T gb The governor time delay constants for thermal power units, hydropower units, and biomass generator units are respectively; T t T is the time constant of the thermal power unit; rs T rh T w1s These are the time constants of the hydroelectric generator units; T wb is the time constant of the bio-generator.
3. The multi-regional coordinated control method applicable to large-scale new energy grid connection according to claim 1, characterized in that, The calculation methods for the supercapacitor energy storage capacity in each region are as follows: Among them, P C,i L represents the supercapacitor energy storage power in region i; -1 Let s be the inverse Laplace transform, s be the Laplace operator, and Δf be the inverse Laplace transform. i K represents the real-time frequency deviation of the power grid in region i. C For supercapacitor energy storage units, T1, T2, T3, T4, T c This is the time constant of the supercapacitor energy storage unit.
4. The multi-regional coordinated control method applicable to large-scale new energy grid connection according to claim 1, characterized in that, The reinforcement learning controller includes two policy networks μ(.;Φ1) and μ(.;Φ2) and two evaluation networks Q(.;θ1) and Q(.;θ2). Each policy network comprises an input layer, a first hidden layer, a second hidden layer, a third hidden layer, and an output layer. Policy networks μ(.;Φ1) and μ(.;Φ2) output action a in state S, where state S represents the regional control deviation of the corresponding region, and action a represents the total control power deviation of the corresponding region. The mathematical expression for the policy network is as follows: Where, x u x1 is the input to the policy network; x2 and x3 are the outputs of the first hidden layer, the second hidden layer, and the third hidden layer, respectively. These are the weight matrices for the first hidden layer, the second hidden layer, and the third hidden layer, respectively. Here are the bias vectors for the first, second, and third hidden layers; in the output action a, tanh(x) = (e x -e -x ) / (e x +e -x ) is the activation function that maps the output to [-1, 1]; A is the maximum action value; The evaluation network consists of an input layer, a first hidden layer, a second hidden layer, and an output layer. The evaluation networks Q(·θ1) and Q(·θ2) are used to evaluate the output action a in state S and output the value Q. The mathematical expression for the evaluation network is as follows: In the formula, y u To evaluate the network input, x1 and x2 are the outputs of the first and second hidden layers, respectively; Q is the output of the output layer. These are the weight matrices for the first hidden layer, the second hidden layer, and the output layer, respectively. y is the bias vector for the first hidden layer, the second hidden layer, and the output layer; relu(y) = max(0,y) is the activation function, which ensures that the output value is positive.
5. A multi-regional coordinated control method applicable to large-scale new energy grid connection according to claim 1, characterized in that, The control deviation of each region is input into the reinforcement learning controller of the corresponding region. The reinforcement learning controller is then iteratively updated using its internal control method, and the total control power of the corresponding region is output. The specific steps include: Initialize two evaluation networks Q(.;θ1) and Q(.;θ2) and two policy networks μ(.;Φ1) and μ(.;Φ2), and set the learning rate α, discount factor γ, number of candidate actions N, number of sub-experience pools m, number of clusters k, maximum action value, and time period T. p Sampling batch b c ; Input the area control deviation (ACE) value of the region at time t and use it as the environmental state S at time t. t ; The action a at time t is output based on the policy network μ(.;Φ1). t ; If time step t is greater than the time step t at the start of training and update start After updating the policy network and evaluation network, the policy network μ(.;Φ1) outputs the action a at time t. t ; Time step t is less than or equal to the time step t from the start of training and update. start Then, the action a at time t is directly output from the policy network μ(.;Φ1). t ; Based on the action a at time t t The total control power P at time t is calculated. t : P t =P t-1 +a t ; Among them, P t For the total power command at time t in the corresponding region, P t-1 This is the total power command for the corresponding region at time t-1.
6. A multi-regional coordinated control method applicable to large-scale new energy grid connection according to claim 5, characterized in that, The specific steps for updating the policy network and evaluation network are as follows: Calculate the reward r at time t t ; The environmental state S at time t-1 t-1 Action a at time t-1 t-1 The reward r at time t-1 t-1 The environmental state S at time t t Constitute a sample (s) t-1 ,a t-1 ,r t-1 ,s t Stored into the c-th sub-experience pool B c ,in if Then sample b c One sample, and for sub-experience pool B c The samples are clustered into k clusters using k-means clustering, which are then used for policy networks and evaluation network updates. if Sample b c One sample is used for updating the policy network and evaluation network; After sampling is completed, the policy network μ(s) is used. t ;Φ1) Generate a set of candidate actions and from the candidate action set Select the best action This optimal action is used to calculate the target value of the evaluation network and does not output any interaction with the environment; Update the evaluation network based on its target value; The updated evaluation network is used to update the strategy network.
7. A multi-regional coordinated control system suitable for large-scale grid connection of new energy sources, characterized in that, The method applied to any one of claims 1 to 6 includes: The model building module is used to build a multi-regional interconnected power grid model, initialize model parameters, and integrate the multi-regional interconnected power grid model into the power grid. The multi-regional interconnected power grid model includes reinforcement learning controllers for each region, wind power plants, photovoltaic power plants, thermal power units, hydropower units, bio-generator units, supercapacitor energy storage units, and loads. The reinforcement learning controllers are connected to the thermal power units, hydropower units, and bio-generator units, respectively. The regional control deviation calculation module is used to calculate the real-time frequency deviation of the power grid and the exchange power between adjacent regions based on the output power of wind power plants, photovoltaic power plants, thermal power units, hydropower units, biomass generator units, supercapacitor energy storage units, and load in each region, and to calculate the regional control deviation of each region. The reinforcement learning controller module is used to input the regional control deviation of each region into the reinforcement learning controller of the corresponding region, and iteratively update the control power of the corresponding region through the internal control method of the reinforcement learning controller. The power command calculation module is used to obtain the allocated power of the thermal power unit, hydropower unit, and bio-power unit based on the total control power of the corresponding area and the power allocation ratio of the thermal power unit, hydropower unit, and bio-power unit, and then obtain the power command of the thermal power unit, hydropower unit, and bio-power unit in the corresponding area. The interaction module is used by thermal power units, hydropower units, and biomass generator units to execute corresponding power commands and interact with the power grid.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device containing the computer-readable storage medium to perform the multi-regional coordinated control method applicable to large-scale new energy grid connection as described in any one of claims 1 to 6.
9. A processor, characterized in that, The processor is used to run a program, wherein the program executes the multi-regional coordinated control method applicable to large-scale new energy grid connection as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Frequency control method and system for interconnected power system containing energy storage resources
CN112531792A
Dynamic interaction adjustment control method suitable for flexible resources of micro-grid
CN117913911A