An island micro-grid frequency control method and system based on reachability perception reinforcement learning
Patent Information
- Application Number
- CN202610834394.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-10
- Publication Date
- 2026-09-22
AI Technical Summary
[0004]首先,强化学习智能体的决策主要依赖于有限的交互轨迹,由于无法覆盖所有可能的不确定扰动场景,其在不同扰动实现下的鲁棒性往往不足,难以保证在复杂信息物理扰动下的频率稳定性
[0107]本发明的技术效果是毋庸置疑的,本发明通过将可达性分析获得的系统动态边界几何特征直接嵌入强化学习智能体的状态空间与奖励函数,使智能体能够显式感知系统在多场景下的演化趋势和不确定性波动范围,从而有效弥补传统强化学习仅依赖有限交互轨迹、对不确定扰动鲁棒性不足的固有缺陷。在此框架下,智能体可依据可达集的中心跟踪并抑制频率偏差趋势,同时依据可达集的半径主动压缩瞬态波动包络,实现对频率动态的双向优化。
Smart Images

Figure CN122801299A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power engineering, specifically to a frequency control method and system for island microgrids based on reachability-aware reinforcement learning. Background Technology
[0002] Networked microgrids (NMs) provide an efficient, clean, and flexible power supply solution for islands by integrating distributed energy sources such as photovoltaics and wind power, as well as energy storage systems. However, the replacement of traditional synchronous generators with numerous inverter interfaces in microgrids leads to a significant reduction in the system's equivalent inertia, making dynamic frequency stability extremely vulnerable. More critically, as a typical cyber-physical system, the tightly coupled nature of microgrids inevitably exposes them to complex cyber-physical disturbances (CPDs). At the physical layer, intermittent generation from renewable energy sources and random load fluctuations cause frequent active power disturbances; at the information layer, communication delays, communication failures, and signal distortions (such as electromagnetic interference, channel noise, and spoofing attacks) severely degrade the accuracy of secondary control signals. The combined effects of these disturbances exhibit high uncertainty, posing a serious challenge to the frequency stability of networked microgrids. Therefore, a robust control method that can effectively cope with uncertain cyber-physical disturbances and ensure the frequency stability of island-based networked microgrids is urgently needed.
[0003] To suppress frequency instability, virtual synchronous generator technology is widely used to dynamically adjust system inertia and damping. However, achieving online adaptive parameter scheduling under highly uncertain operating conditions remains a key challenge. Traditional robust control and fixed-gain proportional-integral control strategies are typically conservative and struggle to maintain optimal dynamic adjustment performance in complex time-varying scenarios. In recent years, data-driven reinforcement learning methods have been attempted for online gain scheduling, such as adaptively adjusting control parameters to cope with communication delays. However, these methods still suffer from the following problems:
[0004] First, the decision-making of reinforcement learning agents mainly relies on a limited number of interaction trajectories. Since they cannot cover all possible uncertain perturbation scenarios, their robustness under different perturbation implementations is often insufficient, and it is difficult to guarantee frequency stability under complex cyber-physical perturbations.
[0005] Secondly, reachability analysis techniques can rigorously quantify dynamic boundaries by calculating the complete set of reachable states of a system under uncertainty. However, existing methods are mainly used for offline verification and lack explicit mechanisms to use these calculated boundaries for real-time parameter scheduling.
[0006] Finally, there is currently a lack of a closed-loop control architecture that can directly embed the reachable set into the decision-making process of reinforcement learning, enabling the agent to perceive the dynamic boundaries of the system, thereby achieving robust online parameter optimization. Summary of the Invention
[0007] The purpose of this invention is to provide a frequency control method for island microgrids based on reachability-aware reinforcement learning, comprising the following steps:
[0008] Step 1) Consider the uncertainties of the island-connected microgrid and construct a continuous dynamic representation of the microgrid;
[0009] Step 2) The agent is trained offline based on the continuous dynamic representation of the networked microgrid to obtain a trained agent; the trained agent takes the geometric features of the predicted reachable set at the current time as input and the optimal frequency control parameters at the next time as output.
[0010] Step 3) Deploy the trained agent on the online controller;
[0011] Use a trained agent to predict a prediction period The optimal frequency control parameters within the time frame are used to generate a sequence of actions arranged in chronological order.
[0012] The controller executes the action sequence in chronological order to achieve frequency control of the island's interconnected microgrid.
[0013] Furthermore, in step 1), the continuous dynamic representation of the interconnected microgrid is as follows:
[0014] (1)
[0015] In the formula, This represents the total state vector of the interconnected microgrid. The differential of the total state vector of the interconnected microgrid. For uncertain disturbances, The state matrix, The input matrix;
[0016] The total state vector of the interconnected microgrid is shown below:
[0017] (2)
[0018] In a networked microgrid with N sub-microgrids, the vector of the i-th sub-microgrid. As shown below:
[0019] (3)
[0020] In the formula, For frequency deviation, The active power deviation of the speed controller. This refers to the active power deviation of the generator. For the offset of the switching power of the tie line, This is to address the control deviation in the automatic power generation control signal area. These are auxiliary variables; i = 1, 2, ..., N;
[0021] The uncertainty disturbance is shown below:
[0022] (4)
[0023] In the formula, Let i be the uncertainty disturbance of the i-th sub-microgrid in the N sub-microgrids of the interconnected microgrid.
[0024] Furthermore, the frequency deviation dynamics of sub-microgrid i are described below:
[0025] (5)
[0026] In the formula, For the frequency deviation of sub-microgrid i, For the equivalent system inertia, This is the equivalent damping coefficient. This refers to the active power deviation of the generator. For aggregated power disturbances within the microgrid, For the offset of the switching power of the tie line;
[0027] The dynamic description of the switching power deviation of the tie line is as follows:
[0028] (6)
[0029] In the formula, The synchronous torque coefficient between sub-microgrids i and j is given. For the frequency deviation of sub-microgrid j;
[0030] The auxiliary variable is dynamically described as follows:
[0031] (7)
[0032] In the formula, For delay parameters; This is the frequency deviation coefficient; As an auxiliary variable;
[0033] The integral dynamic description of the control deviation in the automatic generation control signal area is as follows:
[0034] (8)
[0035] In the formula, The impact of injecting signal distortion into false data on the regional control deviation of automatic power generation control signals;
[0036] The dynamic description of the governor and the steam turbine is as follows:
[0037] (9)
[0038] (10)
[0039] In the formula, This is to address the control deviation in the automatic power generation control signal area. The time constant of the speed controller, For the proportional parameters of the PI controller, These are the integral parameters of the PI controller. The time constant of the steam turbine. This is the droop parameter.
[0040] Furthermore, in step 2), the optimal frequency control parameters include the inertia of the grid-connected microgrid, the droop coefficient, and the integral gain of the PI controller.
[0041] Furthermore, the steps for offline training of the agent include:
[0042] Step 2.1) Construct a probabilistic set of perturbation scenarios based on the type and intensity of cyber-physical perturbations, as shown below:
[0043] (11)
[0044] In the formula, Let j be the perturbation pattern corresponding to the j-th scene, and its probability of occurrence is... Perform agent initialization, setting the initial value of Episode to 0;
[0045] Step 2.2) Allow the agent to train for a number of rounds. ;
[0046] judge If the condition is met, end the training and output the trained agent; otherwise, reset the system parameters and the initial reachable set, and let... ;
[0047] Step 2.3) Let ,judge ( Check if the maximum number of iterations is true. If yes, proceed to step 2.8; otherwise, proceed to step 2.4.
[0048] Step 2.4) Extract the geometric features of the multi-scene reachable set at the current time: extract all The center vector and radius vector of each scenario are concatenated in sequence to form the current state vector of the agent;
[0049] Step 2.5) Sample the continuous actions based on the agent's current state vector, and the agent outputs actions to update the optimal frequency control parameters of the networked microgrid;
[0050] Step 2.6) For each scenario j, perform reachability analysis based on the updated control parameters, calculate the reachable set for the next time period, and use the reachable set at the terminal time as the initial set for the next time step;
[0051] Step 2.7) Calculate the total reward for the n scenarios and cache the data, then return to step 2.3).
[0052] Step 2.8) Update the Actor network and Critic network based on the cached data, clear the cache, and return to step 2.2).
[0053] Furthermore, in step 2.4), the global center vector and radius vector of the multi-scene reachable set at the current time under each scene j are as follows:
[0054] (12)
[0055] (13)
[0056] In the formula, , For each scenario j, extract the key physical variables of each sub-microgrid i from the reachability set. The midpoint and radius of the range of variation;
[0057] The agent's current state vector is shown below:
[0058] (14)
[0059] In the formula, , These are the global center vector and radius vector in scene j, respectively;
[0060] In step 2.5), the continuous actions are as follows:
[0061] (15)
[0062] (16)
[0063] In the formula, For the Actor network, H is the inertia of the interconnected microgrid, R is the droop coefficient, and K is the integral gain of the PI controller;
[0064] The feasible range constraints for the inertia, droop factor, and integral gain of the PI controller in the gridded microgrid are as follows:
[0065] (17)
[0066] In the formula, These are the upper and lower limits for inertia adjustment. These are the upper and lower limits for adjusting the droop parameter. These are the upper and lower limits for adjusting the integral parameters of the PI controller.
[0067] Furthermore, in step 2.6), the reachability analysis steps include:
[0068] The continuous dynamic representation equation of the gridded microgrid is linearized using Taylor expansion to obtain an approximate linear system equation.
[0069] The reachable set at the next time step is determined based on the approximate linear system equations;
[0070] Repeat the steps to construct the reachable set within the control cycle;
[0071] The approximate linear system equations are shown below:
[0072] (18)
[0073] In the formula, This is an approximate coefficient matrix; For generalized perturbation inputs that include linearization error and uncertainty; This is the state deviation of the system state variables relative to the linearization point; This is Minkowski addition.
[0074] The reachable set for the next time step is shown below:
[0075] (19)
[0076] In the formula, For the set that is reachable at the current moment, For time step, The reachable set portion determined by the center point of the input set. For the portion of the reachable set determined by the generalized uncertain input, The center point of the input set;
[0077] The reachable set over a continuous time interval is shown below:
[0078] (20);
[0079] In the formula, This is a convex hull operation;
[0080] The reachable set during the control period is shown below:
[0081] (twenty one)
[0082] In the formula, It is the reachable set within a continuous time interval;
[0083] In step 2.7), the total rewards for the n scenarios are as follows:
[0084] (twenty two)
[0085] In the formula, Let be the probability of scene j. For single-scene rewards under scene j;
[0086] The single-scene reward for scene j is shown below:
[0087] (twenty three)
[0088] In the formula, , These are the center and radius vectors of the frequency deviation, respectively. and Preset weighting coefficients;
[0089] Data caching refers to transferring samples Store in the experience cache.
[0090] Furthermore, in step 2.9), updating the Actor network refers to maximizing the pruning objective function. As shown below:
[0091] (twenty four)
[0092] In the formula, For Actor network parameters, For the expectation of sample k in the batch, To trim hyperparameters, The dominant function;
[0093] Updating the Critic network refers to minimizing the mean square error. As shown below:
[0094] (25)
[0095] In the formula, For the state The estimated value, The regression target is calculated from the actual cumulative discount rewards.
[0096] Furthermore, in step 3), the steps for generating the action sequence arranged in chronological order include:
[0097] Step 3.1) In each prediction period Initially, the real-time operating point of the sampled networked microgrid is used as the initial state, and then... ;
[0098] Step 3.2) Let ,judge If the condition is met, proceed to step 3.7; otherwise, proceed to step 3.3.
[0099] Step 3.3) Based on the current interconnected microgrid model, calculate... Time period available;
[0100] Step 3.4) Extraction The center vector and radius vector of the time-reachable set are input into the trained agent;
[0101] Step 3.5) The agent outputs the optimal continuous action to update the system parameters, that is, to update the optimal frequency control parameters of the networked microgrid model;
[0102] Step 3.6) Add the optimal continuous action to the action sequence arranged in chronological order, and return to step 3.2).
[0103] Step 3.7) Output the entire prediction period An internal sequence of actions arranged in chronological order and a predictable reachable set. A system based on any of the aforementioned island microgrid frequency control methods includes a system model and uncertainty description module, an offline training module for the agent, and an online application module for the agent;
[0104] The system model and uncertainty description module is used to construct an automatic generation control model for the frequency regulation problem of island-connected microgrids;
[0105] The offline training module for intelligent agents is used to obtain trained intelligent agents, enabling the agents to output optimal frequency control parameters based on the current dynamic boundary characteristics of the system, including the microgrid's inertia, droop coefficient, and PI controller integral gain.
[0106] The agent online application module is used to deploy the trained agent on the online controller. It adopts a two-stage strategy of forward prediction and action execution to ensure that the actual frequency trajectory is restricted to the predicted reachable set.
[0107] The technical effectiveness of this invention is undeniable. By directly embedding the dynamic boundary geometric features of the system obtained from reachability analysis into the state space and reward function of the reinforcement learning agent, this invention enables the agent to explicitly perceive the evolution trend and uncertainty fluctuation range of the system in multiple scenarios. This effectively overcomes the inherent shortcomings of traditional reinforcement learning, which relies only on finite interaction trajectories and lacks robustness to uncertain perturbations. Within this framework, the agent can track and suppress frequency deviation trends based on the center of the reachable set, while actively compressing the transient fluctuation envelope based on the radius of the reachable set, achieving bidirectional optimization of frequency dynamics.
[0108] In the agent training phase, this invention employs a multi-scenario training strategy, enabling the agent to learn fully under a set of scenarios encompassing various complex cyber-physical disturbances such as power fluctuations, signal distortion, communication delay variations, and communication failures. This ultimately yields a robust control strategy capable of adapting to different types of uncertain disturbances. The strategy, after offline convergence, exhibits excellent generalization ability, maintaining effective frequency regulation performance even when faced with specific disturbances not encountered during training.
[0109] In the online application phase, this invention employs a two-stage control architecture of "forward prediction – action execution". In the forward prediction phase, the controller, based on real-time sampled states and the system model, pre-synthesizes the action sequence for the entire prediction period through reachability set calculation and agent decision-making. In the execution phase, the controller applies the actions sequentially. This strategy provides a formalized robustness guarantee: as long as the actual disturbance does not exceed the pre-defined zonotope uncertainty boundary set during offline modeling, the system's actual frequency trajectory will always remain within the predicted reachability set, thus ensuring frequency stability.
[0110] In summary, this invention deeply integrates model-driven reachability analysis with data-driven reinforcement learning, which improves robustness to uncertain information physical disturbances while achieving online adaptive optimization of control parameters. This significantly improves the dynamic frequency response performance of gridded microgrids and provides effective technical support for the safe and stable operation of microgrid systems such as those on islands under complex disturbance environments. Attached Figure Description
[0111] Figure 1 This is a framework diagram of microgrid i in the system model of the present invention;
[0112] Figure 2 This is a flowchart of the offline training process for the intelligent agent of the present invention;
[0113] Figure 3 This is a flowchart of the online application of the intelligent agent of the present invention. Detailed Implementation
[0114] The present invention will be further described below with reference to embodiments, but it should not be construed that the scope of the present invention is limited to the following embodiments. Various substitutions and modifications made based on ordinary technical knowledge and common practices in the art without departing from the above-described technical concept of the present invention should be included within the scope of protection of the present invention.
[0115] Example 1:
[0116] See Figures 1 to 3 A frequency control method for island microgrids based on reachability-aware reinforcement learning includes the following steps:
[0117] Step 1) Consider the uncertainties of the island-connected microgrid and construct a continuous dynamic representation of the microgrid;
[0118] Step 2) The agent is trained offline based on the continuous dynamic representation of the networked microgrid to obtain a trained agent; the trained agent takes the geometric features of the predicted reachable set at the current time as input and the optimal frequency control parameters at the next time as output.
[0119] Step 3) Deploy the trained agent on the online controller; use the trained agent to predict for one prediction period. The optimal frequency control parameters within the time frame are used to generate a sequence of actions arranged in chronological order.
[0120] At the predicted time Within this range, the scope of the uncertainty disturbance is fixed and is initialized starting from time 0.
[0121] The initialization settings also include the "real-time status of the system," and these two parts are the initial conditions for performing "reachability analysis"; for example, time. Divide the system into several intervals with t as the unit. At time 0, use the "real-time state of the system" as the starting point for prediction and perform reachability analysis between 0s and ts to obtain the reachable set of the analysis results within this time period. Input the reachable set at time ts into the agent and the agent outputs actions to update the system parameters.
[0122] The prediction of the second time interval takes the "reachable set of ts" as the starting point for prediction. A reachability analysis is performed between ts and 2ts to obtain the reachable set of the analysis results within this time interval. The reachable set at time 2ts is input into the agent, and the agent outputs an action to update the system parameters.
[0123] Subsequent steps follow the same pattern. Uncertainty and disturbance remain constant. Except for time 0, which is a sampled value, the starting point of each prediction time is based on the result of the previous prediction.
[0124] The controller executes the action sequence in chronological order to achieve frequency control of the island's interconnected microgrid.
[0125] When the actual disturbance exceeds the preset Zonotope boundary, the reachable set obtained by forward prediction will not be able to achieve an absolute envelope of the actual frequency trajectory. Although the reinforcement learning agent itself has a certain generalization and adaptability, if there is a significant deviation between the actual disturbance conditions and the training scenario, the original agent's control strategy will be difficult to ensure system stability, and it is necessary to retrain it. To avoid such situations, the disturbance boundary setting of Zonotope can be actively expanded during the offline training phase to cover a wider range of more extreme actual operating conditions, thereby improving the overall robustness of the system in advance.
[0126] Example 2:
[0127] The main structure of this embodiment is the same as that of Embodiment 1. Further, in step 1), the continuous dynamic representation of the networked microgrid is as follows:
[0128] (1)
[0129] In the formula, This represents the total state vector of the interconnected microgrid. The differential of the total state vector of the interconnected microgrid. Uncertainty disturbances include power fluctuations and signal distortion. The state matrix, The input matrix;
[0130] The total state vector of the interconnected microgrid is shown below:
[0131] (2)
[0132] In a networked microgrid with N sub-microgrids, the vector of the i-th sub-microgrid. As shown below:
[0133] (3)
[0134] In the formula, For frequency deviation, The active power deviation of the speed controller. This refers to the active power deviation of the generator. For the offset of the switching power of the tie line, This is to address the control deviation in the automatic power generation control signal area. These are auxiliary variables; i = 1, 2, ..., N;
[0135] The uncertainty disturbance is shown below:
[0136] (4)
[0137] In the formula, Let i be the uncertainty disturbance of the i-th sub-microgrid in the N sub-microgrids of the interconnected microgrid.
[0138] Example 3:
[0139] The main structure of this embodiment is the same as any one of embodiments 1-2. Furthermore, the dynamic description of the frequency deviation of sub-microgrid i is as follows:
[0140] (5)
[0141] In the formula, For the frequency deviation of sub-microgrid i, For the equivalent system inertia, This is the equivalent damping coefficient. This refers to the active power deviation of the generator. For aggregated power disturbances within the microgrid, For the offset of the switching power of the tie line;
[0142] The dynamic description of the switching power deviation of the tie line is as follows:
[0143] (6)
[0144] In the formula, The synchronous torque coefficient between sub-microgrids i and j is given. For the frequency deviation of sub-microgrid j;
[0145] The auxiliary variable is dynamically described as follows:
[0146] (7)
[0147] In the formula, For delay parameters; This is the frequency deviation coefficient; As an auxiliary variable;
[0148] The integral dynamic description of the control deviation in the automatic generation control signal area is as follows:
[0149] (8)
[0150] In the formula, The impact of injecting signal distortion into false data on the regional control deviation of automatic power generation control signals;
[0151] The dynamic description of the governor and the steam turbine is as follows:
[0152] (9)
[0153] (10)
[0154] In the formula, This is to address the control deviation in the automatic power generation control signal area. The time constant of the speed controller, For the proportional parameters of the PI controller, These are the integral parameters of the PI controller. The time constant of the steam turbine. This is the droop parameter.
[0155] Example 4:
[0156] The main structure of this embodiment is the same as any one of embodiments 1-3. Further, in step 2), the optimal frequency control parameters include the inertia of the networked microgrid, the droop coefficient, and the integral gain of the PI controller.
[0157] Example 5:
[0158] The main structure of this embodiment is the same as any one of embodiments 1-4. Furthermore, the steps for offline training of the agent include:
[0159] Step 2.1) Construct a probabilistic set of perturbation scenarios based on the type and intensity of cyber-physical perturbations, as shown below:
[0160] (11)
[0161] In the formula, Let j be the perturbation pattern corresponding to the j-th scene, and its probability of occurrence is... Perform agent initialization, setting the initial value of Episode to 0;
[0162] Step 2.2) Allow the agent to train for a number of rounds. ;
[0163] judge If the condition is met, end the training and output the trained agent; otherwise, reset the system parameters and the initial reachable set, and let... ;
[0164] Step 2.3) Let ,judge ( Check if the maximum number of iterations is true. If yes, proceed to step 2.8; otherwise, proceed to step 2.4.
[0165] Step 2.4) Extract the geometric features of the multi-scene reachable set at the current time: extract all The center vector and radius vector of each scenario are concatenated in sequence to form the current state vector of the agent;
[0166] Step 2.5) Sample the continuous actions based on the agent's current state vector, and the agent outputs actions to update the optimal frequency control parameters of the networked microgrid;
[0167] Step 2.6) For each scenario j, perform reachability analysis based on the updated control parameters, calculate the reachable set for the next time period, and use the reachable set at the terminal time as the initial set for the next time step;
[0168] Step 2.7) Calculate the total reward for the n scenarios and cache the data, then return to step 2.3).
[0169] Step 2.8) Update the Actor network and Critic network based on the cached data, clear the cache, and return to step 2.2).
[0170] Example 6:
[0171] The main structure of this embodiment is the same as any one of embodiments 1-5. Further, in step 2.4), the global center vector and radius vector of the multi-scene reachable set at the current time under each scene j are as follows:
[0172] (12)
[0173] (13)
[0174] In the formula, , For each scenario j, extract the key physical variables of each sub-microgrid i from the reachability set. The midpoint and radius of the range of variation;
[0175] The agent's current state vector is shown below:
[0176] (14)
[0177] In the formula, , These are the global center vector and radius vector in scene j, respectively;
[0178] In step 2.5), the continuous actions are as follows:
[0179] (15)
[0180] (16)
[0181] In the formula, For the Actor network, H is the inertia of the interconnected microgrid, R is the droop coefficient, and K is the integral gain of the PI controller;
[0182] The feasible range constraints for the inertia, droop factor, and integral gain of the PI controller in the gridded microgrid are as follows:
[0183] (17)
[0184] In the formula, These are the upper and lower limits for inertia adjustment. These are the upper and lower limits for adjusting the droop parameter. These are the upper and lower limits for adjusting the integral parameters of the PI controller.
[0185] Example 7:
[0186] The main structure of this embodiment is the same as any one of embodiments 1-6. Further, in step 2.6), the reachability analysis step includes:
[0187] The continuous dynamic representation equation of the gridded microgrid is linearized using Taylor expansion to obtain an approximate linear system equation.
[0188] The reachable set at the next time step is determined based on the approximate linear system equations;
[0189] Repeat the steps to construct the reachable set within the control cycle;
[0190] The approximate linear system equations are shown below:
[0191] (18)
[0192] In the formula, This is an approximate coefficient matrix; For generalized perturbation inputs that include linearization error and uncertainty; This is the state deviation of the system state variables relative to the linearization point; This is Minkowski addition.
[0193] The reachable set for the next time step is shown below:
[0194] (19)
[0195] In the formula, For the set that is reachable at the current moment, For time step, The reachable set portion determined by the center point of the input set. For the portion of the reachable set determined by the generalized uncertain input, The center point of the input set;
[0196] The reachable set over a continuous time interval is shown below:
[0197] (20);
[0198] In the formula, This is a convex hull operation;
[0199] The reachable set during the control period is shown below:
[0200] (twenty one)
[0201] In the formula, It is the reachable set within a continuous time interval;
[0202] In step 2.7), the total rewards for the n scenarios are as follows:
[0203] (twenty two)
[0204] In the formula, Let be the probability of scene j. For single-scene rewards under scene j;
[0205] The single-scene reward for scene j is shown below:
[0206] (twenty three)
[0207] In the formula, , These are the center and radius vectors of the frequency deviation, respectively. and Preset weighting coefficients;
[0208] Data caching refers to transferring samples Store in the experience cache.
[0209] Example 8:
[0210] The main structure of this embodiment is the same as any one of embodiments 1-7. Further, in step 2.9), updating the Actor network means maximizing the pruning objective function. As shown below:
[0211] (twenty four)
[0212] In the formula, For Actor network parameters, For the expectation of sample k in the batch, To trim hyperparameters, The dominant function;
[0213] Updating the Critic network refers to minimizing the mean square error. As shown below:
[0214] (25)
[0215] In the formula, For the state The estimated value, The regression target is calculated from the actual cumulative discount rewards.
[0216] Example 9:
[0217] The main structure of this embodiment is the same as any one of embodiments 1-8. Further, in step 3), the step of generating the action sequence arranged in chronological order includes:
[0218] Step 3.1) In each prediction period Initially, the real-time operating point of the sampled networked microgrid is used as the initial state, and then... ;
[0219] Step 3.2) Let ,judge If the condition is met, proceed to step 3.7; otherwise, proceed to step 3.3.
[0220] Step 3.3) Based on the current interconnected microgrid model, calculate... Time period available;
[0221] Step 3.4) Extraction The center vector and radius vector of the time-reachable set are input into the trained agent;
[0222] Step 3.5) The agent outputs the optimal continuous action to update the system parameters, that is, to update the optimal frequency control parameters of the networked microgrid model;
[0223] Step 3.6) Add the optimal continuous action to the action sequence arranged in chronological order, and return to step 3.2).
[0224] Step 3.7) Output the entire prediction period The action sequence arranged in chronological order and the predictable reachable set.
[0225] Example 10:
[0226] A system based on the island microgrid frequency control method described in any one of Examples 1-9 includes a system model and uncertainty description module, an agent offline training module, and an agent online application module.
[0227] The system model and uncertainty description module is used to construct an automatic generation control model for the frequency regulation problem of island-connected microgrids;
[0228] The offline training module for intelligent agents is used to obtain trained intelligent agents, enabling the agents to output optimal frequency control parameters based on the current dynamic boundary characteristics of the system, including the microgrid's inertia, droop coefficient, and PI controller integral gain.
[0229] The agent online application module is used to deploy the trained agent on the online controller. It adopts a two-stage strategy of forward prediction and action execution to ensure that the actual frequency trajectory is restricted to the predicted reachable set.
[0230] Example 11:
[0231] This invention proposes a robust frequency control method for networked microgrids based on reachability-aware reinforcement learning. This method addresses the frequency instability problem of networked microgrid systems in scenarios such as islands under cyber-physical disturbances by embedding the dynamic boundary geometric features calculated through reachability analysis into the reinforcement learning decision-making process, thereby achieving online adaptive optimization of control parameters.
[0232] This method includes the following main modules:
[0233] 1. System model and uncertainty description module;
[0234] 2. Offline training module for intelligent agents;
[0235] 3. Online application module for intelligent agents.
[0236] Specific implementation steps:
[0237] Step 1: System Model and Uncertainty Description
[0238] This invention addresses the frequency regulation problem of gridded microgrids in an island setting, and studies its automatic generation control model. The considered gridded microgrid system consists of N sub-microgrids interconnected by tie lines. The frequency dynamics of each sub-microgrid i can be described by the following oscillation equation:
[0239]
[0240] in, This is for frequency deviation; For the equivalent system inertia, This is the equivalent damping coefficient; This refers to the generator's active power deviation. Aggregated power disturbances (including power generation and load fluctuations) within a microgrid. For the alternating power deviation of the tie line, its dynamics satisfy:
[0241]
[0242] in, Let be the synchronization torque coefficient between microgrids i and j. A first-order Padé approximation is used to account for the potential delay in area control error (ACE) of the automatic generation control signal. Impact, introducing auxiliary variables :
[0243]
[0244] in For delay parameters; This is the frequency deviation coefficient. Furthermore, the impact of signal distortions such as FDI on ACE is comprehensively considered as follows: The ACE expression is:
[0245]
[0246] The dynamics of the governor and turbine are described as a first-order inertial element, expressed as follows:
[0247]
[0248] Combining the above equations, the continuous dynamics of NMs can be written in state-space form:
[0249]
[0250] The overall state vector is composed of the vectors of each microgrid. It is pieced together to form a whole. .enter Uncertainty disturbances considered in this study, including power fluctuations and signal distortion, are modeled as Zonotope and used in the accessibility analysis.
[0251] Step 2: Offline training of the agent
[0252] The goal of offline training is to obtain a robust control policy that enables the agent to output optimal frequency control parameters, including the microgrid's inertia, droop factor, and PI integral gain, based on the current dynamic boundary characteristics of the system.
[0253] Based on the type and intensity of cyber-physical disturbances, a set of probabilistic disturbance scenarios is discretized and constructed:
[0254]
[0255] Each scene For a specific perturbation mode, their probabilities satisfy... At each control step The system performs reachability analysis for each scenario j, and the reachability set is calculated based on Zonotope uncertainty representation and linearized approximation method. Let the reachability set of the current system state be... The input uncertain set is The system is dynamically linearized using Taylor expansion, resulting in an approximately linear system:
[0256]
[0257] in, This is an approximate coefficient matrix; Let the input be a generalized perturbation containing linearization error and uncertainty. Then the reachable set at the next time step is:
[0258]
[0259] in, Determined by the center point of the input set, Determined by uncertain input. The reachable set within a continuous time interval is obtained through convex hull operations:
[0260]
[0261] The reachable set within the control period is:
[0262]
[0263] For each scenario j, from the reachable set Extracting the key physical variables of each microgrid i ( The range of variation of ). Define the center vector. The radius vector represents the midpoint of the range of variation of this variable. This represents the half-width of the variable's range of variation. By concatenating the centers and radii of all microgrids, the global center vector for scenario j is obtained. and radius vector :
[0264]
[0265] Concatenate the center vectors and radius vectors from all n scenarios in order to form the agent's state vector:
[0266]
[0267] This state explicitly encodes the dynamic evolution trend and uncertainty fluctuation range of the system under multiple scenarios, enabling the agent to possess "reachability awareness" capabilities. The agent employs a proximal policy optimization algorithm, and its Actor network... Based on the current state Sampling obtains continuous motion :
[0268]
[0269] action Includes three adjustable control parameters for each microgrid i:
[0270]
[0271] Where H is the system inertia, R is the droop coefficient, and K is the integral gain of the PI controller. All parameters are constrained within physically feasible limits.
[0272]
[0273] Actions This is applied to the system model to update the corresponding parameters in each dynamic matrix. Then, for each scenario j, the reachable set for the next time period is recalculated, and the reachable set at the terminal time is used as the initial set for the next time period.
[0274] The reward function is designed to simultaneously suppress the offset trend and fluctuation range of frequency deviation. The single-scene reward for scene j is:
[0275]
[0276] in, and These are the center and radius vectors of the frequency deviation, respectively (the frequency-corresponding dimension is extracted from the global center / radius). and
[0277] These are preset weighting coefficients. The total reward is the weighted sum of the probabilities of each scenario:
[0278]
[0279] Transfer sample Store in the experience cache. As the cache accumulates... After the first step, the discounted reward and generalized advantage estimate are calculated, and the Actor network and Critic network are updated using the PPO algorithm:
[0280] Actor network update: Maximize the pruning objective function ;
[0281] Critic network update: Minimize mean square error .
[0282] Repeat the above process until the policy converges, and output the trained robust policy.
[0283] Step 3: Online Application of Intelligent Agents
[0284] The online controller deploys the trained agent and adopts a two-stage strategy of "forward prediction + action execution" to ensure that the actual frequency trajectory is restricted to the predicted reachable set.
[0285] In each forecast period Initially, the real-time operating point of the sampled microgrid is used as the initial state. Evenly divided into multiple control sub-intervals For the k-th sub-interval :
[0286] Based on the current system model, calculate the reachable set within this sub-interval;
[0287] Extract the geometric features (center and radius) of the reachable set and input them into the trained agent;
[0288] The agent outputs the optimal action for that sub-interval;
[0289] Add the action to the action sequence and update the system model parameters with the action;
[0290] Use the terminal reachable set as the initial set for the next sub-interval.
[0291] Repeat the above process to pre-synthesize the entire prediction period. The controllable reachable set and the corresponding action sequence within it.
[0292] In actual operation, the controller executes the pre-synthesized action sequence sequentially in chronological order, adjusting the inertia, droop coefficient, and PI integral gain of each microgrid in real time. Since the effectiveness of the action sequence has been verified based on the model and uncertainty boundary during the forward prediction stage, as long as the actual disturbance does not exceed the pre-set zonotope boundary during offline modeling, the actual frequency trajectory of the system will be strictly limited within the pre-calculated controllable reachable set, thereby ensuring frequency stability.
[0293] Each forecast period After execution, the controller resamples the real-time state and enters the next prediction cycle to achieve rolling optimization control.
Claims
1. A frequency control method for island microgrids based on reachability-aware reinforcement learning, characterized in that, Includes the following steps: Step 1) Consider the uncertainties of the island-connected microgrid and construct a continuous dynamic representation of the microgrid; Step 2) The agent is trained offline based on the continuous dynamic representation of the networked microgrid to obtain a trained agent; the trained agent takes the geometric features of the predicted reachable set at the current time as input and the optimal frequency control parameters at the next time as output. Step 3) Deploy the trained agent on the online controller; Use a trained agent to predict a prediction period The optimal frequency control parameters within the time frame are used to generate a sequence of actions arranged in chronological order. The controller executes the action sequence in chronological order to achieve frequency control of the island's interconnected microgrid.
2. The frequency control method for island microgrids based on reachability-aware reinforcement learning according to claim 1, characterized in that, In step 1), the continuous dynamic representation of the interconnected microgrid is shown below: (1) In the formula, This represents the total state vector of the interconnected microgrid. The differential of the total state vector of the interconnected microgrid. For uncertain disturbances, The state matrix, The input matrix; The total state vector of the interconnected microgrid is shown below: (2) In a networked microgrid with N sub-microgrids, the vector of the i-th sub-microgrid. As shown below: (3) In the formula, For frequency deviation, The active power deviation of the speed controller. This refers to the active power deviation of the generator. For the offset of the switching power of the tie line, This is to address the control deviation in the automatic power generation control signal area. These are auxiliary variables; i = 1, 2, ..., N; The uncertainty disturbance is shown below: (4) In the formula, Let i be the uncertainty disturbance of the i-th sub-microgrid in the N sub-microgrids of the interconnected microgrid.
3. The frequency control method for island microgrids based on reachability-aware reinforcement learning according to claim 2, characterized in that, The frequency deviation dynamics of sub-microgrid i are described below: (5) In the formula, For the frequency deviation of sub-microgrid i, For the equivalent system inertia, This is the equivalent damping coefficient. This refers to the active power deviation of the generator. For aggregated power disturbances within the microgrid, For the offset of the switching power of the tie line; The dynamic description of the switching power deviation of the tie line is as follows: (6) In the formula, The synchronous torque coefficient between sub-microgrids i and j is given. For the frequency deviation of sub-microgrid j; The auxiliary variable is dynamically described as follows: (7) In the formula, For delay parameters; This is the frequency deviation coefficient; As an auxiliary variable; The integral dynamic description of the control deviation in the automatic generation control signal area is as follows: (8) In the formula, The impact of injecting signal distortion into false data on the regional control deviation of automatic power generation control signals; The dynamic description of the governor and the steam turbine is as follows: (9) (10) In the formula, This is to address the control deviation in the automatic power generation control signal area. The time constant of the speed controller, For the proportional parameters of the PI controller, These are the integral parameters of the PI controller. The time constant of the steam turbine. This is the droop parameter.
4. The frequency control method for island microgrids based on reachability-aware reinforcement learning according to claim 1, characterized in that, In step 2), the optimal frequency control parameters include the inertia of the grid-connected microgrid, the droop coefficient, and the integral gain of the PI controller.
5. The frequency control method for island microgrids based on reachability-aware reinforcement learning according to claim 4, characterized in that, The steps for offline training of an agent include: Step 2.1) Construct a probabilistic set of perturbation scenarios based on the type and intensity of cyber-physical perturbations, as shown below: (11) In the formula, Let j be the perturbation pattern corresponding to the j-th scene, and its probability of occurrence is... Perform agent initialization, setting the initial value of Episode to 0; Step 2.2) Allow the agent to train for a number of rounds. ; judge If the condition is met, end the training and output the trained agent; otherwise, reset the system parameters and the initial reachable set, and let... ; Step 2.3) Let ,judge ( Check if the maximum number of iterations is true. If yes, proceed to step 2.8; otherwise, proceed to step 2.
4. Step 2.4) Extract the geometric features of the multi-scene reachable set at the current time: extract all The center vector and radius vector of each scenario are concatenated in sequence to form the current state vector of the agent; Step 2.5) Sample the continuous actions based on the agent's current state vector, and the agent outputs actions to update the optimal frequency control parameters of the networked microgrid; Step 2.6) For each scenario j, perform reachability analysis based on the updated control parameters, calculate the reachable set for the next time period, and use the reachable set at the terminal time as the initial set for the next time step; Step 2.7) Calculate the total reward for the n scenarios and cache the data, then return to step 2.3). Step 2.8) Update the Actor network and Critic network based on the cached data, clear the cache, and return to step 2.2).
6. The frequency control method for island microgrids based on reachability-aware reinforcement learning according to claim 5, characterized in that, In step 2.4), the global center vector and radius vector of the multi-scene reachable set at the current time under each scene j are as follows: (12) (13) In the formula, , For each scenario j, extract the key physical variables of each sub-microgrid i from the reachability set. The midpoint and radius of the range of variation; The agent's current state vector is shown below: (14) In the formula, , These are the global center vector and radius vector in scene j, respectively; In step 2.5), the continuous actions are as follows: (15) (16) In the formula, For the Actor network, H is the inertia of the interconnected microgrid, R is the droop coefficient, and K is the integral gain of the PI controller; The feasible range constraints for the inertia, droop factor, and integral gain of the PI controller in the gridded microgrid are as follows: (17) In the formula, These are the upper and lower limits for inertia adjustment. These are the upper and lower limits for adjusting the droop parameter. These are the upper and lower limits for adjusting the integral parameters of the PI controller.
7. The frequency control method for island microgrids based on reachability-aware reinforcement learning according to claim 6, characterized in that, In step 2.6), the reachability analysis steps include: The continuous dynamic representation equation of the gridded microgrid is linearized using Taylor expansion to obtain an approximate linear system equation. The reachable set at the next time step is determined based on the approximate linear system equations; Repeat the steps to construct the reachable set within the control cycle; The approximate linear system equations are shown below: (18) In the formula, This is an approximate coefficient matrix; For generalized perturbation inputs that include linearization error and uncertainty; This is the state deviation of the system state variables relative to the linearization point; This is Minkowski addition. The reachable set for the next time step is shown below: (19) In the formula, For the set that is reachable at the current moment, For time step, The reachable set portion determined by the center point of the input set. For the portion of the reachable set determined by the generalized uncertain input, The center point of the input set; The reachable set over a continuous time interval is shown below: (20); In the formula, This is a convex hull operation; The reachable set during the control period is shown below: (21) In the formula, It is the reachable set within a continuous time interval; In step 2.7), the total rewards for the n scenarios are as follows: (22) In the formula, Let be the probability of scene j. For single-scene rewards under scene j; The single-scene reward for scene j is shown below: (23) In the formula, , These are the center and radius vectors of the frequency deviation, respectively. and Preset weighting coefficients; Data caching refers to transferring samples Store in the experience cache.
8. The frequency control method for island microgrids based on reachability-aware reinforcement learning according to claim 7, characterized in that, In step 2.9), updating the Actor network refers to maximizing the pruning objective function. As shown below: (24) In the formula, For Actor network parameters, For the expectation of sample k in the batch, To trim hyperparameters, The dominant function; Updating the Critic network refers to minimizing the mean square error. As shown below: (25) In the formula, For the state The estimated value, The regression target is calculated from the actual cumulative discount rewards.
9. The frequency control method for island microgrids based on reachability-aware reinforcement learning according to claim 1, characterized in that, Step 3), generating the action sequence arranged in chronological order, includes the following steps: Step 3.1) In each prediction period Initially, the real-time operating point of the sampled networked microgrid is used as the initial state, and then... ; Step 3.2) Let ,judge If the condition is met, proceed to step 3.7; otherwise, proceed to step 3.
3. Step 3.3) Based on the current interconnected microgrid model, calculate... Time period available; Step 3.4) Extraction The center vector and radius vector of the time-reachable set are input into the trained agent; Step 3.5) The agent outputs the optimal continuous action to update the system parameters, that is, to update the optimal frequency control parameters of the networked microgrid model; Step 3.6) Add the optimal continuous action to the action sequence arranged in chronological order, and return to step 3.2). Step 3.7) Output the entire prediction period The action sequence arranged in chronological order and the predictable reachable set.
10. A system based on the frequency control method for island microgrids according to any one of claims 1-9, characterized in that, It includes a system model and uncertainty description module, an offline training module for intelligent agents, and an online application module for intelligent agents; The system model and uncertainty description module is used to construct an automatic generation control model for the frequency regulation problem of island-connected microgrids; The offline training module for intelligent agents is used to obtain trained intelligent agents, enabling the agents to output optimal frequency control parameters based on the current dynamic boundary characteristics of the system, including the microgrid's inertia, droop coefficient, and PI controller integral gain. The agent online application module is used to deploy the trained agent on the online controller. It adopts a two-stage strategy of forward prediction and action execution to ensure that the actual frequency trajectory is restricted to the predicted reachable set.