Micro-grid control parameter robust optimization method considering uncertainty of system parameters
By constructing the downorder model of the isolated AC microgrid and the adversarial Markov decision-making process, the joint adversarial soft actor-commentator algorithm is used to optimize the control parameters, and the stability and dynamic performance problems of the microgrid system under parameter uncertainty are solved, achieving efficient and robust optimization.
Patent Information
- Application Number
- CN202510384107.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-11
AI Technical Summary
When facing uncertainty in system parameters, the existing microgrid control strategies have low computational efficiency and insufficient damping characteristics, which leads to the impact of the dynamic performance and stability of the system. They rely on theoretical parameters rather than on-site determined parameters and cannot reflect the actual operating status.
The siloed AC microgrid down-order model is constructed, the fast and slow variables are separated using feature analysis method, combined with the singular perturbation theory to reduce the order, and the control parameters are optimized through the adversarial Markov decision-making process and the joint adversarial soft actor-commentator algorithm to ensure the system damping characteristics and robustness.
Improves the damping characteristics and robustness of microgrid systems, ensuring stability and dynamic performance in the face of parameter uncertainty, and provides high-quality operation support for inverter control.
Smart Images

Figure CN120300889A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power grid control, and specifically to a robust optimization method for microgrid control parameters considering system parameter uncertainty. Background Art
[0002] With the aggravation of the global energy crisis and environmental pollution problems, the development and utilization of renewable energy have become an important way to achieve sustainable development. Due to non-renewability and environmental destructiveness, traditional fossil energy has gradually been replaced by clean energy such as wind energy and solar energy. In 2023, clean energy developed vigorously, and compared with the previous year, the installed capacity of renewable energy in the global energy system increased by 50%.
[0003] As a flexible energy management system, the microgrid improves the comprehensive utilization efficiency of new energy, expands the utilization ways of new energy, and gradually becomes one of the most popular new power systems. In the island mode, the microgrid needs to independently provide the support of the internal voltage and frequency of the system. No matter what control strategy is adopted, the key to the stable control of the microgrid system voltage and frequency lies in the control performance of the inverter. Coupled with the fact that the microgrid system is easily affected by various disturbances, it makes the high-quality control of microgrid electric energy complicated and also places higher requirements on the underlying control of the inverter. Therefore, considering the model mismatch problem caused by parameter uncertainty is of great significance for the practical application of existing microgrid control strategies.
[0004] Although in-depth research has been carried out on the stability analysis of the microgrid, there are still some problems in the actual application of the existing technologies. First of all, most of the current research is carried out based on the full-order mathematical model of the microgrid. However, in the field of power systems, with the increasing complexity of large-scale system models, the model order is also increasing continuously, resulting in low efficiency of calculation, solution and simulation, which hinders the research of microgrid control strategies. Secondly, most of the existing parameter optimizations of AC islanded microgrids based on droop control focus on the size of the real part of the characteristic roots to judge the stability of the system. By analyzing the real part trajectory of the dominant characteristic roots of the microgrid, the droop control parameters are optimized to improve the system stability margin. However, the damping characteristics of the system are not fully considered, and at the same time, the information of the imaginary part of the characteristic roots is not fully used. Poor damping characteristics may lead to insufficient damping or even undamped oscillation of the system, which will seriously affect the dynamic performance of the microgrid system after suffering small disturbances, thus destroying the system stability. Finally, the research on microgrid parameter regulation extremely depends on the accuracy of the system topology and parameters. However, the parameters of the microgrid in most research are theoretical parameters rather than on-site determined parameters. Parameter drifts caused by component aging, temperature rise, parasitic capacitance, and inductance parameters make the theoretical parameters of the constructed model deviate from the actual operating parameters of the system and cannot reflect the true operating state of the system. Summary of the Invention
[0005] The objective of the present invention is to provide a robust optimization method for the control parameters of a microgrid considering system parameter uncertainties, including the following steps:
[0006] 1) Construct a reduced-order model of the islanded AC microgrid;
[0007] 2) Based on the reduced-order model of the islanded AC microgrid, construct a robust optimization framework for the control parameters considering system parameter uncertainties;
[0008] 3) Transform the robust optimization framework for the control parameters considering system parameter uncertainties into an adversarial Markov decision process;
[0009] 4) Use the joint adversarial soft actor-critic algorithm to solve the adversarial Markov decision process to obtain the microgrid control parameters.
[0010] Furthermore, the islanded AC microgrid includes an inverter module, a line module, and a load module.
[0011] Furthermore, in step 1), the steps of constructing the reduced-order model of the islanded AC microgrid include:
[0012] 1.1) Construct a full-order model of the islanded AC microgrid;
[0013] 1.2) Use the eigenvalue analysis method to divide the state variables in the full-order model of the islanded AC microgrid into fast and slow variables, that is:
[0014]
[0015] Wherein, is the state variable in the full-order model of the islanded AC microgrid; are the fast and slow variables; is the system state matrix;
[0016] 1.3) Set the perturbation parameter μ, and transform the state variables into the singular perturbation form, that is:
[0017]
[0018] 1.4) Let the perturbation parameter μ = 0 to obtain the quasi-steady-state solution of the fast variables, that is:
[0019]
[0020] 1.5) Substitute the quasi-steady-state solution of the fast variables into Equation (2) to obtain the reduced-order model of the islanded AC microgrid, that is:
[0021]
[0022] Wherein, is the system state matrix after order reduction.
[0023] Furthermore, the full-order model of the islanded AC microgrid is as follows:
[0024]
[0025] where, Δx INV represents the inverter state variable; Δi lineDQ represents the line state variable, and Δi loadDQ represents the load state variable; A sys is the system state matrix.
[0026] Furthermore, in the reduced-order model of the islanded AC microgrid, the state variables of the inverter module are as follows:
[0027] Δx inv = [Δδ ΔP ΔQ Δi od Δi oq (6)
[0028] In the formula, Δδ, ΔP, ΔQ, Δi od , and Δi oq are the change in the inverter voltage phase angle, the change in the inverter output active power, the change in the inverter output reactive power, the change in the d-axis component of the inverter output current, and the change in the q-axis component of the inverter output current, respectively.
[0029] Furthermore, the robust optimization framework minJ of the control parameters considering system parameter uncertainties is as follows:
[0030]
[0031] where N represents the number of inverters, Λ dom represents the set of key characteristic roots that dominate the system dynamic response, ζ(λ) represents the damping ratio corresponding to the key characteristic roots, ζ represents the lower limit of the damping ratio, k represents the droop control parameter of the inverter, k respectively represent the upper and lower limits of the droop control parameter; Δk i,f , and Δk i,v represent the droop parameter deviation of the controller.
[0032] Furthermore, the state space of the adversarial Markov decision process is as follows:
[0033] s t = {ζ min,t , ε t} ∈ S(10)
[0034] In the formula, ε t = ζmin,t - ζ is the minimum damping ratio ζ of the system min,t and the preset lower limit of the damping ratio ζ The difference between them; S is the state space; s t is the state at time t;
[0035] In the adversarial Markov decision process, the action of the adversarial agent is the parameter deviation of the controller coupling impedance, and the action of the main agent is the droop parameter of the controller;
[0036] The action space of the adversarial agent is as follows:
[0037] a o,t ={ΔL c,i,t , ΔR c,i,t}} ∈ A o (11)
[0038]
[0039] In the formula, R c,i,t and L c,i,t represent the coupling impedance in the controller, and represent the approximate theoretical parameters; ΔR c,i,t and ΔL c,i,t represent the parameter deviation of the controller coupling impedance; a o,t is the action of the adversarial agent at time t; A o is the action space of the adversarial agent;
[0040] The action space of the main agent is as follows:
[0041] a p,t ={Δk i,f,t , Δk i,v,t}} ∈ A p (13)
[0042] In the formula, k i,f,t and k i,v,t represent the droop parameters of the controller; Δk i,f,t , Δk i,v,t represent the droop parameter deviation of the controller; a p,t is the action of the main agent at time t; A p is the action space of the main agent;
[0043] The state transition probability P(s t , a p,t , a o,t ) of the adversarial Markov decision process is as follows:
[0044] s t+1 = P(s t , ap,t , a o,t )(14)
[0045] The reward function of the adversarial Markov decision process is as follows:
[0046] r t = r 1,t + r 2,t (15)
[0047]
[0048] r 2,t = α2(ζ min,t - ζ )(17)
[0049] where α1 and α2 are the weights of the reward functions r 1,t and r 2,t .
[0050] Furthermore, in step 4), the steps of solving the adversarial Markov decision process using the joint adversarial soft actor-critic algorithm include:
[0051] 4.1) Initialize the experience pool, update frequency T G , the main policy parameter θ, the adversarial policy parameter ω, and the value function φ;
[0052] 4.2) Sample from the policy π p (·|s) Sample from the adversarial policy π o (·|s)
[0053] 4.3) Input into the environment to obtain the reward r t and the next state s t+1 ;
[0054] 4.4) Add the experience to the experience pool
[0055] 4.5) When t mod T G = 0, enter step 6), when t mod T G ≠ 0, return to step 2) to continue data collection;
[0056] 4.6) Randomly sample a batch B from the experience pool ;
[0057] 4.7) Update the joint Q function;
[0058] 4.8) Update the main policy parameter θ;
[0059] 4.9) Update the adversarial policy parameter ω;
[0060] 4.10) Update the value function φ;
[0061] 4.11) Determine whether θ, ω, and φ converge. If they do not converge, return to step 2). If they converge, output the main policy π p 。
[0062] Furthermore, in step 4.7), the joint Q - function is updated as follows:
[0063]
[0064] In the formula, the policy π=(π p ,π o ); α p ,α o are the Lagrange multipliers of the entropy constraint; α p >0 and α o <0; y(r, s′) is the target value of the Q - function; denotes the expectation; and are generated by the latest policies π p (·|s′) and π o (·|s′); γ is a coefficient.
[0065] Furthermore, the main policy parameter θ and the adversarial policy parameter ω are updated as follows:
[0066]
[0067] In the formula, is the loss function.
[0068] The technical effect of the present invention is beyond doubt. The present invention proposes a method for robust optimization of control parameters considering system parameter uncertainty. Based on adversarial reinforcement learning, it ensures that the parameter optimization scheme for improving the system damping characteristics has strong robustness. First, by linearizing the dynamic equations of the islanded AC microgrid and arranging them into a state-space model, and then combining the singular perturbation theory to reduce the full-order model, an approximate model can be obtained after substituting the theoretical parameters. While simplifying the computational complexity, it ensures the stable consistency before and after model reduction and precisely retains the dynamic response of the key state variables of the system, providing a good learning environment for the subsequent reinforcement learning agent. On this basis, the model mismatch problem caused by system parameter uncertainty and the parameter optimization problem for improving the system damping characteristics are transformed into an adversarial Markov decision process, turning the originally complex problem of robust optimization of control parameters under parameter uncertainty into a mathematical model that can be efficiently solved by adversarial reinforcement learning. The task of the adversarial agent is to simulate the model error caused by parameter uncertainty and create the worst-case scenario by perturbing the system parameters. The goal of the main agent is to ensure that the minimum damping ratio of the system can still be within a reasonable range by adjusting the control parameters of the inverter under the worst-case scenario created by the adversarial agent, maintaining the dynamic performance and stability of the system, and ensuring that the final parameter optimization scheme has strong robustness. Finally, the joint adversarial soft actor-critic algorithm is used to solve the robust optimization problem. By optimizing the objective function based on maximum entropy to enhance the exploration ability of the agent, and at the same time placing the adversarial agent and the main agent in the same zero-sum game to share the value function. The value function is updated based on the Bellman equation and the strategies of the main agent and the adversarial agent are alternately optimized to ensure that the final parameter adjustment scheme can converge quickly while taking into account the dynamic performance, stability, and robustness of the system.
[0069] In summary, based on the theory of adversarial reinforcement learning, the present invention solves the model mismatch problem caused by system parameter uncertainty, ensures the robustness of the control parameter optimization results, improves the damping characteristics of the islanded AC microgrid, guarantees the small-signal stability and dynamic performance of the system, and provides a new optimization method and technical support for the high-quality operation of inverter-based microgrids. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 is a schematic structural diagram of the islanded AC microgrid of the present invention;
[0071] Figure 2 is a flow chart of microgrid model reduction;
[0072] Figure 3 is a schematic diagram of the joint adversarial soft actor-critic algorithm. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] The present invention will be further described below in conjunction with embodiments, but it should not be understood that the above-mentioned subject scope of the present invention is limited to the following embodiments. Without departing from the above technical idea of the present invention, various substitutions and changes made according to ordinary technical knowledge and customary means in the art should be included within the protection scope of the present invention.
[0074] Embodiment 1:
[0075] Refer to Figures 1 to 3 , a robust optimization method for microgrid control parameters considering system parameter uncertainties, includes the following steps:
[0076] 1) Construct a reduced-order model of the islanded AC microgrid;
[0077] 2) Based on the reduced-order model of the islanded AC microgrid, construct a robust optimization framework for control parameters considering system parameter uncertainties;
[0078] 3) Transform the robust optimization framework for control parameters considering system parameter uncertainties into an adversarial Markov decision process;
[0079] 4) Use the joint adversarial soft actor-critic algorithm to solve the adversarial Markov decision process to obtain the microgrid control parameters.
[0080] The islanded AC microgrid includes an inverter module, a line module, and a load module.
[0081] In step 1), the steps of constructing a reduced-order model of the islanded AC microgrid include:
[0082] 1.1) Construct a full-order model of the islanded AC microgrid;
[0083] 1.2) Use the eigenvalue analysis method to divide the state variables in the full-order model of the islanded AC microgrid into fast and slow variables, that is:
[0084]
[0085] In the formula, is the state variable in the full-order model of the islanded AC microgrid; are the fast and slow variables; is the system state matrix;
[0086] 1.3) Set the perturbation parameter μ, and transform the state variable into a singular perturbation form, that is:
[0087]
[0088] 1.4) Let the perturbation parameter μ = 0 to obtain the quasi-steady-state solution of the fast variable, that is:
[0089]
[0090] 1.5) Substitute the quasi-steady state solution of the fast variable into Equation (2) to obtain the reduced-order model of the islanded AC microgrid, that is:
[0091]
[0092] where, is the reduced-order system state matrix.
[0093] The full-order model of the islanded AC microgrid is as follows:
[0094]
[0095] where, Δx INV represents the inverter state variable; Δi lineDQ represents the line state variable, and Δi loadDQ represents the load state variable; A sys is the system state matrix.
[0096] In the reduced-order model of the islanded AC microgrid, the state variables of the inverter module are as follows:
[0097] Δx inv = [Δδ ΔP ΔQ Δi od Δi oq (6)
[0098] where, Δδ, ΔP, ΔQ, Δi od , Δi oq are the change in the inverter voltage phase angle, the change in the inverter output active power, the change in the inverter output reactive power, the change in the d-axis component of the inverter output current, and the change in the q-axis component of the inverter output current, respectively.
[0099] The robust optimization framework of the control parameters considering the system parameter uncertainty is as follows:
[0100]
[0101] where N represents the number of inverters, Λ dom represents the set of key characteristic roots that dominate the system dynamic response, ζ(λ) represents the damping ratio corresponding to the key characteristic roots, ζ represents the lower limit of the damping ratio, k represents the droop control parameter of the inverter, k represents the upper and lower limits of the droop control parameter, respectively; Δk i,f , Δk i,v represent the droop parameter deviation of the controller.
[0102] The state space of the adversarial Markov decision process is as follows:
[0103] st = {ζ min,t , ε t} ∈ S(10)
[0104] where ε t = ζ min,t - ζ is the difference between the minimum damping ratio ζ min,t of the system and the lower limit of the preset damping ratio ζ ; S is the state space; s t is the state at time t;
[0105] In the adversarial Markov decision process, the action of the adversarial agent is the parameter deviation of the controller coupling impedance, and the action of the main agent is the droop parameter of the controller;
[0106] The action space of the adversarial agent is as follows:
[0107] a o,t = {ΔL c,i,t , ΔR c,i,t} ∈ A o (11)
[0108]
[0109] where R c,i,t and L c,i,t represent the coupling impedance in the controller, and represent the approximate theoretical parameters; ΔR c,i,t and ΔL c,i,t represent the parameter deviation of the controller coupling impedance; a o,t is the action of the adversarial agent at time t; A o is the action space of the adversarial agent;
[0110] The action space of the main agent is as follows:
[0111] a p,t = {Δk i,f,t , Δk i,v,t} ∈ A p (13)
[0112] where k i,f,t and k i,v,t represent the droop parameters of the controller; Δk i,f,t , Δk i,v,t represent the droop parameter deviation of the controller; a p,t is the action of the main agent at time t; A p is the action space of the main agent;
[0113] The state transition probability P(s t , ap,t , a o,t ) is as follows:
[0114] s t+1 = P(s t , a p,t , a o,t )(14)
[0115] The reward function of the adversarial Markov decision process is as follows:
[0116] r t = r 1,t + r 2,t (15)
[0117]
[0118] r 2,t = α2(ζ min,t - ζ )(17)
[0119] Where α1 and α2 are the weights of the reward functions r 1,t and r 2,t .
[0120] In step 4), the steps of solving the adversarial Markov decision process using the joint adversarial soft actor-critic algorithm include:
[0121] 4.1) Initialize the experience pool, update frequency T G , the main policy parameter θ, the adversarial policy parameter ω, and the value function φ;
[0122] 4.2) Sample from the policy π p (·|s) Sample from the adversarial policy π o (·|s)
[0123] 4.3) Input into the environment to obtain the reward r t and the next state s t+1 ;
[0124] 4.4) Add the experience to the experience pool
[0125] 4.5) When t mod T G = 0, go to step 6), when t mod T G ≠ 0, return to step 2) to continue data collection;
[0126] 4.6) Randomly sample a batch B from the experience pool .
[0127] 4.7) Update the joint Q function;
[0128] 4.8) Update the main policy parameter θ;
[0129] 4.9) Update the adversarial policy parameter ω;
[0130] 4.10) Update the value function φ;
[0131] 4.11) Determine whether θ, ω, and φ converge. If they do not converge, return to step 2). If they converge, output the main policy π p .
[0132] In step 4.7), the joint Q function is updated as follows:
[0133]
[0134] where the policy π = (π p , π o ); α p , α o are the Lagrange multipliers of the entropy constraint; α p > 0 and α o < 0; y(r, s′) is the target value of the Q function; denotes the expectation; and are generated by the latest policies π p (·|s′) and π o (·|s′); γ is a coefficient.
[0135] The main policy parameter θ and the adversarial policy parameter ω are updated as follows:
[0136]
[0137] where is the loss function.
[0138] Example 2:
[0139] A robust optimization method for microgrid control parameters considering system parameter uncertainties includes the following steps:
[0140] 1) Construct a reduced-order model of the islanded AC microgrid;
[0141] 2) Based on the reduced-order model of the islanded AC microgrid, construct a robust optimization framework for control parameters considering system parameter uncertainties;
[0142] 3) Convert the robust optimization framework for control parameters considering system parameter uncertainties into an adversarial Markov decision process;
[0143] 4) Solve the adversarial Markov decision process using the combined adversarial soft actor-critic algorithm to obtain the microgrid control parameters.
[0144] Embodiment 3:
[0145] A robust optimization method for microgrid control parameters considering system parameter uncertainties, with the technical content being the same as that of Embodiment 2. Further, the islanded AC microgrid includes an inverter module, a line module, and a load module.
[0146] Embodiment 4:
[0147] A robust optimization method for microgrid control parameters considering system parameter uncertainties, with the technical content being the same as any one of Embodiments 2-3. Further, in step 1), the steps of constructing a reduced-order model of the islanded AC microgrid include:
[0148] 1.1) Construct a full-order model of the islanded AC microgrid;
[0149] 1.2) Use the eigenvalue analysis method to divide the state variables in the full-order model of the islanded AC microgrid into fast and slow variables, that is:
[0150]
[0151] 1.3) Set the perturbation parameter μ and transform the state variables into the singular perturbation form, that is:
[0152]
[0153] 1.4) Let the perturbation parameter μ = 0 to obtain the quasi-steady-state solution of the fast variables, that is:
[0154]
[0155] 1.5) Substitute the quasi-steady-state solution of the fast variables into Equation (2) to obtain the reduced-order model of the islanded AC microgrid, that is:
[0156]
[0157] where, is the reduced-order system state matrix.
[0158] Embodiment 5:
[0159] A robust optimization method for microgrid control parameters considering system parameter uncertainties, with the technical content being the same as any one of Embodiments 2-4. Further, the full-order model of the islanded AC microgrid is as follows:
[0160]
[0161] where, Δx INV represents the inverter state variable; Δi lineDQThe state variable representing the line, Δi loadDQ The state variable representing the load; A sys is the state matrix of the system.
[0162] Example 6:
[0163] A robust optimization method for microgrid control parameters considering system parameter uncertainties, with the technical content being the same as any one of Examples 2 - 5. Further, in the reduced - order model of the islanded AC microgrid, the state variables of the inverter module are as follows:
[0164] Δx inv =[Δδ ΔP ΔQ Δi od Δi oq (6)
[0165] In the formula, Δδ, ΔP, ΔQ, Δi od , Δi oq are the change in the inverter voltage phase angle, the change in the inverter output active power, the change in the inverter output reactive power, the change in the d - axis component of the inverter output current, and the change in the q - axis component of the inverter output current.
[0166] Example 7:
[0167] A robust optimization method for microgrid control parameters considering system parameter uncertainties, with the technical content being the same as any one of Examples 2 - 6. Further, the robust optimization framework minJ of the control parameters considering system parameter uncertainties is as follows:
[0168]
[0169] where N represents the number of inverters, Λ dom represents the set of key characteristic roots that dominate the system dynamic response, ζ(λ) represents the damping ratio corresponding to the key characteristic roots, ζ represents the lower limit of the damping ratio, k represents the droop control parameter of the inverter, k respectively represent the upper and lower limits of the droop control parameter; Δk i,f , Δk i,v represent the deviation of the droop parameter of the controller.
[0170] Example 8:
[0171] A robust optimization method for microgrid control parameters considering system parameter uncertainties, with the technical content being the same as any one of Examples 2 - 7. Further, the state space of the adversarial Markov decision process is as follows:
[0172] s t ={ζ min,t ,ε t}∈S(10)
[0173] where ε t = ζ min,t − ζ is the difference between the minimum damping ratio of the system and the lower limit of the preset damping ratio;
[0174] In the adversarial Markov decision process, the action of the adversarial agent is the parameter deviation of the controller coupling impedance, and the action of the principal agent is the droop parameter of the controller;
[0175] The action space of the adversarial agent is as follows:
[0176] a o,t = {ΔL c,i,t , ΔR c,i,t} ∈ A o (11)
[0177]
[0178] where, R c,i,t and L c,i,t represent the coupling impedance in the controller, and represent the approximate theoretical parameters;
[0179] The action space of the principal agent is as follows:
[0180] a p,t = {Δk i,f,t , Δk i,v,t} ∈ A p (13)
[0181] where k i,f,t and k i,v,t represent the droop parameters of the controller;
[0182] The state transition probability of the adversarial Markov decision process is as follows:
[0183] s t+1 = P(s t , a p,t , a o,t )(14)
[0184] The reward function of the adversarial Markov decision process is as follows:
[0185] r t = r 1,t + r 2,t (15)
[0186]
[0187] r 2,t = α2(ζmin,t - ζ )(17)
[0188] where α1 and α2 are the weights of the reward functions r 1,t and r 2,t .
[0189] Embodiment 9:
[0190] A robust optimization method for microgrid control parameters considering system parameter uncertainties, the technical content is the same as any one of Embodiments 2-8. Further, in step 4), the steps of using the joint adversarial soft actor-critic algorithm to solve the adversarial Markov decision process include:
[0191] 4.1) Initialize the experience pool, update frequency T G , the main policy parameter θ, the adversarial policy parameter ω, and the value function φ;
[0192] 4.2) Sample from the policy π p (·|s) Sample from the adversarial policy π o (·|s)
[0193] 4.3) Input into the environment to obtain the reward r t and the next state s t+1 ;
[0194] 4.4) Add the experience to the experience pool
[0195] 4.5) When t mod T G = 0, go to step 6), when t mod T G ≠ 0, return to step 2) to continue data collection;
[0196] 4.6) Randomly sample a batch B from the experience pool ;
[0197] 4.7) Update the joint Q function;
[0198] 4.8) Update the main policy parameter θ;
[0199] 4.9) Update the adversarial policy parameter ω;
[0200] 4.10) Update the value function φ;
[0201] 4.11) Determine whether θ, ω, and φ converge. If they do not converge, return to step 2). If they converge, output the main policy π p .
[0202] Embodiment 10:
[0203] A robust optimization method for microgrid control parameters considering system parameter uncertainties, the technical content is the same as any one of Embodiments 2-9. Further, in step 4.7), the joint Q function is updated as follows:
[0204]
[0205] where π = (π p , π o ); α p , α o are the Lagrange multipliers of the entropy constraint; α p > 0 and α o < 0; y(r, s′) is the target value of the Q function.
[0206] Embodiment 11:
[0207] A robust optimization method for microgrid control parameters considering system parameter uncertainties, the technical content is the same as any one of Embodiments 2-10. Further, the main policy parameter θ and the adversarial policy parameter ω are updated as follows:
[0208]
[0209] where is the loss function.
[0210] Embodiment 12:
[0211] A robust optimization method for microgrid control parameters considering system parameter uncertainties, this method includes the following main modules:
[0212] Microgrid reduced-order model construction module
[0213] Adversarial Markov decision generation module
[0214] Joint adversarial soft actor-critic algorithm solving module
[0215] Specific implementation steps:
[0216] Step 1: Approximate model construction
[0217] First, construct a reduced-order model of the island AC microgrid based on the theoretical parameters, such as Figure 1As shown, the model mainly includes the following three modules: an inverter module, a line module, and a load module. In a microgrid system based on an inverter, its control loop usually adopts a double-closed-loop structure, and the time scale of the outer loop is higher than that of the inner loop. In practical applications, the power loop usually serves as the outer loop, and its time constant is in seconds, while the time constants of the inner current loop and voltage loop are usually in milliseconds. For the circuit loop of the microgrid, the dynamic response time of the LCL filter circuit is generally in milliseconds, and the operating frequency of switching devices such as IGBTs is as high as 10 kHz. Considering the multi-time scale characteristics of the microgrid, it is possible to
[0218] use the singular perturbation method for order reduction.
[0219] For the microgrid system, at a certain steady-state operating point, linearize the dynamic equation of the system and organize it into a small-signal model as shown in Equation (1):
[0220] as follows:
[0221]
[0222] where Δx INV represents the inverter state variables. Before order reduction, a single inverter has 13 state variables, Δi lineDQ represents the line state variables, and Δi loadDQ represents the load state variables.
[0223] A sys is the state matrix of the system. Solving the eigenvalues of this matrix can determine whether the system is stable.
[0224] Next, reduce the order of the full-order model of the microgrid according to the Figure 2 shown process. In the full-order model of Equation (1), the state variables are divided into fast and slow variables through eigenvalue analysis, as shown in Equation (2):
[0225]
[0226] After selecting appropriate perturbation parameters, the model is transformed into a singular perturbation form:
[0227]
[0228] where μ represents the perturbation parameter we selected. Setting the perturbation parameter to 0, the quasi-steady-state solution of the boundary layer model can be obtained, that is, the quasi-steady-state solution of the fast variables of the system, as shown in Equation (4):
[0229]
[0230] Then, substituting the obtained quasi-steady-state solution of the fast variables into Equation (2), the final reduced-order model can be obtained:
[0232]
[0233] In the final reduced-order model, the state variables of the inverter module are reduced from 13 to 5:
[0234] Δx inv = [Δδ ΔP ΔQ Δi od Δi oq (
[0235] where is the reduced-order system state matrix, and finally substituting the theoretical parameters to obtain an approximate model of the AC islanded microgrid.
[0236] Step 2: Construct an adversarial Markov decision process
[0237] Currently, the parameters for modeling the power electronics-dominated microgrid system are not determined on-site. The parameter drift caused by component aging, temperature rise, the parameter uncertainties brought by parasitic capacitance and inductance parameters make the theoretical parameters of the constructed model deviate from the actual operating parameters of the system. The resulting model mismatch problem is equivalent to adding an uncertain part ΔA to the originally determined state matrix A, as shown in Equation (7):
[0238]
[0239] Taking Figure 1 as an example, considering that the system parameters affecting the system state matrix in the reduced-order model are the coupling resistance and inductance R C , L C of the controller, and the line R l , L l , the present invention assumes that the modeling error is caused by the coupling resistance and inductance in the controller. Then, the optimization objectives are clarified in the microgrid model with model mismatch caused by parameter uncertainties. First, minimize the control cost as much as possible. Second, make the system have good damping characteristics by constraining the damping ratio of the key characteristic roots of the system within the desired region:
[0240]
[0241] where N represents the number of inverters, Λ dom represents the set of key characteristic roots that dominate the system dynamic response, ζ(λ) represents the damping ratio corresponding to the key characteristic roots, ζ represents the lower limit of the damping ratio, k represents the droop control parameter of the inverter, k respectively represent the upper and lower limits of the droop control parameter.
[0242] Next, the robust optimization framework of control parameters considering system parameter uncertainties is transformed into an adversarial Markov decision process (AMDP). The specific definitions of the state space, action space, state transition probability, and reward function are as follows:
[0243] 1) State space: In the robust optimization problem of control parameters considering system parameter uncertainties under study, the agent will receive a definite state from the environment. At time step t, the measured state is expressed as:
[0244] s t ={ζ min,t ,ε t}∈S (1
[0245] where ε t =ζ min,t - ζ is the difference between the minimum damping ratio of the system and the lower limit of the preset damping ratio, which can be calculated through the approximate model of the system.
[0246] 2) Action space: In this section, the action generated by the adversarial agent is defined as the parameter deviation of the controller coupling impedance. The goal is to simulate the model error caused by parameter uncertainties and create the worst-case scenario by perturbing the parameters of the system. Therefore, the action space of the adversarial agent is defined as:
[0247] a o,t ={ΔL c,i,t ,ΔR c,i,t}∈A o (1
[0248] The simulation parameters are calculated according to formula (13):
[0249]
[0250] where, R c,i,t and L c,i,t represent the coupling impedance in the controller, and represent the approximate theoretical parameters.
[0251] The action generated by the main agent is defined as the droop parameter of the controller. The goal is to optimize the system dynamic performance and stability by adjusting the control parameters of the inverter under the worst-case scenario created by the adversarial agent. Therefore, the action space of the main agent is defined as:
[0252] a p,t ={Δk i,f,t ,Δk i,v,t}∈A p (1
[0253] where k i,f,t and k i,v,tRepresents the droop parameter of the controller, where the adjustable magnitude of the droop parameter should be within the range defined in Equation (10).
[0254] 3) State transition probability: The system state transition probability is as follows:
[0255] s t+1 = P(s t , a p,t , a o,t )(1
[0256] That is, the next state s t+1 of the system is jointly determined by the current state s t , the agent action a p,t and a o,t .
[0257] 4) Reward function: The reward function r t is used to evaluate the performance of the action a t in the state s t . The reward function is defined to solve the problems in (7)-(10) considering the stochastic environment. Therefore, the reward can be defined as:
[0258] r t = r 1,t + r 2,t (1
[0259]
[0260] r 2,t = α2(ζ min,t - ζ )(1
[0261] The reward function consists of two parts r 1,t and r 2,t . The first part, as defined in Equation (8), aims to minimize the system control cost, i.e., minimize the change in the droop control parameter. The second part, as defined in Equation (9), aims to constrain the system damping ratio within the desired interval to optimize the dynamic performance.
[0262] Step 3: Solve using the Joint Adversarial Soft Actor-Critic Algorithm
[0263] Use the Joint Adversarial Soft Actor-Critic algorithm (JASAC) to solve the above constructed AMDP, where the schematic diagram of JASAC is as shown in Figure 3As shown, JASAC maximizes and optimizes for each agent based on SAC in an alternative way. During training, the agents share system knowledge using a spanning state-action value function, and each agent considers the behavior of the other agent at each step to optimize the policy. Under the framework of AMDP, a maximum optimization framework is adopted for both agents, and the value function is derived based on the SAC algorithm:
[0264]
[0265] where π represents (π p , π o ), and α p , α o are the Lagrange multipliers of the entropy constraint. Since the goal of the adversarial agent is to minimize the expected reward, α p > 0 and α o < 0.
[0266] Although the action spaces of the agents are different, they are in the same game, so the actual value function should be consistent:
[0267]
[0268] Therefore, during adversarial training, the value approximation between agents can be shared to accelerate reaching equilibrium. That is, different from the two independent Q-functions in traditional adversarial reinforcement learning, this method defines a joint Q-function and estimates the Q-value through the Bellman equation:
[0269]
[0270] Q π The shared information in makes the algorithm consider the current policies π p and π o simultaneously during the approximation process, rather than considering them alternately. Then, based on the Bellman equation, the target value y(r, s′) of the Q-function is calculated, and its expression is as follows:
[0271]
[0272] where and are generated by the latest policies π p (·|s′) and π o (·|s′). The parameters of the target Q-network are delayed compared to the learned φ, but φ1 and φ2 can be updated through gradient descent in the MSBE function J Q (φ), where the function J Q (φ) is often calculated from a batch of samples B from :
[0273]
[0274] Based on the joint training and updating of the completed value function, the policies of the main agent and the adversarial agent can be updated. Among them, the policy is learned by optimizing the expected value. In the joint Q-function of this step, the value function is also calculated by the expectations on two action spaces:
[0275]
[0276] To decouple the expectation and the action, the parameter reparameterization technique in SAC is borrowed:
[0277]
[0278] where θ is the parameter of the main agent's policy and ω is the parameter of the adversarial agent's policy. On this basis, the original problem is transformed into:
[0279]
[0280] For the min-max optimization problem on the continuous space, it is very time-consuming to evaluate the equilibrium solution at each step. Therefore, π p and π o are learned alternately in a descending manner. Since θ and ω are independent in the entropy term, the loss functions of the policy parameters θ and ω can be derived:
[0281]
[0282] In equations (28) and (29), the joint Q-function not only considers its own behavior but also the behavior of the opponent, thus providing possible state values. This enables the updated policy to take into account the random behavior of the other party and correct the search direction. The adversarial reinforcement learning training process based on JASAC is summarized in Table 1.
[0283]
[0284]
Claims
1. A robust optimization method for microgrid control parameters considering system parameter uncertainties, characterized in that, It includes the following steps: 1) Construct a reduced-order model of the islanded AC microgrid; 2) Based on the reduced-order model of the islanded AC microgrid, construct a robust optimization framework for control parameters considering system parameter uncertainties; 3) Transform the robust optimization framework for control parameters considering system parameter uncertainties into an adversarial Markov decision process; 4) Use the joint adversarial soft actor-critic algorithm to solve the adversarial Markov decision process to obtain the microgrid control parameters.
2. The robust optimization method for microgrid control parameters considering system parameter uncertainties according to claim 1, characterized in that The islanded AC microgrid includes an inverter module, a line module, and a load module.
3. The robust optimization method for microgrid control parameters considering system parameter uncertainty according to claim 1, characterized in that In step 1), the steps for constructing a reduced-order model of the islanded AC microgrid include: 1.1) Construct a full-order model of the islanded AC microgrid; 1.2) Use the eigenvalue analysis method to divide the state variables in the full-order model of the islanded AC microgrid into fast and slow variables, that is: In the formula, is the state variable in the full-order model of the islanded AC microgrid; are the fast and slow variables; is the system state matrix; 1.3) Set the perturbation parameter μ and transform the state variables into the singular perturbation form, that is: 1.4) Let the perturbation parameter μ = 0 to obtain the quasi-steady-state solution of the fast variables, that is: 1.5) Substitute the quasi-steady-state solution of the fast variables into equation (2) to obtain the reduced-order model of the islanded AC microgrid, that is: In the formula, is the system state matrix after order reduction.
4. The robust optimization method for microgrid control parameters considering system parameter uncertainties according to claim 3, characterized in that The full-order model of the islanded AC microgrid is as follows: Among them, Δx INV represents the inverter state variable; Δi lineDQ represents the state variable of the line, and Δi loadDQ represents the state variable of the load; A sys is the state matrix of the system.
5. The robust optimization method for microgrid control parameters considering system parameter uncertainties according to claim 3, characterized in that In the reduced-order model of the islanded AC microgrid, the state variables of the inverter module are as follows: Δx inv = [Δδ ΔP ΔQ Δi od Δi oq (6) where Δδ, ΔP, ΔQ, Δi od , Δi oq are the change in the inverter voltage phase angle, the change in the inverter output active power, the change in the inverter output reactive power, the change in the d-axis component of the inverter output current, and the change in the q-axis component of the inverter output current, respectively.
6. The robust optimization method for microgrid control parameters considering system parameter uncertainties according to claim 1, characterized in that The robust optimization framework minJ for control parameters considering system parameter uncertainties is as follows: where N represents the number of inverters, Λ dom represents the set of key characteristic roots that dominate the dynamic response of the system, and ζ(λ) represents the damping ratio corresponding to the key characteristic roots. ζ represents the lower limit of the damping ratio, k represents the droop control parameter of the inverter, k respectively represents the upper and lower limits of the droop control parameter; Δk i,f and Δk i,v represent the droop parameter deviation of the controller.
7. The robust optimization method for microgrid control parameters considering system parameter uncertainty according to claim 1, characterized in that The state space of the adversarial Markov decision process is as follows: s t = {ζ min,t , ε t} ∈ S(10) where ε t = ζ min,t - ζ is the difference between the minimum damping ratio ζ min,t of the system and the lower limit ζ of the preset damping ratio; S is the state space; s t is the state at time t; In the adversarial Markov decision process, the action of the adversarial agent is the parameter deviation of the controller coupling impedance, and the action of the main agent is the droop parameter of the controller; The action space of the adversarial agent is as follows: a o,t = {ΔL c,i,t , ΔR c,i,t} ∈ A o (11) where R c,i,t and L c,i,t represent the coupling impedance in the controller, and represent approximate theoretical parameters; ΔR c,i,t and ΔL c,i,t represent the parameter deviations of the controller coupling impedance; a o,t is the action of the adversary agent at time t; A o is the action space of the adversary agent. The action space of the main agent is as follows: a p,t = {Δk i,f,t , Δk i,v,t} ∈ A p (13) where k i,f,t and k i,v,t represent the droop parameters of the controller; Δk i,f,t , Δk i,v,t represent the deviation of the droop parameters of the controller; a p,t is the action of the master agent at time t; A p is the action space of the master agent; The state transition probability P(s t , a p,t , a o,t ) of the adversarial Markov decision process is as follows: s t+1 = P(s t , a p,t , a o,t )(14) The reward function of the adversarial Markov decision process is as follows: r t =r 1,t +r 2,t (15) r 2,t = α2(ζ min,t - ζ )(17) where α1 and α2 are the weights of the reward functions r 1,t and r 2,t .
8. The robust optimization method for microgrid control parameters considering system parameter uncertainty according to claim 1, characterized in that In step 4), the steps for using the joint adversarial soft actor-critic algorithm to solve the adversarial Markov decision process include: 4.1) Initialize the experience pool and update frequency T G , the main policy parameter θ, the adversarial policy parameter ω, and the value function φ; 4.2) Sample from the policy π p (·|s) Sample from the adversarial policy π o (·|s) 4.3) Obtain the input environment and get the reward r t and the next state s t+1 ; 4.4) Add experience to the experience pool 4.5) When t mod T G = 0, proceed to step 6). When t mod T G ≠ 0, return to step 2) to continue data acquisition; 4.6) Randomly sample batch B from the experience pool ; 4.7) Update the joint Q function; 4.8) Update the main policy parameter θ; 4.9) Update the adversarial policy parameter ω; 4.10) Update the value function φ; 4.11) Determine whether θ, ω, and φ converge. If they do not converge, return to step 2). If they converge, output the main policy π p .
9. The robust optimization method for microgrid control parameters considering system parameter uncertainties according to claim 8, characterized in that, In step 4.7), the joint Q-function J Q (φ) is updated as follows: where the policy π = (π p , π o ); α p , α o are the Lagrange multipliers of the entropy constraint; α p > 0 and α o < 0; y(r, s′) is the target value of the Q function; denotes the expectation; and are generated by the latest policies π p (·|s′) and π o (·|s′); γ is a coefficient.
10. The robust optimization method for microgrid control parameters considering system parameter uncertainties according to claim 8, characterized in that The update of the main policy parameter θ and the adversarial policy parameter ω is as follows: In the formula, is the loss function.