Structured feedback risk constraint LQR solving method based on optimal control optimization
By combining the framework of zero-order policy gradient and optimal control problem, the sparse feedback gain matrix is optimized, which solves the convergence speed and stability problems of traditional LQR methods under sparse structure and risk constraints, and realizes faster and more stable control policy generation, which is suitable for distributed control systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-07
AI Technical Summary
Traditional LQR methods struggle to meet sparsity and risk constraints in real-world systems when dealing with control gain matrices with sparse structures, especially in distributed control systems. This results in analytical solutions not being directly obtainable, and insufficient convergence speed and stability.
By combining the Zero Policy Gradient (ZOPG) method with the Optimal Control Problem (OCP) framework, the sparse structure of the matrix is maintained by applying random perturbations to non-zero elements in each gradient update, and second-order information is used to accelerate convergence. The gradient and Hessian matrix are estimated by combining the Lagrangian function and the difference formula, thus achieving the optimization of the sparse feedback gain matrix.
While satisfying sparsity and risk constraints, it significantly improves convergence speed and stability, adapts to the rapid control requirements of distributed control systems, and provides faster and more stable control strategy generation.
Smart Images

Figure CN121806480A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of control optimization technology, specifically to a structured feedback risk constraint LQR solution method based on optimal control optimization. Background Technology
[0002] The Linear Quadratic Regulator (LQR) problem, as one of the most representative optimal control problems in modern control theory, has long played a crucial role in engineering fields such as aerospace, power systems, and robotics. Its core idea lies in designing a linear state feedback controller that minimizes the quadratic cost function involving state variables and control inputs while ensuring system stability. Traditional LQR theory, under the assumption of an exact and unconstrained system model, obtains analytical solutions by solving the algebraic Riccati equations. This approach is not only computationally efficient but also guarantees the global asymptotic stability of the closed-loop system.
[0003] However, with the increasing complexity of engineering applications, the traditional LQR framework faces significant challenges. Real-world systems often require the inclusion of various risk constraints, such as frequency safety constraints in power systems or collision avoidance constraints in robot motion planning. These constraints, typically expressed as conditional risk values or mean-variance measures, disrupt the convexity of the problem, making analytical solutions potentially unavailable through traditional Riccati equations. Furthermore, in numerous practical applications, particularly in distributed control systems, the control gain matrix becomes crucial due to limitations in communication bandwidth and topology. Typically, a specific sparsity pattern needs to be satisfied.
[0004] To overcome these limitations, for control gain matrices with sparse structures, this invention introduces a hybrid algorithm combining ZOPG with the optimal control problem (OCP) framework. In structured feedback problems, Some elements must be zero. The Zero-Order Policy Gradient (ZOPG) method preserves the matrix by applying random perturbations only to non-zero elements in each gradient update. The sparse structure of the ZOPG framework is further accelerated by using the OCP framework to accelerate convergence after ZOPG provides gradient estimation. This hybrid approach preserves the structural adaptability of ZOPG while utilizing second-order information to improve accuracy. Furthermore, the method of this invention converges faster and more stably. Additionally, for... In unconstrained situations, this invention can still employ an OCP-based algorithm to solve the problem. It is worth noting that, since it does not need to consider... By applying structured constraints, this invention can directly derive first-order and second-order information to accelerate convergence. Furthermore, the method of this invention maintains faster and more stable convergence. Summary of the Invention
[0005] To address the aforementioned problems, this invention proposes a structured feedback risk constraint LQR solution method based on optimal control optimization, comprising the following steps: Obtain the initial feedback gain matrix that satisfies the sparsity constraint of the microgrid communication topology. ; The zero-order policy gradient method is adopted, which estimates the gradient direction by using the observed changes in the objective function value through an outer iterative optimization step and an inner sampling estimation sub-step. The strategy is updated based on the optimal control algorithm, and the output is... The sparse feedback gain matrix obtained after rounds of iteration It is deployed directly on the local controller of the microgrid to achieve distributed frequency control that meets risk constraints.
[0006] Furthermore, the zeroth-order policy gradient method includes: Input smooth radius Current feedback gain matrix , and the current feedback gain matrix Structured perturbation matrix of sparse structure and non-zero element quantity ; Each , and As a controller, it performs time-domain simulation of the power system dynamic model to obtain the system trajectory of length T and calculates the corresponding Lagrange function value. as well as ; Calculate the gradient estimate using the two-point difference formula The estimated value of the Hessian matrix was calculated using the three-point difference method. ; Repeat the above process for a total of Next, to Sub-independent estimates and Taking the average yields the final average gradient estimate. Estimation of the mean Hessian matrix .
[0007] Furthermore, policy updates are performed based on the optimal control algorithm OCP, specifically including: Input initial feedback gain matrix Total number of iterations Number of samples per round Adjusting parameters and smooth radius ;for Execute in sequence: a) The average gradient estimate obtained in this round Estimation of the mean Hessian matrix ; b) Constructing the vectorized update direction of the gradient ,in: , ; This is the initial gradient vector; c) Matrixing yields a sparse update matrix. ,according to Update feedback gain matrix It automatically maintains a sparse structure consistent with the communication topology.
[0008] Furthermore, the structured perturbation matrix The generation process includes: determining the feedback gain matrix based on the actual communication link of the microgrid. The sparse pattern is such that element values are sampled from a uniform distribution U[-1,1] at non-zero positions of the sparse pattern, and the entire matrix is normalized using the Frobenius norm to satisfy... This ensures that gradient estimation is always explored in the feasible sparse subspace.
[0009] Furthermore, the zero-order policy gradient method component provides gradient information through a two-point estimation scheme. , .
[0010] Furthermore, in the gradient estimation stage of the zero-order policy gradient method, the computational complexity of the gradient for each sample is O(n). The computational complexity of the Hessian matrix for each sample is O(n). In the policy update phase based on the optimal control algorithm, the computational complexity of the update operation is O(n). .
[0011] Furthermore, the gradient estimate is calculated using the two-point difference formula. : ; The estimated value of the Hessian matrix is calculated using the three-point difference method. : H .
[0012] Furthermore, consider the risk constraint function expression as follows: ; in, Indicates expectation, state , Let be the state penalty matrix. Let be the covariance matrix of the noise. The third-order matrix of noise, The fourth central moment of noise, It is the original risk threshold. Operable standard constraint factors after noise statistical correction.
[0013] Furthermore, the method for handling LQR constraint problems is as follows: By introducing multipliers Let's consider its Lagrangian function: ; in, Lagrange multipliers To augment the state weight matrix, To control the penalty matrix.
[0014] Furthermore, the microgrid local controller deploys a sparse feedback gain matrix. The subsequent operation process includes: Each local controller, based on the sparse feedback gain matrix For each non-zero element corresponding to its communication neighbor, only local state information allowed by its communication topology is collected; Based on local information and the corresponding gain coefficient, local control commands are calculated and output in real time to achieve fully distributed frequency regulation without a central coordinator. Meanwhile, each controller continuously monitors local key state variables, and automatically triggers a mechanism based on the detected abnormal risks exceeding preset safety boundaries. Robust corrective action or switch to backup safety mode.
[0015] Compared with the prior art, the present invention has the following beneficial effects: Compared to directly seeking control strategies Unlike other methods, the method of this invention focuses on deriving the optimal gain matrix. For sparse matrices caused by communication topology constraints, this invention effectively combines the ZOPG algorithm with an algorithm based on the optimal control problem. This scheme uses ZOPG to randomly perturb and sample non-zero elements to estimate the gradient, and estimates the Hessian matrix using the three-point finite difference method, while embedding the OCP framework for iterative refinement. In practical applications, especially in microgrid networks, the optimized final control gain matrix is used... Deployed to the local controllers of each microgrid, each controller according to This invention generates local control signals to achieve distributed frequency stability control. When the generated control strategy fully conforms to the actual communication network topology, the proposed ZOPG-OCP hybrid algorithm significantly improves convergence speed, stability, and robustness to sampling number compared to the traditional stochastic gradient descent method. It can generate effective control strategies faster, meeting the rapid control requirements of microgrids. Furthermore, the convergence of the proposed algorithm has been verified, providing a theoretical guarantee for practical applications. Attached Figure Description
[0016] Figure 1 A radially connected networked microgrid system; Figure 2 For LQR cost trajectory. Detailed Implementation
[0017] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the scope of protection of the present invention.
[0018] Note: Expressing expectations, Representing an n-dimensional real vector space, the operator Let represent the Kronecker product of matrices, and let tr represent the trace of the matrix. express Norm.
[0019] Example 1 Consider a discrete-time linear time-invariant stochastic system: (1); State variables The action is Random noise is It is unrelated to time. Here is the state transition matrix. The input matrix is denoted as .
[0020] Infinite horizon constraint The question takes the following form: (2).
[0021] in, Let be the state penalty matrix. To control the penalty matrix, the matrix and All are positive semi-definite. This indicates... As the deadline The system status and control trajectory, As a parameter of risk tolerance.
[0022] Risk constraints The formulation of the problem is particularly driven by power system applications, such as frequency regulation and power control in microgrid clusters. Specifically, state variables... This could represent frequency deviation, voltage error, or power imbalance in each microgrid; control input actions. It can correspond to generator setpoint adjustment, energy storage commands, or load control signals; objective function Minimize overall operating costs and frequency fluctuations; risk constraints Limit extreme variations in frequency or power to prevent system instability or equipment overload.
[0023] here, This represents a structured feedback set defined by an information exchange diagram, leading to the design of distributed control systems based on information transmitted via communication links. Enforcing the structured strategy is as follows: (3).
[0024] Let me explain. It doesn't have a numerical meaning; rather, it represents a sparse pattern, which is the layout of a real microgrid, such as the matrix in a simulation experiment. , This is the variable to be optimized in this invention, and its sparse structure must satisfy... .
[0025] Use one Represents a node and This method, which does not rely on communication links, is particularly common in power system applications, such as power transmission between regional microgrids.
[0026] The constraints in Equation (2) enable this invention to mitigate the worst-case scenario of very high system variability caused by external load disturbances and imperfect modeling. A potential problem is... Involves relative to past trajectories The expected conditions. This invention has: (4); use The weighted noise statistic is given as follows: (5); The mean vector of the noise. Let be the covariance matrix of the noise. The third-order matrix of noise, The fourth central moment of noise.
[0027] To handle the constraints in (2), multipliers are introduced. Let's consider its Lagrange function. (6); Using matrices This makes the Lagrange quantity follow the original... Quadratic forms with the same objective.
[0028] Based on (6), the dual problem becomes: (7); Note that this invention addresses... Use bounded sets Assume (2) within reasonable boundaries The following is feasible, therefore It is finite. A sufficiently large constant can guarantee that adequate penalties are imposed when constraints are violated.
[0029] Solving (7) is quite difficult, as it requires a gradient method based on the original binary system, which involves both internal and external issues. (Variables) and Updates must be performed alternately in a nested manner, with each iteration depending on the two variables from the previous step. This leads to a highly complex solution process. Furthermore, when determining... At this time, the present invention needs to consider structured feedback. To solve these problems, the present invention considers its minimax correspondence problem, that is: (8); The stationary point (SP) of the original objective function (2) is equivalent to that of the stationary point of the original objective function (2). The optimal That is, the optimal solution of the objective function.
[0030] because yes Since it is a linear function, the present invention can directly find the optimal solution to the internal problem in (8). If the constraints are satisfied, then ,otherwise .
[0031] The KKT stationarity condition of formula (7) is closely related to the KKT stationarity condition of formula (8), which allows the present invention to find the stationary point of the original problem by solving formula (8) instead of formula (7), and also provides convenience for subsequent algorithm design and implementation.
[0032] Control strategy Find an optimal linear feedback gain matrix This minimizes the cost function (2).
[0033] To address the frequency oscillation problem in microgrids under high load fluctuations, this invention proposes a risk-constrained reinforcement learning controller employing a nested optimization framework. Its core is a solution method based on the zero-order policy gradient for optimal control.
[0034] Due to the gain matrix of the power grid controller To address the sparsity constraints of real-world communication topologies, where certain elements must be zero, traditional analytical gradient-based optimization methods are difficult to apply. This invention obtains an initial feedback gain matrix that satisfies the sparsity constraints of the microgrid communication topology as the current strategy. It employs the zero-order policy gradient (ZOPG) method, including an outer iterative optimization step and an inner sampling estimation sub-step. The gradient direction is estimated solely by observing changes in the objective function value. The strategy update step is performed using the optimal control algorithm (OCP), thus avoiding direct differentiation for non-convex, non-differentiable problems.
[0035] The steps and parameters of the zero-order policy gradient method are explained in detail below: (1) Input parameter description: Smooth radius This is a small positive number. It's in the current feedback gain matrix. A small, structured random disturbance is applied. In a power system, this corresponds to a small, tentative adjustment to the controller setpoint (such as voltage or power reference values) that conforms to communication rules.
[0036] Current feedback gain matrix : That is, the linear feedback gain matrix, which maps the system states such as frequency and voltage deviation of each node to control actions such as setpoint adjustment.
[0037] Structured perturbation matrix : is a with Random matrices with identical sparse structure, whose non-zero elements are sampled from a uniform distribution and normalized, guarantee that... It ensures that the update direction of the gradient estimate always lies in a feasible sparse subspace that conforms to the communication topology. Each non-zero element corresponds to an actual existing communication link, and the algorithm only explores possible performance improvements on these links.
[0038] Non-zero dollar quantity Current feedback gain matrix The total number of non-zero elements allowed in the value set is the number of actual communication links. This is used to scale the gradient estimate to obtain an unbiased estimate.
[0039] (2) Single-sample disturbance assessment process: Step 1: First, perform perturbation and evaluation, and then adjust the current feedback gain matrix. With disturbance By adding and subtracting, we obtain the perturbation strategy. Then, in a power system dynamic model as shown in equation (1), respectively using... as well as As a controller, time-domain simulation is performed, generating a length of... The system trajectory represents changes in state such as frequency and voltage. Based on these trajectories, the corresponding Lagrangian function values are calculated. as well as .
[0040] Step 2: Calculate the gradient estimate using the two-point difference formula: .
[0041] And the estimated value of the Hessian matrix is calculated using the three-point difference method: .
[0042] The physical meaning of this formula is: by observing the rate of change in the overall system performance after applying small structured perturbations to the controller parameters, the direction of improvement can be inferred. The estimated gradient... The Hessian matrix H automatically maintains the same position as... and The same sparse structure.
[0043] Step 3: Output Results and Variance Control Strategy: A... The gradient matrix and Hessian matrix estimates of the same dimension and sparse structure are used to guide the following steps. Update.
[0044] The methods described above provide single-step gradient estimation and Hessian matrix estimation, but the estimation variance is relatively large. Building upon this, the present invention further reduces the variance through multiple sampling averaging and employs the Optimal Control Program (OCP) framework for iterative updates, ultimately finding a near-optimal sparse controller that satisfies the risk constraints.
[0045] The steps and parameters of the Optimal Control Algorithm (OCP) are explained in detail below: (1) Input parameter description: Initial feedback gain matrix : An initial sparse gain matrix that satisfies a given communication topology.
[0046] Number of iterations The total number of rounds the algorithm runs.
[0047] Number of samples The number of independent repetitions of gradient estimation in each iteration. This is achieved through averaging. We need to obtain a more reliable gradient direction for this round using independent estimates. and Heisenberg matrix This is the key to reducing variance and achieving stable convergence.
[0048] Adjust parameters By adjusting The size of is used to balance the convergence speed and stability of the algorithm in order to achieve the optimal result.
[0049] Smooth radius : Same as above.
[0050] (2) Execution process of zero-order gradient algorithm: Gradient estimation: for arrive Zero-order gradient estimation is performed, using an independently sampled structured perturbation matrix each time. This yields a gradient estimate. Hessian matrix estimation Then calculate. The average of the estimates is obtained as well as This average gradient and Hessian matrix estimate are closer to the true gradient direction.
[0051] (3) Policy update and topology preservation mechanism: The controller is updated using the OCP algorithm formula: .in yes The matrix form, , This invention employs vectorization techniques to ensure dimensionality matching in the algorithm. Because... and Inherited The sparse structure automatically ensures the new controller in this update. The predetermined communication topology constraints are still met. That is, the gain will only be updated between controller nodes that have a communication link.
[0052] (4) Output deployable sparse gain matrix: the optimal or near-optimal sparse feedback gain matrix obtained after training. This matrix can be directly deployed to the local controllers of each power grid to achieve distributed, risk-aware frequency control.
[0053] The proposed structured feedback risk constraint LQR solution method based on optimal control optimization exhibits superlinear convergence, thus achieving faster convergence than the gradient descent method. Furthermore, this method generally possesses divergence robustness, remaining applicable even when the Hessian matrix is singular or non-positive definite. By adjusting... The size of the value can achieve a balance between convergence speed and stability.
[0054] In the structured feedback risk constraint LQR solution method based on optimal control optimization, the gradient and Hessian matrix for each sample are calculated as follows: and In the OCP algorithm update process, matrix inversion is the most computationally intensive step, requiring... operate.
[0055] The structured feedback risk constraint LQR solution method based on optimal control optimization effectively integrates the ZOPG method and the OCP framework, leveraging the complementary advantages of the two methods. This novel integrated method provides a structured feedback gain matrix... The optimization provides a systematic solution, where the ZOPG components provide gradient information through a two-point estimation format, where... , The OCP module iteratively improves the solution to converge to the optimal solution. .
[0056] Example 2 Stochastic Gradient Descent (SGD) is an efficient optimization method that estimates the gradient direction by randomly sampling a subset of data during iterative updates. This invention combines this method with ZOPG and compares and analyzes it with the OCP algorithm proposed in this invention.
[0057] After proposing a structured feedback risk constraint LQR solution method based on optimal control optimization for structured constraints, this invention provides a rigorous theoretical proof of the algorithm's convergence to ensure its reliable application in real microgrid control systems. Addressing the stringent requirements of microgrid frequency control on controller reliability, stability, and fast convergence, this proof establishes that the algorithm can start from any initial feedback gain matrix satisfying the communication topology. Starting from this point, the theoretical guarantee that the solution can stably approximate the first-order optimal solution satisfying the risk constraints is established. In the proof, an enlarged subset of levels is first defined. This set can be uniformly defined. constant Smoothness constant and neighborhood radius Furthermore, this invention transforms the update rules of the OCP algorithm, facilitating convergence derivation and making it a standard gradient descent form. However, it is important to note that its step size changes, denoted as . Its lower bound is .
[0058] The following conditions must be met: smoothing radius By adjusting parameters Enables OCP adaptive step size , OCP step size Another equivalent expression for it, which is convenient for calculation. The risk threshold is the total number of iterations. ,in The number of samples for the zero-order policy gradient (ZOPG). This is a relatively large positive constant. The core of the proof lies in analyzing the iterative dynamics of combining stochastic gradient estimation with deterministic OCP updates, by introducing a function... The Moro envelope notation (This can be achieved by substituting initial values) get Using this approximately smooth function substitution as an analytical tool, this invention derives the key iterative relationship for the descent of the function value. By analyzing the expected change of the Moro envelope function, it can be proven that under the selected parameters, the noise variance of the gradient estimation is controlled during the iteration process, and the sequence remains stable with high probability. Inside. Further utilization. maximal inequalities and The inequality can be obtained at... After at least [number] iterations, the algorithm achieves [number] results. The probability converges to -Stability point ( -SP). This probability follows Increase and improve, but too much This will reduce the step size and decrease the convergence speed, so a trade-off must be made in practice. This conclusion mathematically guarantees that the method, when faced with communication constraints and random disturbances, can not only avoid divergence but also efficiently and stably generate reliable control strategies, providing a theoretical basis for its engineering application in distributed control systems such as microgrids.
[0059] Sparse feedback gain matrix deployed in microgrid local controller The subsequent operation process includes: Each local controller according to the matrix For each non-zero element corresponding to its communication neighbor, only local state information permissible by its communication topology is collected. Based on the local information and the corresponding gain coefficient, local control commands are calculated and output in real time to achieve fully distributed frequency regulation without a central coordinator. Simultaneously, each controller continuously monitors local key state variables, and when an abnormal risk exceeding a preset safety boundary is detected, it automatically triggers a control command based on... Robust corrective action or switch to backup safety mode.
[0060] Example 3 Simulation Example Simulation and experimental results are presented to demonstrate the effectiveness of the proposed OCP-based ZOPG method. For Under structured constraints, this invention compares the convergence performance of the structured feedback risk constraint LQR solution method based on optimal control optimization and the classical stochastic gradient descent in the context of a real microgrid system.
[0061] When on When applying structured constraints, the design of the structured feedback set often depends on the specific scenario, such as the distribution within a power system. Here, this invention considers a defined set, such as... Figure 1 As shown, consider four different microgrids A1, A2, A3, and A4. Due to geographical or infrastructure limitations, they cannot be directly interconnected. For example, A1 can be directly connected to A2 and A3, but not to A4. For this pre-specified information transmission relationship, the algorithm of this invention directly configures the corresponding topology according to the connection mode. Here, "1" represents a direct connection, and "0" represents that a direct connection is not possible, as shown in the following equation: ; Next, the present invention designs the following model: ; State weights are set as State penalty matrix at time Here, using an identity matrix means giving equal weight to all state deviations such as frequency and voltage deviations, with the control weights set to... Control penalty matrix at time A coefficient of 0.1 signifies a state penalty relative to a coefficient of 1 for Q. The cost of control actions is set with a relatively small weight. This invention can assume that devices in each power grid can adjust their setpoints relatively freely without incurring significant costs. The noise covariance is set as... It can represent random fluctuations in load, with the risk threshold set to... This represents the invention's explicit quantitative requirement for safety margin. For the stochastic gradient descent (SGD) algorithm, the parameter configuration is as follows: perturbation radius... Number of samplings Learning rate For the structured feedback risk constraint LQR solution method based on optimal control optimization, the perturbation radius and sample size are kept consistent with the SGD algorithm, and the adjustment parameter is...
[0062] Figure 2 The comparison results also show that the OCP-based ZOPG solution method, represented by the red line, has a faster convergence speed and better performance. In microgrid applications, it minimizes the average frequency deviation, reducing operating costs, and achieves a better balance between security and economy without increasing additional communication burden. This further verifies the effectiveness of the proposed OCP-based ZOPG solution method in solving the LQR problem. It is worth noting that the OCP-based ZOPG solution method is effective in handling... It performs exceptionally well even on problems without structured constraints. By leveraging the respective strengths of the zero-order gradient algorithm and the OCP algorithm, the results show significant promise for future applications.
[0063] This paper proposes a ZOPG-based solution method based on OCP to solve the risk-constrained LQR problem with structured feedback constraints. This method addresses the optimization of the feedback gain matrix in a strictly sparse mode. A key challenge arises from communication limitations in distributed control systems such as networked microgrids. This algorithm only addresses... Random perturbations are applied to the non-zero elements to preserve their sparsity structure during gradient estimation. Subsequently, second-order information is incorporated into the OCP to accelerate convergence. Simulation results on a radially connected microgrid system validate the effectiveness of the proposed method.
[0064] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.
Claims
1. A structured feedback risk constraint LQR solution method based on optimal control optimization, characterized in that, Includes the following steps: Obtain the initial feedback gain matrix that satisfies the sparsity constraint of the microgrid communication topology. ; The zero-order policy gradient method is adopted, which estimates the gradient direction by using the observed changes in the objective function value through an outer iterative optimization step and an inner sampling estimation sub-step. The strategy is updated based on the optimal control algorithm, and the output is... The sparse feedback gain matrix obtained after rounds of iteration It is deployed directly on the local controller of the microgrid to achieve distributed frequency control that meets risk constraints.
2. The method according to claim 1, characterized in that, The zero-order policy gradient method includes: Input smooth radius Current feedback gain matrix , and the current feedback gain matrix Structured perturbation matrix of sparse structure and non-zero element quantity ; Each , and As a controller, it performs time-domain simulation of the power system dynamic model to obtain the system trajectory of length T and calculates the corresponding Lagrange function value. as well as ; Calculate the gradient estimate using the two-point difference formula The estimated value of the Hessian matrix was calculated using the three-point difference method. ; Repeat the above process for a total of Next, to Sub-independent estimates and Taking the average yields the final average gradient estimate. Estimation of the mean Hessian matrix .
3. The method according to claim 2, characterized in that, Policy updates are based on the Optimal Control Algorithm (OCP), specifically including: Input initial feedback gain matrix Total number of iterations Number of samples per round Adjusting parameters and smooth radius ;for Execute in sequence: a) The average gradient estimate obtained in this round Estimation of the mean Hessian matrix ; b) Constructing the vectorized update direction of the gradient ,in , ; This is the initial gradient vector; c) Matrixing yields a sparse update matrix. ,according to Update feedback gain matrix It automatically maintains a sparse structure consistent with the communication topology.
4. The method according to claim 3, characterized in that, The structured perturbation matrix The generation process includes: determining the feedback gain matrix based on the actual communication link of the microgrid. The sparse pattern is such that element values are sampled from a uniform distribution U[-1,1] at non-zero positions of the sparse pattern, and the entire matrix is normalized using the Frobenius norm to satisfy... .
5. The method according to claim 4, characterized in that, The zero-order policy gradient method component provides gradient information through a two-point estimation scheme. : , 。 6. The method according to claim 5, characterized in that, In the gradient estimation stage of the zero-order policy gradient method, the computational complexity of the gradient for each sample is O(n). The computational complexity of the Hessian matrix for each sample is O(n). In the policy update phase based on the optimal control algorithm, the computational complexity of the update operation is O(n). .
7. The method according to claim 6, characterized in that, Calculate the gradient estimate using the two-point difference formula : ; The estimated value of the Hessian matrix is calculated using the three-point difference method. : H 。 8. The method according to claim 7, characterized in that, Consider the risk constraint function expression as follows: ; in, Indicates expectation, state , Let be the state penalty matrix. Let be the covariance matrix of the noise. The third-order matrix of noise, It is the original risk threshold. Operable standard constraint factors after noise statistical correction.
9. The method according to claim 8, characterized in that, The method for handling LQR constraint problems is as follows: By introducing Lagrange multipliers Consider the Lagrangian function: ; in, To augment the state weight matrix, To control the penalty matrix.
10. The method according to claim 1, characterized in that, Sparse feedback gain matrix deployed in microgrid local controller The subsequent operation process includes: Each local controller, based on the sparse feedback gain matrix For each non-zero element corresponding to its communication neighbor, only local state information allowed by its communication topology is collected; Based on local information and the corresponding gain coefficient, local control commands are calculated and output in real time to achieve fully distributed frequency regulation without a central coordinator. Meanwhile, each controller continuously monitors local key state variables, and automatically triggers a mechanism based on the detected abnormal risks exceeding preset safety boundaries. Robust corrective action or switch to backup safety mode.
Citation Information
Cited By
Control method for adaptive morphing wing based on hessian vector product
CN122365729B