Lyapunov Stable Reinforcement Learning Control Algorithm for Uncertain Systems Based on Worst-Case Study

By employing the worst-case uncertainty-based Lyapunov stable reinforcement learning control algorithm, utilizing ReLU activation networks and robustness conditions, the stability problem of robust RL under uncertainties in dynamic models and state estimation is solved, thus achieving stable control of the robot system.

CN119439743BActive Publication Date: 2025-11-14HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411584075.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-07
Publication Date
2025-11-14
Estimated Expiration
2044-11-07

AI Technical Summary

Technical Problem

Existing robust reinforcement learning methods struggle to provide strict convergence guarantees when faced with uncertainties in the dynamic models and state estimations of robotic systems, leading to compromised system stability. This is especially true when agents explore new or unfamiliar states, resulting in frequent suboptimal or even unsafe behaviors.

Method used

We employ a worst-case-based Lyapunov-stabilized reinforcement learning control algorithm for uncertain systems. By modeling the system dynamics and uncertainty boundary through a ReLU activation network, we determine the robustness conditions and establish a robust guarantee RL. Using mixed-integer linear programming and advanced piecewise linear function representation, we train the controller and Lyapunov function to ensure the system's stability under uncertainty.

Benefits of technology

The system stability was achieved under bounded uncertainties in dynamic model and state estimation. The effectiveness of the method was demonstrated through numerical simulation, and it is applicable to the stability control of complex systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119439743B_ABST
    Figure CN119439743B_ABST
Patent Text Reader

Abstract

This invention discloses a worst-case Lyapunov stable reinforcement learning control algorithm for uncertain systems, comprising the following steps: S1, modeling the system dynamics and uncertainty boundary using a ReLU activation network; S2, determining robustness conditions and using them to pre-determine the area of ​​the attraction domain; S3, determining the robustness guarantee RL under dynamic model uncertainty and state estimation; S4, establishing network parameterization; and S5, performing numerical simulations on an inverted pendulum and a quadcopter UAV. This invention employs the aforementioned worst-case Lyapunov stable reinforcement learning control algorithm for uncertain systems, still accurately finding the most inverse state, thereby forcing its stability under uncertainty. A geometric view of the existence of robust RL solutions is provided to explain robustness and its capability. Numerical simulations on an inverted pendulum and a quadcopter under various uncertainties demonstrate the effectiveness of the proposed method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of automation control technology, and in particular to a Lyapunov-stabilized reinforcement learning control algorithm for uncertainty systems based on worst-case scenarios. Background Technology

[0002] Robotic systems are susceptible to stability degradation and catastrophic failures caused by various uncertainties, including modeling uncertainty and state uncertainty. Existing techniques on this topic include robust control, adaptive control, stochastic control, sliding mode control, and h-∞, etc. However, existing methods may be limited by the models they use to describe uncertainties and system dynamics. RL algorithms, which rely on function approximation through neural networks, have become a highly flexible framework for solving Markov decision processes (MDPs). However, classical RL struggles with robustness to uncertainties, disturbances, or changes in environmental structure. To this end, robust RL is investigated to enhance policy robustness in aspects such as: (1) uncertain dynamic models; (2) external disturbances; (3) controllable noise; and (4) uncertain observations and state estimation. The min-max problem or optimality is then formulated and solved based on the Hamilton-Jacobi-Isaacs equations to obtain a controller that suppresses the effects of uncertainty in the worst case.

[0003] If the small gain theorem is used to construct the objective function to minimize the stabilization's sensitivity to disturbances or uncertainties, disturbances caused by modeling errors can be suppressed if the gain is bounded by the inverse of the sensitivity. The uncertainty set is defined as centered on an error MDP that produces a single-sample trajectory and estimated by sampling. An alternative adversarial action or a perturbation from the opponent is added to the selected action, and the controller is trained as a generative adversarial network (GAN).

[0004] However, robust RL struggles to provide strict convergence guarantees simultaneously under uncertainties in both the dynamic model and state estimation, hindering its reliability in practical applications. Uncertainty can be compensated for by minimizing the expectation of the reward function in some designs. Modeling uncertainties can still lead to suboptimal or even unsafe behavior, especially when the agent explores new or unfamiliar states. Noisy observations can introduce uncertainty into state estimation, potentially compromising system stability.

[0005] The auxiliary loss function based on Lyapunov theory and the barrier function partially solves the convergence problem of the RL method. However, its applicability to uncertain dynamic models and system states has not yet been explored. Summary of the Invention

[0006] The purpose of this invention is to provide a worst-case Lyapunov-based reinforcement learning control algorithm for uncertain systems, used to synthesize neural network controllers to maintain the stability of the system under bounded multiplicative and additive uncertainties in the dynamic model and state estimation.

[0007] To achieve the above objectives, this invention provides a worst-case-based Lyapunov-stabilized reinforcement learning control algorithm for uncertain systems, comprising the following steps:

[0008] Includes the following steps:

[0009] S1. Model the system dynamics and uncertainty boundaries using a ReLU activation network;

[0010] S2. Determine the robustness conditions and use them to predetermine the area of ​​the attraction domain;

[0011] S3. Determine the robustness guarantee RL under the uncertainty of the dynamic model and state estimation;

[0012] S4. Network parameterization setup;

[0013] S5. Numerical simulation of inverted pendulum and quadcopter UAV.

[0014] Preferably, in step S1, the modeling of the system dynamics model and uncertainty boundary is carried out as follows:

[0015] Consider a discrete-time robot system with the following dynamic model:

[0016] (1);

[0017] ∈ and ∈ These represent the robot in t The state and control input at any given time, with the control saturation controlled by the lower bound. or upper boundary Give;

[0018] The dynamic model is given, assuming its approximate value is... The results were obtained through system identification, as shown below:

[0019] (2);

[0020] and These represent the multiplicative uncertainty and additive uncertainty between the actual dynamics and its approximation, respectively. This indicates the uncertainty of the centralized total model. Indicates controller In state The control provided at the location;

[0021] However, system state It is estimated from some estimator, and therefore will be corrupted by noise, as shown below:

[0022] (3);

[0023] in and Let these represent the multiplicative uncertainty and additive uncertainty of the state estimation, respectively. ( ) represents the lumped uncertainty in state estimation;

[0024] therefore, t The predicted state at time +1 is:

[0025] (4);

[0026] in, Indicates controller In state estimation The control provided at the location;

[0027] For an approximate dynamic model ( , The output system estimates the state estimator. Controlling saturation and And with an attractive domain D, find a controller This makes the system subject to model uncertainty. and state uncertainty It stabilizes at the equilibrium point.

[0028] Assumption 1: An approximate dynamic model is given in the form of a ReLU activated neural network. ;

[0029] Assumption 2: Assume uncertainty in the dynamics of the system model. Uncertainty in state estimation Bounded, and its boundaries are denoted as follows: and Then we have:

[0030] (5);

[0031] in, For the infinite norm, and They are all presented in the form of ReLU activated neural networks.

[0032] Preferably, in step S2, the robustness condition is a stability condition under the worst-case scenario of uncertainty, derived from Lyapunov neural control. When uncertainty is not involved, this is achieved by training coupled Lyapunov functions that satisfy the following stability condition. :

[0033] (6);

[0034] (7);

[0035] (8);

[0036] in D To attract the state region, It is in equilibrium. ( )and ( ) belongs to the K-function set, and stability conditions (6) to (8) guarantee that from D A system starting from any state converges to an equilibrium state. .

[0037] Preferably, in step S2, the stability condition under the worst-case uncertainty includes robustness conditions under dynamic model uncertainty and robustness conditions under state estimation uncertainty:

[0038] Model uncertainty in the dynamics Down, t The state at time +1 is not entirely controlled. and current state Therefore, equation (8) must be applied to... t Each possible state at time +1 If both are valid, then:

[0039] (9);

[0040] in, (10);

[0041] Through the The robustness condition (8) of the Lyapunov function at the given position is obtained by Taylor expansion:

[0042] (11);

[0043] Given Based on the randomness of the randomness and on assumption two, a stronger and more sufficient robustness condition is obtained, as follows:

[0044] (12);

[0045] in, and The difference is negligible, considering the worst-case scenario of uncertainties in dynamic modeling, and forcing the Lyapunov value to decrease along the state trajectory;

[0046] Compared with the given ground reality The control input obtained Conversely, controller In fact, in state estimation The input is calculated to determine the control input. Therefore, the stability condition needs to take into account the stochasticity of the state estimation; considering the uncertainty in the worst-case state estimation, in The Lyapunov value at point is obtained through Taylor expansion as follows:

[0047] (13);

[0048] The derivative of the Lyapunov function is obtained by the chain rule, as shown in the following formula:

[0049] (14);

[0050] in, =f( , ), = π ( );

[0051] The robustness condition under uncertainty in state estimation is as follows:

[0052] (15);

[0053] Under assumption two, the stability condition (8) becomes the new robustness condition as follows:

[0054] (16);

[0055] in, Ignored, derivative , and They are , and The piecewise constant in the equation.

[0056] Preferably, in step S3, let This represents the trainable parameters in a Lyapunov network. Let represent the trainable parameters in the controller network. Using robustness conditions (7), (12), and (16), the following loss function is designed to regulate the training of the controller network and the Lyapunov network. The loss function that guarantees the semipositive definition of the Lyapunov network (7) is given as follows:

[0057] (17);

[0058] When only uncertainty is considered in the model dynamics At that time, the Lyapunov value along the trajectory should increase strictly monotonically, and this is guaranteed by satisfying the robustness condition (12). The corresponding loss function is designed as follows:

[0059] (18);

[0060] When only the uncertainty in the state estimation is considered When the robustness condition is met, the loss function (16) is expressed as:

[0061] Based on the robustness condition (16), which represents a system with only uncertainties in the dynamic model, a loss function is constructed:

[0062] (19);

[0063] When both uncertainties mentioned above are considered simultaneously, the robustness condition (8) becomes:

[0064] (20);

[0065] Then, the loss function corresponding to the robustness condition eq is obtained, expressed as:

[0066] (twenty one);

[0067] The objective function and all constraints are expressed as follows: Piecewise linear functions:

[0068] set up This represents the set of all trainable parameters, which are used in each iteration of training with the values ​​from the previous iteration. Attraction domain S The division is as follows:

[0069] set up yi It is if and only if In the If the binary number in the partition is equal to 1, then... The convex combination of the extreme points of the partition, represented as a partition, transforms the optimization problem of the loss function into a mixed-integer linear programming problem by introducing slack variables:

[0070] (22a);

[0071] (22b);

[0072] Among them, the problem coefficient , , and All An explicit function, the optimization variable vector s is composed of x and ...composed of, the optimal value of the objective function (22) can be written as:

[0073] ;

[0074] Using gradients Trainable parameters are obtained through backpropagation. Optimize.

[0075] Preferably, in step S4, the network parameterization establishment process is as follows:

[0076] Using the HL-PWL method, an arbitrary piecewise linear function can be represented by the following expression. F ( x ):

[0077] (twenty three);

[0078] in, It is a coefficient vector. It is the stack of base functions, base functions The stack includes nested layers. All nested functions; coefficient vector Nested functions Related;

[0079] Robustness conditions under HL-PWL representation:

[0080] The positive determination of the Lyapunov function, formula (7) is rewritten as:

[0081] (twenty four);

[0082] in and yes K Function, and with The Lyapunov function increases with the increase of . andK The functions are all rewritten in HL-PWL form (23), and the expression is:

[0083] (25);

[0084] (26);

[0085] in, For the infinite norm, a∈ and ∈ For parameter vectors, and The number of corresponding parameter vectors, upper bound, and lower bound. , ∈ and It is a vector function, including nested functions at nesting levels 0 and 1. It is easy to obtain (i.e., |x|max);

[0086] because It is K The function then has:

[0087] (27);

[0088] (28);

[0089] The robustness condition (6) and positive condition (7) under equilibrium conditions are written as follows:

[0090] (29);

[0091] and

[0092] (30);

[0093] Partition and define vertices for a given attraction domain D, and then transform condition (30) into a 6-form. ,in ,matrix and By superposition V All vector The result of superposition, among which V It is a set of D partitions;

[0094] Similarly, the upper bound condition of equation (24) can be rewritten as follows: Then the upper boundary and the lower realm It is a K-function that satisfies equation (27), that is:

[0095] (31);

[0096] and functions that satisfy (28), namely:

[0097] (32);

[0098] in, and γ Given the grid size of the partition, when considering uncertainties in both dynamic modeling and state estimation, robustness conditions (20) and (14) produce another condition:

[0099] (33);

[0100] Further assume that both the dynamic system and the controller satisfy the following Lipshitz conditions:

[0101] (34);

[0102] (35);

[0103] in and F These represent the upper bounds of the controller and the positive dynamic change, respectively;

[0104] Then, item In state The Taylor expansion generated at this point is:

[0105] (36);

[0106] Through negation And using HL-PWL representation, equation (36) is written as:

[0107] (37);

[0108] in, express The derivative is Then the robustness condition (33) can be written as:

[0109] (38);

[0110] It can be written as the following quadratic inequality:

[0111] (39);

[0112] in It is a column vector. X is a real symmetric matrix; since X is a rank-1 matrix, it has a non-zero eigenvalue that equals The square of the norm, i.e. The remaining eigenvalues ​​are zero, let

[0113] (40);

[0114] but Rewritten as:

[0115] (41);

[0116] in It is an orthogonal identity matrix Therefore, formula (39) can be further rewritten as: (42);

[0117] make Then the following condition is met:

[0118] (43);

[0119] Equation (43) shows that, Convert to standard quadratic form, There must exist a correct solution;

[0120] However, when control is saturated at the lower and upper bounds, if for all , So regardless of the coefficient No matter how it changes, the entire quadratic form is less than zero. Therefore, the necessary condition for the existence of this solution, namely the robust condition (33), is written as:

[0121] (44);

[0122] The aforementioned necessary conditions require the controller to have sufficient capability to exceed the limits set by the controller. arrive By utilizing the uncertainty range and this necessary condition, infeasible states are found, and the attraction domain D is improved, eliminating states that the method cannot guarantee.

[0123] Since system (2) is piecewise linear, control is evaluated at its boundary vertices to obtain each state. reachable vertex set For each vertex, a cone is obtained; then, the union of these cones is constructed to approximate a new cone. ;

[0124] Now, Each vertex in A cone in corresponding space The union of these cones C is The solution space is considered, and the cones corresponding to the four vertex states are aligned. Considering the constraints imposed by formula (30), only four cones are shown. The intersection of these four cones is approximately a linear cone.

[0125] Without losing generality, assume the cone has N Intersection points These intersection points form N planes in the solution space, with normal vectors... The linear cone formed by these N intersection points is represented by calculating each plane:

[0126] (45);

[0127] (46);

[0128] Then, the Lyapunov candidate function The solution for the coefficients is:

[0129] (47);

[0130] in, The coefficients are the HL-PWL representations of the Lyapunov function and its boundary. The attraction domain was initially verified using the solution space (47) before applying the proposed RGRL to learn the controller and the coupled Lyapunov function.

[0131] Therefore, the present invention employs the aforementioned worst-case-based Lyapunov-stabilized reinforcement learning control algorithm for uncertain systems, with the following beneficial effects:

[0132] (1) By modeling the uncertainty of the ReLU activation network and the system dynamics model, this invention can still accurately find the most inverse state, thereby forcing its stability under uncertainty.

[0133] (2) This invention discusses the solvability of the problem of robust controller and Lyapunov function, and derives the necessary condition for determining the attraction domain before solving the problem.

[0134] (3) This invention provides a geometric view of the existence of solutions to robust RL problems to explain robustness and its capability. Numerical simulations of inverted pendulums and quadcopters under various uncertainties prove the effectiveness of the proposed method.

[0135] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0136] Figure 1 This is a schematic diagram of the solution space composed of cones, representing an embodiment of the Lyapunov stable reinforcement learning control algorithm for worst-case uncertain systems according to the present invention.

[0137] Figure 2 This is a simulation result of an inverted pendulum in a precision system using NNC in an embodiment of the worst-case uncertainty system Lyapunov stability reinforcement learning control algorithm of the present invention, where (a) is a trajectory illustration, (b) is the Lyapunov value, and (c) is the control value;

[0138] Figure 3 This invention compares the NNC and RNC under observation uncertainty in an embodiment of the worst-case uncertainty system Lyapunov stable reinforcement learning control algorithm, where (a) is the NCC trajectory, (b) is the RNC trajectory, (c) is the Lyapunov value of the NCC, (d) is the Lyapunov value of the RNC, (e) is the control signal of the NCC, and (f) is the control signal of the RNC.

[0139] Figure 4 This invention compares the NNC and RNC controllers under observation uncertainty in an embodiment of the worst-case uncertainty system Lyapunov stable reinforcement learning control algorithm, where (a) is the trajectory and Lyapunov value generated by the NCC, (b) is the trajectory and Lyapunov value generated by the RNC, (c) is the control signal of the NCC, and (d) is the control signal of the RNC.

[0140] Figure 5 This invention compares the NNC and RNC controllers in the observation uncertainty multiplicative form of the worst-case uncertainty Lyapunov stable reinforcement learning control algorithm embodiment, where (a) is the NCC trajectory, (b) is the RNC trajectory, (c) is the Lyapunov value of the NCC, (d) is the Lyapunov value of the RNC, (e) is the control signal of the NCC, and (f) is the control signal of the RNC.

[0141] Figure 6 This invention compares the NNC and RNC in the form of observation uncertainty multiplication in the worst-case uncertainty Lyapunov stable reinforcement learning control algorithm embodiment, where (a) is the NCC trajectory, (b) is the RNC trajectory, (c) is the NCC control signal, and (d) is the RNC control signal.

[0142] Figure 7 This invention compares the controllers under multiplicative and additive uncertainties in an embodiment of the worst-case uncertain Lyapunov stable reinforcement learning control algorithm for systems, where (a) is... The trajectory below, (b) is The trajectory below, (c) is The Lyapunov value (d) is given by The Lyapunov value (e) is... The control signal below, (f) is The control signal below;

[0143] Figure 8 This invention relates to the convergence efficiency of the compressed space under state space conditions in an embodiment of the worst-case uncertain system Lyapunov stable reinforcement learning control algorithm.

[0144] Figure 9 This invention provides the loss function over the entire state space during the training process of an embodiment of the worst-case uncertainty system Lyapunov stable reinforcement learning control algorithm, where (a) is the loss function over the entire state space after 1 training cycle, (b) is the loss function over the entire state space after 20 training cycles, and (c) is the loss function over the entire state space after 100 training cycles.

[0145] Figure 10 This invention compares the NCC and RNC controllers under observation uncertainty in an embodiment of the Lyapunov stable reinforcement learning control algorithm for worst-case uncertain systems, where (a) is... The controller below, (b) is The controller below;

[0146] Figure 11 This invention presents a comparison of a two-dimensional quadcopter under multiplicative and additive uncertainties in an embodiment of the Lyapunov stable reinforcement learning control algorithm for worst-case uncertain systems, where (a) is... The trajectory below, (b) is The trajectory below, (c) is The Lyapunov value (d) is given by The Lyapunov value (e) is... The control signal below, (f) is The control signals below. Detailed Implementation

[0147] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0148] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.

[0149] Robust reinforcement learning (robustReinforcement Learning (RL) optimizes the expected cumulative reward in the worst case through bi-level optimization. However, when uncertainties are present, it is difficult to obtain strict robustness guarantees, thus limiting the credibility of RL in practical applications. Without involving uncertainties, the stability problem of RL methods has been partially solved by using auxiliary loss functions or barrier functions based on Lyapunov theory. This invention presents a unified approach to stable RL under bounded uncertainty in dynamic models and state estimation.

[0150] As shown in the figure, the Lyapunov-stabilized reinforcement learning control algorithm for uncertain systems based on the worst-case scenario includes the following steps:

[0151] S1. Model the system dynamics and uncertainty boundaries using a ReLU activation network. The specific process is as follows:

[0152] Consider a discrete-time robot system with the following dynamic model:

[0153] (1);

[0154] ∈ and ∈ These represent the robot in t The state and control input at any given time, with the control saturation controlled by the lower bound. or upper boundary Give;

[0155] The dynamic model is given, assuming its approximate value is... The results were obtained through system identification, as shown below:

[0156] (2);

[0157] and These represent the multiplicative uncertainty and additive uncertainty between the actual dynamics and its approximation, respectively. This indicates the uncertainty of the centralized total model. Indicates controller In state The control provided at the location;

[0158] However, system state It is estimated from some estimator, and therefore will be corrupted by noise, as shown below:

[0159] (3);

[0160] in and Let these represent the multiplicative uncertainty and additive uncertainty of the state estimation, respectively. ( ) represents the lumped uncertainty in state estimation;

[0161] therefore, t The predicted state at time +1 is:

[0162] (4);

[0163] in, Indicates controller π In state estimation The control provided at the location;

[0164] For an approximate dynamic model ( , The output system estimates the state estimator. Controlling saturation and And with an attractive domain D, find a controller This makes the system subject to model uncertainty. and state uncertainty It stabilizes at the equilibrium point.

[0165] controller Typically, in the established approximate dynamic model The training is performed in a simulator and then applied to a real-world system. This creates a gap between the simulation and reality in the dynamic model. Furthermore, it is also achieved through state estimation. controller To obtain control input, both of these factors can lead to instability in the closed-loop robot system.

[0166] Assumption 1: An approximate dynamic model is given in the form of a ReLU activated neural network. ;

[0167] Assumption 2: Assume uncertainty in the dynamics of the system model. Uncertainty in state estimation Bounded, and its boundaries are denoted as follows: and Then we have:

[0168] (5);

[0169] in, For the infinite norm, and They are all presented in the form of ReLU activated neural networks.

[0170] The ReLU-activated network allows for flexible modeling of system dynamics and the boundaries of uncertainties, making the proposed RGRL method applicable to many complex systems. More importantly, through Assumptions 1 and 2, the robustness condition can be transformed into a set of linear constraints, and states violating these constraints can be efficiently found using Mixed Integer Linear Programming (MILP). In addition to solving for the robust controller, the attraction domain D is refined because the antagonistic relationship between control saturation and uncertainty can lead to states failing to satisfy the robustness condition.

[0171] S2. Determine the robustness conditions and use them to predetermine the area of ​​the attraction domain;

[0172] The robustness conditions used to design the RGRL loss function are specifically derived to derive the stability conditions under the worst-case uncertainty to regulate the training of the controller and the Lyapunov function.

[0173] The robustness condition is the stability condition under the worst-case scenario of uncertainty, derived from Lyapunov neural control. When uncertainty is not involved, it is achieved by training coupled Lyapunov functions that satisfy the following stability condition. :

[0174] (6);

[0175] (7);

[0176] (8);

[0177] in D To attract domains, It is in equilibrium. ( )and ( ) belongs to the K-function set, and stability conditions (6) to (8) guarantee that from D A system starting from any state converges to an equilibrium state. .

[0178] The following sections discuss the uncertainties of a given dynamic model and the robustness conditions of state estimation.

[0179] This invention explicitly considers the bounds of uncertainty in dynamical modeling and specifically guarantees the worst-case scenario. This contrasts with domain-randomized RL methods, which optimize the expectation of reward over these uncertainties without providing guaranteed robustness. More importantly, the derived robustness conditions are linear over the system state, allowing the use of MILP to find states that violate these conditions.

[0180] Stability conditions under worst-case uncertainty include robustness conditions under dynamic model uncertainty and robustness conditions under state estimation uncertainty:

[0181] Model uncertainty in dynamics Down, t The state at time +1 is not entirely controlled. and current state Therefore, equation (8) must be applied to... t Each possible state at time +1 If both are valid, then:

[0182] (9);

[0183] in, (10);

[0184] Through the The robustness condition (8) of the Lyapunov function at the given position is obtained by Taylor expansion:

[0185] (11);

[0186] Given Based on the randomness of the randomness and on assumption two, a stronger and more sufficient robustness condition is obtained, as follows:

[0187] (12);

[0188] in, and The difference is negligible. The worst-case scenario of uncertainties in the dynamic modeling is considered, and the Lyapunov value is forced to decrease along the state trajectory.

[0189] Note that, given the parameters of the Lyapunov network, the Lyapunov function... derivative Represented as (wrt), state It is a piecewise constant. Furthermore, the uncertainty boundary... and Lyapunov functions in The expression is piecewise linear. Therefore, condition (12) is valid. The expression is piecewise linear.

[0190] Robustness conditions under uncertainty in state estimation

[0191] Compared with the given ground reality The control input obtained Conversely, controller In fact, in state estimation The input is calculated to determine the control input. Therefore, the stability condition needs to take into account the randomness of the state estimation.

[0192] Considering the uncertainty in the worst-case scenario of state estimation, in The Lyapunov value at point is obtained through Taylor expansion as follows:

[0193] (13);

[0194] The derivative of the Lyapunov function is obtained by the chain rule, as shown in the following formula:

[0195] (14);

[0196] in, ;

[0197] The robustness condition under uncertainty in state estimation is as follows:

[0198] (15);

[0199] Under assumption two, the stability condition (8) becomes the new robustness condition as follows:

[0200] (16);

[0201] in, It was ignored.

[0202] Note 3: Derivative , and They are , and The piecewise constant in. Therefore, condition (16) is in The middle is piecewise linear.

[0203] S3. Determine the robustness guarantee RL under the uncertainty of the dynamic model and state estimation;

[0204] Robust Guaranteed RL (RGRL) under uncertainties in dynamic modeling and state estimation is similar to the guaranteed stability RL for controller networks. and Lyapunov network They are trained at the same time.

[0205] Assumption This represents the trainable parameters in a Lyapunov network. Let represent the trainable parameters in the controller network. Using robustness conditions (7), (12), and (16), the following loss function is designed to regulate the training of the controller network and the Lyapunov network. The loss function that guarantees the semipositive definition of the Lyapunov network (7) is given as follows:

[0206] (17);

[0207] When only uncertainty is considered in the model dynamics At that time, the Lyapunov value along the trajectory should increase strictly monotonically, and this is guaranteed by satisfying the robustness condition (12). The corresponding loss function is designed as follows:

[0208] (18);

[0209] When only the uncertainty in the state estimation is considered When the robustness condition is met, the loss function (16) is expressed as:

[0210] Based on the robustness condition (16), which represents a system with only uncertainties in the dynamic model, a loss function is constructed:

[0211] (19);

[0212] When both uncertainties mentioned above are considered simultaneously, the robustness condition (8) becomes:

[0213] (20);

[0214] Then, the loss function corresponding to the robustness condition eq is obtained, expressed as:

[0215] (twenty one);

[0216] Note that calculating the above loss function involves solving for the optimal system state that most violates the robustness condition. Inspired by Lyapunov neural control, the activation functions in the controller network, Lyapunov network, and uncertainty boundary network are "leakage-ReLU", making the controller, Lyapunov function, and uncertainty boundary piecewise linear. Let... Represents the set of all trainable parameters, in each training iteration, when solving for the state that primarily violates the loss function (18), (19), or (21). At that time, the trainable parameters of the previous iteration θ It is fixed and used. This state is also called a "counterexample", and the loss functions (18), (19) and (21) are all... Piecewise linear functions in the model.

[0217] The attraction domain D is decomposed into multiple partitions, by Confirmed. Let's assume... It is one if and only if In the If the binary number in the partition is equal to 1, then... A convex combination of the extreme points of a partition, represented by a slack variable, can be used to find counterexamples of the loss function (18), (19), or (21). This can be transformed into a mixed-integer linear programming (MLIP) problem, as shown below:

[0218] The objective function and all constraints are expressed as follows: Piecewise linear functions:

[0219] set up This represents the set of all trainable parameters, which are used in each iteration of training with the values ​​from the previous iteration. Attraction domain S The division is as follows:

[0220] set up It is if and only if In the If the binary number in the partition is equal to 1, then... The convex combination of the extreme points of the partition, represented as a partition, transforms the optimization problem of the loss function into a mixed-integer linear programming problem by introducing slack variables:

[0221] (22a);

[0222] (22b);

[0223] Among the problem coefficients All Explicit functions, optimizing variable vectors s Depend on and If we consider the composition, then the optimal cost of equations (26a) and (26b) above can be written in closed form: , The gradient is calculated by backpropagating the closed-form expression. .

[0224] Among them, the problem coefficient , , and All The explicit function is considered fixed when searching for "counterexamples". The optimization variable vector s consists of x and ...composed of, the optimal value of the objective function (22) can be written as:

[0225] ;

[0226] Using gradients Trainable parameters are obtained through backpropagation. Optimize.

[0227] Algorithm 1 below summarizes the proposed RGRL, which finds... Fixed "counterexample" state (Step 3 in Algorithm 1) and its application in the found "counterexamples" gradient calculated on renew Iterate between (step 5 in Algorithm 1).

[0228] Algorithm 1 is robust to guarantee RL:

[0229] 1. Initialize the controller and Lyapunov functions ;

[0230] 2. When iterations ≤ maximum value:

[0231] 3. Maximize in equation (22) using the Gurobi solver To find ;

[0232] 4. Calculate the gradient ;

[0233] 5. Update ,in It's the step length.

[0234] Please note that the robust guarantee of RL solutions depends to a great extent on uncertainties, system dynamics, control saturation, attraction domain, and the flexibility of the neural network structure. In fact, the solvability of Problem 1 has been embedded in the MILP problem in (22) and will be addressed in subsequent studies.

[0235] Solubility analysis

[0236] The RGRL proposed in Algorithm 1 involves finding "counterexample" states through MLIP and adjusting the trainable parameters. θ Optimization is required. Due to the antagonistic relationship between the magnitude of uncertainty and control saturation, it is possible that no controller can stabilize certain states and guarantee robustness. Therefore, before using RGRL to find a robust controller and coupled Lyapunov function, the solvability of the region of attraction, given the uncertainty and control saturation, should be carefully examined.

[0237] Since there always exists a PWL function corresponding to a ReLU-activated neural network, the PWL function has been widely applied in many fields such as control systems, signal processing, optimization problems, and circuit design. PWL systems can utilize a wealth of analytical tools. When there is no uncertainty or control saturation, if the PWL system is affine, there exists a piecewise linear Lyapunov function and a stable controller. Furthermore, if the states of the vertices of a partition satisfy the stability condition, then all states in the partition satisfy the stability condition. Let H denote the vertices of the state partition; this step advances this work by considering control saturation and uncertainty.

[0238] S4. Network parameterization is established, the process is as follows:

[0239] To further formulate problem (22) in a more compact form without resorting to slack variables, a high-level canonical representation of the PWL function, known as HL-PWL, is adopted. HL-PWL has been shown to use fewer parameters than the classic PWL while maintaining the same accuracy, expressed as:

[0240] Using the HL-PWL method, an arbitrary piecewise linear function can be represented by the following expression. F ( x ):

[0241] (twenty three);

[0242] in, It is a coefficient vector. It is the stack of base functions, base functions The stack includes nested layers. All nested functions, coefficient vector Nested functions Related;

[0243] Robustness conditions under HL-PWL representation:

[0244] The positive determination of the Lyapunov function, formula (7) is rewritten as:

[0245] (twenty four);

[0246] in and yes K Function, and with The Lyapunov function increases with the increase of . and K The functions are all rewritten in HL-PWL form (23), and the expression is:

[0247] (25);

[0248] (26);

[0249] in, For the infinite norm, a∈ and ∈ For parameter vectors, and The number of corresponding parameter vectors, upper bound, and lower bound. , ∈ and It is a vector function, including nested functions at nesting levels 0 and 1. It is easy to obtain (i.e., |x|max).

[0250] because It is K The function then has:

[0251] (27);

[0252] (28);

[0253] The robustness condition (6) and positive condition (7) under equilibrium conditions are written as follows:

[0254] (29);

[0255] and

[0256] (30);

[0257] Partition and define vertices for a given attraction domain D, and then transform condition (30) into a 6-form. ,in ,matrix and By superposition V All vector The result of superposition, among which V It is a collection of D partitions.

[0258] Similarly, the upper bound condition of equation (24) can be rewritten as follows: Then the upper boundary and the lower realm It is a K-function that satisfies equation (27), that is:

[0259] (31);

[0260] and functions that satisfy (28), namely:

[0261] (32);

[0262] in, and For the PWL system, the mesh size for the partition is determined by analyzing the vertices. The current state is sufficient. Note that in the rest of the discussion, the state... It is assumed to be one of the vertices.

[0263] When considering uncertainties in both dynamic modeling and state estimation, robustness conditions (20) and (14) produce another condition:

[0264] (33);

[0265] Further assume that both the dynamic system and the controller satisfy the following Lipshitz conditions:

[0266] (34);

[0267] (35);

[0268] in and F These represent the upper bounds of the controller and the positive dynamic change, respectively;

[0269] Then, item In state The Taylor expansion generated at this point is:

[0270] (36);

[0271] Through negation And using HL-PWL representation, equation (36) is written as:

[0272] (37);

[0273] in, express The derivative is Then the robustness condition (33) can be written as:

[0274] (38);

[0275] It can be written as the following quadratic inequality:

[0276] (39);

[0277] in, It is a column vector. X is a real symmetric matrix. Since X is a rank-1 matrix, it has a non-zero eigenvalue that equals... The square of the norm, i.e. The remaining eigenvalues ​​are zero, let

[0278] (40);

[0279] but Rewritten as:

[0280] (41);

[0281] in It is an orthogonal identity matrix Therefore, formula (39) can be further rewritten as:

[0282] (42);

[0283] make Then the following condition is met:

[0284] (43);

[0285] Equation (43) shows that, Convert to standard quadratic form, There must exist a correct solution;

[0286] However, when control is saturated at the lower and upper bounds, if for all , So regardless of the coefficient No matter how it changes, the entire quadratic form is less than zero. Therefore, the necessary condition for the existence of this solution, namely the robust condition (33), is written as:

[0287] (44);

[0288] The aforementioned necessary conditions require the controller to have sufficient capability to exceed the limits set by the controller. arrive By utilizing the uncertainty range and this necessary condition, infeasible states can be found, and the attraction domain D can be improved to eliminate states that cannot be guaranteed by this method.

[0289] Since system (2) is piecewise linear, control is evaluated at its boundary vertices to obtain each state. reachable vertex set For each vertex, a cone is obtained; then, the union of these cones is constructed to approximate a new cone. .

[0290] Now, Each vertex in A cone in corresponding space The union of these cones C is The solution space, with Figure 2 For example, considering the cones corresponding to the four vertex states, the cones are aligned, and the constraints imposed by formula (30) are considered. Only four cones are shown, and the intersection of these four cones is approximately a linear cone, such as... Figure 2 As shown by the red dot in the image.

[0291] Without losing generality, assume the cone has N Intersection points (like Figure 1 (As shown by the red dots in the diagram). These intersections form N planes in the solution space, with normal vectors... The linear cone formed by these N intersection points is represented by calculating each plane:

[0292] (45);

[0293] (46);

[0294] Then, the Lyapunov candidate function The solution for the coefficients is:

[0295] (47);

[0296] in, The coefficients are the Lyapunov function and its boundary HL-PWL representation. The attraction domain was initially verified using the solution space (47) before applying the proposed RGRL to learn the controller and the coupled Lyapunov function. So far, the control problem has been linearized for approximation. This planning problem can be used for initial verification before training the neural network, and the characteristics of linear programming can be used to analyze the control problem.

[0297] Determining the attraction area

[0298] In existing literature, the attraction region is often given. D . V Some vertices in the region may not satisfy the necessary condition (44), making the problem unsolvable. Before applying the proposed RGRL, the necessary condition is used to refine the given attraction region. D Given upper and lower bounds for control. and uncertainty boundary and In the attraction area DThe vertices that satisfy the necessary condition (44) form a new set. Then a precise attraction area was obtained. .if The vertices in the equation satisfy the necessary condition (44), then All states satisfy this condition.

[0299] Robust conditional linearized inference

[0300] Each vertex has a reachable domain. From each vertex ,Right now By refining the attraction area any state All of them meet the necessary condition (44).

[0301] Proposition 1: When using the proposed RGRL, when uncertainty boundary and When magnified, the controller becomes more robust, in which For the maximum value, any Both will make RL unsolvable, where:

[0302] (48);

[0303] Proof 1: Since the first term in the necessary condition (44) is positive, while the others are negative, this indicates that the quadratic form forms a second type of cone (also known as a hyperbolic cone) in space. To calculate the cone angle of this cone, the cone angle can be approximated by considering the angle between the cone and the principal axis. Furthermore, since the principal axis of the cone is aligned with the direction of the eigenvector corresponding to the positive eigenvalue, the half angle of the cone is... θ The following relationship can be used to approximate it:

[0304] (49);

[0305] When the uncertainty of the system changes from Increase to At that time, if Then the new half-width It has become the following formula:

[0306] (50);

[0307] Since the tangent function is monotonically increasing, when the uncertainty of the system is proportional to the scaling factor... When magnified, the cone angle of the quadratic shape decreases:

[0308] (51);

[0309] in, and This represents the solution space; furthermore, the size of the cone angle is equivalent to the solution space. The size of the solution space decreases as the cone angle decreases. This makes the proposed RGRL method challenging, but enables the controller to handle increased uncertainty.

[0310] Then the coefficient is obtained. The necessary condition for the maximum value of (44). By reducing the above state space, all vertex states should satisfy the necessary condition. For all , The upper limit of the magnification factor The following conditions must be met:

[0311] (52).

[0312] S5. Numerical simulation of inverted pendulum and quadcopter UAV.

[0313] Numerical simulations of an inverted pendulum and a quadcopter UAV were used to compare the robustness-guaranteed RGRL proposed in this invention with Lyapunov neural network control (NNC), with particular consideration given to multiplicative and additive uncertainties in the dynamic model and state estimation. The forms of multiplicative and additive uncertainties are shown in Equation (3). Training and testing were both performed on an Intel i7 CPU.

[0314] A. Inverted pendulum

[0315] The simulated inverted pendulum is 1 meter long, weighs 1 kilogram, and has a gravitational acceleration set to 9.81. The simulation was performed in Gym, with a torque control range of -100 Nm to 100 Nm. Let... θ This represents the angle between the pendulum and the horizontal x-axis. Let the equilibrium point be... , The region of attraction D in the state space (angle and angular velocity) is specified as... The controller and Lyapunov function were synthesized using the RGRL proposed in this invention, and the resulting controller is called the Robust Neural Controller (RNC). Three scenarios were simulated to illustrate the effectiveness of the RNC in handling various uncertainties.

[0316] In scenario I, only multiplicative observation uncertainty is considered. In Scenario II, multiplicative observation uncertainty is also considered. Uncertainty in the sum-multiplication model Scenario III considers various uncertainties. and .

[0317] 1) Scenario I: Multiplicative observation uncertainty in state estimation:

[0318] In scenario I, multiplicative observation uncertainty Set as The NNC and RNC were derived from Lyapunov neural control and the proposed robust guarantee RL, respectively. The NNC and RNC converged within 3 and 5 hours, respectively. The controller was tested under 30 different initial states, and the obtained trajectories were... Figure 3 The values ​​are visually represented by various colors, with a red star indicating the equilibrium point. In the case of NNC (Non-Normalized Control), the system stabilizes at a point such as... Figure 3 (a) shows the blue star state.

[0319] The differences between the two stars highlight the presence of multiplicative observational uncertainty. In such cases, the NNC cannot effectively stabilize the inverted pendulum at the required equilibrium point. Conversely, as... Figure 3 As shown in (b), the trajectories derived from RNC all converge to the desired equilibrium point. This is demonstrated by examining... Figure 3 The same conclusion can be drawn from the evolution of the Lyapunov function over time shown in (c) and (d). Initially, the Lyapunov values ​​generated by both methods decrease. The Lyapunov value generated by RNC exhibits a consistent monotonically decreasing trend, eventually approaching zero. In contrast, the Lyapunov value generated by NNC shows an increasing trend before converging to a non-zero constant value.

[0320] The control signals corresponding to the two methods are as follows: Figure 3 As shown in (e) and (f), similar to the Lyapunov value, the control signal from the NNC converges to a non-zero value, while the control signal from the RNC converges to zero. The RNC proposed in this invention rigorously regulates the controller training process through a robust loss based on state estimation uncertainty. This method effectively counteracts the effects of this uncertainty, ensuring that the system converges to the desired equilibrium state. The robustness guarantee provided by the RNC can be evaluated by assessing the value function based on robustness constraints.

[0321] Figure 4 (a) and (b) present example trajectories generated by NNC and RNC, respectively, and their corresponding Lyapunov functions. The Lyapunov functions for both methods are visualized as three-dimensional cones, with contour lines projected onto a base plane. These contour lines are centered at the equilibrium point. Red stars represent the initial state, and blue stars represent the convergent state. The trajectories are shown alongside the Lyapunov function values, with the system state progressing along the Lyapunov function cone and converging towards the base. Figure 4When there is observation uncertainty in the state estimation shown, NNC is unlikely to reach the expected equilibrium point.

[0322] 2) Scenario II, Multiplicative Uncertainties in Dynamic Model and State Estimation:

[0323] Uncertainty of the dynamic model Set as To reduce the uncertainty in state estimation Set as .like Figure 5 As shown in (a), the NNC struggles to handle uncertainties in the dynamic model and state estimation, causing the trajectory to gradually deviate from the desired equilibrium point. Therefore, the inverted pendulum continues to rotate without reaching a stationary state. This perpetual motion leads to a continuous increase in the overall energy of the controlled system, resulting in an exponential increase in the Lyapunov value and control saturation, as shown in (a). Figure 5 (c) and Figure 5 As shown in (e).

[0324] In contrast, RNC cleverly mitigates the effects of these two uncertainties, effectively stabilizing the control system and guiding it towards the desired equilibrium state. The Lyapunov value drops to zero, as... Figure 5 As shown in (d). In Figure 5 In (b), all trajectories converge rapidly to a diagonal trajectory, while... Figure 5 In (b), from the initial state The trajectory of the initial state deviates slightly, and then converges to the diagonal.

[0325] This is because when the state is far from the equilibrium point, the multiplicative uncertainty requires a larger control input, such as... Figure 5 As shown in (f). At the beginning of the trajectory, the control is typically saturated, leading to rapid convergence to the diagonal trajectory. Through Figure 5 The monotonically decreasing Lyapunov value in (d) verifies the robustness of the guarantee. The figure shows an example trajectory generated by NNC and RNC and its corresponding Lyapunov function.

[0326] like Figure 6 As shown in (a) and (b), an example trajectory and its corresponding Lyapunov function generated by NNC and RNC, respectively, are presented. The Lyapunov functions for both methods are visualized as 3D surfaces with isomorphic contours projected onto the base plane. These contours are centered at the equilibrium point. Red stars represent the initial state, and blue stars represent the desired equilibrium state. The trajectory is shown next to the Lyapunov function values. The system state proceeds along the Lyapunov function cone and converges towards the base.

[0327] Under the combined influence of uncertainties in the dynamic model and state estimation, the NNC (Non-Normalized Control) proved ineffective in guiding the trajectory of the inverted pendulum along the Lyapunov function surface, resulting in unstable motion with oscillating trajectories and causing the controller output to remain saturated. In contrast, the RNC (Reverse Control) effectively mitigated the effects of these uncertainties, allowing the trajectory to descend smoothly along the Lyapunov surface and reach the target position.

[0328] 3) Scenario III: Multiplicative and Additive Uncertainties in Dynamic Models and State Estimation

[0329] In addition to multiplicative uncertainties, Scenario III also simulates additive uncertainties in the dynamic model and state estimation. Additive uncertainty and Set as follows and 0.01. exist Figure 7 In the figure, the graph corresponds to the performance of the two controllers under additive and multiplicative uncertainties. A single trajectory in the graph corresponds to... Figure 5 The diagram shows a specific trajectory within the trajectory set of the controller. It can be observed that the trajectory under multiplicative uncertainty converges smoothly to the equilibrium point, where both the Lyapunov function and the control function demonstrate the ability to stably converge to the equilibrium state.

[0330] Next, we focus on the trajectory under additive uncertainty; as the trajectory approaches equilibrium, we observe oscillations rather than stable convergence. Furthermore, the Lyapunov function fluctuates near the convergence point, indicating that the controller can only guarantee stable convergence within a certain range. Within this range, we cannot guarantee that the value of the Lyapunov function will continuously decrease.

[0331] Overall, in uncertain systems containing constant terms, the method provided by this invention can guarantee a reduction in overall energy. The final, unguaranteed convergence range depends on the constant terms in the uncertainty.

[0332] B. Refinement of the attraction field

[0333] In this example, the uncertainties in the dynamic model and state estimation Set as The attraction region was refined before applying the proposed RGRL. The results are as follows: Figure 8 As shown, the blue area does not meet the necessary conditions and is therefore excluded, while the refined attraction area is described by the red dashed box.

[0334] C. Robustness Enhancement

[0335] During training, uncertainty was deliberately introduced. The factor amplification results in a second RNC (called the enhanced RNC). Figure 10 In the diagram, blue circles indicate the state where the Lyapunov value increases when using the regular RNC, while red circles indicate the state where the Lyapunov value increases when amplified by the amplification factor. The results show that using the enhanced RNC performs better than the regular one, resulting in a smaller unstable region.

[0336] D . Learning Lyapunov functions

[0337] The learning process of the Lyapunov function is visualized. In RGRL, the Lyapunov function is trained based on the loss in the formula. Formulas (17), (18), and (19) realize the maximization operation in the bilayer optimization by solving the MILP problem. The degree of violation of the robustness condition is analyzed. Figure 9 In this invention, three loss snapshots of the entire state space after 1, 20, and 100 training epochs are described. Initially, the state space is predominantly yellow-green, indicating violations of the robustness conditions of the initial controller network and the Lyapunov network. Initially, "counterexample" states are those far from the equilibrium state. As training progresses, the "counterexample" states gradually approach the equilibrium state. Figure 9 The example in [the document] supports this point; during training, we enforced priority on peripheral states to meet robustness conditions.

[0338] E. Two-dimensional quadcopter

[0339] Simulate a two-dimensional quadcopter with a rotor arm length of 0.25m, a quadcopter mass of 0.486kg, and a moment of inertia of 0.0083. The gravitational acceleration is 9.81. The activation area is designed as follows: The Lyapunov condition was verified in the region, and the NNC and RNC converged within 6 hours and 12 hours, respectively.

[0340] Two scenarios were simulated: scenario VI included only multiple uncertainties in the dynamic model and state estimation, while scenario V included only additive uncertainties. and Set as and In both scenarios, the aircraft launches within a box domain, such as... Figure 11 The black box shown is set to a balanced state. .

[0341] The trajectories obtained by NNC and RNC in scene VI are as follows Figure 11As shown in (a). Under RNC control, the aircraft can stably fly to the equilibrium point shown in red, while NNC cannot stabilize the aircraft and eventually flies out of the black box. This is also evident in the comparison of Lyapunov values. The Lyapunov value of NNC increases from the beginning, while the Lyapunov value of RNC decreases monotonically and converges to zero after 10 seconds. The control signals of NNC and RNC are as follows: Figure 11 As shown in (e), it can be seen from the figure that the NNC cannot stabilize the aircraft, and its control saturates continuously after 2 seconds, while the RNC reaches saturation at the beginning and drops to a constant value after the aircraft stabilizes.

[0342] The trajectories obtained by NNC and RNC in scene V are as follows Figure 11 As shown in (b), in this case, both NNC and RNC are able to bring the aircraft to the desired equilibrium point. However, RNC makes the aircraft stably fly to the equilibrium point shown in red, while NNC fails and cannot stabilize the aircraft at the equilibrium point. This is also evident in the comparison of Lyapunov values. The Lyapunov values ​​of both methods decrease, but the Lyapunov value of the NNC method may increase from time to time, indicating that robustness is not guaranteed. The Lyapunov value of RNC decreases monotonically and converges to zero after 10 seconds. The control signals of NNC and RNC are as follows: Figure 11 As shown in (f), the results show that the output control of NNC has frequent fluctuations, while the output signal of RNC is smooth and has a certain degree of robustness.

[0343] Therefore, this invention employs the aforementioned worst-case-based Lyapunov-stabilized reinforcement learning control algorithm for uncertain systems to synthesize a neural network controller, thereby maintaining the stability of the system under bounded multiplicative and additive uncertainties in the dynamic model and state estimation. To ensure the robustness of the neural network, a set of inverse business losses is proposed to train the Lyapunov function of the neural network. Since maximizing the loss function is transformed into mixed-integer linear programming, the bi-level optimization is solved by a Gurobi solver. A robustness criterion is proposed to guarantee the solvability of the RL problem and is used to determine the possible attraction domain. Guaranteed robustness is demonstrated through tests on a simulated inverted pendulum and a two-dimensional quadcopter, and performance comparisons with existing methods. The robustness criterion strictly requires the Lyapunov value to decrease along the system trajectory, making the solution overly conservative.

[0344] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A Lyapunov-stabilized reinforcement learning control algorithm for worst-case uncertain systems, characterized in that, Includes the following steps: S1. Model the system dynamics and uncertainty boundaries using a ReLU activation network; S2. Determine the robustness conditions and use them to predetermine the attraction domain; S3. Reinforcement learning RGRL that determines robustness guarantees under uncertainties in dynamic models and state estimation; S4. Network parameterization setup; S5. Numerical simulation of inverted pendulum and quadcopter UAV; Assumption 2: Assume the system model is uncertain. and state uncertainty Bounded, and its boundaries are denoted as follows: and Then we have: (5); Where |·| is the infinite norm, and They are all presented in the form of ReLU activated neural networks; In step S2, the robustness condition is the stability condition under the worst-case uncertainty, derived from Lyapunov neural network control. When uncertainty is not involved, it is achieved by training a coupled Lyapunov function that satisfies the following stability condition. V ( ): →R is used to ensure system stability: (6); (7); (8); in D To attract domains, It is in equilibrium. ( )and ( ) belongs to the K-function set, and stability conditions (6) to (8) guarantee that from the region of attraction D A system starting from any state converges to an equilibrium state. ; In step S2, the stability condition under the worst-case uncertainty includes robustness conditions under model uncertainty and robustness conditions under state uncertainty: In model uncertainty Down, t The state at time +1 is not entirely controlled. With robots t state of time Therefore, equation (8) must be applied to... t Predicted state at time +1 If it is valid, then: (9); in, (10); Through the The stability condition (8) of the Lyapunov function at the given location is obtained by Taylor expansion: (11); Given Based on the randomness of the randomness and on assumption two, a stronger and more sufficient robustness condition is obtained, as follows: (12); in, and The difference is negligible, considering the worst-case scenario of uncertainties in dynamic modeling, and forcing the Lyapunov value to decrease along the state trajectory; With a given robot t state of time Gained control Conversely, controller In fact, in state estimation The value is calculated to determine the control estimate. Therefore, the stability condition needs to take into account the stochasticity of the state estimation; considering the uncertainty in the worst-case state estimation, in The Lyapunov value at point is obtained through Taylor expansion as follows: (13); The derivative of the Lyapunov function is obtained by the chain rule, as shown in the following formula: (14); in, ; The robustness condition under state estimation uncertainty is as follows: (15); Under assumption two, the stability condition (8) becomes the new robustness condition as follows: (16); in, Ignored, derivative , and They are states ,control and controller In state Piecewise constants; In step S3, let This represents the trainable parameters in a Lyapunov network. β Let represent the trainable parameters in the controller network. Using conditions (7), (12), and (16), design the following loss function to regulate the training of the controller network and the Lyapunov neural network. The semi-positive definitional loss function that guarantees the stability condition (7) is given as follows: (17); When only model uncertainty is considered in model dynamics δ At that time, the Lyapunov value along the trajectory should increase strictly monotonically, and this is guaranteed by satisfying the robustness condition (12). The corresponding loss function is designed as follows: (18); When only considering the state uncertainty ϵ in the state estimation, the loss function is designed according to the robustness condition (16): (19); When both model uncertainty and state uncertainty are considered, the stability condition (8) becomes: (20); Then, the loss function corresponding to the stability condition formula (20) is obtained, expressed as: (21); The objective function and all constraints are expressed as follows: Piecewise linear functions: set up This represents the set of all trainable parameters, using the values ​​from the previous iteration in each training iteration. Attraction domain D The division is as follows: set up It is if and only if X In the If the binary number in the partition is equal to 1, then... X The convex combination of the extreme points of the partition, represented as a partition, transforms the optimization problem of the loss function into a mixed-integer linear programming problem by introducing slack variables: (22a); (22b); Among them, the problem coefficient , , and All The explicit function, the optimization variable vector s is given by X and If the composition is such that the optimal value of the objective function (22a) is written as: ; Using gradients Backpropagation is performed on the set of trainable parameters. θ Optimize.

2. The Lyapunov-based stable reinforcement learning control algorithm for uncertain systems based on worst-case scenarios as described in claim 1, characterized in that, In step S1, the process of modeling the system dynamics model and uncertainty boundaries is as follows: Consider a discrete-time robot system with the following dynamic model: (1); in, ∈ and ∈ These represent the robot in t The state and control at any given time, with saturation controlled by a lower bound. or upper boundary Give; The dynamic model is Assuming its approximate value The results were obtained through system identification, as shown below: (2); in, and These represent the multiplicative uncertainty and additive uncertainty between the actual dynamics and its approximation, respectively. This indicates the uncertainty of the centralized total model. Indicates controller In state The control provided at the location; However, system state It is estimated from some estimator, and therefore will be corrupted by noise, as shown below: (3); in and These represent the multiplicative and additive uncertainties in the state estimation, respectively. This represents the lumped uncertainty in state estimation; therefore, t The predicted state at time +1 is: (4); in, Indicates controller In state estimation The control provided at the location; For an approximate dynamic model ( , Output the state estimate of the system. Controlling saturation , and attraction domain D This makes the system subject to model uncertainty. and state uncertainty It stabilizes at the equilibrium point. Assumption 1: An approximate dynamic model is given in the form of a ReLU activated neural network. .

3. The Lyapunov-based stable reinforcement learning control algorithm for uncertain systems based on worst-case scenarios as described in claim 2, characterized in that, In step S4, the network parameterization establishment process is as follows: Using the HL-PWL method, an arbitrary piecewise linear function can be represented by the following expression. : (23); in, It is a coefficient vector. It is the stack of base functions, base functions The stack includes nested layers. i All nested functions; coefficient vector With basis functions Related; Robustness conditions under HL-PWL representation: The positive determination of the Lyapunov function, formula (7) is rewritten as: , (24); in and yes K Function, and with The Lyapunov function increases with the increase of . and K The functions were all rewritten in HL-PWL form, with the following expressions: (25); (26); in, For the infinite norm, ∈ and ∈ For the coefficient vector, and The number of corresponding coefficient vectors, upper bound, and lower bound. , and It is a vector function, including nested functions with nesting levels 0 and 1, in set D. The maximum value is It is easy to obtain; because It is K The function then has: (27); (28); Stability conditions (6) and (7) under equilibrium conditions are written as follows: (29); and (30); Partition and define vertices for a given attraction domain D, and then transform condition (30) into ,in ,matrix and By superposition V All vector , The obtained, of which V It is a set of D partitions; Similarly, the upper bound condition of equation (24) can be rewritten as follows: Then the upper boundary and the lower realm It is a K-function that satisfies equation (27), that is: (31); and functions that satisfy (28), namely: (32); in, and Given the grid size of the partition, when considering both the dynamic model and uncertainties in the state, stability conditions (20) and (14) produce another condition: (33); Further assume that both the dynamic system and the controller satisfy the following Lipshitz conditions: (34); (35); in and F These represent the upper bounds of the controller and the positive dynamic change, respectively; Then, item In state The Taylor expansion generated at this point is: (36); Through negation And using HL-PWL representation, equation (36) is written as: (37); in, Representation function right The derivative of is then the stability condition (33) can be written as: (38); Rewritten as the following quadratic inequality: (39); in It is a column vector. X is a real symmetric matrix; since X is a rank-1 matrix, it has a non-zero eigenvalue that equals The square of the absolute value, i.e. The remaining eigenvalues ​​are zero, let (40); but Rewritten as: (41); in It is an orthogonal identity matrix Therefore, formula (39) can be further rewritten as: (42); make Then the following condition is met: (43); Equation (43) shows that, Convert to standard quadratic form, There must exist a correct solution; However, when control is saturated at the lower and upper bounds, if for all , So regardless of the coefficient vector No matter how it changes, the entire quadratic form is less than zero. Therefore, the necessary condition for the existence of the correct solution, namely the stability condition (33), is written as: (44); The aforementioned necessary conditions require that the controller has sufficient capability to exceed the limits set by the controller. arrive By utilizing the uncertainty range and this necessary condition, infeasible states are found, and the attraction domain D is improved, eliminating states that the method cannot guarantee. Since system (2) is piecewise linear, the control at its boundary vertices is evaluated to obtain the vertex set. Find the reachable region of each vertex, control each vertex to obtain a cone, then construct the union of these cones, and obtain a new cone through approximation. ; Now, Each vertex in A cone in corresponding space These cones The union is a The solution space is considered, and the cones corresponding to the four vertices are considered. For a pair of aligned cones, the constraint imposed by formula (30) is considered. Only four cones are shown. The intersection of these four cones is approximately a linear cone. Without losing generality, assume the cone has N Intersection points These intersection points form N planes in the solution space, with normal vectors... The linear cone formed by these N intersection points is represented by calculating each plane: (45); (46); Then, the Lyapunov candidate function The solution for the coefficients is: (47); in, The coefficient vector is the HL-PWL representation of the Lyapunov function and its boundary. The attraction domain was initially verified using the solution space (47) before applying the proposed robust reinforcement learning RGRL to learn the controller and the coupled Lyapunov function.

Citation Information

Patent Citations

  • Device and method for controlling a system having uncentrity in the dynamics thereof

    CN116324635A

  • Reinforcement learning safety control method based on transfer learning

    CN118092185A