Multi-Lyapunov stability constraint-based security reinforcement learning control method and system for average residence time switching system

By constructing a multi-Lyapunov stability criterion for the switching system and training a Lyapunov commentator network, and by optimizing the policy network parameters in conjunction with the average dwell time constraint, the problem of insufficient stability of the switching system during mode switching is solved, and stable and reliable policy learning and control performance optimization are achieved.

CN121832235APending Publication Date: 2026-04-10HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HARBIN INST OF TECH
Filing Date
2026-02-12
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing reinforcement learning methods lack multi-Lyapunov stability criteria in switching systems, making it difficult to guarantee the overall stability and security of the system during mode switching, especially under the constraint of average dwell time, it is difficult to achieve stable and reliable policy learning.

Method used

A multi-Lyapunov stability criterion is constructed by building a Lyapunov function for each subsystem mode and combining it with the average dwell time constraint. A Lyapunov commentator network is trained, and multi-Lyapunov descent constraints and policy entropy constraints are introduced. The Lagrange multiplier method is used to optimize the policy network parameters to ensure the stability of policy updates and exploration capabilities.

Benefits of technology

It enhances the overall stability assurance capability of the switching system, realizes stable and reliable strategy learning in complex switching systems, ensures system stability during the learning and execution phases, and maintains stable closed-loop control performance in multiple application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention provides a security reinforcement learning control method and system for an average residence time switching system based on multiple Lyapunov stability constraints, and belongs to the field of intelligent control and reinforcement learning. The objective of the invention is to solve the problem of how to introduce a multi-Lyapunov stability criterion suitable for a switching system in reinforcement learning and realize stable and reliable strategy learning on the premise of satisfying an average residence time constraint. According to the method, a mode-dependent Lyapunov function is constructed for each subsystem mode, a multi-Lyapunov stability criterion meeting an average residence time condition is introduced, and a multi-Lyapunov descent constraint is explicitly applied in a reinforcement learning strategy updating process; therefore, the mean square index stability of the switching system in the learning stage and the execution stage is realized. The method does not need to depend on an accurate system dynamics model, can safely learn a mode-dependent control strategy only based on the state and control data collected in the system operation process, and has high robustness and engineering applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent control and reinforcement learning technology, and more specifically, to a safe reinforcement learning control method and system for a system based on average residence time switching of multiple Lyapunov stability constraints. Background Technology

[0002] Switching systems are widely used in engineering and technological fields such as robotics, spacecraft, and unmanned aerial vehicles. During operation, they typically need to switch between different system dynamic modes to adapt to changing environmental conditions or perform different control tasks. Due to significant differences in the dynamic characteristics between different subsystems, the mode switching process may introduce discontinuities in system state or control behavior, leading to potential stability risks. To ensure the safety and control performance of switching systems during actual operation, mean residence time constraints are often used to limit the number of switching cycles per unit time, thus maintaining overall system stability. However, in complex, large-scale, or high-dimensional dynamic systems, designing traditional controllers based on accurate models is often difficult. Deep reinforcement learning methods, through data-driven control policy learning, provide a feasible approach to solving the control problems of such highly complex switching systems.

[0003] Existing reinforcement learning methods primarily employ three technical approaches when handling safety constraints. The first is safety-based reinforcement learning methods based on Lagrange relaxation or primal-dual frameworks. These methods assign adaptive Lagrange multipliers to constraints, transforming them into penalty terms added to the objective function. During optimization, they satisfy safety constraints in the desired sense by "penalizing constraint violations." Theoretically, this provides a relatively systematic constraint handling mechanism. However, due to the lack of explicit constraints on the policy update magnitude, the policy update process is highly sensitive to the adjustment of Lagrange multipliers. When the multipliers are not adjusted properly, a single policy update may produce excessively large policy changes, leading to system state oscillations or instantaneous out-of-bounds behavior, making it difficult to guarantee safety throughout the learning process. The second approach is confidence region-based constraint policy optimization methods. These methods explicitly set confidence region constraints during policy optimization. They consider both cumulative reward and safety constraints during updates, keeping the new policy close to the old one, thus limiting the magnitude of policy updates to some extent and reducing the risk of constraint violations due to excessively rapid updates. However, when the initial policy itself is near an unsafe or poor operating point, even with confidence region constraints, the system may still explore near the unsafe region in the early stages of learning, failing to fundamentally guarantee the safety of the entire learning process. Finally, there are safe reinforcement learning methods based on Lyapunov functions. These methods constrain policy updates by designing a single Lyapunov function, ensuring the system remains within the safe region throughout the learning process. However, these methods are only applicable to single-modal systems; for switching systems that frequently switch between multiple subsystems, their stability cannot be guaranteed by a single Lyapunov function.

[0004] On the other hand, existing reinforcement learning methods lack mechanisms for handling the stability of switching systems. Due to the inconsistency of Lyapunov functions in different subsystems within a switching system, a single Lyapunov reduction condition usually cannot guarantee overall stability when switching frequently between subsystems. Furthermore, the stochasticity of reinforcement learning during the exploration phase further amplifies the stability risks caused by switching. When switching systems operate under the constraint of average dwell time, stability analysis generally relies on multi-Lyapunov theory. This method constructs an independent Lyapunov function for each mode and combines the amplification factor of mode switching with dwell time to comprehensively evaluate system stability. However, currently, there is no method to systematically embed multi-Lyapunov stability conditions into a reinforcement learning framework to achieve data-driven, safe, and mode-dependent policy learning. Summary of the Invention

[0005] The technical problem to be solved by this invention is:

[0006] This addresses the problem of how to introduce a multi-Lyapunov stability criterion applicable to switching systems into reinforcement learning, and how to achieve stable and reliable policy learning while satisfying the average dwell time constraint.

[0007] The technical solution adopted by the present invention to solve the above-mentioned technical problems is as follows:

[0008] This invention provides a safe reinforcement learning control method for average dwell time switching systems based on multiple Lyapunov stability constraints, comprising the following steps:

[0009] S100. Establish a switching system model that satisfies the average residence time constraint, define the system state variables, control inputs, subsystem mode sets, and stage cost functions, thereby forming a switching system model for reinforcement learning control.

[0010] S200. For each subsystem mode in the switching system model, construct the corresponding Lyapunov function, derive the intra-mode Lyapunov descent condition and the inter-mode Lyapunov comparison condition, and form the criterion for the overall stability of the switching system under the average residence time constraint.

[0011] S300. Based on the state, action and state transition data collected during system operation, combined with the stability criterion obtained in step S200, the training mode depends on the Lyapunov commentator network to estimate the Lyapunov function value in each subsystem mode, and calculate the Lyapunov decrease within the mode and the Lyapunov comparison quantity when switching modes, i.e., the Lyapunov stability measure.

[0012] S400: Using the Lyapunov stability metric obtained in step S300, the intra-mode Lyapunov descent constraint, inter-mode Lyapunov comparison constraint, and policy entropy constraint are introduced into the policy optimization process. The policy network parameters are updated using the Lagrange multiplier method, thereby maintaining the necessary policy exploration capability while ensuring stability.

[0013] Further, in step S100, let the system state and control input be respectively: system state , indicating at time The system's state vector has dimension 1. Control input , indicating at time The system's control input vector has a dimension of ;

[0014] The system's current operating mode is switched by an external signal. Confirmed, among which , indicating the system's position at time k. Subsystem mode number, This represents the total number of modes in the system.

[0015] S110. Constructing system state transitions and pattern dependencies.

[0016] The state transition relationships of a system under a given pattern are described by a state transition probability distribution. The state transition under the given conditions is determined by the probability transition moments. Description, which is defined as:

[0017]

[0018] Indicates the system is in mode The current state is as follows: The control input is When, transition to the next state. The probability; the probability transition moment in each subsystem mode The modes can take different forms, therefore the dynamic behavior of the system depends on the current mode.

[0019] S120. Construct average residence time constraints.

[0020] External switching signal of the system The average dwell time constraint must be met. The minimum average dwell time for each mode is mathematically expressed as follows:

[0021]

[0022] in, Indicates the time interval Inside, the system is in mode Number of switching times It is a constant representing the upper limit of the number of switching operations in the system under unconstrained conditions;

[0023] S130. Construct the control objective and stage cost function.

[0024] To describe the system's control objectives, a stage cost function is defined. Used to measure at time System status and control input The cost:

[0025]

[0026] in, The square norm represents the system state, and the deviation from the state represents the deviation from the state. This represents the square norm of the control input, and this represents the magnitude of the control input. It is a weighting factor that balances the deviation from the state and the magnitude of the control input;

[0027] The system's control objective is to minimize the cost function of this stage. To achieve effective control of the system state.

[0028] Furthermore, in step S200, when the system is in operation, the current state of the system is... and control input Based on this, for each subsystem pattern Construct the corresponding Lyapunov function The Lyapunov function is a non-negative function, and in the target state... Minimum value is obtained at:

[0029]

[0030] The Lyapunov function is defined as a quadratic form function:

[0031]

[0032] in, Subsystem pattern The corresponding positive definite matrix;

[0033] When the system operates in the same subsystem mode, apply intra-mode descent constraints to the Lyapunov function to ensure that the system state remains convergent in that mode:

[0034]

[0035] in, For pattern The corresponding convergence factor;

[0036] When the system switches modes, inter-mode comparison constraints are applied to the Lyapunov functions under different modes; assuming the system consists of modes... Switch to mode Then the following conditions must be met:

[0037]

[0038] in, For pattern Switch to mode The magnification factor at that time;

[0039] To constrain the cumulative effects caused by multiple mode switching during the system's operation, a time-averaged stability parameter is introduced. :

[0040]

[0041] This time-averaged stability measure is used to characterize the average growth level of the inter-mode amplification factor during the handover process, and together with the intra-mode convergence factor, constitutes the overall stability criterion; at the same time, the system's handover signal must satisfy the average dwell time constraint defined in step S100. .

[0042] Further, in step S300, a Lyapunov commentator network based on a neural network approximates the Lyapunov function in each subsystem mode; including,

[0043] S310, Construction of training objectives,

[0044] Assume the system is in mode The current state is as follows: The control input is Then at the next moment The system status is This forms a set of state transition data. This leads to the construction of training objectives, namely:

[0045]

[0046] in, It is a stage cost function, representing the system's current state. and control input The price paid It is a discount factor, representing the impact of future rewards on the current estimate;

[0047] The Lyapunov function is recursively estimated using the Bellman equation, and the training objective is to optimize the commentator network by minimizing the difference between the estimated and actual values.

[0048] S320, Construct nonnegativity constraints and stability guarantees.

[0049] During training, the Lyapunov commentator network is required to output estimates of the Lyapunov function. Satisfies the nonnegativity constraint, that is:

[0050]

[0051] S330. Optimize the Lyapunov critic network.

[0052] By minimizing the following loss function Conduct training:

[0053] in, These are parameters of the Lyapunov commentator network.

[0054] Furthermore, in step S400, the control strategy is parameterized using an actor-commentator structure; for each subsystem mode Construct the corresponding pattern-related policy network Its parameters are denoted as During system operation, when the switching signal satisfies... At that time, the policy network uses the current system state. Input, output control input Or control the probability distribution of the input And apply this control input to the system;

[0055] During the policy update phase, based on the Lyapunov commentator network's analysis of the Lyapunov function in step S300... and control input common

[0056] in, The stage cost function defined in step S100 at any time step. and This indicates the transition data for the next moment. Indicates the first The upper bound of Lyapunov descent allowed in each mode. This is the lower bound of the policy entropy. This represents the set of switching signals that satisfy the average dwell time constraint.

[0057] The primal-dual optimization method is employed, introducing Lagrange multipliers to incorporate constraints into the objective function; for each subsystem mode... Define the pattern-related constraint residual function for:

[0058] Constructing the Lagrangian objective function for policy optimization :

[0059] in, For the stability constraint Lagrange multipliers of the corresponding mode, The multiplier corresponding to the policy entropy;

[0060] During policy update, policy network parameters The objective function is updated using stochastic gradient descent, and its gradient update... The format is:

[0061] in, For pattern indicator functions;

[0062] The Lagrange multipliers are adaptively adjusted according to the primal-dual update rule, and the update form is as follows:

[0063]

[0064] in, and These are the learning rates corresponding to the Lagrange multipliers;

[0065] Through this update mechanism, when the sample data indicates that the stability constraint or entropy constraint is violated, the weight of the corresponding multiplier is increased, thereby strengthening the constraint effect in subsequent policy updates; when the constraint condition is met, the multiplier weight is gradually reduced to avoid overly conservative policy updates.

[0066] A safe reinforcement learning control system for a switching system based on multiple Lyapunov stability constraints and average residence time, the system having program modules corresponding to the above steps, and executing the steps in the above-described method for safe reinforcement learning control of a switching system based on multiple Lyapunov stability constraints and average residence time.

[0067] A computer-readable storage medium storing a computer program configured to, when invoked by a processor, implement steps of a safe reinforcement learning control method for a system with average dwell time switching based on multiple Lyapunov stability constraints.

[0068] Compared with the prior art, the beneficial effects of the present invention are:

[0069] ① By introducing a multi-Lyapunov stability analysis framework, the overall stability assurance capability of the switching system is improved. This invention addresses the significant differences in dynamics among the subsystems in a switching system by constructing Lyapunov functions for different modes and establishing a multi-Lyapunov stability criterion in conjunction with average dwell time constraints. This effectively suppresses state divergence caused by system discontinuities during mode switching, achieving mean square exponential stability of the switching system in a global sense, thus overcoming the limitations of traditional methods based on a single Lyapunov function in multimodal switching scenarios.

[0070] ② Ensuring system stability during both the learning and execution phases of reinforcement learning policy updates. By introducing both multi-Lyapunov descent constraints and policy entropy constraints into the primal-dual optimization framework, this invention imposes stability constraints on the magnitude of policy changes during policy updates, avoiding the unstable behavior that may occur in the learning phase of traditional reinforcement learning methods. This ensures that the control policy continuously optimizes during the training phase while always meeting stability and safety requirements.

[0071] This invention enhances the engineering applicability of the method in complex switching systems and various application scenarios. The method is applicable to continuous control systems with multiple operating modes and frequent switching behaviors, maintaining stable closed-loop control performance under different switching conditions and control tasks. Closed-loop control verification shows that this method can effectively balance system stability and control performance even in complex dynamic environments, demonstrating significant engineering application value.

[0072] In summary, this invention addresses the insufficient stability guarantees of existing secure reinforcement learning techniques in multi-mode switching systems by proposing a deep reinforcement learning control method based on multiple Lyapunov stability constraints and average dwell time conditions. This method not only effectively solves the problem of inconsistent stability guarantees during the learning and execution processes of switching systems, but also achieves synergistic optimization of stability, security, and control performance without relying on an accurate system model, demonstrating significant technological advancement and engineering application value. Attached Figure Description

[0073] Figure 1 This is a flowchart of a safety reinforcement learning control method for a system with average dwell time switching based on multiple Lyapunov stability constraints, as described in an embodiment of the present invention.

[0074] Figure 2 This is a framework diagram of a safety reinforcement learning control method for a switching system based on multiple Lyapunov stability constraints in an embodiment of the present invention.

[0075] Figure 3 This is a schematic diagram illustrating the control and stability constraint relationship of a switching system satisfying the average dwell time constraint under different subsystem modes in an embodiment of the present invention. Detailed Implementation

[0076] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0077] Specific Implementation Plan 1: Combining Figure 1 and Figure 2 As shown, this invention provides a secure reinforcement learning control method for a switching system based on multiple Lyapunov stability constraints and average residence time, comprising the following steps:

[0078] S100. Establish a switching system model that satisfies the average residence time constraint, define the system state variables, control inputs, subsystem mode sets, and stage cost functions, thereby forming a switching system model for reinforcement learning control.

[0079] Specifically, including,

[0080] Consider a class of discrete-time switched systems that satisfy the average residence time constraint; the system consists of multiple subsystem modes, each corresponding to a different system dynamic characteristic; the system's operating mode switches between different modes according to an external switching signal; specifically, let the system state and control input be: system state , indicating at time The system's state vector has dimension 1. Control input , indicating at time The system's control input vector has a dimension of ;

[0081] The system's current operating mode is switched by an external signal. Confirmed, among which , indicating the system's position at time k. Subsystem mode number, This represents the total number of modes in the system.

[0082] S110. Constructing system state transitions and pattern dependencies.

[0083] The state transition relationships of a system under a given pattern are described by a state transition probability distribution. The state transition under the given conditions is determined by the probability transition moments. Description, which is defined as:

[0084]

[0085] Indicates the system is in mode The current state is as follows: The control input is When, transition to the next state. The probability; the probability transition moment in each subsystem mode The modes can take different forms, therefore the dynamic behavior of the system depends on the current mode.

[0086] S120. Construct average residence time constraints.

[0087] External switching signal of the system The average dwell time constraint must be met. The minimum average residence time for each mode, i.e., the system's residence time in mode. The average time spent in the lower position must be greater than or equal to The mathematical expression for this constraint is as follows:

[0088]

[0089] in, Indicates the time interval Inside, the system is in mode Number of switching times It is a constant representing the upper limit of the number of switching operations in the system under unconstrained conditions;

[0090] This average dwell time constraint ensures that the switching frequency between modes is not too high, thereby avoiding system instability caused by rapid switching.

[0091] S130. Construct the control objective and stage cost function.

[0092] To describe the system's control objectives, a stage cost function is defined. Used to measure at time System status and control input The cost of each stage; the stage cost function can be designed according to the specific needs of the system. For example, in a control task, the following cost function might be used:

[0093]

[0094] in, The square norm represents the system state, and the deviation from the state represents the deviation from the state. This represents the square norm of the control input, and this represents the magnitude of the control input. It is a weighting factor that balances the deviation from the state and the magnitude of the control input;

[0095] The system's control objective is to minimize the cost function of this stage. To achieve effective control of the system state, making the system state as close as possible to the target state, while keeping the control input within an appropriate range;

[0096] S200. For each subsystem mode in the switching system model, construct the corresponding Lyapunov function and derive the intra-mode Lyapunov descent condition and inter-mode Lyapunov comparison condition. Under the average residence time constraint, form the criterion for the overall stability of the switching system, and provide the stability constraint basis for the policy learning in steps S300 and S400.

[0097] Specifically, including,

[0098] Based on the discrete-time switching system model that satisfies the average residence time constraint constructed in step S100, and the current state of the system... With control input Based on this, a mode-dependent Lyapunov function is constructed for each subsystem mode, and multi-Lyapunov stability constraints are established at two levels: intra-mode evolution and inter-mode switching. The obtained stability constraints serve as the basis for constructing the Lyapunov commentator network training objective and stability metric in step S300, and are used for stability constraint calculation in the subsequent policy optimization process.

[0099] During system operation, based on the current system state and control input Based on this, for each subsystem pattern Construct the corresponding Lyapunov function The Lyapunov function is a non-negative function, and in the target state... Minimum value is obtained at:

[0100]

[0101] The Lyapunov function is defined as a quadratic form function:

[0102]

[0103] in, Subsystem pattern The corresponding positive definite matrix is ​​used to characterize the energy level of the system state in this mode;

[0104] When the system operates under the same subsystem mode, an intra-mode descent constraint is applied to the Lyapunov function to ensure that the system state remains convergent under that mode, specifically satisfying:

[0105]

[0106] in, For pattern The corresponding convergence factor is used to constrain the decay rate of the Lyapunov function within the mode;

[0107] When the system switches modes, inter-mode comparison constraints are applied to the Lyapunov functions under different modes; assuming the system consists of modes... Switch to mode Then the following conditions must be met:

[0108]

[0109] in, For pattern Switch to mode The amplification factor at that time is used to limit the change of the upper bound of the Lyapunov function at the instant of switching;

[0110] To constrain the cumulative effects caused by multiple mode switching during the system's operation, a time-averaged stability parameter is introduced. Its definition is:

[0111]

[0112] This time-averaged stability measure is used to characterize the average growth level of the inter-mode amplification factor during the switching process, and together with the intra-mode convergence factor, it constitutes the overall stability criterion.

[0113] At the same time, the system's switching signal must meet the average dwell time constraint defined in step S100. ;

[0114] By combining intra-mode descent constraints, inter-mode comparison constraints, and time-averaged stability and average dwell time constraints, the system maintains overall stability while meeting switching conditions.

[0115] S300. Based on the state, action and state transition data collected during system operation, and combined with the stability conditions obtained in step S200, a mode-dependent Lyapunov commentator network is trained to estimate the Lyapunov function values ​​in each subsystem mode, and to calculate the Lyapunov decrease within the mode and the Lyapunov comparison quantity during mode switching.

[0116] Specifically, including,

[0117] Based on the Lyapunov function and its stability constraints constructed for each subsystem mode in step S200, a Lyapunov commentator network is further constructed. To avoid relying on an accurate system dynamics model, a neural network-based Lyapunov commentator network is used to approximate the Lyapunov function for each subsystem mode. This method can learn and estimate the Lyapunov function in a data-driven manner without an accurate system model, thereby providing stability guarantees during the optimization process.

[0118] S310, Construction of training objectives,

[0119] During system operation, the Lyapunov commentator network is trained through interaction with the environment; specifically, the system collects system state data through interaction with the environment. Control input and state transition data And this data was used to train the Lyapunov commentator network; assuming the system in the pattern The current state is as follows: The control input is Then at the next moment The system status is This forms a set of state transition data. This leads to the construction of training objectives, namely:

[0120]

[0121] in, It is a stage cost function, representing the system's current state. and control input The price paid It is a discount factor, representing the impact of future rewards on the current estimate;

[0122] The Lyapunov function is recursively estimated using the Bellman equation, and the training objective is to optimize the commentator network by minimizing the difference between the estimated and actual values.

[0123] S320, Construct nonnegativity constraints and stability guarantees.

[0124] During training, the Lyapunov commentator network is required to output estimates of the Lyapunov function. Satisfies the nonnegativity constraint, that is:

[0125]

[0126] This nonnegativity constraint ensures that the Lyapunov function will not have a negative value during each training iteration, thus maintaining stability. By introducing the nonnegativity constraint, the Lyapunov commentator network can accurately reflect the stability of the system during the estimation process and effectively constrain the stability change trend of the system state under different modes.

[0127] S330. Optimize the Lyapunov critic network.

[0128] During training, the network parameters are optimized by minimizing the difference between the estimated and actual Lyapunov function values ​​of the network output; specifically, this is achieved by minimizing the following loss function. Conduct training:

[0129]

[0130] in, These are the parameters of the Lyapunov commentator network. The network parameters are updated using gradient descent or other optimization algorithms, so that the estimated value of the Lyapunov function output by the network gradually approaches the true value.

[0131] S400. Using the Lyapunov stability metric obtained in step S300, the intra-mode Lyapunov descent constraint, inter-mode Lyapunov comparison constraint, and policy entropy constraint are introduced into the policy optimization process. The policy network parameters are updated using the Lagrange multiplier method, thereby maintaining the necessary policy exploration capability while ensuring stability.

[0132] Specifically, including,

[0133] Based on the mode-dependent Lyapunov function constructed in step S200, its intra-mode descent constraints, and inter-mode comparison constraints, as well as the Lyapunov function estimate obtained through the Lyapunov critic network in step S300, the control strategy is optimized under the deep reinforcement learning framework, thereby achieving continuous improvement in control performance while ensuring the stability of the switching system.

[0134] The control strategy is parametrically represented using an actor-commentator structure; for each subsystem mode Construct the corresponding pattern-related policy network Its parameters are denoted as During system operation, when the switching signal satisfies... At that time, the policy network uses the current system state. Input, output control input Or control the probability distribution of the input And apply this control input to the system;

[0135] During the policy update phase, based on the Lyapunov commentator network's analysis of the Lyapunov function in step S300... and control input The common estimation results formulate the policy learning problem as a constrained optimization problem in the following form:

[0136]

[0137] in, The stage cost function defined in step S100 at any time step. and This indicates the transition data for the next moment. Indicates the first The upper bound of Lyapunov descent allowed in each mode. This is the lower bound of the policy entropy. This represents the set of switching signals that satisfy the average dwell time constraint.

[0138] To solve the above constrained optimization problem, a primal-dual optimization method is adopted, introducing Lagrange multipliers to incorporate the constraints into the objective function; for each subsystem mode... Define the pattern-related constraint residual function for:

[0139]

[0140] in, and The values ​​are all provided by the Lyapunov commentator network trained in step S300;

[0141] Based on this, a Lagrangian objective function for policy optimization is constructed. Its form is:

[0142]

[0143] in, For the stability constraint Lagrange multipliers of the corresponding mode, This is the multiplier corresponding to the policy entropy, used to balance the requirements of exploration and stability;

[0144] During policy update, policy network parameters Based on the above objective function, the algorithm uses stochastic gradient descent to update the gradient. The format is:

[0145]

[0146] in, This is a mode indicator function used to ensure that the corresponding policy network parameters are updated only when the corresponding mode is activated;

[0147] Meanwhile, the Lagrange multipliers are adaptively adjusted according to the primal-dual update rule, and the update form is as follows:

[0148]

[0149] in, and These are the learning rates corresponding to the Lagrange multipliers;

[0150] Through this update mechanism, when the sample data indicates that the stability constraint or entropy constraint is violated, the weight of the corresponding multiplier is increased, thereby strengthening the constraint effect in subsequent policy updates; when the constraint condition is met, the weight of the multiplier is gradually reduced to avoid the policy update being overly conservative.

[0151] Through the above strategy optimization process based on multiple Lyapunov stability constraints, the strategy network achieves a balance between control performance optimization and stability assurance under the premise of satisfying the average residence time switching condition, providing a stable and optimal control strategy for the closed-loop control in the subsequent step S500.

[0152] S500. Under the condition of fixed strategy network parameters, implement closed-loop control of the system according to the switching sequence that satisfies the average residence time constraint, apply the control strategy obtained in step S400 to control the switching system, and verify the closed-loop control performance and overall stability of the system under different mode switching conditions.

[0153] Specifically, including,

[0154] Based on the optimal control strategy obtained under the Lyapunov stability constraints in step S400, closed-loop control is implemented on the switching system under the action of the switching signal that satisfies the average residence time constraint, thereby verifying the stability and control effect of the proposed control method in actual operation.

[0155] During closed-loop control, the system generates control inputs in real time based on the current state and operating mode, and applies them to the system; by analyzing the system state sequence during closed-loop operation... Continuous observation verifies that the system remains bounded and gradually converges to the target state under different mode switching conditions, and that the system maintains stable operation under different mode switching conditions. Simultaneously, during closed-loop operation, the stage cost function defined in step S100 is used... The cumulative performance index of the system is calculated, and its form is as follows:

[0156]

[0157] Combination Figure 3 As shown, the switching system environment based on the robot ant model has each subsystem corresponding to the robot's dynamic model under different physical operating conditions. The switching mechanism is used to simulate real environmental changes that may occur during the robot's movement, including changes in ground conditions, actuator performance, and load distribution. In this embodiment, the ground friction coefficient varies within the range [0.4, 1.5], the motor gear transmission ratio fluctuates within ±30% of the nominal value, and the body mass distribution is adjusted within ±20%, thus forming multiple different operating modes. The switching between subsystems is controlled by an external switching signal and meets the average dwell time constraint to ensure that each operating mode lasts for a preset time before switching. By observing the changes in the cumulative performance index over time, it can be found that when the optimized strategy obtained in step S400 is used for control, the cumulative performance index of the system can converge in a short time and eventually reach a low stable value, indicating that the generated control strategy achieves effective optimization of control performance while ensuring stable system operation.

[0158] Specific Implementation Method Two: The present invention provides a safety reinforcement learning control system for a system with average dwell time switching based on multiple Lyapunov stability constraints. This system has program modules corresponding to the above steps, and executes the steps in the above-mentioned method for safety reinforcement learning control of a system with average dwell time switching based on multiple Lyapunov stability constraints during runtime.

[0159] The other combinations and connections in this implementation scheme are the same as in Specific Implementation Scheme 1.

[0160] Specific Implementation Method 3: The present invention provides a computer-readable storage medium storing a computer program configured to, when called by a processor, implement the steps of a safe reinforcement learning control method for a system with average dwell time switching based on multiple Lyapunov stability constraints.

[0161] The other combinations and connections in this implementation scheme are the same as in Specific Implementation Scheme 1.

[0162] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.

Claims

1. A method for safety reinforcement learning control of an average dwell time switching system based on multiple Lyapunov stability constraints, characterized in that, Includes the following steps: S100. Establish a switching system model that satisfies the average residence time constraint, define the system state variables, control inputs, subsystem mode sets, and stage cost functions, thereby forming a switching system model for reinforcement learning control. S200. For each subsystem mode in the switching system model, construct the corresponding Lyapunov function, derive the intra-mode Lyapunov descent condition and the inter-mode Lyapunov comparison condition, and form the criterion for the overall stability of the switching system under the average residence time constraint. S300. Based on the state, action and state transition data collected during system operation, combined with the stability criterion obtained in step S200, the training mode depends on the Lyapunov commentator network to estimate the Lyapunov function value in each subsystem mode, and calculate the Lyapunov decrease within the mode and the Lyapunov comparison quantity when switching modes, i.e., the Lyapunov stability measure. S400: Using the Lyapunov stability metric obtained in step S300, the intra-mode Lyapunov descent constraint, inter-mode Lyapunov comparison constraint, and policy entropy constraint are introduced into the policy optimization process. The policy network parameters are updated using the Lagrange multiplier method, thereby maintaining the necessary policy exploration capability while ensuring stability.

2. The method for secure reinforcement learning control of a switching system based on multiple Lyapunov stability constraints and average residence time as described in claim 1, characterized in that: In step S100, let the state and control input of the system be: the state of the system , representing the state vector of the system at time , with dimension ; the control input , representing the control input vector of the system at time , with dimension ; The system is currently running in a mode determined by an external switching signal is determined, wherein represents the number of the subsystem mode in which the system is at time k, is the total number of modes of the system; S110. Constructing system state transitions and pattern dependencies. The state transition relationship of the system in a given mode is described by a state transition probability distribution, and the state transition of the system in the mode is described by a probability transition matrix , which is defined as: The system represents the system in mode The current state is The control input is The probability of transition to the next state The probability transition matrix has a different form in each subsystem mode Thus the dynamic behavior of the system depends on the current mode S120. Construct average residence time constraints. External switching signal of a system The average residence time constraint must be satisfied, set For each mode, the minimum average residence time, the mathematical expression of this constraint is as follows: in, Indicates the time interval Inside, the system is in mode Number of switching times It is a constant representing the upper limit of the number of switching operations in the system under unconstrained conditions; S130. Construct the control objective and stage cost function. To describe the system's control objectives, a stage cost function is defined. Used to measure at time System status and control input The cost: wherein represents the squared norm of the system state, represents the deviation of the state; represents the squared norm of the control input, represents the magnitude of the control input; is a weight factor that trades off state deviation and control input size; The control objective of the system is to achieve efficient control of the system state by minimizing the stage cost function Jk(xk, uk) 3. The method for secure reinforcement learning control of a switching system based on multiple Lyapunov stability constraints and average residence time as described in claim 2, characterized in that: In step S200, when the system is in operation, the current system state is used. and control input Based on this, for each subsystem pattern Construct the corresponding Lyapunov function The Lyapunov function is a non-negative function, and in the target state... Minimum value is obtained at: The Lyapunov function is defined as a quadratic form function: wherein subsystem mode corresponding positive definite matrix When the system operates in the same subsystem mode, apply intra-mode descent constraints to the Lyapunov function to ensure that the system state remains convergent in that mode: wherein is the mode corresponding convergence factor; When the system switches from one mode to another, the inter-mode comparison constraint is imposed on the Lyapunov function in different modes; let the system switch from mode to mode then it is required to satisfy: wherein, is the mode switches to the mode amplification factor at the time In order to constrain the cumulative effect caused by multiple mode switching in the whole operation process of the system, a time-averaged stability quantity is introduced : The time average stability quantity is used to characterize the average growth level of the inter-mode amplification factor in the switching process, and together with the intra-mode convergence factor constitutes an overall stability criterion; at the same time, the switching signal of the system needs to satisfy the average residence time constraint defined in step S100 .

4. The safety reinforcement learning control method for average dwell time switching system based on multiple Lyapunov stability constraints according to claim 3, characterized in that: In step S300, the Lyapunov commentator network based on the neural network approximates the Lyapunov function in each subsystem mode; include, S310, Construction of training objectives, Assume that the system is in mode , the current state is , the control input is , then at the next time , the system state is , thus forming a set of state transition data , and further constructing the training target, that is: wherein, is a stage cost function representing the cost of the system in the current state and control input , is a discount factor representing the influence of future rewards on the current estimate; The Lyapunov function is recursively estimated using the Bellman equation, and the training objective is to optimize the commentator network by minimizing the difference between the estimated and actual values. S320, Construct nonnegativity constraints and stability guarantees. During training, the Lyapunov commentator network is required to output estimates of the Lyapunov function. Satisfies the nonnegativity constraint, that is: S330. Optimize the Lyapunov critic network. By minimizing the following loss function Conduct training: in, These are parameters of the Lyapunov commentator network.

5. The method for secure reinforcement learning control of a switching system based on multiple Lyapunov stability constraints and average residence time as described in claim 4, characterized in that: In step S400, the control strategy is parameterized using an actor-commentator structure; for each subsystem mode Construct the corresponding pattern-related policy network Its parameters are denoted as ; During system operation, when the switching signal satisfies At that time, the policy network uses the current system state. Input, output control input Or control the probability distribution of the input And apply this control input to the system; During the policy update phase, based on the Lyapunov commentator network's analysis of the Lyapunov function in step S300... and control input The common estimation results formulate the policy learning problem as a constrained optimization problem in the following form: in, The stage cost function defined in step S100 at any time step. and This indicates the transition data for the next moment. Indicates the first The upper bound of Lyapunov descent allowed in each mode. This is the lower bound of the policy entropy. This represents the set of switching signals that satisfy the average dwell time constraint. The primal-dual optimization method is employed, introducing Lagrange multipliers to incorporate constraints into the objective function; for each subsystem mode... Define the pattern-related constraint residual function for: Constructing the Lagrangian objective function for policy optimization : in, For the stability constraint Lagrange multipliers of the corresponding mode, The multiplier corresponding to the policy entropy; During policy update, policy network parameters The objective function is updated using stochastic gradient descent, and its gradient update... The format is: in, For pattern indicator functions; The Lagrange multipliers are adaptively adjusted according to the primal-dual update rule, and the update form is as follows: in, and These are the learning rates corresponding to the Lagrange multipliers; Through this update mechanism, when the sample data indicates that the stability constraint or entropy constraint is violated, the weight of the corresponding multiplier is increased, thereby strengthening the constraint effect in subsequent policy updates; when the constraint condition is met, the multiplier weight is gradually reduced to avoid overly conservative policy updates.

6. A safety reinforcement learning control system for a switching system based on average dwell time using multiple Lyapunov stability constraints, characterized in that: The system has a program module corresponding to the steps of any one of the claims 1-5 above, and executes the steps in the above-described reinforcement learning control method for a switching system with average residence time based on multiple Lyapunov stability constraints when it is run.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program configured to, when invoked by a processor, implement the steps of any one of claims 1-5 of the secure reinforcement learning control method for a system with average dwell time switching based on multiple Lyapunov stability constraints.