Human-in-the-loop elastic control method for multi-agent systems based on event-triggered integral reinforcement learning

By adopting event-triggered integral reinforcement learning method in multi-agent systems, the problems of difficulty in decision-making and high energy consumption in the prior art multi-agent systems in emergency situations are solved, and more efficient, safe and robust control effects are achieved.

CN118915466BActive Publication Date: 2025-05-13UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411077224.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-07
Publication Date
2025-05-13
Estimated Expiration
2044-08-07

AI Technical Summary

Technical Problem

The existing multi-agent system humans have difficulty making decisions in emergency situations in loop control methods, and traditional feedback linearization and adaptive dynamic programming methods cannot effectively reduce energy consumption and solve the problem of false data injection attacks.

Method used

Using an integrated reinforcement learning method based on event triggering, a novel human-in-loop elastic control method for multi-agent system is designed by constructing the event-triggered Hamilton-Jacobby-Bellmann equation, and combining evaluation neural networks to estimate the cost function.

Benefits of technology

This method can improve the system's decision-making efficiency in emergency situations, reduce the frequency of information exchange between agents, reduce the cost of computing and communication, and improve the system's security and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118915466B_ABST
    Figure CN118915466B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-agent system human-in-the-loop elastic control method based on event-triggered integral reinforcement learning, which belongs to the field of artificial intelligence collaborative control technology. The method innovatively adopts an integral reinforcement learning method based on an event-triggered mechanism to reduce the frequency of information exchange between agents, reduce the burden of agents, and solve the optimal control problem of multi-agent systems under false input injection attacks; in addition, by integrating human intelligence and decision-making, when a human operator sends a command signal to a non-autonomous leader agent, the system security is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence collaborative control, and specifically relates to a multi-agent system human-in-the-loop elastic control method based on event-triggered integral reinforcement learning. Background Art

[0002] As an important branch of cooperative control, cluster cooperative control has received extensive attention in the past decade due to its representativeness and fundamentality in practical multi-task applications. In most existing cluster cooperative control studies, a traditional and common setting is that each agent (including the leader) in the multi-agent system is autonomous. This setting benefits from the development of artificial intelligence and liberates human participation. Despite providing such indisputable advantages, a fully autonomous setting is ideal to some extent because it is difficult to make decisions in emergency situations, which may lead to serious consequences. Considering the possible emergency situations, a human operator is usually set up to assist the autonomous system in completing the task. Therefore, human-in-the-loop control has been studied. However, most of the current human-in-the-loop control methods ignore the control performance with the goal of minimum energy consumption.

[0003] How to achieve control performance with minimum energy consumption as the goal becomes the key. In the past few years, optimal tracking control has attracted more and more attention. The methods of optimal tracking control have evolved from traditional feedback linearization and object inversion to the current adaptive dynamic programming and integrated reinforcement learning. However, traditional feedback linearization methods and adaptive dynamic programming are based on explicit formulas and require a complete understanding of the system dynamics. At the same time, these methods do not consider the steady-state part of the performance cost. Similar to adaptive dynamic programming, integrated reinforcement learning is one of the most effective learning methods in the field of artificial intelligence. Integrated reinforcement learning is an improvement of reinforcement learning. It uses policy iteration technology. It is an iterative method that solves the Hamilton-Jacobi-Bellman equation by alternately constructing performance indicators and control strategies, and finally converges to the optimal control solution. In the process of integrated reinforcement learning, performance indicators and control strategies are solved simultaneously. However, the control strategies obtained by existing integrated reinforcement learning are all based on time period sampling; and the time periodic sampling mechanism will lead to waste of communication resources and increase of computing cost.

[0004] In addition to control performance, communication security is also a factor that needs to be considered in control schemes for multi-agent systems. The use of communication networks in multi-agent systems gives multi-agents significant advantages in terms of efficiency, design cost, and simplicity. However, this benefit comes at the expense of increased vulnerability to a range of cyber-physical attacks, such as false data injection attacks. Adversaries can transmit false data to initiate attacks by accessing communication channels and injecting faults into the information transmitted from the agents to the control center. Generally speaking, attackers have limited energy and all information transmission cannot be interrupted, so attackers often use sparse attacks. The core of existing solutions to sparse attacks is to design safe state estimation methods that use system measurements and model information to infer the safe state of the system. Safe state estimation methods mainly follow two design ideas: one method is to use multiple sensors to measure the same output and obtain reliable output data through data filtering; the other design idea is to use multiple sensors to measure different outputs, obtain multiple groups of state estimates based on different output combinations, and obtain a reliable estimate from these estimates. However, how to design safe control schemes against false data injection attacks in the field of multi-agent systems has not been fully explored, especially in solving key problems such as optimization and limited resources. On the one hand, while achieving predefined metrics is desirable, minimizing costs is also crucial. On the other hand, reducing computation and communication time to a reasonable level is necessary, given the limited resources available to everyone.

[0005] Therefore, how to achieve optimal human-in-the-loop control based on integral reinforcement learning has become a research focus. Summary of the invention

[0006] In view of the problems existing in the background technology, the purpose of the present invention is to provide a multi-agent system human-in-the-loop elastic control method based on event-triggered integral reinforcement learning. This method innovatively adopts an integral reinforcement learning method based on an event-triggered mechanism to reduce the frequency of information exchange between agents, reduce the burden of agents, and solve the optimal control problem of multi-agent systems under false input injection attacks; in addition, by integrating human intelligence and decision-making, when a human operator sends a command signal to a non-autonomous leader agent, the system security is greatly improved.

[0007] To achieve the above object, the technical solution of the present invention is as follows:

[0008] A human-in-the-loop elastic control method for a multi-agent system based on event-triggered integral reinforcement learning includes the following steps:

[0009] S1. Based on the nonlinear multi-agent system, establish the dynamic model of the i-th follower;

[0010] S2. Give the leader's dynamic model and define the local neighborhood consensus error;

[0011] S3, filtering the unattacked data through the median operator Med[·], obtaining a security preselector based on the unattacked data, and designing a state observer;

[0012] S4, construct the optimal performance index function and establish the Hamilton-Jacobi-Bellman equation;

[0013] S5. Introduce event trigger mechanism and construct event-triggered Hamilton-Jacobi-Bellman equation;

[0014] S6, using the integral reinforcement learning algorithm to solve the event-triggered Hamilton-Jacobi-Bellman equation;

[0015] S7. Use the evaluation neural network to estimate the cost function and obtain the required approximate event-triggered optimal controller.

[0016] Furthermore, in step S1, the nonlinear multi-agent system includes a leader and a plurality of followers, and the dynamic model of the i-th follower is specifically as follows:

[0017]

[0018] in, represents the derivative, x i (t) is the state information of the i-th agent at time t, u i (t) is the control input of the ith agent at time t, f i (x i ) is the known internal function of the ith follower, g i (x i ) is the known input matrix function of the ith follower, y i (t) is the measured signal obtained after transmission, The malicious attack signal injected by the attacker, P i is the system measurement matrix; x i For x i (t) is the abbreviation of x i =[x i,1 ,…,x i,n ] T ∈R n ,u i (t)∈R m , f i (x i )∈R n , g i (x i )∈R n×m , R refers to the real number domain, n and m refer to the dimensions of the matrix, and pi is the number of sensor outputs.

[0019] Furthermore, in step S2, the dynamic model of the leader is:

[0020]

[0021] Where x0(t)∈R n represents the state information of the leader at time t, y0(t) represents the output of the leader at time t, and u0(t) is the control input, which is an unknown bounded variable;

[0022] Define the local neighborhood consistency error δ i (t) is:

[0023]

[0024] Among them, b i represents the containment gain, a ij represents the connection weight between the ith agent and the jth agent, x j (t) represents the state vector of the jth agent at the tth time, N i represents the set of neighbor agents of the ith agent.

[0025] Furthermore, the specific form of the control input u0(t) is,

[0026]

[0027] t1 is the first time threshold, and t2 is the second time threshold.

[0028] Furthermore, in step S3, the control signal y i The elements in (t) are arranged in ascending order to obtain a new vector α i =[α i,1 ,…,α i,pi ] T , α i,1 ≤…≤α i,pi ,

[0029] Then the median operator Med[·] is,

[0030]

[0031] Design the following safety preselector x for the i-th agent: i,1,p (t):

[0032] x i,1,p (t) = Med[y i (t)](4)

[0033] Using the unattacked output data x i,1 (t), the first state of the system, Represents the estimation, for the i-th follower, the following state observer is designed To estimate the state of the system:

[0034]

[0035] in, and K0 represent the state estimation vector and the given positive gain respectively.

[0036] Furthermore, the specific process of step S4 is as follows:

[0037] Construction and local neighborhood consistency error δ i and control input u i Related performance index function V i (δ i ), the specific form is:

[0038]

[0039] Among them, U i (·) represents the utility function, Q ii , R ii and R ij All are constant matrices;

[0040] Based on the local neighborhood consistency error, define V i (δ i )’s Hamiltonian:

[0041]

[0042] l ij is an element in the Laplace matrix, L = DA = [l ij ]∈R N×N , D is the degree matrix, D = diag(d1,…,d N ), represents the diagonal elements of the degree matrix, A is the adjacency matrix, and represents the communication topology of the directed graph. A=[a ij ]∈R N×N , a ij is an element in the adjacency matrix A, and l ij =-a ij ;

[0043] represents the drift dynamics model, b ij =0;

[0044] Based on Bellman's optimality principle, the optimal cost function V i* (δ i ) is in the form of:

[0045]

[0046] And satisfies the following Hamiltonian function:

[0047]

[0048] in, is the optimal cost function V i * (δ i ) About δ i The partial derivative of

[0049] The optimal control strategy The expression is:

[0050]

[0051] Will Substituting into the Hamiltonian function (6), we can get the Hamilton-Jacobi-Bellman equation as follows:

[0052]

[0053] Furthermore, the constant matrix Q ii , R ii and R ij Both are greater than 0.

[0054] Furthermore, the specific process of step S5 is:

[0055] Let the μth triggering time be t μ , and satisfies t μ <t μ+1 , where μ∈N, N is a set of natural numbers, then the sequence of triggering times can be obtained In t μ The state of the sampling at the moment is expressed as

[0056] At two consecutive triggering times t μ and t μ+1 Between, that is, t∈[t μ ,t μ+1 ), there are usually two states and The error between them is defined as the error function, recorded as Its form is:

[0057]

[0058] According to the local neighborhood consistent error formula, define the local neighborhood consistent error based on event triggering for,

[0059]

[0060] The event-triggered optimal control strategy is:

[0061]

[0062] in, Indicated in In the case of i * δ i The partial derivative of

[0063] Therefore, the event-based Hamilton-Jacobi-Bellman equation is,

[0064]

[0065] Furthermore, the specific process of step S6 is:

[0066] S6.1. Selecting the initial permissible events to trigger the optimal control strategy;

[0067] S6.2. For each agent i, calculate the cost function V under the optimal control strategy triggered by the current event i k (δ i (t)),

[0068]

[0069] Where t∈t μ , V i k (0) = 0, k represents the current number of iterations;

[0070] S6.3. According to the cost function event triggering optimal control strategy, the following event-triggered control strategy is obtained:

[0071]

[0072] make Return to step S6.2 until V i k →V i * ,

[0073] Furthermore, the specific process of step S7 is:

[0074] Establish an evaluation neural network, and apply the approximation properties of the evaluation neural network to any continuous function. The cost function V i * (δ i ) can be written as,

[0075] V i * (δ i )=W i T θ(δ i )+ε i (δ i )(11)

[0076] in, is the ideal weight vector, is the activation function, ε i (δ i )∈R is the approximation error, N c ∈Z + is the number of neurons;

[0077] Then V i * (δ i ) for the local neighborhood consistency error δ i The partial derivative of

[0078]

[0079] in, is the activation function θ i δ i The partial derivative of is the approximate error function ε i δ i The partial derivative of

[0080] Substituting formula (12) into formula (9), we get the following optimal control strategy based on event triggering:

[0081]

[0082] in,

[0083] Since the ideal weight vector W i is unknown, and the event-triggered optimal control strategy (13) cannot be directly obtained; therefore, an evaluation neural network is used to estimate the cost function:

[0084]

[0085] in, is an estimate of the ideal weight;

[0086] Similarly, for The corresponding partial derivatives can be obtained,

[0087]

[0088] Therefore, the approximate event-triggered optimal controller is,

[0089]

[0090] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0091] 1. The control method of the present invention introduces an integral reinforcement learning method to construct the integral Bellman equation. The integral reinforcement learning method allows the requirements on the system drift dynamics to be relaxed in the controller design without the need for system identification. In addition, in order to reduce the computational and communication costs, an event-triggered control condition is designed. The integral reinforcement learning algorithm is combined with the event-triggered control framework to address the challenges of multi-agent systems in a novel way, making the learning process more flexible.

[0092] 2. Compared with the existing multi-agent adaptive dynamic programming method, the method of the present invention focuses on the problem of human-in-the-loop resilience control, and the output trajectory of the leader is given as a reference signal. If the control signal of the leader is given by a human operator, this is conducive to achieving a consistent trajectory, with better utility value and safety performance. At the same time, by constructing a pre-selector, safe output data can be stripped from a set of output measurements and used to construct a state observer, thereby successfully avoiding the impact of sparse attacks. Therefore, the designed human-in-the-loop resilience control scheme can operate normally in dangerous environments and has higher robustness than other schemes when facing network attacks. BRIEF DESCRIPTION OF THE DRAWINGS

[0093] Figure 1 Flow chart of the control method of the present invention.

[0094] Figure 2 The state x of the follower of the present invention i,1 With the leader's state x 01 Graph of changes over time.

[0095] Figure 3 The state x of the follower of the present invention i,2 With the leader's state x 02 Graph of changes over time.

[0096] Figure 4 is the observation error e of the five agents of the present invention i,1 Graph.

[0097] Figure 5is the observation error e of the five agents of the present invention i,2 Graph.

[0098] Figure 6 This is the control input trajectory diagram of the five agents of the present invention.

[0099] Figure 7 This is the control input trajectory diagram of the five agents of the present invention.

[0100] Figure 8 Schematic diagram of the trigger sampling intervals of the five intelligent agents of the present invention. DETAILED DESCRIPTION

[0101] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with the implementation modes and the accompanying drawings.

[0102] The control method provided by the present invention first uses a safety preselector to derive safe output information from a set of output measurements, which is then used to construct a state observer; by using integral reinforcement learning technology, the learning system does not require knowledge of drift dynamics. Secondly, a single evaluation neural network is designed to approximate the unknown value function. In order to reduce the computational and communication costs, event-triggered control conditions are designed. In addition, the introduction of the evaluation network successfully avoids the repeated approximation problem that occurs in the traditional dual network. The weight vector in the comment network is adjusted using an empirical delay technique, thereby eliminating the need for restrictive continuous excitation conditions.

[0103] A human-in-the-loop elastic control method for a multi-agent system based on event-triggered integral reinforcement learning. The flowchart of this method is shown in the figure. Figure 1 As shown, the following steps are included:

[0104] S1. Based on a nonlinear multi-agent system, the nonlinear multi-agent system includes a leader and several followers, and a dynamic model of the ith follower is established, specifically,

[0105]

[0106] in, represents the derivative, x i (t) is the state information of the i-th agent at time t, u i (t) is the control input of the ith agent at time t, f i (x i ) is the known internal function of the ith follower, g i (x i ) is the known input matrix function of the ith follower, y i (t) is the measured signal obtained after transmission, The malicious attack signal injected by the attacker, Pi is the system measurement matrix; x i For x i (t) is the abbreviation of x i =[x i,1 ,…,x i,n ] T ∈R n ,u i (t)∈R m , f i (x i )∈R n , g i (x i )∈R n×m , R refers to the real number domain, n and m refer to the dimensions of the matrix, and p i is the number of sensor outputs;

[0107] S2. Give the leader's dynamic model and define the local neighborhood consistency error. The specific process is:

[0108] The dynamic model of leaders is:

[0109]

[0110] Where x0(t)∈R n represents the state information of the leader at time t, y0(t) represents the output of the leader at time t, and u0(t) is the control input, which is an unknown bounded variable;

[0111] Define the local neighborhood consistency error δ i (t) is:

[0112]

[0113] Among them, b i represents the containment gain, a ij represents the connection weight between the ith agent and the jth agent, x j (t) represents the state vector of the jth agent at the tth time, N i represents the set of neighboring agents of the ith agent;

[0114] S3. Filter the unattacked data through the median operator Med[·], obtain the security preselector based on the unattacked data, and design the state observer. The specific process is as follows:

[0115] The status information y i The elements in (t) are arranged in ascending order to obtain a new vector α i =[α i,1 ,…,αi,pi ] T , α i,1 ≤…≤α i,pi ,

[0116] Then the median operator Med[·] is,

[0117]

[0118] Design the following safety preselector x for the i-th agent: i,1,p (t):

[0119] x i,1,p (t) = Med[y i (t)](4)

[0120] Using the unattacked output data x i,1 (t), the first state of the system, Represents the estimation, for the i-th follower, the following state observer is designed To estimate the state of the system:

[0121]

[0122] in, and K0 represent the state estimation vector and the given positive gain respectively.

[0123] The observation error is defined as The error dynamics can be obtained as

[0124]

[0125] S4. Construct the optimal performance index function and establish the Hamilton-Jacobi-Bellman equation. The specific process is:

[0126] Construction and local neighborhood consistency error δ i and control input u i Related performance index function V i (δ i ), the specific form is:

[0127]

[0128] Among them, U i (·) represents the utility function, Q ii , R ii and R ij All are constant matrices, and all three are greater than 0;

[0129] Defining V based on the local neighborhood consistency error i (δ i )’s Hamiltonian:

[0130]

[0131] l ij is the element in the Laplace matrix, L=DA=[l ij ]∈R N×N , D is the degree matrix, D = diag(d1,…,d N ), represents the diagonal elements of the degree matrix, A is the adjacency matrix, and represents the communication topology of the directed graph. A=[a ij ]∈R N×N , a ij is an element in the adjacency matrix A, and l ij =-a ij ; represents the drift dynamics model, b ij =0;

[0132] Based on Bellman's optimality principle, the optimal cost function V i * (δ i ) is in the form of:

[0133]

[0134] And satisfies the following Hamiltonian function:

[0135]

[0136] in, is the optimal cost function V i * (δ i ) About δ i The partial derivative of

[0137] The optimal control strategy The expression is:

[0138]

[0139] Will Substituting into the Hamiltonian function (6), we can get the Hamilton-Jacobi-Bellman equation as follows:

[0140]

[0141] Typically, the optimal controller can be derived by solving the Hamilton-Jacobi-Bellman equation (8). However, the existing integral reinforcement learning method formula (7) It is proposed under the time period sampling mechanism. The control law based on time period sampling usually leads to inefficient use of communication resources between the controlled system and the actuator. More importantly, the control strategy based on time period sampling usually leads to excessive computational load. In addition, due to The drift dynamics model of is unknown, and it is not feasible to directly solve the Hamilton-Jacobi-Bellman equation (8). Therefore, in order to meet these challenges, the present invention designs an integral reinforcement learning method in event-triggered mode to approximate the solution of the Hamilton-Jacobi-Bellman equation (8), thereby eliminating the need for drift dynamics model information.

[0142] S5. Introduce the event trigger mechanism and construct the event-triggered Hamilton-Jacobi-Bellman equation. The specific process is as follows:

[0143] Let the μth triggering time be t μ , and satisfies t μ <t μ+1 , where μ∈N, the sequence of triggering moments can be obtained In t μ The state of the sampling at the moment is expressed as

[0144] At two consecutive triggering times t μ and t μ+1 Between, that is, t∈[t μ ,t μ+1 ), there are usually two states and The error between them is defined as the error function, recorded as Its form is:

[0145]

[0146] According to the local neighborhood consistent error formula, define the local neighborhood consistent error based on event triggering for,

[0147]

[0148] Event trigger error for:

[0149]

[0150] The event-triggered optimal control strategy is:

[0151]

[0152] in, Indicated in In the case of i* δ i The partial derivative of

[0153] Therefore, the event-based Hamilton-Jacobi-Bellman equation is,

[0154]

[0155] By observing formula (9) The expression of : yes Since the event-triggered Hamilton-Jacobi-Bellman equation (10) is a nonlinear partial differential equation, it is difficult to obtain Therefore, in order to realize the event-triggered optimal control strategy, the integral reinforcement learning algorithm is used to solve the event-triggered Hamilton-Jacobi-Bellman equation and the evaluation neural network is used to estimate the cost function;

[0156] S6. Use the integral reinforcement learning algorithm to solve the event-triggered Hamilton-Jacobi-Bellman equation. The specific process is as follows:

[0157] The integral reinforcement learning algorithm is proposed using reinforcement learning technology. Its advantage is that it can reduce the need for knowledge of system drift dynamics during the analysis process. According to the Bellman optimality principle, the integral Bellman equation for i∈N in the time interval [t, t+T] is derived.

[0158]

[0159] S6.1. Selecting the initial permissible events to trigger the optimal control strategy

[0160] S6.2. For each agent i, calculate the cost function V under the optimal control strategy triggered by the current event i k (δ i (t)),

[0161]

[0162] Where t∈t μ , V i k (0) = 0, k represents the current number of iterations;

[0163] S6.3. Update the event-triggered optimal control strategy according to the cost function and obtain the following event-triggered control strategy:

[0164]

[0165] make Return to step 2 until Vi k →V i * ,

[0166] S7. Use the evaluation neural network to estimate the cost function and obtain the required approximate event-triggered optimal control strategy. The specific process is as follows:

[0167] Establish an evaluation neural network, and apply the approximation properties of the evaluation neural network to any continuous function. The cost function V i * (δ i ) can be written as,

[0168] V i * (δ i )=W i T θ(δ i )+ε i (δ i )(11)

[0169] in, is the ideal weight vector, is the activation function, ε i (δ i )∈R is the approximation error, N c ∈Z + is the number of neurons;

[0170] Then V i * (δ i ) for the local neighborhood consistency error δ i The partial derivative of

[0171]

[0172] in, is the function θ i δ i The partial derivative of is the function ε i δ i The partial derivative of

[0173] Substituting formula (12) into formula (9), we get the following optimal control strategy based on event triggering:

[0174]

[0175] in,

[0176] Since the ideal weight vector W iis unknown, and the event-triggered optimal control strategy (13) cannot be directly obtained; therefore, an evaluation neural network is used to estimate the cost function:

[0177]

[0178] in, is an estimate of the ideal weight;

[0179] Similarly, for The corresponding partial derivatives can be obtained,

[0180]

[0181] Therefore, the approximate event-triggered optimal controller is,

[0182]

[0183] Example 1

[0184] The nonlinear multi-agent system of this embodiment includes five agents, including one leader and four followers;

[0185] The dynamics of the ith agent is as follows:

[0186]

[0187] where x i,1 、x i,2 and u i They represent the incompletely measurable system state and control input respectively. The subscripts 1 and 2 are subsystem 1 and subsystem 2 respectively, indicating that the state contains two subsystems.

[0188] Define the dynamics of leadership as

[0189]

[0190] y0=x 0,s

[0191] Where s=1,2, y0 is the system output, and the control input u0 is an unknown bounded variable.

[0192] In addition, the leader’s control input (denoted as u0) is provided by a human operator and is not accessible to all followers. The control input u0 is designed as follows:

[0193]

[0194] The initial state of the selected follower is: x 1,1 (0) = 0.3, x 1,2 (0) = 0.2, x 2,1 (0) = 0.1, x2,2 (0) = 0.4, x 3,1 (0) = 0.1, x 3,2 (0) = 0.2, x 4,1 (0) = 0.1, x 4,2 (0) = 0.3. The initial state of the leader is: x 0,1 (0) = -0.5, x 0,2 (0) = 0.5. Set the initial evaluation network weight vector to The parameters are: c =0.75, N c =5, M = 2, τ = 0.2, ∈ = 1.732. The weight matrix is ​​selected as Q ii =I2,R ii =1, R ij =0.1.

[0195] The nonlinear intelligent agent system in this embodiment is simulated, and the simulation results are as follows: Figure 2-8 shown.

[0196] Figure 2 The state x of the follower of the present invention i,1 With the leader's state x 01 Graph of changes over time. Figure 3 The state x of the follower of the present invention i,2 With the leader's state x 02 A graph of the changes over time. Figure 2 and Figure 3 It can be seen that no matter which follower and which leader, the system status is guaranteed to be consistent after 15 seconds. Figure 4 is the observation error e of the five agents of the present invention i,1 Graph, Figure 5 is the observation error e of the five agents of the present invention i,2 Curve graph (i is 1, 2, 3, 4, 5). The observation error is Figure 4 and Figure 5 It is proved that even the system state that is not completely measurable can still be observed through the state observer (Formula (5)).

[0197] Figure 6 is the control input trajectory diagram of the five intelligent agents of the present invention. Figure 6 It can be seen that by applying the designed weight update law, the evaluation network weights converge to the ideal value after about 10 seconds, and the final convergence value is [0.6567, 1.2623, 1.0173, 0.8403, 0.6775] T Since the limiting sustained excitation condition has been eliminated by incorporating the historical record data into the controller design, there is no need to introduce a probe signal.

[0198] Figure 7 The control input trajectory diagram of the five agents is shown in FIG. 1 . It can be seen from the figure that the control input trajectory finally converges to the vicinity of the zero point, which indicates that the method provided by the present invention effectively stabilizes the system. In addition, Figure 8 The sampling time intervals of the five agents are depicted. Obviously, the minimum sampling interval is greater than 0. This shows that the event-triggered control strategy proposed in the present invention can avoid the occurrence of Zeno behavior.

[0199] In summary, the simulation results show the effectiveness of the control method provided by the present invention, and provide a useful reference for practical applications.

[0200] The above description is only a specific implementation mode of the present invention. Any feature disclosed in this specification, unless otherwise stated, can be replaced by other alternative features that are equivalent or have similar purposes; all the disclosed features, or all the steps in the methods or processes, except for mutually exclusive features and / or steps, can be combined in any way.

Claims

1. A human-in-the-loop elastic control method for a multi-agent system based on event-triggered integral reinforcement learning, characterized in that: The steps include: S1. Based on the nonlinear multi-agent system, establish the dynamic model of the i-th follower; S2. Give the leader's dynamic model and define the local neighborhood consensus error; Among them, the dynamic model of the leader is: Where x0(t)∈R n represents the state information of the leader at time t, y0(t) represents the output of the leader at time t, and u0(t) is the control input; Define the local neighborhood consistency error δ i (t) is: Among them, b i represents the containment gain, a ij represents the connection weight between the ith agent and the jth agent, x i (t) is the state information of the i-th agent at time t, x j (t) represents the state vector of the jth agent at the tth time, N i represents the set of neighboring agents of the ith agent; S3, filtering the unattacked data through the median operator Med[·], obtaining a security preselector based on the unattacked data, and designing a state observer; S4, construct the optimal performance index function and establish the Hamilton-Jacobi-Bellman equation; S5. Introduce event trigger mechanism and construct event-triggered Hamilton-Jacobi-Bellman equation; S6, using the integral reinforcement learning algorithm to solve the event-triggered Hamilton-Jacobi-Bellman equation; S7. Use the evaluation neural network to estimate the cost function and obtain the required approximate event-triggered optimal controller.

2. The method for human-in-the-loop elastic control of a multi-agent system based on event-triggered integral reinforcement learning as claimed in claim 1, characterized in that: In step S1, the nonlinear multi-agent system includes a leader and several followers, and the dynamic model of the i-th follower is as follows: in, represents the derivative, x i (t) is the state information of the i-th agent at time t, u i (t) is the control input of the ith agent at time t, f i (x i ) is the known internal function of the ith follower, g i (x i ) is the known input matrix function of the ith follower, y i (t) is the measured signal obtained after transmission, The malicious attack signal injected by the attacker, P i is the system measurement matrix; x i For x i (t) is the abbreviation of x i =[x i,1 ,…,x i,n ] T ∈R n ,u i (t)∈R m , f i (x i )∈R n , g i (x i )∈R n×m , R refers to the real number domain, n and m refer to the dimensions of the matrix, and p i is the number of sensor outputs.

3. The method for human-in-the-loop elastic control of a multi-agent system based on event-triggered integrated reinforcement learning as claimed in claim 1, characterized in that: The control input u0(t) is an unknown bounded variable, and its specific form is, t1 is the first time threshold, and t2 is the second time threshold.

4. The method for human-in-the-loop elastic control of a multi-agent system based on event-triggered integrated reinforcement learning as claimed in claim 1, characterized in that: In step S3, the measurement signal y i The elements in (t) are arranged in ascending order to obtain a new vector α i =[α i,1 ,…,α i,pi ] T , α i,1 ≤…≤α i,pi , Then the median operator Med[·] is, Design the following safety preselector x for the i-th agent: i,1,p (t): x i,1,p (t)=Med[y i (t)] (4) Using the unattacked output data x i,1 (t), the first state of the system, Represents the estimation, for the i-th follower, the following state observer is designed To estimate the state of the system: in, and K0 represent the state estimation vector and the given positive gain respectively.

5. The method for human-in-the-loop elastic control of a multi-agent system based on event-triggered integrated reinforcement learning as claimed in claim 1, characterized in that: The specific process of step S4 is: Construction and local neighborhood consistency error δ i and control input u i Related performance index function V i (δ i ), the specific form is: Among them, U i (·) represents the utility function, Q ii , R ii and R ij are all constant matrices; u i (t) represents the control input of the ith agent at time t; u j (t) represents the control input of the j-th agent at time t; Based on the local neighborhood consistency error, define V i (δ i )’s Hamiltonian: l ij is the element in the Laplace matrix, L=DA=[l ij ]∈R N×N , D is the degree matrix, D = diag(d1,L,d N ), represents the diagonal elements of the degree matrix, A is the adjacency matrix, and represents the communication topology of the directed graph. A=[a ij ]∈R N ×N , a ij is an element in the adjacency matrix A, and l ij =-a ij ; represents the drift dynamics model, b ij =0; For the jth follower in an unknown state The input matrix function under d i represents the diagonal elements of the degree matrix; Based on Bellman's optimality principle, the optimal cost function V i * (δ i ) is in the form of: And satisfies the following Hamiltonian function: in, is the optimal cost function V i * (δ i ) About δ i The partial derivative of The optimal control strategy The expression is: Will Substituting into the Hamiltonian function (6), we can get the Hamilton-Jacobi-Bellman equation as follows: is the optimal control strategy of the ith agent; Represents the optimal control strategy of the jth agent.

6. The method for human-in-the-loop elastic control of a multi-agent system based on event-triggered integrated reinforcement learning as claimed in claim 5, characterized in that: Constant Matrix Q ii , R ii and R ij Both are greater than 0.

7. The method for human-in-the-loop elastic control of a multi-agent system based on event-triggered integrated reinforcement learning as claimed in claim 5, characterized in that: The specific process of step S5 is: Let the μth triggering time be t μ , and satisfies t μ <t μ+1 , Where μ∈N, N is a set of natural numbers, then the sequence of triggering times can be obtained In t μ The state of the sampling at the moment is expressed as At two consecutive triggering times t μ and t μ+1 Between, that is, t∈[t μ ,t μ+1 ), there are usually two states and The error between them is defined as the error function, recorded as Its form is: According to the local neighborhood consistent error formula, define the local neighborhood consistent error based on event triggering for, The event-triggered optimal control strategy is: in, Indicated in In the case of i * δ i The partial derivative of Therefore, the event-based Hamilton-Jacobi-Bellman equation is, 8. The method for human-in-the-loop elastic control of a multi-agent system based on event-triggered integrated reinforcement learning as claimed in claim 7, characterized in that: The specific process of step S6 is: S6.

1. Selecting the initial permissible events to trigger the optimal control strategy; S6.

2. For each agent i, calculate the cost function V under the optimal control strategy triggered by the current event i k (δ i (t)), Where t∈t μ , V i k (0) = 0, k represents the current number of iterations; S6.

3. According to the cost function event triggering optimal control strategy, the following event-triggered control strategy is obtained: make Return to step S6.2 until V i k →V i * , 9. The method for human-in-the-loop elastic control of a multi-agent system based on event-triggered integrated reinforcement learning as claimed in claim 8, characterized in that: The specific process of step S7 is: Establish an evaluation neural network, and apply the approximation properties of the evaluation neural network to any continuous function. The cost function V i * (δ i ) can be written as, V i * (d i )=W i T θ(δ i )+e i (d i ) (11) in, is the ideal weight vector, is the activation function, ε i (δ i )∈R is the approximation error, N c ∈Z + is the number of neurons; Then V i * (δ i ) for the local neighborhood consistency error δ i The partial derivative of in, is the activation function θ i δ i The partial derivative of is the approximate error function ε i δ i The partial derivative of Substituting formula (12) into formula (9), we get the following optimal control strategy based on event triggering: in, Since the ideal weight vector W i is unknown, and the event-triggered optimal control strategy (13) cannot be directly obtained; therefore, an evaluation neural network is used to estimate the cost function: in, is an estimate of the ideal weight; Similarly, for The corresponding partial derivatives can be obtained, Therefore, the approximate event-triggered optimal controller is,

Citation Information

Patent Citations

  • Model unknown multi-agent consistency control method based on reinforcement learning

    CN112947084A