Optimal consistent control method of random multi-agent system based on dynamic event triggering

By introducing dynamic event triggering mechanism and neural network model in a random multi-agent system, the adaptive identifier and actor-critic algorithm are designed, the optimal consistent control of the system is achieved, the problem of large-scale communication resource occupation under the time-triggering strategy is solved, and the stability and efficiency of the system are improved.

CN120103711APending Publication Date: 2025-06-06NANKAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510332881.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The optimal consistent control strategy of existing random multiagent systems is usually time-triggered, resulting in a large consumption of communication resources and it is difficult to achieve optimal control while taking into account system stability and cost functions.

Method used

An optimal consistent control method based on dynamic event triggering is proposed. By constructing a consistent error dynamic system between the follower state and the navigator state, defining the sampling error and threshold function of the dynamic event trigger, designing an adaptive identifier and neural network model, and updating the control strategy using the actor-critic algorithm structure.

Benefits of technology

The optimal consistent control of the nonlinear random multi-agent system is realized, which avoids the occurrence of Zeno behavior, saves communication resources, reduces energy loss, extends the service life of the controller, and effectively solves the performance uncertainty caused by random perturbations and unknown nonlinear dynamics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120103711A_ABST
    Figure CN120103711A_ABST
Patent Text Reader

Abstract

The invention relates to the field of random multi-agent system control, and particularly discloses an optimal consistent control method of a random multi-agent system based on dynamic event triggering. For a first-order nonlinear random multi-agent system, the method effectively eliminates the uncertain influence caused by random disturbance and unknown nonlinearity in the system by designing an adaptive identifier based on a neural network. Besides, in order to save communication resources, auxiliary dynamic variables are introduced on the basis of static event triggering conditions, a dynamic event triggering mechanism is designed, and Zeno behaviors are avoided. Under the framework of an act-critic algorithm, a weight updating rule of a neural network is designed by performing gradient descent on a positive function related to a Bellman residual error, and the control algorithm is simplified. Finally, the effectiveness of the designed control algorithm is verified through theoretical proof and simulation experiments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of random multi-agent system control, and in particular to an optimal consensus control method for a random multi-agent system based on dynamic event triggering. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] With the rapid development of intelligent technology, the cooperative control of multi-agent systems has been widely used in various fields, including intelligent transportation, drone formation, robot collaboration, etc. Consensus control, as the basis of cooperative control, provides a basic design idea for other cooperative controls such as formation control and inclusive control. It usually refers to designing a control protocol so that the states of all agents tend to the same value.

[0004] In practical applications, systems are often affected by random disturbances. In addition, in many cases, it is necessary to further consider the cost function while considering the stability of the system in order to save resources and reduce costs. Therefore, the optimal consensus control of random multi-agent systems has received more and more attention. However, the current optimal consensus control strategy for random multi-agent systems is usually time-triggered, and control commands are executed regularly at fixed time intervals to ensure the effective update of the system state, but this often takes up a lot of communication resources.

[0005] Different from traditional time-triggered control, event-triggered control only allows the control strategy to be updated when a specific event occurs, which will effectively reduce communication overhead and computational burden while maintaining the stability and response speed of the system. Depending on whether the design of the trigger threshold includes dynamic variables, event-triggered control can be divided into dynamic event triggering and static event triggering. Studies have shown that dynamic event triggering can further expand the trigger threshold compared to static event triggering, thereby saving more communication resources. To this end, the present invention proposes an optimal consensus control method based on dynamic event triggering, which can achieve optimal consensus control of nonlinear random multi-agent systems. Summary of the invention

[0006] The purpose of the present invention is to provide an optimal consensus control method based on dynamic event triggering in order to achieve optimal consensus of a nonlinear random multi-agent system while saving communication resources as much as possible.

[0007] To achieve the above object, the technical solution implemented by the present invention is as follows:

[0008] The present invention is an optimal consensus control method for a random multi-agent system based on dynamic event triggering, comprising the following steps:

[0009] Step 1: Consider a A random multi-agent system consisting of followers and a navigator is proposed. The consistent error between the follower state and the navigator state is constructed, and the consistent error dynamic system is derived.

[0010] Followers in nonlinear stochastic multi-agent systems The kinetic model is expressed as:

[0011] ,

[0012] in Is a follower ( ) status, is the control input, and are all unknown continuous and bounded nonlinear functions, is the control gain and satisfies ,in is a positive constant. In addition, is an independent Wiener process.

[0013] The status of the navigator and its first-order derivative Known.

[0014] The consistent error between the follower state and the leader state is established as:

[0015] ,

[0016] in Is a follower The set of neighbor agents, if ,but ,otherwise If the follower If the information of the pilot can be received, ,otherwise .

[0017] The dynamic equation of the consistent error system is expressed as:

[0018] ,

[0019] in , , .

[0020] Step 2: Define the value function as the performance indicator for evaluating the control effect of the random multi-agent system, and obtain the HJB equation based on the consistent error system and the value function.

[0021] The value function is designed as:

[0022] .

[0023] The HJB equation is expressed as:

[0024] ,

[0025] in is the optimal control strategy, .

[0026] Step 3: Construct a dynamic event trigger sampling error and threshold function. When the sampling error exceeds the trigger threshold, the event is triggered and the control strategy is updated.

[0027] The event-triggered sampling error is designed to be , among which hour, , For followers No. A trigger moment.

[0028] The trigger threshold is designed as:

[0029] ,

[0030] in and is a positive parameter, a positive constant make satisfy , auxiliary dynamic variables The dynamic equation is designed as:

[0031] ,

[0032] in and is a positive constant and satisfies ,also, Initial value of .

[0033] Step 4: Design an adaptive marker to approximate the random terms and unknown nonlinearities in the system.

[0034] Adaptive markers are designed to:

[0035] ,

[0036] in is the state of the marker, represents the estimated error of the marker, It is a parameter that needs to be designed. and are the weight estimates and activation functions of the marker neural network, respectively, and The update rule is designed as:

[0037] ,

[0038] in It is a parameter that needs to be designed.

[0039] Step 5: Construct a two-layer neural network model consisting of a critic neural network and an actor neural network. Use the critic neural network to approximate the value function of each agent, and use the actor neural network to approximate the control strategy function of each agent.

[0040] The value function fitted by the critic neural network is:

[0041] ,

[0042] in is the weight estimate of the critic neural network, is the activation function.

[0043] The control strategy fitted by the actor neural network is:

[0044] ,

[0045] in are the estimated weights of the actor neural network.

[0046] Step 6: Design the weight update rules of the critic neural network and the actor neural network based on the gradient descent algorithm, and obtain the optimal control strategy by adjusting the weight parameters.

[0047] and The update rule is designed as:

[0048] ,

[0049] ,

[0050] in and are the learning rates of the critic neural network and the actor neural network respectively.

[0051] Step 7: Design the Lyapunov function and analyze the stability and error convergence of the system.

[0052] The beneficial effects of the present invention are:

[0053] 1. The present invention realizes the optimal consistent control of random multi-agent systems by applying a control algorithm based on dynamic event triggering;

[0054] 2. The dynamic event triggering condition introduced by the present invention avoids the occurrence of Zeno behavior and saves more communication resources than the static event triggering algorithm, thereby reducing the energy loss of the multi-agent system and extending the service life of the controller;

[0055] 3. The present invention designs an adaptive identifier based on a neural network for a random multi-agent system, which effectively solves the uncertain impact of random disturbances and unknown nonlinear dynamics on system performance;

[0056] 4. The present invention adopts the actor-critic algorithm structure in reinforcement learning and designs a weight update rule by performing gradient descent on a positive function related to the Bellman residual, which not only ensures the stability of the system but also greatly reduces the complexity of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The accompanying drawings, which constitute a part of the specification of the present invention, are used to provide a further understanding of the present invention. The illustrative examples of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0058] Figure 1 This is a flow chart of the present invention for realizing optimal consensus control of a random multi-agent system triggered by dynamic events;

[0059] Figure 2 The present invention is The control block diagram of a follower;

[0060] Figure 3 It is a topological diagram of a first-order random multi-agent system consisting of three agents in a simulation instance;

[0061] Figure 4 is the consistent error convergence diagram of each follower in the simulation instance;

[0062] Figure 5 It is the state trajectory diagram of each agent in the simulation instance;

[0063] Figure 6 It is the convergence diagram of the estimation error of the marker and the weight parameters of the marker neural network in the simulation example; DETAILED DESCRIPTION

[0064] The technical solutions in the embodiments of the present invention will be further described in detail below in conjunction with the accompanying drawings in the embodiments of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those generally understood by those of ordinary skill in the art to which the present invention belongs. However, it should be understood that these practical details should not be used to limit the present invention. That is to say, in some embodiments of the present invention, these practical details are not necessary.

[0065] The present invention is an optimal consensus control method for a random multi-agent system based on dynamic event triggering. The control step flow chart is as follows: Figure 1 shown.

[0066] 1. System Modeling and Problem Description

[0067] contain The topology of a random multi-agent system with 10 followers is a directed graph. To indicate that is a node set, is an edge set, is the adjacency matrix. If the node pair , then there is a slave node To Node information flow, ,otherwise .when When Is a node neighbors, and the nodes The neighbor set of .make ,in ,use To represent the graph The Laplacian matrix of . Define the diagonal matrix Indicates the communication between the leader node and the follower node. If the leader's information can be directly transmitted to the follower node ,but ;otherwise Assuming that at least one follower node is directly connected to the leader node, then .

[0068] Assumption 1: The communication topology graph has a directed spanning tree with the navigator as the root node.

[0069] Consider a class of A random multi-agent system consisting of followers and a leader. The dynamic equation of the follower is:

[0070] ,

[0071] in Is a follower ( ) status, is the control input, and are all unknown continuous and bounded nonlinear functions, is the control gain and satisfies ,in is a positive constant. In addition, is an independent standard Wiener process.

[0072] The status of the navigator and its first-order derivative known;

[0073] Building a following The consistent error between the state of and the navigator state is:

[0074] .

[0075] The dynamic equation of the consistent error system is expressed as:

[0076] ,

[0077] in , , .

[0078] Control objective: Design an optimal consistent control strategy based on dynamic event triggering so that the stochastic agent system (1) achieves optimal consistency and all error signals of the closed-loop control are ultimately uniformly bounded.

[0079] 2. Design of Optimal Consensus Control Algorithm

[0080] For the optimal consensus control problem of stochastic MASs, the goal is to Finding the Optimal Controller , while each agent in the system achieves consistency, minimize the following performance indicator function

[0081] .

[0082] Then, the optimal performance indicator function can be expressed as

[0083] .

[0084] So the corresponding HJB equation can be expressed as

[0085] ,

[0086] in .

[0087] According to the optimal control theory, the optimal controller It can be solved by Get, that is:

[0088] .

[0089] It should be noted that It is based on time triggering, which usually leads to unnecessary resource consumption. In order to avoid this problem, the present invention will introduce a dynamic event triggering mechanism in the next step.

[0090] 3. Design of dynamic event triggering mechanism

[0091] Using a monotonically increasing sequence Indicates Follower's trigger moment, then the sampling state can be expressed as , In order to determine the triggering moment, the following sampling error is defined:

[0092] ,

[0093] When the sampling error exceeds the trigger threshold defined later, the control strategy is updated and the sampling error is reset to 0. According to equation (7), the control strategy based on event triggering can be expressed as:

[0094] .

[0095] Then the corresponding event-triggered HJB equation can be written as:

[0096] .

[0097] Design dynamic event trigger conditions as follows:

[0098] ,

[0099] in and is a positive parameter, a positive constant make satisfy , auxiliary dynamic variables The dynamic equation is designed as:

[0100] ,

[0101] in and is a positive constant and satisfies ,also, Initial value of .

[0102] In order to prove that the designed dynamic event triggering conditions can ensure the stability of the system, the present invention will firstly put forward the following assumptions.

[0103] Assumption 2: Optimal Controller Satisfies the Lipschiz condition, that is, there is a normal number So that the following formula holds

[0104] .

[0105] For Followers , choose the Lyapunov function as . Calculating its infinitesimal operator, we can get:

[0106] .

[0107] According to HJB equation (6), we know that:

[0108] .

[0109] Substituting (15) into (14) we get:

[0110] .

[0111] According to formula (7), we have Combining Young's inequality and Assumption 2, we have . Combined with the auxiliary dynamic variables (12), (16) can be expressed as:

[0112] .

[0113] If the dynamic event triggering condition (11) is satisfied, then equation (17) can be transformed into:

[0114] .

[0115] because , so when choosing appropriate parameters, So we can conclude that if the dynamic event trigger condition is set as (11), then in the controller based on dynamic event trigger Under this condition, the consistent error system (3) can achieve eventual consistent bounded stability.

[0116] 4. Adaptive Marker Design

[0117] Due to the existence of random disturbances and unknown nonlinear dynamics in the system, the following identifier is designed:

[0118] ,

[0119] in is the state of the marker, represents the estimated error of the marker, It is a parameter that needs to be designed. and are the weight estimates and activation functions of the marker neural network, respectively, and The update rule is designed as:

[0120] ,

[0121] in It is a parameter that needs to be designed.

[0122] Combining the marker (19) and the consistent error system (3), we can get the marker estimation error The kinetic equation is:

[0123] ,

[0124] in , , .and The labeler neural network is approximated as ,in is the ideal weight parameter, is the estimation error of the neural network.

[0125] To illustrate the estimation error of the marker and the weight estimation error of the marker neural network In order to understand the boundedness of , the present invention will first introduce the following common assumptions.

[0126] Assumption 3: and are bounded and satisfy and ,in and They are all normal numbers.

[0127] For Followers , choose the Lyapunov function as , combining (19) and (20), we can get The infinitesimal operator of is:

[0128] .

[0129] According to Young's inequality, we have . So (22) can be expressed as:

[0130] .

[0131] because , and , (23) can be written as ,in , So we can conclude that the estimated error of the marker is and the weight estimation error of the marker neural network is semiglobally eventually uniformly bounded.

[0132] 5. Actor-critic network design

[0133] Based on the above analysis, it can be seen that when considering the marker, the marker system (19) can be used to replace the consistent error system (2) in the subsequent discussion. Therefore, replace , replace etc.

[0134] The corresponding HJB equation is

[0135] .

[0136] A two-layer neural network model consisting of a critic neural network and an actor neural network is constructed. The critic neural network is used to approximate the value function of each agent, and the actor neural network is used to approximate the control strategy function of each agent.

[0137] The value function fitted by the critic neural network is expressed as ,in is the ideal weight value of the critic neural network, is the activation function, is the estimation error of the neural network, and there is a positive constant Satisfy . Then the approximate value function is:

[0138] ,

[0139] in is the estimated weight of the critic neural network.

[0140] The control strategy fitted by the actor neural network is expressed as ,in is the ideal weight value of the actor neural network. Then the approximate control strategy is:

[0141] ,

[0142] in are the estimated weights of the actor neural network.

[0143] The weight update rules of the critic neural network and the actor neural network are designed based on the gradient descent algorithm, and the optimal control strategy is obtained by adjusting the weight parameters.

[0144] and The update rule is designed as:

[0145] ,

[0146] ,

[0147] in and are the learning rates of the critic neural network and the actor neural network respectively.

[0148] It should be noted that in order to further illustrate the effectiveness of the weight update rules (27) and (28) of the critic and actor neural networks, the present invention will provide the following analysis.

[0149] Substituting (27) and (28) into the HJB equation, we have

[0150] .

[0151] The Bellman residual is defined as

[0152] ,

[0153] because , so (30) can be further written as . Control strategy is the only solution to the HJB equation, and it is necessary to decay the Bellman residual to 0. Assume There is a unique solution, so we can get the following equation:

[0154] .

[0155] If we choose a positive definite function , then it is obvious that (31) is the same as are equivalent. So the weight update rules of the actor and critic neural networks can be obtained by Perform gradient descent to obtain .

[0156] Combining (26) and (27), The time derivative of

[0157] .

[0158] From (32), we can get that when When the weight update rules (27) and (28) can ensure and . Different from directly calculating the Bellman residual Perform gradient descent, using a positive definite function This can make the design of update rules simpler.

[0159] 6. Stability Analysis

[0160] Combined with the marker, the event-triggered sampling error can be redefined as ,in .

[0161] In order to illustrate the boundedness of the consistency error and the neural network weight estimation error under the optimal control law obtained in the actor-critic framework, the present invention will first introduce the following common assumptions.

[0162] Assumption 4: Activation function of neural network Bounded and satisfied ,in is a normal number.

[0163] For Followers , choose the Lyapunov function as:

[0164] ,

[0165] in , , , The weight estimation errors of the critic and actor neural networks are and .

[0166] Since the control strategy is event-triggered, it is necessary to explain the boundedness of the error system from both the triggering moment and the non-triggering moment.

[0167] Case 1: At the non-event triggering moment, that is Calculation The infinitesimal operator of , then we have:

[0168] .

[0169] because ,in and are all positive vectors, then we know that:

[0170] .

[0171] Inspired by formula (9), we have:

[0172] ,

[0173] .

[0174] Substituting (36) and (37) into (35), we have:

[0175] .

[0176] So formula (34) can be expressed as:

[0177] .

[0178] Combined with the event trigger condition (11), (39) becomes:

[0179] .

[0180] According to the actor and critic neural network weight update rules (27) and (28), we have

[0181] ,

[0182] .

[0183] Substituting (41) and (42) into (40), if and If established,

[0184] .

[0185] By choosing appropriate parameter values If the inequality holds, then

[0186] ,

[0187] ,

[0188] ,

[0189] When it was established, there was That is, when the event is not triggered, the consistent error and the neural network weight estimation error and is bounded.

[0190] Case 2: At the time of event triggering, that is When. Lyapunov function The infinitesimal operator can be written as

[0191] .

[0192] From case 1, we can see that for ,have Taking the limit on both sides of the equation, we have From the above analysis, we can know that the state of the marker is asymptotically convergent, so So when inequalities (44)-(46) do not hold, we have . So, at the time of event triggering, the consistent error , critic neural network weight estimation error and actor neural network weight estimation error All are bounded.

[0193] 7. Zeno Behavior Analysis

[0194] In order to illustrate that the dynamic event triggering mechanism designed by the present invention will not cause the system to have Zeno behavior, the following common assumptions are first introduced.

[0195] Assumption 4: Satisfies the Lipschitz condition, that is, there is a positive constant make Established.

[0196] When Assumption 4 holds, system (19) satisfies

[0197] .

[0198] Through the above analysis, we can see that there is a normal number make Established. , then:

[0199] .

[0200] By taking the derivative of both sides of (44), and knowing ,so According to the dynamic event triggering condition (11), we can know ,then This also shows that , that is, Zeno behavior will not occur.

[0201] 8. Simulation Experiment

[0202] In order to further verify the effectiveness of the proposed optimal consistent control strategy, the present invention will conduct the following numerical simulation experiments. Consider a For a 3-agent multi-agent system, the dynamic equation of each follower is

[0203]

[0204] in It is The status of a follower, . , , . Set their initial values ​​to , , The dynamic equation of the navigator is , and the initial state is set to 1. The communication topology of the multi-agent system is shown in Figure 3 shown.

[0205] In order to obtain the optimal control strategy, the parameters in the identifier (19) and the identifier neural network weight update rule (20) are designed as , , , The learning rates of the actor and critic neural networks are set to , , , , , In addition, the marker neural network, actor neural network, and critic neural network are all configured with 12 neurons, and the activation function is centered evenly distributed in The initial value of the neural network is set to , , , , , , , , The parameters in the dynamic event trigger mechanism are designed as , , , , .

[0206] Figure 4-Figure 6 The simulation results are given. Figure 4 The consistent error convergence diagram of each follower is given, and it can be seen that the consistent errors of the three followers can eventually converge to near 0. Figure 5 The state trajectory diagram of each intelligent agent is displayed, and it can be seen that the states of the three followers can eventually tend to the leader state. Figure 6 It is the convergence diagram of the estimation error of the marker and the weight parameter of the marker neural network. It can be seen that the estimation error of the marker can quickly converge to near 0, indicating that the marker can accurately estimate the random disturbances and unknown nonlinear dynamics in the system. The simulation results further show that the proposed optimal consistent control method based on dynamic event triggering can achieve the expected control objectives.

[0207] The present invention is not limited to the above-mentioned implementation modes, and anyone should be aware of the structural changes made under the enlightenment of the present invention, and all those having the same or similar technical solutions as the present invention fall within the protection scope of the present invention.

[0208] The techniques, shapes, and structural parts not described in detail in the present invention are all well-known techniques.

Claims

1. An optimal consensus control method for a random multi-agent system triggered by dynamic events, characterized in that: The following steps are involved: Step 1: Establish a first-order nonlinear random multi-agent system model with a directed graph topology. The multi-agent system consists of It is composed of agents, constructs the consistent error between the follower state and the leader state, and derives the consistent error dynamic system; Step 2: Define the value function as the performance indicator for evaluating the control effect of the random multi-agent system, and obtain the HJB equation based on the consistent error system and the value function; Step 3: Construct a dynamic event trigger sampling error and threshold function. When the sampling error exceeds the threshold function, the event is triggered and the control strategy is updated. Step 4: Design an adaptive marker; Step 5: Construct a two-layer neural network model including a critic neural network and an actor neural network. Use the critic neural network to approximate the value function of each agent, and use the actor neural network to approximate the control strategy function of each agent. Step 6: Design the weight update rules of the critic neural network and the actor neural network based on the gradient descent algorithm, and obtain the approximate solution of the HJB equation by adjusting the weight parameters; Step 7: Design the Lyapunov function and analyze the stability and error convergence of the system.

2. The optimal consensus control method for a random multi-agent system based on dynamic event triggering according to claim 1 is characterized in that: In the step 1, the follower in the nonlinear random multi-agent system The kinetic model is expressed as: , in Is a follower ( ) status, is the control input, and are all unknown continuous and bounded nonlinear functions, is the control gain and satisfies ,in is a positive constant. In addition, It is an independent Wiener process; The status of the navigator and its first-order derivative known; The consistent error between the follower state and the leader state is established as: , in Is a follower The set of neighbor agents, if ,but ,otherwise , if the follower If the information of the pilot can be received, ,otherwise ; The dynamic equation of the consistent error system is expressed as: , in , , .

3. The optimal consensus control method for a random multi-agent system based on dynamic event triggering according to claim 1 is characterized in that: In the step 2, the value function is designed as: ; The HJB equation is expressed as: in is the optimal control strategy, .

4. The optimal consensus control method for a random multi-agent system based on dynamic event triggering according to claim 1 is characterized in that: In step 3, the event-triggered sampling error is designed to be , among which hour, , For followers No. A trigger moment; The trigger threshold is designed as: , in and is a positive parameter, a positive constant make satisfy , auxiliary dynamic variables The dynamic equation is designed as: , in and is a positive constant and satisfies ,also, Initial value of .

5. The optimal consensus control method for a random multi-agent system based on dynamic event triggering according to claim 1 is characterized in that: In step 4, the adaptive marker is designed as follows: , in is the state of the marker, represents the estimated error of the marker, It is a parameter that needs to be designed. and are the weight estimates and activation functions of the marker neural network, respectively, and The update rule is designed as: 。 6. The optimal consensus control method for a random multi-agent system based on dynamic event triggering according to claim 1 is characterized in that: In step 5, the value function fitted by the critic neural network is: , in is the weight estimate of the critic neural network, is the activation function; The control strategy fitted by the actor neural network is: , in are the estimated weights of the actor neural network.

7. The optimal consensus control method for a random multi-agent system based on dynamic event triggering according to claim 1 is characterized in that: In the step six, and The update rule is designed as: , , in and are the learning rates of the critic neural network and the actor neural network respectively.