Multi-agent cooperative control system and method based on side event triggering
By adopting a collaborative control method based on edge event triggering and a generalized strategy iterative algorithm in multi-agent systems, the problems of communication redundancy and computation complexity in the prior art are solved, and more efficient resource utilization and autonomous decision-making capabilities are achieved.
Patent Information
- Application Number
- CN202510261611.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-27
AI Technical Summary
The existing event-triggered multi-agent collaborative control strategy has communication redundancy problems, resulting in waste of resources. The traditional consensus protocol lacks autonomy and optimization, slow convergence speed and complex calculations.
Adopting multi-agent collaborative control method based on edge event triggering, combining Bellman's optimal method and generalized strategy iterative algorithm, adaptive dynamic programming technology and edge event triggering mechanism are designed to optimize communication efficiency and computational complexity.
It effectively reduces the communication loss and computing burden of multi-agent systems, improves the efficiency of collaborative control and autonomous decision-making capabilities, and achieves better state synergy effects.
Smart Images

Figure CN120215259A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of event triggering, reinforcement learning control, and multi-agent systems, and more specifically, to a multi-agent cooperative control system and method based on edge event triggering. Background Art
[0002] Compared with single-agent systems, multi-agent systems have great advantages in terms of computing power, decision-making ability, and robust performance, and have received extensive attention in the academic and industrial fields. The consensus problem is one of the main forms of multi-agent cooperative control, and its goal is to make all agents in the multi-agent system tend to synchronize, which means that the states of all agents finally reach consistency. Traditional cooperative control strategies are all based on the time-triggered mechanism, that is, the system state is sampled periodically at fixed time, and then the control strategy is updated, which often causes huge communication redundancy and computational burden. To solve this problem, the event-triggered mechanism of aperiodic sampling is proposed, so as to achieve good cooperative effects with fewer control actions and communication resources.
[0003] However, most of the current multi-agent cooperative control strategies based on event triggering design the event-triggered mechanism based on agents, that is, when the state measurement error of an agent exceeds a certain threshold, the communication with all neighbor agents is restored, and the control strategy is updated. This easily leads to the restoration of communication between two agents with similar or the same states at the trigger moment, resulting in communication redundancy and wasting communication resources. To solve this problem, an edge-based event-triggered mechanism is proposed. Under the edge-triggered condition, the event-triggered mechanism acts on the communication edge, and each edge separately judges the trigger condition. Only the adjacent agents with state errors exceeding a certain threshold will restore communication at the edge trigger moment, thus greatly reducing the situation of communication redundancy.
[0004] In addition, most of the traditional multi-agent system cooperative controllers are based on the form of typical consensus protocols, and the agents lack autonomy and optimality. Adaptive dynamic programming technology combines the basic knowledge of dynamic programming, neural networks, and reinforcement learning, and is a typical learning-based control strategy. Value iteration and policy iteration, as two typical iterative learning algorithms of adaptive dynamic programming technology, have problems of slow convergence speed and high computational complexity. Therefore, the generalized policy iteration algorithm is proposed to achieve a compromise between the convergence speed and the computational complexity.
[0005] How to combine the above technologies to develop an intelligent cooperative control method and an edge event-triggered method is very necessary for improving the multi-agent cooperative control effect and communication efficiency. Summary of the Invention
[0006] The present invention provides a collaborative control method for multi-agent systems based on edge event triggering. This method solves the problem of multi-agent collaborative control. The present invention uses the Bellman optimal method to ensure the optimality of the collaborative control strategy, uses the adaptive dynamic programming technology based on the generalized policy iteration algorithm to improve the autonomous decision-making ability, and uses the edge event triggering mechanism to improve the communication efficiency at the point-to-point level between nodes. Details are described below:
[0007] The present invention is implemented by adopting the following technical solutions:
[0008] A collaborative control method for multi-agent systems based on edge event triggering, the method comprising the following steps:
[0009] Establish an optimal collaborative control model for each agent in the multi-agent system;
[0010] Through the Bellman optimal method and the cost function V of the agent i (e i ) for the optimal collaborative control model Calculate the optimal collaborative control strategy for each agent
[0011] Construct an evaluation neural network and use the generalized policy iteration algorithm to trigger the edge event moment for the optimal collaborative control strategy Control and output a multi-control strategy based on edge events
[0012] Construct a Lyapunov function based on the local cost function and establish an edge event triggering rule for the agent;
[0013] Apply the collaborative control strategy of the event and the edge event triggering rule to the multi-agent system.
[0014] Furthermore, the optimal collaborative control strategy for the multi-agent includes:
[0015] 101. Establish a single-agent model according to the following formula:
[0016]
[0017] where f i (x i ) is the internal dynamic information of the system of agent i, g i (x i ) is the control input matrix of agent i, u i is the control strategy, N represents the number of agents, and when i = 0, it means that the agent is the leader, and the rest are followers;
[0018] 102. The optimal cooperative control model of each agent is calculated according to the following formula, that is, the dynamic equation representing the synchronization error is obtained:
[0019]
[0020] Among them, for agent i, e i is the local neighborhood error, F i is the internal dynamic information of the error system, a ij is the element of the adjacency matrix, b i is the element of the reachability matrix; l ii is the element of the Laplacian matrix;
[0021] 103. A quadratic cost function V i (e i (t)) is set for each agent according to the following formula:
[0022]
[0023] Among them, Q, R ii and R ij are positive definite and symmetric constant matrices, used to adjust the components of the system operation cost and control cost in the cost function, N i is the set of neighbor agents of agent i; it can be seen that this cost function includes the control costs of all neighbor agents;
[0024] 104. Based on the optimal cost function, according to the Bellman optimality principle, the optimal cooperative control strategies of each agent are obtained as follows:
[0025]
[0026] Among them, is the partial derivative of the optimal cost function V i * (e i (t)) with respect to e i (t), and the optimal cooperative control strategy of the multi-agent system constructed is
[0027] Furthermore, the edge event-based cooperative control strategy includes:
[0028] 201. A three-layer feedforward neural network is used to construct an evaluation neural network for online approximating the optimal cost function. The output value of the evaluation network will be used to calculate the cooperative control strategy. This output value is characterized by the weights of the evaluation network and the activation function φ i (e i ), as shown below:
[0029]
[0030] Among them, is the approximate cost function with respect to e i partial derivative;
[0031] 202. Construct a generalized policy iteration algorithm with an inner loop for policy evaluation and an outer loop for policy improvement. For agent i, based on the cooperative control policy obtained from the p-th policy improvement, the policy evaluation steps are as follows:
[0032]
[0033] where T0>0 is a fixed integration interval; is the neighborhood cooperation error at the edge trigger moment; p and q are the iteration indices for policy improvement and policy evaluation steps respectively, and V i p,q is the cost function related to the cooperative policy obtained from the q-th policy evaluation. Further, combined with the evaluation neural network structure, the equivalent form of the policy evaluation step is:
[0034]
[0035] 203. The learning error in the policy iteration process is expressed as follows:
[0036]
[0037] where is the utility function, is the weight of the evaluation network corresponding to the current cost function. Based on the normalized gradient descent algorithm, the weight update law of the evaluation network is:
[0038]
[0039] where α i is the learning rate of the evaluation network;
[0040] 204. Calculate and output the edge event-based cooperative control policy using the evaluation weight and the optimal control policy at the edge event trigger moment as follows:
[0041]
[0042] Furthermore, the process of establishing the edge event trigger rule includes:
[0043] 301. For the communication edge ι connecting agents i and j in a multi-agent system, define the edge state δ ι :
[0044]
[0045] Edge triggering time The corresponding edge state can be expressed as:
[0046]
[0047] 302. Establish the following Lyapunov function based on the local cost function;
[0048]
[0049] where V i * (e i ) is the optimal cost function, is the optimal cost function based on the triggering cooperation error, is the weight error of the evaluation network. This function L i (t) is a criterion that comprehensively considers stability, convergence, and event triggering.
[0050] 303. Calculate the stability condition according to the Lyapunov function and obtain the following adaptive event-triggering rule:
[0051]
[0052] where, λ min (R ii ) and λ max (R ii ) represent the minimum and maximum eigenvalues of the matrix R ii respectively. is the Lipschitz constant of the control strategy . σ ι and γ ι are two constant parameters. This triggering rule describes when to generate the required edge event triggering time, that is, to obtain the next triggering time based on the current triggering time The initial triggering time of all communication edges can be considered as The determination of the next triggering time of the communication edge ι of the agent depends on the control strategy of the head node of the current communication edge ι, that is, the agent i, and the edge state error
[0053] Finally, apply the obtained cooperative control strategy and edge event triggering rule to the multi-agent system:
[0054] 401. Load the cooperative control strategy and edge event triggering rules into the multi-agent system.
[0055] 402. Implement edge event-triggered intelligent cooperative control. Under the edge event triggering mechanism, each communication edge in the multi-agent independently calculates the triggering moment. At this time, each agent uses the local edge event trigger and evaluation network to execute the learning process of generalized policy iteration, realizes edge event-triggered intelligent cooperative control, and finally completes the state coordination task of the entire multi-agent system.
[0056] The present invention also provides a non-transitory computer-readable storage medium, on which computer programs and control instructions are stored. When the programs and instructions are executed by each agent, the multi-agent can implement the method described in the present disclosure.
[0057] The beneficial effects of the technical solution provided by the present invention are:
[0058] The cooperative control method of the multi-agent system based on edge event triggering disclosed by the present invention constructs an optimal cooperative control system model for the multi-intelligent system, designs an edge event-triggered intelligent cooperative control technology in combination with the adaptive dynamic programming technology based on the generalized policy iteration algorithm, and at the same time uses the edge event-triggered control protocol to realize data sampling and control update at the point-to-point level between nodes, effectively reducing the communication loss and computational burden of the multi-agent system.
[0059] Aiming at the cooperative consensus problem of the multi-agent system, the present invention uses edge event triggering and adaptive dynamic programming technology based on generalized policy iteration to study the edge event-triggered intelligent cooperative control method under the condition of limited communication resources, which meets the application requirements and development trends of related technologies. Through the retrieval of existing literature and technologies, no similar technical solutions have been found. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 is a flowchart of a cooperative control method for a multi-agent based on edge event triggering according to the present invention;
[0061] Figure 2 is a flowchart of step S10 shown according to an embodiment;
[0062] Figure 3 is a flowchart of step S20 shown according to an embodiment;
[0063] Figure 4 is a flowchart of step S30 shown according to an embodiment;
[0064] Figure 5 is a flowchart of step S40 shown according to an embodiment;
[0065] Figure 6 It is a communication topology structure diagram of a multi-agent system shown according to an embodiment;
[0066] Figure 7 It is the state trajectory of a linear multi-agent system shown according to an embodiment;
[0067] Figure 8 It is the control signal of each agent in a linear multi-agent system shown according to an embodiment;
[0068] Figure 9 It is the triggering moment of each communication edge in a linear multi-agent system shown according to an embodiment;
[0069] Figure 10 It is the state trajectory of a non-linear multi-agent system under the proposed method shown according to an embodiment;
[0070] Figure 11 It is a comparison diagram of the triggering moments of each communication edge in a non-linear multi-agent system shown according to an embodiment; Detailed implementation manners
[0071] To make the objectives, technical solutions and advantages of the present invention clearer, the following further describes in detail the embodiments of the present invention. It should be noted that: unless otherwise stated, the steps, mathematical expressions and numerical values described in these embodiments do not limit the scope of the present disclosure. At the same time, it should be understood that, for the sake of convenience of description, the dimensions of each part shown in the drawings are not drawn according to the actual proportional relationship.
[0072] The present invention provides a multi-agent cooperative control method based on edge event triggering. This method is an edge event-triggered learning control method based on saving communication and computing resources, and specifically includes the following steps:
[0073] First, establish the optimal cooperative control model of each agent A multi-agent system is a typical complex network system. The corresponding cooperative consensus problem aims to achieve the consistency of the states of all agents. In this process, each agent realizes the cooperative task through communication with its neighbor agents. The cooperative error of each agent is determined by the state of the agent itself and the neighbor agents communicating with it. Based on the form of the cooperative error, in one embodiment, the system dynamics of a single agent is established. Further, based on the form of the local neighborhood cooperative error, the neighborhood cooperative error system dynamics is constructed, and then the cooperative control model of the multi-agent system is established. From the perspective of each agent, considering the convergence situation of the cooperative error and the control cost of the neighbor agents including the agent itself, the cost function of the agent is set up. Each agent makes a cooperative control decision by minimizing the designed cost function, and according to the Bellman optimality principle, the optimal cooperative control strategy of the multi-agent cooperative control model is obtained.
[0074] According to the collaborative control requirements of multiple agents, the quadratic cost function V i (e i (t)) can be selected, and then the required optimal collaborative control strategy can be obtained For different control requirements, the cost function needs to be specifically changed. For example, considering the case of control input saturation, the relevant limiting function can be added to the control cost term of the cost function.
[0075] The present invention constructs an evaluation neural network and uses the generalized policy iteration algorithm to obtain a multi-agent collaborative control strategy based on the edge event trigger mechanism For the multi-agent control model, an evaluation neural network is constructed based on a multi-layer feedforward neural network to approximate the cost function of the iterative learning process. At the same time, a generalized policy iteration algorithm including an inner loop for policy evaluation and an outer loop for policy improvement is constructed to iteratively solve the approximate optimal value of the cost function In the generalized policy iteration algorithm, the collaborative control strategy is given by the policy improvement step. In order to obtain the cost function corresponding to the current collaborative control strategy, it is iteratively solved through multiple policy evaluation steps. More specifically, in each policy evaluation step, the cost function corresponding to the current collaborative control strategy is related to the reward accumulation over a fixed time interval and the function value of the value function obtained from the previous policy evaluation after a fixed time. After multiple policy evaluation steps, the obtained value function is used in the policy improvement step to obtain the next collaborative control strategy. Based on the policy evaluation step, the normalized gradient descent algorithm is used to obtain the weight adaptive update law of the evaluation network. Further, the output value of the evaluation network is further used to calculate the collaborative control strategy. Combining the weights of the evaluation network and the optimal collaborative control strategy, the edge-based collaborative control strategy is obtained at the edge event trigger moment This collaborative control strategy is the actual decision result of each agent, that is, the actual output of the agent controller.
[0076] The adaptive dynamic programming technology based on the generalized policy iteration algorithm designed by the present invention can also be implemented using a behavior-evaluation network structure. At this time, the control strategy is directly given by the behavior network instead of being indirectly calculated by the evaluation network; in comparison, the single evaluation network structure is easier to implement and has a smaller computational burden.
[0077] The present invention obtains an edge event trigger rule established on the communication edges of agents. Under the edge event trigger mechanism, for each communication edge in the multi-agent system, a corresponding edge state is defined, and a local event trigger is equipped. The edge state sampling and data transmission between agents are independently carried out on each edge. According to the edge state δ recorded by the local edge event trigger ι, calculate the edge state error. When the edge state error satisfies the corresponding triggering rule, the edge event is triggered, and the corresponding communication edge resumes communication, thereby updating the local neighborhood error and the agent controller. To ensure the optimality and stability of the edge triggering moment, a Lyapunov function based on the local cost function is established. The derivative of the Lyapunov function is obtained to derive the stability condition, and the adaptive edge event triggering rule is inversely derived according to the stability condition. This triggering rule describes when to generate the required edge event triggering moment.
[0078] The present invention applies the obtained cooperative control strategy and edge event triggering rule to a multi-agent system. The cooperative control strategy and edge event triggering rule are loaded into the multi-agent system. Among them, each communication edge independently controls the communication process of the two connected agents according to the edge event generated by the local edge event trigger, and then updates the local neighborhood cooperation error and the cooperative control strategy of the agent. Each agent uses the corresponding evaluation network to execute the iterative learning process based on the generalized policy iteration algorithm to achieve intelligent cooperative control based on edge event triggering.
[0079] The present invention also provides a non-transitory computer-readable storage medium, on which computer programs and control instructions are stored. When the programs and instructions are executed by each controller, the multi-controller system can implement the method described in the present disclosure.
[0080] Embodiment 1
[0081] The embodiment of the present invention aims at the cooperative control problem of multi-agents, based on the generalized policy iteration algorithm and the edge event triggering mechanism. Refer to Figure 1 , and the method includes the following steps:
[0082] S10: Establish a multi-agent optimal cooperative control model;
[0083] Among them, multi-agents is a typical complex network system. The corresponding cooperative consensus problem aims to achieve the consistency of the states of all agents. In this process, each agent realizes the cooperative task through communication with its neighbor agents. The local neighborhood cooperation error of each agent is determined by the state of the agent and its neighbor agents with which it communicates. In one embodiment, first establish the system dynamics of a single agent Further, based on the form of the local neighborhood cooperation error, construct the system dynamics of the neighborhood cooperation error, and then establish a multi-agent cooperative control model The multi-agents are based on a leader-follower structure, where i = 0 represents the leader and i = 1, 2,..., N are followers, and u i is the cooperative control strategy of the followers. From the perspective of each agent, considering the convergence of the cooperation error and the control cost of the neighbor agents including the agent, the cost function V of the agent is established i (ei )。 Each agent makes collaborative control decisions by minimizing the cost function V i (e i ), so as to achieve the convergence of the local domain collaborative error while minimizing the control cost, and finally realize the collaboration of the entire multi-agent system. The optimal cost function V i * (e i ) is the solution of the corresponding Hamilton-Jacobi-Bellman equation. According to the Bellman optimality principle, the optimal collaborative control strategy of the multi-agent collaborative control model is obtained
[0084] S20: Construct an evaluation neural network and use the generalized policy iteration algorithm to obtain a collaborative control strategy based on edge events;
[0085] Among them, for the multi-agent system control model, an evaluation neural network is constructed based on a multi-layer feedforward neural network to approximate the cost function of the iterative learning process Characterized by the weights of the evaluation network And the activation function At the same time, a generalized policy iteration algorithm including a policy evaluation inner loop and a policy improvement outer loop structure is constructed to iteratively solve the approximate optimal value of the cost function. In the generalized policy iteration algorithm, the collaborative control strategy Is given by the policy improvement step. In order to obtain the cost function corresponding to the current collaborative control strategy Iteratively solve through multiple policy evaluation steps. More specifically, in each policy evaluation step, the cost function corresponding to the current collaborative control strategy And the reward accumulation over a fixed time interval And the function value of the value function obtained from the previous policy evaluation after a fixed time Related. After multiple policy evaluation steps, the obtained value function Is used for the policy improvement step to obtain the next collaborative control strategy Based on the policy evaluation step, the weight adaptive update law of the evaluation network is obtained using the normalized gradient descent algorithm Furthermore, the output value of the evaluation network is further used to calculate the collaborative control strategy. Combining the weights of the evaluation network And the optimal collaborative control strategy At the edge event trigger moment Obtain a collaborative control strategy based on edge events This collaborative control strategy is the actual decision result of each agent, that is, the actual output of the agent controller.
[0086] S30: Obtain an adaptive edge event trigger rule;
[0087] Under the edge event triggering mechanism, for each communication edge in the multi-agent system, a corresponding edge state δ is defined ι = x i (t) - x j (t). At the same time, a local event trigger is equipped to independently sample the edge state and transfer data between agents on each edge. According to the edge state recorded by the local edge event trigger Calculate the edge state error When the edge state error satisfies the corresponding triggering rule, the edge event is triggered, and the corresponding communication edge resumes communication, thereby updating the local neighborhood error and the agent controller. To ensure the optimality and stability of the edge triggering moment, a Lyapunov function L i (t) is established. This function comprehensively considers the cost function, the evaluation network weight error, and the influence of edge event triggering, and is a comprehensive evaluation criterion for stability, convergence, and edge event triggering. Derive the stability condition for the Lyapunov function L i (t), and inversely deduce the adaptive edge event triggering rule according to the stability condition. This triggering rule describes when to generate the required edge event triggering moment
[0088] S40: Apply the obtained cooperative control strategy and edge event triggering rule to the multi-agent system;
[0089] Among them, load the cooperative control strategy and edge event triggering rule into the multi-agent system. Among them, each communication edge independently controls the communication process of the two connected agents according to the edge event generated by the local edge event trigger, that is, at the edge event triggering moment Resume the communication of the corresponding edge, and then update the local neighborhood cooperation error of the corresponding agent And the corresponding cooperative control strategy Each agent uses the corresponding evaluation network to execute an iterative learning process based on the generalized policy iteration algorithm to achieve intelligent cooperative control based on edge event triggering, and finally complete the consistency of all agent states.
[0090] Embodiment 2
[0091] Next, in combination with specific calculation formulas, Embodiment 2 further introduces the solution in Embodiment 1, as detailed below:
[0092] First, establish an optimal cooperative control model through Figure 1 the steps in S10.
[0093] S10: Establish an optimal cooperative control model for the multi-agent system;
[0094] In this embodiment, it can be throughFigure 2 The steps in
[0095] Step S10 mainly includes establishing an optimal control model:
[0096]
[0097] where f i (x i ) is the system internal dynamic information of agent i, g i (x i ) is the control input matrix of agent i, u i is the control strategy. N represents the number of agents. Note that the multi - agents involved in the present invention have a leader - follower structure. When i = 0, it means that the agent is the leader, and the rest are followers; the ultimate goal of multi - agent cooperative control based on the leader - follower structure is to make the states of all followers consistent with the state of the leader.
[0098] Step S102, designing a multi - agent cooperative control model. In a multi - agent system, each agent can only communicate with its neighbor agents. During the cooperation process, each agent obtains the corresponding local neighborhood cooperation error by communicating with its neighbor agents, and then adjusts the control strategy to make the local neighborhood cooperation error converge, ultimately achieving the cooperation of the entire multi - agent system. The cooperative control model is established from the perspective of each agent according to the following formula:
[0099]
[0100] where x0 is the leader state. For agent i, the local neighborhood error, F i is the internal dynamic information of the error system, a ij is the element of the adjacency matrix, b i is the element of the reachability matrix; l ii is the element of the Laplacian matrix.
[0101] Step S103, designing the cost function of each agent. In order to achieve a good cooperation effect with as small a control cost as possible, considering the convergence of the cooperation error and the control cost of the neighbor agents including this agent, the cost function V i (e i ) of the agent is established:
[0102]
[0103] Among them, Q and R ii and R ij are positive definite and symmetric constant matrices used to adjust the components of the system operation cost and the control cost in the cost function. More specifically, R ii and R ij respectively represent the proportions of the control costs of the current agent i and the neighbor agent j in the cost function. N i is the set of neighbor agents of agent i. It can be seen that this cost function includes the control costs of all neighbor agents. By minimizing this cost function, each agent can ensure the convergence of the local neighborhood error e i while minimizing the control cost, and finally achieve the collaborative task of the entire multi-agent system.
[0104] Step S104: Obtain the optimal collaborative control strategy. During the collaboration process, each agent expects to minimize its own cost function, that is, to obtain the optimal cost function V i * (e i ), so as to ensure that the entire multi-agent system achieves the fastest collaborative effect with the minimum control cost. Based on the optimal cost function, according to the Bellman optimality principle, the optimal collaborative control strategies of each agent are as follows:
[0105]
[0106] Among them, is the partial derivative of the optimal cost function V i * (e i (t)) with respect to e i (t). The constructed optimal collaborative control strategy of the multi-agent system is
[0107] After obtaining the optimal control model, the evaluation neural network can be continuously constructed through the steps in Figure 1 to approximate the optimal cost function V i * (e i ).
[0108] S20: Construct an evaluation neural network and use the generalized policy iteration algorithm to obtain an edge-event-based collaborative control strategy;
[0109] In this embodiment, the evaluation neural network can be constructed through the steps in Figure 3 and the edge-event-based collaborative control strategy can be obtained. As shown in Figure 3 , step S20 mainly includes:
[0110] Step S201: Construct an evaluation neural network. In this embodiment, a three-layer feedforward neural network is used to construct the evaluation neural network for online approximating the optimal cost function. The output value of the evaluation network will be used to calculate the cooperative control strategy, and this output value is determined by the evaluation weights and the activation function φ i (e i ) as shown below:
[0111]
[0112] where is the approximate cost function with respect to the partial derivative of e i .
[0113] Step S202: Construct a generalized policy iteration algorithm to solve the approximate optimal value of the cost function. To solve the optimal cost function of the agent, a generalized policy iteration algorithm with a policy evaluation inner loop and a policy improvement outer loop structure is constructed. In the generalized policy iteration algorithm, there are two iteration indices p and q, corresponding to the policy improvement outer loop and the policy evaluation inner loop respectively. The cooperative control strategy is given by the policy improvement step. To obtain the cost function corresponding to the current cooperative control strategy it is iteratively solved through multiple policy evaluation steps. More specifically, in each policy evaluation step, the cost function corresponding to the current cooperative control strategy is related to the reward accumulation over a fixed time interval and the function value of the value function obtained from the previous policy evaluation after a fixed time as shown by the following formula:
[0114]
[0115] When the policy evaluation step is executed q max times, the cost function corresponding to the current control strategy is expressed as
[0116] Furthermore, the corresponding cost function is applied to the policy improvement step to obtain the cooperative control strategy for the next policy evaluation loop, expressed as
[0117]
[0118] Based on the constructed evaluation neural network structure, the cost function in the iterative process is expressed as i.e., the cost function related to the cooperative policy obtained in the q-th policy evaluation step The cost function. Therefore, the policy evaluation step in the generalized policy iteration algorithm based on the evaluation neural network approximation can be expressed as:
[0119]
[0120] The corresponding policy improvement step is expressed as:
[0121]
[0122] The generalized policy iteration algorithm continuously updates the cooperative control policy and the corresponding cost function through the double-loop structure of policy evaluation and policy improvement, and finally obtains an approximation of the optimal cost function.
[0123] Step S203, obtain the adaptive update rule of the evaluation network. Based on formula (8), the learning error in the policy iteration process is expressed as follows:
[0124]
[0125]
[0126] It can be regarded as a constant in the previous policy evaluation process. Based on the normalized gradient descent algorithm, the weight update law of the evaluation network is obtained as:
[0127]
[0128] where α i is the learning rate of the evaluation network of the i-th agent, is used for normalization.
[0129] Step S204, obtain the cooperative control policy based on edge events. Using the evaluation network weights and the optimal cooperative control policy shown in formula (4) at the edge event trigger moment calculate and output the cooperative control policy based on edge events
[0130]
[0131] As can be seen from formula (12), the cooperative control policy of agent i is updated at the edge event trigger moment corresponding to the communication edge ι. This cooperative control policy is the actual decision result of each agent in the multi-agent system under the edge event trigger mechanism, and it is also the actual output of each agent controller.
[0132] The cooperative control policy based on edge events is obtained above, but the edge event trigger moment is still unavailable, so continue throughFigure 1 The adaptive edge event triggering rule is obtained in step S30 in , and this rule will be used to calculate the edge event triggering moments corresponding to each communication edge in the multi-agent system.
[0133] S30: Obtain the adaptive edge event triggering rule;
[0134] In this embodiment, the adaptive event triggering rule can be obtained through the steps in Figure 4 As shown in , step S30 mainly includes: Figure 4 As shown in , step S30 mainly includes:
[0135] Step S301, calculate the local sampling error. Compared with the traditional agent-based event triggering mechanism, the present invention constructs an event triggering mechanism based on communication edges, and each communication edge independently judges the event triggering condition and realizes intermittent communication between adjacent agents. Under the edge event triggering mechanism, the following corresponding edge states are defined for each communication edge in the multi-agent system:
[0136]
[0137] Agents i and j are the head node and the tail node of communication edge ι respectively. Corresponding local edge event triggers are equipped on each communication edge, and edge state sampling is independently performed on each edge. According to the triggering moment recorded by the edge event trigger The corresponding edge state is:
[0138]
[0139] According to formulas (13) and (14), the local sampling error is obtained as:
[0140]
[0141] The local sampling error of this communication edge will be used to construct the edge event triggering rule.
[0142] Step S302, establish a Lyapunov function based on the local cost function. To ensure the optimality and stability of the edge triggering moment, a Lyapunov function L i (t) is as follows:
[0143]
[0144] where V i * (e i ) is the optimal cost function, is the optimal cost function based on the triggering cooperation error, is the weight error of the evaluation network. This function L i(t) is a comprehensive evaluation criterion that combines stability, convergence, and event triggering; taking the derivative of this function, according to Lyapunov stability theory, the stability condition that needs to be satisfied is According to this condition, the corresponding edge event triggering rule can be obtained.
[0145] Step S303, obtaining the adaptive edge event triggering rule. Through reverse design based on the above stability condition, the adaptive edge event triggering rule is as follows:
[0146]
[0147] where λ min (R ii ) and λ max (R ii ) respectively represent the minimum and maximum eigenvalues of the matrix R ii . is the Lipschitz constant of the control strategy . σ ι and γ ι are two constant parameters. This triggering rule describes when to generate the required edge event triggering moments, that is, obtaining the next triggering moment based on the current triggering moment The initial triggering moments of all communication edges can be considered as The determination of the next triggering moment of the communication edge ι of the agent depends on the head node of the current communication edge ι, that is, the control strategy of agent i and the edge state error When the edge state error of the communication edge ι satisfies the edge event triggering rule described in formula (16), the two agents connected by it resume communication, and the corresponding head node, that is, the local neighborhood cooperation error of agent i is updated, expressed as:
[0148]
[0149] At the same time, the cooperative control strategy of agent i is updated based on formula (2).
[0150] After obtaining the cooperative control strategy and the edge event triggering rule, they can be continued to be applied to the multi-agent system through Figure 1 the steps in.
[0151] S40: Apply the obtained cooperative control strategy and edge event triggering rule to the multi-agent system;
[0152] In this embodiment, the cooperative control strategy and the edge event triggering rule can be applied to the multi-agent system through Figure 5 the steps in, such as Figure 5As shown, step S40 mainly includes:
[0153] Step S401: Load the cooperative control strategy and edge event triggering rules into the multi-agent system. In some embodiments, the cooperative control strategy and edge event triggering rules can be loaded into the multi-agent system in the form of computer programs and mathematical symbols. At this time, a series of edge triggering moments are formed for each communication edge in the multi-agent system. Each time a new edge triggering moment is formed the communication edge resumes communication, and the cooperative control strategy of the head node agent connected thereto is updated accordingly. Furthermore, a new round of update iteration and learning calculation is completed.
[0154] Step S402: Implement edge event-triggered intelligent cooperative control. Under the edge event triggering mechanism, each communication edge in the multi-agent system independently calculates the triggering moment, and the edge triggering moments of each communication edge are different, that is ι,χ are different communication edges. And when a new edge triggering moment appears for any communication edge with agent i as the head node, the local neighborhood error of the agent and the corresponding cooperative control strategy are updated. At this time, each agent uses the local edge event trigger and the evaluation network to execute the learning process of generalized policy iteration, realizes edge event-triggered intelligent cooperative control, and finally completes the state coordination task of the multi-agent system.
[0155] Embodiment 3
[0156] Next, the feasibility of the solutions in Embodiments 1 and 2 is verified by combining specific experimental data and examples. In this embodiment, a linear continuous-time multi-agent system is considered, as described in detail below:
[0157] First, Figure 6 shows the communication topology structure of the multi-agent system in this embodiment. It can be seen that the cooperative control method of the multi-agent system based on edge event triggering designed by the present invention is applied to a multi-agent system composed of a single leader and three followers. The corresponding Laplacian matrix L and reachability matrix B are expressed as:
[0158]
[0159] Based on Figure 6 the linear continuous-time multi-agent system shown, the designed cooperative control method finally makes the states of all followers consistent with the leader. Then, continue to execute according to Figure 1 the steps S10 - S40 shown.
[0160] According to step S10, first establish the system dynamics of a single agent. The linear multi-agent dynamics are expressed as:
[0161]
[0162] Among them, the leader control quantity is u0 = 0, and the initial state of the leader is x0(0) = [-0.4, 0.6] T , and the initial states of each follower are set as x1(0) = [0.7, -0.3] T , x2(0) = [-1, 0.6] T , x3(0) = [-0.3, -0.9] T . Subsequently, a cooperative control model is established according to formula (2). For all follower agents, the configuration adopted by the cost function is: Q = diag([10, 10]), R ii = 5, i = 1, 2, 3 and R ij = 1, j ∈ N i .
[0163] According to step S20, in this embodiment, 2 nodes are configured in the hidden layer of the evaluation network, and the activation function is set as
[0164]
[0165] During the update process of the evaluation network, the learning rate of the evaluation network is [α i = [0.03, 0.02, 0.01], and the initial weights of the neural network are randomly selected between -1 and 1. The maximum values of the two iteration indexes of the generalized policy iteration algorithm are p max = 5 and q max = 3, and the integration interval of the policy evaluation process is T0 = 0.1s.
[0166] According to step S30, for all communication edges ι = 1, 2, 3, 4 in the multi-agent system, independent edge event triggers are designed respectively, and the parameter configuration of the corresponding adaptive edge event triggering rule is: the Lipschitz constant of the optimal cooperative control strategy of each agent is The amplitudes and decay rates of the exponential terms are [σ ι = [5, 5, 5, 5] and [γ ι = [0.5, 0.5, 0.25, 0.25].
[0167] According to step S40, the linear continuous-time multi-agent system calculates the edge event-based cooperative control strategy according to formula (12) and calculates the new edge event trigger time according to formula (17). Each follower agent realizes the iterative learning process based on the generalized policy iteration algorithm by constructing an evaluation network, and realizes event-triggered control through designing an edge event triggering mechanism, thereby completing edge event-triggered intelligent cooperative control.
[0168] Figure 7is the state trajectory of the linear continuous-time multi-agent system shown in this embodiment, where agent0 is the leader agent and agent1 - agent3 are three follower agents; it can be seen that the states of all followers and the leader's state eventually reach agreement. Therefore, the multi-agent cooperative control method designed by the present invention based on edge event triggering has a good cooperative effect. Figure 8 are the control signals of each follower in the linear continuous-time multi-agent system shown in this embodiment. It can be seen that the control signals of the follower agents generally show a "stepped" shape, indicating that the cooperative control strategy of the agents is affected by the edge event triggering mechanism, making its update process have an aperiodic characteristic. Figure 9 are the triggering moments of each communication edge in the linear continuous-time multi-agent system shown in this embodiment. It can be seen that compared with the 1000 triggering times of traditional time-triggered control, the edge event triggering times of communication edges 1 to 4 are 41, 61, 70, and 66 times respectively, that is, the triggering rates are 4.1%, 6.1%, 7.0%, and 6.6% respectively. Based on the edge event triggering mechanism, the communication edges of the multi-agent system in the method of the present invention each reduce the communication resource loss by 95.9%, 93.9%, 93.0%, and 93.3%.
[0169] Embodiment 4
[0170] Next, on the basis of Embodiment 3, considering two classic edge event triggering control rules, the advantages of the adaptive edge event triggering rule designed by the present invention in ensuring the cooperative control effect of the multi-agent system are verified by comparison. In this embodiment, the application object is a continuous-time nonlinear multi-agent system, and the specific implementation steps are described in detail below:
[0171] First, the communication topology structure of the multi-agent system in this embodiment is the same as that in Embodiment 3, as shown in Figure 6 , and the relevant matrices are the same as those in Embodiment 3. Then, continue to execute according to the steps S10 - S40 shown in Figure 1 .
[0172] According to step S10, first establish the system dynamics of a single agent. The nonlinear multi-agent dynamics are expressed as:
[0173]
[0174] where the initial states of each follower are set as x1(0) = [0.6, 0.4] T , x2(0) = [-1, -1] T , x3(0) = [-0.2, 0.8] T . Subsequently, establish a cooperative control model according to formula (2). The configuration of the cost function of each agent is the same as that in Embodiment 3. In addition, the leader dynamics are:
[0175]
[0176] Among them, the initial state of the leader is x0(0) = [-0.5, 0.5] T .
[0177] According to step S20, in this embodiment, 3 nodes are configured in the hidden layer of the evaluation network, and the activation function is set to
[0178]
[0179] During the update of the evaluation network, the learning rate and initial weights of the evaluation network are set the same as in Embodiment 3. The maximum values of the two iteration indexes of the generalized policy iteration algorithm are p max = 5 and q max = 2 respectively, and the integration interval of the policy evaluation process is T0 = 0.1s.
[0180] According to step S30, for all communication edges ι = 1, 2, 3, 4 in the multi-agent system, independent edge event triggers are designed respectively, and the parameter configurations of the corresponding adaptive edge event triggering rules are as follows: The Lipschitz constant of the optimal cooperative control strategy of each agent is The amplitudes and decay rates of the exponential terms are [σ ι = [2, 2, 2, 2] and [γ ι = [3, 3, 3, 3] respectively.
[0181] According to step S40, the nonlinear continuous-time multi-agent system calculates the edge-event-based cooperative control strategy according to formula (11), and calculates the new edge-event trigger time according to formula (16). Each follower agent realizes the iterative learning process based on the generalized policy iteration algorithm by constructing an evaluation network, and realizes event-triggered control by designing an edge-event trigger mechanism, thereby completing the edge-event-triggered intelligent cooperative control. In addition, considering two classic edge-event trigger rules, denoted as "Classic Method 1" and "Classic Method 2", the corresponding trigger rules are expressed as:
[0182]
[0183] and
[0184]
[0185] Figure 10 represents the state trajectory of the continuous-time nonlinear multi-agent system under the designed cooperative control method shown in this embodiment. It can be seen that the states of all followers and the state of the leader finally reach consistency. Figure 11Indicates the triggering situations of each communication edge under different edge event triggering rules. Table 1 shows the comparison of the cooperative control effects of the multi-agent system under the designed adaptive edge event triggering rule, "Classical Method 1", and "Classical Method 2", respectively in terms of the adjustment time
[0186] It can be clearly seen that the adaptive edge event triggering rule designed in the present invention can achieve better cooperative control effects while ensuring good communication efficiency.
[0187] Table 1 Comparison effects of different methods
[0188]
[0189] Combined with the results of Example 3 and Example 4, the beneficial effects of the edge event-triggered intelligent cooperative control method for saving communication and computing resources disclosed in the embodiments of the present invention are true.
[0190] Example 5
[0191] This embodiment provides a non-transitory computer-readable storage medium (including but not limited to disk memory, optical memory, etc.), on which are stored:
[0192] A computer program, including the adaptive update rule of the evaluation network, the adaptive edge event triggering rule, and the relevant computer code for method implementation.
[0193] Control instructions, mainly the control strategies of each agent.
[0194] When the program and instructions are executed by each agent, the multi-agent system can implement the method described in the first aspect of the present disclosure, which will not be elaborated here.
[0195] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as methods, prototype systems, or computer program products. Therefore, the present disclosure can take the form of a complete software embodiment or an embodiment combining software and hardware.
[0196] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-agent collaborative control method based on edge event triggering, characterized in that: The method comprises the following steps: Establish the optimal cooperative control model for each agent based on multi-agent; The optimal cooperative control strategy for each agent is obtained by calculating the optimal cooperative control model through the Bellman optimal method and the cost function of each agent. Construct an evaluation neural network and use the generalized policy iteration algorithm to find the optimal cooperative control strategy Triggering time of edge event Control, output collaborative control strategy based on edge events Construct a Lyapunov function based on the local cost function and establish the edge event triggering rules of multi-agents; The edge event-based collaborative control strategy and the edge event triggering rules are applied to a multi-agent system.
2. According to the multi-agent collaborative control method based on edge event triggering according to claim 1, it is characterized in that: The optimal cooperative control strategy The output process includes:
101. Establish a single agent model according to the following formula: Where: f i (x i ) is the internal dynamic information of the system of agent i, g i (x i ) is the control input matrix of agent i, u i is the control strategy, N represents the number of agents, when i=0, it means that the agent is the leader, and the rest are followers; 102. Establish the optimal collaborative control model for each agent according to the following formula: Among them: For agent i, e i is the local neighborhood error, F i is the internal dynamic information of the error system, a ij is the adjacency matrix element, b i is the reachable matrix element; l ii is the Laplace matrix element; 103. Calculate the quadratic cost function V of each agent's control cost setting according to the following formula: i (e i (t)): Where: Q, R ii and R ij is a positive definite and symmetric constant matrix used to adjust the components of system operation cost and control cost in the cost function, N i is the set of neighboring agents of agent i; 104. According to the Bellman optimal method, set the quadratic cost function V of the control cost for each agent according to the following formula i (e i (t)) Calculate the optimal collaborative control strategy for each agent in: is the optimal cost function V i * (e i (t)) About e i The partial derivative of (t); 105. According to the optimal collaborative control strategy of each intelligent agent The optimal cooperative control strategy of the constructed multi-agent is 3. According to the multi-agent collaborative control method based on edge event triggering according to claim 1, it is characterized in that: The collaborative control strategy output process based on edge events includes:
201. Use a three-layer feedforward neural network to construct an evaluation neural network to approximate the optimal cost function online and evaluate the output value of the network It will be used to calculate the collaborative control strategy. The output value is determined by the evaluation network weights. and the activation function φ i (e i ) is represented as follows: in: is an approximate cost function About e i The partial derivative of 202. Construct a generalized policy iteration algorithm with an inner loop of strategy evaluation and an outer loop of strategy improvement. For agent i, the collaborative control strategy obtained based on the p-th strategy improvement The strategy evaluation steps are as follows: Where: T0>0 is a fixed integration interval; is the neighborhood coordination error at the edge triggering moment; p and q are the iteration indexes of the strategy improvement and strategy evaluation steps, respectively, V i p,q is the collaborative strategy obtained from the qth strategy evaluation Cost function; further, combined with the evaluation neural network structure, the equivalent form of the strategy evaluation step is obtained:
203. The learning error during the policy iteration process is expressed as follows: in: is the utility function, is the evaluation network weight corresponding to the current cost function; based on the normalized gradient descent algorithm, the evaluation network weight update law is obtained as: Where: α i To evaluate the network learning rate; 204. Using Evaluation Weights and optimal control strategy At the time when the edge event is triggered Compute and output the coordinated control strategy based on edge events as follows:
4. According to the multi-agent collaborative control method based on edge event triggering according to claim 1, it is characterized in that: The edge event triggering rule establishment process:
301. Based on the communication edge ι connecting agents i and j in the multi-agent system, define the edge state δ ι : Furthermore, the edge triggering moment The corresponding edge state is expressed as:
302. Establish the following Lyapunov function based on the local cost function; Where: V i * (e i ) is the optimal cost function, is the optimal cost function based on the triggering coordination error, is the weight error of the evaluation network. This function L i (t) is a criterion that combines stability, convergence, and event triggering; 303. Calculate the stability condition based on the Lyapunov function and obtain the following adaptive event triggering rules: in: λ min (R ii ) and λ max (R ii ) represent the matrix R ii The minimum and maximum eigenvalues of . For control strategy The Lipschitz constant; σ ι and γ ι are two constant parameters; the trigger rule describes when to generate the required edge event trigger time, that is, at the current trigger time Get the next trigger time based on The initial triggering time of all communication edges can be considered as The next triggering time of the agent communication edge ι is determined by the head node of the current communication edge ι, that is, the control strategy of agent i and edge state error 5. According to the multi-agent collaborative control method based on edge event triggering according to claim 1, it is characterized in that: Applying the edge event-based collaborative control strategy and the edge event triggering rule to a multi-agent system process includes:
401. Loading the edge event-based collaborative control strategy and edge event triggering rules into the multi-agent; 402. Implement edge event-triggered intelligent collaborative control. Under the edge event triggering mechanism, each communicating edge in the multi-agent independently calculates the triggering time. At this time, each agent uses the local edge event trigger and the evaluation network to perform the learning process of generalized strategy iteration to implement edge event-triggered intelligent collaborative control and complete the state collaborative task of the entire multi-agent.
6. A computer-readable storage medium, characterized in that: Computer programs and control instructions are stored thereon, and when the programs and instructions are executed by each agent, the multi-agent system can implement the processes of claims 1 to 5.
Citation Information
Cited By
Heterogeneous multi-unmanned aerial vehicle system tracking control method based on event triggering
CN121070048A
Multi-agent cooperative control method based on hybrid strategy iteration
CN121277070A
A multi-agent cooperative control method based on hybrid policy iteration
CN121277070B
Multi-node single-chip microcomputer automatic synchronization control method based on edge collaboration
CN121477736A
Multi-node single-chip microcomputer automatic synchronization control method based on edge collaboration
CN121477736B