An Optimal Consensus Cooperative Control Method for Intelligent Unmanned Cluster Systems under the Influence of Multiple Time Delays
Through the combination of distributed control protocol and neural network, the problem of multi-delay impact in multi-agent systems is solved, and stable and consistent collaborative control under model-free conditions is achieved, reducing control costs and improving training stability.
Patent Information
- Application Number
- CN202211452513.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-11-21
AI Technical Summary
The prior art is difficult to effectively solve the situation where multiple delays exist simultaneously in multi-agent systems, especially under the condition of no model, which leads to system performance degradation or instability.
The distributed control protocol is adopted, and the Actor network and Critic network combine experience pool to realize the state error calculation of agents and network weight updates, and the state information and control information are used for optimal consistency coordinated control under multi-time delay.
The stable consistency of multi-agent systems is achieved under the condition of no model, reducing control costs, improving data utilization and stability of neural network training, and adapting to actual complex environments.
Smart Images

Figure CN115793448B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent unmanned cluster system control, and particularly relates to an optimal consensus cooperative control method for an intelligent unmanned cluster system under the influence of multiple time delays. Background Art
[0002] In recent years, inspired by the collective behavior of biological groups in nature, experts and scholars have applied the consensus of multi-agent systems (MASs) to the cooperative control of complex systems. The consensus problem of multi-agent systems has important application prospects in fields such as swarm control, sensor networks, distributed computing, and group decision-making. The optimal consensus problem is one of the most interesting research topics in multi-agent systems, and its purpose is to design a distributed protocol that enables all agents to achieve synchronization with the minimum energy consumption.
[0003] As a very important branch of machine learning, reinforcement learning (RL) is a very effective controller design method. Its core idea is to reinforce or encourage some actions of an agent by giving the agent certain observations and rewards from the environment, so as to generate the results or goals expected by the experimenter with a higher probability. Policy gradient (PG) is one of the most practical RL algorithms. Its main idea is to model the policy gradient and then update the control policy using the gradient descent method. PG can better adapt to high-dimensional and continuous spaces and actions. Algorithms based on PG have been used in many situations and have played a significant role in the field of multi-agent optimal control.
[0004] Due to the communication limitations between agents and the differences in the computing capabilities of the agents themselves, time delays are often inevitable in practical multi-agent systems, and these time delays can lead to performance degradation or instability. The existing time delay effects on multi-agent systems mainly fall into three categories: input time delay, output time delay, and state time delay. At present, most of the research on the influence of time delays on agents is based on the situation where a certain time delay exists alone, but these three time delays are not independent of each other, and there are many situations in real life where multiple time delays exist simultaneously. Therefore, it is of great significance to study the situation where multiple time delays exist simultaneously.
[0005] Most of the above-mentioned research works are based on single agents, and most of the research on time delays requires knowledge of the precise system model. However, in reality, communication between multi-agents is very common, and it is quite difficult to obtain an accurate system model. Therefore, how to achieve the consensus cooperative control of a multi-agent system under the influence of multiple time delays by a model-free method is an urgent problem to be solved. Summary of the Invention
[0006] Aiming at the deficiencies of the existing technology, the present invention proposes an optimal consensus cooperative control method for an intelligent unmanned cluster system under the influence of multiple time delays. The method includes:
[0007] S1: Obtain the interaction information of the agents in the intelligent unmanned cluster system, and determine the topological structure of the system according to the interaction information;
[0008] S2: Set the initial weight values of the Actor network and the Critic network in the agent;
[0009] S3: The agent sends its own state information to its neighbor agents, and the agent calculates the state error of the agent according to the state information of the neighbor agents and its own state information; wherein, the state information includes state variables and control information;
[0010] S4: Both the Critic network and the Actor network update the network weights at the current moment according to the state error of the agent at the previous moment and the network weights; the Actor network updates the control information according to the network weights of the Actor network at the current moment and the state error of the agent at the previous moment;
[0011] S5: Determine whether the intelligent unmanned cluster system reaches consensus according to the weights of the Critic network. If it reaches consensus, control the agent with the control information of the current Critic network; otherwise, return to step S3.
[0012] Preferably, the state information and control information of the agent are expressed as:
[0013]
[0014] Wherein, represents the state information of the i-th agent at the k-th moment, represents the state variable of the i-th agent from the moment to the k-th moment, represents the control input of the i-th agent from the moment to the k-1-th moment, represents the maximum state time delay, represents the maximum input time delay.
[0015] Preferably, the formula for calculating the state error of the agent is:
[0016]
[0017] Wherein, e i (k) represents the new state information of the i-th agent under the consistency error; represents the neighbor nodes of agent i; a ijDenote the connection weight from agent j to agent i; b i Denote the pinning control on agent i; Denote the state information of the leader at time k.
[0018] Preferably, the process of the Critic network updating the network weight at the next moment includes:
[0019] The Critic network calculates the performance index at the current moment and the performance index at the next moment according to the Critic network weight at the current moment and the agent state error, calculates the first temporal error according to the performance index at the current moment, the performance index at the next moment, the state information of the agent at the current moment and the agent state error at the current moment; calculates the first objective function according to the first temporal error; calculates the network weight at the next moment according to the first objective function and the network weight at the current moment.
[0020] Further, the formula for calculating the performance index is:
[0021]
[0022] where, V(e i (k)) is the performance index; Denote the Critic network weight on agent i at the l - th iteration; s c,i (k) denotes the input of the Critic network on the i - th agent; σ(·) denotes the activation function on the nodes in the network.
[0023] Further, the formula for calculating the first temporal error is:
[0024] κ c,i =V i (e i (k)) - r(e i (k), u i (k), u j (k)) - αV i (e i (k + 1))
[0025] where, κ c,i denotes the first temporal error at the current moment, V i (e i (k)) denotes the performance index of the i - th agent at time k; r(·) is the utility function, u i (k) denotes the control input of the i - th agent at time k, u j (k) denotes the control input of the neighbor agent j of agent i at time k, e i(k) represents the consensus error under the new state information on the i-th agent; u i (k) represents the control input of the i-th agent at time k; α represents the discount factor; V i (e i (k + 1) represents the performance index of the i-th agent at time k + 1.
[0026] Furthermore, the process of the Actor network updating the network weights at the next moment includes:
[0027] Calculating the performance index at the current moment according to the weights of the Critic network and the state error of the agent at the current moment; taking the performance index at the current moment as the second temporal error and calculating the second objective function according to the second temporal error; calculating the network weights at the next moment according to the second objective function and the network weights at the current moment.
[0028] Preferably, the formula for updating the control information is:
[0029]
[0030] Wherein, represents the control input of the i-th agent at time k, represents the weights of the Actor network on agent i during the l-th iteration, s a,i (k) represents the input of the Actor network of the i-th agent at time k.
[0031] Preferably, the formula for the intelligent unmanned cluster system to reach consensus is:
[0032]
[0033] Wherein, represents the weights of the Critic network on agent i during the l-th iteration, represents the weights of the Critic network on agent i during the (l + 1)-th iteration, and ∈ represents the stability threshold.
[0034] The beneficial effects of the present invention are as follows:
[0035] 1. In the multi-agent system of the present invention, input delay and state delay are considered, which is more in line with the actual situation, and the requirements for delay are not strict, which enhances the feasibility.
[0036] 2. The consensus control method of the present invention is model-free and online, does not require knowledge of the precise model of the system, and at the same time, as an online algorithm, it can be better applied to real life.
[0037] 3. The control protocol of the present invention is a distributed control protocol. By enabling the leader to communicate with some agents, the system converges to the expected value, which not only reduces the control cost but also decreases the workload.
[0038] 4. The method of the present invention requires allocating a space in the memory as an experience pool, which improves the data utilization rate and is beneficial to the stability of neural network training. Brief Description of the Drawings
[0039] Figure 1 It is a flowchart of the optimal consensus cooperative control method for the intelligent unmanned cluster system under the influence of multiple time delays in the present invention;
[0040] Figure 2 It is a topology diagram of agents in an embodiment of the present invention;
[0041] Figure 3 It is an evolution diagram of the state of agents in an embodiment of the present invention;
[0042] Figure 4 It is an evolution diagram of the control of agents in an embodiment of the present invention;
[0043] Figure 5 It is an evolution diagram of the network weights of the Critic part of agents in an embodiment of the present invention;
[0044] Figure 6 It is an evolution diagram of the network weights of the Actor part of agents in an embodiment of the present invention. Detailed Embodiment
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0046] The "intelligent unmanned cluster system" and "agent" referred to in the present invention are both multi-agent systems affected by input time delay and state time delay.
[0047] The present invention proposes an optimal consensus cooperative control method for an intelligent unmanned cluster system under the influence of multiple time delays. As Figure 1 shown, the method includes the following contents:
[0048] S1: Obtain the interaction information of agents in the intelligent unmanned cluster system, and determine the topological structure of the system according to the interaction information.
[0049] The obtained interaction information includes the maximum state delay and the maximum input delay of the intelligent unmanned cluster system, as well as the current state delay value and input delay value of the system.
[0050] In the optimal consensus cooperative control of multi - agents, the consensus requires us to consider the difference in states between an agent and its neighbor agents, and between an agent and the leader. The optimality, on the other hand, requires us to consider the issues of the energy consumption and convergence speed of the agents. This makes us involve an index to evaluate the state of the agents.
[0051] In the intelligent unmanned cluster system of the present invention, each agent is affected by the input delay and the state delay, and its system equation can be expressed as:
[0052]
[0053] where, x i (k + 1) represents the state of agent i at time k + 1, which is different from the state before augmentation; x i (k - δ) represents the state with a delay of δ at time k; u i (k - τ) represents the control input with a delay of τ at time k; A δ and B iτ are the system matrices of the agent under state delay and input delay respectively; δ i represents the magnitude of the state delay of agent i, and τ i represents the magnitude of the input delay of agent i.
[0054] The evaluation index for the current performance of the agent is:
[0055]
[0056]
[0057] where, α is the discount factor, α ∈ [0, 1]; r i (e i (n), u i (n), u j (n)) represents the utility function, e i (n) represents the consensus error of agent i at time n, u i (n) represents the control input of agent i at time n, u j (n) represents the control input of neighbor agent j at time n; Q ii 、R ii and R ij are the first, second, and third adjustment matrices respectively, which represent different degrees of emphasis on each variable and are generally set as the identity matrix.
[0058] S2: Set the initial weights of the Actor network and the Critic network in the agent.
[0059] It is assumed that each agent has an Actor network for making control decisions, a Critic network for making evaluations, and an experience pool for storing historical data. The Actor network obtains a control based on the current state information, which can be used as a control protocol. The Critic network scores the control obtained from the Actor and the current environmental information to obtain evaluation information. Both networks are three-layer neural networks, and the weights between the first two layers are fixed. The experience pool is used to store historical data, which can break the temporal correlation between data, thus making the training of the neural network more stable.
[0060] For the training process, the trained Actor network will be directly used to control the intelligent unmanned cluster system using the Actor network, the Critic network, and the experience pool.
[0061] During training, set the initial weights of the Actor network and the Critic network in the agent.
[0062] S3: The agent sends its own state information to neighboring agents, and the agent calculates the agent state error based on the state information of the neighboring agents and its own state information; where the state information includes state variables and control information.
[0063] The intelligent unmanned cluster system of the present invention is a system that considers state delay and input delay. The state input and control input of the agent are represented by x i (k) and u i (k) respectively. The state delay and input delay of each agent are δ i and τ i , and the maximum values of the state delay and input delay are and The state information of the agent consists of state variables and control information at the current time and past times, and the state information can be expressed as:
[0064]
[0065] where, represents the augmented state information of the i-th agent at time k, represents the state variables of the i-th agent from time to time k, represents the control input of the i-th agent from time to time k - 1.
[0066] The agent calculates the agent state error based on the state information of neighboring agents and its own state information, which is expressed as:
[0067]
[0068] where e i (k) represents the new state information on the i-th agent under the consensus error, represents the neighbor nodes of agent i, a ij represents the connection weight from agent j to agent i; b i represents the pinning control on agent i, that is, when the leader is connected to agent i, then b i = 1, otherwise b i = 0; represents the state information of the leader at time k. The leader represents the state that the follower agents ultimately want to reach. In this multi-agent consensus collaborative control scenario, it is expected that the followers can ultimately be consistent with the leader's trajectory, that is, e i (k) = 0.
[0069] The agent records information such as state information and agent state error in the experience pool; the agent retrieves data from the experience pool according to the priority and uses it to update the Critic network and the Actor network. At the same time, the priority of the retrieved data in the experience pool is updated; the priority of the experience pool data is determined by the absolute value of the temporal error. The greater the temporal error of the data, the smaller its priority. Its specific form is: where p is the priority of the data, κ c,i represents the current temporal error of the Actor network. Compared with the conventional prioritized experience replay strategy, the prioritized experience replay strategy of the present invention can accelerate the learning of the optimal strategy.
[0070] The process by which the agent updates the Critic network and the Actor network according to the retrieved data from the experience pool is as follows:
[0071] S4: Both the Critic network and the Actor network update the network weights at the next moment according to the agent state error and network weights at the current moment; the Actor network updates the control information according to the Actor network weights at the current moment and the agent state error at the current moment.
[0072] The process of the Critic network updating the network weights at the next moment includes:
[0073] The Critic network calculates the performance index at the current moment and the performance index at the next moment according to the Critic network weights at the current moment and the agent state error;
[0074]
[0075] V i (e i (k + 1)) = (W′ c,i ) T σ(s c,i (k + 1))
[0076] where, V(e i (k)) represents the performance index of agent i at time k; represents the weight of the Critic network on agent i at the l - th iteration; s c,i (k) represents the input of the Critic network on the i - th agent, and its specific form is s c,i (k) = [e i (k)]; σ(·) represents the activation function on the nodes in the network, and its specific expression is σ(·) = tanh(·); V i (e i (k + 1)) represents the performance index of agent i at time k + 1, and W′ c,i represents the weight of the target network on agent i, and its specific form can be expressed as The target network is a neural network with the same structure as the Actor and Critic networks. By delaying the update of the target network, the process of network training can be made more stable.
[0077] Calculate the first - order temporal error according to the performance index at the current time, the performance index at the next time, the state information of the agent at the current time, and the state error of the agent at the current time:
[0078] κ c,i = V i (e i (k)) - r(e i (k), u i (k), u j (k)) - αV i (e i (k + 1))
[0079] where, κ c,i represents the first - order temporal error at the current time, u i (k) represents the control input of the i - th agent at time k, u j (k) represents the control input of the neighbor agent j of agent i at time k, and α represents the discount factor.
[0080] Calculate the first objective function according to the first - order temporal error:
[0081]
[0082] Among them, E c,i represents the objective function of the Critic network, m represents the number of data taken from the experience pool in one iteration, and the data taken is the consistency error at multiple moments. If the number of data in the experience pool is less than m, multiple data are randomly repeated until m data are taken; represents the temporal error calculated from the d-th data.
[0083] Calculate the network weight at the next moment according to the first objective function and the Critic network weight at the current moment:
[0084]
[0085] Among them, represents the Critic network weight on agent i at the l-th iteration, represents the Critic network weight on agent i at the (l + 1)-th iteration; β c represents the learning rate of the Critic network, E c,i represents the objective function of the Critic network.
[0086] The process of the Actor network updating the network weight at the next moment includes:
[0087] Calculate the performance index at the current moment according to the Actor network weight at the current moment and the agent state error:
[0088]
[0089] Take the performance index at the current moment as the second temporal error and calculate the second objective function according to the second temporal error:
[0090]
[0091] K a,i = V i (e i (k))
[0092] Among them, E a,i represents the objective function of the Critic network, represents the second temporal error calculated from the d-th data.
[0093] Calculate the network weight at the next moment according to the second objective function and the network weight at the current moment:
[0094]
[0095] Among them, Denote the network weights of the Actor on agent \(i\) at the \(l\)-th iteration. Denote the network weights of the Actor on agent \(i\) at the \((l + 1)\)-th iteration, \(\beta\). a Denote the learning rate of the Actor network, \(E\). a,i Denote the objective function of the Critic network.
[0096] The Actor network updates the control information according to the Actor network weights at the current moment and the state error of the agent at the current moment:
[0097]
[0098] Where, Denote the control information of the \(i\)-th agent at time \(k\); Denote the network weights of the Actor on agent \(i\) at the \(l\)-th iteration, which can be regarded as the updated Actor network weights at the current moment; \(s\) a , i \((k)\) denotes the input of the Actor network of the \(i\)-th agent at time \(k\), and its form can be expressed as \([e\) i \((k); u\) i \((k); u\) f \((k)]\).
[0099] S5: Judge whether the intelligent unmanned cluster system reaches consensus according to the weights of the Critic network. If it reaches consensus, control the agents with the control information of the current Critic network. Otherwise, return to step S3.
[0100] Compare the state information of the agent itself at the current moment with the state information of its neighboring agents. If the state difference between each agent and its neighboring agents is less than the threshold, the multi-agent system reaches consensus. Eventually, the network will tend to be stable, and the intelligent unmanned cluster system also reaches consensus. It can be judged whether the intelligent unmanned cluster system reaches consensus through the weights of the Critic network. The formula for the intelligent unmanned cluster system to reach consensus is:
[0101]
[0102] Where, \(\in\) represents the stability threshold, which is set as a small constant.
[0103] When the intelligent unmanned cluster system reaches consensus, control the agents with the control information of the current Critic network, and the agents enter the next state, generating the state information at the next moment.
[0104] Evaluate the present invention:
[0105] Use Matlab for simulation verification. Assume the topological graph of the agents is as Figure 2As shown, a multi-agent system consisting of 5 nodes, where the system nodes consist of a leader and four followers. The maximum state delay and maximum input delay of the agents in the system are both 2, while the state delays and input delays of the 4 agents are 1, 2, 1, 2 and 2, 1, 1, 2 respectively; regarding the parameter values in the system, take Q ii = I, R ii = 1, R ij = 1 and α = 0.85.
[0106] Randomly initialize the weights of the Critic network within [0, 1] and the weights of the Actor network within [-0.1, 0.1]. The reason for choosing the former interval is to ensure that the V function is positive definite, and the latter is to avoid training instability caused by excessive gradients due to too large initial values. The learning rates of the Actor network and the Critic network are respectively selected as β a = β c = 0.01. The size of the data processed simultaneously in one iteration is m = 5, and the target network is updated every 10 iterations.
[0107] From the simulation results, it can be concluded that, as Figure 3 shown, which shows the trajectory evolution process of all agents. It can be seen that as the number of iterations increases, the 4 followers gradually track the leader, and finally at the 25th iteration, they are basically completely consistent, which also means that the consensus error of the agents gradually tends to 0. As Figure 4 shown, which shows the control input change diagram of the 4 followers. It can be clearly seen that all controls finally tend to 0, which indicates that the agents have basically reached consensus and no additional control is required. As Figure 5 , Figure 6 shown, which is the evolution diagram of the weights of the Critic network and the Actor network on the agents. From the evolution trend of the image, it can be seen that the network weight change rate is relatively large at the beginning, and then the network weights tend to be stable until the weights basically do not change, which means that the neural network has converged. At this time, the state error also tends to 0, which also represents that the training effect of the neural network is relatively ideal.
[0108] After verification by simulation experiments, according to the optimal consensus control method under the influence of multi-delays, the states of the agents in the system can be made to tend to be consistent.
[0109] It should be noted that those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0110] The above embodiments are used to further elaborate on the purpose, technical solutions, and advantages of the present invention. It should be understood that the above embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An optimal consensus cooperative control method for an intelligent unmanned cluster system under the influence of multiple time delays, characterized in that, Including: S1: Obtain the interaction information of agents in the intelligent unmanned cluster system, and determine the topological structure of the system according to the interaction information; S2: Set the initial weight values of the Actor network and the Critic network in the agent; S3: The agent sends its own state information to neighboring agents, and the agent calculates the agent state error according to the state information of the neighboring agents and its own state information; among them, the state information includes state variables and control information; the state information and control information of the agent itself are expressed as: Among them, represents the state information of the i-th agent at time k, represents the state variables of the i-th agent from time to time k, represents the control input of the i-th agent from time to time k - 1, represents the maximum state time delay, represents the maximum input time delay; S4: Both the Critic network and the Actor network update the network weights of the next moment according to the agent state error and network weights at the current moment; the Actor network updates the control information according to the Actor network weights at the current moment and the agent state error at the current moment; The process of the Critic network updating the network weights of the next moment includes: The Critic network calculates the performance index at the current moment and the performance index of the next moment according to the Critic network weights at the current moment and the agent state error, calculates the first time series error according to the performance index at the current moment, the performance index of the next moment, the state information of the agent at the current moment, and the agent state error at the current moment; calculates the first objective function according to the first time series error; calculates the network weights of the next moment according to the first objective function and the network weights at the current moment; The process of the Actor network updating the network weights of the next moment includes: Calculates the performance index at the current moment according to the Critic network weights at the current moment and the agent state error; takes the performance index at the current moment as the second time series error and calculates the second objective function according to the second time series error; calculates the network weights of the next moment according to the second objective function and the network weights at the current moment; S5: Judge whether the intelligent unmanned cluster system reaches consensus according to the weights of the Critic network. If it reaches consensus, control the agent with the control information of the current Critic network; otherwise, return to step S3.
2. The optimal consensus cooperative control method for an intelligent unmanned cluster system under the influence of multiple time delays according to claim 1, characterized in that The formula for calculating the agent state error is: Among them, e i (k) represents the new state information on the i-th agent under the consensus error; represents the neighbor nodes of agent i; a ij represents the connection weight from agent j to agent i; b i represents the pinning control on agent i; represents the state information of the leader at time k.
3. The optimal consensus cooperative control method for an intelligent unmanned cluster system under the influence of multiple time delays according to claim 1, wherein The formula for calculating the performance index is: Among them, V(e i (k)) is the performance index; represents the weights of the Critic network on agent i at the l-th iteration; s c,i (k) represents the input of the Critic network on the i-th agent; σ(·) represents the activation function of the nodes in the network.
4. The optimal consensus cooperative control method for an intelligent unmanned cluster system under the influence of multiple time delays according to claim 1, characterized in that The formula for calculating the first time series error is: κ c,i = V i (e i (k)) - r(e i (k), u i (k), u j (k)) - αV i (e i (k + 1)) where, κ c,i represents the first timing error at the current moment, V i (e i (k)) represents the performance index of the i-th agent at time k; r(·) is the utility function; u i (k) represents the control input of the i-th agent at time k, u j (k) represents the control input of the neighbor agent j of agent i at time k, e i (k) represents the consensus error under the new state information of the i-th agent; α represents the discount factor; V i (e i (k + 1) represents the performance index of the i-th agent at time k + 1.
5. The optimal consensus cooperative control method for an intelligent unmanned cluster system under the influence of multiple time delays according to claim 1, wherein, The formula for updating the control information is: Among them, represents the control input of the $i$-th agent at time $k$, represents the weights of the Actor network on agent $i$ at the $l$-th iteration, a,i (k) represents the input of the Actor network of the $i$-th agent at time $k$.
6. The optimal consensus cooperative control method for an intelligent unmanned cluster system under the influence of multiple time delays according to claim 1, wherein The formula for the intelligent unmanned cluster system to reach consensus is: Among them, represents the weights of the Critic network on agent i at the l-th iteration, represents the weights of the Critic network on agent i at the (l + 1)-th iteration, and ∈ represents the stability threshold.
Citation Information
Patent Citations
Multi-agent system optimal consistency control method based on reinforcement learning
CN114755926A
Non-affine multi-agent dynamic event triggering tracking control method under asynchronous framework
CN115327901A