Optimal consistency cooperative control method for intelligent unmanned cluster system
By adopting Actor-Critic neural network and gradient descent algorithm in the intelligent unmanned cluster system, combined with experience playback technology, taking into account the differences in the computational capabilities of the agent, the optimal consistent collaborative control of the intelligent unmanned cluster system is achieved, the problem of inconsistent agent speed is solved, and the system's convergence speed and control effect are improved.
Patent Information
- Application Number
- CN202510208257.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-30
AI Technical Summary
In an intelligent unmanned cluster system, the difference in computing power between agents leads to inconsistent speeds of information interaction and neural network parameter updates, affecting the system's convergence.
An optimal consistent collaborative control method is proposed. By constructing the topological structure and Laplace matrix of the intelligent unmanned cluster system, the agent is divided into leaders and followers, the interaction strategy is updated using the Actor-Critic neural network, and the parameters are trained through gradient descent algorithm and empirical playback technology, the differences in computing power are considered and distributed control is realized.
It effectively solves the problem of inconsistent speed of the agent caused by differences in computing power, improves the system's convergence speed, reduces control costs, and achieves better control effects when the system's accurate model is unknown.
Smart Images

Figure CN120065850A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent unmanned cluster system control, and particularly relates to an optimal consensus cooperative control method for an intelligent unmanned cluster system. Background Art
[0002] In recent years, inspired by the behavior of biological groups, experts and scholars have applied the consensus theory of multi-agent systems (MASs) to the cooperative control of complex systems. The consensus problem of multi-agent systems has broad application prospects in fields such as swarm control, sensor networks, distributed computing, and collective decision-making. Among them, the optimal consensus problem is an important research direction of multi-agent systems, aiming to design a distributed protocol that enables all agents to achieve synchronization with the lowest energy consumption.
[0003] Reinforcement learning (RL) is an important branch of machine learning and is widely used in the design of controllers. Its core idea is to reinforce or encourage certain behaviors of agents through the observations and rewards provided by the environment, thereby increasing the likelihood of achieving the desired results. Policy gradient (PG) is a commonly used reinforcement learning algorithm, and its basic principle is to model the policy gradient and update the control policy by the method of gradient descent. The PG method can better adapt to high-dimensional and continuous state and action spaces and has achieved remarkable results in the field of multi-agent optimal control.
[0004] In the field of intelligent unmanned cluster systems, there are certain communication restrictions between agents, and at the same time, there are differences in the computing capabilities of different agents. However, in existing research, the problem of differences in the computing capabilities of agents is rarely considered, and such differences are very common in real life. Therefore, considering the problem of differences in computing capabilities is more in line with reality and of great significance. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present invention proposes an optimal consensus cooperative control method for an intelligent unmanned cluster system. The method includes: constructing an intelligent unmanned cluster system and determining the topological structure and Laplacian matrix of the system; dividing the agents in the intelligent unmanned cluster system into leaders and followers, where the followers are composed of first-order and second-order hybrid agents; each agent conducts information interaction; reconstructing the local tracking error through the topological structure and the interaction information between agents, and defining a performance index function; using an Actor-Critic neural network to update the interaction strategies between agents, and updating and training the parameters of the Actor-Critic neural network through a gradient descent algorithm and an experience replay technique. When the parameters of the Actor-Critic neural network are stable, the agents in the intelligent unmanned cluster system reach consensus.
[0006] Advantages of the present invention:
[0007] The present invention proposes a policy gradient reinforcement learning algorithm considering computational power differences, which solves the problem that the difference in computational power causes the speed of the agent to be inconsistent when interacting with the environment and updating neural network parameters, and has a certain impact on the convergence of the entire system. The unmanned cluster system of the present invention is a system with an unknown exact model, and the Actor-Critic framework used can better solve the situation of an unknown exact model of the system. The control protocol of the present invention is a distributed control protocol. By enabling the leader to communicate with some agents, the system converges to the expected value, which not only reduces the control cost but also reduces the workload. The present invention introduces an experience pool and uses the experience replay technique to break the correlation between data, enabling the agent to fully interact with the environment and improving the utilization rate of data. Brief Description of the Drawings
[0008] Figure 1 It is the system control flow chart of the embodiment of the present invention;
[0009] Figure 2 It is the communication topology diagram that appears during the system convergence process of the embodiment of the present invention;
[0010] Figure 3 It is the agent displacement state evolution diagram of the embodiment of the present invention;
[0011] Figure 4 It is the agent speed state evolution diagram of the embodiment of the present invention;
[0012] Figure 5 It is the first-order agent control evolution diagram of the embodiment of the present invention;
[0013] Figure 6 It is the second-order agent control evolution diagram of the embodiment of the present invention;
[0014] Figure 7 It is the agent state error evolution diagram of the embodiment of the present invention;
[0015] Figure 8 It is the agent state error evolution diagram of not using the asynchronous policy gradient reinforcement learning method proposed by the present invention. Detailed Embodiments
[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0017] An optimal consensus cooperative control method for an intelligent unmanned cluster system, the method comprising: constructing an intelligent unmanned cluster system, and determining the topological structure and Laplacian matrix of the system; dividing the agents in the intelligent unmanned cluster system into leaders and followers, where the followers consist of first- and second-order heterogeneous agents; each agent performs information interaction; reconstructing the local tracking error through the topological structure and the interaction information between agents, and defining a performance index function; using an Actor-Critic neural network to update the interaction strategy between each agent, and updating and training the parameters of the Actor-Critic neural network through a gradient descent algorithm and an experience replay technique, and when the parameters of the Actor-Critic neural network are stable, the agents in the intelligent unmanned cluster system reach consensus.
[0018] As Figure 1 shown, an optimal consensus cooperative control method for an intelligent unmanned cluster system considering computational power differences, the method including but not limited to the following steps:
[0019] S1: According to the information interaction of the agents in the intelligent unmanned cluster system, determine the topological structure of the system, and at the same time introduce a virtual velocity to reconstruct the heterogeneous multi-agent system into a homogeneous multi-agent system, which is convenient for calculation and subsequent implementation.
[0020] The leader dynamic equation is expressed as:
[0021] The follower dynamic equation is expressed as:
[0022] where, x 0 (t + 1) is the displacement state value of the leader at time t + 1, x 0 (t) is the displacement state value of the leader at time t, v 0 (t + 1) is the velocity state value of the leader at time t + 1, v 0 (t) is the velocity state value of the leader at time t, xi(t) is the displacement state value of follower i at time t, vi(t) is the velocity state value of follower i at time t, ui(t) represents the control input of follower unmanned aircraft i at time t, A, B, C, are different unknown constant matrices, i ∈ β 1 indicates that agent i belongs to a second-order agent, i ∈ β 2 indicates that agent i belongs to a first-order agent.
[0023] Introduce a virtual velocity w i (t) for the first-order agent, and construct the first-order agent as a second-order agent, then the dynamic equation of the original system can be reconstructed as:
[0024]
[0025] where, when \(i\in\beta\) 1 , when \(i\in\beta\) 2 , is the state of the reconstructed follower agent \(i\) at time \(t + 1\), is the state of the reconstructed follower agent \(i\) at time \(t\), is the state of the reconstructed leader agent at time \(t + 1\), is the state of the reconstructed leader agent at time \(t\), \(\Xi,\xi\) i , are unknown constant matrices.
[0026] S2: Divide the intelligent unmanned cluster system into leaders and followers. The followers are composed of first - order and second - order agents. Through the topological structure and the interaction information between agents, reconstruct the local tracking error of the system, and at the same time define the performance index function. In the optimal consensus cooperative control of multi - agents, the consensus requires considering the difference in the states between an agent and its neighbor agents, and between an agent and the leader. While the optimality requires considering the energy consumption and convergence speed of the agents, which makes it necessary to design a performance index to evaluate the state of the agents.
[0027] The local tracking error of the agent is reconstructed as:
[0028]
[0029] The performance index function is expressed as:
[0030] where, \(\varepsilon\) i (t) represents the local state error system of follower \(i\) at time \(t\); \(b\) i represents whether the follower UAV can receive the state information of the leader UAV. \(b\) i is also called the anchoring gain, which is a non - negative positive definite matrix. \(b\) i = 1 means the follower can receive the information of the leader, \(b\) i = 0 means the follower cannot receive the information of the leader; \(a\) ij > 0 means follower \(i\) can receive the state information of follower \(j\), \(a\) ij = 0 means follower \(i\) cannot receive the state information of follower \(j\); \(S\) i represents the adjacent nodes within the same subgroup; \(D\) i represents another subgroup corresponding to the node of agent \(i\). is the utility function, \(Q\) ii and \(R\) ijThey are respectively the set constant matrices. The optimal consensus cooperative control is to find an optimal control strategy to minimize the energy consumption of multi-agent systems during the process of convergence, that is, J i is minimized.
[0031] S3: Adopt an asynchronous reinforcement learning algorithm so that each agent can independently learn and update its policy at different time steps. Use the Actor-Critic neural network to approximate the control policy and the performance metric function respectively. The Critic network evaluates the control actions approximated by the Actor network, and the Actor network adjusts the control actions according to the evaluation of the Critic network.
[0032] The control can be expressed as:
[0033] where, are the weight parameters of the Actor neural network, δ ai is the activation function, h ai (t) is the input vector of the action information and related position information of the follower agent i and its neighbor agents in the Actor network.
[0034] S4: Use gradient descent to update the neural network parameters during training. At the same time, combine the experience pool and adopt the experience replay technique to break the correlation of data and improve the utilization rate of data in the parameter update process. Store the historical data in the experience pool, and the agent retrieves the data from the experience pool when needed to update the Critic and Actor neural networks. When the parameters of the Actor-Critic neural network tend to be stable, the intelligent unmanned cluster system reaches consensus.
[0035] Its update method is:
[0036]
[0037] where, are the weights of the Critic network, β c and β a are the learning rates of the Critic network and the Actor network respectively, E ci and are the objective functions of the Ciric network and the Actor network respectively. After continuous updates, the final neural network will tend to be stable. At this time, the agent can also obtain the optimal control according to the Actor network to make the multi-agent system reach consensus. Compare the current state of the agent and the state information of its neighbor agents. If the state difference of each agent is less than the threshold, the multi-agent system reaches consensus. Finally, the network will tend to be stable, and the intelligent unmanned cluster system also reaches consensus. The condition for judging stability is where ∈ is a relatively small constant set as required. When the network is stable, it is determined whether to end the training phase and directly use the Actor network in actual applications.
[0038] This embodiment considers a multi-agent system composed of n agents. The relationship topology of the multi-agent system can be represented by a directed weighted graph G = (V, E, A). Each agent is a node of the undirected weighted graph G = (V, E, A), where V = {v 1 , v 2 , …, v n} represents the set of nodes, E represents the set of edges, and A = [a ij represents the adjacency matrix, where the matrix element a ij represents the connection weight from agent node i to j. If there is a connection between node i and node j, that is, e ij = (v j , v i ), then a ij > 0; if there is no connection between node i and node j, then a ij = 0. It is stipulated that a ij = 0, that is, the system has no self-loops. The nodes connected to node i are the neighbor nodes of node i. The neighbor nodes of node i are represented by the set N i = {v j ∈ V|(v j , v i ) ∈ E}. The in-degree matrix of the system nodes is where the in-degree of node v i is expressed as The Laplacian matrix of the system topology is where l ij = -a ij , i ≠ j, l ii = ∑ i≠j a ij .
[0039] The system needs to satisfy an assumption: the directed graph G has a spanning tree, and the leader is at least connected to one follower in G.
[0040] To ensure that the present invention satisfies the consistency condition of the consistent collaborative control of the intelligent unmanned cluster system, the following proof is carried out, including:
[0041] Define the agent behavior state value function as:
[0042] V i (ε i (λ)) = r(ε i (λ), u i (λ), uj (λ)) + V i (ε i (λ + 1))
[0043] where is the utility function, Q ii and R ij are respectively the set constant matrices.
[0044] Construct the Lyapunov function as α λ V i (ε i (λ)), this function is positive definite, and its difference
[0045] Δ(α λ V i (ε i (λ))) = -α λ r(ε i (λ), u i (λ), u j (λ)) ≤ 0
[0046] When the assumption of the spanning tree of the directed graph is satisfied, it indicates that the error of the agents can reach asymptotic consensus. In other words, all followers can reach consensus with the leader.
[0047] To verify the effectiveness of the proposed algorithm, Matlab is used for experimental simulation. In this embodiment, the experiment Figure 2 is the system topology graph, a multi-agent system composed of eight nodes. The system nodes consist of one leader and seven followers. Among them, agent 0 is the leader, agents 1, 2, 3, 4 are second-order agents, and agents 5, 6, 7 are first-order agents. Regarding the parameter values in the system, Q ii = I, R ii = 1, R ij = 1. The system matrix ξ 1 = (0, 1.11) T , ξ 2 = (0, 0.82) T , ξ 3 = (0, 0.91) T , ξ 4 = (0, 0.75) T , ξ 5 = diag{0.779, 0.779}, ξ 6 = diag{0.51, 0.51}, ξ 7 = diag{0.86, 0.86}. The learning rates of the Actor network and the Critic network are respectively selected as β c = β a= 0.003, the target network is the same neural network as the Actor and Critic networks, and the time steps for asynchronous update are T1 = T2 = T3 = T4 = 1, T5 = T6 = T7 = 3.
[0048] From the simulation results, as Figure 3 shown, the displacement trajectory evolution diagrams of all agents are presented. The first-order and second-order agents both basically reach full consistency at the 150th round respectively. Figure 4 The velocity-displacement trajectory diagram of the second-order agent is shown. From the figure, it can be seen that the second-order agent finally basically reaches consistency at the 80th round. As Figure 5 、 Figure 6 The control input change diagrams of all followers are shown. It can be clearly seen that all controls finally tend to 0, which indicates that the agents have basically reached consistency and no other additional control is needed. From Figure 7 it can be seen that the trends of the tracking errors and control inputs of all agents are consistent and finally tend to 0. Considering the control inputs and state errors together, the performance function of the agents will finally tend to 0, that is, reaching the optimum. At the same time, to verify the advantages of the proposed algorithm, a set of comparative test diagrams are made. As Figure 8 shown is the tracking error diagram of the agents using the original dynamic programming algorithm. From the figure, it can be seen that the tracking errors basically reach consistency at the 70th round of iteration, while for the agents using the algorithm proposed in the present invention Figure 7 the agents basically reach consistency at the 60th round of iteration. Therefore, the proposed algorithm can accelerate the convergence speed to a certain extent and has a certain improvement effect on the original algorithm.
[0049] Through the verification of simulation experiments, according to the optimal consistency control method of asynchronous policy gradient, the states of the agents in the system can tend to be consistent.
[0050] It should be noted that those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The said program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.
[0051] The above-mentioned embodiments further elaborate on the purpose, technical solutions, and advantages of the present invention. It should be understood that the above-mentioned embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. An optimal consistency collaborative control method for an intelligent unmanned cluster system, characterized in that: include: Build an intelligent unmanned cluster system and determine the system's topological structure and Laplace matrix; divide the agents in the intelligent unmanned cluster system into leaders and followers; each agent exchanges information; reconstruct the local tracking error through the topological structure and the interaction information between agents, and define the performance indicator function; use the Actor-Critic neural network to update the interaction strategy between agents, and update the parameters of the Actor-Critic neural network through the gradient descent algorithm and experience replay technology. When the parameters of the Actor-Critic neural network are stable, the agents in the intelligent unmanned cluster system are consistent.
2. The optimal consistency collaborative control method of an intelligent unmanned cluster system according to claim 1 is characterized in that: The followers consist of a mixture of first-order agents and second-order agents.
3. The optimal consistency collaborative control method of an intelligent unmanned cluster system according to claim 2 is characterized in that: The leader dynamic equation is: The follower dynamic equation is expressed as: Among them, x0(t+1) represents the displacement state value of the leader at time t+1, x0(t) represents the displacement state value of the leader at time t, v0(t+1) represents the speed state value of the leader at time t+1, v0(t) represents the speed state value of the leader at time t, x i (t) represents the displacement state value of follower i at time t, v i (t) represents the speed state value of follower i at time t, u i (t) represents the control input of follower UAV i at time t, are different unknown constant matrices.
4. The optimal consistency collaborative control method of an intelligent unmanned cluster system according to claim 1 is characterized in that: Reconstructing the local tracking error includes: Among them, ε i (t) represents the local state error system of follower i at time t, b i Indicates whether the follower drone can receive the status information of the leader drone, b i is called the anchor gain, which is a non-negative positive definite matrix, b i =1 means the follower can receive the leader's information, b i =0 means the follower cannot receive the leader's information; a ij >0 means that follower i can receive the status information of follower j, a ij = 0 means that follower i cannot receive the status information of follower j; S i represents adjacent nodes in the same subgroup; D i Represents another subgroup of nodes corresponding to agent i.
5. The optimal consistency collaborative control method of an intelligent unmanned cluster system according to claim 1 is characterized in that: The performance indicator function is: Among them, Q ii and R ij are the constant matrices set respectively.
6. The optimal consistency collaborative control method of an intelligent unmanned cluster system according to claim 1 is characterized in that: The Actor-Critic neural network is used to update the interaction strategies between various agents: the actor network outputs the probability distribution of the agent's actions in the current state, guiding the agent to explore different interactive behaviors; the critic network evaluates the value generated by the output actions of the actor network and provides feedback for the strategy optimization of the actor network; through continuous iterative training, the policy gradient algorithm is used to adjust the actor network parameters so that the agent can choose actions that can obtain higher cumulative rewards; at the same time, the critic network is updated according to the time difference error to improve the accuracy of its value evaluation.
7. The optimal consistency collaborative control method of an intelligent unmanned cluster system according to claim 1 is characterized in that: The parameters of the Actor-Critic neural network are updated as follows: in, is the weight of the Critic network, β c and β a are the learning rates of the Critic network and the Actor network, respectively. ci and They are the objective functions of Ciric network and Actor network respectively.
Citation Information
Cited By
Method and system for controlling dynamic event triggering consistency of multi-agent system
CN120722803A
A dynamic event-triggered consensus control method and system for a multi-agent system
CN120722803B
Hot water spherical tank electric heating intelligent control system and method
CN120848646A
A hot water ball tank electric heating intelligent control system and method
CN120848646B
Bipartite cooperative control method of intelligent unmanned cluster system
CN121091893A