A Multi-Agent Formation Control Method Based on a Single Critic Reinforcement Learning Structure

Through the combination of single-commentist reinforcement learning structure and neural network, the problems of multiple calculation errors and long time in traditional methods are solved, and the efficient implementation of multi-agent formation control is achieved.

CN116185020BActive Publication Date: 2025-07-18FUZHOU UNIV

Patent Information

Application Number
CN202310081638.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-19
Publication Date
2025-07-18
Estimated Expiration
2043-01-19

AI Technical Summary

Technical Problem

Traditional actor-critician reinforcement learning structures have problems such as many calculation errors and long calculation time in the multi-agent formation control, which is difficult to effectively solve the problem of solving the Hamilton-Jackby-Belman equation.

Method used

The single-commenter reinforcement learning structure is adopted, the actor network is removed, the critic network update strategy is redesigned, and the unknown nonlinear terms are approximate, and the formation control is carried out through the single-commenter reinforcement learning structure.

Benefits of technology

Effectively reduce calculation time and estimation errors, and ensure the smooth completion of the formation behavior of the nonlinear multi-agent system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116185020B_ABST
    Figure CN116185020B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-agent formation control method based on a single critic reinforcement learning structure, including: constructing a communication structure for each agent in a multi-agent system; constructing a tracking error of an agent relative to a leader agent, and constructing an error describing the agent and the leader as well as the agent and neighbor agents, namely a formation error; constructing a cost function and a value function related to the formation error and an optimal control input based on optimal control; expanding and solving the value function to construct a corresponding HJB equation; solving the partial derivative of the HJB equation with respect to the optimal control to obtain the expression form of the optimal control input with respect to the optimal value function; partitioning the optimal value function to obtain a partitioned form of the optimal control input; introducing a single critic reinforcement learning structure and combining it with a neural network to solve the obtained partitioned optimal value function and the optimal control input. This method is beneficial to reducing the estimation error and the calculation time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multi-agent formation control, and particularly relates to a multi-agent formation control method based on a single critic reinforcement learning structure. Background Art

[0002] A multi-agent system consists of autonomous and interactive entities that share a common environment. Agents can perceive and perform operations based on the environment. Formation control is an application area of multi-agents, including formation application scenarios such as satellites, underwater robots, and drone flights. Through the research of scientists, many formation control methods have been proposed, such as the leader-follower method, the virtual structure method, etc. Among them, the leader-follower formation control method, as a simple and scalable formation control algorithm, is currently widely used in multi-agent formations. Its strategy is to set an agent as the leader and set its movement trajectory, and then design a controller to control other follower agents to track the leader's trajectory.

[0003] Optimal control is an effective method for balancing control performance and control resource consumption. It can achieve control goals by minimizing a cost function. As a method in optimal control, dynamic programming has wide application value. Its basic idea is to decompose the problem of finding the optimal solution of a large problem into solving several small problems. However, the backward solution process and the curse of dimensionality of this method have hindered its further application and development. The emergence of adaptive dynamic programming combines the optimal control method with the structure of reinforcement learning, overcomes the above defects of dynamic programming, and can estimate unknown equations through function approximation. Among them, a method that combines optimal control with the actor-critic reinforcement learning structure can effectively solve the problem of the intractability of the Hamilton-Jacobi-Bellman equation in solving the optimal controller. However, this method involves the iteration of the actor and critic dual networks, which will generate more computational errors and longer computational times.

[0004] Therefore, designing a formation control method for a multi-agent system based on a reinforcement learning structure that reduces computational time and computational errors remains an open problem. To address this problem, the present invention removes the actor network and redesigns the critic network update strategy to enable it to evaluate performance and correct in a timely manner while performing control actions. This method can effectively reduce computational time and estimation errors and ensure the successful completion of the formation behavior of a nonlinear multi-agent system. Summary of the Invention

[0005] The purpose of the present invention is to provide a multi-agent formation control method based on a single critic reinforcement learning structure, which is beneficial to reducing estimation errors and computational time.

[0006] To achieve the above object, the technical solution adopted by the present invention is: a multi-agent formation control method based on a single critic reinforcement learning structure, comprising the following steps:

[0007] Step 1: Based on graph theory in applied mathematics, construct the communication structure of each agent in the multi-agent system. Considering the system as a first-order multi-agent system, each agent only obtains the position information of its neighboring agents; meanwhile, there is a leader agent in the system, and other agents act as followers and move along the trajectory of the leader agent during operation;

[0008] Step 2: For each agent in the system, construct its tracking error relative to the leader agent according to the information of its neighboring agents obtained, and construct the error describing the agent and the leader as well as the agent and its neighboring agents, that is, the formation error, according to the tracking error;

[0009] Step 3: Based on optimal control, construct a cost function and a value function related to the formation error and the optimal control input;

[0010] Step 4: Based on Taylor's formula and the value function obtained in Step 3, expand and solve the value function to obtain the corresponding Hamilton-Jacobi-Bellman equation;

[0011] Step 5: For the Hamilton-Jacobi-Bellman equation obtained in Step 4, solve the partial derivative with respect to the optimal control to obtain the expression form of the optimal control input with respect to the optimal value function;

[0012] Step 6: Divide the optimal value function to obtain its expression form with respect to the formation error and the unknown function, and obtain the divided optimal control input form according to the expression form of the optimal control input in Step 5;

[0013] Step 7: Introduce a single critic reinforcement learning structure and combine it with a neural network to solve the divided optimal value function and the optimal control input obtained in Step 6, where the neural network approximates the unknown non-linear terms in the multi-agent system, and the critic network performs formation control of the agent system and evaluates and improves the effect of the formation control.

[0014] Further, the single critic reinforcement learning structure is used to remove the need for the actor network in the traditional actor-critic reinforcement learning method, thereby effectively reducing the approximation error of the system and reducing the calculation time.

[0015] Further, in Step 1, the model of the multi-agent system is expressed as:

[0016]

[0017] where, xi (t) represents the position of the i-th agent in the system; u i (t) represents the control input of the i-th agent in the system; f i (·) represents an unknown nonlinear function, and it is assumed to be Lipschitz continuous;

[0018] The model of the leader agent is as follows:

[0019]

[0020] where, p l and v l represent the trajectory and velocity of the leader respectively, that is, the desired trajectory and velocity in formation movement; The tracking error of each agent relative to the leader is set as:

[0021] z i = x i - p l - ζ i

[0022] where, ζ i represents the position between the leader agent and the i-th follower agent, and is used to describe the formation shape of the system;

[0023] According to the structure of the tracking error, the formation error form is defined as follows:

[0024]

[0025] where, a ij is the element in the i-th row and j-th column of the adjacency matrix in graph theory; b i is the connection weight parameter between the i-th follower agent and the leader agent; Λ i represents the neighbor set of the i-th agent.

[0026] Furthermore, in step three, combining the defined formation error, the expression form of the cost function is obtained as:

[0027]

[0028] where, C = dia. g {c1, c2,..., c i ,..., c n}; w1 and w2 are two set constants; I m is an identity matrix of an appropriate dimension; is the symbol of the tensor product;

[0029] According to the obtained cost function, the corresponding value function is established, and the optimal control input The corresponding optimal value function is finally obtained as follows:

[0030]

[0031] where τ represents the integration constant.

[0032] Furthermore, in step four, the Hamilton-Jacobi-Bellman equation is established as follows:

[0033]

[0034] Taking the partial derivative of the above equation with respect to , the expression form of the optimal control input is obtained as:

[0035]

[0036] Furthermore, for the unknown nonlinear term f i (x i ) existing in the multi-agent system, a neural network is introduced for approximate estimation:

[0037]

[0038] where represents the ideal neural network weight matrix; S fi (x i ) represents the basis function vector; ∈ fi (x i ) represents the approximation error;

[0039] Since is only used for theoretical analysis but is an unknown matrix in practice, an estimation matrix is introduced for estimation, and the approximated by the neural network identifier is obtained as follows:

[0040]

[0041] According to the obtained approximate function the estimated values of other variables are obtained.

[0042] Furthermore, the optimal value function and the optimal control input are converted into the following expression forms by splitting the parameters:

[0043]

[0044]

[0045] where k i represents a constant term greater than zero; and The expression of

[0046] Furthermore, the expressions of the optimal value function and the optimal control input after introducing the single-critic reinforcement learning structure are as follows:

[0047]

[0048]

[0049] Among them, represents the introduced estimated critic network parameter matrix; S i represents the radial basis function of the neural network; the update law of the critic network parameter matrix is expressed as follows:

[0050]

[0051] Among them, k ci represents the learning rate of the critic network, and the specific expression of φ i is as follows:

[0052]

[0053] Compared with the prior art, the present invention has the following beneficial effects: Aiming at the problems of redundant calculation errors and longer calculation time caused by the iteration of the actor-critic dual network in the multi-agent formation control method based on the traditional actor-critic reinforcement learning structure, the present invention proposes a multi-agent formation control method based on a single-critic reinforcement learning structure. This method removes the actor network and redesigns the update strategy of the critic network, enabling it to evaluate performance and correct in a timely manner while performing control actions. This method can effectively reduce the calculation time and reduce the estimation error, and ensure the successful completion of the formation behavior of the nonlinear multi-agent system. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 is a block diagram of the traditional actor-critic reinforcement learning structure in the prior art;

[0055] Figure 2 is a block diagram of the single-critic reinforcement learning structure in the embodiment of the present invention;

[0056] Figure 3 is a communication topology diagram of the nonlinear multi-agent system in the embodiment of the present invention;

[0057] Figure 4 is a schematic diagram of the multi-agent formation trajectory in the embodiment of the present invention;

[0058] Figure 5 is a schematic diagram of the multi-agent formation speed trajectory in the embodiment of the present invention;

[0059] Figure 6 It is a comparison chart of position errors in the embodiments of the present invention and those of the traditional actor-critic method;

[0060] Figure 7 It is a comparison chart of speed errors in the embodiments of the present invention and those of the traditional actor-critic method;

[0061] Figure 8 It is a comparison chart of the calculation times of the actor-critic and single-critic reinforcement learning structures in the embodiments of the present invention;

[0062] Figure 9 It is a flow chart of the method implementation in the embodiments of the present invention. Detailed implementation manners

[0063] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0064] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0065] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0066] Starting from the formation requirements of a non-linear multi-agent system, for a class of non-linear multi-agent systems, this embodiment provides a multi-agent formation control method based on a single-critic reinforcement learning structure. As Figure 9 shown, this method includes the following steps:

[0067] Step 1: Based on graph theory in applied mathematics, construct the communication structure of each agent in the multi-agent system. Considering the system as a first-order multi-agent system, each agent only obtains the position information of its neighboring agents. At the same time, there is a leader agent in the system, and other agents act as followers and move along the trajectory of the leader agent during operation.

[0068] Step 2: For each agent in the system, construct its tracking error relative to the leader agent according to the information of its neighboring agents obtained, and construct the error describing the agent and the leader agent as well as the agent and its neighboring agents, that is, the formation error, according to the tracking error.

[0069] Step 3: Based on optimal control, construct a cost function and a value function related to the formation error and the optimal control input.

[0070] Step 4: Based on the Taylor formula and the value function obtained in Step 3, expand and solve the value function to obtain the corresponding Hamilton-Jacobi-Bellman (HJB) equation.

[0071] Step 5: For the HJB equation obtained in Step 4, solve the partial derivative with respect to the optimal control to obtain the expression form of the optimal control input with respect to the optimal value function.

[0072] Step 6: Divide the optimal value function to obtain its expression form with respect to the formation error and the unknown function, and according to the expression form of the optimal control input in Step 5, obtain the divided optimal control input form.

[0073] Step 7: Introduce a single-critic reinforcement learning structure and combine it with a neural network to solve the divided optimal value function and the optimal control input obtained in Step 6, where the neural network approximates the unknown non-linear terms in the multi-agent system, the critic network performs the formation control of the agent system, and evaluates and improves the effect of the formation control.

[0074] In this embodiment, the single-critic reinforcement learning structure is as Figure 2 shown. The single-critic reinforcement learning structure can remove the need for the actor network in the traditional actor-critic reinforcement learning method, thereby effectively reducing the approximation error of the system and reducing the calculation time.

[0075] In Step 1, the model of the multi-agent system is expressed as:

[0076]

[0077] where, x i (t) represents the position of the i-th agent in the system; u i (t) represents the control of the i-th agent in the system; f i (·) represents an unknown non-linear function, which is assumed to be Lipschitz continuous here.

[0078] The expected trajectory change of the leader agent is expressed by the following formula:

[0079]

[0080] where, p l and v l represent the trajectory and speed of the leader respectively, that is, the expected trajectory and speed in the formation movement.

[0081] In this embodiment, the communication topology diagram of the non - linear multi - agent system is as follows Figure 3 as shown.

[0082] In step two, according to the constructed leader - follower multi - agent system model, the tracking error of each agent relative to the leader is set as:

[0083] z i = x i - p l - ζ i

[0084] where ζ i represents the position between the leader agent and the i - th follower agent, and is used to describe the formation shape of the system.

[0085] According to the structure of the tracking error, the formation error form is defined as follows:

[0086]

[0087] where a ij is the element in the i - th row and j - th column of the adjacency matrix in graph theory; b i is the connection weight parameter between the i - th follower agent and the leader agent; Λi represents the neighbor set of the i - th agent.

[0088] In step three, based on the knowledge of optimal control and combined with the defined formation error, the expression form of the cost function is obtained as:

[0089]

[0090] where C = diag{c1, c2,..., c i ,..., c n}; w1 and w2 are two set constants; I m is an identity matrix of an appropriate dimension; is the symbol of the tensor product.

[0091] According to the obtained cost function, the corresponding value function is established, and the optimal control input is introduced Finally, the corresponding optimal value function expression is as follows:

[0092]

[0093] where τ represents the integral constant.

[0094] In step four, based on the established optimal value function, the distributed solution is used to obtain the Hamilton - Jacobi - Bellman equation as follows:

[0095]

[0096] Take the partial derivative of the above equation with respect to the optimal control input , and the expression form of the optimal control input is obtained as follows:

[0097]

[0098] It can be seen from this formula that the required optimal control input in this embodiment is a quantity related to the derivative of the optimal value function. However, due to the non-linearity of the multi-agent system and the unknown model, it is actually very difficult to solve the derivative term of the optimal value function, which will also lead to the difficulty in solving the optimal control input.

[0099] Step 5: The neural network algorithm has been proven to have a powerful approximation effect and can approximately estimate non-linear functions. For the unknown non-linear term f i (x i ) existing in the multi-agent system, approximate estimation is carried out by introducing a neural network:

[0100]

[0101] Among them, represents the ideal neural network weight matrix; S fi (x i ) represents the basis function vector; ∈ fi (x i ) represents the approximation error.

[0102] Since is only used for theoretical analysis but is an unknown matrix in practice, an estimation matrix is introduced for estimation, and the approximated by the neural network identifier is obtained as follows:

[0103]

[0104] represents the approximate function of the actual non-linear function f i (x i ) generated by introducing the neural network method.

[0105] According to the obtained approximate function , the estimated values corresponding to some variables related to and etc. can be obtained.

[0106] And the estimation matrix needs to be updated, and its corresponding update law can be expressed in the following form by design:

[0107]

[0108] Among them, T i represents a positive definite matrix; θ i represents a positive constant.

[0109] Step Six: From the approximate variables obtained through the neural network above, the optimal value function is segmented to obtain a segmented expression form of the optimal value function as follows:

[0110]

[0111] Among them, k i represents a constant term greater than zero; and The expression of

[0112] Through this separated value function expression, combined with the optimal control expression form The following separated expression form for the optimal control is obtained:

[0113]

[0114] Step Seven: By introducing a single-critic reinforcement learning structure, the segmented optimal value function and the optimal control input are approximately evaluated to obtain the following expression:

[0115]

[0116]

[0117] Among them, represents the introduced estimated critic parameter matrix; S i represents the radial basis function of the neural network. In the traditional actor-critic reinforcement learning structure, the actor network needs to execute control actions in the controller, while the critic network only needs to evaluate the control actions in the optimal value function and feedback to the actor network for correction. In the present invention, by design, the actor network is removed, and in addition to evaluating control actions, the critic neural network also needs to undertake the responsibilities of the actor network in the traditional method, that is, to execute control actions.

[0118] And the parameter matrix in the critic network also needs to be updated, and the expression form of its update law is:

[0119]

[0120] Among them, k ci represents the learning rate of the critic network, and φ i The specific expression of

[0121]

[0122] To verify that the optimal control input based on the single-critic reinforcement learning structure proposed in this embodiment can achieve the formation movement behavior of the nonlinear multi-agent system, corresponding simulation experiments are carried out here. The expression form of the nonlinear multi-agent system given is as follows:

[0123]

[0124] where h i = -0.7, 0.1, -0.5, 0.1; And the initial positions of the four follower agents are set as x i (0) = [4, 4] T , [-4, 4] T , [4, -4] T , [-4, -4] T .

[0125] The expected movement trajectory set by the leader agent is:

[0126]

[0127] The initial position of the leader agent among them is [0, 0] T .

[0128] The information exchange between agents needs to use the relevant knowledge of graph theory. Matrix A is the communication weight matrix used to describe the communication between follower agents and their neighbor follower agents, and its expression is:

[0129]

[0130] And matrix B is used to represent the communication weight matrix between follower agents and the leader agent, and its expression is:

[0131] B = diag{1, 0, 0, 0}

[0132] Figure 3 Fig. shows the communication topology graph of the nonlinear multi-agent system in the embodiment of the present invention. The multi-agent systems in the embodiment of the present invention and the traditional actor-critic method used for comparison will use this communication topology as the way of communication between agents, so as to facilitate comparison. Figure 4 This is the multi-agent formation trajectory shown in the embodiment of the present invention. It can be seen that the four follower agents can follow the trajectory of the leader agent well for movement. Figure 5 Fig. shows the formation speed trajectory of the multi-agent system in the embodiment of the present invention, and the four follower agents among them can keep up with the speed of the leader agent. Figure 6It is a comparison chart of the position error in the embodiment of the present invention and the position error in the traditional actor-critic method. It can be seen that the position error of the method proposed by the present invention is smaller than that obtained by the traditional method. Figure 7 It shows a comparison chart of the speed error in the embodiment of the present invention and the speed error in the traditional method. It can be seen that the speed errors of the two methods are relatively close. Figure 8 It shows a comparison chart of the calculation time of the actor-critic method and the single-critic reinforcement learning structure in the embodiment of the present invention. It can be seen that the calculation time of the single-critic reinforcement learning structure method proposed by the present invention is shorter than that of the actor-critic method, and as the number of iterations increases, the reduced time is more.

[0133] As described above, it is only a preferred embodiment of the present invention, and it is not a limitation of the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still belong to the protection scope of the technical solution of the present invention.

Claims

1. A multi-agent formation control method based on a single critic reinforcement learning structure, characterized in that It includes the following steps: Step 1: Based on graph theory in applied mathematics, construct the communication structure of each agent in the multi-agent system. Considering the system as a first-order multi-agent system, each agent only obtains the position information of its neighboring agents. At the same time, there is a leader agent in the system, and other agents, as followers, move along the trajectory of the leader agent during operation; Step 2: For each agent in the system, construct its tracking error relative to the leader agent according to the information of its neighboring agents obtained, and construct the error describing the agent and the leader as well as the agent and its neighboring agents, that is, the formation error, according to the tracking error; Step 3: Based on optimal control, construct a cost function and a value function related to the formation error and the optimal control input; Step 4: Based on Taylor's formula and the value function obtained in Step 3, expand and solve the value function to obtain the corresponding Hamilton-Jacobi-Bellman equation; Step 5: For the Hamilton-Jacobi-Bellman equation obtained in Step 4, solve the partial derivative with respect to the optimal control to obtain the expression form of the optimal control input with respect to the optimal value function; Step 6: Divide the optimal value function to obtain its expression form with respect to the formation error and the unknown function, and obtain the divided optimal control input form according to the expression form of the optimal control input in Step 5; Step 7: Introduce a single critic reinforcement learning structure and combine it with a neural network to solve the divided optimal value function and the optimal control input obtained in Step 6, where the neural network approximates the unknown non-linear terms in the multi-agent system, and the critic network performs the formation control of the agent system and evaluates and improves the effect of the formation control; For the unknown nonlinear term f i (x i ) existing in the multi-agent system, approximate estimation is carried out by introducing a neural network: Among them, represents the ideal neural network weight matrix; S fi (x i ) represents the basis function vector; ∈ fi (x i ) represents the approximation error; Since it is only used for theoretical analysis but is an unknown matrix in practice, an estimation matrix is introduced for estimation, and the approximated by the neural network identifier is obtained as follows: Based on the obtained approximate function obtain the estimated values of other variables; The optimal value function and the optimal control input are converted into the following expression forms by dividing the parameters: where k i represents a constant term greater than zero; and The expression of The expressions of the optimal value function and the optimal control input after introducing the single-critic reinforcement learning structure are as follows: Among them, represents the introduced estimated critic network parameter matrix; S i represents the radial basis function of the neural network; the update law of the critic network parameter matrix is expressed as follows: where k ci represents the learning rate of the critic network, and the specific expression of φ i is as follows:

2. The multi-agent formation control method based on a single critic reinforcement learning structure according to claim 1, wherein, The single critic reinforcement learning structure is used to remove the requirement for the actor network in the traditional actor-critic reinforcement learning method, thereby effectively reducing the approximation error of the system and reducing the calculation time.

3. A multi-agent formation control method based on a single critic reinforcement learning structure according to claim 1, characterized in that, In Step 1, the model of the multi-agent system is expressed as: where x i (t) represents the position of the i-th agent in the system; u i (t) represents the control input of the i-th agent in the system; f i (·) represents an unknown nonlinear function, and it is assumed to be Lipschitz continuous; The model of the leader agent is as follows: where p l and v l respectively represent the trajectory and velocity of the leader, i.e., the desired trajectory and velocity during formation movement; Set the tracking error of each agent relative to the leader as: z i = x i - p l - ζ i Among them, ζ i represents the position between the leader agent and the i-th follower agent, and is used to describe the formation shape of the system; According to the structure of the tracking error, define the formation error form as follows: where a ij is the element in the \(i\)-th row and \(j\)-th column of the adjacency matrix in graph theory; \(b\) i is the connection weight parameter between the \(i\)-th follower agent and the leader agent; \(\varLambda\) i represents the neighbor set of the \(i\)-th agent.

4. A multi-agent formation control method based on a single critic reinforcement learning structure according to claim 3, characterized in that, In Step 3, combined with the defined formation error, the expression form of the cost function is obtained as: Among them, C = diag{c1, c2,..., c i ,..., c n}; w1 and w2 are two set constants; I m is an identity matrix of an appropriate dimension; is the symbol of the tensor product; Based on the obtained cost function, a corresponding value function is established, and the optimal control input is introduced Finally, the corresponding optimal value function is expressed as follows: where τ represents the integral constant.

5. A multi-agent formation control method based on a single critic reinforcement learning structure according to claim 4, characterized in that In Step 4, establish the Hamilton-Jacobi-Bellman equation as follows: Taking the partial derivative of the above equation with respect to yields the expression for the optimal control input as follows:

Citation Information

Patent Citations

  • Multi-agent formation control method based on actor-reviewer reinforcement learning and fuzzy logic

    CN111897224A

  • Random nonlinear multi-agent reinforcement learning optimization formation control method

    CN114740710A

Cited By

  • Multi-agent distributed formation control method and system based on event triggering

    CN121348739A