Simplified Reinforcement Learning Formation Control Method for Multi-Agent State Delay Systems
By constructing the communication topology structure of the multi-agent system and simplified reinforcement learning algorithm, combining the fuzzy logic system and the Liyapunov-Krasovsky functional, a simplified reinforcement learning update law and optimal controller are designed, and the computational complexity problem in the multi-agent state time-delay system is solved and stable formation control is achieved.
Patent Information
- Application Number
- CN202310073781.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-19
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-01-19
AI Technical Summary
The prior art is difficult to effectively solve the computational complexity and high computational burden of reinforcement learning update law in multi-agent state time delay systems, and it is difficult to achieve smooth formation control under the condition of state time delays.
By constructing the communication topology structure of multi-agent systems, establishing assumptions, constructing tracking errors and formation errors, combining optimal control theory and simplified reinforcement learning algorithms, using fuzzy logic systems and the Liyapunov-Krasovsky functionals, a simplified reinforcement learning update law and optimal controller are designed to offset the impact of state delays.
The calculation amount of multi-agent formation control is reduced, ensuring stable formation control effect is achieved under state time delay conditions, and reducing calculation complexity and calculation burden.
Smart Images

Figure CN116027790B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-agent formation control, and particularly relates to a simplified reinforcement learning formation control method for a multi-agent state time-delay system. Background Art
[0002] The formation problem of multi-agents has become an attractive topic due to its wide applications in fields such as unmanned aerial vehicles, unmanned ships, and satellite clusters. Among them, the leader-follower algorithm is one of the main methods for formation control and is widely used in multi-agent formation control due to its simplicity and scalability. The main feature of the leader-follower algorithm is that the trajectory of the follower is determined by the trajectory of the leader.
[0003] Optimal control realizes the control objective by minimizing the performance index function to balance performance and resources. However, the Hamilton-Jacobi-Bellman (HJB) equation in optimal control is difficult to solve to obtain the optimal control law when there is non-linearity in the system. To solve this problem, the optimal control theory can be combined with the reinforcement learning algorithm. To solve the problem that reinforcement learning requires a complete understanding of the system dynamics, a fuzzy logic system can be introduced because the fuzzy logic system has been proven to have good approximation performance and learning ability. Although reinforcement learning and the fuzzy logic system can solve the non-linearity in the system, in an actual system, there often exists state time delay. In addition, there are also problems in the reinforcement learning algorithm, such as the difficulty in obtaining the update laws of the actor network and the critic network and the large amount of calculation. Summary of the Invention
[0004] The purpose of the present invention is to provide a simplified reinforcement learning formation control method for a multi-agent state time-delay system, which can reduce the amount of calculation and ensure the smooth execution of the formation control of the multi-agent state time-delay system.
[0005] To achieve the above purpose, the technical solution adopted by the present invention is: a simplified reinforcement learning formation control method for a multi-agent state time-delay system, including the following steps:
[0006] Step 1: Construct the communication topology structure in the multi-agent system through graph theory, where the multi-agents can only obtain the information of adjacent multi-agents;
[0007] Step 2: Construct relevant hypothesis conditions for constraint aiming at the state time-delay information existing in the multi-agent system; it is divided into two hypothesis conditions, namely the time limit of the state time delay and the limit of the norm of the state time-delay function;
[0008] Step 3: According to Step 1 and Step 2, obtain the position information of each agent, construct the tracking error between each follower agent and the leader agent, as well as the error between follower agents, and construct the formation error based on the tracking error and the error between agents;
[0009] Step 4: Utilize the idea of optimal control theory to establish the cost function and value function, obtain the Hamilton-Jacobi-Bellman (HJB) equation according to the obtained value function, and solve the HJB equation to obtain the expression form of the optimal controller;
[0010] Step 5: Aiming at the problem that it is difficult to solve by substituting the obtained expression form of the optimal controller back into the HJB equation, apply a simplified reinforcement learning algorithm combined with a fuzzy logic system to reconstruct the optimal controller;
[0011] Step 6: According to the optimal controller established in Step 5, construct a term for offsetting the state time delay existing in the multi-agent system in combination with the Lyapunov-Krasovskii functional and introduce the optimal controller.
[0012] Furthermore, the specific form of the multi-agent state time-delay system is:
[0013]
[0014] where x i (t) is the position information of the i-th multi-agent; u i (t) is the control input of the i-th agent; p i (·) and g i (·) are smooth vector nonlinear functions with uncertainties; τ i is the unknown time delay;
[0015] Construct the following assumption conditions:
[0016] Assumption 1: For the unknown time delay τ i , there exists a known positive constant τ max such that the condition τ i ≤τ max is satisfied;
[0017] Assumption 2: For the term g i (x i (t)), there exists a smooth vector function satisfying the condition
[0018] The tracking error is expressed as follows:
[0019]
[0020] Among them, \(x_0(t)\) represents the moving trajectory of the leader; is the relative position between the \(i\)-th agent and the leader, which is used to represent the formation shape of multi-agent systems;
[0021] The formation error is expressed as follows:
[0022]
[0023] where \(M\) i represents the neighbor set of the \(i\)-th multi-agent; \(a\) ij represents the element in the \(i\)-th row and \(j\)-th column of the adjacency matrix; \(b\) i represents the connection weight between the \(i\)-th agent and the leader.
[0024] Furthermore, the established value function is expressed as follows:
[0025]
[0026] where \(r(w, u)=w\) T w + u T u represents the cost function;
[0027] According to the established value function and cost function, the distributed Hamilton-Jacobi-Bellman equation, that is, the expression form of the HJB equation is:
[0028]
[0029] Through the obtained HJB equation, the expression form of the optimal controller is solved as:
[0030]
[0031] Furthermore, aiming at the problem that it is difficult to solve by substituting the obtained expression form of the optimal controller back into the HJB equation, a simplified reinforcement learning algorithm is applied to reconstruct the optimal controller in combination with a fuzzy logic system. Specifically:
[0032] For the gradient term Separation is carried out, and the specific expression form after separation is as follows:
[0033]
[0034] According to the specific expression form of the obtained gradient term and combined with the expression form of the optimal controller, the expression of the optimal controller is re-expressed as:
[0035]
[0036] where \(k\) i (t) is a design function; and
[0037] Obtain the gradient term After obtaining the expression form of the optimal controller, the actor-critic reinforcement learning method and the fuzzy logic system are applied to estimate it, and the estimated expression is:
[0038]
[0039]
[0040] Wherein, Represents the estimated optimal parameter matrix, Represents the fuzzy basis function vector, And Are used to approximate the unknown nonlinear term p i (x i ) of the multi-agent system; And Respectively represent the estimated parameter matrices of the critic and actor networks, Performs formation control in the optimal controller, Is used to evaluate the control behavior of the actor network and feedback the evaluation result to the actor;
[0041] The specific update laws of the actor and critic networks are as follows:
[0042]
[0043]
[0044] Wherein, k ci > 0 and k ai > 0 are the learning rates of the critic network and the actor network respectively.
[0045] Furthermore, for the positive function term k i (t) existing in the expression form of the optimal controller, the specific expression form is obtained by combining the Lyapunov-Krasovskii functional;
[0046] In the Lyapunov stability proof, the Lyapunov-Krasovskii functional design function is introduced:
[0047]
[0048] Through this design function, combined with the Lyapunov stability proof, the specific expression form of the positive function term k i (t) used to cancel the state time delay of the multi-agent system in the optimal controller is as follows:
[0049] k if(t) = k i0 + k i1 f(t), k i0 >2
[0050]
[0051] Compared with the prior art, the present invention has the following beneficial effects:
[0052] 1. Aiming at the deficiencies of the reinforcement learning update law obtained from the Bellman residual in the multi-agent formation control method based on reinforcement learning, which is computationally complex and has a large computational burden, the present invention proposes a simplified form for obtaining the reinforcement learning update law, simplifies the form for obtaining the multi-agent update law, and reduces the amount of calculation.
[0053] 2. By introducing the Lyapunov-Krasovskii functional into the Lyapunov stability proof, a term is obtained during the proof process to offset the time delay of the multi-agent state, effectively solving the influence of the state time delay on the formation control. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is the flowchart of the method implementation in the embodiment of the present invention;
[0055] Figure 2 is the communication topology diagram of multi-agents in the embodiment of the present invention;
[0056] Figure 3 is the multi-agent formation performance display diagram in the embodiment of the present invention;
[0057] Figure 4 is the information change diagram of the formation error in the embodiment of the present invention;
[0058] Figure 5 is the information change diagram of the cost function in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0059] The present invention will be further described below in conjunction with the drawings and embodiments.
[0060] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0061] Note that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly dictates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprises" and / or "comprising" are used in this specification, they specify the presence of the stated features, steps, operations, devices, components, and / or combinations thereof.
[0062] Starting from the formation requirements of the multi-agent system, this embodiment constructs a simulation of a multi-agent system with a first-order non-linear state time-delay model. Each multi-agent system can obtain information about its neighboring agents. This embodiment provides a simplified reinforcement learning formation control method for a first-order non-linear state time-delay multi-agent system, as Figure 1 shown, including the following steps:
[0063] Step 1: Establish a communication topology structure among multi-agents through graph theory, where each agent only obtains information about its neighboring agents. The multi-agent communication topology graph in this embodiment is as Figure 2 shown.
[0064] Step 2: Construct relevant hypothesis conditions for constraint in view of the state time-delay information existing in the multi-agent system.
[0065] The multi-agent model, that is, the specific form of the multi-agent state time-delay system, is as follows:
[0066]
[0067] where, x i (t) represents the position information of the i-th agent; u i (t) represents the control input of the i-th agent; p i (·) and g i (·) represent smooth vector non-linear functions with uncertainties; τ i represents the unknown time delay.
[0068] In view of the unknown state time-delay terms existing in the established multi-agent system model, the following hypothesis conditions are proposed:
[0069] Hypothesis 1: For the unknown time delay τ i , there exists a known positive constant τ max , satisfying τ i ≤τ max .
[0070] Hypothesis 2: For the g i (x i (t)) term, there exists a smooth vector function satisfying the condition
[0071] Step 3: According to Steps 1 and 2, obtain the position information of each agent, construct the tracking error between each follower agent and the leader agent, as well as the error between follower agents, and construct the formation error based on the tracking error and the error between agents.
[0072] Based on the established multi-agent system model, the tracking error between the follower agent and the leader agent is obtained as follows:
[0073]
[0074] where, \(x_0(t)\) represents the moving trajectory of the leader; represents the relative position between the \(i\)-th agent and the leader agent, and is used to represent the formation shape of multi-agents.
[0075] Through the established tracking error expression, the expression of the formation error can be established as follows:
[0076]
[0077] where, \(M\) i represents the neighbor set of the \(j\)-th multi-agent; \(m\) j \((t)\) represents the \(j\)-th agent adjacent to the \(i\)-th agent; \(a\) ij represents the element in the \(i\)-th row and \(j\)-th column of the adjacency matrix; \(b\) i represents the connection weight between the \(i\)-th agent and the leader agent.
[0078] Step 4: Utilize the idea of optimal control theory to establish the cost function and the value function, and based on the obtained value function, obtain the Hamilton-Jacobi-Bellman (HJB) equation, and solve the HJB equation to obtain the expression form of the optimal controller.
[0079] According to the optimal control theory, the expression form of the value function is established as follows:
[0080]
[0081] where, \(r(w, u)=w\) T \(w + u\) T \(u\) represents the cost function, which is composed of formation error information and control input information.
[0082] According to the established expression forms of the cost function and the value function, the expression form of the distributed Hamilton-Jacobi-Bellman (HJB) equation can be calculated according to the optimal control theory as:
[0083]
[0084] Among them, represents the optimal value function.
[0085] By taking the partial derivative of the obtained HJB equation with respect to , the expression of can be obtained as:
[0086]
[0087] In order to obtain a more specific expression of , the conventional method is to substitute it back into the HJB equation to find the specific expression of . However, due to the existence of unknown dynamics and state delays in the multi-agent system, it is difficult or even impossible to obtain the expression of .
[0088] Step 5: To address the problem that it is difficult to solve the HJB equation by substituting the obtained expression of the optimal controller back, a simplified reinforcement learning algorithm is applied in combination with a fuzzy logic system to reconstruct the optimal controller.
[0089] For the gradient term , its separated expression is as follows:
[0090]
[0091] where k i (t) represents a design function;
[0092] According to the separated expression of the obtained gradient term , substituting it back into the expression of the gradient term , the following separated expression of the optimal controller can be obtained:
[0093]
[0094] In the separated expressions of the gradient term and the optimal controller, there are dynamic terms p i (x i (t)) of the multi-agent system that are unknown and it is still difficult to directly solve the HJB equation. Therefore, a reinforcement learning method is applied in combination with a fuzzy logic system to obtain an approximate solution of the optimal controller, and the obtained approximate expression is as follows:
[0095]
[0096]
[0097] Among them, denotes the estimated optimal parameter matrix, represents the fuzzy basis function vector, and is used to approximate the unknown non - linear term p i (x i ) of the multi - agent system; and respectively denote the estimated parameter matrices of the critic and actor networks, perform formation control in the optimal controller, evaluate the control behavior of the actor network and feedback the evaluation results to the actor.
[0098] In the conventional reinforcement learning algorithm, and the update laws need to be obtained from the square term of the Bellman residual by using the gradient descent method, and the obtained update laws are relatively complex and have a large amount of calculation. To solve this problem, this patent discloses a simplified method to simplify the derivation of the update laws. First, take the partial derivative of the HJB equation to obtain the following formula:
[0099]
[0100] According to the formula obtained above, a function equivalent to the HJB equation can be established, and this function is used to replace the Bellman residual of the conventional method. The expression of this function is as follows:
[0101]
[0102] According to this function, take its derivative, and finally the update laws of and can be designed and obtained as:
[0103]
[0104]
[0105] where, k ci > 0 and k ai > 0 are the learning rates of the critic and actor networks respectively.
[0106] Step six: Through the reinforcement learning algorithm, the non - linear problem in the system is solved in the optimal controller design, but there is still the influence of unknown state time - delay. To solve this problem, the Lyapunov - Krasovskii functional is introduced in the stability proof, and its design form is as follows:
[0107]
[0108] Through this Lyapunov-Krasovskii functional, the positive function term \(k(t)\) in the optimal controller for canceling the state time-delay term of the multi-agent system is obtained after the stability proof. i The specific expression form of \(k(t)\) is as follows:
[0109] k i (t)=k i0 +k i1 (t), where \(k\) i0 >2(16)
[0110]
[0111] Step 7: To prove that the obtained expression form of the optimal controller can achieve multi-agent leader-follower formation control, a simulation experiment is conducted in this example, and the specific model expression form of the multi-agent system is given as follows:
[0112]
[0113] where \(\alpha\) i = 0.7, -0.1, 0.5, -0.4; \(\beta\) i = 0.8, 0.7, -1.2, -1.4; \(h(x(t))=\rho x\cos(x(t))\); \(h(x(t))=\rho x\sin(x_2(t))\), where \(\rho\) i1 (x i1 (t))=\rho i1 x i1 \cos(x i1 (t)); \(h(x i2 (t))=\rho i2 x i2 \sin(x_2(t)), where \(\rho\) i2 = 0.7, 0.5, 0.6, 0.4, \(\rho\) i1 = 0.8, 0.4, 0.9, 0.5; the time-delay term \(\tau\) i2 = 0.2, 0.4, 0.8, 0.7, and thus \(\tau\) i = 1, and max In addition, the trajectory of the leader is set as follows:
[0114]
[0115]
[0116] And the initial position of the leader is set at \(x_0(0)=[0,0]\) T i T T (0)=[5,4] T , [4, -5] T , [-4, 5] T, [-5, -4] T 。
[0117] According to the relevant knowledge of graph theory, there is also a matrix A representing the communication weight relationship between follower agents and a matrix B describing the communication weight relationship between follower agents and leader agents. The values of matrices A and B are as follows
[0118]
[0119] B = diag{1, 0, 0, 0}
[0120] Figure 2 This is the multi-agent communication topology graph in this embodiment. The communication of the agent system in this embodiment is carried out according to this topology graph. Figure 3 This is the graph showing the performance of multi-agent formation in this embodiment. It can be seen that the embodiments of the present invention can effectively ensure that the system can still achieve the desired formation performance effect under the condition of the existence of state time delay. Figure 4 This is the graph showing the change of formation error information in this embodiment. It can be seen that the present invention can ensure that the formation error of the system is maintained within an acceptable range. Figure 5 This is the graph showing the change of cost function information in this embodiment, which reflects the information change of control input and error during the formation process. Since the reinforcement learning update law of the present invention is simplified, the overall calculation amount and information change can be kept within a small range.
[0121] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still belong to the protection scope of the technical solution of the present invention.
Claims
1. A simplified reinforcement learning formation control method for multi-agent state time-delay systems, characterized in that It includes the following steps: Step 1: Construct the communication topology structure in the multi-agent system through graph theory; Step 2: Construct relevant hypothesis conditions for constraint aiming at the state time-delay information existing in the multi-agent system; Step 3: According to Step 1 and Step 2, obtain the position information of each agent, construct the tracking error between each follower agent and the leader agent, and the error between follower agents, and construct the formation error according to the tracking error and the error between agents; Step 4: Utilize the idea of optimal control theory, establish the cost function and the value function, obtain the HJB equation according to the obtained value function, and solve the HJB equation to get the expression form of the optimal controller; Step 5: Aiming at the problem that it is difficult to solve by substituting the obtained expression form of the optimal controller back into the HJB equation, apply the simplified reinforcement learning algorithm combined with the fuzzy logic system to reconstruct the optimal controller; Step 6: According to the optimal controller established in Step 5, combine the Lyapunov-Krasovskii functional to construct a term for offsetting the state time-delay existing in the multi-agent system and introduce the optimal controller; The specific form of the multi-agent state time-delay system is: where, x i (t) is the position information of the i-th multi-agent; u i (t) is the control input of the i-th agent; p i (·) and g i (·) are smooth vector nonlinear functions with uncertainties; τ i is the unknown time delay; Construct the following hypothesis conditions: Hypothesis 1: For an unknown time delay τ i , there exists a known positive constant τ max , satisfying the condition that τ i ≤τ max ; Hypothesis 2: Regarding item g i (x i (t)), there exists a smooth vector function satisfying the condition The tracking error is expressed as follows: Among them, \(x_0(t)\) represents the moving trajectory of the leader; is the relative position between the \(i\)-th agent and the leader, which is used to represent the formation shape of multi-agent systems; The formation error is expressed as follows: Among them, M i represents the neighbor set of the i-th multi-agent; a ij represents the element in the i-th row and j-th column of the adjacency matrix; b i represents the connection weight between the i-th agent and the leader; The expression form of the established value function is as follows: where r(w, u) = w T w + u T u represents the cost function; According to the established value function and cost function, calculate and obtain the distributed Hamilton-Jacobi-Bellman equation, that is, the expression form of the HJB equation is: Through the obtained HJB equation, solve and obtain the expression form of the optimal controller as: Aiming at the problem that it is difficult to solve by substituting the obtained expression form of the optimal controller back into the HJB equation, apply the simplified reinforcement learning algorithm combined with the fuzzy logic system to reconstruct the optimal controller, specifically: For the gradient term Separation is performed, and the specific expression after separation is as follows: According to the specific expression form of the obtained gradient term, combined with the expression form of the optimal controller, re-express the optimal controller expression as: where k i (t) is a design function; and Obtain the gradient term After obtaining the expression form of the optimal controller, the actor-critic reinforcement learning method and the fuzzy logic system are applied to estimate it, and the estimated expression is: Among them, represents the estimated optimal parameter matrix, represents the fuzzy basis function vector, and are used to approximate the unknown nonlinear term p i (x i ) of the multi-agent system; and respectively represent the estimated parameter matrices of the critic and actor networks, perform formation control in the optimal controller, evaluate the control behavior of the actor network and feedback the evaluation results to the actor; The specific form of the update law of the actor and critic networks is as follows: where k ci > 0 and k ai > 0 are the learning rates of the critic network and the actor network, respectively; For the positive function term k i (t) existing in the expression form of the optimal controller, the specific expression form is obtained by combining the Lyapunov-Krasovskii functional; In the Lyapunov stability proof, introduce the Lyapunov-Krasovskii functional design function: Through this design function, combined with the Lyapunov stability proof, the positive function term k i (t) used to offset the state time delay of the multi-agent system in the optimal controller is obtained, and its specific expression is as follows: k i k(t) = k i0 + k i1 k(t), k i0 > 2
Citation Information
Patent Citations
Multi-agent formation control method based on actor-reviewer reinforcement learning and fuzzy logic
CN111897224A
Distributed control of heterogeneous multi-agent systems
US10983532B1