Multi-agent system formation control method and device based on fuzzy reinforcement learning
By using fuzzy reinforcement learning and the recognizer-actor-critic reinforcement learning algorithm, and dynamically switching fuzzy basis functions, the problems of convergence accuracy and speed in formation control of multi-agent systems are solved, and efficient formation control is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YANSHAN UNIV
- Filing Date
- 2026-02-02
- Publication Date
- 2026-05-08
AI Technical Summary
In existing multi-agent system formation control, the convergence accuracy and speed of the solution process need to be improved, especially when applying optimal control methods, the analytical solution of the HJB equation is difficult to obtain.
A fuzzy reinforcement learning-based approach is adopted. The communication topology is established through graph theory, the neighbor set and formation error are determined, and the gradient decomposition of the HJB equation is approximated by the fuzzy logic system in combination with the optimal control strategy. Furthermore, the discriminator-actor-critic reinforcement learning algorithm is introduced to dynamically switch the fuzzy basis function to achieve a balance between convergence accuracy and speed.
It significantly improves the formation control performance and adaptability of multi-agent systems, solves the problem of the difficulty in finding analytical solutions to the HJB equation, achieves a balance between convergence accuracy and speed, and enhances the control effect in multi-robot environments.
Smart Images

Figure CN121995737A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-agent system formation control technology, and in particular to a multi-agent system formation control method and apparatus based on fuzzy reinforcement learning. Background Technology
[0002] Over the past decade, multi-robot systems, as a typical type of multi-agent system, have demonstrated greater fault tolerance and robustness compared to single-robot systems due to their higher redundancy. They can also accomplish many tasks that a single robot cannot perform independently through collaboration. Robot formation, as a control strategy for achieving cooperative tasks in multi-robot systems, has become an important research direction.
[0003] The leader-follower method is one type of distributed control approach that can be flexibly applied to the formation control of multi-robot systems and is highly user-friendly. Applying optimal control theory to multi-robot formation control allows for optimization between control performance and resource consumption by minimizing the cost function. Traditional optimal control is typically achieved by solving the Hamilton-Jacobi-Bellman (HJB) equations, but due to the nonlinearity of these equations, finding analytical solutions is extremely complex.
[0004] To overcome the difficulties in applying optimal control methods to formation control, existing studies have considered using reinforcement learning and fuzzy logic systems for solving the problem. However, the convergence accuracy and speed of the solution process still need further improvement. Summary of the Invention
[0005] This invention provides a method and apparatus for formation control of a multi-agent system based on fuzzy reinforcement learning, in order to address the problem that the convergence accuracy and speed of the solution process need to be improved.
[0006] In a first aspect, embodiments of the present invention provide a multi-agent system formation control method based on fuzzy reinforcement learning, comprising: Based on graph theory, a communication topology for a multi-agent system is established, and a neighbor set for each agent is determined according to the communication topology. The neighbor set includes all the neighbor agents of the agent. The formation error is determined based on the position status information of each agent, the position status information of each agent's neighboring agents, and the desired trajectory of the lead agent. To address the formation error, an optimal control strategy is introduced. This involves solving the HJB equation to obtain a first optimal value function expression and an optimal virtual controller expression that minimize the first value function. A first fuzzy basis function is then determined based on the formation error to approximate the system uncertainty and estimation error compensation terms after gradient decomposition of the first optimal value function expression. Following the gradient decomposition and approximation of the first optimal value function expression, optimization is performed using the first fuzzy basis function and a recognizer-actor-critic reinforcement learning algorithm. Based on the optimization results, an approximate value of the optimal virtual controller expression is obtained. The virtual error is determined based on the velocity state information of each agent and the approximation of the optimal virtual controller. An optimal control strategy is introduced to address this virtual error. The second optimal value function expression and the optimal controller expression, which minimize the second value function, are obtained by solving the HJB equation. A second fuzzy basis function is determined based on the formation error to approximate the system uncertainty and estimation error compensation terms after gradient decomposition of the second optimal value function expression. The approximation is then performed using the second fuzzy basis function and a recognizer-actor-critic reinforcement learning algorithm, according to the gradient decomposition of the second optimal value function expression. An approximation of the optimal controller expression is obtained based on the optimization result. A disturbance observer is determined based on the approximation of the optimal virtual controller. Finally, the optimal controller for the multi-agent system is obtained based on the approximation of the optimal controller expression and the disturbance observer.
[0007] In one possible implementation, the first fuzzy basis function includes: a first discriminator fuzzy basis function and a first actor-critic fuzzy basis function; Determining the first fuzzy basis function based on the formation error includes: Calculate the norm of the formation error; According to the norm, Determine the switching coefficient; Based on the switching coefficient, the first preset identifier fuzzy basis function, and the second preset identifier fuzzy basis function, according to Determine the fuzzy basis function of the first identifier; Based on the switching coefficients, the first preset actor-critic fuzzy basis function, and the second preset actor-critic fuzzy basis function, according to... Determine the first actor-critic fuzzy basis function; in, The switching coefficient is... Let the norm be... To set a threshold, Let be the fuzzy basis function of the first identifier. The first preset identifier's fuzzy basis function is... , For the first in a multi-agent system The location and state information of each agent. The center of the fuzzy basis function of the first preset identifier. The width of the first preset identifier's fuzzy basis function. The second preset identifier's fuzzy basis function, , The center of the fuzzy basis function of the second preset identifier. The width of the second preset identifier's fuzzy basis function. Let be the first actor-critic fuzzy basis function. Let the first preset actor-critic fuzzy basis function be used. , As the first state variable, Let the first preset actor-critic fuzzy basis function center be , The width of the first preset actor-critic fuzzy basis function. Let the first preset actor-critic fuzzy basis function be used. , The second preset actor-critic fuzzy basis function center, The width of the second preset actor-critic fuzzy basis function.
[0008] In one possible implementation, the gradient decomposition and approximation of the first optimal value function expression is as follows: ; The gradient of the first optimal value function expression. and Let them be two constants greater than 0. The formation error is... Functions designed to avoid singularity issues, It is a positive constant. It is a constant. , , This is the parameter matrix of the first optimal identifier. The first optimal commentator parameter matrix, This refers to the first identifier fuzzy basis function in the first fuzzy basis function. Let be the first actor-critic fuzzy basis function in the first fuzzy basis function. The approximate error is satisfied. ,in and These are two positive constants.
[0009] In one possible implementation, the optimal virtual controller expression is optimized according to the gradient decomposition and approximation form of the first optimal value function expression, based on the first fuzzy basis function and the identifier-actor-critic reinforcement learning algorithm. The approximate value of the optimal virtual controller expression is obtained based on the optimization result, including: The expression for the optimal virtual controller is transformed according to the gradient decomposition and approximation of the first optimal value function expression: ; Based on the gradient decomposition and approximation form of the first optimal value function expression and the transformed optimal virtual controller expression, the expression of the identifier-actor fuzzy logic system module is determined as follows: ; In the formula, This is an approximation of the optimal virtual controller expression. For the identifier parameter vector, For actor parameter vectors; The expression for the identifier-evaluator fuzzy logic system module is: ; In the formula, This is an estimate of the gradient of the first optimal value function expression. For the commentator parameter vector; The parameter vector update law of the fuzzy logic system module of the identifier is as follows: ; In the formula, for The derivative, express of OK Column elements, The learning rate of the fuzzy logic system module of the identifier. express of Row elements, Indicates the formation error Column elements, and It is a positive design constant used to control the update magnitude of weights and avoid over-updating; The parameter vector update law of the fuzzy logic system module is as follows: ; In the formula, for The derivative, Let be the learning rate of the fuzzy logic system module for the commentator, and be a constant greater than 0. To improve training speed, among other things, for An identity matrix of order 1; The parameter vector update law of the actor fuzzy logic system module is as follows: ; In the formula, for The derivative, Let be the learning rate of the actor's fuzzy logic system module, and be a constant greater than 0. Optimize according to the parameter vector update laws to determine the first optimal identifier parameter matrix and the first optimal critic parameter matrix. Substitute the first optimal identifier parameter matrix and the first optimal critic parameter matrix into the transformed optimal virtual controller expression to obtain an approximate value of the optimal virtual controller expression.
[0010] In one possible implementation, determining the virtual error based on the velocity state information of each agent and an approximation of the optimal virtual controller includes: according to Determine the virtual error; in, The virtual error, For the first The velocity state information of each agent This is an approximation of the optimal virtual controller.
[0011] In one possible implementation, the second fuzzy basis function includes: a second identifier fuzzy basis function and a second actor-critic fuzzy basis function; Determining the second fuzzy basis function based on the formation error includes: Calculate the norm of the formation error; According to the norm, Determine the switching coefficient; Based on the switching coefficient, the third preset identifier fuzzy basis function, and the fourth preset identifier fuzzy basis function, according to Determine the fuzzy basis function of the second identifier; Based on the switching coefficients, the third preset actor-critic fuzzy basis function, and the third preset actor-critic fuzzy basis function, according to... Determine the second actor-critic fuzzy basis function; in, The switching coefficient is... Let the norm be... To set a threshold, The second identifier's fuzzy basis function is... The third preset identifier fuzzy basis function, , , , The third preset identifier's fuzzy basis function center, The width of the fuzzy basis function of the third preset identifier. The fourth preset identifier fuzzy basis function, , The center of the fuzzy basis function of the fourth preset identifier, The width of the fuzzy basis function of the fourth preset identifier. Let the second actor-critic fuzzy basis function be , The third preset actor-critic fuzzy basis function, , , The third preset actor-critic fuzzy basis function center, The width of the third preset actor-critic fuzzy basis function. For the fourth preset actor-critic fuzzy basis function, , The fourth preset actor-critic fuzzy basis function center, The width of the fourth preset actor-critic fuzzy basis function.
[0012] In one possible implementation, determining the disturbance observer based on an approximation of the optimal virtual controller includes: according to Determine the disturbance observer; In the formula, for The derivative, For the disturbance observer of Column elements, and Used to control the convergence speed and stability of the estimator, it is a constant greater than 0. , , =1, It is 0.7. It is 0.01. for The derivative, for of Column elements; The optimal controller for a multi-agent system is obtained based on an approximation of the optimal controller expression and the disturbance observer, including: according to To obtain the optimal controller for a multi-agent system; In the formula, The optimal controller for a multi-agent system. This is an approximation of the optimal controller expression. This refers to the disturbance observer.
[0013] Secondly, embodiments of the present invention provide a multi-agent system formation control device based on fuzzy reinforcement learning, comprising: The modeling module is used to establish the communication topology of the multi-agent system based on graph theory, and to determine the neighbor set of each agent based on the communication topology. The neighbor set includes all the neighbor agents of the agent. The processing module is used to determine the formation error based on the position status information of each agent, the position status information of each agent's neighboring agents, and the desired trajectory of the lead agent. The first control module is used to introduce an optimal control strategy for the formation error. It obtains a first optimal value function expression and an optimal virtual controller expression that minimize the first value function by solving the HJB equation. It also determines a first fuzzy basis function based on the formation error to approximate the system uncertainty term and the estimated error compensation term after gradient decomposition of the first optimal value function expression. Furthermore, it optimizes the first optimal value function expression according to the gradient decomposition and approximation form, based on the first fuzzy basis function and the identifier-actor-critic reinforcement learning algorithm, and obtains an approximate value of the optimal virtual controller expression based on the optimization result. The second control module is used to determine the virtual error based on the velocity state information of each agent and the approximation of the optimal virtual controller; introduce an optimal control strategy for the virtual error; obtain the second optimal value function expression and the optimal controller expression that minimize the second value function by solving the HJB equation; determine the second fuzzy basis function based on the formation error to approximate the system uncertainty term and the estimation error compensation term after gradient decomposition of the second optimal value function expression; optimize the second fuzzy basis function and the identifier-actor-critic reinforcement learning algorithm according to the gradient decomposition and approximation form of the second optimal value function expression; obtain the approximation of the optimal controller expression based on the optimization result; determine the disturbance observer based on the approximation of the optimal virtual controller; and obtain the optimal controller of the multi-agent system based on the approximation of the optimal controller expression and the disturbance observer.
[0014] Thirdly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect or any possible implementation thereof.
[0015] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect or any possible implementation thereof.
[0016] In this embodiment of the invention, on the one hand, the first fuzzy basis function and the second fuzzy basis function are determined by the formation error, and the two different sets of "fuzzy rules" are dynamically switched. When the formation error is large, one set of rules is used for rapid convergence, and when the formation error is small, the other set of rules is used for fine adjustment, thereby achieving a balance between convergence accuracy and speed. On the other hand, by utilizing the ability of fuzzy logic systems to approximate continuous functions, the problem of the difficulty in finding analytical solutions to the HJB equation in optimal control methods is solved. Combined with the identifier-actor-critic reinforcement learning algorithm, the challenge of unknown optimal parameter vectors is effectively overcome. The parameter vectors of the identifier fuzzy logic system module, the actor fuzzy logic system module, and the critic fuzzy logic system module are updated in real time by gradient descent, which significantly improves the adaptability and control performance of the algorithm in multi-robot environments. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the implementation of the multi-agent system formation control method based on fuzzy reinforcement learning provided in this embodiment of the invention. Figure 2 This is a flowchart illustrating the implementation of a multi-agent system formation control method based on fuzzy reinforcement learning, provided in another embodiment of the present invention. Figure 3 This is a schematic diagram illustrating the formation control effect provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the formation control effect provided by another embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of a multi-agent system formation control device based on fuzzy reinforcement learning provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0018] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0019] See Figure 1 The document illustrates a flowchart of the implementation of a multi-agent system formation control method based on fuzzy reinforcement learning provided in an embodiment of the present invention, detailed below: Step 101: Based on graph theory, establish the communication topology of the multi-agent system, and determine the neighbor set of each agent according to the communication topology. The neighbor set includes all the neighbor agents of the agent.
[0020] For example, in a multi-robot system, a robot acts as an agent and establishes a communication topology between the multi-robot system through graph theory. Each robot can obtain the state data of its neighboring robots (i.e., neighboring agents).
[0021] In this embodiment, the task of the multi-robot system is to achieve motion control based on the preset trajectory of the lead robot (i.e., the lead agent). The follower robots obtain the state information of their neighboring robots or the lead robot through a communication topology constructed using graph theory, and then perform tracking operations in a specific formation. The completion of the system task is marked by the stability of the formation structure, that is, the relative positions of the robots remain unchanged, and the movement speed of the follower robots gradually approaches that of the lead robot, forming a coordinated and consistent operating state.
[0022] The information acquired by the robot can include the following categories: the status information of neighboring robots. If the follower robot communicates with the navigator robot, the follower robot can acquire the status information of the navigator robot.
[0023] For example, a multi-robot system can take the following specific forms: ; In the formula, for With respect to the time derivative, for With respect to the time derivative, For the first A robot Location status information at any given time. For the first A robot Velocity status information at any given moment For an unknown first continuous function, For an unknown second continuous function, For the first The controller of a robot This is an external disturbance.
[0024] Step 102: Determine the formation error based on the position status information of each agent, the position status information of each agent's neighboring agents, and the desired trajectory of the lead agent.
[0025] In this embodiment, after establishing a multi-agent communication topology and determining the neighbor set of each agent based on the communication topology, the obtained state data of the neighbor agents are used to construct the formation error.
[0026] For example, taking a multi-robot system as an example, the formation error is set as follows: ; In the formula, For formation error, For robots Neighborhood set, The adjacency matrix is the first Line number The elements of the column represent robots. Neighborhood Robot The connection weights, For robots The relative state vector with respect to the lead robot is represented by the formation shape. Neighborhood robots exist Location status information at any given time. Neighborhood robots The relative state vector with respect to the navigating robot. For robots The connection weight with the navigation robot The trajectory of the navigating robot, i.e., the desired trajectory.
[0027] Step 103: Introduce an optimal control strategy for formation error. Obtain the first optimal value function expression and the optimal virtual controller expression that minimize the first value function by solving the HJB equation. Determine the first fuzzy basis function based on the formation error to approximate the system uncertainty term and the estimated error compensation term after gradient decomposition of the first optimal value function expression. Optimize the first optimal value function expression according to the gradient decomposition and approximation form, based on the first fuzzy basis function and the identifier-actor-critic reinforcement learning algorithm. Obtain an approximate value of the optimal virtual controller expression based on the optimization result.
[0028] Combination Figure 2In this embodiment, an optimal control strategy is introduced to derive the substitution function and value function from the calculated formation error. Then, the value function is expanded using the Taylor formula, and the HJB equation is obtained to acquire the expression of the optimal virtual controller (i.e., the optimal virtual controller expression) and the expression of the optimal value function with respect to the formation error (i.e., the first optimal value function expression). Large and small error intervals are defined using the formation error, and different fuzzy basis functions are constructed for each interval. A switching coefficient is defined between the large and small error intervals to achieve a smooth transition. The gradient of the optimal value function is then decomposed into low-order terms for achieving fast convergence of the system in a fixed time; high-order terms for achieving fast convergence of the system in a fixed time; linear terms for balancing the control strength of the system; system uncertainty terms; and estimation error compensation terms. A fuzzy logic system is used to approximate the system uncertainty terms and estimation error compensation terms based on the fuzzy basis functions.
[0029] Building upon this foundation, a Reinforcement Learning Algorithm of Identifier-Actor-Critic is introduced, combined with a fuzzy logic system to form three modules: the Identifier Fuzzy Logic System Module, the Actor Fuzzy Logic System Module, and the Critic Fuzzy Logic System Module. The Identifier Fuzzy Logic System Module, based on formation error, is used to identify unknown nonlinear functions. The Actor Fuzzy Logic System Module, based on the optimal virtual controller expression, is used to execute the control behavior of the multi-agent system. The Critic Fuzzy Logic System Module, based on the optimal value function, is used to evaluate the behavior taken by the Actor Fuzzy Logic System Module, assess control performance, and provide feedback to the Actor Fuzzy Logic System Module.
[0030] Taking a multi-robot system as an example, in this embodiment, an optimal control strategy is introduced to improve the formation control performance of a multi-robot system with a leader-follower structure. The core of this strategy lies in constructing a cost function and minimizing it to obtain the formation controller parameters. This method aims to ensure the desired control effect while minimizing system resource consumption, achieving an effective balance between performance and efficiency. Specifically, the desired control objective is: the relative positions of the members in the multi-robot system remain stable, and the movement speed of the following robots gradually becomes consistent with that of the leader robot. The designed cost function for formation error... As shown below: ; in, , Right now , It is a virtual controller.
[0031] By integrating the above cost function over time, the cumulative cost function, i.e., the value function, is obtained over the integration time interval: .
[0032] The optimal virtual control strategy is: assuming an optimal virtual controller. The expression that minimizes the value function, i.e., the first optimal value function, is: ; Accordingly, the optimal virtual controller expression is: ; in, Equivalent to .
[0033] The value function is expanded using Taylor's formula, i.e., the HJB equation is obtained. The gradient decomposes the optimal value function of formation error (i.e., the expression of the first optimal value function) into lower-order terms, higher-order terms, linear terms, system uncertainty terms, and estimation error compensation terms. The specific decomposition form is shown below: ; In the formula, and Two constants greater than 0; Functions designed to avoid singularity issues, It is a positive constant. It is a constant. , , For system uncertainties, To estimate the error compensation term, It is a state variable (i.e., the first state variable).
[0034] In the optimal control method, since the HJB equation is difficult to obtain an analytical solution, a fuzzy logic system is used to approximate the uncertain terms and estimation error compensation terms of the system after gradient decomposition. On this basis, in order to achieve a balance between convergence accuracy and speed, the large error interval and the small error interval are distinguished according to the formation error, and fuzzy basis functions are designed for the large error interval and the small error interval respectively.
[0035] For example, it can be based on Calculate the norm of the formation error, use the norm of the formation error to characterize the overall magnitude of the error, and then... Defined as a large error interval, Defined as a small error interval, to avoid abrupt changes in the analytical solution during interval switching and to achieve a smooth transition, according to Determine the switching coefficient ,when hour, ,when , This allows for a smooth transition based on fuzzy basis functions designed for large and small error intervals respectively.
[0036] Based on this, the fuzzy basis functions of the fuzzy logic system include: the first identifier fuzzy basis function. And the first actor-critic fuzzy basis function .
[0037] according to Determine the fuzzy basis function of the first identifier; according to Determine the first actor-critic fuzzy basis function.
[0038] in, To set a threshold, This is the first preset fuzzy basis function of the identifier (which can also be understood as the fuzzy basis function of the identifier corresponding to the large error interval). , The center of the fuzzy basis function of the first preset identifier, The width of the fuzzy basis function of the first preset identifier. This is the second preset fuzzy basis function of the identifier (which can also be understood as the fuzzy basis function of the identifier corresponding to the small error interval). , The center of the fuzzy basis function of the second preset identifier. The width of the fuzzy basis function for the second preset recognizer. The first preset actor-critic fuzzy basis function (corresponding to the large error interval). As the center of the first presupposed actor-critic fuzzy basis function, The width of the first preset actor-critic fuzzy basis function. The second preset actor-critic fuzzy basis function (corresponding to the small error interval) is used. For the second presupposed actor-critic fuzzy basis function center, The width of the second preset actor-critic fuzzy basis function.
[0039] After approximation by the fuzzy logic system, the gradient decomposition of the first optimal value function expression is transformed into: ; in, This is the parameter matrix of the first optimal identifier. The first optimal commentator parameter matrix, This refers to the first fuzzy basis function of the first fuzzy basis function. Let the first actor-critic fuzzy basis function be the first fuzzy basis function. The approximate error is satisfied. ,in and These are two positive constants.
[0040] The optimal virtual controller expression is transformed into: ; However, this approximation process involves the problem of unknown optimal parameter vectors. To address this, a recognizer-actor-critic algorithm incorporating reinforcement learning is introduced, and a fuzzy logic system is used to construct the recognizer, actor, and critic modules. The recognizer module identifies system uncertainties, the actor module generates virtual control strategies, and the critic module evaluates the control effect and provides feedback on optimization information. These three modules work synergistically to achieve adaptive improvement in control performance. The specific structures of each module are shown below: The expression for the identifier-actor fuzzy logic system module is: ; In the formula, This is an approximation of the optimal virtual controller expression. For the identifier parameter vector, This is the actor parameter vector.
[0041] The expression for the identifier-evaluator fuzzy logic system module is: ; In the formula, This is an estimate of the gradient of the first optimal value function expression. For the commentator's parameter vector.
[0042] The parameter vector update laws for the fuzzy logic system modules of the identifier, critic, and actor are designed using the gradient descent method. The parameter vectors are updated online, and the specific forms of each parameter vector update law are shown below: The parameter vector update law of the fuzzy logic system module of the identifier: ; In the formula, for The derivative, express of OK Column elements, The learning rate of the fuzzy logic system module of the identifier. express of Row elements, Indicates formation error Column elements, and It is a positive design constant used to control the magnitude of weight updates and avoid over-updates.
[0043] The parameter vector update law of the fuzzy logic system module for critics: ; In the formula, for The derivative, Let be the learning rate of the fuzzy logic system module for the commentator, and be a constant greater than 0. To improve training speed, among other things, for An identity matrix of order 1.
[0044] The parameter vector update law of the actor fuzzy logic system module: ; In the formula, for The derivative, Let be the learning rate of the actor's fuzzy logic system module, and be a constant greater than 0.
[0045] Therefore, optimization is performed based on the update law of each parameter vector to determine the first optimal identifier parameter matrix and the first optimal critic parameter matrix. The first optimal identifier parameter matrix and the first optimal critic parameter matrix are then substituted into the transformed optimal virtual controller expression to obtain an approximate value of the optimal virtual controller expression.
[0046] Step 104: Determine the virtual error based on the velocity state information of each agent and the approximation of the optimal virtual controller. Introduce an optimal control strategy for the virtual error. Obtain the second optimal value function expression and the optimal controller expression that minimize the second value function by solving the HJB equation. Determine the second fuzzy basis function based on the formation error to approximate the system uncertainty term and the estimation error compensation term after gradient decomposition of the second optimal value function expression. Optimize the second optimal value function expression according to the gradient decomposition and approximation form, based on the second fuzzy basis function and the identifier-actor-commentator reinforcement learning algorithm. Obtain the approximation of the optimal controller expression based on the optimization result. Determine the disturbance observer based on the approximation of the optimal virtual controller. Obtain the optimal controller of the multi-agent system based on the approximation of the optimal controller expression and the disturbance observer.
[0047] In this embodiment, the optimal virtual controller obtained in step 103 is used to construct a virtual error; and the HJB equation is solved to obtain the expression of the optimal controller (i.e., the optimal controller expression) and the expression of the optimal value function with respect to the virtual error (i.e., the second optimal value function expression). Similarly, the formation error is used to define large error intervals and small error intervals respectively, and different fuzzy basis functions are constructed for different intervals. At the same time, a switching coefficient is defined between the large error interval and the small error interval to achieve a smooth transition. Then, without considering the disturbance, the gradient of the optimal value function is decomposed into low-order terms to achieve fast convergence of the system in a fixed time; high-order terms to achieve fast convergence of the system in a fixed time; linear terms to balance the control strength of the system; system uncertainty terms and estimation error compensation terms; and the fuzzy logic system is used to approximate the system uncertainty terms and estimation error compensation terms.
[0048] Based on this, similar to the algorithm introduced in step 103, a distinguisher-actor-critic reinforcement learning algorithm is introduced, combined with a fuzzy logic system to form a distinguisher fuzzy logic system module, an actor fuzzy logic system module, and a critic fuzzy logic system module; and a disturbance observer is introduced to observe the disturbance of the system in real time, and the final controller is obtained after considering the disturbance.
[0049] For example, firstly, according to Determine the virtual error.
[0050] in, This is a virtual error. For the first The velocity state information of each agent This is an approximation of the optimal virtual controller.
[0051] Based on this, the cost function is designed as follows: .
[0052] in, , Equivalent to , for the first The controller for a robot / agent.
[0053] Accordingly, the value function is: .
[0054] The optimal control strategy is: assuming an optimal controller. The expression that minimizes the value function, i.e., the second optimal value function, is: ; Accordingly, the optimal controller is: ; in, Equivalent to .
[0055] The value function is expanded using Taylor's formula, i.e., the HJB equation is obtained. The gradient of the optimal value function of the virtual error (i.e., the expression of the second optimal value function) is decomposed into lower-order terms, higher-order terms, linear terms, system uncertainty terms, and estimation error compensation terms. The specific decomposition form is shown below: ; In the formula, and Two constants greater than 0; Functions designed to avoid singularity issues, It is a positive constant; For system uncertainties, To estimate the error compensation term, .
[0056] Similarly, based on the formation error, large and small error intervals are distinguished. Based on this, the fuzzy basis functions of the fuzzy logic system are set, including: the fuzzy basis function of the second identifier. Second actor-critic fuzzy basis function .
[0057] according to Determine the fuzzy basis function of the second identifier; according to Determine the fuzzy basis function for the second actor-critic.
[0058] in, The switching coefficient, Let be the norm of the formation error. To set a threshold, The third preset identifier fuzzy basis function (corresponding to the large error interval). , , The center of the fuzzy basis function of the third preset identifier. The width of the fuzzy basis function for the third preset identifier. This is the fuzzy basis function for the fourth preset identifier (corresponding to the small error interval). , The center of the fuzzy basis function of the fourth preset identifier. The width of the fuzzy basis function for the fourth preset identifier. For the third preset actor-critic fuzzy basis function (corresponding to the large error interval), , For the third presupposed actor-critic fuzzy basis function center, For the third preset actor-critic fuzzy basis function width, The fourth preset actor-critic fuzzy basis function (corresponding to the small error interval) is used. , For the fourth presupposed actor-critic fuzzy basis function center, The width of the fourth preset actor-critic fuzzy basis function.
[0059] After processing by the fuzzy logic system, the gradient decomposition of the second optimal value function expression is transformed into: ; The optimal controller is transformed into: ; In the formula, This is the second optimal identification parameter matrix. This is the second-optimal critic parameter matrix. This refers to the second fuzzy basis function of the second fuzzy basis function. Let be the second actor-critic fuzzy basis function in the second fuzzy basis function, where The approximate error is satisfied. ,in and These are two positive constants.
[0060] Similarly, we introduce a recognition-actor-critic algorithm that integrates reinforcement learning. The specific structure of each module is shown below: The expression for the identifier-actor fuzzy logic system module is: ; In the formula, This is an estimate of the optimal controller. For the identifier parameter vector, This is the actor parameter vector.
[0061] The expression for the identifier-evaluator fuzzy logic system module is: ; In the formula, This is an estimate of the gradient of the second optimal value function expression. For the commentator's parameter vector.
[0062] Similarly, using the gradient descent method, parameter vector update laws are designed for the identifier fuzzy logic system module, the critic fuzzy logic system module, and the actor fuzzy logic system module, and parameter vectors are updated online. The specific forms of each parameter vector update law are shown below: The parameter vector update law of the fuzzy logic system module of the identifier: ; In the formula, for The derivative, express of OK Column elements, The learning rate of the fuzzy logic system module of the identifier. express of Row elements, Indicates virtual error Column elements, and It is a positive design constant used to control the magnitude of weight updates and avoid over-updates.
[0063] The parameter vector update law of the fuzzy logic system module for critics: ; In the formula, for The derivative, Let be the learning rate of the fuzzy logic system module for the commentator, and be a constant greater than 0. To improve training speed, among other things, for An identity matrix of order 1.
[0064] The parameter vector update law of the actor fuzzy logic system module is as follows: ; In the formula, for The derivative, Let be the learning rate of the actor's fuzzy logic system module, and be a constant greater than 0.
[0065] The update rate of the perturbation observer is: ; In the formula, for The derivative, For the perturbation observer of Column elements, and Used to control the convergence speed and stability of the estimator, it is a constant greater than 0. , , =1, It is 0.7. It is 0.01. for The derivative, for of The elements of the column.
[0066] The final determined optimal controller takes the following form: ; In the formula, The optimal controller for a multi-agent system. This is an approximation of the optimal controller expression. For disturbance observers.
[0067] like Figure 3 As shown in the figure, this embodiment uses 4 followers and 1 navigator for MATLAB simulation example.
[0068] In this embodiment, in a specific test case, given the desired trajectory and speed of the navigating robot, the following robot tracks the navigating robot's movements, and its trajectory eventually converges with that of the navigating robot. The specific form of the navigating robot's desired trajectory is shown below: ; according to Figure 3 As can be seen, the four follower robots move in a specific formation, following the trajectory of the navigator robot. In this example, the formation shape is rectangular, and the initial coordinates of the four follower robots are: .
[0069] Here is another set of initial coordinates: .
[0070] Simulation results under two initial states are as follows: Figure 4 As shown, according to Figure 4 It can be seen that the states of the four follower robots and the lead robot eventually converge.
[0071] This invention proposes a multi-agent formation control method based on a recognizer-actor-critic reinforcement learning and fuzzy logic system. This method introduces optimal control into the leader-follower formation control of a multi-agent system. It leverages the ability of fuzzy logic to approximate continuous functions, solving the problem of obtaining analytical solutions to the HJB equations in optimal control. Furthermore, it defines large and small error intervals based on formation errors, constructing different fuzzy basis functions within each interval. A switching coefficient is defined to achieve smooth transitions, addressing the insufficient approximation accuracy of fixed basis functions when errors dynamically change, thus achieving a balance between convergence accuracy and speed. Building upon this, a recognizer-actor-critic reinforcement learning algorithm is used to construct three fuzzy logic system modules: a recognizer module, an actor module, and a critic module. In this framework, the recognizer module identifies uncertainties in the system, the actor module executes control actions, and the critic module evaluates the control behavior and transmits feedback information to the first two modules. Finally, the parameter update rules for the three modules are designed using gradient descent. This method improves control performance while effectively reducing resource consumption, and enhances the adaptability of multi-robot systems to the environment through online learning.
[0072] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0073] The following are device embodiments of the present invention. For details not described in detail, please refer to the corresponding method embodiments described above.
[0074] Figure 5 A schematic diagram of the structure of a multi-agent system formation control device based on fuzzy reinforcement learning provided in an embodiment of the present invention is shown. For ease of explanation, only the parts related to the embodiment of the present invention are shown, and are described in detail below: like Figure 5 As shown, the formation control device for a multi-agent system based on fuzzy reinforcement learning includes: Modeling module 51 is used to establish the communication topology of a multi-agent system based on graph theory, and to determine the neighbor set of each agent based on the communication topology. The neighbor set includes each of the agent's neighbor agents. Processing module 52 is used to determine the formation error based on the position status information of each agent, the position status information of each agent's neighbor agents, and the desired trajectory of the lead agent. The first control module 53 is used to introduce an optimal control strategy for the formation error, obtain a first optimal value function expression and an optimal virtual controller expression that minimize the first value function by solving the HJB equation; determine a first fuzzy basis function based on the formation error to approximate the system uncertainty term and the estimated error compensation term after gradient decomposition of the first optimal value function expression; optimize the first fuzzy basis function and the identifier-actor-critic reinforcement learning algorithm according to the gradient decomposition and approximation form of the first optimal value function expression, and obtain an approximate value of the optimal virtual controller expression based on the optimization result; The second control module 54 is used to determine the virtual error based on the velocity state information of each agent and the approximation of the optimal virtual controller; introduce an optimal control strategy for the virtual error; obtain the second optimal value function expression and the optimal controller expression that minimize the second value function by solving the HJB equation; determine the second fuzzy basis function based on the formation error to approximate the system uncertainty term and the estimation error compensation term after gradient decomposition of the second optimal value function expression; optimize the second fuzzy basis function and the identifier-actor-critic reinforcement learning algorithm according to the gradient decomposition and approximation form of the second optimal value function expression; obtain the approximation of the optimal controller expression based on the optimization result; determine the disturbance observer based on the approximation of the optimal virtual controller; and obtain the optimal controller of the multi-agent system based on the approximation of the optimal controller expression and the disturbance observer.
[0075] Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. For example... Figure 6 As shown, the electronic device 6 of this embodiment includes a processor 60 and a memory 61. The memory 61 stores a computer program 62. When the processor 60 executes the computer program 62, it implements the steps in the various method embodiments described above. Alternatively, when the processor 60 executes the computer program 62, it implements the functions of each module / unit in the various device embodiments described above.
[0076] For example, computer program 62 may be divided into one or more modules / units, which are stored in memory 61 and executed by processor 60 to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of computer program 62 in electronic device 6.
[0077] Electronic device 6 may include, but is not limited to, processor 60 and memory 61. Those skilled in the art will understand that... Figure 6This is merely an example of electronic device 6 and does not constitute a limitation on electronic device 6. It may include more or fewer components than shown, or combine certain components, or different components. For example, electronic device 6 may also include input / output devices, network access devices, buses, etc.
[0078] The processor 60 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0079] The memory 61 can be an internal storage unit of the electronic device 6, such as a hard disk or RAM. The memory 61 can also be an external storage device of the electronic device 6, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory 61 can include both internal and external storage units of the electronic device 6. The memory 61 is used to store the computer program 62 and other programs and data required by the electronic device 6. The memory 61 can also be used to temporarily store data that has been output or will be output.
[0080] For the sake of simplicity and clarity, only the above-described functional modules / units are used as examples. In practical applications, the functions described above can be assigned to different functional modules / units as needed. These modules / units can be implemented in hardware, software, or a combination of both.
[0081] This invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the methods described in the above-described method embodiments.
[0082] Computer programs include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. Computer-readable media can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.
[0083] In the above embodiments, the descriptions of each embodiment have their own emphasis. Parts not detailed or described in a particular embodiment can be referred to in the relevant descriptions of other embodiments. Unless otherwise specified or in conflict with logic, the terminology and / or descriptions between different embodiments are consistent and can be referenced interchangeably. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.
[0084] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A formation control method for multi-agent systems based on fuzzy reinforcement learning, characterized in that, include: Based on graph theory, a communication topology for a multi-agent system is established, and a neighbor set for each agent is determined according to the communication topology. The neighbor set includes all the neighbor agents of the agent. The formation error is determined based on the position status information of each agent, the position status information of each agent's neighboring agents, and the desired trajectory of the lead agent. To address the formation error, an optimal control strategy is introduced. By solving the HJB equation, the expression for the first optimal value function that minimizes the first value function and the expression for the optimal virtual controller are obtained. The first fuzzy basis function is determined based on the formation error to approximate the system uncertainty term and the estimation error compensation term after gradient decomposition of the first optimal value function expression; according to the form after gradient decomposition and approximation of the first optimal value function expression, the first fuzzy basis function and the identifier-actor-critic reinforcement learning algorithm are used for optimization, and the approximate value of the optimal virtual controller expression is obtained based on the optimization result; The virtual error is determined based on the velocity state information of each agent and the approximation of the optimal virtual controller. An optimal control strategy is introduced to address this virtual error. The second optimal value function expression and the optimal controller expression, which minimize the second value function, are obtained by solving the HJB equation. A second fuzzy basis function is determined based on the formation error to approximate the system uncertainty and estimation error compensation terms after gradient decomposition of the second optimal value function expression. The approximation is then performed using the second fuzzy basis function and a recognizer-actor-critic reinforcement learning algorithm, according to the gradient decomposition of the second optimal value function expression. An approximation of the optimal controller expression is obtained based on the optimization result. A disturbance observer is determined based on the approximation of the optimal virtual controller. Finally, the optimal controller for the multi-agent system is obtained based on the approximation of the optimal controller expression and the disturbance observer.
2. The formation control method for multi-agent systems based on fuzzy reinforcement learning according to claim 1, characterized in that, The first fuzzy basis function includes: a first discriminator fuzzy basis function and a first actor-critic fuzzy basis function; Determining the first fuzzy basis function based on the formation error includes: Calculate the norm of the formation error; According to the norm, Determine the switching coefficient; Based on the switching coefficient, the first preset identifier fuzzy basis function, and the second preset identifier fuzzy basis function, according to Determine the fuzzy basis function of the first identifier; Based on the switching coefficients, the first preset actor-critic fuzzy basis function, and the second preset actor-critic fuzzy basis function, according to... Determine the first actor-critic fuzzy basis function; in, The switching coefficient is... Let the norm be... To set a threshold, Let f be the fuzzy basis function of the first identifier. The first preset identifier's fuzzy basis function is... , For the first in a multi-agent system The location and state information of each agent. The center of the fuzzy basis function of the first preset identifier. The width of the first preset identifier's fuzzy basis function. The second preset identifier's fuzzy basis function, , The center of the fuzzy basis function of the second preset identifier. The width of the second preset identifier's fuzzy basis function. Let be the first actor-critic fuzzy basis function. Let the first preset actor-critic fuzzy basis function be . , As the first state variable, Let the first preset actor-critic fuzzy basis function center be , The width of the first preset actor-critic fuzzy basis function. Let the second preset actor-critic fuzzy basis function be used. , The second preset actor-critic fuzzy basis function center, The width of the second preset actor-critic fuzzy basis function.
3. The formation control method for multi-agent systems based on fuzzy reinforcement learning according to claim 1, characterized in that, The gradient decomposition and approximation of the first optimal value function expression is as follows: ; in, The gradient of the first optimal value function expression. and Two constants greater than 0, The formation error is... Functions designed to avoid singularity issues It is a positive constant. It is a constant. , , This is the parameter matrix of the first optimal identifier. The first optimal commentator parameter matrix, This refers to the first identifier fuzzy basis function in the first fuzzy basis function. Let be the first actor-critic fuzzy basis function in the first fuzzy basis function. The approximate error is satisfied. ,in and These are two positive constants.
4. The formation control method for multi-agent systems based on fuzzy reinforcement learning according to claim 3, characterized in that, Based on the gradient decomposition and approximation of the first optimal value function expression, optimization is performed according to the first fuzzy basis function and the identifier-actor-critic reinforcement learning algorithm. Based on the optimization results, an approximate value of the optimal virtual controller expression is obtained, including: The expression for the optimal virtual controller is transformed according to the gradient decomposition and approximation of the first optimal value function expression: ; Based on the gradient decomposition and approximation form of the first optimal value function expression and the transformed optimal virtual controller expression, the expression of the identifier-actor fuzzy logic system module is determined as follows: ; In the formula, This is an approximation of the optimal virtual controller expression. For the identifier parameter vector, For actor parameter vectors; The expression for the identifier-evaluator fuzzy logic system module is: ; In the formula, This is an estimate of the gradient of the first optimal value function expression. For the commentator parameter vector; The parameter vector update law of the fuzzy logic system module of the identifier is as follows: ; In the formula, for The derivative of express of OK Column elements, The learning rate of the fuzzy logic system module of the identifier. express of Row elements, Indicates the formation error Column elements, and It is a positive design constant used to control the update magnitude of weights and avoid over-updating; The parameter vector update law of the fuzzy logic system module is as follows: ; In the formula, for The derivative of Let be the learning rate of the fuzzy logic system module for the commentator, and be a constant greater than 0. To improve training speed, among other things, for An identity matrix of order 1; The parameter vector update law of the actor fuzzy logic system module is as follows: ; In the formula, for The derivative of Let be the learning rate of the actor's fuzzy logic system module, and be a constant greater than 0. Optimize according to the parameter vector update laws to determine the first optimal identifier parameter matrix and the first optimal critic parameter matrix. Substitute the first optimal identifier parameter matrix and the first optimal critic parameter matrix into the transformed optimal virtual controller expression to obtain an approximate value of the optimal virtual controller expression.
5. The formation control method for multi-agent systems based on fuzzy reinforcement learning according to claim 1, characterized in that, The virtual error is determined based on the velocity state information of each agent and the approximation of the optimal virtual controller, including: according to Determine the virtual error; in, The virtual error, For the first The velocity state information of each agent This is an approximation of the optimal virtual controller.
6. The formation control method for multi-agent systems based on fuzzy reinforcement learning according to claim 1, characterized in that, The second fuzzy basis function includes: the second discriminator fuzzy basis function and the second actor-critic fuzzy basis function; Determining the second fuzzy basis function based on the formation error includes: Calculate the norm of the formation error; According to the norm, Determine the switching coefficient; Based on the switching coefficient, the third preset identifier fuzzy basis function, and the fourth preset identifier fuzzy basis function, according to Determine the fuzzy basis function of the second identifier; Based on the switching coefficients, the third preset actor-critic fuzzy basis function, and the third preset actor-critic fuzzy basis function, according to Determine the second actor-critic fuzzy basis function; in, The switching coefficient is... Let the norm be... To set a threshold, The second identifier's fuzzy basis function is... The third preset identifier fuzzy basis function, , , , The third preset identifier's fuzzy basis function center, The width of the fuzzy basis function of the third preset identifier. The fourth preset identifier fuzzy basis function, , The center of the fuzzy basis function of the fourth preset identifier, The width of the fuzzy basis function of the fourth preset identifier. Let the second actor-critic fuzzy basis function be , The third preset actor-critic fuzzy basis function, , , The third preset actor-critic fuzzy basis function center, The width of the third preset actor-critic fuzzy basis function. For the fourth preset actor-critic fuzzy basis function, , The fourth preset actor-critic fuzzy basis function center, The width of the fourth preset actor-critic fuzzy basis function.
7. The formation control method for multi-agent systems based on fuzzy reinforcement learning according to claim 1, characterized in that, Determining the disturbance observer based on an approximation of the optimal virtual controller includes: according to Determine the disturbance observer; In the formula, for The derivative of For the disturbance observer of Column elements, and Used to control the convergence speed and stability of the estimator, it is a constant greater than 0. , , =1, It is 0.
7. It is 0.
01. for The derivative of for of Column elements; The optimal controller for a multi-agent system is obtained based on an approximation of the optimal controller expression and the disturbance observer, including: according to To obtain the optimal controller for a multi-agent system; In the formula, The optimal controller for a multi-agent system. This is an approximation of the optimal controller expression. This refers to the disturbance observer.
8. A multi-agent system formation control device based on fuzzy reinforcement learning, characterized in that, include: The modeling module is used to establish the communication topology of the multi-agent system based on graph theory, and to determine the neighbor set of each agent based on the communication topology. The neighbor set includes all the neighbor agents of the agent. The processing module is used to determine the formation error based on the position status information of each agent, the position status information of each agent's neighboring agents, and the desired trajectory of the lead agent. The first control module is used to introduce an optimal control strategy for the formation error, and obtain the first optimal value function expression and the optimal virtual controller expression that minimize the first value function by solving the HJB equation. The first fuzzy basis function is determined based on the formation error to approximate the system uncertainty term and the estimation error compensation term after gradient decomposition of the first optimal value function expression; according to the form after gradient decomposition and approximation of the first optimal value function expression, the first fuzzy basis function and the identifier-actor-critic reinforcement learning algorithm are used for optimization, and the approximate value of the optimal virtual controller expression is obtained based on the optimization result; The second control module is used to determine the virtual error based on the velocity state information of each agent and the approximation of the optimal virtual controller; introduce an optimal control strategy for the virtual error; obtain the second optimal value function expression and the optimal controller expression that minimize the second value function by solving the HJB equation; determine the second fuzzy basis function based on the formation error to approximate the system uncertainty term and the estimation error compensation term after gradient decomposition of the second optimal value function expression; optimize the second fuzzy basis function and the identifier-actor-critic reinforcement learning algorithm according to the gradient decomposition and approximation form of the second optimal value function expression; obtain the approximation of the optimal controller expression based on the optimization result; determine the disturbance observer based on the approximation of the optimal virtual controller; and obtain the optimal controller of the multi-agent system based on the approximation of the optimal controller expression and the disturbance observer.
9. An electronic device, characterized in that, It includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method as described in any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 6.