Formation strategy for multi-agent system based on hierarchical differential game

By employing a hierarchical differential game approach, the formation control problem of multi-agent systems is divided into a policy layer and a planning layer. By utilizing the Pontryagin minimum principle and rolling optimization, the problems of high computational cost and performance degradation in obstacle environments of rolling differential game are solved, thus achieving efficient formation control.

CN116360265BActive Publication Date: 2026-02-06FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310342269.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-03
Publication Date
2026-02-06
Estimated Expiration
2043-04-03

AI Technical Summary

Technical Problem

Existing rolling differential game methods are computationally expensive and degrade system performance when dealing with formation control of multi-agent systems with unknown obstacles, and it is difficult to guarantee the existence and global convergence of Nash equilibrium solutions.

Method used

A hierarchical differential game approach is adopted to divide the formation control problem into a strategy layer and a planning layer. The local Nash equilibrium solution is solved by using the Pontryagin minimum principle. The strategy is modified in real time through a rolling optimization quadratic programming model to form a hybrid formation strategy, which reduces the computational cost and ensures the unique existence and global convergence of the equilibrium solution.

Benefits of technology

It effectively reduces computational costs, shortens the time for agents to complete formation tasks, improves system performance, and theoretically guarantees the unique existence and global convergence of local Nash equilibrium solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116360265B_ABST
    Figure CN116360265B_ABST
Patent Text Reader

Abstract

The application relates to a kind of multi-agent system formation strategy based on hierarchical differential game: first, the communication topology between agents is established by using graph theory, and agent model and obstacle environment model are established;Secondly, for unknown obstacle environment, in the strategy layer, the multi-agent formation problem with known obstacle constraints is converted into distributed differential game, and a cost function is designed for game participants to establish a game model;Using Pontryagin minimum principle, the existence and uniqueness of local Nash equilibrium solution are analyzed, and the expression form of local Nash equilibrium solution is solved, and the global convergence condition of local Nash equilibrium solution is given;In the planning layer, the local Nash equilibrium solution from the strategy layer is modified in real time by using the quadratic programming model based on rolling optimization, and finally a hybrid formation strategy is formed;In theory, it is guaranteed that after the agent successfully avoids the unknown obstacle, the hybrid formation strategy can converge to the local Nash equilibrium solution.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of multi-agent coordination control, and particularly relates to a multi-agent system formation strategy based on hierarchical differential game. BACKGROUND

[0002] Multi-agent coordination control is a hot topic in the field of multi-agent systems, which requires each agent to complete the designated system task through communication, collaboration and competition. In particular, the formation control of multi-agent systems plays an important role in coordination control, such as cluster light show, cooperative cargo handling, and hunting and tracking. In the above scenarios, the multi-agent system is usually subject to communication and task constraints, so it is necessary to study the distributed formation problem, that is, to design a certain distributed mechanism to deal with the limited communication capability of the multi-agent system to solve the formation problem with obstacle constraints.

[0003] In the problem of distributed formation control, there are many solutions, such as behavior-based methods, model predictive control, reinforcement learning, etc. However, these methods do not fully exploit the dynamic interaction characteristics between agents. Game theory quantifies the decision-making process between agents to ensure that agents can always choose the best behavior to achieve individual or collective goals. In particular, differential game quantifies the dynamic conflict process between agents, which is often used to solve coordination control problems, such as pursuit and evasion problems and formation problems.

[0004] The application of differential game in formation control problem is explored as follows: the non-cooperative differential game model is studied to solve the formation problem in the document (Mylvaganam T, Astolfi A. A differential game approach to formation control for a team of agents with one leader [C] / / 2015 American Control Conference (ACC). IEEE, 2015: 1469-1474.). However, this method assumes that the communication between agents is perfect, in fact, each agent will be subject to communication constraints. In order to solve the problem of communication constraints, the document (Mylvaganam T, Astolfi A. Towards a systematic solution for differential games with limited communication [C] / / 2016 American Control Conference (ACC). IEEE, 2016: 3814-3819.) combines distributed control, optimal control and game theory to form a multi-player distributed differential game, and proves the existence of local Nash equilibrium solution formed by participants. However, in the above-mentioned documents, there are two main problems that have not been solved in formation control. First, the formation strategy does not consider obstacle constraints, that is, the environment is assumed to have no obstacles. Second, the formation control strategy obtained is offline, which will reduce the performance of the control strategy of the agent constrained by the dynamic environment.

[0005] In order to solve the two problems, the document (Lin W, Li C, Qu Z, et al. Distributed formation control with open-loop Nash strategy [J]. Automatica, 2019, 106: 266-273.) proposes a rolling differential game method to solve the problem of dynamic environment in formation control, that is, at each time step, the first control input in the open-loop Nash equilibrium is used to control the agent, and at the next time step, the above process is repeated. The rolling formation strategy can compensate for the time inconsistency of the offline Nash equilibrium. For this reason, rolling differential game has become a popular means to deal with formation control problems in dynamic environments. However, this method needs to re-negotiate at each time step, so it not only has high computational cost and poor real-time performance, but also is difficult to guarantee the existence of equilibrium solution in theory, resulting in a decline in system performance. SUMMARY

[0006] Therefore, in order to solve the problem that the rolling differential game has high calculation cost and low system performance in solving the formation control problem with unknown obstacle constraints, the application designs a multi-agent system formation strategy based on hierarchical differential game.

[0007] In the strategy layer, the formation control problem with known obstacle constraints is converted into a distributed differential game, and the local open-loop Nash equilibrium solution is solved by using the Pontryagin minimum principle. In addition, the uniqueness and global convergence of the Nash equilibrium solution are ensured. Considering the time inconsistency of the offline Nash equilibrium solution, in the planning layer, the local Nash equilibrium solution from the strategy layer is modified in real time by using a quadratic programming model based on rolling optimization, and finally a hybrid formation strategy is formed. It is theoretically proved that the hybrid formation strategy can converge to the local Nash equilibrium solution after the agent successfully avoids the unknown obstacle. This method can reduce the calculation cost and shorten the completion time of the formation task of the agent.

[0008] The application relates to a multi-agent system formation strategy based on hierarchical differential game. In the design process, firstly, a communication topology structure between agents is established by using graph theory, and an agent model and an obstacle environment model are established; secondly, for the unknown obstacle environment, in the strategy layer, a multi-agent formation problem with known obstacle constraints is converted into a distributed differential game, a cost function is designed for game participants, and a game model is established; the existence and uniqueness of the local Nash equilibrium solution are analyzed by using the Pontryagin minimum principle, and the expression form of the local Nash equilibrium solution is solved, and the global convergence condition of the local Nash equilibrium solution is given; further, in the planning layer, the local Nash equilibrium solution from the strategy layer is modified in real time by using a quadratic programming model based on rolling optimization, and finally a hybrid formation strategy is formed. According to the application, it is theoretically ensured that the hybrid formation strategy can converge to the local Nash equilibrium solution after the agent successfully avoids the unknown obstacle. This method can reduce the calculation cost and shorten the completion time of the formation task of the agent.

[0009] The application solves the technical problems and specifically adopts the following technical solutions:

[0010] A multi-agent system formation strategy based on hierarchical differential game comprises the following steps:

[0011] Step S1: the communication relationship between agents in a multi-agent system is established by using graph theory; a first-order linear integrator is established as an agent model; one of the agents is set as a leader, and the rest of the agents are set as followers; only the leader knows the expected terminal position, and the rest of the followers maintain the expected formation distance with the leader; the working environment area of the agent is defined, and the obstacles existing in the environment are regarded as ellipsoids;

[0012] Step S2: At the strategy layer, all agents participating in the task are regarded as game participants for the known obstacles existing in the environment, and the formation problem of the multi-agent system with obstacle constraints is modeled as a distributed differential game, which includes a running cost function and a control cost. The participants and the neighbor participants meet the convergence to the local Nash equilibrium.

[0013] Step S3: At the planning layer, for unknown obstacles in the environment, a quadratic programming model based on rolling optimization is used to modify the local Nash equilibrium solution from the strategy layer in real time, and a hybrid formation strategy is formed.

[0014] Further, in step S1:

[0015] The specific form of the multi-agent dynamic equation is:

[0016]

[0017] In the formula, t is the time scale; is the rate of change of the position information of the ith agent at time t; x i is the position information of the ith agent at time t; u i is the control strategy of the ith agent at time t;

[0018] The state of the multi-agent system at time t is defined as where x1(t) and x N (t) are the position information of the first agent and the position information of the Nth agent at time t, respectively.

[0019] A directed interaction topology graph G(v,ε) of N agents is established, where v={1,...,N} represents the agent set.

[0020] represents the set of edges; e ij The communication relationship between agent i and agent j; e ij ∈ε indicates that agent i can receive information from agent j; the neighbor set of agent i is

[0021] In the formation control problem, considering that only the leader knows the desired target position, and the followers maintain a relative distance to achieve the desired formation, the state error of agent i is defined as:

[0022]

[0023] In the formula, is the rate of change of the state error information of the ith agent at time t; is the state error information of the ith agent at time t; α iand β i Define the i-th agent as the leader (α) i =1,β i =0) or follower (α) i =0,β i =1); x j (t) represents the position information of neighboring agent j of agent i at time t; x d The desired target location; Let be the expected grouping distance between the i-th agent and its neighboring agent j;

[0024] Step S12: Establish the working environment area for the intelligent agent:

[0025] Considering that the obstacles in the agent's working environment are ellipsoidal, define the collision avoidance region C. ik for:

[0026]

[0027] In the formula, R 3 a is the three-dimensional plane in which the intelligent agent is located; i The safe distance for agent i; For the obstacle at time t Location, It is an obstacle The radius; I2 is the unit weight matrix;

[0028] Define the sensing area for:

[0029]

[0030] In the formula, A i It is the sensing range of agent i;

[0031] Define free region for:

[0032]

[0033] The following assumptions are given:

[0034] Assumption 1: The obstacles in the working environment of the intelligent agent are sparse;

[0035] Assumption 2: The initial positions of each agent do not coincide, that is:

[0036] x i (t0)≠x j (t0) (6)

[0037] In the formula, t0 is the initial working time of the agent, x i(t0) and x j (t0) are the positions of the i-th agent and the j-th agent at the initial time t0, respectively;

[0038] Assumption Three: The initial position of each agent and the position of the target point satisfy the following conditions, respectively:

[0039]

[0040] where t f is the end time of the work of the agent, x i (t f ) is the position of the i-th agent at the end time t f ;

[0041] Assumption Four: The desired formation distance satisfies the following condition:

[0042]

[0043] Thus, the communication relationship between the agents in the multi-agent system is established.

[0044] Further, in step S2:

[0045] To achieve the goal that the leader in the multi-agent system can reach the desired target point while the followers maintain the desired distance, the optimal control strategy of each agent needs to be determined, wherein each agent is subject to the dynamic equation (1), the network topology G(v, ε), and the collision region constraint.

[0046] To achieve the control goal of the formation problem of the multi-agent system, the formation problem with known obstacle constraints is modeled as a distributed game problem, and the agents participating in completing the task are defined as game participants.

[0047] The game cost of each participant is designed as follows:

[0048]

[0049] where u i and u -i are the control strategy of the agent at time t and the control strategy set of the neighbor self-agent except the i-th agent, respectively, x(t0) is the initial position information of the multi-agent system at t0, χ i (x(t)) is the running cost of the i-th agent at time t, and R i is the adjustable positive definite weight matrix of the i-th agent.

[0050] The running cost function of the i-th agent is defined as follows:

[0051]

[0052] wherein, and are constants, p i (x) and are the barrier cost function of leader i and the barrier cost function of the i-th follower at time t, respectively, defined as follows:

[0053]

[0054] wherein, is the total number of obstacles in the working environment of the agents;

[0055]

[0056] wherein, a j is the safety distance of agent j;

[0057] The control objective of the formation control problem of the multi-agent system is to design the optimal coordination strategy for each player under the dynamic equation (1), the network topology G(v, ε) and the known collision region , and to make the leader safely reach the desired target point while the followers maintain the relative formation position; the player i and the neighbor players satisfy the convergence to the local Nash equilibrium, i.e., the strategy set at time t satisfies:

[0058]

[0059] wherein, and are the optimal strategy of player i and the optimal strategy set of the neighbor players at time t, respectively, and are the optimal cost and the suboptimal cost of player i relative to the neighbor players, respectively.

[0060] Further, in step S3:

[0061] In the planning layer, the multi-agent working environment with unknown obstacles is considered by using the rolling optimization idea, and on the basis of the strategy layer, the formation control problem is converted into a quadratic programming problem, and the cost function of agent i in the planning layer is defined as:

[0062]

[0063] wherein, is the time at the end of the rolling window, wherein is the initial time of the rolling window; t NP is the rolling time domain; is a position error of the multi-agent system at an initial time of a rolling window, the position error being a difference between an optimal position from a policy layer and an actual position, wherein and are actual position information and optimal position information of the multi-agent system at the initial time of the rolling window, respectively; is an agent i.

[0064] Compared with the prior art, the present application and the preferred schemes thereof are directed to a formation control problem with unknown obstacle constraints, adopt a rolling optimization and differential game method, and propose a hierarchical differential game multi-agent system formation strategy. Compared with the existing rolling differential game method which needs to re-game at each time step, the method can change the strategy of only the agent which needs to avoid the unknown obstacle without affecting the equilibrium solution of the game participants, reduce the calculation cost, and theoretically guarantee the unique existence of the local Nash equilibrium solution and the global convergence. In addition, the running time of the system to complete the desired formation can be reduced. BRIEF DESCRIPTION OF DRAWINGS

[0065] The present application will be further described in detail below in combination with the drawings and specific embodiments:

[0066] Figure 1 is a method principle block diagram of an embodiment of the present application;

[0067] Figure 2 is a multi-agent communication topology graph adopted in a simulation example of an embodiment of the present application;

[0068] Figure 3 is a position deviation display of each agent of a comparative method in a simulation example of an embodiment of the present application Figure 1 ;

[0069] Figure 4 is a position deviation display of each agent of the proposed method in a simulation example of an embodiment of the present application Figure 1 ;

[0070] Figure 5 is a calculation time display graph of the two methods in a simulation example of an embodiment of the present application;

[0071] Figure 6 is a position deviation display of each agent of a comparative method in a simulation example of an embodiment of the present application Figure 2 ;

[0072] Figure 7 is a position deviation display of each agent of the proposed method in a simulation example of an embodiment of the present application Figure 2 . DETAILED DESCRIPTION

[0073] In order to make the features and advantages of the patent more obvious and easy to understand, the following specific examples are described in detail as follows.

[0074] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise indicated, all technical and scientific terms used in this description have the same meaning as commonly understood by one of ordinary skill in the art to which the present application pertains.

[0075] It should be noted that the terms used herein are only for the purpose of describing specific embodiments, and are not intended to limit the exemplary embodiments according to the present application. As used herein, the singular form is intended to include the plural form unless the context clearly indicates otherwise, and it should also be understood that when the terms "comprise" and / or "include" are used in the present description, they indicate the presence of a feature, step, operation, device, component, and / or combination thereof.

[0076] As shown in Figure 1 The present embodiment provides a multi-agent system formation strategy based on hierarchical differential game, and the specific design and verification process is as follows:

[0077] Using graph theory, the communication relationship between agents in the multi-agent system is established; a first-order linear integrator is established as the model of the agent; in addition, one of the agents is set as the leader and the remaining agents are set as followers; only the leader knows the desired end position, and the remaining followers maintain the desired formation distance with the leader; the working environment area of the agent is defined, and the obstacles existing in the environment are regarded as ellipsoids.

[0078] In the strategy layer, all agents participating in the task are regarded as game participants in view of the known obstacles existing in the environment, and since the agents have limited communication capability, the formation problem of the multi-agent system with obstacle constraints is modeled as a distributed differential game; the game model includes a running cost function and a control cost;

[0079] The present embodiment uses the Pontryagin minimum principle to analyze the existence and uniqueness of the local Nash equilibrium solution, and lists the expression form of the local Nash equilibrium solution, and in the case of strong connectivity of the communication topology graph, the global convergence of the local Nash equilibrium solution is ensured;

[0080] In the planning layer, for unknown obstacles in the environment, a quadratic programming model based on rolling optimization is used to modify the local Nash equilibrium solution from the strategy layer in real time, and finally a hybrid formation strategy is formed;

[0081] After the theoretical and simulation demonstration of the present embodiment, when the agent successfully avoids the unknown obstacle, it is theoretically ensured that the hybrid formation strategy can converge to the local Nash equilibrium solution, and in addition, the hybrid strategy can converge to the desired formation shape.

[0082] In the embodiment, the task performed by the intelligent agent is to enable a multi-intelligent agent system subject to communication constraints to achieve a desired formation in an environment with unknown obstacles, reduce the computational cost, and improve the system performance.

[0083] In the embodiment, the multi-intelligent agent dynamic equation has a specific form as follows:

[0084]

[0085] In the formula, t is a time scale; is the rate of change of the position information of the i-th intelligent agent at time t; x i is the position information of the i-th intelligent agent at time t; u i is the control strategy of the i-th intelligent agent at time t;

[0086] The state of the multi-intelligent agent system at time t is defined as where x1(t) and x N are the position information of the first intelligent agent and the position information of the N-th intelligent agent at time t, respectively;

[0087] A directed interaction topology graph G(v,ε) of the N intelligent agents is established, where v={1,...,N} represents the intelligent agent set;

[0088] represents the set of edges; e ij is the communication relationship between intelligent agent i and intelligent agent j; e ij ∈ε indicates that intelligent agent i can receive information from intelligent agent j; the neighbor set of intelligent agent i is

[0089] In the formation control problem, considering that only the leader knows the desired target position and the followers maintain a relative distance to achieve the desired formation, the state error of intelligent agent i is defined as:

[0090]

[0091] In the formula, is the rate of change of the state error information of the i-th intelligent agent at time t; is the state error information of the i-th intelligent agent at time t; α i and β i respectively define whether the i-th intelligent agent is a leader (α i =1, β i =0) or a follower (α i =0, β i =1); x j(t) is the position information of the neighbor agent j of the ith agent at time t; x d is the desired target position; is the desired formation distance between the ith agent and the neighbor agent j;

[0092] The working environment area of the agent is established:

[0093] Considering that the obstacle existing in the working environment of the agent is an ellipsoid, the collision avoidance area is defined is:

[0094]

[0095] In the formula, R 3 is a three-dimensional plane in which the agent is located; a i is the safety distance of the agent i; is the position of the obstacle at time t, is the radius of the obstacle ; I2 is a unit weight matrix;

[0096] The sensing area is defined is:

[0097]

[0098] In the formula, A i is the sensing range of the agent i;

[0099] The free area is defined is:

[0100]

[0101] The following assumptions are given:

[0102] Assumption one: the obstacles in the working environment of the agent are sparse.

[0103] Assumption two: the initial positions of each agent do not coincide, that is:

[0104] x i (t0)≠x j (t0) (6)

[0105] In the formula, t0 is the initial time of the agent working, x i (t0) and x j (t0) are the positions of the ith agent and the jth agent at the initial time t0, respectively;

[0106] Assumption three: the initial positions of each agent and the target point positions respectively satisfy the following conditions:

[0107]

[0108] where t f is the working end time of the agent, x i (t f ) is the position of the i-th agent at the end time t f .

[0109] Assumption four: the desired formation distance satisfies the following condition:

[0110]

[0111] To achieve the goal that the leader can reach the desired target point and the followers can keep the desired distance in the multi-agent system, the optimal control strategy of each agent needs to be determined, where each agent is subject to the dynamic equation (1), the network topology G(v, ε) and the collision region constraint.

[0112] In order to achieve the control goal of the formation problem of the multi-agent system, the formation problem with known obstacle constraints is modeled as a distributed game problem, and in addition, the agents participating in completing the task are defined as game participants.

[0113] The game cost of each participant is designed as follows:

[0114]

[0115] where u i and u -i are the control strategy of the agent at t and the control strategy set of the neighbor self-agent except the i-th agent, x(t0) is the initial position information of the multi-agent system at t0, χ i (x(t)) is the running cost of the i-th agent at t, R i is the adjustable positive definite weight matrix of the i-th agent.

[0116] The running cost function of the i-th agent is defined as follows:

[0117]

[0118] where and are constants, p i (x) and are the obstacle penalty function of the leader i and the obstacle penalty function of the i-th follower at t, respectively, which are defined as follows:

[0119]

[0120] where the total number of obstacles existing in the working environment of the intelligent agent.

[0121]

[0122] where a j is the safety distance of the intelligent agent j.

[0123] Therefore, the control objective of the formation control problem of the multi-agent system is to design the optimal coordination strategy for each agent subject to the dynamic equation (1), the network topology G(v, e) and the known collision region The game players need to design the optimal coordination strategy to make the leader safely reach the desired target point while the followers maintain the relative formation positions. In addition, the players i and the neighbor players can converge to a local Nash equilibrium, i.e., the strategy set satisfies:

[0124]

[0125] where and are the optimal strategy of the player i and the optimal strategy set of the neighbor players at time t, respectively, and are the optimal cost and the suboptimal cost of the player i with respect to the neighbor players, respectively.

[0126] The theorem of the existence of the local Nash equilibrium is given: the admissible control strategy set is given and is assumed to be a compact set. In addition, the optimal position information set of the neighbor players except the player i is assumed to be x -i (t), and satisfies assumptions one to three. Then, there exists the optimal control strategy

[0127] The uniqueness rule of the local Nash equilibrium is given: for the differential game model (9), the joint system with two-point boundary values is:

[0128]

[0129] where is the rate of change of the position information of the multi-agent system at time t; is the joint state rate of change of the multi-agent system at time t; I 3N is a 3N-dimensional identity matrix; is a given certain matrix; x(t) is the position information of the multi-agent system at time t; l(t) is the joint state of the multi-agent system at time t; x(t0) is the position information of the multi-agent system at the initial time; x0 is a certain constant; l(t0) is the joint state of the multi-agent system at the initial time.

[0130] Then if is positive definite, for any initial position information x0of the multi-agent system, there exists a unique solution of equation (14).

[0131] The local Nash equilibrium solution is solved based on the Pontryagin minimum principle. First, the Hamilton function is defined as

[0132]

[0133] where λ i (t) is the joint state of agent i at time t; x * (t) is the optimal position information of the multi-agent system at time t; u i (t) is the control strategy of agent i at time t; u -i (t) is the control strategy set of the neighbor agents of agent i at time t; x(t0) is the initial position information of the multi-agent system;

[0134] is the transpose information of the joint state of agent i at time t; is the rate of change of the position information of agent i at time t.

[0135] Secondly, the local Nash equilibrium solution satisfies the following differential equation set:

[0136]

[0137] The boundary condition is:

[0138]

[0139] where is the optimal strategy of agent i at time t; x(t0) is the initial position information of the multi-agent system; x0is a constant; λ i (t f ) is the joint state of agent i at the end time;

[0140] is the Hamilton function form when the neighbor agents of agent i reach the optimal strategy.

[0141] The global convergence of the local Nash equilibrium solution is analyzed. First, the definition of the global Nash equilibrium solution is given: for the differential game of N agents, if the game cost function satisfies the following inequality, then the N control strategy sets converge to the global Nash equilibrium solution.

[0142]

[0143] where, is the optimal control policy set of all agents except agent i;

[0144] is the optimal cost of agent i with respect to all other agents; is the suboptimal cost of agent i with respect to all other agents.

[0145] Furthermore, there exists a policy such that the following inequality holds,

[0146]

[0147] where, is a certain policy of agent i at time t; is the cost of agent i with respect to all other agents when agent i reaches a certain policy at time t.

[0148] Secondly, the global convergence of local Nash equilibrium solutions is given: Let be the optimal control policy of agent i at time t, assuming that the communication connected graph G(v, ε) of agents is strongly connected, then 1) the local Nash equilibrium solutions corresponding to all sub-connected graphs containing agent i are equal; 2) all local Nash equilibrium solutions corresponding to agents can converge to the global Nash equilibrium solution.

[0149] At the planning layer, the multi-agent working environment with unknown obstacles is considered by using the rolling optimization idea, and the formation control problem is converted into a quadratic programming problem on the basis of the policy layer. The cost function of agent i at the planning layer is defined as:

[0150]

[0151] where, is the time at the end of the rolling window, where is the time at the beginning of the rolling window; t NP is the rolling time domain; is the position error of the multi-agent system at the beginning of the rolling window, which is the difference between the optimal position from the policy layer and the actual position, where and are the actual position and optimal position information of the multi-agent system at the beginning of the rolling window, respectively; is the position error of agent i at the rolling window t, where is the optimal position of agent i at the rolling window t; F i and Qi is the adjustable positive definite weight matrix for the ith agent; is the position error of agent i at the end of the rolling window.

[0152] At the planning layer, agent i adjusts the optimal position information from the strategy layer in real time online using a linear quadratic model based on rolling optimization. When agent i encounters an unknown obstacle, it no longer games with the original neighbor agents, but instead finds a collision-free trajectory close to the optimal position in the strategy layer through the linear quadratic model (20); if agent i does not encounter an unknown obstacle, it continues to use the equilibrium solution of the strategy layer, and finally forms a hybrid formation strategy.

[0153] A theorem of formation convergence is given: consider a multi-agent system composed of N agents (1), which are subject to network communication topology graph G(v, ε) and collision region constraints. It is assumed that agent i encounters an unknown obstacle If there is at least one feasible control strategy for each agent at time t, there exists a matrix such that the following holds:

[0154]

[0155] where, Q i > 0; F i > 0; is the sampling number of the rolling window; Δt is the total sampling number of the rolling window.

[0156] Let be the optimal solution of the following optimal problem, then all agents gradually converge to the desired formation.

[0157]

[0158] where, is the optimal strategy of the ith agent at time t of the rolling window; is the initial time of the rolling window; is the position error rate of the ith agent at time t of the rolling window; is the position error of the multi-agent system at the initial time of the rolling window.

[0159] Based on the above theorem of formation convergence, a corollary is given: for a multi-agent system composed of N agents (1), which are subject to network communication topology graph G(v, ε) and collision region Constraints. Suppose there are g agents encountering unknown obstacles, and if the g agents successfully avoid all unknown obstacles based on the quadratic programming model (20), then the g agents will continue to converge to the local Nash equilibrium solution of the policy layer because the other agents still follow the local Nash equilibrium solution from the policy layer.

[0160] The following is a specific example to demonstrate the effectiveness and superiority of the proposed hierarchical differential game multi-agent system formation strategy in solving the multi-agent formation problem with obstacle constraints.

[0161] To demonstrate that the proposed optimal controller reduces computational cost and improves system performance, this example compares the proposed method with existing rolling differential game methods and conducts simulation experiments. Taking six agents as an example, the specific model expression of agent i is given below:

[0162]

[0163] The communication topology of the six agents is as follows: Figure 2 As shown, agent 1 is defined as the leader, and the other 5 agents are followers. Agent 1's initial position and desired target position are x1(t0) = [-30, -30, -30]. T ,x1(t f = [300, 300, 100] T The initial positions of followers 2 to 6 are: x2(t0) = [-50, -20, -30] T x2(t0) = [-40, -70, -30] T x2(t0) = [-20, -50, -30] T x2(t0) = [-60, -40, -30] T x2(t0) = [-20, -60, -30] T The expected formation distance between follower 2 and leader 1 is: d 12 (t0) = [100, 0, 0] T The expected formation distance between follower 3 and follower 2 is: d 23 (t0) = [-100, -100, 0] T The expected formation distance between follower 4 and follower 3 is:

[0164] d 34 (t0) = [100, -100, 0] T The expected formation distance between follower 5 and follower 4 is: d 45 (t0) = [100, 0, 0] TThe expected formation distance of follower 6 to follower 5 is d 56 (t0) = [100, 100, 0] T The induction radius is A i = 25.

[0165] Figure 3 The agent position deviation diagram of the rolling distributed differential game method (the comparative method) is shown in FIG. 2, Figure 4 The agent position deviation diagram of the hierarchical differential game formation strategy (the proposed method) is shown in FIG. 3, according to Figure 3 and Figure 4 It can be known from the two methods that the comparative method and the proposed method can make the agent converge to the expected formation shape, but according to Figure 5 The calculation time comparison diagram of the two methods shows that, since the proposed method does not affect the equilibrium solution of the game participants, it is not necessary to repeat the game, and the calculation cost of the proposed method is obviously superior to that of the comparative method, so the proposed formation control strategy based on hierarchical differential game can reduce the calculation cost and has timeliness.

[0166] Figure 6 The agent strategy display diagram of the comparative method is shown in FIG. 4, Figure 7 The agent strategy display diagram of the proposed method is shown in FIG. 5, according to Figure 6 and Figure 7 It can be known that, since the comparative method needs to re-game in each rolling time domain, the existence of Nash equilibrium solution in the game process is difficult to guarantee, and it is difficult to guarantee the individual to achieve the optimal, so the time for the comparative method to complete the task is obviously increased compared with the proposed method. In this regard, the proposed method can effectively reduce the time for the agent system to complete the formation task and improve the system performance.

[0167] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.

[0168] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.

[0169] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.

[0170] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.

[0171] The above description merely provides a preferred embodiment of the present application, and is not intended to limit the present application to other forms. Any person skilled in the art can make equivalent changes or modifications to the above-mentioned embodiments without departing from the technical scope of the present application. Therefore, the scope of the technical solutions of the present application should be subject to the scope defined by the appended claims.

[0172] The present patent is not limited to the above-mentioned best mode, and anyone can derive other various forms of multi-agent system formation strategies based on hierarchical differential game from the present patent. Any equivalent changes and modifications made within the scope of the present patent application should be covered by the present patent.

Claims

1. A multi-agent system formation strategy based on hierarchical differential game theory, characterized in that, Includes the following steps: Step S1: Using graph theory, establish the communication relationship between agents in a multi-agent system: establish a first-order linear integrator as the agent model; set one agent as the leader and the rest as followers; only the leader knows the desired destination position, and the rest follow the leader to maintain the desired formation distance; define the working environment area of ​​the agents and treat the obstacles in the environment as ellipsoids; Step S2: At the policy layer, for the known obstacles in the environment, all agents participating in the task are regarded as game players. The formation problem of the multi-agent system with obstacle constraints is modeled as a distributed differential game. The game model includes the running cost function and the control cost. The players and their neighboring players meet the convergence to the local Nash equilibrium. Step S3: In the planning layer, for unknown obstacles in the environment, the local Nash equilibrium solution from the policy layer is modified in real time using a quadratic programming model based on rolling optimization to form a hybrid formation policy; the cost function of agent i in the planning layer is defined as: In the formula, At the end of the scrolling window, where t represents the initial time of the scrolling window; NP For the rolling time domain; The position error of the multi-agent system at the initial moment of the scrolling window is the difference between the optimal position from the policy layer and the actual position. and These are the actual and optimal positions of the multi-agent system at the initial moment of the scrolling window, respectively. Let i be the intelligent agent.

2. The multi-agent system formation strategy based on hierarchical differential game theory according to claim 1, characterized in that: In step S1: The specific form of the multi-agent dynamic equation is as follows: In the formula, t represents the time scale; x is the rate of change of the position information of the i-th agent at time t; i (t) represents the position information of the i-th agent at time t; u i (t) represents the control strategy of the i-th agent at time t; Define the state of the multi-agent system at time t as Where x1(t) and x N (t) represents the position information of the first agent and the Nth agent at time t, respectively; Construct a directed interaction topology graph G(v,ε) with N agents, where v={1,...,N} represents the set of agents; Denotes the set of edges; e ij The communication relationship between agent i and agent j; e ij ∈ε indicates that agent i can receive information from agent j; the neighbor set of agent i is In the formation control problem, considering that only the leader knows the desired target position, and the followers maintain relative distances to achieve the desired formation, the state error of agent i is defined as: In the formula, Let $t$ be the rate of change of the state error information of the $i$-th agent at time $t$. This represents the state error information of the i-th agent at time t; α i and β i Define the i-th agent as the leader: α i =1,β i =0, or a follower: α i =0,β i =1; x j (t) represents the position information of neighboring agent j of agent i at time t; x d The desired target location; Let be the expected grouping distance between the i-th agent and its neighboring agent j; Step S12: Establish the working environment area for the intelligent agent: Considering that the obstacles in the agent's working environment are ellipsoidal, define the collision avoidance area. for: In the formula, R 3 a is the three-dimensional plane in which the intelligent agent is located; i The safe distance for agent i; For the obstacle at time t Location, It is an obstacle The radius; I2 is the unit weight matrix; Define the sensing area for: In the formula, A i It is the sensing range of agent i; Define free region for: The following assumptions are given: Assumption 1: The obstacles in the working environment of the intelligent agent are sparse; Assumption 2: The initial positions of each agent do not coincide, that is: x i (t0)≠x j (t0) (6) In the formula, t0 is the initial working time of the agent, x i (t0) and x j (t0) represents the positions of the i-th agent and the j-th agent at the initial time t0, respectively; Assumption 3: The initial position and target position of each agent satisfy the following conditions: In the formula, t f x is the end time of the agent's work. i (t f Let ) represent the i-th agent at the final time t. f Location; Assumption 4: The desired formation distance satisfies the following conditions: This establishes the communication relationships between agents in a multi-agent system.

3. The multi-agent system formation strategy based on hierarchical differential game theory according to claim 2, characterized in that: In step S2: To achieve the goal of enabling the leader in a multi-agent system to reach the desired target point while maintaining the desired distance among followers, it is necessary to determine the optimal control strategy for each agent, where each agent is subject to the dynamic equation (1), the network topology G(v,ε), and the collision region. constraint; To achieve the control objective of the formation problem in multi-agent systems, the formation problem with known obstacle constraints is modeled as a distributed game problem, and the agents participating in completing the task are defined as game participants. The game cost for each participant is designed as follows: In the formula, u i and u -i Let x(t0) represent the control policy of agent i at time t and the set of control policies of neighboring agents (excluding agent i), respectively. Let x(t0) represent the initial position information of the multi-agent system at time t0. i (x(t)) represents the operating cost of the i-th agent at time t, R i Let be the adjustable positive definite weight matrix for the i-th agent; The operating cost function for the i-th agent is defined as follows: In the formula, and p is a constant i (x) and Let be the obstacle penalty function for the leader i at time t and the obstacle penalty function for the i-th follower, respectively, defined as follows: In the formula, The total number of obstacles in the agent's working environment; In the formula, a j The safe distance for agent j; The control objective of the formation control problem in a multi-agent system is to define each subject's dynamic equation (1), the network topology G(v,ε), and the known collision region. In a constrained game, players design optimal coordination strategies that ensure the leader safely reaches the desired objective while followers maintain relative formation positions. Player i and its neighbors converge to a local Nash equilibrium, i.e., the strategy set at time t. satisfy: In the formula, and Let be the optimal strategy of participant i at time t and the set of optimal strategies of their neighboring participants. and These represent the optimal and suboptimal costs of participant i relative to its neighboring participants, respectively.

Citation Information

Patent Citations

  • Multi-mobile robot formation method based on Q-learning

    CN114047758A

  • Distributed evolutionary game-based model prediction leader-free formation control method

    CN115616913A