A Collision Avoidance Strategy Method for Multi-Agent Systems Based on Robust Differential Games
Through a robust differential game method, combined with artificial potential field method and non-dominant ant colony optimization algorithm, the problems of low task efficiency and poor robustness caused by communication restrictions and external interference in the collision avoidance problem of multi-agents are solved, and the agent can quickly and reliably reach the target point.
Patent Information
- Application Number
- CN202310086217.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-16
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-01-16
AI Technical Summary
The existing differential game methods fail to effectively consider the limited communication capabilities and external interference of the agent in the collision avoidance problem of multiple agents, resulting in the lack of robustness of the collision avoidance strategy and low task completion efficiency.
A method based on robust differential game is adopted, and collision avoidance rules are designed in combination with artificial potential field method, a degree of deviation from the target is introduced, a distributed zero-sum game model is constructed, and an optimal control theory and non-dominant ant colony optimization algorithm are used to solve the optimal feedback gain to ensure the global convergence of local Nash equilibrium.
It reduces the time for the agent to reach the target point, improves the robustness of the collision avoidance strategy, and ensures efficient completion of tasks under a fixed strong connectivity topology diagram.
Smart Images

Figure CN115981163B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-agent coordinated control, and particularly relates to a collision avoidance strategy method for multi-agent systems based on robust differential games. Background Art
[0002] In the past decade, multi-agent systems have attracted much attention due to their loosely coupled network structure, and agents can interact with each other to solve problems that cannot be solved by a single agent. In multi-agent systems, agent collision avoidance is a prerequisite for agents to safely execute tasks collaboratively.
[0003] Game theory is an effective tool for solving multi-agent decision-making, and differential games are widely applied to the field of multi-agent coordinated control. Differential game is the combination of game theory and optimal control. Introducing differential game method into multi-agent coordinated control can fully reflect the dynamic interaction among agents. Compared with distributed optimization algorithms, the differential game method does not require a central coordination mechanism. Agents only need to selfishly optimize their own cost functions and can still converge to the Nash equilibrium in the end, with strict mathematical guarantees. Currently, the methods based on differential games have achieved success in application fields such as solving the pursuit-evasion problem considering wired communication capabilities and the formation problem considering external disturbances. For example, the literature (Lin W, Qu Z, Simaan M A. Nash strategies for pursuit-evasion differential games involving limited observations[J]. IEEE Transactions on Aerospace and Electronic Systems, 2015, 51(2): 1347-1356.) proposed a method for constructing feedback pursuit-evasion strategies, which does not rely on the global state information of agents. The literature (de la Cruz N, Jimenez-Lizarraga M. Finite time robust feedback Nash equilibrium for linear quadratic games[J]. IFAC-PapersOnLine, 2017, 50(1): 11794-11799.) established a centralized differential game model with external disturbances, regarding external disturbances as virtual players that maximize the cost function, but did not consider the limited communication capabilities of agents. The literature (Fu Y, Chai T. Online solution of two-player zero-sum games for continuous-time nonlinear systems with completely unknown dynamics[J]. IEEE transactions on neural networks and learning systems, 2015, 27(12): 2577-2587.) constructed a distributed uncertain zero-sum differential game and obtained a local robust Nash equilibrium, but without strict theoretical guarantees. To achieve the coordination of multi-agent global tasks, the global convergence guarantee of local robust Nash equilibrium is required.Considering that traditional differential game methods do not consider the communication ability limitations of agents and external interference problems when solving the multi-agent collision avoidance problem, the collision avoidance strategy lacks robustness and cannot ensure the efficient and smooth completion of tasks. Therefore, in order to better achieve the safe and efficient completion of tasks by multi-agents, it is necessary to establish a corresponding differential game model for the limited communication ability of agents and the existing external interference problems, so as to improve the robustness of the collision avoidance strategy and minimize the time for agents to complete tasks as much as possible.
[0004] Therefore, in order to solve the problems of low task completion efficiency and poor control performance caused by introducing differential game methods into the collision avoidance problem, it is possible to consider introducing the artificial potential field method to design collision avoidance rules, and considering the interference as a virtual player method that maximizes the cost function. Design a collision avoidance strategy method for multi-agent systems based on robust differential games. The existing solutions based on robust differential games proposed by current technologies mainly focus on the case where the global information of agents is known. For distributed robust differential game methods, they are still rarely applied to multi-agent collision avoidance problems and cannot provide appropriate solutions. Summary of the Invention
[0005] The purpose of the present invention is to overcome the defects and deficiencies existing in the prior art, and provide a collision avoidance strategy method for multi-agent systems based on robust differential games. This method considers the existing differential game methods that only consider the obstacle avoidance goal. Based on the artificial potential field method, it introduces the distance from the target to penalize the degree of deviation of the agent from the target point, weighs the distance between the agent reaching the target point and the distance from the obstacle, and reduces the time for the agent to reach the target point. For the existing external interference problem, the interference and the control strategy form a zero-sum game relationship, and the optimal controller in the worst interference case is solved. Based on the optimal control principle, under the assumption of a fixed strongly connected topological graph, the global convergence of the local Nash equilibrium solution is guaranteed. For the limited communication ability of agents, considering that the traditional method of solving the Riccati equation is no longer applicable, an inverse optimization method based on the best performance index is introduced to construct the optimal feedback strategy, and the optimal feedback gain is solved using the non-dominated ant colony optimization algorithm. This method can reduce the time for agents to reach the target point and can achieve the robustness of the collision avoidance strategy.
[0006] To achieve the above purpose, the technical solution of the present invention is: a collision avoidance strategy method for multi-agent systems based on robust differential games, including the following steps:
[0007] Step S1: Use graph theory to establish the communication relationship between agents in the multi-agent system; regard the agent and its neighbors as game participants, and establish a first-order linear integrator as the model of the agent; define the collision area, sensing area, and free area for the working environment of the agent, and regard the obstacle as an ellipse to encompass all shapes of obstacles;
[0008] Step S2: Design a collision avoidance rule using the artificial potential field method as the running cost function of the agent in the game model;
[0009] Step S3: Regard the collision avoidance problem of a multi-agent system with limited communication capabilities and external interference as a distributed zero-sum differential game problem; establish a distributed robust differential game model, which includes a running cost function, a control cost, an interference cost, and a terminal cost;
[0010] Step S4: Use the optimal control theory to establish a local robust value function, obtain the Hamilton-Jacobi-Isaacs (HJI) equation according to the obtained local robust value function, and solve the HJI equation to obtain the expression form of the optimal controller; analyze the relationship between the optimal control and the local robust Nash equilibrium, as well as the global convergence of the local robust Nash equilibrium;
[0011] Step S5: Solve the local robust Nash equilibrium of the agent by using the inverse optimization method based on the approximate optimal performance index.
[0012] In an embodiment of the present invention, step S1 specifically includes the following steps:
[0013] Step S11: Establish a game participant model:
[0014] The specific form of the dynamic equation of the multi-agent system is:
[0015]
[0016] where t is the time scale; is the change rate of the position information of the i-th agent at time t; x i (t) is the position information of the i-th multi-agent at time t; u i (t) and u j (t) are the control strategies of the i-th agent and the j-th agent at time t, respectively; ω i (t) and ω j (t) are the interference strategies of the i-th agent and the j-th agent at time t, respectively; B ii , B ij , E ii , E ij are constant matrices corresponding to the respective strategies;
[0017] Establish a directed interaction topology graph G(v, ε) of N agents, where v = {v1,..., v N} represents the set of agents; represents the set of edges; e ij represents the communication relationship between agent i and agent j; e ij∈ε means that agent i can receive information from agent j; the neighbor set of agent i is
[0018] Define the local dynamic equation of agent i as:
[0019]
[0020] where is the change rate of the local state information of the i-th agent at time t; is the local state information of the i-th multi-agent at time t; u ij (t) is the inference strategy of agent i for neighbor agent j at time t; ω ij is the interference strategy inferred by the i-th agent for neighbor agent j at time t; are the constant matrices corresponding to the respective strategies;
[0021] Step S12: Establish an obstacle environment model:
[0022] Considering the obstacle as an ellipse, define the collision avoidance area S ik as:
[0023]
[0024] where R 2 is the two-dimensional plane where the agent is located; r i is the safety distance of agent i; c k (t) is the position of obstacle k at time t, is the radius of obstacle k; I2 is the unit weight matrix;
[0025] Define the sensing area D ik as:
[0026]
[0027] where R i is the sensing range of agent i;
[0028] Define the free area M ik as:
[0029]
[0030] In an embodiment of the present invention, in step S2, the following assumption conditions are given:
[0031] Assumption 1: ω i (t) is square integrable, for there exists a constant satisfying the following conditions:
[0032]
[0033] where t f is the running time at the end of agent i; is a certain positive constant; is a positive real number;
[0034] Hypothesis 2: The directed interaction topology graph G(v, ε) is fixed and strongly connected;
[0035] Design a collision avoidance rule based on the artificial potential field method:
[0036]
[0037] where is the penalty function of agent i at time t; is the deviation between the current position of the multi-agent system and the target point position at time t; χ i (0 < χ i < 1) and are constants respectively; is the obstacle penalty function of agent i at time t; is the distance penalty function of agent i at time t;
[0038] The distance penalty function is expressed as follows:
[0039]
[0040] where is the target point position of agent i; is the deviation between the current position of agent i and the target point position at time t;
[0041] To optimize the trajectory of the agent, a distance penalty function is introduced to penalize the deviation degree of the agent from the target point, which is expressed as follows:
[0042]
[0043] where γ i (t) is the deviation angle of agent i at time t, and this deviation angle is the angle between the current position of the agent and the predefined reference trajectory.
[0044] In an embodiment of the present invention, in step S3,
[0045] A distributed robust differential game cost function is established, and the expression form is as follows:
[0046]
[0047] where Can be abbreviated as J i ; t f is the end running time of agent i; is the initial state information of the multi-agent system; u -i (t) and ω -i (t) are the control strategy and interference strategy of neighbor agents other than agent i at time t, respectively; are the optimal control strategy and optimal interference strategy of neighbor agents other than agent i at time t, respectively; is the optimal control strategy of agent i at time t; is the terminal cost, is the running cost of agent i at time t, is the control cost of agent i at time t, is the interference cost of agent i at time t, F ii 、R ii 、R ij 、W ii 、W ij are adjustable positive definite weight matrices, respectively;
[0048] The goal of the collision avoidance problem for the distributed multi-agent system is to design a feedback control strategy for each agent and safely reach the target point within a finite time domain; at the same time, agent i and its neighbor agents can converge to the global Nash equilibrium, that is, the strategy set satisfies:
[0049]
[0050]
[0051] wherein, is the optimal cost of agent i at time t; is the optimal strategy of agent i at time t.
[0052] In an embodiment of the present invention, in step S4,
[0053] The expression form of the local robust value function of the agent i is as follows:
[0054]
[0055] wherein, J i is the distributed robust differential game cost function;
[0056] The optimal control strategy is: assume that the optimal controller u * can minimize the value function, that is, the optimal controller u * can make reach the optimal value function, and the expression form is as follows:
[0057]
[0058] Sufficient conditions for the existence of local robust Nash equilibrium solutions are as follows: Assume that for all and exist and are continuous, then the local robust value function V i satisfies the following distributed robust Hamilton-Jacobi-Isaacs (HJI) equation:
[0059]
[0060] where is the rate of change of the state information of the i-th agent at time t; are the partial derivatives of the local robust value function of agent i with respect to the state deviation of the multi-agent system and with respect to time t, respectively;
[0061] Then the optimal robust control strategy and the worst-case disturbance strategy of agent i at time t are respectively:
[0062]
[0063]
[0064] Regarding the differential game problem as a linear quadratic problem, let:
[0065]
[0066] where D i is the neighbor information matrix of agent i; P i (t) = P i T (t) > 0 is a positive definite matrix at time t;
[0067] Then the optimal robust control strategy and the worst-case disturbance strategy of agent i at time t are respectively:
[0068]
[0069]
[0070] where P i (t) satisfies the following coupled Riccati equation:
[0071]
[0072] where is the rate of change of the matrix P i (t) at time t; P i (t f ) = Fii ;
[0073] Assume that the sufficient conditions for the existence of the local robust Nash equilibrium solution hold, and the neighbor agents have reached the optimal strategy. Let the optimal robust control strategy and the worst interference strategy form of agent i satisfy in the form of the formula, then the strategies of agent i and its neighbor agents will converge to the local robust Nash equilibrium solution;
[0074] Let be the optimal control strategy and the worst interference strategy of agent i relative to its neighbor agents at time t, respectively. The local robust Nash equilibrium can converge to the global Nash equilibrium if and only if the graph G(v, ε) is strongly connected, that is
[0075] In an embodiment of the present invention, in step S5,
[0076] The local robust Nash equilibrium solution of the agent is constructed by the inverse optimization method based on the approximate best performance index, and the optimal feedback gain is solved by using the ant colony optimization algorithm based on non-dominated dominance;
[0077] Construct the local robust feedback control strategy of agent i at time t as:
[0078]
[0079]
[0080] wherein, are the feedback gain matrix corresponding to the constructed control strategy of agent i at time t and the feedback gain matrix corresponding to the constructed interference strategy, respectively; are the constructed control strategy and the interference strategy of agent i at time t, respectively;
[0081] Solving the constructed local optimal robust optimal strategy set is equivalent to solving the optimal feedback gain matrix where, are the local optimal constructed robust strategy and the local optimal constructed interference strategy of agent i at time t, respectively, are the local optimal constructed robust strategy and the local optimal constructed interference strategy of the neighbor agents other than agent i at time t, respectively; are the optimal feedback gain matrix corresponding to the local optimal constructed robust strategy of agent i at time t and the optimal feedback gain matrix corresponding to the local optimal constructed interference strategy, respectively; They are the optimal feedback gain matrix corresponding to the locally optimal construction of the robust strategy and the optimal feedback gain matrix corresponding to the locally optimal construction of the interference strategy for the neighbor agents other than agent i at time t, respectively;
[0082] The constructed approximate best performance index is:
[0083]
[0084] where is the constructed correlation coefficient, defined as follows:
[0085]
[0086]
[0087]
[0088]
[0089]
[0090]
[0091]
[0092] In an embodiment of the present invention, for the local cost function described in step S3, the constructed approximate best performance index can be deformed into the following form:
[0093]
[0094] where
[0095]
[0096]
[0097] In an embodiment of the present invention, for the local cost function described in step S3, the constraint condition of the constructed feedback gain matrix can be obtained, and the problem of solving the optimal feedback gain matrix is transformed into a multi-objective optimization problem;
[0098] The optimal feedback gain matrix constructed is solved by using the ant colony optimization algorithm based on non-dominated dominance, and the corresponding local robust Nash equilibrium solution is obtained;
[0099] For agent i at time t, the multi-objective optimization function is defined as:
[0100]
[0101] where υc > 0 is a constant; the goal of this multi-objective optimization problem is to find the optimal feedback gain set to minimize the function the function value.
[0102] Compared with the prior art, the present invention has the following beneficial effects:
[0103] The present invention and its preferred solution transform the multi-agent collision avoidance problem with limited communication capabilities into a distributed differential game problem for a first-order linear model with external disturbances; considering the existing differential game methods that only consider the obstacle avoidance goal, based on the artificial potential field method, a trajectory optimization goal is introduced to penalize the degree of deviation of the agent from the target point, weighing the distance between the agent reaching the target point and the obstacle, and reducing the time for the agent to reach the target point; for the problem of external disturbances, a zero-sum game relationship is formed between the disturbance and the control strategy to solve the optimal controller under the worst disturbance; based on the optimal control principle, under the assumption of a fixed strongly connected topological graph, the global convergence of the local Nash equilibrium solution is guaranteed; for the limited communication capabilities of the agents, considering that the traditional method of solving the Riccati equation is no longer applicable, an inverse optimization method based on the best performance index is introduced to construct the optimal feedback strategy, and the non-dominated ant colony optimization algorithm is used to solve the optimal feedback gain; this method can reduce the time for the agent to reach the target point and can achieve the robustness of the collision avoidance strategy; in addition, a distributed architecture differential game model is introduced, which has good scalability compared with the centralized architecture. BRIEF DESCRIPTION OF THE DRAWINGS
[0104] The present invention will be further described in detail below with reference to the drawings and specific embodiments. The drawings constituting a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention; in the drawings,
[0105] Figure 1 is the method principle block diagram of the embodiment of the present invention;
[0106] Figure 2 is the schematic diagram of the collision avoidance rules based on the artificial potential field method of the embodiment of the present invention;
[0107] Figure 3 is the display diagram of the task completion time of multi-agents in the simulation example of the present invention;
[0108] Figure 4 is the multi-agent communication topological graph adopted in the simulation example of the present invention;
[0109] Figure 5 is the position error change diagram in the simulation example of the present invention;
[0110] Figure 6It is a diagram showing the collision avoidance performance of multiple agents in the simulation example of the embodiment of the present invention. Detailed implementation mode
[0111] It should be noted that the following detailed description is exemplary and is intended to provide further description of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0112] It should be noted that the terms used herein are only for describing specific implementation modes and are not intended to limit the exemplary implementation modes according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "include" and / or "comprise" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0113] As Figure 1 shown, this embodiment provides a collision avoidance strategy method for a multi-agent system based on robust differential game, including the following steps:
[0114] Establish the communication relationship between agents through graph theory. Agents can obtain the position information of their neighboring agents, establish a first-order linear model as the game participants, model the working environment of the agents, and regard the obstacles as ellipses;
[0115] Design a collision avoidance rule based on the artificial potential field method, as Figure 2 shown, and establish a distributed robust differential game model;
[0116] Introduce an optimal robust control strategy, calculate the cost function and the local robust value function through the obtained position error, and analyze the relationship between the optimal control and the local robust Nash equilibrium solution, as well as the existence and global convergence of the Nash equilibrium solution;
[0117] Introduce an inverse optimization method based on the best performance index and a non-dominated ant colony optimization algorithm to solve the optimal control strategy.
[0118] In this embodiment, the tasks executed by the agents are: under the condition that the multi-agents are affected by external interference and the communication is restricted, make the multi-agents reach the target point without collision from the initial position, and reduce the time for the agents to complete the tasks.
[0119] In this embodiment, the specific form of the multi-agent dynamic equation is:
[0120]
[0121] In the formula, t is the time scale; is the rate of change of the position information of the $i$-th agent at time $t$; $x$ i (t) is the position information of the $i$-th multi-agent at time $t$; $u$ i (t) and $u$ j (t) are the control strategies of the $i$-th agent and the $j$-th agent at time $t$ respectively; $\omega$ i (t) and $\omega$ j (t) are the interference strategies of the $i$-th agent and the $j$-th agent at time $t$ respectively; $B$ ii , $B$ ij , $E$ ii , $E$ ij are constant matrices corresponding to the respective strategies;
[0122] Construct a directed interaction topology graph $G(v, \epsilon)$ of $N$ agents, where $v = \{v_1,..., v$ N $\}$ represents the set of agents; represents the set of edges; $e$ ij represents the communication relationship between agent $i$ and agent $j$; $e$ ij $\in \epsilon$ means that agent $i$ can receive information from agent $j$; the neighbor set of agent $i$ is
[0123] Define the local dynamic equation of agent $i$ as:
[0124]
[0125] In the formula, is the rate of change of the local state information of the $i$-th agent at time $t$; is the local state information of the $i$-th multi-agent at time $t$; $u$ ij (t) is the inference strategy of agent $i$ for neighbor agent $j$ at time $t$; $\omega$ ij is the interference strategy inferred by the $i$-th agent for neighbor agent $j$ at time $t$; are constant matrices corresponding to the respective strategies;
[0126] Construct an obstacle environment model:
[0127] Considering the obstacle as an ellipse, define the collision avoidance area $S$ ik as:
[0128]
[0129] In the formula, $R$ 2 is the two-dimensional plane where the agent is located; $r$ i is the safety distance of agent $i$; $c$ i (t) is the position of obstacle $k$ at time $t$, is the radius of obstacle $k$; $I_2$ is the unit weight matrix;
[0130] Define the sensing area D ik as:
[0131]
[0132] where R i is the sensing range of agent i;
[0133] Define the free area M ik as:
[0134]
[0135] The following assumptions are given:
[0136] Assumption 1: ω i (t) is square-integrable, and for there exists a constant satisfying the following conditions:
[0137]
[0138] where t f is the running time at the end of agent i; is a certain positive constant; is a positive real number;
[0139] Assumption 2: The directed interaction topology graph G(v, ε) is fixed and strongly connected;
[0140] Design a collision avoidance rule based on the artificial potential field method:
[0141]
[0142] where is the penalty function of agent i at time t; is the deviation between the current position of the multi-agent system and the target point position at time t; χ i (0 < χ i < 1) and are constants respectively; is the obstacle penalty function of agent i at time t; is the distance penalty function of agent i at time t;
[0143] The distance penalty function is expressed as follows:
[0144]
[0145] where is the target point position of agent i; is the deviation between the current position of agent i and the target point position at time t;
[0146] To optimize the trajectory of the agent, a distance penalty function is introduced to penalize the deviation of the agent from the target point, which is expressed as follows:
[0147]
[0148] In the formula, γ i (t) is the deviation angle of agent i at time t, which is the angle between the current position of the agent and the predefined reference trajectory;
[0149] A distributed robust differential game cost function is established, and its expression form is as follows:
[0150]
[0151] In the formula, can be abbreviated as J i ; t f is the end running time of agent i; is the initial state information of the multi-agent system; u -i (t) and ω -i (t) are the control strategy and interference strategy of the neighbor agents except agent i at time t, respectively; are the optimal control strategy and optimal interference strategy of the neighbor agents except agent i at time t, respectively; is the optimal control strategy of agent i at time t; is the terminal cost, is the running cost of agent i at time t, is the control cost of agent i at time t, is the interference cost of agent i at time t, F ii , R ii , R ij , W ii , W ij are adjustable positive definite weight matrices, respectively;
[0152] The goal of the collision avoidance problem of the distributed multi-agent system is to design a feedback control strategy for each agent and safely reach the target point within a finite time domain; meanwhile, agent i and its neighbor agents can converge to the global Nash equilibrium, that is, the strategy set satisfies:
[0153]
[0154] In the formula, is the optimal cost of agent i at time t; is the optimal strategy of agent i at time t;
[0155] The expression form of the local robust value function of the agent i is as follows:
[0156]
[0157] In the formula, J i is the distributed robust differential game cost function described in step S3;
[0158] The optimal control strategy is: Let the optimal controller u * be able to minimize the value function, that is, the optimal controller u * can make reach the optimal value function, and the expression form is as follows:
[0159]
[0160] The sufficient condition for the existence of the local robust Nash equilibrium solution is given as: Assume that for all and exist and are continuous, then the local robust value function V i satisfies the following distributed robust Hamilton-Jacobi-Isaacs (HJI) equation:
[0161]
[0162] In the formula, is the change rate of the state information of the i-th agent at time t; are respectively the partial derivative of the local robust value function of agent i with respect to the state deviation of the multi-agent system and the partial derivative with respect to time t;
[0163] Then the optimal robust control strategy and the worst interference strategy of agent i at time t can be written as:
[0164]
[0165]
[0166] Regarding this differential game problem as a linear quadratic problem, let:
[0167]
[0168] In the formula, D i is the neighbor information matrix of agent i; P i (t) = P i T (t) > 0 is a positive definite matrix at time t;
[0169] Then the optimal robust control strategy and the worst interference strategy of agent i at time t are:
[0170]
[0171]
[0172] In the formula, P i (t) satisfies the following coupled Riccati equation:
[0173]
[0174] In the formula, is the rate of change of the matrix P i (t); P i (t f ) = F ii ;
[0175] Assume that the sufficient conditions for the existence of the local robust Nash equilibrium solution hold, and the neighbor agents have reached the optimal strategy. Let the optimal robust control strategy and the worst interference strategy form of agent i satisfy the forms in (17a) and (17b). Then the strategies of agent i and its neighbor agents will converge to the local robust Nash equilibrium solution;
[0176] Let be the optimal control strategy and the worst interference strategy of agent i relative to its neighbor agents at time t, respectively. When and only when the graph G(v, ε) is strongly connected, the local robust Nash equilibrium can converge to the global Nash equilibrium, that is
[0177] In step S5,
[0178] Construct the local robust Nash equilibrium solution of the agent by the inverse optimization method based on the approximate optimal performance index, and use the ant colony optimization algorithm based on non-dominated dominance to solve the optimal feedback gain;
[0179] Construct the local robust feedback control strategy of agent i at time t as:
[0180]
[0181]
[0182] In the formula, are the feedback gain matrix corresponding to the constructed control strategy of agent i and the feedback gain matrix corresponding to the constructed interference strategy at time t, respectively; are the constructed control strategy and the interference strategy of agent i at time t, respectively;
[0183] Solving the constructed local optimal robust optimal strategy set is equivalent to solving the optimal feedback gain matrix wherein, are respectively the local optimal construction robust strategy and the local optimal construction interference strategy of agent i at time t, are respectively the local optimal construction robust strategy and the local optimal construction interference strategy of neighbor agents other than agent i at time t; are respectively the optimal feedback gain matrix corresponding to the local optimal construction robust strategy of agent i at time t and the optimal feedback gain matrix corresponding to the local optimal construction interference strategy; are respectively the optimal feedback gain matrix corresponding to the local optimal construction robust strategy of neighbor agents other than agent i at time t and the optimal feedback gain matrix corresponding to the local optimal construction interference strategy;
[0184] The constructed approximate optimal performance index is:
[0185]
[0186] wherein, is the constructed correlation coefficient, defined as follows:
[0187]
[0188]
[0189]
[0190]
[0191] According to the local cost function described in step S3, the constructed approximate optimal performance index can be transformed into the following form:
[0192]
[0193] wherein,
[0194]
[0195] According to the local cost function described in step S3, the constraint conditions of the constructed feedback gain matrix can be obtained, and the problem of solving the optimal feedback gain matrix is transformed into a multi-objective optimization problem;
[0196] Use the ant colony optimization algorithm based on non-dominated dominance to solve the constructed optimal feedback gain matrix and obtain the corresponding local robust Nash equilibrium solution.
[0197] For agent i at time t, the multi-objective optimization function is defined as:
[0198]
[0199] wherein υ c > 0 is a constant; the objective of this multi-objective optimization problem is to find the optimal feedback gain set to minimize the function function value to the minimum
[0200] In this embodiment, a specific example is given to demonstrate the effectiveness and superiority of the proposed distributed robust differential game method in solving the multi-agent collision avoidance problem
[0201] According to Figure 3 it can be seen that to prove that the obtained optimal controller can reduce the time for the agents to complete the task, a simulation experiment is conducted in this example and compared with the existing centralized differential game method that only considers the obstacle penalty objective. The specific model expression of the multi-agent system is as follows
[0202]
[0203]
[0204] wherein are the change rates of the local state information of the first agent and the second agent at time t respectively are the local state information of the first agent and the second agent at time t respectively; u1(t) and u2(t) are the control strategies of the first agent and the second agent at time t respectively
[0205] The initial positions and target point positions of each agent are: x1(0) = [30, 30] T , x2(0) = [370, 370] T , The sensing radii are R1 = R2 = 48
[0206] The specific forms of the benefit functions of each agent are as follows
[0207]
[0208]
[0209] The model form of the multi-objective is as follows
[0210]
[0211] wherein υ1 = 0.5, υ2 = 0.3, υ3 = 0.2
[0212] According to Figure 3It can be seen that in this simulation case, the centralized differential game method that only considers obstacle penalty is compared with the centralized differential game method that introduces the trajectory optimization objective, that is, the comparison method is compared with the proposed method. Although both methods can make the position deviation of the agent tend to 0, the convergence time of the centralized differential game method that introduces the trajectory optimization objective is 49s, and the convergence time of the centralized differential game method that only considers obstacle penalty is 59s. Therefore, the centralized differential game method that introduces the trajectory optimization objective can reduce the time for the agent to complete the task.
[0213] According to Figure 4 It can be seen that this example provides a directed communication topology diagram of 3 agents. To prove that the obtained optimal controller is robust, this example conducts a simulation experiment and compares it with the distributed differential game method that does not consider interference in the existing cost function. The specific model expression form of the multi-agent system is as follows:
[0214]
[0215]
[0216]
[0217] In the formula. They are respectively the change rates of the local state information of the first agent, the second agent, and the third agent at time t; They are respectively the local state information of the first agent, the second agent, and the third agent at time t; u1(t), u2(t), and u3(t) are respectively the control strategies of the first agent, the second agent, and the third agent at time t; u 13 (t), u 21 (t), u 32 (t), ω 13 (t), ω 21 (t), ω 32 (t) are respectively the inferred control strategies and inferred interference strategies of the first agent, the second agent, and the third agent at time t for the neighboring agents;
[0218] The specific forms of the benefit functions of each agent are as follows:
[0219]
[0220]
[0221]
[0222] The model form of the multi-objective is as follows:
[0223]
[0224] In the formula, υ1 = 0.5, υ2 = 0.2, υ3 = 0.1, υ4 = 0.1, υ5 = 0.1.
[0225] The initial positions of each agent and the target point position are: x1(0) = [30, 30] T , x2(0) = [370, 370] T , x3(0) = [370, 30] T , x3(0) = [30, 370] T , R1 = R2 = 48, and the external disturbance is ω i = sin(t).
[0226] According to Figure 5 it can be seen that the proposed distributed robust differential game method can make the position deviation of each agent tend to 0, indicating that the task is completed. And according to Figure 6 it can be seen that in the compared distributed differential game method, agent 2 collides with the obstacle, and the final position deviation is not 0, indicating that the task is not completed. This example shows that the collision avoidance strategy of the proposed method is robust.
[0227] It should be noted that the present invention is not limited to the content shown in the above exemplary examples, and the present invention can be implemented in other forms without departing from the basic features of the present invention. Therefore, the examples should be regarded as exemplary rather than restrictive, and the scope of the present invention is determined by the appended claims rather than the above description, aiming to include all changes falling within the meaning and scope of the equivalent elements of the claims within the present invention. Without departing from the principle of the present invention, several modifications and improvements made to the present invention should be regarded as within the protection scope of the present invention.
Claims
1. A collision avoidance strategy method for multi-agent systems based on robust differential games, characterized in that, It includes the following steps: Step S1: Using graph theory, establish the communication relationship between agents in the multi-agent system; take the agent and its neighbors as game participants, and establish a first-order linear integrator as the model of the agent; define the collision area, sensing area, and free area for the working environment of the agent, and regard the obstacle as an ellipse to encompass all shapes of obstacles; Step S2: Use the artificial potential field method to design the collision avoidance rule as the running cost function of the agent in the game model; Step S3: Regard the collision avoidance problem of the multi-agent system with limited communication ability and external interference as a distributed zero-sum differential game problem; establish a distributed robust differential game model, which includes the running cost function, control cost, interference cost, and terminal cost; Step S4: Using the optimal control theory, establish a local robust value function, obtain the Hamilton-Jacobi-Isaacs (HJI) equation according to the obtained local robust value function, and solve the HJI equation to get the expression form of the optimal controller; analyze the relationship between the optimal control and the local robust Nash equilibrium, as well as the global convergence of the local robust Nash equilibrium; Step S5: Use the inverse optimization method based on the approximate best performance index to solve the local robust Nash equilibrium of the agent.
2. The collision avoidance strategy method for a multi-agent system based on robust differential game according to claim 1, characterized in that The specific steps of Step S1 include the following steps: Step S11: Establish the game participant model: The specific form of the dynamic equation of the multi-agent system is: where t is the time scale; is the change rate of the position information of the i-th agent at time t; x i (t) is the position information of the i-th multi-agent at time t; u i (t) and u j (t) are the control strategies of the i-th agent and the j-th agent at time t, respectively; ω i (t) and ω j (t) are the interference strategies of the i-th agent and the j-th agent at time t, respectively; B ii 、B ij 、E ii 、E ij are constant matrices corresponding to the respective strategies; Construct a directed interaction topology graph \(G(v, \epsilon)\) of \(N\) agents, where \(v=\{v_1, \ldots, v N \}\) represents the set of agents; \(\epsilon\) represents the set of edges; \(e ij _{ij}\) represents the communication relationship between agent \(i\) and agent \(j\); \(e ij _{ij} \in \epsilon\) indicates that agent \(i\) can receive information from agent \(j\); the neighbor set of agent \(i\) is Define the local dynamic equation of agent i as: wherein, is the change rate of the local state information of the i-th agent at time t; is the local state information of the i-th multi-agent at time t; u ij (t) is the inference strategy of agent i for neighbor agent j at time t; ω ij is the interference strategy inferred by the i-th agent for neighbor agent j at time t; are the constant matrices corresponding to the respective strategies; Step S12: Establish the obstacle environment model: Considering the obstacle is elliptical, define the collision avoidance area S ik as follows: where R 2 is the two-dimensional plane where the agent is located; r i is the safety distance of agent i; c k (t) is the position of obstacle k at time t, is the radius of obstacle k; I2 is the unit weight matrix; Define the induction area D ik as follows: where R i is the sensing range of agent i; Define the free area M ik as follows:
3. A collision avoidance strategy method for a multi-agent system based on robust differential game according to claim 2, characterized in that In Step S2, the following assumptions are given: Hypothesis 1: ω i (t) is square-integrable, for There exists a constant satisfying the following conditions: where \(t\) f is the running time at the end of agent \(i\); is a certain positive constant; is a positive real number; Assumption 2: The directed interaction topology graph G(v, ε) is fixed and strongly connected; Design the collision avoidance rule based on the artificial potential field method: wherein, is the penalty function of agent i at time t; is the deviation between the current position of the multi-agent system and the target point position at time t; χ i (0 < χ i < 1) and are constants respectively; is the obstacle penalty function of agent i at time t; is the distance penalty function of agent i at time t; The representation of the distance penalty function is as follows: wherein, is the target point position of agent i; is the deviation between the current position and the target point position of agent i at time t; To optimize the trajectory of the agent, introduce a distance penalty function to punish the deviation degree of the agent from the target point, which is expressed as follows: where γ i (t) is the deviation angle of agent i at time t, which is the angle between the current position of the agent and the predefined reference trajectory.
4. A collision avoidance strategy method for a multi-agent system based on robust differential game according to claim 3, characterized in that In Step S3, Establish the distributed robust differential game cost function, and its expression form is as follows: wherein, can be abbreviated as J i ; t f is the end running time of the agent i; is the initial state information of the multi-agent system; u -i (t) and ω -i (t) are the control strategy and interference strategy of neighbor agents other than agent i at time t, respectively; are the optimal control strategy and optimal interference strategy of neighbor agents other than agent i at time t, respectively; is the optimal control strategy of agent i at time t; is the terminal cost, is the running cost of agent i at time t, is the control cost of agent i at time t, is the interference cost of agent i at time t, F ii 、R ii 、R ij 、W ii 、W ij are adjustable positive definite weight matrices, respectively; The goal of the collision avoidance problem of the distributed multi-agent system is to design a feedback control strategy for each agent and safely reach the target point within a finite time domain; Meanwhile, agent i and its neighbor agents can converge to the global Nash equilibrium, i.e., the strategy set satisfies: wherein, is the optimal cost of agent i at time t; is the optimal policy of agent i at time t.
5. A collision avoidance strategy method for a multi-agent system based on robust differential game according to claim 4, characterized in that In Step S4, The expression form of the local robust value function of agent i is as follows: where, J i is the distributed robust differential game cost function; The optimal control strategy is as follows: Let the optimal controller be u * which can minimize the value function, that is, the optimal controller u * can make reach the optimal value function, and the expression is as follows: Sufficient conditions for the existence of local robust Nash equilibrium solutions are given as follows: Assume that for all and exist and are continuous, then the local robust value function V i satisfies the following distributed robust Hamilton-Jacobi-Isaacs HJI equation: wherein, is the rate of change of the state information of the i-th agent at time t; are respectively the partial derivatives of the local robust value function of agent i with respect to the state deviation of the multi-agent system and the partial derivative with respect to time t; Then the optimal robust control strategy and the worst interference strategy of agent i at time t are respectively: Regard the differential game problem as a linear quadratic problem, and let: where D i is the neighbor information matrix of agent i; P i (t) = P i T (t) > 0 is a positive definite matrix at time t; Then the optimal robust control strategy and the worst interference strategy of agent i at time t are respectively: where P i (t) satisfies the following coupled Riccati equations: In the formula, is the rate of change of matrix P i (t); P i (t f ) = F ii ; Assume that the sufficient conditions for the existence of the local robust Nash equilibrium solution hold, and the neighboring agents have reached the optimal strategy. Let the optimal robust control strategy of agent i and the form of the worst-case disturbance strategy satisfy in the form of the formula, then the strategies of agent i and the neighboring agents will converge to the local robust Nash equilibrium solution; Let be the optimal control strategy and the worst interference strategy of agent i relative to its neighbor agents at time t, respectively. When and only when the graph G(v, ε) is strongly connected, the local robust Nash equilibrium can converge to the global Nash equilibrium, that is 6. A collision avoidance strategy method for a multi-agent system based on robust differential game according to claim 5, characterized in that In Step S5, The inverse optimization method based on the approximate best performance index constructs the local robust Nash equilibrium solution of the agent, and uses the ant colony optimization algorithm based on non-dominated dominance to solve the optimal feedback gain; Construct the local robust feedback control strategy of agent i at time t as: wherein, are respectively the feedback gain matrix corresponding to the construction control strategy of agent i at time t and the feedback gain matrix corresponding to the construction interference strategy; are respectively the construction control strategy and the interference strategy of agent i at time t; Solving the constructed set of locally optimal robust optimal strategies is equivalent to solving the optimal feedback gain matrix where are respectively the locally optimal constructed robust strategy and the locally optimal constructed interference strategy of agent i at time t, are respectively the locally optimal constructed robust strategy and the locally optimal constructed interference strategy of the neighbor agents other than agent i at time t; are respectively the optimal feedback gain matrix corresponding to the locally optimal constructed robust strategy and the optimal feedback gain matrix corresponding to the locally optimal constructed interference strategy of agent i at time t; are respectively the optimal feedback gain matrix corresponding to the locally optimal constructed robust strategy and the optimal feedback gain matrix corresponding to the locally optimal constructed interference strategy of the neighbor agents other than agent i at time t; Construct the approximate best performance index as: wherein, is the constructed correlation coefficient, which is defined as follows:
7. A collision avoidance strategy method for a multi-agent system based on robust differential game according to claim 6, characterized in that The distributed robust differential game cost function described in Step S3, the constructed approximate best performance index can be deformed into the following form: In the formula, 8. A collision avoidance strategy method for a multi-agent system based on robust differential game according to claim 6, characterized in that, For the distributed robust differential game cost function described in Step S3, the constraint conditions of the constructed feedback gain matrix can be obtained, and the problem of solving the optimal feedback gain matrix is transformed into a multi-objective optimization problem; The optimal feedback gain matrix constructed is solved by using the non-dominated dominance-based ant colony optimization algorithm, and the corresponding local robust Nash equilibrium solution is obtained; For agent i at time t, the multi-objective optimization function is defined as: wherein, υ c > 0 is a constant; the objective of this multi-objective optimization problem is to find the optimal feedback gain set to minimize the function the function value is minimized.
Citation Information
Patent Citations
Unmanned aerial vehicle cluster cooperative confrontation control method simulating eagle pigeon intelligent game
CN112269396A
Optimal Strategies in Security Games
US20130273514A1