A method for finding drone pursuit and escape game based on Nash equilibrium with specified time

By establishing the second-order dynamic equation and communication diagram of the drone intelligent body, a pursuit and fugitive game algorithm that converges within the specified time is designed, which solves the problem of realizing Nash equilibrium in the drone pursuit and fugitive game, and achieves fast, practical and anti-interference Nash equilibrium convergence.

CN117369505BActive Publication Date: 2025-09-02SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311441521.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-01
Publication Date
2025-09-02
Estimated Expiration
2043-11-01

AI Technical Summary

Technical Problem

The prior art is difficult to achieve Nash equilibrium in the pursuit and escape game of drones within a limited time, and it has limitations on the initial speed and position of the agent, and lacks practicality and anti-interference performance.

Method used

By establishing the second-order dynamic equation of the drone agent, defining the communication graph, optimizing the income function, and designing a pursuit and pursuit game algorithm that converges within the specified time, using the Lyapunov stability theory to verify the convergence, combining the control input assumption and the communication graph assumption, the rapid convergence of Nash equilibrium is achieved.

Benefits of technology

The Nash equilibrium of drone pursuit and escape game is achieved within the specified time, which improves the practicality and anti-interference performance of the algorithm, and is not limited by the initial state of the agent, so that the target state can be quickly reached.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117369505B_ABST
    Figure CN117369505B_ABST
Patent Text Reader

Abstract

This invention discloses a method for finding a PEG (Pursuit-Escape Game) based on a prescribed time Nash equilibrium (PTNE). The method aims to solve the game problem between multiple pursuers and evaders in a chase-escape game. The method first establishes a second-order dynamic equation for the drone's motion in the chase-escape game. It then proposes a communication graph for the drone and a payoff function for the pursuer. The Nash equilibrium definition and lemma conditions for the drone chase-escape game are then proposed. A convergence algorithm is designed to achieve the Nash equilibrium within a prescribed time, and the prerequisites for algorithm convergence are given. Finally, the convergence of the algorithm is proven. The distributed algorithm designed in this invention can implement PEG under PTNE by adaptively adjusting the control scheme parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of control systems and intelligent agent game theory, and mainly relates to the development of a method for finding a preset time Nash equilibrium of a pursuit-escape game (PEG) of an unmanned aerial vehicle intelligent agent with second-order dynamics. Background Art

[0002] The pursuit-escape game consists of two groups of agents: the evaders and the pursuers. The pursuers' goal is to capture the evaders through cooperation, while the evaders strive to evade capture. PEG is a classic non-cooperative game. It has attracted widespread attention due to its wide range of applications, such as smart grids, formation control, and spacecraft rendezvous. Notably, most research on finding Nash equilibria in this game focuses on asymptotic and exponential convergence, in which no agent has an incentive to change its behavior. However, in real-world scenarios, achieving a Nash equilibrium within a finite time is valuable and practical. Compared to asymptotic and exponential convergence, finite-time convergence can speed up the Nash equilibrium search algorithm and improve its robustness to interference. Therefore, in the pursuit-escape game, finding a Nash equilibrium within a specified time is a challenging and meaningful problem. Summary of the Invention

[0003] Purpose of the invention: In view of the shortcomings of the existing technology, the present invention develops a method for realizing the pursuit-escape game within a specified time, which can achieve the convergence time of the pursuit-escape game that is predetermined and user-defined.

[0004] Technical solution: In order to achieve the above-mentioned purpose of the invention, the technical solution adopted by the present invention is:

[0005] A method for solving the drone chase-and-escape game based on finding a Nash equilibrium within a specified timeframe is proposed. First, the second-order dynamic equations for the drone agent's motion are established, followed by a definition of the agent's communication graph. The agent's payoff function is optimized, taking into account its velocity and position. Lemma conditions and assumptions are proposed, and based on these conditions, an algorithm for solving the chase-and-escape game within a specified timeframe is proposed, along with prerequisites for algorithm convergence. Convergence is derived and proven based on the assumptions and the Lyapunov function. The method includes the following steps:

[0006] Step 1: Establish a second-order dynamic equation for the drone's motion in a chase-and-escape game. This equation describes the dynamic relationship between the various elements in the system.

[0007] Step 2: Propose a communication graph of drones to determine the communication relationship between each drone;

[0008] Step 3: Based on the second-order dynamic equation of the UAV motion established in Step 1 and the UAV communication graph proposed in Step 2, design the payoff function of the chaser, and consider the position and velocity of the UAV in the payoff function;

[0009] Step 4: Based on the pursuer's payoff function designed in step 3, the Nash equilibrium definition and lemma conditions, control input assumptions and communication graph assumptions of the drone pursuit game are proposed;

[0010] Step 5: Based on the second-order dynamic equations of the drone motion established in Step 1 and the Nash equilibrium definition of the drone pursuit-escape game proposed in Step 4, design a convergence algorithm that reaches the Nash equilibrium of the pursuit-escape game within the specified time;

[0011] Step 6: Based on the control input assumptions and communication graph assumptions proposed in step 4, give the prerequisites for the convergence of the Nash equilibrium algorithm of the pursuit-escape game proposed in step 5;

[0012] Step 7: Based on the control input assumptions and communication graph assumptions proposed in Step 4 and the prerequisites given in Step 6, prove the convergence of the convergence algorithm proposed in Step 5.

[0013] Furthermore, in step 1, a second-order dynamic equation model of the UAV motion is established, specifically:

[0014]

[0015] For drone collection in chase-evasion games It's a collection of chasers, is the set of evaders, where is the control input, position and velocity; if there are m evaders, it can be decomposed into a set problem with m pursuers and one evader. For i∈{1,2,…,n}, define u i , x i , v i is the control input, position and velocity of the ith pursuer, n is the number of pursuers, and u n+1 , x n+1 , v n+1 denote the control input, position, and velocity of the evader, respectively.

[0016] Furthermore, the proposed drone communication graph and matrix operation in step 2 specifically include the following steps:

[0017] Step 2-1. Define the drone communication graph as an undirected graph:

[0018]

[0019] in, represents the point set, ε represents the edge set, Represents the UAV communication diagram;

[0020] Step 2-2. Propose an adjacency matrix for a connected undirected graph:

[0021]

[0022] in, is the adjacency matrix, a kj Is a parameter in the adjacency matrix, indicating the adjacency matrix The kth row and jth column of , if (k,j)∈ε, then a kj =1; otherwise a kj =0;

[0023] Step 2-3. Propose the Laplacian matrix:

[0024]

[0025] in, is the Laplace matrix, l kj is the parameter of the Laplace matrix, indicating the Laplace matrix The kth row and jth column of If k≠j, then l kj =-a kj .

[0026] Furthermore, in step 3, the pursuer's reward function is optimized based on the evader's strategy, and the position and speed of the agent are considered in the reward function, specifically:

[0027]

[0028] in, J i is the payoff function of the ith pursuer, e i is the set of positions and velocities of the ith pursuer, called the strategy set, x i is the position of the i-th pursuer, v i is the speed of the ith pursuer, e -i =(x -i ,v -i ) is the set of positions and velocities of all drones except the i-th pursuer, where x -i =(x1,…,x i-1 ,x i+1 ,…,x n+1 ) is the set of positions of all drones except the i-th pursuer, v -i =(v1,…,v i-1 ,v i+1,…,v n+1 ) is the set of velocities of all drones except the i-th pursuer.

[0029] Furthermore, the Nash equilibrium definition, lemma conditions and assumptions proposed in step 4 include:

[0030] Step 4-1. Propose a Nash equilibrium definition, specifically:

[0031]

[0032] For a game Ω i is the domain of the strategy set of the i-th pursuer in the game, e i is the set of strategies of the ith pursuer, is the optimal strategy set of the i-th pursuer, is the optimal strategy set of all drones except the i-th pursuer. This strategy set It's Game G pe Nash equilibrium;

[0033] Step 4-2. Propose the profit function lemma, specifically:

[0034] For each e i ∈Ω i , the profit function J i (e i ,e -i )=J i (x i ,x -i ,v i ,v -i ) is C in its domain 2 , and is strictly convex and radially unbounded for every object where C m represents a set of m-times continuously differentiable functions;

[0035] Step 4-3. Show that the Nash equilibrium exists and satisfies the following conditions:

[0036] For Nash equilibrium, no subject has the motivation to unilaterally deviate from their behavior. After inspection, the resulting payoff function satisfies the conditions in the payoff function lemma, which shows that Nash equilibrium is the only one that exists,

[0037] Furthermore, the Nash equilibrium satisfies:

[0038]

[0039] in, This represents the gradient of the different benefits obtained by the i-th pursuer according to his own behavior;

[0040] Step 4-4. Propose the communication graph hypothesis, that is, the undirected communication graph of n pursuers and 1 evader is connected, specifically:

[0041]

[0042] diag(a 1(n+1) ,a 2(n+1) ,…,a n(n+1) ) indicates that the element a 1(n+1) ,a 2(n+1) ,...,a n(n+1) The diagonal matrix composed of The Laplace matrix of n chasers and 1 evader is in, is the connection matrix between pursuers and evaders, The Laplace matrix of n chasers is represents a real vector space of q×p dimensions;

[0043] Step 4-5. Propose the control input hypothesis and high gain function of the evader, specifically:

[0044] The control input hypothesis states that the derivative of the evader's control input is bounded, that is, there exists a positive constant satisfy High gain function θ ι (t) is expressed as:

[0045]

[0046]

[0047] Where ι is a positive integer, c is the static gain of the high gain function, T s is the artificially specified convergence time, ι=1,2, c>1, T s >0;

[0048] Step 4-6. Propose the Laplace matrix lemma, specifically:

[0049] Based on the communication graph assumption, we get Is positive and symmetric, we can find a vector δ=(δ1,δ2,…,δ n ) T satisfy The superscript T indicates the transpose of the matrix;

[0050] in, i=1,2,…,n;1 nis a unit vector of length n

[0051] Let δ min =min(δ1,δ2,…,δ n ), get:

[0052]

[0053] in, Representation matrix The transpose of , the adjuvant matrices Γ and Δ are positive definite.

[0054] Furthermore, in step 5, a controller is designed to achieve the convergence algorithm of the pursuit-escape game within a specified time, specifically including the following steps:

[0055]

[0056]

[0057] Where k = (k1θ2, k2), k1, k2 are the static gains of the controller; in addition, y i is the control input u from the pursuer i to the evader n+1 Estimates, It is y i b1, d1 are the static gains in the control input estimator, μ is a parameter selected based on the first-order derivative of the control input, and a is is the adjacency matrix The element in the row and column s corresponding to chaser i, and y n+1 =u n+1 , and k1>0, k2>0, b1≥0, d1>0, μ>0; e i is the set of positions and velocities of the ith pursuer, e -i is the set of positions and velocities of all drones except the i-th pursuer. sign(x) represents the sign function. When x>0, sign(x)=1; when x=0, sign(x)=0; when x<0, sign(x)=-1.

[0058] Furthermore, step 6 gives the prerequisites for algorithm convergence based on the control input assumption and the communication graph assumption:

[0059] For a given time T p =2T s With any initial position x(0) and velocity v(0), the PTNE search algorithm ensures that all pursuers are within T pThe convergence to PTNE is always achieved if the auxiliary parameters β>0, γ>0, ρ>0, b2≥0, d2>0 exist and the following inequalities are satisfied:

[0060]

[0061] β 2 -γ(k1γ+k2β)λ min (Γ)δ min <0

[0062] κ1<0

[0063] κ2<0

[0064]

[0065]

[0066] in, λ min (Γ) represents the minimum eigenvalue of the auxiliary matrix Γ, δ min Representative vector δ=(δ1,δ2,…,δ n ) T The minimum value in .

[0067] Furthermore, in step 7, based on the control input assumptions and communication diagram assumptions and preconditions, the effectiveness of the control algorithm is proved, specifically:

[0068] Step 7-1. Prove that the pursuer's estimate of the evader's control input is s The truth value of the control input of the evader is reached within:

[0069] The error between the estimated value and the true value of the control input of the evader by the pursuer i is defined as

[0070]

[0071] According to the given control algorithm, we can get:

[0072]

[0073] where l is is the Laplace matrix The element in the row and column s corresponding to chaser i;

[0074] The estimated vector form of the control input is defined as The first-order derivative of the estimated vector form of the control input can be obtained as

[0075]

[0076] in Indicated by The stacked vectors, represents the Kronecker product, is the first-order derivative of the evader's control input, and the Lyapunov function is designed. Get the time derivative of V1

[0077]

[0078] Based on the control input assumption and have to:

[0079]

[0080] in

[0081] Based on the Lyapunov function lemma, we can get At time T s Converges to the origin, and when t>T s When y i ≡u n+1 ;

[0082] Step 7-2. Get u n+1 After accurate estimation, at T p =2T s Always establish Nash equilibrium to achieve:

[0083] x i * and v i * represents the Nash equilibrium of PEG, then u i * v i * The derivative of Combined profit function J i (e i ,e -i ) and the gradient of the reward function get:

[0084]

[0085]

[0086] in,

[0087] definition and y=col(y1,y2,…,y n) can be obtained in vector form:

[0088]

[0089]

[0090] Design of Lyapunov function in,

[0091] In the time interval t∈[T s ,2T s ], we can get

[0092]

[0093] When the auxiliary parameters b2≥0, d2>0, the objective function is defined as You can get:

[0094]

[0095] According to the algorithm convergence prerequisite, the auxiliary parameters κ1<0 and and high gain functions We can get:

[0096]

[0097] Based on the Lyapunov function, we can get:

[0098]

[0099]

[0100] in, exp(-b2t) represents the (-b2t) power of the natural base e. Represents a vector The square of the 2-norm;

[0101] Based on V2, we can get:

[0102]

[0103] Therefore, when t→2T s ,get and

[0104] Compared with the prior art, the present invention has the following beneficial effects:

[0105] 1. This invention considers the speed and position information of the intelligent agent in the profit function, and the designed second-order system is closer to the real physical world. At the same time, this also makes the pursuit-escape game problem under PTNE more challenging.

[0106] 2. Compared to existing Nash equilibrium-seeking algorithms based on gradual or fixed-time convergence, this method imposes no restrictions on the agent's initial velocity or position. The convergence time can be pre-programmed to achieve convergence to the target state, improving the practicality of the pursuit-escape game algorithm. Leveraging Lyapunov stability theory, this method derives sufficient conditions for convergence within a specified timeframe. BRIEF DESCRIPTION OF THE DRAWINGS

[0107] Figure 1 This is a schematic diagram of the design steps of the pursuit-escape game algorithm within a specified time of the present invention;

[0108] Figure 2 It is a schematic diagram of the network communication topology simulated by the present invention;

[0109] Figure 3 is a graph showing the position trajectory changes of the agent in the first set of simulations of the present invention;

[0110] Figure 4 is a graph showing the velocity trajectory changes of the agent in the first set of simulations of the present invention;

[0111] Figure 5 is a graph showing the estimated trajectory changes of the agent in the first set of simulations of the present invention;

[0112] Figure 6 is a graph showing the position trajectory changes of the agent in the second set of simulations of the present invention;

[0113] Figure 7 is a graph showing the velocity trajectory changes of the agent in the second set of simulations of the present invention;

[0114] Figure 8 is a graph showing the estimated trajectory changes of the agent in the second set of simulations of the present invention; DETAILED DESCRIPTION

[0115] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments.

[0116] like Figure 1 The pursuit-escape game algorithm is implemented within the specified time shown, including the following steps:

[0117] Step 1: Establish a second-order dynamic equation for the drone's motion in a chase-and-escape game. This equation describes the dynamic relationship between the various elements in the system.

[0118] Step 2: Propose a communication graph of drones to determine the communication relationship between each drone;

[0119] Step 3: Based on the second-order dynamic equation of the UAV motion established in Step 1 and the UAV communication graph proposed in Step 2, design the payoff function of the chaser, and consider the position and velocity of the UAV in the payoff function;

[0120] Step 4: Based on the pursuer's payoff function designed in step 3, the Nash equilibrium definition and lemma conditions, control input assumptions and communication graph assumptions of the drone pursuit game are proposed;

[0121] Step 5: Based on the second-order dynamic equations of the drone motion established in Step 1 and the Nash equilibrium definition of the drone pursuit-escape game proposed in Step 4, design a convergence algorithm that reaches the Nash equilibrium of the pursuit-escape game within the specified time;

[0122] Step 6: Based on the control input assumptions and communication graph assumptions proposed in step 4, give the prerequisites for the convergence of the Nash equilibrium algorithm of the pursuit-escape game proposed in step 5;

[0123] Step 7: Based on the control input assumptions and communication graph assumptions proposed in Step 4 and the prerequisites given in Step 6, prove the convergence of the convergence algorithm proposed in Step 5.

[0124] The following is a specific implementation case using a simulation method, and the effectiveness of the present invention is verified by writing a Matlab program.

[0125] Step A: Consider the second-order dynamic equation of the UAV intelligent body motion as follows:

[0126]

[0127] Step B: The pursuit-escape game problem consists of 1 escaper and 3 pursuers. The network communication topology diagram is as follows Figure 2 As shown, the Laplacian matrix containing 3 chasers and 1 evader is The pursuit-escape matrix is Laplacian matrix between 3 chasers Diagonal matrix Based on this, the auxiliary matrix and λ is obtained by calculation min (Γ)=0.3250,δ min =1.5.

[0128] Step C: Based on the previous discussion, when the position and velocity of the pursuer are the same as those of the evader, that is, x1 = x2 = x3 = x4 and v1 = v2 = v3 = v4, Nash equilibrium is reached. The simulation is performed within the time interval t∈(0s,4s), and the convergence time is set to T p =2Ts =2s.

[0129] Step D: Set the initial state and control input parameters of each UAV agent. Set the evader's control input u4(t) = 2sin(t). Select parameters β = 1, γ = 1, ρ = 2, b2 = 0, d2 = 0.1, and set the control algorithm parameters b1 = 0, d1 = 0, μ = 2, c = 3, k1 = 5.3, k2 = 2.5. Set the sampling period to 0.01s. In the first set of simulation experiments, the initial position, initial velocity, and estimated values ​​of the evader's control input are x(0) = (1, 2, 3, 4) T ,v(0)=(1,2,3,4) T , y(0)=(-1,2,3,0) T .

[0130] Step E: Use Matlab to simulate according to the provided parameters. The simulation results are as follows: Figure 3-5 , we can get the position trajectory, velocity trajectory and estimated trajectory of all agents. Figure 3 It can be seen that the pursuer can catch the evader within the specified time. Figure 4 It can be seen that the speed of all pursuers is consistent with the speed of the evader. Figure 5 It can be seen that within 1s, each agent's estimate of u4 reaches the true value.

[0131] Step F: Change the initial state of each UAV in the simulation experiment to verify the practicality of the control algorithm. In the second set of simulation experiments, the initial position, initial velocity and estimated values ​​of the evader control input of the UAV are x′(0)=(5,10,15,20) T , v′(0)=(5,10,15,20) T , y′(0)=(1,2,3,0) T , the simulation results are as follows Figure 6-8 This result proves that the pursuers all reach the Nash equilibrium within the specified time, thus verifying the effectiveness of the algorithm.

[0132] In summary, this paper proposes a method for finding a Nash equilibrium within a specified time for a drone chase-and-escape game. Numerical verification shows that the PTNE search algorithm converges within a specified time, which is artificially set and independent of the system's initial state.

[0133] The above are only preferred embodiments of the present invention. It should be noted that any equivalent replacements made within the principles of the present invention should be included in the scope of protection of the present invention. Contents not elaborated in detail in the present invention belong to the existing technologies known to those skilled in the art.

Claims

1. A drone pursuit and escape game method based on finding a Nash equilibrium at a specified time, characterized in that: The method comprises the following steps: Step 1: Establish a second-order dynamic equation for the drone's motion in a chase-and-escape game. This equation describes the dynamic relationship between the various elements in the system. Step 2: Propose a communication graph of drones to determine the communication relationship between each drone; Step 3: Based on the second-order dynamic equation of the UAV motion established in Step 1 and the UAV communication graph proposed in Step 2, design the payoff function of the chaser, and consider the position and velocity of the UAV in the payoff function; Step 4: Based on the pursuer's payoff function designed in step 3, the Nash equilibrium definition and lemma conditions, control input assumptions and communication graph assumptions of the drone pursuit game are proposed; Step 5: Based on the second-order dynamic equations of the drone motion established in Step 1 and the Nash equilibrium definition of the drone pursuit-escape game proposed in Step 4, design a convergence algorithm that reaches the Nash equilibrium of the pursuit-escape game within the specified time; Step 6: Based on the control input assumptions and communication graph assumptions proposed in step 4, give the prerequisites for the convergence of the Nash equilibrium algorithm of the pursuit-escape game proposed in step 5; Step 7: Based on the control input assumptions and communication graph assumptions proposed in Step 4 and the preconditions given in Step 6, prove the convergence of the convergence algorithm proposed in Step 5; In step 1, the second-order dynamic equation model of the UAV motion is established, specifically: For the drone set N in the chase-escape game, p ,N e }, N p is the chaser set, N e is a set of evaders, where u(t)={u i ,i∈N},x(t)={x i ,i∈N},v(t)={v i ,i∈N} is the control input, position and velocity; if there are m evaders, it is decomposed into m sets of problems with multiple pursuers and one evader. For i∈{1,2,…,n}, define u i , x i , v i is the control input, position and velocity of the ith pursuer, n is the number of pursuers, and u n+1 , x n+1 , v n+1 denote the control input, position, and velocity of the evader, respectively; The communication diagram and matrix operations of the drone described in step 2 specifically include the following steps: Step 2-1. Define the drone communication graph as an undirected graph: in, represents a set of points, represents the edge set, Represents the UAV communication diagram; Step 2-2. Propose an adjacency matrix for a connected undirected graph: A=(a kj ) (n+1) ×(n+1) Among them, A is the adjacency matrix, a kj Is a parameter in the adjacency matrix, representing the kth row and jth column of the adjacency matrix A. If (k, j)∈ , then a kj =1; otherwise a kj =0; Step 2-3. Propose the Laplacian matrix: L PE =(l kj ) (n+1 )×(n+1) Among them, L PE is the Laplace matrix, l kj is the parameter of the Laplace matrix, indicating the Laplace matrix L PE The kth row and jth column of If k≠j, then l kj =-a kj , Step 3 optimizes the pursuer’s payoff function based on the evader’s strategy, and considers the agent’s position and velocity in the payoff function. Specifically: Where i∈N p , J i is the payoff function of the ith pursuer, e i is the set of positions and velocities of the ith pursuer, called the strategy set, x i is the position of the i-th pursuer, v i is the speed of the ith pursuer, e -i =(x -i ,v -i ) is the set of positions and velocities of all drones except the i-th pursuer, where x -i =(x1,…,x i-1 ,x i+1 ,…,x n+1 ) is the set of positions of all drones except the i-th pursuer, v -i =(v1,…,v i-1 ,v i+1 ,…,v n+1 ) is the set of velocities of all drones except the i-th pursuer.

2. The drone pursuit and escape game method based on finding a Nash equilibrium at a specified time according to claim 1 is characterized in that: Step 4 proposes the definition of Nash equilibrium and the lemma conditions and assumptions, including; Step 4-1. Propose a Nash equilibrium definition, specifically: For a game G pe (N,J i ,Ω i ),Ω i is the domain of the strategy set of the i-th pursuer in the game, e i is the set of strategies of the ith pursuer, is the optimal strategy set of the i-th pursuer, is the optimal strategy set of all drones except the i-th pursuer. This strategy set It's Game G pe Nash equilibrium; Step 4-2. Propose the profit function lemma, specifically: For each i∈N p ,e i ∈Ω i , the profit function J i (e i ,e -i )=J i (x i ,x -i ,v i ,v -i ) is C in its domain 2 , and is strictly convex and radially unbounded for every object where C m represents a set of m-times continuously differentiable functions; Step 4-3. Show that the Nash equilibrium exists and satisfies the following conditions: For Nash equilibrium, no subject has the motivation to unilaterally deviate from their behavior. After inspection, the resulting payoff function satisfies the conditions in the payoff function lemma, which shows that Nash equilibrium is the only one that exists, Furthermore, the Nash equilibrium satisfies: in, This represents the gradient of the different benefits obtained by the i-th pursuer according to his own behavior; Step 4-4. Propose the communication graph hypothesis, that is, the undirected communication graph of n pursuers and 1 evader is connected, specifically: diag(a 1(n+1) ,a 2(n+1) ,…,a n(n+1) ) indicates that the element a 1(n+1) ,a 2(n+1) ,...,a n(n+1) The diagonal matrix composed of n chasers and 1 escaper is denoted as M. The Laplace matrix of n chasers and 1 escaper is Among them, L PE is the connection matrix between the pursuer and the evader, L pe ∈R n×n ,L 1 ∈R n×1 ,L 2 ∈R 1×n , the Laplace matrix of n chasers is L p , L p ∈R n×n , R q×p represents a real vector space of q×p dimensions; Step 4-5. Propose the control input hypothesis and high gain function of the evader, specifically: The control input hypothesis states that the derivative of the evader's control input is bounded, that is, there exists a positive constant satisfy High gain function θ ι (t) is expressed as: Where ι is a positive integer, c is the static gain of the high gain function, T s is the artificially specified convergence time, ι=1,2, c>1, T s >0; Step 4-6. Propose the Laplace matrix lemma, specifically: Based on the communication graph assumption, we get L pe Is positive definite and symmetric, find a vector δ=(δ1,δ2,…,δ n ) T Meet L pe δ=1 n , the superscript T indicates the transpose of the matrix; in, 1 n is a unit vector of length n Let δ min = min(δ1, δ2, …, δ n ), obtain: Among them, (L pe ) T Represents the matrix L pe The transpose of , the adjuvant matrices Γ and Δ are positive definite.

3. The drone pursuit and escape game method based on finding a specified time Nash equilibrium according to claim 1 is characterized in that: In step 5, a controller is designed to achieve the convergence of the pursuit-escape game algorithm within a specified time, specifically including the following steps: Where k = (k1θ2, k2), k1, k2 are the static gains of the controller; in addition, y i is the control input u from the pursuer i to the evader n+1 Estimates, It is y i b1, d1 are the static gains in the control input estimator, μ is a parameter selected based on the first-order derivative of the control input, and a is is the element in the row and column s corresponding to pursuer i in the adjacency matrix A, and y n+1 =u n+1 , and k1>0, k2>0, b1≥0, d1>0, μ>0; e i is the set of positions and velocities of the ith pursuer, e -i is the set of positions and velocities of all drones except the i-th pursuer. sign(x) represents the sign function. When x>0, sign(x)=1; when x=0, sign(x)=0; when x<0, sign(x)=-1.

4. The drone pursuit and escape game method based on finding a Nash equilibrium within a specified time according to claim 3 is characterized in that: Step 6 gives the prerequisites for algorithm convergence based on the control input assumption and communication graph assumption: For a given time T p =2T s With any initial position x(0) and velocity v(0), the PTNE search algorithm ensures that all pursuers are within T p The convergence to PTNE is always achieved if the auxiliary parameters β>0, γ>0, ρ>0, b2≥0, d2>0 exist and the following inequalities are satisfied: β2-γ(k1γ+k2β)λ min (C)d min <0 κ1<0 κ2<0 in, λ min (Γ) represents the minimum eigenvalue of the auxiliary matrix Γ, δ min Representative vector δ=(δ1,δ2,…,δ n ) T The minimum value in .

5. The drone pursuit and escape game method based on finding a Nash equilibrium within a specified time according to claim 1, characterized in that: In step 7, based on the control input assumptions, communication diagram assumptions, and preconditions, the effectiveness of the control algorithm is proved. Specifically: Step 7-1. Prove that the pursuer's estimate of the evader's control input is s The truth value of the control input of the evader is reached within: The error between the estimated value and the true value of the control input of the evader by the pursuer i is defined as According to the given control algorithm, we can obtain: where l is is the Laplace matrix L pe The element in the row and column s corresponding to chaser i; The estimated vector form of the control input is defined as The first-order derivative of the estimated vector form of the control input is in Indicated by The stacked vectors, represents the Kronecker product, is the first-order derivative of the evader's control input, and the Lyapunov function is designed. Get the time derivative of V1 Based on the control input assumption and L p 1 n =0 n have to: in Based on the Lyapunov function lemma, we get At time T s Converges to the origin, and when t>T s When y i ≡u n+1 ; Step 7-2. Get u n+1 After accurate estimation, at T p =2T s Always establish Nash equilibrium to achieve: x i * and v i * represents the Nash equilibrium of PEG, then u i * v i * The derivative of Combined profit function J i (e i ,e -i ) and the gradient of the reward function get: in, definition and y=col(y1,y2,…,y n ) in vector form: Design of Lyapunov function in, In the time interval t∈[T s ,2T s ], we get When the auxiliary parameters b2≥0, d2>0, the objective function is defined as get: According to the algorithm convergence prerequisite, the auxiliary parameters κ1<0 and and high gain functions have to: Based on the Lyapunov function, we get: in, exp(-b2t) represents the (-b2t) power of the natural base e. Represents a vector The square of the 2-norm; Based on V2: Therefore, when t→2T s ,get and

Citation Information

Patent Citations

  • Nash equilibrium specified time search method for intra-group decision consistent multi-group game

    CN114488802A

  • Non-zero sum game unmanned aerial vehicle formation control method based on reinforcement learning

    CN115877871A