Large-scale area pursuit game method with collision avoidance guarantee
By designing and implementing 1-to-1 and 2-to-2 strategies in large-scale regional pursuit and fugitive game, combining the H-J-I equation and the maximum matching algorithm, the problem of invalid losses caused by internal members in the pursuit and fugitive game is solved, and the effect of efficient interception and collision avoidance is achieved.
Patent Information
- Application Number
- CN202510035362.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-13
AI Technical Summary
In the large-scale pursuit and escape game of autonomous systems, the defense party loses ineffectively due to collisions between internal members when performing interception tasks, reducing the overall defensive effect.
The H-J accessibility principle is used to design the evasion set and target set of 1-to-1 and 2-to-2 strategies. The game value function is solved in discrete state space through the H-J-I equation, a defensive party's collision avoidance alliance is constructed, and the maximum matching calculation is performed in the bilateral graph to obtain the optimal control amount to achieve interception and collision avoidance.
It significantly reduces the collision loss rate of defensive members in mission execution, improves the overall defensive effect and interception success rate, and ensures safety among defensive members.
Smart Images

Figure CN119990180A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of control and information technology, and in particular to a large-scale regional pursuit and escape game method with collision avoidance guarantee. Background Art
[0002] The regional pursuit game is a widely studied game scenario, and has a wide range of applications in areas such as robot obstacle avoidance and strategy design for non-human players in video games. With the automation of manufacturing technology and the integration and intelligence of robot systems, a large number of unmanned autonomous entities have been put into pursuit game scenarios in order to use their numerical advantage to suppress opponents. This requires the pursuit game strategy to evolve in the direction of expanding the scale of clusters and enhancing alliance collaboration.
[0003] The regional pursuit game is a bilateral game. The two parties are the attacker and the defender. All members of the attacker and the defender move in a bounded closed convex region Ω. The boundary of the target area, the inside of the boundary is the target area When the attacking team members enter a certain range of the defending team members, they will be captured by the defending team and lost. If the distance between the defending team members is less than the safe distance, they will collide and lose. goal Inward movement, the goal is to reach the boundary The defenders move in Ω, and the goal is to prevent as many attackers as possible from entering the target area Ω by blocking the attackers and capturing them. goal .
[0004] In the pursuit-and-escape game of a large-scale autonomous system, as the number of defenders and attackers increases, the probability of collision between autonomous agents increases significantly, which places higher demands on the defenders. The defenders not only need to have complex maneuvering strategies to effectively intercept the attackers, but also need to avoid collision losses between internal members when executing the strategies. In addition, the diverse strategy combinations of the defenders make task allocation an NP-hard problem with multiple constraints.
[0005] At present, for large-scale pursuit and escape game scenarios, the strategy design of the defender mainly focuses on optimizing the interception effect, but the consideration of collision avoidance constraints is relatively insufficient. This leads to invalid losses caused by collisions between internal members when the defender performs interception tasks in complex scenarios, thereby reducing the overall defense effect. Therefore, in order to solve the problem of lack of collision avoidance constraints in the defender's differential game strategy, it is necessary to integrate the collision avoidance constraints into the defender's differential game strategy while ensuring the optimality of the strategy, and combine it with an efficient task allocation algorithm to achieve effective interception of the attacker and minimize the defender's losses. Summary of the invention
[0006] The present application aims to solve one of the technical problems in the related art at least to some extent.
[0007] To this end, the first purpose of this application is to propose a large-scale regional pursuit and escape game method with collision avoidance guarantee.
[0008] The second purpose of this application is to propose a large-scale regional pursuit and escape game device with collision avoidance guarantee.
[0009] The third objective of the present application is to provide an electronic device.
[0010] A fourth objective of the present application is to provide a computer-readable storage medium.
[0011] A fifth object of the present application is to provide a computer program product.
[0012] To achieve the above-mentioned purpose, the first embodiment of the present application proposes a large-scale regional pursuit and escape game method with collision avoidance guarantee, comprising:
[0013] According to the HJ reachability principle, the avoidance set and target set of the 1-to-1 strategy and the 2-to-2 strategy are designed respectively;
[0014] Substitute the avoidance set and target set of the designed 1-on-1 strategy and 2-on-2 strategy into the HJI equation, solve the value of the game value function in the discrete state space and store it;
[0015] According to the boundary fence information of the 2-on-2 strategy and the distance between the defenders, a collision avoidance alliance of the defenders is constructed, and the 2-on-2 matching result is determined in the collision avoidance alliance;
[0016] A bilateral graph is constructed according to the 2-to-2 matching result and the boundary fence information of the 1-to-1 strategy, and a maximum matching calculation is performed in the bilateral graph to obtain a 1-to-1 matching result;
[0017] The calculated HJI equation value functions of the 2-on-2 strategy and the 1-on-1 strategy are used to obtain the optimal control quantity for the defenders who implement the 2-on-2 strategy and the 1-on-1 strategy respectively.
[0018] Optionally, assume that in a regional pursuit game of multiple autonomous agents, there are M members of the attacker and N members of the defender, forming the attacker set and the defender set respectively. and And both members adopt the first-order integral dynamic model for control, then the motion control model is:
[0019]
[0020] In the formula, and They represent the coordinates of the attacking member i and the defending member j in the two-dimensional Euclidean space, and They represent the positions of the attacking member i and the defending member j at the initial moment, u i and v j They represent the normalized control amount of the attacking member i and the defending member j, respectively, and v A and v D are the maximum speeds of the attacker and defender respectively.
[0021] Optionally, for the regional pursuit game system, the joint state of the attacker and the defender satisfies the following state equation:
[0022]
[0023] The solution trajectory is:
[0024]
[0025] Where x is the joint state of the attacker and the defender, which represents the coordinates of all members of the attacker and the defender in the two-dimensional Euclidean space; and are the joint control quantities of the attacker and the defender respectively; f(x,u,v) is the state transfer function of the system; Indicates that from the initial state x 0 The state at time τ after inputting the control variables u(·) and v(·).
[0026] Optionally, based on the HJ reachability principle, design the avoidance set and target set of the 2-to-2 strategy, including:
[0027] For the 8-dimensional joint state space of 2 defenders and 2 attackers, design the avoidance set A and target set R of the 2-on-2 strategy;
[0028] The avoidance set A is used to describe the set of state subspaces that the state vector should not be in during the game, and its definition includes the following two situations: when a member of the attacking party is intercepted by a member of the defending party, another member of the defending party can ensure to intercept another member of the attacking party; a collision occurs between two members of the attacking party;
[0029] The target set R is used to describe the set of state subspaces that the state vector should be in when the game ends, and its definition includes the following two situations: any member of the attacking party successfully enters the target area and is not intercepted by any member of the defending party; a collision occurs between two members of the defending party.
[0030] Optionally, the step of bringing the avoidance set and target set of the designed 1-to-1 strategy and 2-to-2 strategy into the HJI equation, solving and storing the value of the game value function in the discrete state space, comprises:
[0031] The avoidance set and target set of the 1-to-1 and 2-to-2 strategies are brought into the HJI equation, where the pursuit problem is expressed as a minimization-maximization planning problem, which is expressed as follows:
[0032]
[0033] Where V(x,t) represents the value function, the optimal strategy result at state x and time t; γ is the strategy function; From the initial state x 0 At the beginning, the state trajectory at time τ when the defender adopts strategy γ and the attacker adopts strategy v; the level functions h(x) and l(x) are defined according to the designed avoidance set and target set, and their specific forms are:
[0034]
[0035]
[0036] In the formula, d(x,A) and d(x,R) represent the minimum distances from state x to the avoidance set A and the target set R respectively;
[0037] Through the numerical discretization method, the HJI equation is solved in the joint state space of 1-to-1 and 2-to-2 strategies, and the value function results obtained by numerical calculation are stored in each node of the discrete state space for subsequent strategy execution and optimal control quantity calculation.
[0038] Optionally, the step of constructing a collision avoidance alliance of the defender based on the boundary fence information of the 2-to-2 strategy and the distance between the defender members, and determining the 2-to-2 matching result in the collision avoidance alliance includes:
[0039] Determine whether the defending team members have matched with the attacking team members in the last matching process. If the defending team members have not matched with the attacking team members, the 2-to-2 strategy will not be adopted;
[0040] For defenders who can adopt a 2-on-2 strategy, a collision avoidance alliance is established based on the collision risk between defenders. If there is a collision risk between defenders and two defenders can cooperate to intercept two attackers under the 2-on-2 strategy, they will be added to the collision avoidance alliance.
[0041] For the two defenders in the collision avoidance alliance, calculate the fence based on their positions On the boundary fence, both the chasing and fleeing parties have strategies to prevent the opponent from winning, where the boundary fence is defined as:
[0042]
[0043] Where V(x,t) is the value function, which represents the optimal strategy result of the defender being able to intercept the attacker at state x and time t;
[0044] If the last matched offensive members of two defenders in the collision avoidance alliance are both outside the fence, the corresponding two defenders and two offensive members form a 2-to-2 match result.
[0045] Optionally, constructing a bilateral graph based on the 2-to-2 matching result and the boundary fence information of the 1-to-1 strategy, performing maximum matching calculation in the bilateral graph, and obtaining the 1-to-1 matching result includes:
[0046] Based on the 2-to-2 matching results and the boundary fence of the 1-to-1 strategy, a bilateral graph G is constructed, which is expressed as:
[0047] G=(U,V,E)
[0048] In the formula, U is the set of defender members, which indicates the defender members who are not involved in the 2-to-2 matching; V is the set of attacker members, which indicates the attacker members who are not matched by the 2-to-2 strategy; E is the set of edges in the bilateral graph, which indicates the matching relationship that can be formed between defender members and attacker members;
[0049] In the fence In the state region outside the fence, determine whether there is a defender whose strategy can intercept the attacker, and construct all the edges from the defender to the attacker outside the fence in the bilateral graph;
[0050] Run the bilateral graph maximum matching algorithm to find the maximum matching result M in the bilateral graph G, which makes each defender match with at most one attacker, and maximizes the number of matches.
[0051] According to the maximum matching result M, the offensive target corresponding to each defensive member is determined to obtain a 1-to-1 matching result.
[0052] Optionally, the HJI equation value functions of the 2-to-2 strategy and the 1-to-1 strategy obtained by calculation are used to obtain the optimal control amount for the defenders implementing the 2-to-2 strategy and the 1-to-1 strategy, respectively, including:
[0053] Based on the HJI equation value function V2(x,t) of the 2-to-2 strategy, calculate the partial derivative of the value function with respect to the state x Solve the optimal control amount of the defender under the 2-on-2 strategy, the expression is:
[0054]
[0055] In the formula, u* (x, t) represents the optimal control amount of the defender under the 2-on-2 strategy;
[0056] Based on the HJI equation value function V1(x,t) of the 1-to-1 strategy, calculate the partial derivative of the value function with respect to the state x Solve the optimal control amount of the defender under the 1-to-1 strategy, the expression is:
[0057]
[0058] In the formula, v * (x, t) represents the optimal control amount of the defender under the 1-to-1 strategy.
[0059] To achieve the above-mentioned purpose, the second embodiment of the present application proposes a large-scale regional pursuit and escape game device with collision avoidance guarantee, comprising:
[0060] A design module, used to design the avoidance set and target set of the 1-to-1 strategy and the 2-to-2 strategy respectively according to the HJ reachability principle;
[0061] A solution module is used to bring the avoidance set and target set of the designed 1-on-1 strategy and 2-on-2 strategy into the HJI equation, solve the value of the game value function in the discrete state space and store it;
[0062] The first matching module is used to build a collision avoidance alliance of the defender according to the boundary fence information of the 2-to-2 strategy and the distance between the defender members, and determine the 2-to-2 matching result in the collision avoidance alliance;
[0063] A second matching module is used to construct a bilateral graph according to the 2-to-2 matching result and the boundary fence information of the 1-to-1 strategy, and perform maximum matching calculation in the bilateral graph to obtain a 1-to-1 matching result;
[0064] The obtaining module is used to adopt the calculated HJI equation value functions of the 2-to-2 strategy and the 1-to-1 strategy to respectively obtain the optimal control quantity for the defender members implementing the 2-to-2 strategy and the 1-to-1 strategy.
[0065] To achieve the above-mentioned purpose, the third aspect of the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0066] The memory stores computer-executable instructions;
[0067] The processor executes the computer-executable instructions stored in the memory to implement the method as described in any one of the first aspects.
[0068] To achieve the above-mentioned purpose, the fourth aspect embodiment of the present application proposes a computer-readable storage medium, in which computer-readable storage medium is stored computer execution instructions, and when the computer execution instructions are executed by a processor, they are used to implement the method as described in any one of the first aspects.
[0069] To achieve the above-mentioned purpose, the fifth aspect of the present application proposes a computer program product, which implements any method in the first aspect when executed by a processor.
[0070] This application introduces a 2-to-2 differential game strategy and designs a corresponding matching algorithm to achieve efficient interception of the attacker in a large-scale regional pursuit and escape game problem, and solves the collision avoidance problem between defenders, which has the following beneficial technical effects:
[0071] First, the concept of safety guarantee is successfully introduced into the pursuit-escape game problem, which significantly reduces the collision loss rate of the defenders in the task execution. Compared with the autonomous control method based on deep learning, the local differential game method of this scheme achieves the optimality of control under the constraints of the state equation, while taking into account the collision avoidance requirements in practical applications.
[0072] Secondly, the optimal local game strategy based on HJ reachability provides collision avoidance constraints for the defending agent, ensuring the safety of collisions during the game. On this basis, by designing a constrained bilateral graph maximum matching method, the local differential game strategy is successfully applied to the large-scale pursuit and escape game problem, achieving effective interception of large-scale attacking clusters.
[0073] Finally, the combination of local game strategy and global matching method not only significantly improved the survival rate of defenders, but also ensured a high interception success rate. The overall strategy takes into account the balance between efficient interception and the defender's own safety, while taking into account the collision loss, showing excellent technical effects and practical application value.
[0074] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0076] Figure 1 A flowchart of a large-scale regional pursuit and escape game method with collision avoidance guarantee provided in an embodiment of the present application;
[0077] Figure 2A schematic diagram of a flow chart for offline solving differential game strategies provided in an embodiment of the present application;
[0078] Figure 3 A heat map of a 2D slice of the value function provided in an embodiment of the present application;
[0079] Figure 4 A schematic diagram of the process of online bilateral graph matching and control provided in an embodiment of the present application;
[0080] Figure 5 A schematic diagram of a process of constructing a collision avoidance alliance of the defender for the 2-on-2 strategy provided in an embodiment of the present application;
[0081] FIG6( a ) is a schematic diagram of a 10-on-10 pursuit-and-escape game scenario and results provided in an embodiment of the present application;
[0082] FIG6( b ) is a schematic diagram of a 20-on-20 pursuit-and-escape game scenario and results provided in an embodiment of the present application;
[0083] Figure 7 A schematic structural diagram of a large-scale regional pursuit and escape game device with collision avoidance guarantee provided in an embodiment of the present application. DETAILED DESCRIPTION
[0084] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0085] To address this problem, an embodiment of the present application provides a large-scale regional pursuit and escape game method with collision avoidance guarantee, which is used to enable a large-scale multi-autonomous system to collaboratively intercept another group of invading multi-autonomous systems, while ensuring efficient interception while ensuring that the autonomous system itself suffers as little loss as possible due to collision. Figure 1 A flowchart of a large-scale regional pursuit and escape game method with collision avoidance guarantee provided in an embodiment of the present application.
[0086] Before explaining the technical solution of the present application, the technical scenario of the present application is first explained.
[0087] In the technical scenario of this application, the two parties in the game are the attacker and the defender. All members of the attacker and the defender move in a bounded closed convex area Ω. The straight line segment The boundary of the target area, the inside of the boundary is the target area When the attacking team members enter a certain range of the defending team members, they will be captured by the defending team and lost. If the distance between the defending team members is less than the safe distance, they will collide and lose. goal Inward movement, the goal is to reach the boundary The defenders move in Ω, and the goal is to prevent as many attackers as possible from entering the target area Ω by blocking the attackers and capturing them. goal .
[0088] Assume that in a regional pursuit game of multiple autonomous agents, there are M members of the attacker and N members of the defender, forming the attacker set and the defender set respectively. and Both members use the first-order integral dynamics model for control, and their motion control model is:
[0089]
[0090] In the formula, and They represent the coordinates of the attacking member i and the defending member j in the two-dimensional Euclidean space, and They represent the positions of the attacking member i and the defending member j at the initial moment, u i and v j They represent the normalized control amount of the attacking member i and the defending member j, respectively, and v A and v D are the maximum speeds of the attacker and defender respectively.
[0091] For the regional pursuit game system, the joint state of the attacker and the defender satisfies the following state equation:
[0092]
[0093] The solution trajectory is:
[0094]
[0095] Where x is the joint state of the attacker and the defender, which represents the coordinates of all members of the attacker and the defender in the two-dimensional Euclidean space; and are the joint control quantities of the attacker and the defender respectively; f(x,u,v) is the state transfer function of the system; Indicates that from the initial state x 0 The state at time A after inputting the control variables u(·) and v(·).
[0096] To solve the above problems, Figure 1 As shown, this application adopts the following technical solutions:
[0097] Step 101, according to the HJ reachability principle, design the avoidance set and target set of the 1-to-1 strategy and the 2-to-2 strategy respectively.
[0098] The embodiment of the present application uses the HJ (Hamilton-Jacobi) reachability principle to design avoidance sets and target sets for 1-to-1 and 2-to-2 strategies respectively. The definition of these sets provides clear constraints and goals for the defender and attacker in the pursuit-escape game, ensuring the optimization of the strategy and the security of the task. Figure 2 A schematic diagram of a flow chart for offline solving differential game strategies provided in an embodiment of the present application;
[0099] In the 2-on-2 strategy, the embodiment of the present application designs an avoidance set A and a target set R for the 8-dimensional joint state space of two defenders and two attackers. These two sets respectively define the states that need to be avoided and the states that need to be achieved during the game. The avoidance set A is used to describe the state subspace set that the state vector should not be in during the game, and the target set R is used to describe the state subspace set that the state vector should be in at the end of the game.
[0100] It should be noted that since this game problem is a zero-sum game, the results of solving the value function for the defender and the attacker are opposite. It is advisable to design the avoidance set and target set from the perspective of the attacker.
[0101] From the attacker's perspective, the avoidance set describes the failure states that the attacker needs to avoid, and the target set describes the success states that the attacker hopes to achieve.
[0102] The avoidance set mainly includes the following two situations:
[0103] The state where the attacker is completely intercepted: In a 2-on-2 game, if any member of the attacker is successfully intercepted by a member of the defender, the remaining members of the attacker may also be intercepted by another member of the defender. In this case, the attacker's mission will be completely failed. Therefore, this state must be avoided by the attacker.
[0104] Collision between attackers: If two attackers collide during movement, both of them may lose their ability to act or their strategies may become ineffective. This collision directly affects the performance of the attacker's mission and is another state that needs to be avoided from the attacker's perspective.
[0105] The target set mainly includes the following two situations:
[0106] Any member of the attacking team successfully enters the target area: If a member of the attacking team can successfully enter the target area without being intercepted by the defending team, the attacking team's mission is considered to be completed. This is the primary goal of the attacking team in the game.
[0107] Collision between defenders: If two defenders collide while attempting to intercept, causing the defender to lose the ability to act, the attacker can effectively break through the defense and achieve the mission goal. Therefore, the collision state of the defender is the ideal target state for the attacker.
[0108] It is understandable that the avoidance set clarifies the high-risk states of the attacker in the game from the attacker's perspective. The attacker needs to adopt flexible strategies to avoid these states in order to maximize the success rate of the mission. The target set clarifies the specific conditions for success from the attacker's perspective. The strategic goal of the attacker is to maximize the realization of these states and thus win the game.
[0109] It should be noted that designing the avoidance set and target set from the perspective of the attacker can not only intuitively describe the success and failure conditions of the attacker, but also reversely guide the strategy optimization of the defender through the nature of zero-sum game. The defender needs to take action against the attacker's avoidance set, forcing the attacker to enter the state it needs to avoid, while preventing the attacker from entering the favorable state of the target set. In a 2-on-2 game, the attacker's perspective design can effectively capture the complex dynamic relationship between multiple autonomous agents, allowing the defender to more accurately predict the attacker's behavior and formulate targeted interception strategies.
[0110] For the 1-to-1 strategy, the definitions of the avoidance set and the target set are logically similar to those of the 2-to-2 strategy, but since it only involves a single defender and a single attacker, its state space is simplified to a 4-dimensional joint state space, and this application will not repeat the description.
[0111] It can be understood that the HJI (Hamilton-Jacobi-Isaacs) equation of the 2-on-2 strategy constructed in the embodiment of the present application is an 8th-order PDE, so the joint state space is x=(x A1 ,y A1 ,x A2 ,y A2 ,x D1 ,y D1 ,x D2 ,y D2 ). Let’s assume that both the defender and the attacker use normalized coordinates, that is, x Ai ,x Dj ,y Ai ,y Dj ∈[-1,1]. Due to computational memory limitations, each state space dimension is quantized into 15 equal parts.
[0112] If we define R Ca is the capture radius of the defender, R Co is the collision radius of the defender members, is the Reach-avoid tube of one-to-one game. The Avoid set is the union of the state quantity sets corresponding to the following situations: In the case described by L1, when A1 is captured by D1, A2 is in the 1-to-1 Reach-avoid tube of D2, and D1 and D2 do not collide; in the case described by L2, when A2 is captured by D1, A1 is in the 1-to-1 Reach-avoid tube of D2, and D1 and D2 do not collide; in the case described by L3, when A1 is captured by D2, A2 is in the 1-to-1 Reach-avoid tube of D1, and D1 and D2 do not collide; in the case described by L4, when A2 is captured by D2, A1 is in the 1-to-1 Reach-avoid tube of D1, and D1 and D2 do not collide; in the case described by L5, the distance between A1 and D2 is too close and a collision occurs.
[0113]
[0114] L5={x∈Ω 3 |‖p A1 -p A2 ≤R Co ‖}
[0115] The avoid set is:
[0116] A 22 =L1∪L2∪L3∪L4∪L5
[0117] In the embodiment of the present application, the target set (Reach set) is the union of the state quantity sets corresponding to the following situations: in the case of M1, A1 successfully reaches the target area and is not captured by D1 or D2, and the two attackers do not collide; in the case of M2, A2 successfully reaches the target area and is not captured by D1 or D2, and the two attackers do not collide; in the case of M3, the distance between D1 and D2 is too close and a collision occurs.
[0118] M1={x∈Ω 3 |p A1 ∈Ω goal ∧‖p A1 -p D1 ‖>R Ca ∧‖p A1 -p D2 ‖>R Ca}∩{x∈Ω 3 |‖p A1 -p A2 ‖>R Co}
[0119] M2={x∈Ω3 |p A2 ∈Ω goal ∧‖p A2 -p D1 ‖>R Ca ∧‖p A2 -p D2 ‖>R Ca}∩{x∈Ω 3 |‖p A1 -p A2 ‖>R Co}
[0120] M3={x∈Ω 3 |‖p D1 -p D2 ‖≤R Co}
[0121] R 22 =M1∪M2∪M3
[0122] Step 102, bring the avoidance set and target set of the designed 1-on-1 strategy and 2-on-2 strategy into the HJI equation, solve the value of the game value function in the discrete state space and store it.
[0123] The embodiment of the present application introduces the avoidance set and target set of the 1-to-1 and 2-to-2 strategies into the HJI equation, and formally expresses the pursuit problem as a minimization and maximization planning problem, which is expressed as follows:
[0124]
[0125] Where V(x,t) represents the value function, the optimal strategy result at state x and time t; γ is the strategy function; From the initial state x 0 At the beginning, the state trajectory at time τ when the defender adopts strategy γ and the attacker adopts strategy v; the level functions h(x) and l(x) are defined according to the designed avoidance set and target set, which are the penalty function related to the avoidance set and the reward function related to the target set, respectively.
[0126] In the HJI equation, the level function h(x) is in the form of:
[0127]
[0128] In the formula, d(x,A) represents the minimum distance from state x to avoidance set A, A Cis the complement of the avoidance set, indicating that the state is not in the avoidance set. This function describes the area that the attacker needs to avoid. If the state x is in the avoidance set (such as the defender's successful interception area or the attacker's collision area), its value is positive, indicating a high-risk state; if the state is outside the avoidance set, its value is negative, indicating a low-risk state.
[0129] The form of l(x) is as follows:
[0130]
[0131] In the formula, d(x,R) represents the minimum distance from state x to the avoidance set A and the target set R, and R C is the complement of the target set, indicating the part of the state that has not entered the target area. This function describes the target area of the attacker. If the state x is in the target set (such as the attacker successfully reaches the target area), its value is negative, indicating that the attacker's mission is successful; if the state is outside the target set, its value is positive, indicating that the attacker needs to make further efforts.
[0132] Furthermore, the present application solves the HJI equation in the joint state space of 1-to-1 and 2-to-2 strategies through a numerical discretization method, and stores the value function results obtained by numerical calculation in each node of the discrete state space for subsequent strategy execution and optimal control quantity calculation.
[0133] It can be understood that the numerical discretization method is the key to solving the HJI equation. By discretizing the continuous state space and time dimensions into finite grids, the complex continuous optimization problem is converted into a discrete optimization problem, making the calculation of the value function and the execution of the strategy more feasible. In the actual implementation process, the discretization method or numerical solution tool of the existing technology can be used to complete it, such as the dynamic programming method based on grid division, the fast recursive algorithm, and the mature technology such as the finite difference or finite volume method. These methods have been widely used in related fields and can effectively meet the computing needs of high-dimensional space and complex dynamic systems.
[0134] It should be noted that this application does not limit the specific implementation of the numerical discretization method. Users can choose appropriate discretization techniques or tools according to the specific scenario requirements to achieve a balance between computational efficiency and accuracy. For example, for high-dimensional joint state spaces, hierarchical grid discretization or adaptive grid division can be selected to reduce the computational burden; for the time dimension, a fixed step size or variable step size discretization strategy can be adopted to more flexibly adapt to the dynamic changes of the system. This application focuses on providing support for strategy design and execution through the calculation results of the discretized value function, and the specific discretization implementation method can be adjusted according to actual needs.
[0135] Through the above numerical method, the complex pursuit and escape game problem of the HJI equation is transformed into an optimization problem in the discrete state space. The value function not only describes the optimal strategy results in each state, but also provides a basis for real-time decision-making during the game. When the strategy is executed, the system can quickly determine the control input and optimal path selection of the defender by looking up the stored value function, ensuring the efficient completion of the interception task.
[0136] Step 103, constructing a collision avoidance alliance of the defender according to the boundary fence information of the 2-to-2 strategy and the distances between the defender members, and determining a 2-to-2 matching result in the collision avoidance alliance.
[0137] In the embodiment of the present application, it is first determined whether the defender has matched any attacker in the last matching process. If the defender has not matched any attacker, the 2-to-2 strategy is not adopted and it is directly determined as "no 2-to-2 matching result".
[0138] For defenders who can use a 2v2 strategy, see Figure 4 In the embodiment of the present application, a collision avoidance alliance is constructed based on the collision risk between defenders. If there is a collision risk between defenders, that is, the distance between the two defenders is less than a threshold, and the two defenders can cooperate to intercept the two attackers under a 2-to-2 strategy, they will be added to the collision avoidance alliance.
[0139] Furthermore, for the two defenders in the collision avoidance alliance, the present application calculates the boundary fence based on their positions. On the boundary fence, both the chasing and fleeing parties have strategies to prevent the opponent from winning, where the boundary fence is defined as:
[0140]
[0141] Where V(x, t) is the value function, which indicates the optimal strategy result of the defender to intercept the attacker at state x and time t, and the boundary reflects the dynamic equilibrium point between the attacker and the defender in the state space.
[0142] If the last matched offensive members of two defenders in the collision avoidance alliance are both outside the fence, the corresponding two defenders and two offensive members form a 2-to-2 match result.
[0143] In one possible embodiment, a hyperplane with a value function of 0 is obtained by the difference method as the boundary fence for 2-to-2 matching. When the defenders are at (0, 0.3) and (0, -0.2) and an attacker is at (-0.5, 0), the heat map of the 2D slice of the value function is as follows: Figure 3 As shown, the 1-to-1 and 2-to-2 boundary fences are shown by the purple and blue outlines in the figure.
[0144] Furthermore, the specific matching process is as follows Figure 5 As shown, including:
[0145] First, check whether the last execution was a 1-on-1 strategy or a 2-on-2 strategy. If it was not a 2-on-2 strategy, it will be directly judged as "no 2-on-2 match result", indicating that the current conditions do not meet the requirements of a 2-on-2 match. If the last execution was a 2-on-2 strategy, start calculating the distance between each defender and other defenders. This step is used to determine whether there is a possible collision risk between members.
[0146] Then, based on the calculated distance, determine whether the current member is the nearest neighbor of the defender. If not, it is determined as "no 2-to-2 match result", indicating that the member cannot cooperate with the nearest other defender to implement the 2-to-2 strategy. Further, check whether the relationship between the two defenders and the attackers in the last match is outside the boundary range of the 2-to-2 strategy. If this condition is not met, a valid match cannot be formed and it is determined as "no 2-to-2 match result".
[0147] If two defenders meet all the above conditions, they will form a 2-on-2 match with two attackers. This match is based on the defenders' collision avoidance alliance and fence information, ensuring that the defenders can cooperate to intercept the attackers without collision.
[0148] It can be understood that the calculation process of the matching results ensures the coordination between the defenders and the maximum interception success rate of the attackers, while avoiding the risk of collision between the defenders. The final 2-on-2 matching result is determined based on the combination of all valid matches in the collision avoidance alliance, providing clear guidance for subsequent strategy execution.
[0149] Step 104 , constructing a bilateral graph based on the 2-to-2 matching result and the boundary barrier information of the 1-to-1 strategy, performing maximum matching calculation in the bilateral graph, and obtaining a 1-to-1 matching result.
[0150] For the defenders who have not entered the 2-to-2 strategic alliance and the attackers who have not been matched, the embodiment of the present application determines the 1-to-1 matching relationship between the defenders and the attackers by constructing a bilateral graph and running the maximum matching algorithm. This process is based on the boundary fence information of the 1-to-1 strategy and the maximum matching algorithm. Specifically, it includes:
[0151] First, based on the 2-to-2 matching results and the boundary fence of the 1-to-1 strategy, a bilateral graph G is constructed, which is expressed as:
[0152] G=(U,V,E)
[0153] Where U is the set of defender members, which indicates the defender members who do not participate in the 2-on-2 matching; V is the set of attacker members, which indicates the attacker members who are not matched by the 2-on-2 strategy; E is the set of edges in the bilateral graph, which indicates the matching relationship that can be formed between defender members and attacker members.
[0154] In a bilateral graph, the edge set E is constructed based on the boundary fence In the fence In the outer area of the fence, the defender has a strategy that can effectively intercept the attacker. When constructing the bilateral graph, for each defender, find the attacker that he can intercept outside the fence, and add an edge from the defender to the attacker in the bilateral graph.
[0155] Then, in the constructed bilateral graph, the maximum matching algorithm is run to find the maximum matching result M in the bilateral graph G, which makes each defender match at most one attacker, and the number of matches is maximized.
[0156] Finally, according to the maximum matching result M, each defender is assigned an attacker target, and a 1-to-1 matching result is obtained. Specifically, each matching edge represents the match between a defender and an attacker; the target of each defender is the attacker that matches it, and its strategy is to develop the optimal interception path around this target.
[0157] Through the construction of a bilateral graph and the operation of a maximum matching algorithm, this application can efficiently solve the matching problem of defenders and attackers who do not participate in a 2-to-2 strategy. Compared with direct 1-to-1 strategy optimization, the maximum matching algorithm can ensure the maximum number of matches under given constraints, optimize the overall interception efficiency of the defender, and this method can adapt to a variety of pursuit scenarios, whether the number of attackers is greater than the number of defenders or the number of defenders is greater than the number of attackers. This method provides a systematic solution for the 1-to-1 matching of defenders in the pursuit game, while ensuring the efficient completion of the interception task.
[0158] Step 105, using the calculated HJI equation value functions of the 2-on-2 strategy and the 1-on-1 strategy, respectively, the optimal control amount is obtained for the defender members implementing the 2-on-2 strategy and the 1-on-1 strategy.
[0159] In the embodiment of the present application, for defenders with high collision risks, a 2-to-2 differential game strategy is preferentially adopted, and its multi-member coordination ability is utilized, combined with the HJI equation value function of the 2-to-2 strategy to calculate the optimal control amount. This strategy can intercept multiple attackers at the same time, while effectively avoiding collisions between defenders. For defenders who have not entered the 2-to-2 match, a 1-to-1 differential game strategy is adopted for interception. The HJI equation value function based on the 1-to-1 strategy calculates the optimal control amount, adjusts the motion trajectory of the defenders, and achieves accurate interception of a single attacker. For defenders who have not obtained a matching result, a collision avoidance priority maneuvering strategy is adopted to find a dynamically adjusted interception path while maintaining their own safety.
[0160] Specifically, based on the HJI equation value function V2(x,t) of the 2-on-2 strategy, the partial derivative of the value function with respect to the state x is calculated Solve the optimal control amount of the defender under the 2-on-2 strategy, the expression is:
[0161]
[0162] In the formula, u * (x, t) represents the optimal control amount of the defender under the 2-on-2 strategy. This formula ensures that the defender can formulate the optimal strategy and adjust the movement trajectory in the worst case to achieve joint interception of multiple attackers through the minimization and maximization optimization process.
[0163] Based on the HJI equation value function V1(x,t) of the 1-to-1 strategy, calculate the partial derivative of the value function with respect to the state x Solve the optimal control amount of the defender under the 1-to-1 strategy, the expression is:
[0164]
[0165] In the formula, v * (x, t) represents the optimal control amount of the defender under the 1-to-1 strategy. This formula determines the best interception path for the defender through a maximization optimization process to ensure accurate interception of a single attacker.
[0166] After obtaining the optimal control amount, the defender can dynamically adjust its own movement trajectory and strategy execution according to the calculation results, thereby achieving efficient interception of the attacker and ensuring a safe distance between the defenders. Specifically:
[0167] For the defenders who adopt the 2-on-2 differential game strategy, the defenders can *(x, t) adjusts its trajectory to achieve the best relative position between itself and the attacker, thus achieving collaborative interception. This strategy also ensures that the two defenders can effectively divide the work, avoid mutual interference, and improve interception efficiency.
[0168] For the defenders who adopt the 1-to-1 differential game strategy, the defenders can control the optimal control amount v * (x, t) accurately adjusts its own trajectory and directly aims at the target attacker for interception. This strategy emphasizes the game relationship between a single defender and a single attacker, and is suitable for situations that require individual precise actions.
[0169] For defenders who are not matched with attackers, the system will assign them a maneuver strategy based on the principle of collision avoidance priority. These members will dynamically adjust their paths to prioritize avoiding collisions with other defenders, while monitoring the attackers' movements and looking for new interception opportunities.
[0170] In a possible embodiment, schematic diagrams of the 10 vs. 10 and 20 vs. 20 pursuit and escape game scenarios and results are shown in FIG6(a) and FIG6(b) respectively. The blue track is the track of the defender, the red track is the track of the attacker, and the purple frame line is the boundary of the target area.
[0171] In order to implement the above-mentioned embodiments, the present application also proposes a large-scale regional pursuit and escape game device with collision avoidance guarantee. Figure 7 This is a schematic diagram of the structure of a large-scale regional pursuit and escape game device with collision avoidance guarantee provided in an embodiment of the present application. Figure 7 As shown, the device comprises:
[0172] A design module 100 is used to design the avoidance set and the target set of the 1-to-1 strategy and the 2-to-2 strategy respectively according to the HJ reachability principle;
[0173] A solution module 200 is used to bring the avoidance set and target set of the designed 1-on-1 strategy and 2-on-2 strategy into the HJI equation, solve the value of the game value function in the discrete state space and store it;
[0174] The first matching module 300 is used to build a collision avoidance alliance of the defender according to the boundary fence information of the 2-to-2 strategy and the distance between the defender members, and determine the 2-to-2 matching result in the collision avoidance alliance;
[0175] The second matching module 400 is used to construct a bilateral graph according to the 2-to-2 matching result and the boundary fence information of the 1-to-1 strategy, perform maximum matching calculation in the bilateral graph, and obtain a 1-to-1 matching result;
[0176] The obtaining module 500 is used to use the calculated HJI equation value functions of the 2-on-2 strategy and the 1-on-1 strategy to respectively obtain the optimal control amount for the defender members implementing the 2-on-2 strategy and the 1-on-1 strategy.
[0177] In order to implement the above embodiments, the present application also proposes an electronic device, comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided by the above embodiments.
[0178] In order to implement the above embodiments, the present application also proposes a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the methods provided by the above embodiments.
[0179] In order to implement the above embodiments, the present application also proposes a computer program product, including a computer program, which implements the methods provided by the above embodiments when executed by a processor.
[0180] The collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in this application are in compliance with relevant laws and regulations and do not violate public order and good morals.
[0181] It should be noted that personal information from users should be collected for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. In addition, such collection / sharing should be carried out after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign the agreement / authorization including authorization of relevant user information before the user uses the function. In addition, any necessary steps should be taken to protect and safeguard access to such personal information data and ensure that others who have access to personal information data comply with its privacy policy and procedures.
[0182] The present application is expected to provide an implementation scheme for users to selectively block the use or access of personal information data. That is, the present disclosure is expected to provide hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by limiting data collection and deleting the data. In addition, when applicable, such personal information is de-identified to protect the privacy of the user.
[0183] In the description of the aforementioned embodiments, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0184] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0185] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0186] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0187] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0188] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0189] The storage medium mentioned above may be a read-only memory, a disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
[0190] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this application can be executed in parallel, sequentially or in different orders, as long as the expected results of the technical solution of this application can be achieved, and this document is not limited here.
[0191] The above specific implementations do not constitute a limitation on the protection scope of this application. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of this application should be included in the protection scope of this application.
Claims
1. A large-scale regional pursuit and escape game method with collision avoidance guarantee, characterized in that: The following steps are involved: According to the HJ reachability principle, the avoidance set and target set of the 1-to-1 strategy and the 2-to-2 strategy are designed respectively; Substitute the avoidance set and target set of the designed 1-on-1 strategy and 2-on-2 strategy into the HJI equation, solve the value of the game value function in the discrete state space and store it; According to the boundary fence information of the 2-on-2 strategy and the distance between the defenders, a collision avoidance alliance of the defenders is constructed, and the 2-on-2 matching result is determined in the collision avoidance alliance; A bilateral graph is constructed according to the 2-to-2 matching result and the boundary fence information of the 1-to-1 strategy, and a maximum matching calculation is performed in the bilateral graph to obtain a 1-to-1 matching result; The calculated HJI equation value functions of the 2-on-2 strategy and the 1-on-1 strategy are used to obtain the optimal control quantity for the defenders who implement the 2-on-2 strategy and the 1-on-1 strategy respectively.
2. The method according to claim 1, characterized in that Assume that in a regional pursuit game of multiple autonomous agents, there are M members of the attacker and N members of the defender, forming the attacker set and the defender set respectively. and And both members adopt the first-order integral dynamic model for control, then the motion control model is: In the formula, and They represent the coordinates of the attacking member i and the defending member j in the two-dimensional Euclidean space, and They represent the positions of the attacking member i and the defending member j at the initial moment, u i and v j They represent the normalized control amount of the attacking member i and the defending member j, respectively, and v A and v D are the maximum speeds of the attacker and defender respectively.
3. The method according to claim 2, characterized in that For the regional pursuit game system, the joint state of the attacker and the defender satisfies the following state equation: The solution trajectory is: Where x is the joint state of the attacker and the defender, which represents the coordinates of all members of the attacker and the defender in the two-dimensional Euclidean space; and are the joint control quantities of the attacker and the defender respectively; f(x,u,v) is the state transfer function of the system; Indicates that from the initial state x 0 The state at time τ after inputting the control variables u(·) and v(·).
4. The method according to claim 3, characterized in that According to the HJ reachability principle, the avoidance set and target set of the 2-to-2 strategy are designed, including: For the 8-dimensional joint state space of 2 defenders and 2 attackers, design the avoidance set A and target set R of the 2-on-2 strategy; The avoidance set A is used to describe the set of state subspaces that the state vector should not be in during the game, and its definition includes the following two situations: when a member of the attacking party is intercepted by a member of the defending party, another member of the defending party can ensure to intercept another member of the attacking party; a collision occurs between two members of the attacking party; The target set R is used to describe the set of state subspaces that the state vector should be in when the game ends, and its definition includes the following two situations: any member of the attacking party successfully enters the target area and is not intercepted by any member of the defending party; a collision occurs between two members of the defending party.
5. The method according to claim 4, characterized in that The method of bringing the avoidance set and target set of the designed 1-to-1 strategy and 2-to-2 strategy into the HJI equation, solving the value of the game value function in the discrete state space and storing it includes: The avoidance set and target set of the 1-to-1 and 2-to-2 strategies are brought into the HJI equation, where the pursuit problem is expressed as a minimization-maximization planning problem, which is expressed as follows: Where V(x,t) represents the value function, the optimal strategy result at state x and time t; γ is the strategy function; From the initial state x 0 At the beginning, the state trajectory at time τ when the defender adopts strategy γ and the attacker adopts strategy v; the level functions h(x) and l(x) are defined according to the designed avoidance set and target set, and their specific forms are: In the formula, d(x,A) and d(x,R) represent the minimum distances from state x to the avoidance set A and the target set R respectively; Through the numerical discretization method, the HJI equation is solved in the joint state space of 1-to-1 and 2-to-2 strategies, and the value function results obtained by numerical calculation are stored in each node of the discrete state space for subsequent strategy execution and optimal control quantity calculation.
6. The method according to claim 5, characterized in that The method of constructing a collision avoidance alliance of the defender based on the boundary fence information of the 2-to-2 strategy and the distance between the defender members, and determining the 2-to-2 matching result in the collision avoidance alliance, includes: Determine whether the defending team members have matched with the attacking team members in the last matching process. If the defending team members have not matched with the attacking team members, the 2-to-2 strategy will not be adopted; For defenders who can adopt a 2-on-2 strategy, a collision avoidance alliance is established based on the collision risk between defenders. If there is a collision risk between defenders and two defenders can cooperate to intercept two attackers under the 2-on-2 strategy, they will be added to the collision avoidance alliance. For the two defenders in the collision avoidance alliance, calculate the fence based on their positions On the boundary fence, both the chasing and fleeing parties have strategies to prevent the opponent from winning, where the boundary fence is defined as: Where V(x,t) is the value function, which represents the optimal strategy result of the defender being able to intercept the attacker at state x and time t; If the last matched offensive members of two defenders in the collision avoidance alliance are both outside the fence, the corresponding two defenders and two offensive members form a 2-to-2 match result.
7. The method according to claim 6, characterized in that The method of constructing a bilateral graph based on the 2-to-2 matching result and the boundary fence information of the 1-to-1 strategy, performing a maximum matching calculation in the bilateral graph, and obtaining a 1-to-1 matching result includes: Based on the 2-to-2 matching results and the boundary fence of the 1-to-1 strategy, a bilateral graph G is constructed, which is expressed as: G=(U,V,E) In the formula, U is the set of defender members, which indicates the defender members who are not involved in the 2-to-2 matching; V is the set of attacker members, which indicates the attacker members who are not matched by the 2-to-2 strategy; E is the set of edges in the bilateral graph, which indicates the matching relationship that can be formed between defender members and attacker members; In the fence In the state region outside the fence, determine whether there is a defender whose strategy can intercept the attacker, and construct all the edges from the defender to the attacker outside the fence in the bilateral graph; Run the bilateral graph maximum matching algorithm to find the maximum matching result M in the bilateral graph G, which makes each defender match with at most one attacker, and maximizes the number of matches. According to the maximum matching result M, the offensive target corresponding to each defensive member is determined to obtain a 1-to-1 matching result.
8. The method according to claim 7, characterized in that The HJI equation value functions of the 2-to-2 strategy and the 1-to-1 strategy obtained by calculation are used to obtain the optimal control amount for the defenders implementing the 2-to-2 strategy and the 1-to-1 strategy, respectively, including: Based on the HJI equation value function V2(x,t) of the 2-to-2 strategy, calculate the partial derivative of the value function with respect to the state x Solve the optimal control amount of the defender under the 2-on-2 strategy, the expression is: In the formula, u * (x, t) represents the optimal control amount of the defender under the 2-on-2 strategy; Based on the HJI equation value function V1(x,t) of the 1-to-1 strategy, calculate the partial derivative of the value function with respect to the state x Solve the optimal control amount of the defender under the 1-to-1 strategy, the expression is: In the formula, v * (x, t) represents the optimal control amount of the defender under the 1-to-1 strategy.
9. A large-scale regional pursuit and escape game device with collision avoidance guarantee, characterized in that: include: A design module, used to design the avoidance set and target set of the 1-to-1 strategy and the 2-to-2 strategy respectively according to the HJ reachability principle; A solution module is used to bring the avoidance set and target set of the designed 1-on-1 strategy and 2-on-2 strategy into the HJI equation, solve the value of the game value function in the discrete state space and store it; The first matching module is used to build a collision avoidance alliance of the defender according to the boundary fence information of the 2-to-2 strategy and the distance between the defender members, and determine the 2-to-2 matching result in the collision avoidance alliance; A second matching module is used to construct a bilateral graph according to the 2-to-2 matching result and the boundary fence information of the 1-to-1 strategy, and perform maximum matching calculation in the bilateral graph to obtain a 1-to-1 matching result; The obtaining module is used to adopt the calculated HJI equation value functions of the 2-to-2 strategy and the 1-to-1 strategy to respectively obtain the optimal control quantity for the defender members implementing the 2-to-2 strategy and the 1-to-1 strategy.
10. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 8.