Multi-against-multi attack and defense game method and device in obstacle environment

By integrating differential game theory and deep reinforcement learning methods, expanding the defense victory zone, and designing an interception strategy for drone swarms, this approach solves the problem of intercepting defensive drones in many-to-many attack-defense games under obstacle conditions. It achieves effective interception of attacking drones and protection of the target area under polygonal obstacles.

CN120046494BActive Publication Date: 2026-05-05BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIHANG UNIV
Filing Date
2025-02-19
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

In obstacle-prone environments, existing technologies struggle to effectively address the challenge of how defensive drones in a multi-to-multi attack and defense drone swarm can maximize their interception of attacking drones and protect the target area. This is especially true when polygonal obstacles are present, where learning-based methods may result in poor interception performance and lack performance guarantees.

Method used

By integrating differential game theory and deep reinforcement learning methods, the defensive victory region is expanded. By establishing a many-to-many attack and defense problem model, expanding the defensive victory region under differential game theory and reinforcement learning, and combining neural symbolic algorithms and binary integer programming, an interception strategy for defending against drones is designed.

Benefits of technology

In the attack and defense problem of drone swarms with polygonal obstacles, the defending drones can intercept as many attacking drones as possible, protect the target area to the maximum extent, and improve the success rate of defense.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046494B_ABST
    Figure CN120046494B_ABST
Patent Text Reader

Abstract

This application discloses a method and apparatus for multi-to-multi attack and defense game in an obstacle environment, relating to the field of unmanned aerial vehicle (UAV) control technology. The method includes: establishing a multi-to-multi attack and defense problem model for a UAV swarm; establishing a one-to-one attack and defense subgame problem; expanding the defense victory region under differential game based on the one-to-one attack and defense subgame problem to obtain an expanded defense victory region under differential game; expanding the defense victory region under reinforcement learning to obtain an expanded defense victory region under reinforcement learning; and performing multi-to-multi defense decision based on the multi-to-multi attack and defense problem model, the expanded defense victory region under differential game, and the expanded defense victory region under reinforcement learning to obtain a multi-to-multi defense decision planning result. In the attack and defense problem of a UAV swarm with polygonal obstacles, this application enables the defending UAV to intercept as many attacking UAVs as possible, maximizing the protection of the target area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of unmanned aerial vehicle (UAV) control technology, and in particular to a method and apparatus for many-to-many attack and defense game in an obstacle environment. Background Technology

[0002] Drone swarm attack-defense differential game theory considers an adversarial scenario where a group of evasive (or attacking) drones aims to enter an area protected by a group of tracking (or defending) drones. Compared to chase-and-escape games, attack-defense differential games are more challenging and have broader applications, such as attacking critical facilities or high-value targets. Since attacking drones not only aim to avoid being captured by defending drones but also to approach key targets, while defending drones must not only intercept attacking drones but also protect the target area, this problem is more challenging and practical. In barrier-free environments, many researchers have studied attack-defense differential games, differing in game space, capture conditions, information structure, winning team, speed ratio, dynamics of the attackers and defenders, and target area. However, research on attack-defense differential games in obstacle-prone environments is still limited.

[0003] In recent years, learning-based methods have been frequently used to solve attack-defense differential games. Related technologies have developed an imitation learning framework utilizing graph neural networks and centralized expert algorithms. To defend against a faster escapee, a deep reinforcement learning (RL) method has been proposed, which trains the tracking strategy by fixing an analytical policy for the escapee. However, learning-based methods may encounter interpretability and performance guarantees issues, which are crucial in adversarial scenarios. Furthermore, in attack-defense analysis with obstacles, the results are often conservative. In anti-no-game scenarios, such as drone swarm attack-defense in polygonal obstacle environments, learning-based methods perform poorly in interception. Summary of the Invention

[0004] The purpose of this application is to provide a method and apparatus for many-to-many attack and defense game in an obstacle environment. In the attack and defense problem of drone swarms with polygonal obstacles, the defending drone can intercept as many attacking drones as possible, protect the target area to the maximum extent, and improve the success rate of defense.

[0005] To achieve the above objectives, this application provides the following solution:

[0006] Firstly, this application provides a method for many-to-many attack and defense game in an obstacle environment, including:

[0007] A multi-to-multi attack and defense problem model for a drone swarm is established; the drone swarm includes several attacker drones and several defender drones; the multi-to-multi attack and defense problem model includes a dynamic model of the attack and defense differential game between the two sides and a game scenario model under polygonal obstacles.

[0008] A one-to-one attack-defense subgame problem is established; the one-to-one attack-defense subgame problem includes a one-to-one attack-defense subgame model, a defensive victory region under differential game theory, and a defensive victory region under reinforcement learning; the defensive victory region consists of several defensive victory states; the defensive victory state is the state in which the defender's drone achieves victory;

[0009] Based on the aforementioned one-to-one attack and defense subgame problem, the defensive victory region under the differential game is expanded to obtain the expanded defensive victory region under the differential game.

[0010] The defensive victory region under reinforcement learning is expanded to obtain the extended defensive victory region under reinforcement learning.

[0011] Based on the many-to-many attack and defense problem model, the defensive victory region under the extended differential game and the defensive victory region under the extended reinforcement learning, many-to-many defense decision is made, and the many-to-many defense decision planning result is obtained.

[0012] Optionally, a one-to-one attack and defense subgame problem can be established, specifically including:

[0013] Establish a one-to-one attack and defense subgame model;

[0014] Establish models for defensive victory state, defensive victory zone, attack victory state, and attack victory zone; the attack victory state is the state in which the attacker's drone achieves victory; the attack victory zone is composed of several attack victory states.

[0015] Using differential game theory and reinforcement learning algorithms, a model of the winning state and winning region under the differential game theory and reinforcement learning algorithms is established; the model of the winning state and winning region under the differential game theory and reinforcement learning algorithms includes the defensive winning state under differential game theory, the defensive winning region under differential game theory, the defensive winning state under reinforcement learning, and the defensive winning region under reinforcement learning.

[0016] Establish forward reachable set and backward reachable set region models;

[0017] Using reinforcement learning algorithms, based on the forward reachable set and backward reachable set region model and the defense victory region under reinforcement learning, a new defense victory region under reinforcement learning is constructed; the new defense victory region under reinforcement learning is used to expand to obtain the extended defense victory region under reinforcement learning.

[0018] Optionally, based on the aforementioned one-to-one attack-defense subgame problem, the defensive victory region under the differential game is expanded to obtain the expanded defensive victory region under the differential game, specifically including:

[0019] Based on the minimum Euclidean distance algorithm and the one-to-one attack-defense subgame problem, the defensive victory region under the differential game is expanded to obtain the expanded defensive victory region under the differential game.

[0020] Optionally, the defensive victory region under reinforcement learning is expanded to obtain an expanded defensive victory region under reinforcement learning, specifically including:

[0021] By optimizing the reinforcement learning algorithm using a near-end strategy, the defensive victory region under reinforcement learning is expanded to obtain an expanded defensive victory region under reinforcement learning.

[0022] Optionally, based on the many-to-many attack and defense problem model, the defensive victory region under the extended differential game, and the defensive victory region under the extended reinforcement learning, many-to-many defense decisions are made to obtain many-to-many defense decision planning results, specifically including:

[0023] Based on a many-to-many attack and defense problem model, a defense victory region under extended differential game theory, and a defense victory region under extended reinforcement learning, a subgraph of a drone swarm is constructed. The subgraph includes vertex features and edge features. The vertex features represent the positions of the attacker and defender drones. The edge features indicate that the current states of the attacker and defender drones belong to the defense victory region under extended differential game theory or their union. The union is the union of the defense victory region under extended differential game theory and the defense victory region under extended reinforcement learning.

[0024] Based on the subgraph, a binary integer programming problem model for unmanned aerial vehicle (UAV) swarm task matching is established.

[0025] Solving the binary integer programming problem model yields the many-to-many defense decision programming results.

[0026] Secondly, this application provides a multi-to-multi attack and defense game device under obstacle conditions, including:

[0027] A module for establishing a many-to-many attack and defense problem model is used to establish a many-to-many attack and defense problem model for a drone swarm; the drone swarm includes several attacker drones and several defender drones; the many-to-many attack and defense problem model includes a dynamic model of the attack and defense differential game between the two sides and a game scenario model under polygonal obstacles.

[0028] A module for establishing one-to-one attack and defense subgame problems is used to establish one-to-one attack and defense subgame problems. The one-to-one attack and defense subgame problems include one-to-one attack and defense subgame models, defensive victory regions under differential games, and defensive victory regions under reinforcement learning. The defensive victory regions are composed of several defensive victory states. The defensive victory states are the states in which the defender's drone achieves victory.

[0029] The defensive victory region expansion module under differential game is used to expand the defensive victory region under differential game based on the one-to-one attack and defense subgame problem, so as to obtain the expanded defensive victory region under differential game.

[0030] The reinforcement learning-based defensive victory region extension module is used to extend the defensive victory region under reinforcement learning to obtain an extended reinforcement learning-based defensive victory region.

[0031] The many-to-many defense decision module is used to make many-to-many defense decisions based on the many-to-many attack and defense problem model, the defense victory region under the extended differential game, and the defense victory region under the extended reinforcement learning, and obtain the many-to-many defense decision planning results.

[0032] Optionally, the one-to-one attack-defense subgame problem establishment module is used for:

[0033] Establish a one-to-one attack and defense subgame model;

[0034] Establish models for defensive victory state, defensive victory zone, attack victory state, and attack victory zone; the attack victory state is the state in which the attacker's drone achieves victory; the attack victory zone is composed of several attack victory states.

[0035] Using differential game theory and reinforcement learning algorithms, a model of the winning state and winning region under the differential game theory and reinforcement learning algorithms is established; the model of the winning state and winning region under the differential game theory and reinforcement learning algorithms includes the defensive winning state under differential game theory, the defensive winning region under differential game theory, the defensive winning state under reinforcement learning, and the defensive winning region under reinforcement learning.

[0036] Establish forward reachable set and backward reachable set region models;

[0037] Using reinforcement learning algorithms, based on the forward reachable set and backward reachable set region model and the defense victory region under reinforcement learning, a new defense victory region under reinforcement learning is constructed; the new defense victory region under reinforcement learning is used to expand to obtain the extended defense victory region under reinforcement learning.

[0038] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned multi-to-multi attack and defense game method under obstacle conditions.

[0039] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the aforementioned multi-to-multi attack and defense game method under obstacle conditions.

[0040] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method for many-to-many attack and defense game in an obstacle environment.

[0041] According to the specific embodiments provided in this application, the following technical effects are disclosed:

[0042] This application provides a method and apparatus for many-to-many attack and defense game theory in obstacle-prone environments. By integrating differential game (DG) and deep reinforcement learning (RL) methods, it expands the defensive victory zone, completes the matching of defensive drones with attacking drones, and further selects corresponding defensive maneuver strategies to intercept and counter the attacking drones. Using this method, in the attack and defense problem of drone swarms with polygonal obstacles, defensive drones can intercept as many attacking drones as possible, maximizing the protection of the target area. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This is a diagram illustrating the application environment of a multi-to-multi attack and defense game method under obstacle conditions, as described in one embodiment of this application.

[0045] Figure 2 A flowchart illustrating a multi-to-multi attack and defense game method under an obstacle environment, provided as an embodiment of this application;

[0046] Figure 3 A schematic diagram illustrating the specific process of a multi-to-multi attack and defense game method in an obstacle environment provided in an embodiment of this application;

[0047] Figure 4An initial assignment graph for a defender drone algorithm based on a multiplayer onsite and close-to-goal (MOCG) differential game (ESP) with Euclidean shortest distance extension, provided in an embodiment of this application;

[0048] Figure 5 A diagram showing the strategy results of MOCG-ESP differential game provided in an embodiment of this application;

[0049] Figure 6 An initial allocation diagram of MOCG-ESP many-to-many attack and defense differential game strategy based on neural symbols is provided in one embodiment of this application;

[0050] Figure 7 A diagram illustrating the strategy results of a multi-to-multi attack and defense differential game based on neural symbols in an embodiment of this application;

[0051] Figure 8 This is a schematic diagram of the functional modules of a multi-to-multi attack and defense game device in an obstacle environment, provided in an embodiment of this application.

[0052] Figure 9 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0053] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0054] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0055] The multi-to-multi attack and defense game method in obstacle-based environments provided in this application can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be set up independently, integrated into server 104, or placed in the cloud or on another server. Terminal 102 can send decision requests to be processed to server 104. After receiving the decision requests, server 104 establishes a many-to-many attack and defense problem model for a drone swarm, establishes a one-to-one attack and defense subgame problem, expands the defense victory region under differential game theory to obtain an expanded differential game victory region, expands the defense victory region under reinforcement learning to obtain an expanded reinforcement learning victory region, and performs many-to-many defense decision-making based on the many-to-many attack and defense problem model, the expanded differential game victory region, and the expanded reinforcement learning victory region, obtaining the many-to-many defense decision planning result. Server 104 can feed back the obtained many-to-many defense decision planning result for the decision requests to terminal 102. In addition, in some embodiments, the many-to-many attack and defense game method in the obstacle environment can also be implemented by the server 104 or the terminal 102 alone. For example, the terminal 102 can directly perform many-to-many attack and defense differential game decision processing on the decision request to be processed, or the server 104 can obtain the decision request to be processed from the data storage system and perform many-to-many attack and defense differential game decision processing on the decision request to be processed.

[0056] The terminal 102 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be implemented using a standalone server or a server cluster composed of multiple servers, or it can be a cloud server.

[0057] In one exemplary embodiment, such as Figure 2 and Figure 3 As shown, a method for many-to-many attack and defense game in an obstacle environment is provided. This method is executed by a computer device, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, the method is applied to... Figure 1 Taking server 104 as an example, the explanation includes the following steps 201 to 205.

[0058] in:

[0059] Step 201: Establish a multi-to-multi attack and defense problem model for a drone swarm; the drone swarm includes several attacker drones and several defender drones; the multi-to-multi attack and defense problem model includes a dynamic model of the attack and defense differential game between the two sides and a game scenario model under polygonal obstacles.

[0060] Step 202: Establish a one-to-one attack-defense subgame problem; the one-to-one attack-defense subgame problem includes a one-to-one attack-defense subgame model, a defensive victory region under differential game theory, and a defensive victory region under reinforcement learning; the defensive victory region consists of several defensive victory states; the defensive victory state is the state in which the defender drone wins.

[0061] Step 203: Based on the one-to-one attack and defense subgame problem, the defensive victory region under the differential game is expanded to obtain the expanded defensive victory region under the differential game.

[0062] Step 204: Expand the defensive victory region under reinforcement learning to obtain the expanded defensive victory region under reinforcement learning.

[0063] Step 205: Based on the many-to-many attack and defense problem model, the defensive victory region under the extended differential game and the defensive victory region under the extended reinforcement learning, perform many-to-many defense decision-making and obtain the many-to-many defense decision planning results.

[0064] By implementing steps 201 to 205, and integrating methods based on Differential Game (DG) and Deep Reinforcement Learning (RL), the defensive victory zone is expanded. A neural symbolic algorithm is used to match defensive drones with attacking drones, and corresponding defensive maneuver strategies are selected to intercept and counter the attacking drones. Using this method, in drone swarm attack and defense problems with polygonal obstacles, defensive drones can intercept as many attacking drones as possible, maximizing the protection of the target area.

[0065] In the differential game problem of attack and defense of drone swarms, consider a swarm consisting of attacking drone alliances and defending drone alliances, where the attacking drone alliance ε has N... E The attacker's drone, denoted as... Indicates the Nth E One attacker drone, defending drone alliance Having N P One defender drone, denoted as } is the Nth P A defender drone.

[0066] Both the attacking and defending drones can be described by the following single-integrator dynamic model, that is, the dynamic model of the drone swarm attack and defense differential game is expressed as follows:

[0067]

[0068] in, Indicates that at time t, the defender's drone P i Location, This indicates that at time t, the attacker's drone E j Location, and The defender drone P i and attacker drone E j The initial position; and The defender's drone P at time t i and attacker drone E j The control input satisfies bounded saturation and satisfies... and For the defender drone P i and attacker drone E j The speed magnitude satisfies The i-th defender drone P i With the j-th attacker drone E j The speed ratio α between ij , as shown in the following formula.

[0069]

[0070] To ensure the effectiveness of the pursuit, the i-th defender drone P i With the j-th attacker drone E j The speed ratio between them is greater than 1, i.e., α ij >1.

[0071] The following is a model of a game scenario with polygonal obstacles:

[0072] Let Ω be the region where the game takes place. Let V be the set of polygonal obstacles that exist in the scene, and let V be the set of vertices of all obstacles. obs Within the game's domain Ω, there exists a target domain Ω. goal And it is not connected to any polygonal obstacles. Target area Ω goal It is a convex, non-empty, closed set, which can be represented as in It is the target region description vector, and T is the transpose symbol. Represents a two-dimensional real vector. It is the target region description parameter, I goal This represents the target region index set, where x is the index that constitutes the target region Ω. goal The coordinates of point a m T x+b m ≥0 characterizes a closed half-plane, while the target region Ω goal It is a convex, non-empty region composed of these half-planes. The game region excluding the obstacle and target areas is denoted as the adversarial region Ω. play Its satisfaction Define the free zone as the area that both sides of the game (i.e., the attacker's drone and the defender's drone) can freely traverse, and the free zone satisfies Ω. free =Ω goal ∪Ω play .

[0073] The attacker drone aims to enter the target area, while the defender drone aims to intercept it before it does, thus countering the attacker and protecting the target area. The defender drone carries a radius of r. i The attack zone is defined when the distance between the attacker's drone and the defender's drone is less than the interception radius r. i Furthermore, since there are no obstacles between the two, the defender's drone can intercept the attacker's drone.

[0074] In another exemplary embodiment of this application, step 202 may include steps 301 to 305.

[0075] Step 301: Establish a one-to-one attack-defense subgame model. In the one-to-one attack-defense subgame model, there is only one attacker drone and one defender drone, denoted as E and P respectively. P Indicates the position of the defender drone P, x E Let X represent the position of the attacker's drone E, and r represent the interception radius of the defender's drone. Let the current state information be X = (x... P ,x E The models and objectives of both the attackers and defenders are consistent with those in step 201.

[0076] Step 302: Establish models for defensive victory state, defensive victory region, attack victory state, and attack victory region; the attack victory state is the state in which the attacker's drone achieves victory; the attack victory region is composed of several attack victory states. Details are as follows:

[0077] Victory state: If in a state X there exists a defender drone control input u pIf the defender drone P can win regardless of the attacker drone E's strategy, then this state is called a defense victory state.

[0078] Defensive Victory Zone: If a set of zones X contains states where every state is a defensive victory state, then this set is called a defensive victory zone. The largest defensive victory zone is denoted as M. p (u p ).

[0079] Attack victory state: If in a state X, there exists an attacker drone control input u E If the defender drone E can win regardless of the attacker drone P's strategy, then this state is called the attack victory state.

[0080] Attack Victory Region: If a set of regions X contains states where every state is an attack victory state, then this set is called an attack victory region. The largest attack victory region is denoted as M. E (u E ).

[0081] Step 303: Using differential game theory and reinforcement learning algorithms, establish a model of the winning state and winning region under the differential game theory and reinforcement learning algorithms; the model of the winning state and winning region under the differential game theory and reinforcement learning algorithms includes the defensive winning state under differential game theory, the defensive winning region under differential game theory, the defensive winning state under reinforcement learning, and the defensive winning region under reinforcement learning.

[0082] Defensive Victory State in Differential Game Theory: If a state X can be proven to be a defensive victory state under the differential game algorithm, then it is called a defensive victory state in differential game theory.

[0083] Defensive Victory Region in Differential Game Theory: If a region X can be proven to be a defensive victory region under a differential game algorithm, then it is called a defensive victory region in differential game theory, denoted as X.

[0084] Defensive victory state under reinforcement learning: If a state X can be proven to be a defensive victory state under a reinforcement learning algorithm, then it is called a defensive victory state under reinforcement learning.

[0085] Defensive Victory Region under Reinforcement Learning: If a region X can be proven to be a defensive victory region under a reinforcement learning algorithm, then it is called a defensive victory region under reinforcement learning.

[0086] Step 304: Establish forward reachable set and backward reachable set region models.

[0087] Let the state at time t be X.t (u P ,u E Forward reachable set Backward reachable set Then it satisfies:

[0088]

[0089] That is, the forward reachable set is defined as the set of states that can be reached given a policy u. p The set of states reachable from state X; the backward reachable set is defined as the set of states reachable from the attacker drone E regardless of the attacker drone E's strategy, with control input u. p It is certain that the state set can be reached.

[0090] For a defensive strategy u p If state X and region These are the defensive victory state and the defensive victory zone, respectively. Also a defensive strategy p The defensive victory zone below.

[0091] Step 305: Using a reinforcement learning algorithm, based on the forward reachable set and backward reachable set region model and the defense victory region under reinforcement learning, construct a new defense victory region under reinforcement learning; the new defense victory region under reinforcement learning is used to expand to obtain the extended defense victory region under reinforcement learning.

[0092] Specifically: Defensive Victory Zone Based on Differential Game Theory Using reinforcement learning algorithms and the forward reachable set and backward reachable set established in step 304, the winning scenarios in other regions are analyzed and expanded.

[0093] First, initialize the collection. It is an empty set, and then it starts from the defensive victory region under differential game theory. A limited number of samples were collected from other areas, and the collected samples were recorded as the first sampling set. Second sampling set Due to the defensive victory zone in differential game Points in the nearby area are more likely to be associated with Having the same properties, therefore, in the first sample set The first sample X1 is obtained by sampling, and the sampling follows the rules below:

[0094]

[0095] in,

[0096] Subsequently, a reinforcement learning algorithm was used to verify whether the first sample X1 belongs to the largest defensive victory zone. The verification method involves having the defender drone P use control input u p The attacker drone E uses a reinforcement learning algorithm to test whether it can complete the capture. If the test result is true (i.e., the defender drone P uses control input u), the attacker drone E will capture the target drone. p If the capture can be completed, then based on the forward reachable set and backward reachable set region model established in step 304, the first region is constructed using the forward and backward reachability algorithm on the basis of the original region. And expand the defensive victory zone under differential game theory. First region after construction The conditions and extension methods are as follows:

[0097]

[0098] Repeat sampling until all first sample sets have been traversed. The samples in the data are used to construct a new defensive victory region under reinforcement learning.

[0099] The algorithm is executed until the first sample set. The area in is empty.

[0100] In step 203, based on geometric analysis and using the minimum Euclidean distance, the differential game victory region of the one-to-one attack-defense subgame is constructed, and its corresponding defensive maneuver strategy is proposed. Specifically, step 203 includes: based on the minimum Euclidean distance algorithm and the aforementioned one-to-one attack-defense subgame problem, expanding the defensive victory region under the differential game to obtain the expanded defensive victory region under the differential game. This includes steps 401 to 403.

[0101] Step 401: Establish the Euclidean shortest distance path, the Euclidean shortest distance reachable area, the shortest path area map, and the convex target covering polygon model.

[0102] Euclidean shortest distance path: in the free region Ω free Given any two points x1 and x2, the Euclidean shortest path is denoted as P. ESP (x1, x2), the Euclidean shortest path must be within the free region Ω. free Within the range, and if there are multiple paths, specify any one of them. Let P be the Euclidean shortest path. ESP The length of (x1, x2) is d. ESP (x1, x2). Euclidean shortest path P ESP (x1, x2) can be represented as a series of ordered polygonal obstacle vertex regions plus the starting point x1 and the ending point x2, which can be used to find the Euclidean shortest path P.ESP (x1, x2) is a series of points starting from the starting point x1 and ending at the ending point x2.

[0103] For any two Euclidean shortest distance paths P 1 ESP ,P 2 ESP ,definition P represents 1 ESP It is P 2 ESP The prefix, i.e., P 2 ESP The sequence point set of the middle and front parts and P 1 ESP The vertices in the complete sequence set are identical in both order and vertices.

[0104] Euclidean shortest distance reachable region: This model describes the reachable region Ω. free Let be the set of all points in a given region x whose distance is not greater than l. l is a parameter describing the region reachable by the Euclidean shortest distance, and can be selected as any positive real number depending on the needs of subsequent algorithms.

[0105] Shortest path region map: Due to the Euclidean shortest distance path P ESP (x1, x2) can be described as a series of points starting from the starting point x1 and ending at the ending point x2. Therefore, there exists a series of points with x ∈ Ω. free Having the same second sequence of points, i.e., the Euclidean shortest path P other than the starting point. ESP The second path point in (x1,x2) allows us to define the free region Ω. free Divide the data into two parts, with the condition that x∈Ω free The region of identical second sequence points is called the shortest path region map for x, denoted as SPM(x).

[0106] Convex target covering polygon: For a polygon Its with respect to the point x∈Ω free A polygon is called a convex target-covering polygon if it satisfies the following four properties: 1) Polygon It is convex; 2) 3) 4) Convex target covering polygon The range of the x-direction of a certain point in the middle Defined as a set The angle in the equation is σ(x,y), where σ(x,y) refers to the unit vector from x to y.

[0107] Step 402: Establish a multiplayer onsite and close-to-goal (MOCG) strategy for multiplayer onsite and close-to-goal differential games.

[0108] Under this strategy, there exists a defensive victory zone in differential game theory, denoted as... It includes three cases: (1) based on the position x of the defender drone P P The location x of the attacker's drone E E The center of the pursuit circle of Apollo can be calculated by comparing the velocity ratio α. And the radius of the Apollo circle for pursuit Furthermore, a positive parameter δ > 0 is set for x. A With R as the center, A +δ is the radius. Construct a δ-pursuit Apollo circle, and the area inside the circle is the δ-pursuit Apollo circle area. If the constructed δ-pursuit Apollo circle area has no intersection with the obstacle and the target area, then this state is the defensive victory state under the differential game, and the area composed of this state is the defensive victory area under the differential game. Under this condition, the position of the defender drone P is very close to the attacker drone E and there are no obstacles nearby; (2) the defender drone P can see the entire target area, and the distance to the defending target is less than the distance to the attacking target. The distance to the defending target is the distance between the defender drone P and the target area, and the distance to the attacking target is the distance between the attacker drone E and the target area. That is, the defender drone P can see the entire target area and is closer to the target area than the attacker drone E; (3) the condition (2) cannot be satisfied, but the defender drone P can move to a certain position so that condition (2) is satisfied. For any state There is a control input u in each case. p mo (X) enables the defender's drone P to win.

[0109] In case (1), let the game start time be t. Then, for any time τ, τ≥t, the control input of the defender's drone at time τ is... The intermediate parameter z(τ) is expressed as:

[0110] In case (2), let the start time of the game be t. Then, for any time τ, τ≥t, the control input of the defender's drone at time τ is... in, θ represents the position x of the defender's drone. P (τ) Covering polygons of convex targets The direction angle in the middle, The position of the defender's drone at time τ is x. P (τ) Covering polygons of convex targets The range in the central direction. (x) I (τ),x G (τ) is a convex optimization problem. The optimal solution is the state X of the attacker's drone and the defender's drone at time τ. ij (τ) Input convex optimization problem The output results of this convex optimization problem The specific form is as follows:

[0111]

[0112] Where x and y are 2D real vectors to be optimized, f ij (x,X ij f is the pursuit potential function, and its expression is: f ij (x,X ij )=||xx P ||2-α||xx E ||2-r.

[0113] In scenario (3), the defender drone will first reach the boundary vertex s of a polygonal obstacle that satisfies the conditions of scenario (2), i.e., the control input u of the defender drone. p mo (X) satisfies:

[0114] when At that time, the control input along Find the Euclidean shortest distance path between the defender drone and the boundary vertex s of the polygonal obstacle. for Length;

[0115] when When this happens, situation (3) becomes situation (2), and the strategy in situation (2) is executed.

[0116] Step 403: Based on the extended Euclidean shortest distance MOCG algorithm, construct a field approach target differential game defender drone algorithm (MOCG-ESP strategy) based on the extended Euclidean shortest distance multi-to-multi attack and defense differential game.

[0117] Control input u based on the defender drone p mo (X), the defensive victory region under the differential game corresponding to the MOCG strategy. The region beyond this was expanded, and a field-based approach target differential game defender drone algorithm (MOCG-ESP pursuit strategy) based on Euclidean shortest distance extension was designed. The controller drone control input under this MOCG-ESP pursuit strategy is u.p me (X). The MOCG-ESP pursuit strategy is as follows: for any state... All choose strategy u p me (X)=u p mo (X); For the remaining states, the selection strategy is u. p me (X)=u p eps (X)=σ(x p ,x), where σ(x) p x) refers to the position x of the defender's drone. p A unit vector pointing to x, where x refers to the vector at point P. ESP (x p ,x E In sequence x p The next point, P ESP (x p ,x E (x) represents the position of the defender's drone. p Location x of the attacker's drone E The shortest distance path in the European style.

[0118] The MOCG-ESPpursuit strategy can effectively expand the defensive winning region in differential games, resulting in an expanded defensive winning region in differential games.

[0119] Step 204 specifically includes: optimizing the reinforcement learning algorithm using a near-end strategy to expand the defensive victory region under reinforcement learning, thereby obtaining the expanded defensive victory region under reinforcement learning.

[0120] By employing an offline construction and online verification method, we introduce the Proximal Policy Optimization (PPO) reinforcement learning algorithm to extend the defensive victory region under reinforcement learning, obtaining the extended defensive victory region based on the MOCG-ESP policy. Let X be the extended defensive victory region under reinforcement learning. me This can be seen as a specific strategy proposed in step 205, using a forward and backward reachability algorithm, given at the geometric analysis level.

[0121] Offline construction:

[0122] First, the research areas for both the attacker and defender are initialized.

[0123]

[0124] Contains x E area

[0125] Among them, SPM(x p ) for the defender drone x p The corresponding shortest path region map.

[0126] Then, traverse region P. ESP (x P ,x E )\{x P For all states x, perform the following operation:

[0127]

[0128] The Cut_MOCG() function performs the following operation: Select For any state (x', y') in the region, x' is along P ESP (x P (x) moves. Then, all defensive victory regions in differential games that are not corresponding to MOCG strategies were removed. The state in.

[0129] Online verification:

[0130] When performing online verification, follow these steps: First, initialize:

[0131]

[0132] Subsequently, combined with the offline construction process Iterate through each group For all If If it is established, then

[0133] Ultimately, the defensive win region RL-PWR based on MOCG-ESP policy extended reinforcement learning is obtained, i.e.

[0134] Step 205 may include the following steps 501 to 503:

[0135] Step 501: Based on the many-to-many attack and defense problem model, the defense victory region under the extended differential game, and the defense victory region under the extended reinforcement learning, construct a subgraph of the drone swarm; the subgraph includes vertex features and edge features; the vertex features are the positions of the attacker drone and the defender drone; the edge features indicate that the current state of the attacker drone and the defender drone belongs to the defense victory region under the extended differential game or the union; the union is the union of the defense victory region under the extended differential game and the defense victory region under the extended reinforcement learning.

[0136] Step 502: Based on the subgraph, establish a binary integer programming problem model for drone swarm task matching.

[0137] Step 503: Solve the binary integer programming problem model to obtain the many-to-many defense decision programming result.

[0138] Using the defensive victory regions established in steps 203 and 204 under differential game theory and reinforcement learning, an algorithm is designed for the many-to-many attack-defense problem in step 201 to complete the allocation and interception control decisions of the defender drones against the attacker drones. Through this algorithm, the defending drone coalition can provide the optimal allocation method for the attacking drone coalition under the established binary integer programming problem, thereby achieving the maximum interception of the attacking drone coalition and maximizing the protection of the target area.

[0139] In the design of defender drone strategies in many-to-many attack-defense differential games, a key problem to be solved is the task allocation problem, namely, matching defender drones with attacker drones. To address this problem, two subgraphs are first constructed. in, and These are the sets of points formed by the attacker's drone and the defender's drone, ε d ε' and ε' represent the first and second edge sets, respectively. Let X... ij Let P represent the i-th defender drone. i and the j-th attacker drone E j The current state, then its corresponding edge e ij If the following relationship is satisfied: Then e ij ∈ε d ;if Then e ij ∈ε'.

[0140] After constructing the subgraph, solve the following binary integer programming problem based on the relationships between the edges of the subgraph:

[0141]

[0142] Among them, a ij =1 indicates that the i-th defender drone P i To counter the j-th attacker drone E j Otherwise a ij =0. |ε'| represents the number of edges in the second edge set ε'. Let M be the set of edges. * ∈ε' is the optimal solution to this binary integer programming problem, and M is taken as... d * ←M * ∩ε d .

[0143] The designed MOCG-ESP many-to-many attack and defense differential game strategy based on neural symbols revolves around the above subgraph and binary integer programming problem. Initialization is performed first:

[0144]

[0145] Then, a subgraph is constructed, and the optimal solution M of the binary integer programming problem is obtained. * ∈ε' and M d * ←M * ∩ε d .

[0146] If M d * The number of elements in the set is greater than the set of attacker drones M that can be defeated. defeat If the number of elements in the array is determined, then operation M is executed. defeat ←M d M run ←M,M run For the set of matches to be executed; subsequently, for the assigned defender drones, i.e., for those satisfying e ij ∈M run Defender drones, executing For unassigned defender drones, execute Every Δ steps, remove captured attacker drones and attacker drones that have reached the target area, and update the current status information of both attackers and defenders and the captured set ε. capture Repeat this algorithm until... The results of the many-to-many defense decision planning were obtained.

[0147] Step 201 describes the overall scenario, a many-to-many attack and defense problem with obstacles.

[0148] Step 202 investigates a game subproblem in Step 201, namely the one-to-one attack and defense subproblem. Steps 201-304 contain relevant definitions and models, while Step 205 is based on Steps 201-304 and utilizes a differential game victory domain. However, it doesn't explain how this differential game victory domain was obtained; instead, it uses Step 204 to construct the reinforcement learning victory domain in Step 203. Therefore, Step 205 can be considered a general method for constructing victory domains under reinforcement learning.

[0149] Step 203 investigates how to obtain the defensive victory region under the differential game of the one-to-one attack and defense sub-problem in Step 202, and how to construct the defender's drone's defensive strategy. This step constructs the victory region under this differential game, which was not mentioned in Step 205. Step 402 did not utilize the strategy and region based on the Euclidean shortest distance path; Step 403 incorporates the Euclidean shortest distance path. Through this step, under the scenario satisfying Steps 402 and 403, the defender's drone can effectively intercept the attacker's drone in the one-to-one game sub-problem described in Step 202.

[0150] Since the defensive victory region established in step 203 for the differential game of the one-to-one attack-defense subproblem is limited, a reinforcement learning algorithm is proposed to expand the defensive victory region. The algorithm utilizes a general method for constructing the victory region under reinforcement learning and the defensive victory region under the differential game. This algorithm can expand the region in step 203 where it is uncertain whether the defender drone can successfully intercept, thus obtaining the defensive victory region under reinforcement learning.

[0151] The defensive victory regions established in steps 203 and 204 under the two differential strategies and the defensive victory regions under reinforcement learning are used to design an algorithm for the many-to-many attack and defense problem in step 201, complete the many-to-many defense decision, and improve the defense success rate.

[0152] Consider a many-to-many attack-defense differential game problem with a polygonal obstacle. There are 5 attacking drones (E1, E2, E3, E4, E5) and 5 defending drones (P1, P2, P3, P4, P5), and the target area Ω. goal The left rectangular region is shown below. The initial allocation results obtained using the MOCG-ESP differential game algorithm, the proposed neurosymmetric MOCG-ESP obstacle environment many-to-many attack-defense game method, and the final attack-defense confrontation results are respectively as follows: Figure 4 , Figure 5 , Figure 6 and Figure 7 As shown. By Figures 4-7As can be seen, the designed MOCG-ESP many-to-many attack-defense differential game strategy based on neural symbols can improve the defensive effect, increase the interception success rate, and better protect the important target area Ω. goal .

[0153] This application also provides an application scenario in which the aforementioned multi-to-multi attack-defense game method under obstacle conditions is applied. Specifically, the multi-to-multi attack-defense game method under obstacle conditions provided in this embodiment can be applied to a drone attack-defense game scenario. The drone attack-defense game scenario includes a request generation stage, a multi-to-multi attack-defense differential game processing link, and a drone control stage. The decision request to be processed enters the multi-to-multi attack-defense differential game processing link from the request generation stage, obtains the corresponding multi-to-multi defense decision planning result through human-machine collaboration, and then enters the downstream drone control stage. The multi-to-multi attack-defense game method under obstacle conditions provided in this embodiment belongs to the multi-to-multi attack-defense differential game processing link. Specifically, in the process of handling the many-to-many attack and defense differential game in response to the decision requests to be processed, a many-to-many attack and defense problem model of the drone swarm can be established, and a one-to-one attack and defense subgame problem can be established. Based on the one-to-one attack and defense subgame problem, the defense victory region under the differential game is expanded to obtain the expanded defense victory region under the differential game. The defense victory region under reinforcement learning is also expanded to obtain the expanded defense victory region under reinforcement learning. Based on the many-to-many attack and defense problem model, the expanded defense victory region under the differential game, and the expanded defense victory region under reinforcement learning, many-to-many defense decisions are made to obtain the many-to-many defense decision planning results.

[0154] Based on the same inventive concept, this application also provides a device for implementing the aforementioned method for many-to-many attack and defense game in an obstacle environment. The solution provided by this device is similar to the implementation scheme described in the above method. Therefore, the specific limitations of one or more embodiments of the device for many-to-many attack and defense game in an obstacle environment provided below can be found in the limitations of the method for many-to-many attack and defense game in an obstacle environment described above, and will not be repeated here.

[0155] In one exemplary embodiment, such as Figure 8 As shown, a multi-to-multi attack and defense game device for obstacle environments is provided, comprising the following modules:

[0156] The multi-to-multi attack and defense problem model building module T1 is used to build a multi-to-multi attack and defense problem model of a drone swarm; the drone swarm includes several attacker drones and several defender drones; the multi-to-multi attack and defense problem model includes a dynamic model of the attack and defense differential game between the two sides and a game scenario model under polygonal obstacles;

[0157] The module T2 for establishing a one-to-one attack and defense subgame problem is used to establish a one-to-one attack and defense subgame problem. The one-to-one attack and defense subgame problem includes a one-to-one attack and defense subgame model, a defense victory region under differential game theory, and a defense victory region under reinforcement learning. The defense victory region consists of several defense victory states. The defense victory state is the state in which the defender drone wins.

[0158] The defensive victory region extension module T3 under differential game is used to extend the defensive victory region under differential game based on the one-to-one attack and defense subgame problem, so as to obtain the extended defensive victory region under differential game.

[0159] The reinforcement learning-based defensive victory region extension module T4 is used to extend the defensive victory region under reinforcement learning to obtain an extended reinforcement learning-based defensive victory region.

[0160] The many-to-many defense decision module T5 is used to make many-to-many defense decisions based on the many-to-many attack and defense problem model, the defense victory region under the extended differential game, and the defense victory region under the extended reinforcement learning, and obtain the many-to-many defense decision planning results.

[0161] As an optional implementation, the one-to-one attack-defense subgame problem establishment module T2 is used for:

[0162] Establish a one-to-one attack and defense subgame model;

[0163] Establish models for defensive victory state, defensive victory zone, attack victory state, and attack victory zone; the attack victory state is the state in which the attacker's drone achieves victory; the attack victory zone is composed of several attack victory states.

[0164] Using differential game theory and reinforcement learning algorithms, a model of the winning state and winning region under the differential game theory and reinforcement learning algorithms is established; the model of the winning state and winning region under the differential game theory and reinforcement learning algorithms includes the defensive winning state under differential game theory, the defensive winning region under differential game theory, the defensive winning state under reinforcement learning, and the defensive winning region under reinforcement learning.

[0165] Establish forward reachable set and backward reachable set region models;

[0166] Using reinforcement learning algorithms, based on the forward reachable set and backward reachable set region model and the defense victory region under reinforcement learning, a new defense victory region under reinforcement learning is constructed; the new defense victory region under reinforcement learning is used to expand to obtain the extended defense victory region under reinforcement learning.

[0167] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 9 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs in the non-volatile storage media to run. The database stores decision-making data for UAV many-to-many attack-defense differential game. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When executed by the processor, the computer program implements a many-to-many attack-defense game method under obstacle conditions.

[0168] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0169] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0170] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0171] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0172] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0173] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0174] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0175] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0176] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for many-to-many attack and defense game in an obstacle environment, characterized in that, The multi-to-multi attack and defense game method under the obstacle environment includes: A multi-to-multi attack and defense problem model for a drone swarm is established; the drone swarm includes several attacker drones and several defender drones; the multi-to-multi attack and defense problem model includes a dynamic model of the attack and defense differential game between the two sides and a game scenario model under polygonal obstacles. A one-to-one attack-defense subgame problem is established; the one-to-one attack-defense subgame problem includes a one-to-one attack-defense subgame model, a defensive victory region under differential game theory, and a defensive victory region under reinforcement learning; the defensive victory region consists of several defensive victory states; the defensive victory state is the state in which the defender's drone achieves victory; The study establishes a one-to-one attack-defense subgame problem, specifically including: establishing a one-to-one attack-defense subgame model; establishing models for defensive victory states, defensive victory regions, attack victory states, and attack victory regions; using differential game theory and reinforcement learning algorithms, establishing victory state and victory region models under these algorithms; establishing forward reachable set and backward reachable set region models; using reinforcement learning algorithms, based on the forward reachable set and backward reachable set region models and the defensive victory regions under reinforcement learning, constructing new defensive victory regions under reinforcement learning; and using these new defensive victory regions under reinforcement learning for expansion to obtain extended defensive victory regions under reinforcement learning. Based on the aforementioned one-to-one attack-defense subgame problem, the defensive victory region under the differential game is expanded to obtain the expanded defensive victory region under the differential game. Specifically, this includes: establishing Euclidean shortest distance paths, Euclidean shortest distance reachable regions, shortest path region maps, and convex target coverage polygon models; establishing the MOCG strategy; and constructing a field-approaching target differential game defender UAV algorithm based on the Euclidean shortest distance extended MOCG algorithm, namely the MOCG-ESP strategy, for any state... All chose the strategy as For the remaining states, the selection strategy is as follows: ,in, This represents the defensive victory zone in differential game theory. For the control input of the defender's drone, This refers to the position of the defender's drone. point to unit vector, It refers to In sequence The next point, Position of the defender's drone Location of the attacker's drone The European shortest distance path; The defensive victory region under reinforcement learning is expanded to obtain the extended defensive victory region under reinforcement learning. Based on the many-to-many attack and defense problem model, the defensive victory region under the extended differential game and the defensive victory region under the extended reinforcement learning, many-to-many defense decision is made, and the many-to-many defense decision planning result is obtained.

2. The method for many-to-many attack and defense game in an obstacle environment according to claim 1, characterized in that, The attack victory state is the state in which the attacker's drone achieves victory; the attack victory area is composed of several attack victory states. The victory state and victory region model under the differential game and reinforcement learning algorithm includes the defensive victory state under differential game, the defensive victory region under differential game, the defensive victory state under reinforcement learning, and the defensive victory region under reinforcement learning.

3. The method for many-to-many attack and defense game in an obstacle environment according to claim 1, characterized in that, The defensive victory region under reinforcement learning is expanded to obtain the extended defensive victory region under reinforcement learning, which specifically includes: By optimizing the reinforcement learning algorithm using a near-end strategy, the defensive victory region under reinforcement learning is expanded to obtain an expanded defensive victory region under reinforcement learning.

4. The method for many-to-many attack and defense game in an obstacle environment according to claim 1, characterized in that, Based on a many-to-many attack and defense problem model, the defensive victory region under the extended differential game theory and the defensive victory region under the extended reinforcement learning, many-to-many defense decision-making is performed, and the many-to-many defense decision planning results are obtained, specifically including: Based on a many-to-many attack and defense problem model, a defense victory region under extended differential game theory, and a defense victory region under extended reinforcement learning, a subgraph of a drone swarm is constructed. The subgraph includes vertex features and edge features. The vertex features represent the positions of the attacker and defender drones. The edge features indicate that the current states of the attacker and defender drones belong to the defense victory region under extended differential game theory or their union. The union is the union of the defense victory region under extended differential game theory and the defense victory region under extended reinforcement learning. Based on the subgraph, a binary integer programming problem model for unmanned aerial vehicle (UAV) swarm task matching is established. Solving the binary integer programming problem model yields the many-to-many defense decision programming results.

5. A device for playing a many-to-many attack-defense game under obstacles, based on the method for playing many-to-many attack-defense games under obstacles according to any one of claims 1-4, characterized in that, The multi-to-multi attack and defense game device under obstacle conditions includes: A module for establishing a many-to-many attack and defense problem model is used to establish a many-to-many attack and defense problem model for a drone swarm; the drone swarm includes several attacker drones and several defender drones; the many-to-many attack and defense problem model includes a dynamic model of the attack and defense differential game between the two sides and a game scenario model under polygonal obstacles. A module for establishing one-to-one attack and defense subgame problems is used to establish one-to-one attack and defense subgame problems. The one-to-one attack and defense subgame problems include one-to-one attack and defense subgame models, defensive victory regions under differential games, and defensive victory regions under reinforcement learning. The defensive victory regions are composed of several defensive victory states. The defensive victory states are the states in which the defender's drone achieves victory. The defensive victory region expansion module under differential game is used to expand the defensive victory region under differential game based on the one-to-one attack and defense subgame problem, so as to obtain the expanded defensive victory region under differential game. The reinforcement learning-based defensive victory region extension module is used to extend the defensive victory region under reinforcement learning to obtain an extended reinforcement learning-based defensive victory region. The many-to-many defense decision module is used to make many-to-many defense decisions based on the many-to-many attack and defense problem model, the defense victory region under the extended differential game, and the defense victory region under the extended reinforcement learning, and obtain the many-to-many defense decision planning results.

6. The multi-to-multi attack and defense game device under obstacle conditions according to claim 5, characterized in that, The module for establishing the one-to-one attack and defense subgame problem is used for: Establish a one-to-one attack and defense subgame model; Establish models for defensive victory state, defensive victory zone, attack victory state, and attack victory zone; the attack victory state is the state in which the attacker's drone achieves victory; the attack victory zone is composed of several attack victory states. Using differential game theory and reinforcement learning algorithms, a model of the winning state and winning region under the differential game theory and reinforcement learning algorithms is established; the model of the winning state and winning region under the differential game theory and reinforcement learning algorithms includes the defensive winning state under differential game theory, the defensive winning region under differential game theory, the defensive winning state under reinforcement learning, and the defensive winning region under reinforcement learning. Establish forward reachable set and backward reachable set region models; Using reinforcement learning algorithms, based on the forward reachable set and backward reachable set region model and the defense victory region under reinforcement learning, a new defense victory region under reinforcement learning is constructed; the new defense victory region under reinforcement learning is used to expand to obtain the extended defense victory region under reinforcement learning.

7. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the many-to-many attack and defense game method in an obstacle environment as described in any one of claims 1-4.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the multi-to-multi attack and defense game method in an obstacle environment as described in any one of claims 1-4.

9. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the multi-to-multi attack and defense game method in an obstacle environment as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Intelligent defense decision-making method and device based on reinforcement learning and attack and defense games

    CN110166428A

  • Unmanned aerial vehicle confrontation game training control method based on reinforcement learning

    CN113282100A