Many-to-many attack and defense game method and device in obstacle environment
By combining differential game and deep reinforcement learning methods, the defensive victory area is expanded, and the problem of poor interception effect in the offensive and defense problems of drone clusters in obstacle environments is solved, and efficient interception and target protection are achieved in the polygonal obstacle environment.
Patent Information
- Application Number
- CN202510180127.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-19
AI Technical Summary
In the attack and defense problems of drone clusters with obstacles, the prior art is difficult to effectively intercept attack drones and protect target areas, and learning-based methods have problems of interpretability and performance guarantee.
By integrating differential game and deep reinforcement learning methods, expand the defensive victory area, complete the matching of defensive drones to attack drones, and select corresponding defensive maneuvering strategies to achieve interception and counter-attack attack drones.
In the environment of polygonal obstacles, improve the ability of defensive drones to intercept attack drones, maximize the protection of target areas, and improve the defense success rate.
Smart Images

Figure CN120046494A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of unmanned aerial vehicle control, and in particular, to a multi - to - multi attack - defense game method and device in an obstacle environment. Background Art
[0002] The differential game of unmanned aerial vehicle (UAV) swarm attack - defense considers an adversarial scenario where a group of evading UAVs (or attacking UAVs) aim to enter an area protected by a group of tracking UAVs (or defending UAVs). Compared with the pursuit - evasion game, the attack - defense differential game is more challenging and has a wider range of application scenarios, such as attacking some key facilities or high - value targets. Since the goal of the attacking UAVs is not only to avoid being captured by the defending UAVs but also to approach some key targets, while the defending UAVs not only need to intercept the attacking UAVs but also pay attention to the protection of the target area, this problem is more challenging and practical. In a barrier - free environment, many researchers have studied the attack - defense differential game problem, and the differences lie in aspects such as the game space, capture conditions, information structure, winning teams, speed ratio, dynamics of both attacking and defending sides, and target areas. However, there are not many current studies on the attack - defense differential game problem in an environment with obstacles.
[0003] In recent years, learning - based methods have often been used to solve the attack - defense differential game. Related technologies have developed an imitation learning framework using graph neural networks and a centralized expert algorithm. To defend against a faster evader, related technologies have proposed a deep reinforcement learning (RL) method to train the tracking strategy by fixing an analytical strategy for the evader. However, learning - based methods may encounter problems of interpretability and lack of performance guarantees, which are crucial in adversarial scenarios. In addition, for the analysis of attack - defense problems, in an environment with obstacles, relatively conservative analysis results are often obtained. In the UAV swarm attack - defense problem in a polygonal obstacle environment in the anti - UAV game scenario, learning - based methods have poor interception effects. Summary of the Invention
[0004] The purpose of the present application is to provide a multi - to - multi attack - defense game method and device in an obstacle environment, so that in the UAV swarm attack - defense problem with polygonal obstacles, the defending UAVs can intercept as many attacking UAVs as possible, maximize the protection of the target area, and improve the defense success rate.
[0005] To achieve the above purpose, the present application provides the following solutions:
[0006] In the first aspect, the present application provides a multi - to - multi attack - defense game method in an obstacle environment, including:
[0007] Establish a multi - to - multi attack - defense problem model for an unmanned aerial vehicle (UAV) cluster; the UAV cluster includes a number of attacker UAVs and a number of defender UAVs; the multi - to - multi attack - defense problem model includes the dynamics models of both sides of the attack - defense differential game and the game scenario model under polygonal obstacles.
[0008] Establish a one - to - one attack - defense sub - game problem; the one - to - one attack - defense sub - game problem includes a one - to - one attack - defense sub - game model, the defense victory region under differential game, and the defense victory region under reinforcement learning; the defense victory region consists of a number of defense victory states; the defense victory state is the state where the defender UAV wins.
[0009] Based on the one - to - one attack - defense sub - game problem, expand the defense victory region under differential game to obtain the expanded defense victory region under differential game.
[0010] Expand the defense victory region under reinforcement learning to obtain the expanded defense victory region under reinforcement learning.
[0011] Based on the multi - to - multi attack - defense problem model, the expanded defense victory region under differential game, and the expanded defense victory region under reinforcement learning, make multi - to - multi defense decisions to obtain the multi - to - multi defense decision - making planning result.
[0012] Optionally, establishing a one - to - one attack - defense sub - game problem specifically includes:
[0013] Establish a one - to - one attack - defense sub - game model.
[0014] Establish models for defense victory states, defense victory regions, attack victory states, and attack victory regions; the attack victory state is the state where the attacker UAV wins; the attack victory region consists of a number of attack victory states.
[0015] Using differential game and reinforcement learning algorithms, establish a victory state and victory region model under differential game and reinforcement learning algorithms; the victory state and victory region model under differential game and reinforcement learning algorithms includes the defense victory state under differential game, the defense victory region under differential game, the defense victory state under reinforcement learning, and the defense victory region under reinforcement learning.
[0016] Establish forward reachable set and backward reachable set region models.
[0017] Using the reinforcement learning algorithm, based on the forward reachable set and backward reachable set region models and the defense victory region under reinforcement learning, construct a new defense victory region under reinforcement learning; the new defense victory region under reinforcement learning is used to expand to obtain the expanded defense victory region under reinforcement learning.
[0018] Optionally, based on the one-on-one attack-defense sub-game problem, the defense victory region under differential game is extended to obtain the extended defense victory region under differential game, specifically including:
[0019] Based on the minimum Euclidean distance algorithm and the one-on-one attack-defense sub-game problem, the defense victory region under differential game is extended to obtain the extended defense victory region under differential game.
[0020] Optionally, the defense victory region under reinforcement learning is extended to obtain the extended defense victory region under reinforcement learning, specifically including:
[0021] Using the proximal policy optimization reinforcement learning algorithm, the defense victory region under reinforcement learning is extended to obtain the extended defense victory region under reinforcement learning.
[0022] Optionally, based on the multi-on-multi attack-defense problem model, the extended defense victory region under differential game, and the extended defense victory region under reinforcement learning, multi-on-multi defense decisions are made to obtain the multi-on-multi defense decision planning result, specifically including:
[0023] Based on the multi-on-multi attack-defense problem model, the extended defense victory region under differential game, and the extended defense victory region under reinforcement learning, a sub-graph of the UAV cluster is constructed; the sub-graph includes point features and edge features; the point features are the positions of the attacker UAVs and the defender UAVs; the edge features indicate that the current states of the attacker UAVs and the defender UAVs belong to the extended defense victory region under differential game or the union set; the union set is the union of the extended defense victory region under differential game and the extended defense victory region under reinforcement learning;
[0024] Based on the sub-graph, a binary integer programming problem model for UAV cluster task matching is established;
[0025] The binary integer programming problem model is solved to obtain the multi-on-multi defense decision planning result.
[0026] In a second aspect, the present application provides a multi-on-multi attack-defense game device in an obstacle environment, including:
[0027] A multi-on-multi attack-defense problem model establishment module, configured to establish a multi-on-multi attack-defense problem model of a UAV cluster; the UAV cluster includes a plurality of attacker UAVs and a plurality of defender UAVs; the multi-on-multi attack-defense problem model includes an attack-defense differential game bilateral dynamics model and a game scenario model under a polygonal obstacle;
[0028] One-on-one attack and defense sub-game problem establishment module, used to establish a one-on-one attack and defense sub-game problem; the one-on-one attack and defense sub-game problem includes a one-on-one attack and defense sub-game model, a defense victory region under differential game, and a defense victory region under reinforcement learning; the defense victory region is composed of several defense victory states; the defense victory state is the state where the defender UAV wins.
[0029] Defense victory region expansion module under differential game, used to expand the defense victory region under differential game based on the one-on-one attack and defense sub-game problem to obtain the expanded defense victory region under differential game.
[0030] Defense victory region expansion module under reinforcement learning, used to expand the defense victory region under reinforcement learning to obtain the expanded defense victory region under reinforcement learning.
[0031] Multi-to-multi defense decision-making module, used to make multi-to-multi defense decisions based on the multi-to-multi attack and defense problem model, the expanded defense victory region under differential game, and the expanded defense victory region under reinforcement learning to obtain the multi-to-multi defense decision-making planning result.
[0032] Optionally, the one-on-one attack and defense sub-game problem establishment module is used to:
[0033] Establish a one-on-one attack and defense sub-game model.
[0034] Establish defense victory state, defense victory region, attack victory state, and attack victory region models; the attack victory state is the state where the attacker UAV wins; the attack victory region is composed of several attack victory states.
[0035] Use differential game and reinforcement learning algorithms to establish victory state and victory region models under differential game and reinforcement learning algorithms; the victory state and victory region models under differential game and reinforcement learning algorithms include defense victory states under differential game, defense victory regions under differential game, defense victory states under reinforcement learning, and defense victory regions under reinforcement learning.
[0036] Establish forward reachable set and backward reachable set region models.
[0037] Use the reinforcement learning algorithm to construct a new defense victory region under reinforcement learning based on the forward reachable set and backward reachable set region models and the defense victory region under reinforcement learning; the new defense victory region under reinforcement learning is used to be expanded to obtain the expanded defense victory region under reinforcement learning.
[0038] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the above-mentioned multi-to-multi attack and defense game method in an obstacle environment.
[0039] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the above-mentioned multi-to-multi attack and defense game method in an obstacle environment.
[0040] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the above-mentioned multi-to-multi attack and defense game method in an obstacle environment.
[0041] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application:
[0042] The present application provides a multi-to-multi attack and defense game method and device in an obstacle environment. By integrating methods based on differential game (DG) and deep reinforcement learning (RL), the defensive victory area is expanded, the matching of defensive drones to attacking drones is completed, and a corresponding defensive maneuver strategy is further selected to complete the interception and countermeasure against attacking drones. Using this method, in the problem of attacking and defending an unmanned aerial vehicle (UAV) cluster with polygonal obstacles, the defensive drones can intercept as many attacking drones as possible and protect the target area to the greatest extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0044] Figure 1 It is an application environment diagram of a multi-to-multi attack and defense game method in an obstacle environment in an embodiment of the present application;
[0045] Figure 2 It is a flowchart of a multi-to-multi attack and defense game method in an obstacle environment provided by an embodiment of the present application;
[0046] Figure 3 It is a schematic diagram of the specific process of a multi-to-multi attack and defense game method in an obstacle environment provided by an embodiment of the present application;
[0047] Figure 4Initial allocation diagram of the defender UAV algorithm for the multiplayer onsite and close-to-goal (MOCG) differential game (ESP) based on the extension of the Euclidean shortest distance provided by an embodiment of the present application;
[0048] Figure 5 MOCG-ESP differential game strategy result diagram provided by an embodiment of the present application;
[0049] Figure 6 Initial allocation diagram of the MOCG-ESP multiplayer offensive and defensive differential game strategy based on neuro-symbolic provided by an embodiment of the present application;
[0050] Figure 7 MOCG-ESP multiplayer offensive and defensive differential game strategy result diagram based on neuro-symbolic provided by an embodiment of the present application;
[0051] Figure 8 Schematic diagram of the functional modules of a multiplayer offensive and defensive game device in an obstacle environment provided by an embodiment of the present application;
[0052] Figure 9 Schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0053] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0054] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0055] The multiplayer offensive and defensive game method in an obstacle environment provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set separately, integrated on the server 104, placed on the cloud or other servers. The terminal 102 can send a decision request to be processed to the server 104. After receiving the decision request to be processed, the server 104 establishes a multi-to-multi attack and defense problem model for the UAV cluster, establishes a one-to-one attack and defense sub-game problem, based on the one-to-one attack and defense sub-game problem, expands the defensive victory area under differential game to obtain the expanded defensive victory area under differential game, expands the defensive victory area under reinforcement learning to obtain the expanded defensive victory area under reinforcement learning, and based on the multi-to-multi attack and defense problem model, the expanded defensive victory area under differential game and the expanded defensive victory area under reinforcement learning, makes a multi-to-multi defense decision to obtain the multi-to-multi defense decision planning result. The server 104 can feedback the obtained multi-to-multi defense decision planning result for the decision request to be processed to the terminal 102. In addition, in some embodiments, the multi-to-multi attack and defense game method in the obstacle environment can also be implemented independently by the server 104 or the terminal 102. For example, the terminal 102 can directly perform multi-to-multi attack and defense differential game decision processing for the decision request to be processed, or the server 104 can obtain the decision request to be processed from the data storage system and perform multi-to-multi attack and defense differential game decision processing for the decision request to be processed.
[0056] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0057] In an exemplary embodiment, such as Figure 2 and Figure 3 shown, a multi-to-multi attack and defense game method in an obstacle environment is provided. This method is executed by a computer device, and can specifically be executed independently by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, taking this method applied to Figure 1 the server 104 in as an example for illustration, it includes the following steps 201 to step 205.
[0058] Among them:
[0059] Step 201, establish a many-to-many attack and defense problem model for the UAV cluster; the UAV cluster includes several attacker UAVs and several defender UAVs; the many-to-many attack and defense problem model includes the dynamic models of both sides of the attack and defense differential game and the game scenario model under polygonal obstacles.
[0060] Step 202, establish a one-to-one attack and defense sub-game problem; the one-to-one attack and defense sub-game problem includes a one-to-one attack and defense sub-game model, the defense victory area under differential game, and the defense victory area under reinforcement learning; the defense victory area consists of several defense victory states; the defense victory state is the state where the defender UAV wins.
[0061] Step 203, based on the one-to-one attack and defense sub-game problem, expand the defense victory area under differential game to obtain the expanded defense victory area under differential game.
[0062] Step 204, expand the defense victory area under reinforcement learning to obtain the expanded defense victory area under reinforcement learning.
[0063] Step 205, based on the many-to-many attack and defense problem model, the expanded defense victory area under differential game, and the expanded defense victory area under reinforcement learning, make many-to-many defense decisions to obtain the many-to-many defense decision planning result.
[0064] Implement the above steps 201 to 205. By integrating the methods based on differential game (DG) and deep reinforcement learning (RL), the defense victory area is expanded. The matching of defender UAVs to attacker UAVs is completed through the neuro-symbolic algorithm, and the corresponding defensive maneuver strategy is further selected to complete the interception and countermeasure against the attacker UAVs. Using this method, in the attack and defense problem of UAV clusters with polygonal obstacles, the defender UAVs can intercept as many attacker UAVs as possible and protect the target area to the greatest extent.
[0065] In the attack and defense differential game problem of UAV clusters, considering that there are attacker UAV coalitions and defender UAV coalitions, where the attacker UAV coalition ε has N E attacker UAVs, denoted as representing the N E th attacker UAV, and the defender UAV coalition has N P defender UAVs, denoted as } is the N P th defender UAV.
[0066] The UAVs of both the attacking and defending sides can be described by the following single-integrator dynamic model, that is, the dynamic models of both sides of the differential game of UAV swarm attack and defense are as follows:
[0067]
[0068] where, represents the position of the defender UAV P i at time t, represents the position of the attacker UAV E j at time t, and are the initial positions of the defender UAV P i and the attacker UAV E j respectively; and are the control inputs of the defender UAV P i and the attacker UAV E j at time t, and their control inputs are bounded and saturated, satisfying and are the magnitudes of the velocities of the defender UAV P i and the attacker UAV E j respectively, satisfying The velocity ratio α i between the i-th defender UAV P j and the j-th attacker UAV E ij is expressed as follows.
[0069]
[0070] To ensure the pursuit effect, the velocity ratio between the i-th defender UAV P i and the j-th attacker UAV E j is greater than 1, that is, α ij > 1.
[0071] Next, a game scenario model under polygonal obstacles is established:
[0072] Let the area range of the game be Ω, be the set composed of the polygonal obstacles existing in the scenario, and let the vertex set of all obstacles be V obs . In the game area range Ω, there is a target area Ω goal and it is not connected to any polygonal obstacle. The target area Ω goal is a convex non-empty closed set, which can be expressed as where is the target area description vector, T is the transpose symbol, represents a two-dimensional real vector, is the target area description parameter, I goal represents the target area index set, and x is the coordinate point position that constitutes the target area Ω goal of m T x + b m ≥ 0 depicts a closed half-plane, and the target area Ω goal is a convex non-empty area composed of these half-planes. The game area except for the obstacles and the target area is denoted as the adversarial area Ω play , which satisfies Define the free area where both sides of the game (i.e., the attacker UAV and the defender UAV) can move freely. The free area satisfies Ω free = Ω goal ∪Ω play .
[0073] The purpose of the attacker UAV is to enter the range of the target area. The purpose of the defender UAV is to intercept the attacker UAV before it enters the range of the target area, so as to achieve the countermeasure against the attacking UAV and the protection of the target area. Among them, the defender UAV has an attack area with a radius of r i . When the distance between the attacker UAV and the defender UAV is less than the interception radius r i and there is no obstacle blocking between the two, the defender UAV can intercept the attacker UAV.
[0074] In another exemplary embodiment of the present application, step 202 may include the following steps 301 to step 305.
[0075] Step 301: Establish a one-on-one attack and defense sub-game model. In the one-on-one attack and defense sub-game model, there is only one attacker UAV and one defender UAV, which are denoted as E and P respectively. x P represents the position of the defender UAV P, and x E represents the position of the attacker UAV E, and r represents the interception radius of the defender UAV. Denote the current state information as X = (x P , x E ). The models and purposes of both the attacker and the defender are the same as those in step 201.
[0076] Step 302: Establish models for the defender's victory state, defender's victory area, attacker's victory state, and attacker's victory area; the attacker's victory state is the state where the attacker UAV wins; the attacker's victory area is composed of several attacker's victory states. Specifically as follows:
[0077] Defender's victory state: If in a state X, there exists a defender UAV control input u p, regardless of the strategy adopted by the attacker drone E, if the defender drone P can always achieve victory, then this state is called a defensive victory state.
[0078] Defensive victory region: If a set of regions X, where each state is a defensive victory state, then this set is called a defensive victory region, and the largest defensive victory region is denoted as M p (u p ).
[0079] Attack victory state: If in a state X, there exists an attacker drone control input u E , regardless of the strategy adopted by the attacker drone P, if the defender drone E can always achieve victory, then this state is called an attack victory state.
[0080] Attack victory region: If a set of regions X, where each state is an attack victory state, then this set is called an attack victory region. The largest attack victory region is denoted as M E (u E ).
[0081] Step 303: Use differential game and reinforcement learning algorithms to establish a victory state and victory region model under differential game and reinforcement learning algorithms; the victory state and victory region model under differential game and reinforcement learning algorithms includes the defensive victory state under differential game, the defensive victory region under differential game, the defensive victory state under reinforcement learning, and the defensive victory region under reinforcement learning.
[0082] Defensive victory state under differential game: If a state X can be proven to be a defensive victory state under the differential game algorithm, then it is called the defensive victory state under differential game.
[0083] Defensive victory region under differential game: If a region X can be proven to be a defensive victory region under the differential game algorithm, then it is called the defensive victory region under differential game, and the defensive victory region under differential game is denoted as
[0084] Defensive victory state under reinforcement learning: If a state X can be proven to be a defensive victory state under the reinforcement learning algorithm, then it is called the defensive victory state under reinforcement learning.
[0085] Defensive victory region under reinforcement learning: If a region X can be proven to be a defensive victory region under the reinforcement learning algorithm, then it is called the defensive victory region under reinforcement learning.
[0086] Step 304: Establish a forward reachable set and a backward reachable set region model.
[0087] Denote the state at time t as Xt (u P , u E ), the forward reachable set the backward reachable set Then it satisfies:
[0088]
[0089] That is, the forward reachable set is defined as the set of states that can be reached from state X under the given policy u p . The backward reachable set is defined as the set of states that the defender drone P can surely reach under the control input u p regardless of the policy of the attacker drone E
[0090] For a defense policy u p , if the state X and the region are the defense victory state and the defense victory region respectively, then is also the defense victory region under the defense policy u p .
[0091] Step 305: Using the reinforcement learning algorithm, based on the forward reachable set, the backward reachable set region model, and the defense victory region under reinforcement learning, construct a new defense victory region under reinforcement learning; the new defense victory region under reinforcement learning is used to be expanded to obtain the expanded defense victory region under reinforcement learning
[0092] Specifically: Based on the defense victory region under differential game Using the reinforcement learning algorithm and the forward reachable set and the backward reachable set established in step 304, analyze and expand the winning situations in other regions
[0093] First, initialize the set as an empty set, and then perform a finite number of sample samplings from the defense victory region under differential game and other regions, and record the sampled samples as the first sampling set the second sampling set Since the points in the region near the defense victory region under differential game are more likely to have the same properties as , therefore, sample in the first sampling set to obtain the first sample X 1 , and the sampling is carried out according to the following rules:
[0094]
[0095] Among them,
[0096] Subsequently, the first sample X is tested using a reinforcement learning algorithm 1 to determine whether it belongs to the largest defensive victory region The testing method is to have the defender drone P adopt the control input u p , and the attacker drone E adopt the reinforcement learning algorithm to check whether it can complete the capture. If the test result is true (i.e., the defender drone P can complete the capture with the control input u p ), then based on the forward reachable set and backward reachable set region models established in step 304, the first region is constructed using the forward and backward reachable algorithms on the original region basis and the defensive victory region under differential game is expanded After the construction, the first region satisfies the following conditions and expansion methods:
[0097]
[0098] Repeat sampling until all samples in the first sampling set are traversed, thereby constructing a new defensive victory region under reinforcement learning
[0099] Execute this algorithm until the region in the first sampling set is empty
[0100] In step 203, based on geometric analysis, using the minimum Euclidean distance, the differential game victory domain of the one-on-one attack and defense sub-game is constructed, and its corresponding defensive maneuver strategy is proposed. Then step 203 specifically includes: expanding the defensive victory region under differential game based on the minimum Euclidean distance algorithm and the one-on-one attack and defense sub-game problem to obtain the expanded defensive victory region under differential game. It includes the following steps 401 to step 403
[0101] Step 401: Establish the Euclidean shortest distance path, Euclidean shortest distance reachable region, shortest path region graph, and convex target coverage polygon model
[0102] Euclidean shortest distance path: In the free region Ω free arbitrarily given two points x 1 , x 2 , then its Euclidean shortest distance path is denoted as P ESP (x 1 , x 2 ). This Euclidean shortest distance path must be within the free region Ω free , and if there are multiple paths, any one of them is specified. Denote the length of the Euclidean shortest distance path P ESP (x 1 , x 2 ) as d ESP (x1 , x 2 ). The Euclidean shortest distance path P ESP (x 1 , x 2 ) can be represented as a series of vertex regions of ordered polygonal obstacles plus the starting point x 1 and the ending point x 2 . The Euclidean shortest distance path P ESP (x 1 , x 2 ) can be described as a series of points starting from the starting point x 1 and ending at the ending point x 2 .
[0103] For any two Euclidean shortest distance paths P 1 ESP , P 2 ESP , it is defined that denotes that P 1 ESP is a prefix of P 2 ESP , that is, the set of ordered points in the front part of P 2 ESP has the same vertices and order as those in the complete set of ordered points in P 1 ESP .
[0104] Euclidean shortest distance reachable region: The model describes the set of all points whose distance to a certain point x in the free region Ω free is not greater than l, denoted as l is a description parameter of the Euclidean shortest distance reachable region and can be selected as any positive real number according to the needs of subsequent algorithms.
[0105] Shortest path region graph: Since the Euclidean shortest distance path P ESP (x 1 , x 2 ) can be described as a series of points starting from the starting point x 1 and ending at the ending point x 2 , therefore, there exists a series of points with the same second sequence of points as x ∈ Ω free , that is, except for the starting point, the second path point in the Euclidean shortest distance path P ESP (x 1 , x 2 ). Thus, the free region Ω free can be divided. The region with the same second sequence of points as x ∈ Ω free is called the x shortest path region graph, denoted as SPM(x).
[0106] Convex target covering polygon: For a polygon with respect to a point x ∈ Ω free satisfying the following four properties, it is called a convex target covering polygon: 1) The polygon is convex; 2) 3) 4) The direction range of a certain point x in the convex target covering polygon is defined as the angle in the set where σ(x, y) refers to the unit vector from x to y.
[0107] Step 402: Establish a multiplayer onsite and close-to-goal (MOCG) strategy for the multi-player offensive and defensive differential game.
[0108] Under this strategy, there exists a defensive victory region in the differential game, denoted as which includes three cases: (1) According to the position x of the defender drone P P , the position x of the attacker drone E E and the speed ratio α, the center of the pursuit-evasion Apollonius circle and the radius of the pursuit-evasion Apollonius circle can be calculated. Further, a positive parameter δ > 0 is set, with x A as the center and R A + δ as the radius. Construct a δ-pursuit-evasion Apollonius circle, and the region inside the circle is the δ-pursuit-evasion Apollonius circle region. If there is no intersection between the constructed δ-pursuit-evasion Apollonius circle region and the obstacle and target regions, this state is a defensive victory state in the differential game, and the region composed of such states is the defensive victory region in the differential game. Under this condition, the position of the defender drone P is very close to the attacker drone E and there is no obstacle nearby; (2) The defender drone P can see the entire target region, and the defensive target distance is less than the attacking target distance. The defensive target distance is the distance between the defender drone P and the target region, and the attacking target distance is the distance between the attacker drone E and the target region, that is, the defender drone P can see the entire target region and is closer to the target region than the attacker drone E; (3) Case (2) cannot be satisfied, but the defender drone P can move to a certain position to make case (2) satisfied. For any state there exists a control input u p mo (X) such that the defender drone P can win.
[0109] In case (1), the start time of the game is denoted as t. Then, for any moment τ, τ ≥ t, the control input of the defender drone at the moment τ Among them, the intermediate parameter z(τ) is expressed as
[0110] In case (2), the game start time is denoted as t. Then, for any moment τ, τ≥t, the control input of the defender UAV at moment τ Among them, θ represents the direction angle of the position x P (τ) in the convex target coverage polygon ; is the position x of the defender UAV at moment τ P (τ) in the convex target coverage polygon ; (x I (τ), x G (τ)) is the optimal solution of the convex optimization problem , that is, the state X of the attacker UAV and the defender UAV at moment τ ij (τ) is input into the convex optimization problem to obtain the output result. The specific form of this convex optimization problem is as follows:
[0111]
[0112] Among them, x and y are two-dimensional real vectors to be optimized, and f ij (x, X ij ) is the pursuit-evasion potential function, and its expression is: f ij (x, X ij ) = ||x - x P || 2 - α||x - x E || 2 - r.
[0113] In case (3), the defender UAV will first reach the boundary vertex s of a polygon obstacle that can meet the conditions of case (2). That is, the control input u p mo (X) satisfies:
[0114] When , the control input is along , which is the Euclidean shortest distance path between the defender UAV and the boundary vertex s of the polygon obstacle, is the length of;
[0115] When , case (3) becomes case (2), and then the strategy in case (2) is executed.
[0116] Step 403: Based on the Euclidean shortest distance extended MOCG algorithm, construct the on-site approaching target differential game defender UAV algorithm (MOCG-ESP strategy) for the multi-to-multi attack and defense differential game based on the Euclidean shortest distance extension.
[0117] Based on the control input u of the defender UAV p mo (X), the area outside the defense victory area under the differential game corresponding to the MOCG strategy is extended, and the on-site approaching target differential game defender UAV algorithm (MOCG-ESP pursuit strategy) for the multi-to-multi attack and defense differential game based on the Euclidean shortest distance extension is designed. The control input of the defender UAV under this MOCG-ESP pursuit strategy is u (X). This MOCG-ESP pursuit strategy is as follows: for any state p me , select the strategy u to be u p me (X) = u p mo (X); for the remaining states, select the strategy u p me (X) = u p eps (X) = σ(x p , x), where σ(x p , x) refers to the unit vector pointing from the position x of the defender UAV p to x, and x refers to the next point in P ESP (x p , x E ) according to the sequence x p . P ESP (x p , x E ) is the Euclidean shortest distance path between the position x of the defender UAV p and the position x of the attacker UAV E .
[0118] Through the MOCG-ESP pursuit strategy, the defense victory area under the differential game can be effectively extended to obtain the extended defense victory area under the differential game.
[0119] Step 204 specifically includes: using the proximal policy optimization reinforcement learning algorithm to extend the defense victory area under reinforcement learning to obtain the extended defense victory area under extended reinforcement learning.
[0120] By means of offline construction and online verification, and introducing the Proximal Policy Optimization (PPO) reinforcement learning algorithm, we expand the defensive victory region under reinforcement learning to obtain the defensive victory region under extended reinforcement learning based on the MOCG-ESP strategy, and denote the defensive victory region under extended reinforcement learning as X me . It can be regarded as a specific strategy given at the geometric analysis level by using the forward and backward reachability algorithm proposed in step 205
[0121] Offline construction:
[0122] First, initialize the areas to be studied for both the attacker and the defender
[0123]
[0124] The area containing x E of the region
[0125] where SPM(x p ) is the shortest path area graph corresponding to the defender's drone x p .
[0126] Subsequently, traverse the region P ESP (x P , x E )\{x P} and perform the following operations on all states x therein:
[0127]
[0128] where the following operations are performed in the Cut_MOCG() function: Select any state (x', y') in the region, and x' moves along P (x ESP (x P , x), and then cut off all states that are not in the defensive victory region under the differential game corresponding to the MOCG strategy in .
[0129] Online verification:
[0130] During online verification, proceed as follows. First, initialize:
[0131]
[0132] Subsequently, combine the formed in offline construction and traverse each group For all If there is holds, then
[0133] Finally, the defensive victory region RL-PWR based on the extended reinforcement learning under the MOCG-ESP strategy is obtained, that is
[0134] Step 205 may include the following steps 501 to 503:
[0135] Step 501: Based on the many-to-many attack-defense problem model, the defensive victory region under the extended differential game, and the defensive victory region under the extended reinforcement learning, construct a subgraph of the UAV cluster; the subgraph includes point features and edge features; the point features are the positions of the attacker UAVs and the defender UAVs; the edge features indicate that the current states of the attacker UAVs and the defender UAVs belong to the defensive victory region or the union set under the extended differential game; the union set is the union of the defensive victory region under the extended differential game and the defensive victory region under the extended reinforcement learning.
[0136] Step 502: Based on the subgraph, establish a binary integer programming problem model for UAV cluster task matching.
[0137] Step 503: Solve the binary integer programming problem model to obtain the many-to-many defense decision planning result.
[0138] Using the defensive victory region under the differential game and the defensive victory region under the reinforcement learning established in Steps 203 and 204, design an algorithm for the many-to-many attack-defense problem in Step 201 to complete the allocation and interception control decision of the defender UAVs for the attacker UAVs. Through this algorithm, the defensive UAV coalition can give the optimal allocation method for the attacking UAV coalition under the established binary integer programming problem, so as to achieve the maximum number of interceptions of the attacking UAV coalition and the protection of the target area as much as possible.
[0139] In the design of the defender UAV strategy in the many-to-many attack-defense differential game, an important problem to be solved is the task allocation problem, that is, to solve the problem of matching the defender UAVs with the attacker UAVs. To solve this problem, first construct two subgraphs Among them, and are respectively the point sets composed of the attacker UAVs and the defender UAVs, and ε d , ε' are the first edge set and the second edge set respectively. Let X ij represent the current states of the i-th defender UAV P i and the j-th attacker UAV E j , then the corresponding edge e ij satisfies the following relationship: If then e ij∈ε d ; If then e ij ∈ε'.
[0140] After constructing the sub - graph, according to the relationship of the sub - graph edges, solve the following binary integer programming problem:
[0141]
[0142] where, a ij = 1 means that the i - th defender UAV P i goes to counter the j - th attacker UAV E j , otherwise a ij = 0. |ε'| represents the number of edges in the second edge set ε'. Denote M * ∈ε' as the optimal solution to this binary integer programming problem and take M d * ←M * ∩ε d .
[0143] The designed neuro - symbolic - based MOCG - ESP multi - to - multi attack - defense differential game strategy is carried out around the above sub - graph and binary integer programming problem. First, initialize:
[0144]
[0145] Subsequently, construct the sub - graph and obtain the optimal solution M * ∈ε' of the binary integer programming problem and M d * ←M * ∩ε d .
[0146] If the number of elements in M d * is greater than the number of elements in the set M of attacker UAVs that can be defeated, then perform the operation M defeat ←M defeat ←M d , M run ←M, M run is the matching set to be executed; Subsequently, for the assigned defender UAVs, that is, for the defender UAVs satisfying e ij ∈M run , execute For the unassigned defender UAVs, execute Every Δ steps, eliminate the attacker UAVs that have been captured and the attacker UAVs that have reached the target area, and update the current attack - defense state information and the captured set ε capture . Repeat this algorithm until Obtain the many-to-many defense decision-making planning results.
[0147] Step 201 describes the overall scenario, a many-to-many attack and defense problem with obstacles.
[0148] Step 202 studies a game sub-problem in Step 20201, namely the one-to-one attack and defense sub-problem. Among them, Steps 201 - 304 are related definitions and models. Step 205 is based on Steps 201 - 304 and uses the differential game victory domain, but does not mention how to obtain this differential game victory domain. Instead, it constructs the reinforcement learning victory domain in Step 203 using Step 204. So Step 205 can be regarded as a general method for constructing the victory domain under reinforcement learning.
[0149] Step 203 studies how to obtain the defensive victory area under the differential game of the one-to-one attack and defense sub-problem in Step 202 and how to construct the defense strategy of the defender UAV. In this step, the victory area under this differential game not mentioned in Step 205 is constructed. Step 402 does not use the strategy and area under the Euclidean shortest distance path, and Step 403 adds the Euclidean shortest distance path. Through this step, in the scenarios satisfying Step 402 and Step 403, in the one-to-one game sub-problem described in Step 202, it can be ensured that the defender UAV effectively intercepts the attacker UAV.
[0150] Since the range of the defensive victory area established in the previous Step 203 for the one-to-one attack and defense sub-problem under the differential game is limited, a reinforcement learning algorithm is proposed to expand the defensive victory area. When constructing the algorithm, a general method for constructing the victory domain under reinforcement learning and the defensive victory area under the differential game are used. Through this algorithm, the area where it cannot be determined whether the defender UAV can successfully intercept in Step 203 can be expanded to obtain the defensive victory area under reinforcement learning.
[0151] The two defensive victory areas under the differential game established in Steps 203 and 204 and the defensive victory area under reinforcement learning are used to design an algorithm for the many-to-many attack and defense problem in Step 201, complete the many-to-many defense decision-making, and improve the defense success rate.
[0152] Consider a many-to-many attack and defense differential game problem under a polygonal obstacle, with 5 attacker UAVs (E 1 、E 2 、E 3 、E 4 、E 5 ) and 5 defender UAVs (P 1 、P 2 、P 3 、P 4 、P5 ),the target area Ω goal is the left rectangular area. The initial allocation results obtained by using the MOCG-ESP differential game algorithm and the proposed multi-to-multi attack and defense game method based on neuro-symbolic in the MOCG-ESP obstacle environment, and the final attack and defense confrontation results are respectively as Figure 4 , Figure 5 , Figure 6 and Figure 7 shown. It can be seen from Figures 4 - 7 that the designed multi-to-multi attack and defense differential game strategy based on neuro-symbolic can improve the defense effect, enhance the interception success rate, and better protect the important target area Ω goal .
[0153] The present application also provides an application scenario, which applies the above-mentioned multi-to-multi attack and defense game method in the obstacle environment. Specifically: the multi-to-multi attack and defense game method provided in this embodiment can be applied in the UAV attack and defense game scenario. The UAV attack and defense game scenario includes a request generation link, a multi-to-multi attack and defense differential game processing link, and a UAV control link; the decision request to be processed enters the multi-to-multi attack and defense differential game processing link from the request generation link, and through a human-machine collaboration method, the corresponding multi-to-multi defense decision planning result is obtained and enters the downstream UAV control link. The multi-to-multi attack and defense game method provided in this embodiment belongs to the multi-to-multi attack and defense differential game processing link. Specifically, in the process of the multi-to-multi attack and defense differential game processing link for the decision request to be processed, a multi-to-multi attack and defense problem model of the UAV swarm can be established, a one-to-one attack and defense sub-game problem can be established, based on the one-to-one attack and defense sub-game problem, the defense victory area under the differential game is expanded to obtain the expanded defense victory area under the differential game, the defense victory area under the reinforcement learning is expanded to obtain the expanded defense victory area under the reinforcement learning, and based on the multi-to-multi attack and defense problem model, the expanded defense victory area under the differential game, and the expanded defense victory area under the reinforcement learning, multi-to-multi defense decisions are made to obtain the multi-to-multi defense decision planning result.
[0154] Based on the same inventive concept, the embodiments of the present application also provide a multi-to-multi attack and defense game device in the obstacle environment for implementing the above-mentioned multi-to-multi attack and defense game method in the obstacle environment. The implementation solutions provided by this device to solve problems are similar to those described in the above method. Therefore, the specific limitations in one or more embodiments of the multi-to-multi attack and defense game device in the obstacle environment provided below can refer to the limitations on the multi-to-multi attack and defense game method in the obstacle environment above, and will not be repeated here.
[0155] In an exemplary embodiment, as Figure 8As shown in the figure, a multi - to - multi attack - defense game device in an obstacle environment includes the following modules:
[0156] A multi - to - multi attack - defense problem model establishment module T1, which is used to establish a multi - to - multi attack - defense problem model of an unmanned aerial vehicle (UAV) cluster; the UAV cluster includes several attacker UAVs and several defender UAVs; the multi - to - multi attack - defense problem model includes an attack - defense differential game two - side dynamics model and a game scenario model under a polygon obstacle;
[0157] A one - to - one attack - defense sub - game problem establishment module T2, which is used to establish a one - to - one attack - defense sub - game problem; the one - to - one attack - defense sub - game problem includes a one - to - one attack - defense sub - game model, a defender victory area under differential game, and a defender victory area under reinforcement learning; the defender victory area is composed of several defender victory states; the defender victory state is a state where the defender UAV wins;
[0158] A defender victory area expansion module T3 under differential game, which is used to expand the defender victory area under differential game based on the one - to - one attack - defense sub - game problem to obtain an expanded defender victory area under differential game;
[0159] A defender victory area expansion module T4 under reinforcement learning, which is used to expand the defender victory area under reinforcement learning to obtain an expanded defender victory area under reinforcement learning;
[0160] A multi - to - multi defense decision - making module T5, which is used to make a multi - to - multi defense decision based on the multi - to - multi attack - defense problem model, the expanded defender victory area under differential game, and the expanded defender victory area under reinforcement learning to obtain a multi - to - multi defense decision - making planning result.
[0161] As an optional implementation manner, the one - to - one attack - defense sub - game problem establishment module T2 is used to:
[0162] Establish a one - to - one attack - defense sub - game model;
[0163] Establish models of defender victory states, defender victory areas, attacker victory states, and attacker victory areas; the attacker victory state is a state where the attacker UAV wins; the attacker victory area is composed of several attacker victory states;
[0164] Use differential game and reinforcement learning algorithms to establish a victory state and victory area model under differential game and reinforcement learning algorithms; the victory state and victory area model under differential game and reinforcement learning algorithms includes a defender victory state under differential game, a defender victory area under differential game, a defender victory state under reinforcement learning, and a defender victory area under reinforcement learning;
[0165] Establish a forward reachable set and a backward reachable set area model;
[0166] Using a reinforcement learning algorithm, based on the forward reachable set and the backward reachable set region models and the defensive victory region under reinforcement learning, a new defensive victory region under reinforcement learning is constructed; the new defensive victory region under reinforcement learning is used to be extended to obtain an extended defensive victory region under reinforcement learning.
[0167] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 9 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data for the decision-making process of the multi-agent offensive and defensive differential game of unmanned aerial vehicles. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a multi-agent offensive and defensive game method in an obstacle environment.
[0168] Those skilled in the art can understand that Figure 9 the structure shown in
[0169] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in the above-mentioned method embodiments.
[0170] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by the processor, it implements the steps in the above-mentioned method embodiments.
[0171] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, it implements the steps in the above-mentioned method embodiments.
[0172] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0173] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0174] The databases involved in the various embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the various embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0175] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0176] In this text, specific examples are used to illustrate the principles and implementation modes of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation modes and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A multi-to-multi attack and defense game method in an obstacle environment, characterized in that: The multi-to-multi attack and defense game method under the obstacle environment includes: Establish a multi-to-multi attack and defense problem model of a drone swarm; the drone swarm includes a number of attacker drones and a number of defender drones; the multi-to-multi attack and defense problem model includes a dynamics model of the attack and defense differential game and a game scenario model under polygonal obstacles; A one-to-one attack-defense sub-game problem is established; the one-to-one attack-defense sub-game problem includes a one-to-one attack-defense sub-game model, a defense victory region under a differential game, and a defense victory region under reinforcement learning; the defense victory region is composed of a number of defense victory states; the defense victory state is a state in which the defender UAV wins; Based on the one-to-one attack and defense subgame problem, the defense victory area under the differential game is expanded to obtain the expanded defense victory area under the differential game; Expand the defense victory area under reinforcement learning to obtain the defense victory area under extended reinforcement learning; Based on the many-to-many attack and defense problem model, the defense victory area under the extended differential game and the defense victory area under the extended reinforcement learning, many-to-many defense decision-making is made and the many-to-many defense decision planning results are obtained.
2. The multi-to-multi attack and defense game method under an obstacle environment according to claim 1, characterized in that: Establish a one-on-one attack and defense subgame problem, including: Establish a one-on-one attack and defense subgame model; Establishing a defense victory state, a defense victory area, an attack victory state, and an attack victory area model; the attack victory state is a state in which the attacker's drone achieves victory; the attack victory area is composed of a number of attack victory states; Using differential game and reinforcement learning algorithms, a victory state and victory area model under the differential game and reinforcement learning algorithm is established; the victory state and victory area model under the differential game and reinforcement learning algorithm includes a defense victory state under the differential game, a defense victory area under the differential game, a defense victory state under reinforcement learning, and a defense victory area under reinforcement learning; Establish forward reachable set and backward reachable set region models; Using the reinforcement learning algorithm, a new defense victory area under reinforcement learning is constructed based on the forward reachable set and backward reachable set region models and the defense victory area under reinforcement learning; the new defense victory area under reinforcement learning is used to expand to obtain the extended defense victory area under reinforcement learning.
3. The multi-to-multi attack and defense game method under an obstacle environment according to claim 1, characterized in that: Based on the one-to-one attack and defense sub-game problem, the defense victory area under the differential game is expanded to obtain the expanded defense victory area under the differential game, which specifically includes: Based on the minimum Euclidean distance algorithm and the one-to-one attack and defense sub-game problem, the defense victory area under the differential game is expanded to obtain the expanded defense victory area under the differential game.
4. The multi-to-multi attack and defense game method under an obstacle environment according to claim 1, characterized in that: The defense victory area under reinforcement learning is expanded to obtain the defense victory area under extended reinforcement learning, which specifically includes: The proximal strategy is used to optimize the reinforcement learning algorithm, and the defense victory area under reinforcement learning is expanded to obtain the defense victory area under extended reinforcement learning.
5. The multi-to-multi attack and defense game method under an obstacle environment according to claim 1, characterized in that: Based on the many-to-many attack and defense problem model, the defense victory area under the extended differential game and the defense victory area under the extended reinforcement learning, many-to-many defense decision-making is carried out, and the many-to-many defense decision-making planning results are obtained, including: Based on the many-to-many attack and defense problem model, the defense victory area under the extended differential game, and the defense victory area under the extended reinforcement learning, a subgraph of the drone cluster is constructed; the subgraph includes point features and edge features; the point features are the positions of the attacker drone and the defender drone; the edge features indicate that the current states of the attacker drone and the defender drone belong to the defense victory area or the union under the extended differential game; the union is the union of the defense victory area under the extended differential game and the defense victory area under the extended reinforcement learning; Based on the subgraph, a binary integer programming problem model for drone cluster task matching is established; The binary integer programming problem model is solved to obtain a many-to-many defense decision planning result.
6. A multi-to-multi attack and defense game device under an obstacle environment based on the multi-to-multi attack and defense game method under an obstacle environment according to any one of claims 1 to 5, characterized in that: The multi-to-multi attack and defense game device under the obstacle environment comprises: A module for establishing a many-to-many attack and defense problem model is used to establish a many-to-many attack and defense problem model of a drone cluster; the drone cluster includes a number of attacker drones and a number of defender drones; the many-to-many attack and defense problem model includes a dynamics model of the two sides of the attack and defense differential game and a game scenario model under polygonal obstacles; A one-to-one attack-defense sub-game problem establishment module is used to establish a one-to-one attack-defense sub-game problem; the one-to-one attack-defense sub-game problem includes a one-to-one attack-defense sub-game model, a defense victory region under a differential game, and a defense victory region under reinforcement learning; the defense victory region is composed of a number of defense victory states; the defense victory state is a state in which the defender UAV achieves victory; A defense victory area expansion module under a differential game, used to expand the defense victory area under the differential game based on the one-to-one attack and defense sub-game problem to obtain an expanded defense victory area under the differential game; A defense victory area expansion module under reinforcement learning is used to expand the defense victory area under reinforcement learning to obtain the defense victory area under extended reinforcement learning; The many-to-many defense decision module is used to make many-to-many defense decisions based on the many-to-many attack and defense problem model, the defense victory area under the extended differential game, and the defense victory area under the extended reinforcement learning, and obtain the many-to-many defense decision planning results.
7. The multi-to-multi attack and defense game device in an obstacle environment according to claim 6, characterized in that: The one-to-one attack and defense sub-game problem establishment module is used to: Establish a one-on-one attack and defense subgame model; Establishing a defense victory state, a defense victory area, an attack victory state, and an attack victory area model; the attack victory state is a state in which the attacker's drone achieves victory; the attack victory area is composed of a number of attack victory states; Using differential game and reinforcement learning algorithms, a victory state and victory area model under the differential game and reinforcement learning algorithm is established; the victory state and victory area model under the differential game and reinforcement learning algorithm includes a defense victory state under the differential game, a defense victory area under the differential game, a defense victory state under reinforcement learning, and a defense victory area under reinforcement learning; Establish forward reachable set and backward reachable set region models; Using the reinforcement learning algorithm, a new defense victory area under reinforcement learning is constructed based on the forward reachable set and backward reachable set region models and the defense victory area under reinforcement learning; the new defense victory area under reinforcement learning is used to expand to obtain the extended defense victory area under reinforcement learning.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multi-to-multi attack and defense game method in an obstacle environment as described in any one of claims 1 to 5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multi-to-multi attack and defense game method in an obstacle environment described in any one of claims 1-5 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the multi-to-multi attack and defense game method in an obstacle environment described in any one of claims 1-5 is implemented.
Citation Information
Patent Citations
Intelligent defense decision-making method and device based on reinforcement learning and attack and defense games
CN110166428A
Unmanned aerial vehicle confrontation game training control method based on reinforcement learning
CN113282100A
Non-zero sum game unmanned aerial vehicle formation control method based on reinforcement learning
CN115877871A
Attack and defense game decision-making method based on reinforcement learning
CN115983389A
Unmanned aerial vehicle cooperative air combat decision-making method based on GRU-MAPPO deep reinforcement learning
CN119129413A
Cited By
Defense allocation method and device for multi-agent confrontation game and electronic equipment
CN120725155A
Rectangular area unmanned aerial vehicle hunting method based on differential game and reinforcement learning
CN121325911A
Rectangular region encirclement method for unmanned aerial vehicles based on differential game and reinforcement learning
CN121325911B