Differential game based multi-objective defense decision control method for heterogeneous multi-agent system
By constructing a heterogeneous multi-agent multi-target defense decision and control method based on differential strategies, the problems of high computational resource requirements and insufficient real-time performance in heterogeneous multi-agent collaborative decision and control are solved, realizing efficient multi-agent multi-target defense tasks, which are suitable for battlefield applications in complex environments.
Patent Information
- Application Number
- CN202411278551.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-12
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2044-09-12
AI Technical Summary
Existing technologies for collaborative decision-making and control of heterogeneous multi-agent systems in complex environments suffer from high computational resource requirements and insufficient real-time performance. In particular, the application of differential game methods is limited by simplified theoretical derivation and small-scale agents in multi-agent, multi-target defense tasks.
Based on differential game theory, a heterogeneous multi-agent multi-objective defense decision and control method is constructed. By determining the agent team in a three-dimensional half-space, a multi-objective defense game environment model and a heterogeneous multi-agent dynamic model are constructed. The optimal control strategy and value function for one-to-one attack and defense confrontation are designed. By combining the payoff function and the value function, the task allocation of the multi-agent game scenario is determined, and a rigorous mathematical derivation and analytical solution are achieved.
It improves the success rate and robustness of multi-agent, multi-target defense missions, reduces computational costs, solves the real-time and computational resource consumption problems in multi-agent cooperative operations, and is suitable for high-precision battlefield applications.
Smart Images

Figure CN119250199B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of multi-agent decision control, and more particularly, to a heterogeneous multi-agent multi-objective defense decision control method based on differential game. BACKGROUND
[0002] Multi-agent decision control in the adversarial environment is a hot issue in the current swarm intelligence system research, which refers to the decision and action of intelligent agents in different teams to achieve their team's goals and conflict with intelligent agents in other teams, so as to maximize the team's benefits. Target defense game is a variant of the classic pursuit game, which refers to the attack and defense confrontation of two conflicting teams for one or more target points. Generally speaking, the task of the defense team is to protect these targets, and the task of the attack team is to approach these targets. Multi-agent multi-objective defense decision control has wide application in important resource protection, unmanned battlefield attack and defense, and city security, such as realizing the combat mission of unmanned vehicles, unmanned aerial vehicles, unmanned underwater vehicles, satellites, etc. In dangerous battlefield environments, unmanned devices have been used many times to achieve many important target attack missions. Accordingly, considering the safety of personnel, it is particularly important to use unmanned devices for important target defense. In a dense and complex urban environment, the target defense game-based method can effectively complete the interception of non-cooperative unmanned aerial vehicles and protect people's life and property safety. In future military battlefields, multi-agent target defense games can achieve the use of unmanned vehicles and unmanned aerial vehicles to complete combat missions and achieve the penetration and attack of enemy key targets and the defense of our targets.
[0003] The above research is mainly based on homogeneous multi-agent attack and defense games. With the further development of unmanned system intelligence, especially for complex environments, homogeneous agent platforms gradually appear disadvantages. For example, in the battlefield attack and defense task, unmanned vehicles have small field of view, slow maneuvering, and other shortcomings; and unmanned aerial vehicles also have low payload, low endurance, low detection accuracy, and other defects. Therefore, in order to make up for the defects of a single agent in complex tasks, heterogeneous multi-agent collaborative decision and control becomes an effective solution. For example, in the air-ground cooperative target defense task, through the cooperation and complementation of air units and ground units, the enemy units are maximized intercepted, and the purpose of protecting important resources as much as possible in complex environments is achieved.
[0004] In the adversarial environment, there are some methods for multi-agent decision control of target area defense. At present, the main methods are based on reinforcement learning and differential game. The advantage of the method based on multi-agent reinforcement learning is to deal with randomness and dynamic environment, but the main disadvantage is similar to most machine learning methods, which requires high computing resources, so offline learning methods are generally used. Considering the actual environment that needs real-time decision, online learning is needed to interact with the environment continuously, which further increases the consumption of computing resources. Therefore, at present, the method of reinforcement learning is still far from mature in actual combat tasks. Differential game is a mathematical method for studying the antagonism or cooperation between multi-agents in a dynamic environment, which is usually used to describe decision-making problems in continuous time and space. It is based on differential equations and considers the dynamic changes of system state, which can well describe and solve the control and decision-making problems of target area defense game. Compared with the method of reinforcement learning, the advantage of differential game is that it is based on control theory and differential equations, with strict mathematical proof and analysis; in some cases, it can obtain analytical or semi-analytical solutions, which helps to understand the behavior of agents in depth; it can strictly analyze the optimal strategy and equilibrium state of agents, and is suitable for high-precision battlefield application fields. However, in the general research of this method, in order to simplify the complexity of theoretical derivation and mathematical proof, simple homogeneous agents and small number of agents are usually considered, so there is a great limitation in practical application. In order to improve the success rate and robustness of target defense, heterogeneous multi-agent cooperative combat is considered in practical application, so the method of target defense based on differential game also has certain limitations.
[0005] Therefore, it is necessary to develop a heterogeneous multi-agent multi-target defense decision control method based on differential game.
[0006] The information disclosed in the background section of the present invention is only intended to deepen the understanding of the general background of the present invention, and should not be regarded as acknowledging or implying in any form that the information constitutes prior art known to those skilled in the art. SUMMARY
[0007] The present invention provides a heterogeneous multi-agent multi-target defense decision control method based on differential game, which can improve the success rate of target area defense by taking advantage of air-ground cooperation, and provide strict theoretical derivation and mathematical proof for the target defense game of air units and ground units, and solve the problem of multi-to-multi target defense with large number of agents, which has practicality in solving the actual application of air-ground cooperative combat task.
[0008] In the first aspect, the present disclosure provides a heterogeneous multi-agent multi-target defense decision control method based on differential game, comprising:
[0009] Identify the team of intelligent agents in a three-dimensional half-space, including defenders and attackers;
[0010] Construct a multi-objective defense game environment model and a heterogeneous multi-agent dynamics model;
[0011] Determine the termination conditions and payoff functions for multi-objective offensive and defensive games;
[0012] Based on the multi-objective defense game environment model and the heterogeneous multi-agent dynamics model, the optimal control strategy and corresponding value function for one-to-one offensive and defensive confrontation are constructed based on differential strategies.
[0013] Based on the payoff function and the value function, the task allocation in the multi-agent game scenario is determined.
[0014] Preferably, constructing a multi-objective defense game environment model includes:
[0015] The cumulative distance between the interception point and the target point is used as the task indicator. The attacker's task is to minimize the task indicator, and the defender's task is to maximize the task indicator.
[0016] Transforming the adversarial relationship between attackers and defenders into a non-zero-sum game problem is the multi-objective defense game environment model.
[0017] Preferably, constructing a heterogeneous multi-agent dynamics model includes:
[0018] Determine the position coordinates of the defender, the attacker, and the target point, and then determine the state variables of the system in the differential game;
[0019] Establish the kinematic equations of the heterogeneous multi-agent dynamics model for:
[0020]
[0021] in, Let u represent the initial state of the system in the differential game. D ,u A This represents the state feedback control of the intelligent agent. Indicates the defender's location coordinates. This indicates the attacker's location coordinates. Indicates the position coordinates of the target point. This represents the state variables of a system in a differential game. and Indicates the speed of the defender and the attacker. Indicates defender D i The input of control variables, Indicates attacker A jcontrol variable input, (Θ, φ, ψ) represent the direction angle of the instantaneous velocity of the agent, Θ, φ, ψ ∈ [-π, π).
[0022] Preferably, the termination condition of the multi-objective attack-defense game is:
[0023]
[0024] wherein, represents the termination state of the game.
[0025] Preferably, the payoff function of the multi-objective attack-defense game is:
[0026]
[0027] Preferably, the optimal control strategy of the single-to-single attack-defense confrontation includes the isomorphic case and the heterogeneous case, wherein the isomorphic case is the ground defender against the attacker, and the heterogeneous case is the air defender against the attacker.
[0028] Preferably, the optimal feedback control strategy and the value function of the single-to-single target defense game in the isomorphic case are:
[0029] The victory domain of the attacker is the area inside the circle, and the victory domain of the defender is the area outside the circle, the coordinates of the center C of the area and the radius of the circle are:
[0030]
[0031] wherein,
[0032] The interception point of the maximum payoff of both sides is determined, and then the optimal interception target point is determined;
[0033] The optimal feedback control strategy represented by the direction angle calculated according to the optimal interception target point is:
[0034]
[0035] The value function of the game is:
[0036]
[0037] Preferably, the optimal feedback control strategy and the value function of the single-to-single target defense game in the isomorphic case are:
[0038] The victory domain of the attacker is the area inside the circle, and the victory domain of the defender is the area outside the circle, the coordinates of the center C of the area and the radius of the circle are:
[0039]
[0040] wherein,
[0041] The interception point maximizing the benefits of both parties is determined, and then the optimal interception target point is determined;
[0042] The optimal feedback control strategy represented by the direction angle calculated according to the optimal interception target point is
[0043] The value function of the game is:
[0044]
[0045] Preferably, according to the benefit function and the value function, the task allocation in the multi-agent game scene is determined, and the task allocation in the multi-agent game scene includes:
[0046] The target function of the task allocation is the cumulative value:
[0047]
[0048] Wherein, u ij =1 represents that the defender D i allocates the interception of the attacker A j , and v ij (x) represents the value function of the game confrontation between D i and A j .
[0049] According to the value function, the cost matrix W is defined, wherein W ij =V ij (x) represents the cost of the one-to-one game confrontation between the defender D i and the attacker A j .
[0050] The beneficial effects are as follows:
[0051] (1) The control decision method based on the differential game theory has strict mathematical derivation and analytical solution, and provides the explicit optimal state feedback control of the heterogeneous confrontation parties in the target attack and defense game.
[0052] (2) The cost matrix of the multi-agent task allocation is designed by combining the value function and the decision model of the one-to-one attack and defense game, the scene of the multi-agent multi-target defense game is solved, the algorithm improves the operation efficiency, and has the characteristics of high applicability and strong real-time performance.
[0053] The method of the present application has other characteristics and advantages, which will be apparent or will be described in detail in the accompanying drawings and subsequent specific embodiments incorporated herein, which together serve to explain the specific principles of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0054] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings in which like reference characters refer to like parts throughout the figures, and wherein:
[0055] Figure 1 A flow chart showing steps of a differential game based heterogeneous multi-agent multi-target defense decision control method according to an embodiment of the present application is shown.
[0056] Figure 2 A schematic diagram showing a heterogeneous multi-agent multi-target defense game model according to an embodiment of the present application is shown.
[0057] Figure 3 A schematic diagram showing single vs. single target defense game optimal feedback control under homogeneity according to an embodiment of the present application is shown.
[0058] Figure 4 A schematic diagram showing single vs. single target defense game optimal feedback control under heterogeneity according to an embodiment of the present application is shown.
[0059] Figure 5 A schematic diagram showing task allocation of a heterogeneous multi-agent target defense game according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0060] Preferred embodiments of the present application will be described in greater detail below. While the preferred embodiments of the present application are described below, it is to be understood that the present application can be embodied in various forms without being limited to the embodiments set forth herein.
[0061] For the purpose of understanding the scheme of the embodiments of the present application and their effects, a specific application example is given below. It should be understood by those skilled in the art that the example is only for the purpose of understanding the present application, and any specific details thereof are not intended to limit the present application in any way.
[0062] Example 1
[0063] Figure 1 A flow chart showing steps of a differential game based heterogeneous multi-agent multi-target defense decision control method according to an embodiment of the present application is shown.
[0064] As Figure 1 shown, the differential game based heterogeneous multi-agent multi-target defense decision control method includes:
[0065] Step 101, determining a team of agents in a three-dimensional half-space, including defenders and attackers;
[0066] Step 102, constructing a multi-target defense game environment model and a heterogeneous multi-agent dynamics model;
[0067] Step 103, determining the termination condition and the payoff function of the multi-target attack and defense game;
[0068] Step 104, constructing a single-to-single attack and defense confrontation optimal control strategy and the corresponding value function based on differential games according to the multi-target defense game environment model and the heterogeneous multi-agent dynamics model;
[0069] Step 105, determining the multi-agent game scene task allocation according to the payoff function and the value function.
[0070] In one example, constructing a multi-target defense game environment model includes:
[0071] Taking the cumulative distance of the interception point and the target point as the task index, the task of the attacker is to minimize the task index, and the task of the defender is to maximize the task index;
[0072] The confrontation between the attacker and the defender is converted into a non-zero-sum game problem, that is, a multi-target defense game environment model.
[0073] In one example, constructing a heterogeneous multi-agent dynamics model includes:
[0074] Determining the position coordinates of the defender, the attacker and the target point, and then determining the state variable of the system in the differential game;
[0075] Establishing the kinematics equation of the heterogeneous multi-agent dynamics model is:
[0076]
[0077] wherein, represents the initial state of the system of the differential game, u D represents the state feedback control of the agent, A represents the position coordinates of the defender, represents the position coordinates of the attacker, represents the position coordinates of the target point, represents the state variable of the system in the differential game, and represents the speed of the defender and the attacker, represents the control variable input of the defender D i , represents the control variable input of the attacker A j , (Θ, φ, ψ) respectively represent the direction angle of the instantaneous speed of the agent, θ, φ, ψ ∈ [-π, π).
[0078] In one example, the termination condition of the multi-objective attack-defense game is:
[0079]
[0080] wherein, represents the termination state of the game.
[0081] In one example, the payoff function of the multi-objective attack-defense game is:
[0082]
[0083] In one example, the optimal control strategy of single-to-single attack-defense confrontation includes the isomorphic case and the heterogeneous case, wherein the isomorphic case is the ground defender against the attacker, and the heterogeneous case is the air defender against the attacker.
[0084] In one example, the optimal feedback control strategy and the value function of the single-to-single target defense game in the isomorphic case are:
[0085] The victory area of the attacker is determined as the area inside the circle, and the victory area of the defender is determined as the area outside the circle, the coordinates of the center C of the area and the radius of the circle are:
[0086]
[0087] wherein,
[0088] The interception point of the maximum payoff of both sides is determined, and then the optimal interception target point is determined;
[0089] The optimal feedback control strategy represented by the direction angle calculated according to the optimal interception target point is:
[0090]
[0091] The value function of the game is:
[0092]
[0093] In one example, the optimal feedback control strategy and the value function of the single-to-single target defense game in the isomorphic case are:
[0094] The victory area of the attacker is determined as the area inside the circle, and the victory area of the defender is determined as the area outside the circle, the coordinates of the center C of the area and the radius of the circle are:
[0095]
[0096] wherein,
[0097] The interception point of the maximum payoff of both sides is determined, and then the optimal interception target point is determined;
[0098] The optimal feedback control strategy represented by the direction angle calculated according to the optimal interception target point is
[0099] The value function of the game is:
[0100]
[0101] In one example, according to the payoff function and the value function, determining the task allocation in the multi-agent game scenario includes:
[0102] The objective function of the task allocation is the cumulative value:
[0103]
[0104] Wherein, u ij =1 represents that the defender D i allocates the interception of the attacker A j , represents the value function of the game confrontation between D i and A j .
[0105] According to the value function, the cost matrix W is defined, wherein represents the cost of the single-to-single game confrontation between the defender D i and the attacker A j .
[0106] Specifically, the present application is a multi-agent decision control algorithm for target defense tasks. Based on the single-to-single agent attack-defense game optimal control law based on differential game, a heterogeneous multi-agent decision model and task allocation algorithm are designed, the optimization decision control problem of agent attack-defense game is solved, the operation cost is effectively reduced, the dimension explosion problem caused by the increase of the number of clusters is avoided, and the real-time performance and practicality of the algorithm are improved. The specific technical solutions are as follows:
[0107] Figure 2 A schematic diagram of a heterogeneous multi-agent multi-target defense game model according to one embodiment of the present application is shown.
[0108] (I) Construct a multi-target defense game environment model. For example Figure 2As shown, the agents in the three-dimensional half-space are divided into two teams, one is the defender, and the other is the attacker, and their confrontation revolves around several target (important resource) points on the plane; the task of the attacker is to reach these targets, and if it cannot, it also wants to approach these targets as much as possible; the task of the defender is to prevent the attacker from reaching its target and to keep the attacker away from these targets as much as possible. In addition, the defenders are composed of two-dimensional agents and three-dimensional agents, and the attackers are composed of two-dimensional agents. The cumulative distance of the interception point and the target point is used to evaluate the task index of the defender; obviously, the task of the attacker team is to minimize this index, and the task of the defender team is to maximize this index. In this way, the confrontation between the attacker and the defender is described as a non-zero-sum game problem. The detailed mathematical description is as follows: there are M defenders in space, represented by D i , where i = 1, …, M; at the same time, the other party has N attackers, represented by A j , where j = 1, …, N; there are K targets (green markers), represented by T k , where k = 1, …, K. The attack and defense game occurs in a three-dimensional half-space, and the defender team is heterogeneous, with both ground defenders whose trajectories are limited to the plane and air defenders; while all the trajectories of the attackers are limited to the plane, and all the targets are also in the plane.
[0109] (ii) Construct a heterogeneous multi-agent dynamics model. The position coordinates of the defenders, attackers, and target points are represented by , respectively. The state variable of the system in the differential game is represented by . Considering simple first-order motion, the velocities of the defenders and attackers are represented by and respectively. The control variable input of the defender D i is represented by , and the control variable input of the attacker A j is represented by . The direction angles of the instantaneous velocities of the agents are represented by (Θ, φ, ψ), and Θ, φ, ψ ∈ [-π, π). The agents control the state of the system by changing the instantaneous direction of their velocity. The kinematic equation is as follows:
[0110]
[0111] where represents the initial state of the system of the differential game, and u D , u A represents the state feedback control of the agent. It is assumed that the defender D i targets the attacker A j , and D ithe speed of A j the speed of A Since both the defender and the attacker are confined in a plane, for those defenders, For the attacker, we have
[0112] (Three) Determine the termination condition and value function of the multi-target attack-defense game. When all attackers are intercepted or reach the target, the game ends. The termination set of the game is represented as follows:
[0113]
[0114] Where the termination state of the game is represented by To simplify the model, it is assumed that the distance between the defender and the attacker is 0 at the moment of successful interception; at the same time, it is assumed that each attacker can only specify one target for attack, and multiple attackers can specify the same target, and each defender can only intercept one attacker; in the model, in order to reasonably utilize the interception resources, each attacker is only assigned one interceptor. If the attacker A j designates T k as the target, it will try to reach the target T k , if it cannot reach T k , A j will try to approach T k as much as possible before being intercepted, until it is intercepted by the defender, at which time the payoff function of the game is defined as:
[0115]
[0116] The final payoff of the game is only related to the final state of the system, i.e. the position of each attacker being intercepted; and the final state of the game system is only related to the initial state of the system and the subsequent control; when the initial state is determined, the final payoff of the game can be considered to be only related to the control strategy (u D , u A ) of the defender and the attacker, i.e. both parties of the game affect the final payoff of the game through the control strategy. We define this payoff from the perspective of the defender, the farther the cumulative distance between the interception point and the target, the greater the payoff of the defender, and vice versa. Therefore, the defender wants to maximize the final payoff of the game and will continuously optimize its strategy; similarly, the attacker wants to minimize the final payoff of the game and will also continuously optimize its strategy, and eventually reaches a balance at a certain moment, at which time the payoff is defined as the value function of the game
[0117] (Four) Designing optimal control law for single-to-single attack-defense confrontation based on differential game
[0118] According to the game environment model in step (one) and the agent model in step (two), the control strategy of single-to-single game confrontation is designed as follows. It is divided into isomorphic and heterogeneous cases. The isomorphic case is the ground defender against the attacker, and the heterogeneous case is the air defender against the attacker.
[0119] Figure 3 A schematic diagram of single-to-single target defense game optimal feedback control under isomorphism according to an embodiment of the application is shown.
[0120] For the isomorphic case, as shown in Figure 3 , the opposing parties are ground defender D1, attacker A1, and target T1. At this time, the plane is directly divided into two regions, and the outer region of the circle represents the victory domain of the defense party, and the inner region of the circle represents the victory domain of the attack party. The coordinates of the center C and the radius of the circle are as follows,
[0121]
[0122] If the target is inside the circle, interception will inevitably occur. In order to get as close to the target as possible, the interception point of the maximum benefit of the two parties should be the intersection point of the circle and the straight line of C 11 T1. Therefore, the coordinates of the optimal interception target point "×" are:
[0123]
[0124] Among them, Thus, the optimal feedback control strategy represented by the direction angle can be calculated according to the optimal target point Finally, the value function of the game is:
[0125] Figure 4 A schematic diagram of single-to-single target defense game optimal feedback control under isomorphism according to an embodiment of the application is shown.
[0126] For the isomorphic case, as shown in Figure 4 , the opposing parties are ground defender D1, attacker A1, and target T1. At this time, the plane is directly divided into two regions, and the outer region of the circle represents the victory domain of the defense party, and the inner region of the circle represents the victory domain of the attack party. The coordinates of the center C and the radius of the circle are as follows,
[0127]
[0128] Among them, Likewise, if the target is inside, an interception occurs, and to maximize the payoff of both sides, the interception point should be exactly the point where the circle and C intersect 22 T2 straight line intersection, can be calculated optimal target point Optimal control law And the value function
[0129] (Five) Multi-agent game scene task allocation algorithm design
[0130] According to the objective function in the multi-agent game confrontation scene in step (three) and the value function of the single-to-single agent attack and defense game obtained in step (four), the task allocation algorithm is designed below to realize the multi-target defense game decision in the heterogeneous multi-agent scene.
[0131] Define the cumulative value The objective function of the allocation problem is represented. The bipartite graph matching method: KM algorithm is used to solve this problem, where u ij =1 represents the defender D i to allocate the interceptor A j , The value function of the game confrontation between D i and A j is calculated by the method in step (four). According to the value function in step (four), the cost matrix W is defined. W is an n×n square matrix that will be used in the KM algorithm, where represents the cost of the single-to-single game confrontation between the defender D i and the attacker A j . If M>N, append M-N columns of zero elements to the end of the matrix W. For a specific i' and j', when indicates that the defender D i′ cannot intercept the attacker A j′ , we define W i′j′ as negative infinity. When the defender confirms the target of the attacker in advance, W is uniquely determined.
[0132] Table 1 gives the construction time of the cost matrix W and the running time of the algorithm in different sizes of agent number instances. The calculation is run in Python 3.11.5, with a CPU of 3.60GHz and a RAM of 32GB. The algorithm is in milliseconds, solving the defender allocation problem, and the calculation time remains stable as the size of the agent number increases. Therefore, this method can solve the real-time problem in the multi-agent multi-target defense task.
[0133] Table 1
[0134]
[0135] Figure 5 A schematic diagram of task allocation of a heterogeneous multi-agent target defense game according to an embodiment of the present application is shown.
[0136] Consider 12 agents, attack and defense confrontation around 3 targets; among them, 7 defenders (3 air defenders, 4 ground defenders), 5 attackers. The discrete simulation results are shown in Figure 5 The time step of the simulation is 1 millisecond. Figure 5 The blue markers are the defenders, the blue triangles are the ground defenders, the blue rectangles are the air defenders, and their trajectories are also represented by blue solid lines; the red circles are the attackers, and their trajectories are represented by red solid lines; the star markers are the targets, and the green star markers indicate that no attacker has reached this target, and the red star markers indicate that an attacker has reached this target. The game confrontation occurs in a three-dimensional half-space, where the ground defenders, attackers and targets are located in the xy plane, and the trajectories of these agents are also limited to this plane; the air defenders are above the xy plane, i.e. in the region where z>0 in the xyz space, and their trajectories are also three-dimensional.
[0137] According to the heterogeneous multi-agent task allocation scheme of the present application, the defenders will get different results when using different allocation strategies. Figure 5 In (a), the defenders use the present scheme, and all 5 attackers are intercepted, and the maximum value of the cumulative distance between the interception point and the target is obtained; in Figure 5 In (b), the defenders use other allocation schemes, and 4 attackers are intercepted at the same time, and the cumulative distance between the interception point and the target cannot be maximized. Therefore, the present application can provide the optimal defense strategy for the defenders.
[0138] The present application designs the decision-making tasks and objective functions of both sides for the target defense attack and defense confrontation of air units and ground units, and through strict mathematical derivation and theoretical proof, an explicit state feedback control strategy is obtained that maximizes the benefits of both sides. The present application designs a decision-making model and task allocation algorithm for multi-target multi-agent attack and defense game, decouples the complex multi-agent confrontation scene into multiple one-on-one confrontation scenes, uses the results of one-on-one target defense game to design a cost matrix, solves the multi-agent task allocation problem, and effectively reduces the consumption of computing resources.
[0139] Those skilled in the art will understand that the purpose of the above description of embodiments of the present application is only to exemplarily illustrate the beneficial effects of the embodiments of the present application, and is not intended to limit the embodiments of the present application to any of the examples given.
[0140] Having described various embodiments of the application, it is to be understood that the above description is meant to be illustrative only, and that many modifications and variations of the embodiments are possible without departing from the scope and spirit of the described embodiments. Many modifications and variations of the described embodiments are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims and their equivalents, the described embodiments can be practiced otherwise than as specifically described.
Claims
1. A heterogeneous multi-agent multi-objective defense decision control method based on differential game, characterized in that, The application relates to a method for constructing a multi-target defense game environment model and a heterogeneous multi-agent dynamics model, determining a termination condition and a payoff function of a multi-target attack-defense game, constructing a single-to-single optimal control strategy and a corresponding value function of an attack-defense confrontation based on a differential game, and determining a multi-agent game scene task allocation according to the payoff function and the value function. The application relates to a method for constructing a multi-target defense game environment model and a heterogeneous multi-agent dynamics model, determining a termination condition and a payoff function of a multi-target attack-defense game, constructing a single-to-single optimal control strategy and a corresponding value function of an attack-defense confrontation based on a differential game, and determining a multi-agent game scene task allocation according to the payoff function and the value function. The application relates to a method for constructing a multi-target defense game environment model and a heterogeneous multi-agent dynamics model, determining a termination condition and a payoff function of a multi-target attack-defense game, constructing a single-to-single optimal control strategy and a corresponding value function of an attack-defense confrontation based on a differential game, and determining a multi-agent game scene task allocation according to the payoff function and the value function. The application relates to a method for constructing a multi-target defense game environment model and a heterogeneous multi-agent dynamics model, determining a termination condition and a payoff function of a multi-target attack-defense game, constructing a single-to-single optimal control strategy and a corresponding value function of an attack-defense confrontation based on a differential game, and determining a multi-agent game scene task allocation according to the payoff function and the value function. The application relates to a method for constructing a multi-target defense game environment model and a heterogeneous multi-agent dynamics model, determining a termination condition and a payoff function of a multi-target attack-defense game, constructing a single-to-single optimal control strategy and a corresponding value function of an attack-defense confrontation based on a differential game, and determining a multi-agent game scene task allocation according to the payoff function and the value function. The application relates to a method for constructing a multi-target defense game environment model and a heterogeneous multi-agent dynamics model, determining a termination condition and a payoff function of a multi-target attack-defense game, constructing a single-to-single optimal control strategy and a corresponding value function of an attack-defense confrontation based on a differential game, and determining a multi-agent game scene task allocation according to the payoff function and the value function. The application relates to a method for constructing a multi-target defense game environment model and a heterogeneous multi-agent dynamics model, determining a termination condition and a payoff function of a multi-target attack-defense game, constructing a single-to-single optimal control strategy and a corresponding value function of an attack-defense confrontation based on a differential game, and determining a multi-agent game scene task allocation according to the payoff function and the value function. The application relates to a method for constructing a multi-target defense game environment model and a heterogeneous multi-agent dynamics model, determining a termination condition and a payoff function of a multi-target attack-defense game, constructing a single-to-single optimal control strategy and a corresponding value function of an attack-defense confrontation based on a differential game, and determining a multi-agent game scene task allocation according to the payoff function and the value function. The application relates to a method for constructing a multi-target defense game environment model and a heterogeneous multi-agent dynamics model, determining a termination condition and a payoff function of a multi-target attack-defense game, constructing a single-to-single optimal control strategy and a corresponding value function of an attack-defense confrontation based on a differential game, and determining a multi-agent game scene task allocation according to the payoff function and the value function. The application relates to a method for constructing a multi-target defense game environment model and a heterogeneous multi-agent dynamics model, determining a termination condition and a payoff function of a multi-target attack-defense game, constructing a single-to-single optimal control strategy and a corresponding value function of an attack-defense confrontation based on a differential game, and determining a multi-agent game scene task allocation according to the payoff function and the value function. The application relates to a method for constructing a multi-target defense game environment model and a heterogeneous multi-agent dynamics model, determining a termination condition and a payoff function of a multi-target attack-defense game, constructing a single-to-single optimal control strategy and a corresponding value function of an attack-defense confrontation based on a differential game, and determining a multi-agent game scene task allocation according to the payoff function and the value function. establishing the kinematic equation of the heterogeneous multi-agent dynamics model is: in, This represents the initial state of a system in a differential game. This represents the state feedback control of the intelligent agent. Indicates the defender's location coordinates. This indicates the attacker's location coordinates. Indicates the position coordinates of the target point. This represents the state variables of a system in a differential game. and Indicates the speed of the defender and the attacker. Indicates the defender The input of control variables, Indicates attacker The input of control variables, These represent the direction angles of the agent's instantaneous velocity. ; The application relates to a method for constructing a multi-target defense game environment model and a heterogeneous multi-agent dynamics model, determining a termination condition and a payoff function of a multi-target attack-defense game, constructing a single-to-single 2. The differential game based heterogeneous multi-agent multi-objective defense decision control method according to claim 1, wherein, wherein, represents the termination state of the game.
3. The differential game based heterogeneous multi-agent multi-objective defense decision control method according to claim 1, wherein, 。 4. The differential game based heterogeneous multi-agent multi-objective defense decision control method according to claim 1, wherein, The winning area of the attacker is the area inside the circle, and the winning area of the defender is the area outside the circle, with the center of the circle being the coordinates of the center and the radius of the circle: wherein ; Determining the interception point that maximizes the benefits of both sides, and then determining the optimal interception target point ; wherein .
5. The differential game based heterogeneous multi-agent multi-objective defense decision control method according to claim 1, wherein, The winning area of the attacker is the area inside the circle, and the winning area of the defender is the area outside the circle, with the center of the circle being the coordinates of the point and the radius of the circle. wherein , ; Determining the interception point that maximizes the benefits of both parties, and then determining the optimal interception target point The optimal feedback control strategy represented by the direction angle calculated according to the optimal intercept target point is ; 。 6. The differential game based heterogeneous multi-agent multi-objective defense decision control method according to claim 4 or 5, wherein, wherein, represents a defender de-allocating an interceptor , represents with a value function of a game play According to the value function define the cost matrix where denotes the cost of a single pair single game of the game between the defender and the attacker .