Learning-based and easy-to-generalize large-scale multi-agent collaborative game method
By combining supervised learning and reinforcement learning methods, multi-layer perception machines and graph neural networks are trained, and the generalization ability of large-scale multi-autonomous collaborative game is realized, solving the problem of insufficient generalization ability of existing methods in large-scale scenarios, and improving training efficiency and stability.
Patent Information
- Application Number
- CN202510197366.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-30
AI Technical Summary
Most of the existing multi-autonomous offensive and defensive game methods only use a single supervised learning or reinforcement learning algorithm, which is difficult to effectively generalize in large-scale multi-autonomous collaborative games. The end-to-end reinforcement learning method has a long training time and is easy to not be reduced in the results.
Using a method combining supervised learning and reinforcement learning, through 1VS1 game trajectory sampling and annotation, multi-layer perceptron prediction network and graph neural network strategy network are trained to achieve the generalization ability of large-scale multi-autonomous collaborative games, and to achieve the capture of defense autonomous bodies through sliding mode control.
It realizes the ability to generalize from 1VS1 game scenarios to large-scale game scenarios, and can apply various autonomous dynamic models, reduce the dimension and complexity of the action space, and improve the training efficiency and stability of the algorithm.
Smart Images

Figure CN120068992A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of control and information technology, and particularly to a large-scale multi-agent collaborative game method based on learning and easy generalization. Background Art
[0002] In the past decade, multi-agent systems have developed rapidly. In recent years, some adversarial game problems of multi-agent systems have attracted wide attention due to their extensive applications in civilian and military fields. For example, pursuit-evasion games, target attack-defense games, and border defense problems. In the border defense problem, the goal of the invaders is to send as many invaders as possible to a given border, while the goal of the defender is to stop the invaders.
[0003] In recent years, intelligent solution methods based on deep reinforcement learning have achieved success in the field of decision control, giving birth to numerous algorithms and being widely applied in fields such as cluster systems, attack-defense games, etc. For example, the multi-agent deep deterministic policy gradient (MADDPG) algorithm developed by the OpenAI team in the United States. As a classic machine learning method, supervised learning has important application potential in multi-agent attack-defense game problems, such as policy imitation, assisting in complex reinforcement learning tasks, etc. The core idea of policy imitation is to train a model through a set of labeled data (input-output pairs) to learn the mapping relationship between the input and the expected output. In multi-agent attack-defense problems, supervised learning is usually used for policy imitation learning.
[0004] It should be noted that in the field of multi-agent attack-defense games, most of the existing methods only use a single supervised learning or reinforcement learning algorithm. However, in practical applications, coupling the two methods can bring some advantages lacking in a single method, such as the generalization ability to the agent scale, the generalization ability to the agent dynamics, etc.
[0005] Under the above background, we propose a large-scale multi-agent collaborative game method based on learning and easy generalization. This method can realize the collaborative game of large-scale multi-agents. Its basic idea is to use the supervised learning method to predict the 1VS1 game, so as to be applicable to agent systems with different dynamics; at the same time, based on the reinforcement learning method, using graph neural networks, under the condition of local perception, design a reward function and train a policy network to learn the probability distribution of the defense tasks of the defense agents.
[0006] The difficulties in designing a large-scale multi-agent collaborative game method based on learning and easy generalization are as follows:
[0007] First: The difficulty of solving the multi-agent collaborative game problem increases exponentially with the growth of the number of agents;
[0008] Second: Additionally, the end-to-end reinforcement learning method for large-scale agents requires long-term training and is extremely prone to non-convergence of results;
[0009] Third: Most of the game methods are only applicable to a certain kind of dynamic model and cannot be simply migrated and applied to the cooperative game problems of agent clusters with other dynamic models. Summary of the Invention
[0010] The object of the present invention is to propose a learning-based and easily generalized large-scale multi-agent cooperative game method for realizing the cooperative control of a large-scale multi-agent system representing the defense side, and at the same time playing a game with another multi-agent system representing the offense side, with the goal of capturing all the offense sides so as to achieve the protection of key targets.
[0011] To achieve the above object, the technical solution of the present invention is: A learning-based and easily generalized large-scale multi-agent cooperative game method, specifically including the following steps:
[0012] Step 1: The offense agent and the defense agent conduct a 1VS1 game, sample the game trajectory, collect the position and speed information of both the offense and defense sides, and label the game result (0 represents defense failure, 1 represents defense success) to generate training samples; in the boundary defense task of the agent cluster, collect the position and speed information of both the offense and defense sides as new samples;
[0013] Step 2: Design a prediction network for each defense agent based on a multi-layer perceptron; the input of the prediction network is the new samples generated in Step 1, and the output is the predicted game result; and train the prediction network through the training samples in Step 1 and the supervised learning method.
[0014] Step 3: Design a policy network for each defense agent based on a graph neural network and a multi-layer perceptron; the input of the policy network is the output result of the prediction network, the position and speed information of some offense agents perceived by each defense agent, and the position and speed information of some other defense agents, and the output is the probability distribution of the defense tasks of each defense agent, that is, the defense probability distribution of all the offense agents perceived by each defense agent; train the policy network based on the PPO algorithm, and adopt the method of centralized training and distributed execution.
[0015] Step 4: Based on the probability distribution of the defense tasks of each defense agent, confirm the defense targets of each defense agent according to the sampled defense strategy.
[0016] Step 5: Based on sliding mode control, realize the pursuit and capture of the confirmed defense targets by the defense agents.
[0017] Preferably, the attacking autonomous agent and the defending autonomous agent play a 1VS1 game, sample the game trajectory, collect the position and speed information of the attacking and defending parties, and mark the game results to generate training samples, specifically:
[0018] An attacking autonomous agent and a defending autonomous agent are randomly generated in the game area. The attacking autonomous agent runs based on the shortest path rule and active obstacle avoidance rule to avoid the defending autonomous agent and reach the target position O at the same time. i ,The defensive agent pursues the attacker according to the sliding mode control algorithm; the position and speed information of the attacker and the defender are recorded in real time, and the result of the game is stored as a label.
[0019] Preferably, the characteristic vector X of the position and speed information of the attack and defense includes the speed of the attacker and the defender, the relative position of the attacker and the defender, and the relative distance of the attacker to the defense boundary, which is specifically expressed as:
[0020] X=[V a ,V d ,P (a,t) ,D (a,d) ]
[0021] Where V a With V d Represent the speed of the attacker and defender, P (a,t) Denotes the distance from the attacker to the defense boundary, D (a,d) Indicates the relative position between the defender and the attacker.
[0022] Preferably, the prediction network is a multi-layer perceptron comprising 64 layers, and the activation function adopts a ReLu function.
[0023] Preferably, the prediction network uses a nearest neighbor-based supervised learning algorithm KNN to classify the newly added samples, as follows:
[0024] Step 2.1, calculate the classification labels based on the training samples, using C and D to represent the sets of successful and failed interceptions respectively;
[0025] Step 2.2: Select k samples from the interception success set C and the interception failure set D respectively;
[0026] Determine the average value of the feature vectors of the k samples in the successful interception set and set it as the center point O suc , k samples with O suc The sample data with the farthest distance from O suc The 2-norm of is set to the radius R suc ;
[0027] Determine the average value of the feature vectors of the k samples in the intercept failure set and set it as the center point Ofal , the sample data in the k samples that is farthest from O fal has its 2-norm from O fal set as the radius R fal ;
[0028] Step 2.3. According to the center point O suc and the radius R suc , and the center point O fal and the radius R fal , process the newly added sample data, and calculate the distances d(x I from the center point O suc and O fal for each newly added sample data x I , O suc ) and d(x I , O fal ), and compare them with the radius R suc and R fal respectively for data classification. Specifically:
[0029] Calculate the distances d(x I from the center point O suc and O fal for the feature vector x of the newly added sample I I , O suc ) and d(x I , O fal );
[0030] d(x I , O suc ) = ||x I - O suc ||, d(x I , O fal ) = ||x I - O fal ||
[0031] In the formula, x I represents the feature vector of the newly added sample I, and ||x I - O suc || represents the 2-norm of the difference between two vectors. Compare the distance d(x I , O suc ) with R suc . If d(x I , O suc ) is less than R suc , it indicates that the label of the newly added sample I is classified as interception successful; compare the distance d(x I , O fal ) with R fal . If d(x I , Ofal ) less than R fal It indicates that the label classification of the newly added sample I fails to intercept; in other cases, it is set to not classify temporarily and discard the current sample data.
[0032] Preferably, step 3 is specifically as follows:
[0033] Step 3.1: Define a communication topology graph. Each defense agent is used as a vertex of the communication topology graph, and there are m defense agents M = {1, 2,..., m}; the communication relationship between defense agents is used as an adjacency matrix E = {e ij , i, j ∈ {1, 2,…, m}}, when defense agents i and j can communicate with each other, let e ij = 1, and in other cases, let e ij = 0; finally, the communication topology graph is represented as G = [M, E];
[0034] Step 3.2: Perform an aggregation operation on the feature vectors of each defense agent: Let the feature vector of defense agent j be s j , and the feature vector is s j Specifically, it is the position and speed information of some attacking agents perceived by each defense agent, as well as the position and speed information of some other defense agents; the set of all defenders that can communicate with defense agent i is represented by , then defense agent i obtains a new feature vector through the aggregation operation
[0035] Step 3.3: Activate the new feature vector obtained from the aggregation operation through an activation function. Through K times of aggregation and activation, the final feature vector of each vertex is output;
[0036] Step 3.4 Combine the output result of the prediction model and the final feature vector as the input of the policy network. The policy network is set as a 64-layer multi-layer perceptron. The input dimension is the sum of the feature vector dimensions of each vertex and the number of defense objects perceived by the defense agent, and the output dimension is the number of defense objects perceived by the defense agent. The activation function uses the ReLu function. The policy network is trained through a typical reinforcement learning algorithm, that is, the PPO algorithm, and a centralized training and distributed execution method is adopted.
[0037] Preferably, the reward function of the policy network includes two parts: the first part is the sum of the negative values of the distances between the defense objects of each defense agent and the defense agent, and the second part is that when each attacker is captured, all defense agents increase a fixed reward.
[0038] Preferably, the probability distribution of the defense task of each defense agent D i is expressed in the following form:
[0039] L i = [P(A 1 |D i ), P(A 2 |D i ),..., P(A n |D i )]
[0040] where A 1 to A n represent the perceived attacking agents; the probability distribution L i satisfies the following constraints:
[0041]
[0042] Preferably, in step 4, to confirm the defense target of each defense agent according to the sampled defense strategy, specifically, according to the greedy sampling principle, select one of the perceived attacking parties as the defense target.
[0043] Preferably, step 5 is specifically: according to the position and speed information of the attacking party confirmed in step 4, design a sliding mode controller, use the position error as the sliding mode surface, and design the sliding mode control gain under the condition that the position and speed of the agent are bounded to achieve the capture of the attacking party; the sliding mode controller is expressed as:
[0044]
[0045] In the formula, u i represents the control input signal of the defense agent D i , k' represents the gain coefficient of the sliding mode controller, S di represents the position information of the defense agent D i , represents the position information of the defense target of the defense agent D i confirmed in step 4.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] 1. By coupling the supervised learning and reinforcement learning methods, the game method can be generalized from the 1VS1 game scenario to large-scale game scenarios;
[0048] 2. Through the relevant data such as the game trajectories and game results generated in specific scenarios, train the neural network to estimate the game results based on the supervised learning method, so that the final game algorithm can be generalized and applied to the dynamic models of various agents;
[0049] 3. The action space of the reinforcement learning algorithm is designed as the corresponding defense probability of the observed attacking agent, which reduces the dimension and complexity of the action space and effectively improves the stability of the training efficiency of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 FIG. is a training framework diagram of the supervised learning model for steps 1-2 of the present invention;
[0051] Figure 2 FIG. is a training framework diagram of the reinforcement learning model for steps 3-5 of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0052] The following combines the attached Figure 1-2 , and specifically describes the technical solution of the present invention.
[0053] The present invention studies how to design a large-scale multi-agent collaborative game method based on learning and easy to generalize, which can control multiple agents to achieve defense and attack agents to reach key target areas through collaborative capture. The present invention combines supervised learning and reinforcement learning methods, so that the collaborative game algorithm has strong generalization ability, including generalizing from 1VS1 game to many-to-many game, and from simple integrator dynamics to complex non-linear dynamics. This feature can play an important role in many practical application scenarios, such as game problems based on unmanned aerial vehicle clusters, agent clusters, and unmanned underwater vehicle clusters.
[0054] The present invention is a multi-agent boundary defense method combining supervised learning and reinforcement learning methods, specifically including the following steps:
[0055] Step 1: The attacking agent and the defending agent conduct a 1VS1 game, sample the game trajectory, collect the position and speed information of both the attacking and defending sides, and label the game result (0 represents defense failure, 1 represents defense success) to generate training samples; in the boundary defense task of the agent cluster, collect the position and speed information of both the attacking and defending sides as new samples;
[0056] Step 2: Design a prediction network for each defending agent based on a multi-layer perceptron; the input of the prediction network is the training samples and new samples generated in Step 1, and the output is the predicted game result; and train the prediction network through a supervised learning method;
[0057] Step 3: Design a policy network for each defensive agent based on a graph neural network and a multi-layer perceptron; the input of the policy network is the output result of the prediction network, the position and velocity information of some attacking agents perceived by each defensive agent, and the position and velocity information of some other defensive agents, and the output is the probability distribution of the defensive tasks of each defensive agent, that is, the probability distribution of defending all the perceived attacking agents; train the policy network based on the PPO algorithm, and adopt the method of centralized training and distributed execution;
[0058] Step 4: Based on the probability distribution of the defensive tasks of each defensive agent, confirm the defensive targets of each defensive agent according to the sampled defensive strategy;
[0059] Step 5: Based on sliding mode control, realize the pursuit and capture of the confirmed defensive targets by the defensive agents.
[0060] In the boundary defense task of the agent cluster, there are attacking agents with n agents and defensive agents with m agents. Both the attacking agents and the defensive agents move in a bounded closed region Ω, and the straight line T forms the boundary of the target region, and the target region is included inside the region The movement of the defensive agents is restricted in Ω goal , and the initial positions of the attacking agents are restricted in Ω / Ω goal , that is, in Ω but not in Ω goal . The goal of the attacker is to reach Ω goal , and the goal of the defender is to intercept the attacker at the boundary of the target region Ω goal . Each defender can only perceive the existence of some attackers, and according to the distance relationship between the attacker and the defender, sort the perceived attackers in ascending order of distance.
[0061] In the above step 1, the attacking agent and the defensive agent conduct a 1VS1 game, sample the game trajectory, collect the position and velocity information of both the attacking and defensive sides, and label the game result to generate training samples, specifically:
[0062] Randomly generate an attacking agent and a defensive agent in the game area. The attacking agent runs based on the shortest path rule and the active obstacle avoidance rule to avoid the defensive agent while going to the target position O i , and the defensive agent pursues the attacker according to the sliding mode control algorithm; record the position and velocity information of the attack and defense in real time, and store the result of the game as a label.
[0063] The feature vector X of the position and velocity information of the attack and defense includes the velocities of the attacker and the defender, the relative positions of the attacker and the defender, and the relative distance of the attacker to the defense boundary, and is specifically expressed as:
[0064] X = [V a , V d , P (a,t) , D (a,d)
[0065] Where V a and V d represent the speeds of the attacker and the defender respectively, P (a,t) represents the distance of the attacker to the defense boundary, and D (a,d) represents the relative position between the defender and the attacker.
[0066] In the above step 2, the prediction network is a multi-layer perceptron with 64 layers, and the activation function uses the ReLu function. The prediction network uses the KNN (K-Nearest Neighbor) supervised learning algorithm based on the nearest neighbor to classify new samples, specifically as follows:
[0067] Step 2.1: Calculate the classification labels based on the training samples, and use C and D to represent the sets of successful and failed interceptions respectively;
[0068] Step 2.2: Select k samples from the set C of successful interceptions and the set D of failed interceptions respectively;
[0069] Determine the average value of the feature vectors of the k samples in the set of successful interceptions, and set it as the center point O suc , and the 2-norm of the sample data in the k samples that is farthest from O suc to O suc is set as the radius R suc ;
[0070] Determine the average value of the feature vectors of the k samples in the set of failed interceptions, and set it as the center point O fal , and the 2-norm of the sample data in the k samples that is farthest from O fal to O fal is set as the radius R fal Similarly, define the center point O fal and the radius R fal of the set of failed interceptions;
[0071] Step 2.3: According to the center point O suc and the radius R suc , as well as the center point O fal and the radius R fal , process the new sample data, and calculate the distances d(x I to the center point O suc and O fal for each new sample data x I , O suc ) and d(x I , O fal ), respectively compare with the radii R suc and R fal for data classification. Specifically:
[0072] Calculate the feature vector x of the new sample I through the following formula I with the center point O suc and O fal The distances d(x I , O suc ) and d(x I , O fal );
[0073] d(x I , O suc ) = ||x I - O suc ||, d(x I , O fal ) = ||x I - O fal ||
[0074] In the formula, x I represents the feature vector of the new sample I, and ||x I - O suc || represents the 2-norm of the difference between two vectors. Compare the distance d(x I , O suc ) with R suc . If d(x I , O suc ) is less than R suc , it indicates that the label of the new sample I is classified as interception successful; compare the distance d(x I , O fal ) with R fal . If d(x I , O fal ) is less than R fal , it indicates that the label of the new sample I is classified as interception failed; in other cases, it is set to not classify for the time being and discard the current sample data.
[0075] In step 3 above, based on the strategy after iteration in step 2, the output result and the self-agent perception information are input into the reinforcement learning model after being aggregated by GNN to obtain the final matching probability distribution; specifically, it includes the following steps:
[0076] Step 3.1: Define the communication topology graph. Each defense self-agent is used as a vertex of the graph, and there are m defense self-agents M = {1, 2,..., m}. The communication relationship between defense self-agents is used as the adjacency matrix E = {e ij , i, j ∈ {1, 2,..., m}}. When defense self-agent i and j can communicate with each other, let e ij= 1, and let e be 0 in other cases, ignoring the link of a node communicating with itself. Finally, the graph is represented as G = [M, E]. ij = 0, ignoring the link of a node communicating with itself. Finally, the graph is represented as G = [M, E].
[0077] Step 3.2: Perform an aggregation operation on each defender agent feature vector, that is, calculate the linear combination of the feature vectors of adjacent nodes. For example, let the feature vector of defender agent j be s j , and the feature vector be s j Specifically, it is the position and velocity information of some attacker agents perceived by each defender agent, as well as the position and velocity information of some other defender agents; the set of all defenders that can communicate with defender agent i is represented by . Then, defender agent i can obtain a new feature vector through the aggregation operation
[0078] Step 3.3: Activate the new feature vector obtained from the aggregation operation through an activation function. After K times of aggregation and activation, output the final feature vector of each vertex.
[0079] Step 3.4: Combine the output result of the prediction model and the final feature vector as the input of the policy network. The policy network is set as a 64-layer multi-layer perceptron. The input dimension is the sum of the feature vector dimension of each vertex and the number of defense objects perceived by the defender agent, and the output dimension is the number of defense objects perceived by the defender agent. The activation function uses the ReLu function. The policy network is trained through a typical reinforcement learning algorithm, that is, the PPO algorithm, and the centralized training and distributed execution method is adopted.
[0080] The reward function of the policy network includes two parts: the first part is the sum of the negative values of the distances between each defense object of each defender agent and the defender agent, and the second part is that all defender agents increase a fixed reward when each attacker is captured.
[0081] The probability distribution of the defense task of each defender agent D i is represented in the following form:
[0082] L i = [P(A 1 |D i ), P(A 2 |D i ),..., P(A n |D i )]
[0083] where A 1 to A n represent the perceived attacker agents; the probability distribution L i satisfies the following constraint conditions:
[0084]
[0085] In the above step 4, according to the sampling defense strategy, it is confirmed that the defense target of each defense agent is, according to the greedy sampling principle, to select one of the perceived attackers as the defense target.
[0086] In the above step 5, according to the position and speed information of the attacker confirmed in step 4, a sliding mode controller is designed, with the position error as the sliding mode surface. Under the condition that the position and speed of the agent are bounded, the sliding mode control gain is designed to achieve the capture of the attacker; the sliding mode controller is expressed as:
[0087]
[0088] where u i represents the control input signal of the defense agent D i , k' represents the gain coefficient of the sliding mode controller, and S di represents the position information of the defense agent D i , represents the position information of the defense target of the defense agent D i confirmed in step 4.
[0089] So far, all steps are completed. The present invention studies how to achieve multi-agent boundary defense based on the matching idea. This algorithm can control the intrusion of agents outside the defense areas of multiple agents into the area. For the boundary defense problem of multi-agent systems, the existing mainstream algorithms are to achieve agent control based on the end-to-end method, so as to perform defense task allocation, that is, to assign a defense target according to the distance between the defender and the nearest enemy. And this method extends on the basis of the existing theory,
[0090] and directly generates a global matching scheme. This matching scheme is based on the perception of enemy agents and the analysis of the global situation, and determines the targets that each defense agent should intercept preferentially. Compared with the traditional method of preferentially intercepting the nearest perceived enemy agent, it can improve the overall defense strategy. In addition, the present invention combines the ideas of supervised learning and reinforcement learning, decomposes the task into two parts, reduces the learning difficulty of the reinforcement learning model, speeds up the convergence speed and improves the final interception rate.
[0091] Compared with the prior art, the specific beneficial technical effects of the present invention are as follows:
[0092] A method combining reinforcement learning and supervised learning is proposed, which reduces the learning difficulty of the reinforcement learning algorithm and makes the strategy easier to converge.
[0093] A defense method based on the matching idea is proposed. This method directly generates a global matching scheme, effectively preventing the problem that multiple defenders are paranoid about intercepting the same attacker.
[0094] Using a graph neural network enables the agent to expand the communication range under limited perception, so the strategy still has good performance after generalization.
[0095] Finally, it should be emphasized that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention should be included within the protection scope of the present invention.
Claims
1. A large-scale multi-agent collaborative game method based on learning and easy to generalize, characterized in that: The specific steps include: Step 1: The attacking agent and the defending agent play a 1VS1 game, sample the game trajectory, collect the position and speed information of both the attacker and the defender, and mark the game results to generate training samples; in the border defense task of the agent cluster, collect the position and speed information of both the attacker and the defender as new samples; Step 2: Design a prediction network for each defense agent based on a multi-layer perceptron; the input of the prediction network is the newly added samples generated in step 1, and the output is the predicted game result; and the prediction network is trained through the training samples and supervised learning method in step 1; Step 3: Design a strategy network for each defense agent based on graph neural network and multi-layer perceptron; the input of the strategy network is the output result of the prediction network, the position and speed information of some attacking agents perceived by each defense agent, and the position and speed information of other defense agents; the output is the probability distribution of the defense task of each defense agent; train the strategy network based on the PPO algorithm, and adopt the method of centralized training and distributed execution; Step 4: Based on the probability distribution of the defense tasks of each defense agent, the defense object of each defense agent is confirmed according to the sampling defense strategy; Step 5: Based on sliding mode control, the defense autonomous agent pursues and captures the confirmed defense object.
2. According to claim 1, a large-scale multi-agent collaborative game method based on learning and easy to generalize is characterized by: The attacking agent and the defending agent play a 1VS1 game, sample the game trajectory, collect the position and speed information of the attacking and defending parties, and mark the game results to generate training samples, specifically: An attacking autonomous agent and a defending autonomous agent are randomly generated in the game area. The attacking autonomous agent runs based on the shortest path rule and active obstacle avoidance rule to avoid the defending autonomous agent and reach the target position O at the same time. i ,The defending agent pursues the attacker according to the sliding mode control algorithm; The position and speed information of attack and defense are recorded in real time, and the results of the game are stored as tags.
3. A large-scale multi-agent collaborative game method based on learning and easy to generalize according to claim 2, characterized in that: The characteristic vector X of the position and speed information of the attack and defense includes the speed of the attacker and the defender, the relative position of the attacker and the defender, and the relative distance of the attacker to the defense boundary, which is specifically expressed as: X=[V a ,V d ,P (a,t) ,D (a,d) ] Where V a With V d Represent the speed of the attacker and defender, P (a,t) Denotes the distance from the attacker to the defense boundary, D (a,d) Indicates the relative position between the defender and the attacker.
4. According to claim 1, a large-scale multi-agent collaborative game method based on learning and easy to generalize is characterized by: The prediction network is a multi-layer perceptron with 64 layers, and the activation function adopts the ReLu function.
5. The method of large-scale multi-agent collaborative game based on learning and easy to generalize according to claim 1 is characterized in that: The prediction network uses the nearest neighbor-based supervised learning algorithm KNN to classify the newly added samples, as follows: Step 2.1, calculate the classification labels based on the training samples, using C and D to represent the sets of successful and failed interceptions respectively; Step 2.2: Select k samples from the interception success set C and the interception failure set D respectively; Determine the average value of the feature vectors of the k samples in the successful interception set and set it as the center point O suc , k samples with O suc The sample data with the farthest distance from O suc The 2-norm of is set to the radius R suc ; Determine the average value of the feature vectors of the k samples in the intercept failure set and set it as the center point O fal , k samples with O fal The sample data with the farthest distance from O fal The 2-norm of is set to the radius R fal ; Step 2.3, according to the center point O suc and radius R suc , and the center point O fal and radius R fal , process the newly added sample data and calculate each newly added sample data x I With the center point O suc and O fal The distance d(x I ,O suc ) and d(x I ,O fal ), respectively with radius R suc and R fal Compare and classify the data, specifically: The feature vector x of the newly added sample I is calculated by the following formula I With the center point O suc and O fal The distance d(x I ,O suc ) and d(x I ,O fal ); d(x I ,O suc )=||x I -O suc ||,d(x I ,O fal )=||x I -O fal || In the formula, x I represents the feature vector of the newly added sample I, ||x I -O suc || represents the 2-norm of the difference between two vectors, and the distance d(x I ,O suc ) and R suc For comparison, if d(x I ,O suc ) is less than R suc This means that the label of the newly added sample I is classified as interception success; the distance d(x I ,O fal ) and R fal For comparison, if d(x I ,O fal ) is less than R fal This means that the label of the newly added sample I is classified as interception failure; in other cases, it is set to temporarily not be classified and the current sample data is discarded.
6. The large-scale multi-agent collaborative game method based on learning and easy to generalize according to claim 1 is characterized in that: The step 3 is specifically as follows: Step 3.1, define the communication topology graph, take each defense agent as a vertex of the communication topology graph, there are m defense agents M = {1, 2, ..., m}; take the communication relationship between defense agents as the adjacency matrix E = {e ij ,i,j∈{1,2,…,m}}, when the defense agents i and j can communicate with each other, let e ij =1, in other cases let e ij =0; finally, the communication topology graph is represented as G=[M,E]; Step 3.2: Aggregate the feature vectors of each defense agent: Let the feature vector of defense agent j be s j , the eigenvector is s j Specifically, the position and speed information of some attacking agents perceived by each defense agent, as well as the position and speed information of other defense agents; the set of all defense agents that can communicate with defense agent i is used To represent it, the defense agent i obtains a new feature vector after aggregation operation Step 3.3: The new feature vector obtained by the aggregation operation is activated by the activation function. After K times of aggregation and activation, the final feature vector of each vertex is output; Step 3.4 combines the output of the prediction model and the final feature vector as the input of the policy network. The policy network is set to a 64-layer multi-layer perceptron. The input dimension is the sum of the feature vector dimension of each vertex and the number of defense objects perceived by the defense self-agent. The output dimension is the number of defense objects perceived by the defense self-agent. The activation function uses the ReLu function.
7. The method of large-scale multi-agent collaborative game based on learning and easy to generalize according to claim 1 is characterized in that: The reward function of the strategy network includes two parts: the first part is the sum of the negative values of the distances between the defense object of each defending agent and the defending agent, and the second part is to add a fixed reward to all defending agents when each attacker is captured.
8. The learning-based and easily generalizable large-scale multi-agent collaborative game method according to claim 1 is characterized in that: Each of the defense agents D i The probability distribution of the defense task is expressed as follows: L i =[P(A1|D i ),P(A2|D i ),...,P(A n |D i )] Where A1 to A n represents the perceived attacking agent; probability distribution L i The following constraints are met:
9. The method of large-scale multi-agent collaborative game based on learning and easy to generalize according to claim 1 is characterized in that: In the step 4, the defense object of each defense agent is confirmed according to the sampling defense strategy, and specifically, according to the greedy sampling principle, one of the perceived attackers is selected as the defense object.
10. The learning-based and easily generalizable large-scale multi-agent collaborative game method according to claim 1, characterized in that: The step 5 is specifically as follows: according to the position and speed information of the attacker confirmed in step 4, a sliding mode controller is designed, the position error is used as the sliding mode surface, and under the condition that the position and speed of the subject are bounded, the sliding mode control gain is designed to achieve the capture of the attacker; the sliding mode controller is expressed as: In the formula, u i Denotes the defensive agent D i The control input signal, k' represents the gain coefficient of the sliding mode controller, S di Denotes the defensive agent D i location information, Indicates the defensive agent D confirmed in step 4 i The location information of the defense object.