A multi-agent reinforcement learning method for collaborative decision-making of multiple combat units

Through the multi-agent enhancement learning method for multi-combat units, the problems of low decision-making synergy and difficult to obtain training samples in the game confrontation of the red and blue sides were solved, and the rapid optimization convergence of the enhanced learning model and the improvement of the collaborative decision-making effect were achieved.

CN114358141BActive Publication Date: 2025-05-06CHINA ACAD OF LAUNCH VEHICLE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111530475.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-14
Publication Date
2025-05-06
Estimated Expiration
2041-12-14

AI Technical Summary

Technical Problem

In the prior art, the decision-making synergy of the red and blue sides against multi-combat units is low, and valuable training samples are difficult to obtain.

Method used

A multi-agent enhancement learning method for collaborative decision-making of multiple combat units is proposed, including establishing a multi-agent enhancement learning model, using the post-hoc target conversion method to increase the number of effective training samples, and model training is carried out by building reward functions and massive simulation game confrontations.

Benefits of technology

The number of positive samples in the game confrontation scenarios of the red and blue sides of the multi-combat unit has been effectively improved, and the rapid optimization and convergence of the enhanced learning intelligent model has been achieved, and the decision-making synergy effect and combat decision-making ability have been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114358141B_ABST
    Figure CN114358141B_ABST
Patent Text Reader

Abstract

A multi-agent reinforcement learning method for collaborative decision-making of multiple combat units includes the following steps: for the red-blue game confrontation scenario, a multi-agent reinforcement learning model is established to realize intelligent collaborative decision-making modeling for multiple combat units; the post-goal conversion method is used to increase the number of effective training samples to achieve the optimization convergence of the multi-agent reinforcement learning model; the reward function is constructed based on the team's global task reward and the specific action reward of each combat unit as feedback information; a variety of opponent strategies are generated according to different combat plans, and the multi-agent reinforcement learning model is trained through massive simulated game confrontation using the reward function. The present invention solves the problems existing in the prior art of low coordination of decision-making of multiple combat units in the red-blue game confrontation and difficulty in obtaining valuable training samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology game confrontation and relates to a multi-agent enhanced learning method. Background Art

[0002] Multi-agent deep reinforcement learning combines the collaborative capabilities of multiple agents with the decision-making capabilities of reinforcement learning to solve the collaborative decision-making problem of multiple units in a cluster. It is an emerging research hotspot and application direction in the field of machine learning. It covers many algorithms, rules, and frameworks, and is widely used in real-world fields such as autonomous driving, energy distribution, formation control, trajectory planning, route planning, and social problems. It has extremely high research value and significance. Relevant foreign research institutions have carried out some preliminary basic technical research on multi-agent deep reinforcement learning. Domestic research on this technology, especially its application in the field of military command, has just begun.

[0003] Most of the current intelligent decision-making algorithms use optimization-based and prior knowledge-based methods. They are aimed at the dynamic optimization problem of multiple combat units in the red-blue game confrontation scenario, but have problems such as low decision-making coordination and difficulty in obtaining valuable training samples. Summary of the invention

[0004] The technical problem solved by the present invention is: to overcome the shortcomings of the prior art, to propose a multi-agent reinforcement learning method for collaborative decision-making of multiple combat units, and to solve the problems existing in the prior art such as low coordination of decision-making of multiple combat units in the red and blue game confrontation and difficulty in obtaining valuable training samples.

[0005] The technical solution of the present invention is: a multi-agent enhanced learning method for collaborative decision-making of multiple combat units, comprising the following steps:

[0006] Step 1: For the red-blue game confrontation scenario, a multi-agent reinforcement learning model is established to realize intelligent collaborative decision-making modeling for multiple combat units;

[0007] The construction process of the multi-agent reinforcement learning model is as follows:

[0008] Build a game confrontation scenario between the red and blue teams;

[0009] Analyze the task characteristics and decision points in the red-blue game confrontation scenario, and determine the state space of collaborative task decision points;

[0010] A multi-agent reinforcement learning model is established for collaborative task decision points.

[0011] The method for determining the state space of collaborative task decision points is as follows:

[0012] The overall situation information of the game confrontation scenario and the local observation information of the combat unit are used as state inputs. Default verification is performed by fixing the values ​​of some state inputs, eliminating useless or counterproductive states, and determining the key state space of the mission decision point.

[0013] Step 2: Use the post-hoc target conversion method to increase the number of effective training samples and achieve optimal convergence of the multi-agent reinforcement learning model;

[0014] The specific method of using the post-target conversion method to increase the number of effective training samples is:

[0015] In each round of iterative training, sample data is selected from the experience pool according to the sampling probability value, and the original task goals that the agent failed to achieve in the sample are changed to a state that it can achieve at a certain moment, constructing effective positive samples for model training.

[0016] The calculation formula for the sampling probability value is as follows:

[0017]

[0018] Among them, p i =|δ i |+ε represents the priority of the i-th sample, δ i represents the time difference error of the i-th sample, ε represents random noise to prevent the sampling probability from being 0; α is used to adjust the priority, and P(i) is the sampling probability of the i-th sample data.

[0019] Step 3: Using the team's global task reward as a benchmark and the specific action rewards of each combat unit as feedback information, construct a reward function;

[0020] The method to construct the reward function is:

[0021] According to the situation information at the end of the task decision sequence, calculate the global task reward R task ;

[0022] According to the execution action sequence of each combat unit, calculate the action reward R of each combat unit i ; i represents the serial number of the combat unit, i=1,2,3,……

[0023] According to the global task reward R task and each combat unit's action bonus R i , calculate the collaborative task decision feedback information of each combat unit in the red and blue game confrontation scenario

[0024] Global Mission Rewards R task There are two categories:

[0025] Mission completion reward refers to the Red side completing the combat mission objectives at the end time;

[0026] Damage bonus refers to the number of blue combat units destroyed by the red side in attacks exceeding the number of damages suffered by the red side itself;

[0027] Mission completion rewards and damage rewards are both double values, with different value distribution ranges.

[0028] Each combat unit's action bonus R i It includes three categories:

[0029] Death reward refers to the Red team's combat unit being destroyed by the Blue team, which is a negative reward;

[0030] Ammunition consumption bonus refers to the amount of ammunition consumed by the Red team's combat units, which is a negative bonus;

[0031] Field of view reward refers to the Red team's combat unit being able to detect the Blue team's situation information, which is a positive reward;

[0032] The red team's death reward, ammunition consumption reward, and field of view reward are all double values, with different value distribution ranges.

[0033] The collaborative task decision feedback information of each combat unit in the red-blue game confrontation scenario The calculation formula is:

[0034]

[0035] Among them, η represents the importance of the team's global task reward, η = 0 means that each combat unit only considers the benefits brought by its own actions, and η = 1 means that only the overall benefits of the team are considered.

[0036] Step 4: Generate multiple opponent strategies based on different combat plans, and use the reward function to train the multi-agent reinforcement learning model through massive simulated game confrontation.

[0037] The specific method of using reward functions to train multi-agent reinforcement learning models through massive simulated game confrontation is:

[0038] A blue team strategy library is built based on different combat plans. Every set training cycle, the red team's online decision-making model is used to expand the blue team's strategy library. The reward function is used to complete the evolutionary training of the red team's multi-agent reinforcement learning model through massive simulated game confrontation.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] 1. The present invention uses post-goal conversion to optimize and select samples obtained online in the red-blue game confrontation scenario and generate valuable training samples, which can effectively increase the number of positive samples in the war game scenario with a large search space and realize the rapid optimization convergence of the enhanced learning intelligent model;

[0041] 2. The present invention combines global task rewards with specific action rewards for each combat unit to calculate the feedback of the reinforcement learning model in real time, which is more suitable for multi-combat unit red and blue game confrontation scenarios and improves the synergy effect of the reinforcement learning model;

[0042] 3. The present invention constructs a blue team strategy library based on different combat plans, and uses massive self-game deduction to complete the evolutionary training of the red team's multi-agent reinforcement learning model. By increasing the diversity of opponent strategies and the difficulty of battle, the combat decision-making ability of the reinforcement learning model can be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a flow chart of the method of the present invention;

[0044] Figure 2 It is a model structure diagram of the present invention;

[0045] Figure 3 It is a schematic diagram of the post-target conversion method of the present invention. DETAILED DESCRIPTION

[0046] The present invention proposes a multi-agent enhanced learning method for collaborative decision-making of multiple combat units. Figure 1 As shown, the steps include:

[0047] The first step is to establish a multi-agent reinforcement learning model for the red-blue game confrontation scenario to realize intelligent collaborative decision-making modeling for multiple combat units.

[0048] The construction process of the multi-agent reinforcement learning model is:

[0049] (1.1) Build a game confrontation scenario between the red and blue teams;

[0050] (1.2) Analyze the task characteristics and decision points in the red-blue game confrontation scenario and determine the state space of the task decision points;

[0051] The specific method of state space design is:

[0052] Before building the decision model, the overall situation information of the game confrontation scenario and the local observation information of the combat unit are used as state inputs. By fixing the values ​​of some state inputs, default verification is performed to eliminate useless or counterproductive states and determine the key state space of the task decision point.

[0053] (1.3) For the collaborative task decision points in step (1.2), a multi-agent reinforcement learning model is established to realize collaborative task decision modeling for the red and blue team multi-combat unit game confrontation scenario.

[0054] In the second step, the post-target conversion method is used to increase the number of effective training samples and achieve rapid optimization convergence of the intelligent model.

[0055] The detailed method of generating effective training samples using the post-target conversion method is as follows:

[0056] In each round of iterative training, sample data is first selected from the experience pool according to the following requirements, and the network model is optimized and trained in the process of enhanced learning to achieve higher training efficiency:

[0057]

[0058] Let p i =|δ i |+ε represents the priority of the i-th sample, δ i represents the time difference error (td-error) of the i-th sample, ε represents random noise to prevent the sampling probability from being 0, α is used to adjust the priority (α=0 represents uniform sampling), and P(i) is the sampling probability of the i-th sample data.

[0059] When conducting network training, training samples are selected from the experience pool according to the sampling probability value, and the original task goals that the intelligent agent in the sample failed to achieve are changed to a state that it can achieve at a certain moment. Valid positive samples are constructed for model training, thereby increasing the number of valuable training samples in the red-blue game confrontation scenario with a large search space, and achieving rapid optimization convergence of the intelligent model.

[0060] The third step is to construct a reward function based on the team's global task reward and the specific action rewards of each combat unit as feedback information.

[0061] The calculation steps of the collaborative task decision reward function in the red-blue game confrontation scenario are as follows:

[0062] (3.1) Calculate the global task reward R based on the situation information at the end of the task decision sequence task ;

[0063] (3.2) According to the execution action sequence of each combat unit, calculate the action reward R of each combat unit i ;

[0064] (3.3) Combine the reward values ​​in steps (3.1) and (3.2) to calculate the collaborative task decision feedback information of each combat unit in the red and blue game confrontation scenario

[0065] Global mission rewards include two categories:

[0066] (3.1.1) Mission completion reward, i.e. the Red Team completes the combat mission objectives at the end time;

[0067] (3.1.2) Damage reward, that is, the number of blue units destroyed by the red side is greater than the number of damages suffered by the red side;

[0068] The mission completion reward and damage reward are both double values, but the value distribution range is different.

[0069] For each combat unit in the red-blue game confrontation scenario, its action rewards include three categories:

[0070] (3.2.1) Death reward, that is, the Red Team’s combat unit is destroyed by the Blue Team, which is a negative reward;

[0071] (3.2.2) Ammunition consumption bonus, i.e. the amount of ammunition consumed by the Red Team’s combat units, is a negative bonus;

[0072] (3.2.3) Vision reward, that is, the Red side combat unit can detect the Blue side situation information, which is a positive reward;

[0073] The red team’s death reward, ammunition consumption reward, and field of view reward are all double values, but the value distribution ranges are different.

[0074] Collaborative mission decision feedback information of each combat unit The specific calculation method is:

[0075]

[0076] Let η represent the importance of the team's global task reward, η = 0 means that each combat unit only considers the benefits brought by its own actions, and η = 1 means that only the overall benefits of the team are considered.

[0077] The fourth step is to generate multiple opponent strategies based on different combat plans, and use the reward function to train the multi-agent reinforcement learning model through simulated game confrontation.

[0078] A blue team strategy library is built based on different combat plans, and it is expanded using the red team's online decision-making model at regular training cycles to increase the diversity of the strategy library and the difficulty of the battle. The evolutionary training of the red team's multi-agent reinforcement learning model is completed through massive simulated game confrontations.

[0079] The scenario of the present invention is a collaborative sequential decision-making problem with delayed rewards. In order to solve the task coordination, the present invention uses multi-agent reinforcement learning to first perform reinforcement learning modeling on each combat unit and then conduct centralized training.

[0080] The algorithm framework of the multi-agent reinforcement learning model is as follows Figure 2 As shown. The model is built on the basis of the DDPG algorithm model. The DDPG algorithm uses the Actor-Critic framework for single-step updates, which is faster than the traditional policy gradient and can solve the sequential decision-making problem in the continuous action space. The model includes the Actor policy network and the Critic evaluation network. The Actor network fits the action strategy of the intelligent agent. Its output is not a single action, but the probability distribution of the selected action π(a|s), which indicates the probability of selecting action a in the current state s; the Critic network fits the action value function Q of the intelligent agent. π (s,a) represents the value of taking action a in the current state s.

[0081] At the same time, the present invention adopts post-goal conversion to increase the number of effective training samples, and constructs feedback values ​​by combining the team's global task rewards and the specific action rewards of each combat unit, thereby improving the training efficiency and synergy effect of the multi-agent reinforcement learning model.

[0082] The input of the policy network of the multi-agent reinforcement learning model is the real-time state of the combat scene, that is, the position and survival status of each combat unit of the red side, the observed position of the combat unit of the blue side, terrain information and simulation time. The output of the network is the position of the mobile target and the choice of attack target of the red side. In this scenario, it is hoped that under the premise of limited time, the mapping relationship from state to action can be established through neural network training, and the mobile and firepower allocation plan can be quickly generated online by reinforcement learning method. The input of the evaluation network is the real-time state of the scene and the action information of each agent, and the output is the joint Q value.

[0083] According to the method of the present invention, the global task reward of the team and the specific action reward of each combat unit are calculated to obtain the feedback value of the reinforcement learning training.

[0084] Taking the land warfare game confrontation scenario as an example, the reward for completing the task is to occupy the control point. If the number of blue combat units destroyed by the red side is greater than the number of its own casualties at the end, it will receive a damage reward. By counting the survival status of each combat unit at the end, the number of ammunition consumed, and the number of blue units that can be detected, the individual action reward of each combat unit is calculated to obtain the feedback value of the training of each agent in the multi-agent reinforcement learning model.

[0085] The principle of "centralized training-distributed execution" is adopted. During the model training process, the observation information and execution actions of all intelligent agents are known, which are used to train the evaluation network of each intelligent agent. When the model is executed, each intelligent agent strategy network only generates decision actions based on its own local observation information.

[0086] The specific steps of the multi-agent reinforcement learning model training algorithm are as follows:

[0087] 1) Initialize the main network of each strategy and evaluate the main network Policy Target Network and the evaluation target network and experience pool, the target network is a copy of the main network, o i represents the local observation information of each combat unit, s represents the joint situation information, a represents the joint action, and is the main network weight parameter, and is the weight parameter of the target network; i = 1, 2, 3, ...

[0088] 2) Select the action of each agent's current state: N t is a random noise that follows a normal distribution and is used to increase the agent’s exploration ability;

[0089] 3) Execute actions to obtain the corresponding reward values ​​of each agent, and convert the state-action data (O t ,A t ,R t ,O t+1 ) is stored in the experience pool;

[0090] O t ={o 1,t ,o 2,t ,…,o i,t} represents the joint state at time t, A t ={a 1,t ,a 2,t ,…,a i,t} represents the joint action of each agent at time t, R t = {r 1,t ,r 2,t ,…,r i,t} represents the feedback of each agent at time t, O t+1 ={o 1,t+1 ,o 2,t+1 ,…,o i,t+1} represents the joint state at time t+1.

[0091] 4) When the sample size of the experience pool reaches a certain number, the sample data is processed by post-target conversion (samples are selected based on the sampling probability of td-error) to construct effective positive samples for model training. The loss function L of each agent evaluation network is calculated as follows:

[0092]

[0093] E[] represents the expected function; r i,trepresents the reward value of the i-th agent at the t-th moment; γ represents the decay factor, 0≤γ≤1.

[0094] In the process of sampling the model training samples, the present invention prioritizes the sample data stored in the experience pool to increase the probability of valuable samples being sampled and improve the training efficiency. TD-error is used as a measure of sample importance. The higher the value, the greater the gap between the action value estimate of the evaluation network and the action value target value, and the more valuable the training sample is.

[0095] The scenario of the present invention is not a strict sequential decision problem, or in other words, the scenario is a collaborative sequential decision problem that combines the problems of sparse returns and delayed returns. We can regard this problem as a function optimization problem, whose goal is to find the maximum value of the function and the objective function is to maximize the reward of the collaborative combat task. Since the search space of the red and blue game confrontation scenario is large, it is difficult or even impossible to obtain successful samples using the reinforcement learning method of random exploration. According to the method of the present invention, the original task goal that the agent failed to achieve in the sample data is changed to a state that it can reach at a certain moment by using post-goal conversion, and an effective positive sample is constructed for model training.

[0096] The schematic diagram of the subsequent target conversion is as follows Figure 3 As shown in the figure, although the agent fails to reach the desired target position, the sample data is relabeled through the post-target conversion method, and the current failed decision sequence is converted into a successful decision trajectory, which helps the agent accumulate mobility skill experience and ultimately achieves collaborative mobility strategy learning for the target position.

[0097] The modeling and training of intelligent strategy models in the red-blue game confrontation scenario requires data-driven. The present invention simulates the offensive and defensive confrontation game process in the combat scenario, quickly obtains training samples to improve the learning efficiency of the strategy model, and completes the decision-making ability evolution of the red multi-agent enhanced learning model. The specific steps of the game confrontation training method are as follows:

[0098] 1) Before model training begins, multiple opponent strategies are generated offline based on different combat scenarios using prior knowledge to build a blue team strategy library;

[0099] 2) Randomly extract strategy models from the strategy library, and generate training data through red-blue confrontation in the simulation platform for model iterative training;

[0100] 3) Every certain training cycle, the red team strategy model is used to expand the blue team strategy library online;

[0101] 4) Repeat steps 2) to 4) in a loop to achieve evolutionary training of the intelligent model in a game confrontation scenario.

[0102] In the red-blue game confrontation simulation platform, the method of the present invention is verified based on the coordinated mobile strike decision-making ability of the red side to occupy the control point. The test process is as follows:

[0103] 1) Set up appropriate game confrontation combat scenarios;

[0104] 2) Through simulated confrontation, the multi-agent reinforcement learning model is trained and the adaptability of the Red Army's coordinated mobile strike decision model to typical scenarios is verified. If the model training does not converge, the parameters are adjusted and retrained until the model converges to the next step;

[0105] 3) Conduct verification tests on the method of the present invention under random scenario assumptions;

[0106] 4) In the same typical combat scenario as step 3), each combat unit of the Red Army adopts a single-agent reinforcement learning model. After the model training converges, a verification test is conducted on the model.

[0107] 5) In the same typical combat scenario as step 3), conduct multi-agent reinforcement learning model training and verification experiments without post-target switching;

[0108] 6) The test results of step 3), step 4) and step 5) were statistically compared and analyzed, and it was found that the present invention can well solve the problems of low coordination of decision-making in traditional multi-combat unit game confrontation and difficulty in obtaining valuable training samples.

[0109] Aiming at the requirements of the army's tactical-level task planning, the present invention uses a multi-agent reinforcement learning model to make decisions on the coordinated mobile strike action sequence of the red side in the red-blue game confrontation scenario; uses a sample generation method of post-target conversion to improve the efficiency of reinforcement learning and exploration capabilities, and achieve rapid convergence of the intelligent model; constructs an evaluation parameter that considers the team's global task reward and the specific action reward of each combat unit, and uses this parameter as feedback to effectively improve the collaborative decision-making effect of the intelligent model; uses a massive game confrontation training method to generate a variety of blue side combat strategies offline and online, and quickly generates training samples through massive confrontation deductions to achieve the evolution and upgrade of the combat capability of the intelligent model; in the red-blue game confrontation deduction simulation platform, the effectiveness of the present invention is verified based on the coordinated mobile strike decision-making ability of the red side to occupy the control point. The present invention solves the problems existing in the prior art of low coordination of decision-making of multiple combat units in the red-blue game confrontation and difficulty in obtaining valuable training samples.

[0110] The contents not described in detail in the specification of the present invention belong to the common knowledge of those skilled in the art.

Claims

1. A multi-agent reinforcement learning method for collaborative decision-making of multiple combat units, characterized in that: The steps include: For the red-blue game confrontation scenario, a multi-agent reinforcement learning model is established to realize intelligent collaborative decision-making modeling for multiple combat units; The post-hoc target conversion method is used to increase the number of effective training samples and achieve optimal convergence of the multi-agent reinforcement learning model; The reward function is constructed based on the team's global task reward and the specific action reward of each combat unit as feedback information; Generate multiple opponent strategies based on different combat plans, and use reward functions to train multi-agent reinforcement learning models through simulated game confrontation; The construction process of the multi-agent reinforcement learning model is as follows: Build a game confrontation scenario between the red and blue teams; Analyze the task characteristics and decision points in the red-blue game confrontation scenario, and determine the state space of collaborative task decision points; A multi-agent reinforcement learning model is established for collaborative task decision points; The method for determining the state space of collaborative task decision points is as follows: The overall situation information of the game confrontation scenario and the local observation information of the combat unit are used as state inputs. By fixing the values ​​of some state inputs, default verification is performed to eliminate useless or counterproductive states and determine the key state space of the mission decision point. The specific method of using the post-target conversion method to increase the number of effective training samples is: In each round of iterative training, sample data is selected from the experience pool according to the sampling probability value, and the original task goals that the agent failed to achieve in the sample are changed to the state that it can achieve at a certain moment, so as to construct effective positive samples for model training; The calculation formula for the sampling probability value is as follows: Among them, p i =|δ i |+ε represents the priority of the i-th sample, δ i represents the time difference error of the i-th sample, ε represents random noise to prevent the sampling probability from being 0; α is used to adjust the priority, and P(i) is the sampling probability of the i-th sample data; The specific method of using the reward function to train the multi-agent reinforcement learning model through simulated game confrontation is: A blue team strategy library is constructed according to different combat plans. Every set training cycle, the red team's online decision-making model is used to expand the blue team strategy library. The reward function is used to complete the evolutionary training of the red team's multi-agent reinforcement learning model through simulated game confrontation.

2. The multi-agent enhanced learning method for collaborative decision-making of multiple combat units according to claim 1, characterized in that: The method to construct the reward function is: According to the situation information at the end of the task decision sequence, calculate the global task reward R task ; According to the execution action sequence of each combat unit, calculate the action reward R of each combat unit i ; i represents the serial number of the combat unit, i=1,2,3,…… According to the global task reward R task and each combat unit's action bonus R i , calculate the collaborative task decision feedback information R of each combat unit in the red and blue game confrontation scenario agenti .

3. The multi-agent enhanced learning method for collaborative decision-making of multiple combat units according to claim 2, characterized in that: Global Mission Rewards R task There are two categories: Mission completion reward refers to the Red side completing the combat mission objectives at the end time; Damage bonus refers to the number of blue combat units destroyed by the red side's strikes exceeding the number of damages suffered by the red side itself; Mission completion rewards and damage rewards are both double values, with different value distribution ranges.

4. The multi-agent enhanced learning method for collaborative decision-making of multiple combat units according to claim 2, characterized in that: Each combat unit's action bonus R i There are three categories: Death reward refers to the Red team's combat unit being destroyed by the Blue team, which is a negative reward; Ammunition consumption bonus refers to the amount of ammunition consumed by the Red team's combat units, which is a negative bonus; Vision bonus: refers to the Red team's combat unit being able to detect the Blue team's situation information, which is a positive reward; The red team's death reward, ammunition consumption reward, and field of view reward are all double values, with different value distribution ranges.

5. The multi-agent enhanced learning method for collaborative decision-making of multiple combat units according to claim 2, characterized in that: The collaborative task decision feedback information R of each combat unit in the red-blue game confrontation scenario agenti The calculation formula is: R agenti =ηR task +(1-n)R i Among them, η represents the importance of the team's global task reward, η = 0 means that each combat unit only considers the benefits brought by its own actions, and η = 1 means that only the overall benefits of the team are considered.

Citation Information

Patent Citations

  • Method for training intelligent agent with dynamic award example samples based on reinforcement learning

    CN111582311A

  • Improved PMADDPG multi-unmanned aerial vehicle task decision-making method based on transfer learning

    CN111859541A