An Integrated Cooperative Multi-Agent Deep Reinforcement Learning Method

By integrating multiple policy networks or action value networks for each agent in the cooperative multi-agent system, and using an integrated method of action confidence weight and optimal action voting, the problem of insufficient robustness of the agent's decision-making is solved, significantly improving the task success rate.

CN116468107BActive Publication Date: 2025-06-10GUILIN UNIV OF ELECTRONIC TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310453463.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-25
Publication Date
2025-06-10
Estimated Expiration
2043-04-25

AI Technical Summary

Technical Problem

In cooperative multi-agent systems based on deep reinforcement learning, the robustness of agent decisions is insufficient, resulting in decision failure in some states, affecting the task success rate.

Method used

Multiple policy networks or action value networks are integrated for each agent, so that they make decisions based on the outputs of multiple networks integrated, and network diversity is ensured through different sample training. Specific methods include an integration method based on action confidence weights (ACW) and an integration method based on optimal action voting (OAV).

Benefits of technology

It effectively improves the robustness of agent decision-making, ensures that agents can make correct decisions in multiple states in complex environments, improves task success rate, and is applicable to actor critic-based methods and value-based function decomposition methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116468107B_ABST
    Figure CN116468107B_ABST
Patent Text Reader

Abstract

The present invention discloses an integrated cooperative multi-agent deep reinforcement learning method, which includes the following steps: Step 1, initialize the actor-critic network or the action-value network; Step 2, obtain local observations; Step 3, the cooperative multi-agent system makes decisions in the environment; Step 4, extract transition samples; Step 5, train the actor-critic network or the action-value network; Step 6, repeat Steps 2-5 until the training ends. This method integrates multiple policy networks or action-value networks for each agent, enabling the agent to make decisions based on the output integrated by multiple policy networks or action-value networks, so as to improve the robustness of the agent's decision-making. This method trains the integrated multiple policy networks or action-value networks using different samples for training, ensuring their diversity. It effectively improves the robustness of the agent's decision-making and also has good applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of deep reinforcement learning and cooperative multi-agent, and specifically to a cooperative multi-agent deep reinforcement learning method based on integration. Background Art

[0002] The cooperative multi-agent technology based on deep reinforcement learning has great potential in dealing with sequential decision-making tasks. Many problems in reality can be regarded as sequential decision-making problems of cooperative multi-agent systems. For example, in urban traffic signal control, each agent controls a traffic signal, and all traffic signals on the entire road network are a cooperative multi-agent system. Many studies have shown through simulation experiments (such as PressLight, MPLight, etc.) that the cooperative multi-agent technology based on deep reinforcement learning can effectively improve the effect of urban traffic signal control. In addition, the cooperative multi-agent technology based on deep reinforcement learning can also be used to solve problems such as robot control, power system optimization, and autonomous driving. Therefore, the cooperative multi-agent technology based on deep reinforcement learning is receiving increasing attention. However, in real-world scenarios, the environment is complex and variable, and multiple agents learn simultaneously in the same environment, which is prone to environmental non-stationarity problems, leading to adverse effects on the decisions of cooperative multi-agents in some states. On the other hand, since many methods adopt an experience replay mechanism, the samples for each training are randomly drawn from the experience buffer, which may result in insufficient training of the transition samples in some states, and further lead to the failure of the decisions of cooperative multi-agents in these states. This shows that there is a problem of insufficient robustness in the decisions of cooperative agent systems.

[0003] Cooperative multi - agents based on deep reinforcement learning can be mainly divided into methods based on actor - critic and methods based on value - function decomposition. Methods based on actor - critic (such as MADDPG, COMA, MAAC, etc.) train a local actor network and a centralized critic network for each agent. The actor network is used for the agent's decision - making, and the critic network is used to guide the training of the actor network. The advantage of methods based on actor - critic is that they can handle continuous actions and stochastic policy tasks because the agent makes decisions by sampling according to the action probability distribution output by the actor network. Its disadvantage is that it is easy to fall into a sub - optimal policy and there is a high - variance problem. Methods based on value - function decomposition (such as VDN, QMIX, QTRAN, QPLEX, etc.) train a local action - value function for each agent and approximate a global action - value function through the local action - value functions of all agents. The local action - value function is used for each agent's decision - making, and the approximated global action - value function is used to update the parameters of the neural network in reverse. The advantage of methods based on value - function decomposition is that agents can directly make decisions using the action - value function trained with global information, and the disadvantage is that they cannot handle continuous actions and stochastic policy tasks.

[0004] However, in cooperative multi - agents based on deep reinforcement learning, whether it is the method based on actor - critic or the method based on value - function decomposition, it is impossible to avoid the problem of insufficient robustness of agent decision - making in the cooperative agent system. In some sequential decision - making tasks, the failure of the cooperative multi - agent system to make decisions in certain states is likely to lead to the failure of the entire task. For example, in autonomous driving, this problem may lead to traffic accidents. Therefore, it is necessary to optimize the decision - making of agents in the cooperative agent system. Summary of the Invention

[0005] The object of the present invention is to propose an integrated cooperative multi - agent deep reinforcement learning method for the problem of insufficient robustness of agent decision - making in current cooperative multi - agents based on deep reinforcement learning. This method integrates multiple policy networks or action - value networks for each agent, enabling the agent to make decisions based on the output integrated by multiple policy networks or action - value networks, so as to improve the robustness of agent decision - making. The method trains the integrated multiple policy networks or action - value networks using different samples for training, ensuring their diversity. It effectively improves the robustness of agent decision - making and also has good applicability. It can be used both for methods based on actor - critic and for methods based on value - function decomposition.

[0006] The technical solution for achieving the object of the present invention is as follows:

[0007] An Ensemble-based Decision Optimization for Multi-Agent Deep Reinforcement Learning (EDO for short), the main idea and overall model of the method are as follows:

[0008] 1. Main idea: Integrate multiple policy networks or action-value networks for each agent, so that the agent can make decisions based on the action probability distribution or action value integrated by multiple networks; that is:

[0009] Inside each agent, m policy networks or value networks are integrated. The role of the integration module is to integrate the outputs of the m policy networks or value networks into an overall action probability distribution or action value;

[0010] 2. Overall model: The process of cooperative multi-agent interacting with the environment is as follows: Each agent i obtains its own local observation according to the global state of the environment. Each agent i inputs the local observation into its m policy networks or value networks respectively, and outputs m action probability distributions P i j (o i |·) or action values Then use the ACW or OAV integration method to integrate the m action probability distributions or action values, and obtain the integrated action probability distribution P i (o i |·) or action value Q i (o i |·). Each agent i performs action sampling according to the integrated action probability distribution P i (o i |·) or makes ε-greedy action selection according to the integrated action value Q i (o i |·). The actions selected by all agents form a joint action to transfer the environment to the next state. Each agent obtains the team reward feedback from the environment. Repeat the above process until the end of this round;

[0011] In the EDO model, the m policy networks or value networks integrated by each agent are trained using the transition samples <s, a, r, s'> in the same experience buffer pool. However, each policy network or value network independently randomly samples its own transition samples for training from the experience buffer pool. In the implementation of multi-agent deep reinforcement learning algorithms, the typical capacity of the experience buffer pool is 5000, that is, the experience buffer pool can store at most 5000 transition samples. When the free capacity of the experience buffer pool is not enough to store the latest transition samples collected by the agent, the oldest part of the transition samples in the experience buffer pool is deleted to free up the free capacity. The typical number of randomly sampled transition samples each time is 32, that is, 32 transition samples are randomly sampled from the experience pool for training each time. Since each time during training, the m policy networks or value networks independently sample 32 samples from the 5000 samples in the experience pool for training, the samples used by each policy network or value network during training are likely to be different. Because the m policy networks or value networks each time use different transition samples for training, the parameters of these policy networks or value networks are different, that is, they have diversity;

[0012] Since there is diversity among the m policy networks or value networks, in a certain state, when a certain policy network or value network cannot output the correct action probability distribution or action value, the agent makes normal decisions based on the correct information output by other policy networks or value networks;

[0013] In the integration module of the EDO model, two integration methods are designed: one is the Ensemble Method Based on Action Confidence Weights (ACW for short), and the other is the Ensemble Method Based on Optimal Action Voting (OAV for short):

[0014] In the ACW integration method, the concept of Action Confidence is used. In reinforcement learning, the meaning of Action Confidence is the degree of confidence of the agent when making decisions based on the action probability distribution or action value. When the agent makes a decision, if the preference for a certain action is more obvious, the Action Confidence in this state is greater. If the preference for a specific action is less obvious, the Action Confidence in this state is greater. In the value-based method, the Action Confidence is defined as ψ(Q(o|·)):

[0015] ψ(Q(o|·)) = Q max (o|·) - Q second (o|·) (1)

[0016] Among them, Q(o|·) represents the action values corresponding to all actions under the local observation o, and Q max (o|·) represents the maximum action value under the local observation o, and Q second (o|·) represents the second largest action value under the local observation o. In the policy-based method, the action confidence is defined as ψ(P(o|·))

[0017] ψ(P(o|·)) = P max (o|·) - P second (o|·) (2)

[0018] Among them, P(o|·) represents the action probabilities corresponding to all actions under the local observation o, and P max (o|·) represents the maximum action probability under the local observation o, and P second (o|·) represents the second largest action probability under the local observation o;

[0019] The ACW integration method uses the action confidence of each policy network or value network as its weight, and obtains the integrated action probability distribution P ACW (o|·) or the action value Q ACW (o|·):

[0020]

[0021]

[0022] Among them, m represents the number of policy networks or value networks integrated by the agent;

[0023] In the OAV integration method, the m policy networks or value networks integrated by each agent vote according to the optimal action, and the optimal action a * is defined as the action with the maximum action value or the maximum action probability. In the output of each policy network or value network, there is an optimal action. The idea of the OAV integration method is that if the optimal actions of multiple policy networks or value networks of the agent are the same action then it is considered that choosing this action can bring the maximum future expected reward to the agent, and the action needs to be selected However, in the early stage of training, each policy network or value network has not been fully trained, and the optimal actions output by them have a large randomness. The earlier the training stage, the greater the randomness of the optimal actions output; therefore, it cannot be completely based on for decision-making throughout the training process, but rather as the number of training steps increases, the importance of should be gradually increased;

[0024] Meanwhile, the OAV integration method also needs to consider other situations. When the number of integrations m is large and there are multiple in the same state, only consider the policy network or value network with the most corresponding In addition, when the optimal actions of all integrated policy networks or value networks in a certain state are different, the OAV integration method uses the sum of the outputs of all policy networks or value networks for integration. After analysis, it can be known that the result obtained by using the OAV integration method to integrate m policy networks is:

[0025]

[0026] where m represents the number of policy networks integrated by the agent, P i (o|·) represents the action probability distribution output by the corresponding policy network, k represents the number of corresponding policy networks, P j (o|·) represents the action probability output by the policy network whose optimal action is not , t cur represents the current training step, t max represents the set maximum training step. The result obtained by using the OAV integration method to integrate m policy networks is Q OAV (o|·):

[0027]

[0028] where m represents the number of policy networks integrated by the agent, Q i (o|·) represents the action value distribution output by the corresponding policy network, k represents the number of corresponding policy networks, Q j (o|·) represents the action value whose optimal action is not , t cur represents the current training step, t max represents the set maximum training step;

[0029] The method includes the following steps:

[0030] Step 1, Initialization of the actor-critic network or action value network: Integrate m actor-critic networks or action value networks for each agent and randomly initialize the network parameters;

[0031] Step 2, Obtain local observations: Each agent obtains its own local observations from the environment;

[0032] Step 3. Decision-making of the cooperative multi-agent system in the environment: Each agent inputs its local observation into m policy networks or action-value networks, and uses the integration method based on action confidence weight or the integration method based on optimal action voting to integrate the outputs of the m policy networks or action-value networks into an overall action probability distribution or action value. The agent performs action sampling according to the overall action probability distribution or makes ε-greedy action selection according to the action value. The cooperative multi-agent system obtains the reward feedback from the environment and transfers to the next state. Repeat the above steps until the end of the episode. The specific calculation methods are shown in formulas (1), (2), (3), (4), (5) and (6);

[0033] Step 4. Extract transition samples: Extract transition samples from the experience buffer for the m actor-critic networks or action-value networks integrated for each agent. The transition samples used for each actor-critic network or action-value network are independently and randomly extracted;

[0034] Step 5. Train the actor-critic networks or action-value networks: Use the extracted transition samples to train all the training actor-critic or action-value networks in turn;

[0035] Step 6. Repeat Step 2 - Step 5 until the training is completed.

[0036] Compared with the prior art, the technical solution has the following characteristics:

[0037] 1. By integrating multiple policy networks or action-value networks for each agent in the cooperative multi-agent system, this technical solution can effectively improve the robustness of agent decision-making.

[0038] 2. Based on the characteristics of reinforcement learning, this technical solution proposes two integration methods, the integration method based on action confidence weight (ACW) and the integration method based on optimal action voting (OAV), which can effectively integrate the outputs of multiple policy networks or action-value networks into an overall output for agent decision-making.

[0039] 3. This technical solution has wide applicability and can be applied to all methods in the field of cooperative multi-agents based on deep reinforcement learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a schematic diagram of the architecture and agent decision-making of the embodiment;

[0041] Figure 2 It is an architecture diagram implemented by the embodiment based on the actor-critic method MADDPG;

[0042] Figure 3The architecture diagram implemented in the embodiment based on the value function decomposition method QMIX;

[0043] Figure 4 The schematic diagram of the field of view and shooting range of the agent in the embodiment;

[0044] Figure 5 The schematic diagram of the corridor environment in the embodiment;

[0045] Figure 6 The schematic diagram of the predator-prey environment in the embodiment;

[0046] Figure 7 The schematic diagram of the physical fraud environment in the embodiment. Detailed implementation manners

[0047] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, but it is not a limitation of the present invention.

[0048] Embodiment:

[0049] Refer to Figure 1 , an integrated cooperative multi-agent deep reinforcement learning method, namely EDO. The main idea and overall model of the method are as follows:

[0050] 1. Main idea: Integrate multiple policy networks or action value networks for each agent, so that the agent can make decisions based on the action probability distribution or action value integrated by multiple networks; that is:

[0051] m policy networks or value networks are integrated inside each agent, and the role of the integration module is to integrate the outputs of the m policy networks or value networks into an overall action probability distribution or action value;

[0052] 2. Overall model: The process of the cooperative multi-agent interacting with the environment is as follows: Each agent i obtains its own local observation according to the global state of the environment. Each agent i inputs the local observation into its own m policy networks or value networks respectively, and outputs m action probability distributions P i j (o i |·) or action value Then, the ACW or OAV integration method is used to integrate the m action probability distributions or action values, and the integrated action probability distribution P i (o i |·) or action value Q i (o i |·) is obtained. Each agent i samples actions according to the integrated action probability distribution P i (o i |·) or makes decisions according to the integrated action value Qi (o i Perform ε-greedy action selection. The actions selected by all agents form a joint action that causes the environment to transition to the next state. Each agent receives the team reward feedback from the environment. Repeat the above process until the end of this round;

[0053] In the EDO model, the m policy networks or value networks integrated by each agent are trained using the transition samples <s, a, r, s'> in the same experience buffer pool. However, each policy network or value network independently randomly extracts its own transition samples for training from the experience buffer pool. In the implementation of multi-agent deep reinforcement learning algorithms, the typical capacity of the experience buffer pool is 5000, that is, the experience buffer pool can store at most 5000 transition samples. When the remaining capacity of the experience buffer pool is not enough to store the latest transition samples collected by the agent, the oldest part of the transition samples in the experience buffer pool is deleted to free up the remaining capacity. The typical number of randomly extracted transition samples each time is 32, that is, 32 transition samples are randomly extracted from the experience pool for training each time; Since each time during training, the m policy networks or value networks independently extract 32 samples from the 5000 samples in the experience pool for training, the samples used for training by each policy network or value network are likely to be different. Because the m policy networks or value networks each take different transition samples for training every time, the parameters of these policy networks or value networks are different, that is, they have diversity;

[0054] Since there is diversity among the m policy networks or value networks, in a certain state, when a certain policy network or value network cannot output the correct action probability distribution or action value, the agent makes a normal decision based on the correct information output by other policy networks or value networks;

[0055] In the integration module of the EDO model, two integration methods are designed: one is the integration method based on action confidence weight, namely ACW, and the other is the integration method based on optimal action voting, namely OAV:

[0056] In the ACW integration method, the concept of action confidence is used. In reinforcement learning, the meaning of action confidence is the degree of confidence of the agent when making a decision based on the action probability distribution or action value. When the agent makes a decision, if the preference for a certain action is more obvious, the action confidence in this state is greater. If the preference for a specific action is less obvious, the action confidence in this state is greater. In the value-based method, the action confidence is defined as ψ(Q(o|·)):

[0057] ψ(Q(o|·)) = Q max (o|·) - Q second(o|·) (1)

[0058] Among them, Q(o|·) represents the action values corresponding to all actions under the local observation o, and Q max (o|·) represents the maximum action value under the local observation o, and Q second (o|·) represents the second-largest action value under the local observation o. In the policy-based method, the action confidence is defined as ψ(P(o|·))

[0059] ψ(P(o|·)) = P max (o|·) - P second (o|·) (2)

[0060] Among them, P(o|·) represents the action probabilities corresponding to all actions under the local observation o, and P max (o|·) represents the maximum action probability under the local observation o, and P second (o|·) represents the second-largest action probability under the local observation o;

[0061] The ACW integration method uses the action confidence of each policy network or value network as its weight, and obtains the integrated action probability distribution P ACW (o|·) or the action value Q ACW (o|·):

[0062]

[0063]

[0064] Among them, m represents the number of policy networks or value networks integrated by the agent;

[0065] The following demonstrates the ACW integration method through a specific example: Assume that agent i integrates 3 action value networks, and there are 4 actions to choose from each time a decision is made. In a certain state, the local observation of agent i is o, and the action values output by the 3 action value networks are respectively: Q 1 (o|·) = [0.1, 0.5, 0.2, 0.2], Q 2 (o|·) = [0.2, 0.6, 0.1, 0.1], Q 3 (o|·) = [0.4, 0.3, 0.1, 0.2]. Under the above assumptions, according to formula (1), the action confidences of the 3 action value networks can be calculated as: ψ 1 (Q 1 (o|·)) = 0.5 - 0.2 = 0.3, ψ 2 (Q 2 (o|·)) = 0.6 - 0.2 = 0.4, ψ3 (Q 3 (o|·)) = 0.4 - 0.3 = 0.1, and then according to formula (4), the integrated action value can be calculated as:

[0066] Q ACW (o|·) = 0.3×Q 1 (o|·) + 0.3×Q 2 (o|·) + 0.1×Q 3 (o|·) = [0.13, 0.36, 0.1, 0.11]

[0067] If the agent i makes an action selection according to ε - greedy, the agent i will randomly select an action with a probability of ε and select the action 2 with the maximum action value with a probability of 1 - ε. In this example, when the agent makes a decision based on 3 action value networks simultaneously, it will probably select action 2. If the agent makes a decision only based on the Q 3 action value network, then it will probably select action 1. It can be seen that EDO changes the decision of the agent. If the agent integrates multiple policy networks, the calculation process of the ACW integration method is similar to the above process;

[0068] In the OAV integration method, the m policy networks or value networks integrated by each agent vote according to the optimal action, and the optimal action a * is defined as the action with the maximum action value or the maximum action probability. In the output of each policy network or value network, there exists an optimal action. The idea of the OAV integration method is that if the optimal actions of multiple policy networks or value networks of the agent are all the same action then it is considered that selecting this action can bring the maximum future expected reward to the agent, and the action needs to be selected However, in the early stage of training, each policy network or value network has not been fully trained, and the optimal actions output by them have relatively large randomness. The earlier the training stage is, the greater the randomness of the optimal actions output; Therefore, it cannot be completely based on for decision - making throughout the training process, but rather as the number of training steps increases, the importance of should be gradually increased;

[0069] At the same time, the OAV integration method also needs to consider other situations. When the integration number m is large and there are multiple in the same state, only the corresponding to the most policy networks or value networks is considered. For example, if the agent i integrates 7 policy networks, under a certain local observation, 4 of them have the maximum action probability for a 1 and another 2 policy networks have the maximum action probability for a 3The corresponding action probability is the largest. In this case, a 1 will be used as In addition, when the optimal actions of all integrated policy networks or value networks in a certain state are different, the OAV integration method uses the sum of the outputs of all policy networks or value networks for integration. After analysis, it can be known that when using the OAV integration method to integrate m policy networks, the result obtained is:

[0070]

[0071] where m represents the number of policy networks integrated by the agent, and P i (o|·) represents the action probability distribution output by the corresponding policy network, k represents the number of corresponding policy networks, and P j (o|·) represents the action probability output by the policy network whose optimal action is not , t cur represents the current training step, and t max represents the set maximum training step. When using the OAV integration method to integrate m policy networks, the result obtained is Q OAV (o|·):

[0072]

[0073] where m represents the number of policy networks integrated by the agent, and Q i (o|·) represents the action value distribution output by the corresponding policy network, k represents the number of corresponding policy networks, and Q j (o|·) represents the action value whose optimal action is not , t cur represents the current training step, and t max represents the set maximum training step;

[0074] The method includes the following steps:

[0075] Step 1, Initialization of the actor-critic network or action value network: Integrate m actor-critic networks or action value networks for each agent and randomly initialize the network parameters;

[0076] Step 2, Obtain local observations: Each agent obtains its own local observations from the environment;

[0077] Step 3. Decision-making of the cooperative multi-agent system in the environment: Each agent inputs its local observation into m policy networks or action-value networks, and uses the integration method based on action confidence weight or the integration method based on optimal action voting to integrate the outputs of the m policy networks or action-value networks into an overall action probability distribution or action value. The agent performs action sampling according to the overall action probability distribution or makes an ε-greedy action selection according to the action value. The cooperative multi-agent system obtains the reward feedback from the environment and transfers to the next state. Repeat the above steps until the end of the episode. The specific calculation methods are shown in formulas (1), (2), (3), (4), (5), and (6);

[0078] Step 4. Extract transition samples: Extract transition samples from the experience buffer for the m actor-critic networks or action-value networks integrated for each agent. The transition samples for each actor-critic network or action-value network are independently and randomly extracted;

[0079] Step 5. Train the actor-critic networks or action-value networks: Use the extracted transition samples to train all the training actor-critic or action-value networks in turn;

[0080] Step 6. Repeat Steps 2 - 5 until the training is completed.

[0081] The following introduces the experimental performance of applying EDO to two representative algorithms (MADDPG and QMIX). Figure 2 and Figure 3 respectively show the implementations of EDO on the actor-critic method MADDPG and the value function decomposition method QMIX. Since EDO includes two integration methods, ACW and OAV, the methods tested in the experiment include a total of 6 kinds: ACW-MADDPG, OAV-MADDPG, ACW-QMIX, OAV-QMIX, MADDPG, and QMIX. Among them, ACW-MADDPG and OAV-MADDPG mean combining the EDO method with the ACW and OAV integration methods respectively on the basic framework of the actor-critic-based method, and ACW-QMIX and OAV-QMIX mean combining the EDO method with the ACW and OAV integration methods respectively on the basic framework of the value function decomposition-based method. A total of 3 experimental environments are selected from the StarCraft Multi-Agent Challenge experimental environment (SMAC) and the Multi-Agent Particle experimental environment (MPE) for testing, including: corridor, Physical deception, and Predator-prey experimental environments.

[0082] The StarCraft Multi-Agent Challenge (SMAC) is based on the popular real-time strategy game StarCraft II. In the regular full game of StarCraft II, a player competes against multiple players or the built-in game AI to collect resources, build buildings, and form troops to defeat opponents. Similar to most real-time strategy games, StarCraft II has two main gameplay components: macro management and micro management. Macro management refers to high-level strategic considerations such as economy and resource management. Micro management, on the other hand, refers to the fine-grained control of individual units. To build a rich multi-agent reinforcement learning testbed, SMAC only focuses on micro management, i.e., each unit is independently controlled by an agent that must take actions based on local observations. The local observation is a limited field of view centered on that unit (see Figure 4 ). The goal of the cooperative multi-agent system is to improve the micro-operation and cooperative behavior of agents through training to fight and win against the enemy army centrally controlled by the in-game built-in script AI in challenging combat scenarios. Micro-operation is a very important aspect of the StarCraft II gameplay, with a high skill ceiling, and both amateur and professional players can practice it alone. SMAC provides a set of different experimental scenarios to test the ability of algorithms to handle high-dimensional inputs and the ability of algorithms to learn cooperative behavior in a cooperative multi-agent system under the limitations of partial observability and decentralized execution.

[0083] corridor is an environment in the StarCraft Multi-Agent Challenge experimental environment with a difficulty defined as Super Hard. corridor contains 6 Zealot units, each of which is independently controlled by an agent, and the enemy is 24 Zergling units centrally controlled by the StarCraft II built-in script AI. Both Zealot units and Zergling units have melee attack methods. Each time the environment is initialized, the 24 Zergling units always appear close to each other at a specific location. If there are no Zealot units within the field of view of the Zergling units, they will always move towards one end of the environment until they reach the end. Since both Zealot units and Zergling units use melee attacks in corridor and the number of Zergling units is 6 times that of Zealot units, the multi-agent system must learn to disperse the firepower of the Zergling units to have a chance to win. The corridor environment is as Figure 5 shown.

[0084] Multi-agent particle tasks are a type of simulation task that simplifies agents and some objects in the environment into particles. In such tasks, agents have continuous observations and discrete action spaces and have some basic physical properties such as volume, speed, etc. By setting different agent attributes and goals, etc., a variety of different tasks have been derived.

[0085] Both physical deception and predator-prey belong to the multi-agent particle experimental environment. As Figure 6 shown, in the predator-prey task, there are N cooperative agents with slower movement speeds that need to chase M prey with faster movement speeds in the environment, and there are L large roadblocks blocking the way in the environment. Whenever a cooperative agent collides with a prey, the agent will receive a reward while the prey will be punished. The agent can observe the relative positions and speeds of other agents and prey as well as the positions of roadblocks. The initial positions of the agents, prey, and roadblocks are randomly initialized in each round.

[0086] In physical deception, there are n landmarks in the environment, including a target landmark. There are n cooperative agents that obtain rewards based on the minimum distance of any agent to the target landmark (the closer to the target landmark, the greater the reward), so only one agent needs to reach the target landmark to obtain the maximum reward. At the same time, there is also an opponent in the environment that also hopes to reach the target landmark but doesn't know which of the n landmarks is the target landmark. In addition, the cooperative agents will be punished according to the distance of the opponent to the target landmark, so the cooperative agents need to learn to disperse and cover all landmarks to deceive the opponent. The physical deception environment is as Figure 7 shown.

[0087] Settings of relevant parameters of the experimental environment: The difficulty level of the corridor experimental scenario is set to 7. In the physical deception environment, the number of cooperative agents is 3, the number of roadblocks is 2, and the number of opponents is 1. In the predator-prey environment, the number of cooperative agents is 3, the number of roadblocks is 2, and the number of prey is 1.

[0088] Settings of General Parameters: In the QMIX, ACW-QMIX, and OAV-QMIX algorithms, the agent explores the environment using the ε-greedy exploration strategy, that is, the agent randomly selects an action with a probability of ε and selects the action with the maximum action value with a probability of 1-ε. The initial value of ε is 1 and it decays continuously as the number of training steps increases. The decay stops when training reaches 50,000 steps, and then ε = 0.05 is maintained until the end of training. All algorithms use the experience replay mechanism to save and sample transition samples. The size of the experience buffer pool is 5000, and the number of randomly sampled transition samples each time is 32. The target network of all algorithms is updated once every 200 episodes. The discount factor γ used to balance the immediate reward and the future reward is set to 0.99. When updating the network parameters, the learning rate α is set to 0.0005. In the corridor experimental environment, the maximum number of training steps is 2,000,050, and it is tested every 50,000 training steps. Each time it is tested, the average value of 100 repeated test episodes is taken. In the physical fraud and predator-prey experimental environments, the maximum number of training steps is 2,050,000. It is tested every 10,000 training steps, and each time it is tested, the average value of 100 repeated test episodes is taken. For the experiments of each algorithm in different environments, 3 repeated experiments are carried out with different random seeds, and the average value of the 3 repeated experiments is used for plotting.

[0089] Settings of Hyperparameters of EDO: In the ACW-MADDPG and OAV-MADDPG algorithms, the number of policy networks integrated by each agent is 3, and in the ACW-QMIX and OAV-QMIX algorithms, the number of individual Q networks integrated by each agent is 3.

[0090] Table 1 lists some basic parameter settings in this example:

[0091] Table 1 Experimental Parameter Settings

[0092]

[0093] Table 2 shows the rewards obtained by the multi-agent system and the winning rates achieved when the QMIX, OAV-QMIX, and ACW-QMIX algorithms battle against the built-in script AI in StarCraft II during the entire training process in the corridor experimental environment. The OAV-QMIX and ACW-QMIX algorithms can ultimately reach a reward value above 20, while the original QMIX can only reach a maximum reward value of 13.6. The original QMIX algorithm can only reach a maximum test winning rate of 19%, while both the OAV-QMIX and ACW-QMIX algorithms can reach a test winning rate higher than 80%. The application of EDO has increased the maximum test winning rate of the original QMIX algorithm by more than 60%. From the comparison of the reward values and the winning rates in battles, it can be seen that the application of EDO has significantly improved the performance of the original QMIX algorithm.

[0094] Experimental results in Table 2 in corridor

[0095]

[0096]

[0097] Experimental results in Table 3 in physical fraud

[0098]

[0099] Experimental results in Table 4 in predator-prey

[0100]

[0101] Table 3 and Table 4 respectively show the reward values obtained by the multi-agent system when the MADDPG, OAV-MADDPG, and ACW-MADDPG algorithms are in the physical fraud and predator-prey experimental environments during the entire training process. In the physical fraud experimental environment, both the OAV-MADDPG and ACW-MADDPG algorithms that incorporate the EDO method have achieved reward values higher than those of the basic MADDPG algorithm. In the predator-prey experimental environment, the OAV-MADDPG and ACW-MADDPG algorithms have also widened the gap in reward values compared to the basic MADDPG algorithm. This shows that EDO can effectively improve the performance of the multi-agent system in this experimental environment.

[0102] The above experimental results show that the method in this example (EDO) can effectively solve the problem of insufficient robustness in the decision-making of agents in cooperative multi-agents based on deep reinforcement learning and can improve the stability of the original algorithm. In addition, EDO has strong applicability and can be applied in two mainstream cooperative multi-agent deep reinforcement learning methods (actor-critic-based methods and value function decomposition-based methods), and can be widely used to optimize the decision-making of agents in cooperative multi-agent systems.

Claims

1. An integrated cooperative multi-agent deep reinforcement learning method, characterized in that: The method comprises the following steps: Step 1, initialization of the actor-critic network or the action-value network: Integrate multiple actor-critic networks or action-value networks for each agent, and randomly initialize the network parameters; Step 2, obtain local observations: Each agent obtains its own local observations from the environment; Step 3, decision-making of the cooperative multi-agent system in the environment; Step 4, extract transition samples: Extract transition samples from the experience buffer for the multiple actor-critic networks or action-value networks integrated for each agent. The transition samples for each actor-critic network or action-value network are independently and randomly extracted; Step 5, train the actor-critic network or the action-value network: Use the extracted transition samples to sequentially train all the trained actor-critics or action-value networks; Step 6, repeat Step 2 - Step 5 until the training ends; Among them, multiple policy networks or value networks are integrated inside each agent. The integration module integrates the outputs of these networks into an overall action probability distribution or action value. When the cooperative multi-agent interacts with the environment, each agent obtains local observations according to the global state of the environment, and inputs the local observations into its multiple policy networks or value networks to output multiple action probability distributions or action values. The ACW or OAV integration method is used to integrate the multiple action probability distributions or action values to obtain the integrated action probability distribution or action value. The agent performs action sampling or ε-greedy action selection based on this integrated result. The actions selected by all agents form a joint action to make the environment transfer to the next state. Each agent obtains the team reward feedback from the environment. Repeat the above process until the end of this round; The multiple policy networks or value networks integrated for each agent are trained using the transition samples in the same experience buffer pool, but each network independently and randomly extracts its own transition samples for training from the experience buffer pool. Since the samples used for training each policy network or value network are mostly different, there is diversity among these networks. When a certain policy network or value network cannot output the correct action probability distribution or action value, the agent can make a decision based on the correct information output by other networks.

2. The integrated cooperative multi-agent deep reinforcement learning method according to claim 1, characterized in that, In the ACW integration method, the concept of action confidence is used. In reinforcement learning, the meaning of action confidence is the degree of confidence when the agent makes a decision based on the action probability distribution or action value. When the agent makes a decision, if the preference for a certain action is more obvious, the action confidence in this state is greater. If the preference for a specific action is less obvious, the action confidence in this state is greater. In the value-based method, the action confidence is defined as ψ(Q(o|·)): ψ(Q(o|·)) = Q max (o|·) - Q second (o|·)(1) where Q(o|·) represents the action values corresponding to all actions under the local observation o, and Q max (o|·) represents the maximum action value under the local observation o, and Q second (o|·) represents the second largest action value under the local observation o. In the policy-based method, the action confidence is defined as ψ(P(o|·)) ψ(P(o|·)) = P max (o|·) - P second (o|·)(2) where P(o|·) represents the action probabilities corresponding to all actions under the local observation o, P max (o|·) represents the maximum action probability under the local observation o, P second (o|·) represents the second - largest action probability under the local observation o; The ACW integration method uses the action confidence of each policy network or value network as its weight, and obtains the integrated action probability distribution P by accumulating the action probability distributions or action values multiplied by the weights ACW (o|·) or the action value Q ACW (o|·): where m represents the number of policy networks or value networks integrated for the agent.

3. The integrated cooperative multi-agent deep reinforcement learning method according to claim 1, characterized in that, The OAV integration method also needs to consider other situations. When the integration quantity m is large and there are multiple in the same state, only consider the one corresponding to the most policy networks or value networks. In addition, when the optimal actions of all integrated policy networks or value networks in a certain state are different, the OAV integration method uses the sum of the outputs of all policy networks or value networks for integration. When using the OAV integration method to integrate m policy networks, the result obtained is: Moreover, when all the integrated policy networks or value networks in a certain state have different optimal actions, the OAV integration method uses the sum of the outputs of all policy networks or value networks for integration. When using the OAV integration method to integrate m policy networks, the result obtained is: Among them, m represents the number of policy networks integrated by the agent, and P i (o|·) represents the action probability distribution output by the corresponding policy network, and k represents the number of corresponding policy networks, and P j (o|·) represents that the optimal action is not the action probability output by the policy network, and t cur represents the current training step, and t max represents the set maximum training step. The result obtained by integrating m policy networks using the OAV integration method is Q OAV (o|·): Among them, m represents the number of policy networks integrated by the agent, and Q i (o|·) represents the action value distribution output by the corresponding policy network, and k represents the number of corresponding policy networks, and Q j (o|·) represents that the optimal action is not the action value, and t cur represents the current training step, and t max represents the set maximum training step.

Citation Information

Patent Citations

  • Enemy-friend deep deterministic strategy method and system based on reinforcement learning

    CN112215364A

  • Multi-agent reinforcement learning training method with high sample efficiency

    CN113313209A