A Multi-Agent Reinforcement Learning Method and System Based on Value Decomposition
By designing an efficient value decomposition algorithm in multi-agent reinforcement learning, using parameterized noise and competition-agent network decomposition joint reward value function, the problem of inefficient value decomposition in complex multi-agent environments is solved, and rapid convergence and efficient learning are achieved.
Patent Information
- Application Number
- CN202210301408.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-03-24
AI Technical Summary
In a complex multiagent environment, the combined state-action space of the value function increases exponentially with the increase of the number of agents, resulting in inefficient value decomposition process, difficult to guarantee convergence time, and it is difficult for the agent to effectively perceive the environment and make correct decisions in the early stage of exploration.
A multi-agent reinforcement learning method based on value decomposition is adopted to design an efficient value decomposition algorithm through four networks with different functions (evaluation-agent network, random-agent network, target-agent network and competition-agent network). This algorithm adds parameterized noise to the connection weight of the neural network, allowing the agent to explore new environments or states and speed up exploration efficiency. At the same time, a competition-agent network is introduced to decompose the joint reward value function into a state value function and an advantage value function to reduce the estimation error.
The learning efficiency and reaction ability of agents in multi-agent systems are improved, the error between the estimated joint reward value function and the actual value function is reduced, the convergence speed of the algorithm model is significantly accelerated, and the learning time of agents in the early stages of exploration is reduced.
Smart Images

Figure CN114662639B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to multi-party collaboration processing technology, and in particular to a multi-agent reinforcement learning method and system based on value decomposition. Background Art
[0002] Nowadays, multi-agent reinforcement learning (MARL) is a relatively popular topic. MARL has broad prospects in solving many complex real-world problems such as sensor networks, swarm robotics coordination, and autonomous vehicles. However, in practical applications, MARL faces two major challenges: partial observability and stability. First, when an agent interacts with the environment, the agent cannot observe and make decisions from a global perspective, so it cannot learn the globally optimal policy and can only observe the information within its own field of vision. Second, in a multi-agent environment, agents influence each other. The actions taken by each agent based on its local observations may affect other agents because other agents are also learning, constantly updating their policies, and the actions they take are constantly changing. To address the above problems, in recent years, relevant researchers have proposed a centralized training decentralized execution architecture (CTDE). This architecture allows a single agent to access global information during centralized training to train a joint reward value function Q tot (τ,a), where τ is the global information and a is the action taken by the agent. Such a method takes into account global information and also eliminates the instability between agents. However, due to the partial observability of a single agent, even if Q tot (τ,a) is trained, a single agent cannot obtain the global state and needs to decompose Q tot (τ,a) to individual agents to replace directly learning the global state. Therefore, a series of problems regarding Q totThe method of (τ,a) decomposition is value decomposition. For example, the Chinese patent document with the publication number CN113313267A combines the attention mechanism and value decomposition, and designs a multi-agent reinforcement learning algorithm based on Actor-Critic. The Chinese patent document with the publication number CN12396187A proposes a multi-agent reinforcement learning method based on a dynamic cooperation graph, which solves the problem of the cooperation efficiency between agents in a multi-agent system. The literature "Yang Y, Hao J, Liao B, et al. Qatten: A general framework for cooperative multiagent reinforcement learning[J]. arXiv preprint arXiv:2002.03939, 2020." mentions introducing the multi-head attention mechanism into multi-agent reinforcement learning, providing theoretical support for the previous value decomposition algorithm.
[0003] Google's DeepMind team proposed a value-decomposition network architecture (Value-Decomposition Networks, VDN) based on cooperative MARL (see Sunehag P, Lever G, Gruslys A, et al. Value-Decomposition Networks For Cooperative Multi-Agent Learning Based On Team Reward[C] / / AAMAS. 2018.), which sums up the rewards on each agent to approximately obtain Q tot (τ,a), VDN uses the update method of Deep Q Network (DQN) to update Q through the global reward tot (τ,a), the gradient of a single agent will be backpropagated to the local value function of each agent through Qtot(τ,a) to update them, so that the local reward value function of a single agent can be updated from a global perspective. VDN sums up the local reward value functions of all agents to approximately obtain Q tot (τ,a), that is, it is considered that Q tot (τ,a) and the local reward value function of a single agent is a summation relationship, and Q is approximated by accumulation tot (τ,a). But actually, summation is a very simple relationship. If Q tot (τ,a) and the local reward value function of a single agent is very complex, then VDN will lose its effect.
[0004] The QMIX algorithm (Rashid T, Samvelyan M, Schroeder C, et al. Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning[C] / / International Conference on Machine Learning. PMLR, 2018: 4295-4304.) is a multi-agent reinforcement learning algorithm based on CTDE further proposed on the basis of VDN. This algorithm uses a neural network to approximate Q tot (τ, a), but it is necessary to ensure that all parameters of the neural network must be non-negative, because the neural network hopes to learn local actions that can help itself approximate the global reward. Therefore, the relationship between Q tot (τ, a) and the local reward value function of a single agent must satisfy monotonicity. The situation where the local optimal action selected by each agent is exactly a part of the global optimal action is a special case. Moreover, VDN does not make full use of the advantages of centralized training and ignores any additional state information available during learning. QMIX additionally uses the global state when approximating Q tot (τ, a), so that training can be carried out based on the global state, accelerating the overall training speed. However, whether it is additivity or monotonicity, they actually strictly limit the relationship between Q tot (τ, a) and the local reward value function of a single agent, so that they can only solve a small part of the tasks, because in many tasks, the relationship between Q tot (τ, a) and the local reward value function of a single agent may not be additive or monotonic, so that the approximated Q tot (τ, a) by VDN and QMIX is very different from the real Q tot (τ, a).
[0005] QTRAN (Son K, Kim D, Kang W J, et al. Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning[C] / / International Conference on Machine Learning. PMLR, 2019: 5887-5896.) focuses on releasing the restrictions of additivity and monotonicity to decompose all decomposable tasks. The idea is that as long as it is ensured that the individual optimal action and the joint optimal action are the same, then the local reward value function of a single agent and Qtot (τ, a)'s specific relationship doesn't need to be considered. The QTRAN algorithm directly learns a true global reward and introduces a compensation term to make up for the learned Q tot (τ, a) and the true Q tot (τ, a) to ensure that the learned Q tot (τ, a) and the true Q tot (τ, a) are very close.
[0006] Although the above frameworks all have strong theoretical guarantees, in complex environments such as the StarCraft Multi-Agent Challenge (SMAC) mentioned in the Chinese patent document with publication number CN111632387A, the above algorithms mainly have the following problems:
[0007] In a large-scale collaborative environment, a problem will occur during the process of value function decomposition: the size of the joint state-action space of the value function grows exponentially with the increase in the number of agents, which makes it more difficult to perform value decomposition quickly and effectively, and the convergence time often cannot be guaranteed. The inefficiency of the value decomposition process will have the following two impacts on the agents:
[0008] (1) Due to the complexity of the multi-agent environment, at the initial stage of exploration, agents need to spend a lot of time exploring states that are beneficial to themselves or the system. The number of exploration spaces will increase with the increase in the number of agents. In some multi-agent scenarios with sparse rewards, it is very likely that positive feedback of rewards cannot be obtained for a long time, and agents cannot effectively perceive the scene information and make correct decisions, and the convergence time is difficult to guarantee.
[0009] (2) When agents execute actions according to the policy, if the wrong estimation of several sub-optimal joint actions exceeds the better estimation of a single optimal joint action, it will lead to the underestimation of the reward corresponding to the optimal action, causing agents to choose actions with sub-optimal values, thus resulting in the action evaluation of agents falling into a cycle of local optimality and being unable to decide the best action, prolonging the time for agents to make action decisions.
[0010] Therefore, the key is how to implement a comprehensive and effective value decomposition strategy to achieve fast convergence while reducing the exploration cost. Summary of the Invention
[0011] The purpose of the present invention is to propose a multi-agent reinforcement learning method and system based on value decomposition, which is applied in complex partially observable scenarios, has a relatively fast convergence speed, improves the learning efficiency of agents, and thus improves the reaction ability of multi-agents in complex environments.
[0012] To achieve the above invention purpose, the present invention adopts the following technical solutions:
[0013] In a first aspect, a multi-agent reinforcement learning method based on value decomposition is proposed, including the following steps:
[0014] Obtain the state S of the environment at the current time t t , the initial observation value of each agent Available actions and the reward r corresponding to this action, where i is the serial number of the agent, and the state S t includes the number of agents, role types, and the size of the joint reward Q value function obtained at the previous moment in the current multi-agent scenario;
[0015] For each agent, calculate the value function Q of each action based on the local information τ through the evaluation-agent network i observed i (τ i ), where the local information τ i is the observation value of agent i action reward r and state S t information set;
[0016] Use the random-agent network to add parameterized noise to the state S at the current moment to randomize the weight and bias parameters, and then sum the weights of the trained weight and bias parameters at the end of each round with the Q t (τ i ) based on the local information τ of each agent to obtain the reward value function Q of each agent based on the global information τ i (τ); i ) i (τ);
[0017] The target-agent network calculates the loss function and updates the parameters, and then the random-agent network also updates the noise parameters and calculates the loss function, where the target-agent network is obtained by copying the parameters of the evaluation-agent network at regular intervals;
[0018] Use the competition-agent network to decompose the reward value function Q i (τ) of each agent based on the global information τ into an advantage value function, a state value function, and an action value function;
[0019] Add the decomposition results to obtain the joint reward value function Q tot (τ, a) based on the global information τ, and update the parameters of the competition-agent network and the overall loss function. After the parameter update is completed, use the trained agents to execute actions in the environment.
[0020] Further, the agent includes any one of the following: a hero in a game scenario, a single sensor in a sensor network, a single robot in a robot collaboration scenario, or a car in an autonomous driving scenario.
[0021] Further, the evaluation-agent network includes an MLP for processing local observations and action inputs / outputs, and a GRU recurrent neural network for memorizing historical states and action information: Each agent encodes the local observations and actions and inputs them into the MLP, and then inputs them into the GRU recurrent neural network. The GRU recurrent neural network concatenates the hidden information h at the current time t with the output of the previous layer to generate the hidden information h at the next time t+1 and the input to the next layer. The MLP of the third layer then generates the value function Q for each agent based on the local information t and the output of the previous layer to generate the hidden information h at the next time t+1 and the input to the next layer. The MLP of the third layer then generates the value function Q for each agent based on the local information t+1 (τ i ) that contains the local historical information τ of a single agent i . i
[0022] Further, the random-agent network is a neural network whose weights and biases are randomly perturbed by noise parameters. The noise parameter θ is defined as θ = μ + ∑⊙ε, where μ and ∑ are learnable noise parameter vectors, ε is a vector of zero-mean noise with fixed statistics, and ⊙ represents element-wise multiplication. The output of the noise layer is expressed as y = (μ w +σ w ⊙ε w )x + μ b +σ b ⊙ε b , where x is the state S at the current time as the input, and y is the state S t ′ after being randomly perturbed by noise as the output. The weight parameter term μ t +σ w ⊙ε w is greater than 0.
[0023] Further, the decomposition of the reward value function Q w (τ) for each agent by the competition-agent network includes: inputting the global state S at the current time t into the evaluation-agent network to convert it into a Q-value function that is only affected by the state, i.e., the state value function S i (τ); inputting the action selected by agent i at the current time t into the random-agent network and outputting the action value function C t (τ); subtracting S i (τ) and C i (τ) from the reward value function Q i (τ) for each agent based on the global information τ i (τ) and C i(τ) Obtain the advantage value function A i (τ), sum the S i (τ) of each agent to obtain the joint state value function S tot (τ), sum the C i (τ) to obtain the joint action value function C tot (τ), then multiply the A i (τ) of each agent by a coefficient δ and sum them up to obtain A tot (τ).
[0024] Furthermore, the state value function S i (τ) is only related to the state S at the current time t t or the global information τ, and S tot (τ) is expressed as: The action value function C i (τ) is only related to the action at the current time t or the global information τ, and C tot (τ) is expressed as: The advantage value function A tot (τ) is not only related to the state S at the current time t t but also related to the actions taken by each agent , and A tot (τ) is expressed as: δ > 0, and δ is the parameter weight trained by the random-agent network.
[0025] Furthermore, the joint reward value function Q tot (τ, a) is simplified to the following form:
[0026]
[0027] where A i (τ) = Q i (τ) - S i (τ) - C i (τ).
[0028] Furthermore, the method further includes: storing the global state S at the current time t t and the global state S at the next time t + 1 t+1 , the observation value of each agent the selected action and the reward r into the experience buffer M as a tuple; when the size of the experience buffer M exceeds the specified threshold, take the evaluation-agent network's network parameter update every specified number of times as one episode, group the experience tuples in the experience buffer by episode, and sample from the similar experience tuples in different episodes.
[0029] In a second aspect, a multi-agent reinforcement learning system based on value decomposition is proposed, including:
[0030] An information perception module for obtaining the state S of the environment at the current time t t , the initial observation value of each agent Available actions And the reward r corresponding to the action, where i is the serial number of the agent, and the state S t Includes the number of agents, role types, and the size of the joint reward Q value function obtained at the previous moment in the current multi-agent scenario;
[0031] An evaluation-agent network for calculating, for each agent, the value function Q i Observed based on each action and local information τ i (τ i ), where the local information τ i Is the observation value of agent i Actions Reward r and state S t Information set;
[0032] A random-agent network for adding parameterized noise to the state S at the current moment to randomize the weight and bias parameters, and then summing the weights of the trained weight and bias parameters at the end of each round with the Q t Based on local information τ of each agent i Of i (τ i ) to obtain the reward value function Q i (τ) of each agent based on global information τ;
[0033] A parameter update module, where the target-agent network calculates the loss function and updates the parameters, and then the random-agent network also updates the noise parameters and calculates the loss function. The target-agent network is obtained by copying the parameters of the evaluation-agent network at regular intervals;
[0034] A competitive-agent network for decomposing the reward value function Q i (τ) of each agent based on global information τ into an advantage value function, a state value function, and an action value function;
[0035] A result output module for adding the decomposition results to obtain the joint reward value function Q tot (τ,a) based on global information τ, and updating the parameters of the competitive-agent network and the overall loss function. After the parameter update is completed, the trained agents are used to execute actions in the environment.
[0036] In a third aspect, a computing device is proposed, and the device includes:
[0037] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and when the programs are executed by the processors, the method for multi-agent reinforcement learning based on value decomposition described in the first aspect of the present invention is implemented.
[0038] Compared with the prior art, the present invention has the following beneficial effects: Based on four networks with different functions, the present invention designs a value decomposition algorithm that can efficiently integrate the information of multi-agent systems. This algorithm adds parameterized noise to the connection weights of neural networks, randomizes the weight and bias parameters of the network, so that agents can explore new environments or states, and accelerate the exploration efficiency of agents in complex environments. At the same time, a competition-agent network is introduced before outputting the joint reward value function. This competition structure divides the joint reward value function into a state value function and an advantage value function, and can learn the value of the environmental state without the influence of actions, reducing the error between the estimated joint reward value function and the actual joint reward value function. The method of the present invention double extracts the logical topological relationship between multi-agents through the noise mechanism and the competition-agent network. In complex heterogeneous partially observable scenarios, agents can explore the state information beneficial to the multi-agent system within a limited time, and at the same time avoid agents choosing sub-optimal actions for rewards and ignoring optimal actions for rewards. The reward values obtained by agents are more stable, and the convergence speed of the algorithm model becomes faster. When the convergence speed of the algorithm model becomes faster, agents only need to spend a small amount of time in the initial exploration to explore the state information useful to themselves or the system, can quickly and effectively perceive the environment and make correct decisions, and can effectively select the optimal joint action for rewards to ensure that agents obtain the maximum reward feedback. Good algorithm performance promotes the effective cooperation between agents to ensure the completion of collaborative tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is the overall flowchart of the multi-agent reinforcement learning method based on value decomposition proposed by the present invention;
[0040] Figure 2 It is the schematic diagram of the network structure provided according to the example of the present invention;
[0041] Figure 3 It is the flowchart of the multi-agent reinforcement learning method based on value decomposition in the example of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The following further describes the present invention in detail with reference to the drawings and specific embodiments. It should be noted that the following described embodiments are intended to facilitate the understanding of the idea of the present invention and do not limit it in any way.
[0043] The present invention proposes a multi-agent reinforcement learning method based on value decomposition, which can be applied to complex partially observable scenarios, such as scenarios of video games, sensor networks, swarm coordination of robots, and autonomous driving vehicles. Each agent acts as a different role in different multi-agent scenarios. For example, in a game scenario, it acts as a hero, in a sensor network, it represents each sensor, in a robot cooperation scenario, it represents a single robot, and in an autonomous driving scenario, it acts as a vehicle. Any object that can perceive its surrounding environment and independently make decisions to affect the environment can be abstracted as an agent. The present invention uses a new value decomposition method to improve the cooperation efficiency among agents in a multi-agent system. Refer to Figure 1 , an improved multi-agent reinforcement learning method based on value decomposition, comprising the following steps:
[0044] Step 1: Construct a multi-agent reinforcement learning environment, which includes the following components: multiple agents, an evaluation-agent network and a target-agent network, multiple random-agent networks, a competition-agent network, and an experience buffer M with a capacity of m.
[0045] An agent is an individual that resides in the environment and acts. It can perceive the environment through autonomous operations and execute corresponding actions, continuously adapt to the changes in the environment in the long term, and gradually establish its own pursuit goals to cope with possible environmental changes that may be sensed in the future. Any object that can perceive its surrounding environment and independently make decisions to affect the environment can become an agent.
[0046] The evaluation-agent network is used to evaluate the quality of agent actions and learn agent cooperation strategies, and is mainly composed of a multi-layer perceptron and a gated recurrent unit. Its input is the observation value of each agent at the initial moment and the available actions, and the output result is the Q-value function of a single agent based on local information.
[0047] The role of the parameters of the target-agent network is to obtain the updated parameters of the evaluation-agent network at fixed intervals and copy the obtained parameters, so as to cut off the correlation of parameter data to ensure the convergence of the final algorithm model. Its input and output are the same as those of the evaluation-agent network.
[0048] The random-agent network is a neural network whose weights and biases are randomly perturbed by noise parameters. Its main role is to add randomness to the Q-value function output by the target-agent network. Its input is the global state information at the current moment and the Q-value function output by the target-agent network, and the output is the Q-value function of the agent based on the global state information.
[0049] The competition-agent network decomposes the Q-value function of the agent based on the global state information into an advantage value function, a state value function, and an action value function. Finally, the advantage value function, the state value function, and the action value function are summed respectively to obtain the joint reward value function in the entire multi-agent environment. The input is the global state information at the current moment and the sum of the Q-value functions of all agents based on the global state information, and the output is the joint reward Q-function of the entire system.
[0050] Step 2: Start the multi-agent scenario, initialize each network parameter, and then the agent obtains the state S of the environment at the current moment t t , the initial observation value of each agent available actions and the reward r corresponding to this action, where i is the serial number of the agent. The state S t includes the number of agents, the role type, and the size of the joint reward Q-value function obtained at the previous moment in the current multi-agent scenario. The specific content of the available actions depends on the specific settings of each scenario type. For example, in a game scenario, the actions include: attack, retreat, return to the city, buy equipment, etc. The sensor network includes actions such as sensing, collecting, and analyzing. The robot cooperation scenario contains actions such as starting, communicating, and executing. The actions of the agent in the autonomous driving scenario include: turning left, turning right, moving forward, moving backward, parking, starting, and so on. The reward is the numerical feedback of the agent making a certain specific action, usually also set by the scenario. The size and sign of the numerical value reflect the quality of this action.
[0051] Step 3: Each agent calculates the value function Q i observed based on local information τ i (τ i ) for each action through the evaluation-agent network, where the local information τ i is the observation value obtained by agent i in the previous step action reward r and state S t information set. The evaluation-agent network includes a recurrent layer (Gated Recurrent Unit, GRU) and two multi-layer perceptrons (Multi-layer Perceptron, MLP). The GRU is used to record the hidden information of each agent. The hidden information is the parameter of the input and output of the hidden layer of the evaluation-agent network and is used as the input of the next GRU hidden information. The hidden information of each agent should correspond one by one.
[0052] Step 4: The global state S at the current moment t t and the global state S at the next moment t+1 t+1 , the observation value of each agent selected action The experience tuple of the state s, action a, reward r is stored in the experience buffer M as a tuple.
[0053] Step 5: If the size of the experience buffer M exceeds its maximum capacity m, the evaluation-agent network's network parameter updates are counted as one episode every specified number of times (e.g., every 1000 times). The experience tuples in the experience buffer are grouped by episode, and then similar experience tuples from different episodes are randomly sampled.
[0054] Step 6: For each agent, based on the local information τ i The reward value function Q i (τ i ) and the state S at the current time t t are used as the input to the stochastic-agent network. Add parameterized noise to the state S at the current time to randomize the weight and bias parameters. Then, at the end of each episode, the trained weight and bias parameters are summed with the Q t (τ i ) of each agent based on the local information τ i (τ i ) to output the reward value function Q i (τ) of each agent based on the global information τ.
[0055] Step 7: After the stochastic-agent network outputs the reward value function Q i (τ) of each agent based on the global information τ, the target-agent network calculates the loss function and updates the parameters in the same way as DQN. Then, the stochastic-agent network also updates the noise parameters and calculates the loss function. The target-agent network is obtained by copying the parameters of the evaluation-agent network every p episodes.
[0056] Step 8: Use the reward value function Q i (τ) of each agent based on the global information τ as the input to the competitive-agent network, and decompose the reward value function Q i (τ) of each agent. The specific decomposition method is as follows:
[0057] Step 8.1: Input the global state S at the current time t t into the evaluation-agent network to convert it into a Q-value function that is only affected by the state, i.e., S i (τ). Since the state value function sometimes has little impact on the next state regardless of the action taken in a certain state, it is necessary to separately measure the value generated by the state on the environment.
[0058] Step 8.2: Input the action selected by agent i at the current time t into the stochastic-agent network and output the action value function C i(τ) because the agent is easily trapped in a cycle of locally optimal actions. Once the misestimation of several suboptimal joint actions exceeds the better estimation of a single optimal joint action, it will lead to the underestimation of the optimal action value, causing the agent to choose suboptimal actions and separately evaluate the action value function C i (τ) To reduce the impact brought by the underestimation of action value.
[0059] Step 8.3: Finally, subtract S i (τ) from the reward value function Q i (τ) of each agent based on the global information τ and C i (τ) to obtain the advantage value function A i (τ). Then sum up S i (τ) of each agent to get the joint state value function S tot (τ), sum up C i (τ) to get the joint action value function C tot (τ), and multiply A i (τ) of each agent by a coefficient δ and sum them up to get A tot (τ).
[0060] The role of decomposition is to better adapt to different multi-agent environments. The relationship constraint between the joint reward value function and the local reward value function only depends on the magnitude of the advantage value function, simplifies the value decomposition process, and improves the convergence speed of the value decomposition algorithm.
[0061] The S i (τ), C i (τ) and A i (τ) mentioned in the steps are value functions used by each agent in various actual scenarios to quantify what is a good action or a bad state, and are the expected values of future rewards for the agent to choose a state or an action under a given policy. The Q value function is the overall evaluation of the agent's selection of a specific action in a specific state to produce a specific reward. Generally speaking, the value function is a measurement criterion, and its specific meaning depends on the actual application scenario. For example, in a game scenario, the value function represents the economic return obtained by each game character. In a sensor network, the magnitude of the value function corresponds to the performance of each sensor. In a robot cooperation scenario, the value function represents the contribution degree of each robot to the task completion. In an autonomous driving scenario, the value function represents the number of kilometers traveled by each car, etc.
[0062] Step 9: Add S tot (τ), C tot (τ) and A i (τ) to obtain the joint reward value function Q tot(τ, a), and update the parameters of the competition - agent network and the overall loss function. After the parameter update is completed, use the trained agent to execute actions in the game environment.
[0063] Based on four networks with different functions, the present invention designs a value decomposition algorithm that can efficiently integrate the information of a multi - agent system. This algorithm adds parameterized noise to the connection weights of the neural network, randomizes the weight and bias parameters of the network, so that the agent can explore new environments or states, and speeds up the exploration efficiency of the agent in complex environments. At the same time, a competition - agent network is introduced before outputting the joint reward value function. This competition structure divides the joint reward value function into a state value function and an advantage value function, and can learn the value of the environmental state without the influence of actions, reducing the error between the estimated joint reward value function and the actual joint reward value function.
[0064] As Figure 2 shown, the evaluation - agent network includes an MLP for processing local observations and action input - output, and a GRU recurrent neural network for memorizing historical state and action information: Each agent encodes the local observations and actions and inputs them into the MLP, and then inputs them into the GRU recurrent neural network. The GRU network concatenates the hidden information h t at the current moment t with the output of the previous layer to generate the hidden information h t+1 at the next moment t + 1 and the input of the next layer; the MLP of the third layer then generates the value function Q i (τ i ) of each agent based on local information, which contains the local historical information τ i of a single agent. The target - agent network is obtained by copying the parameters of the evaluation - agent network every p rounds.
[0065] In addition, the system also includes a random - agent network and a competition - agent network.
[0066] In the random - agent network, the Q i (τ i ) output by each agent based on local information of the target - agent network and the state S t at the current moment t are input into the noise layer of the current random - agent network, and finally the value function Q i (τ) of each agent based on global information τ is output. Among them, the noise parameter θ is defined as θ = μ+∑⊙ε, where μ and ∑ are learnable noise parameter vectors, ε is a vector of zero - mean noise with fixed statistics, and ⊙ represents element - wise multiplication. The output result of the noise layer can be expressed as y=(μ w +σ w ⊙ε w )x + μb +σ b ⊙ε b where \(x\) is the state \(S\) at the current moment of the input t and \(y\) is the state \(S\) after being perturbed by noise randomization in the output t '. The weight parameter term \(\mu\) w +σ w ⊙ε w must be greater than 0 because for the entire multi-agent system, it only hopes to obtain beneficial feedback
[0067] In the competitive-agent network, for each agent output by the random-agent network, based on the value function \(Q\) of the global information \(\tau\) i (\(\tau\)) as the input, calculate the state value function \(S\) of each agent in the current state i (\(\tau\)) and the action value function \(C\) corresponding to the current action i (\(\tau\)), and then subtract the state value function \(S\) from the reward value function \(Q\) of each agent i (\(\tau\)) and the action value function \(C\) i (\(\tau\)) to obtain the advantage value function \(A\) generated by each agent taking the current action under the global information \(\tau\) i (\(\tau\)), and sum the \(S\) of each agent i (\(\tau\)) and \(C\) i (\(\tau\)) to obtain the joint state value function \(S\) i (\(\tau\)) and the joint action value function \(C\) tot (\(\tau\)), and then multiply the \(A\) of each agent tot (\(\tau\)) by a coefficient \(\delta\) and sum them to get \(A\) i (\(\tau\)). Finally, calculate the joint loss function and update the parameters. The joint reward value function is: \(Q\) tot (\(\tau, a\)) = \(V\) tot (\(\tau\)) + \(C\) tot (\(\tau\)) + \(A\) tot (\(\tau\)). tot (\(\tau\)).
[0068] Among them, the state value function \(S\) i (\(\tau\)) is only related to the state \(S\) at the current \(t\) moment t or the global information \(\tau\). \(S\) tot (\(\tau\)) can be expressed as: The action value function \(C\) i (\(\tau\)) is only related to the action at the current \(t\) moment or the global information \(\tau\), and \(C\) tot (\(\tau\)) can be expressed as: And the advantage value function \(A\) tot (\(\tau\)) is not only related to the state \(S\) at the current \(t\) moment t but also related to the action taken by each agent is related to, so A tot (τ) can be expressed as: δ > 0, where δ is the parameter weight trained by the random-agent network. Finally, Q tot (τ, a) can be simplified to the following form:
[0069]
[0070]
[0071] where A i (τ) = Q i (τ) - S i (τ) - C i (τ).
[0072] According to an embodiment of the present invention, the multi-agent reinforcement learning method based on value decomposition is applied to the StarCraft II game. As Figure 3 shown, the method includes the following steps:
[0073] S01: Set up the virtual environment of StarCraft II, and install the game and the interface library PySC2 provided by DeepMind for StarCraft II externally in the system.
[0074] S02: Configure the parameters of the game, and start the game scenario through the provided interface. Initialize the parameters of the evaluation-agent network and obtain the state S of the environment at the current moment t t , the initial observation value of each agent the available actions and the reward r corresponding to this action, where i is the serial number of the agent.
[0075] The state S t includes the number of combat units in StarCraft II, the types of combat units, and the size of the joint reward Q value function obtained at the previous moment t - 1. The available actions include: fighting, retreating, purchasing equipment, restoring health, etc. The reward r is the number of gold coins obtained by a combat unit for making a specific action, which can help the combat unit purchase game equipment and thus improve the combat winning rate.
[0076] S03: As Figure 2 shown, the evaluation-agent network takes the observation value and action obtained by each agent based on local observation as inputs. First, it encodes the input information through a single-layer fully connected layer MLP and inputs it into the GRU network. Then, the GRU network concatenates the hidden layer information h t-1 at the previous moment t - 1 and the output of the previous layer to generate the hidden layer information h tand the input of the next layer. The third layer then generates the reward value function Q for each agent in the discrete action space based on the local information τ i of i (τ i ).
[0077] S04: Evaluate - The reward value function Q for each agent output by the agent network based on the local information τ i of i (τ i ) includes the reward r of the agent under the current independent action and the global state S at the next moment t + 1 t+1 , and the observation value of each current agent The action taken The reward r obtained, the observed state S at the current moment t t and the observed state S at the next moment t + 1 t+1 are stored in the experience buffer as an experience tuple every p rounds, and the parameters are updated to the target - agent network every p rounds for sampling during training. Storing experiences every p rounds is to cut off the correlation between consecutive operation sequences, maintain independent and identically distributed, so that the network can converge.
[0078] S05: Before storing the experience tuple, judge whether the capacity of the experience buffer exceeds m. If the capacity of the experience buffer is full, count the network parameter update of the evaluation - agent network every specified number of times (for example, every 1000 times) as one round, group the experience tuples in the experience buffer according to rounds, and then randomly sample from similar experience tuples in different rounds. When choosing an action, not only the current input needs to be input, but also the hidden state needs to be input to the neural network, and the hidden state is related to the previous experience, so it is not possible to randomly extract experiences for learning. Therefore, here experiences from multiple rounds are extracted at once, and then similar tuples at the same position in each round are passed to the neural network at once.
[0079] S06: Use the reward value function Q for each agent output by the target - agent network based on the local information τ i of i (τ i ) and the state S at the current moment t t as the input of the random - agent network, add decomposed Gaussian noise to the fully connected layer of the random - agent network, and the parameters of the noise can be adjusted automatically through parameter update during training, so that the agent can decide when and at what ratio to introduce weights into uncertainty. Then add parameterized noise to the state S at the current moment t t to randomize the weight and bias parameters, and then combine the Q i for each agent based on the local information τ i (τ iPerform weighted summation and output the reward value function Q of each agent based on the global information τ at the end of each round. i (τ), where the global information includes the state S of each agent at the current time t t obtained reward r and action updated noise parameters and the current state S at the current time t t and the state S at time t + 1 t+1 .
[0080] S07: After the random-agent network outputs the reward value function Q of each agent based on the global information τ i (τ), the target-agent network and the random-agent network update their parameters and loss functions in turn. The loss function of the target-agent network is: where ω and ω′ are the parameter sets before and after the update of the target-agent network respectively. The loss function of the random-agent network is where and are the parameter sets before and after the update of the random-agent network respectively, γ is the reward decay factor and 0 < γ < 1, ζ = (μ, ∑) is defined as a set of learnable noise parameter vectors, where ζ and ζ′ are the noise parameter vectors before and after the update respectively.
[0081] S08: Use the reward value function Q of each agent based on the global information τ output by the random-agent network i (τ) as the input of the competitive-agent network. According to the decomposition method described in step 8 of the invention content, decompose the reward value function Q of each agent based on the global information τ i (τ) into the advantage value function A of each agent based on the global information τ i (τ), the state value function S i (τ) and the action value function C i (τ) according to the influence degree of the current state and action on Q i (τ). First, sum the state value functions S of all agents based on the global information τ i (τ) and the action value functions C i (τ) respectively to obtain S tot (τ) and C tot (τ). Then multiply the advantage value function A of each agent based on the global information τ i (τ) by the weight parameter δ trained by the noise network, and δ > 0. Introducing δ is to be consistent with the strategy of greedy action selection. It represents the credit assignment to different individual agents. Then for each agent's A i(τ) is multiplied by a coefficient δ and added to obtain A tot (τ).
[0082] S09: Finally, S tot (τ), C tot (τ) and A i (τ) are added to obtain the joint reward value function Q tot (τ, a) based on the global information τ, and update the parameters of the competitive-agent network and the overall loss function. After the parameter update is completed, the trained agent is used to execute actions in the game environment.
[0083] According to another embodiment of the present invention, a multi-agent reinforcement learning system based on value decomposition is provided, including:
[0084] An information perception module for obtaining the state S of the environment at the current time t t , the initial observation value of each agent The available actions And the reward r corresponding to the action, where i is the serial number of the agent, and the state S t Includes the number of agents, the role type, and the size of the joint reward Q value function obtained at the previous moment in the current multi-agent scenario;
[0085] An evaluation-agent network for calculating, for each agent, the value function Q i Observed based on the local information τ i (τ i ), where the local information τ i Is the observation value of agent i Action Reward r and state S t Information set;
[0086] A random-agent network for adding parameterized noise to the state S at the current moment to randomize the weight and bias parameters, and then summing the weights of the trained weight and bias parameters at the end of each round with the Q t Based on the local information τ of each agent i Of i (τ i ) to obtain the reward value function Q i (τ) of each agent based on the global information τ;
[0087] A parameter update module, the target-agent network calculates the loss function and updates the parameters, and then the random-agent network also updates the noise parameters and calculates the loss function, where the target-agent network is obtained by copying the parameters of the evaluation-agent network every once in a while;
[0088] A competition-agent network for decomposing the reward value function Q(τ) of each agent based on global information τ into an advantage value function, a state value function, and an action value function; i (τ) is decomposed into an advantage value function, a state value function, and an action value function;
[0089] A result output module for adding the decomposition results to obtain a joint reward value function Q(τ, a) based on global information τ, and updating the parameters of the competition-agent network and the overall loss function. After the parameter update is completed, the trained agent is used to execute actions in the environment. tot (τ, a), and update the parameters of the competition-agent network and the overall loss function. After the parameter update is completed, use the trained agent to execute actions in the environment.
[0090] Furthermore, the multi-agent reinforcement learning system further includes: an experience buffer module that stores the global state S at the current time t t and the global state S at the next time t + 1 t+1 , the observation value of each agent the selected action and the reward r as a tuple into the experience buffer M; and
[0091] An experience tuple grouping module that, when the size of the experience buffer M exceeds a specified threshold, takes each specified network parameter update of the evaluation-agent network as one episode, groups the experience tuples in the experience buffer according to episodes, and samples from similar experience tuples in different episodes.
[0092] It should be understood that the multi-agent reinforcement learning system based on value decomposition provided in this embodiment can implement all the functions of the above-mentioned improved multi-agent reinforcement learning method based on value decomposition. Each module of the system can implement the functions of the corresponding steps of the method, and the specific implementation method will not be elaborated here.
[0093] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0094] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 one or more blocks
[0095] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 one or more blocks
[0096] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 one or more blocks
[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, the interaction methods and online scheduling methods of the network information collection and scheduling devices (devices) in the present invention are applicable in various systems. Those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.
Claims
1. A multi-agent reinforcement learning method based on value decomposition, characterized in that, it includes the following steps: Obtain the state S of the environment at the current time t t , the initial observation value of each agent Available actions and the reward r corresponding to this action, where i is the serial number of the agent, and the state S t includes the number of agents, role types, and the size of the joint reward Q-value function obtained at the previous moment in the current multi-agent scenario; For each agent, the value function Q of each action based on the local information τ is calculated through the evaluation-agent network i observed i (τ i ), where the local information τ i is the observed value of agent i action the information set of the reward r and the state S t ; Using a random-agent network for the current state S t Adding parameterized noise to randomize the weight and bias parameters, and then summing the weights of the trained weight and bias parameters at the end of each episode with the Q i of each agent based on local information τ i (τ i ) to obtain the reward value function Q i (τ) for each agent based on global information τ; The target-agent network calculates the loss function and updates the parameters, and then the random-agent network also updates the noise parameters and calculates the loss function, where the target-agent network copies the parameters of the evaluation-agent network every once in a while; Using a competition-agent network, the reward value function Q of each agent based on the global information τ is decomposed into an advantage value function, a state value function, and an action value function. The decomposition of the reward value function Q of each agent by the competition-agent network includes: inputting the global state S at the current moment t into the evaluation-agent network to be transformed into a Q-value function affected only by the state, that is, the state value function S i (τ); inputting the action selected by agent i at the current moment t into the random-agent network and outputting the action value function C i (τ); subtracting S t (τ) and C i (τ) from the reward value function Q of each agent based on the global information τ to obtain the advantage value function A i (τ), summing the S i (τ) of each agent to obtain the joint state value function S i (τ), summing C i (τ) to obtain the joint action value function C i (τ), and then multiplying the A i (τ) of each agent by a coefficient δ and summing them to obtain A tot (τ); i (τ) tot (τ), and then multiplying the A i (τ) of each agent by a coefficient δ and summing them to obtain A tot (τ); Add the decomposition results to obtain the joint reward value function Q based on the global information τ tot (τ,a),Q tot (τ,a) = S tot (τ) + A tot (τ) + C tot (τ), and update the parameters of the competition-agent network and the overall loss function. After the parameter update is completed, use the trained agent to execute actions in the environment.
2. The multi-agent reinforcement learning method based on value decomposition according to claim 1, characterized in that, the agent includes any one of the following: a hero in a game scenario, a single sensor in a sensor network, a single robot in a robot cooperation scenario, or a car in an autonomous driving scenario.
3. The multi-agent reinforcement learning method based on value decomposition according to claim 1, characterized in that, The evaluation-agent network includes an MLP for processing local observations and action inputs / outputs, and a GRU recurrent neural network for memorizing historical state and action information: Each agent encodes the local observations and actions into the MLP and inputs them into the GRU recurrent neural network, and the GRU recurrent neural network concatenates the hidden information h at the current time step t with the output of the previous layer to generate the hidden information h at the next time step t+1 and the input to the next layer; the MLP of the third layer then generates the value function Q for each agent based on the local information t and the input to the next layer; the MLP of the third layer then generates the value function Q for each agent based on the local information t+1 and the input to the next layer; the MLP of the third layer then generates the value function Q for each agent based on the local information i (τ i ), which contains the local historical information τ of a single agent i .
4. The multi-agent reinforcement learning method based on value decomposition according to claim 1, characterized in that, The random-agent network is a neural network in which the weights and biases are randomly perturbed by noise parameters. The noise parameter θ is defined as θ = μ + ∑ ⊙ ε, where μ and ∑ are learnable noise parameter vectors, ε is a vector of zero-mean noise with fixed statistics, and ⊙ represents element-wise multiplication. The output of the noise layer is expressed as y = (μ w + σ w ⊙ ε w )x + μ b + σ b ⊙ ε b , where x is the state S at the current time of the input t , and y is the state S' after being randomly perturbed by noise. The weight parameter term μ t + σ w ⊙ ε w ⊙ ε w is greater than 0.
5. The multi-agent reinforcement learning method based on value decomposition according to claim 1, characterized in that, State value function S i (τ) is only related to the state S at the current time t t or the global information τ, and S tot (τ) is expressed as: Action value function C i (τ) is only related to the action at the current time t or the global information τ, and C tot (τ) is expressed as: And the advantage value function A tot (τ) is not only related to the state S at the current time t t but also related to the actions taken by each agent and A tot (τ) is expressed as: δ > 0, and δ is the parameter weight trained by the random-agent network.
6. The multi-agent reinforcement learning method based on value decomposition according to claim 5, characterized in that, Combined reward value function Q tot (τ,a) is simplified to the following form: Where A i (τ) = Q i (τ) - S i (τ) - C i (τ).
7. The multi-agent reinforcement learning method based on value decomposition according to claim 1, characterized in that, The method further includes: the global state S at the current moment t t and the global state S at the next moment t+1 t+1 , the observation value of each agent the selected action and the reward r are stored as a tuple in the experience buffer M; when the size of the experience buffer M exceeds the specified threshold, the evaluation-agent network updates the network parameters every specified number of times as one episode, groups the experience tuples in the experience buffer according to episodes, and samples from the similar experience tuples in different episodes.
8. A multi-agent reinforcement learning system based on value decomposition, characterized in that, it includes: An information perception module for obtaining the state S of the environment at the current time t t , the initial observation value of each agent Available actions and the reward r corresponding to the action, where i is the serial number of the agent, and the state S t includes the number of agents, the role type, and the magnitude of the joint reward Q-value function obtained at the previous moment in the current multi-agent scenario; Evaluation - Agent Network, which for each agent, calculates for each action based on local information τ i Observed value function Q i (τ i ), where the local information τ i is the observed value of agent i Action Reward r and state S t is the information set of; A random-agent network for the state S at the current moment t Add parameterized noise to randomize the weight and bias parameters, and then sum the weights of the trained weight and bias parameters at the end of each episode with the Q i of each agent based on the local information τ i (τ i ) to obtain the reward value function Q i (τ) for each agent based on the global information τ; A parameter update module, where the target-agent network calculates the loss function and updates the parameters, and then the random-agent network also updates the noise parameters and calculates the loss function, where the target-agent network copies the parameters of the evaluation-agent network every once in a while; A competition-agent network is used to decompose the reward value function Q of each agent based on the global information τ into an advantage value function, a state value function, and an action value function. The competition-agent network decomposes the reward value function Q of each agent as follows: i (τ) includes: inputting the global state S at the current time t into the evaluation-agent network to convert it into a Q-value function that is only affected by the state, i.e., the state value function S i (τ); inputting the action selected by agent i at the current time t into the random-agent network and outputting the action value function C t (τ); subtracting S i (τ) and C i (τ) from the reward value function Q of each agent based on the global information τ to obtain the advantage value function A i (τ); summing up the S i (τ) of each agent to obtain the joint state value function S i (τ); summing up the C i (τ) to obtain the joint action value function C i (τ); then multiplying the A tot (τ) of each agent by a coefficient δ and summing them up to obtain A i (τ); tot (τ); i (τ) and adding them up to obtain A tot (τ); A result output module, which is used to add the decomposition results to obtain a joint reward value function Q based on the global information τ tot (τ,a),Q tot (τ,a) = S tot (τ) + A tot (τ) + C tot (τ), and update the parameters of the competition-agent network and the overall loss function. After the parameter update is completed, use the trained agent to execute actions in the environment.
9. A computer device, characterized in that, it includes: One or more processors; A memory; And one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and when the programs are executed by the processors, they implement the steps of the multi-agent reinforcement learning method based on value decomposition according to any one of claims 1-7.
Citation Information
Patent Citations
Command and control system based on Starcraft II
CN111632387A
Multi-agent reinforcement learning method based on value decomposition and attention mechanism
CN113313267A
Multi-agent cooperation model based on deep reinforcement learning
CN113592101A
Cited By
Systems and methods for dynamics-aware comparison of reward functions
US12468941B2
Systems and methods for dynamics-aware comparison of reward functions
US20230104027A1