A Multi-Agent Reinforcement Learning Collaboration Method Based on Dynamic Graph Communication
By introducing dynamic graph communication model and credit allocation algorithm in multi-agent reinforcement learning, the problem of space search complexity and non-stationarity of cooperative strategy in multi-agent systems is solved, and efficient cooperative decision-making under restricted communication conditions is achieved, which improves cooperation performance and scalability.
Patent Information
- Application Number
- CN202310114762.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-15
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-02-15
AI Technical Summary
With the increase in the scale of multi-agents, the complexity of spatial search of joint cooperative strategies has increased exponentially. Coupled with the non-stationarity brought about by independent decision-making of agents, and the complex coupling relationship between multiple agents, the development of multi-agent reinforcement learning algorithms has greatly limited the development of multi-agent reinforcement learning algorithms.
A multi-agent reinforcement learning collaboration method based on dynamic graph communication is proposed. By introducing a communication model into a credit allocation algorithm, advanced cooperative strategies are learned, and credit allocation between agents is used to use hypernetwork in the training stage, and in the execution stage, the agent makes collaborative decisions based on local information and communication messages.
Real-time dynamic communication diagram between agents under the conditions of restricted communication in real applications is established, and credit allocation is used for centralized hypernetwork, which significantly improves the cooperation performance between agents, reduces communication overhead, and has higher scalability.
Smart Images

Figure CN116306966B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence and multi-agent cooperation, and particularly relates to a multi-agent reinforcement learning cooperation method based on dynamic graph communication. Background Art
[0002] Multi-agent cooperation mainly refers to an agent system containing multiple agents continuously interacting with the environment in an interactive environment to maximize the benefits obtained by the system. Each agent makes independent policy decisions, and they cooperate autonomously to complete the team goals. Multi-agent cooperation technology plays a crucial role in fields such as smart cities, intelligent transportation, vehicle-road cooperation, and drone control, and can be used in tasks such as communication coordination of multiple independent terminals, optimization of resource allocation, and cluster path planning.
[0003] In recent years, multi-agent cooperation methods have made great progress. However, with the increase in the scale of multi-agents, the search complexity of the joint cooperation strategy space increases exponentially. Coupled with the non-stationarity brought about by independent agent decisions and the complex coupling relationships among multiple agents, the development of related algorithms is greatly restricted. Therefore, multi-agent reinforcement learning algorithms, as an effective method for adaptively promoting agent cooperation, have gradually received more and more attention. They can directly perform trial-and-error learning in the training stage using the interaction data between agents and the environment, have strong scalability, and have great development prospects.
[0004] Currently, the main methods for multi-agent cooperation research are generally divided into three categories: (1) Each independently decision-making agent builds a model of the strategies of other agents locally and makes individual decisions based on local interaction information and the modeled strategies. (2) Use a hypernetwork to decompose the overall team reward during the centralized training stage for reasonable credit assignment among agents, thereby implicitly promoting cooperation among agents based on reinforcement learning methods. (3) Enable effective communication among agents, and each agent makes decisions based on local data and communication messages to achieve cooperation. The first category of methods reduces the non-stationarity brought by other dynamic strategies during the agent decision-making process through active modeling. However, as the number of agents increases, the difficulty of modeling will increase exponentially, and it cannot handle complex cooperation tasks. The second category of algorithms guides agent cooperation through the team reward value directly related to the task. By reasonably decomposing the team reward value through a hypernetwork, the joint behavior strategy of the multi-agent system can converge to a cooperative strategy that satisfies monotonicity constraints. The third category of methods promotes agents to collaboratively complete the team goal by communicating, either artificially defining or generating communication messages through a designed specific network. In practical applications, due to the appropriate learning cost and strong generalization ability of the second and third algorithms, they have higher application value in large-scale multi-agent cooperation.
[0005] In recent years, the popular cooperative method of multi-agent reinforcement learning mainly adopts the paradigm of centralized training and distributed execution to train and deploy the agent decision-making model. During the training process, the reward signal obtained by interacting the joint action formed by all agent decisions with the environment is decomposed to achieve the belief distribution among agents. By continuously trial-and-error in the environment, each individual policy network converges to an effective joint cooperation strategy. The decomposition of the reward information depends on a hypernetwork that can obtain the global agent system information during the training phase, and it should have the ability to represent the complete policy space. In the execution phase, the centralized hypernetwork will be removed, and each agent only depends on its own policy network to select actions. Rashid et al. proposed a multi-agent value decomposition framework that integrates the independent Q-functions of each agent through a non-linear hypernetwork with non-negative weights, so as to achieve credit assignment during the backpropagation update process of the reward signal (Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, Shimon Whiteson QMIX: Monotonic Value Function Factorisation for Deep Multi-agent Reinforcement Learning[C] / / Proceedings of the 35th International Conference on Machine Learning, PMLR 80: 4295-4304, 2018.). Since the hypernetwork can obtain the true state of the agent system during the centralized training phase, and the non-linear hypernetwork can represent the monotonicity constraints that are more in line with the individual and joint policies, thus learning more effective cooperation strategies, the design of the non-linear hypernetwork is also one of the research hotspots in the field of multi-agent reinforcement learning. The current research methods promote the model to learn effective cooperation strategies by restricting the distance between the behavior policy and the target policy or using hierarchical hypernetworks.Wang et al. utilized a hierarchical reinforcement learning structure to distribute rewards through two hypernetworks, thereby reducing the difficulty of a single network representing the entire policy space (T Wang, T Gupta, B Peng, A Mahajan, S Whiteson, and C Zhang. 2021. RODE: Learning Roles to Decompose Multi-agent Tasks[C] / / In Proceedings of the International Conference on Learning Representations. OpenReview.). Communication learning is another important research direction in the field of multi-agent reinforcement learning. Yuan et al. proposed a communication mechanism based on variational inference. Agents use local predictions of the Q-functions of teammate agents as communication messages, and at the same time use the messages obtained from other agents as biases for their own Q-functions to achieve stable action value estimation, and introduce communication regularization to reduce communication costs (Lei Yuan, Jianhao Wang, Fuxiang Zhang, Chenghe Wang, Zongzhang Zhang, Yang Yu, and Chongjie Zhang. 2022. Multi-Agent Incentive Communication via Decentralized Teammate Modeling[C] / / In Proceedings of the AAAI conference on artificial intelligence). For some multi-agent collaborative tasks that require advanced cooperation behaviors, it is difficult to achieve the team's goals relying solely on implicit cooperation guidance. Therefore, it is difficult to learn complex cooperation strategies relying solely on the belief assignment of hypernetworks. At the same time, in most realistic multi-agent collaborative application scenarios, there are various communication restrictions, and existing algorithms often have excessive communication overhead and thus cannot be effectively applied. Summary of the Invention
[0006] In view of the above problems, the present invention proposes a multi-agent reinforcement learning collaboration method based on dynamic graph communication. By introducing a communication model into the credit assignment-based algorithm, it aims to learn advanced cooperation strategies; at the same time, it controls the scope and degree of communication based on a dynamic communication graph, aiming to achieve effective communication under more realistic restricted communication conditions, thereby promoting cooperation among agents.
[0007] The method of the present invention can work under the conditions of widespread restricted communication, adaptively communicate for agents according to the requirements of restricted communication domains in real-world applications, and use a hypernetwork for credit assignment among agents during the training phase. During the execution phase, agents make collaborative decisions based on local information and communication messages, effectively reducing the non-stationarity in the joint decision-making process and improving the collaborative performance of multi-agent systems.
[0008] The technical solution of the present invention:
[0009] A multi-agent reinforcement learning collaboration method based on dynamic graph communication, comprising the following steps:
[0010] Step 1: According to the communication restriction conditions of the environment and the agent system, extract the communicable agents within the communication domain of the agents in real time, and establish a communication graph.
[0011] Step 2: Encode the local observation information of the agents according to the communication graph in Step 1, and generate the weights of the communication graph based on it and the corresponding weight generator to control the degree of communication between agents.
[0012] Step 3: Perform communication of the encoded observation information between agents based on the weights of the communication graph in Step 2 and the communication graph in Step 1.
[0013] Step 4: Each agent uses the action value estimation network to complete individual action value estimation according to local interaction data, communication messages, and historical information.
[0014] Step 5: The hypernetwork aggregates all the action value estimations generated in Step 4 and completes the joint action value estimation of the agent system based on global information.
[0015] Step 6: According to the reward value obtained from the interaction between the joint action and the environment, update the parameters of the hypernetwork, and then backpropagate the confidence assignment value of the reward to the action value estimation networks of each agent and update their network parameters.
[0016] Step 7: Repeat Steps 1 to 6 until the action value estimation networks, communication weight generators, and hypernetworks of each agent converge or reach the specified number of training steps. Remove the hypernetwork and apply the final communication weight generator and the action value estimation networks of each agent to the interaction environment for agent system decision-making.
[0017] Further, the specific process of Step 1 is as follows: Establish a communication graph according to the communication domain under the communication restriction in the interaction environment where represents the communication graph, represents the set of agents, w represents the weights of each edge of the communication graph and is initialized to 0, and ε is the set of edges of the communication graph. The process of establishing the communication graph is i ≠ j, If agent j ∈ d i , where d i is the restricted communication domain of agent i.
[0018] Furthermore, the specific process of step 2 is as follows: Use the encoding network to encode the local observation information o of agent i j into the observation encoding e j , and then generate the weights of each edge of the communication graph according to the weight generator.
[0019] If a learnable weight generator is used, first use a linear transformation W to map the observation encoding to a high-dimensional space to enhance the network's expressive ability, and then use a single-layer non-linear network to calculate the communication coefficient c between corresponding communicable agents pairwise ij :
[0020]
[0021] where a(·) represents a single-layer non-linear network, represents the concatenation operation, e i and e j represent any communicable agents i and j respectively, Finally, perform softmax normalization on the weights of all communicable agents of each agent to ensure scalability:
[0022]
[0023] where w ij represents the communication weight between agent i and agent j, LeakyReLU(·) represents the non-linear activation function, and exp(·) represents the exponential symbol.
[0024] If a weight generator using similarity metric is used, then replace the non-linear network a(·) with an inner product similarity metric:
[0025]
[0026] where F is a linear embedding operation that can map the observation encoding to a high-dimensional space.
[0027] Furthermore, the specific process of step 3 is as follows: Generate the communication message of the agent according to the communication graph obtained in step 1 and the communication weight obtained in step 2:
[0028]
[0029] where m i represents the communication message obtained by agent i at the current moment.
[0030] Further, the specific process of step 4 is as follows: Generate the information representation at the current moment based on the communication message obtained in step 3, the local observation data of the agent, and the historical data
[0031]
[0032] where GRU(·) represents the gated recurrent unit neural network and respectively represent the observation information and communication message of agent i at the current time t, represents the historical information of the agent. The action value estimation network then performs action value estimation based on the information representation:
[0033] Q i (a) = Q i (a|e i ,m i ,h i ; θ) (6)
[0034] where a represents the optional action of agent i, and θ represents the parameters of the action value estimation network of the agent
[0035] Further, the specific process of step 5 is as follows: Perform joint action value estimation based on the individual action value estimation obtained in step 4:
[0036] Q tot (a) = mixing((s, Q1(τ1, a1), …, Q n (τ n , a n ))) (7)
[0037] where s represents the overall state of the agent system, Q tot represents the value estimation of the joint action, a represents the joint action of the agent system, and mixing(·) represents the hypernetwork
[0038] Further, the specific process of step 6 is as follows: Calculate the temporal difference loss based on the deviation between the joint action value estimation obtained in step 5 and the actually obtained reward
[0039]
[0040] where represents taking the expectation over the buffer φ, ξ, θ, ψ respectively represent the parameters of the observation encoding network, the weight generator, the action value estimation network, and the hypernetwork, s′ represents the state at the next moment, and a′ represents the joint action value estimation at the next moment Denote the target network of the action value estimation network, and γ denote the discount factor.
[0041] Advantages of the present invention:
[0042] Under the condition of satisfying the restricted communication widely existing in practical applications, the present invention uses the restricted communication domain to establish a real-time dynamic communication graph among agents, and uses a centralized hypernetwork to perform credit assignment among agents during the reinforcement learning process. In this way, on the premise of a small communication overhead, the model can adaptively perform effective communication, significantly improve the cooperation performance among agents, and at the same time has higher scalability.
[0043] First, through the actual communication domain restriction in the interaction environment, the communication graph structure at the current moment is established in real time, effectively reducing the communication overhead and ensuring strong scalability.
[0044] Secondly, a learnable weight generator and a weight generator for similarity measurement are used to generate communication weights in the communication topology, so as to realize adaptive communication control among agents and ensure the real-time effectiveness of communication.
[0045] Finally, in the action value estimation stage, the communication message and the observation information are encoded together as part of the historical information, effectively retaining the important communication information during the cooperation process of agents for accurate action value estimation, and using the hypernetwork to perform credit assignment among agents to ensure effective cooperation among agents. Description of the Drawings
[0046] Figure 1 is the flow chart of the multi-agent reinforcement learning cooperation method based on dynamic graph communication of the present invention;
[0047] Figure 2 is the flow chart of the training steps of the specific embodiment of the present invention;
[0048] Figure 3 is the flow chart of the test steps of the specific embodiment of the present invention;
[0049] Figure 4 is the schematic diagram of the restricted communication domain scenario of the work of the present invention in combination with specific embodiments. Detailed Embodiments
[0050] The present invention will be further described below in conjunction with the drawings and specific embodiments. The present invention includes but is not limited to the following embodiments.
[0051] As Figure 1 shown, the present invention provides a multi-agent reinforcement learning cooperation method based on dynamic graph communication, and its specific implementation process is as follows:
[0052] 1. Generation of communication graph and communication weights, asFigure 4 As shown, in an interactive environment taking a certain game as an example, the communication range of each agent is limited, and its communication range is consistent with its observable range. It can only observe information within a certain range, and a communication graph needs to be established based on the observation domain of the agents in the game. Among them represents the communication graph, represents the set of agents, and w represents the weight of each edge of the communication graph and is initialized to 0. As Figure 2 shown, the process of building the graph is as follows: i≠j, If agent j∈d i , where d i is the restricted communication domain of agent i. Then use the encoding network shown on the left side of Figure 2 to encode the local observation information o j of agent i into the observation encoding e j . Then, according to the weight generator, use the observation encodings of each agent as input to generate the weights of each edge of the communication graph. If a learnable weight generator is used, first use a linear transformation W to map the observation encoding to a high-dimensional space to enhance the network's expressive ability. Subsequently, use a single-layer non-linear network to calculate the communication coefficients between corresponding communicable agents pairwise, and the calculation formula is as follows:
[0053]
[0054] where a(·) represents a single-layer non-linear network, represents the concatenation operation, e i and e j represent any communicable agents i and j respectively, Finally, perform softmax normalization on the weights of all communicable agents of each agent to ensure scalability, and the calculation formula is as follows:
[0055]
[0056] where w ij represents the communication weight between agent i and agent j, LeakyReLU() represents the non-linear activation function, and exp(·) represents the exponential symbol.
[0057] If a weight generator using similarity measurement is used, then replace the above non-linear network with an inner product similarity measurement, and the calculation formula is as follows:
[0058]
[0059] where F is a linear embedding operation that can map the observation encoding to a high-dimensional space.
[0060] 2. Inter-agent communication
[0061] As Figure 2 shown, based on the communication graph obtained in Step 1 and the communication weights obtained in Step 2, generate the communication messages m1, m2, …, m n for the agents, and the calculation formula is as follows:
[0062]
[0063] Again, as Figure 2 shown, send the corresponding messages to the corresponding player units.
[0064] 3. Agents perform action value estimation in the interaction environment
[0065] As Figure 2 shown, based on the communication messages obtained in Step 3, the local observation data and historical data of the agents, use the Figure 2 GRU network in it to generate the information representation at the current moment
[0066]
[0067] where GRU(·) represents the gated recurrent unit neural network, respectively represent the observation information and communication message of agent i at the current t moment, represents the historical information of the agent. Then use the Figure 2 action value estimation network shown below to perform action value estimation based on this:
[0068] Q i (a) = Q i (a|e i ,m i ,h i ; θ) (14)
[0069] where a represents the optional action of agent i, and θ represents the parameters of the action value estimation network of the agent.
[0070] Then, based on the obtained individual action value estimation, perform joint action value estimation:
[0071] Q tot (a) = (s, Q1(τ1, a1), …, Q n (τ n , a n )) (15)
[0072] where s represents the overall state of the agent system, Q tot represents the value estimation of the joint action, and a represents the joint action of the agent system.
[0073] 4. Update the learnable parameters of each network in reverse
[0074] As Figure 2 shown, calculate the temporal difference loss based on the deviation between the estimated joint action value obtained in step 5 and the reward actually obtained from the game system
[0075]
[0076] where denotes taking the expectation over the buffer , φ, ξ, θ, ψ represent the parameters of the observation encoding network, weight generator, action value estimation network, and hypernetwork respectively, s ′ represents the state at the next moment, a ′ represents the estimated joint action value at the next moment represents the target network of the action value estimation network, and γ represents the discount factor
[0077] 5. Perform multi-agent collaborative tasks in the test environment
[0078] As Figure 3 shown, remove the hypernetwork from the converged model trained in the above steps or the n individual networks and weight generator that have reached the specified number of training steps, and apply the final weight generator and each agent's action estimation network to the interactive environment for agent system decision-making. During the execution phase, the agents communicate adaptively according to the dynamic communication graph network and perform collaborative tasks
Claims
1. A multi-agent reinforcement learning collaboration method based on dynamic graph communication, characterized in that, It includes the following steps: Step 1: According to the communication restriction conditions of the environment and the agent system, extract the communicable agents within the agent communication domain in real time and establish a communication graph. Specifically as follows: Establish a communication graph according to the communication domain under communication restrictions in the interaction environment where represents the communication graph, represents the set of agents, w represents the weight of each edge of the communication graph and is initialized to 0, and ε is the set of edges of the communication graph; the process of establishing the communication graph is as follows If agent j ∈ d i , where d i is the restricted communication domain of agent i; Step 2: According to the communication graph in Step 1, encode the local observation information of the agent, and generate the weights of the communication graph based on it and the corresponding weight generator to control the degree of communication between agents. Specifically as follows: Encode the local observation information \(o\) of agent \(i\) using an encoding network j into an observation encoding \(e\) j , and then generate the weights of each edge of the communication graph according to the weight generator; If a learnable weight generator is used, first, a linear transformation W is used to map the observed encoding to a high-dimensional space to enhance the network's expressive power. Subsequently, a single-layer non-linear network is utilized to calculate the communication coefficient c between corresponding communicable agents pairwise ij : where a(·) represents a single-layer non-linear network, represents the concatenation operation, e i and e j represent any communicable agents i and agent j respectively, Finally, the weights of all communicable agents of each agent are softmax-normalized to ensure scalability: where w ij represents the communication weight between agent i and agent j, LeakyReLU() represents the non-linear activation function, and exp(·) represents the exponential symbol; If using a weight generator of similarity measure, replace the non-linear network a(·) with an inner product similarity measure: where F is a linear embedding operation that can map the observation encoding to a high-dimensional space; Step 3: Based on the weights of the communication graph in Step 2 and the communication graph in Step 1, conduct the communication of the observation information encoding between agents. Specifically as follows: Generate the communication messages of the agents: where m i represents the communication message obtained by agent i at the current moment; Step 4: Each agent uses the action value estimation network to complete the individual action value estimation according to the local interaction data, communication messages and historical information. Specifically as follows: Generate the information representation at the current moment based on the communication messages obtained in step 3, the local observation data, and the historical data of the agent where GRU(·) represents the gated recurrent unit recurrent neural network, and represent the observation information and communication message of agent i at the current time t, respectively, represents the historical information of the agent; The action value estimation network conducts action value estimation based on the information representation: Q i (a) = Q i (a|e i ,m i ,h i ; θ) (6) where a represents the optional actions of agent i, and θ represents the parameters of the action value estimation network of the agent; Step 5: The hypernetwork aggregates all the action value estimations generated in Step 4 and completes the joint action value estimation of the agent system based on the global information. Specifically as follows: Q tot (a) = mixing((s, Q1(τ1, a1), …, Q n (τ n , a n ))) (7) where s represents the overall state of the agent system, Q tot represents the value estimation of the joint action, a represents the joint action of the agent system, and mixing(·) represents the hypernetwork; Step 6: According to the reward value obtained from the interaction between the joint action and the environment, update the parameters of the hypernetwork, and then backpropagate the confidence assignment value of the reward to the action value estimation networks of each agent and update their network parameters; Specifically as follows: Calculate the temporal difference loss based on the deviation between the estimated combined action value obtained in step 5 and the actually obtained reward Among them represents the expectation for the buffer φ, ξ, θ, ψ represent the parameters of the observation coding network, weight generator, action value estimation network, and hypernetwork respectively, s′ represents the state at the next moment, and a′ represents the joint action value estimation at the next moment represents the target network of the action value estimation network, and γ represents the discount factor Step 7: Repeat Step 1 to Step 6 until the action value estimation networks, communication weight generators and hypernetworks of each agent converge or reach the specified number of training steps. Remove the hypernetwork and apply the final communication weight generator and the action value estimation networks of each agent to the interaction environment for agent system decision-making.
Citation Information
Patent Citations
Multi-agent consistency cooperative control method under multiple information constraints
CN112596395A
Strategy selection method in complex game scene based on agent communication mechanism
CN113254872A