Multi-agent Gaussian mixture skill learning method based on pattern graph embedding
By introducing Gaussian hybrid variational inference encoder and graph matching technology in multiagent reinforcement learning, the similarity of the spatial distribution of entities is solved, and the existing methods have limited generalization capabilities when different teams form tasks, achieving more efficient coordinated skill learning and task completion.
Patent Information
- Application Number
- CN202510083023.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
AI Technical Summary
The existing multi-agent reinforcement learning methods have limited generalization capabilities when facing tasks composed of different teams, and require retraining the model, which leads to time-consuming and high computational cost.
A Gaussian hybrid offline coordination skills learning method (GOSPE) based on pattern graph embedding is proposed. The coordination skills of the agent are learned through the Gaussian hybrid variational inference encoder, and the similarity of the entity spatial distribution is embedded into the skill representation using graph matching and graph regularization terms to achieve universalization of the coordination pattern.
The generalization ability and coordination efficiency of multi-agent reinforcement learning are improved, and coordination modes can be effectively learned and reused in different task environments, significantly improving task completion efficiency and collaboration between agents.
Smart Images

Figure CN119990244A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-agent reinforcement learning and collaborative control, and specifically relates to a multi-agent Gaussian mixture skill learning method based on pattern graph embedding. Background Art
[0002] In recent years, cooperative multi-agent reinforcement learning (MARL) has made significant progress and breakthroughs in many real-world application scenarios, such as traffic control, multi-robot system control, dynamic spectrum access, and mobile crowd sensing. In the setting of cooperative MARL, the goal is to train a group of agents so that they can work together to achieve a common goal defined by a global reward function. Existing cooperative MARL methods have successfully learned effective coordination strategies between agents by maximizing the global reward b.
[0003] Although these methods have shown excellent performance on certain specific tasks, their generalization ability is relatively limited when they are applied to new tasks with different team compositions (i.e., combinations of different agents and target entities). Specifically, most existing MARL methods are designed only for fixed team compositions for specific tasks, which makes it difficult for the coordination strategies they learn to flexibly adapt to new team compositions. In order to obtain coordination strategies that better match the new team composition, these existing methods often need to retrain the model, which is not only time-consuming but also incurs huge computational costs. Fortunately, we can observe some common coordination patterns when dealing with tasks with different team compositions. In-depth exploration and learning of these common coordination patterns may be the key to improving generalization ability. Multi-task multi-agent reinforcement learning (MT-MARL) is a promising method that aims to discover these common coordination patterns.
[0004] Previous MT-MARL research has focused on solving multi-agent coordination problems in a predefined set of tasks, or fine-tuning pre-trained strategies online to adapt to specific target tasks. However, unpredictable changes in the number of entities often lead to a series of unknown target tasks, which is very common in many real-world scenarios. Taking urban traffic signal control as an example, traffic signals at multiple intersections need to be coordinated. However, damage or updates to traffic lights may cause unpredictable changes in the number of agents, resulting in different unknown tasks. Changes in the number of entities not only increase the complexity of possible configuration combinations, but also further increase the need for MARL generalization capabilities.
[0005] Moreover, in many practical applications, online interaction is not always feasible because deploying multiple agents into the environment can be very expensive. In contrast, offline methods have shown great practical value by learning policies from offline datasets. Current offline multi-task reinforcement learning methods share a unified backbone network between tasks and generalize multi-agent coordination by learning task-invariant representations or discrete skills. However, directly sharing representations or skills between tasks may not fully capture the complex cooperative behaviors between agents, thereby hindering the effective execution of coordination strategies in unknown tasks. Moreover, the cooperative behaviors of agents are often closely linked to their spatial relationships. Existing multi-task MARL methods simply use local neighbor relationships between agents for information sharing and aggregation, but fail to fully consider the overall spatial structure of the agents. Summary of the invention
[0006] To address the above problems, this paper proposes a new offline MT-MARL method called Gaussian mixture offline coordination skill learning (GOSPE) based on pattern graph embedding. In GOSPE, cooperative behaviors are regarded as variables generated by unknown Gaussian mixture distributions, where multiple components enable agents to learn different cooperative skills between tasks. In addition, the present invention finds that common coordination patterns across different tasks show similarities in the entity space distribution. Therefore, the similarity of the entity space is used to guide the emergence of generalized coordination patterns. In GOSPE, graph matching is introduced, and a coordination graph regularization term is derived to embed entity distribution similarity into skill representation learning. Then, a local skill selection module is applied to select the coordination skills that maximize the global benefit.
[0007] The present invention is roughly divided into two parts:
[0008] This paper innovatively proposes an offline multi-agent reinforcement learning model called "Gaussian mixture offline coordination skill learning based on pattern graph embedding" (GOSPE). This model can learn general Gaussian mixture skills from offline datasets of multiple tasks, which can flexibly represent complex cooperative behaviors. This paper also proposes a new coordination graph regularization term, which aims to guide and promote the emergence of generalizable coordination patterns. This innovative design not only enhances the generalization ability of the model, but also improves its coordination efficiency between different tasks.
[0009] To verify the effectiveness of GOSPE, it was comprehensively evaluated in a series of challenging multi-agent coordination environments. Compared with multiple baseline methods, GOSPE showed excellent performance, not only achieving excellent generalization ability on unseen tasks, but also showing excellent coordination performance. This result fully demonstrates the advancement and practicality of GOSPE in the field of multi-agent reinforcement learning.
[0010] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0011] The multi-agent Gaussian mixture skill learning method based on pattern graph embedding includes the following steps:
[0012] Step (1) constructs a Gaussian mixture variational inference encoder.
[0013] In traditional multi-task learning, the behavior of agents is usually highly complex and diverse. In order to capture the inherent laws of these behaviors, this paper first constructs a Gaussian mixture variational inference encoder. The encoder helps agents better understand and perform tasks by extracting coordination skill representations from static multi-task datasets.
[0014] Specifically, the behavior of the agent is considered to be generated by multiple potential coordination skills (e.g., action strategies). Behind each behavior is a mixture model composed of Gaussian distributions, which allows the model to flexibly express complex behavior patterns. In order to efficiently reason and learn these skills, an autoencoder structure is designed to encode and decode the behavior data of the agent. Assume that the cooperative behavior a of agent i i Coordination skills Generated, is the set of latent variables, z i is sampled from the prior skill distribution. We choose Gaussian mixture distribution as the prior distribution to allow more complex expressions, where a cooperative behavior of an agent can be represented as a component of a Gaussian distribution mixture. Therefore, each agent's cooperative behavior a i The generation process of Where p(w i ) and p(y i ) are the latent variables w of agent i i and i The prior distribution of θz and p θa It means that through the parameter θ z and θ a The distribution functions of the respective controls, It means that through the parameter θ z and θ a The distribution function of simultaneous control, where a i The behavior is represented by a set of latent variables w i ,z i and i In the trajectory τ i Generated under the condition, τ i Contains the agent's previous behavior history, the state of the environment, or other information related to its behavior decisions.
[0015] To infer cooperative behavior a i z i ,ω i and i , with a new distribution function Approximate posterior distribution in, and They are w i 、z i and i The variational distribution of, given an input state s and a joint action u, the inferred encoder q with an attention layer βa Encode the state s and joint action u as a coordination skill representation z = [z 1 ,z 2 ,...,z n ] and the component variable w = [w 1 ,w 2 ,...,w n ]. Then each z i and the corresponding w i By skill encoder f βz Encode, infer component labels y i . Skill Decoder g θz From i and i Refactoring skills Action Decoderg θa τ i and z i As input, predict the action p associated with the skill θa (a i |τ i ,z i ) probability.
[0016] The parameters βa, βz, θa, θz can be trained by maximizing the log-evidence lower bound (ELBO) as follows:
[0017]
[0018] in, is a Gaussian prior, p(y) is a uniform prior that allows for different skills, represents the expectation, and KL represents the KL divergence. As the reconstruction term L r , the remaining part of formula (2) is taken as the prior term L pri , that is, L ELBO =L r -L priIn order to obtain the decoupled skill representation, the regularization coefficient ω is introduced and the new inference loss L is formed infer =-L r +ωL pri .
[0019] After training, use the state encoder f βa Get the coordination skill representation z of each agent i i , forming a coordination pattern z=[z 1 ,z 2 ,...,z n ]. By maximizing the logarithmic evidence lower bound (ELBO), the encoder can continuously optimize and extract more accurate skill representations. During training, the skill encoder can learn the coordinated skill representation of each agent, providing a basis for subsequent task execution.
[0020] Step (2) embeds the coordination pattern diagram.
[0021] In a multi-agent system, the coordination behavior between agents is often affected by the spatial distribution of entities in the environment. The spatial distribution of entities usually reflects the relative positions and dynamic changes of the agents in the physical or virtual environment. By comparing these spatial distributions with the coordination patterns of the agents, we can better understand and optimize the behavior of the multi-agent system.
[0022] However, in real applications, the number, type, and configuration of entities vary from task to task. This variation makes it very difficult to directly measure the similarity between entity spatial distributions and coordination patterns. To address this problem, we introduce the Gromov-Wasserstein distance (GWD), a graph matching method that quantifies the similarity between different entity spatial distributions by computing the difference between metric spaces. This method is particularly suitable for multi-agent systems facing dynamically changing numbers of entities.
[0023] For a spatial state s, the spatial distribution can be expressed as a metric space G s =(V s ,D s ,μ s ), where V s is a set of vertices, each vertex represents the coordinate of entity i in s, D s is a metric function, representing the distance between s is the Borel probability metric. The two entity space distributions The Gromov-Wasserstein distance between Said J p (T) measures the spatial distribution of two entities through the coupling matrix T and The weighted sum of the distance differences between , which is calculated as:
[0024]
[0025] Where T is included and The coupling matrix of the corresponding relationship between * Solved by Sinkhorn-Knopp algorithm, from T * Extract the parts related to the intelligent agent from in and Respectively represent s i and j The number of agents in T, R represents a real number set. ij Shows i and j The correspondence between agents in is used to align the coordination skill representation z i and z j .
[0026] Then the following graph regularization term is introduced:
[0027]
[0028] Among them, the sampling batch size is l, and the Euclidean distance ||z is selected i -T ij z j ||2 to measure z i and z j The discreteness of the nearest neighbor matrix W ij ∈R l×l is a weight matrix. It can be seen that the emergence of similar coordination patterns is promoted in scenarios where the spatial distribution of agents shows high similarity.
[0029] Step (3) Distributed skill selection.
[0030] After step (1) coordination skill learning and step (2) graph embedding, in order to enable the agent to flexibly select skills according to different task requirements, the present invention adopts a value-based multi-agent reinforcement learning (MARL) method. Each agent selects the most appropriate coordination skill to complete the task based on its own state information and skill representation. Apply value-based MARL to learn individual value Q i (τ i ,y i ) and the joint action value Q tot (τ,y).y i , where the individual value is the largest Q i (τi ,y i ) determines which component to choose to generate the skill representation z i Then, the action decoder g θa Can be reused to select actions related to skills in decentralized execution and give corresponding observations o i and skill indication z i as input.
[0031] In order to achieve decentralized skill selection, a separate Q network f is introduced βo , the hybrid network g(·|θ following the QMIX framework v ) and the local skill generation network g obtained based on the policy network and the behavior generation network θo To achieve decentralized skill selection. Local skill generation network g θo The skill generation network is a neural network with stacked fully connected layers combined with an attention mechanism, with both the hidden layer dimension and the attention dimension being 64. βo Estimate the Q value Q of the coordination skill generated by taking different components i (τ i ,y i ), the agent can generate its corresponding action based on its local state and observation. This enables each agent to independently select and use skills when performing tasks, while ensuring coordination and consistency among multiple agents. According to y obtained from a separate Q network i Generating coordination skill representations In order to ensure that the state encoder f βa What I Learned and z i The consistency between them introduces a skill alignment loss L align :
[0032]
[0033] By minimizing L align , the local skill generation network g θo Generated Sent to action decoder g θa , select actions related to the skill while performing them in a distributed manner.
[0034] Step (4) Build the GOSPE training system.
[0035] GOSPE's training program is divided into two phases.
[0036] In the first stage, by minimizing the inference loss L infer and graph embedding constraint L g The combination ofβa 、Skill encoder f βz , motion decoder g θa and skill decoder g θz ,The main goal of this stage is to learn accurate ,representation of coordinated skills and lay the foundation for the subsequent ,decentralized skill selection.
[0037] In the second stage, by minimizing L infer and graph embedding constraint L g to train a separate Q network f βo , local skill generation network g θo and super network f βs , super network f βs It is a key component in QMIX and is responsible for generating the weights or parameters of the hybrid network in step (3). The goal of this stage is to optimize the agent's behavior selection so that it can perform efficiently in a variety of tasks.
[0038] Beneficial effects of the present invention:
[0039] The present invention solves the problem of poor performance in multi-agent collaborative tasks due to the lack of effective coordination skill learning methods between agents. GOSPE helps agents learn general coordination skills by introducing Gaussian mixture models and graph embedding technology. It can effectively extract generalizable collaboration patterns from multiple tasks and reuse these patterns in different tasks. Through graph matching and local skill selection modules, GOSPE can achieve decentralized skill selection and improve the adaptability and collaboration efficiency of agents in different task environments.
[0040] By comparing with traditional methods, experimental results show that GOSPE has significant advantages in multi-agent collaborative tasks, especially in task generalization and coordination capabilities. In particular, when dealing with complex tasks with different numbers of entities and spatial distributions, GOSPE can efficiently learn and utilize similar coordination patterns, thereby significantly improving the efficiency of task completion and the collaborative effect between agents.
[0041] The present invention is applicable to multi-agent collaborative tasks, including complex task environments and dynamically changing task settings. It can be used in multiple tasks and perform excellent performance, which verifies the effectiveness and applicability of the method in different task environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is the overall framework diagram of the present invention.
[0043] Figure 2-a 2 is a diagram showing the use of coordination skills in a 3m task in an embodiment.
[0044] Figure 2-bThis is a diagram showing the use of coordination skills in a 4m task in the embodiment.
[0045] Figure 2-c It is a diagram showing the use of coordination skills in the 9m and 10m tasks in the embodiment.
[0046] Figure 2-d This is a diagram showing the use of coordination skills in the 10m and 11m tasks in the embodiment.
[0047] Figure 3 It is a skill demonstration diagram in the embodiment. DETAILED DESCRIPTION
[0048] The specific implementation of the present invention is further described below in conjunction with the accompanying drawings and technical solutions.
[0049] The overall framework of the multi-agent collaborative skill learning method based on Gaussian mixture skill learning and graph pattern embedding in the present invention is as follows: Figure 1 As shown in Figure 1, the input is an offline multi-task dataset, and the output is a complex action decision and its generalization ability in unseen tasks. The overall framework can be roughly divided into three parts: Gaussian mixture skill inference, collaborative pattern graph embedding, and decentralized skill selection. The following takes the task in the "StarCraft II Micro-Operation Challenge (SMAC)" as an example, combined with Figure 2-a , Figure 2-b , Figure 2-c , Figure 2-d and Figure 3 The specific implementation of the present invention is described in detail. The SMAC task requires controlling multiple collaborative units to complete the task of fighting against the built-in AI. The task scenarios are diverse, such as battles between homogeneous units (such as the "3m" task) and battles between heterogeneous units (such as the "2s3z" task). The specific functions of each part are as follows:
[0050] (1) Gaussian mixture skill inference. In the SMAC task, the collaborative behavior of each agent depends on the current task scenario and state. For example, in the "3m" task, the three combat players in the game need to coordinate their division of labor to defeat the other three AI-controlled combat players. Through the Gaussian mixture skill inference module, the potential skill representations behind these behaviors are extracted, such as the skill of avoiding enemy fire and the skill of concentrated fire attack.
[0051] Its input is a sequence of states and joint actions. In the "3m" task, the state encoder receives the current position, health value, and position information of each naval unit. As shown in formula (1), these skill representations are learned using variational inference, and the latent skill expression (such as Figure 2-aSkill 1, Skill 2, and Skill 3 in the game are used to obtain coordinated skill expressions suitable for further skill allocation. Skill 1 represents "avoiding firepower"; Skill 2 represents "concentrating firepower on a specific enemy"; and Skill 3 corresponds to "forward pressure, strengthening and attracting firepower." These skills provide the basis for the execution of subsequent tasks.
[0052] (2) Collaborative pattern graph embedding. In the multi-task environment of SMAC, such as the "10m_vs_11m" task (10 friendly Marines against 11 enemy AI Marines), the spatial distribution of entities and the enemy-friendly interaction relationship are crucial to the success of the task. The present invention uses the Gromov-Wasserstein distance to quantify the similarity of the spatial distribution of entities in different task scenarios. For example, in "3m" and "10m_vs_11m", the overall pattern of unit distribution is similar, reflecting the common strategy of friendly units in cover and attack.
[0053] It receives the state sequence from the offline dataset as input. In the "10m_vs_11m" task, the spatial distribution matrix is constructed by the position and distance of each unit. These state sequences record in detail the position, state and other information of each entity during the task execution. The Gromov-Wasserstein distance is used to calculate the similarity between the spatial distributions of entities, and this similarity is embedded in skill learning through the graph regularization term (such as formula (4)), and finally outputs a skill expression with pattern consistency.
[0054] (3) Decentralized skill selection. Through local observation information and skill expression, the MARL framework based on value function is used to learn individual value functions and joint value functions to provide the agent with value evaluations on different action choices. Subsequently, the hybrid network is used to generate skill-related actions, thereby realizing a flexible skill selection strategy in dynamic entity distribution tasks.
[0055] During the task execution phase, each agent dynamically selects appropriate skills based on the current state to achieve complex collaborative behavior. Figure 3 The specific application of skills in the "10m_vs_11m" mission is shown. In the early stage of the mission, friendly units use skill 3 to form a forward pressure formation to attract enemy firepower; in the middle and late stages of the mission, low-health units switch to skill 1 to avoid damage, while healthy units continue to use skill 3 to maintain firepower suppression. At the same time, skill 2 (concentrated fire attack) is used to carry out targeted strikes on high-threat enemy units, thereby optimizing team combat strategies.
[0056] In summary, this paper proposes an offline multi-task collaborative skill learning method that can generalize to unseen tasks, and verifies its performance advantages and flexibility in complex multi-agent tasks.
Claims
1. A multi-agent Gaussian mixture skill learning method based on pattern graph embedding, characterized in that: The steps include: Step (1) constructing a Gaussian mixture variational inference encoder; Construct a Gaussian mixture variational inference encoder to extract coordination skill representation from a static multi-task dataset; specifically: The behavior of the agent is regarded as generated by multiple potential coordination skills, and an autoencoder structure is designed to encode and decode the behavior data of the agent; assuming that the cooperative behavior a of agent i i Coordination skills Generated, is the set of latent variables, z i It is sampled from the prior skill distribution; We choose Gaussian mixture distribution as the prior distribution, where a cooperative behavior of an agent can be represented as a component of a Gaussian distribution mixture; therefore, each agent’s cooperative behavior a i The generation process of p θz,θa (w i ,y i ,z i ,a i |τ i )=p(w i )p(y i ) θz (z i |w i ,y i ) θa (a i |τ i ,z i ), where p(w i ) and p(y i ) are the latent variables w of agent i i and i The prior distribution of θz and p θa It means that through the parameter θ z and θ a The distribution function of the control is respectively θz,θa It means that through the parameter θ z and θ a The distribution function of simultaneous control, where a i The behavior is represented by a set of latent variables w i ,z i and i In the trajectory τ i Generated under the condition, τ i Contains the agent's previous behavior history, the state of the environment, or other information related to its behavior decisions; To infer cooperative behavior a i z i ,ω i and i , with a new distribution function Approximate posterior distribution in, and They are w i 、z i and i The variational distribution of, given an input state s and a joint action u, the inferred encoder q with an attention layer βa Encode the state s and joint action u as a coordination skill representation z = [z 1 ,z 2 ,...,z n ] and the component variable w = [w 1 ,w 2 ,...,w n ]; then each z i and the corresponding w i By skill encoder f βz Encode, infer component labels y i ; Skill decoder g θz From i and i Refactoring skills Action Decoderg θa τ i and z i As input, predict the action p associated with the skill θa (a i |τ i ,z i ) probability; The parameters βa, βz, θa, θz are trained by maximizing the log-evidence lower bound ELBO, as follows: in, is a Gaussian prior, p(y) is a uniform prior that allows for different skills, represents the expectation, KL represents the KL divergence; replace formula (2) As the reconstruction term L r , the remaining part of formula (2) is taken as the prior term L pri , that is, L ELBO =L r -L pri ; In order to obtain the decoupled skill representation, the regularization coefficient ω is introduced and the new inference loss L is formed infer =-L r +ωL pri ; After training, use the state encoder f βa Get the coordination skill representation z of each agent i i , forming a coordination pattern z=[z 1 ,z 2 ,...,z n ]; By maximizing the logarithmic evidence lower bound, the encoder can continuously optimize and extract more accurate skill representations; during training, the skill encoder learns the coordinated skill representation of each agent, providing a basis for subsequent task execution; Step (2) embedding the coordination pattern diagram; The Gromov-Wasserstein distance is introduced to quantify the similarity between different entity space distributions by calculating the difference between metric spaces; specifically as follows: For a spatial state s, the spatial distribution is represented as a metric space G s =(V s ,D s ,μ s ), where V s is a set of vertices, each vertex represents the coordinate of entity i in s, D s is a metric function, representing the distance between s is the Borel probability metric; the two entity space distributions The Gromov-Wasserstein distance between Said J p (T) measures the spatial distribution of two entities through the coupling matrix T and The weighted sum of the distance differences between , which is calculated as: Where T is included and The coupling matrix of the corresponding relationship between * Solved by Sinkhorn-Knopp algorithm, from T * Extract the parts related to the intelligent agent from in and Respectively represent s i and j The number of agents in, R represents a real number set; T ij Shows i and j The correspondence between agents in is used to align the coordination skill representation z i and z j ; Then the following graph regularization term is introduced: Among them, the sampling batch size is l, and the Euclidean distance ||z is selected i -T ij z j ||2 to measure z i and z j The discreteness of the nearest neighbor matrix W ij ∈R l×l is a weight matrix; Step (3) decentralized skill selection; In order to enable the agents to flexibly select skills according to different task requirements, a value-based multi-agent reinforcement learning (MARL) method is adopted; each agent selects the most appropriate coordination skill to complete the task based on its own state information and skill representation; value-based MARL is used to learn the individual value Q i (τ i ,y i ) and the joint action value Q tot (τ,y).y i , where the individual value is the largest Q i (τ i ,y i ) determines which component to choose to generate the skill representation z i ; Then, the action decoder g θa is reused to select actions related to skills during decentralized execution and give corresponding observations o i and skill indication z i As input; In order to achieve decentralized skill selection, a separate Q network f is introduced βo , the hybrid network g(·|θ following the QMIX framework v ) and the local skill generation network g obtained based on the policy network and the behavior generation network θo To achieve decentralized skill selection; local skill generation network g θo The skill generation network is a neural network with stacked fully connected layers combined with an attention mechanism, with both the hidden layer dimension and the attention dimension being 64; the separate Q network f βo Estimate the Q value Q of the coordination skill generated by taking different components i (τ i ,y i ), the agent generates its corresponding actions based on its local state and observations; according to y obtained from a separate Q network i Generating coordination skill representations In order to ensure that the state encoder f βa What I Learned and z i The consistency between them introduces a skill alignment loss L align : By minimizing L align , the local skill generation network g θo Generated Sent to the action decoder g θa In the process, select actions related to the skill while performing them in a decentralized manner; Step (4) constructing a GOSPE training system; GOSPE’s training program consists of two phases; In the first stage, by minimizing the inference loss L infer and graph embedding constraint L g The combination of βa , Skill Encoderf βz , motion decoder g θa and skill decoder g θz ,The goal of this stage is to learn accurate coordination skill representation and lay the foundation for the subsequent ,distributed skill selection; In the second stage, by minimizing L infer and graph embedding constraint L g to train a separate Q network f βo , local skill generation network g θo and hypernetwork βs , super network f βs Responsible for generating the weights or parameters of the hybrid network in step (3); the goal of this stage is to optimize the agent's behavior selection so that it can perform efficiently in a variety of tasks.