Multi-agent reinforcement learning method, device, equipment, medium and product

By introducing the self-attention mechanism and action embedding method into the multi-agent system, the sample efficiency and scalability of multi-agent reinforcement learning are optimized, the problem of difficulty in capturing the interaction relationship between agents in traditional methods is solved, and efficient adversarial strategy learning and decision-making are achieved.

CN120633759APending Publication Date: 2025-09-12NORTH CHINA UNIVERSITY OF TECHNOLOGY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510943035.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Traditional multi-agent reinforcement learning has low sample efficiency and scalability, making it difficult to effectively capture the complex interactions between agents and the similarities between actions, resulting in the model having difficulty learning adversarial strategies under high-dimensional sparse representation.

Method used

By combining the self-attention mechanism with action embedding, the weighted directed graph of the multi-agent system is processed to generate an attention weight matrix and a direction embedding vector, which captures the interaction relationship between agents and converts high-dimensional discrete actions into low-dimensional continuous actions, thereby optimizing the learning efficiency of the model.

Benefits of technology

It improves the sample efficiency and scalability of multi-agent reinforcement learning, enabling faster and more stable learning and decision-making in complex multi-agent adversarial tasks, significantly improving the learning efficiency and strategy effectiveness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633759A_ABST
    Figure CN120633759A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent reinforcement learning method and device, equipment, a medium and a product, and relates to the field of multi-agent adversarial decision, and the method comprises the steps: carrying out the processing of a weighted directed graph of a multi-agent system, and obtaining a vertex embedding vector of each vertex and an edge embedding vector of each edge; according to the vertex embedding vector, obtaining the similarity between every two adjacent vertexes; processing the similarity matrix by adopting a self-attention mechanism to obtain an attention weight matrix; according to the attention weight and the edge embedding vector between every two adjacent vertexes, calculating a weighted edge embedding vector; according to the weighted edge embedding vector of each edge and the vertex embedding vector of each vertex, obtaining direction embedding vectors corresponding to adjacent vertexes; and obtaining a confrontation strategy according to the direction embedding vectors corresponding to every two adjacent vertexes and the action embedding vectors corresponding to every two adjacent vertexes. According to the method and the device, the sample efficiency and the expandability of multi-agent reinforcement learning can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of multi-agent adversarial decision-making, and in particular to a multi-agent reinforcement learning method, device, equipment, medium and product. Background Art

[0002] Deep reinforcement learning technology has been widely studied in the field of multi-agent adversarial decision-making, and has made significant contributions to research in areas such as multi-agent competition and cooperation, collaborative planning, and formation control. For example, OpenAI proposed a multi-agent deep deterministic policy gradient algorithm based on stochastic games and applied it to scenarios such as collaborative navigation and coordinated roundups. A mean-field multi-agent reinforcement learning algorithm considers large-scale coordinated strikes between red and blue agents, planning the attack or movement direction for each agent. A hierarchical graph attention network and multi-agent reinforcement learning are used to achieve collaborative decision-making in navigation and pursuit scenarios. Curriculum learning is used to improve the generalization of reinforcement learning algorithms by gradually increasing the number of trained agents. A hierarchical multi-agent reinforcement learning framework is used to enable agents to promote cooperation through dynamic reward sharing. Considering the impact of received information on agent strategies, a multi-agent reinforcement learning framework uses feature extraction based on recurrent neural networks and information representation based on a hierarchical attention mechanism to effectively improve agent decision-making efficiency. Considering the complex interactions between agents, a coupled representation method based on cognitive differences is proposed. Deep feature extraction of topological interaction information is achieved through cognitive difference networks and coupled cognitive networks, thereby promoting agent cooperation. However, traditional multi-agent reinforcement learning suffers from low sample efficiency and scalability. Summary of the Invention

[0003] The purpose of this application is to provide a multi-agent reinforcement learning method, device, equipment, medium and product that can improve the sample efficiency and scalability of multi-agent reinforcement learning.

[0004] To achieve the above objectives, this application provides the following solutions:

[0005] In a first aspect, the present application provides a multi-agent reinforcement learning method, comprising:

[0006] The weighted directed graph of the multi-agent system is processed to obtain the vertex embedding vector of each vertex and the edge embedding vector of each edge in the weighted directed graph of the multi-agent system; the weighted directed graph of the multi-agent system is a weighted directed graph constructed with each agent in the multi-agent system as a vertex and the interaction relationship between the agents as an edge;

[0007] Obtaining a similarity matrix based on the vertex embedding vector of each vertex; the similarity matrix includes the similarity between every two adjacent vertices;

[0008] The similarity matrix is ​​processed using a self-attention mechanism to obtain an attention weight matrix; the attention weight matrix includes: the attention weight between every two adjacent vertices;

[0009] Calculate the weighted edge embedding vector of each edge based on the attention weight matrix and the edge embedding vector of each edge;

[0010] According to the weighted edge embedding vector of each edge and the vertex embedding vector of each vertex, the direction embedding vector corresponding to each two adjacent vertices is obtained;

[0011] The adversarial strategy is obtained based on the direction embedding vector corresponding to every two adjacent vertices and the action embedding vector corresponding to every two adjacent vertices.

[0012] In a second aspect, the present application provides a multi-agent reinforcement learning device, comprising:

[0013] The vertex and edge encoding module is used to process the weighted directed graph of the multi-agent system to obtain the vertex embedding vector of each vertex and the edge embedding vector of each edge in the weighted directed graph of the multi-agent system. The weighted directed graph of the multi-agent system is a weighted directed graph constructed with each agent in the multi-agent system as a vertex and the interaction relationship between the agents as an edge.

[0014] A node information aggregation and transmission module is used to obtain a similarity matrix based on the vertex embedding vector of each vertex; the similarity matrix includes the similarity between every two adjacent vertices;

[0015] An attention weight calculation module is used to process the similarity matrix using a self-attention mechanism to obtain an attention weight matrix; the attention weight matrix includes the attention weight between every two adjacent vertices, and calculates the weighted edge embedding vector of each edge based on the attention weight matrix and the edge embedding vector of each edge;

[0016] The direction embedding and strategy generation module is used to obtain the direction embedding vector corresponding to every two adjacent vertices based on the weighted edge embedding vector of each edge and the vertex embedding vector of each vertex, and to obtain the adversarial strategy based on the direction embedding vector corresponding to every two adjacent vertices and the action embedding vector corresponding to every two adjacent vertices.

[0017] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement any of the multi-agent reinforcement learning methods described above.

[0018] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the multi-agent reinforcement learning methods described above.

[0019] In a fifth aspect, the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the multi-agent reinforcement learning methods described above.

[0020] According to the specific embodiments provided in this application, this application has the following technical effects:

[0021] The present application provides a multi-agent reinforcement learning method, apparatus, equipment, medium and product. In traditional multi-agent reinforcement learning, actions are usually directly input into the model in discrete or continuous form. As adversarial tasks become more and more complex, the number of agents required is also increasing. The direct input of a large number of actions easily leads to high-dimensional sparse representation, which makes it difficult to capture the similarities or relationships between actions, making it difficult for the model to learn the potential characteristics of the actions, thereby affecting adversarial strategy learning. The attention mechanism focuses on the interactive information that is more important for task completion and ignores some inefficient exploration actions. Adding action embedding is to convert high-dimensional discrete actions into low-dimensional continuous actions to capture the similarities and relationships between actions. The present application combines the attention mechanism with action embedding to reduce the dimension of the action state space and omit redundant information, thereby improving the sample efficiency and scalability of multi-agent reinforcement learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0023] Figure 1 is a weighted directed graph;

[0024] Figure 2 To obtain the flow chart of interaction embedding;

[0025] Figure 3 It is a flowchart of the interactive network;

[0026] Figure 4 A framework diagram of a multi-agent reinforcement learning method provided in another embodiment of the present application;

[0027] Figure 5 A schematic diagram of a multi-agent reinforcement learning method provided in another embodiment of the present application;

[0028] Figure 6 A flowchart of a multi-agent reinforcement learning method provided in one embodiment of the present application;

[0029] Figure 7 A schematic diagram of the functional modules of a multi-agent reinforcement learning device provided in another embodiment of the present application.

[0030] Figure 8 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0031] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0032] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0033] In an exemplary embodiment, Figure 4 、 Figure 5 and Figure 6 As shown, a multi-agent reinforcement learning method is provided, including:

[0034] Step 201: Process the weighted directed graph of the multi-agent system to obtain vertex embedding vectors for each vertex and edge embedding vectors for each edge in the weighted directed graph of the multi-agent system. The weighted directed graph of the multi-agent system is constructed with each agent in the multi-agent system as a vertex and the interactions between agents as edges. The multi-agent system can be applied to adversarial tasks.

[0035] Step 202: Obtain a similarity matrix based on the vertex embedding vector of each vertex; the similarity matrix includes the similarity between every two adjacent vertices.

[0036] Step 203: Use the self-attention mechanism to process the similarity matrix to obtain an attention weight matrix; the attention weight matrix includes: the attention weight between every two adjacent vertices.

[0037] Step 204: Calculate the weighted edge embedding vector of each edge based on the attention weight matrix and the edge embedding vector of each edge.

[0038] Step 205: Obtain the direction embedding vector corresponding to every two adjacent vertices based on the weighted edge embedding vector of each edge and the vertex embedding vector of each vertex.

[0039] Step 206: Obtain a confrontation strategy based on the direction embedding vector corresponding to every two adjacent vertices and the action embedding vector corresponding to every two adjacent vertices.

[0040] In a multi-agent system, in order to effectively capture the interaction between agents, such as attack, defense, and pursuit, it is usually necessary to consider the specific representation of the agents. This application uses a graph representation method to model the interaction between multiple agents. The agents are represented as vertices in the graph, and the interactions between agents are represented as edges in the graph. Specifically, given a multi-agent system containing N agents, a weighted directed graph G = (V, E) is constructed, such as Figure 1 As shown, V is a vertex set representing an agent, E is an edge set representing the interaction relationship between agents, and k1, k2, k3, and k4 are the weight values ​​of the corresponding edges. In another exemplary embodiment of the present application, the process of constructing a weighted directed graph of a multi-agent system is as follows:

[0041] like Figure 5 As shown in the flowchart of the network model, the data acquisition module first obtains the state information in the environment (observation data of each agent and its surrounding environment, and the observation data of the surrounding environment includes the state of other agents, obstacles, target points, terrain features and connection relationships), and then performs conventional processing on the state information to generate feature data of edges and vertices in the graph structure. Based on the feature data, a weighted directed graph G = (V, E) is generated. The feature data includes: vertex features and edge features. Vertex features are used to describe the attributes of each agent or environmental entity, such as: the state of the agent (position, speed, energy, task progress, etc.). Attributes of environmental entities (obstacle type, target point coordinates, etc.). Edge features are used to describe the relationship between agents or entities, such as: interaction intensity (communication frequency, inverse of distance). The directionality of the edge (such as information flow direction, physical action direction).

[0042] In another exemplary embodiment of the present application, a weighted directed graph of a multi-agent system is processed to obtain a vertex embedding vector of each vertex and an edge embedding vector of each edge in the weighted directed graph of the multi-agent system, specifically including:

[0043] The weighted directed graph of the multi-agent system is input into the vertex encoder to obtain the vertex embedding vector of each vertex in the weighted directed graph of the multi-agent system.

[0044] The weighted directed graph of the multi-agent system is fed into the edge encoder to obtain the edge embedding vector for each edge in the weighted directed graph of the multi-agent system. The vertex embedding vector contains local information about the agent itself, such as its position and health, while the edge embedding vector contains the interaction between the two agents corresponding to the edge.

[0045] In practical applications, both vertex encoders and edge encoders are graph neural networks.

[0046] In a multi-agent system, there are m agents, each corresponding to a vertex. The vertex representation should not only reflect its individual characteristics but also effectively capture the interactions between them. To achieve this goal, a neural network-based vertex embedding method is used to map the vertex features of the agents into a low-dimensional embedding space.

[0047] Let the state of vertex i be represented by S i , the features of the agent are mapped to a low-dimensional embedding space through the embedding function f, i.e., the vertex encoder, to obtain the vertex embedding vector v of vertex i i , specifically expressed as v i =f(S i ), where i∈(1,...,n), and the vertex encoder is a fully connected layer of nonlinear units. The vertex embedding operation captures the feature information of vertex i and represents it as a vector.

[0048] In a multi-agent system, every connection between two agents can be represented as an edge. There are two main types of edges: interaction edges, such as the distance between agents, and action edges, such as an agent's movement and attack. Because interactions are directional, edges can be represented as directed edges.

[0049] In order to extract the features of the edge, we define an edge encoder f edge , the edge encoder transforms the edge state feature h ij Mapped to edge embedding vector e ij , specifically expressed as: e ij =f edge (h ij ), where h ij Represents the feature vector consisting of the state between vertex i and vertex j, such as distance, speed, etc. ij The edge embedding vector representing the edge connecting vertex i and vertex j.

[0050] In another exemplary embodiment of the present application, a similarity matrix is ​​obtained based on the vertex embedding vector of each vertex, specifically including:

[0051] The vertex embedding vectors of each vertex are input into the Graph Convolutional Network (GCN) to obtain a feature matrix; the feature matrix includes the aggregated features corresponding to each edge. The vertex embedding vectors of adjacent vertices are aggregated together through the Graph Convolutional Network, which is expressed as: Among them, Y is the feature matrix, σ() represents the graph convolutional network, V is the matrix composed of the vertex embedding vectors of each vertex, and W is the learnable transformation matrix. is a symmetric normalized adjacency matrix containing self-loops, represents a diagonal matrix containing self-loops, and I represents the identity matrix.

[0052] According to the feature matrix, the query and key value of each vertex are obtained. Through the message transmission mechanism, the query and key value of the agent behavior are calculated: Q i =Y ij W Q , K j =Y ij W K , where Y ij represents the aggregated features corresponding to the edge formed by the connection between vertex i and vertex j, Q i represents the query of vertex i, K j represents the key value of vertex j, W Q , W K is a learnable transformation matrix.

[0053] According to the query and key value of each vertex, the similarity between each two adjacent vertices is calculated to obtain the similarity matrix. In order to calculate the similarity of the propagation relationship between agents, the formula Calculate the similarity between adjacent vertices i and j, where i, j∈(1,...,n), d represents Q i and K j The dimension of the embedding vector of , represents the dot product, K j The transpose of .

[0054] In another exemplary embodiment of the present application, the self-attention mechanism is used to process the similarity matrix to obtain an attention weight matrix; the attention weight matrix includes: the attention weight between each two adjacent vertices, specifically:

[0055] According to the formula Calculate the attention weight A between vertex i and vertex j ij , where vertex i is adjacent to vertex j, exp() represents the exponential function with the natural constant e as the base, sim(Q i ,K j ) represents the similarity between vertex i and vertex j, sim(Q i ,K k ) represents the similarity between vertex i and vertex k, vertex i is adjacent to vertex k, and N represents the total number of all vertices adjacent to vertex i;

[0056] Construct an attention weight matrix based on the attention weights between every two adjacent vertices.

[0057] In another exemplary embodiment of the present application, the weighted edge embedding vector of each edge is calculated based on the attention weight matrix and the edge embedding vector of each edge. Specifically expressed as: ij ′=Aij e ij , where e ij ′ represents the edge embedding vector after weighting the edge connecting vertex i and vertex j.

[0058] In another exemplary embodiment of the present application, the direction embedding vector corresponding to each two adjacent vertices is obtained based on the weighted edge embedding vector of each edge and the vertex embedding vector of each vertex, specifically including:

[0059] For any vertex, the edge embedding vectors of all the weighted edges corresponding to the vertex are averaged and pooled to obtain the aggregated edge embedding representation corresponding to the vertex. Specifically, in a system containing m agents, each vertex i has a logical interaction relationship with the other m-1 vertices. Therefore, by performing an average pooling operation on the edge embedding vectors of the m-1 weighted edges, the aggregated edge embedding representation corresponding to the vertex i can be obtained. The formula is expressed as

[0060] like Figure 2 As shown, based on the aggregated edge embedding representation corresponding to the vertex, the vertex embedding vector of the vertex, and the vertex embedding vector of the target point, the direction representation corresponding to the vertex and the target point is obtained; the target point is any vertex adjacent to the vertex. Specifically, the vertex embedding vector of the vertex and the vertex embedding vector of the target point are concatenated, and then the concatenated vector is added or concatenated with the aggregated edge embedding representation corresponding to the vertex to obtain the direction representation corresponding to the vertex and the target point. In order to further capture the interactive relationship between intelligent agents, the vertex embedding vector of a single intelligent agent is concatenated with the vertex embedding vector of its neighbor.

[0061] The direction representation corresponding to the vertex and the target point is input into the direction encoder to obtain the direction embedding vector corresponding to the vertex and the target point. dir A neural network (which can be a fully connected layer) generates the direction embedding vector d corresponding to vertex i and vertex j ij The direction embedding vector here includes: relative position, distance and angle, etc. The calculation formula is: d ij =f dir (o i ,o j ), where o i , o j They represent the position state features of vertex i and vertex j, such as coordinates and velocity vectors, respectively. The direction encoder is used to capture relationships such as relative position, distance, and angle.

[0062] In another exemplary embodiment of the present application, a confrontation strategy is obtained according to the direction embedding vector corresponding to each two adjacent vertices and the action embedding vector corresponding to each two adjacent vertices, such as Figure 2 As shown, specifically including:

[0063] For any two adjacent vertices, the direction embedding vectors corresponding to the two adjacent vertices and the action embedding vectors corresponding to the two adjacent vertices are concatenated to obtain the interaction embedding vectors corresponding to the two adjacent vertices. Specifically, the action embedding vector e ij act and the direction embedding vector d ij Perform vector splicing to obtain the interaction embedding vector corresponding to vertex i and vertex j Add action embedding vectors to capture high-dimensional discrete actions.

[0064] For any vertex, the interaction embedding vectors corresponding to the vertex and each vertex in the target point set are pooled to obtain the overall entity embedding vector corresponding to the vertex; the target point set includes all vertices adjacent to the vertex. The overall entity embedding vector z of vertex i is obtained by performing Pooling() on the interaction embedding vectors of vertex i and the remaining m-1 vertices. i , the formula is In this way, the action information and direction information of the agent can be integrated into the state representation, enhancing the consideration of high-dimensional spatial relationships in the decision-making process and better capturing the complex relationships in the multi-agent system.

[0065] The overall entity embedding vector corresponding to each vertex is input into the gated recurrent unit to obtain the adversarial strategy. Specifically, the gated recurrent unit (GRU) captures the temporal dependency of the interactive information through the forward propagation process, resets the gate to the previous hidden state h t-1 and the current input information x t That is, the overall entity embedding vectors corresponding to each vertex are spliced, and then the data is scaled by the tanh activation function to obtain the new information h′. Then, the update gate selectively memorizes the information h′ and selectively forgets the hidden state h t-1 , the implementation process is as follows: h t =(1-z)·h t-1 +z·h′, the final output strategy, where h t Represents the final hidden state at the current moment, z represents the update gate, and the strategy is the action distribution or value function value generated after the final hidden state is input into the fully connected layer.

[0066] In practical applications, the specific process of determining the action embedding vector is as follows:

[0067] When making decisions, multiple agents must not only consider their individual states, but also the impact of their action choices on other agents or the environment. This application maps the discrete action space of the agent to a low-dimensional continuous space to generate an action embedding vector This enables the model to capture the similarities and differences between actions, thereby enhancing the learning efficiency of the model.

[0068] First, we extract the original features of discrete actions, such as movement, attack, and defense, from the multi-agent adversarial environment. We define the possible actions of all agents to form an action set. Among them, N a Represents the total number of actions. Each action a in the set m' will be assigned a unique index m'∈{1,2,...,N a In order to map the discrete action set to the continuous space, an embedding matrix is ​​introduced Among them, k is the dimension of the embedding vector, which refers to the dimension occupied by the action in the continuous vector space. Each row of the embedding matrix E represents the embedding vector of an action. Specifically, action a m' The embedding vector e ij act It can be obtained by looking up the m′th row of the embedding matrix: ij act =E[m′], where E[m′] represents the vector of values ​​in the m′th row of the embedding matrix. To ensure uniform distribution of the initialized embedding vector, the embedding matrix is ​​usually initialized with a Gaussian distribution and expressed as: E m′q ~N(0,σ 2 ), where N(0,σ 2 ) means the mean is 0 and the variance is σ 2 Gaussian distribution, σ 2 is the initialization parameter, E m′q Represents the value of the m′th row and qth column of the embedding matrix.

[0069] This application calculates the weights between adjacent vertices through the self-attention mechanism based on vertex embedding vectors. The process of constructing attention weights is called graph interaction network. Figure 3As shown. The graph interaction network proposed in this application is a process of using a neural network for training to generate a numerical matrix |A|×(|A|+|O|), where |A| represents the number of agents, and |O| is the maximum number of given environmental objects (environmental objects refer to entities or elements that exist in the environment where the agent is located and may affect the agent's behavior or task completion. They can be physical entities such as obstacles, tools, target items, or logical entities such as task goals, signal markers, etc.). The graph represents the relationship between agents and between agents and environmental objects. For objects that do not exist in the current environment, their corresponding weights are set to 0. The higher the absolute value of the edge weight between agent A and another agent B or the environment O, the more important B or O is to the completion of the task of agent A. The feature representation of each agent is extracted from the environment and mapped to the vertex embedding space.

[0070] In order to further analyze the relationship between different vertices, this application calculates the attention weight matrix Each element A ij Both represent the attention weight values ​​between the two agents. These attention weights are determined by aggregating information from neighboring vertices using a graph convolutional network, which is based on the query Q generated by each agent in the graph structure. i and key value K j Calculated.

[0071] To further optimize the cooperative and competitive strategies between agents, this application also introduces a self-attention mechanism to normalize the similarity matrix and generate an attention weight matrix A. The attention weight matrix is ​​used to weight the representation relationships of edges, determining the degree of attention each agent pays to other agents. These attention weights are then used to perform a weighted summation of the edge relationships, thereby updating the embedding representation of each agent. The graph interaction network is a process that calculates the attention weight matrix using node embedding representations and the attention mechanism, where the node and edge representations are continuously updated during network training.

[0072] In multi-agent collaborative decision-making and complex task solving, efficient information representation and strategy optimization are core challenges. Figure 4 As shown, this application significantly improves the decision-making efficiency of intelligent agents by integrating multi-scale feature learning and time series modeling methods. Specifically, this application adopts a dual-channel encoding mechanism to process edge attributes and vertex attributes in graph data respectively, and perform embedded representation. In this process, the attention weight matrix constructed based on multi-agent vertex information is applied to the reconstruction of edge features, so that the edge embedding representation between the two intelligent agents forms a new feature expression after the attention weight adjustment. The weighted edge embedding vector is then compressed using an average pooling operation to generate a vector representation reflecting the vertex-neighbor association: Among them, |N(i)| represents the number of adjacent nodes of node i, e ij ′ represents the weighted edge embedding vector.

[0073] This aggregated vector is then fused with the original vertex embedding vector to form an enhanced representation that combines individual attributes and neighborhood relationships. These features are combined and fed into the policy generation unit along with the vertex embedding vector. A dedicated directional encoder extracts directional features of the network topology, and action embedding vectors are introduced to model dynamic interaction patterns in high-dimensional space, ensuring that the decision module can effectively integrate historical state information and environmental action values. During policy inference, a temporal processing mechanism using gated recurrent units is used to capture temporal correlations in decision sequences.

[0074] This application uses a graph convolutional network to aggregate information about agents and their neighbors in the environment, calculates the attention weights between them, and uses a self-attention mechanism to generate a graph interaction network containing a weight matrix. This weights the aggregated edge representations, focusing on information that is more important for task completion. Combining action embedding vectors and direction embedding vectors, the high-dimensional discrete action space is converted into a low-dimensional continuous action space, better capturing the complex interactions between multiple agents and reducing the complexity of the action state space. This method is particularly suitable for complex and dynamic multi-agent adversarial tasks, improving the efficiency and stability of task completion in adversarial scenarios.

[0075] This application is based on graph theory and uses graph neural networks to model the interaction relationships of multiple agents. It combines action embedding vectors and direction embedding vectors to capture discrete action spaces in high-dimensional environments and convert them into low-dimensional, easy-to-process continuous action spaces. It optimizes the representation of the interaction relationships between multiple agents in complex scenarios, represents the agents as vertices in the graph, and the interaction relationships between agents as edges in the graph. It can effectively deal with the dimensionality disaster problem caused by the increase in the number of multiple agents in complex scenarios, and provides a basis for improving the learning efficiency of multi-agent reinforcement learning.

[0076] This application combines the self-attention mechanism with action interaction relationships. First, the vertex encoder is used to obtain the vertex information of multiple agents, and the attention weights are calculated based on the specificity between the vertices to dynamically weight the representation relationship of the edges between the agents. The weighted edge representation relationships are then aggregated and combined with the corresponding vertex embedding vectors. After adding the action embedding vector and the direction embedding vector, an agent embedding vector containing the corresponding weights is generated. This method is particularly suitable for multi-agent confrontation tasks in complex and dynamic scenarios.

[0077] This application also provides an example to verify the effectiveness of the multi-agent reinforcement learning method proposed in this application:

[0078] The comparison algorithm is the best performing (State Of The Art, SOTA) method in the current task, namely the fine-tuned Q-value Mixing Network (QMIX) algorithm and the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm. This application uses a graph structure to characterize the interaction relationship in the multi-agent system and handles complex interaction topology, making it more flexible and convenient to deal with multi-agent system problems. According to the specificity of a single vertex, the attention weight relationship between each agent and the neighboring agents is calculated, focusing on the information that is more important for task completion, so that the agent can achieve the purpose of efficient exploration more quickly. Due to the direct input of a large number of discrete actions in the environment, the rewards in reinforcement learning become sparse. Therefore, the network adds action embedding vectors to the interaction process of the agent, dynamically captures the similarities and differences of the actions, and maps high-dimensional discrete actions to low-dimensional continuous actions, effectively alleviating the dimensionality disaster problem in high-dimensional environments. This application combines graph neural network modeling with self-attention mechanism to add action embedding vectors, so that the model can capture actions in high-dimensional environments, achieve more efficient exploration, and faster learning rate.

[0079] The simulation parameters are set as follows. The experiment was carried out on an Ubuntu 20.04 system with python 3.8 and torch 1.12.1 as the development environment. The simulation verification was carried out on two multi-agent confrontation platforms, namely StarCraft Multi-Agent Challenge (SMAC) and Multi-agent Particle Environment (MPE). Five maps were selected on SMAC for 10 million steps of training. Performance evaluation was performed every 10,000 training steps. This application was compared with other algorithms on different maps. The results are shown in Table 1.

[0080] Table 1 Simulation results of different algorithms in various maps in SMAC environment

[0081]

[0082] On these maps, the present application not only demonstrates strong strategic decision-making capabilities, but also can effectively respond to dynamic changes in high-dimensional space and optimize the strategies of intelligent agents. More importantly, from the perspective of training efficiency, it shows an extremely high convergence speed. In complex scenarios such as corridor, MMM2, and 3s vs 5z, the method proposed in this application requires only 10%-30% of the steps of the baseline method during training, and it has shown obvious signs of convergence and achieved the same or even higher performance as the baseline method. This shows that when dealing with complex multi-agent interaction tasks, the present application not only improves the effectiveness of the strategy, but also significantly improves the sample efficiency, especially in high-dimensional and dynamically changing environments, it can learn quickly and converge stably.

[0083] To further evaluate the effectiveness of our proposed method, we conducted an ablation experiment in an MPE environment, comparing our method with a standard fully connected neural network. The experimental results show that our proposed method exhibits a clear advantage in all scenarios on these maps, particularly on the 6 vs 2 and 7 vs 3 maps, where our network achieves a nearly 100% win rate. The results are shown in Table 2.

[0084] Table 2 Simulation results of various maps in MPE environment

[0085]

[0086] On the 6vs. 2 map, this application achieved a significant improvement in win rate compared to the baseline method, and also converged faster, with the overall difference not being significant. On the 8vs. 4 map, the win rate increased significantly, but due to the increased number of agents in the environment and certain obstacles, the win rate initially increased, then decreased, and finally leveled off after a certain training time, making convergence more difficult. On the 7vs. 3 map, not only did the win rate improve, but the algorithm's convergence speed also increased significantly.

[0087] In summary, the multi-agent reinforcement learning method proposed in this application demonstrates superior performance in multiple experimental scenarios, especially when dealing with high-dimensional complex environments.

[0088] Currently, multi-agent reinforcement learning schemes mainly include three categories: the first is to use a fully connected network to directly process the state of multiple agents, but this method cannot effectively model the interaction relationship between agents, resulting in a significant decline in performance in complex scenarios. As shown in Table 2, the MLP win rate on the 6vs2 map is only 59.27%, while this application reaches 96.36%; the second is to use the deep ensemble method. Although it meets the permutation invariance, it cannot capture the dynamic interaction characteristics between agents. For example, the fine-tuned QMIX performance in the 5m_vs_6m scenario is 12.75% lower than that of this application; the third is the traditional graph convolutional network, whose fixed edge weights make it difficult to adapt to the dynamic changes in adversarial scenarios.

[0089] However, none of these solutions can simultaneously address the following key issues: (1) the curse of dimensionality caused by high-dimensional discrete action spaces; and (2) the need for real-time modeling of interactive relationships in dynamic adversarial scenarios. This application, by combining self-attention edge weighting with action embedding technology, effectively overcomes these technical obstacles and demonstrates significant performance advantages across multiple benchmarks. As shown in Table 1, the win rate in various scenarios increased by 14.6%-76.3%, and the convergence speed increased by up to 67.3%. Therefore, this application is irreplaceable in practical application scenarios that require processing complex multi-agent interactions.

[0090] Based on the same inventive concept, embodiments of the present application also provide a multi-agent reinforcement learning device for implementing the multi-agent reinforcement learning method described above. The solution provided by this device is similar to the solution described in the method described above. Therefore, the specific limitations of one or more multi-agent reinforcement learning device embodiments provided below can be found in the above-mentioned limitations of the multi-agent reinforcement learning method and will not be further elaborated here.

[0091] In an exemplary embodiment, Figure 7 As shown, a multi-agent reinforcement learning device is provided, comprising:

[0092] The vertex and edge encoding module is used to process the weighted directed graph of the multi-agent system to obtain the vertex embedding vector of each vertex and the edge embedding vector of each edge in the weighted directed graph of the multi-agent system. The weighted directed graph of the multi-agent system is a weighted directed graph constructed with each agent in the multi-agent system as a vertex and the interaction relationship between the agents as an edge.

[0093] A node information aggregation and transmission module is used to obtain a similarity matrix based on the vertex embedding vector of each vertex; the similarity matrix includes the similarity between every two adjacent vertices;

[0094] An attention weight calculation module is used to process the similarity matrix using a self-attention mechanism to obtain an attention weight matrix; the attention weight matrix includes the attention weight between every two adjacent vertices, and calculates the weighted edge embedding vector of each edge based on the attention weight matrix and the edge embedding vector of each edge;

[0095] The direction embedding and strategy generation module is used to obtain the direction embedding vector corresponding to every two adjacent vertices based on the weighted edge embedding vector of each edge and the vertex embedding vector of each vertex, and to obtain the adversarial strategy based on the direction embedding vector corresponding to every two adjacent vertices and the action embedding vector corresponding to every two adjacent vertices.

[0096] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 8 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store multi-agent reinforcement learning data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a multi-agent reinforcement learning method is implemented.

[0097] Those skilled in the art will understand that Figure 8 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application and does not constitute a limitation on the computer device to which the solution of the present application is applied. A specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the above-mentioned method embodiments when executing the computer program.

[0098] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the above-mentioned method embodiments when executed by a processor.

[0099] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the above method embodiments are implemented.

[0100] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0101] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0102] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0103] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0104] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A multi-agent reinforcement learning method, characterized in that: The multi-agent reinforcement learning method includes: The weighted directed graph of the multi-agent system is processed to obtain the vertex embedding vector of each vertex and the edge embedding vector of each edge in the weighted directed graph of the multi-agent system; the weighted directed graph of the multi-agent system is a weighted directed graph constructed with each agent in the multi-agent system as a vertex and the interaction relationship between the agents as an edge; Obtaining a similarity matrix based on the vertex embedding vector of each vertex; the similarity matrix includes the similarity between every two adjacent vertices; The similarity matrix is ​​processed using a self-attention mechanism to obtain an attention weight matrix; the attention weight matrix includes: the attention weight between each two adjacent vertices; Calculate the weighted edge embedding vector of each edge based on the attention weight matrix and the edge embedding vector of each edge; According to the weighted edge embedding vector of each edge and the vertex embedding vector of each vertex, the direction embedding vector corresponding to each two adjacent vertices is obtained; The adversarial strategy is obtained based on the direction embedding vector corresponding to every two adjacent vertices and the action embedding vector corresponding to every two adjacent vertices.

2. The multi-agent reinforcement learning method according to claim 1, characterized in that: The weighted directed graph of the multi-agent system is processed to obtain the vertex embedding vector of each vertex and the edge embedding vector of each edge in the weighted directed graph of the multi-agent system, specifically including: Input the weighted directed graph of the multi-agent system into the vertex encoder to obtain the vertex embedding vector of each vertex in the weighted directed graph of the multi-agent system; The weighted directed graph of the multi-agent system is input into the edge encoder to obtain the edge embedding vector of each edge in the weighted directed graph of the multi-agent system.

3. The multi-agent reinforcement learning method according to claim 1, characterized in that: According to the vertex embedding vector of each vertex, the similarity matrix is ​​obtained, which specifically includes: Inputting the vertex embedding vector of each vertex into the graph convolutional network to obtain a feature matrix; the feature matrix includes the aggregated features corresponding to each edge; According to the feature matrix, the query and key value of each vertex are obtained; According to the query and key value of each vertex, the similarity between every two adjacent vertices is calculated to obtain the similarity matrix.

4. The multi-agent reinforcement learning method according to claim 1, characterized in that The self-attention mechanism is used to process the similarity matrix to obtain the attention weight matrix, which is: According to the formula Calculate the attention weight A between vertex i and vertex j ij , where vertex i is adjacent to vertex j, exp() represents the exponential function with the natural constant e as the base, sim(Q i ,K j ) represents the similarity between vertex i and vertex j, sim(Q i ,K k ) represents the similarity between vertex i and vertex k, vertex i is adjacent to vertex k, and N represents the total number of all vertices adjacent to vertex i; Construct an attention weight matrix based on the attention weights between every two adjacent vertices.

5. The multi-agent reinforcement learning method according to claim 1, characterized in that: According to the weighted edge embedding vector of each edge and the vertex embedding vector of each vertex, the direction embedding vector corresponding to each two adjacent vertices is obtained, specifically including: For any vertex, perform an average pooling operation on the weighted edge embedding vectors of all edges corresponding to the vertex to obtain the aggregated edge embedding representation corresponding to the vertex; Obtaining a direction representation corresponding to the vertex and the target point based on the aggregate edge embedding representation corresponding to the vertex, the vertex embedding vector of the vertex, and the vertex embedding vector of the target point; the target point is any vertex adjacent to the vertex; The direction representation corresponding to the vertex and the target point is input into a direction encoder to obtain a direction embedding vector corresponding to the vertex and the target point.

6. The multi-agent reinforcement learning method according to claim 1, characterized in that: The adversarial strategy is obtained based on the direction embedding vector corresponding to each two adjacent vertices and the action embedding vector corresponding to each two adjacent vertices. Specifically, it includes: For any two adjacent vertices, concatenate the direction embedding vectors corresponding to the two adjacent vertices and the action embedding vectors corresponding to the two adjacent vertices to obtain the interaction embedding vectors corresponding to the two adjacent vertices; For any vertex, pool the interaction embedding vectors corresponding to the vertex and each vertex in the target point set to obtain the overall entity embedding vector corresponding to the vertex; the target point set includes all vertices adjacent to the vertex; The overall entity embedding vector corresponding to each vertex is input into the gated recurrent unit to obtain the adversarial strategy.

7. A multi-agent reinforcement learning device, characterized in that: The multi-agent reinforcement learning device comprises: The vertex and edge encoding module is used to process the weighted directed graph of the multi-agent system to obtain the vertex embedding vector of each vertex and the edge embedding vector of each edge in the weighted directed graph of the multi-agent system. The weighted directed graph of the multi-agent system is a weighted directed graph constructed with each agent in the multi-agent system as a vertex and the interaction relationship between the agents as an edge. A node information aggregation and transmission module is used to obtain a similarity matrix based on the vertex embedding vector of each vertex; the similarity matrix includes the similarity between every two adjacent vertices; An attention weight calculation module is used to process the similarity matrix using a self-attention mechanism to obtain an attention weight matrix; the attention weight matrix includes the attention weight between every two adjacent vertices, and calculates the weighted edge embedding vector of each edge based on the attention weight matrix and the edge embedding vector of each edge; The direction embedding and strategy generation module is used to obtain the direction embedding vector corresponding to every two adjacent vertices based on the weighted edge embedding vector of each edge and the vertex embedding vector of each vertex, and to obtain the adversarial strategy based on the direction embedding vector corresponding to every two adjacent vertices and the action embedding vector corresponding to every two adjacent vertices.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multi-agent reinforcement learning method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the multi-agent reinforcement learning method described in any one of claims 1 to 6.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the multi-agent reinforcement learning method described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-robot collaborative navigation method based on hierarchical relation graph learning in dynamic environment

    CN113296502A

  • Power grid multi-agent collaborative optimization method, system and device based on graph partition and medium

    CN117521902A

  • Reinforcement learning knowledge graph reasoning method and system guided by confrontation and attention mechanism

    CN118606485A

  • Knowledge graph embedding method and system based on multi-relation knowledge enhancement graph convolutional network

    CN119089995A

  • Multi-agent cooperation decision-making and training method

    US20200125957A1