Multi-agent deep reinforcement learning training method based on weak connection theory
By applying weak connection theory in deep reinforcement learning of multi-agents, establishing interaction diagrams between agents and calculating connection strengths, generating information interaction weight matrix, the problem of low information utilization efficiency between weak connection agents in multi-agent information interaction is solved, and training efficiency is improved.
Patent Information
- Application Number
- CN202411708332.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When existing multi-agent deep reinforcement learning methods handle information interactions between agents, it is difficult to effectively utilize information interactions between weakly connected agents, resulting in inefficient training.
Based on the weak connection theory, a multi-agent interaction diagram is established, the connection intensity distribution and dominant agents are calculated between agents, the information interaction weight matrix is generated, and information interaction between agents is optimized.
Optimize information interaction through weak connection theory, effectively reduce information interaction redundancy and improve the efficiency of deep reinforcement learning and training for multiple agents.
Smart Images

Figure CN120068984A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep reinforcement learning, and in particular to a multi-agent deep reinforcement learning training method based on the weak connection theory. Background Art
[0002] Reinforcement learning is an important artificial intelligence strategy learning method and an imitation of the human learning process. It learns methods to solve problems in interaction with the environment and has far exceeded human capabilities in dealing with certain real-world problems. However, simply applying single-agent reinforcement learning methods to multi-agent scenarios brings a huge training burden due to the large amount of state information generated. Therefore, it becomes extremely important to improve the utility of interaction information and thus reduce the dimension of the exploration space. The existing multi-agent information interaction methods can be divided into two categories. One is from a global perspective, sharing all agents' observation information globally to strengthen global cooperation among agents; the second is from a correlation perspective, focusing on information interaction and cooperation between strongly correlated agents. Both try to solve the interaction and cooperation problems among agents from different scales, but pay less attention to the interaction between weakly connected agents.
[0003] In the human social system, more valuable interactions often do not occur between two people with a close relationship. Sociological research shows that weak connections can provide diverse and novel information, and can provide more potential job opportunities for individuals. At the same time, we also notice that in social interactions, the initiator of the interaction makes a greater contribution to promoting interaction between individuals.
[0004] Inspired by the research on social network connection dynamics, the present invention aims at how to screen more valuable information from the information provided by other agents. Considering the multi-agent scenario as a small interaction graph network, agents need to interact with information to better complete tasks. However, not all the observed states of friendly agents can provide valuable information. Generally, only some friendly agents can provide valuable information for the current agent. We will establish corresponding multi-agent interaction graph networks in different scenarios and demonstrate the improvement of information interaction quality using the weak connection theory. For this purpose, we designed a multi-agent deep reinforcement learning training method based on the weak connection theory. Summary of the Invention
[0005] The purpose of the present invention is to provide a multi-agent deep reinforcement learning training method based on the weak connection theory to solve the problem that the existing methods focus on information interaction and cooperation between strongly correlated agents, and both try to solve the interaction and cooperation problems among agents from different scales, but pay less attention to the interaction between weakly connected agents as mentioned in the above background art.
[0006] To achieve the above object, the present invention provides the following technical solutions: A multi-agent deep reinforcement learning training method based on the weak connection theory, comprising the following steps:
[0007] Step 1: Establish a multi-agent interaction graph based on the weak connection theory to represent the connection relationship between agents;
[0008] Step 2: Calculate the multi-agent connection strength distribution and the leading agent, and establish a connection strength distribution matrix;
[0009] Step 3: Generate an information interaction weight matrix, thereby optimizing the information interaction between agents and improving the training efficiency of multi-agent deep reinforcement learning.
[0010] Preferably, the establishment of the agent connection relationship in Step 1 is specifically as follows:
[0011] Step 1-1: Abstract a single agent as a node, including the observation range, attack range, coordinates, and direction information of the corresponding agent;
[0012] Step 1-2: Rules for establishing a multi-agent information interaction graph. The specific method for establishing a multi-agent information interaction graph with agents as graph nodes is as follows:
[0013] Among them, the connection between agents: When two agents i and j are within each other's field of vision, there is an edge W between them i,j ; The connection between sub-groups: When there are unconnected sub-graphs in the established interaction connection graph, the two agents i and j with the closest distance between the two sub-groups are most likely to interact, that is, the two groups are connected by an edge W i,j Connect;
[0014] Step 1-3: The weak and strong connections and the leading agent in the multi-agent information interaction graph are specifically as follows:
[0015] Among them, the weak connection between agents: If the connection strength between agents i and j is less than the threshold G, the connection between the two agents is a weak connection; The strong connection between agents: If the connection strength between agents i and j is greater than the threshold G, the connection between the two agents is a strong connection; Leading agent: If agent i has the largest number of connections in the current interaction graph, then agent i is the leading agent in this interaction network.
[0016] Preferably, the specific calculation formula for the establishment in Step 2 is:
[0017] Step 2-1: The calculation formula for the connection strength between agents is as follows:
[0018] S strenth = w i,j / (D i + D j + w i,j – 2);
[0019] Among them, W i,j is the number of edges on the shortest path between agents i and j, D i and D j are the number of connections between agents i and j respectively;
[0020] The calculation formula of the leading agent in Step 2-2 is as follows:
[0021]
[0022] Among them, W k,j is the number of edges on the shortest path between agents k and j, D k is the degree of agent k, and g is the multi-agent interaction graph established;
[0023] The specific definitions in the calculation formulas of the agent connection strength and the leading agent in Step 2-3 are:
[0024] Among them, a matrix of size n*n is established for the agent connection strength distribution matrix, where n is the total number of agents. The value of the element in the i-th row and j-th column of the matrix is the connection strength between agents i and j. When the row number i and the column number j are equal, the corresponding matrix element value is set to 1. Assuming that the calculated leading agent node is k, then all elements in the k-th column of the connection strength matrix are set to 0.
[0025] Preferably, the establishment information in Step 3 is specifically:
[0026] Step 3-1 Determine the connection strength threshold and generate the information interaction weight matrix:
[0027] Among them, a matrix H of the same size as the connection strength matrix is newly established. The connection strength threshold G is set to 0.3. When the element value in the connection strength matrix is greater than 0.3, the corresponding element value in the weight matrix H is set to 0, otherwise, it is set to 1, and the information interaction weight matrix H is obtained;
[0028] Step 3-2 Calculate the high-quality information required by each agent according to the information interaction weight matrix. For each agent i, one strong and weak connectivity information sharing array is:
[0029] α i ={α i,1, α i,2, ...α i,i, ...α i,n,};
[0030] The i-th row of the weight matrix, where the information received by agent i can be expressed as:
[0031]
[0032] Among them, \(O_{-i}\) is the set of observed states of all agents except agent \(i\), and \(\alpha_{-i}\) is the set of actions of all agents except agent \(i\).
[0033] In step 3-3, each agent uses high-quality information as input and combines it with the MAPPO neural network to train and update the multi-agent adversarial model. The initial network input size is \(n\times c\) dimensions, where \(n\) is the number of agents and \(c\) is the dimension of the observed information of a single agent. The output size is \(n\) dimensions, the activation function is the Relu function, and the optimizer selects the Adam optimizer.
[0034] Preferably, the information received by the agent established according to step 3-2 is:
[0035] Derive the new state value function as follows:
[0036]
[0037] Derive the calculation formulas for the joint advantage function and the action value function as follows:
[0038]
[0039] Minimize the loss function as follows:
[0040]
[0041] Among them, \(b\) is the counterfactual baseline, \(\theta\) and \(\omega\) are the parameters of the target network and the critic network respectively, and \(\alpha_{-i}\) -i represents the set of actions other than agent \(i\). The critic network updates the network according to this loss function.
[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0043] 1. Considering the interaction strength relationship between agents, an information interaction graph among multi-agents is established by combining the weak connection theory in human sociology, effectively representing the interaction relationship among multi-agents;
[0044] 2. Define the leading agent in agent interaction, share the observed information of the leading agent, and play the role of the observed information of the leading agent in model training;
[0045] 3. Define the calculation formula for the connection strength distribution, quantify the connection strength between agents, and obtain the information interaction weight matrix therefrom, optimize the information interaction process of agents, effectively reduce information interaction redundancy, and improve the training efficiency of multi-agent deep reinforcement learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1Schematic diagram of multi-agent interaction and strong / weak connection relationships;
[0047] Figure 2 Schematic diagram of the leading agent in the multi-agent interaction graph;
[0048] Figure 3 Network schematic diagram of the multi-agent deep reinforcement learning training method based on the weak connection theory in the present invention;
[0049] Figure 4 StarCraft Figure 1 Schematic diagram of the C3S5Z scenario;
[0050] Figure 5 StarCraft Figure 1 Results of establishing the agent interaction graph at different confrontation moments in C3S5Z;
[0051] Figure 6 StarCraft Figure 1 Training effects compared with different algorithms in C3S5Z;
[0052] Figure 7 This method in StarCraft Figure 1 Confrontation process in C3S5Z, the red side and our side agents. Detailed implementation manners
[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0054] Combined with Figures 1-7 , a multi-agent deep reinforcement learning training method based on the weak connection theory in the illustrated figure includes the following steps:.
[0055] Further, step 1 is to establish a multi-agent interaction graph based on the weak connection theory to represent the connection relationships between agents
[0056] The specific establishment of the agent connection relationship is as follows:
[0057] Step 1-1 abstracts a single agent into a node, including the observation range, attack range, coordinates, and direction information of the corresponding agent;
[0058] Step 1-2 rules for establishing a multi-agent information interaction graph. The specific establishment of the multi-agent information interaction graph with agents as graph nodes is as follows:
[0059] Among them, the edge connection between agents: When two agents i and j are within each other's field of vision, there is an edge connection W between them. i,j Edge connection between sub-groups: When there are unconnected sub-graphs in the established interaction connection graph, the two agents i and j with the shortest distance between the two sub-graphs are most likely to interact, that is, the two groups are connected by the edge W. i,j Connection;
[0060] Steps 1-3 The strong and weak connections and the leading agent in the multi-agent information interaction graph are as follows:
[0061] Among them, the weak connection between agents: If the connection strength between agents i and j is less than the threshold G, the connection between the two agents is a weak connection; the strong connection between agents: If the connection strength between agents i and j is greater than the threshold G, the connection between the two agents is a strong connection; the leading agent: If agent i has the largest number of connections in the current interaction graph, then agent i is the leading agent in this interaction network.
[0062] Furthermore, in step 2, calculate the connection strength distribution and the leading agent of the multi-agent, and establish a connection strength distribution matrix.
[0063] The specific calculation formula for the establishment is as follows:
[0064] The calculation formula for the connection strength between agents in step 2-1 is as follows:
[0065] S strenth = w i,j / (D i + D j + w i,j – 2)
[0066] Among them, W i,j is the number of edges on the shortest path between agents i and j, D i and D j are the number of connections between agents i and j respectively;
[0067] The calculation formula for the leading agent in step 2-2 is as follows:
[0068]
[0069] Among them, W k,j is the number of edges on the shortest path between agents k and j, D k is the degree of agent k, and g is the established multi-agent interaction graph;
[0070] The specific definitions in the calculation formulas for the connection strength of agents and the leading agent in step 2-3 are as follows:
[0071] Among them, to establish the connection strength distribution matrix of agents, a matrix of size n*n is established, where n is the total number of agents. The value of the element in the i-th row and j-th column of the matrix is the connection strength between agent i and agent j. When the row number i is equal to the column number j, the corresponding matrix element value is set to 1. Assuming that the calculated leading agent node is k, then all elements in the k-th column of the connection strength matrix are set to 0.
[0072] Furthermore, step 3 generates the information interaction weight matrix, thereby optimizing the information interaction between agents and improving the training efficiency of multi-agent deep reinforcement learning.
[0073] The specific establishment information is as follows:
[0074] Step 3-1 determines the connection strength threshold and generates the information interaction weight matrix;
[0075] Among them, a matrix H of the same size as the connection strength matrix is newly established. The connection strength threshold G is set to 0.3. When the element value in the connection strength matrix is greater than 0.3, the corresponding element value in the weight matrix H is set to 0; otherwise, it is set to 1, obtaining the information interaction weight matrix H.
[0076] Step 3-2 calculates the high-quality information required by each agent according to the information interaction weight matrix. For each agent i, one strong and weak connectivity information sharing array is:
[0077] α i ={α i,1, α i,2, ...α i,i, ...α i,n,};
[0078] The i-th row of the weight matrix, where the information received by agent i can be expressed as:
[0079]
[0080] Among them, O -i is the set of observation states of all agents except agent i, and ɑ -i is the set of actions of all agents except agent i;
[0081] Step 3-3 takes the high-quality information adopted by each agent as the input, combines it with the MAPPO neural network, and realizes the training and update of the multi-agent confrontation model. The initial network input size is n*c dimensions, where n is the number of agents and c is the dimension of the observation information of a single agent. The output size is n dimensions, the activation function is the Relu function, and the optimizer selects the Adam optimizer.
[0082] Furthermore, according to step 3-2, the information received by the established agent is
[0083] Derive the new state value function as follows:
[0084]
[0085] Derive the calculation formulas for the joint advantage function and the action value function as follows:
[0086]
[0087] Minimize the loss function as follows:
[0088]
[0089] where b is the counterfactual baseline, and ω are the parameters of the target network and the critic network respectively, and ɑ -i represents the set of actions outside agent i. The critic network updates the network according to this loss function.
[0090] A specific implementation example of the algorithm process of the multi-agent deep reinforcement learning training method based on the weak connection theory proposed by the present invention includes a multi-agent interaction graph modeling module and an information interaction optimization module. The latter includes a leading agent and an information interaction weight matrix calculation, an information interaction control module, and a multi-agent action decision generation network. By combining the weak connection interaction theory, the interaction relationship between agents is characterized in the form of a graph, and by calculating the agent connection strength distribution, the connection strength between agents is quantified. Combining the leading agent information, the information interaction weight matrix of each agent is calculated to improve the information interaction efficiency. Through the action decision generation network, the generated information interaction weight matrix is used to transmit information between agents, reduce information redundancy, and improve the model training speed. Corresponding multi-agent interaction graph networks are established in different scenarios and it is shown that using the weak connection theory can improve the quality of information interaction.
[0091] The simulation environment configuration of the example in this article is an INTER i7-9750H processor, with a main frequency of 2.6GHz, 16GB of memory, and an NVIDIA GeForce GTX 3080 graphics card. The software used for training is PyCharm Community 2020.3.2 version.
[0092] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0093] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A multi-agent deep reinforcement learning training method based on weak connection theory, characterized in that: The following steps are involved: Step 1: Establish a multi-agent interaction graph based on weak connection theory to represent the connection relationship between agents; Step 2: Calculate the connection strength distribution of multiple agents and the dominant agent, and establish a connection strength distribution matrix; Step 3 generates an information interaction weight matrix to optimize the information interaction between agents and improve the efficiency of multi-agent deep reinforcement learning training.
2. According to the multi-agent deep reinforcement learning training method based on weak connection theory according to claim 1, it is characterized in that: The establishment of agent connection relationship in step 1 is specifically as follows: Step 1-1 abstracts a single agent into a node, which contains the observation range, attack range, coordinates, and direction information of the corresponding agent; Step 1-2: The rules for establishing a multi-agent information interaction graph are as follows: Among them, the edge between agents: When two agents i and j are within the field of view of each other, there is an edge W between them. i,j ; Edges between subgroups: When there are unconnected subgraphs in the established interactive connection graph, the two agents i and j closest to the two subgraphs are most likely to interact, that is, the two groups are connected by the edge w i,j connect; Steps 1-3 define the strong and weak connections and dominant agents in the multi-agent information interaction graph as follows: Among them, weak connection between agents: if the connection strength between agents i and j is less than the threshold G, the connection between the two agents is a weak connection; strong connection between agents: if the connection strength between agents i and j is greater than the threshold G, the connection between the two agents is a strong connection; dominant agent: if agent i has the largest number of connections in the current interaction graph, then agent i is the dominant agent in this interaction network.
3. According to the multi-agent deep reinforcement learning training method based on weak connection theory in claim 1, it is characterized in that: The specific calculation formula established in step 2 is: Step 2-1 The formula for calculating the connection strength between agents is as follows: S strent =w i,j / (D i +D j +w i,j -2) Among them, W i,j is the number of edges on the shortest path between agents i and j, D i and D j are the number of connections between agents i and j respectively; Step 2-2 The calculation formula for the leading agent is as follows: Among them, W k,j is the number of edges on the shortest path between agents k and j, D k is the degree of agent k, g is the established multi-agent interaction graph; The specific definitions in the calculation formula of agent connection strength and dominant agent in step 2-3 are: Among them, the agent connection strength distribution matrix is established as a matrix of size n*n, where n is the total number of agents, and the value of the matrix element in the i-th row and j-th column is the connection strength between agents i and j. When the number of rows i and the number of columns j are equal, the corresponding matrix element value is set to 1. Assuming that the calculated dominant agent node is k, all elements in the k-th column of the connection strength matrix are set to 0.
4. According to the multi-agent deep reinforcement learning training method based on weak connection theory according to claim 1, it is characterized in that: The establishment information in step 3 is specifically as follows: Step 3-1 Determine the connection strength threshold and generate the information interaction weight matrix: Among them, a new matrix H of the same size as the connection strength matrix is created, and the connection strength threshold G is set to 0.
3. When the element value in the connection strength matrix is greater than 0.3, the element at the corresponding position of the weight matrix H is set to 0, otherwise, it is set to 1, and the information interaction weight matrix H is obtained; Step 3-2 calculates the required high-quality information for each agent based on the information interaction weight matrix. For each agent i, one of the strong and weak connectivity information sharing arrays is: α i ={α i,1 ,α i,2 ,...α i,i, ...α i,n,} The i-th row of the weight matrix, where the information received by agent i can be expressed as: Among them, O -i is the set of observed states of all agents except agent i, α -i is the action set of all agents except agent i; In step 3-3, each agent uses high-quality information as input and combines it with the MAPPO neural network to realize the training and updating of the multi-agent adversarial model. The initial network input size is n*c dimensions, where n is the number of agents and c is the dimension of observation information of a single agent. The output size is n dimensions, the activation function is the Relu function, and the optimizer selects the Adam optimizer.
5. According to the multi-agent deep reinforcement learning training method based on weak connection theory according to claim 4, it is characterized in that: According to the establishment of the agent in step 3-2, the information received is: The new state value function is derived as follows: The calculation formulas for the joint advantage function and action value function are derived as follows: The minimization loss function is as follows: Where b is the counterfactual baseline, and ω are the target network and critic network parameters respectively, α -i Represents the set of actions outside of agent i, and the critic network updates the network according to this loss function.
Citation Information
Cited By
Heavy-load unmanned helicopter cooperative hoisting method based on multi-agent learning
CN121411489A