A Multi-Agent Reinforcement Learning Method for Handling Communication Delays

By integrating real-time and delayed acquired information in multi-agent reinforcement learning, and using dual attention mechanisms and deep neural networks to process latency information, the impact of communication delay on multi-agent reinforcement learning is solved, and the stability and efficient decision-making of the system are achieved.

CN116595373BActive Publication Date: 2025-06-13SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310571611.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-21
Publication Date
2025-06-13
Estimated Expiration
2043-05-21

AI Technical Summary

Technical Problem

The existing multi-agent reinforcement learning methods cannot effectively deal with communication delay problems, resulting in the timeliness and accuracy of information exchange, which in turn affects the decision quality of the agent and the stability of the system.

Method used

By integrating the real-time and delayed information from other agents when training strategies by agents, the dual attention mechanism and deep neural network are used to represent the action value function and action strategy function, and gradually train and optimize the neural network to obtain the action strategy of multi-agent clusters, and the delay information is processed through communication features extraction networks and dual attention networks.

Benefits of technology

It realizes robust processing of communication delay, reduces the impact of delay on agent decisions, improves the stability of the system and the convergence performance of the agent, and can complete tasks efficiently under communication delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116595373B_ABST
    Figure CN116595373B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-agent reinforcement learning method for dealing with communication delay. The feature of this method is that it adopts a communication-based multi-agent reinforcement learning method, uses a communication cache pool to replace delay information, and uses a Transformer as a communication feature extraction network to extract information features. Then, combined with a hard attention mechanism and a soft attention mechanism, it further completes information integration. Finally, action decisions are made according to the integrated information. Compared with the prior art, the present invention better solves the problem of communication delay in the multi-agent communication process, realizes a communication protocol that is robust to delay, reduces the impact of communication delay on the agent's decision-making, enables efficient task completion under communication delay conditions, and enables multi-agent reinforcement learning to be applied to task scenarios with stronger real-world communication constraints.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-agent reinforcement learning, and particularly to a multi-agent reinforcement learning method in an environment with communication delay. Background Art

[0002] Multi-agent reinforcement learning originated from reinforcement learning and has received extensive attention in recent years due to its potential for autonomous decision-making in complex and dynamic environments. The origin of multi-agent reinforcement learning can be traced back to game theory. In the past decade, the progress of deep learning has promoted the development of multi-agent reinforcement learning algorithms. Multi-agent reinforcement learning focuses on the interaction and cooperation between multiple agents. Multiple agents learn to optimize their behaviors based on the feedback from the environment and the behaviors of other agents, rather than just based on the feedback from the environment as in single-agent reinforcement learning. However, in multi-agent reinforcement learning, the behavior of each agent will affect the future behaviors of the environment and other agents and the rewards to be obtained, resulting in the problem of environmental non-stationarity, which is also the main challenge faced by multi-agent reinforcement learning. Communication is an important part of multi-agent reinforcement learning because it enables agents to coordinate their actions and share information about the environment and other agents. Effective communication can improve performance and accelerate the convergence speed of learning algorithms. However, in the communication process of the real environment, communication delay may be caused due to reasons such as network congestion, data packet loss, or bandwidth limitation, which will affect the timeliness and accuracy of information exchange, resulting in sub-optimal decisions or even system instability. However, current multi-agent reinforcement learning methods mainly focus on improving communication efficiency and convergence performance, rarely considering real-world constraints such as data packet loss or communication delay, and cannot learn efficiently and achieve good convergence performance in the case of communication delay during the communication process. Summary of the Invention

[0003] Object of the Invention: In order to solve the defect that the existing communication-based multi-agent reinforcement learning cannot handle the communication delay problem, the present invention provides a multi-agent reinforcement learning method for handling communication delay.

[0004] In the present invention, agents are allowed to use additional information from other agents when training strategies, and the real-time acquired additional information and the delayed acquired additional information are integrated and processed to implement a communication protocol that is robust to delay, and reduce the impact of communication delay on agent decision-making. Based on the integrated additional information, the action value function and the action policy function are represented by using a deep neural network, and the neural network is gradually trained and optimized to obtain the action policy of the multi-agent cluster. The training process of the policy neural network of the multi-agent cluster includes the following steps:

[0005] S1. Initialize the neural network optimizer gradient and the experience replay pool

[0006] S2. Construct the task environment and randomly initialize the positions of all agents and their corresponding communication buffer pools

[0007] S3. The agent obtains the local observation state based on its own local observation of the environment;

[0008] S4. The agent generates communication content containing hidden features through the communication content generation network, and then sends the communication content to other teammates in the form of broadcasting;

[0009] S5. The agent determines the real-time communication teammate set and the delayed communication teammate set according to the distance between each other;

[0010] S6. Determine the real-time communication information and the delayed communication information based on the real-time communication teammate set and the delayed communication teammate set;

[0011] S7. Based on the randomly determined delay hops, the delayed information is received by the designated agent at a future time after multi-hop delay;

[0012] S8. The agent replaces the delayed information at the current moment with the information in the communication buffer pool;

[0013] S9. The agent updates the communication buffer pool according to the currently received information and the replacement information of the delayed information;

[0014] S10. The agent receives the delayed information from the past, ignores the outdated information and processes the valid delayed information;

[0015] S11. The agent integrates all communication information through the communication feature extraction network;

[0016] S12. The agent calculates the communication information weight for the integrated information through the dual attention network;

[0017] S13. The agent calculates the action policy and the action policy value function according to the integrated information and the communication information weight;

[0018] S14. The agent executes actions to interact with the environment and obtains the corresponding rewards and the local observation state of the next state;

[0019] S15. Save the experience to the experience replay pool

[0020] S16. Repeat steps S3 - S15 until the maximum number of single training rounds is reached or all agents complete the task.

[0021] S17. Use the experience replay pool Calculate the gradient of the loss function for the training samples and optimize the entire network using a neural network optimizer.

[0022] S18. Repeat steps S2 - S17 to continuously optimize the network until the network converges or reaches the maximum number of training epochs.

[0023] Furthermore, in the aforementioned S1, the initialization of the neural network optimizer gradient is expressed as: dθ = 0, dφ = 0, where θ and φ are the parameters to be learned for the action policy network and the value network, respectively.

[0024] Furthermore, in the aforementioned S3, at time t, the environment is in state S t , and the local observation of the agent is where represents the set of agents composed of N agents.

[0025] Furthermore, in the aforementioned S4, a Multi - Layer Perceptron (MLP) is used as the communication content generation network, and the communication content containing hidden features generated by it is

[0026] Furthermore, in the aforementioned S5, by comparing the actual distance between agents with the predefined real - time communication range r c , to determine whether there is a delay in communication between agents, thereby determining the real - time communication teammate set and the delayed communication teammate set.

[0027] Furthermore, in the aforementioned S5, for any agent , the real - time communication teammate set C i and the delayed communication teammate set D i can be respectively expressed as:

[0028]

[0029] where is used to represent the distance between agent i and agent j.

[0030] Furthermore, in the aforementioned S6, any agent receives real - time information from real - time communication teammate n at time t, and at the same time lacks the delayed information from delayed communication teammate n′. The two kinds of information can be expressed as:

[0031]

[0032] Furthermore, in the aforementioned S7, the hop count of the delayed information missing by any agent at time t is from a [1, T delayDetermined by a random number within the range:

[0033]

[0034] where T delay represents the maximum communication delay existing in the communication process.

[0035] Furthermore, in the aforementioned S8, any agent at time t uses the communication information of the corresponding agent saved in the communication cache pool to replace the missing delay information

[0036] Furthermore, in the aforementioned S9, any agent at time t saves the received real-time information and the alternative information of the delay information as new communication cache information:

[0037]

[0038] Furthermore, in the aforementioned S10, any agent at time t receives the delay information sent by teammate n' at the past time t' and ignores the outdated information whose delay hop count is greater than the delay information reception time window T r , and its processing process can be expressed as:

[0039]

[0040] Furthermore, in the aforementioned S11, the Transformer encoding network is used as the communication feature extraction network to extract features and integrate information for the information received by any agent at time t, and the integrated information can be expressed as:

[0041]

[0042] Furthermore, in the aforementioned S12, a hard attention network based on a bidirectional long short-term memory network (Bi-directional Long-ShortTerm Memory, Bi-LSTM) is used to output the hard attention weight (0 or 1) to screen the integrated information worthy of attention, and its specific process can be expressed as:

[0043]

[0044] where f(·) represents the fully connected layer network, and Gumbel(·) represents the Gumbel-Softmax function.

[0045] Further, in the aforementioned S12, a soft attention network based on a soft attention mechanism is used to calculate soft attention weights The specific process can be expressed as follows:

[0046]

[0047] where Softmax(·) represents the Softmax function, and W q , W k represent the query vector transformation matrix and the key vector transformation matrix to be learned, respectively.

[0048] Further, in the aforementioned S13, a fully connected network is used as the action output network and the value network, and action values and value functions are output according to the integrated information and communication information weights. The specific process can be expressed as follows:

[0049]

[0050] where π(·) and V(·) represent the action output network and the value network, respectively.

[0051] Further, in the aforementioned S14, any agent executes an action at time t to interact with the environment and obtain corresponding rewards and the local observation state of the next state

[0052] Further, in the aforementioned S15, the experience of any agent at time t is saved to the experience replay pool

[0053] Further, in the aforementioned S17, the loss function gradient is as follows:

[0054]

[0055] where is the loss function, θ and φ are the parameters to be optimized of the action output network and the value network respectively, represents the probability distribution function of selecting action at the given local observation state τ is the time length of a single training sample, and β is a hyperparameter used to balance the roles of the two parts of the loss function. comes from the samples stored in the experience replay pool

[0056] ​Compared with the prior art, the advantages of the present invention are as follows:

[0057] 1. The method of the present invention adopts a dual attention mechanism, so it can be applied to multi-agent task scenarios with different numbers of agents and different difficulties. Even in highly complex task scenarios involving many agents, this method can reduce the attention and calculation of unnecessary information through the application of the attention mechanism, show a faster convergence speed during training, and achieve better convergence performance and scalability at the same time.

[0058] 2. The delay information integration network specially designed for communication delay in the method of the present invention enables the system to effectively compensate and integrate delay information with dynamically changing quantities, thereby reducing the interference of delay information on the system and showing robustness and adaptability to communication delay. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 is the network framework diagram of the method of the present invention;

[0060] Figure 2 is the flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0061] The following introduces a preferred embodiment of the present invention with reference to the specification to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.

[0062] Embodiment: The embodiment of the present invention provides a multi-agent reinforcement learning method for processing communication delay. The embodiment of the present invention applies the method to a grid environment for multi-agent search, which is composed of several grids, including several agents with limited field of view and a target. The agents make action decisions based on their own local observations and information from teammates, and the task is considered completed when all agents reach the target position. Define the observation space of the agent in the search scenario as the grid coordinates where the agent is located and all grid information within the agent's field of view. Define the agent action space as up, down, left, right, stop. Define the reward that the agent can obtain as determined by its own position and the target position, etc. The specific steps are as follows:

[0063] S1. Use the RMSprop optimizer as the neural network optimizer, initialize its gradient to 0, and initialize the experience replay pool at the same time

[0064] S2. Build a search task grid environment, randomly initialize the positions of all agents and the target, and the communication cache pool of all agents

[0065] S3. All agents obtain local observation O according to the observation space rules defined within the task environment t , which includes the grid coordinates where the agent itself is located and all grid information within the agent's field of view;

[0066] S4. Any agent At time t, inputs the local observation into the communication content generation network composed of MLP networks to generate preliminary communication content and broadcasts it to all teammates;

[0067] S5. Any agent At time t, determines whether it belongs to the real-time communication teammate set C or the delayed communication teammate set D c based on whether the distance i from it to other agent j is less than the real-time communication range r i .

[0068] S6. Any agent At time t, receives real-time information from real-time communication teammate n while missing the delayed information from delayed communication teammate n′

[0069] S7. The delayed information missing by any agent at time t will be received by agent i after multi-hop delay with a delay hop count of ;

[0070] S8. Any agent At time t, uses the communication information of the delayed teammate saved in the communication cache pool to replace the missing delayed information

[0071] S9. Any agent At time t, uses the received real-time information and the replacement information of the delayed information to save as new communication cache information

[0072] S10. The agent receives delayed information from the past, ignores the outdated information among them and processes the valid delayed information;

[0073] S11. Any agent At time t, uses the Transformer encoding network to extract and integrate features from all the received information to obtain integrated information

[0074] S12. Any agent At time t, use a dual attention network for the integrated information Calculate the communication information weights

[0075] S13. Any agent At time t, according to the integrated information Calculate the communication information weights Calculate the action value And the value function

[0076] S14. Any agent At time t, according to the action value Determine the specific actions, i.e., up, down, left, right, stop, and interact with the environment. The environment feedbacks a reward according to the result of this interaction Meanwhile, the environment enters the next state S t+1 , and agent i also obtains the next local observation state accordingly

[0077] S15. Save the experience of any agent At time t To the experience replay pool

[0078] S16. Repeat steps S3 - S15 until reaching the maximum number of times for a single training episode or all agents reach the target position.

[0079] S17. Use the training samples in the experience replay pool To calculate the gradient of the loss function and use the RMSprop optimizer to optimize the entire network.

[0080] S18. Repeat steps S2 - S17 to continuously optimize the network until the network converges or reaches the maximum number of training rounds.

[0081] In the grid environment of the multi - agent search task, this method enables the multi - agent reinforcement learning model to be robust to communication delays, can reduce the wrong decisions caused by communication delays, and realizes efficient task completion under communication delays.

Claims

1. A multi-agent reinforcement learning method for handling communication delays, the learning method comprises the following steps: S1. Initialize the gradients of the neural network optimizer and the experience replay pool S2. Construct the task environment and randomly initialize the positions of all agents and their corresponding communication cache pools S3. The agent obtains a local observation state based on its own local observation of the environment; S4. The agent generates communication content containing hidden features through a communication content generation network, and then sends the communication content to other teammates in a broadcast form; S5. The agent determines a set of real-time communication teammates and a set of delayed communication teammates according to the distance between each other; S6. Determine real-time communication information and delayed communication information based on the set of real-time communication teammates and the set of delayed communication teammates; S7. Based on a randomly determined number of delay hops, the delayed information is received by the designated agent at a future time after multiple-hop delays; S8. The agent replaces the delayed information at the current moment with the information in the communication cache pool; S9. The agent updates the communication cache pool according to the information received at the current moment and the replacement information of the delayed information; S10. The agent receives delayed information from the past, ignores the outdated information therein and processes the valid delayed information; S11. The agent integrates all communication information through a communication feature extraction network; S12. The agent calculates the communication information weight for the integrated information through a dual attention network; S13. The agent calculates an action policy and an action policy value function according to the integrated information and the communication information weight; S14. The agent executes actions to interact with the environment and obtains corresponding rewards and the local observation state of the next state; S15. Save the experience to the experience replay pool S16. Repeat steps S3 - S15 until the maximum number of single training rounds is reached or all agents complete the task; S17. Use the experience replay pool to calculate the gradient of the loss function using the training samples in it, and use the neural network optimizer to optimize the entire network; S18. Repeat steps S2 - S17 to continuously optimize the network until the network converges or the maximum number of training rounds is reached.

2. A multi-agent reinforcement learning method for handling communication delays according to claim 1, wherein, in S1, the gradient initialization of the neural network optimizer is expressed as: dθ = 0, dφ = 0, where θ and φ are the parameters to be learned of the action policy network and the value network respectively; In S3, at time t, the environment is in state S t , the local observation of the agent is where represents the set of agents composed of N agents.

3. A multi-agent reinforcement learning method for handling communication delays according to claim 2, wherein, In S4, a Multi-Layer Perceptron (MLP) is used as the communication content generation network, and the communication content containing hidden features generated by it is In S5, by comparing the actual distance between agents with the predefined real-time communication range r c , to determine whether there is a delay in communication between agents, so as to determine the set of real-time communication teammates and the set of delayed communication teammates.

4. A multi-agent reinforcement learning method for handling communication delays according to claim 3, wherein, In S5, any agent 's real-time communication teammate set C i and the delayed communication teammate set D i can be respectively expressed as: Among them is used to represent the distance between agent i and agent j; In S6, any agent receives real-time information from real-time communication teammate n at time t, while missing the delayed information from delayed communication teammate n ′ The two kinds of information can be expressed as:

5. A multi-agent reinforcement learning method for handling communication delays according to claim 4, wherein, In the S7, any agent The delay information missing at time t The delay hop count of Is determined by a random number within the range of [1, T drlay : Among which T delay represents the maximum communication delay existing in the communication process; Among the S8, any agent at time t adopts the communication information of the corresponding agent saved in the communication cache pool to replace the missing latency information 6. A multi-agent reinforcement learning method for handling communication delays according to claim 5, wherein, In the above S9, any agent at time t saves the received real-time information and the substitute information of the delayed information as new communication cache information: In S10, any agent receives at time t the delayed information ′ sent by teammate n ′ at the past time t and ignores the outdated information with a delay hop count greater than the delay information reception time window T r The processing process can be expressed as:

7. A multi-agent reinforcement learning method for handling communication delays according to claim 6, wherein, In S11, a Transformer encoding network is used as a communication feature extraction network to extract features and integrate information from the information received by any agent at time t. The integrated information can be expressed as: In the S12, a hard attention network based on a bidirectional long short-term memory network (Bi-LSTM) is used to output hard attention weights Or 1) to filter out the integrated information worthy of attention, and its specific process can be expressed as: where f(·) represents a fully connected layer network and Gumbel(·) represents a Gumbel-Softmax function.

8. A multi-agent reinforcement learning method for handling communication delays according to claim 7, wherein, In the above S12, a soft attention network based on a soft attention mechanism is used to calculate soft attention weights The specific process can be expressed as follows: where Softmax(·) represents the Softmax function, and W q , W k represent the query vector transformation matrix and the key vector transformation matrix to be learned, respectively; In S13, a fully connected network is used as the action output network and the value network, and action values are output according to the integrated information and the communication information weight and value function The specific process can be expressed as: where π(·) and V(·) represent an action output network and a value network respectively.

9. A multi-agent reinforcement learning method for handling communication delays according to claim 8, wherein, In S14, any agent executes an action at time t interacts with the environment and obtains the corresponding reward and the local observation state of the next state In S15, any agent the experience at time t is saved to the experience replay pool 10. A multi-agent reinforcement learning method for handling communication delays according to claim 9, It is characterized in that In the S17, the gradient of the loss function is as follows: where is the loss function, and θ and φ are the parameters to be optimized for the action output network and the value network respectively, represents the probability distribution function of selecting action when given the local observation state , τ is the time length of a training sample, and β is a hyperparameter used to balance the roles of the two parts of the loss function, which comes from the samples stored in the experience replay pool .

Citation Information

Patent Citations

  • Multi-agent reinforcement learning method and system based on dynamic hierarchical communication network

    CN113919485A

  • Multi-agent cooperative communication strategy training system and method based on teammate perception

    CN114757092A