Multi-agent reinforcement learning method and system based on region division and storage medium

By employing a region-based approach and a leader agent coordination method, the problem of agents lacking real-time policy awareness in multi-agent systems is addressed, improving communication efficiency and system scalability, and enabling adaptation to large-scale environmental changes.

CN120996130APending Publication Date: 2025-11-21SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510881196.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing multi-agent reinforcement learning systems suffer from the problem of agents lacking real-time policy awareness capabilities from other agents in large-scale environments, resulting in low communication and coordination efficiency.

Method used

The agent is divided into several regions by a region partitioning method. A leader agent is selected in each region. The leader agent receives information from other regions and coordinates the actions of the follower agents. The MADDPG algorithm is used to update the policy network parameters. The region partitioning is optimized by combining K-means, DBSCAN and density peak clustering algorithms. Multilayer perceptron and multi-head attention mechanism are used to process information. The communication quality is monitored in real time and a smooth transition strategy is implemented.

Benefits of technology

It effectively maintains the ability to perceive the real-time policies of other intelligent agents, improves communication efficiency and system scalability, adapts to changes in the number of intelligent agents and task scale, and reduces system complexity and communication overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996130A_ABST
    Figure CN120996130A_ABST
Patent Text Reader

Abstract

The invention provides a multi-agent reinforcement learning method and system based on region division and a storage medium, and belongs to the technical field of reinforcement learning, and the method comprises the steps: initializing an analogue simulation environment, and initializing state data, task targets and global reward functions of all agents; dividing all agents into a plurality of areas based on the state data, and selecting one leader agent in each area; receiving information of leader agents in other areas and observation data of follower agents in the areas through the leader agent in each area, and obtaining an action instruction of each follower agent in the area according to the information and the observation data; updating the strategy network parameters of the leader agent; and iteratively training until a preset ending condition is reached, and completing multi-agent training. According to the invention, the leader agent receives the information of the leader agents in other areas to make a decision, and the perception capability of real-time strategies of other agents can be effectively maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of reinforcement learning technology, and in particular to a multi-agent reinforcement learning method, system, and storage medium based on region partitioning. Background Technology

[0002] Multi-agent reinforcement learning (MARL) has shown great potential in solving decision-making problems in multi-agent environments, with applications spanning numerous fields such as robot control, autonomous driving, resource management, and games. However, applying MARL systems to large-scale environments faces many challenges, particularly in the communication, coordination, and scalability issues among agents. A core challenge lies in the curse of dimensionality, where the state space and action space grow exponentially with the number of agents, severely limiting the effectiveness of traditional reinforcement learning algorithms in large-scale scenarios.

[0003] Currently, the Centralized Training with Decentralized Execution (CTDE) architecture is widely adopted. It allows agents to share information during the training phase but act independently during the execution phase. This results in agents lacking the ability to perceive the real-time policies of other agents in environments where agent policies are interdependent.

[0004] Therefore, existing technologies have shortcomings and need to be improved and developed. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a multi-agent reinforcement learning method, system and storage medium based on region division, which addresses the above-mentioned deficiencies of the prior art and aims to solve the problem that agents in the prior art lack the ability to perceive the real-time strategies of other agents.

[0006] The technical solution adopted by this invention to solve the technical problem is as follows:

[0007] In a first aspect, embodiments of the present invention provide a multi-agent reinforcement learning method based on region partitioning, the method comprising:

[0008] Initialize the simulation environment, and initialize the state data, task objectives, and global reward function of all agents;

[0009] Based on state data, all agents are divided into several regions, and a leader agent is selected for each region.

[0010] The leader agent in each region receives information from the leader agents in other regions and observation data from the follower agents in the region, and obtains action instructions for each follower agent in the region based on the information and the observation data.

[0011] After each follower agent in the region executes the action command, the environment provides feedback on the global reward value and the new environment state. The state transition samples are stored in the shared experience pool. When the sampling conditions are met, batch data is extracted from the experience pool, and the MADDPG algorithm is used to update the policy network parameters of the leader agent.

[0012] The training is carried out iteratively until the preset termination condition is met, thus completing the training of the multi-agent system.

[0013] In one implementation, the state data includes location data; based on the state data, all agents are divided into several regions, including:

[0014] Obtain the preset communication radius for each agent;

[0015] Determine if the number of clusters can be obtained;

[0016] If the number of clusters is available, K-means clustering is performed based on the number of clusters to obtain several initial clusters. DBSCAN clustering optimization is then performed on all the initial clusters to obtain several regions.

[0017] If the number of clusters cannot be obtained, the Euclidean distance is calculated based on the position data of any two agents. A distance matrix is ​​constructed based on all the Euclidean distances and the communication radius. The DBSCAN clustering algorithm is then used to cluster the data based on the distance matrix to obtain several regions.

[0018] In one implementation, K-means clustering is performed based on the number of clusters to obtain several initial clusters, including:

[0019] Randomly select a number of agents from each cluster as the initial cluster centers;

[0020] Perform distance correction calculations for each agent:

[0021] Calculate the Euclidean distance from each agent to the center of each initial cluster based on the location data of each agent;

[0022] If the original Euclidean distance is less than or equal to the communication radius, then the original Euclidean distance is used as the output distance.

[0023] If the original Euclidean distance is greater than the communication radius, the original Euclidean distance is multiplied by the preset penalty factor to obtain the output distance.

[0024] Assign each agent to the cluster with the smallest output distance, recalculate the geometric center position of each cluster, and set the agent closest to the geometric center position as the new cluster center;

[0025] The process of iteratively performing distance correction calculations and setting new cluster centers continues until the iteration termination condition is met, resulting in several initial clusters.

[0026] In one implementation, a leader agent is selected for each region, including:

[0027] The agent closest to the geometric center of the region is selected as the leader agent.

[0028] In one implementation, federated learning can also be used to enhance communication and cooperation between regions. The leader agents in each region can share model parameters instead of raw data, thereby protecting the agents' privacy.

[0029] In one implementation, outputting reinforcement learning based on the information and the observation data includes:

[0030] Based on the information received from the leader agents of other regions and the state data of the leader agents of this region, calculate the mutual information between regions;

[0031] By using a multilayer perceptron within the leader's intelligent body to process all mutual information, several inter-regional cooperative feature vectors are obtained.

[0032] The communication trust score at the current moment is obtained by utilizing the prior network within the leader's intelligent body based on the observed data and the communication trust score at the previous moment.

[0033] The observed data and the communication trust score are processed by the encoder in the leader agent to obtain feature vectors of several follower agents.

[0034] The feature vectors of all follower agents are processed using the multi-head attention mechanism within the leader agent to obtain the second fusion feature matrix;

[0035] By using the policy module within the leader agent to process all the inter-regional cooperation feature vectors and the second fused feature matrix, action instructions for each follower agent within the region are obtained.

[0036] In one implementation, before receiving information from leader agents in other regions, the method further includes:

[0037] The information to be sent in the leader's intelligent body at the sending end is compressed according to different levels of abstraction to obtain multiple candidate transmission information;

[0038] Select the candidate transmission information with the highest level of abstraction and send it.

[0039] In one embodiment, the method further includes:

[0040] Real-time monitoring of the position change rate of all agents;

[0041] Real-time monitoring of the speed of all follower agents and the agent density in each region;

[0042] When the speed of any follower agent exceeds a preset speed threshold, the communication quality and connection stability between the follower agent and the leader agent in its region are evaluated. If there is low communication quality and / or unstable connection, the region to which the follower agent belongs is updated, and a smooth transition strategy is used for communication transition.

[0043] When the agent density in any region exceeds a preset density threshold, the communication quality and connection stability between all follower agents in that region and the leader of that region are evaluated. If there are cases of low communication quality and / or unstable connection, the region to which the follower agents in that region belong is updated, and a smooth transition strategy is used to perform a communication transition.

[0044] In one implementation, a smooth transition strategy is used for communication transition, including:

[0045] Obtain the preset transition period duration and divide the transition period duration into n time steps;

[0046] Obtain a preset first communication frequency and a second communication frequency. The first communication frequency is the communication frequency with the original leader agent within the first time step, and the second communication frequency is the communication frequency with the newly assigned leader agent within the first time step. The sum of the first communication frequency and the second communication frequency is 1.

[0047] Obtain a preset ratio, and starting from the second time step, iteratively increase the second communication frequency and decrease the first communication frequency based on the preset ratio until the nth time step, at which point the communication transition is completed.

[0048] Secondly, embodiments of the present invention also provide a multi-agent reinforcement learning system based on region partitioning, the system comprising:

[0049] The initialization module is used to initialize the simulation environment and initialize the state data, task objectives, and global reward function of all agents.

[0050] The selection module is used to divide all agents into several regions based on state data, and select a leader agent for each region.

[0051] The decision-making module is used to receive information from leader agents in other regions and observation data from follower agents in each region through the leader agent in each region, and to obtain action instructions for each follower agent in the region based on the information and the observation data.

[0052] The parameter update module is used to update the global reward value and new environmental state after each follower agent in the region executes the action command, store the state transition samples in the shared experience pool, and extract batch data from the experience pool when the sampling conditions are met, and use the MADDPG algorithm to update the decision network parameters of the leader agent.

[0053] The iterative module is used to iteratively train the multi-agent system until the preset termination condition is met.

[0054] Thirdly, embodiments of the present invention also provide a computer-readable storage medium storing a region-partition-based multi-agent reinforcement learning program, which can be executed to implement the steps of the region-partition-based multi-agent reinforcement learning method as described above.

[0055] The beneficial effects of this invention are as follows: This invention initializes the simulation environment and the state data, task objectives, and global reward function of all agents; based on the state data, all agents are divided into several regions, with one leader agent selected in each region; the leader agent in each region receives information from leader agents in other regions and observation data from follower agents within the region, and obtains action instructions for each follower agent based on the information and observation data; the policy network parameters of the leader agent are updated; training is iteratively performed until a preset termination condition is met, completing the training of multiple agents. This invention, by having the leader agent receive information from leader agents in other regions to make decisions, can effectively maintain the ability to perceive the real-time policies of other agents. Attached Figure Description

[0056] Figure 1 This is a flowchart of a preferred embodiment of the multi-agent reinforcement learning method based on region partitioning in this invention.

[0057] Figure 2 This is a schematic diagram of the region division process in this invention.

[0058] Figure 3 This is a schematic diagram of the leader intelligent agent data processing flow in this invention.

[0059] Figure 4 This is a schematic diagram illustrating the application of the present invention in a drone scenario.

[0060] Figure 5 This is a schematic diagram of a preferred embodiment of the multi-agent reinforcement learning system based on region partitioning in this invention. Detailed Implementation

[0061] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0062] Multi-agent reinforcement learning (MARL) has shown great potential in solving decision-making problems in multi-agent environments, with applications spanning numerous fields such as robot control, autonomous driving, resource management, and games. However, applying MARL systems to large-scale environments faces many challenges, particularly in the communication, coordination, and scalability issues among agents. A core challenge lies in the curse of dimensionality, where the state space and action space grow exponentially with the number of agents, severely limiting the effectiveness of traditional reinforcement learning algorithms in large-scale scenarios.

[0063] Currently, the Centralized Training with Decentralized Execution (CTDE) architecture is widely adopted. It allows agents to share information during the training phase and act independently during the execution phase. This results in agents lacking the ability to perceive the real-time policies of other agents in environments where agent policies are interdependent.

[0064] To address the aforementioned shortcomings of existing technologies, this invention provides a multi-agent reinforcement learning method, system, and storage medium based on region partitioning. The method includes: initializing a simulation environment and initializing the state data, task objectives, and global reward function of all agents; dividing all agents into several regions based on the state data, selecting a leader agent in each region; receiving information from leader agents in other regions and observation data from follower agents within each region, and obtaining action instructions for each follower agent within the region based on the information and observation data; updating the policy network parameters of the leader agent; and iteratively training until a preset termination condition is met, thus completing the multi-agent training. This invention enables the leader agent to make decisions by receiving information from leader agents in other regions, effectively maintaining the ability to perceive the real-time policies of other agents.

[0065] Please see Figure 1The multi-agent reinforcement learning method based on region partitioning described in this embodiment of the invention includes the following steps:

[0066] Step S100: Initialize the simulation environment and initialize the state data, task objectives, and global reward function of all agents.

[0067] Specifically, initializing the simulation environment includes: setting hardware configuration parameters for each agent, such as communication radius and maximum transmission power; setting communication parameters for each agent, including communication frequency and data packet size; and setting model hyperparameters and weight parameters for each agent. State data includes the agent's position data, movement speed, task urgency, signal strength, and simulated data acquired by sensors. Simulated data acquired by sensors may include simulated laser point cloud data, camera image data, IMU inertial measurement data, radar data, etc.

[0068] Please see Figure 1 The multi-agent reinforcement learning method based on region partitioning described in this embodiment of the invention further includes the following steps:

[0069] Step S200: Divide all agents into several regions based on state data, and select a leader agent for each region.

[0070] Specifically, before training, regions are divided based on the initial state data, and a leader agent is selected for each region. Subsequently, regions can exchange information through the leader agents, which can effectively improve communication efficiency.

[0071] In one implementation, the state data includes location data; based on the state data, all agents are divided into several regions, including:

[0072] Obtain the preset communication radius for each agent;

[0073] Determine if the number of clusters can be obtained;

[0074] If the number of clusters is available, K-means clustering is performed based on the number of clusters to obtain several initial clusters. DBSCAN clustering optimization is then performed on all the initial clusters to obtain several regions.

[0075] If the number of clusters cannot be obtained, the Euclidean distance is calculated based on the position data of any two agents. A distance matrix is ​​constructed based on all the Euclidean distances and the communication radius. The DBSCAN clustering algorithm is then used to cluster the data based on the distance matrix to obtain several regions.

[0076] Specifically, this invention supports two different methods for dividing regions, such as Figure 2As shown. For scenarios where the number of clusters is predetermined, the K-means algorithm is first used for preliminary clustering. By iteratively optimizing the cluster centers, each agent is assigned to the nearest cluster, thus forming multiple initial clusters. Based on the preliminary K-means clustering results, the DBSCAN algorithm is further used for optimization. For scenarios where the number of clusters is not predetermined, the Euclidean distance is directly calculated based on the initial position data of any two agents. A distance matrix is ​​constructed based on all Euclidean distances and the communication radius. The DBSCAN clustering algorithm is then used to cluster based on this distance matrix, resulting in several regions.

[0077] In one implementation, K-means clustering is performed based on the number of clusters to obtain several initial clusters, including:

[0078] Randomly select a number of agents from each cluster as the initial cluster centers;

[0079] Perform distance correction calculations for each agent:

[0080] Calculate the Euclidean distance from each agent to the center of each initial cluster based on the location data of each agent;

[0081] If the original Euclidean distance is less than or equal to the communication radius, then the original Euclidean distance is used as the output distance.

[0082] If the original Euclidean distance is greater than the communication radius, the original Euclidean distance is multiplied by the preset penalty factor to obtain the output distance.

[0083] Assign each agent to the cluster with the smallest output distance, recalculate the geometric center position of each cluster, and set the agent closest to the geometric center position as the new cluster center;

[0084] The process of iteratively performing distance correction calculations and setting new cluster centers continues until the iteration termination condition is met, resulting in several initial clusters.

[0085] Specifically, during the clustering process, the distance calculation from an agent to the cluster center is based not only on location information but also on a comprehensive consideration of the communication radius. This ensures communication connectivity and transmission reliability between agents within a region and the cluster center (i.e., potential leaders). When the distance between agents exceeds the communication radius, the distance is amplified, thereby reducing the probability that agents that are not reachable by communication will be assigned to the same cluster, making the region partitioning more consistent with actual communication constraints.

[0086] In one implementation, DBSCAN clustering optimization is performed on all the initial clusters to obtain several regions, including:

[0087] Based on the communication radius of each agent, set the neighborhood radius parameter for each agent;

[0088] The neighborhood range is determined based on the neighborhood radius parameter of each agent. If the neighborhood range of an agent contains no less than a preset number of agents, the agent is marked as a core point. If an agent is not marked as a core point, but its neighborhood range is covered by the core point, it is marked as a boundary point. Otherwise, it is marked as an isolated point.

[0089] Perform reachability checks on boundary points and isolated points:

[0090] When there is no original cluster core point within the neighborhood of the agent, its original cluster affiliation is terminated.

[0091] If there are multiple cluster core points in its neighborhood, it will be assigned to the new cluster that covers the most core points.

[0092] After completing the reachability check, several regions are obtained.

[0093] Specifically, the DBSCAN algorithm of this invention can effectively identify and process data distributions with different density characteristics, and has better adaptability to irregularly shaped or unevenly dense agent distributions. By setting appropriate spatial radius parameters (related to agent communication radius), the initially divided agents within clusters are re-examined and adjusted. For agents located at the boundary in K-means clustering, whose communication with the cluster center may be limited, and agents with overlapping communication areas with other clusters, the DBSCAN algorithm reassigns them to more suitable clusters based on the neighborhood relationship determined by the communication radius. This step aims to refine the region division, enhance the communication connectivity between agents within the region, and avoid situations where the leader agent cannot effectively cover all followers due to unreasonable region division. Through the above fusion clustering process, agents within each cluster are relatively concentrated and communication is more convenient.

[0094] In one implementation, the entire agent is divided into several regions based on state data, and the method further includes:

[0095] If the number of clusters is available, K-means clustering is performed based on the number of clusters to obtain several initial clusters;

[0096] DBSCAN clustering optimization is performed on all the initial clusters to obtain several intermediate clusters;

[0097] The density peak clustering algorithm was used to optimize all the intermediate clusters to obtain several regions.

[0098] Specifically, this invention can also support the introduction of the Density Peak Clustering (DPC) algorithm to further optimize intermediate clusters. The DPC algorithm can effectively identify cluster centers with different densities, and may be more suitable for scenarios where the distribution of agents has significant density differences.

[0099] In addition to the aforementioned methods for region partitioning, spectral clustering or hierarchical clustering algorithms can also be used to divide regions. Spectral clustering algorithms construct a similarity matrix and use spectral decomposition techniques from graph theory to identify cluster structures in the data, which may be more effective for handling complex-shaped agent distributions; hierarchical clustering algorithms, on the other hand, gradually merge or split clusters by constructing nested cluster hierarchies, and can adapt to groups of agents of different densities and sizes.

[0100] In one implementation, a leader agent is selected for each region, including:

[0101] The agent closest to the geometric center of the region is selected as the leader agent.

[0102] Specifically, since the impact of communication radius has been fully considered in the clustering process, the leader agent can communicate directly with most follower agents within its communication radius, ensuring effective information transmission and coordinated command.

[0103] Please see Figure 1 The multi-agent reinforcement learning method based on region partitioning described in this embodiment of the invention further includes the following steps:

[0104] Step S300: The leader agent in each region receives information from the leader agents in other regions and observation data from the follower agents in the region, and obtains action instructions for each follower agent in the region based on the information and the observation data.

[0105] Specifically, each region is managed by a leader agent, which is responsible for collecting information from follower agents and coordinating actions within the region. Leaders in different regions also communicate with each other to coordinate the actions of the entire system. This hierarchical structure reduces unnecessary communication and improves the efficiency of reinforcement learning.

[0106] In one implementation, obtaining action instructions for each follower agent within the region based on the information and the observation data includes:

[0107] Based on the information received from the leader agents of other regions and the state data of the leader agents of this region, calculate the mutual information between regions;

[0108] By using a multilayer perceptron within the leader's intelligent body to process all mutual information, several inter-regional cooperative feature vectors are obtained.

[0109] The communication trust score at the current moment is obtained by utilizing the prior network within the leader's intelligent body based on the observed data and the communication trust score at the previous moment.

[0110] The observed data and the communication trust score are processed by the encoder in the leader agent to obtain feature vectors of several follower agents.

[0111] The feature vectors of all follower agents are processed using the multi-head attention mechanism within the leader agent to obtain the second fusion feature matrix;

[0112] By using the policy module within the leader agent to process all the inter-regional cooperation feature vectors and the second fused feature matrix, action instructions for each follower agent within the region are obtained.

[0113] Specifically, the information to be sent within the leader agent at the sending end is compressed according to different levels of abstraction, resulting in multiple candidate transmission messages. The candidate message with the highest level of abstraction is selected for transmission. Detailed information is transmitted only when further collaboration or to solve specific problems is required. This hierarchical abstraction and compression of communication information reduces the amount of data in inter-regional communication and accelerates information transmission. Existing communication methods such as DIAL, CommNet, and BiCNet, while simplifying the communication process to some extent, suffer from limitations in discrete message transmission, information loss, increased system complexity, and low communication efficiency, making them unsuitable for large-scale multi-agent systems. This invention, by compressing information, effectively improves communication efficiency, adaptively adjusts communication content, transmits on demand, ensures information validity, and reduces system complexity.

[0114] Each leader agent determines communication priority based on the magnitude of mutual information. The formula for calculating mutual information is as follows:

[0115]

[0116] In the formula, S i and S j S represents the states of region i and region j respectively. i and S j It is a matrix containing environmental state information, including the agent's position, velocity, etc.; p(S i ,S j ) is a joint probability distribution, which describes the situation where regions i and j are simultaneously located in (S). i ,S j The probability of state p(S) reflects the correlation between the states of the two regions.i ) indicates that region i is in state S i The probability, p(S) j ) indicates that region j is in state S j The probability of mutual information is calculated using this formula. The mutual information between region i and region j can be used to quantify the correlation between them, thus determining the priority queue for inter-region communication. When the mutual information is greater than a threshold, the communication request is placed in the high-priority queue; otherwise, it is placed in the low-priority queue. The calculation of mutual information comprehensively considers factors such as the state correlation between the current region and its neighboring regions, and task coordination requirements.

[0117] Incorporating mutual information into training allows leader agents to prioritize tasks and effectively apply this element in real-world missions. For example, in a multi-agent rescue mission, when a leader agent in one area detects that a fire's spread may affect neighboring areas, it calculates mutual information with neighboring leader agents and prioritizes communication requests based on the magnitude of this mutual information. High-priority requests receive priority access to communication resources, ensuring timely delivery of critical information. Furthermore, to guarantee real-time communication, a dynamic time-slot allocation mechanism is employed for inter-area communication. The leader agent dynamically adjusts the length and frequency of communication time slots with other leader agents based on the communication priority queue and current communication load.

[0118] By utilizing a multilayer perceptron within the leader's intelligent body to process all mutual information, several inter-regional cooperative feature vectors are obtained. The multilayer perceptron, through its internal weights and bias parameters, performs linear transformations and nonlinear activation functions on the input mutual information to extract higher-level feature representations. These feature representations can capture the complex relationships and mutual influences between communication information among different leaders.

[0119] The prior network is used to evaluate the communication state of the agent and predict the reliability of communication, such as communication delays and failures. It can be an LSTM network. Based on the observed data and the communication trust score of the previous time step, it obtains the communication trust score (CTS) for the current time step. The observed data of the follower agent includes signal strength, packet loss rate, movement speed and task urgency, sensor data, etc.

[0120] The formula for communication trust scoring is:

[0121] CTS t =α·CTS t-1 +(1-α)·(β·S t +γ·(1-P t )+δ·V t +ε·E t );

[0122] Among them, CTS t The current communication trust score is given by α, where α is the weighting coefficient of historical CTS, and S is the weighting coefficient. t P represents the current signal strength. t V represents the packet loss rate. t E represents the speed of the agent's movement. t Let β, γ, δ, and ε be the indicator variables for task urgency, and let β, γ, δ, and ε be the weighting coefficients of each factor. Through this dynamic update mechanism, the Communication Trust Score (CTS) can reflect the communication quality and task relevance of each follower agent in real time. Based on the real-time CTS value, the leader agent dynamically adjusts the communication frequency and data volume with follower agents. For followers with lower CTS, the transmission of non-critical data is appropriately reduced, focusing on the core information required for task execution; while for followers with higher CTS, a higher communication frequency can be maintained to ensure the timeliness and accuracy of information. Through the refined management of the Communication Trust Score (CTS), agents can dynamically adjust the communication frequency and data volume according to communication quality and task requirements. This not only reduces unnecessary communication overhead but also ensures the timely delivery of critical information. The localized collaboration strategy for intra-area communication effectively reduces the communication burden on the leader agent, making information exchange between agents more efficient.

[0123] In one implementation, deep neural networks can also be used to automatically adjust the weights and update strategies of the CTS to adapt to complex and ever-changing communication environments.

[0124] The observation information and communication trust score (CTS) of the agent are integrated and encoded using an encoder. First, the observation information of the follower agent is represented as a matrix F∈R. n×∣f∣ Here, R represents the set of real numbers, n is the number of follower agents, |f| is the dimension of the features, and F is an n×|f| dimensional real matrix. Then, the communication trust score vector is multiplied element-wise with the follower agent's feature matrix F to form the first fused feature matrix F′, denoted as F′=F⊙CTS, where ⊙ denotes element-wise multiplication, and the communication trust score vector is expanded to the same dimension as F.

[0125] The first fused feature matrix is ​​input into the multi-head attention mechanism. This multi-head attention mechanism contains h independent attention heads, and for each attention head, three sets of trainable weight matrices (Q = W) are used. Q F′,K=W K F′,V=W V F′) is linearly transformed onto the input first fusion feature matrix F′ to obtain the Q matrix, K matrix, and V matrix for each attention head. For each attention head, the output is obtained according to the following formula:

[0126]

[0127] Where, d k This is the dimension of the attention head, primarily serving a normalization function to prevent the dot product from becoming too large, ensuring the attention score remains within an appropriate range, thus allowing for a more rational allocation of attention weights. The application of multi-head attention mechanisms further improves the efficiency of information filtering and integration, enabling the leader agent to quickly focus on important information.

[0128] Finally, the outputs of all attention heads are concatenated and transformed using a linear transformation matrix W. O To merge:

[0129] MultiHead(Q,K,V)=W O [head1,head2,…,head h ];

[0130] Where h is the number of attention heads, head i It is the output of the i-th attention head.

[0131] The second fusion feature matrix output by the multi-head attention mechanism module can be represented as m n×∣f∣ Here, m represents the information matrix between the leader agent and the follower agents, n represents the number of agents, and |f| represents the feature dimension. The second fused feature matrix, processed by a multi-head attention mechanism, integrates information from multiple attention heads, enabling it to capture the complex relationships and features between agents. This matrix primarily contains information from the follower agents after feature extraction and attention weighting.

[0132] The policy network is the core decision-making module in a multi-agent reinforcement learning system. It receives information from various sources, including a fused feature matrix processed by a multi-head attention mechanism and a second fused feature matrix. This information is fed into the policy network to generate the action probability distribution for each follower agent. Through learning and optimization, the network can output the optimal action selection policy for each agent based on current communication characteristics and task requirements.

[0133] The data processing flow of the leader intelligent agent is as follows Figure 3 As shown. Mutual information Processed by a multilayer perceptron. Observation information (O i O j O k ...O n The second fusion feature matrix (m) is processed sequentially by a prior network, an encoder, and a multi-head attention mechanism. i ,m j ,m k ...m nThen, the policy network fuses mutual information and the second fused feature matrix to obtain the action instructions (a) of the follower agent. i ,a j ,a k ...a n ).

[0134] In one implementation, the method further includes:

[0135] Real-time monitoring of the speed of all follower agents and the agent density in each region;

[0136] When the speed of any follower agent exceeds a preset speed threshold, the communication quality and connection stability between the follower agent and the leader agent in its region are evaluated. If there is low communication quality and / or unstable connection, the region to which the follower agent belongs is updated, and a smooth transition strategy is used for communication transition.

[0137] When the agent density in any region exceeds a preset density threshold, the communication quality and connection stability between all follower agents in that region and the leader of that region are evaluated. If there are cases of low communication quality and / or unstable connection, the region to which the follower agents in that region belong is updated, and a smooth transition strategy is used to perform a communication transition.

[0138] Specifically, once the initial region division is completed, it is necessary to monitor the speed of the follower agents and the agent density of the region in real time, and adjust the region as needed to ensure communication quality.

[0139] When a follower agent's speed exceeds a preset speed threshold, if there are instances of low communication quality and / or unstable connections between the follower agent with abnormal speed and the leader agent in its region, it means that the region to which the follower agent belongs needs to be updated to ensure communication quality. During the region update process, an incremental update strategy is adopted. Specifically, the affected region is defined as a preset multiple of the communication radius, centered on the follower agent with abnormal speed. This affected region covers all potential communication objects. The follower agent with abnormal speed is removed from its original cluster, and the distance from it to the center of each cluster within the affected region is calculated. Candidate clusters that meet the communication radius constraint are selected, and the cluster with the best communication quality with the leader agent within the candidate cluster is chosen as the new region to which the follower agent belongs.

[0140] When the agent density in any region exceeds a preset density threshold, the communication quality and connection stability between all follower agents in that region and their respective leaders are evaluated. If low communication quality and / or unstable connections exist, the regions belonging to the follower agents with low communication quality and / or unstable connections are updated. During the region update process, an incremental update strategy is adopted. Specifically, the affected region is defined as a preset multiple of the communication radius, centered on the follower agent with low communication quality and / or unstable connection. This affected region covers all potential communication objects. The follower agent with low communication quality and / or unstable connection is removed from the original cluster, and its distance to the center of each cluster within the affected region is calculated. Candidate clusters that meet the communication radius constraint are selected, and the cluster with the best communication quality with the leader agent within the candidate cluster is chosen as the new region. Criteria for determining low communication quality include: packet loss rate exceeding a preset packet loss rate threshold, latency exceeding a preset latency threshold, and bit error rate exceeding a preset bit error rate threshold. Criteria for determining unstable connections include: connection hold time below a preset hold time threshold and disconnection frequency exceeding a preset disconnection frequency threshold.

[0141] This invention uses a dynamic region division method to dynamically adjust according to the number of agents and task requirements, enabling the system to easily cope with the increase in the number of agents and the expansion of the task scale, while maintaining system stability.

[0142] In one implementation, a smooth transition strategy is used for communication transition, including:

[0143] Obtain the preset transition period duration and divide the transition period duration into n time steps;

[0144] Obtain a preset first communication frequency and a second communication frequency. The first communication frequency is the communication frequency with the original leader agent within the first time step, and the second communication frequency is the communication frequency with the newly assigned leader agent within the first time step. The sum of the first communication frequency and the second communication frequency is 1.

[0145] Obtain a preset ratio, and starting from the second time step, iteratively increase the second communication frequency and decrease the first communication frequency based on the preset ratio until the nth time step, at which point the communication transition is completed.

[0146] Specifically, to avoid communication chaos and coordination breakdown among agents due to sudden changes in region division during the region division update, a smooth transition strategy for region division is designed. After the new region division is determined, all communication connections between agents and the original leader are not immediately severed; instead, a transition period is established. During this transition period, agents gradually increase the frequency of information interaction with the newly assigned leader while gradually reducing their communication dependence on the original leader.

[0147] Let the total duration of the transition period be T, and divide it into n time steps. When currently at the t-th time step (t = 1, 2, ..., n), the second communication frequency with the new leader agent is f. new (t), the first communication frequency for communicating with the original leader agent is f. old (t). At the initial time (t=1), the frequency of communication with the new leader agent is the second communication frequency f. new (0), usually set to a lower value, such as f new (0) = 0.2, and the first communication frequency with the original leader agent is:

[0148] f old (1)=f original -f new (0) = 0.8. f original The original communication frequency when the follower agent has not updated the region is set to 1.

[0149] For each time step, the communication frequency with the new leader increases by a preset ratio k, while the communication frequency with the original leader decreases accordingly. This process is represented as:

[0150] f new (t)=f new (t-1)+k·f original ;

[0151] f old (t)=f original -f new (t);

[0152] Where k is a preset ratio, representing the communication frequency adjustment step size, which can be determined according to the transition period and the desired adjustment speed.

[0153] This approach ensures the stable operation of the multi-agent system during the region partitioning update process, reducing system oscillations and performance fluctuations caused by region partitioning adjustments.

[0154] Please see Figure 1 The multi-agent reinforcement learning method based on region partitioning described in this embodiment of the invention further includes the following steps:

[0155] In step S400, after each follower agent in the region executes the action command, the environment provides feedback on the global reward value and the new environment state, generates state transition samples and stores them in the experience pool. When the sampling conditions are met, batch data is extracted from the experience pool, and the policy network parameters of the leader agent are updated using the MADDPG (Multi-Agent Deep Deterministic Policy Gradient) algorithm.

[0156] Specifically, the state transition samples include the old state S t Combined action a t Global reward r t and new state s t+1 The joint action is a sequence of actions performed by each follower.

[0157] Please see Figure 1 The multi-agent reinforcement learning method based on region partitioning described in this embodiment of the invention further includes the following steps:

[0158] Step S500: Iterate the training until the preset termination condition is met to complete the training of the multi-agent system.

[0159] Specifically, the termination condition can be any one of the following: number of iterations, training time, or performance threshold metric.

[0160] Trained multi-agent systems can be applied to various fields, such as drones, autonomous driving, and robot control. In the drone domain, each drone is treated as an agent. After dividing the drone swarm into regions, a leader agent in each region makes decisions, interacts with follower agents within the region, and also interacts with leader agents outside the region. An application diagram in the drone domain is shown below. Figure 4 As shown; in the field of autonomous driving, each vehicle is treated as an intelligent agent. After dividing the vehicle into regions, the leader intelligent agent in each region makes decisions, interacts with the follower intelligent agents within the region, and interacts with the leader intelligent agent outside the region. In the field of robot control, each robot is treated as an intelligent agent. After dividing the robot into regions, the leader intelligent agent in each region makes decisions, interacts with the follower intelligent agents within the region, and interacts with the leader intelligent agent outside the region.

[0161] To improve the reliability of inter-regional communication and prevent system-wide communication paralysis due to the failure of a leader agent or the interruption of a communication link, this invention also designs redundant communication paths and fault-tolerance mechanisms. Each region's leader agent establishes communication connections with multiple neighboring region leader agents, forming a mesh network topology. When a communication link fails, communication information is automatically forwarded through other available links. Furthermore, the leader agent periodically exchanges heartbeat signals and status summary information with other leader agents to monitor the health of communication links in real time. Upon detecting a communication anomaly, fault-tolerance mechanisms are immediately activated, such as rerouting communication requests and temporarily elevating the communication privileges of neighboring leaders, ensuring uninterrupted inter-regional communication and stable system operation.

[0162] In real-world task execution, the communication architecture is not static but dynamically adjusted based on the operational status of the multi-agent system and task requirements. First, regions are divided according to the region division method of this invention, and a leader agent is selected for each region to make decisions. During task execution, system metrics are monitored in real time, including task complexity, task urgency, and system load. If the task complexity is below a preset complexity threshold, the communication mode can be switched to a centralized communication mode, where only one leader agent makes decisions. If the task urgency is above a preset threshold, it means a rapid response is needed, and the system can be switched to a distributed mode, allowing each agent to make its own decisions. If the system load is above a preset load threshold, it means the load on the leader agent needs to be reduced, and the system can be switched back to a distributed mode, allowing each agent to make its own decisions. For example, in an autonomous driving scenario, multiple vehicles need to maintain a certain formation. Initially, decision-making and communication are based on regional divisions. When road conditions are simple and traffic flow is low, a centralized communication mode can be switched to, with a lead vehicle coordinating speed and direction. When traffic is congested or an emergency occurs, requiring rapid response and local decision-making, a distributed communication mode is suitable, allowing each vehicle to make autonomous decisions based on its own sensor data and local information from other vehicles.

[0163] Furthermore, this invention utilizes the LSTM algorithm, taking sequential data such as communication latency, network congestion, and data transmission success rate as input during the training process. By learning patterns in these time-series data, the LSTM model can predict the optimal packet size under the current system state, balancing data transmission efficiency and network resource consumption. By establishing a mapping model between communication parameters and system performance, the system performance under different communication parameter configurations is predicted. During actual operation, based on the current system state and task requirements, optimal communication parameter configuration suggestions are obtained from the mapping model, and relevant parameters in the communication architecture, such as communication frequency, packet size, and hyperparameters of the multi-head attention mechanism, are adjusted in real time. This data-driven optimization method enables the communication architecture to continuously adapt to changing environments and task requirements, continuously improving communication efficiency and system performance.

[0164] In one implementation, the method further includes:

[0165] Training is performed using the decentralized communication architecture of blockchain.

[0166] Specifically, in this approach, each agent participates in the recording and verification of communication information, effectively improving the security and transparency of communication.

[0167] In summary, this invention initializes the state data, task objectives, and global reward function of all agents in a simulation environment; divides all agents into several regions, selecting a leader agent in each region; the leader agent in each region receives information from leader agents in other regions and observation data from follower agents within the region, and obtains action instructions for each follower agent based on the information and observation data; updates the policy network parameters of the leader agent; and iteratively trains until a preset termination condition is met, thus completing the training of multiple agents. This invention, by having the leader agent receive information from leader agents in other regions to make decisions, effectively maintains the ability to perceive the real-time policies of other agents.

[0168] In one embodiment, such as Figure 5 As shown, based on the above-described region-partition-based multi-agent reinforcement learning method, this invention also provides a region-partition-based multi-agent reinforcement learning system, the system comprising:

[0169] Initialization module 100 is used to initialize the simulation environment and initialize the state data, task objectives and global reward function of all agents;

[0170] Module 200 is selected to divide all agents into several regions based on state data, and select a leader agent for each region.

[0171] The decision module 300 is used to receive information from the leader agents in other areas and observation data from the follower agents in each area through the leader agent in each area, and to obtain action instructions for each follower agent in the area based on the information and the observation data.

[0172] The parameter update module 400 is used to update the global reward value and new environment state after each follower agent in the region executes the action command, store the state transition samples in the shared experience pool, and extract batch data from the experience pool when the sampling conditions are met, and use the MADDPG algorithm to update the policy network parameters of the leader agent.

[0173] The iteration module 500 is used to iteratively train the multi-agent system until the preset termination condition is met.

[0174] It should be noted that the foregoing explanation of the implementation of the region-based multi-agent reinforcement learning method also applies to the region-based multi-agent reinforcement learning system of this embodiment, and will not be repeated here.

[0175] This invention also provides a computer-readable storage medium storing a region-partition-based multi-agent reinforcement learning program. When executed by a processor, the region-partition-based multi-agent reinforcement learning program implements the steps of any region-partition-based multi-agent reinforcement learning method provided in this invention.

[0176] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0177] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the above device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0178] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0179] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0180] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / terminal devices and methods can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely illustrative. For instance, the division of modules or units described above is only a logical functional division, and in actual implementation, it can be divided in other ways. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.

[0181] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not mean that the essence of the corresponding technical solutions deviates from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A multi-agent reinforcement learning method based on region partitioning, characterized in that, The method includes: Initialize the simulation environment, and initialize the state data, task objectives, and global reward function of all agents; Based on state data, all agents are divided into several regions, and a leader agent is selected for each region. The leader agent in each region receives information from the leader agents in other regions and observation data from the follower agents in the region, and obtains action instructions for each follower agent in the region based on the information and the observation data. After each follower agent in the region executes the action command, the environment provides feedback on the global reward value and the new environment state. The state transition samples are stored in the shared experience pool. When the sampling conditions are met, batch data is extracted from the experience pool, and the MADDPG algorithm is used to update the policy network parameters of the leader agent. The training is carried out iteratively until the preset termination condition is met, thus completing the training of the multi-agent system.

2. The multi-agent reinforcement learning method based on region partitioning according to claim 1, characterized in that, The status data includes location data; Based on state data, all agents are divided into several regions, including: Obtain the preset communication radius for each agent; Determine if the number of clusters can be obtained; If the number of clusters is available, K-means clustering is performed based on the number of clusters to obtain several initial clusters. DBSCAN clustering optimization is then performed on all the initial clusters to obtain several regions. If the number of clusters cannot be obtained, the Euclidean distance is calculated based on the position data of any two agents. A distance matrix is ​​constructed based on all the Euclidean distances and the communication radius. The DBSCAN clustering algorithm is then used to cluster the data based on the distance matrix to obtain several regions.

3. The multi-agent reinforcement learning method based on region partitioning according to claim 2, characterized in that, K-means clustering is performed based on the number of clusters, resulting in several initial clusters, including: Randomly select a number of agents from each cluster as the initial cluster centers; Perform distance correction calculations for each agent: Calculate the Euclidean distance from each agent to the center of each initial cluster based on the location data of each agent; If the original Euclidean distance is less than or equal to the communication radius, then the original Euclidean distance is used as the output distance. If the original Euclidean distance is greater than the communication radius, the original Euclidean distance is multiplied by the preset penalty factor to obtain the output distance. Assign each agent to the cluster with the smallest output distance, recalculate the geometric center position of each cluster, and set the agent closest to the geometric center position as the new cluster center; The process of iteratively performing distance correction calculations and setting new cluster centers continues until the iteration termination condition is met, resulting in several initial clusters.

4. The multi-agent reinforcement learning method based on region partitioning according to claim 1, characterized in that, One leader agent is selected for each region, including: The agent closest to the geometric center of the region is selected as the leader agent.

5. The multi-agent reinforcement learning method based on region partitioning according to claim 1, characterized in that, Based on the information and the observation data, action instructions for each follower agent within the region are obtained, including: Based on the information received from the leader agents of other regions and the state data of the leader agents of this region, calculate the mutual information between regions; By using a multilayer perceptron within the leader's intelligent body to process all mutual information, several inter-regional cooperative feature vectors are obtained. The communication trust score at the current moment is obtained by utilizing the prior network within the leader's intelligent body based on the observed data and the communication trust score at the previous moment. The observed data and the communication trust score are processed by the encoder in the leader agent to obtain feature vectors of several follower agents. The feature vectors of all follower agents are processed using the multi-head attention mechanism within the leader agent to obtain the second fusion feature matrix; By using the policy module within the leader agent to process all the inter-regional cooperation feature vectors and the second fused feature matrix, action instructions for each follower agent within the region are obtained.

6. The multi-agent reinforcement learning method based on region partitioning according to claim 1, characterized in that, Before receiving information from leader agents in other regions, the process also includes: The information to be sent in the leader's intelligent body at the sending end is compressed according to different levels of abstraction to obtain multiple candidate transmission information; Select the candidate transmission information with the highest level of abstraction and send it.

7. The multi-agent reinforcement learning method based on region partitioning according to claim 1, characterized in that, The method further includes: Real-time monitoring of the speed of all follower agents and the agent density in each region; When the speed of any follower agent exceeds a preset speed threshold, the communication quality and connection stability between the follower agent and the leader agent in its region are evaluated. If there is low communication quality and / or unstable connection, the region to which the follower agent belongs is updated, and a smooth transition strategy is used for communication transition. When the agent density in any region exceeds a preset density threshold, the communication quality and connection stability between all follower agents in that region and the leader of that region are evaluated. If there are cases of low communication quality and / or unstable connection, the region to which the follower agents in that region belong is updated, and a smooth transition strategy is used to perform a communication transition.

8. The multi-agent reinforcement learning method based on region partitioning according to claim 7, characterized in that, Communication transitions are performed using smooth transition strategies, including: Obtain the preset transition period duration and divide the transition period duration into n time steps; Obtain a preset first communication frequency and a second communication frequency. The first communication frequency is the communication frequency with the original leader agent within the first time step, and the second communication frequency is the communication frequency with the newly assigned leader agent within the first time step. The sum of the first communication frequency and the second communication frequency is 1. Obtain a preset ratio, and starting from the second time step, iteratively increase the second communication frequency and decrease the first communication frequency based on the preset ratio until the nth time step, at which point the communication transition is completed.

9. A multi-agent reinforcement learning system based on region partitioning, characterized in that, include: The initialization module is used to initialize the simulation environment and initialize the state data, task objectives, and global reward function of all agents. The selection module is used to divide all agents into several regions based on state data, and select a leader agent for each region. The decision-making module is used to receive information from leader agents in other regions and observation data from follower agents in each region through the leader agent in each region, and to obtain action instructions for each follower agent in the region based on the information and the observation data. The parameter update module is used to update the policy network parameters of the leader agent after each follower agent in the region executes the action command, and the environment feeds back the global reward value and the new environment state. The state transition samples are stored in the shared experience pool. When the sampling conditions are met, batch data is extracted from the experience pool and the MADDPG algorithm is used to update the policy network parameters of the leader agent. The iterative module is used to iteratively train the multi-agent system until the preset termination condition is met.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a region-partition-based multi-agent reinforcement learning program, which, when executed by a processor, implements the steps of the region-partition-based multi-agent reinforcement learning method as described in any one of claims 1-8.

Citation Information

Cited By

  • Multi-agent learning method, device and equipment based on dynamic weighted average field

    CN121809582A