Multi-agent collaborative decision-making method based on teammate representation and incentive communication

Through the multi-agent collaborative decision-making method of teammate representation and excitation communication, a multi-layer perceptron and gated loop unit is used to generate historical trajectory representations, and customized excitation messages are generated in combination with the teammate representation module and message generation module, which solves the problem of information asymmetry in the multi-agent system and realizes efficient and stable collaborative decision-making and expansion capabilities.

CN120337980AActive Publication Date: 2025-07-18NO 15 INST OF CHINA ELECTRONICS TECH GRP

Patent Information

Application Number
CN202510812835.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

The existing multi-agent collaborative decision-making methods are difficult to overcome information asymmetry in a locally observable environment, resulting in high communication redundancy, long delay, unstable training, and difficult to expand to large-scale agent systems.

Method used

Through teammate representation and excitation communication, a multi-layer perceptron and gated loop unit is used to generate historical trajectory representations, a customized excitation message is generated by combining teammate representation modules and message generation modules, and a hybrid neural network is used to fusion individual decision-making, and dynamic sparse communication and QMIX framework are used to ensure system stability and scalability.

Benefits of technology

It realizes collaborative decision-making with high precision and low latency, reduces communication redundancy, improves the scalability and training stability of the system, and is suitable for distributed scenarios with resource limitations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337980A_ABST
    Figure CN120337980A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent collaborative decision-making method based on teammate characterization and incentive communication, and relates to the field of multi-agent reinforcement learning, and the method comprises the steps: carrying out the processing of each agent through a multi-layer perceptron and a gating circulation unit based on a current observation value and a previous moment action, and generating historical track characterization; generating a local action value function through the local action value function branch; generating a message sent to a teammate agent through a teammate characterization module and a message generation module; summing the messages sent by all teammate intelligent agents, and adding the summed messages into a local action value function as an incentive item; and fusing the corrected local action value functions of all the agents through a hybrid neural network to obtain a team total value function and generate a collaborative decision action. Through teammate behavior prediction and a dynamic attention communication mechanism, multi-agent cooperation efficiency and decision precision are improved, low-redundancy sparse communication is realized, and system training stability and large-scale expansion capability are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multi-agent reinforcement learning, and more specifically, to a multi-agent collaborative decision-making method based on teammate representation and incentive communication. Background Art

[0002] The present invention focuses on the field of multi-agent collaborative decision-making, and is corely applied to practical scenarios such as cooperative reconnaissance of UAV swarms and formation flight that require efficient distributed cooperation.

[0003] In a partially observable environment, existing methods are difficult to overcome the problem of information asymmetry among agents. Although exchanging raw or shallowly processed observation information (such as sensor data) through communication can expand the individual perception range, this method increases the local observation dimension and state space of each agent, improves the complexity of the policy network and the training difficulty. Teammates need to infer the intentions and behavior patterns of other agents from the received miscellaneous information by themselves, which not only prolongs the training convergence time, but also causes delays in the decision-making process, and it is difficult to achieve fast and accurate goal-oriented cooperation.

[0004] Existing communication mechanisms generally lack pertinence and selectivity. Common methods such as broadcast or fixed-topology communication cause agents to send information to a large number or even all teammates, including a lot of content that is of little or no value to the decision-making of specific teammates. This kind of extensive information transmission generates significant communication redundancy, resulting in unnecessary bandwidth occupation, energy consumption and communication delay. Especially in practical application scenarios with limited bandwidth resources or a large number of agents (such as large-scale UAV formations, distributed sensor networks), this high-overhead and inefficient communication mode has become the main bottleneck restricting the system scale, real-time performance and deployment feasibility.

[0005] To improve the collaborative effect, some existing technologies adopt a centralized decision-making or tightly coupled communication architecture, sacrificing the decentralization and scalability of the system, and it is difficult to adapt to large-scale agent groups. At the same time, at the training level, how to effectively coordinate individual optimization and the overall team goal, and maintain the stability of the training process after introducing the communication module, still faces severe challenges. Its lack of a robust and scalable value function decomposition and optimization mechanism makes the algorithm prone to problems such as training oscillation, difficult convergence or poor final performance when facing complex tasks and high-dimensional agent spaces, limiting the practicality of the method.

[0006] Therefore, how to design a multi-agent collaborative decision-making method based on teammate representation and incentive communication, which can achieve high-precision and low-latency collaborative decision-making under partially observable conditions, while having high communication efficiency, strong scalability and training stability, is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0007] In view of this, the present invention provides a multi-agent collaborative decision-making method based on teammate representation and incentive communication, which directly optimizes the individual decision-making efficiency by accurately predicting the behavior of teammates and generating customized incentive messages; combines a dynamic sparse communication mechanism to reduce redundant overhead; and relies on an extensible collaborative framework to ensure the stability and scalability of the system, realizing efficient teamwork.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] A multi-agent collaborative decision-making method based on teammate representation and incentive communication, comprising performing the following steps for each agent i in a multi-agent system:

[0010] S1. Based on the current observation and the previous action , process through a multi-layer perceptron and a gated recurrent unit to generate a historical trajectory representation ;

[0011] S2. Input the historical trajectory representation into the local action value function branch, and generate a local action value function through a multi-layer perceptron; and input the historical trajectory representation into the teammate representation module and the message generation module for processing to generate a message sent to the teammate agent;

[0012] S3. Sum the messages sent by all teammate agents, and add the summation result as an incentive term to the local action value function to obtain a corrected local action value function ;

[0013] S4. Fuse the corrected local action value functions of all agents through a hybrid neural network to obtain a team overall value function and generate a collaborative decision-making action.

[0014] Preferably, the S1 includes:

[0015] Concatenate the current observation and the previous action into an input vector , and process through a multi-layer perceptron and a gated recurrent unit in sequence to obtain the hidden state of the gated recurrent unit, and the hidden state is the historical trajectory representation .

[0016] Preferably, in the S2, the processing of the teammate representation module includes:

[0017] According to the historical trajectory representation and the teammate number , calculate the multi-dimensional Gaussian distribution through the encoder ; among them, the dimension of the multi-dimensional Gaussian distribution is the same as the dimension of the agent action space, representing the action selection probability of the teammate agent in each dimension; represents the action mean vector, represents the action variance vector;

[0018] Based on the multi-dimensional Gaussian distribution , sample to obtain the teammate representation vector .

[0019] Preferably, in the teammate representation module, the maximization of mutual information strategy is adopted for optimization, including:

[0020] Define the mutual information objective, ;

[0021] Introduce a variational autoencoder to calculate the variational distribution , with the KL divergence as the loss function:

[0022]

[0023] Among them, includes the encoder parameters and the variational parameters , represents the expectation, represents the KL divergence, and p represents the conditional distribution.

[0024] Preferably, in S2, the message generation module processing includes:

[0025] According to the historical trajectory representation and the teammate representation vector , through the processing of the fully connected layer and the combination of the attention mechanism, generate the communication weight ;

[0026] Multilayer perceptron process the historical trajectory representation and the teammate representation vector to generate the message vector , and calculate the message sent to the teammate agent.

[0027] Preferably, in the message generation module, generating the communication weight includes:

[0028] Calculate the query vector and the key vector , where, 、 are weight coefficients;

[0029] Calculate the normalized weight:

[0030]

[0031] Among them, represents the scaling factor, represents the query vector transpose of, represents the key vector of the i-th agent to the m-th agent.

[0032] Preferably, the generating communication weight , further includes defining an entropy regularization loss:

[0033]

[0034] Among them, represents the message generation module parameter.

[0035] Preferably, the generating communication weight , further includes setting a global communication threshold β

[0036] ;

[0037] When , cut off the communication link from agent i to agent j, and renormalize the weights of the remaining communication links to ensure that the sum of communication weights holds.

[0038] Preferably, in the S3, the corrected local action value function is expressed as:

[0039]

[0040] Among them, represents the message sent by agent j to agent i.

[0041] Preferably, in the S4, the hybrid neural network adopts the QMIX framework, including:

[0042] A hypernetwork for receiving the global state and generating the weight parameters of a fully connected neural network;

[0043] A fully connected network for inputting the corrected local action value functions of all agents and outputting the team's overall value function , satisfying the monotonicity constraint .

[0044] Through the above technical solutions, compared with the prior art, the technical solutions of the present invention have the following beneficial effects:

[0045] 1. This method constructs a teammate behavior prediction model and generates customized incentive messages. The intelligent body can accurately understand the intentions of teammates and directly integrate the incentive information into the local decision-making function. While avoiding increasing the complexity of individual strategies, this mechanism can effectively guide team behavior coordination and significantly improve the accuracy and response timeliness of multi-agent collaborative decision-making in a local observation environment.

[0046] 2. A dynamic attention mechanism is adopted to generate different communication weights for different teammates. Combining entropy optimization and threshold control strategies, the intelligent body only sends targeted incentive messages to key collaboration objects and automatically cuts off inefficient communication links. This design significantly reduces communication redundancy and bandwidth consumption while ensuring the effectiveness of team collaboration, and is applicable to resource-constrained distributed scenarios.

[0047] 3. Based on the value function decomposition framework, individual decisions and team goals are integrated, and a centralized training and decentralized execution mechanism is used to coordinate global optimization and local strategies. The hybrid network structure strictly maintains the monotonic correlation between individual values and team values, ensuring that the algorithm has stable convergence characteristics and large-scale expansion capabilities in complex multi-agent systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0049] Figure 1 Schematic flow chart of a multi-agent collaborative decision-making method based on teammate representation and incentive communication provided by an embodiment of the present invention;

[0050] Figure 2 Schematic flow chart of a multi-agent collaborative decision-making method applied to an unmanned aerial vehicle cluster provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0052] Embodiment 1;

[0053] Such as Figure 1As shown in the figure, this embodiment provides a multi-agent collaborative decision-making method based on teammate representation and incentive communication, including performing the following steps for each agent i in the multi-agent system:

[0054] S1. Based on the current observation value and the previous moment's action , through processing by a multi-layer perceptron and a gated recurrent unit, generate a historical trajectory representation ;

[0055] S2. Input the historical trajectory representation into the local action value function branch, and generate a local action value function through a multi-layer perceptron; and input the historical trajectory representation into the teammate representation module and the message generation module for processing, and generate a message sent to the teammate agent;

[0056] S3. Sum up the messages sent by all teammate agents, and add the summation result as an incentive term to the local action value function to obtain a corrected local action value function ;

[0057] S4. Through a hybrid neural network, fuse the corrected local action value functions of all agents to obtain a team overall value function and generate a collaborative decision-making action.

[0058] This method significantly improves the accuracy and real-time performance of team collaboration by accurately predicting the behavior intentions of teammates and generating customized incentive messages, and integrating key information into the individual decision-making process in an efficient form; at the same time, it uses a dynamic attention mechanism to achieve targeted screening and sparse control of the communication link, and greatly reduces the communication overhead while ensuring the collaborative efficiency; finally, relying on an extensible value function decomposition framework and a centralized training and decentralized execution mechanism, it ensures that the system has stable and efficient collaborative decision-making capabilities in large-scale distributed scenarios.

[0059] The following further elaborates on each step in the above embodiment;

[0060] In S1 of this embodiment, based on the current observation value and the previous moment's action , through processing by a multi-layer perceptron and a gated recurrent unit, generate a historical trajectory representation ; specifically including:

[0061] Concatenate the current observation value and the previous moment's action into an input vector , processed by a multi-layer perceptron and a gated recurrent unit in sequence to obtain the hidden state of the gated recurrent unit, where the hidden state is a historical trajectory representation .

[0062] In step S2 of this embodiment, the historical trajectory representation is input into the local action value function branch, and a local action value function is generated through a multi-layer perceptron ; and the historical trajectory representation is input into the teammate representation module and the message generation module for processing to generate a message sent to the teammate agent ;

[0063] Among them, the processing of the teammate representation module includes:

[0064] According to the historical trajectory representation and the teammate number , a multi-dimensional Gaussian distribution is calculated through an encoder ; among them, the dimension of the multi-dimensional Gaussian distribution is the same as the dimension of the agent action space, indicating the action selection probability of the teammate agent in each dimension; represents the action mean vector, represents the action variance vector;

[0065] Based on the multi-dimensional Gaussian distribution , a teammate representation variable of agent i for agent j is sampled , as a prediction of the teammate's behavior, and is used to guide the generation of the information transmitted from agent i to agent j. By modeling the teammate to predict the teammate's behavior, the agents can clearly understand each other's intentions.

[0066] Furthermore, in order to make the predicted teammate behavior of the established teammate model more consistent with the actual actions taken by the teammate, that is, the prediction is more accurate, this embodiment uses the method of maximizing the mutual information to update and optimize the teammate model. The mutual information can calculate the correlation between two random variables, which can be understood as the amount of information about another random variable Y contained in a random variable X. The definition of the mutual information between random variable X and random variable Y is shown in the following formula:

[0067]

[0068] The larger its value means the higher the degree of association between the two variables. When the two random variables have no relationship, their mutual information is 0. Therefore, it can be used to measure the prediction quality of an agent's behavior towards another agent.

[0069] In this embodiment, the predicted value of the teammate's behavior, that is, the distribution sampling of the teammate modeling model is regarded as a random variable, and the observed teammate behavior It is regarded as another random variable, and then the mutual information between them is calculated to represent the correlation between the two, and the teammate model is updated and optimized by maximizing the mutual information.

[0070] The mutual information between the two can be expressed as the following formula:

[0071]

[0072] Among them, the entropy represents the uncertainty of the teammate representation variable when only the historical trajectory of agent i and the teammate number j are known. The entropy is the uncertainty of the representation variable when , the teammate number, and the true action of the teammate are known. The difference between the two is the correlation between the predicted representation variable and the true action. The larger the difference, the more accurate the prediction. If the entropy is relatively and cannot provide any additional information, that is, when the two are equal, the value of the mutual information will be 0. Maximizing the mutual information is equivalent to reducing the deviation between the predicted value and the true action value of the teammate representation model under the condition of knowing the local information of agent i, so that the teammate representation model of the agent can more accurately understand and predict the behavior of the teammate, and thus can better achieve cooperation between teams.

[0073] When calculating the mutual information, it is necessary to calculate the conditional distribution , but it is difficult to calculate the conditional distribution in reality. And maximizing the mutual information can be achieved by maximizing the lower bound of the mutual information. In this embodiment, variational inference is used to calculate the lower bound of the mutual information: a variational autoencoder with a parameter of is introduced to calculate the variational distribution , and this variational distribution is used to approximately calculate the conditional distribution . The variational encoder is essentially a recurrent neural network. Then the KL divergence is used to measure the similarity between the two distributions, and the negative value of the expectation of the KL divergence is used as the lower bound of the mutual information, as shown in the following formula:

[0074]

[0075] As can be seen from the above formula, maximizing the mutual information is equivalent to minimizing the KL divergence. Therefore, in this embodiment, the KL divergence is used as the loss function of the teammate representation module for the update and optimization of the representation module:

[0076]

[0077] Among them, includes the encoder parameters and variational parameters , represents the expectation, represents the KL divergence, p represents the conditional distribution. Through continuous training and updating, the teammate representation model will be continuously optimized, and the predicted value will also approach the real action.

[0078] Furthermore, the message generation module processes as follows:

[0079] According to the historical trajectory representation and the teammate representation vector , through processing by a fully connected layer and combining with the attention mechanism, generate the communication weight ;

[0080] Perform multi-layer perceptron processing on the historical trajectory representation and the teammate representation vector to generate the message vector , and calculate the message sent to the teammate agent .

[0081] Furthermore, generating the communication weight includes:

[0082] Calculate the query vector and the key vector , where , are weight coefficients;

[0083] Calculate the normalized weight:

[0084]

[0085] where represents the scaling factor, represents the transpose of the query vector , represents the key vector of the i-th agent for the m-th agent.

[0086] After introducing the communication weight, in order to make the agents more concentrated when communicating, that is, only send information to other necessary agents, discard most of the invalid communications, and reduce the communication overhead to achieve efficient and sparse communication. In this embodiment, a sparsity regularization term is further introduced to achieve efficient communication by optimizing the entropy of the categorical distribution composed of the communication weights. The entropy loss function is as follows:

[0087]

[0088] In the above formula, are the parameters of the message generator. By minimizing the entropy loss, the message generation module can generate communication weights with lower uncertainty.

[0089] In addition, in order to further cut off the useless redundant communication links existing between agents and reduce the communication cost, this embodiment proposes the concept of a threshold for communication weights, that is, for a single agent, the sum of its communication weights as an information source is 1, that is , set the communication threshold , when , the communication link from i to j will be cut off. By setting the communication threshold, the overhead of multi-agent communication is further reduced, the efficiency of information interaction is improved, and thus the collaborative decision-making level of multi-agents is enhanced.

[0090] In step S3 of this embodiment, the messages sent by all teammate agents are summed, and the summation result is added as an incentive term to the local action value function to obtain the corrected local action value function ; the corrected local action value function is expressed as:

[0091]

[0092] where represents the message sent by agent j to agent i.

[0093] In step S4 of this embodiment, the corrected local action value functions of all agents are fused through a hybrid neural network to obtain the overall team value function and generate collaborative decision-making actions; specifically, the hybrid neural network adopts the QMIX framework, including:

[0094] A hypernetwork for receiving the global state and generating the weight parameters of a fully connected neural network;

[0095] A fully connected network for inputting the corrected local action value functions of all agents and outputting the overall team value function , satisfying the monotonicity constraint .

[0096] The collaborative decision-making method provided by this embodiment generates accurate teammate representation vectors through a teammate behavior prediction model, and dynamically generates customized incentive messages and sparse communication weights based on the attention mechanism; directly integrates the incentive information into the individual value function to optimize the local decision, and at the same time combines entropy regularization and threshold control to achieve efficient communication link pruning; finally, relies on the QMIX framework to integrate individual decisions into the team goal, and significantly improves the collaborative efficiency and scalability of the multi-agent system in a partially observable environment while ensuring the training stability.

[0097] Embodiment 2;

[0098] Aiming at the problems faced by the UAV swarm in forest fire reconnaissance, such as limited dynamic environment perception, tight communication bandwidth, and low collaborative decision-making efficiency, this embodiment applies a multi-agent collaborative decision-making method based on teammate representation and incentive communication to a 4-UAV swarm, realizing a closed-loop framework of multi-source perception data fusion → probabilistic representation of teammate behavior → attention-weighted communication → global monotonic optimization decision-making, effectively solving the key challenges of collaborative path planning and obstacle avoidance in the complex environment of the fire field.

[0099] In this embodiment, 4 UAVs are required to cooperate in performing the forest fire reconnaissance mission. The goal is to dynamically cover the fire field area, locate the fire point in real time and avoid the thick smoke area. The observation values are the UAV's own position / GPS signal strength / visible fire point coordinates / surrounding smoke concentration. The actions are the flight direction (8 azimuth angles) and speed gear (3 gears). The bandwidth is limited, and low-value messages need to be dynamically filtered.

[0100] As Figure 2 shown, its specific implementation steps include:

[0101] 1) Generation of historical trajectory representation;

[0102] Each UAV splices the current observation values (such as fire point coordinates, smoke concentration) with the previous moment's action into a vector. Through processing by a two-layer MLP (128-dimensional hidden layer) and GRU (64-dimensional hidden state), the historical trajectory representation is output; when the UAV detects a high smoke concentration on a certain side of the fire field multiple times, the historical trajectory representation will encode the high smoke risk mode on the east side. The sequential memory of the GRU enables the UAV to automatically avoid the eastward path in subsequent decisions, improving the survival rate.

[0103] 2) Teammate intention modeling and communication;

[0104] Based on the historical trajectory representation and the teammate number , a Gaussian distribution is generated, and the teammate representation vector is obtained by sampling; and the key teammates are screened through the communication weight . When the weight is lower than the threshold β, the communication link is cut off to generate the message ; when the UAV discovers a new fire point, only high-weight messages ( > 0.7) are sent to the teammates close to this area, ignoring the teammates far from the fire field, reducing some redundant communications.

[0105] 3) Local decision correction;

[0106] The local action value function evaluates the action benefit, receives the teammate messages and sums them to obtain the incentive term, and corrects it to ; The drone originally planned to fly northeast received a thick smoke warning ( =-0.8), and the correction value decreased from +1.5 to +0.7, triggering a turn to the northwest safety path.

[0107] 4) Global collaborative decision-making;

[0108] The QMIX hybrid network inputs all correction value functions and the global fire scene state, and outputs the team value function and generates collaborative actions that satisfy the monotonicity constraint of the monotonicity constraint. The decision result dynamically assigns the roles of the drones: Drone 1 covers the unreconnoitered area, Drone 2 supports the fire point, and Drones 3 / 4 block the diffusion path, shortening the task completion time by 35%.

[0109] In this embodiment, the engineering effectiveness of the technical solution is verified. The teammate intention modeling improves the collaborative coverage rate by 40%, the attention communication mechanism reduces the bandwidth consumption by 60%, and the QMIX global optimization ensures the consistency of the fire extinguishing target. This system provides a patentable distributed decision-making paradigm for disaster rescue scenarios, with both real-time performance and robustness.

[0110] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0111] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multi-agent collaborative decision-making method based on teammate representation and incentive communication, characterized in that Including performing the following steps for each agent i in the multi-agent system: S1. Based on the current observation value and the action at the previous moment , it is processed by a multi-layer perceptron and a gated recurrent unit to generate a historical trajectory representation ; S2. Input the historical trajectory representation into the local action value function branch, and generate the local action value function through a multi-layer perceptron ; And represent the historical trajectory Input it to the teammate representation module and the message generation module for processing, and generate a message to be sent to the teammate agent ; S3. Sum the messages sent by all teammate agents, and add the summation result as an incentive term to the local action value function to obtain the corrected local action value function ; S4. Fusion of the corrected local action value functions of all agents through a hybrid neural network to obtain the overall team value function And generate collaborative decision-making actions.

2. The multi-agent collaborative decision-making method based on teammate representation and incentive communication according to claim 1, wherein The S1 includes: Concatenate the current observation and the action at the previous moment to form an input vector , and process it through a multi-layer perceptron and a gated recurrent unit in sequence to obtain the hidden state of the gated recurrent unit, where the hidden state is a representation of the historical trajectory .

3. A multi-agent collaborative decision-making method based on teammate representation and incentive communication according to claim 1, characterized in that In the S2, the processing of the teammate representation module includes: According to the historical trajectory representation and the teammate number , calculate the multi-dimensional Gaussian distribution through the encoder ; among them, the dimension of the multi-dimensional Gaussian distribution is the same as the dimension of the agent action space, representing the action selection probability of the teammate agent in each dimension; represents the action mean vector, represents the action variance vector; Based on the multi-dimensional Gaussian distribution , sample to obtain the teammate representation vector .

4. The multi-agent collaborative decision-making method based on teammate representation and incentive communication according to claim 3, wherein, In the teammate representation module, the optimization using the maximum mutual information strategy includes: Define the mutual information objective, ; Introduce a variational autoencoder to calculate the variational distribution , with the KL divergence as the loss function: ; Among them, including encoder parameters and variational parameters , denotes expectation, denotes KL divergence, and p denotes conditional distribution.

5. The multi-agent collaborative decision-making method based on teammate representation and incentive communication according to claim 1, wherein, In the S2, the processing of the message generation module includes: According to the historical trajectory representation and the teammate representation vector , through the processing of the fully connected layer and the combination of the attention mechanism, generate the communication weight ; Characterize the historical trajectory and the teammate representation vector through multi-layer perceptron processing to generate a message vector , and calculate the message sent to the teammate agent .

6. The multi-agent collaborative decision-making method based on teammate representation and incentive communication according to claim 5, wherein, In the message generation module, a communication weight is generated including: Calculate the query vector and the key vector , where , are weight coefficients; Calculating the normalized weight: ; Among them, represents the scaling factor, represents the query vector transpose, represents the key vector of the i-th agent with respect to the m-th agent.

7. A multi-agent collaborative decision-making method based on teammate representation and incentive communication according to claim 6, characterized in that Said generation of communication weights , further comprising defining an entropy regularization loss: ; Among them, represents the message generation module parameter.

8. A multi-agent collaborative decision-making method based on teammate representation and incentive communication according to claim 6, characterized in that The generation of communication weights , further includes setting a global communication threshold β ; When occurs, cut off the communication link from agent i to agent j and renormalize the weights of the remaining communication links to ensure that the sum of the communication weights holds.

9. A multi-agent collaborative decision-making method based on teammate representation and incentive communication according to claim 1, characterized in that In S3, the corrected local action value function is expressed as: ; Among them, represents the message sent by agent j to agent i.

10. A multi-agent collaborative decision-making method based on teammate representation and incentive communication according to claim 1, characterized in that, In the S4, the hybrid neural network adopts the QMIX framework, including: A hypernetwork for receiving the global state and generating the weight parameters of the fully connected neural network; A fully connected network, which is used to input the corrected local action value functions of all agents and output the overall team value function , satisfying the monotonicity constraint .

Citation Information

Patent Citations

  • Multi-agent communication cooperation method

    CN113435475A

  • Multi-agent cooperative communication strategy training system and method based on teammate perception

    CN114757092A

  • Multi-agent reinforcement learning algorithm based on multi-head attention mechanism communication

    CN116341611A

  • Multi-agent collaborative exploration method for constructing intrinsic individual rewards based on information gain

    CN120087410A

  • Distributed control of heterogeneous multi-agent systems

    US10983532B1

Cited By

  • Mirror image-based agent training system

    CN121303182A

  • Multi-agent collaborative decision-making method and device

    CN122021785A

  • Multi-agent collaborative decision-making method and device thereof

    CN122021785B