Method and device for evaluating agent contribution degree of multi-agent system, storage medium, electronic equipment and computer program product
By combining MADDPG reinforcement learning with graph neural networks, a graph structure for a multi-agent system is constructed, which solves the problems of insufficient dynamic adjustment and contribution evaluation in multi-agent cooperative systems. This achieves efficient agent cooperation and contribution evaluation, and improves the system's adaptability and cooperative efficiency.
Patent Information
- Application Number
- CN202511058119.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-07-30
AI Technical Summary
Existing technologies in multi-agent cooperative systems suffer from insufficient dynamic adjustment and a lack of effective contribution evaluation. The application of graph neural networks in multi-agent cooperative systems also suffers from insufficient dynamic adjustment and a lack of effective contribution evaluation.
By combining MADDPG reinforcement learning with graph neural networks, a graph structure of a multi-agent system is constructed. The graph neural network model is used to evaluate the contribution of the agents. A centralized training and decentralized execution approach is adopted, combined with a pre-set reward function and value network, to realize the strategy game and coordination among the agents.
It improves the adaptability and collaborative efficiency of multi-agent systems to dynamic task changes, enables refined contribution assessment, and ensures the fairness of reward distribution and the improvement of overall system performance.
Smart Images

Figure CN120912003A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of multi-agent system, in particular to an evaluation method and device for agent contribution degree of a multi-agent system, a storage medium, an electronic device and a computer program product. BACKGROUND
[0002] The cooperative task of a multi-agent system usually involves complex interaction relationships and dynamic environmental changes. For example, in a rescue search scenario, multiple agents need to quickly locate the target in an unknown environment while avoiding interference with each other; in a logistics sorting scenario, agents need to dynamically adjust the division of labor according to real-time task requirements to improve overall efficiency.
[0003] Graph neural networks (GNNs) have shown significant advantages in modeling complex relationship networks. Graph neural network models capture the interaction relationships between agents through graph structures and aggregate local and global information using a message passing mechanism. For example, in social network analysis or traffic prediction, GNNs have been proven to be effective in modeling dynamic topological relationships.
[0004] However, the inventors of the present application have found that the application of GNNs in multi-agent cooperative systems has problems of insufficient dynamic adjustment and lack of effective contribution degree evaluation. For example, the graph structure of a multi-agent system needs to be dynamically adjusted according to task requirements or environmental changes, which puts higher requirements on the real-time performance of GNNs. It is also a current research difficulty to convert the embedding vectors output by GNNs into an interpretable "contribution degree" indicator and combine it with the task objective function.
[0005] The contents of the background section only represent the knowledge of the discloser and do not necessarily represent the prior art in the field. SUMMARY
[0006] According to an aspect of the present application, an evaluation method for agent contribution degree of a multi-agent system is provided. The evaluation method comprises: determining a graph neural network model corresponding to the multi-agent system according to received environment state information, wherein the environment state information at least includes agent state information, task requirements and environmental information; training an action network and the graph neural network model according to the environment state information; determining action information of nodes of the graph neural network model through the trained action network according to the environment state information; and determining the contribution degree of the agent through the trained graph neural network model according to the environment state information and the action information.
[0007] According to some embodiments of the present application, the determining of the graph neural network model corresponding to the multi-agent collaborative system according to the received environment state information comprises: determining a node model according to the agent state information and the task demand; determining an edge structure according to the agent state information, the task demand, and the environment information; and determining the graph neural network model according to the node model and the edge structure.
[0008] According to some embodiments of the present application, the training of the action network and the graph neural network model according to the environment state information comprises: performing the training of the action network step and the training of the graph neural network model step at least once. The training of the action network step comprises: determining training action information of the agent by the action network according to the environment state information; determining a value evaluation parameter of the agent by the value network according to the training action information, the environment state information, and a preset reward function; updating the environment state information according to the value evaluation parameter; and training the action network according to the updated environment state information. The training of the graph neural network model step comprises: determining an experience trajectory according to the environment state information, the training action information, and the updated environment state information; determining a global task performance according to the experience trajectory; and training the graph neural network model according to the global task performance.
[0009] According to some embodiments of the present application, the preset reward function comprises a preset cooperation function and a preset game competition function.
[0010] According to some embodiments of the present application, the determining of the contribution degree of the agent by the trained graph neural network model according to the environment state information and the action information comprises: mapping the environment state information and the action information to a high-dimensional embedding space by a first layer of the trained graph neural network to obtain a first feature vector of a node; determining neighbor information of the node according to the environment state information; determining an updated feature vector of the node according to the first feature vector and the neighbor information by a subsequent layer of the trained graph neural network; and determining the contribution degree of the agent according to the updated feature vector by a linear layer of the trained graph neural network.
[0011] According to another aspect of the present application, the present application also provides an evaluation device for a contribution degree of an agent of a multi-agent system. The evaluation device comprises a processing module. The processing module determines a graph neural network model corresponding to the multi-agent system according to received environment state information, wherein the environment state information at least comprises agent state information, a task demand, and environment information. The processing module trains an action network and the graph neural network model according to the environment state information. The processing module determines action information of a node of the graph neural network model by a trained action network according to the environment state information. The processing module determines a contribution degree of the agent by a trained graph neural network model according to the environment state information and the action information.
[0012] According to some embodiments of the present application, the processing module determines the node model according to the agent state information and the task requirement; the processing module determines the edge structure according to the agent state information, the task requirement and the environment information; and the processing module determines the graph neural network model according to the node model and the edge structure.
[0013] According to another aspect of the present application, the present application further provides a non-volatile computer readable storage medium, having stored thereon a computer program, which, when executed by a processor, enables the evaluation method as described above.
[0014] According to another aspect of the present application, the present application further provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs, which, when executed by the one or more processors, enable the one or more processors to implement the evaluation method as described above.
[0015] According to another aspect of the present application, the present application further provides a computer program product, comprising: a computer program stored on a computer readable storage medium; the computer program comprising program instructions, which, when executed by a computer, cause the computer to perform the evaluation method as described above. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative effort on the basis of these drawings.
[0017] Figure 1 A flowchart of an evaluation method 1000 according to an embodiment of the present application is shown;
[0018] Figure 2 A flowchart of step S110 according to an embodiment of the present application is shown;
[0019] Figure 3 A flowchart of step S120 according to an embodiment of the present application is shown;
[0020] Figure 4 A flowchart of step S121 according to an embodiment of the present application is shown;
[0021] Figure 5 A flowchart of step S122 according to an embodiment of the present application is shown;
[0022] Figure 6 A flowchart of step S140 according to an embodiment of the present application is shown;
[0023] Figure 7 Fig. 1 shows a structural schematic diagram of an evaluation device according to an embodiment of the present application;
[0024] Figure 8 Fig. 2 shows a schematic diagram of an agent action network according to an embodiment of the present application.
[0025] Legend of reference signs:
[0026] evaluation device 20; processing module 21. DETAILED DESCRIPTION
[0027] Example embodiments now will be described more fully hereinafter with reference to the accompanying drawings. Example embodiments, however, can be implemented in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of example embodiments to those skilled in the art. Like reference numerals refer to like elements throughout the several views.
[0028] The described features, structures, or characteristics can be combined in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the disclosure. One skilled in the relevant art will recognize, however, that the technology can be practiced without one or more of the specific details, or with other methods, components, materials, and so forth. In these instances, well-known structures, methods, devices, materials, and so forth have not been described in detail in order to avoid obscuring aspects of the technology.
[0029] Furthermore, the term "comprising" and "including" and their variants are intended to cover both non-exclusive and exclusive inclusion. For example, a process, method, system, product, or apparatus that comprises a list of steps or elements is not necessarily limited to those steps or elements, but can optionally include additional steps or elements not expressly listed or inherent to such process, method, system, product, or apparatus.
[0030] The terms "first", "second", and the like, in the description and in the claims of the present application and in the above description of the drawings merely denote different objects and do not imply a specific order.
[0031] The technical solutions in the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort fall within the scope of the present application.
[0032] The English and its English full name and corresponding Chinese interpretation involved in the present application are as follows:
[0033] RL, Reinforcement Learning, reinforcement learning;
[0034] MARL, Multi-Agent Reinforcement Learning, multi-agent reinforcement learning;
[0035] CTDE, Centralized Training with Decentralized Execution, centralized training and decentralized execution;
[0036] MADDPG, Multi-Agent Deep Deterministic Policy Gradient, multi-agent deep deterministic policy gradient.
[0037] The cooperative task of multi-agent system usually involves complex interaction and dynamic environmental changes. For example, in the rescue search scene, multiple agents need to quickly locate the target in an unknown environment while avoiding interference with each other; in the logistics sorting scene, agents need to dynamically adjust the division of labor according to real-time task requirements to improve overall efficiency.
[0038] Traditional multi-agent cooperation methods have the problems of strategy game and non-stationarity (for example, the strategy game in a multi-agent environment can lead to non-stationarity, that is, the strategy change of other agents will affect the decision-making process of the current agent, leading to unstable training), contradiction between local observation and global cooperation (for example, agents usually can only obtain local observation information and are difficult to fully perceive the global state, which limits their collaboration ability in complex tasks), and fairness limitations of contribution evaluation and reward allocation (for example, in a multi-agent system, how to accurately evaluate the contribution of each agent to the task and allocate rewards or resources accordingly is a long-standing problem. Traditional methods often cannot reflect the actual role of the agent, resulting in efficiency loss.)
[0039] To solve the above problems, the prior art attempts to optimize by the following methods: for example, using centralized control and distributed execution, part of the system uses a centralized controller to coordinate agent behavior, but the inventors of the present application found that this method has poor flexibility in dynamic environments and is difficult to adapt to the cooperation needs of heterogeneous agents (agents with different abilities or task types).
[0040] For example, in the context of reinforcement learning, multi-agent reinforcement learning (MARL) is a subfield that focuses on training multiple agents to collaborate or compete in a shared environment. MARL can be categorized into centralized training and decentralized execution (CTDE) and decentralized training and decentralized execution (DTDE) frameworks. For example, MADDPG (Multi-Agent Deep Deterministic Policy Gradient) allows the Critic network to access global information to optimize the policy, but it still struggles to model complex dependencies between multiple agents.
[0041] For example, in the context of reinforcement learning, multi-agent reinforcement learning (MARL) is a subfield that focuses on training multiple agents to collaborate or compete in a shared environment. MARL can be categorized into centralized training and decentralized execution (CTDE) and decentralized training and decentralized execution (DTDE) frameworks. For example, MADDPG (Multi-Agent Deep Deterministic Policy Gradient) allows the Critic network to access global information to optimize the policy, but it still struggles to model complex dependencies between multiple agents.
[0042] For example, in the context of reinforcement learning, multi-agent reinforcement learning (MARL) is a subfield that focuses on training multiple agents to collaborate or compete in a shared environment. MARL can be categorized into centralized training and decentralized execution (CTDE) and decentralized training and decentralized execution (DTDE) frameworks. For example, MADDPG (Multi-Agent Deep Deterministic Policy Gradient) allows the Critic network to access global information to optimize the policy, but it still struggles to model complex dependencies between multiple agents.
[0043] Graph Neural Networks (GNNs) have shown great potential in modeling complex relational networks. GNN models capture the interaction between agents through graph structures and aggregate local and global information using message passing mechanisms. For example, in social network analysis or traffic prediction, GNNs have been proven to be effective in modeling dynamic topological relationships.
[0044] However, the inventors of the present application have found that the application of GNNs in multi-agent collaborative systems has problems of insufficient dynamic adjustment and lack of effective contribution evaluation. For example, the graph structure of a multi-agent system needs to be dynamically adjusted according to task requirements or environmental changes, which puts higher requirements on the real-time performance of GNNs. Converting the embedding vectors output by GNNs into an interpretable "contribution" indicator and combining it with the task objective function is also a difficult point in current research.
[0045] According to an aspect of the present application, the present application provides a method 1000 for evaluating the contribution of an agent in a multi-agent system. The evaluation method 1000 can be executed by a computer system, which can be a host or server with data processing capabilities.
[0046] Referring to Figure 1 , the evaluation method 1000 can include steps S110-S140.
[0047] In step S110, the computer system determines the graph neural network model corresponding to the multi-agent collaborative system according to the received environmental state information.
[0048] According to an example embodiment, the multi-agent system can be a system capable of accomplishing high-difficulty tasks in a dynamic and uncertain environment through the cooperative work of multiple agents. The agents can be heterogeneous agents such as unmanned aerial vehicles, unmanned vehicles, and carrying robots, and the multi-agent system can be applicable to complex scenarios such as rescue search and logistics sorting.
[0049] The environment state information can be the state of the multi-agent system and the information of the environment in which the multi-agent system is located. The environment state information at least includes agent state information, task requirements, and environment information.
[0050] The agent state information can be information describing the state of all agents, such as speed, acceleration, power, and the like.
[0051] The task requirements can be the state of the target task of the multi-agent system, such as the task type, the task priority, the task deadline, the task completion progress, and the required resources (load, power, sensor type) of the task.
[0052] The environment information can be the information of the environment in which the multi-agent system is located. For example, the environment information can include the observation position, the position and size of the obstacle, the terrain feature, the communication condition, and the position and motion state of the dynamic object.
[0053] The environment state information can constitute a global state matrix, and each row of the global state matrix corresponds to an observation position, a speed, a power, and a task information of an agent.
[0054] The multi-agent system can be abstracted as a graph structure. The agents and tasks and the like can be nodes of the graph structure, and the interaction relationship between the agents or between the agents and the tasks can be edges of the graph structure.
[0055] The graph neural network model can be a graph structure of the multi-agent system constructed according to the environment state information.
[0056] The computer system can construct the graph neural network model according to the nodes of the graph structure and the edges of the graph structure.
[0057] In step S120, the computer system trains the action network and the graph neural network model according to the environment state information.
[0058] According to an example embodiment, the action network can be an Actor network. The computer system can configure each agent with an Actor network and a Critic network based on the MADDPG algorithm. The computer system can train the trained action network through the MADDPG algorithm. The training phase adopts a centralized training mode, allowing the Critic network to access global information, i.e., all environment state information. The execution phase is decentralized execution, and each agent independently determines the training action information of the agent according to local observation and policy. The local observation can be the environment state information of the agent itself.
[0059] The Critic network takes the training action information of all agents and the global state as input, and outputs the value evaluation of the corresponding agent, solving the non-stationary problem in the multi-agent environment. Each Actor network adopts a policy gradient update mechanism, which optimizes the decision policy by maximizing the expected return while ensuring continuous action output.
[0060] The computer system can formalize the multi-agent task cooperation problem as a Markov game with constraints. Each agent interacts with other agents based on local observation during the learning process and optimizes the policy by maximizing its cumulative revenue. In a collaborative scenario, there can be cooperation or competition between agents, for example, when performing a search task, reasonable avoidance and cooperation are required. The centralized Critic structure of MADDPG can coordinate the strategies of each agent and promote teamwork to achieve the task goal. The computer system can update the environment state information according to the value evaluation parameter output by the Critic network.
[0061] The computer system can determine an experience trajectory according to the environment state information, the training action information, and the updated environment state information. The computer system can determine the global task performance according to the experience trajectory. The computer system can train the graph neural network model according to the global task performance.
[0062] In step S130, the computer system determines the action information of the nodes of the graph neural network model according to the environment state information through the trained action network.
[0063] According to an example embodiment, the action information of the nodes can be the execution action information of the agent. The computer system can determine the action information of each node (i.e., agent) according to the environment state information through the trained Actor network.
[0064] In step S140, the computer system determines the contribution degree of the agent according to the environment state information and the action information through the trained graph neural network model.
[0065] According to an example embodiment, the contribution degree can be a quantitative indicator of the positive influence of the agent on the completion of the target task.
[0066] For example, the computer system can map the environment state information and the action information to a high-dimensional embedding space through a first layer of the graph neural network to obtain a first feature vector of the node; the computer system determines neighbor information of the node according to the environment state information; the computer system determines an updated feature vector of the node according to the first feature vector and the neighbor information through a subsequent layer of the trained graph neural network; and the computer system determines the contribution degree of the agent according to the updated feature vector through a linear layer of the trained graph neural network.
[0067] Through the above embodiments, the technical scheme of the present application determines the graph neural network model corresponding to the multi-agent collaborative system through the environment state information. The technical scheme of the present application trains the action network and the graph neural network model through the environment state information. The technical scheme of the present application determines the action information of the node of the graph neural network model through the trained action network through the environment state information. The technical scheme of the present application determines the contribution degree of the agent through the trained graph neural network model through the environment state information and the action information.
[0068] The evaluation method provided by the present application combines MADDPG reinforcement learning and graph neural network organically, realizes strategy game and coordination among agents through MADDPG, solves the non-stationary problem in multi-agent game through centralized Critic, dynamically evaluates the contribution degree of each agent from the cluster level through the graph neural network, and enhances the perception ability of the system to the overall structure. The present application not only supports the cooperation of heterogeneous agents, but also improves the adaptability of the multi-agent system to dynamic task changes.
[0069] Optionally, referring to Figure 2 , step S110 can include steps S111-S113.
[0070] In step S111, the computer system determines a node model according to the agent state information and the task demand.
[0071] According to an example embodiment, the node model can be a node set of a graph structure. The computer system can take entities such as agents and tasks as nodes of the graph structure, and all the nodes form a node model. Each node includes node features, which can be represented as a feature vector according to the agent state information and the task demand.
[0072] In step S112, the computer system determines an edge structure according to the agent state information, the task demand, and the environment information.
[0073] According to an example embodiment, the edge structure can be a set of edges of the graph structure. The computer system can take the interaction relationship between the agents or between the agents and the tasks as the edges of the graph structure. The entire set of edges forms the edge structure. Each edge includes edge features, which can be represented as a feature vector according to the agent state information, the task requirements, and the environment information.
[0074] In step S113, the computer system determines the graph neural network model according to the node model and the edge structure.
[0075] According to an example embodiment, the computer system can construct the graph neural network model of the multi-agent system according to the node model and the edge structure.
[0076] For example, the graph neural network model can include a node set V and an edge set E.
[0077] The node set V represents all agents, and the edge set E represents effective interaction relationships such as communication links, distance proximity, task cooperation dependencies, and the like between agents. The graph structure of the graph neural network model can be stored in an adjacency matrix or an edge list. The graph structure can be used for the input of the graph neural network model. The agent action space is represented by a continuous vector, and each agent has an independent set of action variables A, such as moving speed and direction, pick-and-place actions, and the like. i .
[0078] Optionally, referring to Figure 3 , step S120 can include step S121 and step S122.
[0079] Step S121 is a step of executing a training action network by the computer system. The computer system executes step S121 at least once.
[0080] Referring to Figure 4 , step S121 can include steps S1211-S1214.
[0081] In step S1211, the computer system determines the training action information of the agent by the action network according to the environment state information.
[0082] According to an example embodiment, the training action information can be the execution action information of the agent in the training action network process.
[0083] The computer system can input the environment state information to the action network, and the action network outputs the training action information.
[0084] In step S1212, the computer system determines the value evaluation parameter of the agent by the value network according to the training action information, the environment state information, and a preset reward function.
[0085] According to an example embodiment, the preset reward function can be a quantitative indicator for the agent to complete the training action information. The value network can be a Critic network.
[0086] Optionally, the preset reward function can include a preset cooperation function and a preset game competition function. The preset cooperation function can be a preset reward function for a cooperative type task for the target task. The preset game competition function can be a preset reward function for a game competition type task for the target task. The value evaluation parameter can be a parameter for evaluating the value of the agent completing the target task.
[0087] The computer system can update the state of the agent according to the received training action information, and output the immediate reward of each agent.
[0088] For example, the preset cooperation function can include a global corresponding joint preset cooperation function and an individual corresponding individual preset cooperation function. In a multi-agent cooperation scenario, the goals of multiple agents are consistent, and generally maximize a global reward, and the joint preset cooperation function can be:
[0089]
[0090] wherein, is a set of agents. The training action information of each agent i is a i , the environment state information is s, and the overall training action information is a=(a1,...,a N ). R(s,a) is a global reward function, and the global reward function is a single and unified reward signal given by the environment state-action sequence.
[0091] The joint preset cooperation function can be that all agents receive the same reward, encouraging them to complete the task together.
[0092] The individual preset cooperation function can weight the reward of the individual agent. The individual preset cooperation function can be:
[0093]
[0094] The individual preset cooperation function can be applicable to cooperative tasks with different individual capabilities or responsibilities.
[0095] For another example, the preset game competition function can include a preset zero-sum game function and a preset partial competition function. In a game scenario, the goals of the agents are opposite, and the reward function design should reflect the zero-sum or partial zero-sum characteristics.
[0096] The preset zero-sum game function can be:
[0097] r i (t)=-rj (t);
[0098] r i (t) is a preset zero-sum game function of the agent i, r j (t) is a preset zero-sum game function of the agent j, and the preset zero-sum game function is mainly used in confrontation scenes such as Go, competitive games, and escape scenes.
[0099] The preset partial competition function can be:
[0100] r i (t) = R i (s t ,a t ) - λ·R j (s t ,a t ), λ ∈ [0, 1];
[0101] Wherein, R i (s t ,a t ) is the own benefit of the agent i; R j (s t ,a t ) is the benefit of the opponent (agent j) of the agent i; λ is a weight, which can control the degree of "harm others to benefit oneself".
[0102] In the case of λ = 1 and R i = R j , the preset partial competition function degenerates into the preset zero-sum game function; and in the case of λ = 1, the agent i is a pure self-interest agent.
[0103] According to an example embodiment, the computer system can input the training action information and the environment state information into the value network, and the value network outputs the value evaluation parameter.
[0104] In step S1213, the computer system updates the environment state information according to the value evaluation parameter.
[0105] According to an example embodiment, the environment state information is updated after the agent executes the training action information and has an impact on the environment state information. The computer system can update the environment state information according to the value evaluation parameter.
[0106] In step S1214, the computer system trains the action network according to the updated environment state information.
[0107] According to an example embodiment, the Actor network can adopt a policy gradient update mechanism to optimize the decision strategy by maximizing the expected return while ensuring continuous action output.
[0108] The computer system can input the updated environment state information into the action network, thereby performing the next round of training.
[0109] The computer system can also set a replay buffer structure ReplayBuffer to save historical interaction records of multi-agent states, actions, rewards, and next states for offline training.
[0110] Step S122 is for the computer system to perform the step of training the graph neural network model.
[0111] Referring to Figure 5 , step S122 can include steps S1221-S1223.
[0112] In step S1221, the computer system determines an experience trajectory according to the environment state information, the training action information, and the updated environment state information.
[0113] According to an example embodiment, the experience trajectory can be a state and action sequence of the agent. The experience trajectory can include the environment state information, the training action information, and the updated environment state information, and the experience trajectory can also output an immediate reward according to the training action information of the agent.
[0114] In step S1222, the computer system determines a global task performance according to the experience trajectory.
[0115] According to an example embodiment, the global task performance can be a whole contribution parameter of the experience trajectory of the agent to the target task.
[0116] The computer system can output the global task performance according to the experience trajectory through the graph neural network.
[0117] In step S1223, the computer system trains the graph neural network model according to the global task performance.
[0118] According to an example embodiment, the computer system trains the graph neural network model with the global task performance as a label to adjust the graph neural network model parameters, so that the contribution degree of each agent output by the graph neural network model is more consistent with the actual contribution. The computer system can also use the contribution degree as an auxiliary reward or an input of the Critic to promote the strategy to learn in the direction of global optimality.
[0119] The graph neural network model parameters can be trainable variables that are automatically learned and updated through a backpropagation algorithm during training. The graph neural network model parameters can include node feature updates, weight matrices, bias vectors, attention weights, gating parameters, residual mappings, and output layer parameters, etc.
[0120] For example, the computer system can determine the node feature update of each layer of the graph neural network model according to the following formula:
[0121]
[0122] in, Let L be the weight matrix of the l-th layer of the graph neural network model. Let d be the real number field and W be the number of d. (l) The dimension. b (l) Let σ be the bias vector, and σ(·) be the nonlinear activation function. This represents the set of adjacent nodes of node v. This is the node feature update for each layer of the graph neural network model, that is, the updated state of the current node i (i.e., the current agent i) at layer l+1. The hidden state after updating the current node i (i.e. the current agent i).
[0123] bias vector b (l) Bias terms can be introduced into each layer of a graph neural network model to enhance its expressive power. The computer system can determine the bias vector according to the following formula:
[0124]
[0125] In graph attention networks, the importance of neighboring nodes is modeled through the attention mechanism, and the computer system can determine the attention weights according to the following formula:
[0126]
[0127] Where a is a learnable attention vector; ||·|| denotes the feature concatenation operation. LeakyReLU is an activation function. W is a trainable parameter matrix; The hidden state of the current node i (i.e., the current agent i); This represents the hidden state of the neighboring nodes of the current node i (i.e., the current agent i).
[0128] By introducing a gating mechanism into a gated graph neural network model to control the flow of information, the computer system can determine the gating parameters according to the following formula:
[0129]
[0130] Where ⊙ represents the Hadamard product, z v To update the door, is The degree of retention; m v for Aggregated data; h v r represents the hidden state of the current node i (i.e., the current agent i); vTo reset the door, the current node i (i.e., the current agent i) updates the hidden state before the update to the hidden state after the update. The current node i (i.e., the current agent i) candidate hidden state, that is, the hidden state that integrates the neighbor information of the current node i; h' v The current node i (i.e., the current agent i) updates the hidden state before the update to the hidden state after the update. v The current node i (i.e., the current agent i) updates the hidden state before the update to the hidden state after the update. z The current node i (i.e., the current agent i) updates the hidden state before the update to the hidden state after the update. r The current node i (i.e., the current agent i) updates the hidden state before the update to the hidden state after the update. h The current node i (i.e., the current agent i) updates the hidden state before the update to the hidden state after the update. z The current node i (i.e., the current agent i) updates the hidden state before the update to the hidden state after the update. r The current node i (i.e., the current agent i) updates the hidden state before the update to the hidden state after the update. h The current node i (i.e., the current agent i) updates the hidden state before the update to the hidden state after the update.
[0131]
[0132] Wherein, Indicates a certain graph convolution or message passing function, that is, the residual mapping parameter Θ (l) The parameter set of the layer.
[0133] The output layer is usually a fully connected classification or regression layer, and the computer system can determine the output layer formula according to the following formula:
[0134]
[0135] Wherein, W out is the weight matrix of the output layer, b out is the bias vector of the output layer, L represents the last layer, is the output layer parameter.
[0136] For example, see Figure 8 The agent action network can be the decision probability distribution of the agent, and the agent action network integrates the graph neural network model and the Actor-Critic network structure.
[0137] The computer system can initialize the environment. The computer system randomly arranges the agents and target / task points according to the task scene. Assign an initial role or ability parameter to each agent.
[0138] The computer system can receive the environment state data in real time, and in each training round, the multi-agent executes the strategy in the environment, and each agent generates training action information from the local observation and the information passed by the GNN (environment state data). The multi-agent system receives all agent actions, updates the environment state information and outputs the instant reward of each agent. The updated environment state information and instant reward are stored in the replay buffer.
[0139] The computer system can update the Critic network parameters and the Actor network parameters from the replay buffer every fixed number of steps. For example, the computer system can update the Critic network parameters by iteratively updating the Critic network by minimizing the joint temporal difference error of the multi-agent. The computer system can improve the policy performance of the Actor network by policy gradient. The updating of the Critic network parameters and the Actor network parameters can use a target network and a soft update mechanism to ensure learning stability.
[0140] The computer system can train the graph neural network model synchronously or alternately. The computer system can calculate the global task performance using the experience trajectory collected in the current round. The GNN is trained and the GNN parameters are adjusted using the global task performance as a label, so that the contribution of each agent output is more consistent with the actual contribution. The computer system can also use the contribution as an auxiliary reward or input to the Critic to promote the learning of the policy of the Actor network in the direction of global optimality.
[0141] The computer system can perform steps S121 and S122 multiple times until the graph neural network model converges (for example, the loss function of the graph neural network model is lower than a threshold value). The training process uses a centralized global Critic perspective and graph structure information fusion, so that each agent can consider the overall interests while maintaining decentralized execution capability, and effectively collaborate to achieve complex tasks.
[0142] After training, only the Actor policy and the graph neural network model are retained. In actual execution, each agent uses the trained Actor network to select action information according to the current environment state information and the graph neural network model embedding, and continuously updates the local contribution evaluation through the graph neural network model to guide the task selection or fine-tuning of the collaboration strategy.
[0143] Through the above embodiments, the technical scheme of the present application can set an Actor network for each agent i, and each agent considers the influence of the behavior of other agents when updating the policy, thereby alleviating the high variance and instability problems of traditional RL in a multi-agent environment.
[0144] The technical scheme of the present application can use the centralized Critic network of MADDPG, and each agent considers global information when training the policy, which significantly improves the collaboration efficiency. The graph neural network model can enhance team awareness and fuse the correlation information between multi-agents into the decision-making, so that the collaborative decision-making is more accurate and consistent. Experiments show that the information fusion mechanism combined with GNN can significantly improve the multi-agent cooperation effect and convergence speed.
[0145] The technical solution of the application uses the task completion rate and the resource utilization efficiency global performance of the agent after completing the task as the supervision target, and trains the graph neural network model by minimizing the difference between the contribution degree prediction and the actual performance.
[0146] The graph neural network model can be trained simultaneously with the main RL cycle, so that the contribution degree evaluation can assist in strategy optimization. If necessary, the graph neural network can also combine an attention mechanism to dynamically adjust the edge weight according to the task urgency or the scarcity of agents, reflecting the relative importance of different agents in the current task.
[0147] Optionally, referring to Figure 6 , step S140 can include steps S141-S144.
[0148] In step S141, the computer system maps the environment state information and the action information to a high-dimensional embedding space through the first layer of the trained graph neural network to obtain a first feature vector of the node.
[0149] According to an example embodiment, the first feature vector can be the feature vector of the node after the environment state information and the action information are mapped to the high-dimensional embedding space.
[0150] The graph neural network model can use multi-layer convolution or an attention mechanism for feature aggregation.
[0151] For example, the computer system can take the current position of the agent, the state features, and the task information as the input of the first layer of the graph neural network. The first layer of the graph neural network maps the input features of each node, i.e., the local observation o i and the task state, to a high-dimensional embedding space to obtain a first feature vector.
[0152] In step S142, the computer system determines the neighbor information of the node according to the environment state information.
[0153] According to an example embodiment, the neighbor information can be the state information of the neighboring nodes of a node. For example, the node information can include the state, position, task demand, and environment information of the neighboring nodes.
[0154] In step S143, the computer system determines the aggregated feature vector of the node according to the first feature vector and the neighbor information through the subsequent layers of the trained graph neural network.
[0155] According to an example embodiment, the aggregated feature vector can be the node feature after multi-layer convolution or attention mechanism aggregation of the node.
[0156] For example, the subsequent layers of the graph neural network can update the node through the neighbor information and the first feature vector aggregation.
[0157] In step S144, the computer system determines the contribution degree of the agent according to the updated feature vector by a linear layer of the trained graph neural network.
[0158] According to an example embodiment, the computer system can convert the feature vector into a score c of the contribution degree by the linear layer i .
[0159] Through the above-mentioned embodiments, the technical scheme of the present application utilizes the graph neural network to model the cluster structure and the contribution of the agent, and realizes the fine contribution degree evaluation. Compared with the traditional simple accumulation or equal distribution of rewards, the technical scheme of the present application can allocate more fair and reasonable benefits according to the actual role of the agent in the multi-agent system, and avoid the efficiency loss caused by one-sided rewards. The accurate evaluation result helps the system to effectively motivate and adjust different agents, and further improves the overall performance of the multi-agent.
[0160] According to another aspect of the present application, the present application also provides an evaluation device 20 of the contribution degree of an agent of a multi-agent system. The evaluation device 20 can execute the above-mentioned evaluation method 1000. Referring to Figure 7 , the evaluation device 20 comprises a processing module 21.
[0161] According to an example embodiment, the processing module 21 determines a graph neural network model corresponding to the multi-agent system according to the received environment state information, wherein the environment state information at least includes agent state information, task demand and environment information.
[0162] The processing module 21 trains the action network and the graph neural network model according to the environment state information.
[0163] The processing module 21 determines the action information of the node of the graph neural network model according to the environment state information by the trained action network.
[0164] The processing module 21 determines the contribution degree of the agent according to the environment state information and the action information by the trained graph neural network model.
[0165] The environment state information, the graph neural network model, the action network, the action information and the contribution degree have been described in the above-mentioned evaluation method 1000, and thus will not be described again.
[0166] Through the above embodiment, the technical scheme of the application determines the graph neural network model corresponding to the multi-agent collaborative system through the environment state information. The technical scheme of the application trains the action network and the graph neural network model through the environment state information. The technical scheme of the application determines the action information of the node of the graph neural network model through the trained action network through the environment state information. The technical scheme of the application determines the contribution degree of the agent through the trained graph neural network model through the environment state information and the action information.
[0167] The application combines MADDPG reinforcement learning and graph neural network organically, realizes the strategy game and coordination among agents through MADDPG, solves the non-stationary problem in multi-agent game through centralized Critic, simultaneously evaluates the contribution degree of each agent from the cluster level through the graph neural network, and enhances the perception ability of the system to the overall structure. The application supports the cooperation of heterogeneous agents, and improves the adaptability of the multi-agent system to dynamic task changes.
[0168] Optionally, the processing module 21 determines the node model according to the agent state information and the task demand.
[0169] The processing module 21 determines the edge structure according to the agent state information, the task demand and the environment information.
[0170] The processing module 21 determines the graph neural network model according to the node model and the edge structure.
[0171] The node model and the edge structure have been described in the evaluation method 1000 in the above, and thus will not be described again.
[0172] Optionally, the processing module 21 executes the training action network step at least once.
[0173] The processing module 21 determines the training action information of the agent through the action network according to the environment state information.
[0174] The processing module 21 determines the value evaluation parameter of the agent through the value network according to the training action information, the environment state information and the preset reward function.
[0175] The processing module 21 updates the environment state information according to the value evaluation parameter.
[0176] The processing module 21 trains the action network according to the updated environment state information.
[0177] The processing module 21 executes the training graph neural network model step at least once.
[0178] The processing module 21 determines the experience trajectory according to the environment state information, the training action information and the updated environment state information.
[0179] The processing module 21 determines the experience trajectory according to the experience trajectory.
[0180] The processing module 21 trains the graph neural network model according to the global task performance.
[0181] The training action information, the value network, the value evaluation parameter, the experience trajectory and the evaluation method 1000 have been described above, and thus will not be described again.
[0182] Optionally, the processing module 21 maps the environment state information and the action information to a high-dimensional embedding space by the first layer of the trained graph neural network to obtain the first feature vector of the node.
[0183] The processing module 21 determines the neighbor information of the node according to the environment state information.
[0184] The processing module 21 determines the updated feature vector of the node according to the first feature vector and the neighbor information by the subsequent layer of the trained graph neural network.
[0185] The processing module 21 determines the contribution degree of the agent according to the updated feature vector by the linear layer of the trained graph neural network.
[0186] The first feature vector, the neighbor information and the updated feature vector have been described above in the evaluation method 1000, and thus will not be described again.
[0187] According to another aspect of the present application, the present application also provides a non-volatile computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the evaluation method as described above.
[0188] According to another aspect of the present application, the present application also provides an electronic device, which includes one or more processors, and a storage device configured to store one or more programs, and the one or more programs, when executed by the one or more processors, enable the one or more processors to implement the evaluation method as described above.
[0189] According to another aspect of the present application, the present application also provides a computer program product, which includes a computer program stored on a computer readable storage medium, and the computer program includes program instructions, and the program instructions, when executed by a computer, enable the computer to execute the evaluation method as described above.
[0190] Finally, it should be noted that the above only describes the preferred embodiments of the present application and is not intended to limit the present application. Although the present application is described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions of the foregoing embodiments or make equivalent replacements to some technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for evaluating the contribution of an agent of a multi-agent system, characterized in that, The evaluation method comprises: According to the received environment state information, determine the corresponding graph neural network model of the multi-agent system, wherein the environment state information at least includes agent state information, task demand and environment information; According to the environment state information, train the action network and the graph neural network model; According to the environment state information, determine the action information of the node of the graph neural network model through the trained action network; According to the environment state information and the action information, determine the contribution degree of the agent through the trained graph neural network model.
2. The evaluation method according to claim 1, characterized in that According to the received environment state information, determine the corresponding graph neural network model of the multi-agent system, comprising: According to the agent state information and the task demand, determine the node model; According to the agent state information, the task demand and the environment information, determine the edge structure; According to the node model and the edge structure, determine the graph neural network model.
3. The evaluation method according to claim 1, characterized in that According to the environment state information, train the action network and the graph neural network model, comprising: At least one training action network step, comprising: According to the environment state information, determine the training action information of the agent through the action network; According to the training action information, the environment state information and the preset reward function, determine the value evaluation parameter of the agent through the value network; According to the value evaluation parameter, update the environment state information; According to the updated environment state information, train the action network; At least one training graph neural network model step, comprising: According to the environment state information, the training action information and the updated environment state information, determine the experience trajectory; According to the experience trajectory, determine the global task performance; According to the global task performance, train the graph neural network model.
4. The evaluation method according to claim 3, characterized in that The preset reward function includes a preset cooperation function and a preset game competition function.
5. The evaluation method according to claim 1, characterized in that According to the environment state information and the action information, determine the contribution degree of the agent through the trained graph neural network model, comprising: Map the environment state information and the action information to a high-dimensional embedding space through the first layer of the trained graph neural network to obtain the first feature vector of the node; According to the environment state information, determine the neighbor information of the node; According to the first feature vector and the neighbor information, determine the updated feature vector of the node through the subsequent layer of the trained graph neural network; According to the updated feature vector, determine the contribution degree of the agent through the linear layer of the trained graph neural network.
6. A device for evaluating the contribution of agents in a multi-agent system, characterized in that, The evaluation device comprises: A processing module, according to the received environment state information, determine the corresponding graph neural network model of the multi-agent system, wherein the environment state information at least includes agent state information, task demand and environment information; The processing module trains the action network and the graph neural network model according to the environment state information; The processing module determines the action information of the node of the graph neural network model according to the environment state information by using the trained action network. The processing module determines the contribution degree of the agent according to the environment state information and the action information by using the trained graph neural network model.
7. The evaluation device according to claim 6, characterized in that The processing module determines the node model according to the agent state information and the task requirement. The processing module determines the edge structure according to the agent state information, the task requirement and the environment information. The processing module determines the graph neural network model according to the node model and the edge structure.
8. A non-transitory computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by a processor to implement the evaluation method of any one of claims 1-5.
9. An electronic device, comprising: Comprise: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, so that the one or more processors implement the evaluation method of any one of claims 1-5.
10. A computer program product, characterised in that, The computer program comprises program instructions stored on a computer readable storage medium, when the program instructions are executed by a computer, the computer executes the evaluation method of any one of claims 1-5.
Citation Information
Patent Citations
Multi-agent information fusion method and device, electronic equipment and readable storage medium
CN114139637A
Method and device for evaluating contribution rate of element system of network information system
CN114611990A
Near-end strategy optimization method based on graph convolutional neural network
CN115983373A
Multi-agent collaborative learning heterogeneous convergence network resource scheduling method and device
CN116647564A
Cooperative multi-unmanned-system contribution evaluation and decision-making method, product, medium and equipment
CN118551825A
Cited By
Factory management and control system based on artificial intelligence and big data
CN121581515A
Credit distribution multi-agent cooperative training method and device for land battle heterogeneous marshalling
CN121960557A