Multi-agent cooperation method based on heterogeneous investment reinforcement learning
By introducing heterogeneous investment in multi-agent reinforcement learning, the selfish individual phenomenon is solved, cooperation efficiency and system benefits are improved, and a higher level of agent relationship feature extraction and cooperation incentives are achieved.
Patent Information
- Application Number
- CN202510308874.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
AI Technical Summary
There is a selfish individual phenomenon in multi-agent reinforcement learning, which makes cooperative behavior difficult to sustain and overall cooperation inefficient, especially in scenarios of limited resource allocation and group task execution.
The N-person trust game mechanism for heterogeneous investment is introduced, modeling agent relationships through hypergraphs, designing reward structures, using heterogeneous investment to encourage cooperation of trustworthy individuals, and using multi-layer hypergraph convolutional neural network to extract deep-level feature information, and combining trust game and heterogeneous investment mechanism to optimize agent decision-making.
It has improved the level of cooperation in multi-agent systems, promoted long-term cooperative behavior, and enhanced the stability and overall benefits of the group.
Smart Images

Figure CN120258080A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multi-agent cooperation in multi-agent reinforcement learning, and specifically relates to a multi-agent cooperation method based on heterogeneous investment reinforcement learning, which is a reinforcement learning method that uses heterogeneous investment trust games to improve the cooperation level of cooperative multi-agents. Background Art
[0002] Reinforcement learning is one of the core technologies in the field of machine learning, aiming to guide the behavior adjustment and policy update of an agent through continuous interaction between the agent and the environment, based on the reward signal feedback from the environment, and ultimately achieving the goal of maximum cumulative reward or completing a specific task. Under the reinforcement learning framework, the agent continuously optimizes its decision-making through exploration and learning, and is applicable to various complex dynamic environments, especially achieving remarkable results in fields such as games, robot control, and autonomous driving.
[0003] However, traditional reinforcement learning usually focuses on the interaction between a single agent and the environment. In practical problems, many tasks require multiple agents to cooperate or compete to complete together. For this reason, the concept of multi-agent reinforcement learning (MARL) emerged. This technology extends reinforcement learning to a multi-agent system (MAS), enabling multiple agents to achieve complex group behaviors through interaction and learning in a shared environment. In recent years, the research on MARL has been continuously deepened, and many classic algorithms have emerged, such as MADDPG, VDN, ROMA, and IRPI. These algorithms have their own characteristics and are widely used in UAV formations, group decision-making, game theory, and multi-robot systems, fully demonstrating the importance and practical value of multi-agent reinforcement learning in complex group cooperation.
[0004] Although MARL technology can achieve group cooperation by training agents, since reinforcement learning is centered around reward-driven, agents usually focus more on their own immediate interests when making decisions, while ignoring long-term collective interests. Such agents are called selfish individuals. When there are a large number of "selfish individuals" in the system, the cooperation behavior between agents is difficult to sustain, resulting in low overall cooperation efficiency and ultimately weakening the long-term benefits of the group. This problem is relatively common in multi-agent reinforcement learning, especially prominent in scenarios involving limited resource allocation, group task execution, etc. Therefore, how to effectively motivate agents to achieve long-term cooperation behavior has become one of the core challenges in MARL research. To solve this difficult problem, researchers have proposed various improvement schemes, and among them, introducing the trust mechanism is an effective and practical strategy.
[0005] Trust is one of the key factors maintaining cooperation in human society. The establishment of trust can effectively promote cooperation among individuals, reduce opportunistic behaviors, and enhance group stability. In recent years, researchers have conducted in-depth explorations on trust mechanisms and their applications in multi-agent systems. For example, Granatyr et al. proposed a trust modeling method applicable to MAS in "Trust and reputation models formultiagent systems", and Haoran et al. designed a trust model for multi-agent systems in "Multi-AgentTrust Evaluation Model based on Reinforcement Learning", and by introducing a reputation system to evaluate the impact of agent behavior on overall cooperation, the decision-making weights of agents are dynamically adjusted. These studies indicate that trust models can effectively improve the cooperation level in multi-agent reinforcement learning and alleviate the negative effects brought by individual selfishness.
[0006] In the theoretical framework of trust research, the Trust Game is one of the common paradigms and has become a popular topic in academia in the past decade. The trust game usually involves two types of roles: investors and the invested, where the investor decides whether to make an investment, and the invested decides whether to return to the investor. To extend the trust game to a more complex group environment, Abbass et al. proposed the N-player Trust Game (NTG) in "The N-player trust game and its replicator dynamics", which can characterize the trust decision-making behaviors of multiple parties. Based on the combination of reinforcement learning and the trust game, Zheng et al. studied how reinforcement learning promotes the accumulation of trust and the evolution of cooperation among participants through repeated trust games in "Decoding trust:Areinforcementlearning perspective". Li et al. proposed a dynamic trust model in "Dynamic trust game model between venture capitalistsand entrepreneurs based on reinforcement learning theory" and further extended it to a multi-agent environment, constructing a multi-extreme dynamic trust game model to systematically study the dynamic evolution process of trust among agents in a complex social environment.
[0007] Although many achievements have been made in the combination of trust games and multi-agent reinforcement learning, most existing studies adopt the homogeneous investment hypothesis, that is, the investment amount of investors in all investees is fixed. However, in the real society, the decisions of investors are often heterogeneous, that is, when facing different investees, the investment amount of investors usually adjusts dynamically. Heterogeneous investment reflects the behavioral characteristics of investors adjusting the investment scale according to the credibility of the investees, and can more accurately simulate the diverse investment decisions caused by differences in trust levels in reality. This mechanism is of great significance for motivating cooperation and improving group benefits. Therefore, introducing the concept of heterogeneous investment can not only more realistically simulate the social trust environment, but also provide new ideas for solving the cooperation dilemma in multi-agent reinforcement learning. This kind of heterogeneous investment has been studied more in the public goods game and has been proven to significantly improve the level of group cooperation. Cao, Wu-Jie, and Wang et al. respectively found through systematic experimental analysis in "The evolutionary public goods game on scale-free networks with heterogeneous investment", "Role of investment heterogeneity in the cooperation on spatial public goods game", and "Heterogeneous investments promote cooperation in evolutionary public goods games" that heterogeneous investment strategies can effectively motivate cooperation among individuals and promote the maximization of group benefits. However, in the cross-field of trust games and multi-agent reinforcement learning, there are few studies and related work on heterogeneous investment. Therefore, it is of great academic value and practical significance to deeply explore how heterogeneous investment combines with multi-agent reinforcement learning under the framework of trust games and study its impact on the cooperation behavior of agents and system performance. Summary of the Invention
[0008] In order to solve the problem of selfish individuals often appearing in multi-agent reinforcement learning and improve the level of cooperation in multi-agent systems, the present invention attempts to add N-person trust games to design the reward structure of agents. At the same time, in order to further promote the generation of cooperation, the present invention uses N-person trust games with heterogeneous investments to encourage the emergence of trustworthy cooperative individuals and uses hypergraphs to model the relationships between agents to extract more eigenvalues. This method uses hypergraphs to more reasonably describe the relationships between agents, while using the characteristics of heterogeneous investment N-person trust games to give more benefits to trusted individuals, encouraging agents to cooperate to become trustworthy individuals to improve the level of cooperation in multi-agent systems. Compared with ordinary MARL algorithms, the setting of experiments in unreasonable uniform mixed groups is abandoned, and real-valued hypergraphs are used to model the relationships between agents to extract deeper feature information. At the same time, heterogeneous investments are used instead of homogeneous investments to avoid the situation where selfish individuals can still obtain more rewards when they choose to betray, thereby improving the level of cooperation in multi-agent systems.
[0009] The technical solution of the present invention is as follows:
[0010] A multi-agent cooperation method based on heterogeneous investment reinforcement learning includes the following steps:
[0011] Step 1: Initialize an experience pool. The interactive experience includes (s, o, τ, a, r, s′, o′), where s is the current environment information and s′ is the next environment information; a is the joint action, which is composed of the action a of each agent u. u Composition; r is the agent reward. Input the number of agents n, the predetermined number of hyperedges m, and initialize the agent network.
[0012] Step 2: Each agent receives the environment state s t , select action a according to the current strategy π u , according to the observed value o u And historical action recordsτ u Generate the action-value function Q for each agent u u Values, and based on the observed values, form the adjacency matrix M of the hypergraph H of the agent relationship H Among them, the observation value o u is the characteristics and environment information of multiple agents. The generation process of the hypergraph matrix is as follows:
[0013] The joint observations of the agent are processed by the activation functions ReLU and Linear functions commonly used in two neural networks to generate an H1 matrix. Then, the mean avg(H1) of the H1 matrix is obtained and multiplied by an n-order identity matrix to obtain the H2 matrix. The adjacency matrix M of the hypergraph H of concatenating H1 and H2 is obtained. H .
[0014] H1 = ReLU(Linear(Z)) (1)
[0015] H2 = I n × avg(H1) (2)
[0016] M H = H1 | H2 (3)
[0017] Where Z is the joint observation value, which is composed of the observation value o of each agent u u , avg is the average operation, and I n is the n-order identity matrix.
[0018] Step 3: Agent action a u On the one hand, it acts on the environment and generates an environmental reward r according to the environmental rules env , on the other hand, the agent conducts an N-person trust game of heterogeneous investment and calculates its reward r according to its mechanism tg . The two rewards together constitute the agent's reward r, where the reward r = r env + r tg .
[0019] Step 4: The vector Q composed of the Q of each agent u and M generated in Step 1 H First, pass through a multi-layer hypergraph convolutional layer (HGCN) to generate Q′. Q′ is the preliminary joint action value after integrating the information of each agent obtained through the convolution operation.
[0020] Q′ = HGCN(HGCN(Q, abs(W1)), W2) (4)
[0021] Where W i , i = 1, 2 represents the hyperedge weight matrices of different layers, abs is the absolute value operation. In order to capture local information and aggregate global information, two layers of HGCN are used to extract higher-level features. The HGCN operation is as follows:
[0022]
[0023] Where x (l) is the input signal, x (l+1) is the output signal, D(H), W (H) and B(H) are the vertex degree matrix, hyperedge weight matrix and hyperedge degree matrix of the hypergraph H respectively, and P is the weight matrix between the l-th layer and the l + 1-th layer, which is set as the n-dimensional identity matrix. The input signal here is Q, and Q′ is obtained through two layers of HGCN operations.
[0024] Step 5: Q′ passes through the elu and sum operations of the global state module to generate the joint action value function Q that integrates the global state informationtot Specific generation process is as follows:
[0025] Q″ = elu(abs(MLP(s t )) ⊙ Q′ + MLP(s t )) (6)
[0026] Q tot = sum(abs(MLP(s t )) ⊙ Q″)) + MLP(s t ) (7)
[0027] Where MLP is a multi - layer perceptron, a common artificial neural network, used to map a set of input vectors to a set of output vectors; ⊙ is the symbol for element - wise multiplication.
[0028] Step 6: The agent reward r and the joint action value function Q tot Participate in the calculation of the error function to optimize the algorithm model and narrow the gap between the predicted value and the actual value. The error function L(θ) is as follows:
[0029]
[0030] Where y tot is as follows:
[0031] y tot = r + γmax a′ Q tot (b′, a′|θ - ) (9)
[0032] Where r is the agent joint reward, Q tot is the joint action function, b represents the joint action - observation history value of all agents, a represents the agent joint action, b′ represents the joint action history observation value after the execution of joint action a, a′ represents the joint action with the maximum Q value among all possible joint actions under the history b′, and the parameters of the target network are θ - .
[0033] Step 7: Determine whether the algorithm has reached the preset number of training times. If not, each agent u gets the next action a u , and enters the loop operation of Step 1. If so, directly end the training.
[0034] Advantages of the present invention: By using hypergraphs to model agent relationships, the present invention can extract deeper relationship features, and design rewards for agents using the trust game of heterogeneous investment, enabling agents to make decisions based on the trust level of other agents and encouraging agents to cooperate to become trustworthy individuals to improve the cooperation level in multi - agent systems. Brief Description of the Drawings
[0035] Figure 1 is the flowchart of the method of the present invention.
[0036] Figure 2 is the process of the heterogeneous investment N-person trust game.
[0037] Figure 3 is the overall data flow. Detailed Embodiment
[0038] The following further describes the detailed embodiment of the present invention in conjunction with the drawings.
[0039] As Figure 1 shown in the main flowchart of the present invention, a multi-agent cooperation method based on heterogeneous investment reinforcement learning includes the following steps, where the agent is specifically a combat unit in StarCraft II:
[0040] The first step: The combat unit first initializes the experience pool to store learning experiences.
[0041] The second step: Initialize each network of the combat unit.
[0042] The third step: The combat unit contacts the environment and receives the environmental state to obtain its observation value. Among them, the observation value is the characteristics of friendly and enemy forces in the StarCraft II environment, including blood volume, type, armor, distance, etc., as well as the movement direction and its own characteristics.
[0043] The fourth step: The combat unit calculates the combat unit behavior according to the observation value and the strategy, and generates a corresponding hypergraph according to formulas (1), (2), and (3) using all the combat unit observation values.
[0044] The fifth step: The combat unit takes actions according to the actions output in the fourth step.
[0045] The sixth step: The actions of the combat unit interact with the environment to generate an environmental reward r env , and then its actions will obtain a trust game reward r according to the N-person trust game mechanism of heterogeneous investment tg .
[0046] The seventh step: The combat unit stores the interaction experience with the environment this time into the experience pool.
[0047] The eighth step: Update the parameters of the network according to the loss function of formula (8).
[0048] The ninth step: If the training process is not completed, repeat the process from the third step to the eighth step.
[0049] Figure 2Describes the specific process of the N-person trust game for heterogeneous investment. The N-person trust game mainly involves two roles: the investor and the investee. The actions of the investor are to invest or not to invest, and the actions of the investee are to return or not to return, where investment and return are regarded as cooperative actions, and non-investment and non-return are regarded as non-cooperative behaviors. In an N-person group, one investor will be selected, and the remaining individuals will be the investees. The investor first decides whether to invest in the investee. If the investor chooses not to invest, the investee will not receive any income, and the result of its return or non-return is the same, and the income of both parties is zero, entering the next round of the game or the game ends. If the investor chooses to invest, the investment amount iv will be prepared, and then the proportion of trustworthy individuals who chose to cooperate in the previous round among the investees will be observed Then the actual investment amount I is:
[0050]
[0051] Where N tw is the number of trustworthy individuals, and N ut is the number of untrustworthy individuals.
[0052] This part of the investment amount will be amplified by a fixed gain factor G greater than 1, and the investee will obtain this amplified amount. Then the income p of the investee before choosing to return investee is:
[0053]
[0054] After obtaining the amplified investment income, the investee can choose whether to return part of the income to the investor. If the investee chooses not to return, the investor will not receive any return and bear all the investment losses. If the investee chooses to return a part of the income in proportion g, the investor will receive this part of the return. Then the income of the investor can be expressed as:
[0055] p invest or = g × p investee × N return - I (12)
[0056] The income of the investee who chooses to return is:
[0057]
[0058] Where N return is the number of investees who choose to return. Finally, the settlement is made and then enter the next round of the game, or end.
[0059] In addition, each individual records their previous round of behavior and calculates the proportion of their cooperative behavior. The final payoff will be weighted and adjusted according to this proportion. This mechanism promotes long-term cooperation among individuals to a certain extent, encourages more investors to choose to return, improves the payoff expectations of investors, and forms a virtuous reciprocal cycle. At the same time, this dynamic feedback mechanism also enables the individual's behavior history to have a continuous impact on future decisions.
[0060] Figure 3 Describes the overall data flow. First, the Q value is generated based on all agent observations o and history τ, and then the adjacency matrix M of the hypergraph is generated according to the joint observation Z using formulas (1), (2), and (3). H . The generated M H and the Q value generate Q’, Q”, and Q through the hypergraph convolutional network and the global state module according to formulas (4), (5), (6), and (7). tot In addition, the agent generates an action a according to the observation and the policy. u The action a u generates an environmental reward r when acting on the environment, env and generates a game reward r according to the heterogeneous investment N-person trust game mechanism. tg The two parts of the reward constitute the agent reward r. The reward r and the joint action function Q tot constitute a loss function according to formula (8) to update the network parameters.
Claims
1. A multi-agent cooperation method based on heterogeneous investment reinforcement learning, characterized in that It includes the following steps: Step 1: Initialize an experience pool. The interaction experience includes (s, o, τ, a, r, s′, o′), where s is the current environment information, s′ is the next-step environment information; a is the joint action, which is composed of the actions a of each agent u; r is the agent reward; input the number of agents n and the predetermined number of hyperedges m, and initialize the agent network. u ; Step 2: Each agent receives the environment state s t , select action a according to the current strategy π u , according to the observed value o u And historical action recordsτ u Generate the action-value function Q for each agent u u Values, and based on the observed values, form the adjacency matrix M of the hypergraph H of the agent relationship H ; Among them, the observed value o u is the characteristics and environment information of multiple agents; the generation process of the hypergraph matrix is as follows: Process the joint observations of the agent through the activation functions ReLU and Linear to generate an H1 matrix. Then, obtain the mean avg(H1) of the H1 matrix and multiply it by an n-order identity matrix to get an H2 matrix. Concatenate the adjacency matrix M of the hypergraph H of H1 and H2 H ; H1 = ReLU(Linear(Z)) (1) H2 = I n × avg(H1)(2) M H = H1|H2(3) Among them, Z is the combined observation value, which is composed of the observation value o of each agent u u constitutes, avg is the operation of taking the average value, I n is the n-order identity matrix; Step 3: Agent action a u On the one hand, it acts on the environment and generates an environmental reward r according to the environmental rules env , and on the other hand, the agent conducts an N-person trust game of heterogeneous investment and calculates its reward r according to its mechanism tg ; The two rewards together constitute the agent's reward r, where the reward r = r env + r tg ; Step 4: The Q of each agent u The vector Q formed and M generated in Step 1 H First, it passes through the multi-layer hypergraph convolutional layer HGCN to generate Q', where Q' is the preliminary joint action value after integrating the information of each agent obtained through the convolutional operation; Q′ = HGCN(HGCN(Q, abs(W1)), W2) (4) Among them, two layers of HGCN are used, and W i , where i = 1, 2 represents different layer hyperedge weight matrices, and abs is the absolute value operation; the HGCN operation is as follows: where x (l) is the input signal, x (l+1) is the output signal, D(H), W (H) and B(H) are the matrix of vertex degrees, the hyperedge weight matrix, and the hyperedge degree matrix of the hypergraph H respectively, P is the weight matrix between the l-th layer and the (l + 1)-th layer, which is set to be the n-dimensional identity matrix; that is, the input signal is Q, and Q' is obtained through two layers of HGCN operations; Step 5: Q' generates the joint action value function Q that integrates the global state information through the elu and sum operations of the global state module tot ; The specific generation process is as follows: Q″ = elu(abs(MLP(s t )) ⊙ Q′ + MLP(s t )) (6) Q tot = sum(abs(MLP(s t ))⊙Q″))+MLP(s t ) (7) where MLP is a multi-layer perceptron, a common artificial neural network used to map a set of input vectors to a set of output vectors; ⊙ is the symbol for element-wise multiplication; Step 6: Agent reward r and joint action value function Q tot Participate in the calculation of the error function; the error function L(θ) is shown as follows: where y tot is as follows: y tot = r + γmax a′ Q tot (b′, a′|θ - ) (9) where r is the joint reward of the agents, Q tot is the joint action function, b represents the joint action-observation history of all agents, a represents the joint action of the agents, b′ represents the joint action history observation after the joint action a is executed, a′ represents the joint action with the largest Q value among all possible joint actions under the history b′, and the parameters of the target network are θ - ; Step 7: Determine whether the algorithm has reached the preset number of training times; if not, each agent u obtains the next action a u , enter the loop operation of Step 1, if so, directly end the training.
2. The multi-agent cooperation method based on heterogeneous investment reinforcement learning according to claim 1, wherein, The specific process of the N-person trust game with heterogeneous investment is as follows: The N-person trust game involves two roles: the investor and the investee; the actions of the investor are to invest or not to invest, and the actions of the investee are to return or not to return. Among them, investment and return are regarded as cooperative actions, and non-investment and non-return are regarded as non-cooperative behaviors; In an N-person group, an investor is selected, and the remaining individuals are the investees. The investor first decides whether to invest in the investees. If the investor chooses not to invest, the investees will not receive any returns, and the outcome of whether they return or not is the same. The payoffs of both parties are zero, and the game enters the next round or ends. If the investor chooses to invest, they will prepare an investment amount iv and then observe the proportion of trustworthy individuals who chose to cooperate in the previous round among the investees. Then its actual investment amount I is: Among them, N tw is the number of trustworthy individuals, and N ut is the number of untrustworthy individuals; The investment amount is amplified by a fixed gain factor G greater than 1, and the investor will receive this amplified amount. Then the investor selects the return p before the return. investee It is: After the investee obtains the amplified investment income, it chooses whether to return part of the income to the investor; if the investee chooses not to return, the investor will not receive any return and bear all investment losses; if the investee chooses to return a part of the income at a ratio of g, the investor receives this part of the return; then the income of the investor is expressed as: p investor = g × p investee × N return - I(12) The income of the investee who chooses to return is: where N return is the number of investors selected for return; final settlement then proceeds to the next round of the game, or ends; In addition, each individual will record their previous round of behavior and calculate the proportion of their cooperative behavior, and the final income is weighted and adjusted according to this proportion.