Multi-agent intelligent decision-making method and device, electronic equipment and storage medium

By using the empathy and gift-giving modules in a pre-trained multi-agent reinforcement learning model, the social relationships between agents are inferred, solving the problems of group cooperation and exploitation in multi-agent systems, and realizing friendly behavior and strategy selection.

CN120975256APending Publication Date: 2025-11-18BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410605990.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-15
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In multi-agent reinforcement learning, existing technologies struggle to achieve group cooperation and are easily exploited by others, making it difficult for agents to coordinate effectively.

Method used

By using a pre-trained multi-agent reinforcement learning model, the social relationships between agents are inferred using the empathy module and the gift-giving module. This helps determine decision-making strategies and guide agent behavior to promote group cooperation and avoid exploitation.

Benefits of technology

It achieves the goal of promoting agents to choose prosocial and friendly behaviors in mixed-motivation games, accurately identifying the friendly abilities of others, avoiding exploitation, and deriving decision-making strategies that can both promote group cooperation and avoid being taken advantage of by others.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120975256A_ABST
    Figure CN120975256A_ABST
Patent Text Reader

Abstract

The invention provides a multi-agent intelligent decision-making method and device, electronic equipment and a storage medium, to-be-processed scene information is obtained, and the to-be-processed scene information comprises a plurality of agents, geographical environment information of the agents, and barrier distribution information and resource distribution information existing in the geographical environment; and inputting the to-be-processed scene information into a pre-trained multi-agent reinforcement learning model to obtain a multi-agent decision strategy which is output by the multi-agent reinforcement learning model and corresponds to the to-be-processed scene information, and the multi-agent reinforcement learning model deduces a social relationship between agents through to-be-processed scene information, and determines the multi-agent decision strategy based on the social relationship. The multi-agent decision-making strategy which can promote group cooperation and can avoid being stripped by other agents can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-agent technology, and in particular to a multi-agent intelligent decision-making method, apparatus, electronic device, and storage medium. Background Technology

[0002] The essence of human intelligence is social intelligence. Most human activities involve social groups of multiple people, and solving large, complex problems requires coordination among multiple organizations. To effectively simulate human intelligent activities, multi-agent reinforcement learning has become a trend in the field of agent research.

[0003] According to relevant technologies, multi-agent reinforcement learning is often achieved through centralized training and distributed execution, or distributed training and distributed execution. With centralized training and distributed execution, the conflicting interests of individual agents often make it difficult to achieve coordinated behavior within the group; with distributed training and distributed execution, agents are prone to convergence to a suboptimal equilibrium, making effective cooperation difficult.

[0004] Therefore, finding a multi-agent intelligent decision-making method that can both promote group cooperation and avoid being exploited by others has become a research hotspot. Summary of the Invention

[0005] This invention provides a multi-agent intelligent decision-making method, apparatus, electronic device, and storage medium, which enables the development of multi-agent decision-making strategies that promote group cooperation while avoiding exploitation by other agents.

[0006] This invention provides a multi-agent intelligent decision-making method, the method comprising: acquiring scene information to be processed, wherein the scene information to be processed includes multiple agents, geographical environment information of the agents, obstacle distribution information and resource distribution information in the geographical environment; inputting the scene information to be processed into a pre-trained multi-agent reinforcement learning model to obtain a multi-agent decision-making strategy output by the multi-agent reinforcement learning model corresponding to the scene information to be processed, wherein the multi-agent reinforcement learning model infers the social relationships between agents through the scene information to be processed, and determines the multi-agent decision-making strategy based on the social relationships.

[0007] According to a multi-agent intelligent decision-making method provided by the present invention, after obtaining the multi-agent decision-making strategy output by the multi-agent reinforcement learning model corresponding to the scene information to be processed, the method further includes: guiding the behavior of each agent based on the multi-agent decision-making strategy.

[0008] According to a multi-agent intelligent decision-making method provided by the present invention, the multi-agent reinforcement learning model includes an empathy module and a gift-giving module; the multi-agent reinforcement learning model is pre-trained in the following manner: acquiring a training dataset, wherein the training dataset includes training scene information; inputting the training scene information into the empathy module to obtain the social relationships between multiple training agents corresponding to the training scene information output by the empathy module; inputting the social relationships between the multiple training agents into the gift-giving module to obtain the decision strategies of the multiple training agents corresponding to the social relationships output by the gift-giving module; determining the benefits of the training agents based on the decision strategies of the multiple training agents, and iteratively optimizing the multi-agent reinforcement learning model with maximizing the benefits of the training agents as the optimization objective, until the multi-agent reinforcement learning model converges, thereby obtaining a trained multi-agent reinforcement learning model.

[0009] According to a multi-agent intelligent decision-making method provided by the present invention, the empathy module includes an adversary modeling module and a social relationship reasoning module. The step of inputting the training scenario information into the empathy module to obtain the social relationships among multiple training agents corresponding to the training scenario information, as output by the empathy module, specifically includes: inputting the training scenario information into the adversary modeling module to obtain the predicted behavior of the current training agent towards other training agents, and the evaluation index of the predicted behavior of the current training agent towards other training agents, as output by the adversary modeling module; inputting the predicted behavior of the current training agent towards other training agents, and the evaluation index of the predicted behavior of the current training agent towards other training agents, into the social relationship reasoning module to obtain the social relationships among multiple training agents corresponding to the training scenario information, as output by the social relationship reasoning module.

[0010] According to a multi-agent intelligent decision-making method provided by the present invention, the adversary modeling module includes an egocentric value network, an egocentric policy network, and an observation transformation network, wherein the observation transformation network and the egocentric policy network are connected in series and then connected in parallel with the egocentric value network; the step of inputting the training scenario information into the adversary modeling module to obtain the predicted behavior of the current training agent against other training agents and the evaluation index of the predicted behavior of the current training agent against other training agents output by the adversary modeling module specifically includes: inputting the training scenario information into the observation transformation network and the egocentric policy network respectively. The self-centered value network is described, wherein the training scenario information is input into the observation transformation network to obtain the observation content of other training agents inferred by the current training agent, output by the observation transformation network; the observation content of other training agents inferred by the current training agent is input into the self-centered policy network to obtain the predicted behavior of the current training agent towards other training agents, output by the self-centered policy network; and the training scenario information is input into the self-centered value network to obtain the evaluation index of the predicted behavior of the current training agent towards other training agents, output by the self-centered value network.

[0011] According to a multi-agent intelligent decision-making method provided by the present invention, the multi-agent reinforcement learning model includes an empathy module and a gift-giving module; the step of inputting the scene information to be processed into the pre-trained multi-agent reinforcement learning model to obtain the multi-agent decision-making strategy output by the multi-agent reinforcement learning model corresponding to the scene information to be processed specifically includes: inputting the scene information to be processed into the empathy module to obtain the social relationships between multiple agents corresponding to the scene information to be processed, output by the empathy module; and inputting the social relationships between the multiple agents into the gift-giving module to obtain the multi-agent decision-making strategy output by the gift-giving module corresponding to the social relationships between the multiple agents.

[0012] The present invention also provides a multi-agent intelligent decision-making device, the device comprising: an acquisition module for acquiring scene information to be processed, wherein the scene information to be processed includes multiple agents, geographical environment information of the agents, obstacle distribution information and resource distribution information in the geographical environment; and a processing module for inputting the scene information to be processed into a pre-trained multi-agent reinforcement learning model to obtain a multi-agent decision-making strategy output by the multi-agent reinforcement learning model corresponding to the scene information to be processed, wherein the multi-agent reinforcement learning model infers the social relationships between agents through the scene information to be processed and determines the multi-agent decision-making strategy based on the social relationships.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the multi-agent intelligent decision-making method as described above.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multi-agent intelligent decision-making method as described above.

[0015] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the multi-agent intelligent decision-making method as described above.

[0016] This invention provides a multi-agent intelligent decision-making method, apparatus, electronic device, and storage medium. The method acquires scene information to be processed; inputs this scene information into a pre-trained multi-agent reinforcement learning model to obtain a multi-agent decision-making strategy output by the model corresponding to the scene information. The multi-agent reinforcement learning model infers social relationships between agents based on the scene information and determines the multi-agent decision-making strategy based on these relationships. By inferring social relationships between agents in a mixed-motivation game, it can encourage agents to choose prosocial and friendly behaviors. Furthermore, by enabling agents to recognize whether others are friendly towards them, it can more accurately choose gift-giving strategies, avoiding exploitation by other agents. This achieves a multi-agent decision-making strategy that both promotes group cooperation and avoids exploitation by other agents. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 This is one of the flowcharts of the intelligent decision-making method for multiple agents provided by the present invention;

[0019] Figure 2 This is the second flowchart of the intelligent decision-making method for multiple agents provided by the present invention;

[0020] Figure 3 This is a flowchart illustrating the pre-trained multi-agent reinforcement learning model provided by the present invention.

[0021] Figure 4This is a schematic diagram of the structure of the multi-agent reinforcement learning model provided by the present invention;

[0022] Figure 5 This invention provides a flowchart illustrating the process of inputting scene information to be processed into a pre-trained multi-agent reinforcement learning model to obtain the multi-agent decision-making strategy output by the multi-agent reinforcement learning model corresponding to the scene information to be processed.

[0023] Figure 6 This is a schematic diagram of the structure of the multi-agent intelligent decision-making device provided by the present invention;

[0024] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0026] The multi-agent intelligent decision-making method provided by this invention can achieve intelligent decision-making in mixed-motivation games, promote cooperation, and increase collective rewards; at the same time, it can flexibly adjust strategies to avoid exploitation by others (other agents). Specifically, the multi-agent intelligent decision-making method improves the interpretability of multi-agent group behavior by integrating and modeling the gifting and empathy mechanisms of human society.

[0027] Figure 1 This is one of the flowcharts of the intelligent decision-making method for multiple agents provided by the present invention.

[0028] The following will combine Figure 1 The process of the multi-agent intelligent decision-making method provided by the present invention will be described.

[0029] In an exemplary embodiment of the present invention, combined with Figure 1 As can be seen, the intelligent decision-making method of multi-agent systems may include steps 110 and 120, which will be described in detail below.

[0030] In step 110, the scene information to be processed is obtained.

[0031] In one embodiment, the scenario information to be processed may include multiple agents, geographical environment information of the agents, obstacle distribution information, and resource distribution information in the geographical environment. The scenario information to be processed can be considered as a scenario involving multiple agents making decisions. By obtaining a multi-agent decision-making strategy, the actions of multiple agents can be guided based on this strategy, so that the actions of the agents guided by the decision-making strategy can both promote group cooperation and avoid exploitation by others.

[0032] In step 120, the scene information to be processed is input into a pre-trained multi-agent reinforcement learning model to obtain the multi-agent decision-making strategy output by the multi-agent reinforcement learning model corresponding to the scene information to be processed.

[0033] In one embodiment, the scene information to be processed can be input into a pre-trained multi-agent reinforcement learning model, thereby obtaining a multi-agent decision-making strategy output by the multi-agent reinforcement learning model corresponding to the scene information to be processed. The actions of the agents guided by this decision-making strategy can both promote group cooperation and prevent exploitation by others.

[0034] In another embodiment, the multi-agent reinforcement learning model can infer the social relationships between agents by considering the multiple agents in the scene to be processed, the geographical environment information of the agents, the distribution of obstacles in the geographical environment, and the distribution of resources. That is, by inferring the social relationships between agents in a mixed-motivation game, it can encourage agents to choose prosocial and friendly behaviors. Based on these social relationships, the model determines multi-agent decision-making strategies, enabling agents to more accurately choose gift-giving strategies and avoid being exploited by other agents by equipping them with the ability to recognize whether others are friendly to them. This achieves the goal of deriving multi-agent decision-making strategies that both promote group cooperation and avoid exploitation by other agents.

[0035] This invention provides a multi-agent intelligent decision-making method that acquires scenario information to be processed; inputs this scenario information into a pre-trained multi-agent reinforcement learning model to obtain a multi-agent decision-making strategy output by the multi-agent reinforcement learning model corresponding to the scenario information. The multi-agent reinforcement learning model infers the social relationships between agents based on the scenario information and determines the multi-agent decision-making strategy based on these social relationships. By inferring the social relationships between agents in a mixed-motivation game, it can encourage agents to choose prosocial and friendly behaviors. Furthermore, by enabling agents to recognize whether others are friendly to them, it can more accurately choose gift-giving strategies and avoid being exploited by other agents. This achieves a multi-agent decision-making strategy that both promotes group cooperation and avoids exploitation by other agents.

[0036] Figure 2 This is the second flowchart of the intelligent decision-making method for multiple agents provided by this invention.

[0037] The following will combine Figure 2 The process of another multi-agent intelligent decision-making method is explained.

[0038] In an exemplary embodiment of the present invention, combined with Figure 2 As can be seen, the intelligent decision-making method of multi-agents may include steps 210 to 230, wherein steps 210 to 220 are the same as or similar to steps 110 to 120 respectively. For the specific implementation and beneficial effects, please refer to the previous description. In this embodiment, they will not be repeated. Step 230 will be introduced below.

[0039] In step 230, the behavior of each agent is guided based on the multi-agent decision-making strategy.

[0040] In one embodiment, the actions of agents at each time step can be guided based on the multi-agent decision-making strategy output by the multi-agent reinforcement learning model, so that the actions of the agents can both promote group cooperation and avoid being exploited by others (other agents).

[0041] Figure 3 This is a flowchart illustrating the pre-trained multi-agent reinforcement learning model provided by the present invention.

[0042] The following will combine Figure 3 The process of pre-training a multi-agent reinforcement learning model is explained.

[0043] In an exemplary embodiment of the present invention, the multi-agent reinforcement learning model may include an empathy module and a gifting module; combined with Figure 3 As can be seen, the pre-trained multi-agent reinforcement learning model may include steps 310 to 340, and each step will be described below.

[0044] In step 310, the training dataset is obtained.

[0045] The training dataset includes training scenario information.

[0046] In one embodiment, the training scene information may include multiple training agents, geographical environment information of the training agents, obstacle distribution information, and resource distribution information in the geographical environment of the training agents. It is understood that the information types contained in the training scene information are the same as or similar to the information types contained in the scene information to be processed.

[0047] In step 320, the training scenario information is input into the empathy module to obtain the social relationships between multiple training agents corresponding to the training scenario information, as output by the empathy module.

[0048] In step 330, the social relationships between multiple training agents are input into the gift-giving module to obtain the decision-making strategies of the multiple training agents corresponding to the social relationships between the multiple training agents, which are output by the gift-giving module.

[0049] In one embodiment, training scenario information can be input into the empathy module, thereby obtaining the social relationships between multiple training agents corresponding to the training scenario information, output by the empathy module. Further, the social relationships between the multiple training agents are then input into the gift-giving module, thereby obtaining the decision-making strategies of the multiple training agents corresponding to the social relationships between them, output by the gift-giving module.

[0050] Among them, the multi-agent reinforcement learning model can achieve perspective-taking through the empathy module, analyzing and inferring the social relationships among multiple trained agents. Furthermore, the gift-giving module can determine the gift-giving strategy based on the inferred social relationships, giving more rewards to agents with higher ratings, thereby obtaining the multi-agent decision-making strategy based on the gift-giving strategy.

[0051] In step 340, based on the decision-making strategy of the multi-training agent, the reward of the training agent is determined, and the multi-agent reinforcement learning model is iteratively optimized with the goal of maximizing the reward of the training agent, until the multi-agent reinforcement learning model converges, and the trained multi-agent reinforcement learning model is obtained.

[0052] In one embodiment, the payoff of the training agents can be determined based on the decision-making strategies of the multi-trained agents. Then, with maximizing the payoff of the training agents as the optimization objective, the empathy and gift-giving modules in the multi-agent reinforcement learning model are iteratively optimized until the model converges, thus obtaining a well-trained multi-agent reinforcement learning model. It is understood that the aforementioned pre-training ensures that the well-trained multi-agent reinforcement learning model can infer social relationships between agents in mixed-motivation games, promote agents to choose prosocial and friendly behaviors, and, by enabling agents to recognize whether others are friendly to them, more accurately choose gift-giving strategies, avoiding exploitation by other agents. This achieves the goal of deriving multi-agent decision-making strategies that both promote group cooperation and avoid exploitation by other agents.

[0053] In yet another exemplary embodiment of the present invention, the empathy module may include an adversary modeling module and a social relationship reasoning module. The process of inputting training scenario information into the empathy module to obtain the social relationships between multiple trained agents corresponding to the training scenario information, as output by the empathy module, can be implemented in the following manner:

[0054] The training scenario information is input into the opponent modeling module to obtain the predicted behavior of the current training agent against other training agents, as well as the evaluation index of the predicted behavior of the current training agent against other training agents.

[0055] The predicted behavior of the current training agent towards other training agents, as well as the evaluation index of the predicted behavior of the current training agent towards other training agents, are input into the social relationship reasoning module to obtain the social relationships between multiple training agents corresponding to the training scenario information output by the social relationship reasoning module.

[0056] In one embodiment, training scenario information can be input into the adversary modeling module, thereby obtaining the predicted behavior of the current training agent towards other training agents, as well as the evaluation index of the predicted behavior of the current training agent towards other training agents, output by the adversary modeling module. Specifically, the adversary modeling module can utilize empathy theory to achieve perspective-taking from others (other agents or other training agents), and can also simulate the perspectives of others through an observation-transformation network, thereby inferring the strategies of others using its own model.

[0057] In another embodiment, the predicted behavior of the current training agent towards other training agents, as well as the evaluation index of the predicted behavior of the current training agent towards other training agents, can be input into the social relationship reasoning module, thereby obtaining the social relationships between multiple training agents corresponding to the training scenario information output by the social relationship reasoning module.

[0058] In the social relationship reasoning module, counterfactual reasoning can be used to compare the impact of others' (other agents or other trained agents) actual actions on oneself with one's own psychological expectations of others, thereby obtaining social relationships with others. Specifically, this is achieved by using the current trained agent's predicted behavior towards other trained agents, and evaluating the current trained agent's predicted behavior towards other trained agents. A simple example is: if an agent is believed to have shown a friendly attitude but chooses a disadvantageous action, then a lower rating is given to this relationship. Subsequently, a gift-giving strategy can be determined based on the inferred social relationship, giving more rewards to agents with higher ratings. This allows for more accurate selection of gift-giving strategies, avoiding exploitation by other agents.

[0059] In application, the social relationship reasoning module uses the results obtained from the opponent modeling module—namely, predictions of the value of others' strategies and actions—to perform counterfactual reasoning. It calculates a "counterfactual baseline" by multiplying the probability of others taking each possible action by the value that action generates for the agent. Then, it compares the value of others' actual actions to the agent with this counterfactual baseline to obtain the agent's "social relationship" with others. After normalization, this social relationship is used as a weight in gifting to allocate the agent's rewards to others; the better the social relationship, the higher the weight, and the greater the reward allocated.

[0060] In yet another exemplary embodiment of the present invention, the description continues using the previously described embodiments as examples. The adversary modeling module may include a self-centered value network, a self-centered policy network, and an observation transformation network, wherein the observation transformation network and the self-centered policy network are connected in series and then connected in parallel with the self-centered value network.

[0061] The process of inputting training scenario information into the adversary modeling module to obtain the predicted behavior of the current training agent against other training agents, as well as the evaluation index of the predicted behavior of the current training agent against other training agents, can be achieved in the following way:

[0062] The training scenario information is input into the observation-transformation network and the egocentric value network, respectively.

[0063] The training scenario information is input into the observation transformation network to obtain the observation content of other training agents inferred by the current training agent.

[0064] The observations of other training agents inferred by the current training agent are input into an egocentric policy network to obtain the predicted behavior of the current training agent towards other training agents, output by the egocentric policy network; and

[0065] The training scenario information is input into the egocentric value network to obtain the evaluation index of the current training agent's predicted behavior towards other training agents, which is output by the egocentric value network.

[0066] In one embodiment, since the observation transformation network and the egocentric policy network are connected in series and then connected in parallel with the egocentric value network, training scenario information can be input into the observation transformation network and the egocentric value network respectively.

[0067] In the line where training scene information is input into the observation transformation network, the training scene information can be input into the observation transformation network to obtain the observation content of other training agents inferred by the current training agent. Then, the observation content of other training agents inferred by the current training agent is input into the egocentric policy network to obtain the predicted behavior of the current training agent towards other training agents output by the egocentric policy network.

[0068] In the process of inputting training scenario information into the egocentric value network, the training scenario information can be input into the egocentric value network, thereby obtaining the evaluation index of the current training agent's predicted behavior towards other training agents, which is output by the egocentric value network.

[0069] The agent models the policies and values ​​of others in the environment using two neural networks: a self-centered policy network and a self-centered value network. Both networks are optimized based on the agent's own reward function within the environment. To achieve this empathetic modeling, an observation transformation network is also configured to predict the current visual observations of others. This addresses the problem of partially observable environments, thus enabling more accurate adversary modeling. All three neural networks can be trained using offline reinforcement learning (RL) algorithms.

[0070] Figure 4 This is a schematic diagram of the structure of the multi-agent reinforcement learning model provided by the present invention.

[0071] Combination Figure 4 As can be seen, the multi-agent reinforcement learning model includes an empathy module and a gift-giving module. The empathy module further includes an adversary modeling module and a social relationship reasoning module. The working relationships between these modules have already been described above and will not be repeated here.

[0072] In yet another embodiment, Figure 4As not shown in the diagram, the adversary modeling module may further include a self-centered value network, a self-centered policy network, and an observation transformation network. The observation transformation network and the self-centered policy network are connected in series and then connected in parallel with the self-centered value network. The detailed working principle of the adversary modeling module has been described above and will not be repeated here.

[0073] Figure 5 This invention provides a flowchart illustrating the process of inputting scene information to be processed into a pre-trained multi-agent reinforcement learning model to obtain the multi-agent decision-making strategy output by the multi-agent reinforcement learning model corresponding to the scene information to be processed.

[0074] The following will combine Figure 5 The process of inputting the scene information to be processed into a pre-trained multi-agent reinforcement learning model and obtaining the multi-agent decision-making strategy output by the multi-agent reinforcement learning model corresponding to the scene information to be processed is explained.

[0075] In an exemplary embodiment of the present invention, the multi-agent reinforcement learning model may include an empathy module and a gift-giving module. Combined with... Figure 5 As can be seen, inputting the scene information to be processed into a pre-trained multi-agent reinforcement learning model to obtain the multi-agent decision policy output by the multi-agent reinforcement learning model corresponding to the scene information to be processed can include steps 510 and 520, which will be described in detail below.

[0076] In step 510, the scene information to be processed is input into the empathy module to obtain the social relationships between multiple intelligent agents corresponding to the scene information to be processed, which are output by the empathy module.

[0077] In step 520, the social relationships between multiple agents are input into the gift-giving module to obtain the decision-making strategies of the multi-agents corresponding to the social relationships between the multiple agents, which are output by the gift-giving module.

[0078] In one embodiment, the scenario information to be processed can be input into the empathy module, thereby obtaining the social relationships between multiple agents corresponding to the scenario information output by the empathy module. Further, the social relationships between the multiple agents are then input into the gift-giving module, thereby obtaining the multi-agent decision-making strategies corresponding to the social relationships output by the gift-giving module. By inferring the social relationships between agents through the empathy module in a mixed-motivation game, agents can be encouraged to choose prosocial and friendly behaviors. Furthermore, the gift-giving module enables agents to recognize whether others are friendly towards them, allowing for more accurate selection of gift-giving strategies and avoiding exploitation by other agents. This achieves the goal of deriving multi-agent decision-making strategies that both promote group cooperation and avoid exploitation by other agents.

[0079] It should be noted that the processing of the scenario information to be processed through the adversary modeling module and the social relationship reasoning module in the empathy module, as well as the processing through the egocentric value network, egocentric policy network, and observation transformation network in the adversary modeling module, are the same as or similar to the pre-training process described above, and will not be repeated in this embodiment.

[0080] As described above, the multi-agent intelligent decision-making method provided by this invention acquires information about a scenario to be processed; inputs this information into a pre-trained multi-agent reinforcement learning model to obtain a multi-agent decision-making strategy output by the model corresponding to the scenario information. The multi-agent reinforcement learning model infers the social relationships between agents based on the scenario information and determines the multi-agent decision-making strategy based on these relationships. By inferring social relationships between agents in a mixed-motivation game, it can encourage agents to choose prosocial and friendly behaviors. Furthermore, by enabling agents to recognize whether others are friendly towards them, it can more accurately choose gift-giving strategies and avoid being exploited by other agents. This achieves the goal of deriving multi-agent decision-making strategies that both promote group cooperation and avoid exploitation by other agents.

[0081] Based on the same concept, the present invention also provides a multi-agent intelligent decision-making device.

[0082] The intelligent decision-making device for multiple agents provided by the present invention is described below. The intelligent decision-making device for multiple agents described below can be referred to in correspondence with the intelligent decision-making method for multiple agents described above.

[0083] Figure 6 This is a schematic diagram of the structure of the multi-agent intelligent decision-making device provided by the present invention.

[0084] In an exemplary embodiment of the present invention, combined with Figure 6As can be seen, the intelligent decision-making device for multiple agents may include an acquisition module 610 and a processing module 620, and each module will be described in detail below.

[0085] The acquisition module 610 can be configured to acquire scene information to be processed, wherein the scene information to be processed includes multiple intelligent agents, the geographical environment information of the intelligent agents, the obstacle distribution information and resource distribution information in the geographical environment;

[0086] The processing module 620 can be configured to input the scene information to be processed into a pre-trained multi-agent reinforcement learning model to obtain a multi-agent decision-making strategy output by the multi-agent reinforcement learning model corresponding to the scene information to be processed. The multi-agent reinforcement learning model infers the social relationship between agents through the scene information to be processed and determines the multi-agent decision-making strategy based on the social relationship.

[0087] In an exemplary embodiment of the present invention, the processing module 620 may further be configured to:

[0088] Based on the multi-agent decision-making strategy, the behavior of each agent is guided.

[0089] In an exemplary embodiment of the present invention, the multi-agent reinforcement learning model includes an empathy module and a gift-giving module;

[0090] The processing module 620 can pre-train the multi-agent reinforcement learning model in the following manner:

[0091] Obtain a training dataset, wherein the training dataset includes training scenario information;

[0092] The training scenario information is input into the empathy module to obtain the social relationships between multiple training agents corresponding to the training scenario information, as output by the empathy module.

[0093] The social relationships among the multiple training agents are input into the gifting module to obtain the decision-making strategies of the multiple training agents output by the gifting module, which correspond to the social relationships among the multiple training agents.

[0094] Based on the decision-making strategy of the multi-agent training agent, the reward of the training agent is determined, and the multi-agent reinforcement learning model is iteratively optimized with the goal of maximizing the reward of the training agent, until the multi-agent reinforcement learning model converges, thus obtaining the trained multi-agent reinforcement learning model.

[0095] In an exemplary embodiment of the present invention, the empathy module includes an adversary modeling module and a social relationship reasoning module;

[0096] The processing module 620 can input the training scene information into the empathy module in the following way to obtain the social relationships between multiple training agents corresponding to the training scene information, as output by the empathy module:

[0097] The training scenario information is input into the opponent modeling module to obtain the predicted behavior of the current training agent against other training agents, and the evaluation index of the predicted behavior of the current training agent against other training agents.

[0098] The predicted behavior of the current training agent towards other training agents, and the evaluation index of the predicted behavior of the current training agent towards other training agents, are input into the social relationship reasoning module to obtain the social relationships between multiple training agents corresponding to the training scenario information output by the social relationship reasoning module.

[0099] In an exemplary embodiment of the present invention, the adversary modeling module includes an egocentric value network, an egocentric strategy network, and an observation transformation network, wherein the observation transformation network and the egocentric strategy network are connected in series and then connected in parallel with the egocentric value network.

[0100] The processing module 620 can input the training scenario information into the opponent modeling module in the following manner to obtain the predicted behavior of the current training agent against other training agents, and the evaluation index of the predicted behavior of the current training agent against other training agents, as output by the opponent modeling module:

[0101] The training scenario information is input into the observation-transformation network and the egocentric value network, respectively.

[0102] The training scenario information is input into the observation conversion network to obtain the observation content of other training agents inferred by the current training agent, as output by the observation conversion network.

[0103] The observations of other training agents inferred by the current training agent are input into the egocentric policy network to obtain the predicted behavior of the current training agent towards other training agents, output by the egocentric policy network; and

[0104] The training scenario information is input into the egocentric value network to obtain the evaluation index of the current training agent's predicted behavior towards other training agents, which is output by the egocentric value network.

[0105] In an exemplary embodiment of the present invention, the multi-agent reinforcement learning model includes an empathy module and a gift-giving module;

[0106] The processing module 620 can input the scene information to be processed into a pre-trained multi-agent reinforcement learning model in the following way to obtain the multi-agent decision-making strategy output by the multi-agent reinforcement learning model corresponding to the scene information to be processed:

[0107] The scene information to be processed is input into the empathy module to obtain the social relationships between multiple agents corresponding to the scene information to be processed, which are output by the empathy module.

[0108] The social relationships among the multiple agents are input into the gift-giving module to obtain the decision-making strategy of the multi-agent corresponding to the social relationships among the multiple agents, which is output by the gift-giving module.

[0109] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7 As shown, the electronic device may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740, wherein the processor 710, communications interface 720, and memory 730 communicate with each other via the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute a multi-agent intelligent decision-making method. This method includes: acquiring scene information to be processed, wherein the scene information to be processed includes multiple agents, geographical environment information of the agents, obstacle distribution information, and resource distribution information in the geographical environment; inputting the scene information to be processed into a pre-trained multi-agent reinforcement learning model to obtain a multi-agent decision-making strategy output by the multi-agent reinforcement learning model corresponding to the scene information to be processed, wherein the multi-agent reinforcement learning model infers the social relationships between agents based on the scene information to be processed and determines the multi-agent decision-making strategy based on the social relationships.

[0110] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0111] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multi-agent intelligent decision-making method provided by the above methods. The method includes: acquiring scene information to be processed, wherein the scene information to be processed includes multiple agents, geographical environment information of the agents, obstacle distribution information and resource distribution information in the geographical environment; inputting the scene information to be processed into a pre-trained multi-agent reinforcement learning model to obtain a multi-agent decision-making strategy output by the multi-agent reinforcement learning model corresponding to the scene information to be processed, wherein the multi-agent reinforcement learning model infers the social relationships between agents through the scene information to be processed and determines the multi-agent decision-making strategy based on the social relationships.

[0112] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements a multi-agent intelligent decision-making method provided by the above methods. The method includes: acquiring scene information to be processed, wherein the scene information to be processed includes multiple agents, geographical environment information of the agents, obstacle distribution information and resource distribution information in the geographical environment; inputting the scene information to be processed into a pre-trained multi-agent reinforcement learning model to obtain a multi-agent decision-making strategy output by the multi-agent reinforcement learning model corresponding to the scene information to be processed, wherein the multi-agent reinforcement learning model infers the social relationships between agents through the scene information to be processed and determines the multi-agent decision-making strategy based on the social relationships.

[0113] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0114] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0115] It is further understood that although the operations are described in a specific order in the accompanying drawings in the embodiments of the present invention, this should not be construed as requiring these operations to be performed in the specific order or serial order shown, or requiring all the operations shown to obtain the desired result. In certain environments, multitasking and parallel processing may be advantageous.

[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-agent intelligent decision-making method, characterized in that, The method includes: Acquire scene information to be processed, wherein the scene information to be processed includes multiple intelligent agents, the geographical environment information of the intelligent agents, the obstacle distribution information and resource distribution information existing in the geographical environment; The scene information to be processed is input into a pre-trained multi-agent reinforcement learning model to obtain the multi-agent decision-making strategy output by the multi-agent reinforcement learning model corresponding to the scene information to be processed. The multi-agent reinforcement learning model infers the social relationship between agents through the scene information to be processed and determines the multi-agent decision-making strategy based on the social relationship.

2. The multi-agent intelligent decision-making method according to claim 1, characterized in that, After obtaining the multi-agent decision policy output by the multi-agent reinforcement learning model corresponding to the scene information to be processed, the method further includes: Based on the multi-agent decision-making strategy, the behavior of each agent is guided.

3. The multi-agent intelligent decision-making method according to claim 1, characterized in that, The multi-agent reinforcement learning model includes an empathy module and a gift-giving module; The multi-agent reinforcement learning model is pre-trained using the following method: Obtain a training dataset, wherein the training dataset includes training scenario information; The training scenario information is input into the empathy module to obtain the social relationships between multiple training agents corresponding to the training scenario information, as output by the empathy module. The social relationships among the multiple training agents are input into the gifting module to obtain the decision-making strategies of the multiple training agents output by the gifting module, which correspond to the social relationships among the multiple training agents. Based on the decision-making strategy of the multi-agent training agent, the reward of the training agent is determined, and the multi-agent reinforcement learning model is iteratively optimized with the goal of maximizing the reward of the training agent, until the multi-agent reinforcement learning model converges, thus obtaining the trained multi-agent reinforcement learning model.

4. The multi-agent intelligent decision-making method according to claim 3, characterized in that, The empathy module includes an adversary modeling module and a social relationship reasoning module; The step of inputting the training scenario information into the empathy module to obtain the social relationships between multiple training agents corresponding to the training scenario information, as output by the empathy module, specifically includes: The training scenario information is input into the opponent modeling module to obtain the predicted behavior of the current training agent against other training agents, and the evaluation index of the predicted behavior of the current training agent against other training agents. The predicted behavior of the current training agent towards other training agents, and the evaluation index of the predicted behavior of the current training agent towards other training agents, are input into the social relationship reasoning module to obtain the social relationships between multiple training agents corresponding to the training scenario information output by the social relationship reasoning module.

5. The multi-agent intelligent decision-making method according to claim 4, characterized in that, The adversary modeling module includes an egocentric value network, an egocentric strategy network, and an observation transformation network, wherein the observation transformation network and the egocentric strategy network are connected in series and then connected in parallel with the egocentric value network. The step of inputting the training scenario information into the opponent modeling module to obtain the predicted behavior of the current training agent against other training agents, and the evaluation index of the predicted behavior of the current training agent against other training agents, specifically includes: The training scenario information is input into the observation-transformation network and the egocentric value network, respectively. The training scenario information is input into the observation conversion network to obtain the observation content of other training agents inferred by the current training agent, as output by the observation conversion network. The observations of other training agents inferred by the current training agent are input into the egocentric policy network to obtain the predicted behavior of the current training agent towards other training agents, output by the egocentric policy network; and The training scenario information is input into the egocentric value network to obtain the evaluation index of the current training agent's predicted behavior towards other training agents, which is output by the egocentric value network.

6. The multi-agent intelligent decision-making method according to claim 1, characterized in that, The multi-agent reinforcement learning model includes an empathy module and a gift-giving module; The step of inputting the scene information to be processed into a pre-trained multi-agent reinforcement learning model to obtain the multi-agent decision-making policy output by the multi-agent reinforcement learning model corresponding to the scene information to be processed specifically includes: The scene information to be processed is input into the empathy module to obtain the social relationships between multiple agents corresponding to the scene information to be processed, which are output by the empathy module. The social relationships among the multiple agents are input into the gift-giving module to obtain the decision-making strategy of the multi-agent corresponding to the social relationships among the multiple agents, which is output by the gift-giving module.

7. A multi-agent intelligent decision-making device, characterized in that, The device includes: The acquisition module is used to acquire scene information to be processed, wherein the scene information to be processed includes multiple intelligent agents, the geographical environment information of the intelligent agents, the obstacle distribution information and resource distribution information in the geographical environment; The processing module is used to input the scene information to be processed into a pre-trained multi-agent reinforcement learning model to obtain the multi-agent decision-making strategy output by the multi-agent reinforcement learning model corresponding to the scene information to be processed. The multi-agent reinforcement learning model infers the social relationship between agents through the scene information to be processed and determines the multi-agent decision-making strategy based on the social relationship.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the multi-agent intelligent decision-making method as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the multi-agent intelligent decision-making method as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the multi-agent intelligent decision-making method as described in any one of claims 1 to 6.