Interpretability Analysis Method, Device and Electronic Device for Multi-Agent Decision Making

By converting the real-time state data of multiple agents into a causal graph network, the contribution of each agent's state characteristics to decision-making is determined, and a multi-grained visual path is generated, which solves the problem of unclear causal relationships in multi-agent decision-making, and a more transparent and efficient decision-making process is achieved.

CN119962693BActive Publication Date: 2025-07-29INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510438581.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-29
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The causal relationships in existing multi-agent decisions are unclear and lack of explanatory ability, resulting in opaque decision-making process.

Method used

By obtaining real-time state data and decision data of multiple agents, the pre-trained time series encoder is used to convert it into a causal graph network, the contribution of each agent's state characteristics to decision-making is determined, and a multi-grained visual decision path is generated.

Benefits of technology

It improves the interpretability of decision causality, enhances the transparency of the decision process and system trust, can detect and punish lazy agents, and improves the overall effectiveness of multi-agent system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962693B_ABST
    Figure CN119962693B_ABST
Patent Text Reader

Abstract

The present invention provides an interpretable analysis method, device, and electronic device for multi-agent decision-making, which relates to the field of artificial intelligence technology. The method includes: acquiring real-time state data and real-time decision-making data of multiple agents during the execution of tasks; inputting the real-time state data into a pre-trained time series encoder, and converting the real-time state data into a causal graph network through the pre-trained time series encoder; the causal graph network includes causal relationships between the agents and between each agent and the task objective; determining the contribution of each state feature of each agent to the decision-making according to the real-time state data and the real-time decision-making data; generating a multi-granularity visual decision-making path according to the causal graph network and the contribution of each state feature of each agent to the decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an interpretable analysis method, device, and electronic device for multi-agent decision-making. Background Art

[0002] Multi-agent decision-making studies the behavior of multiple interacting agent groups in a system. In collaborative multi-agent reinforcement learning, a group of agents learns strategies through training to solve tasks that require a team to achieve a common goal. In recent years, there has been increasing attention to the various possibilities brought by better understanding causal relationships for machine learning. Explicit causal modeling in artificial intelligence is crucial for achieving general intelligence. For example, in medical diagnosis and treatment plan recommendation, doctors and patients need to understand and trust the decision-making basis of the system; in financial risk assessment and investment decision-making, financial institutions need to clarify the judgment logic of the system to meet regulatory requirements and internal audit standards; in optimal decision-making for unmanned driving, algorithm designers need to clarify the decision-making logic of driving strategies to ensure the safety and reliability of driving strategies.

[0003] The input information that can be obtained during the multi-agent decision-making process includes the state characteristics of the agents themselves, the surrounding environment, and the task objectives, as well as the decision-making behaviors of the agents. Interpretable causal reasoning expects to establish a causal relationship between the agent's behavior and the task objective. Currently, there are some methods for causal explanation, such as using clustering methods to cluster the decision-making behaviors of different agents, analyzing different categories of behaviors, and giving the causal relationship of the decision-making. However, although different categories of decision-making behaviors can assist in analyzing the role of agents in the decision-making process, this method lacks clear causal relationships, resulting in unclear interpretable causal relationships. Summary of the Invention

[0004] The present invention provides an interpretable analysis method, device, electronic device, non-transitory computer-readable storage medium, and computer program product for multi-agent decision-making, aiming to solve the defect of unclear causal relationships in existing technologies and achieve the effect of improving the interpretability of decision-making causal relationships.

[0005] The present invention provides an interpretable analysis method for multi-agent decision-making, including the following steps.

[0006] Obtain real-time state data and real-time decision data of multiple agents during the execution of tasks;

[0007] Input the real-time state data into a pre-trained time series encoder, and convert the real-time state data into a causal graph network through the pre-trained time series encoder; the causal graph network includes the causal relationships between the agents and between the agents and the task objective;

[0008] Determine the contribution of each state feature of each agent to the decision respectively according to the real-time state data and the real-time decision data;

[0009] Generate a multi-granularity visual decision path according to the causal graph network and the contribution of each state feature of each agent to the decision respectively.

[0010] According to an interpretable analysis method for multi-agent decision-making provided by the present invention, the training steps of the pre-trained time series encoder include:

[0011] Obtain multiple segments of sample state data; each segment of the sample state data contains the state data of multiple sample agents under a time series;

[0012] Divide each segment of the sample state data into sequence segments including multiple preset time steps respectively;

[0013] In each round of iteration, select multiple sequence segments and input them into the time series encoder to be trained, output the predicted causal graph network, input the predicted causal graph network and the selected sequence segments into the decoder to be trained, and output the predicted state data at the next moment corresponding to each selected sequence segment;

[0014] According to the difference between the predicted state data at the next moment and the real state data at the next moment in the sample state data, adjust the parameters of the time series encoder to be trained and the decoder to be trained, enter the next round of iteration, and stop the iteration until the trained time series encoder is obtained.

[0015] According to an interpretable analysis method for multi-agent decision-making provided by the present invention, the obtaining of multiple segments of sample state data includes:

[0016] Perform multi-round dynamic decision-making on the multiple sample agents through a pre-trained multi-agent reinforcement learning model;

[0017] Obtain multiple segments of sample state data according to the state data of the multiple sample agents in each round of dynamic decision-making process.

[0018] According to an interpretable analysis method for multi-agent decision-making provided by the present invention, the determining of the contribution of each state feature of each agent to the decision respectively according to the real-time state data and the real-time decision data includes:

[0019] Respectively use each state feature in the real-time state data as the target state feature, and determine the state feature subsets that can be formed by the state features other than the target state feature;

[0020] Based on the real-time decision data, determine the first decision corresponding to each of the state feature subsets, and the second decision corresponding to each of the state feature subsets plus the target state feature;

[0021] Based on the differences between the first decision and the second decision corresponding to each of the state feature subsets, determine the contribution of the target state feature to the decision, and obtain the contribution of each state feature of each of the agents to the decision.

[0022] According to an interpretable analysis method for multi-agent decision-making provided by the present invention, the determining the contribution of the target state feature to the decision based on the differences between the first decision and the second decision corresponding to each of the state feature subsets includes:

[0023] For each of the state feature subsets, calculate the marginal contribution of the target state feature to the state feature subset according to the difference between the corresponding first decision and the second decision;

[0024] Calculate the weight of each of the state feature subsets, and perform a weighted sum of the corresponding marginal contributions according to the weights of the state feature subsets to obtain the contribution of the target state feature to the decision.

[0025] According to an interpretable analysis method for multi-agent decision-making provided by the present invention, the generating a multi-granularity visual decision path based on the causal graph network and the contribution of each state feature of each of the agents to the decision includes:

[0026] Determine important agents according to the causal graph network;

[0027] Determine the important state features corresponding to each of the agents according to the contribution of each state feature of each of the agents to the decision;

[0028] Generate a multi-granularity visual decision path according to the important agents and the important state features.

[0029] The present invention also provides an interpretable analysis device for multi-agent decision-making, including the following modules:

[0030] A data acquisition module, configured to acquire real-time state data and real-time decision data of multiple agents during task execution;

[0031] A coarse-grained causal discovery module, configured to input the real-time state data into a pre-trained time series encoder, and convert the real-time state data into a causal graph network through the pre-trained time series encoder; the causal graph network includes causal relationships between the agents and between the agents and the task objective;

[0032] A fine-grained feature importance analysis module, configured to determine the contribution of each state feature of each of the agents to the decision respectively according to the real-time state data and the real-time decision data;

[0033] A decision path visualization module, configured to generate a multi-granularity visual decision path according to the causal graph network and the contribution.

[0034] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the interpretable analysis method for multi-agent decision-making as described in any one of the above is implemented.

[0035] The present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the interpretable analysis method for multi-agent decision-making as described in any one of the above is implemented.

[0036] The present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the interpretable analysis method for multi-agent decision-making as described in any one of the above is implemented.

[0037] The interpretable analysis method, device, electronic device, non-transitory computer-readable storage medium, and computer program product for multi-agent decision-making provided by the present invention input the real-time state data of multiple agents during the task execution process into a pre-trained time series encoder. The real-time state data is converted into a causal graph network by the pre-trained time series encoder. The causal graph network includes the causal relationships between agents and between agents and task objectives, so as to coarsely reflect the influence of agents on the decision. Then, according to the real-time state data and the real-time decision data, the contribution of each state feature of each agent to the decision is determined respectively. According to the causal graph network and the contribution of each state feature of each agent to the decision, a multi-granularity visual decision path is generated, which can finely reflect the influence of the state features of agents on the decision, realizing the multi-granularity causal interpretable analysis of agent decision-making and improving the interpretable ability of decision causality. Description of the Drawings

[0038] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0039] Figure 1It is one of the schematic flowcharts of the interpretable analysis method for multi-agent decision-making provided by the present invention.

[0040] Figure 2 It is the schematic diagram of the auto-encoding process in the interpretable analysis method for multi-agent decision-making provided by the present invention.

[0041] Figure 3 It is the second schematic flowchart of the interpretable analysis method for multi-agent decision-making provided by the present invention.

[0042] Figure 4 It is the schematic structural diagram of the interpretable analysis device for multi-agent decision-making provided by the present invention.

[0043] Figure 5 It is the schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0044] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0045] The following will describe Figures 1 - 5 the interpretable analysis method, device, electronic device, non-transitory computer-readable storage medium, and computer program product for multi-agent decision-making of the present invention.

[0046] Figure 1 It is one of the schematic flowcharts of the interpretable analysis method for multi-agent decision-making provided by the present invention. As Figure 1 shown, the method includes the following:

[0047] Step 102, obtaining real-time state data and real-time decision data of multiple agents during the execution of tasks.

[0048] Among them, the real-time state data refers to the real-time state data of the agent during the execution of the task. The state data may include the state of the agent itself, the surrounding environment, and the task objective, etc. The real-time decision data refers to the real-time decision data of the agent during the execution of the task. The decision data refers to the action taken by the agent according to the state through the reinforcement learning algorithm.

[0049] Specifically, multiple agents execute tasks through the reinforcement learning algorithm, and obtain real-time state data and real-time decision data during the execution of the tasks.

[0050] Step 104, input the real-time status data into a pre-trained time series encoder, and convert the real-time status data into a causal graph network through the pre-trained time series encoder; the causal graph network contains the causal relationships between agents and between agents and task objectives.

[0051] Among them, the task objective can be one or more.

[0052] In one embodiment, the causal graph network can be an N×M-dimensional matrix. N represents the number of agents. M represents the number of task objectives.

[0053] In one embodiment, the pre-trained time series encoder can be composed of a multi-layer attention Transformer network.

[0054] In one embodiment, during the process of multi-agent task execution, a real-time status data sequence segment of a preset number of time steps can be input into the pre-trained time series encoder in real time, and the real-time status data sequence segment is converted into a causal graph network through the pre-trained time series encoder to obtain a real-time causal graph network that changes over time.

[0055] For example: if the preset number is 4, then a real-time status data sequence segment of 4 preset time steps is input into the pre-trained time series encoder each time. That is, the real-time status data at the current moment and the previous 3 moments are input into the pre-trained time series encoder, and the time interval between adjacent moments is equal to the preset time step.

[0056] In one embodiment, the pre-trained time series encoder can extract key features from the real-time status data and convert the real-time status data into a causal graph network according to the key features.

[0057] Step 106, determine the contribution of each state feature of each agent to the decision according to the real-time status data and the real-time decision data.

[0058] Step 108, generate a multi-granularity visual decision path according to the causal graph network and the contribution of each state feature of each agent to the decision.

[0059] In one embodiment, important agents can be determined according to the causal graph network, important state features corresponding to each agent can be determined according to the contribution of each state feature of each agent to the decision, and a multi-granularity visual decision path can be generated according to the important agents and the important state features.

[0060] The above-mentioned interpretable analysis method for multi-agent decision-making inputs the real-time state data of multiple agents during the task execution process into a pre-trained time series encoder. The pre-trained time series encoder converts the real-time state data into a causal graph network, which contains the causal relationships between agents and between agents and task goals, thus being able to coarsely reflect the influence of agents on decisions. Then, based on the real-time state data and real-time decision data, the contribution of each state feature of each agent to the decision is determined. According to the causal graph network and the contribution of each state feature of each agent to the decision, a multi-granularity visual decision path is generated, which can finely reflect the influence of the state features of agents on decisions, realizing multi-granularity causal interpretable analysis of agent decisions and improving the interpretability of decision causality. In addition, it also improves the transparency of the decision-making process: by introducing a time-dynamic causal graph network, the system can provide comprehensive agent causal relationships and clear decision paths, so that users can understand and verify the basis of each decision, enhancing trust in the system. It can also enhance the interpretability of the system: using the time-dynamic causal graph network to extract important agents, display the decision path and basis, making each decision-making process can be clearly explained and understood, thus providing a detailed explanation of the complex decision-making process. It can also enhance the multi-agent decision-making ability: when the credit assignment to agents is incorrect and some agents in the team become lazy, the lazy agents learn sub-optimal strategies and do not cooperate to achieve the common goal of the team. Multi-agent causal relationship estimation can be used to detect and punish lazy agents, thereby improving the overall efficiency of multi-agents.

[0061] In one embodiment, the training steps of the pre-trained time series encoder include: obtaining multiple segments of sample state data; each segment of sample state data contains the state data of multiple sample agents under a time series; respectively dividing each segment of sample state data into sequence segments containing multiple preset time steps; in each round of iteration, selecting multiple sequence segments and inputting them into the time series encoder to be trained, outputting a predicted causal graph network, inputting the predicted causal graph network and the selected sequence segments into the decoder to be trained, and outputting the predicted state data of the next moment corresponding to each selected sequence segment; according to the difference between the predicted state data of the next moment and the real state data of the next moment in the sample state data, adjusting the parameters of the time series encoder to be trained and the decoder to be trained, entering the next round of iteration, until the iteration stops, and obtaining the trained time series encoder.

[0062] Wherein, the next moment refers to the next moment after the end moment of the time period corresponding to the input sequence segment.

[0063] In one embodiment, the time series encoder to be trained can be composed of a multi-layer attention Transformer network.

[0064] In one embodiment, the decoder to be trained can be composed of a Recurrent Neural Network (RNN).

[0065] In one embodiment, the decoder can decode the state data in the sequence segment according to the context information of the sequence segment and in combination with the causal graph network to obtain the predicted state data at the next moment.

[0066] In one embodiment, S segments of sample state data are obtained, and each segment of sample state data contains the state data of N sample agents in a time series with T time steps. . Each segment of sample state data is respectively divided into sequence segments containing multiple preset time steps. For example, if the preset time step t' is 4, the sample state data corresponding to time 1 to time 4, time 2 to time 5, and time 3 to time 6, etc. in each segment of sample state data can be used as sequence segments. The time series encoder to be trained can convert the input sequence segment into a causal graph network . The predicted causal graph network and the selected sequence segment from time t to before time t + t' are input into the decoder to be trained, and the predicted state data at time t + t' is output , where represents the current interference factor.

[0067] In one embodiment, the total loss function for each round of iteration is:

[0068]

[0069] where L represents the total loss function. s represents the order of the sample state data. S represents the number of segments of the sample state data. t represents the starting time of the selected sequence segment. T represents the total time length of each segment of sample state data (i.e., the number of time steps included). t' represents the preset time step. represents the loss function. represents the true state data at the next moment. represents the causal graph network output by the time series encoder. represents the selected sequence segment. represents the predicted state data at the next moment output by the decoder. represents the sparse penalty, which is used to make the causal graph network as sparse as possible. represents the parameters of the decoder to be trained. represents the parameters of the time series encoder to be trained.

[0070] As Figure 2 shown, it is a schematic diagram of the entire auto - encoding process. The selected sequence fragment is input into the time - series encoder to be trained . The causal graph network G is output. Then, the causal graph network G and the selected sequence fragment are input into the decoder to be trained , and the predicted state data at the next moment is output .

[0071] In the above - mentioned embodiment, according to the sequence fragment in the sample state data being input into the time - series encoder to be trained to predict the causal graph network, and then through the decoder, decoding is performed based on the predicted causal graph network and the sequence fragment to obtain the predicted state data at the next moment. Then, according to the difference between the predicted state data at the next moment and the real - state data, the time - series encoder and the decoder are iteratively trained. Finally, a time - series encoder for generating the causal graph network can be obtained, which can coarsely reflect the influence of the agent on the decision - making.

[0072] In one embodiment, obtaining multiple segments of sample state data includes: through a pre - trained multi - agent reinforcement learning model, performing multiple rounds of dynamic decision - making on multiple sample agents; and obtaining multiple segments of sample state data according to the state data of multiple sample agents during each round of dynamic decision - making.

[0073] Among them, the pre - trained multi - agent reinforcement learning model can adopt reinforcement learning algorithms such as MAPPO. The system continuously interacts with the environment and gradually optimizes its decision - making process to achieve the expected goal. The state data during each round of dynamic decision - making respectively corresponds to a segment of sample state data.

[0074] In the above - mentioned embodiment, through a pre - trained multi - agent reinforcement learning model, performing multiple rounds of dynamic decision - making on multiple sample agents, and according to the state data of multiple sample agents during each round of dynamic decision - making, multiple segments of sample state data can be obtained for training the time - series encoder.

[0075] In one embodiment, determining the contribution of each state feature of each agent to the decision - making according to the real - time state data and real - time decision data includes: taking each state feature in the real - time state data as the target state feature respectively, and determining the state - feature subsets that can be composed of the state features other than the target state feature; according to the real - time decision data, determining the first decision corresponding to each state - feature subset respectively, and the second decision corresponding to each state - feature subset plus the target state feature; and determining the contribution of the target state feature to the decision - making according to the difference between the first decision and the second decision corresponding to each state - feature subset, so as to obtain the contribution of each state feature of each agent to the decision - making.

[0076] For example, the state features of an agent include A, B, C, and D. Taking A as the target state feature, the state feature subsets that can be formed by each state feature except A include {}, {B}, {C}, {D}, {B, C}, {C, D}, {B, D}, and {B, C, D}. For the state feature subset {B, C}, the corresponding first decision is f({B, C}), and the corresponding second decision after adding the target state feature A is f({A, B, C}).

[0077] In one embodiment, if there is no decision corresponding to the state feature subset in the real-time decision data, the state data corresponding to the state feature subset is input into the multi-agent reinforcement learning model to obtain the corresponding decision.

[0078] In the above embodiment, each state feature in the real-time state data is respectively used as the target state feature, the state feature subsets that can be formed by each state feature except the target state feature are determined, according to the real-time decision data, the first decision corresponding to each state feature subset and the second decision corresponding to each state feature subset after adding the target state feature are determined, and according to the difference between the first decision and the second decision corresponding to each state feature subset, the contribution of the target state feature to the decision is determined, so that the contribution of each state feature of each agent to the decision can be accurately obtained, and thus the influence of the state features of the agent on the decision can be reflected in a fine-grained manner.

[0079] In one embodiment, determining the contribution of the target state feature to the decision according to the difference between the first decision and the second decision corresponding to each state feature subset includes: respectively for each state feature subset, calculating the marginal contribution of the target state feature to the state feature subset according to the difference between the corresponding first decision and the second decision; calculating the weight of each state feature subset, and according to the weights of each state feature subset, performing weighted summation on the corresponding marginal contributions to obtain the contribution of the target state feature to the decision.

[0080] In one embodiment, the weight of the state feature subset can be calculated according to the total number of features and the number of features in the state feature subset.

[0081] In the above embodiment, respectively for each state feature subset, calculating the marginal contribution of the target state feature to the state feature subset according to the difference between the corresponding first decision and the second decision, calculating the weight of each state feature subset, and according to the weights of each state feature subset, performing weighted summation on the corresponding marginal contributions to obtain the contribution of the target state feature to the decision, so that the contribution of each state feature of each agent to the decision can be accurately obtained, and thus the influence of the state features of the agent on the decision can be reflected in a fine-grained manner.

[0082] In one embodiment, a multi-granularity visual decision path is generated according to the causal graph network and the contribution of each state feature of each agent to the decision respectively, including: determining important agents according to the causal graph network; determining important state features corresponding to each agent according to the contribution of each state feature of each agent to the decision respectively; and generating a multi-granularity visual decision path according to the important agents and the important state features.

[0083] Among them, an important agent refers to an agent that has a significant impact on achieving the task goal. An important state feature refers to a state feature that has a significant impact on the agent achieving the task goal.

[0084] In one embodiment, the important agents can be highlighted in the multi-granularity visual decision path.

[0085] In one embodiment, decision trees corresponding to each agent can be generated according to the important state features corresponding to each agent, and the decision trees corresponding to each agent are displayed in the multi-granularity visual decision path.

[0086] For example: if the important state feature of a certain agent is speed, the corresponding decision tree can be if the speed is greater than the preset threshold, go left, and if the speed is less than or equal to the preset threshold, go right.

[0087] In the above embodiment, important agents are determined according to the causal graph network, important state features corresponding to each agent are determined according to the contribution of each state feature of each agent to the decision respectively; a multi-granularity visual decision path is generated according to the important agents and the important state features, realizing multi-granularity causal interpretable analysis of agent decision-making and improving the interpretability of decision-making causality.

[0088] As Figure 3 shown, it is the second flowchart of the interpretable analysis method for multi-agent decision-making provided by the present invention. The following steps are included:

[0089] Step 302, multi-agent reinforcement learning decision generation.

[0090] Step 304, generation of an auto-encoded agent causal graph network.

[0091] Step 306, extraction of fine-grained key features for feature contribution analysis.

[0092] Step 308, multi-granularity visualization of multi-agent decisions.

[0093] Next, an interpretable analysis device for multi-agent decision-making provided by the present invention is described. The interpretable analysis device for multi-agent decision-making described below can be mutually referred to with the interpretable analysis method for multi-agent decision-making described above.

[0094] As shown Figure 4 in the figure, an interpretable analysis device 400 for multi-agent decision-making is provided, including the following modules:

[0095] A data acquisition module 402, configured to acquire real-time state data and real-time decision data of multiple agents during task execution.

[0096] A coarse-grained causal discovery module 404, configured to input the real-time state data into a pre-trained time series encoder, and convert the real-time state data into a causal graph network through the pre-trained time series encoder; the causal graph network includes causal relationships between agents and between agents and task objectives.

[0097] A fine-grained feature importance analysis module 406, configured to determine the contribution of each state feature of each agent to the decision according to the real-time state data and the real-time decision data.

[0098] A decision path visualization module 408, configured to generate a multi-granularity visual decision path according to the causal graph network and the contribution.

[0099] In one embodiment, the data acquisition module 402 is further configured to acquire multiple segments of sample state data; each segment of sample state data includes state data of multiple sample agents under a time series; each segment of sample state data is respectively divided into sequence segments including multiple preset time steps.

[0100] The coarse-grained causal discovery module 404 is further configured to, in each round of iteration, select multiple sequence segments and input them into a time series encoder to be trained, output a predicted causal graph network, input the predicted causal graph network and the selected sequence segments into a decoder to be trained, and output predicted state data of the next moment corresponding to each selected sequence segment; according to the difference between the predicted state data of the next moment and the real state data of the next moment in the sample state data, adjust the parameters of the time series encoder to be trained and the decoder to be trained, enter the next round of iteration, and stop iterating until the trained time series encoder is obtained.

[0101] In one embodiment, the data acquisition module 402 is further configured to perform multi-round dynamic decision-making on multiple sample agents through a pre-trained multi-agent reinforcement learning model; and obtain multiple segments of sample state data according to the state data of the multiple sample agents during each round of dynamic decision-making.

[0102] In one embodiment, the fine-grained feature importance analysis module 406 is further configured to use each state feature in the real-time state data as a target state feature respectively, and determine state feature subsets that can be formed by the respective state features other than the target state feature; according to the real-time decision data, determine the first decision corresponding to each state feature subset respectively, and the second decision corresponding to each state feature subset after adding the target state feature; according to the difference between the first decision and the second decision corresponding to each state feature subset respectively, determine the contribution of the target state feature to the decision, and obtain the contribution of each state feature of each agent to the decision.

[0103] In one embodiment, the fine-grained feature importance analysis module 406 is further configured to, for each state feature subset respectively, calculate the marginal contribution of the target state feature to the state feature subset according to the difference between the corresponding first decision and second decision; calculate the weight of each state feature subset, and perform weighted summation on the corresponding marginal contributions according to the weights of the state feature subsets, to obtain the contribution of the target state feature to the decision.

[0104] In one embodiment, the decision path visualization module 408 is further configured to determine important agents according to the causal graph network; determine the important state features corresponding to each agent according to the contribution of each state feature of each agent to the decision; generate a multi-granularity visualized decision path according to the important agents and the important state features.

[0105] Figure 5 An entity structure diagram of an electronic device is illustrated, such as Figure 5 shown. The electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 complete mutual communication through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute an interpretable analysis method for multi-agent decision-making. The method includes: obtaining real-time state data and real-time decision data of multiple agents during the execution of tasks; inputting the real-time state data into a pre-trained time series encoder, and converting the real-time state data into a causal graph network through the pre-trained time series encoder; the causal graph network includes causal relationships between agents and between agents and task objectives; according to the real-time state data and the real-time decision data, determine the contribution of each state feature of each agent to the decision; generate a multi-granularity visualized decision path according to the causal graph network and the contribution of each state feature of each agent to the decision.

[0106] In addition, when the logical instructions in the above-mentioned memory 530 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods according to the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0107] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the interpretable analysis method for multi-agent decision-making provided by the above-mentioned various methods. The method includes: obtaining real-time state data and real-time decision data of multiple agents during the execution of tasks; inputting the real-time state data into a pre-trained time series encoder, and converting the real-time state data into a causal graph network through the pre-trained time series encoder; the causal graph network contains causal relationships between agents and between agents and task goals; determining the contribution of each state feature of each agent to the decision respectively according to the real-time state data and the real-time decision data; generating a multi-granularity visual decision path according to the causal graph network and the contribution of each state feature of each agent to the decision.

[0108] In yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the interpretable analysis method for multi-agent decision-making provided by the above-mentioned various methods. The method includes: obtaining real-time state data and real-time decision data of multiple agents during the execution of tasks; inputting the real-time state data into a pre-trained time series encoder, and converting the real-time state data into a causal graph network through the pre-trained time series encoder; the causal graph network contains causal relationships between agents and between agents and task goals; determining the contribution of each state feature of each agent to the decision respectively according to the real-time state data and the real-time decision data; generating a multi-granularity visual decision path according to the causal graph network and the contribution of each state feature of each agent to the decision.

[0109] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0110] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An interpretable analysis method for multi-agent decision-making, characterized in that, Including: Obtain the real-time status data and real-time decision-making data of multiple agents during the execution of tasks; Input the real-time status data into a pre-trained time series encoder, and convert the real-time status data into a causal graph network through the pre-trained time series encoder; the causal graph network contains the causal relationships between the agents and between the agents and the task objectives; Determine the contribution of each state feature of each agent to the decision-making according to the real-time status data and the real-time decision-making data; Generate a multi-granularity visual decision-making path according to the causal graph network and the contribution of each state feature of each agent to the decision-making, including: Determine important agents according to the causal graph network; determine the important state features corresponding to each agent according to the contribution of each state feature of each agent to the decision-making; generate a multi-granularity visual decision-making path according to the important agents and the important state features; wherein, the important state feature of the important agent is speed; The training steps of the pre-trained time series encoder include: Obtain multiple segments of sample status data; each segment of the sample status data contains the status data of multiple sample agents under a time series; Divide each segment of the sample status data into sequence segments including multiple preset time steps respectively; In each round of iteration, select multiple sequence segments and input them into the time series encoder to be trained, output the predicted causal graph network, input the predicted causal graph network and the selected sequence segments into the decoder to be trained, and output the predicted status data of the next moment corresponding to each selected sequence segment; Adjust the parameters of the time series encoder to be trained and the decoder to be trained according to the difference between the predicted status data of the next moment and the real status data of the next moment in the sample status data, enter the next round of iteration, and stop the iteration until the trained time series encoder is obtained.

2. The interpretable analysis method for multi-agent decision-making according to claim 1, wherein The obtaining of multiple segments of sample status data includes: Perform multi-round dynamic decision-making on the multiple sample agents through a pre-trained multi-agent reinforcement learning model; Obtain multiple segments of sample status data according to the status data of the multiple sample agents during each round of dynamic decision-making.

3. The interpretable analysis method for multi-agent decision-making according to claim 1, characterized in that The determining of the contribution of each state feature of each agent to the decision-making according to the real-time status data and the real-time decision-making data includes: Respectively use each state feature in the real-time status data as the target state feature, and determine the state feature subsets that can be formed by the state features other than the target state feature; Determine the first decision corresponding to each state feature subset and the second decision corresponding to each state feature subset plus the target state feature according to the real-time decision-making data; Determine the contribution of the target state feature to the decision-making according to the difference between the first decision and the second decision corresponding to each state feature subset, and obtain the contribution of each state feature of each agent to the decision-making.

4. The interpretable analysis method for multi-agent decision-making according to claim 3, characterized in that Determining the contribution of the target state feature to the decision based on the differences between the first decision and the second decision corresponding to each of the state feature subsets respectively includes: For each of the state feature subsets, calculating the marginal contribution of the target state feature to the state feature subset according to the difference between the corresponding first decision and the second decision; Calculating the weights of each of the state feature subsets, and performing weighted summation on the corresponding marginal contributions according to the weights of each of the state feature subsets to obtain the contribution of the target state feature to the decision.

5. An interpretable analysis device for multi-agent decision-making, characterized in that, Including: A data acquisition module, configured to acquire real-time state data and real-time decision data of multiple agents during the execution of tasks; A coarse-grained causal discovery module, configured to input the real-time state data into a pre-trained time series encoder, and convert the real-time state data into a causal graph network through the pre-trained time series encoder; the causal graph network includes causal relationships between the agents and between the agents and the task objective; A fine-grained feature importance analysis module, configured to determine the contribution of each state feature of each agent to the decision according to the real-time state data and the real-time decision data; A decision path visualization module, configured to generate a multi-granularity visualization decision path according to the causal graph network and the contribution, where, including: Determining important agents according to the causal graph network; determining important state features corresponding to each agent according to the contribution of each state feature of each agent to the decision; generating a multi-granularity visualization decision path according to the important agents and the important state features; where the important state feature of the important agent is speed; The training steps of the pre-trained time series encoder include: Obtaining multiple segments of sample state data; each segment of the sample state data includes state data of multiple sample agents under a time series; Dividing each segment of the sample state data into sequence segments including multiple preset time steps respectively; In each round of iteration, selecting multiple of the sequence segments and inputting them into the time series encoder to be trained, outputting a predicted causal graph network, inputting the predicted causal graph network and the selected sequence segments into the decoder to be trained, and outputting predicted state data of the next moment corresponding to each of the selected sequence segments; Adjusting the parameters of the time series encoder to be trained and the decoder to be trained according to the difference between the predicted state data of the next moment and the real state data of the next moment in the sample state data, entering the next round of iteration until the iteration stops, and obtaining a trained time series encoder.

6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the interpretable analysis method for multi-agent decision-making according to any one of claims 1 to 4.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the interpretable analysis method for multi-agent decision-making according to any one of claims 1 to 4.

8. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the interpretable analysis method for multi-agent decision-making according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Driving situation reasoning method based on metadata driving and causal analysis theory

    CN117217314A

  • Multi-agent decision-making method and device, electronic equipment and storage medium

    CN118036645A