Multi-agent control method and device for causal experience playback, equipment and medium
Through the multi-agent control method of causal experience replay, a causal graph is constructed and action vector weight value is assigned, which solves the problems of low sample data utilization and unexplainable decision-making in traditional methods, and improves the efficiency and interpretability of multi-agent cluster formation control.
Patent Information
- Application Number
- CN202411893295.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-20
AI Technical Summary
The traditional multi-agent cluster formation control method requires a large number of samples to learn effective strategies. The utilization rate of sample data is low and it is difficult to explain its decision-making process and results, which limits the application of multi-agent cluster formation control method.
By obtaining the empirical data set collected when multiple agents perform cluster formation tasks, a causal graph is constructed to identify the causal relationship between actions and rewards, and assign weight values to each action vector, update the action vector subset and train the control strategy model to optimize the execution of multi-agent cluster formation tasks.
It enhances the deep learning ability of the agent, improves the interpretability of the agent's decision-making, improves the efficiency of sample data utilization, and optimizes the execution effect of multi-agent cluster formation tasks.
Smart Images

Figure CN119937542A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent robot technology, and in particular to a multi-agent control method, device, equipment and medium for causal experience playback. Background Art
[0002] A multi-agent system (MAS) is a system composed of multiple agents, which can be physical entities such as drones, autonomous vehicles or robots, or virtual entities such as software agents in the network. Agents can act and make decisions autonomously, and have the capabilities of perception, communication, and learning. Multi-agent swarm formation control (also known as collective formation control or collaborative formation control) allows multiple agents to complete the formation control task of large or complex objects through mutual cooperation and collaboration. This technology has broad application prospects in logistics, warehousing, disaster relief, agriculture, industrial automation and other fields. Traditional multi-agent swarm formation control methods require a large number of samples to learn effective strategies, the utilization rate of sample data is low, and it is difficult to explain its decision-making process and results, which limits the application of multi-agent swarm formation control methods. Summary of the invention
[0003] The present invention provides a multi-agent control method, device, equipment and medium for causal experience playback, which can enhance the deep learning ability of the agent and improve the interpretability of the agent's decision-making.
[0004] The present invention provides a multi-agent control method for causal experience playback, the method comprising: Acquire an experience data set collected when multiple agents perform a cluster formation task, wherein the experience data set includes an action vector subset and a reward value subset; Based on the action vector subset and the reward value subset, generating a causal graph, the causal graph including causal relationships between action vectors in the action vector subset and reward values in the reward value subset; Determining a weight value of each of the action vectors according to a result of adjusting each of the action vectors, wherein the weight value is used to indicate the degree of influence of the action vector on the causal relationship; The action vector subset is updated according to each of the weight values, and the preset control strategy model is trained using the updated experience data set of the action vector subset; Based on the control strategy generated by the trained control strategy model, the multiple intelligent agents are controlled to perform cluster formation tasks.
[0005] Optionally, generating a causal graph based on the action vector subset and the reward value subset, wherein the causal graph includes a causal relationship between the action vectors in the action vector subset and the reward values in the reward value subset, includes: Constructing an initial causal graph and generating a plurality of nodes in the initial causal graph, each of the nodes including the action vector and a reward value corresponding to the action vector; generating an edge between every two of the nodes, the edge serving as a causal relationship between the two nodes; Based on the result of detecting the causal relationship corresponding to each edge, screening each edge and determining the type and direction of the screened edge; Determining a weight value of each edge based on the type and direction of each edge; Determine a preferred edge based on a result of comparing a preset threshold with the weight value of each edge; The initial causal graph is updated according to each of the preferred edges to obtain an updated causal graph.
[0006] Optionally, determining a weight value of each action vector according to a result of adjusting each action vector, wherein the weight value is used to represent a degree of influence of the action vector on the causal relationship, includes: adjusting an action vector of a target node among the plurality of nodes; Determine an average processing effect value based on the action vectors before and after adjustment, wherein the average processing effect value is used to indicate the degree to which the action vector of the target node affects the reward value of the target node; Based on the average processing effect value, a weight value of the action vector of the target node is determined.
[0007] Optionally, updating the action vector subset according to each of the weight values, and training a preset control strategy model using the updated experience data set of the action vector subset, includes: Determine the priority of the experience data group based on the weight value of the action vector and the error value of the reward value, wherein the experience data group is composed of experience data including the action vector and the reward value collected in a cluster formation task; Using each of the experience data groups to train the control strategy model, and based on the priority of each of the experience data groups, determining a sampling probability of each of the experience data groups being sampled when training the control strategy model; When training the control strategy model, an importance sampling weight is determined based on a functional relationship between the amount of cache space storing the experience data set and the sampling probability and is used to adjust the sampling probability corresponding to a target experience data set among the multiple experience data sets.
[0008] Optionally, the control strategy generated based on the trained control strategy model controls the plurality of intelligent agents to perform the cluster formation task, including: When the multiple agents perform a cluster formation task, multiple action strategies are generated through a local control strategy model corresponding to each of the agents, wherein each of the agents performs an action corresponding to a target action strategy among the multiple action strategies; Updating the agent state and agent action corresponding to each of the action strategies; Generate a current state vector based on each updated agent state, and generate a current action vector based on each updated agent action; Using the current state vector and the current action vector as inputs of the trained control strategy model, and outputting a first output value through the trained control strategy model; Using the sample state vector and the sample action vector in the experience data set as inputs of the trained control strategy model, and outputting a second output value through the trained control strategy model; Based on the loss function relationship between the first output value and the second output value, updating the parameters of the trained control strategy model; The control strategy model after parameter update is used to generate a control strategy, and the plurality of intelligent agents are controlled according to the control strategy to perform the cluster formation task.
[0009] Optionally, the multi-agent control method of causal experience playback further comprises the step of generating the action vector subset; The generating the action vector subset comprises: Generate a time series based on the behavior data of the agent in the experience data set; Dividing the time series into a plurality of subsequences; Generating multiple clusters and merging each of the subsequences into a target cluster in the multiple clusters to obtain a subsequence after clustering processing; The action vector subset is generated, and each subsequence after the clustering process is used as an action vector in the action vector subset.
[0010] Optionally, the behavior vector corresponding to each time point in the time series includes multiple variables; The step of dividing the time series into a plurality of subsequences comprises: Based on the multiple variables, generate a precision matrix, each element in the precision matrix is used to represent the conditional independence relationship between each of the variables; Dividing the time series into multiple time segments and generating multiple cluster centers; Determine the similarity relationship between the time segment and the cluster center based on the current precision matrix; Based on the similarity relationship between the time segments and the cluster centers, each of the time segments is assigned to a cluster corresponding to the target cluster center; The cluster corresponding to each of the cluster centers is used as the subsequence, and the multiple subsequences are used as the result of the clustering process; Based on the result of the clustering process, updating the current precision matrix; The updated precision matrix is used as the current precision matrix, and the step of determining the similarity relationship between the time segment and the cluster center according to the current precision matrix is returned to be executed until the clustering process reaches a preset end condition, and a plurality of the subsequences are output.
[0011] The present invention also provides a multi-agent control device for causal experience playback, the device comprising: An acquisition module is used to acquire an experience data set collected when multiple agents perform cluster formation tasks, wherein the experience data set includes an action vector subset and a reward value subset; A generating module, configured to generate a causal graph based on the action vector subset and the reward value subset, wherein the causal graph includes a causal relationship between the action vectors in the action vector subset and the reward values in the reward value subset; A calculation module, used for determining a weight value of each of the action vectors according to a result of adjusting each of the action vectors, wherein the weight value is used for indicating the degree of influence of the action vector on the causal relationship; A training module, used for updating the action vector subset according to each of the weight values, and training a preset control strategy model using the updated experience data set of the action vector subset; The control module is used to control the multiple intelligent agents to perform cluster formation tasks based on the control strategy generated by the trained control strategy model.
[0012] The present invention also provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the multi-agent control method of causal experience playback as described in any of the above items.
[0013] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the multi-agent control method of causal experience playback as described in any of the above items.
[0014] The present invention has at least the following beneficial effects: In the technical solution of the present application, firstly, by extracting action vectors and reward values from the experience data set, a causal graph is constructed to identify the causal relationship between action and reward, and this is used as the basis for decision-making to enhance the interpretability of the results. Secondly, by assigning a weight value to each action vector, the degree of influence of each action on the causal relationship is quantified. This weight assignment enables the model to pay more attention to those actions that are more critical to the success of the task, thereby improving the utilization efficiency of the sample data. Then, by updating the action vector subset and training the control strategy model with the updated data set, the experience with important causal influence is trained first, helping the model to quickly identify the corresponding causal relationship. Finally, the control strategy model is trained with the updated action vector subset to generate a more effective control strategy, thereby optimizing the execution of the multi-agent cluster formation task, which helps to improve the application scope and effect of the multi-agent cluster formation control method. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings are used to provide a further understanding of the technical solution of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the technical solution of the present invention and do not constitute a limitation on the technical solution of the present invention.
[0016] Figure 1 is a flowchart of the steps of the multi-agent control method for causal experience playback provided in this embodiment; Figure 2 is a flowchart of step S102 in the multi-agent control method for causal experience playback provided in this embodiment; Figure 3 is a flowchart of step S103 in the multi-agent control method for causal experience playback provided in this embodiment; Figure 4 is a flowchart of step S104 in the multi-agent control method for causal experience playback provided in this embodiment; Figure 5 is a flowchart of step S105 in the multi-agent control method for causal experience playback provided in this embodiment; Figure 6 is another step flow chart of the multi-agent control method for causal experience playback provided in this embodiment; Figure 7 is a flowchart of step S602 in the multi-agent control method for causal experience playback provided in this embodiment; Figure 8 It is a flow chart of the steps of a multi-agent control method for realizing causal experience replay in an application scenario; Fig. 9 It is a causal experience replay flow chart of a multi-agent control method for realizing causal experience replay in an application scenario; Fig.10 It is a first effect schematic diagram of a multi-agent control method for realizing causal experience playback in an application scenario; Fig.11 It is a schematic diagram of the second effect of a multi-agent control method for realizing causal experience playback in an application scenario; Fig.12 It is a schematic diagram of the third effect of a multi-agent control method for realizing causal experience playback in an application scenario; Fig.13 It is a schematic diagram of the fourth effect of a multi-agent control method for realizing causal experience playback in an application scenario; Fig.14 is a schematic diagram of the structure of a multi-agent control device for causal experience playback provided in this embodiment; Fig.15 It is a schematic diagram of the structure of the electronic device provided in this embodiment. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0018] Before describing the embodiments of the technical solution of the present application, the technical terms involved in the technical solution of the present application are first explained.
[0019] A multi-agent system (MAS) is a system composed of multiple agents that can act and make decisions autonomously and have the capabilities of perception, communication, and learning. Cluster formation control (also known as collective formation control or collaborative formation control) refers to the task of completing formation control of large or complex objects by multiple agents through mutual cooperation and collaboration. In the cluster formation control task, each agent is responsible for different formation control positions and operations, and completes the formation control task together through collaboration and information sharing.
[0020] In the related technical field, an architecture called Deep Implicit Coordination Graphs (DICG) was proposed to solve the coordination problem in Multi-agent Reinforcement Learning (MARL). By implicitly learning the coordination relationship between agents, the efficiency and effectiveness of multi-agent tasks are improved. However, due to the high model complexity and implicit coordination characteristics of DICG, the method has poor interpretability and it is difficult to intuitively understand how the model makes specific decisions. DICG does not distinguish between inputs and learns both non-causal and causal factors in the input, which will affect the overall performance of the model and may lead to insufficient generalization of the model in different scenarios. In addition, since the strategy of each agent will change continuously, DICG may require more training steps to achieve stable results, which may be a disadvantage in resource-constrained scenarios.
[0021] The researchers in this application found that DICG learns the coordination relationship between agents in an implicit way, rather than through explicit rules or constraints. Although this implicit learning method can capture the complex interactions between agents, it also makes the internal working mechanism of the model more difficult to explain. Moreover, DICG considers both causal and non-causal factors when inputting, and does not reconstruct the causal mechanism with invariance. These unprocessed data may have a negative impact on the control model of cluster formation control, resulting in a decrease in model performance. In addition, DICG introduces an implicit coordination graph, which is designed so that the model can dynamically adjust the coordination strategy between agents. However, this also increases the complexity of the algorithm, especially when it is necessary to dynamically adjust the interaction of multiple agents. This complexity will increase the uncertainty of the training process, making it difficult for the model to converge stably within a reasonable time.
[0022] In order to solve the defects in the existing technical solutions, the technical solution of the present application proposes a multi-agent control method, device, equipment and medium for causal experience playback. First, by extracting action vectors and reward values from the experience data set, a causal graph is constructed to identify the causal relationship between action and reward, and this is used as the basis for decision-making to enhance the interpretability of the results. Secondly, by assigning a weight value to each action vector, the degree of influence of each action on the causal relationship is quantified. This weight assignment enables the model to pay more attention to those actions that are more critical to the success of the task, thereby improving the utilization efficiency of sample data. Then, by updating the action vector subset and training the control strategy model with the updated data set, the experience with important causal influence is trained first, helping the model to quickly identify the corresponding causal relationship. Finally, the control strategy model is trained with the updated action vector subset to generate a more effective control strategy, thereby optimizing the execution of the multi-agent cluster formation task, which helps to improve the application scope and effect of the multi-agent cluster formation control method. The embodiments provided by the technical solution of the present application are as follows.
[0023] Please refer to Figure 1 , Figure 1 It is a flowchart of the steps of the multi-agent control method of causal experience replay.
[0024] This embodiment provides a multi-agent control method for causal experience playback, the method comprising: S101. Obtain an experience data set collected when multiple intelligent agents perform a cluster formation task, where the experience data set includes an action vector subset and a reward value subset.
[0025] S102. Generate a causal graph based on the action vector subset and the reward value subset, where the causal graph includes a causal relationship between the action vectors in the action vector subset and the reward values in the reward value subset.
[0026] S103 . Determine a weight value of each action vector according to a result of adjusting each action vector, where the weight value is used to indicate the degree of influence of the action vector on the causal relationship.
[0027] S104, updating the action vector subset according to each weight value, and using the updated experience data set of the action vector subset to train the preset control strategy model.
[0028] S105. Based on the control strategy generated by the trained control strategy model, control multiple intelligent agents to perform cluster formation tasks.
[0029] In step S102 of some embodiments, the PC algorithm (Pearl's Causal Inference) is used to construct a causal graph, or the RFCI (Really Fast Causal Inference) algorithm is used to generate a partial ancestral graph (PAG).
[0030] In some embodiments, the experience data set is saved by the experience buffer, and the experience buffer can be used to store the states of all agents at different times. , the experience buffer stores all the experience quintuples of the agent:
[0031] in, For the intelligent agent The state, action, and reward corresponding to each moment. for The state corresponding to the moment, for The status termination signal at the moment indicates whether the formation has reached its target point.
[0032] During training, each agent uses experience-based and Loss, trained by prioritizing experience replay using a batch of experiences from the experience buffer that contains both the data before and after the intervention.
[0033] Please refer to Figure 2 , Figure 2 It is a flow chart of step S102 in the multi-agent control method of causal experience playback.
[0034] In some embodiments, step S102 includes: S201, constructing an initial causal graph and generating a plurality of nodes in the initial causal graph, each node including an action vector and a reward value corresponding to the action vector.
[0035] S202: Generate an edge between every two nodes, where the edge serves as a causal relationship between the two nodes.
[0036] S203: Based on the result of detecting the causal relationship corresponding to each edge, screen each edge and determine the type and direction of the screened edge.
[0037] S204: Determine a weight value of each edge based on the type and direction of each edge.
[0038] S205 : Determine a preferred edge based on the result of comparing the preset threshold with the weight value of each edge.
[0039] S206. Update the initial causal graph according to each preferred edge to obtain an updated causal graph.
[0040] It can be understood that, in this embodiment, firstly, an observation data set containing all representative subsequences is constructed, which is recorded as in It is action vectors, Is the reward value corresponding to the action vector. Starting from all the variables in the dataset, a completely undirected graph is constructed, where each node represents a variable in the observed dataset D.
[0041] The conditional independence test is used to check whether there is a direct causal relationship between two nodes. The next two variables and If they are independent, it is believed that there is no direct causal relationship between them. The functional formula is as follows:
[0042] if and Conditionally independent of , then remove and The edge between.
[0043] After completing all necessary conditional independence tests and removing irrelevant edges, RFCI determines the direction of the retained edges through v-structure detection: if there is a node ,satisfy and and Conditionally independent of , then a v-structure is formed, indicating yes and Common descendants:
[0044] In the v-structure, It is called the tip of the v-structure, which represents the simultaneous and impact.
[0045] Finally, a partial ancestor graph is generated, with the following edge types and directions: :Determine causal relationship, indicating yes The direct cause.
[0046] : Indicates that there are unobserved confounding factors, leading to and There is a correlation between them.
[0047] :The direction of causality is uncertain, indicating may be 's cause, but it has not yet been fully determined.
[0048] If there is uncertainty about the direction of causation, e.g. , the uncertainty symbol is retained. Further causal inference may be performed based on more data to determine the final causal direction.
[0049] The pruning phase is performed after the initial partial ancestry graph is generated, with the goal of reducing the number of edges in the graph so that the final causal graph contains the most significant causal relationships rather than all possible associations.
[0050] The importance of an edge is evaluated by the average treatment effect (ATE) of the edge. An edge with a larger ATE indicates that it has a significant impact on the reward and should be retained:
[0051] For each edge in the causal graph, according to The value of is used to classify edges into different categories, such as high importance edges (significant causality), medium importance edges, and low importance edges (weak causality).
[0052] Define a threshold To decide whether to keep an edge. Based on experience, data size, and model requirements, set , if the edge weight , then delete the edge:
[0053] Importance below threshold The edges may usually be caused by noise or weakly correlated factors and do not make a significant contribution to the causal relationship.
[0054] To prevent the causal graph from being pruned too aggressively at different stages, a dynamically adjusted threshold can be used during the pruning process. For example, in the early stages of training, a lower threshold can be set. , in order to maintain more potential causal edges; as the training progresses, the threshold is gradually increased to , making the causal diagram more concise and focusing on the most important relationships.
[0055] Please refer to Figure 3 , Figure 3 It is a flow chart of step S103 in the multi-agent control method of causal experience playback.
[0056] In some embodiments, step S103 includes: S301. Adjust the action vector of a target node among multiple nodes.
[0057] S302. Determine an average processing effect value based on the action vectors before and after adjustment, wherein the average processing effect value is used to represent the degree to which the action vector of the target node affects the reward value of the target node.
[0058] S303: Determine a weight value of the action vector of the target node based on the average processing effect value.
[0059] It is understandable that this embodiment uses a causal intervention mechanism to measure the degree of influence of the causal variables in the causal graph on the reward, directly controlling a causal variable rather than passively observing it. Apply control and change its value to , to evaluate the direct causal impact of the variable, the specific expression is as follows,
[0060] Then, the effect of the intervention on the causal variable on reward was assessed by the average treatment effect size (ATE). , its ATE is as follows,
[0061] in, and are the rewards before and after the intervention, The larger it is, the more significant the impact of the intervention on the reward is, indicating that the action vector has an important causal effect on the strategy.
[0062] Please refer to Figure 4 , Figure 4 It is a flow chart of step S104 in the multi-agent control method of causal experience playback.
[0063] In some embodiments, step S104 includes: S401. Determine the priority of the experience data group based on the weight value of the action vector and the error value of the reward value, wherein the experience data group is composed of experience data including action vectors and reward values collected in a cluster formation task.
[0064] S402: Use each experience data group to train the control strategy model, and based on the priority of each experience data group, determine the sampling probability of each experience data group being sampled when training the control strategy model.
[0065] S403: When training the control strategy model, based on the functional relationship between the amount of cache space for storing the experience data set and the sampling probability, an importance sampling weight is determined and used to adjust the sampling probability corresponding to the target experience data set among the multiple experience data sets.
[0066] It is understandable that in order to give priority to learning experiences with stronger causal influence during the training of the control strategy model, the mechanism of prioritizing experience playback is adopted in this embodiment. First, the priority of the experience data group is determined, among which Greater experience sets higher priority, and at the same time, comprehensive consideration Error, for experience , its priority for:
[0067] in, For experience of Error and mean treatment effects, is a hyperparameter used to balance the impact of TD error and individual treatment effects. is a small positive number to avoid the priority being zero. This determines the sampling probability of each experience data set being sampled when training the control strategy model. The sampling probability is determined by the following calculation, so its sampling probability is:
[0068] in, Controls the effect of priority on sampling probability.
[0069] Since priority experience replay is non-uniform sampling, the sampling probability of some experiences is higher. In order to correct the deviation caused by this non-uniform sampling, the importance sampling weight is introduced. When determining the importance sampling weight, the following formula is used:
[0070] in, is the size of the experience buffer, It’s experience The sampling probability of It is a hyperparameter that controls the degree of importance sampling, and is usually increased from a small value to 1 to ensure that the deviation is corrected more and more fully as the training process progresses.
[0071] It is understandable that by adopting the causal intervention mechanism and according to the results of the causal intervention, higher weights are given to experience fragments with strong causal influence to improve the model's learning efficiency of important causal factors, and then the priority experience replay mechanism is adopted to improve the training efficiency of deep reinforcement learning by giving priority to replaying experiences with strong causal relationships with reward results.
[0072] Please refer to Figure 5 , Figure 5 It is a flow chart of step S105 in the multi-agent control method of causal experience playback.
[0073] In some embodiments, step S105 includes: S501. When multiple agents perform a cluster formation task, multiple action strategies are generated through the local control strategy model corresponding to each agent, wherein each agent performs an action corresponding to a target action strategy among the multiple action strategies.
[0074] S502: Update the agent state and agent action corresponding to each action strategy.
[0075] S503, generating a current state vector according to each updated agent state, and generating a current action vector according to each updated agent action.
[0076] S504: Use the current state vector and the current action vector as inputs of the trained control strategy model, and output a first output value through the trained control strategy model.
[0077] S505: Using the sample state vector and the sample action vector in the experience data set as inputs of the trained control strategy model, and outputting a second output value through the trained control strategy model.
[0078] S506: Based on the loss function relationship between the first output value and the second output value, update the parameters of the trained control strategy model.
[0079] S507: Generate a control strategy using the control strategy model after parameter update, and control multiple agents to perform cluster formation tasks according to the control strategy.
[0080] In some embodiments, a multi-agent formation is controlled by a MADDPG-based control strategy model. In the MADDPG-based control strategy model, a Critic and Actor network is constructed. For the Critic, each agent has its own online evaluation network. and target evaluation network ; For Actor, each agent has its own online strategy network and target strategy network Using the centralized training method, when calculating the forward propagation of the Critic, the states of all agents are spliced into a state vector , concatenate the actions of all agents into action vectors ;Will As the input of the online evaluation network, the output is a one-dimensional Value, that is , using the global information of the agent in the environment to "centralize" the training of its own evaluation network, and then using the samples in the experience buffer to obtain , Next, we construct the loss function and use the time series difference (TD) to approximate the optimal Q value. Its expression is as follows:
[0081] in, Here, the goal evaluation network is used to calculate the action taken by the agent in the next state. It should be noted that the input of each agent's goal evaluation network only contains the local state information of the agent itself.
[0082] At the same time, in order to correct the deviation caused by non-uniform sampling, the importance sampling weight is introduced , and the total loss function is obtained:
[0083] Update parameters by gradient descent :
[0084] Each agent generates its own action strategy using a local control strategy model based on the DDPG algorithm. It constructs its own Critic and Actor networks respectively. Its input is the global state information obtained by the entire formation system from the environment. The output of the Critic is the Q value of the state-action pair, and the output of the Actor is the deterministic action of the agent.
[0085] When calculating the forward propagation of its own Actor, each agent only uses its own local observation vector As input to the online policy network, it outputs a deterministic action Same as Critic, adding importance sampling weights , solve the loss function and calculate the gradient with respect to the parameters, and then use gradient descent to update the parameters. The loss function and gradient are as follows:
[0086]
[0087] It can be understood that this embodiment implements centralized training and distributed execution, and shares global information during training, so that the intelligent agent can obtain more data from other intelligent agents during training, thereby making better decisions in an environment where other intelligent agents exist, and effectively reducing instability problems caused by mutual influence between intelligent agents.
[0088] Please refer to Figure 6 , Figure 6 It is another step flow chart of the multi-agent control method of causal experience replay.
[0089] In some embodiments, the specific implementation of generating the action vector subset is achieved by the following steps: S601. Generate a time series based on the behavior data of the intelligent agent in the experience data set.
[0090] S602: Divide the time series into multiple subsequences.
[0091] S603: Generate multiple clusters and merge each subsequence into a target cluster in the multiple clusters to obtain a subsequence after clustering.
[0092] S604: Generate an action vector subset, and use each subsequence after clustering as an action vector in the action vector subset.
[0093] In some embodiments, the behavior vector corresponding to each time point in the time series includes multiple variables.
[0094] In a specific embodiment, the behavior data in the experience data set is first constructed into a multidimensional time series, denoted as ,in Indicates Then, the time series is divided into subsequences with conditional independence structure, and the K-means clustering algorithm is used to further integrate the subsequences initially divided by TICC, so as to more effectively represent and utilize the information of these subsequences in causal experience playback.
[0095] It is understandable that after the initial segmentation, a large number of subsequences may be obtained. In order to simplify the representation of these fragments, the K-means algorithm is further used to cluster these subsequences into a few representative fragments, which will serve as nodes of the causal graph.
[0096] Methods for integrating subsequences using the K-means clustering algorithm include: First select Initial cluster centers , then, for each subsequence segment , find the nearest cluster center ,
[0097] After all subsequences are assigned, the mean of each cluster is calculated as the new cluster center according to the currently assigned cluster:
[0098] When the change of cluster center is less than the threshold or the maximum number of iterations is reached, the clustering process is stopped.
[0099] Ultimately, the K-means clustering algorithm will further cluster the initially segmented subsequences into a few representative sequences. These sequences will serve as core elements in causal inference, helping the model to more efficiently capture and analyze causal relationships in time series during causal inference.
[0100] Please refer to Figure 7 , Figure 7 It is a flow chart of step S102 in the multi-agent control method of causal experience playback.
[0101] In some embodiments, step S602 includes: S701. Generate a precision matrix based on multiple variables, where each element in the precision matrix is used to represent the conditional independence relationship between each variable.
[0102] S702: Divide the time series into multiple time segments and generate multiple cluster centers.
[0103] S703: Determine the similarity relationship between the time segment and the cluster center according to the current precision matrix.
[0104] S704: Based on the similarity relationship between the time segments and the cluster centers, each time segment is assigned to a cluster corresponding to the target cluster center.
[0105] S705: taking the cluster corresponding to each cluster center as a subsequence, and taking the multiple subsequences as the result of the clustering process.
[0106] S706: Based on the result of the clustering process, the current precision matrix is updated.
[0107] S707. Use the updated precision matrix as the current precision matrix, return to the step of determining the similarity relationship between the time segment and the cluster center according to the current precision matrix, until the clustering process reaches a preset end condition, and output multiple subsequences.
[0108] In some embodiments, the above steps S701-S707 are implemented through TICC. Specifically, the conditional independence relationship in the time series segments is represented by the precision matrix, so as to discover different behavior patterns. The elements within are:
[0109] in, The precision matrix If , indicating the variable and are independent given the other variables.
[0110] At the same time, TICC uses the Expectation-Maximization (EM) algorithm to segment the time series. The specific steps are as follows: Step E: Fix the current value of the precision matrix and assign each time segment to the nearest cluster based on similarity.
[0111] Step M: Fix the cluster assignment results, update the precision matrix to maximize the log-likelihood, and use regularization to ensure the sparsity of the precision matrix:
[0112] in, is the covariance matrix, is a regularization parameter used to control sparsity. After several iterations of E and M steps, TICC divides the time series into multiple subsequences, each of which corresponds to a cluster.
[0113] It can be understood that the above embodiment combines the action sequences in the experience data set into a multivariate time series, segments it and constructs a causal graph based on it, explores the causal relationship between different actions of the intelligent agent and rewards, makes the model ignore actions that are not related to rewards, and enhances the interpretability of the model results.
[0114] See also Figure 8 , Figure 8 It is a flow chart of the steps of a multi-agent control method for realizing causal experience replay in an application scenario.
[0115] In this application scenario, the control strategy model of the multi-agent control method that implements causal experience replay is a control strategy model based on MADDPG. First, initialize the agent, evaluation system, and policy network. The agent is the entity that performs actions, the evaluation system is used to evaluate the performance of the agent, and the policy network is used to generate the action of the agent. Therefore, it is necessary to define the state space, action space, and policy function of each agent for the construction of the Critic and Actor networks in MADDPG. , its state space Contains its own state information and the perception information of the local environment. The specific expression is as follows:
[0116] in,( ) is the current position of the agent in the environment, is the current speed of the agent, is the moving direction of the agent, The environment information perceived by the agent based on the camera is obtained through a feature extraction network.
[0117] The action space represents the set of actions that the agent can perform. The specific expression is as follows:
[0118] in, Respectively represent the adjustment actions of speed and direction.
[0119] The strategy function is used to Choose the right action It is optimized through training to maximize long-term rewards. The specific expression is as follows:
[0120] in For intelligent agents The policy parameters are optimized independently for each agent.
[0121] In the process of training multi-agent formation control, a reward function needs to be defined to meet the following factors: avoiding obstacles, approaching the target point, maintaining a safe distance from other agents, maintaining the formation control formation, and maintaining speed consistency. Behaviors that meet the above factors will be rewarded, while behaviors that do not meet the above factors will be punished.
[0122] Obstacle Avoidance Rewards: This part of the reward is designed to ensure that the agent can avoid obstacles and maintain a safe distance from other agents:
[0123] in, is the importance weight of obstacle avoidance, It is an intelligent agent distances to obstacles and other agents, is the safety distance threshold, the minimum distance that the agent needs to maintain from obstacles and other agents. When a collision occurs, the agent will be given an instantaneous penalty.
[0124] Rewards for approaching the target location: During formation control, it is necessary to calculate the distance to the target point to ensure that the entire cluster can complete the formation control task:
[0125] in, is the weight coefficient for approaching the target location, is the center position of the formation, The target location for the formation control mission. When reaching the target location, an instant reward will be given.
[0126] Formation Keeping Rewards: Agent The distance between the agents should be close to the preset ideal distance. To ensure the stability of the formation control formation:
[0127] in, is the importance weight of maintaining the formation control formation, It is an intelligent agent and The actual distance between It is an intelligent agent and Ideal spacing in formation.
[0128] Speed Consistency Bonus: To avoid the formation breaking, the speeds of the agents need to be consistent. This reward is used to ensure that the agents The speed of is matched with other agents:
[0129] in, is the importance weight of speed consistency, and Agents and speed.
[0130] Combining all the rewards, we can get the agent The reward function for
[0131] Then, an experience buffer is set up to store the states of all agents at different times and save the experience dataset, which is used for training the MADDPG network.
[0132] During the specific execution, the agent generates and executes actions according to the current policy network. Update the state and state termination signal d at the next moment: According to the action of the agent, update the state of the environment and generate the state termination signal d to determine whether the termination condition is reached. d satisfies the termination condition: Check whether the state termination signal d meets the preset termination condition. If it is satisfied, the process ends; if it is not satisfied, continue to execute. Take out the time series from the experience buffer and split it: Take out the time series data from the buffer storing the agent's experience and split it as needed for subsequent analysis. Generate a causal graph based on the segmented time series: Use the segmented time series data to construct a causal graph to identify the causal relationship between variables. Intervene the causal variables in the causal graph and calculate the average treatment effect ATE: In the causal graph, intervene in the identified causal variables, and calculate the average treatment effect (ATE) before and after the intervention to evaluate the impact of the variable on the result. Prioritize experience replay sampling based on TD error and ATE to update the network: Use the temporal difference (TD) error and ATE to guide the experience replay, and give priority to the experience that is most helpful for learning to replay, so as to update the policy network. End: The end point of the process, indicating that the learning process of the agent is completed.
[0133] See also Fig. 9 , Fig. 9 The present invention is a causal experience replay flow chart of a multi-agent control method for realizing causal experience replay in an application scenario.
[0134] like Fig. 9 As shown in the figure, the training process of the causal reinforcement control strategy model includes: defining the state of the environment in which the agent is located, which is perceptible to the agent. The agent performs actions in the environment and interacts with the environment, which will lead to changes in the state. The agent decides its actions based on the policy network. The policy network is a learning model that outputs the optimal action based on the current state. The experience of the agent's interaction with the environment is stored, including state, action, reward, and next state. Data is extracted from the experience buffer and a causal graph is generated to analyze the causal relationship between different states and actions. Based on the causal graph, certain variables are intervened to evaluate their impact on the results. The evaluation network is used to evaluate the quality of the agent's actions in a specific state and help the agent learn a better strategy. The experience that is most helpful for learning is preferentially selected for replay based on the TD error and ATE (average treatment effect) to update the policy network.
[0135] See also Figures 10 to 13 In this embodiment, the effect of multi-agents performing formation control tasks is as follows: Figures 10 to 13 shown.
[0136] Please refer to Fig.14 , Fig.14 It is a schematic diagram of the structure of a multi-agent control device for causal experience playback.
[0137] This embodiment also provides a multi-agent control device for causal experience playback, the device comprising: The acquisition module 801 is used to acquire an experience data set collected when multiple agents perform cluster formation tasks, and the experience data set includes an action vector subset and a reward value subset.
[0138] The generating module 802 is used to generate a causal graph based on the action vector subset and the reward value subset, wherein the causal graph includes the causal relationship between the action vectors in the action vector subset and the reward values in the reward value subset.
[0139] The operation module 803 is used to determine the weight value of each action vector according to the result of adjusting each action vector, and the weight value is used to represent the degree of influence of the action vector on the causal relationship.
[0140] The training module 804 is used to update the action vector subset according to each weight value, and train the preset control strategy model using the experience data set after the action vector subset is updated.
[0141] The control module 805 is used to control multiple agents to perform cluster formation tasks based on the control strategy generated by the trained control strategy model.
[0142] It will be appreciated by those skilled in the art that all or some of the steps and devices in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or transient medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. As is well known to those skilled in the art, communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0143] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0144] An embodiment of the present application also provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements any of the above multi-agent control methods for causal experience playback.
[0145] refer to Fig.15 , Fig.15 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes: The processor 901 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application; The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device or a random access memory (RAM). The memory 902 can store operating devices and other applications. When the technical solution provided in the embodiments of this specification is implemented by software or firmware, the relevant program code is stored in the memory 902, and the processor 901 calls and executes the multi-agent control method of causal experience playback in the embodiments of this application; Input / output interface 903, used to implement information input and output; Communication interface 904, used to realize communication interaction between the device and other devices, which can be realized through wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.); A bus 905 that transmits information between various components of the device (e.g., the processor 901, the memory 902, the input / output interface 903, and the communication interface 904); The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .
[0146] It can be understood that the contents of the above method embodiments are all applicable to the electronic device embodiment, the functions specifically implemented by the electronic device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0147] An embodiment of the present application also provides a computer-readable storage medium, which stores a program executable by a processor. When the program executable by the processor is executed by the processor, it is used to implement the multi-agent control method of causal experience playback as described in any one of the above-mentioned specific embodiments.
[0148] An embodiment of the present application also provides a computer program product, including a computer program or computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer program or computer instructions from the computer-readable storage medium, and the processor executes the computer program or computer instructions, so that the computer device executes the multi-agent control method for causal experience playback as described in any of the previous embodiments.
[0149] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0150] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein, for example. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, device, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. It should be understood that in the present application, "at least one (item)" refers to one or more, and "a plurality" refers to two or more.
[0151] In the several embodiments provided in the present application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0152] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0153] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0154] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store program codes.
[0155] Although the description of the present application has been quite detailed and specifically describes several described embodiments, it is not intended to be limited to any of these details or embodiments or any particular embodiment, but should be regarded as providing a broad possible interpretation of these claims by reference to the attached claims, taking into account the prior art, so as to effectively cover the intended scope of the present application. In addition, the above description of the present application is based on the embodiments foreseeable by the inventor, and its purpose is to provide a useful description, and those non-substantial changes to the present application that have not yet been foreseen may still represent equivalent changes to the present application.
Claims
1. A multi-agent control method for causal experience playback, characterized in that: The method comprises: Acquire an experience data set collected when multiple agents perform a cluster formation task, wherein the experience data set includes an action vector subset and a reward value subset; Based on the action vector subset and the reward value subset, generating a causal graph, the causal graph including causal relationships between action vectors in the action vector subset and reward values in the reward value subset; Determining a weight value of each of the action vectors according to a result of adjusting each of the action vectors, wherein the weight value is used to indicate the degree of influence of the action vector on the causal relationship; The action vector subset is updated according to each of the weight values, and the preset control strategy model is trained using the updated experience data set of the action vector subset; Based on the control strategy generated by the trained control strategy model, the multiple intelligent agents are controlled to perform cluster formation tasks.
2. A multi-agent control method for causal experience playback according to claim 1, characterized in that: The generating a causal graph based on the action vector subset and the reward value subset, wherein the causal graph includes a causal relationship between the action vectors in the action vector subset and the reward values in the reward value subset, comprises: Constructing an initial causal graph and generating a plurality of nodes in the initial causal graph, each of the nodes including the action vector and a reward value corresponding to the action vector; generating an edge between every two of the nodes, the edge serving as a causal relationship between the two nodes; Based on the result of detecting the causal relationship corresponding to each edge, screening each edge and determining the type and direction of the screened edge; Determining a weight value of each edge based on the type and direction of each edge; Determine a preferred edge based on a result of comparing a preset threshold with the weight value of each edge; The initial causal graph is updated according to each of the preferred edges to obtain an updated causal graph.
3. A multi-agent control method for causal experience playback according to claim 2, characterized in that: The step of determining a weight value of each action vector according to a result of adjusting each action vector, wherein the weight value is used to indicate the degree of influence of the action vector on the causal relationship, includes: adjusting an action vector of a target node among the plurality of nodes; Determine an average processing effect value based on the action vectors before and after adjustment, wherein the average processing effect value is used to indicate the degree to which the action vector of the target node affects the reward value of the target node; Based on the average processing effect value, a weight value of the action vector of the target node is determined.
4. The multi-agent control method of causal experience playback according to claim 1, characterized in that: The updating of the action vector subset according to each of the weight values and the training of the preset control strategy model using the updated experience data set of the action vector subset include: Determine the priority of the experience data group based on the weight value of the action vector and the error value of the reward value, wherein the experience data group is composed of experience data including the action vector and the reward value collected in a cluster formation task; Using each of the experience data groups to train the control strategy model, and based on the priority of each of the experience data groups, determining a sampling probability of each of the experience data groups being sampled when training the control strategy model; When training the control strategy model, an importance sampling weight is determined based on a functional relationship between the amount of cache space storing the experience data set and the sampling probability and is used to adjust the sampling probability corresponding to a target experience data set among the multiple experience data sets.
5. The multi-agent control method of causal experience playback according to claim 1, characterized in that: The control strategy generated based on the trained control strategy model controls the multiple agents to perform the cluster formation task, including: When the multiple agents perform a cluster formation task, multiple action strategies are generated through a local control strategy model corresponding to each of the agents, wherein each of the agents performs an action corresponding to a target action strategy among the multiple action strategies; Updating the agent state and agent action corresponding to each of the action strategies; Generate a current state vector based on each updated agent state, and generate a current action vector based on each updated agent action; Using the current state vector and the current action vector as inputs of the trained control strategy model, and outputting a first output value through the trained control strategy model; Using the sample state vector and the sample action vector in the experience data set as inputs of the trained control strategy model, and outputting a second output value through the trained control strategy model; Based on the loss function relationship between the first output value and the second output value, updating the parameters of the trained control strategy model; The control strategy model after parameter update is used to generate a control strategy, and the plurality of intelligent agents are controlled according to the control strategy to perform the cluster formation task.
6. The multi-agent control method of causal experience playback according to claim 1, characterized in that: The method further comprises the step of generating a subset of the action vectors; The generating the action vector subset comprises: Generate a time series based on the behavior data of the agent in the experience data set; Dividing the time series into a plurality of subsequences; Generating multiple clusters and merging each of the subsequences into a target cluster in the multiple clusters to obtain a subsequence after clustering processing; The action vector subset is generated, and each subsequence after the clustering process is used as an action vector in the action vector subset.
7. A multi-agent control method for causal experience playback according to claim 6, characterized in that: The behavior vector corresponding to each time point in the time series includes multiple variables; The step of dividing the time series into a plurality of subsequences comprises: Based on the multiple variables, generate a precision matrix, each element in the precision matrix is used to represent the conditional independence relationship between each of the variables; Dividing the time series into multiple time segments and generating multiple cluster centers; Determine the similarity relationship between the time segment and the cluster center based on the current precision matrix; Based on the similarity relationship between the time segments and the cluster centers, each of the time segments is assigned to a cluster corresponding to the target cluster center; The cluster corresponding to each of the cluster centers is used as the subsequence, and the multiple subsequences are used as the result of the clustering process; Based on the result of the clustering process, updating the current precision matrix; The updated precision matrix is used as the current precision matrix, and the step of determining the similarity relationship between the time segment and the cluster center according to the current precision matrix is returned to be executed until the clustering process reaches a preset end condition, and a plurality of the subsequences are output.
8. A multi-agent control device for causal experience playback, characterized in that: The device comprises: An acquisition module is used to acquire an experience data set collected when multiple agents perform cluster formation tasks, wherein the experience data set includes an action vector subset and a reward value subset; A generating module, configured to generate a causal graph based on the action vector subset and the reward value subset, wherein the causal graph includes a causal relationship between the action vectors in the action vector subset and the reward values in the reward value subset; A calculation module, used for determining a weight value of each of the action vectors according to a result of adjusting each of the action vectors, wherein the weight value is used for indicating the degree of influence of the action vector on the causal relationship; A training module, used for updating the action vector subset according to each of the weight values, and training a preset control strategy model using the updated experience data set of the action vector subset; The control module is used to control the multiple intelligent agents to perform cluster formation tasks based on the control strategy generated by the trained control strategy model.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the multi-agent control method of causal experience playback as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the multi-agent control method of causal experience playback described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Distributed multi-agent deterministic strategy control method for large complex system
CN112418349A
Multi-agent federated reinforcement learning-based vehicle-road collaborative control system and method under complex intersection
WO2024016386A1