Fast Distributed Task Allocation Method and Device Based on Cooperative Communication Strategy
By constructing a counterfactual strategy and multi-agent reinforcement learning environment, and using counterfactual reward and gating mechanisms to optimize communication actions, the problem of communication conflicts in distributed task allocation is solved, and faster convergence and lower communication bandwidth requirements are achieved.
Patent Information
- Application Number
- CN202510511719.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The distributed task allocation method faces the problems of communication conflicts and conflict resolution difficulties in multi-agent systems, resulting in slow convergence speed and high communication bandwidth demand.
A fast distributed task allocation method based on collaborative communication strategy is adopted, and by constructing a counterfactual strategy and multi-agent reinforcement learning environment, optimizing communication actions using counterfactual rewards and gating mechanisms to reduce communication conflicts and competition.
It significantly reduces communication conflicts and competition, improves the convergence speed and performance of task allocation, and reduces the communication bandwidth requirements in actual network environments.
Smart Images

Figure CN120075298B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of distributed task allocation, and particularly to a fast distributed task allocation method and device based on a cooperative communication strategy. Background Art
[0002] Multi-agent systems have been applied in practical scenarios such as search and rescue, environmental monitoring, security monitoring, and power inspection. They face the crucial task allocation problem, which determines their ability to cooperate to complete difficult and complex tasks. The goal of task allocation is to allocate tasks to agents without any conflicts, while minimizing the total cost or maximizing the global benefit. However, the task allocation problem has been proven to be an NP-hard problem, and it is very difficult to solve due to the interaction of multiple constraints such as time, space, and communication.
[0003] Early studies adopted centralized methods, designating a specific agent as a central node to centrally handle task allocation. The central node faces huge computational and communication burdens, resulting in problems such as state explosion, communication difficulties, and single-point failures in centralized methods. On the other hand, distributed task allocation methods can distribute the computational load to numerous nodes, and these nodes exchange decision variables through communication to achieve task allocation in a team form. In the computational stage, each agent independently uses a decentralized allocation method to select appropriate tasks; in the communication stage, agents exchange the allocation information maintained by other agents and use consistency conflict resolution rules to ensure that the allocation results between adjacent agents are consistent. Compared with centralized methods, agents only need to communicate with adjacent agents; the optimization of the task allocation problem is dispersed among all agents, thus reducing the single-point computational load. However, distributed allocation methods face new challenges because they rely closely on information exchange between agents, which means that the convergence speed will be severely affected by communication problems such as data loss caused by collisions and delays caused by channel contention. Summary of the Invention
[0004] Based on this, it is necessary to provide a fast distributed task allocation method and device based on a cooperative communication strategy for the above technical problems.
[0005] A fast distributed task allocation method based on a cooperative communication strategy, the method includes:
[0006] Model the agent observations according to the task description.
[0007] Model the actions as an adaptive gating for communication between agents during distributed task allocation.
[0008] Construct a counterfactual strategy and generate counterfactual actions of the agents according to the counterfactual strategy.
[0009] Construct a multi-agent reinforcement learning environment for communication-aware distributed task allocation, and define it as a partially observable Markov decision process; define it as a tuple , where is the state, is the observation, is the action, is the state transition, is the reward.
[0010] Replace the original action of each agent with a counterfactual action, and determine the counterfactual reward according to the reward between the state-action obtained after replacement and the original state-action;
[0011] According to the multi-agent reinforcement learning environment, use the multi-agent proximal policy optimization method with a gating mechanism for communication actions to learn a collaborative communication strategy; among them, counterfactual rewards are used during the learning process to evaluate the global task allocation value, the consistency of the final allocation, and communication conflicts.
[0012] Multiple agents adopt a collaborative communication strategy during the execution of distributed task allocation to obtain a distributed task allocation result.
[0013] A fast distributed task allocation device based on a collaborative communication strategy, the device includes:[[]]
[0014] The collaborative communication strategy optimization module is used to model the agent observation according to the task description; model the action as an adaptive gating for communication between agents during distributed task allocation; construct a counterfactual strategy, and generate counterfactual actions for agents according to the counterfactual strategy; construct a multi-agent reinforcement learning environment for communication-aware distributed task allocation, and define it as a partially observable Markov decision process; define it as a tuple , where is the state, is the observation, is the action, is the state transition, is the reward; replace the original action of each agent with a counterfactual action, and determine the counterfactual reward according to the reward between the state-action obtained after replacement and the original state-action; according to the multi-agent reinforcement learning environment, use the multi-agent proximal policy optimization method with a gating mechanism for communication actions to learn a collaborative communication strategy; among them, counterfactual rewards are used during the learning process to evaluate the global task allocation value, the consistency of the final allocation, and communication conflicts.
[0015] The fast distributed task allocation module is used for multiple agents to adopt a collaborative communication strategy during the execution of distributed task allocation to obtain a distributed task allocation result.
[0016] The above-mentioned fast distributed task allocation method and device based on the cooperative communication strategy. In the method, the observation, action, and reward mechanisms simultaneously consider the task-related and communication-related characteristics of the MAC layer. Through the centralized training and decentralized execution scheme, MAPPO is extended with counterfactual rewards to improve the learning speed and final performance. The learned cooperative communication strategy significantly reduces communication conflicts and competition, can be applied to any distributed task allocation algorithm, thereby accelerating convergence and improving performance. The trained communication strategy can effectively reduce the communication bandwidth required in the actual distributed network environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic flowchart of the fast distributed task allocation method based on the cooperative communication strategy in one embodiment;
[0018] Figure 2 It is an overall architecture diagram of the fast distributed task allocation method based on the cooperative communication strategy in one embodiment;
[0019] Figure 3 It is a schematic diagram of the performance comparison of several methods in different scenarios in one embodiment, where Figure 3 (a) is a schematic diagram of the comparison of task allocation values of several algorithms in the scenario of 20 agents and 100 tasks, Figure 3 (b) is a schematic diagram of the comparison of average rewards of several algorithms in the scenario of 20 agents and 100 tasks, Figure 3 (c) is a schematic diagram of the comparison of the number of message losses of several algorithms in the scenario of 20 agents and 100 tasks, Figure 3 (d) is a schematic diagram of the comparison of task allocation values of several algorithms in the scenario of 50 agents and 250 tasks, Figure 3 (e) is a schematic diagram of the comparison of average rewards of several algorithms in the scenario of 50 agents and 250 tasks, Figure 3 (f) is a schematic diagram of the comparison of the number of message losses of several algorithms in the scenario of 50 agents and 250 tasks;
[0020] Figure 4 It is a time slot diagram of the data received by the agents before and after applying the task allocation method in one embodiment, where Figure 4 (a) is the time slot diagram of the data received by the agents before applying the task allocation method, Figure 4 (b) is the time slot diagram of the data received by the agents after applying the task allocation method;
[0021] Figure 5 It is a schematic diagram of the transmission plan of three groups of experience trajectories and the change of credit assignment values over time under different agents and different learning methods in another embodiment, where Figure 5(a) is the transmission plan for three groups of empirical trajectories, Figure 5 (b) is a schematic diagram comparing the credit assignment values for different agents under the CF-MAPPO method, Figure 5 (c) is a schematic diagram comparing the credit assignment values for different agents under the QMIX method, Figure 5 (d) is a schematic diagram comparing the credit assignment values for different agents under the VDN method;
[0022] Figure 6 is a schematic diagram of the statistical data of SGA rewards and the required communication bandwidth before and after applying the communication strategy in an embodiment, where Figure 6 (a) is a schematic diagram of the statistical data of SGA rewards before and after applying the communication strategy, Figure 6 (b) is a schematic diagram of the statistical data of the required communication bandwidth. Detailed implementation manners
[0023] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0024] In one embodiment, as Figure 1 shown, a fast distributed task allocation method based on a collaborative communication strategy is provided, and the method includes the following steps:
[0025] Step 100: Model the agent observations according to the task description.
[0026] Specifically, the task is a communication-aware task allocation task for a multi-agent system. The multi-agent system consists of multiple agents, where the agents can be, but are not limited to, drones, unmanned vehicles, robots, etc.
[0027] The multi-agent system can be applied in actual scenarios such as search and rescue, environmental monitoring, security monitoring, and power inspection. For example: A multi-agent system performs a search and rescue task in a wild environment.
[0028] Modeled the key issues in the real-world communication network based on the CSMA / CA protocol, such as message collision and channel contention.
[0029] Model the agent observations according to the task description to incorporate features such as Bron-Kerbosch (BK), value of message (VoM), and channel access quality, and also include a normalization step.
[0030] Step 102: Model the actions as an adaptive gating of the communication between agents during the distributed task allocation.
[0031] Specifically, the action based on the gating mechanism enables the agent to adaptively select the data transmission time, thereby coordinating the communication between agents to improve the algorithm performance.
[0032] Step 104: Construct a counterfactual policy and generate the counterfactual actions of the agent according to the counterfactual policy.
[0033] Step 106: Construct a multi-agent reinforcement learning environment for communication-aware distributed task allocation, and define it as a partially observable Markov decision process; define it as a tuple , where is the state, is the observation, is the action, is the state transition, is the reward.
[0034] Specifically, the performance of the distributed task allocation method may be affected by the actual communication process. Agents can learn communication policies to coordinate communication times, reduce data collision rates, and improve channel utilization, thereby accelerating the convergence rate of the allocation algorithm. For this purpose, a multi-agent reinforcement learning environment for communication-aware distributed task allocation is constructed and defined as a partially observable Markov decision process (POMDP).
[0035] The state transition function describes the mapping of the environment from state to the next state , usually defined as:
[0036] ;
[0037] where and represent the global states at times t and t +1 respectively, and represents the joint communication actions of all agents.
[0038] Step 108: Replace the original action of each agent with the counterfactual action, and determine the counterfactual reward according to the reward between the state-action obtained after replacement and the original state-action.
[0039] Specifically, the counterfactual reward is used to evaluate the global task allocation value, the consistency of the final allocation, and communication conflicts.
[0040] The actions and policies obtained according to the counterfactual reward.
[0041] Step 110: According to the multi-agent reinforcement learning environment, use the multi-agent proximal policy optimization method with a gating mechanism for communication actions to learn a collaborative communication policy; during the learning process, use counterfactual rewards to evaluate the global task allocation value, the consistency of the final allocation, and communication conflicts.
[0042] Specifically, the communication policy learned through multi-agent proximal policy optimization (MAPPO) is significantly superior to other learning methods, such as MTD3, QMIX, and VDN, in many metrics including allocation value, learning speed, and communication quality.
[0043] Step 112: Multiple agents adopt a collaborative communication policy during the execution of distributed task allocation to obtain a distributed task allocation result.
[0044] Specifically, applying the collaborative communication policy to five classic allocation algorithms: Consensus-Based Auction Algorithm (CBAA), Consensus-Based Bundle Algorithm (CBBA), Hybrid Information and Plan Consensus (HIPC), Distributed Genetic Algorithm (DGA), and Distributed Heuristic Task Allocation (DHBA) shows a consistent and significant improvement in coordination efficiency and final solution performance under various task settings.
[0045] In the above fast distributed task allocation method based on a collaborative communication policy, the observation, action, and reward mechanisms in the method simultaneously consider the task-related and communication-related characteristics of the MAC layer; through the centralized training and decentralized execution scheme, MAPPO is extended with counterfactual rewards to improve the learning speed and final performance. The learned collaborative communication policy significantly reduces communication conflicts and competition, can be applied to any distributed task allocation algorithm, thereby accelerating convergence and improving performance, and the trained communication policy can effectively reduce the communication bandwidth required in the actual distributed network environment.
[0046] In one embodiment, the agent observations in step 100 include: Bron-Kerbosch features, message value features, channel access features; where the channel access features are used to calculate the channel access priority, enabling the agent to focus on the number of times it accesses the communication channel, thereby preventing excessive channel access; the channel access priority can be calculated through the arctangent function, which shows a significant change at the maximum backoff count, thus quickly forming an ordered and continuous backoff sequence for the agent; the expression of the channel access priority is:
[0047] ;
[0048] where, is the channel access priority, is the weighting factor, representing the maximum number of backoffs within a contention period, is the number of backoffs that the agent has made in the channel.
[0049] In one embodiment, step 102 includes: at the t th step, the agent i inputs its local observation into the Actor network to obtain a communication action , and the communication action determines whether to send a message; when , the agent does not use the communication channel at the t th step, and when , the agent uses the communication channel to send a message.
[0050] In one embodiment, the expression of the counterfactual reward in step 108 is:
[0051] ;
[0052] where, represents the counterfactual reward of the agent i at the t th step, R (·) represents the reward function used to evaluate the state-action values of all agents, represents the joint observation of all agents at the t th step, represents the joint communication action of all agents at the t th step, represents the communication action of the agent i at the t th step, represents the estimated communication action of the agent i at the t th step.
[0053] In one embodiment, the estimation process of the counterfactual reward in step 108 includes: sampling the depth network to fit the reward function used to evaluate the state-action values of all agents; taking the joint state-action in the experience trajectory as the input of the depth network, and calculating the training gradient based on minimizing the TD error to update the parameters of the depth network; the update expression of the depth network parameters is:
[0054] ;
[0055] ;
[0056] ;
[0057] where, Indicates depth The parameters of the network, Indicates the objective function, Indicates the hyperparameter, Indicates that the agent at the t +1 step joint observation adopts the value of the joint communication action, Indicates the joint observation of all agents at the t +1 step, Indicates the joint communication action of all agents at the t +1 step, Indicates the global shared reward, Indicates at the t step during which the agent i The change in the average task conflict, 、 Respectively indicate at the t step and the t +1 step at the beginning, the agent i The average task conflict with other agents.
[0058] After constructing the deep network, use the joint observation value and joint action at the t step to approximately calculate the counterfactual reward of the agent i ; The calculation process of the counterfactual reward estimate is as follows:
[0059] ;
[0060] ;
[0061] ;
[0062] Among them, Indicates the counterfactual reward estimate, and Respectively indicate for the original reward and the reward of the deep network estimate.
[0063] Specifically, in order to capture the complex cooperation relationship between agents in this environment and achieve reasonable credit assignment, this application proposes a fast distributed task allocation method based on a cooperative communication strategy. This method is a MAPPO extended algorithm based on the concept of differential reward, abbreviated as: CF-MAPPO. This method uses counterfactual actions to replace the original actions of agents to infer their contributions to the cooperation relationship, thereby clearly allocating the global shared reward. The framework of CF-MAPPO is as Figure 2 shown. Any agent iThe original action is replaced by a counterfactual action. Drawing on the concept of differential reward, the difference between the state-action obtained after replacement and the original state-action is used as the reward it assigns. The replacement, that is, the counterfactual reward:
[0064] ;
[0065] Among them, R (·) represents the reward function used to evaluate the state-action values of all agents.
[0066] In addition, as Figure 2 shown, a deep network is used to fit R (·) and estimate the counterfactual reward to predict the action-state value. To train the deep network, the joint state-action in the experience trajectory is used as the input, and the training gradient is calculated based on minimizing the TD error to update the value network parameters , that is:
[0067] ;
[0068] ;
[0069] After constructing the value function, the joint observation and the joint action are used to approximately calculate the counterfactual reward i of the agent ; The calculation process of the counterfactual reward is as follows:
[0070] ;
[0071] ;
[0072] ;
[0073] Among them, represents the estimated value of the counterfactual reward, and respectively represent the estimated values of the value network for the original reward and the reward .
[0074] In one embodiment, the multi-agent proximal policy optimization method with a gating mechanism communication action in step 110 adopts a centralized training - decentralized execution method to train the shared policy as the Actor network and train the shared value as the Critic network; the update expression of the Critic network parameters is:
[0075] ;
[0076] Among them, represents the Critic network parameters, represents the state value function, represents the global shared reward.
[0077] The update expression of the Actor network parameters is:
[0078] ;
[0079] ;
[0080] ;
[0081] Among them, represents the Actor network parameters, and respectively represent the advantage and entropy of the policy, represents the action-state value of the agent, represents the weighting factor of the advantage, represents the advantage evaluation function, represents the advantage function, represents the hyperparameter.
[0082] Specifically, at the beginning of each step, the agent decides whether to transmit data according to the communication action . During the communication process, the agent autonomously occupies the channel according to the communication policy to avoid data conflicts. During the calculation process, the agent resolves task conflicts based on the received messages and the consistency conflict resolution rules. During the communication process, the agent autonomously occupies the channel according to the communication policy to avoid data conflicts. During the calculation process, the agent resolves task conflicts based on the received messages and the consistency conflict resolution rules. Introduce a cooperation relationship to coordinate communication. This can be achieved by establishing a common goal among agents based on the communication behavior of the gate mechanism. The global shared reward r t can be set to reduce the global task conflict within a fixed time period, that is:
[0083] ;
[0084] Among them, represents the change in the average task conflict of agent t during the i step, represents the average task conflict between agent t and other agents at the beginning of the i step, that is , represents the task list of the agent i . represents the task list of the agent j .
[0085] In one embodiment, during the centralized training phase, the policy advantage is calculated using the training samples with counterfactual rewards as follows:
[0086] ;
[0087] wherein, represents the policy advantage, represents a hyperparameter, represents the agent i at the t -th step of the counterfactual reward.
[0088] In one embodiment, the expression of the weighting factor of the advantage is:
[0089] ;
[0090] wherein, , respectively represent the policy distributions of the agent before and after updating , and represents the KL divergence.
[0091] Specifically, during the communication process of the distributed algorithm, the smaller the amount of data loss caused by communication, the more the number of task conflicts between agents is reduced, and thus the greater the reward feedback obtained from the system. Then, a multi-agent proximal policy optimization algorithm with a gating mechanism communication action is proposed to learn the collaborative communication strategy. First, centralized training and decentralized execution (CTDE) is used to share the state-action information of all agents, helping them identify cooperative relationships and eliminate potential conflicts between policies, thereby improving the learning of collaborative strategies between agents. Second, counterfactual rewards are constructed to distinguish the contributions of different agent policies to teamwork, and the MAPPO method based on counterfactual rewards is proposed to achieve more accurate policy updates. Since the communication action based on the gating mechanism is a discrete action, the proximal policy optimization (PPO) algorithm based on discrete actions is used for training. The method of centralized training - decentralized execution is adopted to train the shared policy as the behavior network and the shared value as the evaluation network. Each agent i utilizes the local observation to obtain the action , so as to achieve the decentralized execution of different actions. This method helps agents identify cooperative relationships and eliminate potential conflicts between policies by sharing state and reward information between agents. During the decentralized execution phase, the agenti At step t its local observations are input into the behavior network to obtain an action and thus determine whether the agent sends data at step t . This constructs the global shared reward obtained by all agents through joint actions t and environmental feedback at step . Meanwhile, each agent i inputs its local observations into the value network to obtain the state value t corresponding to step . Therefore, each agent i can collect training samples to construct its corresponding experience trajectory and store it in a buffer for subsequent training. On the other hand, in the centralized training phase, the main goal is to use the collected experience trajectories to update the value network and the policy network. For the value network, the experience trajectories of all agents are used to update its parameters to fit the state value function. Then, we calculate the training gradient by minimizing the mean squared error between the state value function and the global shared reward . Based on this error, the value network parameters are updated in the same way as shown in the above Critic network parameter update expression, which applies to all agents. For the Actor network, the Generalized Advantage Estimation (GAE) method is usually used to update its parameters, i.e.:
[0092] ;
[0093] where represents the action-state value of the agent. It can be estimated by adding the global shared reward to .
[0094] To avoid too large a difference in the policy distribution before and after parameter update, the KL divergence is introduced as a weighted factor of the advantage , and a truncation function is used to set an upper limit for . This can jointly limit the range of the policy gradient and thus ensure the stability of the policy update, i.e.:
[0095] ;
[0096] where , , respectively represent the agent at the time of update The policy distributions before and after.
[0097] Finally, update the Actor network parameters as follows :
[0098] ;
[0099] where and represent the advantage and entropy of the policy respectively. It should be noted that the above advantage depends on the globally shared reward .
[0100] However, a unified reward shared globally is difficult to distinguish the contributions of different policies. In particular, this may lead to high reward values for ineffective communication policies, ultimately affecting the update of the Actor network.
[0101] In one of the embodiments, the method adopted for distributed task allocation in step 112 is a consensus-based auction algorithm, a consensus-based bundling algorithm, hybrid information and plan consensus, a distributed genetic algorithm, or a distributed heuristic task allocation.
[0102] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0103] In a verification implementation example, extensive simulation experiments were conducted to evaluate the effectiveness of communication strategy optimization based on multi-agent reinforcement learning (MARL) in distributed task allocation. When evaluating these algorithms, the existing simple internal simulator code was used and extended to train the strategies. The simulator modeled key issues in real-world communication networks based on the CSMA / CA protocol, such as message collisions and channel contention. Under simulation conditions, MAPPO and CF-MAPPO were used to learn communication strategies based on a gating mechanism, and performance was evaluated using metrics such as the conflict ratio, assigned task value (SGA reward), total data conflict, and average reward. To further verify the applicability of the proposed method, agent data was recorded and communication strategies were trained on a self-organizing network simulation platform server (clrsp) for hybrid virtual reality. The platform constructs networks with different topologies by simulating the distances and communication ranges of communication nodes. It also runs a five-layer network communication protocol on each node to replicate real-world networks and communication environments. The distributed task allocation algorithm program runs as an independent thread, with each agent corresponding to a node on the server. The distributed task allocation method is mainly evaluated from the following four aspects:
[0104] 1) Proportion of conflicting tasks: The ratio of the number of conflicting tasks to the total number of tasks when the algorithm converges or terminates.
[0105] 2) Assigned task value: The ratio of the total reward of the assigned tasks to the total reward obtained by centralized task allocation when the algorithm converges or terminates.
[0106] 3) Total data conflict: The sum of the data conflicts of all agents in the environment in one round.
[0107] 4) Average reward: The environmental reward generated by reinforcement learning represents the change in the average number of task conflicts in this scenario.
[0108] To verify the effectiveness of the proposed method, scenarios with 20 agents and 100 tasks and 50 agents and 250 tasks were considered, and these were conducted under 5 different random seeds. The learning speed and final performance of the learned strategies were compared with three well-known multi-agent reinforcement learning methods: MTD3, QMIX, and VDN. The performance comparisons of these methods in different scenarios are as Figure 3 shown, where Figure 3 (a) is a schematic diagram of the comparison of the task allocation values of several algorithms in the scenario of 20 agents and 100 tasks, Figure 3 (b) is a schematic diagram of the comparison of the average rewards of several algorithms in the scenario of 20 agents and 100 tasks, Figure 3 (c) is a schematic diagram of the comparison of the number of message losses of several algorithms in the scenario of 20 agents and 100 tasks, Figure 3(d) is a schematic diagram comparing the task assignment values of several algorithms in the scenario of 50 agents and 250 tasks. Figure 3 (e) is a schematic diagram comparing the average rewards of several algorithms in the scenario of 50 agents and 250 tasks. Figure 3 (f) is a schematic diagram comparing the number of message losses of several algorithms in the scenario of 50 agents and 250 tasks. It can be seen from Figure 3 that in all scenarios, the proposed method significantly outperforms the value decomposition methods QMIX and VDN. In particular, the target value of task assignment ranks the highest after convergence, and MTD3 is the lowest (0.68 vs. 0.43). Similar trends can also be found in the comparison of cumulative rewards (0.78 vs. 0.2) and the number of message losses (7 vs. 25). This further demonstrates the advantage of restricting the maximum difference between the old and new policies, although the counterfactual reward improves the accuracy of the actor network. In contrast, the simple decomposition of the global state value in QMIX and VDN ignores the strong dependencies between adjacent agents, resulting in poor training performance in all scenarios. To demonstrate the effectiveness of the communication strategy, a time-slot graph of the data received by the agents is used to depict the transmission and reception states of all agents before and after training; Figure 4 is the time-slot graph of the data received by the agents before and after applying the proposed method for task assignment, where Figure 4 (a) is the time-slot graph of the data received by the agents before applying the proposed method for task assignment, Figure 4 (b) is the time-slot graph of the data received by the agents after applying the proposed method for task assignment. In the scenario with 20 agents, the cooperative communication strategy significantly reduces data conflicts and channel competition, thus improving performance. Credit assignment is crucial for determining whether the actor network of the agents can be correctly updated. To analyze this, we used training samples from three sets of empirical trajectories in a scenario with five agents, and these samples contain random actions. The credit assignment values of the CF-MAPPO, QMIX, and VDN algorithms were visualized and processed using min-max normalization. Figure 5 shows the transmission plans of the three sets of empirical trajectories and the variation of the credit assignment values over time for different agents and different learning methods, where Figure 5 (a) is the transmission plan of the three sets of empirical trajectories, Figure 5 (b) is a schematic diagram comparing the credit assignment values for different agents under the CF-MAPPO method, Figure 5 (c) is a schematic diagram comparing the credit assignment values for different agents under the QMIX method, Figure 5 (d) is a schematic diagram comparing the credit assignment values for different agents under the VDN method.
[0109] It can be seen that, compared with QMIX and VDN, the CF-MAPPO method proposed in this application can more accurately evaluate the contributions of different agents under their local policies, and both its maximum value (0.3 compared with 1.0 and 0.95) and minimum value (0.0 compared with 0.3 and 0.4) are different. This is because the CF-MAPPO method determines the allocated credit based on counterfactual rewards, while QMIX and VDN mainly calculate the credit based on the value function and the corresponding states of each agent. More specifically, as shown in Figure 5, at step 10, the two hidden nodes (Agent 2 and Agent 4) did not simultaneously select the channel occupancy strategy, thus reducing the incidence of communication conflicts and obtaining higher counterfactual rewards.
[0110] To verify the applicability of the actions based on the gating mechanism to different allocation methods, training and testing were carried out by replacing various allocation algorithms. The communication strategy was combined with several classic distributed algorithms including CBBA, CBAA, HIPC, DGA, and DHBA for experiments, and these algorithms have different scales. Monte Carlo simulations show that, compared with the methods without the communication strategy, the allocation algorithms combined with the communication strategy reduce the average task conflict rate by up to 29%, 40%, and 21% respectively in the scenarios of 10, 20, and 50 agents, and reduce the total data loss rate by up to 42%, 80%, and 89% respectively. Among them, the redundancy-based central method DHBA has a slight reduction in the average task conflict rate after combining the communication strategy, reducing by 18%, 7%, and 3% respectively in the scenarios of 10, 20, and 50 agents. In contrast, after incorporating the communication strategy, CBAA has a more significant reduction in the average task conflict rate, reducing by 17%, 40%, and 17% respectively for the scenarios of 10, 20, and 50 agents. In summary, the proposed communication strategy can adapt to various distributed task allocation algorithms.
[0111] To verify the effectiveness of the learning framework under real network conditions, a distributed algorithm with a communication process is simulated on a self-organizing network simulation platform server that closely resembles real network conditions, and data is collected distributively to train the communication strategy. This platform can not only construct the topological structure of a real self-organizing network but also simulate data collisions, losses, and errors through the MAC layer network protocol, thus reproducing the distributed communication scenario. On this platform, the experimental settings are as follows: (1) The goal is to reduce the bandwidth required for communication between agents and optimize the reward design; (2) Since it is difficult to set a unified termination time for the distributed task allocation algorithm, by default, the algorithm running time for each node is fixed at 1.5 seconds. (3) The distributed CBBA algorithm is used for node operations, and experiments are conducted under conditions where the total communication bandwidth is 4.5 Mbps, 9.5 Mbps, and 15.2 Mbps respectively. The trained strategy is verified through 50 Monte Carlo simulations (30 agents in each group). Figure 6 Schematic diagram of statistical data of SGA rewards and required communication bandwidth before and after applying the communication strategy in an embodiment, where Figure 6 (a) Schematic diagram of statistical data of SGA rewards before and after applying the communication strategy, Figure 6 (b) Schematic diagram of statistical data of required communication bandwidth. It can be seen from Figure 6 that before and after applying the communication strategy, the SGA reward of CBBA is hardly affected, while the communication bandwidth required for data transmission is significantly reduced. Under the bandwidth conditions of 4.5 Mbps, 9.5 Mbps, and 15.2 Mbps, the bandwidth required for the CBBA algorithm to send data is reduced by approximately 43%, 48%, and 42% respectively. These results indicate that the gated communication strategy learning framework is applicable to a nearly real network environment and can significantly reduce the required communication bandwidth without affecting the algorithm performance.
[0112] In an embodiment, a fast distributed task allocation device based on a cooperative communication strategy is provided, including: a cooperative communication strategy optimization module and a fast distributed task allocation module, where:
[0113] The cooperative communication strategy optimization module is used to model the agent observations according to the task description; model the actions as an adaptive gate for communication between agents during distributed task allocation; construct a counterfactual strategy and generate counterfactual actions of the agents according to the counterfactual strategy; construct a multi-agent reinforcement learning environment for communication-aware distributed task allocation and define it as a partially observable Markov decision process; define it as a tuple where is the state, is the observation, is the action, is the state transition, As a reward; the original action of each agent is replaced by a counterfactual action, and the counterfactual reward is determined according to the reward between the state-action obtained after replacement and the original state-action; according to the multi-agent reinforcement learning environment, a multi-agent proximal policy optimization method with a gating mechanism for communication actions is used to learn a collaborative communication policy; among them, the counterfactual reward is used to evaluate the global task allocation value, the consistency of the final allocation, and communication conflicts during the learning process.
[0114] A fast distributed task allocation module, which is used for multiple agents to adopt a collaborative communication policy during the execution of distributed task allocation to obtain a distributed task allocation result.
[0115] In one embodiment, the agent observations in the collaborative communication policy optimization module include: Bron-Kerbosch features, message value features, and channel access features; among them, the channel access features are used to calculate the channel access priority, so that the agent pays attention to the number of times it accesses the communication channel, thereby preventing excessive access to the channel; the channel access priority is as shown in the expression of the above-mentioned channel access priority.
[0116] In one embodiment, the collaborative communication policy optimization module is also used in the t th step, the agent i inputs its local observation value into the Actor network to obtain a communication action , and the communication action determines whether to send a message; when , the agent does not use the communication channel in the t th step, and when , the agent uses the communication channel to send a message.
[0117] In one embodiment, the counterfactual reward in the collaborative communication policy optimization module is as shown in the expression of the above-mentioned counterfactual reward.
[0118] In one embodiment, the estimation process of the counterfactual reward in the collaborative communication policy optimization module includes: sampling the depth network to fit the reward function used to evaluate the state-action values of all agents; using the joint state-action in the experience trajectory as the input of the depth network, and calculating the training gradient based on minimizing the TD error to update the parameters of the depth network; the update method of the depth network parameters is as shown in the expression of the update of the above-mentioned depth network parameters.
[0119] After constructing the depth network, use the joint observation value and joint action in the t th step to approximately calculate the agent iCounterfactual rewards. The specific calculation process of the counterfactual rewards is as shown in the calculation process of the counterfactual rewards above.
[0120] In one embodiment, the multi-agent proximal policy optimization method with gating mechanism communication actions in the collaborative communication policy optimization module uses the centralized training - decentralized execution method to train the shared policy as the Actor network and train the shared value as the Critic network; the parameter update method of the Critic network is as shown in the above Critic network parameter update expression; the parameter update method of the Actor network is as shown in the above Actor network parameter update expression.
[0121] In one embodiment, in the centralized training stage of the multi-agent proximal policy optimization method with gating mechanism communication actions in the collaborative communication policy optimization module, the training samples with counterfactual rewards are used to calculate the policy advantage; the policy advantage is as shown in the above expression of the policy advantage.
[0122] In one embodiment, the weighting factor of the advantage is as shown in the above expression of the weighting factor of the advantage.
[0123] In one embodiment, the method adopted for distributed task allocation in the fast distributed task allocation module is the consensus-based auction algorithm, the consensus-based bundling algorithm, the hybrid information and plan consensus, the distributed genetic algorithm, or the distributed heuristic task allocation.
[0124] For the specific limitations of the fast distributed task allocation device based on the collaborative communication policy, reference can be made to the limitations of the fast distributed task allocation method based on the collaborative communication policy in the above text, which will not be elaborated here. Each module in the above fast distributed task allocation device based on the collaborative communication policy can be implemented in whole or in part through software, hardware, and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.
[0125] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0126] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A fast distributed task allocation method based on a cooperative communication strategy, characterized in that, The method includes: Modeling the agent observations according to the task description; Modeling the actions as an adaptive gating for inter-agent communication during distributed task allocation; Constructing a counterfactual policy and generating counterfactual actions of the agent according to the counterfactual policy; Construct a multi-agent reinforcement learning environment for communication-aware distributed task allocation and define it as a partially observable Markov decision process; define it as a tuple , where is the state, is the observation, is the action, is the state transition, is the reward; Replacing the original action of each agent with the counterfactual action, and determining the counterfactual reward according to the reward between the state-action obtained after replacement and the original state-action; According to the multi-agent reinforcement learning environment, using a multi-agent proximal policy optimization method with communication actions of a gating mechanism to learn a collaborative communication policy; wherein the counterfactual reward is used to evaluate the global task allocation value, the consistency of the final allocation, and communication conflicts during the learning process; the multi-agent proximal policy optimization method with communication actions of a gating mechanism uses a centralized training - decentralized execution method to train a shared policy as the Actor network and train a shared value as the Critic network; Multiple agents adopt the collaborative communication policy during the execution of distributed task allocation to obtain a distributed task allocation result; Among them, modeling the actions as an adaptive gating for inter-agent communication during distributed task allocation includes: At step t , the agent i inputs its local observation into the Actor network to obtain a communication action , and the communication action determines whether to send a message; when , the agent does not use the communication channel at step t , and when , the agent uses the communication channel to send a message; The counterfactual reward is: Among them, represents the counterfactual reward of the agent i at the t th step, represents the reward function used to evaluate the state-action values of all agents, represents the joint observation of all agents at the t th step, represents the joint communication action of all agents at the t th step, represents the communication action of the agent i at the t th step, represents the estimated communication action of the agent i at the t th step.
2. The fast distributed task allocation method based on a cooperative communication strategy according to claim 1, wherein Modeling the agent observations according to the task description, and the agent observations in the step include: Bron-Kerbosch features, message value features, channel access features; wherein the channel access features are used to calculate the channel access priority, enabling the agent to pay attention to the number of times it accesses the communication channel, thereby preventing excessive access to the channel; the channel access priority is: wherein, is the channel access priority, is the weighting factor, representing the maximum number of backoffs within a contention period, is the number of backoffs that the agent has made in the channel.
3. The fast distributed task allocation method based on a cooperative communication strategy according to claim 1, wherein The estimation process of the counterfactual reward includes: Sampling depth Q The network fitting is used to evaluate the reward function of all agent state-action values; the joint state-action in the experience trajectory is used as the input of the deep Q network, and the training gradient is calculated based on minimizing the TD error to update the parameters of the deep Q network; the update expression of the deep Q network parameter is: Among them, represents the parameters of the depth Q network, represents the objective function, represents the hyperparameters, r t represents the globally shared reward, represents the value of the joint communication action adopted by the agent for the joint observation at the t +1 step, represents the joint observation of all agents at the t +1 step, represents the joint communication action of all agents at the t +1 step, represents the change in the average task conflict of the agent t during the i step, and respectively represent the average task conflict between the agent t at the beginning of the t step and the i +1 step and other agents, n represents the number of agents; Build depth Q After the network, use the t step joint observations and joint actions to approximately calculate the i agent's counterfactual reward; the calculation process of the counterfactual reward estimate is as follows: Among them, represents the counterfactual reward estimate, and respectively represent the deep and network estimates for the original reward Q and the reward.
4. The fast distributed task allocation method based on a cooperative communication strategy according to claim 1, characterized in that The update method of the Critic network parameters is: Among them, represents the Critic network parameters, represents the state value function, r t represents the globally shared reward; The update method of the Actor network parameters is: Among them, represents the Actor network parameters, and respectively represent the advantage and entropy of the policy, represents the action-state value of the agent, ω represents the weighting factor of the advantage, represents a hyperparameter, represents the advantage evaluation function, represents the advantage function, represents a hyperparameter, represents the agent i at the t step communication action, represents the agent i at the t step local observation.
5. The fast distributed task allocation method based on a collaborative communication strategy according to claim 4, wherein In the centralized training stage, using the training samples added with the counterfactual reward to calculate the policy advantage as: Among them, represents the strategic advantage, represents a hyperparameter, represents the agent i at the t step's counterfactual reward.
6. The fast distributed task allocation method based on a cooperative communication strategy according to claim 4, characterized in that The weighting factor of the advantage is: Among them, , respectively represent the policy distributions of the agent before and after updating , and represents the KL divergence.
7. The fast distributed task allocation method based on a cooperative communication strategy according to claim 1, characterized in that The method adopted for distributed task allocation is a consensus-based auction algorithm, a consensus-based bundling algorithm, hybrid information and plan consensus, a distributed genetic algorithm, or a distributed heuristic task allocation.
8. A fast distributed task allocation device based on a cooperative communication strategy, characterized in that, The device includes: A collaborative communication strategy optimization module, which is used to model the agent observations according to the task description; model the actions as an adaptive gating for communication among agents during distributed task allocation; construct a counterfactual strategy and generate counterfactual actions of the agents according to the counterfactual strategy; construct a multi-agent reinforcement learning environment for communication-aware distributed task allocation and define it as a partially observable Markov decision process, which is defined as a tuple , where is the state,[[]]END]] is the observation,[[]]END]] is the action,[[]]END]] is the state transition,[[]]END]] is the reward; replace the original action of each agent with the counterfactual action, and determine the counterfactual reward according to the reward between the state-action obtained after replacement and the original state-action; according to the multi-agent reinforcement learning environment, use a multi-agent proximal policy optimization method with communication actions of a gating mechanism to learn the collaborative communication strategy; wherein the counterfactual reward is used to evaluate the global task allocation value, the consistency of the final allocation and the communication conflict during the learning process; the multi-agent proximal policy optimization method with communication actions of a gating mechanism uses a centralized training - decentralized execution method to train the shared policy as the Actor network and train the shared value as the Critic network; A fast distributed task allocation module, which is used for multiple agents to adopt the collaborative communication policy during the execution of distributed task allocation to obtain a distributed task allocation result; Among them, the collaborative communication strategy optimization module is further configured to, in the t th step, the agent i inputs its local observation value into the Actor network to obtain a communication action , and the communication action determines whether to send a message; when , the agent does not use the communication channel in the t th step, and when , the agent uses the communication channel to send a message; The counterfactual reward in the collaborative communication policy optimization module is: Among them, represents the counterfactual reward of the agent i at the t step, represents the reward function used to evaluate the state-action values of all agents, represents the joint observation of all agents at the t step, represents the joint communication action of all agents at the t step, represents the communication action of the agent i at the t step, represents the estimated communication action of the agent i at the t step.
Citation Information
Patent Citations
Multi-machine collaborative air combat planning method and system based on deep reinforcement learning
CN112861442A
Unmanned aerial vehicle cluster near-end strategy optimization collaborative confrontation method based on anti-fact baseline
CN118778678A