Rapid distributed task allocation method and device based on cooperative communication strategy
By adopting a distributed task allocation method based on collaborative communication strategy in multi-agent systems, communication conflicts and competition problems are solved, and faster convergence speed and higher performance are achieved.
Patent Information
- Application Number
- CN202510511719.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The distributed task allocation method faces communication conflicts and competition problems in multi-agent systems, resulting in slow convergence speed and poor performance.
A fast distributed task allocation method based on collaborative communication strategy is adopted to build a counterfactual strategy and a multi-agent reinforcement learning environment to learn communication actions with a gated mechanism to reduce communication conflicts and competition.
Significantly reduces communication conflicts and competition, improves the convergence speed and performance of task allocation, and reduces the communication bandwidth required in the actual distributed network environment.
Smart Images

Figure CN120075298A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of distributed task allocation, and particularly to a fast distributed task allocation method and device based on a cooperative communication strategy. Background Art
[0002] Multi-agent systems have been applied in practical scenarios such as search and rescue, environmental monitoring, security monitoring, and power inspection. They face the crucial task allocation problem, which determines their ability to cooperate to complete difficult and complex tasks. The goal of task allocation is to allocate tasks to agents without any conflicts, while minimizing the total cost or maximizing the global benefit. However, the task allocation problem has been proven to be an NP-hard problem and is very difficult to solve due to the interaction of multiple constraints such as time, space, and communication.
[0003] Early research adopted a centralized method, designating a specific agent as a central node to centrally process task allocation. The central node faces huge computational and communication burdens, resulting in problems such as state explosion, communication difficulties, and single-point failures in the centralized method. On the other hand, distributed task allocation methods can distribute the computational load to numerous nodes, and these nodes exchange decision variables through communication to achieve task allocation in a team form. In the calculation stage, each agent independently uses a decentralized allocation method to select appropriate tasks; in the communication stage, agents exchange the allocation information maintained by other agents and use consistency conflict resolution rules to ensure that the allocation results between adjacent agents are consistent. Compared with the centralized method, agents only need to communicate with adjacent agents; the optimization of the task allocation problem is dispersed among all agents, thereby reducing the single-point computational load. However, distributed allocation methods face new challenges because they rely closely on information exchange between agents, which means that the convergence speed will be severely affected by communication problems such as data loss caused by collisions and delays caused by channel competition. Summary of the Invention
[0004] Based on this, it is necessary to provide a fast distributed task allocation method and device based on a cooperative communication strategy for the above technical problems.
[0005] A fast distributed task allocation method based on a cooperative communication strategy, the method includes: Model the agent observations according to the task description.
[0006] Model the actions as an adaptive gating for communication between agents during distributed task allocation.
[0007] Construct a counterfactual strategy and generate counterfactual actions of the agents according to the counterfactual strategy.
[0008] Construct a multi-agent reinforcement learning environment for communication-aware distributed task allocation and define it as a partially observable Markov decision process; define it as a tuple , where is the state, is the observation, is the action, is the state transition, is the reward.
[0009] Replace the original action of each agent with a counterfactual action, and determine the counterfactual reward according to the reward between the state-action obtained after replacement and the original state-action; According to the multi-agent reinforcement learning environment, adopt a multi-agent proximal policy optimization method with a gating mechanism for communication actions to learn a collaborative communication strategy; among them, the counterfactual reward is used to evaluate the global task allocation value, the consistency of the final allocation, and communication conflicts during the learning process.
[0010] Multiple agents adopt a collaborative communication strategy during the execution of distributed task allocation to obtain the distributed task allocation result.
[0011] A fast distributed task allocation device based on a collaborative communication strategy, the device includes: The collaborative communication strategy optimization module is used to model the agent observation according to the task description; model the action as an adaptive gate for communication between agents during distributed task allocation; construct a counterfactual strategy and generate counterfactual actions for agents according to the counterfactual strategy; construct a multi-agent reinforcement learning environment for communication-aware distributed task allocation and define it as a partially observable Markov decision process; define it as a tuple , where is the state, is the observation, is the action, is the state transition, is the reward; replace the original action of each agent with a counterfactual action, and determine the counterfactual reward according to the reward between the state-action obtained after replacement and the original state-action; according to the multi-agent reinforcement learning environment, adopt a multi-agent proximal policy optimization method with a gating mechanism for communication actions to learn a collaborative communication strategy; among them, the counterfactual reward is used to evaluate the global task allocation value, the consistency of the final allocation, and communication conflicts during the learning process.
[0012] The fast distributed task allocation module is used for multiple agents to adopt a collaborative communication strategy during the execution of distributed task allocation to obtain the distributed task allocation result.
[0013] The above-mentioned fast distributed task allocation method and device based on the cooperative communication strategy. In the method, the observation, action, and reward mechanisms simultaneously consider the task-related and communication-related characteristics of the MAC layer. Through the centralized training and decentralized execution scheme, MAPPO is extended with counterfactual rewards to improve the learning speed and final performance. The learned cooperative communication strategy significantly reduces communication conflicts and competition and can be applied to any distributed task allocation algorithm, thereby accelerating convergence and improving performance. The trained communication strategy can effectively reduce the communication bandwidth required in the actual distributed network environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 FIG. is a schematic flowchart of a fast distributed task allocation method based on a cooperative communication strategy in an embodiment; Figure 2 FIG. is an overall architecture diagram of a fast distributed task allocation method based on a cooperative communication strategy in an embodiment; Figure 3 FIG. is a schematic diagram of performance comparison of several methods in different scenarios in an embodiment, where Figure 3 (a) is a schematic diagram of the comparison of task allocation values of several algorithms in a scenario of 20 agents and 100 tasks, Figure 3 (b) is a schematic diagram of the comparison of average rewards of several algorithms in a scenario of 20 agents and 100 tasks, Figure 3 (c) is a schematic diagram of the comparison of the number of message losses of several algorithms in a scenario of 20 agents and 100 tasks, Figure 3 (d) is a schematic diagram of the comparison of task allocation values of several algorithms in a scenario of 50 agents and 250 tasks, Figure 3 (e) is a schematic diagram of the comparison of average rewards of several algorithms in a scenario of 50 agents and 250 tasks, Figure 3 (f) is a schematic diagram of the comparison of the number of message losses of several algorithms in a scenario of 50 agents and 250 tasks; Figure 4 FIG. is a time slot diagram of data received by agents before and after applying the task allocation method in another embodiment, where Figure 4 (a) is a time slot diagram of data received by agents before applying the task allocation method, Figure 4 (b) is a time slot diagram of data received by agents after applying the task allocation method; Figure 5 FIG. is a schematic diagram of the transmission plans of three groups of experience trajectories and the variation of credit assignment values over time under different agents and different learning methods in another embodiment, where Figure 5 (a) is the transmission plan of three groups of experience trajectories, Figure 5(b)Schematic diagram of the comparison of credit assignment values for different agents under the CF-MAPPO method, Figure 5 (c)Schematic diagram of the comparison of credit assignment values for different agents under the QMIX method, Figure 5 (d)Schematic diagram of the comparison of credit assignment values for different agents under the VDN method; Figure 6 Schematic diagram of the statistical data of SGA rewards and required communication bandwidth before and after applying the communication strategy in an embodiment, where Figure 6 (a)Schematic diagram of the statistical data of SGA rewards before and after applying the communication strategy, Figure 6 (b)Schematic diagram of the statistical data of the required communication bandwidth. Detailed implementation manners
[0015] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0016] In one embodiment, as Figure 1 shown, a fast distributed task allocation method based on a cooperative communication strategy is provided, and the method includes the following steps: Step 100: Model the agent observations according to the task description.
[0017] Specifically, the task is a communication-aware task allocation task for a multi-agent system. The multi-agent system is composed of multiple agents, where the agents can be, but are not limited to, unmanned aerial vehicles, unmanned vehicles, robots, etc.
[0018] The multi-agent system can be applied in actual scenarios such as search and rescue, environmental monitoring, security monitoring, and power inspection. For example: A multi-agent system performs a search and rescue task in a wild environment.
[0019] Modeled the key issues in the real-world communication network based on the CSMA / CA protocol, such as message collision and channel contention.
[0020] Model the agent observations according to the task description to incorporate features such as Bron-Kerbosch (BK), value of message (VoM), and channel access quality, and also include a normalization step.
[0021] Step 102: Model the actions as an adaptive gating of the communication between agents during the distributed task allocation.
[0022] Specifically, the actions based on the gating mechanism enable the agents to adaptively select the data transmission time, thereby coordinating the communication between agents to improve the algorithm performance.
[0023] Step 104: Construct a counterfactual strategy and generate a counterfactual action for the agent according to the counterfactual strategy.
[0024] Step 106: Construct a multi-agent reinforcement learning environment for communication-aware distributed task allocation, and define it as a partially observable Markov decision process; define it as a tuple , where is the state, is the observation, is the action, is the state transition, is the reward.
[0025] Specifically, the performance of the distributed task allocation method may be affected by the actual communication process. The agent can learn a communication strategy to coordinate the communication time, reduce the data collision rate and improve the channel utilization rate, thereby accelerating the convergence speed of the allocation algorithm. For this purpose, a multi-agent reinforcement learning environment for communication-aware distributed task allocation is constructed and defined as a partially observable Markov decision process (POMDP).
[0026] The state transition function describes the mapping of the environment from state to the next state , and is usually defined as: ; where and represent the global states at times t and t +1 respectively, and represents the joint communication action of all agents.
[0027] Step 108: Replace the original action of each agent with the counterfactual action, and determine the counterfactual reward according to the reward between the state-action obtained after replacement and the original state-action.
[0028] Specifically, the counterfactual reward is used to evaluate the global task allocation value, the consistency of the final allocation and the communication conflict.
[0029] The actions and policies obtained according to the counterfactual reward.
[0030] Step 110: According to the multi-agent reinforcement learning environment, use the multi-agent proximal policy optimization method with a gating mechanism for communication actions to learn a collaborative communication strategy; where the counterfactual reward is used to evaluate the global task allocation value, the consistency of the final allocation and the communication conflict during the learning process.
[0031] Specifically, the communication policy learned through Multi-Agent Proximal Policy Optimization (MAPPO) significantly outperforms other learning methods, such as MTD3, QMIX, and VDN, in many metrics including allocated value, learning speed, and communication quality.
[0032] Step 112: Multiple agents adopt a cooperative communication policy during the execution of distributed task allocation to obtain a distributed task allocation result.
[0033] Specifically, applying the cooperative communication policy to five classic allocation algorithms: Consensus-Based Auction Algorithm (CBAA), Consensus-Based Bundle Algorithm (CBBA), Hybrid Information and Plan Consensus (HIPC), Distributed Genetic Algorithm (DGA), and Distributed Heuristic Task Allocation (DHBA) shows a consistent and significant improvement in coordination efficiency and final solution performance under various task settings.
[0034] In the above fast distributed task allocation method based on the cooperative communication policy, the observation, action, and reward mechanisms in the method simultaneously consider the task-related and communication-related characteristics of the MAC layer; through the centralized training and decentralized execution scheme, MAPPO is extended with counterfactual rewards to improve the learning speed and final performance. The learned cooperative communication policy significantly reduces communication conflicts and competition, can be applied to any distributed task allocation algorithm, thereby accelerating convergence and improving performance, and the trained communication policy can effectively reduce the communication bandwidth required in the actual distributed network environment.
[0035] In one embodiment, the agent observation in step 100 includes: Bron-Kerbosch features, message value features, channel access features; where the channel access features are used to calculate the channel access priority, enabling the agent to focus on the number of times it accesses the communication channel, thereby preventing excessive channel access; the channel access priority can be calculated through the arctangent function, which shows a significant change at the maximum backoff count, thus quickly forming an ordered and continuous backoff sequence for the agent; the expression of the channel access priority is: ; where, is the channel access priority, is the weighting factor, representing the maximum number of backoffs within a contention period, is the number of backoffs the agent has made in the channel.
[0036] In one embodiment, step 102 includes: at the t th step, the agent i inputs its local observation value into the Actor network to obtain the communication action , the communication action determines whether to send a message; when occurs, the agent does not use the communication channel at step t . When occurs, the agent uses the communication channel to send a message.
[0037] In one embodiment, the expression of the counterfactual reward in step 108 is: ; where represents the counterfactual reward of the agent i at step t , R (·) represents the reward function used to evaluate the state-action values of all agents, represents the joint observation of all agents at step t , represents the joint communication action of all agents at step t , represents the communication action of the agent i at step t , represents the estimated communication action of the agent i at step t .
[0038] In one embodiment, the estimation process of the counterfactual reward in step 108 includes: sampling the depth network to fit the reward function used to evaluate the state-action values of all agents; taking the joint state-action in the experience trajectory as the input of the depth network, and calculating the training gradient based on minimizing the TD error to update the parameters of the depth network; the update expression of the depth network parameters is: ; ; ; where represents the parameters of the depth network, represents the objective function, represents the hyperparameter, represents the value of the joint observation of the agent at step t + 1 with the joint communication action, represents the joint observation of all agents at step t + 1, represents the joint communication action of all agents at step t + 1, Represents the global shared reward, Indicates during the t step, the change in the average task conflict of the i agent, , respectively represent the average task conflict between the t step and the t +1 step at the start between the i agent and other agents.
[0039] After constructing the deep network, use the joint observation value and joint action at the t step to approximately calculate the counterfactual reward of the i agent; the calculation process of the counterfactual reward estimate is as follows: ; ; ; Among them, represents the counterfactual reward estimate, and respectively represent the deep network estimates for the original reward and the reward .
[0040] Specifically, in order to capture the complex cooperation relationship between agents in this environment and achieve reasonable credit assignment, this application proposes a fast distributed task allocation method based on a cooperative communication strategy. This method is a MAPPO extended algorithm based on the concept of differential rewards, abbreviated as: CF-MAPPO. This method replaces the original actions of the agents with counterfactual actions to infer their contributions to the cooperation relationship, thereby clearly allocating the global shared reward. The framework of CF-MAPPO is as Figure 2 shown. The original action of any i agent is replaced by a counterfactual action. Drawing on the concept of differential rewards, the difference between the state-action obtained after replacement and the original state-action is used as a substitute for the reward allocated to it, that is, the counterfactual reward: ; Among them, R (·) represents the reward function used to evaluate the state-action values of all agents.
[0041] In addition, as Figure 2 shown, use the deep network to fit R (·) and estimate the counterfactual reward to predict the action-state value. To train the deep The network takes the joint state-action in the experience trajectory as input and calculates the training gradient based on minimizing the TD error to update the value network parameters That is: ; ; After constructing the value function, the joint observation value and the joint action are used to approximately calculate the counterfactual reward i of the agent ; The calculation process of the counterfactual reward is as follows: ; ; ; Among them, represents the counterfactual reward estimate value, and respectively represent the value network estimates for the original reward and the reward .
[0042] In one embodiment, the multi-agent proximal policy optimization method with a gating mechanism communication action in step 110 uses a centralized training - decentralized execution method to train the shared policy as the Actor network and train the shared value as the Critic network; The update expression of the Critic network parameters is: ; Among them, represents the Critic network parameters, represents the state value function, represents the global shared reward.
[0043] The update expression of the Actor network parameters is: ; ; ; Among them, represents the Actor network parameters, and respectively represent the advantage and entropy of the policy, represents the action-state value of the agent, represents the weighting factor of the advantage, represents the advantage evaluation function, represents the advantage function, denotes a hyperparameter.
[0044] Specifically, at the beginning of each step, the agent determines whether to transmit data according to the communication action. During the communication process, the agent autonomously occupies the channel according to the communication strategy to avoid data conflicts. During the calculation process, the agent resolves task conflicts based on the received messages and the consistency conflict resolution rules. During the communication process, the agent autonomously occupies the channel according to the communication strategy to avoid data conflicts. During the calculation process, the agent resolves task conflicts based on the received messages and the consistency conflict resolution rules. A cooperation relationship is introduced to coordinate communication. This can be achieved by establishing a common goal among agents through communication behaviors based on the gate mechanism. The global shared reward r t can be set to reduce the global task conflict within a fixed time period, that is: ; where denotes the change in the average task conflict of agent t during the i -th step, denotes the average task conflict between agent t and other agents at the beginning of the i -th step, that is , denotes the task list of agent i , denotes the task list of agent j .
[0045] In one embodiment, during the centralized training phase, the policy advantage is calculated using the training samples with the added counterfactual reward as: ; where denotes the policy advantage, denotes a hyperparameter, denotes the counterfactual reward of agent i at the t -th step.
[0046] In one embodiment, the expression of the weighted factor of the advantage is: ; where , respectively denote the policy distributions of the agent before and after updating , denotes the KL divergence.
[0047] Specifically, during the communication process of the distributed algorithm, the smaller the amount of data loss caused by communication, the more the number of task conflicts among agents decreases, and thus the greater the reward feedback obtained from the system. Then, a multi-agent proximal policy optimization algorithm with a gating mechanism for communication actions is proposed to learn collaborative communication strategies. First, centralized training and decentralized execution (CTDE) are used to share the state-action information of all agents, helping them identify cooperative relationships and eliminate potential conflicts between policies, thereby improving the learning of collaborative strategies among agents. Second, counterfactual rewards are constructed to distinguish the contributions of different agent policies to teamwork, and the MAPPO method based on counterfactual rewards is proposed to achieve more accurate policy updates. Since the communication action based on the gating mechanism is a discrete action, the proximal policy optimization (PPO) algorithm based on discrete actions is used for training. The method of centralized training - decentralized execution is adopted to train the shared policy as the behavior network and the shared value as the evaluation network. Each agent i uses the local observation to obtain an action , so as to achieve the decentralized execution of different actions. This method helps agents identify cooperative relationships and eliminates potential conflicts between policies by sharing state and reward information among agents. In the decentralized execution stage, the agent i at the t th step inputs its local observation value into the behavior network to obtain an action , thereby determining whether the agent sends data at the t th step. This constructs the global shared reward obtained by all agents through joint actions t and environmental feedback at the th step. At the same time, each agent i inputs its local observation value into the value network to obtain the state value t corresponding to the th step. Therefore, each agent i can collect training samples to construct its corresponding empirical trajectory , and store it in the buffer for subsequent training. On the other hand, in the centralized training stage, the main goal is to update the value network and the policy network using the collected empirical trajectories. For the value network, the empirical trajectories of all agents are used to update its parameters to fit the state value function. Then, we calculate the training gradient by minimizing the mean square error between the state value function and the global shared reward . Based on this error, the value network parameters The update method is as shown in the above Critic network parameter update expression, which applies to all agents. For the Actor network, the Generalized Advantage Estimation (GAE) method is usually used to update its parameters, that is: ; where represents the action-state value of the agent. It can be estimated by adding the global shared reward to .
[0048] To avoid too large a difference in the policy distribution before and after parameter update, the KL divergence is introduced as a weighted factor of the advantage , and the truncation function is used to set an upper limit for . This can jointly limit the range of the policy gradient, thus ensuring the stability of the policy update, that is: ; where , , respectively represent the policy distributions of the agent before and after updating .
[0049] Finally, the Actor network parameters are updated in the following way : ; where and respectively represent the advantage and entropy of the policy. It should be noted that the above advantage depends on the global shared reward .
[0050] However, a unified reward shared globally is difficult to distinguish the contributions of different policies. In particular, this may lead to high reward values for ineffective communication policies, ultimately affecting the update of the Actor network.
[0051] In one of the embodiments, the method adopted for distributed task allocation in step 112 is a consensus-based auction algorithm, a consensus-based bundling algorithm, hybrid information and plan consensus, a distributed genetic algorithm, or a distributed heuristic task allocation.
[0052] It should be understood that although each step in the flowchart of Figure 1 is shown in sequence according to the indication of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 1At least a part of the steps may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed and completed at the same moment, but can be executed at different moments, and the execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turns with at least a part of other steps or sub-steps or stages of other steps.
[0053] In a validation embodiment, extensive simulation experiments were conducted to evaluate the effectiveness of communication strategy optimization based on multi-agent reinforcement learning (MARL) in distributed task allocation. When evaluating these algorithms, the existing simple internal simulator code was used and extended to train the strategy. The simulator modeled key issues in real-world communication networks based on the CSMA / CA protocol, such as message collision and channel contention. Under simulation conditions, MAPPO and CF-MAPPO were used to learn communication strategies based on a gating mechanism, and performance was evaluated using metrics such as the conflict ratio, assigned task value (SGA reward), total data conflict, and average reward. To further verify the applicability of the proposed method, agent data was recorded and communication strategies were trained on a self-organizing network simulation platform server (clrsp) for hybrid virtual reality. The platform constructs networks with different topologies by simulating the distances and communication ranges of communication nodes. It also runs a five-layer network communication protocol on each node to replicate real-world networks and communication environments. The distributed task allocation algorithm program runs as an independent thread, and each agent corresponds to a node on the server. The distributed task allocation method is mainly evaluated from the following four aspects: 1) Conflict task ratio: The ratio of the number of conflict tasks to the total number of tasks when the algorithm converges or terminates.
[0054] 2) Assigned task value: The ratio of the total reward of the assigned tasks to the total reward obtained by centralized task allocation when the algorithm converges or terminates.
[0055] 3) Total data conflict: The sum of data conflicts of all agents in the environment in one round.
[0056] 4) Average reward: The environmental reward generated by reinforcement learning represents the change in the average number of task conflicts in this scenario.
[0057] To verify the effectiveness of the proposed method, scenarios with 20 agents and 100 tasks and 50 agents and 250 tasks were considered, and this was done under 5 different random seeds. The learning speed and final performance of the learned strategies were compared with three well-known multi-agent reinforcement learning methods: MTD3, QMIX, and VDN. The performance comparison of several methods in different scenarios is as Figure 3 shown, where Figure 3(a)Schematic diagram of the task assignment value comparison of several algorithms in the scenario of 20 agents and 100 tasks, Figure 3 (b)Schematic diagram of the average reward comparison of several algorithms in the scenario of 20 agents and 100 tasks, Figure 3 (c)Schematic diagram of the message loss quantity comparison of several algorithms in the scenario of 20 agents and 100 tasks, Figure 3 (d)Schematic diagram of the task assignment value comparison of several algorithms in the scenario of 50 agents and 250 tasks, Figure 3 (e)Schematic diagram of the average reward comparison of several algorithms in the scenario of 50 agents and 250 tasks, Figure 3 (f)Schematic diagram of the message loss quantity comparison of several algorithms in the scenario of 50 agents and 250 tasks. It can be seen from Figure 3 that in all scenarios, the proposed method is significantly better than the value decomposition methods QMIX and VDN. In particular, the target value of task assignment ranks the highest after convergence, and MTD3 is the lowest (0.68 vs. 0.43). Similar trends can also be found in the comparison of cumulative rewards (0.78 vs. 0.2) and message loss quantities (7 vs. 25). This further shows the advantage of restricting the maximum difference between the old and new policies, although the counterfactual reward improves the accuracy of the actor network. In contrast, the simple decomposition of the global state value in QMIX and VDN ignores the strong dependencies between adjacent agents, resulting in poor training performance in all scenarios. To demonstrate the effectiveness of the communication strategy, a time-slot graph of the data received by the agents is used to depict the transmission and reception states of all agents before and after training; Figure 4 is the time-slot graph of the data received by the agents before and after applying the proposed method for task assignment, where Figure 4 (a)is the time-slot graph of the data received by the agents before applying the proposed method for task assignment, Figure 4 (b)is the time-slot graph of the data received by the agents after applying the proposed method for task assignment. In the scenario with 20 agents, the cooperative communication strategy significantly reduces data conflicts and channel competition, thus improving the performance. Credit assignment is crucial for determining whether the actor network of the agents can be correctly updated. To analyze this, we used training samples from three sets of empirical trajectories in a scenario with five agents, and these samples contain random actions. The credit assignment values of the CF-MAPPO, QMIX, and VDN algorithms were visualized and processed using min-max normalization. Figure 5 Shows the transmission plans of the three sets of empirical trajectories and the variation of the credit assignment values over time under different agents and different learning methods, where Figure 5 (a)is the transmission plan of the three sets of empirical trajectories, Figure 5(b)Schematic diagram for comparing the credit assignment values of different agents under the CF-MAPPO method, Figure 5 (c)Schematic diagram for comparing the credit assignment values of different agents under the QMIX method, Figure 5 (d)Schematic diagram for comparing the credit assignment values of different agents under the VDN method.
[0058] It can be seen that compared with QMIX and VDN, the CF-MAPPO method proposed in this application can more accurately evaluate the contributions of different agents under their local policies, and both its maximum value (0.3 compared with 1.0 and 0.95) and minimum value (0.0 compared with 0.3 and 0.4) are different. This is because the CF-MAPPO method determines the assigned credit based on counterfactual rewards, while QMIX and VDN mainly calculate the credit based on the value function and the corresponding states of each agent. More specifically, as shown in Figure 5, at step 10, the two hidden nodes (Agent 2 and Agent 4) did not simultaneously select the channel occupancy strategy, thereby reducing the occurrence rate of communication conflicts and obtaining higher counterfactual rewards.
[0059] To verify the applicability of the actions based on the gating mechanism to different allocation methods, training and testing were carried out by replacing various allocation algorithms. The communication strategy was combined with several classic distributed algorithms including CBBA, CBAA, HIPC, DGA, and DHBA for experiments, and these algorithms have different scales. Monte Carlo simulations show that compared with the methods without the communication strategy, the allocation algorithms combined with the communication strategy reduce the average task conflict rate by up to 29%, 40%, and 21% respectively in the scenarios of 10, 20, and 50 agents, and reduce the total data loss rate by up to 42%, 80%, and 89% respectively. Among them, for the redundancy-based central method DHBA, after combining the communication strategy, the average task conflict rate decreases slightly, by 18%, 7%, and 3% respectively in the scenarios of 10, 20, and 50 agents. In contrast, after incorporating the communication strategy, CBAA has a more significant reduction in the average task conflict rate, by 17%, 40%, and 17% respectively for the scenarios of 10, 20, and 50 agents. In summary, the proposed communication strategy can adapt to various distributed task allocation algorithms.
[0060] To verify the effectiveness of the learning framework under real network conditions, a distributed algorithm with a communication process is simulated on a self-organizing network simulation platform server that closely resembles real network conditions, and data is collected distributively to train the communication strategy. This platform can not only construct the topological structure of a real self-organizing network but also simulate data collisions, losses, and errors through the MAC layer network protocol to reproduce the distributed communication scenario. On this platform, the experimental settings are as follows: (1) The goal is to reduce the bandwidth required for communication between agents and optimize the reward design; (2) Since it is difficult to set a unified termination time for the distributed task allocation algorithm, by default, the algorithm running time for each node is fixed at 1.5 seconds. (3) The distributed CBBA algorithm is used for node operations, and experiments are conducted under conditions where the total communication bandwidths are 4.5 Mbps, 9.5 Mbps, and 15.2 Mbps respectively. The trained strategy is verified through 50 Monte Carlo simulations (30 agents in each group). Figure 6 It is a schematic diagram of the statistical data of SGA rewards and the required communication bandwidth before and after applying the communication strategy in an embodiment, where Figure 6 (a) is a schematic diagram of the statistical data of SGA rewards before and after applying the communication strategy, Figure 6 (b) is a schematic diagram of the statistical data of the required communication bandwidth. It can be seen from Figure 6 that before and after applying the communication strategy, the SGA reward of CBBA is hardly affected, while the communication bandwidth required for data transmission is significantly reduced. Under the bandwidth conditions of 4.5 Mbps, 9.5 Mbps, and 15.2 Mbps, the bandwidth required for the CBBA algorithm to send data is reduced by approximately 43%, 48%, and 42% respectively. These results indicate that the gated communication strategy learning framework is applicable to a nearly real network environment and can significantly reduce the required communication bandwidth without affecting the algorithm performance.
[0061] In an embodiment, a fast distributed task allocation device based on a collaborative communication strategy is provided, including: a collaborative communication strategy optimization module and a fast distributed task allocation module, where: The collaborative communication strategy optimization module is used to model the agent observations according to the task description; model the actions as an adaptive gating for communication between agents during distributed task allocation; construct a counterfactual strategy and generate counterfactual actions for the agents according to the counterfactual strategy; construct a multi-agent reinforcement learning environment for communication-aware distributed task allocation and define it as a partially observable Markov decision process; define it as a tuple where is the state, is the observation, is the action, is the state transition, As a reward; the original action of each agent is replaced by a counterfactual action, and the counterfactual reward is determined according to the reward between the state-action obtained after replacement and the original state-action; according to the multi-agent reinforcement learning environment, a multi-agent proximal policy optimization method with a gating mechanism for communication actions is used to learn a collaborative communication strategy; among them, the counterfactual reward is used to evaluate the global task allocation value, the consistency of the final allocation, and communication conflicts during the learning process.
[0062] A fast distributed task allocation module, which is used for multiple agents to adopt a collaborative communication strategy during the execution of distributed task allocation to obtain a distributed task allocation result.
[0063] In one embodiment, the agent observations in the collaborative communication strategy optimization module include: Bron-Kerbosch features, message value features, and channel access features; among them, the channel access features are used to calculate the channel access priority, so that the agent pays attention to the number of times it accesses the communication channel, thereby preventing excessive access to the channel; the channel access priority is as shown in the expression of the above channel access priority.
[0064] In one embodiment, the collaborative communication strategy optimization module is also used in the t step, the agent i inputs its local observation value into the Actor network to obtain a communication action , and the communication action determines whether to send a message; when , the agent does not use the communication channel in the t step, and when , the agent uses the communication channel to send a message.
[0065] In one embodiment, the counterfactual reward in the collaborative communication strategy optimization module is as shown in the expression of the above counterfactual reward.
[0066] In one embodiment, the estimation process of the counterfactual reward in the collaborative communication strategy optimization module includes: sampling the depth network to fit the reward function used to evaluate the state-action values of all agents; using the joint state-action in the experience trajectory as the input of the depth network, and calculating the training gradient based on minimizing the TD error to update the parameters of the depth network; the update method of the depth network parameters is as shown in the expression of the update of the above depth network parameters.
[0067] After constructing the depth network, use the joint observation value and joint action in the t step to approximately calculate the agent iCounterfactual reward. The specific calculation process of the counterfactual reward is as shown in the above calculation process of the counterfactual reward.
[0068] In one embodiment, the multi-agent proximal policy optimization method with gated mechanism communication actions in the collaborative communication policy optimization module uses the centralized training - decentralized execution method to train the shared policy as the Actor network and train the shared value as the Critic network; the parameter update method of the Critic network is as shown in the above Critic network parameter update expression; the parameter update method of the Actor network is as shown in the above Actor network parameter update expression.
[0069] In one embodiment, in the centralized training stage of the multi-agent proximal policy optimization method with gated mechanism communication actions in the collaborative communication policy optimization module, the training samples added with counterfactual rewards are used to calculate the policy advantage; the policy advantage is as shown in the above expression of the policy advantage.
[0070] In one embodiment, the weighting factor of the advantage is as shown in the above expression of the weighting factor of the advantage.
[0071] In one embodiment, the method adopted for distributed task allocation in the fast distributed task allocation module is the consensus-based auction algorithm, the consensus-based bundling algorithm, the hybrid information and plan consensus, the distributed genetic algorithm, or the distributed heuristic task allocation.
[0072] For the specific limitations of the fast distributed task allocation device based on the collaborative communication policy, reference can be made to the limitations of the fast distributed task allocation method based on the collaborative communication policy in the above text, which will not be elaborated here. Each module in the above fast distributed task allocation device based on the collaborative communication policy can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0073] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0074] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A fast distributed task allocation method based on collaborative communication strategy, characterized in that: The method comprises: Model the agent observations based on the task description; Modeling actions as adaptive gating of inter-agent communication during distributed task allocation; Constructing a counterfactual strategy and generating counterfactual actions of the agent according to the counterfactual strategy; Construct a multi-agent reinforcement learning environment for communication-aware distributed task allocation, which is defined as a partially observable Markov decision process; defined as a tuple ,in For status, For observation, For action, For state transfer, For reward; The original action of each agent is replaced by a counterfactual action, and the counterfactual reward is determined based on the reward between the state-action obtained after the replacement and the original state-action; According to the multi-agent reinforcement learning environment, a multi-agent proximal strategy optimization method with a gating mechanism communication action is used to learn a collaborative communication strategy; wherein counterfactual rewards are used in the learning process to evaluate the global task allocation value, the consistency of the final allocation, and the communication conflict; Multiple intelligent agents adopt the cooperative communication strategy in the process of executing distributed task allocation to obtain distributed task allocation results.
2. The fast distributed task allocation method based on collaborative communication strategy according to claim 1 is characterized in that: According to the task description, the agent observation is modeled. The agent observation in the step includes: Bron-Kerbosch feature, message value feature, channel access feature; wherein the channel access feature is used to calculate the channel access priority, so that the agent pays attention to the number of times it accesses the communication channel, thereby preventing excessive access to the channel; the channel access priority is: ; in, is the channel access priority, is a weighted factor graph, indicating the maximum number of backoffs in a contention cycle. is the number of backoffs the agent has performed in the channel.
3. The fast distributed task allocation method based on collaborative communication strategy according to claim 1 is characterized in that: Modeling actions as adaptive gating of inter-agent communication during distributed task allocation includes: In the t Step 1, Agent i Its local observation Input to the Actor network to get the communication action , the communication action determines whether to send a message; when When the agent is in t The communication channel is not used. When , the agent uses the communication channel to send messages.
4. The fast distributed task allocation method based on collaborative communication strategy according to claim 1 is characterized in that: The counterfactual reward is: ; in, Representing an Agent i In the t The counterfactual reward for the step R (·) represents the reward function used to evaluate the state-action values of all agents, Indicates that all agents in t The joint observation of the step Indicates that all agents in t The joint communication action of the step, Representing an Agent i In the t The communication action of the step, Represents the estimated agent i In the t The communication action of the step.
5. The fast distributed task allocation method based on collaborative communication strategy according to claim 4 is characterized in that: The estimation process of the counterfactual reward includes: Sampling Depth The network fits a reward function for evaluating the state-action values of all agents; the joint state-action in the experience trajectory is used as the depth The network input is used to calculate the training gradient based on minimizing the TD error and update the depth. Parameters of the network; depth The network parameter update expression is: ; ; ; in, Indicates depth The parameters of the network, represents the objective function, represents the hyperparameter, represents the global shared reward, Indicates that the agent t The joint observation of step +1 uses the value of the joint communication action, Indicates that all agents in t +1 step of joint observation, Indicates that all agents in t +1 step of joint communication action, Indicated in t Agent during step i The change in average task conflict, , Respectively expressed in t Step and t +1 step at the beginning of the agent i Average task conflicts with other agents, n represents the number of agents; Build Depth After the network, use the t step to approximate the agent by combining observations and joint actions i The counterfactual reward of ; the calculation process of the counterfactual reward estimate is as follows: ; ; ; in, represents the counterfactual reward estimate, and Respectively represent the original rewards and rewards Depth Network estimate.
6. The fast distributed task allocation method based on collaborative communication strategy according to claim 1 is characterized in that: The multi-agent proximal policy optimization method with gating mechanism communication actions adopts a centralized training-decentralized execution method to train the shared strategy as the Actor network and the shared value as the Critic network; The critic network parameter update method is: ; in, represents the Critic network parameters, represents the state value function, represents the global shared reward; Actor network parameters are updated as follows: ; ; ; in, Represents the Actor network parameters, and denote the advantage and entropy of the strategy respectively, represents the action-state value of the agent, represents the weighting factor of advantage, represents the hyperparameter, represents the advantage evaluation function, represents the advantage function, represents the hyperparameter, Representing an Agent i In the t The communication action of the step, Representing an Agent i In the t Local observation of the step.
7. The fast distributed task allocation method based on collaborative communication strategy according to claim 6 is characterized in that: In the centralized training phase, the strategy advantage is calculated using training samples with added counterfactual rewards: ; in, Indicates strategic advantage, represents a hyperparameter, Representing an Agent i In the t The counterfactual reward for the step.
8. The fast distributed task allocation method based on collaborative communication strategy according to claim 6 is characterized in that: The weighting factor for the advantage is: ; in, , Respectively represent the agent in updating Strategy distribution before and after, represents the KL divergence.
9. The fast distributed task allocation method based on collaborative communication strategy according to claim 1 is characterized in that: The methods used for distributed task allocation are consensus-based auction algorithms, consensus-based bundling algorithms, hybrid information and planning consensus, distributed genetic algorithms, or distributed heuristic task allocation.
10. A fast distributed task allocation device based on collaborative communication strategy, characterized in that: The device comprises: The collaborative communication strategy optimization module is used to model agent observations based on task descriptions; model actions as adaptive gating of inter-agent communication during distributed task allocation; construct counterfactual strategies and generate counterfactual actions of agents based on the counterfactual strategies; construct a multi-agent reinforcement learning environment for communication-aware distributed task allocation, which is defined as a partially observable Markov decision process; and define it as a tuple ,in For status, For observation, For action, For state transfer, as reward; the original action of each agent is replaced by a counterfactual action, and the counterfactual reward is determined according to the reward between the state-action obtained after the replacement and the original state-action; according to the multi-agent reinforcement learning environment, a multi-agent proximal strategy optimization method with a gating mechanism communication action is used to learn a cooperative communication strategy; wherein the counterfactual reward is used in the learning process to evaluate the global task allocation value, the consistency of the final allocation and the communication conflict; The fast distributed task allocation module is used for multiple intelligent agents to adopt the cooperative communication strategy in the process of executing distributed task allocation to obtain distributed task allocation results.
Citation Information
Patent Citations
Multi-machine collaborative air combat planning method and system based on deep reinforcement learning
CN112861442A
Unmanned aerial vehicle cluster near-end strategy optimization collaborative confrontation method based on anti-fact baseline
CN118778678A
Unmanned aerial vehicle cooperative air combat decision-making method based on GRU-MAPPO deep reinforcement learning
CN119129413A
Cited By
Heterogeneous aircraft cluster multi-scale cross-domain autonomous confrontation decision-making method
CN120353241A
Display panel production scheduling method and system based on agent cooperation
CN120562842A