Communication sensing task allocation method and device based on multi-target asynchronous strategy
By transforming the distributed task allocation problem in the drone swarm into asynchronous constraints and dispersed parts, Markov decision-making process can be observed, and the communication strategy is optimized using the multi-objective coupled PPO method, the problem of excessive communication overhead and delay in the drone swarm is solved, and multi-objective asynchronous decision-making and efficient communication are achieved.
Patent Information
- Application Number
- CN202510512193.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The prior art is difficult to effectively optimize communication strategies in distributed task allocation in drone clusters, resulting in excessive communication overhead and delays, and it is difficult to achieve multi-objective asynchronous decision-making.
The communication perception task allocation method based on multi-objective asynchronous strategy is adopted, and the Markov decision-making process can be observed by transforming the communication strategy optimization problem of the multi-agent system into the asynchronous constraint dispersed part, an asynchronous constraint multi-agent reinforcement learning environment is constructed, and a multi-objective coupled PPO method is used to optimize the multi-objective asynchronous communication strategy.
It realizes that under the constraints of a given bandwidth requirement, optimizes the communication strategy of the drone cluster, reduces bandwidth overhead and task conflicts, and improves communication efficiency and allocation reliability.
Smart Images

Figure CN120029326A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of task allocation, and in particular to a communication-aware task allocation method and device based on a multi-target asynchronous strategy. Background Art
[0002] Distributed task allocation in drone swarms is very sensitive to high bandwidth requirements and frequent inter-drone communications. Combining reinforcement learning with traditional distributed task allocation shows great potential in improving algorithm performance and optimizing communication. However, existing research relies on ideal bandwidth assumptions and unrealistic time synchronization, making training and validation impractical under real conditions.
[0003] With the breakthrough of artificial intelligence technology, combining reinforcement learning with traditional distributed task allocation methods has shown great potential in improving algorithm performance and optimizing communication. Existing research is difficult to implement on actual drones for the following reasons: 1) Communication strategy optimization methods based on multi-agent reinforcement learning assume that all drones execute decisions synchronously to obtain consistent rewards. However, in real-world drone swarms, drones are unable to achieve synchronous decision-making capabilities due to differences in performance, sensor configuration, and processing power. Frequent communication and synchronization time will lead to huge communication overhead and delay, making synchronous multi-agent reinforcement learning (MARL) training challenging. 2) In the distributed task allocation problem, the group must optimize multiple objectives at the same time. It needs to reduce the task conflict rate between drones and minimize the communication overhead of the entire system, which makes it difficult to apply traditional multi-agent reinforcement learning methods that only optimize a single objective. Summary of the invention
[0004] Based on this, it is necessary to provide a communication-aware task allocation method and device based on a multi-objective asynchronous strategy to address the above technical problems.
[0005] A communication perception task allocation method based on a multi-objective asynchronous strategy is proposed. The method is applicable to a UAV cluster to realize communication perception task allocation. Each UAV in the UAV cluster deploys an agent, and the UAV cluster constitutes a multi-agent system. The method includes: According to the bandwidth constraint, the communication strategy optimization problem in the distributed task allocation algorithm of multi-agent system is transformed into an asynchronous constrained decentralized partially observable Markov decision process.
[0006] Based on the asynchronous constrained decentralized partially observable Markov decision process, an asynchronous constrained multi-agent reinforcement learning environment considering the communication process is constructed.
[0007] According to the asynchronous constrained multi-agent reinforcement learning environment, the multi-objective coupling PPO method is used to optimize the multi-objective asynchronous communication strategy in the distributed task allocation method to obtain the final multi-objective asynchronous communication strategy; the multi-objective coupling PPO method refers to a method that extends the Lagrangian relaxation method to the multi-agent proximal strategy optimization algorithm, collects data asynchronously in a distributed manner, and then uses the Lagrangian function to centrally optimize the communication strategy under the given bandwidth demand constraint, thereby achieving simultaneous optimization of multiple objectives and obtaining the final multi-objective asynchronous communication strategy.
[0008] Each UAV adopts the final multi-target asynchronous communication strategy to complete the UAV cluster communication perception task allocation.
[0009] A communication perception task allocation device based on a multi-objective asynchronous strategy is suitable for implementing communication perception task allocation in a drone cluster. Each drone in the drone cluster deploys an agent, and the drone cluster constitutes a multi-agent system. The device includes: The problem modeling module is used to transform the communication strategy optimization problem in the distributed task allocation algorithm of the multi-agent system into an asynchronous constrained decentralized partially observable Markov decision process according to the bandwidth constraint.
[0010] The asynchronous constrained multi-agent reinforcement learning environment building module is used to disperse the partially observable Markov decision process according to the asynchronous constraints and build an asynchronous constrained multi-agent reinforcement learning environment that takes into account the communication process.
[0011] The multi-objective asynchronous communication strategy optimization module is used to optimize the multi-objective asynchronous communication strategy in the distributed task allocation method according to the asynchronous constrained multi-agent reinforcement learning environment, and obtain the final multi-objective asynchronous communication strategy; the multi-objective coupling PPO method refers to a method of extending the Lagrangian relaxation method to the multi-agent proximal strategy optimization algorithm, asynchronously collecting data in a distributed manner, and then using the Lagrangian function to centrally optimize the communication strategy under the given bandwidth demand constraint, to achieve simultaneous optimization of multiple objectives, and obtain the final multi-objective asynchronous communication strategy.
[0012] Based on the communication perception task allocation module, each UAV adopts the final multi-target asynchronous communication strategy to complete the UAV cluster communication perception task allocation.
[0013] The communication perception task allocation method and device based on multi-objective asynchronous strategy, according to the bandwidth constraint, transforms the communication strategy optimization problem in the distributed task allocation algorithm of the multi-agent system into an asynchronous constraint decentralized partially observable Markov decision process; according to the asynchronous constraint decentralized partially observable Markov decision process, constructs an asynchronous constraint multi-agent reinforcement learning environment considering the communication process; according to the asynchronous constraint multi-agent reinforcement learning environment, the multi-objective coupling PPO method is used to optimize the multi-objective asynchronous communication strategy in the distributed task allocation method to obtain the final multi-objective asynchronous communication strategy; each drone adopts the final multi-objective asynchronous communication strategy to complete the drone cluster communication perception task allocation. This method uses the multi-objective coupling PPO method to extend the concept of constrained reinforcement learning to the multi-agent proximal strategy optimization algorithm, and collects data asynchronously in a distributed manner; uses the Lagrangian function to centrally optimize the communication strategy under the given bandwidth demand constraint, balances the communication demand and task conflict, and achieves the simultaneous optimization of multiple goals, while reducing bandwidth overhead and minimizing task conflicts, and achieving a better balance between communication efficiency and allocation reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 A flow chart of a communication-aware task allocation method based on a multi-target asynchronous strategy in one embodiment; Figure 2 FIG. 1 is a schematic diagram of synchronous and asynchronous training data processing in another embodiment, wherein Figure 2 (a) is a schematic diagram of synchronous training data processing. Figure 2 (b) is a schematic diagram of asynchronous training data processing based on serial connection; Figure 3 Schematic diagram of the results of a verification experiment in another embodiment, wherein Figure 3 (a) to Figure 3 (c) Schematic diagram of mission conflict rate, bandwidth requirement and average communication times in the scenario of 30 drones. Figure 3 (d) to Figure 3 (f) Schematic diagram of mission conflict rate, bandwidth requirements and average communication times in a scenario with 40 drones. DETAILED DESCRIPTION
[0015] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0016] In one embodiment, Figure 1As shown, a communication perception task allocation method based on a multi-objective asynchronous strategy is provided. The method is applicable to a UAV cluster to realize communication perception task allocation. Each UAV in the UAV cluster deploys an agent, and the UAV cluster constitutes a multi-agent system. The method includes the following steps: Step 100: According to the bandwidth constraint, the communication strategy optimization problem in the distributed task allocation algorithm of the multi-agent system is transformed into an asynchronous constrained decentralized partially observable Markov decision process.
[0017] Specifically, a multi-objective asynchronous multi-agent reinforcement learning method is used to solve the communication-aware task allocation problem based on a multi-objective asynchronous strategy. The asynchronous allocation algorithm for integrated communication is formalized as an asynchronous partially observable Markov decision process, and the original learning objective is extended to adapt to asynchronous requirements.
[0018] The asynchronous constrained decentralized partially observable Markov decision process (ACDEC-POMDP) is used as the model of this problem, which is defined as the tuple ,in express n A collection of intelligent agents, S Represents the global state, and Represents all agents at time step t The set of observations and actions. is the state transition function. R represents the reward function, which provides feedback from the environment based on the action. Represents the discount factor. Represents multiple auxiliary constraints that are used to limit the updates of the agent's policy.
[0019] The Bron-Kerbosch feature (BK feature) and the message value feature (VoM feature) are part of the observation, and also include a normalization step. On this basis, a channel access feature is proposed, which calculates the channel access priority so that the agent pays attention to the number of times it accesses the communication channel. The channel access priority can be calculated by the inverse tangent function, which shows a significant change at the maximum number of transmissions and can quickly form an ordered and continuous abandonment sequence for the agent. The specific calculation formula is as follows: ; Among them, the adjustment factor Indicates the maximum number of transmissions in a contention cycle, and β is the number of transmissions of the agent in the channel.
[0020] The adaptive gating mechanism between agents is the action, and the shared reward reflects the changes in the global task conflict.
[0021] Step 102: Based on the asynchronous constrained decentralized partially observable Markov decision process, an asynchronous constrained multi-agent reinforcement learning environment considering the communication process is constructed.
[0022] Step 104: According to the asynchronous constrained multi-agent reinforcement learning environment, a multi-objective coupling PPO method is used to optimize the multi-objective asynchronous communication strategy in the distributed task allocation method to obtain a final multi-objective asynchronous communication strategy; the multi-objective coupling PPO method refers to a method of extending the Lagrangian relaxation method to a multi-agent proximal strategy optimization algorithm, asynchronously collecting data in a distributed manner, and then using the Lagrangian function to centrally optimize the communication strategy under a given bandwidth demand constraint, thereby achieving simultaneous optimization of multiple objectives and obtaining the final multi-objective asynchronous communication strategy.
[0023] Specifically, the performance of distributed task allocation methods may be affected by the actual communication process. The key elements of ACDEC-POMDP are now introduced to solve the communication strategy learning problem in asynchronous algorithms. In order to enable agents to utilize the state information of other agents for better strategy collaboration, a standard learning paradigm widely used in multi-agent reinforcement learning, namely centralized training and decentralized execution (CTDE), is adopted. The method is divided into two phases: centralized training phase and decentralized execution phase. In the centralized training phase, the centralized evaluation network evaluates the global state and processes the data at each time step. In the decentralized execution phase, each agent obtains its local observations at different global times and executes its actions independently.
[0024] In distributed task allocation algorithms, the optimization of communication strategies can be viewed as a multi-objective optimization problem, which aims to maximize the agent's global task reward while minimizing the communication overhead. To solve this problem, the concept of constrained reinforcement learning is extended to a multi-agent proximal policy optimization algorithm, and a multi-objective coupled PPO method is proposed. This method collects data asynchronously in a distributed manner; then, a Lagrangian function is used to centrally optimize the communication strategy under a given bandwidth requirement constraint, thereby achieving simultaneous optimization of multiple objectives.
[0025] Step 106: Each UAV adopts the final multi-target asynchronous communication strategy to complete the UAV cluster communication perception task allocation.
[0026] In the above-mentioned communication-aware task allocation method based on multi-objective asynchronous strategy, the method transforms the communication strategy optimization problem in the distributed task allocation algorithm of the multi-agent system into an asynchronous constrained decentralized partially observable Markov decision process according to the bandwidth constraint; constructs an asynchronous constrained multi-agent reinforcement learning environment considering the communication process according to the asynchronous constrained decentralized partially observable Markov decision process; according to the asynchronous constrained multi-agent reinforcement learning environment, the multi-objective coupled PPO method is used to optimize the multi-objective asynchronous communication strategy in the distributed task allocation method to obtain the final multi-objective asynchronous communication strategy; each drone adopts the final multi-objective asynchronous communication strategy to complete the drone cluster communication perception task allocation. This method adopts the multi-objective coupled PPO method to simultaneously reduce bandwidth overhead and minimize task conflicts, achieving a better balance between communication efficiency and allocation reliability.
[0027] In one embodiment, the expression of the tuple of the asynchronous constrained decentralized partially observable Markov decision process in step 100 is: ; in, A tuple representing an asynchronous constrained decentralized partially observable Markov decision process, express n A collection of intelligent agents, S Represents the global state, and Represents all agents at time step t The set of observations and actions, represents the state transition function, represents the new global state, R represents the reward function, represents the discount factor, Represents multiple auxiliary constraints, which are used to limit the update of the agent strategy. Indicates Auxiliary constraints, .
[0028] Asynchronous Constrained Decentralized Partially Observable Markov Decision Process allows agents to make decisions asynchronously, with each agent's goal being to maximize its expected discounted return. , and learn a measurement while satisfying all auxiliary constraints; the expression of the auxiliary constraints is: ; in, represents the auxiliary constraint function, Representing an Agent j At time step t Rewards received, Represents the time stept The discount factor when .
[0029] Specifically, on the one hand, the asynchronously constrained decentralized partially observable Markov decision process allows the agent to make decisions asynchronously. , A subset of agents need to make decisions, while the other agents do not need to make decisions due to environmental reasons. The set of observations Through the function get. Each agent in According to its observed value Select its action , where the action set is . Through the state transfer function A new environment state can be obtained . Agent Collection The reward provided by the environment after executing its strategy Can be obtained through the reward function .
[0030] On the other hand, each agent j The goal is to learn strategies , the strategy aims to maximize the expected discounted return , and satisfy all auxiliary constraints shown in the expressions of the above auxiliary constraints.
[0031] In one embodiment, in step 100, in an asynchronous constrained decentralized partially observable Markov decision process: j At its local time step The action when is defined as a communication strategy based on the gating mechanism, which enables the agent to adaptively select the timing of data transmission, thereby coordinating communication to improve algorithm performance; j The local time step The global shared reward at is defined as the sum of the changes in the average number of task conflicts of all agents; observations include: Bron-Kerbosch characteristics, message value characteristics, and channel access characteristics.
[0032] In one embodiment, step 104 includes: each agent independently and asynchronously collects its interaction data in an asynchronous constrained multi-agent reinforcement learning environment, and during training, the interaction data from different agents are connected into trajectories according to the time relationship; the interaction data includes global state, local observation, action and reward; wherein the action is a communication strategy based on a gating mechanism; based on the original MAPPO algorithm, in order to ensure that the gating mechanism strategy satisfies the constraint conditions, the Lagrangian multiplier is introduced, and the Lagrangian function is defined; the Lagrangian function expression is: ; in, represents the Lagrangian function, represents the Lagrange multiplier, Representing an Agent j Actor network parameters, represents the environmental reward at the current time step, represents the bandwidth penalty threshold, Representing an Agent In the The total amount of data that needs to be sent in each time slot is Representing an Agent j The action at the current time step, Represents the time step t The discount factor when .
[0033] The Lagrangian function is used to centrally optimize the communication strategy under the given bandwidth requirement constraint to achieve simultaneous optimization of multiple objectives and obtain the final multi-objective asynchronous communication strategy.
[0034] Specifically, in the centralized training-decentralized execution framework, training data needs to be collected from each agent, including interaction data between the agent and the environment, such as global state, local observations, actions, rewards, etc. These data are then connected in chronological order into trajectories for centralized training. In synchronous training scenarios, all agents make decisions and collect interaction data at the same time. However, in asynchronous training scenarios, some agents may find it difficult to make decisions within the same global time step, which poses a challenge to processing and training interaction data. In asynchronous distributed task allocation algorithms, communication strategies mainly affect subsequent computational processes and the number of task conflicts in the next round of the algorithm. In other words, the decisions made by an agent within a certain time step only affect the environmental rewards before the next time step. Therefore, the serial method, which is widely used in asynchronous multi-agent learning, is used to process asynchronous data. Under the serial method, each agent collects its interaction data independently and asynchronously, and during training, data from different agents are serialized into trajectories based on the temporal relationship. As Figure 2As shown, the data from Agent 1 and Agent 2 at different time steps are concatenated based on their temporal relationship, and their respective trajectories are as follows: ; .
[0035] On the one hand, since the communication strategy executed by the agent in the distributed task allocation algorithm has little effect on the next stage of calculation, this mechanism can correctly evaluate the impact of different agent strategies. On the other hand, the series method is more flexible and stable in engineering implementation, making it easier to expand the algorithm to real network environments. Figure 2 As shown in (a), the asynchronous training data processing based on the series is as follows Figure 2 (b) as shown.
[0036] By splicing different training trajectory data to supplement the global state data, stable strategy training can be achieved under the CDTE framework.
[0037] In one embodiment, a Lagrangian function is used to centrally optimize the communication strategy under a given bandwidth requirement constraint to achieve simultaneous optimization of multiple objectives and obtain a final multi-objective asynchronous communication strategy, including: defining the dual function of the original problem as:
[0038] To find satisfaction The optimal strategy parameter is the objective, in the Lagrange multiplier at the time step t When is a fixed value, each agent constructs a new reward to train its policy; the new reward is: ; in, Representing an Agent At time step New rewards for Representing an Agent At time step Rewards, Indicates that at time step t The Lagrange multiplier for , Representing an Agent In the The total amount of data that needs to be sent in each time slot is Representing an Agent j The action at the current time step.
[0039] The training gradient is calculated by minimizing the mean square error between the state value function and the new reward, and the Critic network parameters are updated according to the gradient; the update expression of the Critic network parameters is: ; in, is the Critic network parameter, is the state value function.
[0040] For the Actor network, define the policy advantage value, calculate the training gradient used to maximize the policy advantage value, and update the Actor network parameters according to the gradient; the update method of the Actor network parameters is: ; ; ; in, Represents the Actor network parameters, represents the advantage of a state-action pair, represents the advantage of the weighted state-action pair, represents the weighting factor, represents a hyperparameter.
[0041] In one embodiment, the weighting factor is expressed as: ; in, represents the KL divergence, and Respectively indicate update The agent’s policy distribution before and after.
[0042] Specifically, the agent j At its local time step The action is defined as . It is defined as a communication strategy based on a gating mechanism that enables agents to adaptively choose the timing of data transmission, thereby coordinating communication to improve algorithm performance. When the intelligent agent Its local observation Enter the Actor Network for Communication Strategies , which determines whether to send a message. j At time step t The global state is defined as , which is equivalent to the joint observation of all agents In the asynchronous learning scenario, assuming that at the global time step t When the intelligent agent j The allocation algorithm for is complete and a decision must be made, while the agent The allocation algorithm for has not yet completed, so no decision can be made. , Agent Observed value is assigned the value of the zero vector, and is assigned the current observation vector, which means that only the observations of the agent making the decision are considered valid. and the next global state Through the state transfer function Get. In the agent The local time step When . A cooperative relationship is introduced to coordinate the communication and define the reward function. This can be achieved by building a common goal between agents based on a gated communication strategy. The common goal of the communication strategy should minimize the global task conflict within a fixed period. In addition, the change of task conflict of a single agent is affected by the communication strategy of other agents and cannot fully reflect the effectiveness of communication. Therefore, the change of the average number of task conflicts of all agents is added up to form a global shared reward: ; in, Indicates the time step Internal, intelligent The change in the average number of task conflicts is recorded as , Indicates the time step At the beginning, the agent The average number of task conflicts with other agents, ,in, Representing an Agent The task list, Represents another agent k The more conflicts there are between the UAVs, The larger the value.
[0043] In one embodiment, the optimal policy parameters are obtained under the bandwidth requirement constraint. and Lagrange multipliers The process includes: finding the dual function The saddle point and the problem of optimal Lagrange multipliers are transformed into: ; in, represents the optimal Lagrange multiplier.
[0044] By taking the derivative of the Lagrangian function, t The Lagrange multiplier is updated at every moment; t The update expression of the Lagrange multiplier at the moment is: ; Among them, represents the Lagrange multiplier at time represents the Lagrange multiplier at time represents a hyperparameter.
[0045] Adopt a penalty value network to estimate the penalty value of the agent under different observations, and update the parameters of the penalty value network through temporal difference. The update expression of the penalty value network parameters is: ; Among them, represents the penalty value network parameters, represents the penalty value network.
[0046] Calculate the action penalty value of each agent at different time steps according to the penalty value network, and update the Lagrange multiplier according to the action penalty value; the update expression of the Lagrange multiplier is: ; Among them, represents the agent j at time step t of the action penalty value.
[0047] Specifically, based on the bandwidth penalty constraint, the communication strategy optimization problem in the distributed task allocation algorithm becomes solving an asynchronous and constrained Markov Decision Process. Its formulation is as follows: ; Among them, represents the bandwidth penalty threshold.
[0048] To solve this problem, the Lagrangian relaxation method is extended to the multi-agent proximal policy optimization (MAPPO) algorithm, and a multi-objective coupled PPO method (abbreviation: MOC-PPO) is proposed for training multi-objective asynchronous policies. The original MAPPO trains the shared Actor network and value network in a centralized training and distributed execution manner. Each agent j uses its local observation to select an action , and then it is used for the distributed execution of different policies. This method helps agents identify cooperative relationships and resolve potential conflicts between policies by sharing state and reward information among agents. On this basis, to ensure that the gated mechanism policy obtained by training meets the constraint conditions, the Lagrange multiplier , and define the Lagrangian function as shown in the above Lagrangian function expression.
[0049] The dual problem of the original problem can be defined as: , assuming At time step t When fixed, define In order to find the optimal strategy parameters , so that Therefore, when When fixed, each agent constructs a new reward To train its strategy. Therefore, the centralized Critic network Used to estimate the state value function By minimizing the state value function With rewards The mean square error between is used to calculate the training gradient, which is used to update the Critic network parameters , update the Critic network parameters using the above Critic network parameter update expression .
[0050] For an Actor network, update its parameters by defining the advantage of a state-action pair , which is calculated based on the action value of the agent at each step: .
[0051] To avoid a large difference in the policy distribution before and after the parameter update, the KL divergence is used as a weighting factor for the advantage value. . At the same time Perform truncation to set an upper limit on the advantage value: ; in, represents the KL divergence between the new and old strategy distributions, and Respectively indicate update The policy distribution of the agent before and after. We calculate the value used to maximize the policy advantage The training gradient of the Actor network is used to update the Actor network parameters. , the update method of Actor network parameters is shown in the above Actor network parameter update expression.
[0052] In order to achieve simultaneous optimization of multiple objectives, it is necessary to maximize Increase the data transmission bandwidth penalty Weight In order to find the optimal parameters, a reinforcement learning method combined with Lagrangian dual optimization is used. It is necessary to find the dual function Saddle points and optimal Lagrange multipliers , which is equivalent to solving: ; By taking the derivative of the Lagrangian function, we can use the above Update expression of Lagrange multiplier at time instant t Moment Lagrange multiplier.
[0053] Introducing a penalty value network To help solve the optimal Lagrange multiplier The penalty network is used to estimate the penalty value of the agent under different observations, which is recorded as: The penalty value network parameters are updated by defining time differences, and the specific update method is shown in the penalty value network parameter update expression above.
[0054] By calculating the action penalty value of each agent at different time steps , This can be updated using the Lagrange multiplier update expression above.
[0055] Based on this method, in the process of solving the original dual problem, the optimal strategy parameters can be obtained under the bandwidth requirement constraint. and Lagrange multipliers .
[0056] In one embodiment, the expression of the required bandwidth constraint based on time division multiple access is: ; in, represents the upper limit of the probability that the agent sends a message at each step, represents the entropy of the Gaussian distribution, represents the average number of bits sent by the agent, Indicates the maximum transmission rate of TDMA. n represents the average number of symbols sent per second, Indicates the number of symbols contained in each message, F represents the sampling frequency of the system, N represents the number of agents, m Indicates a message; represents the bandwidth penalty threshold, Representing an Agent In the t The total amount of data that needs to be sent in each time slot is Representing an Agent j The action at the current time step.
[0057] Specifically, in the case of the required bandwidth constraint based on time division multiple access, the specific implementation of MOC-PPO includes: First, the required bandwidth constraint based on time division multiple access (TDMA): In order to minimize the communication overhead, the threshold of the transmission probability is estimated based on the upper bound. Then, this constraint is relaxed using the Lagrangian method.
[0058] For the convenience of analysis, we assume that the distribution of message m follows Gaussian distribution and its entropy is H ( m ) is represented by TDMA. In a network protocol based on TDMA, the maximum number of time slots is T, the occupancy time of each time slot is d milliseconds, and the maximum amount of data in each time slot is Z bytes. Therefore, the maximum transmission rate of TDMA is Bits per second. According to Shannon's theorem, in order to ensure that no information is lost during the message transmission process, the average number of bits that the agent sends Must meet: .
[0059] , and H ( m ) is: ,in n is the average number of symbols sent per second.
[0060] Alternatively, the mean of the message distribution can be used μ and variance σ To calculate the entropy H(m) of the Gaussian distribution: ; Using the inequality relationship, we can deduce:
[0061] Assume that under the gating mechanism, the probability of each agent sending a message at each time step is p , and each message contains L symbols. The sampling frequency of the system is set to F In a system consisting of N agents connected by a graph, the average number of symbols n It can be calculated as: , here it is assumed that in the connected graph, each agent has an adjacency relationship with the other half of the agents in the communication channel. Next, we can derive the upper limit of the probability of each agent sending a message at each step : ; In the TDMA communication protocol, only one agent can send a message in each time slot, and the sampling frequency can be calculated as: Based on the previous calculation, the maximum transfer rate is given by the following formula: .
[0062] Combining these, we can deduce that the upper limit of the probability of the agent sending data in the TDMA protocol is: ; in, It represents the upper limit of the probability of the agent sending data in the TDMA protocol.
[0063] Based on the derived upper limit of the probability of sending data, the bandwidth occupancy penalty for the gating mechanism of the agent in a specific period can be described as: ; in, represents the bandwidth penalty threshold, and Representing an Agent j In the t A time slot is penalized when it occupies bandwidth. Representing an Agent j In the t The total amount of data that needs to be sent in each time slot.
[0064] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0065] In a verification example, a hardware-in-the-loop (HIL) simulation platform was built to verify the performance of the multi-objective coupled PPO method (MOC-PPO) to meet the challenge of deploying the strategy into actual communication scenarios. The HIL simulation platform mainly consists of two parts: a virtual training platform and an actual verification platform. The virtual training platform consists of a self-organizing network simulation server platform (SONSP) and a training host equipped with a GPU. The training host will use an asynchronous training method for closed-loop asynchronous training. SONSP contains three communication nodes: virtual nodes, bridge nodes, and actual nodes. Virtual nodes are network devices built into SONSP that run the same 5-layer communication protocol as actual nodes. Each thread accesses the virtual node on the server through any available port on the training host. Inter-node communication on the server is used to simulate self-organizing network communication between drones. The verification platform consists of a hybrid network communication terminal, a bridge communication terminal, and an onboard collaborative controller. The hybrid network communication terminal acts as an actual node in SONSP. As a bridge node in SONSP, the bridge communication terminal acts as an intermediary between actual nodes and virtual nodes. The physical layer protocol runs on the actual terminal device, while the other four network protocol layers run on the SONSP. The system's cluster network consists of N actual nodes and X virtual nodes, forming a hardware-in-the-loop (HIL) network simulation system to simulate cluster network communication. This N+X approach provides good scalability for communication nodes. The onboard collaborative controller is designed to execute the task scheduler when the drone is running. In order to verify the effectiveness of MOC-PPO, the strategy was trained in task allocation scenarios with 30 communication nodes and 40 communication nodes. MACPO was used as a comparison method and tested in multiple scenarios. MACPO is another advanced constrained multi-agent reinforcement learning method, in which the constraints are also designed as bandwidth requirement constraints. Under the same conditions, the ablation version of MOC-PPO was used for communication constraint optimization, which degenerated into the MAPPO algorithm under asynchronous training. The evaluation indicators include task conflict, required bandwidth, and average number of data transmissions. The experimental results are shown in Figure 3 As shown in Figure 2, it is compared with other methods. MOC-PPO trains an effective strategy that can balance communication overhead and task conflict rate. Figure 3 As shown in the communication scenarios of 30 nodes and 40 nodes, the average mission conflict rate of ACBBA optimized by MOC-PPO was reduced by 1% and 62.5% respectively compared with the original ACBBA, while the communication demand was reduced by 84% and 68% respectively, and the average communication frequency was reduced by 82% and 72% respectively. The results show that MOC-PPO significantly reduces the communication demand and frequency, while reducing the mission conflict rate. However, the performance of MACPO is not ideal. In the scenario of 30 drones, although the communication demand is significantly reduced, the mission conflict rate has increased to 15% ( Figure 3 (a) to Figure 3 (c)). In the 40-drone scenario, the mission conflict rate decreased, but the communication demand increased by 8% ( Figure 3 (d) to Figure 3 (f)). MACPO lacks a penalty network to approximate the TDMA bandwidth demand constraint, which makes it difficult to balance communication penalty and task conflict during training. Similarly, the ablation of bandwidth constraint optimization in MOC-PPO also has difficulty training an effective communication strategy, and the average task conflict rate increases to 15% and 13%, respectively. Such a high increase in task conflict rate will hinder the multi-UAV system from completing the specified task, highlighting the importance of bandwidth constraint optimization. Overall, MOC-PPO successfully balances communication demand and task conflict, achieving simultaneous optimization of multiple objectives under constraints.
[0066] As mentioned before, these observations have local characteristics and are not affected by the number of drones. By constructing a strategy that only depends on the local characteristics of neighboring drones and shares the same network parameters, a general communication strategy that is applicable to drones of different sizes can be trained. The optimized communication strategy was tested using the N + X mode in a HIL experimental scenario, and the communication strategy trained in the 30-drone scenario was applied to the 40 and 50-drone scenarios. The experimental results show that the communication strategy reduces the task conflict rate of ACBBA and significantly reduces the required communication bandwidth in all scenarios. In the scenarios of 30, 40, and 50 drones, the bandwidth requirements are reduced by 84%, 62.5%, and 44.4% compared to ACBBA, respectively. These results show that communication strategies trained in small-scale environments can also improve the performance of task allocation algorithms in large-scale or even real environments.
[0067] A hardware-in-the-loop (HIL) platform was built based on the Self-Organizing Network Simulation Platform Server (SONSP) for training and verification to address the difficulty of deploying policies trained in simulation environments to actual communication scenarios. In addition, the performance of MOC-PPO and MACPO under different drone scale configurations was carefully evaluated. Experimental results show that MOC-PPO can reduce bandwidth requirements by up to 84% and reduce mission conflict rates by up to 62.5%. In addition, communication policies trained under small-scale conditions can be extended to larger-scale environments. To address the gap between simulation and reality, the trained policies were verified using a hardware-in-the-loop (HIL) platform, demonstrating scalability to swarms of 30 to 50 drones. Extensive experimental comparisons with state-of-the-art methods (such as MACPO) show that MOC-PPO achieves a better balance between communication efficiency and allocation reliability.
[0068] In one embodiment, a communication perception task allocation device based on a multi-objective asynchronous strategy is provided, which is suitable for implementing communication perception task allocation in a drone cluster; each drone in the drone cluster deploys an agent, and the drone cluster constitutes a multi-agent system, and the device includes: The problem modeling module is used to transform the communication strategy optimization problem in the distributed task allocation algorithm of the multi-agent system into an asynchronous constrained decentralized partially observable Markov decision process according to the bandwidth constraint.
[0069] The asynchronous constrained multi-agent reinforcement learning environment building module is used to disperse the partially observable Markov decision process according to the asynchronous constraints and build an asynchronous constrained multi-agent reinforcement learning environment that takes into account the communication process.
[0070] The multi-objective asynchronous communication strategy optimization module is used to optimize the multi-objective asynchronous communication strategy in the distributed task allocation method according to the asynchronous constrained multi-agent reinforcement learning environment, and obtain the final multi-objective asynchronous communication strategy; the multi-objective coupling PPO method refers to a method of extending the Lagrangian relaxation method to the multi-agent proximal strategy optimization algorithm, asynchronously collecting data in a distributed manner, and then using the Lagrangian function to centrally optimize the communication strategy under the given bandwidth demand constraint, to achieve simultaneous optimization of multiple objectives, and obtain the final multi-objective asynchronous communication strategy.
[0071] Based on the communication perception task allocation module, each UAV adopts the final multi-target asynchronous communication strategy to complete the UAV cluster communication perception task allocation.
[0072] In one embodiment, the tuple of the asynchronous constrained decentralized partially observable Markov decision process in the problem modeling module is: ; in, A tuple representing an asynchronous constrained decentralized partially observable Markov decision process, express n A collection of intelligent agents, S Represents the global state, and Represents all agents at time step t The set of observations and actions, represents the state transition function, represents the new global state, R represents the reward function, γ ∈[0,1] represents the discount factor, Represents multiple auxiliary constraints, which are used to limit the update of the agent strategy. Indicates i Auxiliary constraints, .
[0073] Asynchronous Constrained Decentralized Partially Observable Markov Decision Process allows agents to make decisions asynchronously, with each agent's goal being to maximize its expected discounted return. , and learn a measurement while satisfying all auxiliary constraints; the auxiliary constraints are shown in the auxiliary constraint expressions above.
[0074] In one embodiment, in the problem modeling module, in an asynchronous constrained decentralized partially observable Markov decision process: j At its local time step The action when is defined as a communication strategy based on the gating mechanism, which enables the agent to adaptively select the timing of data transmission, thereby coordinating communication to improve algorithm performance; j The local time step The global shared reward at is defined as the sum of the changes in the average number of task conflicts of all agents; observations include: Bron-Kerbosch characteristics, message value characteristics, and channel access characteristics.
[0075] In one of the embodiments, the multi-objective asynchronous communication strategy optimization module is also used for each agent to independently and asynchronously collect its interaction data in an asynchronous constrained multi-agent reinforcement learning environment. During training, the interaction data from different agents are connected in series into trajectories according to the time relationship; the interaction data includes global state, local observation, action and reward; wherein the action is a communication strategy based on a gating mechanism; on the basis of the original MAPPO algorithm, in order to ensure that the gating mechanism strategy satisfies the constraints, the Lagrangian multiplier is introduced, and the Lagrangian function shown in the above Lagrangian function expression is defined; the Lagrangian function is used to centrally optimize the communication strategy under the given bandwidth demand constraint, so as to achieve the simultaneous optimization of multiple objectives and obtain the final multi-objective asynchronous communication strategy.
[0076] In one embodiment, a Lagrangian function is used to centrally optimize the communication strategy under a given bandwidth requirement constraint to achieve simultaneous optimization of multiple objectives and obtain a final multi-objective asynchronous communication strategy, including: defining the dual function of the original problem as: .
[0077] To find satisfaction The optimal strategy parameter is the objective, in the Lagrange multiplier at the time step tWhen is a fixed value, each agent constructs a new reward to train its strategy; the new reward is shown in the above new reward expression; the training gradient is calculated by minimizing the mean square error between the state value function and the new reward, and the Critic network parameters are updated according to the gradient; the Critic network parameters are updated as shown in the above update expression of the Critic network parameters; for the Actor network, the policy advantage value is defined, and the training gradient used to maximize the policy advantage value is calculated, and the Actor network parameters are updated according to the gradient; the Actor network parameters are updated as shown in the above update expression of the Actor network parameters.
[0078] In one of the embodiments, the weighting factors in the multi-objective asynchronous communication strategy optimization module are as shown in the above weighting factor expression.
[0079] In one embodiment, the multi-objective asynchronous communication strategy optimization module obtains the optimal strategy parameters under the bandwidth requirement constraint. and Lagrange multipliers The process includes: finding the dual function The saddle point and the problem of optimal Lagrange multipliers are transformed into: ; in, represents the optimal Lagrange multiplier.
[0080] By taking the derivative of the Lagrangian function, t The Lagrange multiplier is updated at each moment, t The update method of the Lagrange multiplier at each moment is as described above The update expression of the value is shown in Figure 2. Using the penalty value network Estimate the penalty value of the agent under different observations, update the penalty network parameters through time difference, and the penalty network parameter updating method is shown in the above penalty network parameter updating expression; calculate the action penalty value of each agent at different time steps according to the penalty network, and update the Lagrange multiplier according to the action penalty value; the Lagrange multiplier updating method is described in the above Lagrange multiplier updating expression.
[0081] In one of the embodiments, the required bandwidth constraint based on time division multiple access in the multi-objective asynchronous communication strategy optimization module is as shown in the required bandwidth constraint expression based on time division multiple access mentioned above.
[0082] For the specific limitations of the communication-aware task allocation device based on a multi-target asynchronous strategy, please refer to the limitations of the communication-aware task allocation method based on a multi-target asynchronous strategy in the above text, which will not be repeated here. Each module in the above-mentioned communication-aware task allocation device based on a multi-target asynchronous strategy can be implemented in whole or in part through software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0083] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0084] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A communication-aware task allocation method based on a multi-objective asynchronous strategy, characterized in that: The method is applicable to a drone cluster to realize communication perception task allocation, wherein each drone in the drone cluster deploys an agent, and the drone cluster constitutes a multi-agent system, and the method comprises: According to the bandwidth constraint, the communication strategy optimization problem in the distributed task allocation algorithm of the multi-agent system is transformed into an asynchronous constrained decentralized partially observable Markov decision process; According to the asynchronous constrained decentralized partially observable Markov decision process, an asynchronous constrained multi-agent reinforcement learning environment considering the communication process is constructed; According to the asynchronous constrained multi-agent reinforcement learning environment, a multi-objective coupling PPO method is used to optimize the multi-objective asynchronous communication strategy in the distributed task allocation method to obtain a final multi-objective asynchronous communication strategy; the multi-objective coupling PPO method refers to a method of extending the Lagrangian relaxation method to a multi-agent proximal strategy optimization algorithm, asynchronously collecting data in a distributed manner, and then using the Lagrangian function to centrally optimize the communication strategy under a given bandwidth demand constraint, thereby achieving simultaneous optimization of multiple objectives and obtaining a final multi-objective asynchronous communication strategy; Each UAV adopts the final multi-target asynchronous communication strategy to complete the UAV cluster communication perception task allocation.
2. The communication-aware task allocation method based on a multi-objective asynchronous strategy according to claim 1 is characterized in that: The tuple of the asynchronous constrained decentralized partially observable Markov decision process is: in, A tuple representing an asynchronous constrained decentralized partially observable Markov decision process, express n A collection of intelligent agents, S Represents the global state, and Represents all agents at time step t The set of observations and actions, represents the state transition function, represents the new global state, R represents the reward function, represents the discount factor, Represents multiple auxiliary constraints, which are used to limit the update of the agent strategy. Indicates Auxiliary constraints, ; The asynchronous constrained decentralized partially observable Markov decision process allows agents to make decisions asynchronously, with each agent's goal being to maximize the expected discounted return , and learn a measurement while satisfying all auxiliary constraints; the auxiliary constraints are: in, represents the auxiliary constraint function, Representing an Agent j At time step t Rewards received, Represents the time step t The discount factor when .
3. The communication-aware task allocation method based on a multi-objective asynchronous strategy according to claim 2 is characterized in that: In the asynchronous constrained decentralized partially observable Markov decision process: The agent j At its local time step The action at is defined as a communication strategy based on a gating mechanism, which enables the agent to adaptively select the timing of data transmission, thereby coordinating communication to improve algorithm performance; In the intelligent body j The local time step The global shared reward at is defined as the sum of the changes in the average number of task conflicts of all agents; Observations include: Bron-Kerbosch characteristics, message value characteristics, and channel access characteristics.
4. The communication-aware task allocation method based on a multi-objective asynchronous strategy according to claim 1 is characterized in that: According to the asynchronous constrained multi-agent reinforcement learning environment, a multi-objective coupling PPO method is used to optimize the multi-objective asynchronous communication strategy in the distributed task allocation method to obtain a final multi-objective asynchronous communication strategy; the multi-objective coupling PPO method refers to a method of extending the Lagrangian relaxation method to a multi-agent proximal strategy optimization algorithm, collecting data asynchronously in a distributed manner, and then using the Lagrangian function to centrally optimize the communication strategy under a given bandwidth demand constraint, achieving simultaneous optimization of multiple objectives, and obtaining a final multi-objective asynchronous communication strategy, including: In the asynchronous constrained multi-agent reinforcement learning environment, each agent independently and asynchronously collects its interaction data, and during training, the interaction data from different agents are concatenated into trajectories according to the time relationship; the interaction data includes global state, local observation, action and reward; wherein the action is a communication strategy based on a gating mechanism; On the basis of the original MAPPO algorithm, in order to ensure that the gating mechanism strategy meets the constraints, the Lagrange multiplier is introduced, and the Lagrange function is defined as: in, represents the Lagrangian function, represents the Lagrange multiplier, Representing an Agent Actor network parameters, represents the environmental reward at the current time step, represents the bandwidth penalty threshold, Representing an Agent In the The total amount of data that needs to be sent in each time slot is Representing an Agent The action at the current time step, Represents the time step t Discount factor when The Lagrangian function is used to centrally optimize the communication strategy under the given bandwidth requirement constraint to achieve simultaneous optimization of multiple objectives and obtain the final multi-objective asynchronous communication strategy.
5. The communication-aware task allocation method based on a multi-objective asynchronous strategy according to claim 4 is characterized in that: The Lagrangian function is used to centrally optimize the communication strategy under the given bandwidth demand constraint to achieve simultaneous optimization of multiple objectives and obtain the final multi-objective asynchronous communication strategy, including: The dual function of the original problem is defined as: To find satisfaction The optimal strategy parameter is the objective, in the Lagrange multiplier at the time step t When is a fixed value, each agent constructs a new reward to train its strategy; the new reward is: in, Representing an Agent At time step New rewards for Representing an Agent At time step Rewards, Indicates that at time step t The Lagrange multiplier for , Representing an Agent In the The total amount of data that needs to be sent in each time slot is Representing an Agent The action at the current time step; The training gradient is calculated by minimizing the mean square error between the state value function and the new reward, and the critic network parameters are updated according to the gradient; the critic network parameters are updated as follows: in, represents the Critic network parameters, represents the state value function; For the Actor network, a policy advantage value is defined, and the training gradient for maximizing the policy advantage value is calculated, and the Actor network parameters are updated according to the gradient; the updating method of the Actor network parameters is: in, Represents the Actor network parameters, represents the advantage of a state-action pair, represents the advantage of the weighted state-action pair, represents the weighting factor, represents a hyperparameter.
6. The communication-aware task allocation method based on a multi-objective asynchronous strategy according to claim 5 is characterized in that: The weighting factors are: in, represents the KL divergence, and Respectively indicate update The agent’s policy distribution before and after.
7. The communication-aware task allocation method based on multi-objective asynchronous strategy according to claim 5 is characterized in that: Obtaining optimal policy parameters under bandwidth demand constraints and Lagrange multipliers The process includes: We will find the dual function The saddle point and the problem of optimal Lagrange multipliers are transformed into: in, represents the optimal Lagrange multiplier; By taking the derivative of the Lagrangian function, t The Lagrange multiplier is updated at every moment; t The update method of the Lagrange multiplier at each moment is: in, express The time Lagrange multiplier, express t The time Lagrange multiplier, represents a hyperparameter; Penalty value network Estimate the penalty value of the agent under different observations, and update the penalty network parameters through time difference. The penalty network parameter update method is: in, represents the penalty value network parameter, represents the penalty value network; The action penalty value of each agent at different time steps is calculated according to the penalty value network, and the Lagrange multiplier is updated according to the action penalty value; the Lagrange multiplier is updated as follows: in, Representing an Agent At time step The action penalty value.
8. The communication-aware task allocation method based on a multi-objective asynchronous strategy according to claim 1 is characterized in that: The bandwidth constraints required based on time division multiple access include: in, represents the upper limit of the probability that the agent sends a message at each step, represents the entropy of the Gaussian distribution, represents the average number of bits sent by the agent, Indicates the maximum transmission rate of TDMA. Indicates the number of symbols contained in each message, F represents the sampling frequency of the system, N represents the number of agents, and m represents the message; represents the bandwidth penalty threshold, Representing an Agent In the The total amount of data that needs to be sent in each time slot is Representing an Agent j The action at the current time step.
9. A communication-aware task allocation device based on a multi-objective asynchronous strategy, characterized in that: The device is suitable for implementing communication perception task allocation in a drone cluster, wherein each drone in the drone cluster deploys an agent, and the drone cluster constitutes a multi-agent system, and the device includes: A problem modeling module is used to transform the communication strategy optimization problem in the distributed task allocation algorithm of the multi-agent system into an asynchronous constrained decentralized partially observable Markov decision process according to bandwidth constraints; An asynchronous constrained multi-agent reinforcement learning environment construction module is used to construct an asynchronous constrained multi-agent reinforcement learning environment taking into account the communication process according to the asynchronous constraint decentralized partially observable Markov decision process; A multi-objective asynchronous communication strategy optimization module is used to optimize the multi-objective asynchronous communication strategy in the distributed task allocation method using a multi-objective coupling PPO method according to the asynchronous constraint multi-agent reinforcement learning environment to obtain a final multi-objective asynchronous communication strategy; the multi-objective coupling PPO method refers to a method of extending the Lagrangian relaxation method to a multi-agent proximal strategy optimization algorithm, asynchronously collecting data in a distributed manner, and then using a Lagrangian function to centrally optimize the communication strategy under a given bandwidth demand constraint, thereby achieving simultaneous optimization of multiple objectives and obtaining a final multi-objective asynchronous communication strategy; Based on the communication perception task allocation module, each drone adopts the final multi-target asynchronous communication strategy to complete the drone cluster communication perception task allocation.
Citation Information
Patent Citations
Active power distribution network distributed optimization method, system and device and storage medium
CN116629461A
Virtual power plant distributed resource collaborative optimization scheduling method and system
CN117610813A
Cited By
Differentiated task-oriented sensing integrated multi-unmanned aerial vehicle sensing method and system
CN120353255A