Multi-unmanned aerial vehicle cooperative operation dynamic task allocation method based on posterior probability prediction
By constructing a two-layer allocation framework of UAV benefit function and Bayesian update mechanism, the problems of real-time performance and resource utilization efficiency in UAV collaborative task allocation in dynamic environments are solved, and fast and accurate task allocation and resource optimization are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEBEI UNIV OF TECH
- Filing Date
- 2026-03-12
- Publication Date
- 2026-05-19
AI Technical Summary
Existing UAV collaborative task allocation methods suffer from poor real-time performance, weak adaptability, and low resource utilization efficiency in dynamic and uncertain environments.
A multi-UAV collaborative task allocation method based on posterior probability prediction is adopted. By constructing a UAV benefit function and combining a greedy strategy and a Bayesian update mechanism, a two-layer allocation framework for initial task allocation and re-allocation is realized. Real-time situational information is used to dynamically update action decisions and optimize resource utilization.
It enables fast and accurate task allocation in highly dynamic task environments, improves resource utilization and task completion success rate, meets real-time requirements, and avoids redundant iterations and waste of load resources.
Smart Images

Figure CN122064129A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of UAV swarm collaborative operation technology, and in particular, it is a dynamic task allocation method for multi-UAV collaborative operation based on posterior probability prediction. Background Technology
[0002] Target dynamic programming is a key research direction in current UAV collaborative task allocation. Existing methods are mainly divided into two categories: centralized optimization algorithms and distributed collaborative algorithms. Centralized optimization algorithms, such as genetic algorithms and particle swarm optimization algorithms, seek the optimal or near-optimal task allocation scheme by globally searching the solution space. Although they can theoretically achieve good global performance, their computational complexity increases exponentially with the number of UAVs and targets, making it difficult to meet the real-time decision-making requirements of highly dynamic task environments. Moreover, after the allocation scheme is determined, there is a lack of effective online adjustment mechanisms to cope with sudden changes in the environmental situation, such as target movement, task failure, or the appearance of new targets. Distributed collaborative methods, such as auction algorithms and contract network protocols, achieve task allocation through local negotiation and bidding among drones. While these methods reduce computational burden, they typically rely on simple bidding rules and local information, making them prone to getting trapped in local optima. Furthermore, they lack effective modeling and processing capabilities for uncertainties in the task environment (such as the probability of successful action decisions and changes in target health values). Moreover, most existing methods treat these uncertainties as fixed parameters or make simple assumptions, failing to deeply integrate them into action decisions and dynamic environmental simulations. This results in poor decision-making performance and low resource utilization efficiency in practical applications.
[0003] Therefore, this invention proposes a dynamic task allocation method for multi-UAV collaborative operations based on posterior probability prediction. The method aims to efficiently allocate tasks to targets in a collaborative manner. While ensuring the real-time nature of decision-making, it makes full use of real-time feedback information from the task environment. By updating the posterior probability, it can extrapolate and evaluate the constantly changing environmental situation and dynamically adjust task allocation and action strategies accordingly, thereby achieving continuous optimization of the overall operational efficiency of the UAV swarm. Summary of the Invention
[0004] To address the shortcomings of existing technologies, the technical problem this invention aims to solve is to provide a multi-UAV collaborative task allocation method based on posterior probability prediction, which solves the problems of poor real-time performance, weak adaptability, and low resource utilization efficiency of existing methods when facing dynamic and uncertain task environments.
[0005] The present invention solves the aforementioned technical problem by adopting the following technical solution: A method for allocating collaborative tasks among multiple unmanned aerial vehicles (UAVs) based on posterior probability prediction, characterized by the following steps: Step 1: Construct the benefit function for the drone; (1) (2) (3) In the formula, The benefit function for the UAV in the initial mission allocation phase. , Let be the benefit functions for the free UAV and the UAV with assigned tasks during the task reassignment phase, respectively. , , , The first The drone carried out the first The distance benefit, angle benefit, priority benefit, and survival benefit of each objective task. , , ~ , , These are the weighting coefficients; Step 2: Initial task allocation and assignment; According to the benefit function The efficiency of drones is calculated, and a greedy strategy is used for initial task allocation. Drones with assigned tasks select action strategies based on prior probabilities and execute the first round of tasks. After each round of tasks is completed, the health value of the target task is updated by formula (8), the priority of the target task is updated by formula (9), and the load of the assigned task drone is updated by formula (10). (8) (9) (10) in, , The first , After the first iteration The health value of each target task. For the first The effectiveness of drones in achieving the target mission under various operational strategies. For the first After the first iteration Prioritization of each target task It is a proportionality constant. , The first , After the first iteration The payload capacity of the drones assigned to the mission. For the first Effective action rate of drones under various action strategies; If the health value of all target tasks is zero, the task ends; otherwise, proceed to step three. Step 3: Task reassignment and assignment; For a free unmanned aerial vehicle with a non-zero payload, according to the benefit function Calculate the benefits and use a greedy strategy to allocate target tasks, select action strategies based on prior probabilities and execute target tasks; For drones with assigned tasks whose workload is not zero after the previous round of operations, the benefit function is used. Calculate the benefits; based on the benefits of the assigned task UAV to the current target task, calculate the posterior probability of the assigned task UAV under each action strategy according to Equation (12), select the action strategy according to the posterior probability and execute the target task; (12) In the formula, For the assigned mission drones in the Action Strategies The posterior probability is given below. For the assigned mission drones in the Action Strategies The prior probability is as follows. For the assigned mission drones in the Action Strategies The conditional probability under the following conditions, For the assigned mission drones in the Action Strategies The prior probability is as follows. For the assigned mission drones in the Action Strategies The posterior probability is given below. Total number of action strategies; After the current round of tasks is completed, the health value and priority of the target tasks are updated according to equations (8) and (9), and the load of the assigned task drones is updated according to equation (10). If the health value of all target tasks is zero, the task ends. Otherwise, the posterior probability of the assigned task drones under each action strategy after the current round of iteration is used as the prior probability of the next round of iteration. The above process is repeated to redistribute and operate the next round of tasks, and this process is repeated until the health value of all target tasks is zero, and the entire task is completed.
[0006] Furthermore, the formula for calculating distance benefits is as follows: (4) In the formula, For the first The drone to the The straight-line flight distance of each target mission. For the first Free drones to the first The straight-line flight distance of each target mission. For free drones; The formula for calculating the angle benefit is: (5) In the formula, For the first The drone and the first The difference in heading angle for each target mission. Indicates taking the absolute value; The formula for calculating priority benefits is: (6) In the formula, For the current moment Prioritization of each target task For the current moment The health value of each target task. For the current moment Prioritization of each target task The total number of target tasks; The formula for calculating survival benefits is: (7) In the formula, For the previous moment Prioritization of each target task The success rate of consuming a single load to improve the performance of the target task.
[0007] Furthermore, the conditional probability calculation formula for the assigned task drone under different action strategies is as follows: (11) In the formula, , and For bandwidth, For Gaussian kernel function, , The first , The benefits of the drones with assigned tasks after rounds of iteration. , The first , After the first iteration The drone carried out the first Prioritization and effectiveness of each objective task. , The first , After the first iteration The drone carried out the first The survival benefits of each target task.
[0008] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention proposes a two-layer allocation framework combining initial task allocation and reassignment, achieving both speed and accuracy in task allocation and meeting the real-time requirements of highly dynamic task environments. In the initial allocation phase, the angular and distance benefits between the UAV and the target are considered, enabling rapid initial allocation. In the reassignment phase, the survival and priority benefits of the target are comprehensively considered to ensure that the allocation results better reflect the actual environmental conditions.
[0009] 2. An action decision-making mechanism based on posterior probability updates was established. This mechanism uses prior probability as the baseline strategy, combines real-time situational information of the mission environment, dynamically updates the posterior probability, and makes action decisions based on the posterior probability. This enables the UAV to dynamically adjust the selection probability of different action strategies according to the real-time environmental conditions, allowing the action strategy to be self-optimized and adjusted based on historical benefit values, preventing ineffective load consumption, and improving the utilization rate of UAV load resources.
[0010] 3. A closed-loop feedback architecture based on "initial allocation - task execution - situational inference - dynamic reallocation" is proposed, which enables the UAV swarm to learn from each action decision and continuously update its understanding of the mission environment and action decisions through probabilistic inference using real-time feedback information, thereby achieving adaptive optimization of action decisions and resource allocation, and maximizing the efficiency of resource utilization and the success rate of mission completion. Attached Figure Description
[0011] Figure 1 This is an overall flowchart of the present invention; Figure 2 A comparison chart of the average number of iterations for different methods; Figure 3 This is a comparison chart of the average remaining payload of drones using different methods. Detailed Implementation
[0012] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments, but this does not limit the scope of protection of this application.
[0013] like Figure 1 As shown, this invention provides a multi-UAV cooperative task allocation method based on posterior probability prediction, comprising the following steps: Step 1: Construct the benefit function for the drone; In the mission space The definition of a drone swarm The set of target tasks to be assigned in collaborative operations is , For the first A drone, The total number of drones, For the first One target task, The total number of target tasks; the current set of drone payloads is denoted as... The set of target task health values at the current moment is denoted as , For the current moment The payload capacity of the drone. For the current moment The health value of the first objective task, i.e., the task's life value; the... This action strategy is denoted as , , and They represent the first time in the second month. The payload consumed by the drone under this action strategy, the effectiveness value generated for the target mission (i.e. the amount of damage to the target mission), and the effective action rate.
[0014] The benefits of a drone performing a target mission include distance benefits, angle benefits, priority benefits, and survival benefits. Therefore, the benefit function of the drone at different stages can be expressed as: (1) (2) (3) In the formula, The benefit function for the UAV in the initial mission allocation phase. , Let be the benefit functions for the free UAV and the UAV with assigned tasks during the task reassignment phase, respectively. , , , The first The drone carried out the first The distance benefit, angle benefit, priority benefit, and survival benefit of each objective task. , , ~ , , These are the weighting coefficients; The formula for calculating distance benefits is: (4) In the formula, For the first The drone to the The straight-line flight distance of each target mission. For the first Free drones to the first The straight-line flight distance of each target mission. For free drones; The formula for calculating the angle benefit is: (5) In the formula, For the first The drone and the first The difference in heading angle for each target mission. Indicates taking the absolute value; The formula for calculating priority benefits is: (6) In the formula, For the current moment Prioritization of each target task It is a proportionality constant. For the current moment Prioritizing each target task; The formula for calculating survival benefits is: (7) In the formula, For the previous moment Prioritization of each target task The success rate of consuming a single load to improve the performance of the target task.
[0015] Step 2: Initial task allocation and assignment; In the initial task allocation phase, the drone's benefit is calculated based on the drone's benefit function (Equation (1)) in the initial task allocation phase. A greedy strategy is adopted for initial task allocation, so that the drone and the target task can establish a fast, efficient and high-quality initial task pairing, laying a good foundation for subsequent dynamic optimization. The drone with the assigned task selects an action strategy based on the prior probability and executes the first round of tasks. Each task performed by the UAV is considered as one iteration. During the task execution, the UAV consumes its own load, and the target task also consumes health value. The task is only redistributed and operated when the health value of the target task is zero and the load of the UAV assigned to the target task is not zero; or when the health value of the target task is not zero and the load of the UAV assigned to the target task is zero, the target task participates in the task redistribution and operation. Therefore, after each round of tasks is completed, the health value of the target task is updated by formula (8) according to the feedback information, the priority of the target task is updated by formula (9), and the load of the UAV assigned to the task is updated by formula (10). Furthermore, the distance benefit, angle benefit, priority benefit and survival benefit after the current iteration are updated by formulas (4)-(7). (8) (9) (10) in, , The first , After the first iteration The health value of each target task. For the first After the first iteration Prioritization of each target task , The first , After the first iteration The payload capacity of the drones assigned to the mission; If the health value of all target tasks is zero, the task ends; otherwise, proceed to step three.
[0016] Step 3: Task reassignment and assignment; For free drones with non-zero payload (including drones not assigned tasks during the initial task allocation phase and drones assigned tasks corresponding to target tasks with zero health value after the first round of operations), according to the benefit function... Update its benefits and use a greedy strategy to allocate target tasks, select action strategies based on prior probabilities and execute target tasks; For drones with assigned tasks whose workload is not zero after the previous round of operations, the benefit function is used. The effectiveness of the drones under different action strategies is updated. The level of effectiveness reflects the efficiency of the drones under different action strategies. To quantify this difference and support subsequent decision-making, conditional probabilities corresponding to each action strategy are defined based on the effectiveness of the drones assigned to tasks, the priority effectiveness of the drones performing the target tasks, and the survival effectiveness. The conditional probabilities are then solved using a Gaussian kernel function. The Gaussian kernel function can construct complex nonlinear relationship models by measuring the similarity between data points, quantifying the similarity of changes in task environment parameters in real time, and providing a reliable likelihood estimation basis for posterior probability updates. Therefore, the conditional probabilities of the drones assigned to tasks under different action strategies are defined as follows: (11) In the formula, For the assigned mission drones in the Action Strategies The conditional probability under the following conditions; , and Bandwidth is a core smoothing parameter in kernel density estimation, which adjusts the range of influence of each data point on kernel density estimation. For Gaussian kernel function, , The first , The benefits of the drones with assigned tasks after rounds of iteration. , The first , After the first iteration The drone carried out the first Prioritization and effectiveness of each objective task. , The first , After the first iteration The drone carried out the first The survival benefits of each target task; To enable UAVs to have online learning and adaptive optimization capabilities in selecting action strategies for target tasks, for the action strategy selection in the current iteration, based on the benefits of the UAVs assigned to the current target task, the posterior probabilities of the UAVs assigned to the current task under each action strategy are dynamically updated using Bayesian inference rules. Based on the posterior probabilities, the action strategy is selected and the target task is executed. (12) In the formula, For the assigned mission drones in the Action Strategies The posterior probability is given below. For the assigned mission drones in the Action Strategies The prior probability is as follows. For the assigned mission drones in the Action Strategies The prior probability is as follows. For the assigned mission drones in the Action Strategies The posterior probability is given below. Total number of action strategies; After the current round of tasks is completed, the health value and priority of the target tasks are updated according to equations (8) and (9), and the load of the assigned task drones is updated according to equation (10). If the health value of all target tasks is zero, the task ends. Otherwise, the posterior probability of the assigned task drones under each action strategy after the current round of iteration is used as the prior probability of the next round of iteration. The above process is repeated to redistribute and operate the next round of tasks, and this process is repeated until the health value of all target tasks is zero, and the entire task is completed.
[0017] Example This embodiment takes a drone collaborative operation simulation scenario with a space size of 50m×50m×50m as an example. Six isomorphic drones and three target points are set up. The initial states of the drones and targets are shown in Tables 1 and 2. Each drone has three action strategies, as shown in Table 3.
[0018] Table 1 Initial State of the UAV
[0019] Table 2 Initial State of the Target
[0020] Table 3 Action Strategy Information
[0021] To verify the superiority of the method of this invention, it was compared with a genetic algorithm. Using the average number of iterations and the average remaining payload of the UAV as evaluation indicators, multiple random individual experiments were conducted. The experimental results are as follows: Figure 2 , 3 As shown in the figure. The results show that the method of the present invention only requires an average of 6 iterations to converge, while the genetic algorithm converges in an average of 14 iterations. Furthermore, the average remaining load of the method of the present invention is much higher than that of the genetic algorithm. Therefore, it has significant advantages in terms of convergence speed and resource utilization. This is due to the hierarchical optimization mechanism of the present invention, which combines a greedy strategy with Bayesian updates. The greedy strategy provides high-quality initial task allocation, greatly reducing the search space; while the Bayesian rules dynamically update the UAV's action strategy according to the real-time task environment, effectively avoiding redundant iterations and waste of load resources.
[0022] Any aspects not covered in this invention are applicable to existing technologies.
Claims
1. A method for allocating collaborative tasks among multiple unmanned aerial vehicles (UAVs) based on posterior probability prediction, characterized in that, Includes the following steps: Step 1: Construct the benefit function for the drone; (1) (2) (3) In the formula, The benefit function for the UAV in the initial mission allocation phase. , Let be the benefit functions for the free UAV and the UAV with assigned tasks during the task reassignment phase, respectively. , , , The first The drone carried out the first The distance benefit, angle benefit, priority benefit, and survival benefit of each objective task. , , ~ , , These are the weighting coefficients; Step 2: Initial task allocation and assignment; According to the benefit function Calculate the benefits of drones and use a greedy strategy for initial task allocation; The assigned drone selects its action strategy based on prior probabilities and executes its first mission. After each round of tasks is completed, the health value of the target task is updated by formula (8), the priority of the target task is updated by formula (9), and the load of the assigned task drone is updated by formula (10). (8) (9) (10) in, , The first , After the first iteration The health value of each target task. For the first The effectiveness of drones in achieving the target mission under various operational strategies. For the first After the first iteration Prioritization of each target task It is a proportionality constant. , The first , After the first iteration The payload capacity of the drones assigned to the mission. For the first Effective action rate of drones under various action strategies; If the health value of all target tasks is zero, the task ends; otherwise, proceed to step three. Step 3: Task reassignment and assignment; For a free unmanned aerial vehicle with a non-zero payload, according to the benefit function Calculate the benefits and use a greedy strategy to allocate target tasks, select action strategies based on prior probabilities and execute target tasks; For drones with assigned tasks whose workload is not zero after the previous round of operations, the benefit function is used. Calculate the benefits; based on the benefits of the assigned task UAV to the current target task, calculate the posterior probability of the assigned task UAV under each action strategy according to Equation (12), select the action strategy according to the posterior probability and execute the target task; (12) In the formula, For the assigned mission drones in the Action Strategies The posterior probability is given below. For the assigned mission drones in the Action Strategies The prior probability is as follows. For the assigned mission drones in the Action Strategies The conditional probability under the following conditions, For the assigned mission drones in the Action Strategies The prior probability is as follows. For the assigned mission drones in the Action Strategies The posterior probability is given below. Total number of action strategies; After the current round of tasks is completed, the health value and priority of the target tasks are updated according to equations (8) and (9), and the load of the assigned task drones is updated according to equation (10). If the health value of all target tasks is zero, the task ends. Otherwise, the posterior probability of the assigned task drones under each action strategy after the current round of iteration is used as the prior probability of the next round of iteration. The above process is repeated to redistribute and operate the next round of tasks, and this process is repeated until the health value of all target tasks is zero, and the entire task is completed.
2. The multi-UAV cooperative task allocation method based on posterior probability prediction according to claim 1, characterized in that, The formula for calculating distance benefits is: (4) In the formula, For the first The drone to the The straight-line flight distance of each target mission. For the first Free drones to the first The straight-line flight distance of each target mission. For free drones; The formula for calculating the angle benefit is: (5) In the formula, For the first The drone and the first The difference in heading angle for each target mission. Indicates taking the absolute value; The formula for calculating priority benefits is: (6) In the formula, For the current moment Prioritization of each target task For the current moment The health value of each target task. For the current moment Prioritization of each target task The total number of target tasks; The formula for calculating survival benefits is: (7) In the formula, For the previous moment Prioritization of each target task The success rate of consuming a single load to improve the performance of the target task.
3. The multi-UAV cooperative task allocation method based on posterior probability prediction according to claim 1 or 2, characterized in that, The formula for calculating the conditional probability of a task-assigned drone under different action strategies is as follows: (11) In the formula, , and For bandwidth, For Gaussian kernel function, , The first , The benefits of the drones with assigned tasks after rounds of iteration. , The first , After the first iteration The drone carried out the first Prioritization and effectiveness of each objective task. , The first , After the first iteration The drone carried out the first The survival benefits of each target task.