Heterogeneous unmanned aerial vehicle cluster intelligent task planning method

By combining the quantum annealing algorithm with reinforcement learning, the problem of task allocation and execution strategy optimization of heterogeneous drone clusters in dynamic environments was solved, efficient and accurate task matching and cluster collaboration were achieved, and the flexibility and efficiency of task execution were improved.

CN120746201APending Publication Date: 2025-10-03NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Patent Information

Application Number
CN202511164907.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In the task execution scenario of heterogeneous drone clusters, due to the significant differences in performance parameters such as endurance, payload capacity, flight speed, communication range, etc. among each drone, and the diverse types of tasks, dynamically changing priorities, specific resource requirements and time window constraints, traditional task planning algorithms are difficult to efficiently handle large-scale task allocation, quickly converge to the global optimal solution, and lack flexible real-time response and strategy optimization capabilities in dynamic environments.

Method used

A method combining quantum annealing algorithm and reinforcement learning is adopted. The quantum annealing algorithm is used to perform preliminary task allocation, and reinforcement learning is used for real-time adjustment. Combined with distributed communication and coordination mechanism, efficient task planning of drone clusters is achieved.

Benefits of technology

It achieves efficient processing of large-scale task allocation, has strong dynamic adaptability, improves the matching accuracy between tasks and drones, and enhances the cluster collaboration efficiency, solving the problems of insufficient computational complexity and flexibility of traditional algorithms in large-scale heterogeneous cluster scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120746201A_ABST
    Figure CN120746201A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of heterogeneous unmanned aerial vehicle cluster task planning and intelligent scheduling, and discloses a heterogeneous unmanned aerial vehicle cluster intelligent task planning method, which is characterized by comprising task and unmanned aerial vehicle information initialization, task preliminary allocation based on a quantum annealing algorithm, real-time adjustment based on reinforcement learning, and a communication and cooperation mechanism. Executing and monitoring a task; the super-strong parallel computing capability and the quantum tunneling effect of a quantum annealing algorithm are utilized to quickly process a large-scale task allocation problem and find an approximate range of a globally optimal solution, and then an unmanned aerial vehicle cluster continuously learns and adjusts a task execution strategy according to real-time environment feedback in a task execution process by means of reinforcement learning, so that the task execution efficiency is improved. Therefore, the method adapts to dynamically changing scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of heterogeneous UAV swarm task planning and intelligent scheduling, and more specifically, to a heterogeneous UAV swarm intelligent task planning method. Background Art

[0002] In heterogeneous drone swarm mission execution scenarios, due to significant differences in performance parameters such as endurance, payload capacity, flight speed, and communication range among drones, and the often diverse types of missions with dynamically changing priorities, specific resource requirements, and time window constraints, mission planning faces multiple challenges, including efficient large-scale task allocation, rapid global optimal solution resolution, and real-time adjustment in dynamic environments. Traditional mission planning algorithms either struggle to efficiently handle the complex constraints of large-scale heterogeneous swarms and quickly converge to a global optimal solution; or they lack the flexible real-time response and strategy optimization capabilities to cope with sudden environmental changes, equipment failures, or dynamic adjustments to mission objectives. Consequently, they are unable to meet the requirements for efficient collaborative mission execution in heterogeneous drone swarms in complex and dynamic scenarios.

[0003] Therefore, the present invention provides a heterogeneous UAV cluster intelligent task planning method to improve the above technical problems. Summary of the Invention

[0004] The disclosed embodiments aim to address the shortcomings of the existing technology and provide a method for intelligent task planning of heterogeneous drone clusters. By integrating the quantum annealing algorithm with reinforcement learning, the present invention solves the problem of task allocation and execution strategy optimization in large-scale, dynamic scenarios for clusters composed of drones with different performance parameters.

[0005] The above technical objectives of the present invention are achieved through the following technical solutions: A method for intelligent task planning of heterogeneous UAV swarms, comprising the following steps:

[0006] S1. Mission and UAV information initialization: Quantify the performance parameters of each UAV in the heterogeneous UAV cluster, such as endurance, payload capacity, flight speed, and communication range. At the same time, determine the type, location, priority, resource requirements and other related information of the mission to form a structured initialization data matrix.

[0007] S2. Preliminary Task Allocation Based on Quantum Annealing Algorithm: A spatiotemporal graph structure consisting of drone nodes, task nodes, and the associated edges between them is constructed. The task allocation problem is transformed into a single-person traveling salesman problem. By constructing the target Hamiltonian and simulating the quantum annealing evolution, a preliminary allocation plan for drone swarm tasks is obtained.

[0008] S3, Real-time Adjustment Based on Reinforcement Learning: Define a state vector containing the drone's state and environment information, determine the actions the drone can perform, design a multi-dimensional reward function, and optimize the task execution strategy through learning and decision-making processes to adapt to dynamic environmental changes;

[0009] S4, Communication and Collaboration Mechanism: UAVs share status information, mission execution status, and environmental data through wireless communication networks, and use distributed consistency algorithms to ensure cluster decision-making consensus;

[0010] S5. Mission execution and monitoring: The drone executes the mission according to the planned strategy, monitors its own status, mission progress and environmental interference in real time, and triggers a re-planning mechanism when the mission cannot be executed as planned.

[0011] As a preferred technical solution of the present invention, in S1, the endurance of the UAV is described by the remaining flight time, which comprehensively considers the influence of factors such as the current remaining battery power of the UAV, the average power consumption under standard flight conditions, and the flight altitude and payload weight on the endurance;

[0012] The quantitative formula for the drone's endurance is:

[0013]

[0014] in, Indicates the The remaining flight time of the drone, The current remaining battery power of the drone. is its average power consumption under standard flight conditions, The impact of factors such as flight altitude and payload weight on flight endurance is comprehensively considered, and the value range is 0.6-1.2.

[0015] As a preferred technical solution of the present invention, in S1, the payload capacity of the UAV is defined by the maximum payload weight and the compatibility of the types of equipment that can be carried. The compatibility takes into account the importance of each type of equipment required for the mission and the compatibility of the UAV with these equipment.

[0016] The quantitative formula for the payload capacity of a UAV is:

[0017]

[0018] in, For the The comprehensive payload capacity score of the drone, is the maximum weight it can physically bear, is the total number of equipment types required for the task, For the The importance coefficient of the equipment type, Indicates the The first drone and the Compatibility coefficient of this type of equipment.

[0019] As a preferred technical solution of the present invention, in S2, the target Hamiltonian is composed of a target term reflecting the minimization of task completion time and a constraint term reflecting the task execution order constraint, and the weight of the constraint term can be dynamically adjusted according to the strictness of the constraint;

[0020] The formula for constructing the target Hamiltonian is:

[0021]

[0022] in, is the total target Hamiltonian, To reflect the goal of minimizing task completion time, To reflect the constraints of the task execution order, is the constraint strength coefficient.

[0023] As a preferred technical solution of the present invention, in S2, in the process of converting the task allocation problem into a problem similar to the single-person traveling salesman problem, it is necessary to establish a task and UAV compatibility evaluation model. This model comprehensively considers the matching degree between the UAV payload capacity and the task equipment requirements, the matching degree between the UAV endurance and the time required to complete the task, and the matching between the task priority and the time window.

[0024] The evaluation formula for the compatibility between the mission and the drone is:

[0025]

[0026] in, Indicates the The first drone and the The fitness score of each task, 、 、 is the weight coefficient, For the The minimum required value of the equipment load for each task, For the The current position of the drone to The straight-line distance between the task locations, is the time window matching coefficient.

[0027] As a preferred technical solution of the present invention, in S3, the reward function includes a basic completion reward linked to task priority and completion quality, a time-efficiency reward to encourage on-time or early completion of tasks, a resource-saving reward that reflects the efficient use of resources such as power, a collaborative reward that encourages drones and clusters to maintain strategic consistency, and a penalty item to suppress failures, task failures, or dangerous behaviors.

[0028] The reward function is calculated as:

[0029]

[0030] in, For the The total reward obtained by the drone in the current decision step, As a basic completion reward, For time-limited rewards, Rewards for resource conservation, To reward collaboration, For penalty items.

[0031] As a preferred technical solution of the present invention, in S3, the feasibility of the UAV switching mission target needs to consider factors such as the adaptability of the UAV to the new mission, the distance from the UAV to the new mission, the current flight speed and the remaining flight time.

[0032] As a preferred technical solution of the present invention, in S3, when the UAV adjusts its flight speed, it is necessary to strike a balance between energy consumption and timeliness, and the adjusted speed does not exceed the maximum speed limit of the UAV and is not lower than the minimum safe speed;

[0033] The formula for adjusting the flight speed of the drone is:

[0034]

[0035] in, is the adjusted flight speed, is the speed adjustment factor, is the task priority difference, when Speeding up is allowed to complete high-priority tasks first, while ensuring that the adjusted speed does not exceed the maximum speed limit of the drone. and not less than the minimum safe speed .

[0036] As a preferred technical solution of the present invention, in S4, the communication between drones adopts time division multiple access, and each drone is assigned an independent communication time slot. The length of the time slot is related to the amount of status information data to be transmitted by the drone. Drones with large amounts of information obtain longer transmission time.

[0037] As a preferred technical solution of the present invention, in S5, the triggering conditions of the re-planning mechanism include the remaining flight time of the drone being less than a certain proportion of the estimated time of the current mission, or the remaining power being less than a certain multiple of the estimated power required for the mission;

[0038] The triggering conditions for the replanning mechanism are:

[0039] or

[0040] in, The remaining flight time of the drone. is the estimated execution time of the current task, The remaining battery power of the drone. Estimate the amount of power required for the task.

[0041] In summary, the present invention has the following beneficial effects:

[0042] First, it efficiently handles large-scale task allocation. Leveraging the quantum annealing algorithm's powerful parallel computing capabilities and quantum tunneling effects, it can rapidly discover the relationships and constraints between tasks and drones, quickly converging to the approximate range of the global optimal solution. This addresses the computational complexity and tendency of traditional algorithms to fall into local optimality in large-scale heterogeneous cluster scenarios.

[0043] Second, it has strong dynamic adaptability. Through reinforcement learning, drones can continuously optimize their execution strategies based on real-time environmental feedback (such as sudden threats, equipment failures, and changes in mission objectives), adjusting flight paths, mission objectives, or speed in real time. This overcomes the lack of flexibility of traditional static planning in dynamic scenarios.

[0044] Third, it improves the matching accuracy between tasks and drones. By precisely quantifying drone performance parameters such as endurance, payload, speed, and communication range, as well as mission priority, resource requirements, and time windows, combined with an adaptability assessment model, it achieves precise matching between heterogeneous drones and tasks, avoiding resource waste or mission failures caused by performance mismatches.

[0045] Fourth, it enhances cluster collaboration. Through layered communication protocols, distributed consensus algorithms, and multi-hop routing mechanisms, it ensures efficient sharing of status, mission, and environmental data between drones, enabling rapid decision-making consensus. This addresses the issues of delayed information exchange and poor collaboration in traditional clusters. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 A flowchart of a method for intelligent task planning of a heterogeneous UAV swarm provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0047] The present application is described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but are not intended to limit the present application in any form. It should be noted that those skilled in the art may make several variations and improvements without departing from the scope of the present application. These all fall within the scope of protection of the present application.

[0048] The disclosed embodiments aim to address the problem of task allocation and execution strategy optimization in large-scale, dynamic scenarios for swarms of drones composed of drones with varying performance parameters (such as endurance, payload capacity, flight speed, and communication range). To address this, the disclosed embodiments propose an intelligent task planning method for heterogeneous drone swarms. This method leverages the ultra-parallel computing power and quantum tunneling effects of the quantum annealing algorithm to rapidly process large-scale task allocation and identify the approximate range of the global optimal solution. Furthermore, reinforcement learning enables the drone swarm to continuously learn and adjust its task execution strategy based on real-time environmental feedback during task execution, adapting to dynamically changing scenarios.

[0049] Please refer to Figure 1 , Figure 1 The following is a flow chart of the intelligent task planning method for heterogeneous drone swarms according to an embodiment of the present disclosure. The overall process mainly includes the following five steps:

[0050] Step 1: Initialize the mission and drone information.

[0051] Clarify the performance parameters of each drone in a heterogeneous drone cluster, including endurance, payload capacity, flight speed, and communication range, and determine mission-related information, including mission type, location, priority, and resource requirements.

[0052] During the mission and drone information initialization phase, the performance parameters of each drone in the heterogeneous drone cluster must first be quantified. For endurance, the remaining flight time is used to describe it, and the calculation formula is:

[0053]

[0054] in, Indicates the The remaining flight time of the drone, The current remaining power of the drone (unit: Wh), is its average power consumption in standard flight conditions (unit: W), The impact of factors such as flight altitude and payload weight on flight endurance is comprehensively considered, with a value range of 0.6-1.2. When the drone has a large payload or flies at a high altitude, the value of this factor is relatively small.

[0055] The load capacity is defined by the maximum payload weight and the compatibility of the type of equipment that can be carried. The formula is:

[0056]

[0057] Where, For the The comprehensive payload capacity score of the drone, is the maximum weight it can physically carry (unit: kg), is the total number of equipment types required for the task, For the Importance coefficient of type equipment ( ), Indicates the The first drone and the Compatibility coefficient of the device class (0 means incompatible, 1 means fully compatible).

[0058] In terms of flight speed, it is necessary to consider the speed characteristics under different flight modes. The characterization formula is:

[0059]

[0060] in, For the The actual flight speed of the drone in the current environment (unit: m / s), is its standard cruising speed under undisturbed conditions, is the speed adjustment coefficient of the drone (the value range is -0.3 to 0.3, negative values ​​indicate deceleration, positive values ​​indicate acceleration), is the environmental resistance correction factor (calculated based on wind speed and air density. The greater the wind speed, the higher the air density. Larger values ​​range from 0 to 0.5).

[0061] The quantitative formula for communication range is:

[0062]

[0063] here, It is The actual effective communication range of the UAV (unit: km), The maximum communication distance under ideal conditions is is the interference attenuation coefficient (value range is 0.02-0.1), is the distance between the UAV and the nearest electromagnetic interference source (unit: km), The communication enhancement factor brought by the antenna configuration (typically 1.0-2.5).

[0064] In determining the mission-related information, the mission type needs to be marked in a matching manner with the UAV capability, and the mission location is marked by the three-dimensional coordinates of latitude, longitude and altitude. To accurately locate, 、 Respectively The longitude and latitude of the mission target (unit: degrees), is its altitude (unit: m). Task priority uses a dynamic scoring mechanism, the formula is:

[0065]

[0066] Where, For the The overall priority rating of each task (range 1-10), 、 、 is the weight coefficient ( ), The urgency of the task's completion deadline (1-5 points, the higher the score, the more urgent it is). The strategic importance of the mission objective (1-5 points), Indicates the time decay coefficient of task information (0.5-1.0, the more easily outdated the task is, the closer the value is to 1.0).

[0067] Resource requirements are defined by the required power, device type, and time window, using the formula:

[0068]

[0069] in, To complete the The estimated power required for each task (unit: Wh), is a collection of equipment types necessary to perform a task. The time interval allowed for the task to execute (unit: seconds, starting from the system startup moment). After accurately quantifying the drone's performance parameters and mission information through the above formula, a structured initialization data matrix is ​​formed, providing basic data support for subsequent task allocation and strategy adjustment.

[0070] Step 2: Preliminary task allocation based on quantum annealing algorithm.

[0071] After initializing mission and drone information, to more efficiently address the complex problem of task allocation for heterogeneous drone swarms, it's necessary to leverage appropriate algorithmic models to deeply explore and characterize the relationships and potential constraints between tasks and drones. Physics-inspired spatiotemporal graph neural networks can fully integrate the spatiotemporal characteristics of the physical world with the practical needs of drone swarm mission planning. By constructing a spatiotemporal graph structure consisting of drone nodes, mission nodes, and the associated edges between them, they encode drone performance parameters, mission attribute information, and their dynamic changes over time and space into the feature vectors of the graph nodes and edges. This lays the foundation for subsequently transforming the task allocation problem into a traveling salesman-like problem and efficiently solving it, achieving a smooth transition from information initialization to problem transformation.

[0072] (1) Problem transformation:

[0073] The task allocation problem for a swarm of heterogeneous drones is transformed into a problem similar to the single-person traveling salesman problem. By adding virtual locations and other methods, each drone is ensured to have a task to perform. The combination of tasks and drones is then transformed into a form that can be represented by a quantum bit system.

[0074] In the problem transformation process, we first need to establish a task and UAV compatibility evaluation model to quantify the matching relationship between the two. The formula is:

[0075]

[0076] in, Indicates the The first drone and the The suitability score of each task (range 0-10), 、 、 is the weight coefficient ( ), For the The minimum required value of equipment load for each task (unit: kg), For the The current position of the drone to The straight-line distance between the task locations (unit: m), is the time window matching coefficient (when the UAV can If the task location is reached within the interval, it is 1; otherwise, it is 0.3).

[0077] To transform the problem into a traveling salesman problem, a path network consisting of actual tasks and virtual tasks needs to be constructed. The setting of virtual tasks must meet the load balancing constraint. The formula for calculating the number of virtual tasks is:

[0078]

[0079] in, is the total number of virtual tasks, is the total number of drones in the cluster, is the actual number of tasks to be assigned, For the The total resource consumption of each actual task (comprehensive power consumption, time, etc.), is the average resource carrying capacity of the swarm drones (consisting of all drones and weighted calculation), It is a rounding function to ensure that the mission resource requirements do not exceed the UAV carrying capacity.

[0080] The location coordinates of the virtual task are set as:

[0081]

[0082] in, For the The initial position coordinates of the UAVs are determined to ensure that the virtual mission is in the center of the cluster.

[0083] To achieve the characterization of the quantum bit system, it is necessary to define the encoding rules of the quantum state and map the mission-drone combination into the superposition state of the quantum bit. The formula is:

[0084]

[0085] in, is the quantum state of the system, is the probability amplitude (satisfying ), is the quantum ground state of the drone, is the quantum ground state of the task (including virtual tasks).

[0086] At the same time, the constraint operator is introduced:

[0087]

[0088] in, ( is the unit operator), which is used to ensure that each drone performs at least one task and each actual task is only performed by one drone. This encoding method transforms the task allocation problem into a ground state solution problem of the quantum system.

[0089] (2) Construct the target Hamiltonian:

[0090] According to the goal of task assignment (minimizing task completion time) and constraints (task execution order constraints), the corresponding target Hamiltonian is constructed.

[0091] When constructing the target Hamiltonian, it is necessary to comprehensively consider the optimization objectives and constraints of the task allocation and convert them into the energy function of the quantum system. The target Hamiltonian consists of two parts: the target term and the constraint term, and the formula is:

[0092]

[0093] in, is the total target Hamiltonian, To reflect the goal of minimizing task completion time, To reflect the constraints of the task execution order, is the constraint strength coefficient (the value range is 10-100, and it is dynamically adjusted according to the strictness of the constraint. The stricter the constraint, the larger the value).

[0094] Target Item The construction of is centered on minimizing the total time to complete the task, and the calculation formula is:

[0095]

[0096] Where, For the The total number of tasks assigned to each drone (including virtual tasks), For the The drone from Mission location to The distance between the task locations (when hour, is the distance from the initial position of the UAV to the first mission position), For the The drone carried out The estimated time of each task (the execution time of the virtual task is 0), is a binary variable ( Indicates drone Execute tasks, means not executed).

[0097] Constraints Used to ensure the order of task execution, for task pairs with pre-dependencies (such as tasks Must be on task Executed before), its expression is:

[0098]

[0099] in, is a set of task pairs with order constraints, Indicates a task For the task The pre-tasks is a sequential indicator variable (when the task The completion time is less than the task At the start time of ,otherwise ), is the constraint weight (set based on the importance of the task dependency, ranging from 1 to 5, with larger values ​​for critical dependencies). This constraint ensures that violations of the order constraint will cause the system energy to increase, thus being suppressed during quantum annealing.

[0100] In order to adapt the representation of the quantum bit system, the above variables need to be mapped to quantum operators. , expressed using the Pauli operator as ,in, For the The drone corresponds to Pauli operator (eigenvalues ​​are After replacing the variables, the target Hamiltonian is converted into a combination of quantum operators, which can be directly used in the evolution process of the quantum annealing algorithm to achieve the optimization goal and constraint satisfaction of task allocation by minimizing the system energy.

[0101] (3) Simulated quantum annealing evolution:

[0102] Using the quantum annealing algorithm, the system is allowed to evolve adiabatically to solve the system quantum state corresponding to the optimal solution, thereby obtaining a preliminary allocation plan for drone cluster tasks and determining the tasks that each drone is roughly responsible for.

[0103] The simulated quantum annealing evolution process starts with the initial quantum state of the system, and achieves a smooth transition of the energy function by adiabatically adjusting the external field parameters, and finally converges to the ground state of the target Hamiltonian. First, the initial temperature and evolution time parameters need to be set. The initial temperature The calculation formula is:

[0104]

[0105] in, is the maximum possible energy difference of the system (determined by the extreme value range of the task completion time and the constraint term in the target Hamiltonian), is the Boltzmann constant, is the total number of possible quantum states of the system (equal to the product of the number of drones and the total number of missions). The initial temperature must ensure that the system can traverse enough quantum states to avoid falling into a local optimum.

[0106] During the evolution process, the change of the system Hamiltonian with time satisfies the adiabatic theorem, which is expressed as:

[0107]

[0108] Where, is the evolution time ( ), is the adiabatic evolution scheduling function (using form, is a nonlinear adjustment coefficient, with a value of 1.5-2.0 to enhance the efficiency of late evolution). is the initial Hamiltonian (using uniformly distributed random quantum states to ensure the maximum initial entropy of the system), is the target Hamiltonian (i.e., the energy function constructed to include the task allocation objectives and constraints).

[0109] In each step of evolution, the adiabatic update probability of the quantum state needs to be calculated, and the formula is:

[0110]

[0111] in, is the time step (usually ), and They are and The expected value of the system energy at time t, this probability ensures that the quantum state transitions in the direction with smaller energy gradient, which meets the requirements of adiabatic evolution. is the differential of the system Hamiltonian, is the time differential. It is the quantum state of the system at time t, which is used to describe the state of the quantum system at that moment.

[0112] When the evolution time reaches the total duration (Depend on Sure, is the reduced Planck constant, is the minimum energy level difference of the target Hamiltonian), the system quantum state Converges to the ground state. At this time, by measuring the probability distribution of the quantum state, the ground state with the largest probability amplitude is extracted. , that is, get the The drone is responsible for The resulting task allocation is generated. For virtual tasks, if the proportion of virtual tasks assigned to a particular drone exceeds 30%, load balancing fine-tuning is triggered, and some actual tasks are reallocated to that drone to ensure efficient resource utilization. Through this evolutionary process, the final task allocation solution meets both global optimality requirements and the matching conditions between drone performance and task constraints.

[0113] Step 3: Real-time adjustment based on reinforcement learning.

[0114] The initial task allocation scheme derived from the quantum annealing algorithm provides a foundational framework for swarm mission execution. However, in real-world scenarios, drones may face sudden environmental changes, equipment failures, or dynamic adjustments to mission objectives, requiring dynamic response capabilities. A real-time adjustment mechanism based on reinforcement learning leverages the environmental feedback data continuously generated by the swarm during mission execution. By optimizing decision-making strategies through online learning, it overcomes the limitations of the quantum annealing algorithm in handling dynamic problems, achieving a precise transition from global optimal initial allocation to local dynamic adaptation, ensuring the swarm maintains efficient mission execution in complex and dynamic environments.

[0115] (1) Status definition:

[0116] Determine the state variables in reinforcement learning, including the current position of the drone, remaining battery power, progress of completed tasks, as well as target state changes and emerging threat areas in the mission environment, and combine this information into a state vector to represent the current system state.

[0117] The state definition requires integrating the drone's own state with the dynamic information of the environment to form a high-dimensional state vector that can be directly processed by the reinforcement learning model. First, the current position of the drone is quantified, and a relative coordinate system is used to reduce the computational complexity of the absolute position. The formula is:

[0118]

[0119] in, For the UAV relative to its current mission target The three-dimensional coordinate vector of (unit: km), is the real-time latitude, longitude and altitude of the drone, The coordinates of the current mission target. The relative position can more intuitively reflect the spatial relationship between the UAV and the mission point.

[0120] The remaining power is normalized to eliminate the differences in battery capacity between different drones. The calculation formula is:

[0121]

[0122] Where, is the normalized remaining power (range 0-1), is the rated total capacity of the drone battery (unit: Wh), The remaining power of the drone battery. is the energy consumption coefficient based on the current load and flight altitude (with the initialization phase to ensure that the power level representation can reflect the actual battery life.

[0123] The progress of executed tasks is integrated into the completion status of multiple tasks by weighted summation. The formula is:

[0124]

[0125] in, For the The mission progress score of each drone (range 0-1), The total number of missions assigned to the UAV, For the The priority rating of each task, The time consumed in executing the task (unit: s). The estimated execution time at the time of initial assignment is weighted by priority to make the progress more closely aligned with the importance of the task.

[0126] The target state change in the environment is marked by a binary vector, and the formula is:

[0127] ,in

[0128] in, The length is equal to the actual number of tasks The vector can quickly identify the dynamic changes of task objectives through 0-1 marking, providing trigger signals for strategy adjustment.

[0129] The emerging threat area is quantified by the safety distance, the formula is:

[0130]

[0131] in, For the The safety factor vector of each drone and each threat area (each element ranges from 0 to 1), For drones and The shortest distance to the threat area (unit: km), The impact radius of the threat area (unit: km). When the safety factor is less than 0.3, it is judged as high risk and needs to be avoided first.

[0132] The final state vector integrates all the above information and is expressed as:

[0133]

[0134] This vector contains features in five dimensions: spatial position, energy state, task progress, target stability, and threat risk. It not only retains the consistency of the quantitative parameters in the initialization phase, but also reflects the key changes in the dynamic environment in real time, providing comprehensive state input for the decision-making process of reinforcement learning.

[0135] (2) Action definition:

[0136] Define the actions the drone can take, including changing the flight path, switching mission objectives, and adjusting flight speed.

[0137] Action definition requires quantizing and encoding the executable operations of the drone to form a discrete or continuous action space that adapts to the state vector. Changing the flight path is achieved by combining the steering angle and the path offset, as shown in the formula:

[0138]

[0139] in, For the The steering angle of the drone (unit: degree, value range , positive value indicates clockwise direction, negative value indicates counterclockwise direction). is the path offset distance (unit: m, value range ), this action must meet the safety distance constraint between the new path and the threat area, that is, the distance from any point on the offset path to the threat area is not less than ( is the threat area influence radius).

[0140] Switching mission objectives requires considering the dynamic matching of mission priorities and UAV capabilities. The feasibility of the action is determined by the switching coefficient, which is:

[0141]

[0142] Where, For the UAV from the current mission Switch to new task feasibility coefficient (range 0-1), Score the suitability of the drone for the new task (using the suitability evaluation model from the problem transformation phase), is the minimum value (avoiding the denominator to be 0), is the distance from the UAV to the new mission, is the current flight speed, is the remaining flight time, when The switching action is determined to be feasible. is the i-th UAV and the current mission The fitness score is based on the fitness evaluation model of the problem transformation stage (i.e. The same calculation method is used).

[0143] Adjusting the flight speed requires a balance between energy consumption and timeliness. The speed correction formula is:

[0144]

[0145] in, is the adjusted flight speed (unit: m / s), is the speed adjustment factor (value range , negative values ​​correspond to deceleration, positive values ​​correspond to acceleration), The task priority difference (the difference between the current task priority and the next task priority in the queue, range ),when Speeding up is allowed to complete high-priority tasks first, while ensuring that the adjusted speed does not exceed the maximum speed limit of the drone. and not less than the minimum safe speed .

[0146] All actions are integrated into an action space using one-hot encoding or continuous vectors. Path and speed adjustments are continuous actions, while task switching is a discrete action (selected only from the set of feasible tasks). After an action is executed, it must be synchronized to the cluster via a communication mechanism to ensure that other drones update their environmental state based on the action changes, providing a consistent action feedback basis for collaborative decision-making.

[0147] (3) Reward function design:

[0148] Design a reasonable reward function to motivate the drone to take actions that are conducive to mission completion. Give positive rewards for completing the mission on time, negative rewards for encountering failures or mission failures, and provide different levels of rewards based on mission priority and completion quality.

[0149] The reward function design needs to integrate multi-dimensional indicators such as task completion efficiency, resource consumption, and environmental adaptability to form a dynamic feedback mechanism to guide the optimization of drone strategies. The total reward is composed of the basic completion reward. , Time Reward , Resource Conservation Rewards , Collaboration Rewards and penalties It consists of five parts, and the formula is:

[0150]

[0151] in, For the The total reward obtained by the drone in the current decision step. The specific calculation method of each reward and penalty is as follows:

[0152] The basic completion reward is directly linked to the task priority and completion quality. The formula is:

[0153]

[0154] Where, For the current task Priority score (range 1-10), The quality coefficient of the task completion (calculated based on the accuracy index required by the task, ranging from 0.6 to 1.0, and 1.0 when the task is fully met), It is a completion mark (1 when the task is completed, 0 when it is not completed) to ensure that the completion of high-priority tasks can obtain higher basic rewards.

[0155] Time-efficiency rewards are used to motivate drones to complete tasks on time or ahead of schedule. The calculation formula is:

[0156]

[0157] in, For the task The deadline, For drones Actual completion time, The estimated execution time of the task. When the actual completion time is earlier than the deadline, the reward increases linearly with the increase in the lead time; if the task is completed beyond the deadline, the time reward is 0.

[0158] Resource conservation rewards reflect the efficient use of resources such as electricity. The formula is:

[0159]

[0160] Where, For the task The estimated power demand, For drones The actual amount of power consumed to perform the task. is the energy consumption correction factor (consistent with the initialization stage), and the reward cap is set to 5 to avoid excessive resource conservation leading to a decrease in task completion quality.

[0161] The collaborative reward encourages the drone to maintain policy consistency with the cluster, which is expressed as:

[0162]

[0163] in, is the collaborative weight (value range is 0.5-1.0), For drones The set of neighboring machines within the communication range of and UAVs and neighboring machines The local decision-making reward promotes cluster behavior coordination by narrowing the reward difference with neighboring machines.

[0164] The penalty term is used to suppress failures, task failures, or dangerous behaviors, and is formulated as:

[0165]

[0166] Where, is the penalty weight (value range is 1.0-2.0), is the fault flag (1 when a fault occurs, otherwise 0), To mark the completion (take 0 if not completed), For drones With the The safety factor of the threat area (taken from the state vector ), The threat level coefficient (2.0 for high-risk threats and 1.0 for low-risk threats) ensures that malfunctions, mission failures, and entering dangerous areas are significantly penalized.

[0167] Through the dynamic balance of the above-mentioned multi-dimensional rewards and penalties, the reward function can not only motivate drones to efficiently complete high-priority tasks, but also guide them to rationally utilize resources, avoid risks and maintain cluster coordination, ultimately achieving the goal of global task optimization in a dynamic environment.

[0168] (4) Learning and decision-making:

[0169] Based on its current state, the drone selects actions to execute the mission using a reinforcement learning algorithm. During mission execution, it continuously updates its policy network based on new states and rewards, gradually learning the optimal mission execution strategy. When the environment changes (including new missions and drone malfunctions), it can quickly adjust its mission execution method.

[0170] The learning and decision-making process uses a deep reinforcement learning framework to achieve adaptive decision-making in a dynamic environment through the coordinated optimization of the policy network and the value network. The policy network uses a two-layer LSTM structure to extract state sequence features and output action probability distribution. The formula is:

[0171]

[0172] in, In state Next select action The probability distribution of are policy network parameters, 、 are the weights and biases of the first layer of LSTM, 、 The Softmax function ensures that the sum of the probabilities is 1, which facilitates action selection based on greedy strategies or random exploration.

[0173] The value network is used to evaluate the expected cumulative reward of the state-action pair. It uses a structure that combines a convolutional neural network with a fully connected layer. The calculation formula is:

[0174]

[0175] Where, Status Next action Q value, is the value network parameter, To extract the convolutional features of the state vector, 、 and 、 are the weights and biases of the two-layer fully connected network respectively. The ReLU function introduces nonlinearity to fit the complex reward function.

[0176] The parameter update adopts the dominant Actor-Critic algorithm, and the loss function of the policy network is:

[0177]

[0178] in, is the advantage function, is the state value function (calculated by the mean output of the value network), and the loss function is minimized by gradient descent to update the strategy towards the direction of high-advantage actions. The parameter update target of the value network is:

[0179]

[0180] Where, is the reward obtained at the current step, is the discount factor (ranging from 0.9 to 0.99, balancing immediate rewards and future rewards), The new state after executing the action is obtained by minimizing the mean square error so that the Q value approaches the actual cumulative reward.

[0181] In order to cope with dynamic changes in the environment, an experience replay mechanism is introduced to store Samples, using priority sampling strategy to improve learning efficiency, the sampling probability is:

[0182]

[0183] in, For the The priority of each sample ( is the TD error, is the minimum value), is the priority weight (ranging from 0.4 to 0.6), so that high error samples are used more frequently for training, and k is the sample number.

[0184] At the same time, set the target network to regularly synchronize the main network parameters, synchronization cycle satisfy:

[0185]

[0186] Ensure the stability of the learning process. When a new task is added or a drone fails (determined by the state synchronization information of the communication network), the policy network will trigger a rapid exploration mechanism to temporarily increase the probability of random action selection (from Increase to ), which lasts for 50 decision steps before returning to normal, allowing the system to quickly adapt to environmental changes and update its strategy. Through this mechanism, drone swarms can continuously optimize their decision-making strategies in dynamic scenarios, achieving efficient and robust mission execution.

[0187] Step 4: Communication and coordination mechanism.

[0188] Drones exchange information through wireless communication networks and share their status information, mission execution status, and environmental perception data.

[0189] The communication and coordination mechanism is based on a distributed network architecture, and achieves state synchronization and decision-making consensus of the drone cluster through a layered information interaction protocol. Information interaction adopts time division multiple access, and each drone is assigned an independent communication time slot. The time slot length calculation formula is:

[0190]

[0191] in, For the The communication time slot length of each drone (unit: ms), is the basic time slot length (fixed at 50ms), is the information complexity coefficient (range 0.5-1.0), is the amount of state information data to be transmitted by the UAV (unit: kB), To maximize the amount of information data in the cluster, ensure that drones with large amounts of information have longer transmission time.

[0192] The status information is encapsulated in a structured data frame format, which includes fields such as drone ID, timestamp, location coordinates, remaining battery power, mission progress, and environmental threats. Data integrity verification is achieved through a hash function, and the formula is:

[0193]

[0194] Where, is the hash value of the data frame, is the original status information, For the The receiver verifies whether the data has been tampered with by comparing the hash value. SHA is a secure hash algorithm used to perform hash operations on data to achieve data integrity verification.

[0195] In order to solve the information island problem caused by limited communication range, a multi-hop routing mechanism is adopted. The routing node selection is based on the comprehensive judgment of communication quality and residual energy. The routing weight formula is:

[0196]

[0197] in, For drones As a drone The weight of the routing node (range 0-1), is the weight coefficient (value 0.6, giving priority to communication quality), For drones Receiving drones Signal strength indication (unit: dBm), is the maximum signal strength, For drones The node with the highest weight is selected as the next hop routing.

[0198] The sharing of environmental perception data adopts an event-triggered mechanism. When a new threat area or a change in the mission target is detected, the trigger node broadcasts the updated information to the cluster. The broadcast range is dynamically adjusted according to the formula:

[0199]

[0200] Where, The broadcast range when the event is triggered (unit: km), is the actual communication range of the triggering node, To prioritize tasks affected by events, events corresponding to high-priority tasks will have their broadcast range expanded to ensure timely dissemination of critical information.

[0201] Through the above mechanism, drone swarms can achieve efficient information interaction in a dynamic environment, provide a consistent state cognition basis for collaborative decision-making, and balance communication overhead and information timeliness.

[0202] A distributed consensus algorithm is used to ensure that the drone cluster maintains coordination and consistency during mission planning and execution, and can reach consensus on adjustments and decisions on task allocation.

[0203] The distributed consensus algorithm achieves global convergence of cluster decision-making through information interaction between neighboring nodes. It uses an improved weighted average consensus protocol, and the decision variables of each drone gradually converge with iteration. When the algorithm is initialized, each drone generates an initial decision vector based on local information. , including key parameters such as task assignment weight and action selection probability, the iterative update formula is:

[0204]

[0205] in, and Respectively Frame and The drone The decision vector of the iteration, is the set of neighbors within its communication range, is the neighbor weight (satisfying and ), the weight value is positively correlated with the communication quality and is calculated by normalizing the signal strength: ,ensuring that neighbors with high communication quality have a greater influence on decision making.

[0206] In order to accelerate consistency convergence and cope with dynamic topology changes, an adaptive damping factor is introduced , the formula is:

[0207]

[0208] Where, is the basic damping coefficient (value 0.8), is the cluster average decision vector, is the decision difference threshold (value is 0.1). When the deviation between the UAV decision and the cluster average is small, the damping factor is reduced to accelerate convergence. When the deviation is large, it is increased to suppress oscillation.

[0209] The decision consensus is determined by the variance convergence criterion. When the standard deviation of the cluster decision vector satisfies the following formula, it is determined that the consensus is reached:

[0210]

[0211] Where N is the total number of drones in the drone swarm, is the convergence accuracy threshold (value is 0.05). For key decisions such as task allocation adjustment, more stringent convergence conditions must be met. , and introduced the Byzantine fault tolerance mechanism to eliminate outliers through majority voting: if the decision of a drone differs from that of more than 60% of its neighbors by more than a threshold, it will be marked as an abnormal node and temporarily isolated, and will rejoin the consensus process after it updates to a reasonable range.

[0212] Through the above mechanism, the cluster can maintain decision consistency when the topology changes dynamically (such as drone failures and communication interruptions). The decision convergence time increases logarithmically with the growth of the cluster size, ensuring the collaborative efficiency of large-scale heterogeneous drone clusters in the task planning and execution process.

[0213] Step 5: Task execution and monitoring.

[0214] The drone performs the mission according to the mission plan and adjusted strategy, and monitors its own status and mission progress in real time during the execution process.

[0215] During the mission execution phase, the drone completes the mission process using a pre-set waypoint sequence and dynamic adjustment mechanism based on a preliminary allocation plan generated by a quantum annealing algorithm and a real-time strategy optimized by reinforcement learning. The execution path is smoothed using a piecewise Bezier curve, and the trajectory parameters between adjacent waypoints are calculated using the following formula:

[0216]

[0217] in, For parameters The corresponding trajectory point coordinates, The four control points of the Bezier curve (including the current position, the next task point and two intermediate transition points), is the number of combinations ( ), reducing energy consumption and flight time by smoothing the trajectory. It is the index of the Bezier curve control point, with values ​​of 0, 1, 2, and 3, corresponding to 4 control points. express to the kth power.

[0218] Self-status monitoring covers three dimensions: power, posture, and device status. The real-time sampling frequency is set to 10Hz, and the abnormal status judgment formula is:

[0219]

[0220]

[0221] in, For the The instantaneous energy consumption deviation of the UAV, for Remaining power at all times, is the average power consumption, The energy consumption abnormality threshold (value is 0.2Wh / s); is the attitude angle deviation, are roll, pitch, and yaw angles, is the standard attitude angle, is the attitude abnormality threshold (value is 5 degrees), and a fault warning is triggered when any condition is met. It is the time interval between two adjacent samplings in self-state monitoring.

[0222] Task progress monitoring uses a dynamic progress bar and time threshold double verification. The progress calculation update formula is:

[0223]

[0224] in, for The task progress at the moment (range 0-1), is the monitoring time interval (unit: s), is the estimated execution time of the current task, is the task weight (positively correlated with priority, When the actual execution time exceeds 1.5 times the estimated time and the progress is less than 70%, it is considered a task delay, triggering the reinforcement learning strategy adjustment mechanism.

[0225] Environmental interference monitoring is achieved through multi-sensor fusion. The formula for determining route deviation caused by sudden strong winds is:

[0226]

[0227] in, is the horizontal distance between the actual position and the planned position (unit: km), is the real-time latitude and longitude, is the planned latitude and longitude, is the theoretical flight distance within this time period. When the distance exceeds 10% of the theoretical distance, path replanning is initiated.

[0228] All monitoring data is uploaded to the cluster shared database in real time through encrypted communication, and Kalman filtering is used for noise suppression. The filter update formula is:

[0229]

[0230]

[0231] in, for The estimated state value at time t, is the sensor measurement value, is the observation matrix, is the Kalman gain, is the forecast error covariance, To measure noise covariance and ensure the accuracy and stability of monitoring data. is the transposed matrix of the observation matrix H. Through the above multi-dimensional monitoring mechanism, the drone cluster can detect anomalies in time during the mission execution and trigger corresponding adjustments or replanning processes to ensure the efficient advancement of the mission.

[0232] If it is found that the mission cannot be executed as planned (including encountering sudden strong winds that cause the drone to deviate from the route or the mission target to disappear), timely feedback will be given to other drones in the cluster or the control center to trigger the re-planning mechanism.

[0233] When an abnormality occurs during mission execution, the abnormality determination module first triggers the feedback mechanism through multi-dimensional threshold detection. For route deviations caused by sudden strong winds, a dynamic threshold determination formula is used:

[0234]

[0235] in, for arrive The cumulative deviation distance at the moment (unit: m), is the actual horizontal velocity component of the UAV, is the planned velocity component, is the deviation tolerance coefficient (value is 0.15). When the cumulative deviation exceeds 15% of the theoretical flight distance, it is judged as a serious deviation. Is the integral variable, representing the time from the initial moment Any time between t and the current time t. is the integration variable The differential of .

[0236] The determination of mission target disappearance is achieved through continuous frame image recognition and lidar data fusion. The confidence calculation formula is:

[0237]

[0238] Where, For the The confidence level of the mission target's existence (range 0-1), is the weight coefficient (value is 0.7), is the visual recognition confidence, is the lidar detection confidence, when If the target disappears for three sampling periods, it is considered as disappearing.

[0239] Abnormal information feedback adopts a hierarchical broadcast mechanism, and the priority is determined by the scope of abnormal impact:

[0240]

[0241] in, is the broadcast priority (range 1-10), The number of drones affected by this exception is 7, and cross-cluster relay forwarding is triggered when the priority is higher than 7, ensuring that the control center and related drones receive the information within 1 second.

[0242] The triggering conditions for the replanning mechanism include dual constraints of time and resources:

[0243] or

[0244] When the remaining flight time of the drone is less than 30% of the estimated time of the current mission, or the remaining battery power is less than 1.2 times the mission requirement, re-planning is forced to be triggered. is the time that the i-th UAV has been used to perform the mission.

[0245] After re-planning and starting, the improved quantum annealing algorithm is used to accelerate the solution, and the initial temperature adjustment formula is:

[0246]

[0247] in, is the initial temperature for emergency planning, The remaining available time, For the emergency response benchmark time (fixed at 60s), the convergence time is shortened by lowering the initial temperature to ensure that the temporary allocation plan is output within 2s, while retaining the real-time adjustment interface of reinforcement learning to cope with continuous dynamic changes.

[0248] Through the above mechanism, the delay time of the entire process from abnormal situation detection to re-planning can be effectively controlled, and the task completion rate of the new plan decreases less than that of the original plan, while ensuring safety and maximizing cluster efficiency.

[0249] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A method for intelligent task planning of heterogeneous UAV swarms, characterized by: The method comprises the following steps: S1. Mission and UAV Information Initialization: Quantify and characterize the endurance, payload capacity, flight speed, communication range, and other relevant performance parameters of each UAV in the heterogeneous UAV cluster. At the same time, determine the mission type, location, priority, resource requirements, and other related information to form a structured initialization data matrix. S2. Preliminary Task Allocation Based on Quantum Annealing Algorithm: A spatiotemporal graph structure consisting of drone nodes, task nodes, and the associated edges between them is constructed. The task allocation problem is transformed into a single-person traveling salesman problem. By constructing the target Hamiltonian and simulating the quantum annealing evolution, a preliminary allocation plan for drone swarm tasks is obtained. S3, Real-time Adjustment Based on Reinforcement Learning: Define a state vector containing the drone's state and environment information, determine the actions the drone can perform, design a multi-dimensional reward function, and optimize the task execution strategy through learning and decision-making processes to adapt to dynamic environmental changes; S4, Communication and Collaboration Mechanism: UAVs share status information, mission execution status, and environmental data through wireless communication networks, and use distributed consistency algorithms to ensure cluster decision-making consensus; S5. Mission execution and monitoring: The drone executes the mission according to the planned strategy, monitors its own status, mission progress and environmental interference in real time, and triggers a re-planning mechanism when the mission cannot be executed as planned.

2. The method for intelligent task planning of heterogeneous UAV swarm according to claim 1, characterized in that: In S1, the drone's endurance is described by the remaining flight time, which takes into account the drone's current remaining battery power, average power consumption under standard flight conditions, and the impact of flight altitude and payload weight on endurance. The quantitative formula for the drone's endurance is: ; in, Indicates the The remaining flight time of the drone, The current remaining battery power of the drone. is its average power consumption under standard flight conditions, The impact of flight altitude and payload weight on flight endurance is comprehensively considered, and the value range is 0.6-1.

2.

3. The method for intelligent task planning of heterogeneous UAV swarm according to claim 1, characterized in that: In S1, the payload capacity of the drone is defined by the maximum payload weight and the compatibility of the types of equipment that can be carried. The compatibility takes into account the importance of each type of equipment required for the mission and the compatibility of the drone with these equipment. The quantitative formula for the payload capacity of a UAV is: ; in, For the The comprehensive payload capacity score of the drone, is the maximum weight it can physically bear, is the total number of equipment types required for the task, For the The importance coefficient of the equipment type, Indicates the The first drone and the Compatibility coefficient of this type of equipment.

4. The method for intelligent task planning of heterogeneous UAV swarm according to claim 1, characterized in that: In S2, the target Hamiltonian consists of a target term reflecting the minimization of task completion time and a constraint term reflecting the task execution order constraint, and the weight of the constraint term can be dynamically adjusted according to the strictness of the constraint; The formula for constructing the target Hamiltonian is: ; in, is the total target Hamiltonian, To reflect the goal of minimizing task completion time, To reflect the constraints of the task execution order, is the constraint strength coefficient.

5. The method for intelligent task planning of heterogeneous UAV swarm according to claim 1, characterized in that: In S2, when converting the task allocation problem into a problem similar to the single-person traveling salesman problem, it is necessary to establish a task-UAV compatibility evaluation model. This model comprehensively considers the matching degree between the UAV payload capacity and the mission equipment requirements, the matching degree between the UAV endurance and the time required to complete the task, and the matching between the task priority and the time window. The evaluation formula for the compatibility between the mission and the drone is: ; in, Indicates the The first drone and the The fitness score of each task, 、 、 is the weight coefficient, For the The minimum required value of the equipment load for each task, For the The current position of the drone to The straight-line distance between the task locations, is the time window matching coefficient.

6. The method for intelligent task planning of heterogeneous UAV swarm according to claim 1, characterized in that: In S3, the reward function includes a basic completion reward linked to task priority and completion quality, a time-efficiency reward to encourage on-time or early completion of tasks, a resource-saving reward that reflects efficient power utilization, a collaborative reward that encourages drones and swarms to maintain strategic consistency, and a penalty term to inhibit failures, task failures, or dangerous behaviors. The reward function is calculated as: ; in, For the The total reward obtained by the drone in the current decision step, As a basic completion reward, For time-limited rewards, Rewards for resource conservation, To reward collaboration, For penalty items.

7. The method for intelligent task planning of heterogeneous UAV swarm according to claim 1, characterized in that: In S3, the feasibility of the UAV switching mission objectives needs to consider the adaptability of the UAV to the new mission, the distance from the UAV to the new mission, the current flight speed and the remaining flight time.

8. The method for intelligent task planning of heterogeneous UAV swarm according to claim 1, characterized in that: In S3, when the drone adjusts its flight speed, it must strike a balance between energy consumption and timeliness, and the adjusted speed must not exceed the drone's maximum speed limit and must not be lower than the minimum safe speed; The formula for adjusting the flight speed of the drone is: ; in, is the adjusted flight speed, is the speed adjustment factor, is the task priority difference, when Speeding up is allowed to complete high-priority tasks first, while ensuring that the adjusted speed does not exceed the maximum speed limit of the drone. and not less than the minimum safe speed .

9. The method for intelligent task planning of heterogeneous UAV swarm according to claim 1, characterized in that: In S4, the communication between drones adopts time division multiple access. Each drone is assigned an independent communication time slot. The length of the time slot is related to the amount of status information data to be transmitted by the drone. Drones with large amounts of information have longer transmission time.

10. The method for intelligent task planning of heterogeneous UAV swarm according to claim 1, characterized in that: In S5, the triggering conditions of the re-planning mechanism include the remaining flight time of the drone being less than a certain proportion of the estimated time of the current mission, or the remaining battery power being less than a certain multiple of the estimated power required for the mission; The triggering conditions for the replanning mechanism are: or ; in, The remaining flight time of the drone. is the estimated execution time of the current task, The remaining battery power of the drone. Estimated power required for the task, is the time that the i-th UAV has been used to perform the mission.

Citation Information

Patent Citations

  • Unmanned aerial vehicle group task allocation method based on quantum annealing algorithm model

    CN117764188A

  • Multi-robot layered formation control method and system based on deep reinforcement learning

    CN118377304A

  • Unmanned aerial vehicle cluster cooperative combat method and system

    CN119126828A

  • Unmanned aerial vehicle route planning method and system based on deep reinforcement learning

    CN120178933A

  • Intelligent transportation methods and systems

    WO2023096968A1

Cited By

  • Distributed intelligent relay dynamic deployment and user scheduling method and system

    CN121077548A

  • Remote unmanned aerial vehicle supervision control system and control method thereof

    CN121165768A

  • Unmanned aerial vehicle group distributed dynamic collaboration method and device

    CN121477979A

  • Distributed dynamic coordination method and device for unmanned aerial vehicle group

    CN121477979B

  • Near field communication matching processing method and system for unmanned aerial vehicle in communication area without public network

    CN121486790A