A Multi-UAV Multi-Target Mission Planning Method, System and Electronic Device
By combining graph neural network and deep reinforcement learning technology, integrating constraints and real-time feedback mechanisms, the complexity and dynamic problems in multi-UAV task planning are solved, and efficient and flexible task planning and execution are achieved.
Patent Information
- Application Number
- CN202510346085.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-24
AI Technical Summary
Existing multi-UAV mission planning methods are difficult to achieve efficient task allocation, path planning, resource management and collaborative optimization in complex dynamic environments, especially when mission complexity and number of drones increase.
Graph neural network and deep reinforcement learning technology are used to generate multi-UAV task allocation graphs, integrate battery power, environmental information and resource consumption constraints, update node characteristics through graph neural network, and optimize task allocation strategies using reinforcement learning to adjust task allocation plans in real time.
It realizes efficient and flexible multi-UAV mission planning, which can dynamically adapt in complex dynamic environments, optimize resource utilization, and improve task execution efficiency and effectiveness.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention relates to the field of unmanned aerial vehicles, and particularly to a multi - UAV multi - target mission planning method, system, and electronic device. Background Art
[0002] With the rapid development of technology, unmanned aerial vehicles are increasingly widely used in various industries. Especially in the low - altitude economy field, multi - UAV systems have become an important tool for achieving multi - target missions due to their high efficiency, flexibility, and low cost. The low - altitude economy covers multiple application scenarios such as environmental monitoring, logistics distribution, agricultural spraying, urban security, infrastructure inspection, disaster emergency response, and entertainment activities. In these applications, multi - UAV collaborative work can significantly improve the efficiency and effectiveness of task execution and meet complex and changing work requirements.
[0003] Existing multi - UAV mission planning methods mainly focus on task allocation, path planning, and resource management. However, with the increase in task complexity and the number of UAVs, traditional methods face many challenges:
[0004] 1) Complexity of task allocation: In multi - UAV multi - target missions, how to efficiently allocate tasks to different UAVs to make full use of their respective resources and capabilities is a complex problem. Traditional task allocation methods often rely on predefined rules or simple optimization algorithms and are difficult to adapt to dynamically changing task requirements and environmental conditions.
[0005] 2) Dynamics of path planning: The low - altitude environment is complex and changeable. The existence of factors such as weather changes, terrain obstacles, and other dynamic obstacles makes path planning more difficult. Existing path planning algorithms such as the A* algorithm and Dijkstra algorithm perform well in static environments but often lack real - time adjustment capabilities in dynamic environments and cannot effectively handle emergencies.
[0006] 3) Optimization of resource management: Multi - UAV systems need to reasonably manage resources such as battery power, communication bandwidth, and load capacity to ensure the smooth completion of tasks. Traditional methods often lack comprehensiveness and real - time nature in resource management and are difficult to achieve optimal resource allocation in complex tasks.
[0007] 4) Real - time nature of collaborative work: In multi - UAV systems, the collaborative work between UAVs is crucial for the success of tasks. However, existing methods have deficiencies in achieving real - time collaboration and information sharing, resulting in delays or conflicts during task execution.
[0008] To address the above problems, in recent years, advanced artificial intelligence technologies such as graph neural networks and deep reinforcement learning have been introduced into multi - UAV mission planning. These technologies model the complex relationships between UAVs and missions, learn optimization strategies, and achieve more efficient and flexible task allocation and path planning. However, existing research still has the following deficiencies:
[0009] 1) Insufficient comprehensive consideration of constraints: In the process of task allocation and path planning, existing methods often fail to fully consider constraints such as the battery power of UAVs, environmental information, and resource consumption, resulting in problems with feasibility and efficiency in the actual application of task allocation schemes.
[0010] 2) Limited real - time feedback adjustment ability: In a dynamic environment, the task execution situation and environmental conditions change continuously. Existing methods lack an effective real - time feedback mechanism and cannot dynamically adjust the task allocation scheme according to real - time information, affecting the overall execution effect of the task.
[0011] 3) Insufficient multi - agent collaborative optimization: Although multi - agent reinforcement learning can theoretically optimize task allocation strategies through the cooperation and competition of multiple agents, in actual applications, how to efficiently achieve collaborative optimization among multiple UAVs remains an urgent problem to be solved. Summary of the Invention
[0012] The purpose of the embodiments of the present invention is to provide a multi - UAV multi - target mission planning method, system, and electronic device to achieve efficient and flexible multi - UAV mission planning. The specific technical solutions are as follows:
[0013] In the first aspect, the embodiments of the present invention provide a multi - UAV multi - target mission planning method, and the method includes:
[0014] Generate a multi - UAV task allocation graph; wherein, the nodes in the multi - UAV task allocation graph represent multiple UAVs and multiple tasks, and the edges in the multi - UAV task allocation graph represent the relationships between tasks and UAVs;
[0015] Collect and pre - process the state information and environmental information of UAVs. The state information includes: battery power, position information, and the environmental information includes: weather conditions, terrain features, and obstacle distribution;
[0016] Obtain constraints and integrate the constraints into the task allocation model;
[0017] Based on the constraints integrated in the task allocation model, the state information of each UAV, and the environmental information, use a graph neural network to update the node features of the multi - UAV task allocation graph and generate the first task allocation strategy for each UAV;
[0018] Based on the rewards and punishments for the mission execution of each drone for feedback, a reinforcement learning algorithm is used to optimize the first mission allocation strategy, and a second mission allocation strategy for each drone is generated;
[0019] According to the environmental information and the state information of each drone, the first mission allocation strategy and / or the second mission allocation strategy are adjusted in real time to generate a third mission allocation strategy for each drone.
[0020] In one embodiment of the present invention, the constraint conditions include battery power constraint, environmental information constraint, and resource consumption constraint;
[0021] The battery power constraint is used to ensure that the mission allocation will not cause the drone's power to run out, and the battery power constraint is expressed by the following formula:
[0022] ;
[0023] Where represents the remaining power of the th drone, is the minimum remaining power required after the drone executes the mission;
[0024] The environmental information constraint is used to optimize the flight path and mission execution plan of the drone. Among them, the following path planning algorithm is used to optimize the flight path of the drone:
[0025] ;
[0026] Where is the flight path of the th drone, is star algorithm, used to calculate the optimal flight path, represents the current environmental information, represents the mission executed by the th drone;
[0027] The resource consumption constraint is used to optimize the mission allocation of the drone to improve the resource utilization efficiency, and the resource utilization efficiency is expressed by the following formula:
[0028] ;
[0029] Where is the resource utilization efficiency, is the number of drones, is the th drone's mission quantity, is the th drone's total resource consumption.
[0030] In one embodiment of the present invention, the graph neural network is a graph attention network. The use of the graph neural network to update the node features of the multi-UAV task allocation graph includes:
[0031] Initializing the features of each node in the multi-UAV task allocation graph according to the status information or task requirements of the UAVs;
[0032] For each neighbor node of any target node in the multi-UAV task allocation graph, calculating the attention weight of the neighbor node to the target node;
[0033] Using the graph attention network to update the node features.
[0034] In one embodiment of the present invention, the graph neural network is one of a graph convolutional network, a graph isomorphism network, and a graph sampling and aggregation generator GraphSAGE.
[0035] In one embodiment of the present invention, the reinforcement learning algorithm is a deep reinforcement learning algorithm. The feedback of rewards and punishments for task execution by each UAV is used, and the reinforcement learning algorithm is used to optimize the first task allocation strategy, including:
[0036] Training each UAV as an independent agent, and selecting actions according to the current state of the UAV when performing tasks;
[0037] Determining the reward value and the next state after the UAV executes the action;
[0038] Recording and constructing a trajectory according to the current state, action, reward, and next state of each time period;
[0039] Determining the cumulative return according to the trajectory corresponding to each time period;
[0040] Updating the first task allocation strategy according to the cumulative return and maximizing the expected return.
[0041] In one embodiment of the present invention, the real-time adjustment of the first task allocation strategy and / or the second task allocation strategy includes:
[0042] The first task allocation strategy and / or the second task allocation strategy is adjusted in real time by using at least one of the following methods:
[0043] When a preset condition is triggered during the task execution of the first UAV, the task of the first UAV is assigned to the second UAV, and the preset condition indicates that the task execution of the UAV is blocked;
[0044] Adjust the task priority according to the state of the drone, the resource situation of the drone, and the task urgency level;
[0045] Adjust the flight path and / or task of the drone according to the current environmental information.
[0046] In a second aspect, an embodiment of the present invention provides a multi-drone multi-target task planning system, which includes a server and drones;
[0047] The server is used to perform multi-drone multi-target task planning according to the multi-drone multi-target task planning method of any one of the above to generate a task allocation strategy, and send the task allocation strategy to the drones;
[0048] The drones are used to receive the task allocation strategy and execute tasks according to the task allocation strategy.
[0049] In a third aspect, an embodiment of the present invention provides a multi-drone multi-target task planning device, which includes:
[0050] A task allocation graph generation module, which is used to generate a multi-drone task allocation graph; wherein, the nodes in the multi-drone task allocation graph represent multiple drones and multiple tasks, and the edges in the multi-drone task allocation graph represent the relationships between tasks and drones;
[0051] An information collection and processing module, which is used to collect and preprocess the state information and environmental information of the drones. The state information includes: battery power, position information, and the environmental information includes: weather conditions, terrain features, and obstacle distribution. The state information and environmental information are used for graph neural network processing;
[0052] A constraint condition acquisition and integration module, which is used to acquire constraint conditions and integrate the constraint conditions into the task allocation model;
[0053] A first task allocation strategy generation module, which is used to update the node features of the multi-drone task allocation graph using a graph neural network based on the constraint conditions integrated in the task allocation model, the state information of each drone, and the environmental information, and generate a first task allocation strategy for each drone;
[0054] A second task allocation strategy generation module, which is used to feedback the rewards and punishments for task execution by each drone, and optimize the first task allocation strategy using a reinforcement learning algorithm to generate a second task allocation strategy for each drone;
[0055] The third task allocation strategy generation module is configured to adjust the first task allocation strategy and / or the second task allocation strategy in real time according to the environmental information and the status information of each drone, and generate a third task allocation strategy for each drone.
[0056] In one embodiment of the present invention, the constraint conditions include battery power constraint, environmental information constraint, and resource consumption constraint;
[0057] The battery power constraint is used to ensure that the task allocation will not cause the drone to run out of power. The battery power constraint is expressed by the following formula:
[0058] ;
[0059] Wherein, represents the remaining power of the th drone, is the minimum remaining power required after the drone executes the task;
[0060] The environmental information constraint is used to optimize the flight path and task execution plan of the drone. Among them, the following path planning algorithm is used to optimize the flight path of the drone:
[0061] ;
[0062] Wherein, is the flight path of the th drone, is star algorithm, used to calculate the optimal flight path, represents the current environmental information, represents the th task executed by the drone;
[0063] The resource consumption constraint is used to optimize the task allocation of the drone to improve the resource utilization efficiency. The resource utilization efficiency is expressed by the following formula:
[0064] ;
[0065] Wherein, is the resource utilization efficiency, is the number of drones, is the th number of tasks executed by the drone, is the th total amount of resources consumed by the drone.
[0066] Beneficial effects of the embodiments of the present invention:
[0067] The present invention solves the problems of complex task allocation, path planning, resource management, and collaborative optimization in multi-UAV multi-target task planning by combining graph neural network and deep reinforcement learning techniques. Compared with traditional methods, the present invention has the following advantages:
[0068] 1) Efficient task allocation: Through the deep modeling of the relationship between tasks and UAVs by the graph neural network, a task allocation scheme can be accurately generated, and on this basis, the task execution strategy can be optimized through reinforcement learning.
[0069] 2) Strong dynamic adaptability: Through the real-time feedback mechanism, the present invention can dynamically adjust the task allocation scheme according to environmental changes and task execution status to ensure the efficient completion of tasks in complex dynamic environments.
[0070] 3) Optimization of resource utilization: By integrating constraints such as battery power and environmental information, resource consumption is effectively optimized, avoiding excessive consumption of battery power or other resources by UAVs, and improving the feasibility and efficiency of task execution.
[0071] 4) Flexibility and scalability: The task planning method and system of the present invention can be flexibly applied to various low-altitude economic scenarios, such as logistics distribution, environmental monitoring, agricultural spraying, security patrol, etc., and can adapt to changes in different task complexities and the number of UAVs. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.
[0073] Figure 1 It is a schematic flow chart of a multi-UAV multi-target task planning method provided by an embodiment of the present invention;
[0074] Figure 2 It is a schematic structural diagram of a multi-UAV multi-target task planning system provided by an embodiment of the present invention;
[0075] Figure 3 It is a schematic structural diagram of a multi-UAV multi-target task planning device provided by an embodiment of the present invention;
[0076] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0077] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art based on the present invention belong to the scope of protection of the present invention.
[0078] See Figure 1 , which is a schematic flowchart of a multi-UAV multi-target task planning method provided by an embodiment of the present application. This method can be applied to electronic devices such as servers and intelligent terminals. This method includes steps S11-S16.
[0079] S11, generate a multi-UAV task assignment graph.
[0080] Among them, the nodes in the multi-UAV task assignment graph represent multiple UAVs and multiple tasks, and the edges in the multi-UAV task assignment graph represent the relationships between tasks and UAVs.
[0081] This process is the process of graph modeling. In the embodiments of the present invention, first, a multi-UAV task assignment graph needs to be constructed, with UAVs and tasks represented as nodes in the graph, and the relationships between tasks and UAVs are connected by edges in the graph.
[0082] The structure of the graph model is as follows:
[0083] UAV node: Each UAV node contains the current status information of the UAV, such as position information, battery power, flight speed, etc.
[0084] Task node: Each task node contains the basic information of the task, such as task location, task priority, task type, etc.
[0085] Establishment of edges: The edges between task nodes and UAV nodes represent the potential association degree between tasks and UAVs, and the weights of the edges can be calculated based on factors such as distance and task priority.
[0086] S12, collect and preprocess environmental information, UAV status information, and task information.
[0087] In the embodiments of the present invention, the following data is collected and processed to ensure that all data can be correctly input into the graph neural network:
[0088] UAV status data: battery power, position information, speed information, etc.;
[0089] Environmental information data: weather conditions, obstacle positions, and terrain information, etc.;
[0090] Task information data: task location, task priority, task urgency, etc.
[0091] These data need to be cleaned, denoised, and normalized, and then converted into a format suitable for input to the graph neural network. For example, the battery power of the drone can be normalized according to the maximum power to ensure the consistency of data input.
[0092] Specifically, to ensure the accuracy of the data, a data matrix can be constructed based on the normalized data, and a standard data matrix can also be constructed for the data under normal conditions. Subsequently, the correlation values between the data matrix and the standard data matrix are calculated respectively according to the Pearson coefficient. When the correlation value is within the preset range, it is determined that the normalized data meets the requirements; otherwise, further processing is required, such as performing a trace check, etc. Or, based on the data matrix, the normalized data is used as the gray value, and the pixel points corresponding to the gray value are filled according to the positions of these data in the data matrix, thereby obtaining the gray pixel map of these data. Then, the gray pixel map is input into the neural network model, and the confidence level is calculated through the softmax function. When the confidence level is within the preset range, it is determined that these data meet the requirements; otherwise, further processing is performed, and so on, which is not limited here.
[0093] S13, obtain the constraint conditions and integrate the constraint conditions into the graph neural network.
[0094] In the embodiment of the present invention, the following constraint conditions are considered to ensure that the task allocation scheme is feasible and efficient in actual execution:
[0095] 1) Battery power constraint, according to the remaining power of each drone , ensure that the task allocation does not cause the drone's power to run out. The constraint condition is expressed as:
[0096] ;
[0097] Among them, represents the remaining power of the th drone, is the minimum remaining power required for the drone after executing the task;
[0098] 2) Environmental information constraint, according to the current environmental information (such as weather, terrain, obstacles, etc.), optimize the flight path and task execution plan of the drone, and the path planning is based on an appropriate path planning algorithm;
[0099] ;
[0100] Among them, is the flight path of the th drone, is the star algorithm, used to calculate the optimal flight path, Represents the current environmental information, Represents the task executed by the
[0101] 3) Resource consumption constraint. Considering the resource consumption of the drones (such as fuel, communication bandwidth, etc.), optimize the task allocation to maximize the resource utilization efficiency, and the resource utilization efficiency is expressed as:
[0102] ;
[0103] Among them, is the number of drones, is the number of tasks executed by the th drone, is the total amount of resources consumed by the
[0104] S14. Based on the constraint conditions, the status information and task information of each drone, and the environmental information, use a graph neural network to update the node features of the multi-drone task allocation graph, and generate the first task allocation strategy for each drone.
[0105] This step is the step processed by the graph neural network. In this step, the graph neural network is used to process the multi-drone task allocation graph. Specifically, the graph neural network can be a graph attention network, and step S14 can be specifically executed through the following steps.
[0106] S141. Initialize the features of each node in the multi-drone task allocation graph according to the status of the drone (such as battery power, location, task type, etc.) or the requirements of the task (such as task priority, resource requirements, etc.);
[0107] S142. For each neighbor node , calculate its attention weight for node , specifically including:
[0108] First, concatenate the features of node and the remaining battery power:
[0109] ;
[0110] Among them, respectively represent the feature vectors of node , represents the remaining battery power of node , represents the vector concatenation operation;
[0111] Then, use a weight matrix Perform a linear transformation on the concatenated features:
[0112] ;
[0113] Among them, is a learnable weight matrix, and a is a learnable parameter vector, representing the process of converting the concatenated features into a scalar.
[0114] Obtain the original attention score After that, the LeakyReLU activation function is usually used for non-linear transformation to obtain the final attention weights:
[0115] ;
[0116] Next, by adjusting the weights of the attention scores, ensure that drones (nodes) with lower battery levels have less impact on task allocation. Specifically, through a weight factor to weighted adjust the attention scores:
[0117] ;
[0118] Among them, is the adjustment factor of the battery level, indicating the degree of influence of the battery level on the attention weights, and can be defined as:
[0119] ;
[0120] Among them, is the maximum battery level of this type of drone (i.e., the maximum power under normal conditions);
[0121] Finally, use the softmax function to normalize the attention weights of all neighbor nodes:
[0122] ;
[0123] S143. Use the graph structure, attention mechanism, and combine with actual constraints to update the node features:
[0124] ;
[0125] Among them, is a function responsible for combining the environmental information with the features of the current node and adjusting the update of the node features, is also a function that considers the resource consumption of each node and adjusts the update of the node features, and are weight coefficients used to adjust the influence degree of each constraint condition in the feature update, Represents a node At the layer feature Represents a node The set of neighbor nodes of For node For node At the layer attention weight For the layer weight matrix, used to transform the features of neighbor nodes Represents a node At the layer feature Represents an activation function, used for non-linear transformation of features
[0126] Through graph attention operation, each node in the graph will update its features according to the information of adjacent nodes, so as to capture the complex relationship between drones and tasks, and generate a preliminary task allocation plan for each drone, that is, the first task allocation strategy
[0127] In some embodiments of the present invention, the graph neural network can also be one of a graph convolutional network, a graph isomorphism network, and a graph sampling and aggregation generator GraphSAGE
[0128] S15. Based on the feedback of rewards and punishments for each drone's task execution, use a reinforcement learning algorithm to optimize the first task allocation strategy and generate a second task allocation strategy
[0129] This step is the process of reinforcement learning optimization. Specifically, optimize the task allocation plan through deep reinforcement learning. Under the multi-agent reinforcement learning framework, design a reward function to guide the drone to select the optimal task allocation strategy. The specific steps of this process are as follows
[0130] S151. Define the environment, state, and action space of reinforcement learning
[0131] 1) State : In each time step t, the state of each drone
[0132] 2) Action : The actions that the drone can take
[0133] 3) Reward : Give corresponding rewards according to the action results of each drone
[0134] S152. Initialize the policy and define a parameterized policy , such as defining the first task allocation strategy as a parameterized policy, using a deep neural network architecture to approximate the policy, and the input of the neural network is the current state The output is the probability distribution of each action in this state; the deep neural network architecture includes but is not limited to multi-layer perceptrons and convolutional neural networks.
[0135] S153. Multi-agent interaction and cooperation. Each drone is trained as an independent agent and selects corresponding actions according to the current state when performing tasks; there is both cooperation and competition among multiple drones, and each drone considers the strategies and task states of other drones when selecting actions.
[0136] S154. Interaction with the environment to collect experience, specifically including:
[0137] 1) Each drone selects an action according to the policy at the current moment.
[0138] 2) After the drone executes the action, the environment feedbacks a new state and a reward .
[0139] 3) Record the state, action, reward, and next state at each time step to form a trajectory:
[0140] S155. Calculate the cumulative return , at each time step , calculate the cumulative return starting from this time step according to the obtained reward :
[0141] ;
[0142] where is the discount factor, which controls the influence of future rewards on the current decision, is the reward value at the th moment, is the task termination time;
[0143] S156. Use the policy gradient method to update the policy parameters , with the goal of maximizing the expected return , and the policy update rule is:
[0144] ;
[0145] where is the expected return of the policy, is the policy parameter, is the policy, is the action trajectory of the drone;
[0146] S156. Continuously optimize the strategy through multiple interactions and trainings .
[0147] S16. According to the environmental information, the status information of each drone, and the task information, adjust the first task allocation strategy or the second task allocation strategy in real time to generate a third task allocation strategy.
[0148] This step is the process of real-time feedback adjustment. During the task execution, the system adjusts the task allocation plan according to the real-time feedback information. For example, data such as the battery power of the drone, power consumption, task execution progress, and environmental changes will be continuously monitored to adjust the task allocation in real time. The specific adjustment strategies are as follows:
[0149] Dynamically optimize the task allocation strategy. Based on the real-time feedback information, dynamically adjusting the task allocation strategy includes:
[0150] 1) Reallocate tasks: If the task execution of some drones is blocked (such as insufficient battery power, flight path blocked by obstacles, etc.), reallocate the tasks to other drones.
[0151] 2) Adjust priorities: Dynamically adjust the priorities of tasks according to the urgency of the tasks, resource conditions, or the status of the drones.
[0152] 3) Path and action adjustment: If the environment changes (such as weather changes, obstacle movement, etc.), adjust the flight path or action plan of the drone according to the new environmental information.
[0153] In an embodiment of the present invention, further optimization can be performed. Real-time feedback adjustment is usually achieved through the following methods:
[0154] 1) Dynamic adjustment of reinforcement learning: Through a real-time reward and punishment mechanism, use the reinforcement learning algorithm to continuously adjust the task allocation strategy. After each adjustment, provide feedback according to the task completion degree and efficiency to further optimize the decision-making.
[0155] 2) Prediction and real-time correction: The system can use a prediction model to estimate the task completion time and resource requirements, and correct these predictions in combination with real-time feedback.
[0156] 3) Heuristic algorithm: Quickly adjust the task allocation through simple rules or heuristic methods (such as based on information such as the remaining time of the task and the status of the drone).
[0157] Next, in combination with three different application scenarios, the multi-drone multi-target task planning method provided by the embodiment of the present invention will be described.
[0158] In the following embodiments, three specific application scenarios of the multi - UAV multi - target mission planning method are presented. These scenarios involve different mission types, environmental conditions, and resource constraints, demonstrating the broad adaptability and high efficiency of the method of the present invention.
[0159] Scenario 1: UAV swarm performing mapping tasks
[0160] Mission description: In this scenario, multiple UAVs need to cooperate to perform mapping tasks. The goal is to conduct high - precision topographic mapping of a certain area. There are multiple target locations distributed within the mission area, and each target location needs to be scanned in detail by a UAV.
[0161] Specific operation steps:
[0162] Graph modeling: The system first constructs a task assignment graph. The UAV nodes contain status information such as the position and battery level of each UAV, and the task nodes contain information such as target locations and task priorities.
[0163] Data collection and pre - processing: Collect information such as the battery level, flight speed, position, task priority, and terrain features of each UAV. Adjust the weights of the task nodes according to weather conditions (such as wind speed and visibility).
[0164] Constraint analysis and integration: Integrate the battery level constraint and flight path constraint to ensure that the UAV task assignment conforms to the endurance ability. The priority of the task nodes is dynamically adjusted according to the complexity of the target terrain.
[0165] Graph neural network processing: Update the connection between UAVs and tasks through graph convolution operations and generate a preliminary task assignment plan. The system will consider task priorities, flight paths, and environmental constraints to optimize the assignment results.
[0166] Reinforcement learning optimization: Use multi - agent reinforcement learning to optimize the task assignment plan. Each UAV adjusts its flight path according to the state feedback of task execution to maximize the overall task completion efficiency.
[0167] Real - time feedback adjustment: Adjust the task assignment according to the real - time battery level and flight progress. When the battery level of a UAV is approaching depletion, the system will re - assign tasks to ensure that each UAV completes an appropriate task.
[0168] Technical effect: Through the method of the present invention, UAVs can efficiently complete mapping tasks in complex terrains and dynamic environments, and adjust the task assignment plan in real time to ensure the efficiency and accuracy of mapping.
[0169] Scenario 2: Mission planning of UAV swarms in post - disaster rescue
[0170] Task description: After a natural disaster occurs, multiple drones need to cooperate to perform post-disaster search and rescue tasks. The tasks include assessing key areas in the disaster area, searching for trapped people, and providing real-time data feedback to the ground command center.
[0171] Specific operation steps:
[0172] Graph modeling: The drone nodes include information such as the battery power, flight time, and current position of the drones; the task nodes include the key positions in the disaster area and the task priorities (such as rescue personnel, danger area assessment, etc.).
[0173] Data collection and preprocessing: Collect environmental information (such as wind speed, weather conditions, obstacles in the area, etc.), and preprocess this data into a format suitable for graph neural networks. Dynamically adjust the weights of the task nodes according to factors such as weather, building distribution, and severity of the disaster.
[0174] Constraint condition analysis and integration: Battery power, power consumption, communication bandwidth, etc. will all be analyzed as constraint conditions. The flight path is also affected by obstacles and weather conditions, so environmental information needs to be integrated into the task allocation model to ensure that the drones can execute tasks smoothly in a complex environment.
[0175] Graph neural network processing: Update the task allocation plan for each drone through a graph neural network, and dynamically consider the battery power, task priority, and environmental changes of the drones. The task allocation plan will be adjusted in real time to ensure that all tasks can be completed as efficiently as possible in the post-disaster environment.
[0176] Reinforcement learning optimization: The system optimizes the task allocation strategy through multi-agent reinforcement learning. Each drone acts as an agent and adjusts its task execution strategy according to the tasks it performs, environmental changes, and feedback information to maximize the task completion efficiency.
[0177] Real-time feedback adjustment: According to the real-time situation in the disaster area (such as changes in the disaster area environment, changes in the task execution status, etc.), the system will adjust the task allocation plan in real time. Through the feedback mechanism, the drones can dynamically optimize the task allocation according to the real-time battery power, task urgency, and environmental changes.
[0178] Technical effect: Through the method of the present invention, multiple drones can cooperate efficiently in the post-disaster rescue scenario, ensure that each task is completed according to the priority and resource limitations, improve the rescue efficiency, and help the command center obtain key data of the disaster area in a timely manner.
[0179] Scenario 3: A group of drones perform cargo delivery tasks
[0180] Task description: In the urban goods delivery scenario, multiple drones need to cooperate to complete the delivery tasks of multiple goods. The tasks include transporting the goods from the distribution center to different destinations according to the real-time order requirements.
[0181] Specific operation steps:
[0182] Graph modeling: The system constructs a task assignment graph, where each node in the graph represents a drone or a delivery task. The task nodes include information such as the target address, the weight of the goods, and the task time requirements. The drone nodes include information such as the battery power, flight speed, and position of the drones.
[0183] Data collection and preprocessing: Collect order information, delivery requirements, real-time traffic conditions, etc., and preprocess them into a format suitable for input to the graph neural network. Preprocess the resource consumption situation of the drones, traffic conditions, etc., to ensure the real-time and accuracy of the data.
[0184] Constraint analysis and integration: Multiple constraints need to be considered during the task assignment process, such as battery power, traffic restrictions, time limitations, etc. Each task node will dynamically adjust its weight according to the delivery requirements (such as timeliness, type of goods, etc.).
[0185] Graph neural network processing: Through graph convolution operations, the graph neural network generates a task assignment plan for each drone. Each drone selects and assigns tasks based on its own status information and environmental information, minimizing task delays and power consumption as much as possible.
[0186] Reinforcement learning optimization: Optimize the task assignment strategy through reinforcement learning. Each drone continuously adjusts its strategy based on the feedback of historical task executions to maximize the delivery efficiency and reduce resource consumption.
[0187] Real-time feedback adjustment: Real-time monitor information such as the battery power of the drones, the progress of task completion, and traffic conditions. The system will dynamically adjust the task assignment based on this real-time data to ensure that each task can be completed on time.
[0188] Technical effect: Through the method of the present invention, the drones can efficiently complete the goods delivery tasks in the urban environment, optimize the path planning and task assignment, reduce the delivery time, improve the resource utilization rate, reduce task conflicts, and maximize the system efficiency.
[0189] The present invention shows broad application potential in multiple fields and is widely applicable to multiple application scenarios in the low-altitude economy, such as environmental monitoring, logistics distribution, agricultural spraying, urban security, infrastructure inspection, disaster emergency response, etc., including but not limited to the above application scenarios.
[0190] Based on the same inventive concept, the embodiment of the present invention also provides a multi-drone multi-target task planning system, asFigure 2 As shown, the system includes a server 21 and a drone 22;
[0191] The server 21 is used to perform multi-drone multi-target task planning according to the multi-drone multi-target task planning method given in any of the above embodiments, so as to generate a task assignment strategy and send the task assignment strategy to the drone;
[0192] The drone 22 is used to receive the task assignment strategy and execute the task according to the task assignment strategy.
[0193] Based on the same inventive concept, an embodiment of the present invention also provides a multi-drone multi-target task planning device, as Figure 3 shown, the device includes:
[0194] A task assignment graph generation module 31, configured to generate a multi-drone task assignment graph; wherein, the nodes in the multi-drone task assignment graph represent multiple drones and multiple tasks, and the edges in the multi-drone task assignment graph represent the relationship between the tasks and the drones;
[0195] An information collection and processing module 32, configured to collect and preprocess the state information and environmental information of the drones. The state information includes: battery power and position information, and the environmental information includes: weather conditions, terrain features, and obstacle distribution. The state information and environmental information are used for graph neural network processing;
[0196] A constraint condition acquisition and integration module 33, configured to acquire constraint conditions and integrate the constraint conditions into the task assignment model;
[0197] A first task assignment strategy generation module 34, configured to update the node features of the multi-drone task assignment graph using a graph neural network based on the constraint conditions integrated in the task assignment model, the state information of each drone, and the environmental information, and generate a first task assignment strategy;
[0198] A second task assignment strategy generation module 35, configured to feedback the rewards and punishments for task execution by each drone, and optimize the first task assignment strategy using a reinforcement learning algorithm to generate a second task assignment strategy;
[0199] A third task assignment strategy generation module 36, configured to adjust the first task assignment strategy and / or the second task assignment strategy in real time according to the environmental information and the state information of each drone, and generate a third task assignment strategy.
[0200] An embodiment of the present invention also provides an electronic device, as Figure 4As shown, it includes a processor 41, a communication interface 42, a memory 43, and a communication bus 44. Among them, the processor 41, the communication interface 42, and the memory 43 complete mutual communication through the communication bus 44.
[0201] The memory 43 is used to store computer programs;
[0202] When the processor 41 is used to execute the program stored on the memory 43, the steps of any of the above multi-UAV multi-target mission planning methods are implemented. The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0203] The communication interface is used for communication between the above electronic device and other devices.
[0204] The memory can include a Random Access Memory (RAM), or can also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located far from the aforementioned processor.
[0205] The above processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0206] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0207] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article, or device comprising the element.
[0208] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.
[0209] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included within the protection scope of the present invention.
Claims
1. A multi-UAV multi-objective collaborative task planning method, characterized in that: The method comprises: Generate a multi-UAV task allocation graph; wherein the nodes in the multi-UAV task allocation graph represent multiple UAVs and multiple tasks, and the edges in the multi-UAV task allocation graph represent the relationship between tasks and UAVs; Collect and pre-process environmental information, drone status information, and mission information; Obtaining constraints and integrating the constraints into the graph neural network; Based on the constraint conditions, the state information and task information of each drone, and the environmental information, a graph neural network is used to update node features of the multi-drone task allocation graph to generate a first task allocation strategy; Based on the feedback of rewards and penalties for task execution by each UAV, the first task allocation strategy is optimized by using a reinforcement learning algorithm to generate a second task allocation strategy; According to the environmental information, the status information of each UAV, and the task information, the first task allocation strategy or the second task allocation strategy is adjusted in real time to generate a third task allocation strategy; The constraints include battery power constraints, environmental information constraints, and resource consumption constraints; The battery power constraint is used to ensure that the task allocation will not cause the drone to run out of power. The battery power constraint is expressed by the following formula: AND i ≥E min Among them, E i represents the remaining power of the ith drone, E min The minimum remaining power required for the drone to perform a mission; The environmental information constraints are used to optimize the flight path and mission execution plan of the UAV, wherein the flight path of the UAV is optimized using the following path planning algorithm: Path i =A * (Environment,Task i ) Among them, Path i is the flight path of the ith UAV, A * It is the A-star algorithm, which is used to calculate the optimal flight path. Environment represents the current environment information. Task i represents the task performed by the i-th UAV; The resource consumption constraint is used to optimize the task allocation of the UAV to improve the resource utilization efficiency. The resource utilization efficiency is expressed by the following formula: Among them, N is the number of drones, Tasks i The number of tasks performed for the i-th UAV, Resources i is the total amount of resources consumed by the i-th drone; The step of adjusting the first task allocation strategy and / or the second task allocation strategy in real time includes: The first task allocation strategy and / or the second task allocation strategy are adjusted in real time in at least one of the following ways: When a preset condition is triggered during the execution of a task by the first drone, the task of the first drone is assigned to the second drone, wherein the preset condition indicates that the execution of the task by the drone is blocked; Adjust task priorities based on the status of the drone, the drone’s resources, and the urgency of the task; Adjust the drone’s flight path and / or mission based on current environmental information.
2. The method according to claim 1, characterized in that The graph neural network is a graph attention network, and the use of the graph neural network to update node features of the multi-UAV task allocation graph includes: Initialize each node feature in the multi-UAV task allocation graph according to the state information or task requirements of the UAV; For each neighbor node of any target node in the multi-UAV task allocation graph, calculating the attention weight of the neighbor node to the target node; Use the graph attention network to update node features.
3. The method according to claim 1, characterized in that The graph neural network is one of a graph convolutional network, a graph isomorphism network, a graph sampling and aggregation generator GraphSAGE.
4. The method according to claim 1, characterized in that: The reinforcement learning algorithm is a deep reinforcement learning algorithm. The feedback of rewards and penalties for task execution by each drone is performed, and the reinforcement learning algorithm is used to optimize the first task allocation strategy, including: Train each drone as an independent agent, selecting actions based on the drone’s current state when performing a task; Determine the reward value and next state of the drone after executing the action; Record and construct trajectories based on the current state, action, reward, and next state for each time period; Determine the cumulative return based on the trajectory corresponding to each time period; The first task allocation strategy is updated according to the accumulated reward and the maximized expected reward.
5. A multi-UAV multi-objective collaborative mission planning system, characterized in that: The system includes a server and a drone; The server is used to perform multi-UAV multi-objective collaborative task planning according to the method according to any one of claims 1 to 4 to generate a task allocation strategy, and send the task allocation strategy to the UAV; The drone is used to receive the task allocation strategy and execute the task according to the task allocation strategy.
6. A multi-UAV multi-target collaborative task planning device, characterized in that: The device comprises: A task allocation graph generation module is used to generate a multi-UAV task allocation graph; wherein the nodes in the multi-UAV task allocation graph represent multiple UAVs and multiple tasks, and the edges in the multi-UAV task allocation graph represent the relationship between tasks and UAVs; Information collection and processing module, used to collect and pre-process environmental information, drone status information and mission information; A constraint acquisition and integration module, used to acquire constraint conditions and integrate the constraint conditions into the graph neural network; A first task allocation strategy generation module is used to use a graph neural network to update node features of the multi-UAV task allocation graph based on the constraint conditions, the state information of each UAV and the environmental information, and generate a first task allocation strategy; A second task allocation strategy generation module is used to generate a second task allocation strategy by optimizing the first task allocation strategy using a reinforcement learning algorithm through feedback of rewards and penalties for task execution by each UAV; A third task allocation strategy generating module, configured to adjust the first task allocation strategy or the second task allocation strategy in real time according to the environmental information, the status information of each UAV and the task information, and generate a third task allocation strategy; The constraints include battery power constraints, environmental information constraints, and resource consumption constraints; The battery power constraint is used to ensure that the task allocation will not cause the drone to run out of power. The battery power constraint is expressed by the following formula: AND i ≥E min Among them, E i represents the remaining power of the ith drone, E min The minimum remaining power required for the drone to perform a mission; The environmental information constraints are used to optimize the flight path and mission execution plan of the UAV, wherein the flight path of the UAV is optimized using the following path planning algorithm: Path i =A*(Environment,Task i ) Among them, Path i is the flight path of the ith UAV, A * It is the A-star algorithm, which is used to calculate the optimal flight path. Environment represents the current environment information. Task i represents the task performed by the i-th UAV; The resource consumption constraint is used to optimize the task allocation of the UAV to improve the resource utilization efficiency. The resource utilization efficiency is expressed by the following formula: Among them, N is the number of drones, Task i The number of tasks performed for the i-th UAV, Resources i is the total amount of resources consumed by the i-th drone; The step of adjusting the first task allocation strategy and / or the second task allocation strategy in real time includes: The first task allocation strategy and / or the second task allocation strategy are adjusted in real time in at least one of the following ways: When a preset condition is triggered during the execution of a task by the first drone, the task of the first drone is assigned to the second drone, wherein the preset condition indicates that the execution of the task by the drone is blocked; Adjust task priorities based on the status of the drone, the drone’s resources, and the urgency of the task; Adjust the drone’s flight path and / or mission based on current environmental information.
7. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1 to 4 when executing a program stored in a memory.
Citation Information
Patent Citations
Hybrid graph attention reinforcement learning method for task migration of multi-unmanned aerial vehicle cluster
CN118966648A
Oil depot tank field inspection robot task allocation method and system based on Internet of Things
CN119417192A