Heterogeneous capacity constraint unmanned aerial vehicle path planning method and device based on deep reinforcement learning, and storage medium
By employing a path planning method based on deep reinforcement learning, and utilizing Markov decision processes and a two-stage decoder network, the efficiency and quality issues in heterogeneous UAV path planning are addressed. This approach achieves efficient and intelligent path optimization, adapting to complex constraints and the differences between heterogeneous UAVs.
Patent Information
- Application Number
- CN202511189814.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies suffer from low efficiency, poor solution quality, and difficulty in adapting to complex constraints when dealing with path planning for heterogeneous UAVs in large-scale, three-dimensional spaces. In particular, they lack systematic modeling when dealing with differences in payload capacity and flight performance of heterogeneous UAVs and when performing multi-objective optimization.
A path planning method based on deep reinforcement learning is adopted. By establishing a Markov decision process model, an encoder network based on an attention mechanism is used to extract the global dependencies between task nodes, and a two-stage decoder network is used to decouple the selection decisions of the UAV and task nodes. An end-to-end policy network is designed for path planning.
It significantly improves the efficiency and quality of path planning solutions, can flexibly adapt to dynamically changing real-world scenarios, achieves more accurate and robust path solutions, and optimizes the utilization of UAV resources.
Smart Images

Figure CN120970655A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the interdisciplinary fields of artificial intelligence and aerospace technology, and in particular to a heterogeneous capacity-constrained UAV path planning method, device and storage medium based on deep reinforcement learning. Background Technology
[0002] With the widespread application of drones in urban logistics, emergency response, and air traffic, how to efficiently schedule a fleet of multiple drones to complete a series of tasks in three-dimensional space has become a key technical challenge. This problem can be regarded as a complex variant of the classic Vehicle Routing Problem (VRP) in three-dimensional space, with heterogeneous fleets and multiple constraints.
[0003] Currently, methods for solving such path planning problems are mainly divided into traditional optimization algorithms and emerging deep learning methods. Traditional optimization algorithms can be further divided into exact algorithms and heuristic algorithms. Exact algorithms, represented by Mixed Integer Linear Programming (MILP), can theoretically guarantee finding the optimal solution, but their computational complexity increases exponentially with the problem size (such as the number of task points or drones), resulting in excessively long solution times and making them difficult to apply to real-world scenarios.
[0004] Heuristic and metaheuristic algorithms, such as Ant Colony Optimization (ACO), genetic algorithms, or Variable Neighborhood Search (VNS), while offering fast solution speeds, also have significant drawbacks. These algorithms largely rely on manually designed search rules, lacking adaptability and exhibiting insufficient generalization ability when facing dynamically changing and complex scenarios. Furthermore, they are prone to getting trapped in local optima, and the quality of their solutions is highly sensitive to parameter settings and initial solutions, failing to guarantee global optimization performance and stability.
[0005] In recent years, Deep Reinforcement Learning (DRL) has been explored for solving path optimization problems due to its powerful autonomous decision-making and policy generation capabilities. However, most existing research remains limited to simplified ideal scenarios, such as planning paths for homogeneous vehicles in a two-dimensional plane. These methods are insufficient in handling complex constraints in the real world and generally lack systematic modeling of three-dimensional spatial characteristics, UAV heterogeneity (such as different payload capacities and flight performance), and multi-objective optimization (such as achieving a balance between minimizing total time and balancing the workload of each UAV mission).
[0006] Therefore, existing technologies have significant shortcomings in terms of solution efficiency, solution quality, generalization ability, and adaptability to complex real-world constraints. There is an urgent need for a new method that can efficiently and intelligently solve the path planning problem of heterogeneous UAVs in three-dimensional space. Summary of the Invention
[0007] The purpose of this application is to provide a path planning method based on deep reinforcement learning to overcome the technical defects of existing path planning methods in dealing with the scheduling problem of heterogeneous UAVs in large-scale, three-dimensional space, such as low solution efficiency, poor solution quality, and difficulty in adapting to complex constraints.
[0008] According to one embodiment of this application, a heterogeneous capacity-constrained UAV path planning method based on deep reinforcement learning is proposed, comprising:
[0009] A Markov decision process model is established for planning the paths of multiple heterogeneous UAVs. The system state of the Markov decision process model includes the state information of multiple task nodes to be planned and the state information of multiple heterogeneous UAVs.
[0010] By using an encoder network based on an attention mechanism, feature extraction and fusion are performed on the state information of the multiple task nodes to generate task node embeddings that represent the global dependencies between task nodes.
[0011] In each decision step of the path planning, a two-stage decoder network is used to perform the following decisions: using the UAV selection decoder in the two-stage decoder network, the current state information of the multiple heterogeneous UAVs and historical paths are used to select the current execution UAV; using the node selection decoder in the two-stage decoder network, the target task node is selected for the execution UAV from all unserved task nodes based on the current state information of the selected execution UAV and the task node embedding.
[0012] After each decision step, the system state in the Markov decision process model is updated according to the selected execution drone and target task node, and the selected target task node is marked as served.
[0013] After all mission nodes have been served, a complete flight path scheme covering all mission nodes is generated.
[0014] In some implementations, the UAV's status information includes the UAV's three-dimensional spatial information, remaining payload capacity, and cumulative flight time since the start of the mission;
[0015] The status information of a task node includes its three-dimensional spatial information, task requirements, and service status.
[0016] In some embodiments, before performing feature extraction and fusion of the state information of the plurality of task nodes through the encoder network, the method further includes:
[0017] The requirements of each task node are normalized according to the payload capacity of each of the multiple heterogeneous UAVs to generate a set of enhanced node features that characterize the relative load level of the task node for different UAVs. The enhanced node features are used for feature extraction and fusion to generate the task node embedding.
[0018] In some implementations, the attention-based encoder network is a Transformer encoder structure. The encoder network calculates and fuses the global dependencies between task nodes through a multi-head self-attention sublayer, and performs nonlinear feature transformation on the features processed by the multi-head self-attention sublayer through a feedforward network sublayer to generate the task node embedding.
[0019] In some implementations, a drone selection decoder in a two-stage decoder network is used to select the current executing drone based on the task node embedding, the current state information of the multiple heterogeneous drones, and their historical paths, including:
[0020] All current state information of the drones is processed through the first feedforward neural network to generate drone state feature embeddings;
[0021] The embedded sequence of each task node visited by the drone is aggregated through pooling operations to generate path context feature embeddings that characterize the drone's historical paths.
[0022] The UAV state feature embedding and the path context feature embedding are concatenated to form a combined feature, and the combined feature is input into a second feedforward neural network for nonlinear transformation processing;
[0023] The processing result of the second feedforward neural network is converted into the first probability distribution;
[0024] The drone to execute the current decision step is selected based on the first probability distribution.
[0025] In some implementations, a node selection decoder is used in a two-stage decoder network to select a target task node for the executing drone from all unserved task nodes based on the current state information of the selected executing drone and the task node embedding, including:
[0026] The global embedding of the integrated graph structure, the embedding of the task node visited by the UAV in the last time, and the remaining payload capacity of the UAV are used to construct a contextual information vector to characterize the current decision-making situation.
[0027] Through a multi-head attention mechanism, the constructed context information vector is interactively processed with the embedding of each task node to generate a context representation with enhanced perception of the features of each task node.
[0028] A scaled bilinear attention mechanism is used to calculate the compatibility score between the generated context representation and the embeddings of each task node;
[0029] Mask the compatibility scores corresponding to the serviced task nodes;
[0030] The compatibility score after masking is converted into a second probability distribution;
[0031] The final target task node is determined based on the second probability distribution.
[0032] In some embodiments, the method further includes:
[0033] The policy network, consisting of the encoder network and the two-stage decoder network, is trained end-to-end using a policy gradient algorithm with baseline.
[0034] In some implementations, the objective function of the policy gradient algorithm is to minimize the maximum flight time for the multiple heterogeneous UAVs to complete all tasks, or to minimize the total flight time of the multiple heterogeneous UAVs.
[0035] In some implementations, the baseline is implemented through a separate baseline network, wherein:
[0036] The baseline network selects the action with the highest probability at each decision step to construct a reference path;
[0037] The policy gradient algorithm uses the performance of the reference path as a benchmark to evaluate the current decision quality of the policy network, and updates the parameters of the policy network accordingly.
[0038] According to one embodiment of this application, an electronic device is provided, the device including a memory and a processor, the memory being used to store computer instructions executable on the processor, and the processor being used to implement the method as described in any of the preceding claims when executing the computer instructions.
[0039] According to one embodiment of this application, a computer-readable storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the method as described in any of the preceding claims.
[0040] The proposed heterogeneous capacity-constrained UAV path planning scheme based on deep reinforcement learning overcomes the shortcomings of traditional heuristic algorithms, such as reliance on manual rules and poor generalization ability, by systematically modeling the complex path planning problem as a Markov decision process and solving it end-to-end. This allows the model to learn universal decision strategies and flexibly adapt to dynamically changing real-world scenarios. By employing an attention-based encoder network, the scheme effectively captures the global dependencies between all task nodes, giving the model a global perspective during decision-making. Compared to the local search of traditional algorithms, this significantly improves the global quality of the path planning scheme. A unique two-stage decoder network decouples the complex decisions of selecting the UAV and selecting task nodes, significantly reducing the decision difficulty at each step and generating more accurate, robust, and high-quality path plans. Furthermore, this scheme avoids the exponential computational complexity of exact algorithms, demonstrating strong solution efficiency and scalability when dealing with large-scale problems.
[0041] Furthermore, in various embodiments of this application, the task requirements are normalized according to the payload capacity of each UAV before feature extraction, enabling the model to inherently perceive and quantify the capability differences between UAVs, thereby achieving a more balanced and reasonable load allocation. In the dual-stage decoder network, the UAV selection decoder, by fusing the current state information of the UAV with the contextual features of historical paths, can more comprehensively evaluate the overall state of each UAV, thereby making more strategic scheduling decisions and optimizing the resource utilization of the entire fleet. The node selection decoder, by constructing a context vector containing global and local information and using a multi-head attention mechanism for enhancement processing, makes the node selection decision highly consistent with the current situation, significantly improving the accuracy of each step of the path construction. In terms of the training mechanism, an independent baseline network is used for reference, providing a stable and high-quality benchmark for the parameter update of the policy network, effectively accelerating model convergence and improving the performance of the final policy.
[0042] Other features and advantages of the technical solution proposed in this application are described below. Attached Figure Description
[0043] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this specification and, together with the description, serve to explain the principles of this specification.
[0044] Figure 1 This is a flowchart illustrating a heterogeneous capacity-constrained UAV path planning method based on deep reinforcement learning according to an embodiment of this application.
[0045] Figure 2 This is a schematic diagram of the overall framework of the reinforcement learning policy network used in an exemplary embodiment of this application.
[0046] Figure 3 This is a detailed structural diagram of the policy network in an exemplary embodiment of this application.
[0047] Figure 4 This is a detailed structural diagram of the encoder network in an exemplary embodiment of this application.
[0048] Figure 5 This is a detailed structural diagram of a drone selection decoder module in an exemplary embodiment of this application.
[0049] Figure 6 This is a detailed structural diagram of the node selection decoder module in an exemplary embodiment of this application.
[0050] Figure 7 This is a schematic diagram of the structure of an electronic device shown in at least one embodiment of this application. Detailed Implementation
[0051] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0052] The core concept of this application is to provide an end-to-end intelligent decision-making method based on deep reinforcement learning for the complex heterogeneous capacity-constrained UAV path planning problem in three-dimensional space. This application systematically models the path planning problem as a Markov decision process and designs a policy network based on an attention mechanism as the solver. This policy network employs a unique encoder-dual decoder structure. The encoder captures the global dependencies between all nodes in the task environment, and the dual decoder structure decouples the complex composite decision of choosing which UAV to select from which task node to choose into two ordered and more manageable sub-decisions. In this way, the model can learn efficient decision-making strategies, thereby quickly and intelligently generating high-quality path planning schemes without the need for manual rule design.
[0053] To better understand this application, the following will refer to... Figure 1 The process shown in the figure, together with other accompanying drawings, will be used to elaborate on the technical solution of this application.
[0054] Reference Figure 1 , Figure 1This is a flowchart illustrating a heterogeneous capacity-constrained UAV path planning method based on deep reinforcement learning according to an embodiment of this application. The method includes the following steps S101 to S105.
[0055] Step S101: Establish a Markov decision process model for planning the paths of multiple heterogeneous UAVs. The system state of the Markov decision process model includes the state information of multiple task nodes to be planned and the state information of multiple heterogeneous UAVs.
[0056] In this step, the Heterogeneous Capacity-Constrained Unmanned Aerial Vehicle Path Planning Problem (HCUAVRP) is formally constructed as a Markov Decision Process (MDP) model. Heterogeneous drones refer to a fleet composed of multiple drones with different performance parameters, such as payload capacity, endurance, or flight speed. A Markov Decision Process is a standard mathematical framework in reinforcement learning used to describe sequential decision problems, typically consisting of a quadruple of state space (S), action space (A), state transition rules (t), and reward function (r). By establishing this model, the complex path planning task can be transformed into a process where an agent (i.e., a decision model) learns to make optimal actions in a series of states to maximize cumulative reward.
[0057] In some implementations, the status information of the UAV may include the UAV's three-dimensional spatial information, remaining payload capacity, and cumulative flight time since the start of the mission; the status information of the mission node may include the mission node's three-dimensional spatial information, mission requirements, and service status.
[0058] The drone's status information is primarily acquired and updated through its onboard sensors and flight control system. Three-dimensional spatial information can be obtained in real-time via GPS or other positioning modules. Remaining payload capacity can be calculated by subtracting the required amount of completed tasks from the initial payload, or directly measured by payload sensors. Cumulative flight time is recorded by an onboard timer or flight log. Task node status information can originate from a higher-level task scheduling or order management system. This system provides a list of all pending customer locations (e.g., all customer delivery addresses), including precise three-dimensional geographic coordinates (longitude, latitude, altitude) and specific task requirements (e.g., weight, volume, or quantity of delivered packages). All task nodes are initially set to "not serviced."
[0059] When building a Markov Decision Process (MDP) model based on the acquired initial data, the state space (S) of the model can be defined as all possible combinations of UAV states and task node states. Then, the decision-making action of assigning a specific unserved task node to a particular UAV is defined as the action space (A). Next, based on the UAV's flight performance and the physical rules of task execution, the state transition rule (t) is explicitly defined as how the system evolves to the next state S' after performing action A in state S. Through this construction process, the original, unstructured task data can be successfully transformed into a standardized MDP model that reinforcement learning algorithms can understand and process, laying the foundation for subsequent intelligent solutions.
[0060] Specifically, at any given decision point in time, the system state of the Markov Decision Process (MDP) model can completely record the current 3D coordinates of each drone, the remaining cargo capacity, the flight time, and the 3D coordinates, task requirements, and service status of each task node. The action space of this MDP model can be defined as selecting a drone and assigning it an unserved task node in the current state. The state transition rules define how the system state is updated after performing an action, such as updating the drone's position to the target task node's position and reducing its remaining capacity accordingly. The reward function can be designed based on the optimization objective; for example, in a scenario where the objective is to minimize the longest flight time among all drones, the reward can be set as a negative value of that maximum flight time.
[0061] Step S102: Through an encoder network based on an attention mechanism, feature extraction and fusion are performed on the state information of the multiple task nodes to generate task node embeddings that represent the global dependencies between task nodes.
[0062] According to this embodiment, after establishing the Markov decision process model, an attention-based encoder network is used to understand the global information of the current task. This encoder network receives the state information (such as 3D coordinates, requirements, etc.) of all task nodes as input, and through internal calculations, outputs a high-dimensional feature vector for each task node, i.e., the task node embedding. The task node embedding can include not only the node's own attributes, but also the complex spatial relationships and dependencies between the node and all other nodes, providing global context information for subsequent decision-making.
[0063] To better address the challenges posed by heterogeneous drones, some embodiments of the method may include a preprocessing step before feature extraction: normalizing the demand of each task node based on the payload capacity of each of the multiple heterogeneous drones to generate a set of enhanced node features characterizing the relative load level of the task node with respect to different drones. These enhanced node features are then used for feature extraction and fusion to generate the task node embedding. This embodiment ensures that the features input to the encoder network inherently contain heterogeneous information. For example, for the same task point with a demand of 10, the relative load differs between a drone with a payload capacity of 20 and a drone with a capacity of 50. Through the normalization process of this embodiment, the model can perceive this difference from the outset, thereby providing a more refined basis for subsequent decision-making.
[0064] In some implementations, the attention-based encoder network is a Transformer encoder structure. This encoder network calculates and fuses global dependencies between task nodes through a multi-head self-attention sublayer, and performs a nonlinear feature transformation on the features processed by the multi-head self-attention sublayer through a feedforward network sublayer to generate the task node embeddings. (Refer to...) Figure 4 The encoder network comprises an input projection layer, multiple stacked attention modules, and a final output layer. Each attention module contains a multi-head self-attention (MHA) sublayer and a feedforward (FF) sublayer. The MHA mechanism allows the model to simultaneously focus on the relationships between nodes from different perspectives, thereby capturing richer global information; the feedforward network further processes the extracted features non-linearly, enhancing the model's expressive power. Furthermore, residual connections and layer normalization techniques can be used between modules to ensure the stability and efficiency of network training.
[0065] Step S103: In each decision step of path planning, a two-stage decoder network is used to select the current executing drone and the target task node.
[0066] According to this embodiment, after the encoder network generates a global understanding of the entire task (i.e., task node embedding), the decoder network is responsible for generating specific actions at each decision step based on the current state and global information. This application proposes a unique two-stage decoder network, referring to... Figure 3 The network includes a drone selection decoder and a node selection decoder to break down the complex decision of assigning tasks to whom into two simpler, ordered decisions: first, determining who performs the task, and then determining what task to perform. This greatly improves the accuracy and efficiency of decision-making.
[0067] First, in the first stage of the decision-making process, the UAV selection decoder in the two-stage decoder network is used to select the current execution UAV based on the task node embedding, the current state information of the multiple heterogeneous UAVs, and their historical paths.
[0068] In some implementations, this stage may include: processing the current state information of all UAVs through a first feedforward neural network to generate UAV state feature embeddings; aggregating the embedded sequences of task nodes visited by each UAV through a pooling operation to generate path context feature embeddings representing the historical paths of the UAV; concatenating the UAV state feature embeddings with the path context feature embeddings to form a combined feature, and inputting the combined feature into a second feedforward neural network for nonlinear transformation processing; converting the processing result of the second feedforward neural network into a first probability distribution; and selecting the UAV to execute the current decision step based on the first probability distribution.
[0069] Reference Figure 5 This drone selection decoder comprehensively considers multiple aspects of information to make decisions. It can analyze the current real-time state of all drones (such as location, remaining capacity, etc.), as well as the historical path information that each drone has already traversed (such as the embedding of visited task nodes). It can also use the embedding of all task nodes to obtain a global understanding of the entire task environment. By fusing the drone's own state, historical trajectory, and perception of the overall task situation, the decoder can comprehensively evaluate which drone is the best choice to execute the next task, and finally output a probability distribution representing the probability of selecting each drone, i.e., the first probability distribution.
[0070] After obtaining the first probability distribution of the selected drone, the specific drone can be determined based on this distribution. In some implementations, the drone corresponding to the highest probability value in the first probability distribution can be selected as the executing drone; in other implementations, the first probability distribution can be used as a sampling basis to determine the executing drone through random sampling. The former is often referred to as a greedy strategy, characterized by always selecting the option with the highest current probability, resulting in fast decision-making and making it more suitable for scenarios with high computational time requirements; the latter is called a sampling strategy, which explores a wider solution space by introducing a certain degree of randomness, thus often obtaining higher quality and more diverse path solutions. In practical applications, a suitable decoding strategy can be selected based on different emphases on solution speed and optimality of the solution.
[0071] Next, in the second stage of the decision-making process, the node selection decoder in the two-stage decoder network is used to select a target task node for the execution drone from all unserved task nodes based on the current state information of the selected execution drone and the task node embedding.
[0072] In some implementations, this stage may include: constructing a context information vector representing the current decision context by integrating the global embedding of the graph structure, the embedding of the task nodes previously visited by the executing UAV, and the remaining payload capacity of the executing UAV. The global embedding of the graph structure is a single feature vector obtained by aggregating all task node embeddings output by the encoder network, which can serve as a holistic representation of the macroscopic characteristics of the entire task graph; interactively processing the constructed context information vector with the embeddings of each task node through a multi-head attention mechanism to generate a context representation with enhanced perception of the features of each task node; calculating the compatibility score between the generated context representation and the embeddings of each task node using a scaled bilinear attention mechanism; masking the compatibility scores corresponding to the served task nodes; converting the masked compatibility scores into a second probability distribution; and determining the final target task node based on the second probability distribution.
[0073] Reference Figure 6 After identifying the drone to perform the task, the decoder constructs a customized decision context vector, which may contain global information and the drone's specific state (where it was last located, how much capacity remains), etc. Then, the decoder uses an attention mechanism to calculate the matching degree (compatibility score) between this decision context and all candidate task nodes. To ensure the legitimacy of the path, the decoder uses a masking mechanism to invalidate the scores of nodes that have already been served, thus preventing duplicate visits. Finally, the decoder outputs the probability distribution used to select each unserved node, i.e., the second probability distribution.
[0074] Similarly, after obtaining the second probability distribution of the selected node, in some embodiments, the task node corresponding to the highest probability value in the second probability distribution can be selected as the target task node; in other embodiments, the second probability distribution can be used as a sampling basis to determine the target task node by random sampling.
[0075] This embodiment proposes to use a policy gradient algorithm with baseline to train the policy network consisting of the encoder network and the two-stage decoder network end-to-end, so as to further improve the decision quality of the encoder network and the two-stage decoder network.
[0076] Policy gradient algorithms can be used in reinforcement learning. They continuously try different paths (generating path schemes) and adjust network parameters based on the results (rewards), increasing the probability that the policy network will generate better policies. This implementation uses a policy gradient algorithm with a baseline to further optimize the training process. By introducing an evaluation criterion (baseline), parameter updates are performed more stably and efficiently.
[0077] In some implementations, the baseline is implemented through an independent baseline network, wherein: the baseline network selects the action with the highest probability at each decision step to construct a reference path; and the policy gradient algorithm uses the performance of the reference path as a benchmark to evaluate the current decision quality of the policy network in order to update the parameters of the policy network.
[0078] According to this implementation, a baseline network that always employs a greedy strategy can be maintained. This ensures that during training, the policy network is only positively incentivized to update if its generated solution is better than that of the baseline network. This mechanism significantly improves training stability and efficiency. During training, Monte Carlo methods can be used for parameter updates, and statistical methods such as paired t-tests can be used to evaluate the performance difference between the policy network and the baseline network to determine when to update the baseline network's parameters.
[0079] The training objective can be set according to actual needs. In some implementations, the optimization objective function of the policy gradient algorithm is to minimize the maximum flight time of the multiple heterogeneous UAVs to complete all tasks, or to minimize the total flight time of the multiple heterogeneous UAVs. The former (Min-Max) focuses on the overall task completion efficiency, ensuring that no UAV becomes a bottleneck; the latter (Min-Sum) focuses on the total system cost, such as total energy consumption.
[0080] Step S104: After each decision step, update the system state in the Markov decision process model according to the selected execution drone and target task node, and mark the selected target task node as served.
[0081] This step enables state transitions in a Markov decision process model. After completing a decision (i.e., determining that UAV i will fly to node j), the system state can be updated to prepare for the next decision. Updates may include: updating the 3D spatial information of UAV i to the coordinates of the target task node j; updating the remaining payload capacity of UAV i according to the requirements of the target task node j; updating the cumulative flight time of UAV i; and marking the service status of node j as "serviced".
[0082] Step S105: After all task nodes have been served, generate a complete flight path scheme covering all task nodes.
[0083] The decision-making process proposed in this embodiment is an iterative loop. The system can repeatedly execute S103 (decision) and S104 (state update) until the service status of all task nodes is marked as served. When this termination condition is met, the entire decision-making process ends. All decision sequences recorded during this process (which UAV visited which node, and the order of visits) can be combined to obtain a complete flight path plan covering all task nodes and planned for the entire UAV fleet.
[0084] Reference Figure 2 This demonstrates the overall framework of a reinforcement learning policy network employed in an exemplary embodiment of this application. The framework takes a specific problem instance as input, and its internal encoder and two-stage decoder (including a drone selection decoder and a node selection decoder) can adjust according to the current system state (S). t ), to make sequential decisions to generate specific actions (a t This action serves two purposes: firstly, it gradually constructs and expands intermediate solutions; secondly, it changes the current system state (S) according to the state transition rules. t Update the state to the next moment, and repeat this process until the final complete path is generated.
[0085] To further clarify the internal working mechanism of this plan, the following will refer to... Figures 3 to 6 The network architecture is shown, and a simplified end-to-end example is used to exemplify the data processing flow of a specific application example according to this application.
[0086] Assume the mission scenario is as follows: A fleet consists of two heterogeneous UAVs (UAV-A, payload 50; UAV-B, payload 30). UAV-A and UAV-B need to start from the same origin (warehouse) to provide services to four mission nodes (Node-1 to Node-4). Each node has its own three-dimensional coordinates and demand.
[0087] At the start of path planning, the state information (3D coordinates, requirements, etc.) of all task nodes and the heterogeneous information of the UAV (payload capacity) are first input into the policy network. (Refer to...) Figure 3 and Figure 4The data first enters the encoder network. The encoder network performs normalized preprocessing on the raw features (especially demand) of each task node to reflect the capacity differences between UAVs. Subsequently, the enhanced features undergo deep feature extraction through a multi-layered stacked attention module. In each module, a multi-head self-attention mechanism captures the complex spatial dependencies between all task nodes, while a feedforward network performs nonlinear transformations. Finally, the encoder outputs two sets of key information: one set is the task node embedding generated for each task node, containing global relational information; the other set is the global embedding, representing the overall macro-level situation of the task, obtained by aggregating (e.g., averaging) all node embeddings.
[0088] Next, the system proceeds to the first decision-making step. (Refer to...) Figure 3 and Figure 5 The system first selects a drone to perform the task; this decision is made by the drone selection decoder. This decoder receives and processes two inputs: one is the current state of the drones (currently, both UAV-A and UAV-B are in the warehouse with full payloads); the other is their historical path context (currently, both are empty paths). The drone selection decoder fuses these two pieces of information and combines them with global context information from the encoder, ultimately outputting a first probability distribution through a function such as softmax, for example, {UAV-A: 0.7, UAV-B: 0.3}. Based on this probability, the system selects UAV-A as the current drone to perform the task.
[0089] After identifying the drone as a UAV-A, the system can select a target mission node for it. (Refer to...) Figure 3 and Figure 6 This decision is made by the node selection decoder. The node selection decoder first constructs a contextualized contextual information vector, which integrates the global embedding of the graph structure, the current state of UAV-A (location in the warehouse, remaining payload 50), and other information. Then, this contextual vector undergoes multi-head attention interaction and compatibility calculation with the embeddings of all four candidate task nodes to obtain corresponding matching scores. Since all nodes are unserved, the masking mechanism is ineffective at this stage. Finally, the scores are converted into a probability distribution (i.e., a second probability distribution) using, for example, a softmax function, such as {Node-1: 0.2, Node-2: 0.5, Node-3: 0.2, Node-4: 0.1}. Based on this probability, the system selects Node-2 as the first target task node for UAV-A.
[0090] At this point, the first decision-making step is complete, and the generated path segment is "warehouse -> Node-2", with UAV-A as the executor. Subsequently, the system status is updated: UAV-A's location becomes Node-2, its remaining payload is reduced by Node-2's demand, its cumulative flight time increases, and Node-2's service status is marked as "serviced".
[0091] The system then proceeds to the second decision-making step and repeats the above process. When selecting a UAV, the UAV selection decoder receives the updated system status (UAV-A is at Node-2, UAV-B is still in the warehouse). When selecting the next node for the selected UAV, the node selection decoder receives the updated context and, after calculating the compatibility score, invalidates the score of Node-2 through a masking mechanism to ensure it is not selected repeatedly. This loop continues until all four task nodes are marked as serviced, at which point the entire path planning process ends, and the final output is the complete flight path sequence for UAV-A and UAV-B respectively.
[0092] In summary, this embodiment provides a complete method for heterogeneous capacity-constrained UAV path planning based on deep reinforcement learning. This method models the problem as a Markov decision process and solves it using an encoder-dual-decoder policy network based on an attention mechanism. Its core lies in the dual-decoder design, which decomposes complex composite decisions into two ordered sub-problems, thereby effectively reducing the difficulty of solving the problem and improving the quality of decision-making. The embodiments disclosed in this application, from problem modeling and network structure design to model training and decoding, constitute a systematic end-to-end solution capable of efficiently and intelligently generating high-quality path schemes for multi-UAV cooperative tasks in complex real-world scenarios.
[0093] Those skilled in the art will understand that the specific network structures and algorithms described in this application are merely illustrative and not restrictive. For example, some components in the encoder and decoder can be replaced with other neural networks capable of sequence processing, such as LSTM or GRU; the attention mechanism can also employ other variations known in the art; and the training process can also employ other policy gradient optimization algorithms such as REINFORCE and PPO. These modifications, as long as they substantially utilize the core concepts of this application (e.g., modeling the problem as an MDP and employing an encoder-dual decoder architecture for sequential decision-making), should fall within the protection scope of this application.
[0094] Those skilled in the art will understand that any embodiment of this application can be provided as a method, system, or computer program product. Therefore, this application can take the form of entirely hardware, entirely software, or a combination of hardware and software. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.
[0095] Figure 7 This is a schematic diagram of an electronic device provided according to one embodiment of this application. The electronic device may be a server, a personal computer, a mobile terminal, or other device with data processing capabilities. As shown, the device includes a processor and a memory. The memory stores computer instructions that can run on the processor. The processor, when executing the computer instructions, implements the heterogeneous capacity-constrained UAV path planning method based on deep reinforcement learning described in any embodiment of this application.
[0096] One embodiment of this application also provides a computer-readable storage medium on which a computer program is stored. This computer-readable storage medium includes all forms of non-volatile memory, media, and devices, such as semiconductor storage devices (e.g., EPROM, flash memory devices), magnetic disks, magneto-optical disks, and CD-ROMs and DVD-ROMs. When the program is executed by a processor, it can implement the heterogeneous capacity-constrained UAV path planning method based on deep reinforcement learning described in any embodiment of this application.
[0097] The processing and logic flow described in this specification can be implemented by one or more programmable processors executing one or more computer programs, or by special-purpose logic circuits (such as FPGAs or ASICs). The processor used to execute the computer program can be a general-purpose or special-purpose microprocessor. Typically, the processor receives instructions and data from read-only memory or random access memory. The basic components of a computer include a processor for executing instructions and one or more storage devices for storing instructions and data.
[0098] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of the invention. Certain features described in multiple embodiments may also be implemented in combination in a single embodiment; conversely, various features described in a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Similarly, although the operational flows in the accompanying drawings are depicted in a specific order, this does not require that they be performed in the indicated order or serially to achieve the desired result; in some cases, multitasking or parallel processing may be equally feasible or more advantageous.
[0099] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A heterogeneous capacity-constrained UAV path planning method based on deep reinforcement learning, characterized in that, include: A Markov decision process model is established for planning the paths of multiple heterogeneous UAVs. The system state of the Markov decision process model includes the state information of multiple task nodes to be planned and the state information of multiple heterogeneous UAVs. By using an encoder network based on an attention mechanism, feature extraction and fusion are performed on the state information of the multiple task nodes to generate task node embeddings that represent the global dependencies between task nodes. In each decision step of path planning, a two-stage decoder network is used to perform the following decision: using the UAV selection decoder in the two-stage decoder network, the current execution UAV is selected based on the task node embedding, the current state information of the multiple heterogeneous UAVs and the historical path; The node selection decoder in the two-stage decoder network selects a target task node for the execution drone from all unserved task nodes based on the current state information of the selected execution drone and the task node embedding. After each decision step, the system state in the Markov decision process model is updated according to the selected execution drone and target task node, and the selected target task node is marked as served. After all mission nodes have been served, a complete flight path scheme covering all mission nodes is generated.
2. The method according to claim 1, characterized in that: The status information of the UAV includes the UAV's three-dimensional spatial information, remaining payload capacity, and cumulative flight time since the start of the mission; The status information of a task node includes its three-dimensional spatial information, task requirements, and service status.
3. The method according to claim 1, characterized in that, Before performing feature extraction and fusion of the state information of the multiple task nodes through the encoder network, the method further includes: The requirements of each task node are normalized according to the payload capacity of each of the multiple heterogeneous UAVs to generate a set of enhanced node features that characterize the relative load level of the task node for different UAVs. The enhanced node features are used for feature extraction and fusion to generate the task node embedding.
4. The method according to claim 1, characterized in that, The attention-based encoder network is a Transformer encoder structure. The encoder network calculates and fuses the global dependencies between task nodes through a multi-head self-attention sublayer, and performs nonlinear feature transformation on the features processed by the multi-head self-attention sublayer through a feedforward network sublayer to generate the task node embedding.
5. The method according to claim 1, characterized in that, Using the drone selection decoder in a two-stage decoder network, based on the task node embedding, the current state information of the multiple heterogeneous drones, and their historical paths, the current executing drone is selected, including: All current state information of the drones is processed through the first feedforward neural network to generate drone state feature embeddings; The embedded sequence of each task node visited by the drone is aggregated through pooling operations to generate path context feature embeddings that characterize the drone's historical paths. The UAV state feature embedding and the path context feature embedding are concatenated to form a combined feature, and the combined feature is input into a second feedforward neural network for nonlinear transformation processing; The processing result of the second feedforward neural network is converted into the first probability distribution; The drone to execute the current decision step is selected based on the first probability distribution.
6. The method according to claim 1, characterized in that, Utilizing a node selection decoder in a two-stage decoder network, based on the current state information of the selected execution drone and the task node embedding, a target task node is selected for the execution drone from all unserved task nodes, including: The global embedding of the integrated graph structure, the embedding of the task node visited by the execution drone in the last time, and the remaining payload capacity of the execution drone are used to construct a contextual information vector to characterize the current decision-making situation. Through a multi-head attention mechanism, the constructed context information vector is interactively processed with the embedding of each task node to generate a context representation with enhanced perception of the features of each task node. A scaled bilinear attention mechanism is used to calculate the compatibility score between the generated context representation and the embeddings of each task node; Mask the compatibility scores corresponding to the serviced task nodes; The compatibility score after masking is converted into a second probability distribution; The final target task node is determined based on the second probability distribution.
7. The method according to claim 1, characterized in that, The method further includes: The policy network, consisting of the encoder network and the two-stage decoder network, is trained end-to-end using a policy gradient algorithm with baseline.
8. The method according to claim 7, characterized in that, The objective function of the policy gradient algorithm is to minimize the maximum flight time for the multiple heterogeneous UAVs to complete all tasks, or to minimize the total flight time of the multiple heterogeneous UAVs.
9. The method according to claim 7, characterized in that, The baseline is implemented through an independent baseline network, wherein: The baseline network selects the action with the highest probability at each decision step to construct a reference path; The policy gradient algorithm uses the performance of the reference path as a benchmark to evaluate the current decision quality of the policy network, and updates the parameters of the policy network accordingly.
10. An electronic device, characterized in that, The device includes a memory and a processor, the memory being used to store computer instructions executable on the processor, and the processor being used to implement the method of any one of claims 1 to 9 when executing the computer instructions.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the method described in any one of claims 1 to 9.
Citation Information
Cited By
Unmanned platform path planning method, device, equipment, medium and program product
CN121783172A
Unmanned platform path planning method, device, equipment, medium and program product
CN121783172B
Intelligent robot whole body motion training method based on residual motion learning
CN122113994A
Three-dimensional path planning method of fig seedling patrol unmanned aerial vehicle
CN122329338A
A 3D path planning method for a drone used for patrolling fig seedlings
CN122329338B