Path planning method for drone interruption scenarios based on DDQN with attention mechanism

By constructing a comprehensive objective function and distribution network graph through the DDQN network based on the attention mechanism, the applicability problem of drone path planning in dynamic environments is solved, and a high success rate and efficient path planning are achieved in complex interruption scenarios.

CN120576777BActive Publication Date: 2025-10-03HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511083215.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-10-03
Estimated Expiration
2045-08-04

AI Technical Summary

Technical Problem

Existing drone path planning methods have low applicability in real environments and are difficult to cope with the interference of multi-source dynamic environmental factors, resulting in flight interruptions.

Method used

The DDQN network based on the attention mechanism is used to construct a comprehensive objective function and a distribution network graph. The graph attention network and the multi-layer perceptron are combined for feature extraction. The interruption features and the spatial topological relationship are dynamically weighted through the adaptive attention mechanism to form the optimal path planning strategy.

Benefits of technology

It significantly improves the success rate and adaptability of drone path planning in complex interruption scenarios and multi-source dynamic environments, reduces interruption frequency, optimizes path length and energy consumption, and improves the robustness and efficiency of path planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120576777B_ABST
    Figure CN120576777B_ABST
Patent Text Reader

Abstract

This invention provides a method for drone path planning in interruption scenarios based on an attention-based DDQN (Double Quantization Network), which relates to the technical field of path planning. The method constructs a comprehensive objective function and a distribution network graph based on node location information, drone energy consumption per unit distance, and the comprehensive interruption risk between nodes. A graph attention network and a multilayer perceptron are used to extract features from the distribution network graph and the current state, respectively, to obtain current graph embedding features and current state features. These extracted features are then concatenated and input into a main network to obtain a predicted Q value, reward value, and optimal action. The optimal action is sent to the drone, and the next state is fed back. The main network is then recycled to continuously obtain the optimal action, ultimately forming an optimal path planning strategy. The adaptive attention mechanism dynamically weights the interruption features and spatial topological relationships, significantly improving the drone's path planning success rate and adaptability under interference from multiple dynamic environmental factors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of path planning technology, and in particular to a DDQN (Digital Queries Network)-based attention mechanism path planning method for unmanned aerial vehicle (UAV) interruption scenarios. Background Art

[0002] Existing path planning methods often use Dijkstra algorithm and The Dijkstra algorithm constructs a graph model, starting from a starting point and finding the shortest path to each node one by one, gradually expanding until it reaches the end point. Its core is to find the shortest path from the source node to each node in increasing order of path length. The algorithm introduces a heuristic function based on the Dijkstra algorithm, and guides the search direction through the evaluation function to improve the search efficiency.

[0003] The above methods are all for path planning in static environments. In real environments, the flight of drones is also affected by the external dynamic environment, which may cause the drone flight to be interrupted. When in a changing dynamic environment, the drone flight path obtained by the existing path planning method is no longer applicable. Therefore, the existing drone path planning method has low applicability in real environments and is difficult to cope with the interference of multi-source dynamic environmental factors. Summary of the Invention

[0004] The problem to be solved by the present invention is that the existing UAV path planning method has low applicability in real environments and is difficult to cope with the interference of multi-source dynamic environmental factors.

[0005] To solve the above problems, in the first aspect, the present invention provides a drone interruption scenario path planning method based on DDQN with an attention mechanism, comprising:

[0006] Based on node location information, drone energy consumption per unit distance, and the comprehensive interruption risk between nodes, a comprehensive objective function and a distribution network diagram are constructed; the comprehensive interruption risk includes the risk of interruption due to bad weather, the risk of equipment failure, and the risk of communication interference;

[0007] A graph attention network is used to extract features from the distribution network graph, and a multi-layer perceptron is used to extract features from the current state of the drone, obtaining the current graph embedding features and current state features respectively.

[0008] The current graph embedding features and current state features are concatenated and input into the trained main network to obtain the current step predicted Q value, current reward value, and current optimal action; the DDQN network based on the attention mechanism includes the main network, and the current reward value is calculated based on the reward function constructed according to the comprehensive objective function;

[0009] Send the current optimal action to the drone and receive the next state feedback from the drone;

[0010] The distribution network diagram is updated based on the current optimal action and the next state, and the optimal action for each subsequent step is continuously obtained using the main network, ultimately forming the optimal path planning strategy.

[0011] In a second aspect, the present invention also provides a drone interruption scenario path planning system based on DDQN with an attention mechanism, comprising:

[0012] A construction module is used to construct a comprehensive objective function and a distribution network diagram based on node location information, drone energy consumption per unit distance, and the comprehensive interruption risk between nodes; the comprehensive interruption risk includes the interruption risk of bad weather, the risk of equipment failure, and the risk of communication interference;

[0013] The extraction module is used to extract features from the distribution network graph using a graph attention network and a multi-layer perceptron to extract features from the current state of the drone, obtaining the current graph embedding features and current state features respectively;

[0014] The optimization module is used to splice the current graph embedding features and the current state features, and input them into the trained main network to obtain the current step predicted Q value, the current reward value and the current optimal action; the DDQN network based on the attention mechanism includes the main network, and the current reward value is calculated based on the reward function constructed according to the comprehensive objective function;

[0015] The interaction module is used to send the current optimal action to the drone and receive the next state feedback from the drone;

[0016] The update loop module is used to update the distribution network diagram based on the current optimal action and the next state, and use the main network to continuously obtain the optimal action for each subsequent step, and finally form the optimal path planning strategy.

[0017] In a third aspect, the present invention provides an electronic device comprising a memory and a processor;

[0018] The memory is used to store computer programs;

[0019] The processor is used to implement the drone interruption scenario path planning method based on DDQN of the attention mechanism as described in the first aspect when executing the computer program.

[0020] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the drone interruption scenario path planning method based on DDQN of the attention mechanism as described in the first aspect is implemented.

[0021] This paper provides a DDQN-based attention mechanism-based path planning method for drone interruption scenarios. Compared with the existing technology, it has the following advantages:

[0022] Based on the node location information, the energy consumption per unit distance of the drone and the comprehensive interruption risk between nodes, a comprehensive objective function and a distribution network graph are constructed. Among them, the comprehensive interruption risk includes the interruption risk of bad weather, the risk of equipment failure and the risk of communication interference. In this way, the impact of the external dynamic environment is quantified and taken into consideration in the objective function. A graph attention network is used to extract features from the distribution network graph, focusing on the spatial relationship between nodes and edges in the distribution network to improve the global rationality of path planning. A multi-layer perceptron is used to extract features from the current state of the drone, and the current graph embedding features and current state features are obtained respectively. Then, the current graph embedding features and current state features are spliced ​​to enhance the environmental representation ability and input into the training set. In the main network of the DDQN network based on the attention mechanism, the predicted Q value of the current step, the current reward value and the current optimal action are obtained; the current reward value is the value calculated by the reward function constructed according to the comprehensive objective function, which is used to guide the DDQN network to autonomously learn in a good direction; the current optimal action is sent to the drone, and the next state feedback from the drone is received. The distribution network diagram is updated according to the current optimal action and the next state, and the main network is used to continuously obtain the optimal action for each subsequent step, and finally the optimal path planning strategy is formed. The interruption characteristics and spatial topological relationships are dynamically weighted by the adaptive attention mechanism, which significantly improves the path planning success rate and adaptability of the drone under complex interruption scenarios and interference from multi-source dynamic environmental factors. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0024] Figure 1 A flowchart of a DDQN-based drone interruption scenario path planning method based on an attention mechanism provided by an embodiment of the present invention;

[0025] Figure 2 A schematic diagram of a drone delivery network provided by an embodiment of the present invention;

[0026] Figure 3 A schematic diagram of a network structure in a DDQN network based on an attention mechanism provided by an embodiment of the present invention;

[0027] Figure 4Schematic diagram of information interaction for a DDQN-based drone interruption scenario path planning method based on an attention mechanism according to an embodiment of the present invention;

[0028] Figure 5 A schematic diagram of a curve showing rewards obtained by a drone in an interruption scenario provided by an embodiment of the present invention;

[0029] Figure 6 A schematic diagram of a success rate curve for drone path planning in an interruption scenario provided by an embodiment of the present invention;

[0030] Figure 7 Schematic diagram of reward curves for three different algorithms under the same environment provided by an embodiment of the present invention;

[0031] Figure 8 Schematic diagram of the success rate curves of three different algorithms provided by the embodiment of the present invention under the same environment;

[0032] Figure 9 Schematic diagram of the path curves obtained by simulating and planning the same scenario using the three algorithms provided in the embodiments of the present invention;

[0033] Figure 10 A schematic diagram of the structure of a drone interruption scenario path planning system based on DDQN with an attention mechanism provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0034] To make the objectives, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application are clearly and completely described. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0035] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0036] like Figure 1 As shown, the embodiment of the present application provides a DDQN-based drone interruption scenario path planning method based on the attention mechanism, including:

[0037] S1: Construct a comprehensive objective function and distribution network diagram based on node location information, drone energy consumption per unit distance, and comprehensive interruption risk between nodes; comprehensive interruption risk includes the interruption risk of bad weather, the risk of equipment failure, and the risk of communication interference.

[0038] S2: A graph attention network is used to extract features from the distribution network graph, and a multi-layer perceptron is used to extract features from the current state of the drone, obtaining current graph embedding features and current state features, respectively. The current state includes the drone's location, power level, and battery status, and the current distribution network graph includes node location information, task completion status, interruption status, and flight time.

[0039] S3: Concatenate the current graph embedding features and the current state features and input them into the trained main network to obtain the current step predicted Q value, current reward value, and current optimal action; the DDQN network based on the attention mechanism includes the main network, and the current reward value is calculated based on the reward function constructed according to the comprehensive objective function.

[0040] S4: Send the current optimal action to the drone and receive the next state feedback from the drone, obtain a tuple and save it in the playback cache; the tuple includes the current state, the current optimal action, the current reward value and the next state.

[0041] S5: Update the distribution network diagram based on the current optimal action and the next state, and use the main network to continuously obtain the optimal action for each subsequent step, and finally form the optimal path planning strategy.

[0042] In this optional embodiment, a comprehensive objective function and a distribution network graph are constructed based on node location information, energy consumption per unit distance of the drone, and comprehensive interruption risk between nodes. The comprehensive interruption risk includes the interruption risk due to bad weather, the risk of equipment failure, and the risk of communication interference. In this way, the impact of the external dynamic environment is quantified and taken into consideration in the objective function. A graph attention network is used to extract features from the distribution network graph, focusing on the spatial relationship between nodes and edges in the distribution network to improve the global rationality of path planning. A multi-layer perceptron is used to extract features from the current state of the drone to obtain the current graph embedding features and current state features respectively. The current graph embedding features and current state features are then concatenated to enhance the environmental representation capability and output. The algorithm is fed into the main network of the trained DDQN network based on the attention mechanism to obtain the predicted Q value of the current step, the current reward value and the current optimal action. The current reward value is calculated based on the reward function constructed according to the comprehensive objective function, which is used to guide the DDQN network to autonomously learn in a good direction. The current optimal action is sent to the drone, and the next state of the drone is fed back. The distribution network diagram is updated according to the current optimal action and the next state. The main network is used to continuously obtain the optimal action for each subsequent step, and finally the optimal path planning strategy is formed. The adaptive attention mechanism dynamically weights the interruption characteristics and spatial topological relationships, which significantly improves the path planning success rate and adaptability of the drone under complex interruption scenarios and interference from multi-source dynamic environmental factors.

[0043] Each step is described in detail below.

[0044] S1: Construct a comprehensive objective function and distribution network diagram based on node location information, drone energy consumption per unit distance, and comprehensive interruption risk between nodes; comprehensive interruption risk includes the interruption risk of bad weather, the risk of equipment failure, and the risk of communication interference.

[0045] Specifically, the drone-based delivery system utilizes rechargeable multi-rotor drones dispatched from a warehouse. The drones depart from the warehouse and deliver packages to designated demand points, recharging at charging stations as needed. After completing their delivery mission, they return to the warehouse. During these missions, drones may experience interruptions. The primary objective is to efficiently plan drone flight paths, avoiding routes with high interruption risk to mitigate the adverse effects of interruptions and thereby improve the robustness and adaptability of the delivery system. Figure 2 The drone delivery network diagram is shown. If a single purpose is used as the optimization goal, the shortest delivery path, minimized energy consumption, and minimized interruption risk can be achieved. The objective function of the single optimization goal is as follows.

[0046] Shortest delivery path objective function:

[0047]

[0048] Where P represents the set of all nodes that the drone needs to visit, including warehouses, customer demand points, and charging stations; represents the Euclidean distance between two nodes, and Represent the coordinate values ​​of node i in the x direction and y direction respectively, and Represent the coordinate values ​​of node j in the x direction and y direction respectively.

[0049] Minimize energy consumption objective function:

[0050] Where, is the energy consumption of the UAV when flying from node i to node j, Indicates the energy consumption per unit distance.

[0051] Minimize the interruption risk objective function:

[0052] Where, represents the comprehensive interruption risk of the UAV when flying from node i to node j, ranging from 0 to 1.

[0053] The objective function of a single optimization goal defines the optimization direction from different dimensions. To comprehensively consider these aspects, it is necessary to construct a comprehensive objective function, as shown below.

[0054]

[0055] in, 、 and They represent the weight coefficients of delivery path length, energy consumption and comprehensive interruption risk respectively, + + =1.

[0056] In addition, the UAV power limit constraint needs to be met. The constraints of the comprehensive objective function include: the energy consumption of the path completed by the UAV after a single charge is less than the maximum battery capacity of the UAV.

[0057] In addition, the calculation process for the comprehensive interruption risk is as follows.

[0058]

[0059] in, 、 and They represent the weight coefficients of the interruption risk caused by bad weather, the risk of equipment failure, and the risk of communication interference, respectively.

[0060] The interruption risk factors due to bad weather include wind speed, precipitation and visibility. The interruption risk due to bad weather is:

[0061]

[0062] in, and Respectively represent the current wind speed, current rainfall and current visibility, Indicates the maximum wind speed that the drone can withstand. Indicates the maximum rainfall that the drone can safely fly. Indicates the minimum visibility for safe flight of drones. 、 and represent the weight coefficients of wind speed, precipitation and visibility respectively;

[0063] The risk of equipment failure is calculated using historical failure rates and current environmental stresses. Risk of equipment failure:

[0064]

[0065] in, represents the historical failure rate, Indicates the current environmental stress, such as too high or too low temperature will affect the motor efficiency, E s = (current temperature - 25) / 10; Indicates the current flight time. Indicates the maximum battery life, 、 and Represent the weight coefficients of historical failure rate, flight time and current environmental stress respectively;

[0066] The communication situation between the drone and the ground control center or customer demand point will affect the completion of the mission. The communication interference risk is quantified by the communication signal quality, communication interference level and flight distance. Communication interference risk:

[0067]

[0068] in, and Respectively represent the current communication signal strength, current communication interference strength and current flight distance, Indicates the minimum signal strength for secure communication, Indicates the maximum interference intensity that the drone can resist. Indicates the maximum communication coverage distance of the drone, 、 and They represent the weight coefficients of communication signal strength, communication interference strength and flight distance respectively.

[0069] It should be noted that the state transition of the drone may be a normal transition or an interruption transition into a waiting state. The specific situation needs to be judged based on the comprehensive interruption risk of the drone when performing action a. For normal transition, the drone performs action a as planned and transitions to the next state. , then the transition probability is:

[0070]

[0071] in, As the indicator function, to ensure sufficient power, the UAV is required to have a power of is greater than the energy consumption of the drone flying from node i to node j The state transition probability matrix is ​​as follows:

[0072]

[0073] Among them, the elements in the state transition probability matrix Indicates status To status The transition probability, is the state transition probability matrix of the nth step, is the state transition probability matrix of the n-1th step.

[0074] For the interruption transfer case, the UAV triggers the interruption detection when performing action a. If the comprehensive interruption risk exceeds the interruption risk threshold, it enters the interruption state W (waiting), that is, ,in, It represents the comprehensive interruption risk of the UAV entering the interruption state (W) when encountering an interruption during the process of traveling from node i to node j.

[0075] In addition, in order to better search the state and action of the drone in the subsequent optimization steps, the state space and action space are constructed.

[0076] The state space includes all possible states of the drone. Including the current position, remaining power and mission completion status. The location information p includes the warehouse , customer demand points and charging stations The battery state b includes three levels: high, medium, and low. The task state t includes 0 and 1, indicating unfinished and completed, respectively. The state space is:

[0077]

[0078] The action space includes all actions that the drone can perform, including going to another demand point, charging at a charging station, and entering a waiting state after encountering an interruption. Indicates the process from node i to node j; waiting indicates that the drone enters the waiting state after encountering an interruption; when the drone performs an interruption detection while executing action a, if the comprehensive interruption risk is greater than the interruption risk threshold, it is determined that the drone encounters an interruption and enters the waiting state. Action space:

[0079]

[0080] S2: A graph attention network is used to extract features from the distribution network graph, and a multi-layer perceptron is used to extract features from the current state of the drone, obtaining current graph embedding features and current state features, respectively. The current state includes the drone's location, power level, and battery status, and the current distribution network graph includes node location information, task completion status, interruption status, and flight time.

[0081] Specifically, the algorithm first extracts features from state information: the physical spatial layout, node attributes, and path states of the drone delivery system are converted into computable graph data. The graph structure of the delivery network is then encoded using a Graph Attention Network (GAT) to obtain graph embedding features. The GAT's attention mechanism dynamically weights important nodes and edges (for example, reducing the weight of edges on high-risk paths), enabling the algorithm to more accurately learn routing strategies that meet multi-objective optimization goals (shortest paths, low energy consumption, and low risk). Other drone states, such as location, power level, and battery status, are processed using a Multi-Layer Perceptron (MLP) to obtain corresponding state features.

[0082] S3: The current graph embedding features and current state features are concatenated and input into the trained main network to obtain the current step predicted Q value, current reward value and current optimal action; the DDQN network based on the attention mechanism includes the main network, and the current reward value is the value calculated by the reward function constructed based on the comprehensive objective function. The main network includes an input layer, an attention layer, two hidden layers and an output layer, such as Figure 3 shown.

[0083] Specifically, an attention mechanism is introduced to enhance the ability to process complex environmental information. Unlike conventional attention mechanisms, which focus on the relationship between external queries and inputs, this attention mechanism focuses on exploring the internal connections between various state information within the drone delivery environment. This state information includes location, battery level, mission completion status, interruption status, flight time, and the graph structure of the delivery network. This information is crucial for the drone to make appropriate decisions in different scenarios. This step includes the following.

[0084] S31: embed the current graph features and current state features in the input layer (such as Figure 3 The comprehensive feature vector is obtained by splicing.

[0085] S32: Input the comprehensive feature vector into the attention layer to obtain the query vector, key vector and value vector.

[0086] S33: Perform attention mechanism processing based on the query vector, key vector and value vector to obtain a comprehensive feature representation.

[0087] The graph embedding features are concatenated with the state features to produce a comprehensive feature vector. This vector is fed into the attention layer, where a cubic linear transformation is performed to generate a query vector (Query), a key vector (Key), and a value vector (Value). The dot product of the query vector Q and the key vector K is then calculated, scaled by dividing by the square root of the key vector dimension, and then normalized using the Softmax function to produce the attention weights. These weights reflect the importance of different parts of the comprehensive feature vector to the current decision. The value vector V is weighted and summed using the attention weights to produce the feature representation processed by the attention mechanism. Comprehensive feature representation:

[0088]

[0089]

[0090]

[0091]

[0092] Among them, Q, K and V represent the query vector, key vector and value vector respectively, d represents the dimension of the key vector K, represents the transpose of the key vector K, 、 and denote the parameter matrices of the query vector, the key vector, and the value vector, respectively. represents the combined feature vector after concatenation, g represents the current graph embedding feature, and z represents the current state feature.

[0093] This processing method enables the model to dynamically focus on information that is more critical to path planning and interruption management, allowing the drone delivery neural network to more accurately capture interruption factors when dealing with environmental interruption factors. In addition, an attention layer is introduced into the neural network structure to capture node and edge information. Figure 3 As shown in the figure, the state vector describing the interruption scenario serves as the input to the neural network, and the network output is the Q value. Through the attention layer, the state vector of each node and edge is converted into a vector that considers information from all nodes and edges. The network uses two hidden layers with ReLU functions to facilitate nonlinear feature extraction and enhance the processing capabilities of deep neural networks (DNNs) in high-dimensional function spaces.

[0094] S34: Use two hidden layers and the output layer to process the comprehensive feature representation to obtain the current step predicted Q value, current reward value and current optimal action, such as Figure 4 Stages 2 and 3 are shown. The main network predicts the reward value for each possible action and then selects the action with the largest reward value as the current optimal action.

[0095] The reward function is

[0096]

[0097] in, 、 and They represent the path length of executing action a in state s, the energy consumption of executing action a in state s, and the comprehensive interruption risk of executing action a in state s, respectively, where s∈state space S and a∈action space A.

[0098] S4: Send the current optimal action to the drone and receive the next state feedback from the drone, obtain a tuple and save it in the playback cache; the tuple includes the current state, the current optimal action, the current reward value and the next state, such as Figure 4 , as shown by the information interaction between stage two and stage three.

[0099] S5: Update the distribution network diagram based on the current optimal action and the next state, and use the main network to continuously obtain the optimal action for each subsequent step, and finally form the optimal path planning strategy.

[0100] Compared with traditional deep reinforcement learning algorithms, this invention dynamically weights interruption features and spatial topological relationships through an adaptive attention mechanism, significantly improving the success rate of drone path planning in complex interruption scenarios.

[0101] In an optional embodiment of the present application, the DDQN network based on the attention mechanism also includes a target network. The DDQN-based drone interruption scenario path planning method based on the attention mechanism also includes network training and network updating. The steps of network training and network updating are basically the same, such as Figure 4 In addition, during the initialization phase (e.g. Figure 4 Phase 1 in the algorithm is used to initialize the state of the algorithm's operating environment and incorporate interrupt scenarios. Figure 4 Phase 2 in the training process involves interaction with the environment, full exploration of the environment, and collection of interaction data. Figure 4 During training, the parameters of the primary network are optimized through gradient descent to minimize the loss function. Simultaneously, the parameters of the primary network are periodically synchronized with the target network. During updates, the network is updated using the collected interaction data and the selected actions. The steps are as follows.

[0102] S61: Randomly call multiple tuples from the replay cache; the tuple formed after each drone feedback is received is saved in the replay cache, and the tuple includes the current state, the current optimal action, the current reward value and the next state.

[0103] S62: Input the current state in the tuple into the main network to obtain the current step predicted Q value.

[0104] S63: Input the current optimal action in the tuple into the target network, combine the current reward value and discount factor, and obtain the current step target Q value.

[0105] Specifically, the current step target Q value is

[0106]

[0107] in, represents the parameters of the target network, Indicates the current state, represents the current optimal action, Represents the current reward value, Indicates the next step status. represents the next optimal action, Represents the preset discount factor.

[0108] S64: Substitute the current step predicted Q value and the current step target Q value into the loss function to obtain the loss value.

[0109] Specifically, the loss function is

[0110]

[0111] Among them, m represents the number of training times, Indicates the current step predicted Q value output by the main network, Indicates the current step target Q value output by the target network, Indicates the preset demarcation parameter, usually set to 1; Represents the parameters of the main network, represents the reward value for the next step. This loss function combines the advantages of both the mean squared error (MES) and the mean absolute error (MAE). This loss function stabilizes the training process, avoiding the vulnerability of MSE to outliers and ensuring consistent convergence even under fluctuating environmental conditions.

[0112] S65: Perform back propagation based on the loss values ​​corresponding to multiple tuples to update the parameters of the main network.

[0113] S66: Periodically copy the updated parameters of the primary network to the target network to update the parameters of the target network.

[0114] Through network update, a tuple can be obtained , then enters a new state and repeats the above process. Since there is a strong correlation between adjacent tuple transfers, an experience cache mechanism is introduced to eliminate the correlation, improve the convergence speed and improve data utilization.

[0115] Specific experimental examples

[0116] Scenario Setting and Implementation Steps,The experimental simulation is carried out in a two-dimensional area where the,UAV delivery system operates.,The system consists of drones, warehouses, customer demand points, and,charging stations.,The environmental configuration parameters are shown in Table 1.

[0117] Table 1 Environment configuration parameters

[0118]

[0119] To ensure more efficient training and generate feasible solutions, the masking mechanism mentioned above can effectively restrict certain actions under certain conditions. This masking mechanism sets the logarithmic probability of an infeasible action to -∞, or forces a solution to be generated when certain conditions are met. The specific masking mechanism is as follows:

[0120] (1) When a customer demand point is accessed, the customer demand point is masked so that it cannot be accessed again;

[0121] (2) Continuous access to charging stations is not allowed. After the charging station is visited in the previous step, the charging station will be masked in the node that the drone visits next;

[0122] (3) When the drone battery level is less than the first threshold, the customer demand point is masked in the node that the drone will visit next. For example, when the drone battery level is less than 10%, it is not allowed to access the demand point. In addition, when the drone battery level is greater than the second threshold, the charging station is masked in the node that the drone will visit next. For example, when the drone battery level exceeds 70%, the charging station is masked.

[0123] After a lot of experiments, the reasonable ADDQN (DDQN network based on attention mechanism) algorithm parameters are: discount factor ( ) is 0.99, the learning rate (α) is 0.0003, the exploration rate is 1.0, the exploration rate decay is 0.995, the experience buffer pool capacity is 50000, the batch size is 128, and the maximum step size is 40. In addition, to ensure the feasibility of the simulation, the following assumptions are made:

[0124] (1) Only one drone is used to perform the mission. The drone carries all packages during the delivery process, without returning to the warehouse to load additional goods. The weight of all customer packages does not exceed the maximum load capacity of the drone.

[0125] (2) When the drone arrives at the charging station, the battery will be restored to 100% immediately.

[0126] (3) The UAV flies at a constant speed throughout the mission.

[0127] (4) Each charging station can serve multiple demand points.

[0128] In the experiment, the system was trained for 2000 rounds using the ADDQN algorithm. During the training process, the success rate of path planning was calculated every 50 rounds. Figure 5 and Figure 6 The curves of rewards and success rates of drones in the interruption scenario are shown respectively. Figure 5 As shown in the figure, in the first 250 rounds of the initial training, the reward value obtained by the drone fluctuated greatly and was relatively low. During this period, the drone has not yet learned the correct path planning strategy, mainly due to the variability and complexity brought by the interruption environment. After 1000 rounds of training, the reward value increased significantly and gradually reached the maximum value. It can be seen that the performance of the ADDQN algorithm is not completely stable throughout the training process. Factors such as the randomness of the environment and the balance between exploration and utilization may cause differences in the rewards obtained in different rounds. However, the overall upward and stable trend reflects the learning effect of the ADDQN algorithm. In addition, the trend of the success rate curve is consistent with Figure 6 The reward value trends shown in are highly similar. As training progresses, the success rate gradually increases, eventually approaching the maximum value of 90%.

[0129] To further evaluate the success rate, the model was saved every 200 rounds, and then the saved model was tested 20 times. The success rate was calculated and shown in Table 2. It can be seen that the ADDQN algorithm can complete path planning in the interruption scenario with a high success rate.

[0130] In addition, ADDQN is compared with several other deep reinforcement learning algorithms (DDQN, DQN), and the data is plotted into charts for comparison. Figure 7 and Figure 8 The reward curves and success rate curves for three different algorithms under the same environment are shown. As can be seen, the reward curves of DDQN and DQN ultimately converge to the same maximum value, but at slightly different rates. Specifically, DDQN's cumulative reward is more stable and converges faster than DQN. DDQN reaches its maximum value after 800 training rounds, while DQN requires 1200 rounds and exhibits greater fluctuations. This is attributed to DDQN's dual-network mechanism, which effectively addresses the problem of overestimation of the value function. In comparison, ADDQN's final convergence value is higher than both DDQN and DQN. Furthermore, ADDQN's success rate reaches 90% after 2000 rounds, significantly higher than both DDQN and DQN. Comparing the above graphs shows that the ADDQN algorithm demonstrates superior performance, demonstrating the effectiveness of the attention mechanism.

[0131] Table 2 ADDQN success rate evaluation

[0132]

[0133] The same success rate test was performed on the DDQN and DQN algorithms and compared, with the results shown in Table 3. ADDQN achieved a success rate of 70% in the early rounds and maintained above 90% in the later stages, consistently outperforming both DDQN and DQN.

[0134] Table 3 Success rate evaluation of three algorithms

[0135]

[0136] In addition, the same scene was simulated 20 times using these three algorithms, and the simulated path curves are as follows: Figure 9 In addition, the average path length and reward value of 20 simulations are shown in Table 4. For the three different paths, the path planned by ADDQN is shorter and smoother than that of other algorithms.

[0137] As shown in Table 5, a comparative analysis of the three algorithms in interruption scenario simulations reveals significant differences in their robustness to dynamic disturbances. ADDQN experienced the lowest interruption frequency, 234 times, accounting for 11.7% of the total test count. In contrast, DDQN experienced 303 interruptions, while DQN performed the worst, experiencing 403 interruptions. Therefore, it can be inferred that ADDQN is superior to the other two algorithms in mitigating interruptions.

[0138] Table 4 Simulation results of three algorithms

[0139]

[0140] Table 5 Interruption frequency of three algorithms

[0141]

[0142] This paper proposes a dual-deep Q-network (ADDQN)-based path planning method for low-altitude UAV logistics disruption scenarios. By integrating MDP modeling, graph attention network (GAT), and adaptive attention mechanism, it demonstrates significant advantages in path optimization under dynamic disruption environments. Specifically:

[0143] 1. Path planning success rate and robustness are significantly improved

[0144] Compared with traditional deep reinforcement learning algorithms, this invention uses an adaptive attention mechanism to dynamically weight interruption features and spatial topological relationships, significantly improving the success rate of drone path planning in complex interruption scenarios. Experimental data shows that:

[0145] (1) ADDQN achieves a 90% success rate after 2000 training rounds, while DDQN and DQN achieve success rates of 85% and 80% respectively.

[0146] (2) When encountering dynamic interruptions such as weather changes and equipment failures, ADDQN dynamically adjusts decision weights, and the interruption frequency is only 234 times (accounting for 11.7% of the total test), which is 32%-42% less than DQN (403 times) and DDQN (303 times).

[0147] 2. Better path optimization efficiency and energy consumption control

[0148] (1) By encoding the distribution network topology using GAT and focusing on key nodes with the attention mechanism, ADDQN optimizes path length and energy consumption:

[0149] (2) The average path length is shortened by 12.6% compared with DDQN (39.21 km for ADDQN and 44.88 km for DDQN) and by 29.4% compared with DQN (55.57 km for DQN), effectively reducing the flight distance.

[0150] (3) Combined with the mask mechanism, the energy utilization of ADDQN is improved, avoiding energy waste caused by redundant charging or circuitous paths.

[0151] 3. Enhanced algorithm convergence speed and environmental adaptability

[0152] (1) Although the reward fluctuates due to interruption randomness in the initial training phase, the reward value of ADDQN stabilizes near the maximum value after 1000 rounds.

[0153] Compared with DDQN's dual-network mechanism, ADDQN further reduces the problem of overestimation of the value function through the attention mechanism, allowing policy optimization to focus more on the balance between interruption risk and path cost, and exhibiting stronger generalization capabilities in high-dimensional state spaces.

[0154] In summary, compared with the existing technology, the present invention has the following beneficial effects:

[0155] 1. By integrating the attention mechanism and the dual deep Q network (DDQN), an adaptive attention mechanism is adopted to calculate the impact weight of different environmental conditions (such as wind speed, battery power, and communication signal) on decision-making in real time, giving priority to high-risk interruption factors.

[0156] 2. Use GAT to explicitly encode the spatial relationships between nodes (such as customer demand points and charging stations) and edges in the logistics network to improve the global rationality of path planning; combine it with a multi-layer perceptron (MLP) to extract features of the drone's own status (position, battery level), and integrate them with graph structure features to enhance the environment representation capability.

[0157] 3. A dual-network structure is adopted to separate action selection and Q-value evaluation, alleviating the Q-value overestimation problem of traditional DQN. The loss function replaces the mean squared error (MSE) to balance the accuracy of small errors and the robustness of large errors, improving training stability. The experience cache and action mask mechanism are combined to avoid invalid actions (such as repeatedly visiting customer demand points that have already been delivered) and accelerate convergence.

[0158] 4. Design a composite reward function to comprehensively optimize path length, energy consumption, and interruption risk to achieve a multi-objective balance; based on a dynamic interruption risk threshold, determine interruptions and adjust strategies.

[0159] 5. Tested in a custom drone logistics environment, the proposed method performs better than traditional DQN and DDQN in terms of path smoothness, mission success rate, and interruption robustness.

[0160] like Figure 10 As shown, the embodiment of the present application provides a drone interruption scenario path planning system based on DDQN with an attention mechanism, including:

[0161] Construction module 10 is used to construct a comprehensive objective function and a distribution network diagram based on node location information, energy consumption per unit distance of drones, and comprehensive interruption risks between nodes; comprehensive interruption risks include interruption risks due to bad weather, equipment failure risks, and communication interference risks.

[0162] The extraction module 20 is used to extract features from the distribution network graph using a graph attention network and to extract features from the current state of the drone using a multi-layer perceptron to obtain current graph embedding features and current state features respectively.

[0163] The optimization module 30 is used to splice the current graph embedding features and the current state features, input them into the trained main network, and obtain the current step predicted Q value, the current reward value and the current optimal action; the DDQN network based on the attention mechanism includes the main network, and the current reward value is the value calculated based on the reward function constructed according to the comprehensive objective function.

[0164] The interaction module 40 is used to send the current optimal action to the drone and receive the next state feedback from the drone.

[0165] The update loop module 50 is used to update the distribution network diagram according to the current optimal action and the next step state, and continuously obtain the optimal action of each subsequent step using the main network, and finally form the optimal path planning strategy.

[0166] In this embodiment, the beneficial effects of the drone interruption scenario path planning system based on DDQN with attention mechanism are similar to the beneficial effects of the drone interruption scenario path planning method based on DDQN with attention mechanism mentioned above, and will not be repeated here.

[0167] An electronic device provided in an embodiment of the present application includes a memory and a processor; the memory is used to store a computer program; the processor is used to implement the above-mentioned DDQN-based drone interruption scenario path planning method when executing the computer program.

[0168] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the above-described DDQN-based drone interruption scenario path planning method is implemented.

[0169] In this embodiment, the beneficial effects of the electronic device and the computer-readable storage medium are similar to the beneficial effects of the above-mentioned DDQN-based attention mechanism-based drone interruption scenario path planning method, and will not be repeated here.

[0170] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0171] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A DDQN-based attention mechanism-based UAV interruption scenario path planning method, characterized by: include: Based on the node location information, the energy consumption per unit distance of the drone, and the comprehensive interruption risk between nodes, a comprehensive objective function and a distribution network diagram are constructed; Comprehensive disruption risks include disruption risks due to severe weather, equipment failure, and communications interference; A graph attention network is used to extract features from the distribution network graph, and a multi-layer perceptron is used to extract features from the current state of the drone, obtaining the current graph embedding features and current state features respectively. The current graph embedding features and current state features are concatenated and input into the trained main network to obtain the current step predicted Q value, current reward value, and current optimal action; the DDQN network based on the attention mechanism includes the main network, and the current reward value is calculated based on the reward function constructed according to the comprehensive objective function; Send the current optimal action to the drone and receive the next state feedback from the drone; The distribution network diagram is updated based on the current optimal action and the next state, and the optimal action for each subsequent step is continuously obtained using the main network, ultimately forming the optimal path planning strategy.

2. The DDQN-based drone interruption scenario path planning method according to claim 1 is characterized in that: The main network includes an input layer, an attention layer, two hidden layers and an output layer; The current graph embedding feature and the current state feature are spliced ​​and input into the trained main network to obtain the current step predicted Q value, current reward value and current optimal action. The current graph embedding features and the current state features are concatenated at the input layer to obtain a comprehensive feature vector; Input the comprehensive feature vector into the attention layer to obtain the query vector, key vector and value vector; According to the query vector, key vector and value vector, the attention mechanism is processed to obtain a comprehensive feature representation; The comprehensive feature representation is processed using two hidden layers and the output layer to obtain the current step predicted Q value, the current reward value, and the current optimal action.

3. The DDQN-based drone interruption scenario path planning method according to claim 1 is characterized in that: The DDQN network based on the attention mechanism also includes a target network; The method further comprises: Randomly call multiple tuples from the replay cache. Each tuple generated after receiving drone feedback is saved in the replay cache. The tuple includes the current state, the current optimal action, the current reward value, and the next state. Input the current state in the tuple into the main network to obtain the current step predicted Q value; Input the current optimal action in the tuple into the target network, combine the current reward value and discount factor to obtain the current step target Q value; Substitute the current step predicted Q value and the current step target Q value into the loss function to obtain the loss value; Backpropagate based on the loss values ​​corresponding to multiple tuples to update the parameters of the main network; The updated parameters of the master network are periodically copied to the target network to update the parameters of the target network.

4. The DDQN-based drone interruption scenario path planning method according to claim 1 is characterized in that: The comprehensive objective function is Where P represents the set of all nodes that the drone needs to visit, including warehouses, customer demand points, and charging stations; Represents the Euclidean distance between two nodes; Indicates the energy consumption per unit distance; represents the comprehensive interruption risk when the UAV flies from node i to node j; 、 and They represent the weight coefficients of delivery path length, energy consumption and comprehensive interruption risk respectively, + + =1; The constraints of the comprehensive objective function include: the energy consumption of the route completed by the UAV after a single charge is less than the maximum battery capacity of the UAV; The interruption risk due to bad weather is in, and Respectively represent the current wind speed, current rainfall and current visibility, Indicates the maximum wind speed that the drone can withstand. Indicates the maximum rainfall that the drone can safely fly. Indicates the minimum visibility for safe flight of drones. 、 and represent the weight coefficients of wind speed, precipitation and visibility respectively; The risk of equipment failure is in, represents the historical failure rate, Indicates the current environmental stress, Indicates the current flight time. Indicates the maximum battery life, 、 and Represent the weight coefficients of historical failure rate, flight time and current environmental stress respectively; The communication interference risk is in, Respectively represent the current communication signal strength, current communication interference strength and current flight distance, Indicates the minimum signal strength for secure communication, Indicates the maximum interference intensity that the drone can resist. Indicates the maximum communication coverage distance of the drone, 、 and Represent the weight coefficients of communication signal strength, communication interference strength and flight distance respectively; The combined disruption risk is in, 、 and They represent the weight coefficients of the interruption risk caused by bad weather, the risk of equipment failure, and the risk of communication interference, respectively.

5. The DDQN-based drone interruption scenario path planning method according to claim 4 is characterized in that: The reward function is in, 、 and They represent the path length of executing action a in state s, the energy consumption of executing action a in state s, and the comprehensive interruption risk of executing action a in state s, respectively, s∈state space S, a∈action space A; The state space is Among them, the location information p includes the warehouse , customer demand points and charging stations The battery status b includes three levels: high, medium, and low. The task status t includes 0 and 1, indicating unfinished and completed, respectively. The action space is in, represents the journey from node i to node j; waiting indicates that the drone enters the waiting state after encountering an interruption; when the drone performs an interruption detection while executing action a, if the comprehensive interruption risk is greater than the interruption risk threshold, it is determined that the drone encounters an interruption and enters the waiting state.

6. The DDQN-based drone interruption scenario path planning method according to claim 2 is characterized in that: The comprehensive features are expressed as Among them, Q, K and V represent the query vector, key vector and value vector respectively, d represents the dimension of the key vector K, represents the transpose of the key vector K, 、 and denote the parameter matrices of the query vector, the key vector, and the value vector, respectively. represents the combined feature vector after concatenation, g represents the current graph embedding feature, and z represents the current state feature.

7. The DDQN-based drone interruption scenario path planning method according to claim 3 is characterized in that: The current step target Q value is in, represents the parameters of the target network, Indicates the current state, represents the current optimal action, Represents the current reward value, Indicates the next step status. represents the next optimal action, represents the preset discount factor; The loss function is Among them, m represents the number of training times, Indicates the current step predicted Q value output by the main network, Indicates the current step target Q value output by the target network, Indicates the preset demarcation parameter, Represents the parameters of the main network, Indicates the next reward value.

8. A DDQN-based drone interruption scenario path planning system based on attention mechanism, characterized by: include: A construction module is used to construct a comprehensive objective function and a distribution network graph based on node location information, energy consumption per unit distance of drones, and comprehensive interruption risk between nodes; Comprehensive disruption risks include disruption risks due to severe weather, equipment failure, and communications interference; The extraction module is used to extract features from the distribution network graph using a graph attention network and a multi-layer perceptron to extract features from the current state of the drone, obtaining the current graph embedding features and current state features respectively; The optimization module is used to splice the current graph embedding features and the current state features, and input them into the trained main network to obtain the current step predicted Q value, the current reward value and the current optimal action; the DDQN network based on the attention mechanism includes the main network, and the current reward value is calculated based on the reward function constructed according to the comprehensive objective function; The interaction module is used to send the current optimal action to the drone and receive the next state feedback from the drone; The update loop module is used to update the distribution network diagram based on the current optimal action and the next state, and use the main network to continuously obtain the optimal action for each subsequent step, and finally form the optimal path planning strategy.

9. An electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is configured to implement the drone interruption scenario path planning method based on DDQN with attention mechanism as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by the processor, the drone interruption scenario path planning method based on the attention mechanism DDQN according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Unmanned aerial vehicle rapid path planning method and system based on deep learning

    CN116449860A

  • Mobile robot dynamic path planning method based on improved DDQN algorithm

    CN119952697A