A dynamic topology adaptive unmanned aerial vehicle route planning method and system based on trajectory perception

CN122813835APending Publication Date: 2026-09-25HUNAN NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610915700.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]针对现有技术的以上缺陷或改进需求,本发明提供了一种基于轨迹感知的动态拓扑自适应无人机航线规划方法和系统,其目的在于,解决现有两阶段启发式方法存在的分簇拓扑与飞行航线深度割裂,无人机可能需要绕远路去访问分布极不合理的簇头,极易陷入局部最优的技术问题,以及无法应对无线电衰落灾难,从而引发部分节点能量快速耗尽的现象的技术问题,以及现有基于深度强化学习的联合优化方法存在严重的维度灾难的技术问题,以及由于极易在相邻迭代的“拆分”与“合并”动作之间陷入死循环,从而导致模型产生严重的拓扑震荡问题,使得算法难以有效收敛的技术问题

Benefits of technology

[0097]1、本发明由于采用了步骤(6-1)到步骤(6-3),在中高层强化学习框架内引入轨迹感知机制,该机制将中层网络规划的空间飞行距离和物理代价自适应转化为惩罚项,并反向传递至高层智能体的广义优势估计中,以驱动地面的分簇拓扑主动改变物理形状从而缩短空中飞行航线。因此,本发明能够解决现有传统的两阶段启发式方法存在的分簇拓扑与飞行航线深度割裂、容易陷入局部最优的技术问题;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122813835A_ABST
    Figure CN122813835A_ABST
Patent Text Reader

Abstract

The application discloses a kind of dynamic topology adaptive unmanned aerial vehicle route planning methods based on trajectory perception, it utilizes graph neural network in macroscopic topological evolution layer, combines dynamic mask and taboo table, outputs legal topological variable dimension action, from bottom layer to suppress state space expansion and decision shock;In middle layer route planning layer, introduce the Transformer pointer network with filling mask mechanism to process dynamic change cluster feature, autoregressive generation non-repeated Hamilton path;In bottom layer physical optimization layer, according to the route, construct multilayer state space directed acyclic graph, under the strict communication fading and flight energy consumption constraint, accurately elect the cluster head coordinate sequence with the lowest global energy consumption using A* algorithm;Finally, the physical energy consumption and topological over-standard punishment accurately calculated in bottom layer are adaptively converted into joint reward, and are reversely transmitted to high-level intelligent agent to update network weight, drive ground topological active reshaping physical boundary to shorten air route.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of interdisciplinary technology of Internet of Things communication and artificial intelligence, and more specifically, relates to a method and system for dynamic topology adaptive UAV route planning based on trajectory perception. Background Technology

[0002] In recent years, using unmanned aerial vehicles (UAVs) as mobile data collectors to assist wireless sensor networks (WSNs) in data collection has become an important research direction in the field of IoT data acquisition because it can effectively alleviate the energy constraints of ground nodes and improve network lifespan.

[0003] Existing data acquisition methods for UAV-assisted wireless sensor networks mainly fall into two categories: one is a two-stage heuristic method, which typically adopts a "clustering first, planning later" strategy. That is, firstly, clustering algorithms are used to group ground nodes and select cluster heads (CHs), and then path planning algorithms are used to generate flight paths for UAVs to fly over each cluster head in sequence. The other is a joint optimization method based on deep reinforcement learning, which attempts to introduce deep reinforcement learning for global optimization. The agent directly outputs the cluster partitioning at the node level and relies on the reward function to adaptively generate network topology and flight paths.

[0004] However, the aforementioned existing methods all have some significant drawbacks: First, the two-stage heuristic methods treat ground clustering and flight path planning as two independent problems. Ground clustering does not consider the flight cost of the UAV, and the UAV cannot reverse the ground network topology when planning its flight path. This results in a deep disconnect between the clustering topology and the flight path, potentially causing the UAV to take a longer route to visit cluster heads with extremely unreasonable distributions, making it highly susceptible to local optima. Second, the two-stage heuristic methods typically presuppose a fixed number of clusters. In actual physical space, once the physical radius of a cluster exceeds the free-space communication threshold, communication between boundary nodes and cluster heads will trigger severe multipath fading. Due to the lack of dynamic topology adaptation in the system, this problem becomes more severe. First, the existing reinforcement learning-based joint optimization methods lack the ability to cope with radio fading disasters, leading to the rapid depletion of energy in some nodes. Second, while these methods attempt global optimization, the direct output of node-level cluster partitioning results in an exponential explosion of the action space as the number of nodes increases in complex and ever-changing large-scale IoT scenarios, causing a severe curse of dimensionality problem. Third, the methods lack underlying physical constraints and historical memory error correction mechanisms, making it easy for reinforcement learning agents to get stuck in a dead loop between "split" and "merge" actions in adjacent iterations, leading to severe topological oscillations and making it difficult for the algorithm to converge effectively. Summary of the Invention

[0005] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a dynamic topology adaptive UAV flight path planning method and system based on trajectory perception. Its purpose is to solve the technical problems of existing two-stage heuristic methods, such as the deep disconnect between cluster topology and flight path, which may cause UAVs to take detours to visit poorly distributed cluster heads, easily leading to local optima; the inability to cope with radio fading disasters, resulting in rapid energy depletion of some nodes; the severe dimensionality curse of existing deep reinforcement learning-based joint optimization methods; and the problem of the model easily getting stuck in infinite loops between adjacent iterations of "splitting" and "merging," leading to severe topological oscillations and making effective convergence difficult.

[0006] To achieve the above objectives, according to one aspect of the present invention, a dynamic topology adaptive UAV route planning method based on trajectory awareness is provided, comprising the following steps:

[0007] (1) Obtain the original environmental data in the target area, extract the initial geographic coordinate set of all the corresponding nodes from the original environmental data, and perform feature engineering on the initial geographic coordinate set to obtain the heterogeneous graph state tensor HST.

[0008] (2) Input the heterogeneous graph state tensor obtained in step (1) into the graph neural network to extract the topology of all nodes. Based on the topology, the PPO algorithm is optimized using the near-end strategy to obtain the macroscopic topological evolution action. ;

[0009] (3) The initial cluster label set obtained in step (1) and cluster center coordinates Used as the current cluster partitioning structure Based on the macroscopic topological evolution action obtained in step (2) and the current cluster partitioning structure And use the environmental physical actuator to redistribute or regroup all nodes extracted from the original environmental data in step (1) to obtain the cluster partitioning structure;

[0010] (4) Extract all cluster-level sequence features from the cluster partitioning structure obtained in step (3), and process all cluster-level sequence features using the Transformer pointer network to obtain a non-repeating flight access sequence. ,in The cluster label for the drone to access the k-th cluster;

[0011] (5) Based on the non-repeating flight access sequence obtained in step (4) A multi-level state search space is constructed, and the A* algorithm is used to perform multi-level nested loop optimization iteration within this multi-level state search space to obtain the optimal cluster head (CH) physical coordinate sequence. and global energy consumption ,in This represents the total number of clusters within the target area;

[0012] (6) The global energy consumption obtained from step (5) Generate a joint reward and use the joint reward to synchronously update the network weights of the graph neural network in step (2) and the Transformer pointer network in step (4) respectively, so as to obtain the updated graph neural network and the Transformer pointer network respectively;

[0013] (7) Determine whether the updated graph neural network and the updated Transformer pointer network in step (6) have reached convergence, or whether the number of iterations has reached the preset maximum number of training rounds. If so, stop training and set the preset UAV flight path starting point and the optimal cluster head physical coordinate sequence obtained in step (5). The preset drone flight path endpoints are spliced ​​together according to the order of time access to obtain the final drone flight path. The process ends; otherwise, return to step (1).

[0014] Preferably, step (1) specifically includes the following sub-steps:

[0015] (1-1) Obtain the original environmental data within the target area, and extract the initial geographic coordinates of each node corresponding to the original environmental data. The initial geographic coordinates of all nodes corresponding to the original environmental data constitute the initial geographic coordinate set. , where N represents the total number of all nodes corresponding to the original environment data, and i∈[1,N];

[0016] (1-2) Perform initial clustering processing on the initial geographic coordinate set V obtained in step (1-1) to obtain the initial cluster label set. The cluster center coordinates of the corresponding cluster are used as reference target points for UAV flight access and data collection.

[0017] (1-3) The initial cluster label set obtained from step (1-2) Get the number of nodes in the k-th cluster And based on the number of nodes Obtain the node load variance within the target area ;

[0018] (1-4) The initial cluster label set obtained from step (1-2) The cluster center coordinates corresponding to each cluster, and the initial geographic coordinate set obtained in step (1-1). Get the i-th node and its corresponding cluster center. First distance The distance from the i-th node to its corresponding cluster center Adjacent cluster centers The second distance And based on this second distance Distance from the first The ratio is used to obtain the edge degree index of the i-th node. Where A and B are both ∈ [1, K];

[0019] (1-5) Obtain the edge degree index from all N nodes All nodes smaller than a preset threshold are designated as boundary nodes, and all boundary nodes constitute a boundary node feature list. ;

[0020] (1-6) Using the geographic coordinate set obtained in step (1-1) The edge degree index of the i-th node obtained in step (1-4) Constructing a sparse adjacency matrix ;

[0021] (1-7) Using the initial geographic coordinate set obtained in step (1-1) Obtain the spatial distance between any two nodes within the target area. The boundary node feature list BF obtained in step (1-5) and the number of clusters obtained in step (1-2) are compared. and the load variance obtained in steps (1-3) A fusion encapsulation process is performed to obtain the heterogeneous graph state tensor.

[0022] Preferably, the initial clustering process in step (1-2) uses the K-means algorithm, based on a preset number of clusters. Divide all N nodes into There are 3 clusters, each with an initial cluster label. An initial cluster label is assigned to the i-th node. Then, the coordinates of the cluster center corresponding to the k-th cluster are obtained. The initial cluster labels corresponding to all nodes constitute the initial cluster label set. , where k∈[1, the total number of clusters K obtained from the initial clustering process in step (1-2)], and the initial value of K is 5% to 15% of the total number of all nodes corresponding to the original environment data;

[0023] Node load variance in steps (1-3) The following formula is used for calculation:

[0024] ;

[0025] in This represents the average number of nodes contained in all clusters, and has... =N / K;

[0026] The edge index in steps (1-4) is calculated using the following formula:

[0027] ;

[0028] The sparse adjacency matrix constructed in steps (1-6) The dimension is , of which Line 1 Column elements Used to characterize the The node and the first The topological connection relationship between the nth nodes, if the nth node... The node and the first If the nth node corresponds to the same cluster, and the spatial distance between them is less than a preset communication distance threshold, then the nth node in the sparse adjacency matrix... Line 1 Column elements If the first The node and the first If the nth node belongs to different clusters, but both are boundary nodes, and their spatial distance is less than a preset communication distance threshold, then the nth node in the sparse adjacency matrix... Line number Column elements If the first The node and the first If the nth node does not satisfy any of the above conditions, then the nth node in the sparse adjacency matrix... Line number Column elements , where j∈[1,N].

[0029] Preferably, the graph neural network is a multi-layer graph attention network (GAT).

[0030] Step (2) specifically includes the following sub-steps:

[0031] (2-1) Input the heterogeneous graph state tensor obtained in step (1-7) into the graph neural network and perform message passing processing to obtain the hidden feature vector of each node. ;

[0032] (2-2) Hidden feature vectors of all nodes obtained in step (2-1) Global average pooling is performed to extract the global topological latent vectors of all nodes. And compare it with the number of clusters obtained in step (1-2). and the network load variance obtained in steps (1-3) Concatenate the data to obtain the joint feature vector of all nodes. ;

[0033] (2-3) Combine the joint feature vectors of all nodes obtained in step (2-2) The Actor network in the PPO algorithm is input for fully connected layer operations to obtain the log probability distribution of the five-dimensional original macroscopic topological evolution action of each node. ;

[0034] (2-4) Log probability distribution of the five-dimensional primitive macroscopic topological evolution action obtained in step (2-3) Masking and filtering are performed to obtain the probability distribution of actual macroscopic topological evolution actions. ;

[0035] (2-5) Using a pre-defined taboo list, analyze the probability distribution of the actual macroscopic topological evolutionary actions obtained in step (2-4). Oscillation correction is performed to generate the final macroscopic topological evolution action. .

[0036] Preferably, the five-dimensional original macroscopic topological evolution actions include two categories: the first category is node-level actions, including moving boundary nodes, exchanging boundary nodes, and maintaining the status quo; the second category is cluster-level actions, including merging two adjacent clusters and splitting overloaded clusters.

[0037] Step (3) specifically includes the following sub-steps:

[0038] (3-1) The initial cluster label set obtained in step (1) coordinates of cluster center Used as the current cluster partitioning structure ;

[0039] (3-2) Based on the macroscopic topological evolution action obtained in step (2) The current cluster partitioning structure obtained in step (3-1) Perform dynamic dimension transformation to obtain the updated cluster count. ;

[0040] (3-3) Based on the macroscopic topological evolution action obtained in step (2) The boundary nodes obtained in steps (1-5) are then subjected to topology reassignment to obtain the updated cluster label set. ;

[0041] (3-4) The updated cluster label set obtained in step (3-3) and the initial geographic coordinate set obtained in step (1-1) Perform a centroid update to obtain the updated cluster center coordinates. and the number of nodes in each cluster ;

[0042] (3-5) The updated cluster count obtained in step (3-2) The updated cluster label set obtained in step (3-3) and the updated cluster center coordinates obtained in steps (3-4) The data is packaged and encapsulated to obtain the cluster partitioning structure.

[0043] Preferably, step (4) specifically includes the following sub-steps:

[0044] (4-1) Extract the geometric centroid coordinates of each cluster from the cluster partitioning structure updated in step (3-5). and the number of nodes included Both are then input into a linear projection layer to obtain the initial feature sequence. ,in This represents the node embedding vector of the k-th cluster;

[0045] Specifically, the node embedding vector of the k-th cluster The following formula is used for calculation:

[0046] ;

[0047] in, Represents a shared linear projection layer. Representing the The cluster center of each cluster, Representing the The number of nodes contained in a cluster These are learnable bias parameters;

[0048] (4-2) The initial feature sequence obtained in step (4-1) Zero-vector padding is performed to obtain the cluster feature tensor. And based on the cluster feature tensor The padding mask is generated by using the index positions of the valid elements (which refer to the node embedding vectors corresponding to each cluster in the updated cluster partitioning structure obtained in step (4-1)) and the padding placeholders (which refer to the zero vectors added during the zero vector supplementation process). ,in This indicates the preset maximum number of clusters;

[0049] (4-3) The cluster feature tensor obtained in step (4-2) Input a Transformer encoder to compute cluster feature tensors using the multi-head self-attention mechanism MHSA. The spatial dependency weights between each pair of clusters are used to determine the spatial dependency weights of the cluster feature tensor. We perform weighted feature fusion on the effective elements in the dataset to obtain the updated local node embedding vector for each cluster, and then concatenate the updated node embedding vectors for all clusters to obtain the cluster feature tensor. The corresponding encoded hidden state matrix Eh ,in This indicates the preset maximum number of clusters. This represents the dimension of the hidden layer features output by the Transformer encoder;

[0050] (4-4) Use a pointer decoder to perform autoregressive sequence decoding on the hidden state matrix Eh obtained in step (4-3) to generate a non-repeating flight access sequence covering all K clusters. ,in It is the cluster label of the k-th cluster visited by the drone;

[0051] Specifically, this step involves first initializing a visited state vector with all zeros as a visited mask, used to track and record the UAV's historical visit footprints during flight path planning in real time; then, in the current first decoding step, the pointer decoder is used to encode the hidden state matrix obtained in step (4-3). Attention is calculated to obtain the attention score matrix for each cluster. Then, using the padding mask and visited mask obtained in step (4-2), the attention score matrix is ​​masked, assigning negative infinity to the scores of virtual clusters representing the padding region and visited clusters, to obtain the masked attention score matrix. Subsequently, the masked attention score matrix is ​​normalized to obtain the transition probability corresponding to the first decoding step. ,in The cluster label of the first target cluster to be accessed is The probability of transition is calculated, and based on this transition probability, a maximum probability greedy selection or random sampling strategy is adopted to select the target cluster to be visited by the UAV from the clusters whose corresponding state position is 0 in the already visited mask; then, the state position corresponding to the selected target cluster in the already visited mask is set to 1; then, for the 2nd decoding step, the 3rd decoding step, ... the Kth decoding step, the above operation is repeated until all K decoding steps have been processed, and the sequentially selected clusters are obtained. The target clusters constitute a non-repeating flight access sequence.

[0052] Preferably, step (5) specifically includes the following sub-steps:

[0053] (5-1) Based on the non-repeating flight access sequence generated in step (4-4) The preset drone flight path starting point and initial cluster label set will be used. All nodes within all clusters and the preset drone flight path endpoints are hierarchically expanded and topologically associated along the vertical dimension to construct a directed acyclic graph (DAG). The vertical topology of this DAG is determined by... It consists of several levels, where level 0 is the starting point of the drone flight path, and levels 1 through 2 are... The layers correspond to the non-repeating flight access sequence in sequence. Each cluster in, the first The layer marks the end point of the drone's flight path;

[0054] (5-2) Initialize the priority queue of the A* algorithm, take the node of layer 0 of the DAG obtained in step (5-1) as the initial node to be evaluated, and set the historical actual cost of the initial node to be evaluated. The sequence of non-repeating flight visits to the initial node to be evaluated is obtained based on Euclidean distance. The shortest theoretical flight distance from the cluster center of all clusters to the end of the UAV's flight path is calculated. This shortest theoretical flight distance is then multiplied by a preset unit flight energy consumption coefficient for the UAV to obtain the ideal flight energy consumption of the UAV. This ideal flight energy consumption is then used as the initial heuristic energy consumption estimate for the initial node to be evaluated. Using the formula Obtain the comprehensive evaluation value of the initial node to be evaluated. And add the initial node to be evaluated to the priority queue;

[0055] (5-3) Check if the priority queue is empty. If it is, it means that the globally optimal path has not been found, and the process ends; otherwise, obtain the comprehensive evaluation value from the priority queue. The smallest node is used as the current traversed node. Then proceed to step (5-4);

[0056] (5-4) Determine the currently traversed node Does it belong to the first DAG graph? Layer, i.e., determining the currently traversed node Is this the destination of the drone's flight path? If so, it means that the globally optimal path has been found, and then proceed to step (5-15); otherwise, proceed to step (5-5).

[0057] (5-5) Get the current traversed node The level number in the DAG diagram And extract the first element from the DAG graph. All nodes contained in the layer are used as a candidate cluster head set, and the total number of candidate cluster heads in this candidate cluster head set is obtained. ;

[0058] (5-6) Set up a candidate cluster head traversal counter ;

[0059] (5-7) Determine the counter Is it greater than the total number of candidate cluster heads? If yes, return to step (5-3); otherwise, extract the first cluster head from the candidate cluster head set. The node is selected as the current candidate cluster head. Then proceed to steps (5-8);

[0060] (5-8) Obtain the first digit in the DAG graph. In addition to the candidate cluster head, the layer All other nodes and their total number Set a node traversal counter and initialize the cumulative communication energy consumption within the current layer. ;

[0061] (5-9) Determine the node traversal counter Is it greater than the total number of other nodes? If yes, proceed to step (5-12); otherwise, obtain the first element in the DAG graph. The first in the layer Each node is selected, and the process proceeds to step (5-10).

[0062] (5-10) Calculate the first digit of the DAG graph. The first in the layer Each node and candidate cluster head Spatial communication distance between Determine the spatial communication distance Is it less than or equal to the multipath fading threshold? If so, then the first [signal] is calculated based on the free-space signal propagation model. Each node and candidate cluster head Energy consumption for communication between And accumulate the communication energy consumption to the first In the cumulative communication energy consumption of the layer, that is Then proceed to step (5-11); otherwise, calculate the first step based on the multipath fading signal propagation model. Each node and candidate cluster head Energy consumption for communication between And accumulate the communication energy consumption to the first In the cumulative communication energy consumption of the layer, that is Then proceed to step (5-11), where the multipath fading threshold is... , This represents the power amplifier energy consumption coefficient under the free-space signal propagation model. This represents the power amplifier energy consumption coefficient under the multipath fading signal propagation model;

[0063] (5-11) Set a node traversal counter And return to steps (5-9);

[0064] (5-12) Based on the drone's current traversal nodes Fly to candidate cluster head Spatial transfer distance acquisition mechanical flight energy consumption To obtain the drone in the candidate cluster head Hovering power consumption during wireless communication above According to the energy consumption of this mechanical flight Hovering energy consumption and the first one obtained in step (5-10) Cumulative communication energy consumption of the layer Get the nodes from the current traversal Transfer to candidate cluster head Weighted step cost :

[0065] ;

[0066] in, and These are dimensionless weighted coefficients. The value range is from 0.1 to 0.9. The value range is from 0.1 to 0.9;

[0067] (5-13) Set the currently traversed node The actual historical cost The weighted step cost obtained in step (5-12) Perform summation to obtain the number of nodes reaching the candidate cluster head. The actual cost ; Obtain the candidate cluster head Sequential non-repeating flight access sequence The shortest theoretical flight distance from the cluster center of all remaining cluster labels to the end of the UAV's flight path is calculated. This shortest theoretical flight distance is then multiplied by a preset UAV unit flight energy consumption coefficient to obtain the UAV's ideal flight energy consumption, which is then used as the candidate cluster head. Heuristic estimation of energy consumption Using the formula Obtain the candidate cluster head Comprehensive assessment value , will candidate cluster heads Add it to the priority queue and proceed to step (5-14).

[0068] (5-14) Set up candidate cluster head counters And return to steps (5-7);

[0069] (5-15) Backtrack along the globally optimal path in the priority queue in the DAG graph to obtain the optimal cluster head physical coordinate sequence. and will the The historical actual cost corresponding to the layer node is used as the global energy consumption. Output.

[0070] Preferably, step (6) specifically includes the following sub-steps:

[0071] (6-1) The global energy consumption obtained from step (5-15) Obtain basic reward amount ;

[0072] (6-2) Calculate the node load variance within the target region based on the updated cluster partitioning structure obtained in step (3). Extract the boundary nodes within each cluster in the updated cluster partitioning structure, and calculate the spatial distance from each boundary node to the cluster center coordinates corresponding to that cluster. Then, select the boundary nodes whose spatial distances are greater than the multipath fading threshold obtained in step (5-10). All distance overflows are summed, and the sum is used as a penalty for exceeding the cluster radius limit within the target area. and node load variance Cluster radius exceeding the limit penalty and the basic reward obtained in step (6-1) Perform weighted synthesis to obtain joint rewards ;

[0073] (6-3) The combined reward obtained from step (6-2) Constructing generalized advantage estimation And based on this generalized advantage, estimate In step (2), the network weights of the neural network are updated. To obtain the updated graph neural network;

[0074] (6-4) Use the joint reward obtained in step (6-2) to update the network weights of the Transformer pointer network in step (4) to obtain the updated Transformer pointer network.

[0075] Preferably, the base reward amount is calculated using the following formula:

[0076]

[0077] in, The baseline energy consumption is determined through the following process: First, the initial clustering structure obtained in step (1-2) is used as the current cluster partitioning structure, and all cluster-level sequence features are extracted from this cluster partitioning structure. All extracted cluster-level sequence features are then input into the Transformer pointer network in step (4) for processing to obtain a non-repeating flight access sequence. Subsequently, based on this non-repeating flight access sequence, the preset UAV flight path starting point and the initial cluster label set corresponding to the initial clustering structure are used. All nodes within all clusters and the preset UAV flight path endpoints are hierarchically expanded and topologically associated along the vertical dimension to construct a directed acyclic graph. Finally, the non-repeating flight access sequence is input into the A* algorithm in step (5) to perform multi-layer nested loop optimization iteration processing within the directed acyclic graph to obtain the baseline energy consumption. ;

[0078] Step (6-3) is as follows:

[0079] First, the joint feature vector of all nodes obtained in step (2-2) is... Input the value network from the PPO algorithm to obtain the baseline state value. Then, the updated cluster partitioning structure obtained in step (3-5) is processed using the methods described in steps (1) to (2-2) above to obtain the joint feature vector of all nodes after the update. And update the joint feature vector of all nodes Input the Critic network to obtain the updated baseline state value. ;

[0080] Then, according to the joint reward Baseline state value Updated baseline state value Obtaining timing difference error The timing difference error The calculation formula is:

[0081] ;

[0082] in This indicates the preset discount factor;

[0083] Subsequently, timing difference error was utilized. Discount Factor and the preset generalized advantage estimation decay factor Perform exponentially weighted summation to obtain the generalized dominance estimate. ;

[0084] Subsequently, this generalized advantage is used to estimate Constructing the policy pruning objective function of the PPO algorithm :

[0085] ;

[0086] in, Represents the expectation operator. This indicates that in graph neural networks, the weights... The output is the ratio of the probability of the macroscopic topological evolution action obtained in step (2) to the initial probability corresponding to the actual macroscopic topological evolution action in the probability distribution of the actual macroscopic topological evolution action obtained in step (2-4). This indicates the preset cropping threshold (its value ranges from 0.10 to 0.30, preferably 0.20). This means limiting the ratio to The cutoff function within the interval;

[0087] Finally, the objective function of this strategy is pruned. Regarding network weights in graph neural networks Calculate the partial derivative to obtain the policy gradient. Based on the policy gradient And combine the gradient ascent optimizer to optimize the network weights of the graph neural network. Perform backpropagation to update the network weights of the graph neural network. This is used to obtain the updated graph neural network.

[0088] According to another aspect of the present invention, a trajectory-aware dynamic topology adaptive UAV route planning system is provided, comprising the following modules:

[0089] The first module is used to acquire the original environmental data within the target area, extract the initial geographic coordinate set composed of the initial geographic coordinates of all the corresponding nodes from the original environmental data, and perform feature engineering processing on the initial geographic coordinate set to obtain the heterogeneous graph state tensor (HST).

[0090] The second module is used to input the heterogeneous graph state tensor obtained from the first module into the graph neural network to extract the topology of all nodes. Based on this topology, the PPO algorithm is optimized using a proximal strategy to obtain the macroscopic topological evolution actions. ;

[0091] The third module is used to process the initial cluster label set obtained in the first module. and cluster center coordinates Used as the current cluster partitioning structure Based on the macroscopic topological evolution actions obtained from the second module and the current cluster partitioning structure Furthermore, the environmental physical actuator is used to redistribute or reorganize all nodes extracted from the original environmental data in the first module to obtain the cluster partitioning structure.

[0092] The fourth module extracts all cluster-level sequence features from the cluster partitioning structure obtained in the third module, and processes these features using a Transformer pointer network to obtain a non-repeating flight access sequence. ,in The cluster label for the drone to access the k-th cluster;

[0093] The fifth module is used to determine the non-repeating flight access sequence obtained from the fourth module. A multi-level state search space is constructed, and the A* algorithm is used to perform multi-level nested loop optimization iteration within this multi-level state search space to obtain the optimal cluster head (CH) physical coordinate sequence. and global energy consumption ,in This represents the total number of clusters within the target area;

[0094] The sixth module is used to calculate the global energy consumption based on the fifth module. A joint reward is generated, and this joint reward is used to synchronously update the network weights of the graph neural network in the second module and the Transformer pointer network in the fourth module, respectively, to obtain the updated graph neural network and the Transformer pointer network.

[0095] The seventh module is used to determine whether the graph neural network updated in the sixth module and the updated Transformer pointer network have reached convergence, or whether the number of iterations has reached the preset maximum number of training rounds. If so, training is stopped, and the preset UAV flight path starting point and the optimal cluster head physical coordinate sequence obtained in the fifth module are used. The system splices together the preset drone flight path endpoints according to the order of time access to obtain the final drone flight path. The process ends when the drone flight path is accessed; otherwise, it returns to the first module.

[0096] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects:

[0097] 1. This invention, by employing steps (6-1) to (6-3), introduces a trajectory awareness mechanism within the mid-to-high-level reinforcement learning framework. This mechanism adaptively transforms the spatial flight distance and physical cost planned by the mid-level network into penalty terms, which are then passed back to the generalized advantage estimation of the high-level agent. This drives the ground cluster topology to actively change its physical shape, thereby shortening the air flight path. Therefore, this invention can solve the technical problems of existing traditional two-stage heuristic methods, such as the deep separation between cluster topology and flight path, and the tendency to get trapped in local optima.

[0098] 2. By employing steps (3-1) to (3-5), this invention utilizes an environmental physical actuator for dynamic dimensionality adjustment, allowing for adaptive cluster splitting and merging adjustments in the ground cluster topology. This enables the system to autonomously avoid multipath fading regions in radio communication, effectively mitigating the rapid energy depletion of some nodes and extending system endurance. Therefore, this invention solves the technical problems of existing traditional two-stage heuristic methods, such as lack of dynamic topology adaptation capabilities and inability to effectively cope with radio fading disasters.

[0099] 3. By employing steps (2-3) to (2-4), this invention reduces the traditional node-level cluster partitioning action to a five-dimensional macroscopic topological evolution action. Furthermore, it utilizes dynamic action masks to filter illegal states in the physical space and shrink the search boundary. This design prevents the drastic expansion of the decision state space at the algorithm's underlying level. Therefore, this invention can solve the technical problems of the curse of dimensionality and the exponential explosion of the action space faced by existing joint optimization methods based on deep reinforcement learning.

[0100] 4. By employing steps (2-5), this invention introduces a tabu table with a short-term memory structure to perform de-oscillation correction on the probability distribution of actual macroscopic topological evolution actions. This mechanism can track and restrict recently generated reciprocating inverse actions, prompting the model to explore entirely new effective topological solution spaces. Therefore, this invention can solve the technical problems of topological oscillation and model convergence difficulties in existing joint optimization methods based on deep reinforcement learning.

[0101] 5. By employing steps (4-1) to (4-4), this invention introduces zero-vector padding and padding mask techniques at the route planning layer, along with a dynamic visited mask, enabling the Transformer pointer network to handle cluster feature set sequences with dynamically changing lengths. This approach shields the interference of virtual placeholders at the underlying level and avoids repeated node visits, thereby ensuring the robustness of the algorithm in complex IoT environments.

[0102] 6. This invention, by employing steps (5-1) to (5-15), constructs a directed acyclic graph of a vertically multi-layered state space, utilizes the A* algorithm to perform multi-layered nested loop optimization within the graph, and explicitly converts the theoretical shortest flight distance into physical energy consumption through a preset unit flight energy consumption coefficient. This approach overcomes the drawbacks of uncontrollable energy consumption caused by traditional algorithms relying on random or black-box elections, achieving the election of the optimal cluster head physical coordinate sequence with the lowest global energy consumption;

[0103] 7. The implementation of this invention is simple. The entire system adopts a highly decoupled and logically closed-loop three-layer joint optimization architecture design, which makes the interaction boundary between microscopic physical computation and macroscopic policy evolution very clear. Its low computational complexity and memory overhead make it very suitable for lightweight deployment and online training on general IoT base station gateways or embedded platforms of commercial drones;

[0104] 8. This invention has wide applicability. The dynamic topology adaptive variable-dimensional adjustment mechanism, mask constraint framework, and multi-layer nested path optimization architecture proposed in this invention have strong mathematical topology generalization ability. It is not only applicable to data acquisition route planning in wireless sensor networks, but can also be extended to various complex intelligent scheduling scenarios such as UAV trajectory scheduling in mobile edge computing (MEC), multi-agent logistics distribution in smart cities, disaster search and rescue, and multi-dimensional spatial dynamic routing control. Attached Figure Description

[0105] Figure 1 This is a schematic diagram of the overall process of the dynamic topology adaptive UAV route planning method based on trajectory perception of the present invention;

[0106] Figure 2 This is a schematic diagram of the architecture of the high-level graph neural network and PPO agent of the present invention, as well as the masking process.

[0107] Figure 3 This is a schematic diagram of the structure of generating flight paths by combining the mid-layer Transformer pointer network with dynamic filling masks in this invention;

[0108] Figure 4 This is a schematic diagram of the construction of the multi-layer state search space and cluster head election of the underlying A* algorithm of this invention. Detailed Implementation

[0109] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0110] The technical terms of this invention will be explained and described below:

[0111] Tag list: It is a structure with short-term memory used to record solutions (or moves / actions) that the algorithm has visited in recent search processes and prohibits the revisiting of these recorded states in the next few iterations. The main purpose of the tabu list is to prevent the search process from getting stuck in local optima and to avoid meaningless loops in the solution space, thereby guiding the algorithm to explore a wider new region.

[0112] The basic idea of ​​this invention is to provide a highly decoupled and logically closed-loop three-layer joint optimization architecture to synergistically reduce the overall network energy consumption. Addressing the problem of fragmented clustering and flight paths in traditional methods, this invention utilizes a graph neural network combined with dynamic masks and tabu lists at the macro-topology evolution layer to output legitimate topology dimension-changing actions, curbing state space expansion and decision oscillations from the bottom layer. At the mid-level flight path planning layer, a Transformer pointer network with a filling mask mechanism is introduced to handle dynamically changing cluster features, generating non-repeating Hamiltonian paths through autoregression. At the bottom-level physical optimization layer, a multi-layer directed acyclic graph of state space is constructed based on the flight path. The A* algorithm is used to accurately select the cluster head coordinate sequence with the lowest global energy consumption under stringent communication fading and flight energy consumption constraints. Finally, the system adaptively transforms the precisely calculated physical energy consumption and topology overrun penalties at the bottom layer into joint rewards, which are then passed back to the higher-level agents to update network weights. This drives the ground topology to proactively reshape physical boundaries to shorten air routes, maximizing end-to-end system endurance.

[0113] like Figure 1 As shown, this invention provides a trajectory-aware dynamic topology adaptive UAV route planning method, comprising the following steps:

[0114] (1) Obtain the original environmental data in the target area, extract the initial geographic coordinate set of all the corresponding nodes from the original environmental data, and perform feature engineering on the initial geographic coordinate set to obtain the heterogeneous state tensor (HST).

[0115] This step specifically includes the following sub-steps:

[0116] (1-1) Obtain the original environmental data within the target area, and extract the initial geographic coordinates of each node corresponding to the original environmental data. The initial geographic coordinates of all nodes corresponding to the original environmental data constitute the initial geographic coordinate set. , where N represents the total number of all nodes corresponding to the original environment data, and i∈[1,N].

[0117] Specifically, nodes can be IoT terminals such as temperature and humidity meters or pressure sensors; the initial geographic coordinates of nodes can be obtained using the BeiDou Navigation Satellite System or the Global Positioning System (GPS), or through Received Signal Strength Indicator (RSSI) ranging technology.

[0118] (1-2) Perform initial clustering processing on the initial geographic coordinate set V obtained in step (1-1) to obtain the initial cluster label set. The cluster center coordinates of the corresponding cluster are used as reference target points for UAV flight access and data collection.

[0119] In classic wireless sensor network clustering routing protocols (such as the industry's most famous LEACH protocol), the first phase of network operation is explicitly defined as the "Cluster Establishment / Initialization Phase," and the clusters generated in this phase are called "Initial Clustering."

[0120] Specifically, the initial clustering process in this step uses the K-means algorithm, based on the preset number of clusters. (Its initial value is 5% to 15% of the total number of nodes corresponding to the original environment data, preferably 10%) Divide all N nodes into There are 3 clusters (each with an initial cluster label), and an initial cluster label is assigned to the i-th node. Then, the cluster center coordinates corresponding to the k-th cluster are obtained. The initial cluster labels corresponding to all nodes constitute the initial cluster label set. , where k∈[1, the total number of clusters K obtained from the initial clustering process in step (1-2).

[0121] (1-3) The initial cluster label set obtained from step (1-2) Get the number of nodes in the k-th cluster And based on the number of nodes Obtain the node load variance within the target area (It is used to characterize the network balance).

[0122] Specifically, node load variance The following formula is used for calculation:

[0123] ;

[0124] in This represents the average number of nodes contained in all clusters, and has... =N / K;

[0125] (1-4) The initial cluster label set obtained from step (1-2) The cluster center coordinates corresponding to each cluster, and the initial geographic coordinate set obtained in step (1-1). Get the i-th node and its corresponding cluster center. First distance The distance from the i-th node to its corresponding cluster center Adjacent cluster centers The second distance And based on this second distance Distance from the first The ratio is used to obtain the margin score of the i-th node. Where A and B both ∈ [1, K];

[0126] Specifically, the edge index is calculated using the following formula:

[0127] ;

[0128] The advantage of steps (1-4) is that by quantifying the degree of node detachment at the cluster boundary, it provides a clear microscopic physical basis for subsequent high-level agents to adjust the topology boundary.

[0129] (1-5) Obtain the edge degree index from all N nodes All nodes smaller than a preset threshold are selected as boundary nodes, and all boundary nodes constitute a boundary node feature list. ;

[0130] Specifically, the preset threshold value ranges from 1.00 to 1.20, preferably 1.10.

[0131] The advantage of steps (1-5) is that by filtering out key topological change triggers through physical rules, the neural network can focus on processing boundary nodes, thereby effectively reducing the complexity of subsequent calculations and improving the accuracy of topological evolution.

[0132] (1-6) Using the geographic coordinate set obtained in step (1-1) The edge degree index of the i-th node obtained in step (1-4) Constructing a sparse adjacency matrix (It is used to characterize the sensor network topology within the target area).

[0133] Specifically, the constructed sparse adjacency matrix The dimension is , of which Line 1 Column elements Used to characterize the The node and the first The topological connection relationship between the nth nodes, if the nth node... The node and the first If the nth node corresponds to the same cluster, and the spatial distance between them is less than a preset communication distance threshold (its value ranges from 1.5×r to 3.5×r, preferably 2.2×r, where r is the average sensing radius of the node), then the nth node in the sparse adjacency matrix... Line number Column elements If the first The node and the first If the nth node belongs to different clusters, but both are boundary nodes, and their spatial distance is less than a preset communication distance threshold, then the nth node in the sparse adjacency matrix... Line number Column elements If the first The node and the first If the nth node does not satisfy any of the above conditions, then the nth node in the sparse adjacency matrix... Line number Column elements , where j∈[1,N];

[0134] (1-7) Using the initial geographic coordinate set obtained in step (1-1) Obtain the spatial distance between any two nodes within the target area. The boundary node feature list BF obtained in step (1-5) and the number of clusters obtained in step (1-2) are compared. and the load variance obtained in steps (1-3) A fusion encapsulation process is performed to obtain the heterogeneous graph state tensor (which is used to characterize the topology of all nodes).

[0135] The advantage of step (1-7) is that it transforms discrete physical coordinates into high-dimensional graph feature tensors (HSTs), enabling graph neural networks (GNNs) to simultaneously capture the individual attributes of nodes and the topological distribution of the entire network.

[0136] (2) Input the heterogeneous graph state tensor obtained in step (1) into a graph neural network (e.g., Figure 2 As shown in the figure, the topology of all nodes is extracted, and based on this topology, the macroscopic topology evolution action is obtained using the Proximal Policy Optimization (PPO) algorithm. (It is used to change the topology of all nodes, thereby reducing the cost of subsequent drone data collection flights.)

[0137] The graph neural network used in this invention is a multi-layer graph attention network (GAT).

[0138] This step specifically includes the following sub-steps:

[0139] (2-1) Input the heterogeneous graph state tensor obtained in step (1-7) into the graph neural network and perform message passing processing to obtain the hidden feature vector of each node. .

[0140] Specifically, graph neural networks can automatically allocate attention weights based on the edge degree differences between nodes, thus strengthening the focus on boundary nodes.

[0141] (2-2) Hidden feature vectors of all nodes obtained in step (2-1) Global average pooling (Readout / Pooling) is performed to extract the global topological latent vectors of all nodes. And compare it with the number of clusters obtained in step (1-2). and the network load variance obtained in steps (1-3) Concatenate the data to obtain the joint feature vector of all nodes. .

[0142] (2-3) Combine the joint feature vectors of all nodes obtained in step (2-2) The action policy network (i.e., the actor network) in the PPO algorithm is input into a fully connected layer for computation to obtain the log probability distribution of the original five-dimensional macroscopic topological evolution action of each node. .

[0143] Specifically, the five-dimensional original macroscopic topological evolution actions include two categories: the first category is node-level actions, including moving boundary nodes, swapping boundary nodes, and maintaining the status quo (No-op); the second category is cluster-level actions, including merging two adjacent clusters and splitting overloaded clusters. For the first category of node-level actions, sampling is performed directly based on the output probability of the node; for the second category of cluster-level actions, average pooling or majority voting mechanisms are used to aggregate the probabilities of actions corresponding to all boundary nodes within the cluster to determine whether the cluster to which the node belongs should be merged or split.

[0144] (2-4) Log probability distribution of the five-dimensional primitive macroscopic topological evolution action obtained in step (2-3) Mask filtering is performed to obtain the probability distribution of actual macroscopic topological evolution actions. .

[0145] Specifically, since the original macroscopic topological evolution action probabilities output by the neural network may violate objective constraints in physical space, a dynamic mask is constructed to set the probability of illegal actions to zero: if the current cluster number has reached the preset maximum allowed cluster number. If the current number of clusters has decreased to the preset minimum allowed number of clusters, then the probability of the "split" action will be set to zero using a mask. If the boundary node feature list extracted in step (1-5) is empty (i.e., there are no boundary nodes in the entire network that meet the conditions), then the probabilities of the "Move" and "Swap" actions are set to zero. After completing the masking filtering of illegal actions, the probabilities of the remaining legal macroscopic topology evolution actions are re-normalized to ensure that the sum of the distribution of macroscopic topology evolution actions is 1.

[0146] The advantage of steps (2-4) is that by using hard constraints based on physical rules, the search space of reinforcement learning is effectively reduced, preventing the generation of illegal actions.

[0147] (2-5) Using a pre-defined tabu list, analyze the probability distribution of the actual macroscopic topological evolutionary actions obtained in step (2-4). De-oscillation correction is performed to generate the final macroscopic topological evolution action. .

[0148] Specifically, the taboo list records the most recent (Its value ranges from 5 to 12, preferably 7) the reverse action within the step (such as merging immediately after splitting);

[0149] The advantage of step (2-5) is that it eliminates the problem of back-and-forth oscillations in the topological evolution process from a logical perspective, thus accelerating the convergence of the model.

[0150] The advantage of the above sub-steps (2-1) to (2-5) is that by combining the graph neural network with the PPO algorithm and cascading dynamic action mask and tabu table correction at the action output end, macroscopic topological evolution decision-making under the constraints of underlying physical laws is realized, effectively preventing action dimension explosion.

[0151] (3) The initial cluster label set obtained in step (1) and cluster center coordinates Used as the current cluster partitioning structure Based on the macroscopic topological evolution action obtained in step (2) and the current cluster partitioning structure The Environmental Physical Executor is used to redistribute or reorganize all nodes extracted from the original environmental data in step (1) to obtain the cluster partitioning structure.

[0152] This step specifically includes the following sub-steps:

[0153] (3-1) The initial cluster label set obtained in step (1) coordinates of cluster center Used as the current cluster partitioning structure ;

[0154] (3-2) Based on the macroscopic topological evolution action obtained in step (2) The current cluster partitioning structure obtained in step (3-1) Perform dynamic dimension transformation to obtain the updated cluster count. .

[0155] Specifically, dynamic dimensionality adjustment refers to dynamically adjusting the number of clusters based on macroscopic topological evolution actions to obtain an updated number of clusters. If the macroscopic topological evolution action is "splitting overloaded clusters," then the local K-means algorithm is used to divide the clusters corresponding to the macroscopic topological evolution action into two sub-clusters, and the updated number of clusters is obtained in this case. +1; If the macro-topological evolution action is "merge adjacent clusters", then the two clusters corresponding to the macro-topological evolution action will be merged into one large cluster, and the updated cluster count will be... -1.

[0156] (3-3) Based on the macroscopic topological evolution action obtained in step (2) The boundary nodes obtained in steps (1-5) are then subjected to topological affiliation reallocation to obtain the updated cluster label set. .

[0157] Specifically, if the macro-topology evolution action is "moving boundary nodes" or "swapping boundary nodes", the environmental physics actuator will directly modify the corresponding node's affiliation cluster label in the underlying database.

[0158] The advantage of step (3-3) is that the environmental physical actuator accurately implements the macroscopic instructions of the high-level agent, and completes the reconstruction of the logical topology without changing the physical hardware location.

[0159] (3-4) The updated cluster label set obtained in step (3-3) and the initial geographic coordinate set obtained in step (1-1) Perform a centroid update to obtain the updated cluster center coordinates. and the number of nodes in each cluster .

[0160] (3-5) The updated cluster count obtained in step (3-2) The updated cluster label set obtained in step (3-3) and the updated cluster center coordinates obtained in steps (3-4) The data is packaged and encapsulated to obtain the cluster partitioning structure.

[0161] The advantage of the above sub-steps (3-1) to (3-5) is that they provide a dynamic variable-dimensional mechanism that can adaptively perform cluster splitting and merging, enabling the system to autonomously avoid the multipath fading region of radio communication, thereby alleviating the phenomenon of rapid energy depletion of nodes.

[0162] (4) Extract all cluster-level sequence features from the cluster partitioning structure obtained in step (3), and use a Transformer pointer network (such as...) Figure 3 (As shown) All cluster-level sequence features are processed to obtain a non-repeating flight access sequence. ,in The cluster label for the drone to access the k-th cluster;

[0163] This step specifically includes the following sub-steps:

[0164] (4-1) Extract the geometric centroid coordinates of each cluster from the cluster partitioning structure updated in step (3-5). and the number of nodes included Both are then input into a linear projection layer to obtain the initial feature sequence. ,in This represents the node embedding vector of the k-th cluster.

[0165] Specifically, the node embedding vector of the k-th cluster The following formula is used for calculation:

[0166] ;

[0167] in, Represents a shared linear projection layer. Representing the The cluster center of each cluster, Representing the The number of nodes contained in a cluster These are learnable bias parameters.

[0168] (4-2) The initial feature sequence obtained in step (4-1) Zero-vector padding is performed to obtain the cluster feature tensor. And based on the cluster feature tensor The padding mask is generated by the index positions of the valid elements (which refer to the node embedding vectors corresponding to each cluster in the updated cluster partitioning structure obtained in step (4-1)) and the padding placeholders (which refer to the zero vectors added during the zero vector supplementation process). ,in This indicates the preset maximum number of clusters;

[0169] Specifically, the padding mask is a tensor related to the cluster features. The corresponding binary state tensor has a mask value of 0 (or False) for the index position of the valid element and a mask value of 1 (or True) for the index position of the padding placeholder. This padding mask is used to force the attention weight corresponding to the padding position to be converted to negative infinity in the subsequent multi-head self-attention calculation process, thereby shielding the virtual zero vector from the interference of feature fusion at the underlying level.

[0170] The advantage of step (4-2) is that by generating a padding mask corresponding to the zero vector, the interference of virtual placeholders on attention feature fusion is forcibly shielded at the bottom layer, ensuring the accuracy of feature extraction when the sequence length changes dynamically.

[0171] (4-3) The cluster feature tensor obtained in step (4-2) Input the Transformer encoder to compute the cluster feature tensor using a multi-head self-attention (MHSA) mechanism. The spatial dependency weights between each pair of clusters are used to determine the spatial dependency weights of the cluster feature tensor. We perform weighted feature fusion on the effective elements in the dataset to obtain the updated local node embedding vector for each cluster, and then concatenate the updated node embedding vectors for all clusters to obtain the cluster feature tensor. The corresponding encoded hidden state matrix Eh ,in This indicates the preset maximum number of clusters. This represents the dimension of the hidden layer features output by the Transformer encoder.

[0172] (4-4) The hidden state matrix Eh obtained in step (4-3) is processed by the pointer decoder to perform autoregressive sequence decoding to generate a non-repeating flight access sequence covering all K clusters. ,in It is the cluster label of the kth cluster visited by the drone.

[0173] Specifically, this step involves first initializing a visited state vector with all zeros as a visited mask, used to track and record the UAV's historical visit footprints during flight path planning in real time; then, in the current first decoding step, the pointer decoder is used to encode the hidden state matrix obtained in step (4-3). Attention is calculated to obtain the attention score matrix for each cluster. Then, using the filling mask and visited mask obtained in step (4-2), the attention score matrix is ​​masked, assigning negative infinity (i.e., ...) to the scores of the virtual clusters representing the filling region and the visited clusters. This process masks invalid paths to obtain a masked attention score matrix. Subsequently, the masked attention score matrix is ​​normalized to obtain the transition probabilities corresponding to the first decoding step. ,in The cluster label of the first target cluster to be accessed is ( The updated cluster label set obtained in step (3-3) The probability of a transition is calculated, and based on this transition probability, a maximum probability greedy selection or stochastic sampling strategy is used to select the target cluster to be visited by the UAV from the clusters whose corresponding state position is 0 in the visited mask. Then, the state position corresponding to the selected target cluster in the visited mask is set to 1. This process is repeated for the 2nd, 3rd, ... Kth decoding steps until all K decoding steps have been processed, resulting in the sequentially selected target clusters. The target clusters constitute a non-repeating flight access sequence.

[0174] The advantage of step (4-4) is that by introducing and dynamically maintaining the visited mask inside the autoregressive decoding, the possibility of the UAV repeatedly visiting the same cluster is eliminated from the bottom layer of the graph theory algorithm. Without any post-processing overhead, it ensures that the network directly generates a completely legal Hamiltonian path without local dead loops, thereby eliminating the waste of maneuvering energy caused by overlapping flight paths of the UAV.

[0175] The advantage of the above sub-steps (4-1) to (4-4) is that they innovatively combine zero vector padding, padding mask and visited mask techniques, enabling the middle layer Transformer pointer network to adapt to changes in input dimension caused by dynamic splitting and merging of the underlying topology, thereby improving the robustness of the system in the ever-changing Internet of Things environment.

[0176] (5) Based on the non-repeating flight access sequence obtained in step (4) Construct a multi-level state search space and utilize the A* algorithm (such as...) Figure 4 As shown, a multi-layered nested loop optimization iteration process is performed within this multi-layered state search space to obtain the optimal cluster head (CH) physical coordinate sequence. and global energy consumption ,in This indicates the total number of clusters within the target area.

[0177] This step specifically includes the following sub-steps:

[0178] (5-1) Based on the non-repeating flight access sequence generated in step (4-4) The preset drone flight path starting point and initial cluster label set will be used. All nodes within all clusters and the preset drone flight path endpoints are hierarchically expanded and topologically associated along the vertical dimension to construct a Directed Acyclic Graph (DAG). The vertical topology of this DAG is determined by... It consists of several levels, where level 0 is the starting point of the drone flight path, and levels 1 through 2 are... The layers correspond to the non-repeating flight access sequence in sequence. Each cluster in, the first The layer represents the endpoint of the drone's flight path.

[0179] (5-2) Initialize the priority queue (Open List) of the A* algorithm, take the node of layer 0 of the DAG graph obtained in step (5-1) (i.e. the starting point of the UAV flight path) as the initial node to be evaluated, and set the historical actual cost of the initial node to be evaluated. The sequence of non-repeating flight visits to the initial node to be evaluated is obtained based on Euclidean distance. The shortest theoretical flight distance from the cluster center of all clusters to the end of the UAV's flight path is calculated. This shortest theoretical flight distance is then multiplied by a preset unit flight energy consumption coefficient for the UAV to obtain the ideal flight energy consumption of the UAV. This ideal flight energy consumption is then used as the initial heuristic energy consumption estimate for the initial node to be evaluated. Using the formula Obtain the comprehensive evaluation value of the initial node to be evaluated. And add the initial node to be evaluated to the priority queue.

[0180] (5-3) Check if the priority queue is empty. If it is, it means that the globally optimal path has not been found, and the process ends; otherwise, obtain the comprehensive evaluation value from the priority queue. The smallest node is used as the current traversed node. Then proceed to step (5-4).

[0181] (5-4) Determine the currently traversed node Does it belong to the first DAG graph? Layer (i.e., determining the currently traversed node) (Is this the destination of the drone's flight path?) If yes, it means that the globally optimal path has been found, and then proceed to step (5-15); otherwise, proceed to step (5-5).

[0182] (5-5) Get the current traversed node The level number in the DAG diagram And extract the first element from the DAG graph. All nodes contained in the next layer are used as a candidate cluster head set, and the total number of candidate cluster heads in this set is obtained. .

[0183] (5-6) Set up a candidate cluster head traversal counter .

[0184] (5-7) Determine the counter Is it greater than the total number of candidate cluster heads? If it is (representing the first) If the candidate cluster head set of the layer has been traversed, then return to step (5-3); otherwise, extract the first... The node is selected as the current candidate cluster head. Then proceed to steps (5-8).

[0185] (5-8) Obtain the first digit in the DAG graph. In addition to the candidate cluster head, the layer All other nodes and their total number Set a node traversal counter and initialize the cumulative communication energy consumption within the current layer. .

[0186] (5-9) Determine the node traversal counter Is it greater than the total number of other nodes? If yes, proceed to step (5-12); otherwise, obtain the first element in the DAG graph. The first in the layer Each node is selected, and the process proceeds to step (5-10).

[0187] (5-10) Calculate the first digit of the DAG graph. The first in the layer Each node and candidate cluster head Spatial communication distance between Determine the spatial communication distance Is it less than or equal to the multipath fading threshold? If so, then the first signal propagation step is calculated based on the free space signal propagation model. Each node and candidate cluster head Energy consumption for communication between And accumulate the communication energy consumption to the first In the cumulative communication energy consumption of the layer, that is Then proceed to step (5-11); otherwise, calculate the first step based on the multipath fading signal propagation model. Each node and candidate cluster head Energy consumption for communication between And accumulate the communication energy consumption to the first In the cumulative communication energy consumption of the layer, that is Then proceed to step (5-11), where the multipath fading threshold is... , This represents the power amplifier energy consumption coefficient under the free-space signal propagation model. This represents the power amplifier energy consumption coefficient under the multipath fading signal propagation model;

[0188] (5-11) Set a node traversal counter Then return to steps (5-9).

[0189] (5-12) Based on the drone's current traversal nodes Fly to candidate cluster head Spatial transfer distance acquisition mechanical flight energy consumption To obtain the drone in the candidate cluster head Hovering power consumption during wireless communication above According to the energy consumption of this mechanical flight Hovering energy consumption and the first one obtained in step (5-10) Cumulative communication energy consumption of the layer Get the nodes from the current traversal Transfer to candidate cluster head Weighted step cost .

[0190] Specifically, weighted step cost The following formula is used for calculation:

[0191] ;

[0192] in, and These are dimensionless weighted coefficients. The value range is from 0.1 to 0.9, preferably 0.5. The value range is from 0.1 to 0.9, preferably 0.5.

[0193] (5-13) Set the currently traversed node The actual historical cost The weighted step cost obtained in step (5-12) Perform summation to obtain the number of nodes reaching the candidate cluster head. The actual cost ; Obtain the candidate cluster head Sequential non-repeating flight access sequence The shortest theoretical flight distance from the cluster center of all remaining cluster labels to the end of the UAV's flight path is calculated. This shortest theoretical flight distance is then multiplied by a preset UAV unit flight energy consumption coefficient to obtain the UAV's ideal flight energy consumption, which is then used as the candidate cluster head. Heuristic estimation of energy consumption Using the formula Obtain the candidate cluster head Comprehensive assessment value , will candidate cluster heads Add it to the priority queue and proceed to step (5-14).

[0194] (5-14) Set up candidate cluster head counters Then return to steps (5-7).

[0195] (5-15) Backtrack along the globally optimal path in the priority queue in the DAG graph to obtain the optimal cluster head physical coordinate sequence. and the endpoint of the drone's flight path (i.e., the first...) The historical actual cost corresponding to the layer node is used as the global energy consumption. Output.

[0196] The advantage of the above sub-steps (5-1) to (5-15) is that by constructing a directed acyclic graph of a multi-layer state space and using the A* algorithm to perform precise optimization under the constraints of multipath fading communication model and mechanical flight energy consumption, the uncontrollable energy consumption caused by the reliance on random or black-box elections in traditional algorithms is overcome, and the optimal cluster head physical coordinate sequence election with the lowest global energy consumption is achieved.

[0197] (6) The global energy consumption obtained from step (5) A joint reward is generated, and the joint reward is used to synchronously update the network weights of the graph neural network in step (2) and the Transformer pointer network in step (4) respectively, so as to obtain the updated graph neural network and the Transformer pointer network respectively.

[0198] This step specifically includes the following sub-steps:

[0199] (6-1) The global energy consumption obtained from step (5-15) Obtain basic reward amount .

[0200] Specifically, the base reward is calculated using the following formula:

[0201]

[0202] in, The baseline energy consumption is determined through the following process: First, the initial clustering structure obtained in step (1-2) is used as the current cluster partitioning structure, and all cluster-level sequence features are extracted from this cluster partitioning structure. All extracted cluster-level sequence features are then input into the Transformer pointer network in step (4) for processing to obtain a non-repeating flight access sequence. Subsequently, based on this non-repeating flight access sequence, the preset UAV flight path starting point and the initial cluster label set corresponding to the initial clustering structure are used. All nodes within all clusters and the preset UAV flight path endpoints are hierarchically expanded and topologically associated along the vertical dimension to construct a directed acyclic graph. Finally, the non-repeating flight access sequence is input into the A* algorithm in step (5) to perform multi-layer nested loop optimization iteration processing within the directed acyclic graph to obtain the baseline energy consumption. .

[0203] (6-2) Calculate the node load variance within the target region based on the updated cluster partitioning structure obtained in step (3). Extract the boundary nodes within each cluster in the updated cluster partitioning structure, and calculate the spatial distance from each boundary node to the cluster center coordinates corresponding to that cluster. Then, select the boundary nodes whose spatial distances are greater than the multipath fading threshold obtained in step (5-10). All distance overflows are summed, and the sum is used as a penalty for exceeding the cluster radius limit within the target area. and node load variance Cluster radius exceeding the limit penalty and the basic reward obtained in step (6-1) Perform weighted synthesis to obtain joint rewards .

[0204] (6-3) The combined reward obtained from step (6-2) Constructing generalized advantage estimation (Generalized Advantage Estimation, or GAE for short), and estimates based on this generalized advantage. In step (2), the network weights of the neural network are updated. This is used to obtain the updated graph neural network.

[0205] This step is specifically as follows:

[0206] First, the joint feature vector of all nodes obtained in step (2-2) is... Input the value network (i.e., the Critic network) in the PPO algorithm to obtain the baseline state value. Then, the updated cluster partitioning structure obtained in step (3-5) is processed using the methods described in steps (1) to (2-2) above to obtain the joint feature vector of all nodes after the update. And update the joint feature vector of all nodes Input the Critic network to obtain the updated baseline state value. ;

[0207] Then, according to the joint reward Baseline state value Updated baseline state value Obtaining timing difference error The timing difference error The calculation formula is:

[0208] ;

[0209] in This represents a preset discount factor, the value of which ranges from 0.90 to 0.99, preferably 0.99;

[0210] Subsequently, timing difference error was utilized. Discount Factor and the preset generalized advantage estimation decay factor (The value ranges from 0.90 to 0.99, preferably 0.95) An exponentially weighted sum is performed to obtain the generalized dominance estimate. ;

[0211] Subsequently, this generalized advantage is used to estimate Constructing the policy pruning objective function of the PPO algorithm :

[0212] ;

[0213] in, Represents the expectation operator. This indicates that the graph neural network has network weights The output is the ratio of the probability of the macroscopic topological evolution action obtained in step (2) to the initial probability corresponding to the actual macroscopic topological evolution action in the probability distribution of the actual macroscopic topological evolution action obtained in step (2-4). This indicates the preset cropping threshold (its value ranges from 0.10 to 0.30, preferably 0.20). This means limiting the ratio to The cutoff function within the interval;

[0214] Finally, the objective function of this strategy is pruned. Regarding network weights in graph neural networks Calculate the partial derivative to obtain the policy gradient. Based on the policy gradient And combine the gradient ascent optimizer to optimize the network weights of the graph neural network. Perform backpropagation to update the network weights of the graph neural network. This is used to obtain the updated graph neural network.

[0215] (6-4) Use the joint reward obtained in step (6-2) to update the network weights of the Transformer pointer network in step (4) to obtain the updated Transformer pointer network.

[0216] The advantage of the sub-steps (6-1) to (6-4) above lies in the introduction of a trajectory-aware mechanism for vertical crossover, which adaptively transforms the total network energy consumption and topology overrun penalty calculated by the underlying A* algorithm into a joint reward, which in turn guides the weight update of the mid-to-high-level network. This mechanism can drive the ground topology to actively reshape the physical boundary to shorten the air route, breaking the technical barrier of clustering and routing separation in the traditional two-stage heuristic method.

[0217] (7) Determine whether the updated graph neural network and the updated Transformer pointer network in step (6) have reached convergence, or whether the number of iterations has reached the preset maximum number of training rounds (the value ranges from 10,000 to 50,000, preferably 30,000). If so, stop training and set the preset UAV flight path starting point and the optimal cluster head physical coordinate sequence obtained in step (5). The preset drone flight path endpoints are spliced ​​together according to the order of time access to obtain the final drone flight path. The process ends; otherwise, return to step (1).

[0218] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A trajectory-aware, dynamic topology-adaptive UAV route planning method, characterized in that, Includes the following steps: (1) Obtain the original environmental data in the target area, extract the initial geographic coordinate set of all the corresponding nodes from the original environmental data, and perform feature engineering on the initial geographic coordinate set to obtain the heterogeneous graph state tensor HST. (2) Input the heterogeneous graph state tensor obtained in step (1) into the graph neural network to extract the topology of all nodes. Based on the topology, the PPO algorithm is optimized using the near-end strategy to obtain the macroscopic topological evolution action. ; (3) The initial cluster label set obtained in step (1) and cluster center coordinates Used as the current cluster partitioning structure Based on the macroscopic topological evolution action obtained in step (2) and the current cluster partitioning structure And use the environmental physical actuator to redistribute or regroup all nodes extracted from the original environmental data in step (1) to obtain the cluster partitioning structure; (4) Extract all cluster-level sequence features from the cluster partitioning structure obtained in step (3), and process all cluster-level sequence features using a Transformer pointer network to obtain a non-repeating flight access sequence. ,in The cluster label for the drone to access the k-th cluster; (5) Based on the non-repeating flight access sequence obtained in step (4) A multi-level state search space is constructed, and the A* algorithm is used to perform multi-level nested loop optimization iteration within this multi-level state search space to obtain the optimal cluster head (CH) physical coordinate sequence. and global energy consumption ,in This represents the total number of clusters within the target area; (6) The global energy consumption obtained from step (5) Generate a joint reward and use the joint reward to synchronously update the network weights of the graph neural network in step (2) and the Transformer pointer network in step (4) respectively, so as to obtain the updated graph neural network and the Transformer pointer network respectively; (7) Determine whether the updated graph neural network and the updated Transformer pointer network in step (6) have reached convergence, or whether the number of iterations has reached the preset maximum number of training rounds. If so, stop training and set the preset UAV flight path starting point and the optimal cluster head physical coordinate sequence obtained in step (5). The preset drone flight path endpoints are spliced ​​together according to the order of time access to obtain the final drone flight path. The process ends; otherwise, return to step (1).

2. The trajectory-aware dynamic topology adaptive UAV route planning method according to claim 1, characterized in that, Step (1) specifically includes the following sub-steps: (1-1) Obtain the original environmental data within the target area, and extract the initial geographic coordinates of each node corresponding to the original environmental data. The initial geographic coordinates of all nodes corresponding to the original environmental data constitute the initial geographic coordinate set. , where N represents the total number of all nodes corresponding to the original environment data, and i∈[1,N]; (1-2) Perform initial clustering processing on the initial geographic coordinate set V obtained in step (1-1) to obtain the initial cluster label set. The cluster center coordinates of the cluster and its corresponding cluster are used as reference target points for UAV flight access and data collection; (1-3) The initial cluster label set obtained from step (1-2) Get the number of nodes in the k-th cluster And based on the number of nodes Obtain the node load variance within the target area ; (1-4) The initial cluster label set obtained from step (1-2) The cluster center coordinates corresponding to each cluster, and the initial geographic coordinate set obtained in step (1-1). Get the i-th node and its corresponding cluster center. First distance The distance from the i-th node to its corresponding cluster center Adjacent cluster centers The second distance And based on this second distance Distance from the first The ratio is used to obtain the edge degree index of the i-th node. Where A and B both ∈ [1, K]; (1-5) Obtain the edge degree index from all N nodes All nodes smaller than a preset threshold are designated as boundary nodes, and all boundary nodes constitute a boundary node feature list. ; (1-6) Using the geographic coordinate set obtained in step (1-1) The edge degree index of the i-th node obtained in steps (1-4) Constructing a sparse adjacency matrix ; (1-7) Using the initial geographic coordinate set obtained in step (1-1) Obtain the spatial distance between any two nodes within the target area. The boundary node feature list BF obtained in step (1-5) and the number of clusters obtained in step (1-2) are compared. and the load variance obtained in steps (1-3) A fusion encapsulation process is performed to obtain the heterogeneous graph state tensor.

3. The trajectory-aware dynamic topology adaptive UAV route planning method according to claim 2, characterized in that, The initial clustering process in steps (1-2) uses the K-means algorithm, based on the preset number of clusters. Divide all N nodes into There are 3 clusters, each with an initial cluster label. An initial cluster label is assigned to the i-th node. Then, the coordinates of the cluster center corresponding to the k-th cluster are obtained. The initial cluster labels corresponding to all nodes constitute the initial cluster label set. , where k∈[1, the total number of clusters K obtained from the initial clustering process in step (1-2)], and the initial value of K is 5% to 15% of the total number of all nodes corresponding to the original environment data; Node load variance in steps (1-3) The following formula is used for calculation: ; in This represents the average number of nodes contained in all clusters, and has... =N / K; The edge index in steps (1-4) is calculated using the following formula: ; The sparse adjacency matrix constructed in steps (1-6) The dimension is , of which Line 1 Column elements Used to characterize the The node and the first The topological connection relationship between the nth nodes, if the nth node... The node and the first If the nth node corresponds to the same cluster, and the spatial distance between them is less than a preset communication distance threshold, then the nth node in the sparse adjacency matrix... Line 1 Column elements If the first The node and the first If the nth node belongs to different clusters, but both are boundary nodes, and their spatial distance is less than a preset communication distance threshold, then the nth node in the sparse adjacency matrix... Line 1 Column elements If the first The node and the first If the nth node does not satisfy any of the above conditions, then the nth node in the sparse adjacency matrix... Line 1 Column elements , where j∈[1,N].

4. The trajectory-aware dynamic topology adaptive UAV route planning method according to any one of claims 1 to 3, characterized in that, Graph neural networks are multi-layer graph attention networks (GAT). Step (2) specifically includes the following sub-steps: (2-1) Input the heterogeneous graph state tensor obtained in step (1-7) into the graph neural network and perform message passing processing to obtain the hidden feature vector of each node. ; (2-2) Hidden feature vectors of all nodes obtained in step (2-1) Global average pooling is performed to extract the global topological latent vectors of all nodes. And compare it with the number of clusters obtained in step (1-2). and the network load variance obtained in steps (1-3) The nodes are concatenated to obtain the joint feature vector of all nodes. ; (2-3) Combine the joint feature vectors of all nodes obtained in step (2-2) The Actor network in the PPO algorithm is input for fully connected layer operations to obtain the log probability distribution of the five-dimensional original macroscopic topological evolution action of each node. ; (2-4) Log probability distribution of the five-dimensional primitive macroscopic topological evolution action obtained in step (2-3) Masking and filtering are performed to obtain the probability distribution of actual macroscopic topological evolution actions. ; (2-5) Using a pre-defined taboo list, analyze the probability distribution of the actual macroscopic topological evolutionary actions obtained in step (2-4). Oscillation correction is performed to generate the final macroscopic topological evolution action. .

5. The trajectory-aware dynamic topology adaptive UAV route planning method according to claim 4, characterized in that, The five-dimensional primitive macroscopic topological evolution actions include two categories: the first category is node-level actions, including moving boundary nodes, exchanging boundary nodes, and maintaining the status quo; the second category is cluster-level actions, including merging two adjacent clusters and splitting overloaded clusters. Step (3) specifically includes the following sub-steps: (3-1) The initial cluster label set obtained in step (1) coordinates of cluster center Used as the current cluster partitioning structure ; (3-2) Based on the macroscopic topological evolution action obtained in step (2) The current cluster partitioning structure obtained in step (3-1) Perform dynamic dimension transformation to obtain the updated cluster count. ; (3-3) Based on the macroscopic topological evolution action obtained in step (2) The boundary nodes obtained in steps (1-5) are then subjected to topology reassignment to obtain the updated cluster label set. ; (3-4) The updated cluster label set obtained in step (3-3) and the initial geographic coordinate set obtained in step (1-1) Perform a centroid update to obtain the updated cluster center coordinates. and the number of nodes in each cluster ; (3-5) The updated cluster count obtained in step (3-2) The updated cluster label set obtained in step (3-3) and the updated cluster center coordinates obtained in steps (3-4) The data is packaged and encapsulated to obtain the cluster partitioning structure.

6. The trajectory-aware dynamic topology adaptive UAV route planning method according to claim 5, characterized in that, Step (4) specifically includes the following sub-steps: (4-1) Extract the geometric centroid coordinates of each cluster from the cluster partitioning structure updated in step (3-5). and the number of nodes included Both are then input into a linear projection layer to obtain the initial feature sequence. ,in This represents the node embedding vector of the k-th cluster; Specifically, the node embedding vector of the k-th cluster The following formula is used for calculation: ; in, Represents a shared linear projection layer. Representing the The cluster center of each cluster, Representing the The number of nodes contained in a cluster These are learnable bias parameters; (4-2) The initial feature sequence obtained in step (4-1) Zero-vector padding is performed to obtain the cluster feature tensor. And based on the cluster feature tensor The padding mask is generated by the index positions of the valid elements and the padding placeholders. ,in This indicates the preset maximum number of clusters; (4-3) The cluster feature tensor obtained in step (4-2) Input a Transformer encoder to compute cluster feature tensors using the multi-head self-attention mechanism MHSA. The spatial dependency weights between each pair of clusters are used to determine the spatial dependency weights of the cluster feature tensor. We perform weighted feature fusion on the effective elements in the dataset to obtain the updated local node embedding vector for each cluster, and then concatenate the updated node embedding vectors for all clusters to obtain the cluster feature tensor. The corresponding encoded hidden state matrix Eh ,in This indicates the preset maximum number of clusters. This represents the dimension of the hidden layer features output by the Transformer encoder; (4-4) The hidden state matrix Eh obtained in step (4-3) is processed by autoregressive sequence decoding using a pointer decoder to generate a non-repeating flight access sequence covering all K clusters. ,in It is the cluster label of the k-th cluster visited by the drone; Specifically, this step involves first initializing a visited state vector with all zeros as a visited mask, used to track and record the UAV's historical visit footprints in flight path planning in real time; then, in the current first decoding step, the pointer decoder is used to encode the hidden state matrix obtained in step (4-3). Attention is calculated to obtain the attention score matrix for each cluster. Then, using the padding mask and visited mask obtained in step (4-2), the attention score matrix is ​​masked, assigning negative infinity to the scores of virtual clusters representing the padding region and visited clusters, to obtain the masked attention score matrix. Subsequently, the masked attention score matrix is ​​normalized to obtain the transition probability corresponding to the first decoding step. ,in The cluster label of the first target cluster to be accessed is The probability of transition is calculated, and based on this transition probability, a maximum probability greedy selection or random sampling strategy is adopted to select the target cluster to be visited by the UAV from the clusters whose corresponding state position is 0 in the already visited mask; then, the state position corresponding to the selected target cluster in the already visited mask is set to 1; then, for the 2nd decoding step, the 3rd decoding step, ... the Kth decoding step, the above operation is repeated until all K decoding steps have been processed, and the sequentially selected clusters are obtained. The target clusters constitute a non-repeating flight access sequence.

7. The trajectory-aware dynamic topology adaptive UAV route planning method according to claim 6, characterized in that, Step (5) specifically includes the following sub-steps: (5-1) Based on the non-repeating flight access sequence generated in step (4-4) The preset drone flight path starting point and initial cluster label set will be used. All nodes within all clusters and the preset drone flight path endpoints are hierarchically expanded and topologically associated along the vertical dimension to construct a directed acyclic graph (DAG). The vertical topology of this DAG is determined by... It consists of several levels, where level 0 is the starting point of the drone flight path, and levels 1 through 2 are... The layers correspond to the non-repeating flight access sequence in sequence. Each cluster in, the first The layer marks the end point of the drone's flight path; (5-2) Initialize the priority queue of the A* algorithm, take the node of layer 0 of the DAG obtained in step (5-1) as the initial node to be evaluated, and set the historical actual cost of the initial node to be evaluated. The sequence of non-repeating flight visits to the initial node to be evaluated is obtained based on Euclidean distance. The shortest theoretical flight distance from the cluster center of all clusters to the end of the UAV's flight path is calculated. This shortest theoretical flight distance is then multiplied by a preset unit flight energy consumption coefficient for the UAV to obtain the ideal flight energy consumption of the UAV. This ideal flight energy consumption is then used as the initial heuristic energy consumption estimate for the initial node to be evaluated. Using the formula Obtain the comprehensive evaluation value of the initial node to be evaluated. And add the initial node to be evaluated to the priority queue; (5-3) Check if the priority queue is empty. If it is, it means that the global optimal path has not been found and the process ends. Otherwise, obtain the comprehensive evaluation value from the priority queue. The smallest node is used as the current traversed node. Then proceed to step (5-4); (5-4) Determine the currently traversed node Does it belong to the first DAG graph? Layer, i.e., determining the currently traversed node Is this the destination of the drone's flight path? If so, it means that the globally optimal path has been found, and then proceed to step (5-15); otherwise, proceed to step (5-5). (5-5) Get the current traversed node The level number in the DAG diagram And extract the first element from the DAG graph. All nodes contained in the layer are used as a candidate cluster head set, and the total number of candidate cluster heads in this candidate cluster head set is obtained. ; (5-6) Set up a candidate cluster head traversal counter ; (5-7) Determine the counter Is it greater than the total number of candidate cluster heads? If yes, return to step (5-3); otherwise, extract the first cluster head from the candidate cluster head set. The node is selected as the current candidate cluster head. Then proceed to steps (5-8); (5-8) Obtain the first digit in the DAG graph. In addition to the candidate cluster head, the layer All other nodes and their total number Set a node traversal counter and initialize the cumulative communication energy consumption within the current layer. ; (5-9) Determine the node traversal counter Is it greater than the total number of other nodes? If yes, proceed to step (5-12); otherwise, obtain the first element in the DAG graph. The first in the layer Each node is selected, and the process proceeds to step (5-10). (5-10) Calculate the first digit of the DAG graph. The first in the layer Each node and candidate cluster head Spatial communication distance between Determine the spatial communication distance Is it less than or equal to the multipath fading threshold? If so, then the first [signal] is calculated based on the free-space signal propagation model. Each node and candidate cluster head Energy consumption for communication between And accumulate the communication energy consumption to the first In the cumulative communication energy consumption of the layer, that is Then proceed to step (5-11); otherwise, calculate the first step based on the multipath fading signal propagation model. Each node and candidate cluster head Energy consumption for communication between And accumulate the communication energy consumption to the first In the cumulative communication energy consumption of the layer, that is Then proceed to step (5-11), where the multipath fading threshold is... , This represents the power amplifier energy consumption coefficient under the free-space signal propagation model. This represents the power amplifier energy consumption coefficient under the multipath fading signal propagation model; (5-11) Set a node traversal counter And return to steps (5-9); (5-12) Based on the drone's current traversal nodes Fly to candidate cluster head Spatial transfer distance acquisition mechanical flight energy consumption To obtain the drone in the candidate cluster head Hovering power consumption during wireless communication above According to the energy consumption of this mechanical flight Hovering energy consumption and the first one obtained in step (5-10) Cumulative communication power consumption of the layer Get the nodes from the current traversal Transfer to candidate cluster head Weighted step cost : ; in, and These are dimensionless weighted coefficients. The value range is from 0.1 to 0.

9. The value range is from 0.1 to 0.9; (5-13) Set the currently traversed node The actual historical cost The weighted step cost obtained in step (5-12) Perform summation to obtain the number of nodes reaching the candidate cluster head. The actual cost ; Obtain the candidate cluster head Sequential non-repeating flight access sequence The shortest theoretical flight distance from the cluster center of all remaining cluster labels to the end of the UAV's flight path is calculated. This shortest theoretical flight distance is then multiplied by a preset UAV unit flight energy consumption coefficient to obtain the UAV's ideal flight energy consumption, which is then used as the candidate cluster head. Heuristic estimation of energy consumption Using the formula Obtain the candidate cluster head Comprehensive assessment value , will candidate cluster heads Add it to the priority queue and proceed to step (5-14). (5-14) Set up candidate cluster head counters And return to steps (5-7); (5-15) Backtrack along the globally optimal path in the priority queue in the DAG graph to obtain the optimal cluster head physical coordinate sequence. and will the The historical actual cost corresponding to the layer node is used as the global energy consumption. Output.

8. The trajectory-aware dynamic topology adaptive UAV route planning method according to claim 7, characterized in that, Step (6) specifically includes the following sub-steps: (6-1) The global energy consumption obtained from step (5-15) Obtain basic reward amount ; (6-2) Calculate the node load variance within the target region based on the updated cluster partitioning structure obtained in step (3). Extract the boundary nodes within each cluster in the updated cluster partitioning structure, and calculate the spatial distance from each boundary node to the cluster center coordinates corresponding to that cluster. Then, select the boundary nodes whose spatial distances are greater than the multipath fading threshold obtained in step (5-10). All distance overflows are summed, and the sum is used as a penalty for exceeding the cluster radius limit within the target area. and node load variance Cluster radius exceeding the limit penalty and the basic reward obtained in step (6-1) Perform weighted synthesis to obtain joint rewards ; (6-3) The combined reward obtained from step (6-2) Constructing generalized advantage estimation And based on this generalized advantage, estimate In step (2), update the network weights of the neural network. To obtain the updated graph neural network; (6-4) Use the joint reward obtained in step (6-2) to update the network weights of the Transformer pointer network in step (4) to obtain the updated Transformer pointer network.

9. The trajectory-aware dynamic topology adaptive UAV route planning method according to claim 8, characterized in that, The base reward amount is calculated using the following formula: ; in, The baseline energy consumption is determined through the following process: First, the initial clustering structure obtained in step (1-2) is used as the current cluster partitioning structure, and all cluster-level sequence features are extracted from this cluster partitioning structure. All extracted cluster-level sequence features are then input into the Transformer pointer network in step (4) for processing to obtain a non-repeating flight access sequence. Subsequently, based on this non-repeating flight access sequence, the preset UAV flight path starting point and the initial cluster label set corresponding to the initial clustering structure are used. All nodes within all clusters and the preset UAV flight path endpoints are hierarchically expanded and topologically associated along the vertical dimension to construct a directed acyclic graph. Finally, the non-repeating flight access sequence is input into the A* algorithm in step (5) to perform multi-layer nested loop optimization iteration processing within the directed acyclic graph to obtain the baseline energy consumption. ; Step (6-3) is as follows: First, the joint feature vector of all nodes obtained in step (2-2) is... Input the value network from the PPO algorithm to obtain the baseline state value. Then, the updated cluster partitioning structure obtained in step (3-5) is processed using the methods described in steps (1) to (2-2) above to obtain the joint feature vector of all nodes after the update. And update the joint feature vector of all nodes Input the Critic network to obtain the updated baseline state value. ; Then, according to the joint reward Baseline state value Updated baseline state value Obtaining timing difference error The timing difference error The calculation formula is: ; in This indicates the preset discount factor; Subsequently, timing difference error was utilized. Discount Factor and the preset generalized advantage estimation decay factor Perform exponentially weighted summation to obtain the generalized dominance estimate. ; Subsequently, this generalized advantage is used to estimate Constructing the policy pruning objective function of the PPO algorithm : ; in, Represents the expectation operator. This indicates that in graph neural networks, the weights... The output is the ratio of the probability of the macroscopic topological evolution action obtained in step (2) to the initial probability corresponding to the actual macroscopic topological evolution action in the probability distribution of the actual macroscopic topological evolution action obtained in step (2-4). This indicates the preset cropping threshold. This means limiting the ratio to The cutoff function within the interval; Finally, the objective function of this strategy is pruned. Regarding network weights in graph neural networks Calculate the partial derivative to obtain the policy gradient. Based on the policy gradient And combine the gradient ascent optimizer to optimize the network weights of the graph neural network. Perform backpropagation to update the network weights of the graph neural network. This is used to obtain the updated graph neural network.

10. A trajectory-aware, dynamic topology-adaptive unmanned aerial vehicle (UAV) route planning system, characterized in that, Includes the following modules: The first module is used to acquire the original environmental data within the target area, extract the initial geographic coordinate set composed of the initial geographic coordinates of all the corresponding nodes from the original environmental data, and perform feature engineering processing on the initial geographic coordinate set to obtain the heterogeneous graph state tensor (HST). The second module is used to input the heterogeneous graph state tensor obtained from the first module into the graph neural network to extract the topology of all nodes. Based on this topology, the PPO algorithm is optimized using a proximal strategy to obtain the macroscopic topological evolution actions. ; The third module is used to process the initial cluster label set obtained in the first module. and cluster center coordinates Used as the current cluster partitioning structure Based on the macroscopic topological evolution actions obtained from the second module and the current cluster partitioning structure Furthermore, the environmental physical actuator is used to redistribute or reorganize all nodes extracted from the original environmental data in the first module to obtain the cluster partitioning structure. The fourth module extracts all cluster-level sequence features from the cluster partitioning structure obtained in the third module, and processes these features using a Transformer pointer network to obtain a non-repeating flight access sequence. ,in The cluster label for the drone to access the k-th cluster; The fifth module is used to determine the non-repeating flight access sequence obtained from the fourth module. A multi-level state search space is constructed, and the A* algorithm is used to perform multi-level nested loop optimization iteration within this multi-level state search space to obtain the optimal cluster head (CH) physical coordinate sequence. and global energy consumption ,in This represents the total number of clusters within the target area; The sixth module is used to calculate the global energy consumption based on the fifth module. A joint reward is generated, and the joint reward is used to synchronously update the network weights of the graph neural network in the second module and the Transformer pointer network in the fourth module, respectively, so as to obtain the updated graph neural network and the Transformer pointer network. The seventh module is used to determine whether the graph neural network updated in the sixth module and the updated Transformer pointer network have reached convergence, or whether the number of iterations has reached the preset maximum number of training rounds. If so, training is stopped, and the preset UAV flight path starting point and the optimal cluster head physical coordinate sequence obtained in the fifth module are used. The system splices together the preset drone flight path endpoints according to the order of time access to obtain the final drone flight path. The process ends when the drone flight path is accessed; otherwise, it returns to the first module.