Unmanned aerial vehicle network hierarchical routing method with enhanced communication

By employing communication-aware clustering and reinforcement learning-based routing methods, cluster head nodes are dynamically selected and inter-cluster routing and forwarding are optimized. This solves the problems of topology changes and uneven energy consumption in UAV networks under highly dynamic environments, and achieves efficient and stable communication services.

CN121126480APending Publication Date: 2025-12-12BEIHANG UNIV

Patent Information

Application Number
CN202511150777.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Unmanned aerial vehicle (UAV) networks face challenges such as frequent changes in network topology, high probability of link interruption, uneven energy consumption, and low learning efficiency in highly dynamic environments. Existing routing protocols are unable to guarantee stability and efficiency.

Method used

The algorithm employs a communication-aware clustering algorithm and a reinforcement learning strategy. It dynamically divides clusters and selects cluster head nodes using the fuzzy C-means algorithm, and optimizes inter-cluster routing and forwarding strategies by combining reinforcement learning methods, updating routing paths in real time to adapt to dynamic changes.

Benefits of technology

It enables efficient, stable, and low-latency communication of UAV networks in complex environments, improving the network's adaptability and resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121126480A_ABST
    Figure CN121126480A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of unmanned aerial vehicle ad hoc networks and intelligent routing, in particular to a communication-enhanced unmanned aerial vehicle network hierarchical routing method, which comprises the following steps of: initializing a network; performing dynamic clustering and cluster head election; performing inter-cluster routing decision; judging whether a clustering updating condition or a routing strategy retraining condition is met or not according to the data transmission effect; and when the clustering updating condition or the routing strategy retraining condition is satisfied, triggering clustering updating or routing strategy retraining. According to the method, firstly, through a communication perception clustering algorithm, all unmanned aerial vehicle nodes are automatically divided into a plurality of clusters according to multi-dimensional information such as positions, signal-to-noise ratios, sight distance conditions, residual energy and the number of neighbors, cluster head nodes are dynamically selected so as to maintain the stability and communication quality of a clustering structure, and each cluster head node is divided into multiple clusters through a reinforcement learning method; and in combination with multi-dimensional state perception, a forwarding decision of a next hop cluster head is generated in real time, and efficient data transmission of inter-cluster multi-hop paths is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of unmanned aerial vehicle ad hoc network and intelligent routing, and particularly relates to a layered routing method for unmanned aerial vehicle network with enhanced sensing. BACKGROUND

[0002] Unmanned aerial vehicle ad hoc network communication has important application prospects in dynamic applications such as emergency communication, environmental monitoring and Internet of Things integration, but the existing technology faces many challenges; traditional unmanned aerial vehicle network routing protocols (such as on-demand distance vector routing, optimized link state routing, etc.) are difficult to ensure stability and efficiency in a highly dynamic environment, and mainly have the following outstanding problems:

[0003] 1. High dynamicity and topology fluctuation: the high-speed movement of the unmanned aerial vehicle group in three-dimensional space leads to frequent changes in network topology and high link interruption probability; existing routing protocols and cluster methods based on static topology cannot adapt to dynamic changes and are difficult to update routes and maintain reliable cluster structures in time.

[0004] 2. Link sensing and lack of multi-dimensional indicators: most existing protocols mainly rely on node distance or ID for cluster head selection and routing decision, ignoring key communication indicators such as channel quality, energy margin, line-of-sight conditions, etc., which can easily cause routing failure and uneven energy consumption.

[0005] 3. Learning efficiency and scalability constraints: although existing multi-agent learning methods have certain adaptability, they face problems such as large state synchronization overhead, slow convergence and high routing load in large-scale networks, limiting their application in high-density unmanned aerial vehicle networks.

[0006] Chinese Patent Publication No. CN114641049A discloses a fuzzy logic-based unmanned aerial vehicle ad hoc network layered routing method, which includes two stages of clustering and routing. First, in the clustering stage, a clustering algorithm is used to divide the nodes in the network into different clusters, and a backbone network is constructed according to the clustering results. Then, in the routing stage, an on-demand routing mechanism is used to start the routing discovery process to find the path to the destination node. As can be seen, the existing on-demand distance vector routing protocol in unmanned aerial vehicle ad hoc network communication technology only considers the shortest path selection of path hop count, ignoring node energy, link quality and dynamic topology characteristics, resulting in low data transmission efficiency in complex environments; although the optimized link state routing maintains global topology information by periodic broadcasting, the frequent update mechanism generates a large control overhead, reducing the available bandwidth of the system; the geographic position stateless routing relies on accurate position information and is prone to failure in occlusion and limited line-of-sight environments. SUMMARY

[0007] To address this issue, the present invention provides a sensory-enhanced hierarchical routing method for unmanned aerial vehicle (UAV) networks, which overcomes the problem of low data transmission efficiency in existing UAV networks under highly dynamic scenarios due to the neglect of node energy, link quality, and dynamic topology characteristics.

[0008] To achieve the above objectives, the present invention provides a sensory-enhanced hierarchical routing method for unmanned aerial vehicle (UAV) networks, comprising:

[0009] Obtain real-time status information of each drone node within the drone network;

[0010] Based on the real-time status information and the fuzzy C-means algorithm, a communication-aware clustering process is performed to select the cluster head node of each cluster;

[0011] The inter-cluster routing and forwarding strategy is executed based on the real-time multi-dimensional state information perceived by the cluster head node;

[0012] The execution results of the inter-cluster routing and forwarding strategy are evaluated according to the phased reward mechanism, and the inter-cluster routing and forwarding strategy is optimized based on the evaluation results.

[0013] The process of executing the inter-cluster routing and forwarding strategy includes selecting the next-hop cluster head node based on the real-time multi-dimensional state information, and transmitting data packets to the next-hop cluster head node.

[0014] Furthermore, the communication-aware clustering process includes:

[0015] Based on the communication sensing distance that integrates spatial location, signal-to-noise ratio, and line-of-sight condition parameters, a clustering iteration process is performed for any clustering period to dynamically divide UAV nodes into multiple clusters;

[0016] After the clustering iteration of any clustering cycle is completed, the cluster head node of each cluster is dynamically selected based on the number of neighboring nodes in the cluster.

[0017] When the new cluster center position during the iteration process meets the convergence condition, the clustering iteration is considered complete.

[0018] Furthermore, the clustering iterative process includes:

[0019] The initial cluster size C is determined based on node density and average velocity;

[0020] The coordinates of C UAV nodes are randomly selected from the node set as the initial cluster center locations;

[0021] Calculate the communication sensing distance from each UAV node to each cluster center;

[0022] The fuzzy membership degree of all UAV nodes to each cluster is calculated based on the communication sensing distance;

[0023] A new cluster center location is generated by weighting the node coordinates based on the fuzzy membership degree.

[0024] Each drone node is assigned to a corresponding cluster based on its maximum membership degree.

[0025] Furthermore, the process of inter-cluster routing and forwarding strategy includes:

[0026] The real-time multidimensional state information perceived by each cluster head node is input into the reinforcement learning policy network to select the next hop cluster head node;

[0027] Along the cluster head path generated by the next hop cluster head node, data packets are transmitted according to the selection order of each cluster head node in the cluster head path, and real-time status information of each UAV node in the cluster head path is periodically broadcast to each node in the cluster.

[0028] The real-time status information includes real-time location coordinates, remaining energy value, and link quality indicators.

[0029] Furthermore, the formula for calculating the communication sensing distance is as follows:

[0030]

[0031] in, Let i be the communication sensing distance between UAV i and UAV j;

[0032] Let j be the spatial location of the cluster head node;

[0033] p i The spatial location of UAV node i;

[0034] Let be the Euclidean distance from the i-th UAV to the cluster head node j;

[0035] Γ ij Let be the signal-to-noise ratio from the i-th UAV to the j-th cluster head node;

[0036] ξ ij For sight distance condition parameters;

[0037] β is the signal-to-noise ratio weighting coefficient;

[0038] γ is the line-of-sight penalty coefficient.

[0039] Furthermore, the formula for calculating the fuzzy membership degree is as follows:

[0040]

[0041] Where, μ ij For fuzzy membership degree;

[0042] m is the ambiguity parameter;

[0043] Let represent the communication sensing distance between UAV i and UAV k.

[0044] Furthermore, the steps for dynamically selecting the cluster head node for each cluster include:

[0045] A new cluster head node is elected based on the number of neighbors of each node within the cluster;

[0046] If multiple nodes have the same number of neighbors, select the node with the shortest average distance to its farthest neighbor as the cluster head node.

[0047] Furthermore, the steps for electing a new cluster head node based on the number of neighbors of nodes within the cluster include:

[0048] The node with the most neighbors in the cluster is elected as the new cluster head node.

[0049] Replace the cluster head node generated in the previous clustering cycle with the new cluster head node.

[0050] Furthermore, the steps for determining the real-time multidimensional state information perceived by the cluster head node include:

[0051] Configure each cluster head node as a reinforcement learning agent;

[0052] Each intelligent agent periodically perceives the local network environment and generates a four-dimensional state vector as real-time multi-dimensional state information.

[0053] The four-dimensional state vector includes:

[0054] Spatial perception metrics used to measure the proximity of cluster head nodes to destination cluster head nodes and neighbor density.

[0055] Link quality metrics used to reflect the communication quality between cluster head nodes and their neighbors;

[0056] A sustainability metric used to comprehensively measure node energy consumption status and link lifetime;

[0057] Routing urgency metrics used to characterize the remaining lifetime of data packets.

[0058] Furthermore, the process of optimizing the inter-cluster routing and forwarding strategy based on the evaluation results includes:

[0059] The weighting coefficients are adjusted based on the changes in the real-time multidimensional state information.

[0060] Compared with existing technologies, the beneficial effects of this invention are as follows: In this embodiment, the network first uses a communication-aware clustering algorithm to automatically divide all UAV nodes into multiple clusters based on multi-dimensional information such as location, signal-to-noise ratio, line-of-sight conditions, remaining energy, and number of neighbors, and dynamically selects cluster head nodes to maintain the stability of the clustering structure and communication quality. Subsequently, each cluster head node uses reinforcement learning methods combined with multi-dimensional state awareness to generate forwarding decisions for the next-hop cluster head in real time, achieving efficient data transmission through multi-hop paths between clusters.

[0061] Furthermore, the hierarchical routing method in this embodiment, which possesses dynamic perception capabilities and reinforcement learning adaptive optimization capabilities, enables efficient, stable, and low-latency communication services for UAV networks in complex environments. Attached Figure Description

[0062] Figure 1 This is a flowchart illustrating the sensory enhancement-based hierarchical routing method for unmanned aerial vehicle networks according to an embodiment of the present invention.

[0063] Figure 2 This is a flowchart illustrating clustering, cluster head election, and updating in an embodiment of the present invention;

[0064] Figure 3 This is a schematic diagram comparing the average latency of different routing methods for unmanned aerial vehicle (UAV) networks under different node scales according to an embodiment of the present invention.

[0065] Figure 4 This is a schematic diagram comparing the data packet success rate of different routing methods in the UAV network at different node scales according to an embodiment of the present invention. Detailed Implementation

[0066] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0067] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0068] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.

[0069] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0070] Please see Figure 1 The diagram shown is a flowchart illustrating a sensory-enhanced hierarchical routing method for unmanned aerial vehicle (UAV) networks according to an embodiment of the present invention. The present invention provides a sensory-enhanced hierarchical routing method for UAV networks, comprising:

[0071] Step S1: Input the basic information of each drone node in the drone network and start the network;

[0072] Each drone node periodically broadcasts its location, energy, and link status information;

[0073] Step S2: Based on real-time status information, the communication-aware clustering process is executed by a communication-aware clustering algorithm that integrates fuzzy C-means algorithm to complete cluster division and cluster head selection;

[0074] Step S3: The cluster head node perceives real-time multi-dimensional state information and uses a reinforcement learning strategy network to output the next-hop cluster head selection.

[0075] The data packets are forwarded hop-by-hop along the selected cluster head path, and necessary status information is broadcast within the cluster;

[0076] Step S4: Evaluate the forwarding effect according to the phased reward mechanism, and determine whether to trigger cluster update and routing strategy retraining based on the forwarding effect and network dynamic changes.

[0077] When the conditions are met, return to the state awareness stage to form a continuous closed-loop optimization process;

[0078] The closed-loop optimization process includes:

[0079] If the clustering update conditions are met, the dynamic clustering and cluster head election steps are re-executed.

[0080] When the retraining conditions for the routing policy are met, the parameters of the reinforcement learning policy network are updated online.

[0081] If the cluster update condition and the routing policy retraining condition are not met, proceed to the next routing decision cycle;

[0082] The process of entering the next routing decision cycle includes: performing a communication-aware clustering process, or inputting real-time multi-dimensional state information of the next cluster head node into the reinforcement learning policy network to output the selection decision of the next hop cluster head node.

[0083] The final output is a hierarchical routing method with enhanced synergy, enabling UAV networks to provide adaptive and low-latency communication services in complex mobile environments.

[0084] This process can be executed during the network initialization phase or iterated and updated repeatedly during the dynamic operation phase to ensure the communication robustness and resource utilization efficiency of the drone swarm during mission execution.

[0085] In this embodiment, the basic information is real-time status information, and the cluster update condition is when the update period T is reached. up The system automatically triggers a reselection process. Real-time state information of UAV nodes is collected by the inner cluster layer and output to the upper layer. The cluster head node acts as a gateway, providing normalized state input to the inter-cluster layer. This enables the inter-cluster layer to optimize inter-cluster paths based on reinforcement learning, achieving global load balancing and low-latency transmission. The cluster head, as a reinforcement learning agent, determines the next hop based on the state vector and optimizes the strategy through a phased reward mechanism. The network first uses a communication-aware clustering algorithm to automatically divide all UAV nodes into multiple clusters based on multi-dimensional information such as location, signal-to-noise ratio, line-of-sight conditions, remaining energy, and number of neighbors, and dynamically selects cluster head nodes to maintain the stability of the cluster structure and communication quality. Subsequently, each cluster head node uses reinforcement learning methods, combined with multi-dimensional state awareness, to generate the next-hop cluster head forwarding decision in real time, achieving efficient data transmission through multi-hop inter-cluster paths.

[0086] See Figure 2 As shown, it is a flowchart of clustering, cluster head election and updating in an embodiment of the present invention;

[0087] This embodiment considers a mobile ad hoc network consisting of multiple drones, assuming the drone set is . Where N is the total number of drone nodes. Let P be the set of all drones. The spatial location set of the drones is P = {p1, p2, ..., p...} N}, where p i Let λ represent the three-dimensional spatial coordinates of the i-th UAV. Each UAV obtains real-time location information through a global navigation and positioning module and periodically broadcasts status information, including key indicators such as coordinates, remaining energy, and link quality, to support cluster construction and neighbor discovery. Since the density and mobility constraints of the UAV cluster directly affect the cluster size, the system determines the cluster size based on the node density λ. N and average speed The number of clusters C is adaptively determined to improve cluster adaptability. Let the UAV's transmit power be P. T The unit is watts; the received noise power is N0, also in watts; the minimum signal-to-noise ratio threshold is Γ. θ The power of the received signal is the ratio of the noise power, representing the minimum reliable demodulation threshold required for the communication link. The minimum safe distance between UAVs is d0, representing the lower limit of the physical distance between nodes. Due to interference constraints, the maximum communication distance d of the UAVs is... T Must meet:

[0088]

[0089] Where α is the path loss exponent, a parameter describing the change in wireless signal propagation attenuation with distance, and λ is the node density. N The larger the value, the stronger the interference, and the smaller the maximum communication distance. Meanwhile, to ensure communication stability, let the minimum dwell time of the UAV within the cluster be T. min The average speed is The minimum cluster radius R derived from the mobility constraints is... mobility for:

[0090]

[0091] Therefore, the maximum permissible radius R of the drone cluster c It is determined by both signal-to-noise ratio and mobility:

[0092]

[0093] in, This is determined by the node density λ. N The interference term is jointly determined by the minimum safe distance d0 and the path loss exponent α. Given the network coverage area S, the number of clusters C corresponding to the node density partition and the velocity partition can be derived from the above cluster radius as follows:

[0094] C = C k When λ N ∈[λ k―1 ,λ k )and

[0095] Let the node density range of the unmanned aerial vehicle be [λ]. k―1 ,λ k ), which represents the upper and lower boundaries of the node density partition interval; the average velocity interval is [v k-1 ,v kThe two ranges are pre-defined based on actual deployment experience and are usually divided into three levels (low, medium, and high). For example, in this embodiment, the node density range is divided according to the number of nodes per square kilometer, such as 0–20, 20–40, and above 40; the speed range is divided according to the average flight speed, such as 0–5 m / s, 5–15 m / s, and 15–25 m / s.

[0096] Each interval corresponds to a cluster size scalar C. k It is used to dynamically match different network sizes and mobility conditions, with the number of clusters as a scalar C. k The standard value for the number of clusters is determined by combining the node density interval and the velocity interval. Specifically, the node density and node velocity are first divided into several intervals (e.g., density intervals: 0–20, 20–40, 40+; velocity intervals: 0–5m / s, 5–15m / s, 15–25m / s). Then, Cartesian combinations are performed on these two dimensions, and the optimal number of clusters is calculated for each combination. An example of the optimal method is shown in Table 1.

[0097] Refer to Table 1, which shows the mapping relationship between the upper and lower limits of the interval and the number of clusters. The number of clusters is scalar C. k The configuration table can be determined through empirical statistics. The system can quickly complete the partition mapping by looking up the table during runtime without real-time calculation. This method improves the flexibility and adaptability of cluster partitioning through the above interval division.

[0098] Table 1. Mapping Relationship Between Interval Upper and Lower Limits and Cluster Quantity

[0099] Node density interval (nodes / km 2 ) Speed interval (m / s) [C k ]]> 0-20 0-5 2 0-20 5-15 3 0-20 15-25 4 20-40 0-5 3 20-40 5-15 4 20-40 15-25 5 ≥40 0-5 5 ≥40 5-15 6 ≥40 15-25 7

[0100] Specifically, the communication-aware clustering process includes:

[0101] Based on the communication sensing distance that integrates spatial location, signal-to-noise ratio, and line-of-sight condition parameters, a clustering iteration process is performed for any clustering period to dynamically divide UAV nodes into multiple clusters;

[0102] After the clustering iteration of any clustering cycle is completed, the cluster head node of each cluster is dynamically selected based on the number of neighboring nodes in the cluster.

[0103] When the new cluster center position during the iteration process meets the convergence condition, the clustering iteration is considered complete.

[0104] Specifically, the process of inter-cluster routing and forwarding strategy includes:

[0105] The real-time multidimensional state information perceived by each cluster head node is input into the reinforcement learning policy network to select the next hop cluster head node;

[0106] Along the cluster head path generated by the next hop cluster head node, data packets are transmitted according to the selection order of each cluster head node in the cluster head path, and real-time status information of each UAV node in the cluster head path is periodically broadcast to each node in the cluster.

[0107] The real-time status information includes real-time location coordinates, remaining energy value, and link quality indicators.

[0108] The clustering process in this embodiment iterates through the following steps:

[0109] Initialize cluster center positions: Select the coordinates of C nodes from the node set using a random selection method as the initial cluster centers;

[0110] Calculate the communication sensing distance: Calculate the communication sensing distance from each node to each cluster center one by one;

[0111] Update the membership matrix: Using communication-aware distance, calculate the fuzzy membership degree μ of all nodes according to the fuzzy C-means formula. ij ;

[0112] Update cluster center coordinates: Calculate the new cluster center coordinates according to the weighted average formula;

[0113]

[0114] Where, p j These are the updated cluster center coordinates;

[0115] Convergence assessment: Convergence here means that all C cluster head nodes have been independently determined, and the positional relationships between the nodes satisfy the constraints of the cluster head node update formula. If convergence is achieved, the termination condition is met, and the algorithm ends; otherwise, it returns to the step of recalculating the communication sensing distance. Convergence is determined by independently selecting cluster head nodes and ensuring that the distances between nodes satisfy the update formula.

[0116] Once the iteration is complete, each drone is assigned to the corresponding cluster based on its maximum membership degree:

[0117]

[0118] Among them, cluster(p i Let be the cluster number to which the i-th drone belongs. After clustering is completed, a new cluster head, headj, is elected based on the number of neighbors of the nodes within the cluster.

[0119]

[0120] If multiple nodes have the same number of neighbors, the node with the shortest average distance to its farthest neighbor is selected as the cluster head. To adapt to dynamic network changes, the system sets a periodic cluster update cycle T. up Upon expiration, an automatic re-election process is triggered, re-executing clustering and cluster head election to ensure the continued rationality of the cluster structure and communication reliability. Through the above mechanism, clustering and cluster head election not only consider node density, mobility, signal-to-noise ratio, and line-of-sight conditions, but also utilize fuzzy clustering and adaptive update mechanisms to improve the stability and self-organization capability of the network topology, significantly enhancing the collaborative communication efficiency of UAV networks in dynamic environments.

[0121] Specifically, the steps for determining the initial cluster size based on node density and average velocity include:

[0122] Based on the node density range and velocity range, a table lookup mapping is performed.

[0123] Density and velocity are each divided into three levels: low, medium, and high, with each level corresponding to a predefined scalar number of clusters.

[0124] Once the number of clusters is determined, the system enters the communication-aware clustering phase. To comprehensively measure the physical proximity and communication quality between UAVs, a communication-aware distance is introduced.

[0125] The formula for calculating the communication sensing distance is:

[0126]

[0127] in, Let i be the communication sensing distance between UAV node i and UAV node j.

[0128] This refers to the location of the cluster head node, specifically the spatial location of the UAV cluster head node j.

[0129] p i The spatial location of UAV node i;

[0130] Let be the Euclidean distance from the i-th UAV to the cluster head j;

[0131] Γ ij Let be the signal-to-noise ratio from the i-th drone to the j-th cluster head;

[0132] ξ ij ξ is a line-of-sight condition parameter; for line-of-sight communication, ξ ij =1; for non-line-of-sight communication, ξ ij =0;

[0133] β is the signal-to-noise ratio weighting coefficient, set to 0.5, used to adjust the degree of influence of signal-to-noise ratio on communication distance;

[0134] γ is the line-of-sight penalty coefficient, set to 0.5, used to increase the distance penalty under non-line-of-sight conditions.

[0135] Specifically, the formula for calculating fuzzy membership degree is:

[0136]

[0137] Where, μ ij For fuzzy membership degree;

[0138] m is the ambiguity parameter, set to 2;

[0139] Let represent the communication sensing distance between UAV i and UAV k.

[0140] Specifically, the steps for dynamically selecting the cluster head node for each cluster include:

[0141] A new cluster head node is elected based on the number of neighbors of each node within the cluster;

[0142] If multiple nodes have the same number of neighbors, select the node with the shortest average distance to its farthest neighbor as the cluster head node.

[0143] Specifically, the steps for electing a new cluster head node based on the number of neighbors of a node within the cluster include:

[0144] The node with the most neighbors in the cluster is elected as the new cluster head node.

[0145] Replace the cluster head node generated in the previous clustering cycle with the new cluster head node.

[0146] In this embodiment, all UAVs are automatically divided into clusters based on node density, movement speed, link signal-to-noise ratio, and line-of-sight conditions using an improved fuzzy C-means algorithm. The cluster head is dynamically selected by comprehensively considering the number of neighboring nodes, signal-to-noise ratio, and line-of-sight conditions, ensuring stable clustering even under high-speed dynamic changes in the network.

[0147] The process of inter-cluster routing and forwarding policy includes:

[0148] The real-time multidimensional state information perceived by each cluster head node is input into the reinforcement learning policy network to select the next hop cluster head node;

[0149] Along the cluster head path generated by the next hop cluster head node, data packets are transmitted according to the selection order of each cluster head node in the cluster head path, and real-time status information of each UAV node in the cluster head path is periodically broadcast to each node in the cluster.

[0150] The real-time status information includes real-time location coordinates, remaining energy value, and link quality indicators.

[0151] Specifically, the steps to determine the real-time multidimensional state information perceived by the cluster head node include:

[0152] Configure each cluster head node as a reinforcement learning agent;

[0153] Each agent periodically perceives the local network environment and generates a four-dimensional state vector as real-time multi-dimensional state information.

[0154] In this embodiment, each cluster head node acts as a reinforcement learning agent, perceiving a multi-dimensional state vector including spatial proximity, link quality, node sustainability, and task urgency. Combined with a phased reward mechanism, it evaluates hop-by-hop actions and transmission results, and continuously trains routing strategies through optimization of the policy network and value network.

[0155] In this embodiment, after completing the communication-aware clustering and cluster head election of the UAV network, the present invention employs a reinforcement learning-based inter-cluster routing and forwarding mechanism to improve the adaptability of the routing strategy and the overall performance of the network. This mechanism achieves intelligent optimization of multi-hop paths through the reinforcement learning agent's perception of the environment, action selection, and feedback learning.

[0156] In this mechanism, each cluster head node is treated as an agent, periodically sensing the local network state and autonomously deciding which cluster head to forward data to. The node's state vector s t This represents the current node's perception state, used as input to the decision network. It is described by four types of indicators, forming a four-dimensional state vector:

[0157] s t = A (c),S Q (c),S S (c),S U (c)>

[0158] Among them, S A (c) represents the spatial awareness index of the current cluster head node c, S Q (c) represents the link quality index of the current cluster head node c, S S (c) represents the sustainability index of the current cluster head node c, S U (c) is the routing urgency indicator for the current cluster head node c.

[0159] Spatial perception index S A (c) Used to measure the proximity between the cluster head and the destination cluster head, as well as neighbor density, defined as:

[0160]

[0161] Among them, D cd ∈[0,1],D cd ​This represents the normalized distance from the current cluster head to the target cluster head, where 0 indicates complete overlap and 1 indicates the farthest distance. N represents the number of neighboring cluster heads of the current cluster head node, i.e., the number of cluster head nodes that can be directly connected within the communication range; N is the total number of cluster head nodes in the network. The percentage of neighbor cluster heads is used to measure the local density of a node; α1 is the distance-neighbor density balance coefficient, α1 = 0.5, used to adjust the influence of the two parts.

[0162] Link quality index S Q (c) Used to reflect the communication quality between the cluster head and its neighbors, defined as:

[0163] S Q (c)=α2·Γ c +(1―α2)·ξ c

[0164] Among them, Γ c ξ represents the average signal-to-noise ratio between the current cluster head and its neighboring cluster heads; a higher value indicates better communication quality. c α is the line-of-sight availability, with a value range of [0,1], where 1 represents complete line-of-sight and 0 represents complete non-line-of-sight; α2 is the balance coefficient between signal-to-noise ratio and line-of-sight availability, which determines the weight of the influence of signal-to-noise ratio and line-of-sight, and α2=0.5.

[0165] Sustainability Indicator S S A comprehensive assessment of node energy consumption and link lifetime:

[0166]

[0167] Among them, E c ∈[0,1], E c τ represents the percentage of energy remaining at a node, where 0 indicates depletion and 1 indicates full energy. c T represents the average remaining communication lifetime between the node and its neighboring cluster heads. max α3 is the maximum reference communication lifetime; α3 is the balance coefficient between remaining energy consumption and average remaining communication lifetime, which determines the weight of the influence of energy consumption and link lifetime, α3=0.5.

[0168] Route urgency index S U (c) Used to characterize the remaining lifetime of the data packet.

[0169] The aforementioned indicators collectively constitute the state input, guiding node decision-making. During action selection, the agent selects the next-hop cluster head from the set of neighboring cluster heads based on the current state, completing the forwarding. Actions are sampled from a probability distribution generated by the policy network to dynamically explore different paths; the specific selection process is achieved by the policy network outputting the probability distribution π = {neighbor1:neighbor2:...:neighbor...} of candidate nodes.n The agent samples according to this probability distribution, thus ensuring forwarding to better nodes while maintaining a certain degree of exploratory nature to adapt to dynamic environmental changes.

[0170] Referring to Table 2, which is an example of the state characteristics of neighboring cluster head nodes, when a cluster head node A senses three neighboring cluster heads B, C, and D, it exhibits different behaviors in terms of spatial proximity, link quality, sustainability, and routing urgency.

[0171] Table 2. Examples of state characteristics representation of neighboring cluster head nodes.

[0172]

[0173]

[0174] After inputting the policy network, the output probability distribution is: The probability of choosing C as the next hop is the highest (0.39), so A forwards the data to C with the highest probability.

[0175] Specifically, optimizing the inter-cluster routing and forwarding strategy based on the evaluation results includes:

[0176] The weighting coefficients are adjusted based on the changes in real-time multidimensional state information.

[0177] To ensure the effectiveness of the learning process, a phased reward mechanism is designed in this embodiment. This phased reward mechanism includes hop-by-hop rewards and termination rewards. Hop-by-hop rewards are used to provide real-time feedback on the immediate improvement of the routing path by the current forwarding decision after each action is executed. They are calculated primarily based on changes in indicators such as spatial proximity, link quality, sustainability, and task urgency, guiding the agent to continuously optimize route selection within a local scope. Termination rewards are used to provide a one-time positive or negative reward based on the overall result when the data transmission task is successfully completed or fails and times out, evaluating the comprehensive effect of the routing strategy from a global perspective. By combining the short-term feedback of hop-by-hop rewards and the long-term guidance of termination rewards, the system can achieve adaptive adjustment and stable convergence of routing strategies in dynamic network environments. Hop-by-hop rewards are used to provide immediate feedback on the value of the current action; at time step t, the obtained hop-by-hop reward r... t Defined as:

[0178] r t =w A ·ΔS A +w Q ·ΔS Q +w S ·ΔS S +w U ·S U

[0179] in: ΔS A This represents the change in spatial proximity; a positive value indicates an increase in proximity. This represents the change in link quality; a positive value indicates improved communication conditions. For sustainability changes, positive values ​​indicate improvements in energy and lifespan; S U It reflects the urgency of the task and directly reflects the time pressure of this step. Weight parameter w A ,w Q ,w S ,w U The set values ​​are related to the reinforcement learning environment and are determined based on the number of drones and the frequency of task distribution. Generally, they are set to 0.25 in the initial simulation, and can be adjusted to 0.2, 0.4, 0.13, and 0.27 later based on the simulation results, i.e., the evaluation results. This is used to control the influence of the four indicators on the reward, while also satisfying the following constraints:

[0180] w A +w Q +w S +w U =1

[0181] The weighting parameter w in this embodiment A ,w Q ,w S ,w U The set value is related to the reinforcement learning environment and is determined based on the number of drones and the frequency of task distribution. Generally, it is set to 0.25 in the initial simulation. During system operation, the weights are adjusted every 50 decision cycles based on the normalized cumulative improvement rate of the four indicators. Let the normalized cumulative improvement rates of the four indicators at the i-th adjustment time be... The updated formula is:

[0182]

[0183] And after the update, w is satisfied. A +w Q +w S +w U =1. Example: After running for 50 cycles under initial conditions, the average improvement rate of spatial proximity was statistically obtained. Average link quality improvement rate Average rate of improvement in sustainability Average task urgency Substituting into the formula, we get:

[0184]

[0185] When a data task transmission succeeds or fails, the system will provide a one-time termination reward. T :

[0186]

[0187] Among them, R max A positive number >0 indicates the absolute value of the reward for success or failure. The total cumulative reward consists of the hop-by-hop reward and the termination reward mentioned above: one part is the immediate reward obtained by the agent at each step during data forwarding, and the other part is the termination reward generated when the task is finally completed or failed. All rewards are weighted and accumulated over time according to a certain discount factor to balance the short-term and long-term routing effects. Typically, R is designed... max =20.

[0188] During the training phase, the agent continuously learns through joint optimization by the policy network and the value network, using proximal policy optimization to update the policy. The policy network outputs the next-hop selection probability, while the value network estimates the long-term reward of the current state. Each interaction generates a quadruple of state, action, reward, and next state, used to update network parameters. After training, the cluster head node directly utilizes the learned policy in actual operation to generate next-hop forwarding decisions based on real-time states, achieving continuous adaptation to dynamic topologies and efficient transmission. Through this mechanism, supported by multi-dimensional perception, dynamic decision-making, and phased feedback, the system effectively improves the routing robustness, timeliness, and resource utilization efficiency of the UAV network.

[0189] In this embodiment, a hierarchical routing method with dynamic perception and reinforcement learning adaptive optimization capabilities is used to achieve efficient, stable, and low-latency communication services for UAV networks in complex environments.

[0190] See Figure 3 and Figure 4 As shown, Figure 3 This is a schematic diagram comparing the average latency of different routing methods for unmanned aerial vehicle (UAV) networks under different node scales according to an embodiment of the present invention. Figure 4This diagram illustrates the comparison of packet success rates for different routing methods in a UAV network under varying node sizes, as described in this embodiment of the invention. The performance advantages of the hierarchical routing method in a UAV network environment are verified through simulation experiments. The experimental conditions are as follows: average flight speed of UAV nodes is 15 m / s, simulation area size is 1500 m × 1500 m, data generation rate is 15 KB / s, simulation duration is 60 seconds, and the number of nodes increases incrementally from 20 to 200. To verify the effectiveness, three existing typical routing methods are selected for comparison: on-demand distance vector routing, a classic on-demand path discovery protocol, widely studied in UAV mobile networks; on-demand distance vector routing based on K-means clustering, combining clustering and traditional on-demand mechanisms, representative in clustered ad hoc networks; and a UAV network routing method based on reinforcement learning, employing tabular Q-learning to update the neighbor selection strategy, a commonly used intelligent routing baseline in recent years.

[0191] Related simulation results show that the communication-aware fuzzy clustering and reinforcement learning-based hierarchical routing strategy described in this invention exhibits higher data reception rate and lower end-to-end latency under different node densities and mobility conditions. Detailed data can be found in the embodiments section and accompanying drawings. Figure 3 , Figure 4 Under the conditions of 160 nodes and an average speed of 15 m / s, the average latency of this method is reduced by about 44% and the average data reception rate is increased by about 36% compared with on-demand distance vector routing. In addition, the simulation framework used is based on the network simulation tool NS-3 commonly used in the field. The parameter settings and comparison baselines are in line with industry standards, which can effectively support the authenticity and repeatability of the technical effects of this invention.

[0192] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

[0193] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A layered routing method for unmanned aerial vehicle (UAV) networks with enhanced sensory perception, characterized in that, include: Obtain real-time status information of each drone node within the drone network; Based on the real-time status information and the fuzzy C-means algorithm, a communication-aware clustering process is performed to select the cluster head node of each cluster; The inter-cluster routing and forwarding strategy is executed based on the real-time multi-dimensional state information perceived by the cluster head node; The execution results of the inter-cluster routing and forwarding strategy are evaluated according to the phased reward mechanism, and the inter-cluster routing and forwarding strategy is optimized based on the evaluation results. The process of executing the inter-cluster routing and forwarding strategy includes selecting the next-hop cluster head node based on the real-time multi-dimensional state information, and transmitting data packets to the next-hop cluster head node.

2. The sensory-enhanced hierarchical routing method for unmanned aerial vehicle networks according to claim 1, characterized in that, The communication-aware clustering process includes: Based on the communication sensing distance that integrates spatial location, signal-to-noise ratio, and line-of-sight condition parameters, a clustering iteration process is performed for any clustering period to dynamically divide UAV nodes into multiple clusters; After the clustering iteration of any clustering cycle is completed, the cluster head node of each cluster is dynamically selected based on the number of neighboring nodes in the cluster. When the new cluster center position during the iteration process meets the convergence condition, the clustering iteration is considered complete.

3. The sensory-enhanced hierarchical routing method for unmanned aerial vehicle networks according to claim 2, characterized in that, The clustering iterative process includes: The initial cluster size C is determined based on node density and average velocity; The coordinates of C UAV nodes are randomly selected from the node set as the initial cluster center locations; Calculate the communication sensing distance from each UAV node to each cluster center; The fuzzy membership degree of all UAV nodes to each cluster is calculated based on the communication sensing distance; A new cluster center location is generated by weighting the node coordinates based on the fuzzy membership degree. Each drone node is assigned to a corresponding cluster based on its maximum membership degree.

4. The sensory-enhanced hierarchical routing method for unmanned aerial vehicle networks according to claim 1, characterized in that, The process of inter-cluster routing and forwarding policy includes: The real-time multidimensional state information perceived by each cluster head node is input into the reinforcement learning policy network to select the next hop cluster head node; Along the cluster head path generated by the next hop cluster head node, data packets are transmitted according to the selection order of each cluster head node in the cluster head path, and real-time status information of each UAV node in the cluster head path is periodically broadcast to each node in the cluster. The real-time status information includes real-time location coordinates, remaining energy value, and link quality indicators.

5. The sensory-enhanced hierarchical routing method for unmanned aerial vehicle networks according to claim 3, characterized in that, The formula for calculating the communication sensing distance is: in, Let i be the communication sensing distance between UAV i and UAV j; Let j be the spatial location of the cluster head node; p i The spatial location of UAV node i; Let be the Euclidean distance from the i-th UAV to the cluster head node j; Γ ij Let be the signal-to-noise ratio from the i-th UAV to the j-th cluster head node; ξ ij For sight distance condition parameters; β is the signal-to-noise ratio weighting coefficient; γ is the line-of-sight penalty coefficient.

6. The sensory-enhanced hierarchical routing method for unmanned aerial vehicle networks according to claim 5, characterized in that, The formula for calculating the fuzzy membership degree is: Where, μ ij For fuzzy membership degree; m is the ambiguity parameter; Let represent the communication sensing distance between UAV i and UAV k.

7. The sensory-enhanced hierarchical routing method for unmanned aerial vehicle networks according to claim 2, characterized in that, The steps for dynamically selecting the cluster head node for each cluster include: A new cluster head node is elected based on the number of neighbors of each node within the cluster; If multiple nodes have the same number of neighbors, select the node with the shortest average distance to its farthest neighbor as the cluster head node.

8. The sensory-enhanced hierarchical routing method for unmanned aerial vehicle networks according to claim 7, characterized in that, The steps for electing a new cluster head node based on the number of neighbors of a node within the cluster include: The node with the most neighbors in the cluster is elected as the new cluster head node. Replace the cluster head node generated in the previous clustering cycle with the new cluster head node.

9. The sensory-enhanced hierarchical routing method for unmanned aerial vehicle networks according to claim 4, characterized in that, The steps to determine the real-time multidimensional state information perceived by the cluster head node include: Configure each cluster head node as a reinforcement learning agent; Each intelligent agent periodically perceives the local network environment and generates a four-dimensional state vector as real-time multi-dimensional state information. The four-dimensional state vector includes: Spatial perception metrics used to measure the proximity of cluster head nodes to destination cluster head nodes and neighbor density. Link quality metrics used to reflect the communication quality between cluster head nodes and their neighbors; A sustainability metric used to comprehensively measure node energy consumption status and link lifetime; Routing urgency metrics used to characterize the remaining lifetime of data packets.

10. The sensory-enhanced hierarchical routing method for unmanned aerial vehicle networks according to claim 1, characterized in that, The process of optimizing the inter-cluster routing and forwarding strategy based on the evaluation results includes: The weighting coefficients are adjusted based on the changes in the real-time multidimensional state information.

Citation Information

Patent Citations

  • Unmanned aerial vehicle ad hoc network hierarchical routing method based on fuzzy logic

    CN114641049A

Cited By

  • Low-altitude intelligent networking self-adaptive clustering networking system oriented to multi-dimensional resource collaboration

    CN121815368A