Self-adaptive interference sensing routing method and system for dynamic unmanned aerial vehicle network
Optimizing the routing protocol of the drone network through the Q-learning algorithm and multi-index fusion reward function, solving the problems of susceptibility to interference and link dynamics of the drone network, and achieving the best routing decisions with high efficiency, low latency and low overhead.
Patent Information
- Application Number
- CN202510708958.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-29
AI Technical Summary
When facing the open characteristics and high dynamics of drone communication, the routing protocols of existing drone networks are susceptible to interference and frequently change in communication links, affecting network performance.
Using the routing protocol framework based on the Q-learning algorithm, the adaptive interference-aware routing of dynamic drone networks is achieved by building a multi-index fusion reward function and adaptive learning mechanism, combining link interference, stability and distance factors, and optimizing learning rate and discount factors.
It improves the adaptability of the drone network in the interfering environment, improves the packet delivery rate, reduces end-to-end delay and energy consumption, and enhances the adaptability of the routing protocol in high-speed environments.
Smart Images

Figure CN120499776A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of unmanned aerial vehicle network communication technology, and in particular relates to an adaptive interference-aware routing method and system for dynamic unmanned aerial vehicle networks. Background Art
[0002] At present, drone swarm technology has become an important development trend in the drone field. That is, the use of multiple drones carrying various equipment to form a network can greatly improve the capabilities of payload, perception, and information processing, thereby achieving large-scale, high-efficiency, and multi-task operation goals.
[0003] When a single drone is expanded into a cluster, efficient and reliable end-to-end communication is a key condition for achieving collaborative operation of drone clusters. A routing protocol is a network protocol that ensures that data packets are transmitted from the source address to the destination address. At present, routing protocols suitable for drone networks are mainly divided into two categories: topology-based routing and location-based routing, such as on-demand distance vector routing (AODV), optimized link state routing (OLSR), greedy perimeter stateless routing (GPSR) and other traditional routing protocols. Based on traditional protocols, some papers have proposed improved methods, such as location-based optimized link state routing (P-OLSR), mobility prediction-based virtual routing (MPVR) algorithm, etc. However, the existing routing protocols for drone networks still have the following problems:
[0004] (1) The open nature of drone communications makes them more susceptible to interference than traditional communication systems. Once a drone is attacked by jamming signals, its communication quality may be significantly degraded, thus affecting the normal execution of the mission.
[0005] (2) The high dynamics and spatial freedom of UAVs will lead to frequent changes in communication links, thus affecting the overall network performance. Summary of the Invention
[0006] The purpose of the present invention is to provide an adaptive interference-aware routing method and system for dynamic UAV networks to solve the above problems.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] In a first aspect, the present invention provides an adaptive interference-aware routing method for dynamic UAV networks, comprising:
[0009] A routing protocol framework based on the Q-learning algorithm is established, with data packets as the agent and the drone network as the environment. The packet forwarding decision is mapped into the state-action pair of Q-learning through the state transition mechanism, and a Q-value iterative model including the learning rate α and the discount factor γ is constructed.
[0010] A reward function that integrates multiple indicators is constructed. By using three indicators, namely the expected number of transmissions (ETX) that calculates link quality based on the signal-to-interference-noise ratio (SINR), the link duration predicted by node motion trajectories, and the distance difference that reflects the proximity to the destination node, the dynamic characteristics of the UAV network are converted into quantifiable objectives for reinforcement learning.
[0011] Based on the iterative model and the quantifiable goals reflected by the three indicators, the learning rate α is updated with the negative exponential function of the link duration, the discount factor γ is adjusted by the complementary function of the link stability, and the Softmax strategy with temperature coefficient τ decay is used for action selection to form a closed-loop parameter optimization.
[0012] Furthermore, the construction of a Q-value iterative model including a learning rate α and a discount factor γ includes:
[0013] By representing the data packet as an agent, the entire network is regarded as the environment, and the state space State consists of all drones used to forward the data packet; the selection of the next hop is regarded as an action Action, and the action space includes all available neighbor nodes; the state transition is equivalent to the packet forwarding decision between drones; that is, any data packet in the network environment has a state S t , represents the current drone node; the set of all actions that the data packet can take is the action space A, which consists of all available neighbor nodes; if the data packet is in S t An action a in A taken in state t , that is, a neighbor node is selected for forwarding, and the current state will be transferred to the next state S t+1 ; At the same time as the state transfer, the environment will feedback a reward R(S t ,a t ); Ultimately, the goal of the algorithm is to continuously try to learn the optimal strategy π in the environment, maximize the cumulative reward obtained by long-term execution of the strategy, and obtain the state S according to the strategy t The optimal action a to be performed next t =π(S t );
[0014] The object estimated by evaluating the policy π is the value function Q, which is calculated by the following formula:
[0015]
[0016] where Q t (s t ,a t ) is in state s t Take action a t The Q value when t+1is the Q value at time t+1; α∈(0,1) represents the learning rate, and γ∈(0,1) represents the discount factor; Represents all possible actions a in the next state ′ ∈A t+1 The maximum Q value, where A t+1 is state s t+1 The action space in the Hello message is used to define and calculate the reward R(s) for different QoS indicators. t ,a t ); All drone nodes will maintain a neighbor table and a Q-value table by regularly exchanging Hello messages with their neighbors; the former is used to update the action space and calculate rewards, and the latter is used to determine routing decisions.
[0017] Furthermore, the expected transmission time ETX of the link quality calculated based on the signal-to-interference-and-noise ratio, the link duration predicted by the node motion trajectory, and the distance difference reflecting the proximity to the destination node include three indicators:
[0018] First, the BER is calculated using the mapping relationship between bit error rate and signal to interference and noise ratio;
[0019]
[0020] Among them, SINR represents the signal-to-interference-and-noise ratio at the receiving node, γ is a constant related to the bandwidth; Q(z) represents the probability that the random variable X in the Gaussian distribution X~N(0,1) is greater than the threshold z, that is,
[0021]
[0022] Therefore, the transmission success rate PDR of a data packet of size lbits can be calculated as:
[0023]
[0024] Among them, BER i represents the bit error rate of the i-th bit at the receiving node;
[0025] Finally, the expected number of transmission times ETX of the data packet is obtained
[0026]
[0027] PDR f 、PDR r Respectively represent the forward and backward delivery success rates of data packets;
[0028] After obtaining its own and its neighbors’ location information, the node estimates the link duration; d ij(t) represents the relative distance between nodes i and j at time t, and R is the communication coverage radius of the node. When nodes i and j receive each other's Hello message at time t2, they extract the corresponding node location information from the message. Then, nodes i and j estimate the link duration between themselves and each other based on the information.
[0029] First, the distance between nodes i and j at time t is calculated according to the Euclidean distance formula
[0030]
[0031] Secondly, the relative distance change between nodes i and j is calculated by the following formula,
[0032] Δd=d ij (t2)-d ij (t1)(7)
[0033] Finally, the node estimates the link duration as
[0034]
[0035] Among them, T max =R / V min , V min is the minimum speed of the UAV node;
[0036] The sign of Δd reflects the trend of the relative distance between nodes. When Δd is positive, the relative distance between two nodes gradually increases. From time t1 to t2, the relative speed between nodes is Δd / (t2-t1). Therefore, the expected time when the relative distance between two nodes exceeds the communication coverage is That is, the link duration; when Δd is negative, the two nodes are moving towards each other; the two nodes are moving at V min The time required to move the relative distance R at a relative speed is recorded as T max ; At time t2, d ij (t2) The ratio of R multiplied by T max Used as a predictor of link duration;
[0037] The distance metric in the reward function is,
[0038] Δd ij =d iD -d jD (9)
[0039] where d iD d jD Represent the Euclidean distance between nodes i and j and the destination node respectively.
[0040] Furthermore, the reward function includes:
[0041]
[0042] Among them, ETX is the estimated expected number of transmissions, which is used to measure the degree of interference between nodes; t LD (i,j) represents the duration of the link between nodes, Δd ij It is the difference between the distances of the node and the destination node, indicating that selecting a node closer to the destination node as a relay will obtain a greater reward; A, B, and C are weight parameters of various routing metrics, respectively, which satisfy A+B+C=1.
[0043] Furthermore, there are two special cases of reward function design, as shown in formula (11),
[0044]
[0045] When the next hop node is the destination, the current node will receive the maximum reward R max To ensure that the route can eventually lead to the destination node; when all the adjacent nodes of the node are farther away from the destination node, the routing process falls into the local optimum; at this time, the current node should obtain the minimum reward R min .
[0046] Furthermore, based on the iterative model and the quantifiable goals reflected by the three indicators, the learning rate α is updated as a negative exponential function of the link duration, the discount factor γ is adjusted by a complementary function of the link stability, and a softmax strategy with a temperature coefficient τ decay is used for action selection, forming a closed-loop parameter optimization, including:
[0047] An adaptive learning mechanism is introduced to dynamically adjust the values of α and γ using link stability;
[0048] When the link between nodes is more unstable, that is, the link duration is shorter, the Q value needs to be updated faster, that is, a larger learning rate needs to be applied; the update formula for defining the learning rate α is
[0049]
[0050] Among them, t LD (i, j) is calculated by formula (8), i, j represent the current state s t and select action a t The corresponding pair of adjacent nodes;
[0051] When the link between nodes is more unstable, a smaller γ needs to be applied; the update formula of the discount factor γ is defined as follows
[0052]
[0053] Among them, t LD (j,k) is calculated by formula (8), j and k represent the next state s t+1 And the action with the maximum Q value in the next state, that is The corresponding pair of adjacent nodes;
[0054] Softmax method is used
[0055]
[0056] Among them, τ is a positive constant; in the iterative process, τ is updated by τ(t) = τ(0) / log2(1+t); initially, due to the large size of τ, the strategy probability π(s t ,a t ) are equal, which means that the strategy is more inclined to randomly select actions; as the time step increases, τ decreases, and the strategy probability π(s t ,a t ) gradually increases, which means that a more stable link is preferred.
[0057] Furthermore, the protocol flow executed by each drone is as follows:
[0058] First, drones maintain neighbor tables and interaction status information through Hello messages. When receiving a Hello message from another node, if the sender is new, it is added to the neighbor list and the corresponding status information is updated. Routing metrics such as the expected number of transmissions and link duration are then calculated. If no Hello messages are received from any other drones, neighbor nodes with failed links are removed based on the estimated remaining link duration.
[0059] Secondly, each drone node will execute a distributed Q-learning process after updating its neighbor information through a Hello message. When the destination node is a neighbor of the current node, the current node will receive the maximum reward and update the Q value. When the node's neighbor table does not contain the destination node, the node will gain experience based on the behavior strategy and calculate the reward value. The node will then adaptively adjust the learning rate and discount factor and iteratively update the Q-table.
[0060] Finally, when the drone generates or receives a data packet, it queries the Q-table at the current moment and selects the neighbor node with the largest Q value as the next-hop relay node.
[0061] In a second aspect, the present invention provides an adaptive interference-aware routing system for dynamic UAV networks, comprising:
[0062] The framework construction module is used to establish a routing protocol framework based on the Q-learning algorithm. It uses data packets as intelligent agents and the drone network as the environment. It maps data packet forwarding decisions into Q-learning state-action pairs through a state transition mechanism and constructs a Q-value iterative model that includes a learning rate α and a discount factor γ.
[0063] The reward function construction module is used to construct a reward function that integrates multiple indicators. It uses three indicators: the expected transmission time (ETX) that calculates link quality based on the signal-to-interference-noise ratio, the link duration predicted by the node's motion trajectory, and the distance difference that reflects the proximity to the destination node. This module transforms the dynamic characteristics of the UAV network into a quantifiable goal for reinforcement learning.
[0064] The optimization output module is used to update the learning rate α as a negative exponential function of the link duration based on the iterative model and the quantifiable goals reflected by the three indicators. The discount factor γ is adjusted by the complementary function of the link stability. The softmax strategy with temperature coefficient τ decay is used for action selection to form a closed-loop parameter optimization.
[0065] In a third aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the adaptive interference-aware routing method for dynamic drone networks are implemented.
[0066] In a fourth aspect, the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the adaptive interference-aware routing method for dynamic drone networks.
[0067] Compared with the prior art, the present invention has the following technical effects:
[0068] This paper models the routing process of drone networks using a Markov decision process, analyzes a routing protocol framework based on the Q-learning algorithm, and optimizes the algorithm flow to suit the characteristics of swarm drone networks. This enables drones to execute routing decisions in a distributed and autonomous manner, improving the algorithm's convergence speed and stability. This allows drones to more quickly and accurately select the optimal relay node, effectively improving packet delivery rates and reducing end-to-end latency.
[0069] This method quantitatively analyzes and evaluates factors such as link interference and link stability based on drone perception information, and designs a reward function based on these factors. This improves the adaptability of drone networks in interference environments, enabling routing paths to effectively avoid areas of strong interference. Furthermore, a comparison reveals that the paths established by this method are more efficient, resulting in better overall performance, such as latency and energy consumption.
[0070] The present invention uses the link stability parameters obtained by analysis to design an adaptive learning mechanism suitable for dynamic topology, which can adaptively adjust behavior strategies and learning parameters, thereby effectively improving the success rate of data packet delivery and enhancing the adaptability of routing protocols in high-speed environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 This is a routing protocol framework diagram of the present invention.
[0072] Figure 2 Flowchart of the present invention.
[0073] Figure 3 A diagram of the routing case established for the present invention.
[0074] Figure 4 A comparison chart of the data packet delivery rates of the present invention and the comparison routing protocol.
[0075] Figure 5 This is a comparison chart of the end-to-end delay between the present invention and the comparison routing protocol.
[0076] Figure 6 The figure is a comparison of energy consumption between the present invention and the comparison routing protocol. DETAILED DESCRIPTION
[0077] The present invention is further described below with reference to the accompanying drawings:
[0078] Example 1, please refer to Figure 1 The present invention provides an adaptive interference-aware routing method for dynamic UAV networks, comprising:
[0079] A routing protocol framework based on the Q-learning algorithm is established, with data packets as the agent and the drone network as the environment. The packet forwarding decision is mapped into the state-action pair of Q-learning through the state transition mechanism, and a Q-value iterative model including the learning rate α and the discount factor γ is constructed.
[0080] A reward function that integrates multiple indicators is constructed. By using three indicators, namely the expected number of transmissions (ETX) that calculates link quality based on the signal-to-interference-noise ratio (SINR), the link duration predicted by node motion trajectories, and the distance difference that reflects the proximity to the destination node, the dynamic characteristics of the UAV network are converted into quantifiable objectives for reinforcement learning.
[0081] Based on the iterative model and the quantifiable goals reflected by the three indicators, the learning rate α is updated with the negative exponential function of the link duration, the discount factor γ is adjusted by the complementary function of the link stability, and the Softmax strategy with temperature coefficient τ decay is used for action selection to form a closed-loop parameter optimization.
[0082] The present invention addresses the shortcomings of the existing methods and proposes a fast, adaptive, interference-aware routing protocol for dynamic UAV networks. This protocol enables UAVs to make routing decisions in a distributed and autonomous manner, and comprehensively considers the impact of factors such as link interference and link stability on the UAV network, thereby achieving optimal routing with high efficiency, low latency, and low overhead.
[0083] In embodiment 2, the present invention provides an adaptive interference-aware routing method for a dynamic UAV network, comprising:
[0084] S1: Establish a routing protocol framework based on the Q-learning algorithm
[0085] S2: Design reward function based on actual needs
[0086] S3: Designing an adaptive learning mechanism based on drone network characteristics
[0087] Specifically, in step S1, by representing the data packet as an agent, the entire network can be regarded as the environment, and the state space consists of all drones used to forward the data packet. The selection of the next hop is regarded as an action, and the action space includes all available neighbor nodes. Therefore, the state transition is equivalent to the packet forwarding decision between drones. That is, any data packet in the network environment has a state of S. t , represents the current drone node. The set of all actions that the data packet can take is the action space A, which consists of all available neighbor nodes. If the data packet takes an action a in A in the state St t , that is, a neighbor node is selected for forwarding, and the current state will be transferred to the next state S t+1 At the same time as the state transfer, the environment will feedback a reward R(S t ,a t ). Ultimately, the goal of the algorithm is to continuously try to learn an optimal policy (policy) π in the environment to maximize the cumulative reward obtained by long-term execution of the policy. According to the policy, the state S can be obtained. t The optimal action a to be performed next t =π(S t ).
[0088] The object estimated by evaluating the policy π is the value function Q, which is calculated by the following formula:
[0089]
[0090] where Q t (s t ,a t) is in state s t Take action a t The Q value when t+1 is the Q value at time t+1. α∈(0,1) represents the learning rate, and γ∈(0,1) represents the discount factor. Represents all possible actions a in the next state ′ ∈A t+1 The maximum Q value, where A t+1 is state s t+1 For different QoS indicators, the reward R(s) can be defined and calculated by the information embedded in the Hello message. t ,a t All drone nodes will maintain a neighbor table and a Q-table by regularly exchanging Hello messages with their neighbors. The former is used to update the action space and calculate rewards, while the latter is used to determine routing decisions.
[0091] Specifically, in step S2, the design of the reward function takes into account the influence of factors such as link quality, stability and distance.
[0092] In order to better characterize the link quality, the present invention uses the Signal-to-Interference-plus-Noise Rate (SINR) to calculate the expected transmission time (ETX) of a data packet.
[0093] First, the BER is calculated using the mapping relationship between the bit error rate (BER) and the signal to interference and noise ratio.
[0094]
[0095] Where SINR represents the signal-to-interference-noise ratio at the receiving node, and γ is a constant related to bandwidth. Q(z) represents the probability that the random variable X in the Gaussian distribution X~N(0,1) is greater than the threshold z, that is,
[0096]
[0097] Therefore, the transmission success rate (Packet Delivery Ratio, PDR) of a data packet of size 1bits can be calculated as follows:
[0098]
[0099] Among them, BER i represents the bit error rate of the i-th bit at the receiving node.
[0100] Finally, the expected transmission times (ETX) of the data packet can be obtained.
[0101]
[0102] PDR f 、PDR r Respectively represent the forward and backward delivery success rates of data packets.
[0103] Since drones are highly maneuverable aircraft, drone networks face the challenge of frequent link changes. This paper adopts a link duration estimation method to measure the stability of links between nodes.
[0104] After obtaining the location information of itself and its neighbors, a node can estimate the link duration. For example, nodes i and j are a pair of neighboring nodes. ij (t) represents the relative distance between nodes i and j at time t, and R is the communication coverage radius of the nodes. When nodes i and j receive each other's Hello message at time t2, they extract the corresponding node location information from the message. Nodes i and j can then estimate the link duration between themselves and the other node based on this information.
[0105] First, the distance between nodes i and j at time t can be calculated according to the Euclidean distance formula:
[0106]
[0107] Secondly, the relative distance change between nodes i and j can be calculated by the following formula,
[0108] Δd=d ij (t2)-d ij (t1).#(7)
[0109] Finally, the node can estimate the link duration as,
[0110]
[0111] Among them, T max =R / V min , V min is the minimum speed of the UAV node.
[0112] The sign of Δd reflects the trend of the relative distance between nodes. When Δd is positive, the relative distance between two nodes gradually increases. From time t1 to t2, the relative speed between the nodes is Δd / (t2-t1). Therefore, the expected time for the relative distance between two nodes to exceed the communication coverage range is That is, the link duration. When Δd is negative, the two nodes are moving towards each other. min The time required to move the relative distance R at a relative speed is recorded as T maxAt time t2, d ij (t2) The ratio of R multiplied by T max Used as a prediction of link duration.
[0113] The distance metric in the reward function is,
[0114] Δd ij =d iD -d jD .#(9)
[0115] where d iD d jD Represents the Euclidean distance between nodes i and j and the destination node. If node j has a larger Δd ij , which means that node j is closer to the destination node, and will generate a greater reward accordingly.
[0116] Taking into account the impact of the above factors on routing decisions, the reward function designed by the present invention is as follows:
[0117]
[0118] Among them, ETX is the estimated expected number of transmissions, which is used to measure the degree of interference between nodes; t LD (i,j) represents the duration of the link between nodes, Δd ij It is the difference between the distance between the node and the destination node, indicating that selecting a node closer to the destination node as a relay will receive a greater reward. A, B, and C are the weight parameters of various routing metrics, which satisfy A+B+C=1.
[0119] In addition, in order to make the routing strategy goal-oriented and effectively avoid falling into the local optimum, the design of the reward function needs to consider two special cases, as shown in formula (11):
[0120]
[0121] When the next hop node is the destination, the current node will receive the maximum reward R max To ensure that the route can eventually lead to the destination node. When all the neighboring nodes of a node are farther away from the destination node, the routing process falls into a local optimum. At this time, the current node should obtain the minimum reward R min , in order to avoid the occurrence of routing holes. Normally, the reward will be calculated by formula (10).
[0122] Specifically, in step S3, due to the high dynamics of the swarm UAV network, the location, status, neighbors, and other information of the UAVs will change frequently. In order to better adapt to the dynamics of the swarm UAV network and improve network performance, the present invention introduces an adaptive learning mechanism based on the traditional Q-learning algorithm, and uses link stability to dynamically adjust the values of α and γ in formula (1).
[0123] When the link between nodes is more unstable, that is, the link duration is shorter, the Q value needs to be updated faster, that is, a larger learning rate needs to be applied. Therefore, the update formula for defining the learning rate α is,
[0124]
[0125] Among them, t LD (i, j) is calculated by formula (8), i, j represent the current state s t and select action a t A pair of adjacent nodes.
[0126] Since in formula (1) Represents all possible actions a in the next state ′ ∈A t+1 The maximum Q value of , so the larger the discount factor γ, the greater the weight of future rewards in the iteration. Therefore, when the link between nodes is more unstable, a smaller γ needs to be applied to weaken the impact of future rewards. The update formula of the discount factor γ is defined as follows,
[0127]
[0128] Among them, t LD (j,k) is calculated by formula (8), j and k represent the next state s t+1 And the action with the maximum Q value in the next state, that is A pair of adjacent nodes.
[0129] In addition, the choice of behavioral strategy also needs to consider the impact of node dynamics. When generating experience samples, the behavioral strategy can choose the next action by randomly selecting a new path or greedily selecting the action with the maximum Q value. If only greedy strategies are used, it may fall into the local optimal solution and fail to find a better strategy. On the other hand, if only random strategies are used, it may take a lot of time to explore useless actions and fail to obtain sufficient cumulative rewards. In cluster drone networks, due to the frequent updates of the action space, conventional behavioral strategies cannot fully match the dynamic characteristics of the network. Therefore, the present invention adopts the Softmax method,
[0130]
[0131] Where τ is a positive constant. During the iteration process, τ is updated by τ(t) = τ(0) / log2(1+t). Initially, since τ is large, the strategy probability π(s t ,a t ) are almost equal, which means that the strategy is more inclined to randomly select actions; as the time step increases, τ decreases, and the strategy probability π(s t ,a t ), which means that the strategy will reduce exploration behavior and tend to choose more stable links.
[0132] Specifically, the protocol process executed by each drone is as follows:
[0133] First, drones maintain a neighbor table and exchange state information through Hello messages. When a Hello message is received from another node, if the sender is new, it is added to the neighbor list and the corresponding state information is updated. Routing metrics such as the expected number of transmissions and link duration are then calculated. If no Hello messages are received from any other drones, neighbor nodes with failed links are removed based on the estimated remaining link duration.
[0134] Secondly, each drone node will perform a distributed Q-learning process after updating the neighbor information through the Hello message. When the destination node is a neighbor of the current node, the current node will obtain the maximum reward and update the Q value according to formula (1). When the node's neighbor table does not contain the destination node, the node will gain experience based on the behavior strategy and calculate the reward value according to formula (11). The node will then adaptively adjust the learning rate and discount factor and iteratively update the Q-table. It is worth noting that the iteration will stop when any of the following conditions are met: 1) The maximum number of iterations k is reached max 2)
[0135] Finally, when the drone generates or receives a data packet, it queries the Q-table at the current moment and selects the neighbor node with the largest Q value as the next-hop relay node.
[0136] In another embodiment of the present invention, an adaptive interference-aware routing system for dynamic UAV networks is provided, which can be used to implement the above-mentioned adaptive interference-aware routing method for dynamic UAV networks. Specifically, the system includes:
[0137] The framework construction module is used to establish a routing protocol framework based on the Q-learning algorithm. It uses data packets as intelligent agents and the drone network as the environment. It maps data packet forwarding decisions into Q-learning state-action pairs through a state transition mechanism and constructs a Q-value iterative model that includes a learning rate α and a discount factor γ.
[0138] The reward function construction module is used to construct a reward function that integrates multiple indicators. It uses three indicators: the expected transmission time (ETX) that calculates link quality based on the signal-to-interference-noise ratio, the link duration predicted by the node's motion trajectory, and the distance difference that reflects the proximity to the destination node. This module transforms the dynamic characteristics of the UAV network into a quantifiable goal for reinforcement learning.
[0139] The optimization output module is used to update the learning rate α as a negative exponential function of the link duration based on the iterative model and the quantifiable goals reflected by the three indicators. The discount factor γ is adjusted by the complementary function of the link stability. The softmax strategy with temperature coefficient τ decay is used for action selection to form a closed-loop parameter optimization.
[0140] The module division in the embodiments of the present invention is illustrative and represents only one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in various embodiments of the present invention may be integrated into a single processor, exist physically as separate modules, or two or more modules may be integrated into a single module. The integrated modules may be implemented in either hardware or software functional modules.
[0141] In another embodiment of the present invention, a computer device is provided, which includes a processor and a memory, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used for the operation of the adaptive interference-aware routing method for dynamic drone networks.
[0142] In another embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device, used to store programs and data. It is understood that the computer-readable storage medium herein may include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for being loaded and executed by a processor. These instructions may be one or more computer programs (including program code). It should be noted that the computer-readable storage medium herein may be high-speed RAM memory or non-volatile memory, such as at least one disk drive. The processor may load and execute the one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the adaptive interference-aware routing method for dynamic drone networks described in the above-mentioned embodiment.
[0143] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0144] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0145] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0146] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. An adaptive interference-aware routing method for dynamic UAV networks, characterized by: include: A routing protocol framework based on the Q-learning algorithm is established, with data packets as the agent and the drone network as the environment. The packet forwarding decision is mapped into the state-action pair of Q-learning through the state transition mechanism, and a Q-value iterative model including the learning rate α and the discount factor γ is constructed. A reward function that integrates multiple indicators is constructed. By using three indicators, namely the expected number of transmissions (ETX) that calculates link quality based on the signal-to-interference-noise ratio (SINR), the link duration predicted by node motion trajectories, and the distance difference that reflects the proximity to the destination node, the dynamic characteristics of the UAV network are converted into quantifiable objectives for reinforcement learning. Based on the iterative model and the quantifiable goals reflected by the three indicators, the learning rate α is updated with the negative exponential function of the link duration, the discount factor γ is adjusted by the complementary function of the link stability, and the Softmax strategy with temperature coefficient τ decay is used for action selection to form a closed-loop parameter optimization.
2. The adaptive interference-aware routing method for dynamic UAV networks according to claim 1 is characterized in that: The construction of a Q-value iterative model including a learning rate α and a discount factor γ includes: By representing the data packet as an agent, the entire network is regarded as the environment, and the state space State consists of all drones used to forward the data packet; the selection of the next hop is regarded as an action Action, and the action space includes all available neighbor nodes; the state transition is equivalent to the packet forwarding decision between drones; that is, any data packet in the network environment has a state S t , represents the current drone node; the set of all actions that the data packet can take is the action space A, which consists of all available neighbor nodes; if the data packet is in S t An action a in A taken in state t , that is, a neighbor node is selected for forwarding, and the current state will be transferred to the next state S t+1 ; At the same time as the state transfer, the environment will feedback a reward R(S t ,a t ); Ultimately, the goal of the algorithm is to continuously try to learn the optimal strategy π in the environment, maximize the cumulative reward obtained by long-term execution of the strategy, and obtain the state S according to the strategy t The optimal action a to be performed next t =π(S t ); The object estimated by evaluating the policy π is the value function Q, which is calculated by the following formula: where Q t (s t ,a t ) is in state s t Take action a t The Q value when t+1 is the Q value at time t+1; α∈(0,1) represents the learning rate, and γ∈(0,1) represents the discount factor; Represents all possible actions a in the next state ′ ∈A t+1 The maximum Q value, where A t+1 is state s t+1 The action space in the Hello message is used to define and calculate the reward R(s) for different QoS indicators. t ,a t ); All drone nodes will maintain a neighbor table and a Q-value table by regularly exchanging Hello messages with their neighbors; the former is used to update the action space and calculate rewards, and the latter is used to determine routing decisions.
3. The adaptive interference-aware routing method for dynamic UAV networks according to claim 1 is characterized in that: The expected transmission times (ETX) used to calculate link quality based on the signal-to-interference-and-noise ratio, the link duration predicted by the node motion trajectory, and the distance difference reflecting the proximity to the destination node are three indicators, including: First, the BER is calculated using the mapping relationship between bit error rate and signal to interference and noise ratio; Among them, SINR represents the signal-to-interference-noise ratio at the receiving node, γ is a constant related to the bandwidth; Q(z) represents the probability that the random variable X in the Gaussian distribution X~N(0,1) is greater than the threshold z, that is, Therefore, the transmission success rate PDR of a data packet of size lbits can be calculated as: Among them, BER i represents the bit error rate of the i-th bit at the receiving node; Finally, the expected number of transmission times ETX of the data packet is obtained PDR f 、PDR r Respectively represent the forward and backward delivery success rates of data packets; After obtaining its own and its neighbors’ location information, the node estimates the link duration; d ij (t) represents the relative distance between nodes i and j at time t, and R is the communication coverage radius of the node. When nodes i and j receive each other's Hello message at time t2, they extract the corresponding node location information from the message. Then, nodes i and j estimate the link duration between themselves and each other based on the information. First, the distance between nodes i and j at time t is calculated according to the Euclidean distance formula Secondly, the relative distance change between nodes i and j is calculated by the following formula, Δd=d ij (t2)-d ij (t1)(7) Finally, the node estimates the link duration as Among them, T max =R / V min , V min is the minimum speed of the UAV node; The sign of Δd reflects the trend of the relative distance between nodes. When Δd is positive, the relative distance between two nodes gradually increases. From time t1 to t2, the relative speed between nodes is Δd / (t2-t1). Therefore, the expected time when the relative distance between two nodes exceeds the communication coverage is That is, the link duration; when Δd is negative, the two nodes are moving towards each other; the two nodes are moving at V min The time required to move the relative distance R at a relative speed is recorded as T max ; At time t2, d ij (t2) The ratio of R multiplied by T max Used as a predictor of link duration; The distance metric in the reward function is, Δd ij =d iD -d jD (9) where d iD d jD Represent the Euclidean distance between nodes i and j and the destination node respectively.
4. The adaptive interference-aware routing method for dynamic UAV networks according to claim 3 is characterized in that: The reward function includes: Among them, ETX is the estimated expected number of transmissions, which is used to measure the degree of interference between nodes; t LD (i,j) represents the duration of the link between nodes, Δd ij It is the difference between the distances of the node and the destination node, indicating that selecting a node closer to the destination node as a relay will obtain a greater reward; A, B, and C are weight parameters of various routing metrics, respectively, which satisfy A+B+C=1.
5. The adaptive interference-aware routing method for dynamic UAV networks according to claim 4 is characterized in that: There are two special cases of reward function design, as shown in formula (11), When the next hop node is the destination, the current node will receive the maximum reward R max To ensure that the route can eventually lead to the destination node; When all the neighboring nodes of a node are farther away from the destination node, the routing process falls into a local optimum; at this time, the current node should obtain the minimum reward R min .
6. The adaptive interference-aware routing method for dynamic UAV networks according to claim 5 is characterized in that: Based on the iterative model and the quantifiable goals reflected by the three indicators, the learning rate α is updated as a negative exponential function of the link duration, the discount factor γ is adjusted by the complementary function of the link stability, and the softmax strategy with temperature coefficient τ decay is used for action selection to form a closed-loop parameter optimization, including: An adaptive learning mechanism is introduced to dynamically adjust the values of α and γ using link stability; When the link between nodes is more unstable, that is, the link duration is shorter, the Q value needs to be updated faster, that is, a larger learning rate needs to be applied; the update formula for defining the learning rate α is Among them, t LD (i, j) is calculated by formula (8), i, j represent the current state s t and select action a t The corresponding pair of adjacent nodes; When the link between nodes is more unstable, a smaller γ needs to be applied; the update formula of the discount factor γ is defined as follows Among them, t LD (j,k) is calculated by formula (8), j and k represent the next state s t+1 And the action with the maximum Q value in the next state, that is The corresponding pair of adjacent nodes; Softmax method is used Among them, τ is a positive constant; in the iterative process, τ is updated by τ(t) = τ(0) / log2(1+t); initially, due to the large size of τ, the strategy probability π(s t ,a t ) are equal, which means that the strategy is more inclined to randomly select actions; as the time step increases, τ decreases, and the strategy probability π(s t ,a t ) gradually increases, which means that a more stable link is preferred.
7. The adaptive interference-aware routing method for dynamic UAV networks according to claim 6 is characterized in that: The protocol flow executed by each drone is as follows: First, drones maintain neighbor tables and interaction status information through Hello messages. When receiving a Hello message from another node, if the sender is new, it is added to the neighbor list and the corresponding status information is updated. Then, routing metrics such as the expected number of transmissions and link duration are calculated. If no Hello message is received from any other UAV, the neighbor node whose link has failed is removed based on the estimated remaining link duration; Secondly, each drone node will perform a distributed Q-learning process after updating its neighbor information through Hello messages; When the destination node is a neighbor of the current node, the current node will receive the maximum reward and update the Q value; When the node's neighbor table does not contain the destination node, the node will gain experience based on the behavior strategy and calculate the reward value. Then the node will adaptively adjust the learning rate and discount factor and iteratively update the Q-table. Finally, when the drone generates or receives a data packet, it queries the Q-table at the current moment and selects the neighbor node with the largest Q value as the next-hop relay node.
8. Adaptive interference-aware routing system for dynamic UAV networks, characterized by: include: The framework construction module is used to establish a routing protocol framework based on the Q-learning algorithm. It uses data packets as intelligent agents and the drone network as the environment. It maps data packet forwarding decisions into Q-learning state-action pairs through a state transition mechanism and constructs a Q-value iterative model that includes a learning rate α and a discount factor γ. The reward function construction module is used to construct a reward function that integrates multiple indicators. It uses three indicators: the expected transmission time (ETX) that calculates link quality based on the signal-to-interference-noise ratio, the link duration predicted by the node's motion trajectory, and the distance difference that reflects the proximity to the destination node. This module transforms the dynamic characteristics of the UAV network into a quantifiable goal for reinforcement learning. The optimization output module is used to update the learning rate α as a negative exponential function of the link duration based on the iterative model and the quantifiable goals reflected by the three indicators. The discount factor γ is adjusted by the complementary function of the link stability. The softmax strategy with temperature coefficient τ decay is used for action selection to form a closed-loop parameter optimization.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the adaptive interference-aware routing method for dynamic drone networks as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the adaptive interference-aware routing method for a dynamic drone network as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Software-defined network routing method based on multi-agent reinforcement learning
CN113556287A
Flight ad hoc network distributed routing method based on multi-agent reinforcement learning
CN119316336A
Multi-robot trajectory planning method
WO2022241808A1