Unmanned aerial vehicle network routing recovery method based on reinforcement learning and simulation test method
Patent Information
- Application Number
- CN202310688871.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-12
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-06-12
AI Technical Summary
[0005]针对现有技术中存在的问题,本发明提供了一种基于强化学习的无人机网络路由恢复方法,以解决无人机的路由路径恢复问题,同时还提供了相应的仿真测试方法
[0052]本发明一种基于强化学习的无人机网络路由恢复方法,用于解决无人机网络因蓄意攻击等情况导致路由中断后的恢复问题。本发明无人机网络路由恢复方法以最小化路由端到端总时延为目标,将重新规划路由和恢复问题表述为最小化目标函数问题,并利用强化学习算法进行求解,综合考虑了节点的状态信息,本发明方法不需要额外采集状态数据,所使用的数据均是路由过程中采样的数据,计算与能耗成本低,可以在较短的时间内恢复路由路径。同时,为了研究存在攻击的网络中路由路径规划与恢复问题,本发明还设计了一种基于节点重要性排序的蓄意攻击模型,用于测试所述基于强化学习的无人机网络路由恢复方法。
Smart Images

Figure CN116614855B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of UAV routing technology, and specifically relates to a UAV network routing recovery method and simulation testing method based on reinforcement learning. Background Technology
[0002] In next-generation mobile communication technologies, drones play a crucial role in fully connected and seamless communication coverage. Compared to traditional aircraft, drones offer advantages such as autonomous flight, high adaptability, strong payload capacity, good repeatability, and low cost, making them widely used in fields such as national defense, reconnaissance, real-time monitoring, surveillance, sampling, search and rescue, agriculture, manufacturing, and environmental monitoring.
[0003] While drones have wide applications in various fields, they also pose several security risks. First and foremost is the issue of privacy breaches. Some drones are equipped with high-definition cameras that can capture aerial images; if these images are misused, they could lead to user privacy leaks. Secondly, drones may collide with other aircraft or birds during flight, causing airspace safety issues. Thirdly, drones carry numerous sensors and data acquisition devices; if this data is stolen, sensitive information could be leaked. Finally, the drone's flight control system could be maliciously attacked by hackers, leading to malfunction or control. Therefore, to ensure drone safety, it is necessary to strengthen research on drone security technologies and applications, and improve the safety performance and reliability of drones.
[0004] Drone routing algorithms are a crucial research area for ensuring their security. In scenarios requiring multiple drones to collaborate on tasks, routing algorithms can coordinate data transmission and reception among drones and plan the shortest and most secure data transmission path to improve the efficiency and security of multi-drone collaboration. The distributed topology of drone networks makes them vulnerable to attacks that could disrupt routing; therefore, designing efficient route recovery algorithms is of great significance. Summary of the Invention
[0005] To address the problems existing in the prior art, this invention provides a method for UAV network route recovery based on reinforcement learning to solve the routing path recovery problem of UAVs, and also provides a corresponding simulation test method.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] A method for unmanned aerial vehicle (UAV) network route recovery based on reinforcement learning, characterized by the following steps:
[0008] Step 1) Obtain the network architecture of the drone swarm and describe the scenario for route planning and recovery based on the network architecture, including: determining the link status between drone nodes, the source drone node and destination drone node corresponding to the interrupted route path, as well as the normal drone node and the attacked drone node.
[0009] Assume that both the source drone node and the destination drone node are normal drone nodes;
[0010] Step 2) Based on the communication range of each drone node, select drone nodes within its communication range from its neighboring drone nodes and use them as candidate forwarding drone nodes.
[0011] Based on the candidate forwarding drone node set of each normal drone node, all routing paths from the source drone node to the destination drone node in the network architecture are analyzed and used as candidate routing paths.
[0012] Step 3) Based on the free space path loss model between drone nodes, obtain the maximum permissible transmission rate of a single hop between each drone node and the next-hop drone node in each candidate route path. Then, combine the communication distance between each drone node and the next-hop drone node and the length of the data packets waiting to be transmitted in the next-hop drone node to calculate the single-hop data transmission delay and obtain the end-to-end delay model from the source drone node to the destination drone node.
[0013] Step 4) Transform the routing planning and recovery problem from the source UAV node to the destination UAV node into the objective function problem of minimizing the end-to-end delay model;
[0014] Step 5) Model the minimization objective function problem using a Markov decision process, and use a reinforcement learning algorithm to solve the minimization objective function problem modeled as a Markov decision process to obtain the optimal strategy with the minimum end-to-end delay value. Based on the routing path selected by the optimal strategy, restore the communication between the source UAV node and the destination UAV node.
[0015] Based on the above solutions, further improvements or preferred solutions include:
[0016] Furthermore, in step 1), N UAV nodes are randomly and uniformly deployed in the network architecture of the UAV swarm, denoted as G = (U, E), where:
[0017] U={u i ,i=1,2,…,N},u i Represents the i-th drone node;
[0018] E={e ij,i=1,2,…,N,j=1,2,…,N},e ij ∈{0,1} is a binary variable representing the unmanned aerial vehicle node u. i and drone node u j The link status between them, if e ij =1 indicates that there is a communication link between the two drone nodes. ij =0 indicates that there is no communication link between the two drone nodes;
[0019] Using the set F = {f i {i = 1, 2, ..., N} represents the unmanned aerial vehicle (UAV) node u. i Is it under attack? If it is under attack, then f i =1, otherwise f i =0.
[0020] Furthermore, in step 2), the communication range of each UAV node is limited to a maximum radius of O. max and minimum radius O min Inside the hollow sphere;
[0021] drone node u i The position is represented by coordinates (x) i ,y i ,z i ) indicates that when the drone node u i and u j The Euclidean distance d between them ij Greater than 0 max When the value is d, the two drone nodes cannot communicate; while when d ij Less than O min When the value is equal, a collision will occur; therefore, each drone node u i The set of candidate forwarding drone nodes Θ i ={u j |O min <d ij <O max}
[0022] Furthermore, in step 3), the end-to-end delay model is defined as:
[0023]
[0024] Where, p sd ={u s →u d} represents the source drone node u s to the destination drone node u d A complete route path, p ij and These represent the complete routing path p. sd UAV node u i to u j The single-hop path and time delay, Γ(p sd ) represents the total latency of a complete routing path.
[0025] Furthermore, in step 4), the objective function for minimizing the end-to-end delay model is expressed as:
[0026]
[0027]
[0028] u s ,u d ∈U
[0029]
[0030] Wherein, D(u) s →u d P) represents the routing scheme. For source drone node u s to the destination drone node u d Let b be the set of candidate routing paths, and b be the b-th candidate routing path. This represents the empty set.
[0031] Furthermore, reinforcement learning algorithms are used to solve for the optimal policy π that minimizes the end-to-end delay. * At that time, it corresponds to finding the optimal behavioral value function value Q* of the corresponding UAV node in each state of the Markov decision process, so that it executes the action a corresponding to Q*. t * Proceed to the next state;
[0032] Each drone node maintains a Q-table to store each state-behavior pair (s) of that drone node. t ,a t The value function Q(s) t ,a t The value function Q(s) t ,a t Iterative updates are performed using the following formula:
[0033] Q′(s t ,a t )=Q(s t ,a t )+αδ t E(s t ,a t )
[0034] Where, Q′(st ,a t ) represents the updated Q(s) t ,a t ), where α represents the learning rate, δ t Represents the TD error, E(s) t ,a t ) is a qualification record;
[0035] δ t Defined as: δ t =r t+1 +γQ(s t+1 ,a t+1 )-Q(s t ,a t ), r t+1 This represents the reward value function, where γ represents the discount factor;
[0036] The qualification trace is defined as: E(s) t ,a t ): λ∈(0,1) is the degeneracy parameter, 1(s t =s,a t =a) is a conditional expression; when s t =s0,a t The value is 1 when a = 0, and 0 otherwise.
[0037] The optimal behavior value function value Q* is the state-behavior pair (s) t ,a t The maximum value function Q) max .
[0038] Furthermore, each state s t The following uses the ε-greedy method to select action a t :
[0039]
[0040] Where ε∈(0,1) is the probability of the exploration action, Q max The above formula represents the maximum value function value, where an action is randomly selected from the action space with a preset probability value ε, or selected with a probability value of 1-ε to generate the maximum behavioral value function value Q. max The action.
[0041] Furthermore, the Markov decision process consists of a quintuple <S,A,P,r,γ>, where:
[0042] (a) State space S: each state s t =(d ij ,LP j ) t∈S, where s t Indicates the drone node u at time t i The state, d ij For drone node u i to u j Euclidean distance d ij LP j For drone node u j The length of the queued data packets to be transmitted will be LP j Defined as: LP j =η j ×l j η j and l j Representing the unmanned aerial vehicle (UAV) node u j The number and size of data packets waiting to be transmitted;
[0043] (b) Action Space A: Definition in, Representing drone node u i with u k For a single jump action targeting the next jump, the action space size corresponds to the set Θ. i The number of drone nodes in the middle; a t =A(s) t ), a t Indicates the drone node u i In state s t The action to be performed;
[0044] (c) Transition probability P ss' : Represents drone node u i In state s t Next, execute action a t When entering state s t+1 The probability, P ss' =ε or 1-ε;
[0045] (d) Reward function r: The objective is to minimize the total delay of the routing path from the source drone node to the destination drone node. Let r be... t+1 For the current state s t Take action a t Instantaneous reward, r t+1 Represented as: Among them, H k It is a marker indicating the end of a route, and is defined as
[0046] (e) Discount factor γ: γ∈[0,1], the larger the value of γ, the more it focuses on long-term returns.
[0047] A simulation test method for a UAV network route recovery method based on reinforcement learning is characterized by designing a node importance ranking mechanism that considers node degree and link importance, and establishing a deliberate attack model based on the mechanism. The deliberate attack model is used to simulate a deliberate attack launched against a UAV swarm to determine the attacked UAV.
[0048] The importance of the node is defined as follows:
[0049]
[0050] in, Indicates the drone node u i For link e ij Significant contribution, k i Indicates the drone node u i The degree of the nodes.
[0051] Compared with the prior art, the technical solution provided by the present invention has the following beneficial effects:
[0052] This invention discloses a reinforcement learning-based route recovery method for unmanned aerial vehicle (UAV) networks, addressing the recovery problem after route interruption caused by deliberate attacks or other unauthorized actions. The method aims to minimize the total end-to-end routing latency, formulating the route replanning and recovery problem as a minimization of an objective function, which is then solved using a reinforcement learning algorithm. Taking into account node state information, this method does not require additional state data collection; all data used is sampled during the routing process, resulting in low computational and energy costs and enabling route path recovery in a short time. Furthermore, to investigate route planning and recovery in attacked networks, this invention also designs a deliberate attack model based on node importance ranking to test the reinforcement learning-based UAV network route recovery method. Attached Figure Description
[0053] Figure 1 This is a schematic diagram of the unmanned aerial vehicle (UAV) network architecture involved in this invention;
[0054] Figure 2 This is a schematic diagram of the network node importance ranking topology involved in this invention;
[0055] Figure 3 This is a schematic diagram of the distance calculation method between drones involved in this invention;
[0056] Figure 4 This is a simulation result of the relationship between routing path delay and the number of UAVs in this invention;
[0057] Figure 5 The figure shows the experimental simulation results of the relationship between the number of steps and the number of iterations in this invention;
[0058] Figure 6 The figure shows the experimental simulation results of the relationship between the average number of steps and distance and the number of jumps in this invention. Detailed Implementation
[0059] The invention will now be described in further detail with reference to the accompanying drawings.
[0060] Figure 1 This paper demonstrates a specific embodiment of a drone network routing planning and recovery scenario under deliberate attack.
[0061] Figure 2 The diagram illustrates a topology used to calculate the importance ranking of nodes in a drone network. Specifically, based on the definition of the importance formula, the importance values of drone nodes u2 and u5 can be calculated as L, respectively. u2 =26.9 and L u5 =52.15, L u5 >L u2 This indicates that u5 is more important than u2, which is consistent with reality.
[0062] Figure 3 The diagram illustrates a scenario of the Euclidean distance calculation method between drones provided by this invention.
[0063] exist Figure 4-6 In this invention, the reinforcement learning algorithm Sarsa(λ) is used to solve the routing path recovery problem of a drone network under deliberate attack. It is compared and analyzed with two other reinforcement learning algorithms, Sarsa and Q-learning, as detailed below:
[0064] Figure 4 This paper illustrates the relationship between routing latency and the number of drones in this invention. The number of drones increases sequentially from 10 to 40, and the end-to-end routing latency is recorded for each different number of drones. It can be seen that the average system latency increases exponentially with the number of drones, because increasing the number of drones leads to a rapid increase in network complexity. When a drone in the original routing path is attacked, the path recovery time becomes longer. The algorithm proposed in this invention has lower latency than the other two reinforcement learning methods, meaning it can better solve the routing recovery problem.
[0065] Figure 5 The relationship between the number of steps and iterations in this invention is illustrated. In the reinforcement learning-based UAV network route recovery method under deliberate attack, the convergence of the proposed algorithm is characterized by the variation of the step count curve with the iteration process (number of UAVs N=20). It can be observed that the curve of the reinforcement learning algorithm Sarsa(λ) used in this invention has less fluctuation. This demonstrates that the proposed algorithm not only effectively solves the routing problem but also converges faster.
[0066] Figure 6 The relationship between average steps and distance and hop count in the method of this invention is shown. The hop count of the UAV routing path increases sequentially from 2, 3, to 4. The average steps and distance of the simulation algorithm under different hop counts are recorded. It can be seen that the routing cost increases with the increase of hop count. The reinforcement learning algorithm Sarsa(λ) used in this invention performs better than the other two comparison methods.
[0067] Example 1:
[0068] A method for unmanned aerial vehicle (UAV) network route recovery based on reinforcement learning includes the following steps:
[0069] Step 1) Obtain the network architecture of the drone swarm whose routing is interrupted due to a deliberate attack. Describe the scenario for routing planning and recovery through the network architecture, including: determining the link status between drone nodes, the source drone node and destination drone node corresponding to the interrupted routing path, as well as normal drone nodes and attacked drone nodes (each drone in the drone swarm corresponds to a node in its network architecture).
[0070] In this embodiment, both the source drone node and the destination drone node are assumed to be normal drone nodes.
[0071] Step 2) Based on the communication range of each drone node, select drone nodes within its communication range from its neighboring drone nodes and use them as candidate forwarding drone nodes.
[0072] Based on the candidate forwarding drone node set of each normal drone node, all routing paths from the source drone node to the destination drone node in the network architecture are analyzed and used as candidate routing paths.
[0073] Step 3) Based on the free space path loss model between drone nodes, obtain the maximum permissible transmission rate of a single hop between each drone node and the next-hop drone node in each candidate route path. Then, combine the communication distance between each drone node and the next-hop drone node and the length of the data packets waiting to be transmitted in the next-hop drone node to calculate the single-hop data transmission delay and obtain the end-to-end delay model from the source drone node to the destination drone node.
[0074] Step 4) Transform the routing planning and recovery problem from the source UAV node to the destination UAV node into the objective function problem of minimizing the end-to-end delay model;
[0075] Step 5) Model the minimization objective function problem using a Markov decision process, and solve the minimization objective function problem modeled as a Markov decision process using a reinforcement learning algorithm to obtain the optimal strategy with the minimum end-to-end delay value. Restore communication between the source UAV node and the destination UAV node based on the routing path selected by the optimal strategy.
[0076] In the above process:
[0077] In step 1), N UAV nodes are randomly and uniformly deployed in the network architecture of the UAV swarm, denoted as G = (U, E), where:
[0078] U={u i ,i=1,2,…,N},u i Represents the i-th drone node;
[0079] E={e ij ,i=1,2,…,N,j=1,2,…,N},e ij ∈{0,1} is a binary variable representing the unmanned aerial vehicle node u. i and drone node u j The link status between them, if e ij =1 indicates that a communication link exists between the two drone nodes. Figure 1 The line is represented by an undirected dashed line; if e ij =0 indicates that there is no communication link between the two drone nodes;
[0080] Using the set F = {f i {i = 1, 2, ..., N} represents the unmanned aerial vehicle (UAV) node u. i Is it under attack? If it is under attack, then f i =1, otherwise f i =0.
[0081] In step 2), the communication range of each UAV node is limited to a maximum radius of O. max and minimum radius O min Inside the hollow sphere;
[0082] drone node u i The position is represented by coordinates (x, y). i ,y i ,z i ) indicates that when the drone node u i and u j The Euclidean distance d between them ij Greater than 0 max When the value is d, the two drone nodes cannot communicate; while when d ij Less than O minWhen the value is equal, a collision will occur; therefore, each drone node u i The set of candidate forwarding drone nodes Θ i ={u j |O min <d ij <O max}.by Figure 1 Taking the candidate forwarding drone set of drone node u4 in the network structure as an example: Θ4={u3,u5,u6}.
[0083] In step 3), the end-to-end delay model is defined as follows:
[0084]
[0085] Where, p sd ={u s →u d} represents the source drone node u s to the destination drone node u d A complete route path, p ij and These represent the complete routing path p. sd UAV node u i to u j The single-hop path and time delay, Γ(p sd The total delay of a complete routing path is denoted as . Defined as: Where d ij and R ij They represent u respectively i and u j The distance between them and the maximum permissible transmission rate, where ν represents the signal transmission speed, LP j R represents the length of the data packets waiting to be transmitted. ij Defined as R ij =B ij log2(1+SNR ij ), where B ij For signal transmission bandwidth, SNR ij It is the signal-to-noise ratio, and the calculation formula is: Among them, P ij , Representing the unmanned aerial vehicle (UAV) node u i and u j The transmission power and noise power between them, and the free space path loss are: PL ij =20log 10 (d ij )+20log 10 (g)-147.55, where g represents the frequency.
[0086] In step 4), the objective function for minimizing the end-to-end delay model is expressed as:
[0087]
[0088]
[0089] u s ,u d ∈U
[0090]
[0091] Wherein, D(u) s →u d P) represents the routing scheme. For source drone node u s to the destination drone node u d Let b be the set of candidate routing paths, and b be the b-th candidate routing path. This represents the empty set.
[0092] Step 5) In reinforcement learning, policy π guides the agent to choose action a in state s. Based on the model established by the Markov decision process and the determined policy π, a routing path can be obtained. Therefore, the minimization objective function problem in step 4) can be transformed into finding the optimal policy π that minimizes the total end-to-end delay. * .
[0093] The optimal policy π that minimizes end-to-end delay is found using reinforcement learning algorithms. * When the optimal behavior value function value Q* of the corresponding UAV node in each state of the Markov decision process is found, it executes the action at* corresponding to Q* and moves to the next state;
[0094] Each drone node maintains a Q-table to store each state-behavior pair (s) of that drone node. t ,a t The value function Q(s) t ,a t The value function Q(s) t ,a t Iterative updates are performed using the following formula:
[0095] Q′(s t ,a t )=Q(s t ,a t )+αδ t E(s t ,a t )
[0096] Where, Q′(st ,a t ) represents the updated Q(s) t ,a t ), where α represents the learning rate, δ t Represents the TD (Temporal Difference) error, E(s) t ,a t ) is a qualification record;
[0097] δ t Defined as: δ t =r t+1 +γQ(s t+1 ,a t+1 )-Q(s t ,a t ), r t+1 This represents the reward value function, where γ represents the discount factor;
[0098] This embodiment uses cumulative eligibility trace updates, where the eligibility trace is represented as follows:
[0099] E(s t ,a t ): λ∈(0,1) is the degeneracy parameter, 1(s t =s,a t =a) is a conditional expression; when s t =s0,a t The value is 1 when a = 0, otherwise the value is 0;
[0100] The optimal behavior value function value Q* is the state-behavior pair (s) t ,a t The maximum value function Q) max .
[0101] In addition, to balance exploration and exploitation during the continuous interaction process of reinforcement learning and avoid getting trapped in local optima, a mechanism was designed for each state s. t The following uses the ε-greedy method to select action a t :
[0102]
[0103] Where ε∈(0,1) is the probability of the exploration action, Q max The above formula represents the maximum value function value, where an action is randomly selected from the action space with a preset probability value ε, or selected with a probability value of 1-ε to generate the maximum behavioral value function value Q. max The action.
[0104] Step 5) designs the Markov decision process to consist of a quintuple <S,A,P,r,γ>, specifically:
[0105] (a) State space S: In order for the agent to perceive the environmental state, two routing-related parameters are used to form a state space set, namely the Euclidean distance d. ij and queue length of data packets waiting to be transmitted (LP) j That is, each state s t =(d ij ,LP j ) t ∈S, where:
[0106] s t Indicates the drone node u at time t i The state, d ij For drone node u i to u j Euclidean distance d ij LP j For drone node u j The length of the queued data packets to be transmitted will be LP j Defined as: LP j =η j ×l j η j and l j Representing the unmanned aerial vehicle (UAV) node u j The number and size of data packets waiting to be transmitted;
[0107] (b) Action Space A: Definition in, Representing drone node u i with u k For a single jump action targeting the next jump, the action space size corresponds to the set Θ. i The number of drone nodes in the middle; a t =A(s) t ), a t Indicates the drone node u i In state s t Next, select the action to be executed from the action space S;
[0108] (c) Transition probability P ss' : Represents drone node u i In state s t Next, execute action a t When entering state s t+1 The probability, P ss' =ε or 1-ε;
[0109] (d) Reward value function r: To minimize the total delay of the routing path from the source drone node to the destination drone node, let r... t+1 For the current state s t Take action a t The instantaneous reward is represented as: Among them, H k It is a marker indicating the end of a route, and is defined as
[0110] (e) Discount factor γ: γ is used to consider the immediate reward at the current moment as well as the immediate reward at future moments. γ∈[0,1], and the larger the value of γ, the more emphasis is placed on long-term reward. Reward is the reward obtained after performing an action or a series of actions. The immediate reward is the reward that can be obtained immediately after performing the current action, while the long-term reward is the reward obtained after performing a series of actions. The principle of reinforcement learning is to learn to obtain the maximum long-term reward through a series of optimal actions.
[0111] Example 2:
[0112] To study the routing path planning and recovery problem in attacked networks, this embodiment provides a simulation test method considering intentional attacks, used to test the UAV network routing recovery method based on reinforcement learning as described in Embodiment 1.
[0113] In this testing method, a node importance ranking mechanism that considers node degree and link importance is designed, and a deliberate attack model is established based on this mechanism. The deliberate attack model is used to simulate a deliberate attack on a drone swarm to determine the attacked drone.
[0114] The importance of the node is defined as follows:
[0115]
[0116] in, Indicates the drone node u i For link e ij Significant contribution, k i Indicates the drone node u i The degree of the node. Defined as: Where, k j Indicate u j degree, k i Indicate u i The degree, Indicate u i and u j Link e between ij The importance of Defined as Where Z = (k i -m-1)(k j -m-1), where m indicates that the drone network topology contains link e. ij The number of triangles.
[0117] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should be considered within the scope of protection of the present invention.
Claims
1. A method for unmanned aerial vehicle (UAV) network route recovery based on reinforcement learning, characterized in that, Includes the following steps: Step 1) Obtain the network architecture of the drone swarm and describe the scenario for route planning and recovery based on the network architecture, including: determining the link status between drone nodes, the source drone node and destination drone node corresponding to the interrupted route path, as well as the normal drone node and the attacked drone node. Assume that both the source drone node and the destination drone node are normal drone nodes; Step 2) Based on the communication range of each drone node, select drone nodes within its communication range from its neighboring drone nodes and use them as candidate forwarding drone nodes. Based on the candidate forwarding drone node set of each normal drone node, all routing paths from the source drone node to the destination drone node in the network architecture are analyzed and used as candidate routing paths. Step 3) Based on the free space path loss model between drone nodes, obtain the maximum permissible transmission rate of a single hop between each drone node and the next-hop drone node in each candidate route path. Then, combine the communication distance between each drone node and the next-hop drone node and the length of the data packets waiting to be transmitted in the next-hop drone node to calculate the single-hop data transmission delay and obtain the end-to-end delay model from the source drone node to the destination drone node. Step 4) Transform the routing planning and recovery problem from the source UAV node to the destination UAV node into the objective function problem of minimizing the end-to-end delay model; Step 5) Model the minimization objective function problem using a Markov decision process, and use a reinforcement learning algorithm to solve the minimization objective function problem modeled as a Markov decision process to obtain the optimal strategy with the minimum end-to-end latency. Based on the routing path selected by the optimal strategy, restore the communication between the source UAV node and the destination UAV node.
2. The method for unmanned aerial vehicle (UAV) network route recovery based on reinforcement learning according to claim 1, characterized in that: In step 1), there is a set The drone nodes are randomly and uniformly deployed in the network architecture of the drone swarm, denoted as [network architecture name missing]. ,in: , Representing the One drone node; , It is a binary variable representing the drone node. and drone nodes The link status between them, if This indicates that a communication link exists between the two drone nodes. This indicates that there is no communication link between the two drone nodes; Use sets Indicates drone node Has it been attacked? If attacked, then ,on the contrary .
3. The method for unmanned aerial vehicle (UAV) network route recovery based on reinforcement learning according to claim 2, characterized in that: In step 2), the communication range of each drone node is limited to a maximum radius of... and minimum radius is Inside the hollow sphere; drone nodes Position using coordinates It indicates that when the drone node and Euclidean distance between Greater than When the value is [value], the two drone nodes cannot communicate; while when [value] Less than When the value is equal, a collision will occur; therefore, each drone node Candidate forwarding drone node set .
4. The method for unmanned aerial vehicle (UAV) network route recovery based on reinforcement learning according to claim 3, characterized in that: In step 3), the end-to-end delay model is defined as follows: in, Indicates the source drone node To the destination drone node A complete routing path, and These represent the complete routing paths. Chinese drone node arrive The single-hop path is delayed in time. The total latency of a complete routing path.
5. The method for unmanned aerial vehicle (UAV) network route recovery based on reinforcement learning according to claim 4, characterized in that: In step 4), the objective function for minimizing the end-to-end delay model is expressed as: in, Indicates the routing scheme. Source drone node to the destination drone node The set of candidate routing paths Indicates the first There are 10 candidate routing paths, where Ø represents an empty set.
6. The method for unmanned aerial vehicle (UAV) network route recovery based on reinforcement learning according to claim 5, characterized in that: Solve the optimal policy that minimizes end-to-end latency using reinforcement learning algorithms. At that time, it corresponds to finding the optimal behavioral value function value of the corresponding UAV node in each state of the Markov decision process. Q* To enable its execution and Q* corresponding actions * Proceed to the next state; Each drone node maintains a Q-table to store each state-behavior pair of that drone node. , The value function of ) Q ( , The value function Q ( , Iterative updates are performed using the following formula: in Indicates the updated Q ( , ), Indicates the learning rate. Represents TD error, For qualification traces; Will Defined as: , Represents the reward value function, Represents the discount factor; Define eligibility trace as: : , It is a degradation parameter. To evaluate the expression, when The value is 1 if the condition is met, and 0 otherwise. The optimal behavior value function value Q* That is, state-behavior pairs ( , Maximum value function .
7. The method for unmanned aerial vehicle (UAV) network route recovery based on reinforcement learning according to claim 6, characterized in that, Each state The following is adopted -greedy method to select actions : in, To explore the probability of actions, The above formula represents the maximum value function value, expressed as a preset probability value. Randomly select an action from the action space, or select an action with a probability value of 1- Select the function value that produces the maximum behavioral value The action.
8. The method for unmanned aerial vehicle (UAV) network route recovery based on reinforcement learning according to claim 7, characterized in that, The Markov decision process consists of a quintuple. Composition, in which: (a) State space Each state =( , ) t ∈ S ,in, express t Time Drone Node state, For drone nodes arrive European distance , For drone nodes The length of the data packets waiting to be transmitted will be... Defined as: , and Representing drone nodes The number and size of data packets waiting to be transmitted; (b) Action space :definition ,in, Representing drone nodes by A single jump action for the next jump target, the action space size corresponds to the set. The number of drone nodes in China; =A ( ), Indicates drone node In state The action to be performed; (c) Transition probability : Represents a drone node In state Next action Entering state The probability, = or 1- ; (d) Reward value function The objective is to minimize the total latency of the routing path from the source drone node to the destination drone node. Current state Take action below Instant reward, Represented as: ,in, It is a marker indicating the end of a route, and is defined as ; (e) Discount factor : , The larger the value, the more emphasis is placed on long-term returns.
9. A simulation testing method for the UAV network route recovery method based on reinforcement learning as described in any one of claims 1-8, characterized in that, The design considers node importance ranking mechanism based on node degree and link importance, and establishes a deliberate attack model based on this mechanism. The deliberate attack model is used to simulate a deliberate attack on a drone swarm to identify the attacked drones. The importance of the node is defined as follows: in, Indicates drone node For the link Significant contributions Indicates drone node The degree of the nodes.
Citation Information
Patent Citations
Multi-unmanned aerial vehicle joint task allocation and flight path planning method under limited communication
CN115981369A
Distributed unmanned aerial vehicle ad hoc network routing method based on deep reinforcement learning
CN116234073A