A Multi-Agent Routing Method for Unmanned Aerial Vehicle Self-Organizing Networks

By treating drone nodes as agents and using reinforcement learning models to perform multi-agent routing methods for drone self-organized networks, the delay and overhead problems of traditional routing protocols in drone networks are solved, and efficient and reliable communication is achieved.

CN116113008BActive Publication Date: 2025-07-1810TH RES INST OF CETC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211545965.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-05
Publication Date
2025-07-18
Estimated Expiration
2042-12-05

AI Technical Summary

Technical Problem

The routing protocols of traditional drone ad hoc networks cannot meet the communication needs of low latency, high reliability, scalability and low network overhead, especially in the high-speed movement of drones and unstable links, and cannot sense the topological state in time and make correct routing decisions.

Method used

The multi-agent routing method is adopted to treat the drone nodes as agents, and data packet forwarding and routing decisions are made through reinforcement learning models. The agent obtains link stability and hop information through message interaction, and independently makes the optimal routing decisions, without relying on the whole network topology information.

Benefits of technology

It reduces communication delay and network overhead, improves the reliability and resilience of the drone network, ensures efficient and reliable communication, and adapts to the high-speed movement and topological changes of the drone network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116113008B_ABST
    Figure CN116113008B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-agent routing method for an unmanned aerial vehicle (UAV) ad-hoc network. Each UAV with a forwarding function is regarded as an agent. When an agent generates a data packet with a destination node d, it starts a reinforcement learning model to obtain a routing action for data distribution. The agent that receives the data packet replies to the sending agent with an ACK packet according to the requirements of the reinforcement learning model. The packet carries reward information to help the sending agent complete the iteration of reinforcement learning and detects whether it is the destination node of the data packet. If it is not the destination node, it starts the reinforcement learning algorithm to continue forwarding the data packet. The data packet finally reaches the destination node through the hop-by-hop forwarding of several agents, completing the transmission of the data packet. Each agent perceives the network environment by continuously interacting with other agents through packets, so that when it needs to forward a data packet, it can make an optimal routing decision through the trained reinforcement learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of communication technologies, and particularly relates to a multi-agent routing method for an unmanned aerial vehicle (UAV) ad-hoc network. Background Art

[0002] In recent years, UAVs have developed rapidly due to their broad application prospects and are widely used in military, surveying and mapping, navigation, communication relay and other fields. The payload of a single UAV is limited and its functions are relatively single. To make up for the deficiencies of a single UAV, UAVs act in a cluster manner, form an ad-hoc network through interaction, and cooperate to complete tasks together. The UAV ad-hoc network is a fast-moving ad-hoc network that does not rely on a ground control station and is constructed when needed. Compared with traditional large UAVs, the UAV ad-hoc network has better task execution capabilities, better fault tolerance and robustness, and higher economic affordability. Therefore, the UAV networking technology has become a hot topic in the current development of UAVs.

[0003] As the requirements for the performance of the UAV network in tasks such as sensing, navigation, and rescue are getting higher and higher, it is necessary to ensure the efficient and reliable communication of the UAV network. Currently, the communication requirements of the UAV network include low latency, high reliability, scalability, high adaptability, and low network overhead. The low latency requires the UAV network to maintain a low transmission latency according to the task scenario; the high reliability requires the UAVs to directly establish a reliable transmission network to improve the success rate of data packet delivery; the scalability requires the UAV network to adjust the cluster scale with the change of task requirements on the premise of ensuring normal communication; the high adaptability requires the UAV network to repair the network and maintain normal communication when encountering problems such as topology changes and node failures; the low network overhead requires the UAV network to reduce the communication overhead and improve the bandwidth utilization efficiency.

[0004] Routing protocol is a core issue in the research of UAV ad hoc network communication. Currently, the Optimized Link State Routing (OLSR) protocol is commonly used in UAV networks. The OLSR protocol requires nodes to periodically send HELLO messages and TC messages to obtain the global network topology information. As a proactive routing protocol, OLSR has the advantage of fast routing convergence. When UAVs are performing tasks, they need to move at high speed. The network links may be intermittently interrupted due to the movement of UAV nodes, and the channel quality is not stable. When the sending periods of HELLO messages and TC messages are long, the OLSR routing protocol cannot timely perceive the topology state and cannot make correct routing decisions; when the sending periods of HELLO messages and TC messages are short, the control messages will occupy valuable communication resources and cannot meet the low communication overhead requirements of UAV networks. When the scale of the UAV network is large, the OLSR protocol takes a long time to obtain the global topology, which also increases the transmission delay of messages. In summary, traditional routing protocols cannot well adapt to the communication scenarios of UAV networks.

[0005] Moreover, since drones need to move at high speed when performing tasks, the links in the drone ad-hoc network may be because there is no one Due to the movement of UAV nodes, intermittent interruptions occur, and the channel quality is not stable. Traditional routing protocols need to send a large number of control messages to timely perceive the states of all links in the network. A large number of control messages will increase the network overhead, occupy valuable communication resources, and do not meet the low-overhead communication requirements of UAV networks.

[0006] Traditional routing protocols need to perceive the entire network topology before making routing decisions. In large-scale UAV ad hoc networks, traditional routing protocols take a lot of time to complete the perception of the entire network topology, which causes data messages not to be forwarded in a timely manner, increasing the communication delay of data messages and not meeting the low-delay communication requirements of UAV networks. Summary of the Invention

[0007] The purpose of the present invention is to disclose a multi-agent routing method for UAV ad hoc networks to overcome the problems of the prior art. In this method, UAVs with routing forwarding functions are regarded as agents. Agents obtain information such as link stability, the number of hops required to reach each node, and message queuing situations through message interactions. Agents continuously interact with other agents through messages to complete the perception of the network environment, and can make optimal routing decisions through a trained reinforcement learning model when data messages need to be forwarded.

[0008] The purpose of the present invention is achieved through the following technical solutions:

[0009] A multi-agent routing method for an unmanned aerial vehicle (UAV) ad-hoc network. In the multi-agent routing method, each UAV with a forwarding function is regarded as an agent. The multi-agent routing method for the UAV ad-hoc network includes: when an agent generates a data packet with a destination node d, starting a reinforcement learning model to obtain a routing action for data distribution; the agent receiving the data packet replies to the sending agent with an ACK packet according to the requirements of the reinforcement learning model, and the packet carries reward information to help the sending agent complete the iteration of reinforcement learning, and detects whether it is the destination node of the data packet; if it is not the destination node, it starts the reinforcement learning algorithm to continue forwarding the data packet; the data packet finally reaches the destination node through the hop-by-hop forwarding of several agents, completing the transmission of the data packet, and each agent completes the perception of the network environment by continuously interacting with other agents.

[0010] According to a preferred embodiment, the routing action includes broadcasting or unicasting the data packet to a preset agent node.

[0011] According to a preferred embodiment, the reinforcement learning model is established in the following manner:

[0012] An agent obtains a reward by interacting with the environment to correct its action, which is represented as a Markov decision process or MDP process. The MDP process is described as a triple <A, S, R>, where A represents the set of actions, S represents the set of environmental states, and R represents the feedback of the environment to the agent after the agent gives an action.

[0013] At time t, the agent is in state s t where s t ∈S; according to the policy, it makes an action a t , a t ∈A, and receives a reward r t from the environment, r t ∈R, and enters the next state s t+1 . The agent completes the decision-making by repeatedly executing the foregoing MDP process.

[0014] According to a preferred embodiment, let S i (d, j) represent the state where agent i forwards a data packet with destination node d through node j. Among them, the state S i (d, j) includes two variables Q and C,

[0015] where the variable Q i (d, j) represents the estimated value of the number of hops required for agent i to forward a data packet with destination node d through node j; C i (d, j) represents the evaluation value of the link quality for agent i to forward a data packet with destination node d through node j. The maximum evaluation value is 1, and the minimum is 0.

[0016] According to a preferred embodiment, the action of the agent is set to A i (d, j), A i (d, j) represents that the agent forwards a message with the destination node d to node j;

[0017] When the agent needs to forward a message with the destination node d, it extracts all Q i in the state of (d, j) i (d, j) and C i (d, j) information. The agent i first selects the next-hop node j according to the following formula min , and obtains the action A i (d, j min ). If several js are obtained min , then broadcast the message

[0018]

[0019] In the formula, J is the set of all nodes in the network, and α and β are factors that balance the expected hop count Q i (d, j) and the link quality C i (d, j) importance.

[0020] According to a preferred embodiment, when the agent i makes the action A i (d, j), for the state variables Q i (d, j) and C i (d, j), the rewards are as follows:

[0021] When the agent i forwards the nth data message with the destination node d to node j, each data message has a corresponding sequence number. After receiving the message, the agent j immediately replies to the agent i with an ACK message, and the message carries the rewards Q ACK (d) and C ACK (d),

[0022] Q ACK (d) and C ACK (d) are obtained as shown in the following formula. In the formula, K is the set of all nodes in the UAV network,

[0023] Q ACK (d) = Q j (d, k ack )

[0024] C ACK (d) = C j (d, k ack )

[0025]

[0026] If the node that receives the data packet is the destination node, the replied Q ACK value is 0, and the C ACK value is 1. After agent i receives the ACK packet replied by agent j, it obtains Q ACK and C ACK , and the return function of the state variable Q i (d, j) is as follows:

[0027]

[0028] In the formula, the parameter γ represents the learning efficiency of the agent, where 0 < γ < 1;

[0029] After agent i forwards the packet, it starts timing. If agent j is in a congested state, then when the time elapsed after agent i forwards and receives the ACK exceeds T, the derivation formula of T is as follows:

[0030] T = T D + T ACK + T C + T S

[0031] Among them, T D represents the maximum queuing time of the packet when agent j is not in a congested state, T S represents the transmission delay of the nth data packet sent by agent i, T ACK represents the transmission delay of agent j replying ACK, T C represents the propagation delay between agent i and agent j;

[0032] If agent i receives the ACK packet within T time after forwarding the packet, it is considered that node j is not in a congested state and the link passing through node j is relatively stable, and C i (d, j) is iterated according to the following formula:

[0033]

[0034] If agent i does not receive the ACK packet within T time, it is considered that node j is in a congested state. At this time, C i (d, j) is iterated according to the following formula:

[0035] The main solution of the present invention and its various further alternative solutions described above can be freely combined to form multiple solutions, all of which are solutions that can be adopted and claimed by the present invention. Those skilled in the art can understand that there are various combinations according to the prior art and common general knowledge after understanding the solution of the present invention, and all of them are the technical solutions to be protected by the present invention, and will not be enumerated here.

[0036] Advantages of the present invention:

[0037] The present invention proposes a multi-agent routing method for an unmanned aerial vehicle (UAV) ad-hoc network. In the routing method, each UAV node is regarded as an agent, and the agent makes routing decisions independently, which is in line with the distributed architecture of the UAV ad-hoc network. The agent obtains network information through message interaction and makes routing decisions according to the reinforcement learning algorithm. Routing decisions can be made without obtaining the network-wide topology, greatly reducing the communication delay, reducing the communication overhead, and ensuring the efficient and reliable communication of the UAV network. This is mainly reflected in:

[0038] (1) The routing method in the present invention adopts a multi-agent architecture. Each UAV node with forwarding function is regarded as an agent. The agent independently collects information and makes decisions without relying on the control of a central node and land facilities, which is suitable for the distributed architecture of the UAV ad-hoc network; the failure of a single node will not affect the communication of the network. Therefore, the UAV network adopting the multi-agent routing method will not have the problem of single-point failure, improving the resilience of the network.

[0039] (2) In the present invention, the UAV node does not need to obtain the network-wide topology and can complete routing decisions through message interaction. Therefore, it does not need to frequently send a large number of control messages like traditional routing protocols, effectively reducing the network overhead and saving precious communication resources.

[0040] (3) The present invention reflects the congestion degree and stability of the link through the link quality in the state variable. When making routing decisions, the agent will avoid sending data messages through links with poor link quality, avoiding the emergence of congested links in the network, reducing the queuing delay of data messages, and at the same time improving the transmission success rate of data messages, enabling the UAV network to communicate reliably and efficiently.

[0041] (4) The present invention proposes a new routing action decision mechanism. As long as the link quality reaches a relatively high level, it can meet the communication requirements of the UAV node. Fluctuations in the link quality at a relatively high level will not have a great impact on the routing decisions of the agent. The routing decision mechanism in the present invention reflects the characteristics of the link quality and will not react violently to slight changes in the link quality at a relatively high level, optimizing the routing decisions.

[0042] (5) The present invention proposes a reinforcement learning convergence acceleration mechanism, enabling the agent to obtain the rewards corresponding to all nodes on the path through one message interaction, accelerating the acquisition of environmental information by the agent and realizing the acceleration of the agent's reinforcement learning convergence. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a schematic diagram of the reinforcement learning process of the reinforcement learning model of the present invention;

[0044] Figure 2 It is the ACK message format under the convergence acceleration mechanism;

[0045] Figure 3 It is the schematic diagram of the topology of the UAV ad hoc network;

[0046] Figure 4 It is the schematic diagram of the transmission situation of Message No. 1 in the network;

[0047] Figure 5 It is the schematic diagram of the transmission situation of Message No. 2 in the network;

[0048] Figure 6 It is the schematic diagram of the transmission path of the message after the method converges. Detailed implementation manners

[0049] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0050] The present invention discloses a multi-agent routing method for a UAV ad hoc network, in which each UAV with a forwarding function is regarded as an agent in the multi-agent routing method.

[0051] The multi-agent routing method for the UAV ad hoc network of the present invention includes: when an agent generates a data message with a destination node d, start the reinforcement learning model to obtain a routing action (broadcast or unicast to a preset agent node) for data distribution.

[0052] The agent that receives the data message replies to the sending agent an ACK message according to the requirements of the reinforcement learning model. The message carries reward information (Q ACK and C ACK ) to help the sending agent complete the iteration of reinforcement learning and detect whether it is the destination node of the data message; if it is not the destination node, start the reinforcement learning algorithm to continue forwarding the data message; the data message finally reaches the destination node through the hop-by-hop forwarding of several agents, completing the transmission of the data message. Each agent completes the perception of the network environment by continuously interacting with other agents through messages.

[0053] The agent continuously perceives the network environment by interacting with other agents through messages. When it needs to forward data messages, it can make the optimal routing decision through the trained reinforcement learning model.

[0054] In the traditional reinforcement learning routing algorithm, after the UAV node forwards the data message with the destination node d, it can only obtain the reward about node d for the agent to learn. If the node needs to obtain information about multiple destination nodes, it needs to send multiple messages, which not only brings a huge overhead to the UAV network, but also takes a lot of time to complete the algorithm convergence, and does not meet the requirements of the UAV network for low latency and low overhead.

[0055] Specifically, the present invention proposes a reinforcement learning convergence acceleration mechanism: when the agent receives a message with the destination node d, it needs to immediately reply to the sender with an ACK message. If the node knows the path to node d, it generates Q ACK and C ACK corresponding to each node passed by the path and replies them to the sender through the ACK message. After the UAV node A forwards the data message with the destination node D to node B under this mechanism, the ACK message replied by node B is as Figure 2 shown. Node B knows that its path to the destination node C is B→C→D, and then fills the IDs of the nodes passed by the path and the corresponding Q ACK and C ACK into the ACK message in sequence. If node A receives this ACK message, it can know that the ACK message is sent by node B, know that its path to the destination node is A→B→C→D, and complete the iteration of the state variables Q ACK (D) and C ACK (D) in the agent through Q A (D,B) and C A (D,B). Through Q ACK (C) and C ACK (C), complete the iteration of Q A (C,B) and C A (C,B). Through Q ACK (B) and C ACK (B), complete the iteration of Q A (B,B) and C A (B,B).

[0056] Through the above mechanism, when the node forwards the message to the destination node, it can obtain the reward about each node on the link, accelerate the agent's acquisition of environmental information, and achieve the acceleration of the convergence of the reinforcement learning algorithm.

[0057] Furthermore, the reinforcement learning model of the present invention is established in the following way:

[0058] Figure 1 The reinforcement learning process is given. The agent modifies its actions by obtaining rewards through interaction with the environment. Its behavioral steps can be represented as a Markov Decision Process (MDP), and the MDP process can be described as a triple <A, S, R>, where A represents the set of actions, S represents the set of states, and R represents the feedback from the environment to the agent after the agent takes an action. At time t, the agent is in state s t (s t ∈ S) makes an action a t (a t ∈ A) according to the policy, and receives a reward r t (r t ∈ R), and enters the next state s t+1 , and the above process is looped to complete the decision-making.

[0059] (1) State

[0060] In the UAV ad hoc network, the nodes move at high speeds, the topology changes violently, and the communication environment is complex. In order for the agent to make correct routing decisions, appropriate state variables need to be determined. There is a low-latency communication requirement in the UAV ad hoc network, and the delay for a data packet to reach the destination node can be divided into transmission delay, processing delay, queuing delay, and propagation delay. The propagation delay of the data packet can be ignored. The transmission delay and processing delay required for the data packet to reach the destination node are determined by the number of hops for the service to reach the destination node. The fewer the number of hops for the service to reach the destination node, the less time is required to transmit and process the data packet. The queuing delay is determined by the congestion situation of the link. When the number of packets waiting in the queues of each forwarding node on the link exceeds a certain amount, the node cannot process the queued packets in a timely manner, and the link will become congested. The more serious the congestion situation, the greater the queuing delay. There is a high-reliability communication requirement in the UAV ad hoc network, which requires that the links through which the data packets pass have a certain stability to ensure that the data packets will not be lost during the transmission process.

[0061] In summary, we use S i (d, j) to represent the state of agent i forwarding a data packet with destination d through node j. The state S i (d, j) includes two variables Q and C. Among them, the variable Q i (d, j) represents the estimated value of the number of hops required for agent i to forward a data packet with destination d through node j. The larger Q i (d, j) is, the more expected hops agent i needs to reach the destination through node j. C i (d, j) represents the evaluation value of the link quality for agent i to forward a data packet with destination d through node j. The maximum value of the evaluation value is 1, and the minimum value is 0. C iThe larger (d, j) is, the higher the stability of the link through which the agent reaches the destination via node j, the lower the congestion level, and the higher the link quality.

[0062] (2) Action

[0063] Selecting the next action is the main goal of reinforcement learning. In the multi-agent routing method, the action of the agent is set to A i (d, j), A i (d, j) represents that the agent forwards the packet with the destination node d to node j. When the agent needs to forward a packet with the destination node d, it extracts all the Q i in the state of (d, j) i (d, j) and C i information of (d, j). First, agent i selects the next-hop node j according to the following formula min , and gets the action A i (d, j min ). If several js are obtained min , the packet is broadcast.

[0064]

[0065] J in the formula is the set of all nodes in the network, and α and β are factors that balance the importance of the expected hop count Q i (d, j) and the link quality C i (d, j). This formula aims to select a routing action with as small an expected hop count as possible and a relatively high link quality.

[0066] According to the characteristics of the network, when the stability of the link reaches a relatively high value, the demand for reliable communication can be met, and the impact of the link stability fluctuating at a relatively high level on routing is not significant; when the congestion level of the link is at a relatively low level, the node can process the queued packets in time without causing excessive queuing delay, and the impact of the link congestion level fluctuating at a relatively low level on routing is not significant.

[0067] In summary, when the link quality is at a relatively high level, fluctuations will not have a significant impact on the selection of actions, and when the link quality is at a relatively low level, it will significantly affect the selection of routing actions. In order to reflect the characteristics of the link quality in the action selection strategy, the variable C i (d, j) representing the link quality is located in the denominator of the inverse proportional function. According to the derivative property of the inverse proportional function, when C i (d, j) fluctuates at a relatively large value, the value of the inverse proportional function will not change significantly. When C i (d, j) is at a relatively low value, the inverse proportional function will show a relatively high value, significantly affecting the selection of routing actions; when C i(d, j) placed in the denominator part of the inverse proportional function can reflect the characteristics of the link quality. The n0 in the formula adjusts the sensitivity of the agent to the link quality. The smaller n0 is, the more sensitive the agent is to the link quality.

[0068] To prevent the agent from falling into local optimum during the decision-making process, this method adds an ε-greedy mechanism during the process of determining the routing action. After the agent selects node j as the next hop, there is a probability of ε (0 < ε < 1) to broadcast the data packet to enhance the agent's perception ability of the environment.

[0069] (3) Reward

[0070] After the agent takes a decision-making action, it receives the reward feedback from the environment, and the cumulative reward value is used as the basis for the next decision. When agent i takes action A i (d, j), for the state variables Q i (d, j) and C i (d, j), the reward is as follows: When agent i forwards the nth data packet with the destination node d to node j, each data packet has a corresponding sequence number. After agent j receives the packet, it immediately sends an ACK packet back to agent i, and the packet carries the reward Q ACK (d) and C ACK (d), Q ACK (d) and C ACK (d). The acquisition of Q

[0071] Q ACK (d) = Q j (d, k ack )

[0072] C ACK (d) = C j (d, k ack )

[0073]

[0074] If the node that receives the data packet is the destination node, the replied Q ACK value is 0, and the C ACK value is 1. After agent i receives the ACK packet replied by agent j, it obtains Q ACK and C ACK , and the reward function of the state variable Q i (d, j) is as follows:

[0075]

[0076] The parameter γ (0 < γ < 1) in the formula reflects the learning efficiency of the agent. The larger γ is, the faster the agent learns the rewards of the nodes. However, an overly large γ will cause the agent to react violently to changes in the network, which is not conducive to the agent making optimal routing decisions.

[0077] Traditional routing algorithms do not consider the congestion level of the link. In this paper, C i (d,j) is used to measure the congestion level and stability of the link. It is assumed that the queue length that the UAV node can process in a timely manner is q. When it exceeds q, the node is considered to be in a congested state. Let the average length of the message be L bits, and the processing speed of each node for the message be v bits / s. The maximum queuing time T of the message in the non-congested state can be obtained D :

[0078]

[0079] Meanwhile, the transmission delay of the nth data packet sent by agent i is T S , the transmission delay T ACK of agent j to reply ACK, and the propagation delay T C between agent i and agent j. After agent i forwards the message and starts timing, if agent j is in a congested state, the time elapsed for agent i to receive ACK after forwarding must exceed T. The derivation formula of T is as follows:

[0080] T = T D + T ACK + T C + T S

[0081] If agent i receives the ACK message within the time T after forwarding the message, it is considered that node j is not in a congested state and the link passing through node j is relatively stable. C i (d,j) is iterated according to the following formula:

[0082]

[0083] If agent i does not receive the ACK message within the time T, it is considered that node j is in a congested state, or the link stability is low resulting in message loss. Even if the ACK message is received later, the C ACK (d) in it is not received. At this time, C i (d,j) is iterated according to the following formula:

[0084]

[0085] Through the above reward process, C i (d,j) reflects the link stability and the congestion level of the link, and embodies the quality of the link.

[0086] Application Example

[0087] A typical drone network is as follows Figure 3 shown. In the initialization phase of the network, each node has no knowledge of the network state, and all Q values are set to the maximum value Q max , and all link qualities are set to the maximum value of 1.

[0088] Node A generates a data packet for the first time with the destination node G, and the sequence number of the packet is 1. According to the reinforcement learning algorithm, several equal j values are calculated min , then the broadcast packet is selected. After nodes B and C receive the packet, they immediately reply with ACK packets. Since B and C also have no knowledge of G, Q ACK (g) is Q max , and C ACK (G) is 1. Nodes B and C also select the broadcast packet. Nodes D, E, and F in the network successively receive the packet and reply with ACK, and continue to select the broadcast. Finally, node G receives the packets from E and F. Since G is the destination node of the packet, Q ACK (G) in the ACK packet is set to 1, and C ACK (G) is set to 1. After node E receives the packet, it learns that the path to the destination node is E→G, and completes the training of the agent through Q ACK (G) and C ACK (G). The same applies to node F. The transmission of the No. 1 packet in the network is as follows Figure 4 shown.

[0089] Node A generates a data packet for the second time with the destination node G, and the packet sequence number is 2. According to the reinforcement learning algorithm, only the broadcast packet can be selected. Nodes B, C, and D also select the broadcast packet. After node F receives the packet from node C, according to the fast convergence mechanism, on the path from F to G, nodes G and F are passed. According to Q F (G,j) and C F (G,j), the corresponding Q ACK (G) and C ACK (G) are generated. According to Q F (F,j) and C F (F,j), the corresponding Q ACK (F) and C ACK(F) and reply to node C through an ACK message. Node C learns that the path from itself to the destination node F is C→F→G through the ACK message, and at the same time collects the rewards for nodes G and F, accelerating the training of the agent and speeding up the convergence of the reinforcement learning algorithm. Since node E cannot receive the data message from node D due to unstable links and cannot reply with an ACK, node D makes Q according to the reward process of reinforcement learning after waiting for time T D (G,E) is iterated according to the following formula:

[0090]

[0091] Since F has learned the knowledge about node G, it directly forwards the data message to node G through the reinforcement learning algorithm to complete the transmission of the data message. The transmission of the No. 2 message in the network is as Figure 5 shown.

[0092] The present invention proposes a multi-agent routing method for an unmanned aerial vehicle (UAV) ad hoc network. In the routing method, each UAV node is regarded as an agent, and the agents make routing decisions independently, which is in line with the distributed architecture of the UAV ad hoc network. The agents obtain network information through message interaction and make routing decisions according to the reinforcement learning algorithm. Routing decisions can be made without obtaining the whole network topology, greatly reducing the communication delay, reducing the communication overhead, and ensuring the efficient and reliable communication of the UAV network. It is mainly reflected in:

[0093] (1) The routing method in the present invention adopts a multi-agent architecture. Each UAV node with a forwarding function is regarded as an agent. The agents collect information and make decisions independently, without relying on the control of a central node and land facilities, and is adapted to the distributed architecture of the UAV ad hoc network; the failure of a single node will not affect the communication of the network. Therefore, the UAV network adopting the multi-agent routing method will not have the problem of single point of failure, improving the resilience of the network.

[0094] (2) In the present invention, UAV nodes do not need to obtain the whole network topology and can complete routing decisions through message interaction. Therefore, it is not necessary to frequently send a large number of control messages like traditional routing protocols, effectively reducing the network overhead and saving precious communication resources.

[0095] (3) The present invention shows the congestion degree and stability of the link through the link quality in the state variable. When making routing decisions, the agents will avoid sending data messages through links with poor link quality, avoiding the emergence of congested links in the network, reducing the queuing delay of data messages, and at the same time improving the success rate of data message transmission, enabling the UAV network to communicate reliably and efficiently

[0096] (4) The present invention proposes a new routing action decision-making mechanism. As long as the link quality reaches a relatively high level, it can meet the communication requirements of the UAV nodes. Fluctuations in the link quality at a relatively high level will not have a great impact on the routing decision of the agent. The routing decision-making mechanism in the present invention reflects the characteristics of the link quality and will not react violently to slight changes in the link quality at a relatively high level, thus optimizing the routing decision.

[0097] (5) The present invention proposes a reinforcement learning convergence acceleration mechanism, enabling the agent to obtain the rewards corresponding to all nodes on the path through a single message interaction, accelerating the acquisition of environmental information by the agent, and achieving the acceleration of the agent's reinforcement learning convergence.

[0098] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A multi-agent routing method for an unmanned aerial vehicle self-organizing network, characterized in that, In the multi-agent routing method, each unmanned aerial vehicle (UAV) with forwarding function is regarded as an agent. The multi-agent routing method for UAV ad-hoc network includes: When an agent generates a data packet with the destination node d, it starts the reinforcement learning model to obtain a routing action and conducts data distribution; The agent that receives the data packet replies an ACK packet to the sender agent according to the requirements of the reinforcement learning model. The packet carries reward information to help the sender agent complete the iteration of reinforcement learning and detects whether it is the destination node of the data packet. If it is not the destination node, it starts the reinforcement learning algorithm to continue forwarding the data packet; The data packet finally reaches the destination node through the hop-by-hop forwarding of several agents, completing the transmission of the data packet. Each agent completes the perception of the network environment by continuously interacting with other agents through packets; The reinforcement learning model is established in the following way: The agent interacts with the environment and corrects its actions through the obtained immediate reward. This process is established as a Markov decision process MDP, which is abstracted as a standard triple: <A, S, R>. A represents the set of actions of the agent, S represents the set of standard environmental states, and R represents the immediate reward value obtained after the agent interacts with the environment; At time t, the agent is in state s t where s t ∈ S; make an action a t according to the policy, a t ∈ A, and receive a reward r t from the environment, r t ∈ R, and enter the next state s t+1 The agent completes the decision-making by repeatedly executing the foregoing MDP process; Let S i (d, j) represents the state of agent i forwarding a data packet with destination d through node j, where the state S i (d, j) includes two variables Q and C Among them, variable Q i (d, j) represents the estimated number of hops required for agent i to forward a data packet with destination d through node j; C i (d, j) represents the evaluation value of the link quality for agent i to forward a data packet with destination d through node j. The maximum evaluation value is 1 and the minimum is 0; The action of the agent is set to A i (d, j), A i (d, j) represents that the agent forwards a packet with the destination node d to node j; When the agent needs to forward a packet with the destination node d, all S are extracted i Q in the (d, j) state i (d, j) and C i (d, j) information. Agent i first selects the next-hop node j according to the following formula min , and obtains the action A i (d, j min ). If several js are obtained min , the packet is broadcast In the formula, J is the set of all nodes in the network, and α and β are factors that balance the expected hop count Q i (d, j) and the link quality C i (d, j) factors of importance; When agent i takes action A i (d,j), for state variables Q i (d,j) and C i (d,j), the rewards are as follows: When agent i forwards the nth data packet with the destination node d to node j, each data packet has a corresponding sequence number. After receiving the packet, agent j immediately replies to agent i with an ACK packet, and the packet carries the return Q regarding node d ACK (d) and C ACK (d), Q ACK (d) and C ACK (d) is obtained as shown in the following formula, where K is the set of all nodes in the UAV network, Q ACK (d) = Q j (d, k ack ) C ACK (d) = C j (d, k ack ) If the node that receives the data packet is the destination node, the replied Q ACK value is 0, and the C ACK value is 1. After agent i receives the ACK packet replied by agent j, it obtains Q ACK and C ACK . The reward function of the state variable Q i (d,j) is as follows: The parameter γ in the formula represents the learning efficiency of the agent, where 0 < γ < 1; After the agent i forwards the packet, it starts timing. If the agent j is in a congested state, the time passed by the agent i after forwarding and receiving the ACK exceeds T. The derivation formula of T is as follows: T = T D + T ACK + T C + T S Among them, T D represents the maximum queuing time of packets when agent j is not in a congested state, T S represents the transmission delay of the nth data packet sent by agent i, T ACK represents the transmission delay of agent j to reply ACK, T C represents the propagation delay between agent i and agent j; If agent i receives an ACK packet within T time after forwarding a packet, it is considered that node j is not in a congested state and the link passing through node j is relatively stable, C i (d,j) is iterated according to the following formula: If agent i does not receive the ACK packet within time T, it is considered that node j is in a congested state, and at this time C i (d,j) is iterated according to the following formula:

2. The multi-agent routing method according to claim 1, characterized in that, The routing action includes broadcasting or unicasting the data packet to a preset agent node.

Citation Information

Patent Citations

  • Multi-agent reinforcement learning routing algorithm based on geographic position

    CN112804726A

  • Interference and mobility management in UAV-assisted wireless networks

    US20170118688A1