A Cooperative Autonomous Data Routing Method for Multi-UAV Systems
Through the multi-head attention pruning mechanism and multi-agent reinforcement learning algorithm, the data transmission challenges of multi-UAV systems in complex environments are solved, efficient and stable data routing strategies are realized, and the system's task execution capabilities in complex environments are improved.
Patent Information
- Application Number
- CN202510670147.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-23
AI Technical Summary
In a complex and changeable communication environment, multi-UAV systems face the problems of frequent changes in network topology, unstable communication, high data transmission delay and low throughput, and are difficult to meet the transmission needs of high concurrent and multi-source heterogeneous data.
Using the multi-head attention pruning mechanism and multi-agent reinforcement learning algorithm, a collaborative autonomous data routing method for multi-UAV systems is constructed through action space compression and collaborative decision-making, including inter-node accessibility assessment, deployment of policy networks and value networks, and Markov decision-making process model for multi-UAV routing collaborative decision-making, and optimize data routing strategies.
It significantly improves the data transmission stability and efficiency of multi-UAV systems, can quickly adapt to environmental changes, and improves the system's task execution capabilities in complex environments.
Smart Images

Figure CN120200956B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a data routing method, in particular to a collaborative autonomous data routing method for a multi-UAV system. Background Art
[0002] The information provided in this part is only background information related to the present disclosure, and it is not necessarily prior art.
[0003] With the rapid development of the information age, multi-UAV (Unmanned Aerial Vehicle) systems are increasingly widely used in fields such as disaster emergency and environmental monitoring. The collaborative operation of multi-UAVs can significantly improve the efficiency and effectiveness of task execution due to its advantages of flexibility, high efficiency, wide coverage, and strong adaptability. However, in the actual operation process of multi-UAV systems, they face challenges brought by complex and changeable communication environments and system limitations. On the one hand, during the flight of UAVs, affected by factors such as airflows and terrain, the network topology structure changes frequently; at the same time, the communication between UAVs and with the ground control station is vulnerable to problems such as electromagnetic interference and signal occlusion, seriously affecting the stability and reliability of data transmission. On the other hand, when multi-UAV systems execute tasks, they often need to transmit a large amount of real-time data, such as high-definition images, video streams, sensor monitoring data, etc. Traditional data routing methods are difficult to meet the requirements of low latency and high throughput when dealing with high-concurrency, multi-source heterogeneous data, and are prone to problems such as data congestion and low routing efficiency, thus affecting the collaborative execution of overall tasks of multi-UAV systems.
[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] Object of the Invention: The technical problem to be solved by the present invention is to provide a collaborative autonomous data routing method for a multi-UAV system in view of the deficiencies of the prior art.
[0006] To solve the above technical problem, the present invention discloses a collaborative autonomous data routing method for a multi-UAV system, including the following steps:
[0007] Step 1, in the multi-UAV system, based on a multi-head attention pruning mechanism, perform action space compression to obtain an effective neighbor set and an effective action set for each UAV;
[0008] Step 2: Deploy the policy network in each UAV, deploy the value network in the control center, and based on the obtained valid neighbor set and valid action set in Step 1, use the policy network and value network to make independent decisions for individual UAVs and collaborative decisions for the UAV cluster, obtaining the preliminary decisions for each UAV.
[0009] Step 3: Construct a multi-UAV routing collaborative decision-making Markov decision process model to optimize the preliminary decisions of all UAVs, and complete the collaborative autonomous data routing for the multi-UAV system.
[0010] Furthermore, obtaining the valid neighbor set and valid action set for each UAV in Step 1 includes the following steps:
[0011] Step 1-1: In the multi-UAV system, let the source node i.e., the UAV broadcast detection data packets to the surrounding area to collect the status information of neighbor nodes i.e., the UAV .
[0012] Step 1-2: Encode the status information of neighbor nodes into query vectors, key vectors, and value vectors.
[0013] Step 1-3: Perform attention feature extraction based on the query vectors, key vectors, and value vectors, and conduct node reachability assessment through the multi-head attention mechanism.
[0014] Step 1-4: Mask high-risk nodes based on the node pruning mechanism, and generate a valid neighbor set and a valid action set according to the node reachability assessment results.
[0015] Step 1-5: Update the local view according to the valid neighbor set and valid action set, where the local view is the network topology structure maintained inside the source node after screening the valid neighbor set.
[0016] Furthermore, encoding the status information of neighbor nodes into query vectors, key vectors, and value vectors in Step 1-2 includes the following steps:
[0017] ;
[0018] Among them, , and are used to represent the query vector of the source node , the key vector of the neighbor node , and the value vector of the neighbor node respectively; the query vector represents the query request of node to other nodes, reflecting node Query requirement for potential routing paths; key vector Representative node The characteristics or status of, used to match with the query vector to determine the node For node Degree of attention; value vector Contains node Details of, used for node Decision-making; , And Represent learnable weight matrices.
[0019] Furthermore, the reachability assessment between nodes described in steps 1 - 3 includes the following steps:
[0020] Let the reachability assessment function between node and node be , expressed as follows:
[0021] ;
[0022] Where is the attention function, and the calculation method is as follows:
[0023] ;
[0024] Among them, is the softmax function, is the dimension of the key vector;
[0025] According to the k different dimensions that determine the reachability between node and node , the multi-head attention mechanism is introduced for multi-perspective parallel calculation. Then, the reachability assessment function between node and node is expressed as:
[0026] ;
[0027] Among them, represents the learnable weight matrix of the output layer, represents the operation of concatenating multiple vectors, represents the k-th attention head.
[0028] Furthermore, the k-th attention head , and the calculation method is as follows:
[0029] ;
[0030] Among them, , and represent the learnable weight matrix of the k-th attention head.
[0031] Furthermore, generating the effective neighbor set and the effective action set according to the node reachability evaluation result described in steps 1-4 includes the following steps:
[0032] According to the reachability evaluation function from node to node , perform reachability judgment between nodes, which is expressed as follows:
[0033] ;
[0034] Among them, is the node pruning function, is the threshold of reachability. If the value of the reachability evaluation function between node and node is less than the threshold, the value of the node pruning function is 0, and it is judged as unreachable, indicating that node needs to prune node ;
[0035] The neighbor node set after pruning is the effective neighbor set , which is expressed as follows:
[0036] ;
[0037] The effective action set of node is , which is expressed as follows:
[0038] .
[0039] Furthermore, using the policy network and the value network for independent decision-making of a single UAV and collaborative decision-making of a UAV cluster described in step 2 includes the following steps:
[0040] Step 2-1, assume that each UAV selects an action according to the current policy network, obtains a reward after interacting with the environment, and its next state is ; Store the global experience in the experience replay buffer. The update process of the experience replay buffer is as follows:
[0041] ;
[0042] Step 2-2: Set up target networks corresponding to the policy network and the value network, train and update them to obtain the trained policy network and value network;
[0043] Among them, the policy network update method is expressed as follows:
[0044] ;
[0045] Among them, represents the parameters of the policy network, represents the policy gradient, which is used to update the parameters of the policy network , represents the expected value sampled from the experience replay buffer according to the policy ; represents the action taken in the state ; represents the estimated advantage function, which is expressed as follows:
[0046] ;
[0047] Among them, is the action-value function, which represents the expected return of taking the action in the state and following the policy ; is the state-value function, which represents the expected return of following the policy in the state ;
[0048] The update of the state-value function , that is, the value network update method is expressed as follows:
[0049] ;
[0050] Among them, is the parameter of the value network, represents the value function gradient, which is used to update the parameters of the value network, represents the state distribution under the policy , represents the state distribution under the policy sampled from, is the value function of the target network, which is expressed as follows:
[0051] ;
[0052] Among them, is the immediate reward, is the discount factor, is the state transition probability;
[0053] The method for updating the parameters of the target network is as follows:
[0054] ;
[0055] Among them, is an adjustable parameter used to control the speed of updating the parameters of the target network;
[0056] Step 2-3, the trained policy network selects the maximum action according to the effective neighbor set and the effective action set , as the preliminary decision, is expressed as follows:
[0057] ;
[0058] Among them, the trained policy network is , indicating that the action selected in the state is , are the parameters of the policy network.
[0059] Further, the optimization of the preliminary decisions for all drones described in Step 3 includes:
[0060] Each drone makes a routing decision based on local observations. Through the reward function, each drone independently selects the next-hop node, which is expressed as follows:
[0061] ;
[0062] Among them, represents the joint probability distribution of the next state and the reward given the current state and the action , represents the joint probability distribution of the next state and the reward given the current state and the action and all previous states and actions.
[0063] Further, the said reward , is calculated through the reward function of each drone . The reward function is expressed as follows:
[0064] ;
[0065] Among them, , and are adjustable parameters, is the forward propagation reward of the data packet, is the environmental reward, is the minimum hop count reward, is the delay reward.
[0066] Furthermore, the forward propagation reward of the data packet , which is used to indicate whether the data has reached the sink node, i.e., the end point of the data packet transmission, and at the same time, a step penalty is imposed to limit the behavior of the nodes to retransmit data back and forth. The calculation method is as follows:
[0067] ;
[0068] where is the positive reward for successful data arrival, is the penalty based on the time step;
[0069] The environmental reward , the calculation method is as follows:
[0070] ;
[0071] where represents the wind speed magnitude at node , represents the environmental noise intensity within the area where the node is located, and respectively represent the weight factors that control the influence of wind speed and environmental noise intensity on stability;
[0072] The minimum hop count reward , the calculation method is as follows:
[0073] ;
[0074] where is the shortest path length from the next-hop node of the decision to the sink node ;
[0075] The delay reward , the calculation method is as follows:
[0076] ;
[0077] where is an adjustable parameter, is the normalized propagation delay, and its value range is [0, 1]. The calculation method is as follows:
[0078] ;
[0079] Among them, and respectively represent the maximum and minimum propagation delays among nodes.
[0080] Beneficial effects:
[0081] 1. The present invention fully considers the problems faced by multi-UAV systems, such as high-dimensional action spaces and low efficiency of traditional learning algorithms. Through an action space compression mechanism, unnecessary action exploration is effectively reduced, and the speed of routing decision-making is significantly improved. Combining an experience replay mechanism, a multi-head attention mechanism, and a target network soft update technique, the learning stability and efficiency of the algorithm are significantly enhanced, enabling the routing strategy to quickly adapt to environmental changes.
[0082] 2. In terms of the routing collaborative decision-making strategy of the present invention, the carefully designed reward function comprehensively considers various factors such as forward propagation of data packets, environmental factors, minimum hops, and delay, guiding the UAV agent to make better routing decisions. It provides reliable data transmission support for the collaborative operation of multi-UAV systems in complex environments, enhances the execution ability of the system in various tasks (such as complex environment monitoring, emergency communication guarantee, collaborative combat, etc.), and significantly improves the overall performance and application value of multi-UAV systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] The following further describes the present invention in detail with reference to the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0084] Figure 1 is an architecture diagram of a multi-UAV system based on hierarchical control according to the present invention.
[0085] Figure 2 is a schematic diagram of a multi-head attention mechanism according to the present invention.
[0086] Figure 3 is a schematic diagram of the process of obtaining a dynamic local UAV cluster view according to the present invention.
[0087] Figure 4 is an architecture diagram of a MAL-MHA algorithm according to the present invention.
[0088] Figure 5 is a training flowchart of a MAL-MHA algorithm according to the present invention.
[0089] Figure 6 is a flowchart of UAV routing implementation according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0090] The overall idea of the present invention is as follows: In view of the numerous challenges existing in data transmission in a multi-UAV system under a complex and changeable network environment, such as unstable network topology caused by frequent node movement, vulnerable communication links, etc., an innovative data routing optimization system is constructed based on the unique attributes of UAVs. On this basis, by introducing a multi-head attention pruning mechanism and a multi-agent reinforcement learning algorithm based on action state space constraints, a highly intelligent collaborative autonomous data routing scheme is proposed. The present invention fully considers the problems faced by the multi-UAV system, such as high-dimensional action space and low efficiency of traditional learning algorithms. Through an action space compression mechanism, unnecessary action exploration is effectively reduced, and the speed of routing decision-making is greatly improved. Combining the experience replay mechanism, the multi-head attention mechanism and the target network soft update technology significantly enhances the learning stability and efficiency of the algorithm, enabling the routing strategy to quickly adapt to environmental changes. In terms of the routing collaborative decision-making strategy, a carefully designed reward function comprehensively considers various factors such as positive propagation of data packets, environmental factors, minimum hops and delay, guiding the UAV agent to make better routing decisions. The present invention aims to provide reliable data transmission support for the collaborative operation of the multi-UAV system in a complex environment, enhance the execution ability of the system in various tasks (such as complex environment monitoring, emergency communication guarantee, collaborative combat, etc.), and significantly improve the overall performance and application value of the multi-UAV system.
[0091] The present invention proposes a collaborative autonomous data routing scheme for a multi-UAV system. Based on the multi-head attention pruning mechanism, an action space compression mechanism for a multi-agent reinforcement learning (MARL) algorithm is constructed, and on this basis, a multi-UAV routing collaborative decision-making algorithm based on action state space constraint MARL is constructed for data transmission and sharing among multiple UAVs.
[0092] As Figure 1 shown, the overall technical solution of the present invention is as follows: For a multi-UAV system, a hierarchical control architecture for the multi-UAV system is designed, and all UAVs are divided into a control center UAV-R and ordinary UAV nodes UAV-L. Among them, UAV-L is responsible for managing the entire cluster network, and UAV-R is responsible for performing collaborative navigation tasks such as intelligence reconnaissance, data relay, and data collection. Among UAV-Ls, through self-organizing control, data routing transmission is carried out according to actual business objectives to achieve intelligent data sharing and distributed storage. The present invention constructs a collaborative autonomous data routing scheme for a multi-UAV system based on multi-agent reinforcement learning. Based on the multi-head attention pruning mechanism, an action space compression mechanism for the MARL algorithm is constructed, and on this basis, a multi-UAV routing collaborative decision-making algorithm based on action state space constraint MARL is constructed for data transmission and sharing among multiple UAVs. The specific steps are as follows:
[0093] (1)Action space compression mechanism based on multi-head attention pruning mechanism
[0094] (1.1)Node reachability evaluation based on multi-head attention mechanism
[0095] Based on the idea of the multi-head attention mechanism in Transformer, this invention explores the local network structure through the multi-head attention mechanism in the multi-UAV network routing decision problem, and continuously updates the UAV node state and the view of neighbor nodes that the node can reach. Different from the existing research on the fusion of attention mechanism and MARL, this invention targets the action space of UAV nodes, not just the communication between UAVs.
[0096] For the multi-UAV network architecture, design the node to node The reachability evaluation function between them is , and the formula of the reachability evaluation function based on the attention mechanism is as follows:
[0097]
[0098] Among them, 、 and are used to represent the query vector, key vector and value vector of node respectively. The query vector represents the query request of node for other nodes, reflecting the query demand of node for potential routing paths. The key vector represents the characteristics or status of node , which can include information such as the location, energy level, and communication range of the node. The key vector is used to match with the query vector to determine the attention degree of node to node . The value vector contains the detailed information of node , including the current load of the node, historical communication success rate, etc. After determining the reachability between nodes, this information will be used by node for decision-making.
[0099] Specifically, 、 and are calculated as follows:
[0100]
[0101] Among them, 、 and Represents a learnable weight matrix.
[0102] In formula (1), the attention function is calculated as follows:
[0103] where is the dimension of the key vector.
[0104] Considering that the reachability between nodes to is jointly determined by multiple dimensions, a multi-head attention mechanism is introduced for parallel multi-perspective calculation; under the multi-head attention mechanism, the reachability evaluation function between nodes to is expressed as: is represented as:
[0105]
[0106] where represents the learnable weight matrix of the output layer, The calculation method is as follows:
[0107] where , and represent the learnable weight matrix of the k-th attention head. The overall internal structure of the multi-head attention mechanism of the present invention is as Figure 2 shown.
[0108] (1.2) UAV node pruning mechanism
[0109] After evaluating the reachability between nodes, if nodes and are unreachable, then node needs to prune node . The present invention defines the node pruning function as , and its calculation process can be represented by the following formula:
[0110] ;
[0111] where is the threshold of reachability. When transmitting data in a dynamically changing airspace environment, the stability of the routing link must be ensured. When the value of the node reachability evaluation function is less than the threshold, it is considered that the nodes are unreachable.
[0112] The set of neighbor nodes after pruning is given by the formula:
[0113]
[0114] Thus, node Set of valid actions can be expressed as:
[0115]
[0116] Through the node pruning mechanism, the multi-UAV system can dynamically update the UAV nodes Set of valid actions. The set of valid actions refers to the set of valid actions for a node to forward data packets to surrounding nodes. Specifically, an action refers to the behavior of the current node to forward data packets to other nodes.
[0117] (1.3)Dynamic topology awareness of UAV nodes
[0118] Based on the above research, UAV nodes can dynamically obtain a local network view, perceive network topology changes in real time in the UAV environment, and screen reliable neighbor nodes. The specific process is as Figure 3 shown. First, the source node broadcasts detection data packets to the surrounding area, collects neighbor node status information, encodes the neighbor node status information into query (Q), key (K), and value (V) vectors. Then, attention feature extraction is performed, and the reachability weight between nodes is calculated through the multi-head attention mechanism. Finally, based on the node pruning mechanism, high-risk nodes are masked to generate a set of valid neighbors , and the local view is updated periodically in combination with environmental changes. The local view is the network topology structure maintained inside the node according to the set of valid neighbors after the node screens the set of valid neighbors, so as to achieve adaptive optimization of routing decisions in the dynamic airspace environment.
[0119] Multi-UAV routing collaborative decision-making algorithm based on action state space constraint MARL
[0120] (2.1)Multi-UAV routing collaborative decision-making algorithm
[0121] Traditional reinforcement learning algorithms are directly affected by the dimensions of the state space and action space, which will directly affect the complexity of the algorithm. For complex airspace environments, when the dimensions of the state space and action space increase, the number of state-action pairs to be explored will grow exponentially. This makes the algorithm take too much time to learn effective strategies and difficult to implement in practical applications. In addition, traditional reinforcement learning methods usually use a trial-and-error approach to learn and converge to the optimal strategy, which is particularly time-consuming in dynamic and complex airspace environments. The present invention proposes a collaborative routing decision algorithm based on the multi-head attention mechanism (Multi-Agent Routing Learning with Multi-Head Attention, MRL-MHA), which can capture key features in the learning environment faster by processing the attention calculations of different heads in parallel and improve learning efficiency. The architecture of the MAL-MHA algorithm is as Figure 4 shown.
[0122] The MRL-MHA algorithm proposed in the present invention is a multi-agent reinforcement learning algorithm based on multi-agent proximal policy optimization (MAPPO). MAPPO is extended and optimized for multi-agent environments on the basis of the Actor-Critic network architecture. Due to the use of the CTDE structure, all UAV agents share information during training and make independent decisions during execution. Under this structure, each UAV agent has its own policy network Actor network and value network Critic network, which are responsible for policy update and value evaluation. Different from traditional MAPPO, the MRL-MHA algorithm allows agents to share information during the training phase. The Actor network is responsible for interacting with the environment and learning new policies in a policy gradient learning manner under the guidance of the value function of the Critic network. The Critic network maintains the learning value function through the data collected by the Actor network interacting with the environment, is used to judge the quality of the current action, and further helps the Actor to update the policy.
[0123] In the MRL-MHA algorithm, the policy gradient update formula is as follows:
[0124]
[0125] where, represents the parameters used to update the policy network, represents the policy gradient, represents the expected value sampled from the experience replay according to the policy , represents the policy network, represents the action taken in the state , represents the estimated advantage function, is defined as follows:
[0126] ;
[0127] Among them, represents the action-value function, which represents the expected return when taking action in state and following . While represents the state-value function, which represents the expected return when following in state .
[0128] In the MRL-MHA algorithm, the update formula of the state-value function is as follows:
[0129] ;
[0130] Among them, are the adjustable parameters of the value network, represents the value function gradient, which is used to update the parameters of the value network, represents the state distribution under policy , represents the state distribution according to policy is the expected value obtained by sampling, is the target value function, and its formula is expressed as follows:
[0131] ;
[0132] Among them, is the immediate reward, is the discount factor, is the state transition probability.
[0133] Thus, MRL-MHA is a reinforcement learning algorithm based on MAPPO that allows all UAV nodes to share local network information during the central training process. Each UAV node maintains an Actor network to interact with the environment, and there is a Critic network with complete observational data of the local network during training for value evaluation to guide the Actor network to update its policy.
[0134] The MRL-MHA algorithm is a reinforcement learning framework for multi-agent collaborative tasks, which combines techniques such as experience replay mechanism, multi-head attention mechanism, and soft update of target network; constructs a loop of environment perception, action decision, value evaluation, and policy update to achieve efficient value evaluation and policy optimization. The training process of the MRL-MHA algorithm is as Figure 5 shown.
[0135] (2.1.1)Experience Replay Mechanism
[0136] Each agent selects an action according to the current policy network (Actor network) and obtains a reward after interacting with the environment and the next state . The MRL-MHA algorithm stores the global experience in the experience replay buffer, breaking the temporal correlation of the data; subsequently, through random sampling training, it improves the data utilization rate and stabilizes the learning process. The update process of the experience replay buffer is as follows:
[0137]
[0138] (2.1.2)Multi-Head Attention Mechanism
[0139] The MRL-MHA algorithm introduces the multi-head attention mechanism as described above. Each UAV agent maintains a decentralized Actor network and realizes action selection through the output maximization policy network. Formula (14) represents selecting the action that maximizes the output of the policy network in state s :
[0140]
[0141] When the agent makes an action decision through the policy network, it first dynamically evaluates the node reachability based on the multi-head attention mechanism, eliminates invalid action options through the node pruning mechanism, and combines network topology awareness to analyze the environmental state in real time. In the action decision stage, the system adopts the invalid action masking technology to ensure the effectiveness of the decision, and at the same time updates the probability distribution of the policy network to maintain the stable update of the policy network.
[0142] (2.1.3)Soft Update of Target Network
[0143] Both the Critic and the Actor use target networks to stabilize the training process. The target network is a copy of the Critic and Actorc networks, and its parameters are stably trained through soft updates. The parameters of the target network will gradually approach the main network but are not exactly the same, avoiding oscillations in Q-value estimation and policy updates. The update method is as follows:
[0144]
[0145] where is a tunable parameter much smaller than 1, used to control the update speed of the target network parameters
[0146] (2.2)Multi-UAV Routing Cooperative Decision-Making Strategy
[0147] In the Markov decision process model of multi-UAV routing collaborative decision-making studied in the present invention, the design of the reward function comprehensively considers factors such as whether the data is correctly transmitted towards the sink node (the end point of all data packet transmissions), the data transmission delay, and the link transmission quality. The reward function is defined as:
[0148]
[0149] wherein, are all adjustable parameters, is defined as follows:
[0150] (2.2.1) Forward propagation reward of data packet :
[0151] is used to indicate whether the data reaches the sink node, and at the same time, step penalty is carried out to limit the behavior of the nodes to send and receive data back and forth. The formula is as follows:
[0152]
[0153] wherein, is the positive reward for the successful arrival of the data, is the penalty based on the time step.
[0154] (2.2.2) Environment reward :
[0155] Environment reward considers the airspace environment factors (such as wind speed, environmental noise intensity), which will affect the routing performance of the nodes. Therefore, the routing quality can be optimized by encouraging the UAV nodes to adapt to the changes of the dynamic environment. The environment reward is defined as follows:
[0156]
[0157] wherein, represents the wind speed at node , represents the environmental noise intensity in the area where the node is located, while and respectively represent the weight factors controlling the influence of wind speed and environmental noise intensity on stability.
[0158] (2.2.3) Minimum hop count reward :
[0159] The minimum hop count of the routing path from the sink node directly guides the convergence direction of the reinforcement learning algorithm. The smaller the minimum hop count is, the higher the routing reliability is. It is defined as follows:
[0160]
[0161] Among them, is the shortest path length from the next-hop node of the decision to the aggregation node .
[0162] (2.2.4) Delay Reward :
[0163] Only consider the propagation delay. Since the distance between UAV nodes is relatively far, the propagation delay is usually large and needs to be standardized. The standardization process is as follows:
[0164]
[0165] Among them, and respectively represent the maximum and minimum propagation delays between nodes, The standardized propagation delay, with a value range of [0,1]. The delay reward is defined as follows:
[0166]
[0167] Among them, is an adjustable parameter, The larger it is, the greater the impact of the propagation delay on the reward function.
[0168] In the multi-UAV routing collaborative decision-making Markov decision process model, each UAV agent makes a routing decision based on local observations (such as the shortest hop count, delay information of neighbor nodes, etc.). Each UAV agent independently selects the next-hop node. Guided by the reward function, the UAV agents can collaboratively optimize the global routing performance.
[0169] The multi-UAV routing collaborative decision-making Markov decision process model satisfies the Markov property, that is, the current state and action completely determine the next state and reward , which is expressed by the formula as follows:
[0170]
[0171] Among them, represents the joint probability distribution of the next state and reward given the current state and action , Denotes the next state and action under the condition of the given current state and all previous states and actions, and the joint probability distribution of the reward.
[0172] Example:
[0173] According to the above algorithm design, the pseudo-code of the core algorithm of MRL-MHA is shown in Table 1:
[0174] Table 1 Pseudo-code table of the MRL-MHA algorithm proposed by the present invention
[0175]
[0176] Taking a multi-UAV system with 20 nodes as an example, as Figure 6 shown, first, the multi-UAV nodes are constructed into a multi-UAV system with hierarchical control. The UAV in the control layer distributes the routing task structure to the UAV in the execution layer, and the UAV nodes in the execution layer perform data routing. The specific training method of the routing strategy is as shown in the algorithm in Table 1 finally. Figure 6 The process of data routing and forwarding under 20 UAV nodes shown in step 3 in it is that the 0th node sends a data packet to the aggregation node, and the intermediate nodes are responsible for selecting the optimal neighbor node for forwarding, and finally the data routing process is completed. In summary, using MRL-MHA, based on the preset multi-UAV routing cooperation decision strategy, autonomous data routing between multi-UAVs in the constrained action state space can be realized.
[0177] In specific implementation, the present application provides a computer storage medium and a corresponding data processing unit. Among them, the computer storage medium can store a computer program, and when the computer program is executed by the data processing unit, it can run the invention content of a collaborative autonomous data routing method for a multi-UAV system provided by the present invention and some or all of the steps in each embodiment. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0178] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of a computer program and its corresponding general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a computer program, that is, a software product. The computer program software product can be stored in a storage medium, including several instructions for causing a device (which can be a personal computer, a server, a single-chip microcomputer, an MCU or a network device, etc.) containing a data processing unit to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0179] The present invention provides an idea and method for a collaborative autonomous data routing method for a multi-UAV system. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and retouches can be made, and these improvements and retouches should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by the prior art.
Claims
1. A collaborative autonomous data routing method for a multi-UAV system, characterized in that It includes the following steps: Step 1, in the multi-UAV system, based on the multi-head attention pruning mechanism, perform action space compression to obtain the effective neighbor set and effective action set of each UAV; Step 2, deploy a policy network in each UAV, deploy a value network in the control center, and based on the effective neighbor set and effective action set obtained in Step 1, use the policy network and value network to make independent decisions for individual UAVs and collaborative decisions for the UAV cluster, and obtain the preliminary decisions of each UAV; Step 3, construct a multi-UAV routing collaborative decision-making Markov decision process model to optimize the preliminary decisions of all UAVs and complete the collaborative autonomous data routing for the multi-UAV system; Among them, obtaining the effective neighbor set and effective action set of each UAV in Step 1 includes the following steps: Step 1-1, in the multi-UAV system, assume that the source node i.e., the UAV broadcasts detection data packets to the surroundings to collect the status information of neighbor nodes i.e., the UAV ; Step 1-2, encode the state information of neighbor nodes into query vectors, key vectors, and value vectors; Step 1-3, perform attention feature extraction based on the query vectors, key vectors, and value vectors, and perform node reachability evaluation through the multi-head attention mechanism; Step 1-4, shield high-risk nodes based on the node pruning mechanism, and generate an effective neighbor set and effective action set according to the node reachability evaluation results; Step 1-5, update the local view according to the effective neighbor set and effective action set, where the local view is the network topology structure maintained inside the source node after screening the effective neighbor set; Encoding the state information of neighbor nodes into query vectors, key vectors, and value vectors in Step 1-2 includes the following steps: ; Among them, , and are respectively used to represent the query vector of the source node , the key vector and value vector of the neighbor node ; the query vector represents the query request of node for other nodes, reflecting the query demand of node for potential routing paths; the key vector represents the feature or state of node , which is used to match with the query vector to determine the attention degree of node to node ; the value vector contains the detailed information of node , which is used for the decision-making of node ; , and represent learnable weight matrices; Performing node reachability evaluation in Step 1-3 includes the following steps: Let the node to node The reachability evaluation function between them is , which is expressed as follows: ; Among them is the attention function, and the calculation method is as follows: ; Among them, is the softmax function, is the dimension of the key vector; According to the nodes to the nodes Among the k different dimensions that determine reachability, a multi-head attention mechanism is introduced for multi-perspective parallel computing. Then, the reachability evaluation function from the node to the node is expressed as: ; Among them, represents the learnable weight matrix of the output layer, represents the operation of concatenating multiple vectors, represents the k-th attention head; The k-th attention head , and the calculation method is as follows: ; Among them, , and represent the learnable weight matrix of the k-th attention head; Generating an effective neighbor set and effective action set according to the node reachability evaluation results in Step 1-4 includes the following steps: According to the node to the node between the reachability evaluation function , perform reachability judgment between nodes, which is expressed as follows: ; Among them, is the node pruning function, is the threshold of reachability. If the and the node the value of the reachability evaluation function between them is less than the threshold, the value of the node pruning function is 0, and it is judged as unreachable, indicating that the node needs to prune the node ; The set of neighbor nodes after pruning, i.e., the valid neighbor set, is , which is shown as follows: ; Node The set of valid actions for is as follows: 。 2. The collaborative autonomous data routing method for a multi-UAV system according to claim 1, wherein Using the policy network and value network to make independent decisions for individual UAVs and collaborative decisions for the UAV cluster in Step 2 includes the following steps: Step 2-1, set each drone Select an action according to the current policy network as , and obtain a reward of after interacting with the environment. Its next state is ; Store the global experience in the experience replay buffer. The update process of the experience replay buffer is as follows: ; Step 2-2, set target networks corresponding to the policy network and value network, perform training and update to obtain the trained policy network and value network; Among them, the policy network update method is expressed as follows: ; Among them, represents the parameters of the policy network, represents the policy gradient, which is used to update the parameters of the policy network , represents the expected value sampled according to the policy from the experience replay buffer and represents the action taken in the state , represents the estimated advantage function, which is expressed as follows: ; Among them, is the action value function, representing the expected return when taking action in state and following policy ; is the state value function, representing the expected return when following policy in state ; State value function The update of, that is, the value network update method is expressed as follows: ; Among them, are the parameters of the value network, represents the value function gradient, which is used to update the parameters of the value network, represents the state distribution under the policy , represents the state distribution obtained by sampling according to the policy , and is the expected value obtained by sampling, is the value function of the target network, which is expressed as follows: ; Among them, is the immediate reward, is the discount factor, is the state transition probability; The parameter update method of the target network is as follows: ; Among them, is an adjustable parameter used to control the speed of target network parameter update; Step 2-3: The trained policy network selects the maximum action based on the set of valid neighbors and the set of valid actions, which is used as the preliminary decision, expressed as follows: , as the preliminary decision, is expressed as follows: ; Among them, the trained policy network is , indicating that in state , the selected action is , are the parameters of the policy network.
3. A collaborative autonomous data routing method for a multi-UAV system according to claim 2, characterized in that, Optimizing the preliminary decisions of all UAVs in Step 3 includes: Each UAV makes a routing decision based on local observations. Through the reward function, each UAV independently selects the next-hop node, which is expressed as follows: ; Among them, represents the joint probability distribution of the next state and the reward under the condition of the given current state and the action . represents the joint probability distribution of the next state and the reward under the condition of the given current state and the action , as well as all previous states and actions.
4. The collaborative autonomous data routing method for a multi-UAV system according to claim 3, wherein The described reward , through each drone 's reward function is calculated, and the reward function is expressed as follows: ; Among them, , and are adjustable parameters, is the forward propagation reward of the data packet, is the environmental reward, is the minimum hop count reward, is the delay reward.
5. A collaborative autonomous data routing method for a multi-UAV system according to claim 4, characterized in that Forward propagation reward of data packet , which is used to indicate whether the data has reached the sink node, i.e., the end point of data packet transmission, and at the same time, step penalty is carried out to limit the behavior of nodes to transmit data back and forth. The calculation method is as follows: ; Among them, is the positive reward for successful data arrival, is the penalty based on time steps; Environmental reward , and the calculation method is as follows: ; Among them, represents the wind speed magnitude at the node , represents the environmental noise intensity within the area where the node is located, and respectively represent the weight factors that control the influence of wind speed and environmental noise intensity on stability; Minimum hop count reward , and the calculation method is as follows: ; Among them, is the next-hop node from the decision to the aggregation node of the shortest path length; Time-delay reward , and the calculation method is as follows: ; Among them, is an adjustable parameter, is the normalized propagation delay, and its value range is [0, 1]. The calculation method is as follows: ; Among them, and respectively represent the maximum and minimum propagation delays between nodes.
Citation Information
Patent Citations
Networking method and system for software-defined wireless ad hoc network controlled by unmanned aerial vehicles
CN110167032A
Unmanned aerial vehicle cluster network intelligent multi-hop routing method based on multi-agent cooperation
CN114499648A