Cooperative autonomous data routing method for multi-unmanned aerial vehicle system
By adopting multi-head attention pruning mechanism and multi-agent reinforcement learning algorithm in multi-unmanned aerial systems, data routing decisions between drones are optimized, and the problem of inefficient data routing in complex communication environments is solved, and efficient and stable data transmission is achieved.
Patent Information
- Application Number
- CN202510670147.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-23
AI Technical Summary
Multi-UAV systems face problems such as network topology changes, electromagnetic interference and data transmission delay in complex communication environments, resulting in inefficient data routing and affecting coordinated task execution.
The multi-head attention pruning mechanism and multi-agent reinforcement learning algorithm based on action state space constraints are adopted to optimize data routing decisions between drones through action space compression, experience playback mechanism, multi-head attention mechanism and target network soft update technology.
It significantly improves the speed and learning efficiency of routing decisions, enhances the stability of the algorithm, enables routing strategies to quickly adapt to environmental changes, and improves the data transmission stability and overall performance of multi-UAV systems.
Smart Images

Figure CN120200956A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a data routing method, in particular to a collaborative autonomous data routing method for a multi-UAV system. Background Art
[0002] The information provided in this part is only background information related to the present disclosure, and it is not necessarily prior art.
[0003] With the rapid development of the information age, multi-UAV (Unmanned Aerial Vehicle) systems are increasingly widely used in fields such as disaster emergency and environmental monitoring. The collaborative operation of multi-UAVs can significantly improve the task execution efficiency and effect due to its advantages of flexibility, high efficiency, wide coverage, and strong adaptability. However, in the actual operation process of multi-UAV systems, they face challenges brought by complex and changeable communication environments and system self-limitations. On the one hand, during the flight of UAVs, affected by factors such as airflows and terrain, the network topology structure changes frequently; at the same time, the communication between UAVs and with the ground control station is vulnerable to problems such as electromagnetic interference and signal occlusion, seriously affecting the stability and reliability of data transmission. On the other hand, when multi-UAV systems execute tasks, they often need to transmit a large amount of real-time data, such as high-definition images, video streams, sensor monitoring data, etc. Traditional data routing methods are difficult to meet the requirements of low latency and high throughput when dealing with high-concurrency, multi-source heterogeneous data, and are prone to problems such as data congestion and low routing efficiency, thereby affecting the collaborative execution of the overall tasks of multi-UAV systems.
[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] Object of the Invention: The technical problem to be solved by the present invention is to provide a collaborative autonomous data routing method for a multi-UAV system in view of the deficiencies of the prior art.
[0006] To solve the above technical problem, the present invention discloses a collaborative autonomous data routing method for a multi-UAV system, including the following steps: Step 1, in the multi-UAV system, based on the multi-head attention pruning mechanism, perform action space compression to obtain the effective neighbor set and effective action set of each UAV; Step 2, deploy a policy network in each UAV, deploy a value network in the control center, and based on the effective neighbor set and effective action set obtained in Step 1, use the policy network and value network to perform independent decision-making for a single UAV and collaborative decision-making for the UAV cluster, and obtain the preliminary decision of each UAV; Step 3: Construct a Markov decision process model for multi-UAV routing collaborative decision-making, optimize the initial decisions of all UAVs, and complete collaborative autonomous data routing for multi-UAV systems.
[0007] Further, obtaining the valid neighbor set and valid action set of each drone in step 1 includes the following steps: Step 1-1: Assume that in the multi-UAV system, the source node UAV Broadcast detection packets to the surroundings to collect neighbor nodes UAV Status information; Step 1-2, encode the state information of the neighbor node into a query vector, a key vector and a value vector; Steps 1-3: extract attention features based on the query vector, key vector, and value vector, and evaluate the reachability between nodes through a multi-head attention mechanism; Steps 1-4: shield high-risk nodes based on the node pruning mechanism, and generate a valid neighbor set and a valid action set based on the inter-node reachability evaluation results; Step 1-5, updating the local view according to the valid neighbor set and the valid action set, wherein the local view is the network topology structure maintained inside the source node after the source node screens the valid neighbor set.
[0008] Furthermore, encoding the state information of the neighbor nodes into a query vector, a key vector and a value vector in step 1-2 includes the following steps: ; in, , and Used to represent the source node The query vector and neighbor nodes The key vector and value vector of Representative Node Query requests to other nodes, reflecting the node Query requirements for potential routing paths; key vector Representative Node The features or states of the node are used to match the query vector to determine the node For Node The attention of; value vector Contains nodes Details for nodes decision-making; , and represents the learnable weight matrix.
[0009] Further, the reachability evaluation between nodes described in steps 1-3 includes the following steps: Let the reachability evaluation function between node and node be , which is expressed as follows: ; where is the attention function, and its calculation method is as follows: ; where is the softmax function, is the dimension of the key vector; According to the k different dimensions that determine the reachability between node and node , the multi-head attention mechanism is introduced for multi-perspective parallel calculation. Then, the reachability evaluation function between node and node is expressed as: ; where represents the learnable weight matrix of the output layer, represents the operation of concatenating multiple vectors, represents the k-th attention head.
[0010] Further, the k-th attention head , and its calculation method is as follows: ; where , and represent the learnable weight matrix of the k-th attention head.
[0011] Further, generating the effective neighbor set and the effective action set according to the node reachability evaluation result described in steps 1-4 includes the following steps: According to the reachability evaluation function between node and node , the reachability judgment between nodes is performed, which is expressed as follows: ; where is the node pruning function, is the threshold of reachability. If the reachability evaluation function between node and node When the value is less than the threshold, the node pruning function value is 0, and it is judged as unreachable, indicating the node needs to prune the node ; The set of neighbor nodes after pruning, i.e., the effective neighbor set, is , which is shown as follows: ; The node has an effective action set of , which is shown as follows: .
[0012] Furthermore, the independent decision-making of a single UAV and the collaborative decision-making of a UAV swarm using the policy network and the value network in step 2 include the following steps: Step 2-1, assume that each UAV selects an action according to the current policy network as , obtains a reward of after interacting with the environment, and its next state is ; Store the global experience in the experience replay buffer, and the experience replay buffer is updated as follows: ; Step 2-2, set the target networks corresponding to the policy network and the value network, train and update them to obtain the trained policy network and value network; Among them, the policy network update method is shown as follows: ; Among them, represents the parameters of the policy network, represents the policy gradient, which is used to update the parameters of the policy network , represents the expected value sampled from the experience replay buffer according to the policy , represents the action taken in the state , represents the estimated advantage function, which is shown as follows: ; Among them, is the action value function, which represents the expected return of taking the action in the state and following the policy , is the state value function, which represents following the policy in the state Expected return; State value function The update of, that is, the value network update method is expressed as follows: ; Wherein, are the parameters of the value network, represents the value function gradient, which is used to update the parameters of the value network, represents the state distribution under the policy , represents the state distribution according to the policy under The expected value obtained by sampling, is the value function of the target network, which is expressed as follows: ; Wherein, is the immediate reward, is the discount factor, is the state transition probability; The parameter update method of the target network is as follows: ; Wherein, is an adjustable parameter, which is used to control the update speed of the target network parameters; Step 2-3, the trained policy network selects the maximum action according to the effective neighbor set and the effective action set , as the preliminary decision, which is expressed as follows: ; Wherein, the trained policy network is , which means that the action selected under the state is , are the parameters of the policy network.
[0013] Further, the optimization of the preliminary decisions for all drones described in step 3 includes: Each drone makes a routing decision based on local observations. Through the reward function, each drone independently selects the next-hop node, which is expressed as follows: ; Wherein, represents the joint probability distribution of the next state and the reward given the current state and the action , represents the current state and the action and the conditional joint probability distribution of the next state, given all previous states and actions and reward .
[0014] Furthermore, the reward is calculated by the reward function of each drone , and the reward function is expressed as follows: ; ; where , and are adjustable parameters, is the forward propagation reward of the data packet, is the environmental reward, is the minimum hop count reward, is the delay reward
[0015] Furthermore, the forward propagation reward of the data packet , which is used to indicate whether the data has reached the sink node, i.e., the destination of the data packet transmission, and at the same time imposes a step penalty to limit the behavior of the nodes to transmit data back and forth, is calculated as follows: ; where is the positive reward for successful data arrival, is the penalty based on the time step; The environmental reward is calculated as follows: ; where represents the wind speed magnitude at node , represents the environmental noise intensity within the area where the node is located, and respectively represent the weight factors that control the influence of wind speed and environmental noise intensity on stability; The minimum hop count reward is calculated as follows: ; where is the shortest path length from the next-hop node of the decision to the sink node ; The delay reward is calculated as follows: ; where is an adjustable parameter, is the normalized propagation delay, with a value range of [0, 1], and the calculation method is as follows: ; wherein, and respectively represent the maximum and minimum propagation delays between nodes.
[0016] Beneficial effects: 1. The present invention fully considers the problems faced by multi-UAV systems, such as high-dimensional action space and low efficiency of traditional learning algorithms. Through the action space compression mechanism, unnecessary action exploration is effectively reduced, and the speed of routing decision-making is greatly improved. Combining the experience replay mechanism, multi-head attention mechanism, and target network soft update technology, the learning stability and efficiency of the algorithm are significantly enhanced, enabling the routing strategy to quickly adapt to environmental changes.
[0017] 2. In terms of the routing collaborative decision-making strategy of the present invention, the carefully designed reward function comprehensively considers various factors such as positive propagation of data packets, environmental factors, minimum hop count, and delay, guiding the UAV agent to make better routing decisions. It provides reliable data transmission support for the collaborative operation of multi-UAV systems in complex environments, enhances the execution ability of the system in various tasks (such as complex environment monitoring, emergency communication guarantee, collaborative combat, etc.), and significantly improves the overall performance and application value of multi-UAV systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The following further describes the present invention in detail with reference to the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0019] Figure 1 is an architecture diagram of a multi-UAV system based on hierarchical control according to the present invention.
[0020] Figure 2 is a schematic diagram of a multi-head attention mechanism according to the present invention.
[0021] Figure 3 is a schematic diagram of the process of obtaining a dynamic local UAV cluster view according to the present invention.
[0022] Figure 4 is an architecture diagram of a MAL-MHA algorithm according to the present invention.
[0023] Figure 5 is a training flowchart of a MAL-MHA algorithm according to the present invention.
[0024] Figure 6 is a flowchart of UAV routing implementation according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The general idea of the present invention is as follows: In view of the many challenges existing in data transmission in a multi-UAV system in a complex and changeable network environment, such as the instability of the network topology caused by frequent node movement, the vulnerability of communication links to interference, etc., an innovative data routing optimization system is constructed based on the unique attributes of UAVs. On this basis, by introducing a multi-head attention pruning mechanism and a multi-agent reinforcement learning algorithm based on action state space constraints, a highly intelligent collaborative autonomous data routing scheme is proposed. The present invention fully considers the problems faced by multi-UAV systems, such as high dimensionality of the action space and low efficiency of traditional learning algorithms. Through the action space compression mechanism, unnecessary action exploration is effectively reduced, and the speed of routing decision-making is greatly improved. Combining the experience replay mechanism, the multi-head attention mechanism and the target network soft update technology, the learning stability and efficiency of the algorithm are significantly enhanced, enabling the routing strategy to quickly adapt to environmental changes. In terms of the routing collaborative decision-making strategy, the carefully designed reward function comprehensively considers various factors such as the forward propagation of data packets, environmental factors, minimum hop count and delay, guiding the UAV agents to make better routing decisions. The purpose of the present invention is to provide reliable data transmission support for the collaborative operation of multi-UAV systems in complex environments, enhance the execution ability of the system in various tasks (such as complex environment monitoring, emergency communication guarantee, collaborative combat, etc.), and significantly improve the overall performance and application value of multi-UAV systems.
[0026] The present invention proposes a collaborative autonomous data routing scheme for multi-UAV systems. Based on the multi-head attention pruning mechanism, an action space compression mechanism for a multi-agent reinforcement learning (Multi-Agent-Reinforcement-Learning, MARL) algorithm is constructed, and on this basis, a multi-UAV routing collaborative decision algorithm based on MARL with action state space constraints is constructed for data transmission and sharing between multi-UAVs.
[0027] such as Figure 1As shown in the figure, the overall technical solution of the present invention is as follows: For a multi-UAV system, a hierarchical control architecture for the multi-UAV system is designed, and all UAVs are divided into a control center UAV-R and ordinary UAV nodes UAV-L. Among them, UAV-L is responsible for managing the entire cluster network, and UAV-R is responsible for executing collaborative navigation tasks such as intelligence reconnaissance, data relay, and data collection. Among UAV-Ls, through self-organizing control, according to the actual business objectives, data routing transmission is carried out to achieve intelligent data sharing and distributed storage. The present invention constructs a collaborative autonomous data routing scheme for a multi-UAV system based on multi-agent reinforcement learning, constructs an action space compression mechanism for the MARL algorithm based on the multi-head attention pruning mechanism, and on this basis, constructs a multi-UAV routing collaborative decision-making algorithm based on action-state space constraint MARL for data transmission and sharing among multiple UAVs. The specific steps are as follows: (1) Action space compression mechanism based on multi-head attention pruning mechanism (1.1) Node reachability evaluation based on multi-head attention mechanism Based on the idea of the multi-head attention mechanism in Transformer, in the multi-UAV network routing decision problem, the local network structure is explored through the multi-head attention mechanism, and the UAV node state and the view of neighbor nodes that the node can reach are continuously updated. Different from the existing research on the fusion of attention mechanism and MARL, the present invention is aimed at the action space of UAV nodes, not just the communication between UAVs.
[0028] For the multi-UAV network architecture, design nodes to node The reachability evaluation function between is , and the formula of the reachability evaluation function based on the attention mechanism is as follows:
[0029] Among them, , and are used to represent the query vector, key vector and value vector of node respectively. The query vector represents the query request of node for other nodes, reflecting the query demand of node for potential routing paths. The key vector represents the characteristics or state of node , which can include information such as the location, energy level, and communication range of the node. The key vector is used to match with the query vector to determine the attention degree of node to node . The value vector Contains node details , including the current load of the node, historical communication success rate, etc. These details will be used by the node for decision-making.
[0030] Specifically, , and are calculated as follows:
[0031] Among them, , and represent learnable weight matrices.
[0032] The calculation method of the attention function in formula (1) is as follows: Among them, is the dimension of the key vector.
[0033] Considering that the reachability between nodes to is jointly determined by multiple dimensions, a multi-head attention mechanism is introduced for multi-perspective parallel computing; under the multi-head attention mechanism, the reachability evaluation function between nodes to is expressed as:
[0034] Among them, represents the learnable weight matrix of the output layer, is calculated as follows: Among them, , and represent the learnable weight matrices of the k-th attention head. The overall internal structure of the multi-head attention mechanism of the present invention is as Figure 2 shown.
[0035] (1.2) UAV node pruning mechanism After evaluating the reachability between nodes, if nodes and nodes are unreachable, then nodes need to prune nodes . The present invention defines the node pruning function as , and its calculation process can be expressed by the following formula: ; Among them, is the threshold of reachability. When transmitting data in a dynamically changing airspace environment, the stability of the routing link must be ensured. When the value of the node reachability evaluation function is less than the threshold, the node is considered unreachable.
[0036] The set of neighbor nodes after pruning The formula is as follows:
[0037] Thus, the set of effective actions of node can be expressed as:
[0038] Through the node pruning mechanism, the multi-UAV system can dynamically update the set of effective actions of UAV nodes The set of effective actions refers to the set of effective actions for a node to forward data packets to surrounding nodes. Specifically, an action refers to the behavior of the current node to forward data packets to other nodes.
[0039] (1.3)UAV Node Dynamic Topology Awareness Based on the above research, UAV nodes can dynamically obtain a local network view, and in the UAV environment, they can perceive network topology changes in real time and screen reliable neighbor nodes. The specific process is as Figure 3 shown. First, the source node broadcasts detection data packets to the surrounding area, collects neighbor node status information, encodes the neighbor node status information into query (Q), key (K), and value (V) vectors, and then performs attention feature extraction and calculates the reachability weights between nodes through the multi-head attention mechanism. Finally, based on the node pruning mechanism, high-risk nodes are masked to generate an effective neighbor set , and the local view is updated periodically in combination with environmental changes. The local view is the network topology structure maintained inside the node according to the effective neighbor set after the node screens the effective neighbor set, so as to achieve the adaptive optimization of routing decisions in the dynamic airspace environment.
[0040] Multi-UAV Routing Cooperative Decision Algorithm Based on Action State Space Constrained MARL (2.1)Multi-UAV Routing Cooperative Decision Algorithm Traditional reinforcement learning algorithms are directly affected by the dimensions of the state space and action space, which will directly affect the complexity of the algorithm. For complex airspace environments, when the dimensions of the state space and action space increase, the number of state-action pairs to be explored will increase exponentially. This makes the algorithm take too much time to learn effective strategies and is difficult to implement in practical applications. In addition, traditional reinforcement learning methods usually use a trial-and-error approach to learn and converge to the optimal strategy, which is particularly time-consuming in dynamic and complex airspace environments. The present invention proposes a collaborative routing decision algorithm based on the multi-head attention mechanism (Multi-Agent Routing Learning with Multi-Head Attention, MRL-MHA), which can capture key features in the learning environment faster and improve learning efficiency by parallel processing of attention calculations for different heads. The architecture of the MAL-MHA algorithm is as Figure 4 shown.
[0041] The MRL-MHA algorithm proposed by the present invention is a multi-agent reinforcement learning algorithm based on multi-agent proximal policy optimization (MAPPO). MAPPO is extended and optimized for multi-agent environments on the basis of the Actor-Critic network architecture. Due to the use of the CTDE structure, all UAV agents share information during training and make independent decisions during execution. In this structure, each UAV agent has its own policy network Actor network and value network Critic network, which are responsible for policy update and value evaluation. Different from traditional MAPPO, the MRL-MHA algorithm allows agents to share information during the training phase. The Actor network is responsible for interacting with the environment and learning new policies in the way of policy gradient learning under the guidance of the value function of the Critic network. The Critic network maintains the learning value function through the data collected by the Actor network interacting with the environment, which is used to judge the pros and cons of the current action and further help the Actor to update the policy.
[0042] In the MRL-MHA algorithm, the policy gradient update formula is as follows:
[0043] where, represents the parameters used to update the policy network, represents the policy gradient, represents the expected value sampled from the experience replay according to the policy , represents the policy network, represents the action taken in the state , represents the estimated advantage function, which is defined as follows: ; Among them, represents the action value function, which represents the expected return when taking action in state and following . While represents the state value function, which represents the expected return when following in state .
[0044] In the MRL-MHA algorithm, the update formula of the state value function is as follows: ; Among them, is the adjustable parameter of the value network, represents the value function gradient, which is used to update the parameters of the value network, represents the state distribution under policy , represents the state distribution sampled according to policy , is the expected value obtained by sampling, is the target value function, and its formula is expressed as follows: ; Among them, is the immediate reward, is the discount factor, is the state transition probability.
[0045] Therefore, MRL-MHA is a reinforcement learning algorithm based on MAPPO that allows all UAV nodes to share local network information during the central training process. Each UAV node maintains an Actor network to interact with the environment, and there is a Critic network with complete observation data of the local network during training for value evaluation to guide the Actor network to update the policy.
[0046] The MRL-MHA algorithm is a reinforcement learning framework for multi-agent collaborative tasks, which combines techniques such as experience replay mechanism, multi-head attention mechanism, and target network soft update; constructs a loop of environment perception, action decision-making, value evaluation, and policy update to achieve efficient value evaluation and policy optimization. The training process of the MRL-MHA algorithm is as Figure 5 shown.
[0047] (2.1.1) Experience Replay Mechanism Each agent selects an action according to the current policy network (Actor network), obtains a reward after interacting with the environment and the next state . The MRL-MHA algorithm stores the global experience in the experience replay buffer, breaking the temporal correlation of the data; subsequently, through random sampling training, it improves the data utilization rate and stabilizes the learning process. The experience replay buffer is updated as follows:
[0048] (2.1.2)Multi-Head Attention Mechanism The MRL-MHA algorithm introduces the multi-head attention mechanism as described above. Each UAV agent maintains a decentralized Actor network , and realizes action selection through the output maximization policy network. Equation (14) represents selecting the action that maximizes the output of the policy network in state s :
[0049] When the agent makes action decisions through the policy network, it first dynamically evaluates the node reachability based on the multi-head attention mechanism, eliminates invalid action options through the node pruning mechanism, and combines network topology awareness to analyze the environmental state in real time. In the action decision-making stage, the system adopts an invalid action masking technique to ensure the effectiveness of the decision-making, and at the same time updates the probability distribution of the policy network to maintain the stable update of the policy network.
[0050] (2.1.3)Soft Update of the Target Network Both the Critic and the Actor use the target network to stabilize the training process. The target network is a copy of the Critic and Actorc networks, and its parameters are stably trained through soft updates. The parameters of the target network will gradually approach the main network but are not exactly the same, avoiding oscillations in Q-value estimation and policy updates. The update method is as follows:
[0051] where is a tunable parameter much smaller than 1, used to control the speed of target network parameter update.
[0052] (2.2)Multi-UAV Routing Cooperative Decision-Making Strategy In the multi-UAV routing cooperative decision-making Markov decision process model studied in the present invention, the design of the reward function comprehensively considers factors such as whether the data is correctly transmitted towards the convergence node (the end point of all data packet transmissions), data transmission delay, and link transmission quality. The reward function is defined as:
[0053] where They are all adjustable parameters, and their definitions are as follows: (2.2.1) Forward Propagation Reward of Data Packet : It is used to indicate whether the data has reached the sink node, and at the same time, step penalty is imposed to limit the behavior of nodes to send data back and forth between nodes. The formula is as follows:
[0054] Among them, is the positive reward for the successful arrival of data, is the penalty based on the time step.
[0055] (2.2.2) Environment Reward : The environment reward considers airspace environmental factors (such as wind speed, environmental noise intensity), which will affect the routing performance of nodes. Therefore, the routing quality can be optimized by encouraging UAV nodes to adapt to the changes in the dynamic environment. The environment reward is defined as follows:
[0056] Among them, represents the wind speed magnitude at node , represents the environmental noise intensity in the area where the node is located, while and represent the weight factors that control the influence of wind speed and environmental noise intensity on stability, respectively.
[0057] (2.2.3) Minimum Hop Count Reward : The minimum hop count of the routing path to the sink node directly guides the convergence direction of the reinforcement learning algorithm. The smaller the minimum hop count, the higher the routing reliability. The definition is as follows:
[0058] Among them, is the shortest path length from the next-hop node of the decision to the sink node .
[0059] (2.2.4) Delay Reward : Only the propagation delay is considered. Because the distance between UAV nodes is relatively far, the propagation delay is usually large and needs to be standardized. The standardization process is as follows:
[0060] Among them, and respectively represent the maximum and minimum propagation delays between nodes, The normalized propagation delay, whose value range is [0,1]. The delay reward is defined as follows:
[0061] Among them, is an adjustable parameter, The larger it is, the greater the impact of the propagation delay on the reward function.
[0062] In the multi-UAV routing collaborative decision-making Markov decision process model, each UAV agent makes a routing decision based on local observations (such as the shortest hop count and delay information of neighbor nodes). Each UAV agent independently selects the next-hop node. Guided by the reward function, the UAV agents can collaboratively optimize the global routing performance.
[0063] The multi-UAV routing collaborative decision-making Markov decision process model satisfies the Markov property, that is, the current state and the action completely determine the next state and the reward , which is expressed by the formula as follows:
[0064] Among them, represents the joint probability distribution of the next state and the reward under the condition of the given current state and the action , represents the joint probability distribution of the next state and the reward as well as all previous states and actions under the condition of the given current state and the reward .
[0065] Embodiment: According to the above algorithm design, the pseudo-code of the core algorithm of MRL-MHA is shown in Table 1: Table 1 Pseudo-code table of the MRL-MHA algorithm proposed by the present invention
[0066] Taking a multi-UAV system with 20 nodes as an example, as Figure 6As shown, first, multiple UAV nodes are constructed into a multi-UAV system with hierarchical control. The UAV in the control layer distributes the routing task structure to the UAVs in the execution layer, and the UAV nodes in the execution layer perform data routing. The specific training method of the routing strategy is shown in the algorithm of Table 1 finally. Figure 6 The data routing and forwarding process under 20 UAV nodes shown in step 3 in Figure 6 is that the 0th node sends data packets to the aggregation node, and the intermediate nodes are responsible for selecting the optimal neighbor nodes for forwarding, finally completing the data routing process. In summary, using MRL-MHA and based on the preset multi-UAV routing collaborative decision-making strategy, autonomous data routing between multiple UAVs in the constrained action state space can be achieved.
[0067] In specific implementation, the present application provides a computer storage medium and a corresponding data processing unit. Among them, the computer storage medium can store a computer program, and when the computer program is executed by the data processing unit, it can run the content of the invention of a collaborative autonomous data routing method for a multi-UAV system provided by the present invention and some or all of the steps in each embodiment. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0068] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of a computer program and its corresponding general hardware platform. Based on such an understanding, the essence of the technical solutions in the embodiments of the present invention, or the part that contributes to the prior art, can be embodied in the form of a computer program, that is, a software product. The computer program software product can be stored in the storage medium, including several instructions for causing a device (which can be a personal computer, a server, a single-chip microcomputer, an MCU, or a network device, etc.) including a data processing unit to execute the methods described in each embodiment or some parts of the embodiments of the present invention.
[0069] The present invention provides the idea and method of a collaborative autonomous data routing method for a multi-UAV system. There are many methods and ways to specifically implement this technical solution. The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and retouches can be made, and these improvements and retouches should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by the prior art.
Claims
1. A collaborative autonomous data routing method for a multi-UAV system, characterized in that, It includes the following steps: Step 1, in the multi-UAV system, based on the multi-head attention pruning mechanism, perform action space compression to obtain the effective neighbor set and effective action set of each UAV; Step 2, deploy a policy network in each UAV, deploy a value network in the control center, and based on the effective neighbor set and effective action set obtained in Step 1, use the policy network and value network to make independent decisions for individual UAVs and collaborative decisions for the UAV cluster to obtain the preliminary decisions of each UAV; Step 3, construct a multi-UAV routing collaborative decision-making Markov decision process model to optimize the preliminary decisions of all UAVs and complete the collaborative autonomous data routing for the multi-UAV system.
2. The collaborative autonomous data routing method for a multi-UAV system according to claim 1, wherein The obtaining of the effective neighbor set and effective action set of each UAV in Step 1 includes the following steps: Step 1-1, in the multi-UAV system, assume that the source node i.e., the UAV broadcasts detection data packets to the surrounding area to collect the status information of neighbor nodes i.e., the UAV ; Step 1-2, encode the state information of neighbor nodes into query vectors, key vectors, and value vectors; Step 1-3, perform attention feature extraction based on the query vectors, key vectors, and value vectors, and through the multi-head attention mechanism, perform node reachability evaluation; Step 1-4, shield high-risk nodes based on the node pruning mechanism, and generate an effective neighbor set and effective action set according to the node reachability evaluation result; Step 1-5, update the local view according to the effective neighbor set and effective action set, where the local view is the network topology structure maintained inside the source node after screening the effective neighbor set.
3. A collaborative autonomous data routing method for a multi-UAV system according to claim 2, characterized in that, The encoding of the state information of neighbor nodes into query vectors, key vectors, and value vectors in Step 1-2 includes the following steps: ; Among them, , and are respectively used to represent the query vector of the source node , the key vector and the value vector of the neighbor node ; the query vector represents the query request of the node for other nodes, reflecting the query demand of the node for potential routing paths; the key vector represents the characteristics or status of the node , which is used to match with the query vector to determine the attention degree of the node to the node ; the value vector contains the detailed information of the node , which is used for the decision-making of the node ; , and represent learnable weight matrices.
4. A collaborative autonomous data routing method for a multi-UAV system according to claim 3, characterized in that, The performing of node reachability evaluation in Step 1-3 includes the following steps: Let the node to node The reachability evaluation function between is , which is expressed as follows: ; wherein is an attention function, and the calculation method is as follows: ; Among them, is the softmax function, is the dimension of the key vector; According to the nodes to the nodes Among the k different dimensions that determine reachability, the multi-head attention mechanism is introduced for multi-perspective parallel computing. Then the reachability evaluation function from the nodes to the nodes is expressed as: ; Among them, represents the learnable weight matrix of the output layer, represents the operation of concatenating multiple vectors, represents the k-th attention head.
5. The collaborative autonomous data routing method for a multi-UAV system according to claim 4, wherein The k-th attention head , and the calculation method is as follows: ; Among them, , and represent the learnable weight matrices of the k-th attention head.
6. The collaborative autonomous data routing method for a multi-UAV system according to claim 5, wherein The generating of the effective neighbor set and effective action set according to the node reachability evaluation result in Step 1-4 includes the following steps: According to the node to the node between the reachability evaluation function , the reachability judgment between nodes is carried out, which is expressed as follows: ; Among them, is the node pruning function, is the threshold of reachability. If the and the between the reachability evaluation function value is less than the threshold, the node pruning function value is 0, and it is judged as unreachable, indicating that the needs to prune the node; The set of neighbor nodes after pruning, i.e., the effective neighbor set, is , which is shown as follows: ; Node The set of valid actions for is as follows: 。 7. A collaborative autonomous data routing method for a multi-UAV system according to claim 6, characterized in that, The using of the policy network and value network to make independent decisions for individual UAVs and collaborative decisions for the UAV cluster in Step 2 includes the following steps: Step 2-1, assume each drone selects an action according to the current policy network as , and obtains a reward of after interacting with the environment, and its next state is ; store the global experience in the experience replay buffer, and the update process of the experience replay buffer is as follows: ; Step 2-2, set target networks corresponding to the policy network and value network, perform training and update to obtain the trained policy network and value network; Among them, the policy network update method is expressed as follows: ; Among them, represents the parameters of the policy network, represents the policy gradient, which is used to update the parameters of the policy network , represents the expected value sampled from the experience replay buffer according to the policy , represents the action taken in the state , represents the estimated advantage function, which is expressed as follows: ; Among them, is the action value function, representing the expected return when taking action in state and following policy ; is the state value function, representing the expected return when following policy in state ; State value function The update of, that is, the value network update method is expressed as follows: ; wherein, are the parameters of the value network, represents the value function gradient for updating the parameters of the value network, represents the state distribution under the policy and represents the expected value sampled according to the state distribution under the policy and is the value function of the target network, expressed as follows: ; Among them, is the immediate reward, is the discount factor, is the state transition probability; The parameter update method of the target network is as follows: ; Among them, is an adjustable parameter used to control the speed of updating the target network parameters; Step 2-3: The trained policy network selects the maximum action based on the set of valid neighbors and the set of valid actions , as the preliminary decision, which is expressed as follows: ; Among them, the trained policy network is , indicating that the action is selected in the state as , are the parameters of the policy network.
8. The collaborative autonomous data routing method for a multi-UAV system according to claim 7, wherein The optimizing of the preliminary decisions of all UAVs in Step 3 includes: Each UAV makes a routing decision based on local observations. Through the reward function, each UAV independently selects the next-hop node, which is expressed as follows: ; Among them, represents the joint probability distribution of the next state and the reward under the condition of a given current state and an action . represents the joint probability distribution of the next state and the reward under the condition of a given current state and an action, as well as all previous states and actions .
9. The collaborative autonomous data routing method for a multi-UAV system according to claim 8, wherein, The described reward , through each drone 's reward function is calculated, and the reward function is expressed as follows: ; Among them, , and are adjustable parameters, is the forward propagation reward of the data packet, is the environmental reward, is the minimum hop count reward, is the delay reward.
10. A collaborative autonomous data routing method for a multi-UAV system according to claim 9, characterized in that Forward propagation reward of data packet , which is used to indicate whether the data has reached the sink node, i.e., the end point of data packet transmission, and at the same time, step penalty is carried out to limit the behavior of nodes to transmit data back and forth. The calculation method is as follows: ; Among them, is the positive reward for successful data arrival, is the penalty based on the time step; Environmental reward , the calculation method is as follows: ; Among them, represents the wind speed magnitude at the node , and represents the environmental noise intensity within the area where the node is located, and respectively represent the weight factors that control the influence of wind speed and environmental noise intensity on stability; Minimum hop count reward , and the calculation method is as follows: ; Among them, is the shortest path length from the next-hop node of the decision to the aggregation node ; Time delay reward , the calculation method is as follows: ; Among them, is an adjustable parameter, is the normalized propagation delay, and its value range is [0, 1]. The calculation method is as follows: ; Among them, and respectively represent the maximum and minimum propagation delays between nodes.
Citation Information
Patent Citations
Networking method and system for software-defined wireless ad hoc network controlled by unmanned aerial vehicles
CN110167032A
Unmanned aerial vehicle cluster network intelligent multi-hop routing method based on multi-agent cooperation
CN114499648A
Path planning method of unmanned aerial vehicle and unmanned vehicle heterogeneous cooperative system
CN114779758A
Air-ground cooperative service migration method based on deep reinforcement learning
CN116248688A
Terminal direct connection satellite computing power network task scheduling and routing optimization method
CN117676711A