Unmanned aerial vehicle cluster intelligent opportunistic routing method based on multi-agent deep reinforcement learning

Through the distributed intelligent routing method of multi-agent reinforcement learning, the problem of drone communication network dynamically adjusting routing paths in complex terrain environments is solved, efficient data relay and network optimization of drone clusters in complex environments is realized, and real-time and reliability of data transmission are improved.

CN120358564APending Publication Date: 2025-07-22NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510146990.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Traditional drone communication networks are difficult to dynamically adjust routing paths in complex terrain environments, which makes it difficult to ensure the real-time and reliability of data transmission. Especially in dynamic environments where drone cluster node distributions change rapidly, traditional routing protocols are prone to delay decision-making or failure.

Method used

Using a distributed intelligent routing method based on multi-agent reinforcement learning, the distributed partial observation Markov decision-making process model and multi-agent near-end strategy optimization reinforcement learning algorithm are realized to achieve efficient data relay for drone clusters under complex terrain.

Benefits of technology

In complex terrain environments, the data transmission efficiency and reliability of the drone cluster are significantly improved, and can quickly adapt to dynamic environment changes, optimize routing paths, and improve the end-to-end transmission efficiency of the network and the successful transmission rate of data packets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358564A_ABST
    Figure CN120358564A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle cluster intelligent opportunistic routing method based on multi-agent deep reinforcement learning, and aims to optimize a data transmission path in an unmanned aerial vehicle cluster and improve communication efficiency. The method comprises the following steps: firstly, establishing a communication channel model and a motion model of an unmanned aerial vehicle cluster, and constructing a distributed partial observation Markov decision process model; based on the model, a multi-agent near-end strategy optimization reinforcement learning algorithm is adopted to carry out routing decision training, and after training is completed, a collaborative strategy of the agents is loaded into an airborne computer of the unmanned aerial vehicle, so that intelligent opportunistic routing is realized. In the training process, the unmanned aerial vehicle selects a proper next-hop node to optimize data forwarding by continuously collecting neighbor information and making a local decision. Through the method, the unmanned aerial vehicle cluster can dynamically adapt to a complex environment, and the end-to-end transmission efficiency of a network and the successful transmission rate of data packets are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of unmanned aerial vehicle (UAV) ad-hoc communication networks, and particularly to a distributed intelligent routing method that uses multi-agent reinforcement learning to achieve efficient data relaying of UAV swarms in complex terrain environments. Background Art

[0002] With the rapid development of UAV technology, UAV swarms have shown great potential and wide applications in fields such as military reconnaissance, disaster relief, and logistics distribution. In military reconnaissance, UAV swarms can perform large-scale and continuous monitoring tasks, providing important support for battlefield situation awareness; in disaster relief, UAV swarms can quickly reach the affected areas for environmental monitoring and resource distribution; in logistics distribution, UAV swarms can efficiently complete material transportation in remote areas. However, these application scenarios often involve complex terrain environments, such as mountainous areas, forests, and densely populated urban areas, which pose severe challenges to UAV communication networks.

[0003] In complex terrain environments, obstacles can cause interruptions in communication links, and signal attenuation and multipath effects seriously affect the transmission efficiency and reliability of data. Especially in dynamic environments, due to the strong mobility and rapid change of node distribution of UAV swarms, traditional rule-based routing protocols are difficult to adapt. These protocols usually rely on predefined routing strategies or centralized computing, and are prone to decision-making delays or routing failures when facing frequent changes in network topology. In addition, traditional routing methods lack the ability of global optimization in complex scenarios and are difficult to dynamically adjust routing paths to adapt to changes in node states, resulting in difficulties in guaranteeing the real-time performance and reliability of data transmission.

[0004] Therefore, there is an urgent need in current research for an intelligent routing method that can combine distributed cooperation and adaptive optimization characteristics to meet the communication requirements of UAV swarms in complex terrains. Multi-agent reinforcement learning, due to its adaptive learning ability and distributed decision-making advantages in dynamic environments, has become an ideal technical path to solve this problem. By introducing intelligent learning algorithms, UAV swarms can adjust routing strategies in real time in dynamic environments, achieving the efficiency, stability, and robustness of data transmission, and providing new technical ideas for UAV communication networks in complex scenarios. Summary of the Invention

[0005] Object of the Invention: The present invention aims to provide a distributed intelligent routing method for UAV swarms based on multi-agent reinforcement learning. By introducing the multi-agent proximal policy optimization algorithm, efficient data relaying and network optimization of UAV swarms in complex terrain environments are achieved.

[0006] Technical Solution: To achieve the object of the present invention, the technical solution adopted by the present invention is: an intelligent opportunistic routing method for UAV clusters based on multi-agent reinforcement learning, specifically including the following steps:

[0007] Step 1: Given the components of the intelligent UAV cluster and the communication environment and communication channel model of the UAV cluster;

[0008] Step 2: Set the UAV cluster motion model and the representation of the UAV cluster network topology;

[0009] Step 3: Construct a distributed partially observable Markov decision process model;

[0010] Step 4: According to the established distributed partially observable Markov decision process model, establish a multi-agent intelligent opportunistic routing method based on the multi-agent proximal policy optimization reinforcement learning algorithm;

[0011] Step 5: Train the multi-agent intelligent opportunistic routing method based on the multi-agent proximal policy optimization reinforcement learning algorithm. After training, load the cooperative intelligent opportunistic routing policy of each agent into the on-board computer of the UAV and run the intelligent opportunistic routing method.

[0012] The specific steps for given the components of the intelligent UAV cluster and the communication channel model and communication environment of the UAV cluster are as follows:

[0013] Step 1-1: The overall structure of the intelligent UAV cluster includes the UAV cluster and the ground mobile command center. The ground mobile command center can train the reinforcement learning algorithm carried by the UAV cluster, be responsible for supervising and commanding the activities of the UAV cluster, and execute tasks in complex terrain environments;

[0014] Step 1-2: Model and classify the communication channels between UAVs. According to whether there is a direct line-of-sight path between the transmitter and the receiver, it is divided into the line-of-sight transmission and non-line-of-sight transmission channel attenuation models;

[0015] Step 1-3: If there is no occlusion between UAVs, then use the free space path loss model to model the channel attenuation Lf:

[0016]

[0017] where d ij represents the distance between UAV i and UAV j, and λ is the wavelength of the communication signal;

[0018] Step 1-4: If there are mountain obstacles blocking between UAVs, then use the non-line-of-sight transmission channel attenuation model. The additional attenuation Lo caused by the obstacles is:

[0019]

[0020] where v is the Fresnel diffraction parameter, and its definition is:

[0021]

[0022] Among them, h p is the peak height of the obstacle, that is, the vertical distance from the obstacle to the straight line connecting the two UAV nodes. d1 is the distance from UAV i to the perpendicular line where the obstacle is located, d2 is the distance from UAV j to the obstacle plane, and λ is the wavelength of the communication signal;

[0023] Step 1-5: The attenuation value of the channel attenuation with obstacles is:

[0024] L ij = Lf ij + Lo ij (4)

[0025] Step 1-6: The calculation formula for the received power of the communication signal between UAVs is as follows:

[0026] Pr j = Pt i - L ij + Gt i + Gr j (5)

[0027] Among them, Pr j is the signal power received by UAV j from UAV i, Pt i is the transmission power of UAV i, Gf i is the transmission antenna gain of UAV i, Gr j is the receiving antenna gain of UAV j;

[0028] Step 1-7: When receiving a signal, the thermal noise interference of UAV j is:

[0029] N j = k·T j ·B·F (6)

[0030] Among them, N j is the thermal noise power of UAV j, k is the Boltzmann constant, T j is the temperature of the receiver of UAV j, B is the communication bandwidth of the UAV, and F is the noise figure of the receiver of UAV j;

[0031] Step 1-8: The signal-to-interference-plus-noise ratio of the signal received by UAV j is:

[0032]

[0033] Among them, I j is the sum of the received powers of the signals transmitted by the remaining nodes at UAV j during the process of UAV i transmitting a signal to UAV j;

[0034] Step 1-9: The condition for UAV j to successfully receive the signal from UAV i is as follows:

[0035] SINR j ≥SINR th and Pr j ≥Pr th (8)

[0036] where SINR th is the received signal-to-noise ratio threshold, and Pr th is the received power threshold.

[0037] The specific steps for setting the UAV swarm motion model and representing the UAV swarm network topology are as follows:

[0038] Step 2-1: Assume that the intelligent UAV swarm contains N rotor UAVs, and each UAV is denoted as {n1, n2, …, n N};

[0039] Step 2-2: Set the environmental range in which the intelligent UAV swarm operates. The moving ranges of the UAVs and the ground mobile command and control center are from (0, 0, 0) to (x max , y max , z max ), where x max , y max , z max represent the maximum values of the mission area on the x, y, and z axes respectively;

[0040] Step 2-3: Set the position of UAV i as q i =(x i , y i , z i ), the speed as v i =(v x,i , v y,i , v z,i ), and set the position of the ground mobile command center as q c =(x c , y c , z c ), v c =(v x,c , v y,c , v z,c );

[0041] Step 2-4: Set the digital elevation model of the mission area, represented by the matrix M. M is a two-dimensional discrete matrix, and the sampling intervals along the x-axis and y-axis are Δx and Δy respectively.

[0042] The specific steps for constructing the distributed partially observable Markov decision process model are as follows:

[0043] Step 3-1: Set the set of agents in the current Markov decision process model as Each UAV is an independent agent;

[0044] Step 3-2: Set the global state S as the set of local observation states o of each UAV agent i i :

[0045] S = {o1, o2,..., o n} (9)

[0046] Step 3-3: Set the local observation o of each UAV agent i i as:

[0047] o i = {Info c , Info pkt (i), Info nbr (i), Info link (i)} (10)

[0048] where Info c includes the position and speed of the ground mobile command center; Info pkt (i) is the set of data packet information to be sent by UAV i, including the source node ID, the upstream node ID, the network layer queue length, the traffic priority, and the data packet timestamp; Info nbr (i) is the set of neighbor node information of UAV i, including the node ID, the predicted position, the running speed, the distance to the ground mobile command center, and the neighbor busy degree; Info link (i) is the set of evaluations of the neighbor link state of UAV i, including the received power redundancy and the SINR redundancy;

[0049] The definition of the received power redundancy of the link between UAV nodes i and j is:

[0050] Pr r,ij = min(Pr i - Pr th , Pr j - Pr th ) (11)

[0051] The definition of the SINR redundancy of the link between UAV nodes i and j is:

[0052] SINR r,ij = min(SINR i - SINR th , SINR j - SINR th ) (12)

[0053] The positions of the agent inputs are all normalized to obtain the normalized positions:

[0054]

[0055] where q is the original position coordinate, q min is the minimum coordinate value, and q max is the maximum coordinate value;

[0056] Step 3-4: Set the output vector of each agent i as zi, which is a 1*N vector. The output action passes through the softmax normalization function, representing the probability of currently selecting each neighbor agent as the next-hop node for data forwarding:

[0057] a i = Softmax(z i ) (14)

[0058] Step 3-5: The action selection process is that the drone samples according to the action probability distribution until the cumulative probability exceeds the threshold a th ∈[0, 1], and selects the forwarding node set Tx:

[0059]

[0060] Step 3-6: Set the reward function given by the environment to the agent. The reward is divided into three parts and is defined as follows:

[0061] r i,1 = min(P r,ij , SINR r,ij ) - α·dis(q c - q j ) - r f δ fail (16)

[0062] The first part is the single-hop forwarding reward of the drone. α is the scaling coefficient, dis() is the drone distance calculation function, r f is the reward for forwarding failure, and δ fail is the forwarding failure flag, which is 1 if the forwarding fails and 0 otherwise. The second part is the end-to-end forwarding reward of the drone and is defined as follows:

[0063]

[0064] where r route is the total reward for successful data packet forwarding, and m is the total number of forwarding hops. The third part is the delay penalty and is defined as follows:

[0065] r i,3 = -η·delay ij(18)

[0066] where η is the delay penalty adjustment coefficient, and delay ij is the transmission delay from node i to node j. The total reward function is defined as follows:

[0067] R l = r i,1 + r i,2 + r i,3 。 (19)

[0068] According to the established distributed partially observable Markov decision process model, the specific steps of the multi-agent intelligent opportunistic routing method based on the multi-agent proximal policy optimization reinforcement learning algorithm are as follows:

[0069] Step 4-1: In the ground mobile command center, configure the Actor network and the Critic network of the multi-agent proximal policy optimization reinforcement learning algorithm. The Actor network will be loaded for the UAV agent to make action decisions, and the Critic network is located in the ground mobile command center for algorithm training;

[0070] Step 4-2: Define the loss function in the reinforcement learning policy training process:

[0071]

[0072] where L actor is the loss function of the Actor network, and ρ i,θ = π θ (a i ) / π θold (a i ) is the ratio of the old and new policies, π is the policy network, is the generalized advantage estimation, y(t) = r t + - γVφ(s t-1 ), is the optimization target value, where γ is the attenuation factor, and V φ (st) is the state value estimated by the Critic network, and θ and φ represent the neural network parameters of the Actor network and the Critic network respectively;

[0073] Step 4-3: In the training stage, utilize the computing power resources of the ground mobile command center to run the virtual environment, save the running records to the replay buffer, and periodically perform mini-batch sampling on the data in the buffer. Calculate the loss function based on the sampled data, train the Critic network and the Actor network, update the neural network parameters, centrally evaluate the actions of each agent and optimize the policy parameters until the algorithm converges and achieves a better cumulative reward;

[0074] Step 4-4: The training termination conditions of the algorithm include two conditions, and either one can terminate the training: The algorithm has converged and achieved a good cumulative reward, and the reward value has not increased significantly for a period of time, or the algorithm training has reached the preset number of rounds or training duration.

[0075] Train the multi-agent intelligent opportunity routing method based on the multi-agent proximal policy optimization reinforcement learning algorithm. After the training is completed, load the cooperative intelligent opportunity routing policy of each agent into the on-board computer of the UAV. The specific steps of running the intelligent opportunity routing method are as follows:

[0076] Step 5-1: After the algorithm training is completed, obtain the parameters in the Actor network, and pack and transfer its parameters into the on-board computer of the UAV for use by the routing protocol of the UAV;

[0077] Step 5-2: After the UAV cluster operates in the air, continuously collect neighbor information, and make independent decisions on the collected partial observation data, make routing decisions on the data packets to be forwarded, and forward the data packets to the next-hop node to support the normal operation of the UAV cluster network. Brief Description of the Drawings

[0078] The drawings here are incorporated into the specification and form a part of this specification, showing the embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0079] Figure 1 A step diagram showing the multi-agent intelligent opportunity routing method for UAV clusters based on multi-agent reinforcement learning in an exemplary embodiment of the disclosure;

[0080] Figure 2 A structural diagram showing the UAV cluster based on the multi-agent proximal policy optimization reinforcement learning algorithm in an exemplary embodiment of the disclosure;

[0081] Figure 3 A graph showing the average reward curve during the training process of the multi-agent proximal policy optimization reinforcement learning algorithm in Scenario 1 in an exemplary embodiment of the disclosure;

[0082] Figure 4 A graph showing the average reward curve during the training process of the multi-agent proximal policy optimization reinforcement learning algorithm in Scenario 2 in an exemplary embodiment of the disclosure. Detailed Embodiments

[0083] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.

[0084] In addition, the accompanying drawings are only schematic illustrations of the embodiments of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities.

[0085] A method for intelligent opportunistic routing of an unmanned aerial vehicle (UAV) cluster based on multi-agent reinforcement learning is provided in this example embodiment. The reference diagram of this method is as Figure 1 shown and specifically includes the following steps:

[0086] Step 1: Specify the components of the intelligent UAV cluster and the communication environment and communication channel model of the UAV cluster;

[0087] Step 2: Set the motion model of the UAV cluster and the network topology representation of the UAV cluster;

[0088] Step 3: Construct a distributed partially observable Markov decision process model;

[0089] Step 4: Based on the established distributed partially observable Markov decision process model, establish a multi-agent intelligent opportunistic routing method based on the multi-agent proximal policy optimization reinforcement learning algorithm;

[0090] Step 5: Train the multi-agent intelligent opportunistic routing method based on the multi-agent proximal policy optimization reinforcement learning algorithm. After training is completed, load the cooperative intelligent opportunistic routing policy of each agent into the on-board computer of the UAV and run the intelligent opportunistic routing method.

[0091] The specific steps for specifying the components of the intelligent UAV cluster and the communication channel model and communication environment of the UAV cluster are as follows:

[0092] Step 1-1: The overall structure of the intelligent UAV cluster includes the UAV cluster and the ground mobile command center. The ground mobile command center can train the reinforcement learning algorithm carried by the UAV cluster, be responsible for supervising and commanding the activities of the UAV cluster, and execute tasks in complex terrain environments;

[0093] Step 1-2: Model and classify the communication channels between UAV clusters, and divide them into line-of-sight transmission and non-line-of-sight transmission channel attenuation models according to whether there is a direct line-of-sight path between the transmitter and the receiver;

[0094] Step 1-3: If there is no occlusion between UAVs, then use the free space path loss model to model the channel attenuation Lf:

[0095]

[0096] where d ij represents the distance between UAV i and UAV j, and λ is the wavelength of the communication signal;

[0097] Step 1-4: If there are mountain obstacles between UAVs, then use the non-line-of-sight transmission channel attenuation model, and the additional attenuation Lo caused by the obstacle is:

[0098]

[0099] where v is the Fresnel diffraction parameter, and its definition is:

[0100]

[0101] where, h p is the peak height of the obstacle, that is, the vertical distance from the obstacle to the straight line connecting the two UAV nodes, d1 is the distance from UAV i to the perpendicular line where the obstacle is located, d2 is the distance from UAV j to the obstacle plane, and λ is the wavelength of the communication signal;

[0102] Step 1-5: The attenuation value of the channel attenuation with obstacles is:

[0103] L ij = Lf ij + Lo ij (24)

[0104] Step 1-6: The calculation formula for the received power of the communication signal between UAVs is as follows:

[0105] Pr j = Pt i - L ij + Gt i + Gr j (25)

[0106] where Pr j is the signal power received by UAV j from UAV i, Pt i is the transmission power of UAV i, Gt i is the transmission antenna gain of UAV i, and Gr j is the receiving antenna gain of UAV j;

[0107] Step 1-7: When receiving a signal, the thermal noise interference of UAV j is:

[0108] N j = k·T j ·B·F (26)

[0109] where N j is the thermal noise power of UAV j, k is the Boltzmann constant, T j is the temperature of the receiver of UAV j, B is the communication bandwidth of the UAV, and F is the noise figure of the receiver of UAV j;

[0110] Step 1-8: The signal-to-interference-plus-noise ratio of the signal received by UAV j is:

[0111]

[0112] where I j is the sum of the received powers of the signals transmitted by the remaining nodes at UAV j during the process of UAV i transmitting a signal to UAV j;

[0113] Step 1-9: The condition for determining that UAV j can successfully receive the signal from UAV i is:

[0114] SINR j ≥SINR th and Pr j ≥Pr th (28)

[0115] where SINR th is the received signal-to-interference-plus-noise ratio threshold, and Pr th is the received power threshold.

[0116] The specific steps for setting the UAV cluster motion model and the UAV cluster network topology representation are as follows:

[0117] Step 2-1: Assume that the intelligent UAV cluster contains N rotor UAVs, and each UAV is denoted as {n1, n2,..., n N};

[0118] Step 2-2: Set the environmental range in which the intelligent UAV cluster operates. The moving ranges of the UAVs and the ground mobile command and control center are from (0, 0, 0) to (x max , y max , z max ), where x max , y max , z max represent the maximum values of the mission area on the x, y, and z axes respectively;

[0119] Step 2-3: Set the position of UAV i as q i =(x i , y i , z i ), and the speed as v i =(v x,i , v y,i , v z,i ). Set the position of the ground mobile command center as q c =(x c , y c , z c ), and V c =(v x,c , v y,c , v z,c );

[0120] Step 2-4: Set the digital elevation model of the mission area, represented by matrix M. M is a two-dimensional discrete matrix, and the sampling intervals along the x-axis and y-axis are Δx and Δy respectively.

[0121] The specific steps to construct the distributed partially observable Markov decision process model are as follows:

[0122] Step 3-1: Set the set of agents in the current Markov decision process model as Each UAV is an independent agent;

[0123] Step 3-2: Set the global state S as the set of local observation states o i of each UAV agent i:

[0124] S = {o1, o2,..., o n} (29)

[0125] Step 3-3: Set the local observation o i of each UAV agent i as:

[0126] o i = {Info c , Info pkt (i), Info nbr (i), Info link (i)} (30)

[0127] where Info c includes the position and speed of the ground mobile command center; Info pkt (i) is the set of packet information to be sent by UAV i, including the source node ID, the upstream node ID, the network layer queue length, the traffic priority, and the packet timestamp; Info nbr(i) is the set of neighbor node information of UAV i, including node ID, predicted position, running speed, distance to the ground mobile command center, and neighbor busyness level; Info link (i) is the set of evaluations of the neighbor link status by UAV i, including received power redundancy and SINR redundancy;

[0128] The definition of the received power redundancy of the link between UAV nodes i and j is:

[0129] Pr r,,ij =min(Pr i -Pr th ,Pr j -Pr th ) (31)

[0130] The definition of the SINR redundancy of the link between UAV nodes i and j is:

[0131] SINR r,ij =min(SINR i -SINR th ,SINR j -SINR th ) (32)

[0132] The positions input by the agent need to be normalized to obtain the normalized position:

[0133]

[0134] where q is the original position coordinate, q min is the minimum coordinate value, and q max is the maximum coordinate value;

[0135] Step 3-4: Set the output vector of each agent i as z i , which is a 1*N vector. The output action passes through the softmax normalization function and represents the probability of currently selecting each neighbor agent as the next-hop node for data forwarding:

[0136] a i =Softmax(z i ) (34)

[0137] Step 3-5: The action selection process is for the UAV to sample according to the action probability distribution until the cumulative probability exceeds the threshold a th ∈[0, 1], and select the forwarding node set Tx:

[0138]

[0139] Step 3-6: Set the reward function for the agent in the environment. The reward is divided into three parts and is defined as follows:

[0140] r i,1 = min(Pr r,ij , SINR r,ij ) - α·dis(q c - q j ) - r f S fail (36)

[0141] The first part is the single-drone relay reward. α is the scaling factor, dis() is the function for calculating the distance of the drone, r f is the reward for relay failure, δ fail is the relay failure flag, which is 1 if the relay fails and 0 otherwise. The second part is the end-to-end relay reward for the drone and is defined as follows:

[0142]

[0143] where r route is the total reward for successful data packet relay, and m is the total number of relay hops. The third part is the delay penalty and is defined as follows:

[0144] r i,3 = -η·delay ij (38)

[0145] where η is the delay penalty adjustment coefficient, and delay ij is the transmission delay from node i to node j. The total reward function is defined as follows:

[0146] R i = r i,1 + r l,2 + r i,3 . (39)

[0147] According to the established distributed partially observable Markov decision process model, a multi-agent intelligent opportunistic routing method based on the multi-agent proximal policy optimization reinforcement learning algorithm is established. The framework diagram of the algorithm combined with the intelligent drone cluster is as Figure 2 shown. The specific steps of the algorithm operation are as follows:

[0148] Step 4-1: In the ground mobile command and control center, configure the Actor network and Critic network of the multi-agent proximal policy optimization reinforcement learning algorithm. The Actor network will be loaded for the drone agent to make action decisions, and the Critic network is located in the ground mobile command and control center for algorithm training;

[0149] Step 4-2: Define the loss function in the reinforcement learning policy training process:

[0150]

[0151] where L actor is the loss function of the Actor network, and ρ i,θ = π θ (a i ) / π θold (a i ) is the ratio of the old and new policies, π is the policy network, is the generalized advantage estimation, y(t) = r t + γVφ(s t-1 ), which is the optimization target value, where γ is the attenuation factor, and V φ (s t ) is the state value estimated by the Critic network. θ and φ represent the neural network parameters of the Actor network and the Critic network respectively;

[0152] Step 4-3: In the training phase, utilize the computing power resources of the ground mobile command center to run the virtual environment, save the running records to the replay buffer, and periodically perform mini-batch sampling on the data in the buffer. Calculate the loss function based on the sampled data, train the Critic network and the Actor network, update the neural network parameters, centrally evaluate the actions of each agent and optimize the policy parameters, and use the Adam optimizer to update the network parameters until the algorithm converges and obtains a good cumulative reward;

[0153] Step 4-4: The training termination conditions of the algorithm include two conditions, and either one can terminate the training: the algorithm has converged and obtained a good cumulative reward, the reward value has not increased significantly within a period of time, or the algorithm training has reached the preset number of rounds or training duration.

[0154] Train the multi-agent intelligent opportunity routing method based on the multi-agent proximal policy optimization reinforcement learning algorithm. After the training is completed, load the collaborative intelligent opportunity routing policy of each agent into the on-board computer of the unmanned aerial vehicle. The specific steps for running the intelligent opportunity routing method are as follows:

[0155] Step 5-1: After the algorithm training is completed, obtain the parameters in the Actor network, pack its parameters and transfer them into the on-board computer of the unmanned aerial vehicle for use by the routing protocol of the unmanned aerial vehicle;

[0156] Step 5-2: After the unmanned aerial vehicle cluster operates in the air, continuously collect neighbor information, make independent decisions on the collected partial observation data, make routing decisions on the data packets to be forwarded, and forward the data packets to the next-hop node to support the normal operation of the unmanned aerial vehicle cluster network.

[0157] The operation effect of the method proposed by the present invention in two different-scale UAV cluster scenarios. Scenario 1 is the scenario of a cluster composed of 20 UAVs. The convergence of this method is as Figure 3 shown. The results show that this routing method can quickly adapt to the dynamic environment in a small-scale UAV cluster and provide efficient communication performance.

[0158] Scenario 2 is the scenario of a UAV cluster of 100 UAVs. Figure 2 This is the convergence result of this method, which proves that this method can still maintain high performance in a large-scale UAV cluster, demonstrating its advantages in a complex communication environment.

[0159] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example" or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0160] Those skilled in the art will readily conceive of other implementations of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.

Claims

1. An intelligent opportunistic routing method for UAV swarms based on multi-agent deep reinforcement learning, specifically including the following steps: Step 1: Given the components of the intelligent UAV swarm and the communication environment and communication channel model of the UAV swarm, the specific steps are as follows: Step 1-1: The overall structure of the intelligent UAV swarm includes the UAV swarm and the ground mobile command center. The ground mobile command center can train the reinforcement learning algorithm carried by the UAV swarm, be responsible for supervising and commanding the activities of the UAV swarm, and execute tasks in complex terrain environments; Step 1-2: Model and classify the communication channels between UAVs. According to whether there is a direct line-of-sight path between the transmitter and the receiver, it is divided into the line-of-sight transmission and non-line-of-sight transmission channel attenuation models; Step 1-3: If there is no occlusion between UAVs, then use the free space path loss model to model the channel attenuation Lf: where d ij represents the distance between the UAV i and the UAV j, and λ is the wavelength of the communication signal; Step 1-4: If there are mountain obstacles blocking between UAVs, then use the non-line-of-sight transmission channel attenuation model, and the additional attenuation Lo caused by the obstacle is: where v is the Fresnel diffraction parameter, and its definition is: where h p is the peak height of the obstacle, that is, the vertical distance from the obstacle to the straight line connecting the two UAV nodes, d1 is the distance from UAV i to the perpendicular line where the obstacle is located, d2 is the distance from UAV j to the obstacle plane, and λ is the wavelength of the communication signal; Step 1-5: The attenuation value of the channel attenuation with obstacles is: L ij = Lf ij + Lo ij (4) Step 1-6: The calculation formula for the received power of the communication signal between UAVs is as follows: Pr j = Pt i - L ij + Gt i + Gr j (5) Among them, Pr j is the signal power received by UAV j from UAV i, Pt i is the transmission power of UAV i, Gt i is the transmission antenna gain of UAV i, Gr j is the receiving antenna gain of UAV j; Step 1-7: When receiving the signal, the thermal noise interference of UAV j is: N j = k·T j ·B·F (6) where N j is the thermal noise power of the drone j, k is the Boltzmann constant, T j is the temperature of the receiver of the drone j, B is the communication bandwidth of the drone, and F is the noise figure of the receiver of the drone j; Step 1-8: The signal-to-interference-plus-noise ratio of UAV j receiving the signal is: Among which I j is the sum of the received powers of the signals transmitted by the remaining nodes at UAV j during the process of UAV i transmitting a signal to UAV j; Step 1-9: The condition for judging that UAV j can successfully receive the signal of UAV i is: SINR j ≥SINR th and Pr j ≥Pr th (8) where SINR th is the received signal-to-noise ratio threshold, and Pr th is the received power threshold; Step 2: Set the UAV swarm motion model and the UAV swarm network topology representation. The specific steps are: Step 2-1: Assume that there are N rotor drones in the intelligent drone cluster, and each drone is denoted as {n1, n2, …, n N}; Step 2-2: Set the environmental range in which the intelligent UAV cluster operates. The movement ranges of the UAVs and the ground mobile command and control center are from (0, 0, 0) to (x max , y max , z max ), where x max , y max , and z max respectively represent the maximum values of the mission area on the x, y, and z axes; Step 2-3: Set the position of UAV i as q i =(x i , y i , z i ), and the speed as v i =(v x,i , v y,i , v z,i ). Set the position of the ground mobile command center as q c =(x c , y c , z c ), v c =(v x,c , v y,c , v z,c ); Step 2-4: Set the digital elevation model of the mission area, represented by the matrix M. M is a two-dimensional discrete matrix, and the sampling intervals along the x-axis and y-axis are Δx and Δy respectively; Step 3: Build a distributed partially observable Markov decision process model. The specific steps are as follows: Step 3-1: Set the agent set in the current Markov decision process model as Each unmanned aerial vehicle is an independent agent; Step 3-2: Set the global state S as the set of local observation states oi of each UAV agent i: S = {o1, o2,..., o n} (9) Step 3-3: Set the local observation o of each UAV agent i i as follows: o i ={Info c , Info pkt (i), Info nbr (i), Info link (i)} (10) Among which Info c including the position and speed of the ground mobile command center; Info pkt (i) is the set of data packet information to be sent by UAV i, including the source node ID, the ID of the upstream node, the network layer queue length, the traffic priority, and the data packet timestamp; Info nbr (i) is the set of neighbor node information of UAV i, including the node ID, the predicted position, the running speed, the distance to the ground mobile command center, and the neighbor busyness; Info link (i) is the set of evaluations of the neighbor link status by UAV i, including the received power redundancy and the SINR redundancy; The definition of the received power redundancy of the link between UAV nodes i and j is: Pr r,ij = min(Pr i - Pr th , Pr j - Pr th ) (11) The definition of the SINR redundancy between UAV nodes i and j is: SINR r,ij = min(SINR i - SINR th , SINR j - SINR th ) (12) The input positions of the agents need to be normalized to obtain the normalized positions: where q is the original position coordinate, q min is the minimum value of the coordinate, q max is the maximum value of the coordinate; Step 3-4: Set the output vector of each agent i as z i , which is a 1*N vector. The output action passes through the softmax normalization function and represents the probability of currently selecting each neighbor agent as the next-hop node for data forwarding: a i = Softmax(z i ) (14) Step 3-5: In the action selection process, the UAV samples according to the action probability distribution until the cumulative probability exceeds the threshold a th ∈ [0, 1], and selects the forwarding node set Tx: Step 3-6: Set the reward function given by the environment to the agents. The reward is divided into three parts, and the definition is as follows: r i,1 = min(Pr r,ij , SINR r,ij ) - α·dis(q c - q j ) - r f δ fail (16) The first part is the single - drone relay forwarding reward, α is the scaling coefficient, dis() is the function for calculating the drone distance, r f is the reward for relay failure, δ fail is the relay failure flag, which is 1 if the relay fails and 0 otherwise. The second part is the drone end - to - end relay reward, which is defined as follows: where r route is the total reward for successful data packet forwarding, m is the total number of forwarding hops, and the third part is the delay penalty: r i,3 = -η·delay ij (18) where η is the delay penalty adjustment coefficient, delay ij is the transmission delay from node i to node j, and the total reward function is defined as follows: R i =r i,1 +r i,2 +r i,3 (19) Step 4: According to the established distributed partially observable Markov decision process model, establish a multi-agent intelligent opportunistic routing method based on the multi-agent proximal policy optimization reinforcement learning algorithm. The specific steps are as follows: Step 4-1: In the ground mobile command center, configure the Actor network and Critic network of the multi-agent proximal policy optimization reinforcement learning algorithm. The Actor network will be loaded for the UAV agent to make action decisions, and the Critic network is located in the ground mobile command center for algorithm training; Step 4-2: Define the loss function in the reinforcement learning policy training process: where L actor is the loss function of the Actor network, ρ i,θ = π θ (a i ) / π θold (a i ) is the ratio of the old and new policies, π is the policy network, is the generalized advantage estimation, y(t) = r t + γVφ(s t-1 ), is the optimization target value, where γ is the attenuation factor, V φ (s t ) is the state value estimated by the Critic network, and θ and φ represent the neural network parameters of the Actor network and the Critic network respectively; Step 4-3: In the training phase, utilize the computing power resources of the ground mobile command center to run the virtual environment, save the running records to the replay cache, and periodically perform mini-batch sampling on the data in the cache. Calculate the loss function based on the sampled data, train the Critic network and the Actor network, update the neural network parameters, centrally evaluate the actions of each agent, and optimize the policy parameters until the algorithm converges and achieves a good cumulative reward; Step 4-4: There are two termination conditions for the training of the algorithm. Training can be terminated as long as any one of them is met: the algorithm has converged and achieved a good cumulative reward, the reward value has not increased significantly within a certain period of time, or the algorithm training has reached the preset number of rounds or training duration; Step 5: Train the multi-agent intelligent opportunity routing method based on the multi-agent proximal policy optimization reinforcement learning algorithm. After training is completed, load the collaborative intelligent opportunity routing policy of each agent into the on-board computer of the drone and run the intelligent opportunity routing method. The specific steps are as follows: Step 5-1: After the algorithm training is completed, obtain the parameters in the Actor network, package and transfer its parameters to the on-board computer of the drone for use by the routing protocol of the drone; Step 5-2: After the drone cluster operates in the air, continuously collect neighbor information, make independent decisions on part of the observed data collected, make routing decisions on the data packets to be forwarded, and forward the data packets to the next-hop node to support the normal operation of the drone cluster network.

Citation Information

Cited By

  • Structural design system and method based on unified aircraft data format

    CN120724790A

  • System and method for structural design based on unified aircraft data format

    CN120724790B

  • Multi-network converged communication method and system based on dynamic weight self-adaption

    CN121645391A