An intelligent trajectory planning method for a relay unmanned aerial vehicle in a near-space communication network

The method addresses the challenge of relay drone trajectory planning in near-space communication networks by using a multi-agent reinforcement learning framework to optimize network performance and energy efficiency in complex and dynamic scenarios.

CN120129016BActive Publication Date: 2025-07-15NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510618214.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-07-15
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the trajectory planning of complex adjacent space communication network relay drones, especially in dynamic scenarios, and it is difficult to achieve multi-objective joint optimization and independent decision-making.

Method used

The multi-agent reinforcement learning method based on graph attention network (GAT), gated recurrent unit network (GRU) and soft actor-criticist algorithm (SAC) is adopted. By constructing a multi-agent Markov decision-making process (MDP) model, relay drone trajectory planning is carried out, and multi-modal information of the network topology graph is extracted in combination with graph attention network. The gated recurrent unit network is used to process time dependencies, and the soft actor-criticist algorithm realizes independent decision-making.

Benefits of technology

In complex network topology and dynamic scenarios, independent decision-making and multi-objective collaborative optimization have been achieved, which improves the robustness and overall performance of the network, optimizes resource utilization and energy consumption, and improves the service quality of ground users and the total capacity of the communication network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120129016B_ABST
    Figure CN120129016B_ABST
Patent Text Reader

Abstract

The present invention provides an intelligent trajectory planning method for a near-space communication network relay unmanned aerial vehicle, which includes: Step 1, designing a method for generating a communication network topology graph: constructing a directed graph called a fully connected graph, and adding directed edges E from the base station node to each relay unmanned aerial vehicle, between each relay unmanned aerial vehicle, and from each relay unmanned aerial vehicle to each ground user, where the edges between each relay unmanned aerial vehicle have two directions; Step 2, converting the trajectory planning process of the stratospheric airship into a sequential decision-making process and constructing a multi-agent Markov decision-making process model; Step 3, designing a policy network and a value network and constructing a virtual simulation environment for the trajectory planning of the relay unmanned aerial vehicle; Step 4, training the agent based on the soft actor-critic algorithm to find the optimal policy. The method of the present invention can plan the position of the relay unmanned aerial vehicle in real time, adjust the communication link, and maintain network connectivity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of unmanned aerial vehicles and integrated communication and sensing technologies, and particularly relates to an intelligent trajectory planning method for a relay unmanned aerial vehicle in a near-space communication network. Background Art

[0002] A near-space communication (NS-COM) network is a network that uses near-space platforms (such as high-altitude balloons, stratospheric airships, high-altitude long-endurance relay unmanned aerial vehicles, etc.) at an altitude of 20 km - 100 km above the ground for communication. It has the characteristics of a large coverage area, a long coverage time, and a low node handover frequency, and can provide highly reliable and continuous communication services for areas lacking ground base stations.

[0003] In an air-based station (ABS) near-space communication network based on a stratospheric airship platform, a relay unmanned aerial vehicle is used as a relay node to construct a non-terrestrial network (NTN) of air-based station - relay unmanned aerial vehicle - ground user. This network consists of 1 air-based station, M relay unmanned aerial vehicles, and N ground users. Each ground user multiplexes frequency band resources by means of time division multiple access (TDMA). It can further expand the coverage area of the communication network, improve the service quality of ground users at the edge of the coverage area of the air-based station, and has high flexibility in emergency tasks that require quick response. Summary of the Invention

[0004] Object of the Invention: The technical problem to be solved by the present invention is to provide an intelligent trajectory planning method for a relay unmanned aerial vehicle in a near-space communication network in view of the deficiencies of the prior art, including the following steps:

[0005] Step 1: Design a method for generating a communication network topology graph: First, construct a directed graph called a complete graph , where the nodes V include 1 air-based station, M relay unmanned aerial vehicles, and N ground users; add directed edges E from the base station node to each relay unmanned aerial vehicle, between each pair of relay unmanned aerial vehicles, and from each relay unmanned aerial vehicle to each ground user, where the edges between each pair of relay unmanned aerial vehicles have two directions;

[0006] Step 2: Convert the trajectory planning process of the stratospheric airship into a sequential decision-making process, and construct a multi-agent Markov Decision Process (MDP) model to describe the trajectory planning task of the relay UAV; discretize the time-continuous planning process. Within each time slice, each agent (referring to the constructed multi-agent) observes the environmental state, makes decisions, and receives rewards in sequence; at the beginning of the multi-agent Markov decision process, each relay UAV starts from an arbitrary position in a rectangular plane area with a length of A and a width of C and flies at a constant altitude; if any relay UAV flies out of the boundary, or when the total number of time slices reaches the upper limit, the multi-agent Markov decision process terminates; according to the task requirements, design the state space, action space, and the reward function regarding the states of two adjacent time slices and and as the optimization objective;

[0007] Step 3: Design a policy network and a value network. The structures of the policy network and the value network are the same, both including a graph attention network, a gated recurrent unit network, and a multi-layer perceptron; the role of the policy network is to output the corresponding policy according to the input state; the role of the value network is to output the action value function value of the state and action combination according to the input state and action combination, and evaluate the reward effect of performing the output action in the current state; construct a virtual simulation environment for the relay UAV trajectory planning to realize the generation of the communication network topology graph, the random movement of the space-based base station and ground users, the interaction between multi-agents and the environment, and the state transition;

[0008] Step 4: In the virtual simulation environment of Step 3, train the agent to find the optimal policy based on the Soft Actor-Critic algorithm, and adopt a multi-agent reinforcement learning architecture of distributed training and distributed execution. Each agent includes an independent Actor network and a Critic network. Among them, Actor represents the actor, and selects the next forward direction according to the state of the environment and its own policy, but the direction may not be correct; Critic represents the critic, and calculates a value function through the choices made by the actor and the feedback of the environment for evaluation; the training and execution processes are carried out inside each agent.

[0009] Step 1 includes:

[0010] Step 1-1: Node V includes all nodes in the communication network. Traverse all ground users. For each ground user, use the depth-first search algorithm to search for L multi-hop communication links from the space-based base station via the relay UAV to ground user n, and abstract the th multi-hop communication link as a directed acyclic graph with only a single source and a single sink , where v and e represent specific nodes and edges respectively, ; The number of nodes and the number of edges included are respectively and ;

[0011] Step 1-2, traverse all multi-hop communication links of the ground user n, and abstract the link from the i-th node to the j-th node in each multi-hop communication link as an edge . The feature of the edge is the corresponding link capacity . The feature of the i-th node includes the capacity of the edges flowing into and the capacity of the edges flowing out of , that is ; ;

[0012] According to the line-of-sight transmission model, calculate the channel capacity according to the following formula:

[0013] ,

[0014] ,

[0015] where represents the channel bandwidth, represents the channel gain from the k-th node to the j-th node , is an intermediate parameter in the calculation formula. To simplify the expression, it is written out separately. represents the spatial coordinate of the k-th node , represents the transmit power of the k-th node ;

[0016] When , the corresponding is the signal power at the receiving end, represents the variance of Gaussian white noise;

[0017] When , the corresponding represents the noise power caused by co-channel interference;

[0018] Let the capacity of the -th multi-hop communication link from the space-based base station to the ground user be equal to the minimum value of each link therein, which is expressed as:

[0019] ,

[0020] Select the multi-hop communication link with the largest capacity The multi-hop communication link actually used is:

[0021] ,

[0022] in Refers to returning the corresponding ;

[0023] The multi-hop communication links actually used by each ground user are combined, and , get the actual link connection diagram; The node and edge features of The corresponding node features and edge features are spliced together;

[0024] Steps 1-3, for ground users All feasible multi-hop communication links are combined to form the multi-hop communication link graph of all ground user n ,make (put all multi-hop communication links together), Node features and edge features for Corresponding node features The mean and edge features of The mean of :

[0025] ,

[0026] ;

[0027] Steps 1-4, for all ground users Take and merge to form a communication network topology diagram , ; Node features and edge features Depend on The corresponding node features and edge features are spliced together:

[0028] ,

[0029] ,

[0030] in represents the i-th node feature of the n-th ground user; Represents the edge feature of the edge from i to j of the nth ground user;

[0031] Steps 1-5, if you As the input of the neural network, the following preprocessing is first performed: The nodes and edges of are complemented according to the directed graph, and the value 0 is filled into the corresponding positions of the node features and edge features.

[0032] Step 2 includes: The state space is divided into two parts: where is the graph state, is the state of the relay UAV, and the communication network topology graph includes node features, edge features, and an adjacency matrix.

[0033] All the states of the relay UAVs in the one-dimensional array structure are:

[0034] ,

[0035] where represents the current time, and respectively represent the abscissa and ordinate of the position of the m-th relay UAV, and respectively represent the velocity in the x-direction and the velocity in the y-direction of the m-th relay UAV, represents the propulsion power consumption of the m-th relay UAV.

[0036] The observation space of the agent is equal to the state space, and the action space where respectively represent the modulus and the direction angle of the velocity increment of the m-th relay UAV;

[0037] The reward function ;

[0038] where is the global reward, is the local reward:

[0039] ,

[0040] ,

[0041] where, is the global minimum capacity reward term, which is positive and is proportional to the minimum value of the capacity when each ground user adopts the optimal link : , ;

[0042] is the global total capacity reward term, which is positive and is related to the actual links of all ground users The capacity in is directly proportional to the sum of , ;

[0043] is the local average capacity reward term, which is positive and is directly proportional to the sum of the capacities passing through this relay UAV node in , that is, it is directly proportional to the characteristics of this node: , ;

[0044] is the UAV propulsion power consumption reward term, which is negative, and its absolute value is directly proportional to the propulsion power consumption of this UAV: , ;

[0045] is the site boundary reward term, which is negative, and its absolute value is negatively correlated with the distance from this UAV to the nearest boundary, and when this distance is higher than the threshold, this reward term is 0.

[0046] Step 3 includes:

[0047] Step 3-1: Read the location information of the space-based base station, relay UAVs, and ground users to generate a communication network topology graph and the flight state of the relay UAVs , and splice them into the state at the current moment t ;

[0048] Step 3-2: Input the communication network topology graph into the graph attention network of the policy network, and extract new graph node features through message passing and aggregation;

[0049] Step 3-3: Traverse all agents. For each agent, organize the extracted graph node features into a one-dimensional vector, splice it with the flight state of the relay UAV, and then pass it through a gated recurrent unit network and a fully connected layer, and output the decision actions expressed in the form of the mean and standard deviation of a Gaussian distribution, and sample the decision actions corresponding to each agent from the mean and standard deviation ;

[0050] Step 3-4: Apply the decision actions to the corresponding relay UAVs to cause the self-state transition of the relay UAVs; update the position information of each node at the next moment, generate a new communication network topology graph, splice it into the state at the next moment , and calculate the rewards obtained by each agent according to the reward function .

[0051] Step 4 includes:

[0052] Step 4-1, initialize the parameters of the agent neural network;

[0053] Step 4-2, before the start of each round of agent training, first read the position information of the empty base station, relay UAVs, and ground users, generate a communication network topology , and randomly initialize the flight states of each relay UAV , and splice them into the initial state ;

[0054] Step 4-3, in each time step of each round of training, traverse all agents, and each agent makes corresponding decision actions , update the positions and flight states of each node after interacting with the environment, generate a new communication network topology , the current state transfers to the next moment state , and obtain the reward ; when any relay UAV crosses the boundary, or the cumulative number of time steps reaches the preset upper limit, the multi-agent Markov decision process terminates; store the interaction data group into the replay area, represents the current state, represents the decision action, is the decision action of the m-th agent, is the reward, is the state of the next moment;

[0055] Step 4-4, randomly read the interaction data from the replay area, traverse all agents, and for each agent, calculate the loss function of the actor-critic algorithm neural network of each agent according to the neural network update rule of the soft actor-critic algorithm, and backpropagate the neural network to update the network parameters in the way of gradient descent.

[0056] The soft actor-critic algorithm optimizes the Critic network by the method of gradient descent on the loss function , and optimizes the Actor network by the method of gradient descent on the loss function , and The calculation formulas of are:

[0057] ,

[0058] ,

[0059] where, represents the reward obtained in the t-th step; represents the discount factor, and its value is between 0 and 1; is the target value function. The SAC algorithm adopted by the present invention uses an off-policy action value function learning mode. In this learning mode, there are action value functions with the same structure. and the target action value function , The output value of is used for updating the action value function. The output value of is used to evaluate the expected reward after taking an action in a state. Using the following formula, the internal parameters of are used to update the internal parameters of : :

[0060] ,

[0061] where the update rate is a constant;

[0062] is the action value function, which is used to evaluate the expected reward after taking an action in the state ; after taking the action The term represents the policy entropy. Its introduction can improve the policy exploration ability and help the training converge to the global optimum.

[0063] The present invention also provides an electronic device, including a processor and a memory. The memory stores program code. When the program code is executed by the processor, the processor is caused to execute the steps of the method.

[0064] The present invention also provides a storage medium, storing a computer program or instruction. When the computer program or instruction runs on a computer, the steps of the method are executed.

[0065] The present invention is applied to the following scenario: Based on the topology map of a communication network, multiple agents are trained using a multi-agent reinforcement learning algorithm. Each agent is a flying relay drone in the network for trajectory planning. By planning its flight trajectory, the quality of service (QoS) of multiple ground users is optimized, and the propulsion energy consumption of the relay drones in the task is reduced. Compared with the prior art, the advantages of the present invention include:

[0066] It is applicable to dealing with complex communication network topologies. Since the service area of this network has a large spatial scale, from the airborne base station to the ground users, it needs to be connected through multi-hop communication links. Moreover, the number of ground users is large, and the same relay UAV acts as a node in several links almost simultaneously, so the network topology is relatively complex. The algorithm proposed in the present invention uses a Graph Attention Network (GAT) to fuse and extract multi-modal information of each node and each edge of the network topology graph, including position, uplink capacity, downlink capacity, etc., providing a basis for the agent's decision-making;

[0067] It is applicable to dynamic scenarios where the node positions change over time. Since the airborne base station is based on a stratospheric airship platform, it is easily affected by wind field disturbances and has a long response time to control commands. Therefore, it has random movement in the horizontal direction. Coupled with the movement of ground users and the constraint of the maximum communication distance, the connection situation of communication links and the channel capacity both change accordingly. The algorithm proposed in the present invention uses a Gate Recurrent Unit (GRU) to extract the time series information during the task process to handle the time-dependent relationship of the network topology structure, helping the agent to plan in dynamic scenarios;

[0068] The present invention has high intelligence, realizes multi-objective joint optimization, and achieves overall optimality. The algorithm proposed in the present invention is based on the Soft Actor-Critic (SAC) algorithm, explores the optimal strategy through training iteration, and realizes autonomous decision-making without relying on external instructions. During the training phase, by reasonably allocating the weights of each item of the reward function, a balance is achieved among multiple optimization objectives, and the optimal strategy that maximizes the total reward is obtained through training. Using a cooperative multi-agent reinforcement learning algorithm, each agent does not make decisions depending on its own state, but observes its own and other agents' states and makes collaborative decisions, thereby realizing the optimization of the overall goal.

[0069] Aiming at the problems of complex network topology and dynamic changes in node positions in the near-space communication network relayed by relay UAVs, considering the joint optimization of multiple objectives, the present invention proposes an intelligent trajectory planning method based on a Graph Attention Network (GAT), a Gate Recurrent Unit (GRU), and a Soft Actor-Critic (SAC) algorithm, effectively solving the limitations of traditional quadratic optimization methods under dynamic constraints and the deficiencies of existing deep reinforcement learning methods in dealing with complex network topologies, and realizing autonomous decision-making and multi-objective collaborative optimization in complex dynamic scenarios.

[0070] Beneficial effects: (1) Adaptability to complex network topologies: The present invention uses a Graph Attention Network (GAT) to efficiently fuse and extract multi-modal information (such as position, uplink capacity, downlink capacity, etc.) of each node and edge in the network topology graph, significantly improving the decision-making ability of the agent in a complex multi-hop communication link environment.

[0071] (2) Robustness in dynamic scenarios: For the dynamic scenario where the positions of space-based base stations and ground users change over time, the present invention uses a Gated Recurrent Unit Network (GRU) to extract time series information and effectively handle the time-dependent relationship of the network topology. In the dynamic environment of the movement of space-based base stations and ground users, the algorithm of the present invention can plan the position of relay drones in real time, adjust the communication link, and maintain network connectivity.

[0072] (3) Intelligence of multi-objective joint optimization: The present invention realizes multi-agent collaborative decision-making based on the Multi-Agent Soft Actor-Critic algorithm (MASAC), and achieves a balance among multiple optimization objectives by reasonably allocating the weights of the reward function. The training results show that the algorithm of the present invention realizes joint optimization on multiple objectives such as resource utilization rate, energy consumption, and communication quality. In addition, by using a cooperative multi-agent reinforcement learning algorithm, each agent further improves the overall performance of the network through collaborative decision-making. During the training process, the quality of service of each ground user and the total capacity of the communication network are improved, the power consumption-capacity ratio is reduced, and the energy utilization efficiency is enhanced. Description of the Drawings

[0073] Figure 1 It is a schematic diagram for generating a communication network topology graph.

[0074] Figure 2 It is a schematic diagram of a multi-agent MDP model.

[0075] Figure 3 It is a structure diagram of a policy network.

[0076] Figure 4 It is a structure diagram of a value network.

[0077] Figure 5 It is a graph of the minimum capacity during the training process.

[0078] Figure 6 It is a graph of the total network capacity during the training process.

[0079] Figure 7 It is a graph of the propulsion power consumption during the training process.

[0080] Figure 8 It is the actual link connection diagram of the first step in a specific case.

[0081] Figure 9 It is the actual link connection diagram of the 34th step in a specific case.

[0082] Figure 10 It is the actual link connection diagram of the 115th step in a specific case.

[0083] Figure 11 It is the actual link connection diagram of the 168th step in a specific case.

[0084] Figure 12 It is the actual link connection diagram of the 180th step in a specific case. Specific implementation manners

[0085] The present invention will be further specifically described below in conjunction with the accompanying drawings and specific implementation manners, and the above and / or other advantages of the present invention will become clearer.

[0086] The present invention provides an intelligent trajectory planning method for a near-space communication network relay unmanned aerial vehicle, including the following steps:

[0087] Step 1. Design a method for generating a communication network topology diagram:

[0088] Step 1-1. First, construct a directed graph called a fully connected graph , where the nodes include all nodes in the communication network, and the nodes include 1 space-based base station, M relay unmanned aerial vehicles, and N ground users; add directed edges E from the base station node to each relay unmanned aerial vehicle, between each relay unmanned aerial vehicle, and from each relay unmanned aerial vehicle to each ground user, where the edges between each relay unmanned aerial vehicle have two directions; traverse all ground users, and for each ground user, use the depth-first search (DFS) algorithm to search for the of multi-hop communication links from the space-based base station via several relay unmanned aerial vehicles to the ground user and abstract the multi-hop communication link as a directed acyclic graph with only a single source and a single sink . The number of nodes and edges included in and respectively.

[0089] Step 1-2. Traverse all multi-hop communication links of the ground user n, and abstract the link from the i-th node to the j-th node in each multi-hop communication link as an edge , and the feature of the edge is the corresponding link capacity , and the feature of the i-th node includes the inflow The capacity of the edge and the outflow The capacity of the edge , that is ;

[0090] According to the line-of-sight transmission model, the channel capacity is calculated according to the following formula:

[0091] ,

[0092] ,

[0093] Select the multi-hop communication link with the largest capacity among them as the actually adopted multi-hop communication link:

[0094] ;

[0095] Among them refers to returning the corresponding ;

[0096] Take the union of the actually adopted multi-hop communication links of each ground user, and let , the actual link connection graph can be obtained. The node features and edge features of are spliced by the corresponding node features and edge features.

[0097] Step 1-3, for the ground user Take the union of all feasible multi-hop communication links to form the multi-hop communication link graph of all ground users n , let , The node features and edge features of are The mean values of the corresponding node features and edge features, that is , .

[0098] Step 1-4, for all ground users Take the union to form the communication network topology graph . The node features and edge features of are spliced by the corresponding node features and edge features of , that is:

[0099] ,

[0100] .

[0101] Step 1-5, if is used as the input of the neural network, preprocessing is also required, and the nodes and edges of are arranged according to the fully connected graph Complete it and fill the value 0 into the corresponding positions of the node features and edge features.

[0102] The entire process of Step 1 is as Figure 1 shown: First, construct a fully connected graph from the nodes , then traverse all the multi-hop communication links of each ground user, and select the multi-hop communication link with the largest capacity for each ground user as the actually adopted multi-hop communication link for each ground user. Then take the union of the actually adopted multi-hop communication links of each ground user to form the actual link connection graph . Each ground user takes the union of all the feasible multi-hop communication links to form the graph of all the multi-hop communication links of ground user n , and then take the union of those of all the ground users to form the communication network topology graph . . Figure 1 In , the blue five-pointed star represents the stratospheric airship base station, the red dot represents the relay UAV, the green cross represents the ground user, and the black arrow represents the actual link between each node

[0103] Step 2: Convert the trajectory planning process of the stratospheric airship into a sequential decision-making process, construct a multi-agent Markov decision process (MDP) model to describe the relay UAV trajectory planning task; discretize the time-continuous planning process. Within each time slice, each agent (referring to the constructed multi-agent) observes the environmental state, makes a decision, and receives a reward in sequence; at the beginning of the MDP, each relay UAV starts from an arbitrary position in a rectangular plane area with length A and width C and flies at a fixed altitude; if any relay UAV flies out of the boundary, or when the total number of time slices reaches the upper limit, the MDP terminates; according to the task requirements, design the state space, action space, and the reward function regarding the states and of two adjacent time slices as the optimization objective;

[0104] State space , the communication network topology graph contains node features, edge features, and an adjacency matrix. The flight states of all relay UAVs in a one-dimensional array structure:

[0105] ,

[0106] where represents the current time, , represent the position coordinates of the m-th relay UAV, , represents the speed of the m-th relay UAV, represents the propulsion power consumption of the m-th relay UAV. To enable the agent to have a global perception of the state of the communication network and thus achieve the overall optimality of the optimization goal, the agent can observe the global state. Therefore, the observation space of the agent is equal to the state space. The action space , is the modulus and direction angle of the speed increment of the m-th relay UAV.

[0107] Reward function , which is divided into global reward and local reward: The global reward acts on all agents and is related to the overall performance of multiple agents. , including related to the minimum value of the capacities of each ground user and related to the sum of the capacities of all ground users. .

[0108] The local reward acts on each agent respectively and is related to the performance of the relay UAV node affected by the decision of this agent. Among them, depends on the mean value of the smaller value among the inflow and outflow traffic of this node in each ground user graph described in Step 1, that is, the minimum value of the characteristics of this node; is related to the power consumption of this relay UAV; is related to whether the relay UAV goes out of bounds and its distance to the boundary.

[0109] The entire process of Step 2 is as Figure 2 shown. In each time slice, each agent observes the environmental state, makes a decision, and receives a reward in turn. The environment calculates the global reward and local reward according to the actions made by the agents and inputs them to each agent, and the process repeats.

[0110] Step 3, design the policy network and value network. The structures of both networks include a graph attention network, a gated recurrent unit network, and a multi-layer perceptron. Among them, the role of the policy network is to output the corresponding policy according to the input state; the role of the value network is to output the action value function value of the state-action combination according to the input state-action combination, and evaluate the reward effect of executing this action in this state. Build a virtual simulation environment for relay UAV trajectory planning on a computer to realize the generation of the communication network topology graph, simulate the random movement of the space-based base station and ground users, the interaction between multiple agents and the environment, and state transition, etc.

[0111] Step 3-1, read the position information of the space-based base station, relay UAVs, and ground users, and generate the communication network topology graph and the flight state of the relay UAVs , concatenated into a state .

[0112] Step 3-2: Feed the graph attention network of the input policy network, and through message passing and aggregation, extract new graph node features.

[0113] Step 3-3: Traverse all agents. For each agent, organize the extracted graph node features into a one-dimensional vector, concatenate it with the flight state of the relay UAV, and then pass it through a gated recurrent unit network and a fully connected layer, and output the decision action expressed in the form of the mean and standard deviation of a Gaussian distribution. The corresponding decision action of the agent is sampled from the mean and standard deviation .

[0114] Step 3-4: Apply the to the corresponding relay UAV to cause the state transition of the relay UAV itself. Update the position information of each node at the next moment, generate a new communication network topology graph, and concatenate it into a state . Calculate the rewards obtained by each agent according to the reward function .

[0115] The entire process of Step 3 is as shown in Figure 3 , Figure 4 shown. Figure 3 shows the structure of the policy network. The graph attention network of the input policy network extracts new graph node features, organizes the extracted graph node features into a one-dimensional vector, concatenates it with the flight state of the relay UAV, and then passes it through a gated recurrent unit network and a fully connected layer, and outputs the decision action;

[0116] Figure 4 shows the structure of the value network, which is somewhat similar to the policy network but also has differences. The graph attention network of the input value network extracts new graph node features, organizes the extracted graph node features into a one-dimensional vector, concatenates it with the flight state of the relay UAV and the decision action, and then passes it through a gated recurrent unit network and a fully connected layer, and calculates and outputs the reward.

[0117] Step 4: In the simulation environment of Step 3, train the agents based on the Soft Actor-Critic (SAC) algorithm to find the optimal policy. Select a multi-agent reinforcement learning architecture of distributed training and distributed execution. Each agent includes an independent Actor network and a Critic network, and the training and execution processes are carried out inside each agent. The SAC algorithm optimizes the Critic network by using the gradient descent method for the loss function The Actor network is optimized by the gradient descent method. and The calculation method is as follows:

[0118] ,

[0119] ;

[0120] Step 4-1, initialize the parameters of the agent neural network;

[0121] Step 4-2, before the start of each round of agent training, first read the position information of the empty base station, relay UAVs, and ground users, generate a communication network topology map , and randomly initialize the flight states of each relay UAV , and splice them into the initial state ;

[0122] Step 4-3, at each time step of each round of training, traverse all agents, and each agent makes corresponding decisions , update the positions and flight states of each node after interacting with the environment, generate a new communication network topology map , the environmental state transfers to , and obtain the reward . When any relay UAV crosses the boundary, or the cumulative number of time steps reaches the preset upper limit, the MDP process of this round terminates. Store the interaction data group in the replay area;

[0123] Step 4-4, randomly read several pieces of interaction data from the replay area. Traverse all agents. For each agent, according to the neural network update rule of the SAC algorithm, calculate the loss function of each network, and update the network parameters by backpropagation of the neural network in the way of gradient descent.

[0124] The following is a specific application case of the method of the present invention: training in a virtual simulation environment. The maximum number of steps allowed in each round is 180, the time interval for each step is 1 minute, and the simulation site is a square area. The communication network includes 1 empty base station node, 3 relay nodes, and 10 ground user nodes. Among them, the empty base station node is based on a stratospheric airship platform and flies at a constant altitude of 20 km; the relay nodes are based on UAV platforms and fly at a constant altitude of 2 km; the ground user nodes are all on the ground at an altitude of 0 km. The positions of the empty base station and the ground users will have random displacements to simulate the phenomenon of the airship being disturbed by the stratospheric wind field and the random movement of the ground users. The relay UAVs are controlled by agents and satisfy a maximum speed of 25 m / s and a maximum acceleration The kinematic constraints are assumed to have an on-board energy source sufficient to support a 180-minute flight. The communication device has a transmit power of 1 W, a white noise power of -20 dBm, uses a directional beam technology, has a signal gain of 20 dB, a noise gain of 9.5 dB, and a maximum transmission distance of 50 km. The data transmission rate per unit bandwidth, i.e., the spectral density, is used to measure the capacity.

[0125] During the training process, the change curves of each performance index are as Figure 5 , Figure 6 , Figure 7 shown. Figure 5 is the minimum capacity curve graph during the training process; Figure 6 is the total network capacity curve graph during the training process; Figure 7 is the propulsion power consumption curve graph during the training process; the abscissa represents the number of training rounds, the ordinate represents the reward, and the ordinate value has no unit and the dimension is 1. As the number of training rounds increases, the capacity of the ground user with the minimum capacity and the total capacity achieved by the entire network both increase, and the propulsion energy consumption of the UAV decreases, effectively improving the network efficiency and taking into account the quality of service of each ground user.

[0126] The trajectory planning algorithm obtained through sufficient training is subjected to simulation testing. Figure 8 , Figure 9 , Figure 10 , Figure 11 , Figure 12 shows the actual link connection diagram of the communication network at different times, which can show the positions of each node and the link connection conditions. Among them, the blue five-pointed star represents the space-based base station, the red dot represents the relay UAV, and the green cross represents the ground user. The black arrow represents the actual link between each node , and the red arrow marks the multi-hop communication link corresponding to the ground user with the minimum capacity of the multi-hop communication link. MAXF represents the capacity of this multi-hop communication link, Agent_0 Reward represents the reward obtained by the first agent, Agent_1 Reward represents the reward obtained by the second agent, and Agent_2 Reward represents the reward obtained by the third agent. Actions represents the actions performed by the corresponding agent. The dimensions of Agent_0 Reward, Agent_1 Reward, Agent_2 Reward, Actions, and MAXF are all 1 and have no unit. Figures 8 to 12 The number next to the red dot in

[0127] Step 1, as Figure 8 shown: Agent_0 Reward: 28.93, Actions: -0.68, -0.44;

[0128] Agent_1 Reward: 32.17, Actions: 1.0, 0.97;

[0129] Agent_2 Reward: 32.78, Actions: -0.97, -1.0;

[0130] MAXF: 0.01;

[0131] Step 34, as Figure 9 shown: Agent_0 Reward: 76.64, Actions: -0.62, 0.11;

[0132] Agent_1 Reward: 52.02, Actions: 1.0, -0.43;

[0133] Agent_2 Reward: 27.9, Actions: -1.0, -0.98;

[0134] MAXF: 0.014;

[0135] Step 115, as Figure 10 shown: Agent_0 Reward: 30.39, Actions: 0.51, 0.37;

[0136] Agent_1 Reward: 60.95, Actions: 0.99, 0.55;

[0137] Agent_2 Reward: 34.59, Actions: -1.0, -0.8;

[0138] MAXF: 0.017;

[0139] Step 168, as Figure 11 shown: Agent_0 Reward: 43.55, Actions: 0.21, 0.24;

[0140] Agent_1 Reward: 70.13, Actions: 0.98, -0.55;

[0141] Agent_2 Reward: 55.38, Actions: -0.99, 0.37;

[0142] MAXF: 0.021;

[0143] Step 180, as Figure 12Shown: Agent_0 Reward: 46.04, Actions: -0.43, 0.39;

[0144] Agent_1 Reward: 77.54, Actions: 0.97, -0.39;

[0145] Agent_2 Reward: 59.01, Actions: -0.98, 0.0;

[0146] MAXF: 0.026;

[0147] In the initial state, the positions of each node are randomly distributed. It can be seen that as the simulation time progresses, each agent realizes the optimization of the network capacity in the dynamic scenario of the changes in the positions of each node and the link connection status by real-time planning the movement trajectory of the relay UAV.

[0148] Therefore, the present invention adopts a relay UAV intelligent trajectory planning method based on the topology map of the near-space communication network, realizes the joint trajectory planning of each relay UAV node in the wireless communication network, optimizes the service quality of all ground users and the total capacity of the communication network, and optimizes the power consumption-capacity ratio, with high energy utilization rate.

[0149] The present invention provides a relay UAV intelligent trajectory planning method for a near-space communication network. There are many methods and ways to specifically implement this technical solution. The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by existing technologies.

Claims

1. An intelligent trajectory planning method for a relay unmanned aerial vehicle in a near-space communication network, characterized in that, It includes the following steps: Step 1. Design a method for generating a communication network topology diagram: First, construct a directed graph called a fully connected graph , where the nodes include all nodes in the communication network, and the nodes include 1 empty base station, M relay drones, and N ground users; add directed edges E from the base station node to each relay drone, between each pair of relay drones, and from each relay drone to each ground user, where the edges between each pair of relay drones have two directions; Step 2: Convert the trajectory planning process of the stratospheric airship into a sequential decision-making process, and construct a multi-agent Markov decision process model to describe the trajectory planning task of the relay UAV; discretize the time-continuous planning process. Within each time slice, each agent observes the environmental state, makes a decision, and receives a reward in turn; at the beginning of the multi-agent Markov decision process, each relay UAV starts from an arbitrary position in a rectangular plane area with a length of A and a width of C and flies at a constant altitude; if any relay UAV flies out of the boundary or the total number of time slices reaches the upper limit, the multi-agent Markov decision process terminates; according to the task requirements, design the state space, action space, and the reward function regarding the states of two adjacent time slices as the optimization objective; and ​ Step 3: Design a policy network and a value network. The policy network and the value network have the same structure, both including a graph attention network, a gated recurrent unit network, and a multi-layer perceptron; Among them, the role of the policy network is to output corresponding policies according to the input state; the role of the value network is to output the action value function value of the state and action combination according to the input state and action combination, and judge the reward effect of executing the output action in the current state; construct a virtual simulation environment for relay UAV trajectory planning to realize the generation of communication network topology maps, the random movement of space-based base stations and ground users, the interaction between multiple agents and the environment, and state transition; Step 4: In the virtual simulation environment of Step 3, train the agent to find the optimal policy based on the soft actor-critic algorithm, and adopt a multi-agent reinforcement learning architecture of distributed training and distributed execution. Each agent includes an independent Actor network and a Critic network, where Actor represents the actor and Critic represents the critic; The training and execution processes are carried out inside each agent; Step 3 includes: Step 3-1, read the location information of the empty base station, relay UAV, and ground user, and generate a communication network topology diagram and the flight state of the relay UAV , and splice them into the state at the current moment t ; Step 3-2, input the communication network topology diagram into the graph attention network of the policy network, and extract new graph node features through message passing and aggregation; Step 3-3: Traverse all agents. For each agent, organize the extracted graph node features into a one-dimensional vector, concatenate it with the flight state of the relay UAV, and then pass it through a gated recurrent unit network and a fully connected layer to output a decision action expressed in the form of the mean and standard deviation of a Gaussian distribution. Sample from the mean and standard deviation to obtain the corresponding decision actions for each agent ; Step 3-4, apply the decision-making action to the corresponding relay UAV to transfer the state of the relay UAV itself; update the position information of each node at the next moment, generate a new communication network topology map, and splice it into the state at the next moment , calculate the rewards obtained by each agent according to the reward function .

2. The method according to claim 1, wherein Step 1 includes: Step 1-1, traverse all ground users. For each ground user, use the depth-first search algorithm to search for multi-hop communication links from the aerial base station via relay UAVs to ground user n, and abstract the th multi-hop communication link as a directed acyclic graph with only a single source and a single sink , where v and e represent specific nodes and edges respectively, ; The number of nodes and the number of edges included are and respectively; Step 1-2, traverse all multi-hop communication links of ground user n, and abstract the link from the i-th node to the j-th node in each multi-hop communication link as an edge . The feature of edge is the corresponding link capacity . The feature of the i-th node includes the capacity of the incoming edges and the capacity of the outgoing edges, that is . That is ; ​ According to the line-of-sight transmission model, calculate the channel capacity according to the following formula: , , Among them represents the channel bandwidth, represents from the k-th node to the j-th node channel gain; is an intermediate parameter in the calculation formula, represents the spatial coordinates of the k-th node ; represents the transmission power of the k-th node ; When correspondingly is the signal power at the receiving end, represents the variance of Gaussian white noise; When correspondingly represents the noise power caused by co-channel interference; Let the capacity of the th multi-hop communication link from the empty base station to the ground user be equal to the minimum value of each link therein, expressed as: , Select the multi-hop communication link with the largest capacity among them as the multi-hop communication link actually adopted: , Among them means to return the corresponding ; Take the union of the multi-hop communication links actually adopted by each ground user, and let to obtain the actual link connection diagram; The node features and edge features of are spliced by the corresponding node features and edge features; Step 1-3, for ground users Take the union of the multi-hop communication link graphs of each one to form all the multi-hop communication link graphs of ground user n , let , The node features of and edge features be The mean of the corresponding node features and the mean of the edge features : , ; Steps 1-4, for all ground users' Take the union to form a communication network topology graph , ; The node features of and edge features are concatenated by the corresponding node features and edge features: , , wherein represents the i-th node feature of the n-th ground user; represents the edge feature of the edge from i to j of the n-th ground user; Step 1-5, if [[ ]] is used as the input of the neural network, the following preprocessing is first performed: the nodes and edges of [[ ]] are supplemented according to the directed graph [[ ]], and the value 0 is filled in the corresponding positions of the node features and edge features.​​​​​​ 3. The method according to claim 2, wherein Step 2 includes: state space which is divided into two parts: where is the graph state, is the relay UAV state, and the communication network topology graph includes node features, edge features, and an adjacency matrix.

4. The method according to claim 3, characterized in that, Step 2 further includes: All the states of the relay UAVs in the one-dimensional array structure are: , Among them represents the current time, and respectively represent the abscissa and ordinate of the position of the m-th relay UAV, and respectively represent the x-direction speed and y-direction speed of the m-th relay UAV, represents the propulsion power consumption of the m-th relay UAV.

5. The method according to claim 4, characterized in that, Step 2 further includes: the observation space of the agent is equal to the state space, and the action space , where respectively represent the modulus and the direction angle of the speed increment of the m-th relay UAV; Reward function ; wherein is the global reward, is the local reward: , , Among them, is the global minimum capacity reward term, which is a positive value and is proportional to the minimum value of the capacity when each ground user adopts the optimal link ; is proportional to: , ; It is a global total capacity reward item, which is a positive value and is proportional to the sum of the capacities of all terrestrial users in the actual link : , ;​ is the local average capacity reward term, which is a positive value, and the formula is: , ; It is a power consumption reward item for the drone, which is a negative value. The formula is: , ; It is a site boundary reward item.

6. The method according to claim 5, wherein Step 4 includes: Step 4-1: Initialize the neural network parameters of the agent; Step 4-2, before the start of each round of agent training, first read the location information of the empty base station, relay drones, and ground users to generate a communication network topology map , and randomly initialize the flight states of each relay drone , and splice them into the initial state ; Step 4-3: At each time step of each round of training, traverse all agents, and each agent makes corresponding decision actions , update the positions and flight states of each node after interacting with the environment, and generate a new communication network topology , the current state transfers to the state at the next moment , and obtains a reward ; when any relay UAV crosses the boundary, or the cumulative number of time steps reaches the preset upper limit, the multi-agent Markov decision process terminates; store the interaction data group in the replay area, represents the current state, represents the decision action, is the decision action of the m-th agent, is the reward, is the state at the next moment; Step 4-4: Randomly read the interaction data from the replay area, traverse all agents. For each agent, calculate the loss function of the actor-critic algorithm neural network of each agent according to the neural network update rule of the soft actor-critic algorithm, and backpropagate the neural network to update the network parameters in a gradient descent manner.

7. The method according to claim 6, characterized in that, 4-4 includes: The soft actor-critic algorithm optimizes the Critic network by means of gradient descent on the loss function and optimizes the Actor network by means of gradient descent on the loss function , and the calculation formula of and is as follows: , , Among them, represents the reward obtained at the t-th step; represents the discount factor, with a value between 0 and 1; is the target value function, adopting the learning mode of the action value function under the off-policy, and there are action value functions with the same structure and the target action value function , The output value of is used for the update of the action value function, The output value of is used to evaluate the expected reward after taking an action in the state. By the following formula, using The internal parameters of Update The internal parameters of : , where the update rate is a constant; is the action value function, used to evaluate the expected reward after taking action in state ; The item represents the policy entropy.

8. An electronic device, characterized in that, It includes a processor and a memory. The memory stores program code. When the program code is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 7.

9. A storage medium, characterized in that, It stores a computer program or instruction. When the computer program or instruction runs on a computer, it executes the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method for optimizing network traffic scheduling based on near-end strategy optimization algorithm

    CN115550268A

  • Distributed multi-unmanned aerial vehicle relay network coverage method

    CN116017479A