Near space communication network relay unmanned aerial vehicle intelligent trajectory planning method
By applying multi-agent reinforcement learning methods of graph attention network, gated recurrent unit network and soft actor critic algorithm in adjacent space communication networks, communication service quality and energy consumption management problems under complex topological structures and dynamic node locations are solved, and efficient trajectory planning and optimization effects are achieved.
Patent Information
- Application Number
- CN202510618214.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The prior art is difficult to effectively deal with complex topological structures and dynamic node locations in adjacent spatial communication networks, resulting in difficult communication service quality and energy consumption management.
Multi-agent reinforcement learning methods based on graph attention network (GAT), gated recurrent unit network (GRU) and soft actor critic algorithm (SAC) are adopted to carry out intelligent trajectory planning of relay drones to optimize the topology and node location of the communication network.
It realizes efficient adaptation to the topology of complex communication networks and robust decision-making in dynamic scenarios, optimizes the service quality of ground users and the total capacity of communication networks, and reduces the propulsion energy consumption of relay drones.
Smart Images

Figure CN120129016A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of unmanned aerial vehicles and integrated communication and sensing technologies, and particularly to an intelligent trajectory planning method for a relay unmanned aerial vehicle in a near-space communication network. Background Art
[0002] A near-space communication (NS-COM) network is a network that uses near-space platforms (such as high-altitude balloons, stratospheric airships, high-altitude long-endurance relay unmanned aerial vehicles, etc.) at an altitude of 20 km - 100 km above the ground for communication. It has the characteristics of a large coverage area, a long coverage time, and a low node switching frequency, and can provide high-reliability and continuous communication services for areas lacking ground base stations.
[0003] In an aerial-based station (ABS) near-space communication network based on a stratospheric airship platform, a relay unmanned aerial vehicle is used as a relay node to construct a non-terrestrial network (NTN) of an aerial-based station - relay unmanned aerial vehicle - ground user. This network consists of 1 aerial-based station, M relay unmanned aerial vehicles, and N ground users. Each ground user multiplexes frequency band resources by means of time division multiple access (TDMA). It can further expand the coverage area of the communication network, improve the service quality of ground users at the edge of the coverage area of the aerial-based station, and has high flexibility in emergency tasks that require rapid response. Summary of the Invention
[0004] Object of the Invention: The technical problem to be solved by the present invention is to provide an intelligent trajectory planning method for a relay unmanned aerial vehicle in a near-space communication network in view of the deficiencies of the prior art, including the following steps: Step 1: Design a method for generating a communication network topology graph: First, construct a directed graph called a complete graph , where the nodes V include 1 aerial-based station, M relay unmanned aerial vehicles, and N ground users; add directed edges E from the base station node to each relay unmanned aerial vehicle, between each relay unmanned aerial vehicle, and from each relay unmanned aerial vehicle to each ground user, where the edges between each relay unmanned aerial vehicle have two directions; Step 2: Convert the trajectory planning process of the stratospheric airship into a sequential decision-making process, and construct a multi-agent Markov Decision Process (MDP) model to describe the trajectory planning task of the relay UAV; discretize the time-continuous planning process. Within each time slice, each agent (referring to the constructed multi-agent) observes the environmental state, makes decisions, and receives rewards in turn; at the beginning of the multi-agent Markov decision process, each relay UAV starts from an arbitrary position in a rectangular plane area with a length of A and a width of C and flies at a constant altitude; if any relay UAV flies out of the boundary or the total number of time slices reaches the upper limit, the multi-agent Markov decision process terminates; according to the task requirements, design the state space, action space, and the reward function regarding the states at two adjacent time slices and as the optimization objective. Step 3: Design a policy network and a value network. The structures of the policy network and the value network are the same, and both include a graph attention network, a gated recurrent unit network, and a multi-layer perceptron; the role of the policy network is to output the corresponding policy according to the input state; the role of the value network is to output the action value function value of the state and action combination according to the input state and action combination, and evaluate the reward effect of executing the output action in the current state; construct a virtual simulation environment for the relay UAV trajectory planning to realize the generation of the communication network topology graph, the random movement of the space-based base station and the ground users, the interaction between the multi-agent and the environment, and the state transition. Step 4: In the virtual simulation environment of Step 3, train the agent to find the optimal policy based on the soft actor-critic algorithm, and adopt a multi-agent reinforcement learning architecture of distributed training and distributed execution. Each agent includes an independent Actor network and a Critic network. Among them, Actor represents the actor, and selects the next forward direction according to the state of the environment and its own policy, but the direction may not be correct; Critic represents the critic, and calculates a value function through the choices made by the actor and the feedback of the environment for evaluation; the training and execution processes are carried out inside each agent.
[0005] Step 1 includes: Step 1-1: Node V includes all nodes in the communication network. Traverse all ground users. For each ground user, use the depth-first search algorithm to search for L multi-hop communication links from the space-based base station via the relay UAV to the ground user n, and abstract the th multi-hop communication link as a directed acyclic graph with only a single source and a single sink , where v and e represent specific nodes and edges respectively, ; The number of nodes and the number of edges included in and ; Step 1-2: Traverse all multi-hop communication links of ground user n, and abstract the link from the i-th node to the j-th node as an edge . The feature of edge is the corresponding link capacity . The feature of the i-th node includes the capacity of the edges flowing into and the capacity of the edges flowing out of , that is ; ; That is ; According to the line-of-sight transmission model, calculate the channel capacity according to the following formula: , , where represents the channel bandwidth, represents the channel gain from the k-th node to the j-th node , is an intermediate parameter in the calculation formula. To simplify the expression, it is written separately represents the spatial coordinate of the k-th node , represents the transmit power of the k-th node ; When , the corresponding is the signal power at the receiving end, represents the variance of Gaussian white noise; When , the corresponding represents the noise power caused by co-channel interference; Let the capacity of the -th multi-hop communication link from the space-based base station to the ground user be equal to the minimum value of each link therein, which is expressed as: , Select the multi-hop communication link with the largest capacity as the actually adopted multi-hop communication link: , where refers to returning the corresponding ; Take the union of the multi-hop communication links actually adopted by each ground user, and let to obtain the actual link connection diagram; The node features and edge features of are concatenated by the corresponding node features and edge features; Steps 1-3, for the ground user Take the union of all feasible multi-hop communication links to form all multi-hop communication link diagrams of ground user n Let (Combine all multi-hop communication links together), The node features of and the edge features are The mean value of the corresponding node features and the mean value of the edge features : , ; Steps 1-4, for all ground users Take the union to form the communication network topology diagram , ; The node features of and the edge features are composed of the corresponding node features and edge features concatenated by : , , where represents the i-th node feature of the n-th ground user; represents the edge feature of the edge from i to j of the n-th ground user; Step 1-5, if is used as the input of the neural network, the following preprocessing is performed first: the nodes and edges of are completed according to the directed graph , and the value 0 is filled in the corresponding positions of the node features and edge features.
[0006] Step 2 includes: the state space is divided into two parts: , where is the graph state, is the relay UAV state, and the communication network topology diagram includes node features, edge features and adjacency matrix.
[0007] All relay UAV states in the one-dimensional array structure are: , Among them represents the current time, and respectively represent the abscissa and ordinate of the position of the m-th relay UAV, and respectively represent the velocity in the x-direction and the velocity in the y-direction of the m-th relay UAV, represents the propulsion power consumption of the m-th relay UAV.
[0008] The observation space of the agent is equal to the state space, and the action space , among which respectively represent the modulus and direction angle of the velocity increment of the m-th relay UAV; Reward function ; Among which is the global reward, is the local reward: , , Among them, is the global minimum capacity reward term, which is positive and is proportional to the minimum value of the capacity when each ground user adopts the optimal link : , ; is the global total capacity reward term, which is positive and is proportional to the sum of the capacities in the actual link of all ground users : , ; is the local average capacity reward term, which is positive and is proportional to the sum of the capacities passing through this relay UAV node in , that is, proportional to the characteristics of this node: , ; is the UAV propulsion power consumption reward term, which is negative, and its absolute value is proportional to the propulsion power consumption of this UAV: , ; is the site boundary reward term, which is negative, and its absolute value is negatively correlated with the distance of this UAV to the nearest boundary, and when this distance is higher than the threshold, this reward term is 0.
[0009] Step 3 includes: Step 3-1, read the position information of the sky-based base station, relay UAVs, and ground users, and generate a communication network topology map With the flight state of the relay UAV , splice into the state at the current moment t ; Step 3-2: Input the communication network topology graph into the graph attention network of the policy network, and extract new graph node features through message passing and aggregation; Step 3-3: Traverse all agents. For each agent, organize the extracted graph node features into a one-dimensional vector, splice it with the flight state of the relay UAV, and then pass it through a gated recurrent unit network and a fully connected layer to output a decision action expressed in the form of the mean and standard deviation of a Gaussian distribution. The corresponding decision actions of each agent are sampled from the mean and standard deviation ; Step 3-4: Apply the decision action to the corresponding relay UAV to cause the self-state transition of the relay UAV; update the position information of each node at the next moment, generate a new communication network topology graph, and splice it into the state at the next moment , and calculate the rewards obtained by each agent according to the reward function .
[0010] Step 4 includes: Step 4-1: Initialize the neural network parameters of the agent; Step 4-2: Before the start of each round of agent training, first read the position information of the sky base station, relay UAV, and ground user to generate a communication network topology graph , and randomly initialize the flight state of each relay UAV , and splice it into the initial state ; Step 4-3: In each time step of each round of training, traverse all agents, and each agent makes a corresponding decision action , update the positions and flight states of each node after interacting with the environment, generate a new communication network topology graph , the current state transfers to the state at the next moment , and obtain rewards ; when any relay UAV crosses the boundary, or the cumulative number of time steps reaches the preset upper limit, the multi-agent Markov decision process terminates; store the interaction data group in the replay area, represents the current state, represents the decision action, is the decision action of the m-th agent, is the reward, is the state at the next moment; Step 4-4: Randomly read interaction data from the playback area, traverse all agents, and for each agent, calculate the loss function of the actor-critic algorithm neural network of each agent according to the neural network update rule of the soft actor-critic algorithm, and update the network parameters by backpropagating the neural network in a gradient descent manner.
[0011] The soft actor-critic algorithm optimizes the Critic network by the method of gradient descent, and optimizes the Actor network by the method of gradient descent. and The calculation formula is: , , where represents the reward obtained at the t-th step; represents the discount factor, with a value between 0 and 1; is the target value function. The SAC algorithm adopted in the present invention uses an off-policy action value function learning mode. In this learning mode, there are action value functions and target action value functions with the same structure. The output value of is used for the update of the action value function, and the output value of is used to evaluate the expected reward after taking an action in the state. The internal parameters of are updated as follows: , where the update rate is a constant; is the action value function used to evaluate the expected reward after taking an action in the state ; The
[0012] present invention also provides an electronic device, including a processor and a memory. The memory stores program codes. When the program codes are executed by the processor, the processor executes the steps of the above method.
[0013] The present invention also provides a storage medium storing a computer program or instruction. When the computer program or instruction runs on a computer, it executes the steps of the above method.
[0014] The present invention is applied to the following scenarios: Based on the topology diagram of a communication network, multiple agents are trained using a multi-agent reinforcement learning algorithm. Each agent performs trajectory planning for a flight relay unmanned aerial vehicle (UAV) in the network. By planning its flight trajectory, the quality of service (QoS) of multiple ground users is optimized, and the propulsion energy consumption of the relay UAVs in the mission is reduced. Compared with the prior art, the advantages of the present invention include: It is suitable for dealing with complex communication network topologies. Since the service area of this network has a large spatial scale, the connection from the space-based base station to the ground users requires multiple-hop communication links. Moreover, the number of ground users is large, and the same relay UAV serves as a node in almost several links simultaneously, so the network topology is relatively complex. The algorithm proposed by the present invention uses a Graph Attention Network (GAT) to fuse and extract multi-modal information of each node and each edge in the network topology diagram, including position, uplink capacity, downlink capacity, etc., to provide a basis for agent decision-making; It is suitable for dynamic scenarios where the node positions change over time. Since the space-based base station is based on a stratospheric airship platform, it is easily affected by wind field disturbances and has a long response time to control commands. Therefore, it has random movement in the horizontal direction. Coupled with the movement of ground users and the constraint of the maximum communication distance, the connection situation of communication links and the channel capacity both change accordingly. The algorithm proposed by the present invention uses a Gate Recurrent Unit (GRU) to extract the time series information during the mission process to handle the time-dependent relationship of the network topology structure and help the agent to plan in a dynamic scenario; The present invention has high intelligence, realizes multi-objective joint optimization, and achieves overall optimality. The algorithm proposed by the present invention is based on the Soft Actor-Critic (SAC) algorithm. Through training and iterative exploration of the optimal strategy, it realizes autonomous decision-making without relying on external instructions. During the training phase, by reasonably allocating the weights of each item in the reward function, a balance is achieved among multiple optimization objectives, and the optimal strategy that maximizes the total reward is obtained through training. Using a cooperative multi-agent reinforcement learning algorithm, each agent does not make decisions depending on its own state, but observes its own and other agents' states and makes collaborative decisions, thereby realizing the optimization of the overall goal.
[0015] Aiming at the problems of complex network topology structure and dynamic changes in node positions in the near-space communication network relayed by relay UAVs, considering the joint optimization of multiple objectives, an intelligent trajectory planning method based on Graph Attention Network (GAT), Gated Recurrent Unit Network (GRU) and Soft Actor-Critic algorithm (SAC) is proposed, which effectively solves the limitations of traditional quadratic optimization methods under dynamic constraints and the deficiencies of existing deep reinforcement learning methods in dealing with complex network topologies, and realizes autonomous decision-making and multi-objective collaborative optimization in complex dynamic scenarios.
[0016] Beneficial effects: (1) Adaptability to complex network topology: The present invention uses Graph Attention Network (GAT) to efficiently fuse and extract multi-modal information (such as position, uplink capacity, downlink capacity, etc.) of each node and edge in the network topology graph, significantly improving the decision-making ability of the agent in the complex multi-hop communication link environment.
[0017] (2) Robustness in dynamic scenarios: For the dynamic scenario where the positions of space-based base stations and ground users change over time, the present invention uses Gated Recurrent Unit Network (GRU) to extract time series information and effectively handle the time-dependent relationship of the network topology structure. In the dynamic environment of the movement of space-based base stations and ground users, the algorithm of the present invention can plan the position of relay UAVs in real time, adjust the communication link, and maintain network connectivity.
[0018] (3) Intelligence of multi-objective joint optimization: The present invention realizes multi-agent collaborative decision-making based on the multi-agent Soft Actor-Critic algorithm (MASAC), and achieves a balance among multiple optimization objectives by reasonably allocating the weights of the reward function. The training results show that the algorithm of the present invention realizes joint optimization on multiple objectives such as resource utilization rate, energy consumption and communication quality. In addition, by using the cooperative multi-agent reinforcement learning algorithm, each agent improves the overall performance of the network through collaborative decision-making. During the training process, the quality of service of each ground user and the total capacity of the communication network are improved, the power consumption-capacity ratio is reduced, and the energy utilization efficiency is enhanced. Description of the Drawings
[0019] Figure 1 It is a schematic diagram for generating a communication network topology graph.
[0020] Figure 2 It is a schematic diagram of a multi-agent MDP model.
[0021] Figure 3 It is a structure diagram of a policy network.
[0022] Figure 4 It is a structure diagram of a value network.
[0023] Figure 5 It is a minimum capacity curve graph during the training process.
[0024] Figure 6 It is the curve graph of the total network capacity during the training process.
[0025] Figure 7 It is the curve graph of the power consumption during the training process.
[0026] Figure 8 It is the actual link connection diagram of the first step in a specific case.
[0027] Figure 9 It is the actual link connection diagram of the 34th step in a specific case.
[0028] Figure 10 It is the actual link connection diagram of the 115th step in a specific case.
[0029] Figure 11 It is the actual link connection diagram of the 168th step in a specific case.
[0030] Figure 12 It is the actual link connection diagram of the 180th step in a specific case. Detailed implementation manners
[0031] The following further specifically describes the present invention in conjunction with the accompanying drawings and specific implementation manners, and the above and / or other advantages of the present invention will become clearer.
[0032] The present invention provides an intelligent trajectory planning method for a relay unmanned aerial vehicle in a near-space communication network, including the following steps: Step 1, design a generation method for a communication network topology graph: Step 1-1, first construct a directed graph called a fully connected graph , where the nodes include all nodes in the communication network, and the nodes include 1 space-based base station, M relay unmanned aerial vehicles, and N ground users; add directed edges E from the base station node to each relay unmanned aerial vehicle, between each relay unmanned aerial vehicle, and from each relay unmanned aerial vehicle to each ground user, where the edges between each relay unmanned aerial vehicle have two directions; traverse all ground users, and for each ground user, use the Depth First Search (DFS) algorithm to search for the multi-hop communication links from the space-based base station via several relay unmanned aerial vehicles to the ground user and abstract the multi-hop communication link into a directed acyclic graph with only a single source and a single sink . The number of nodes and edges included in and respectively.
[0033] Step 1-2: Traverse all the multi-hop communication links of the ground user n, and abstract the link from the i-th node to the j-th node in each multi-hop communication link as an edge . The feature of the edge is the corresponding link capacity . The feature of the i-th node includes the capacity of the edges flowing into it and the capacity of the edges flowing out of it, that is ; According to the line-of-sight transmission model, calculate the channel capacity according to the following formula: , , Select the multi-hop communication link with the largest capacity among them as the actually adopted multi-hop communication link: ; where refers to returning the corresponding ; Take the union of the actually adopted multi-hop communication links of each ground user, and let , and the actual link connection graph can be obtained. The node features and edge features of are spliced by the corresponding node features and edge features.
[0034] Step 1-3: Take the union of all the feasible multi-hop communication links of the ground user to form all the multi-hop communication link graphs of the ground user n, and let , The node features and edge features of are the means of the corresponding node features and edge features, that is , .
[0035] Step 1-4: Take the union of of all ground users to form the communication network topology graph . The node features and edge features of are spliced by the corresponding node features and edge features, that is: , .
[0036] Steps 1 - 5, if is used as the input of the neural network, preprocessing is also required. The nodes and edges are completed according to the fully connected graph , and the value 0 is filled into the corresponding positions of the node features and edge features.
[0037] The entire process of Step 1 is as Figure 1 shown: First, a fully connected graph is constructed from the nodes . Then, all multi - hop communication links of each ground user are traversed, and the multi - hop communication link with the largest capacity of each ground user is selected as the actual multi - hop communication link adopted by each ground user. Then, the union of the actual multi - hop communication links of each ground user is taken to form the actual link connection graph . Each ground user takes the union of all feasible multi - hop communication links to form the multi - hop communication link graph of all ground users n. Then, the union of of all ground users is taken to form the communication network topology graph . Figure 1 In , the blue five - pointed star represents the high - altitude base station, the red dot represents the relay UAV, the green cross represents the ground user, and the black arrow represents the actual link between each node
[0038] Step 2. Convert the trajectory planning process of the stratospheric airship into a sequential decision - making process, and construct a multi - agent Markov decision process (MDP) model to describe the relay UAV trajectory planning task; discretize the time - continuous planning process. In each time slice, each agent (referring to the constructed multi - agent) observes the environmental state, makes a decision, and receives a reward in turn; when MDP starts, each relay UAV starts from an arbitrary position in a rectangular plane area with length A and width C and size , and flies at a fixed altitude; if any relay UAV flies out of the boundary, or when the total number of time slices reaches the upper limit, MDP terminates; according to the task requirements, design the state space, action space, and the reward function regarding the states and of two adjacent time slices as the optimization objective; State space , the communication network topology graph includes node features, edge features, and an adjacency matrix. The flight states of all relay UAVs in a one - dimensional array structure: , where represents the current time, , represents the position coordinates of the m-th relay UAV, , represents the speed of the m-th relay UAV, represents the propulsion power consumption of the m-th relay UAV. To enable the agent to have a global perception of the state of the communication network and thus achieve the overall optimality of the optimization goal, the agent can observe the global state. Therefore, the observation space of the agent is equal to the state space. The action space , is the modulus and direction angle of the speed increment of the m-th relay UAV.
[0039] Reward function , which is divided into global reward and local reward: The global reward acts on all agents and is related to the overall performance of the multi-agent system. , including related to the minimum value of the capacities of each ground user and related to the sum of the capacities of all ground users. The local reward is:
[0040] The local reward acts on each agent separately and is related to the performance of the relay UAV node affected by the agent's decision. Among them, depends on the mean of the smaller value of the incoming and outgoing traffic of the node in the graph of each ground user described in step 1, that is, the minimum value of the node characteristics; is related to the power consumption of the relay UAV; is related to whether the relay UAV goes out of bounds and its distance to the boundary.
[0041] The entire process of step 2 is as shown in Figure 2 . In each time slice, each agent observes the environmental state, makes a decision, and receives a reward in turn. The environment calculates the global reward and local reward according to the actions taken by the agent and inputs them to each agent, and this process repeats.
[0042] Step 3: Design the policy network and value network. The structures of both networks include a graph attention network, a gated recurrent unit network, and a multi-layer perceptron. The role of the policy network is to output the corresponding policy according to the input state; the role of the value network is to output the action value function value of the state-action combination according to the input state-action combination, and evaluate the reward effect of executing this action in this state. Build a virtual simulation environment for the trajectory planning of relay UAVs on a computer to realize the generation of the communication network topology graph, simulate the random movement of the space-based base station and ground users, the interaction between multi-agents and the environment, and the state transition, etc.
[0043] Step 3-1: Read the position information of the space-based base station, relay UAVs, and ground users to generate the communication network topology graph With the flight state of the relay UAV , splice into a state .
[0044] Step 3-2: Input the graph attention network of the policy network, and extract new graph node features through message passing and aggregation. Input the graph attention network of the policy network, and extract new graph node features through message passing and aggregation.
[0045] Step 3-3: Traverse all agents. For each agent, organize the extracted graph node features into a one-dimensional vector, splice it with the flight state of the relay UAV, and then output the decision-making action expressed in the form of the mean and standard deviation of the Gaussian distribution through the gated recurrent unit network and the fully connected layer. The corresponding decision-making action of the agent is sampled from the mean and standard deviation. .
[0046] Step 3-4: Apply it to the corresponding relay UAV to cause the relay UAV to transfer its own state. Update the position information of each node at the next moment, generate a new communication network topology graph, and splice it into a state Act on the corresponding relay UAV to cause the relay UAV to transfer its own state. Update the position information of each node at the next moment, generate a new communication network topology graph, and splice it into a state . Calculate the rewards obtained by each agent according to the reward function .
[0047] The entire process of Step 3 is as Figure 3 , Figure 4 shown. Figure 3 Shows the structure of the policy network, Input the graph attention network of the policy network, extract new graph node features, organize the extracted graph node features into a one-dimensional vector, splice it with the flight state of the relay UAV, and then output the decision-making action through the gated recurrent unit network and the fully connected layer; Figure 4 Shows the structure of the value network, which is somewhat similar to the policy network but also different. Input the graph attention network of the value network, extract new graph node features, organize the extracted graph node features into a one-dimensional vector, splice it with the flight state of the relay UAV and the decision-making action, and then calculate and output the reward through the gated recurrent unit network and the fully connected layer.
[0048] Step 4: In the simulation environment of Step 3, train the agents based on the Soft Actor-Critic (SAC) algorithm to find the optimal policy. Select a multi-agent reinforcement learning architecture with distributed training and distributed execution. Each agent includes an independent Actor network and a Critic network, and the training and execution processes are carried out inside each agent. The SAC algorithm optimizes the Critic network by using the method of gradient descent on the loss function Gradient descent method to optimize the Critic network, and by using the loss function The method of gradient descent is used to optimize the Actor network. and The calculation method is as follows: , ; Step 4-1, initialize the parameters of the agent neural network; Step 4-2, before the start of each round of agent training, first read the position information of the empty base station, relay UAVs, and ground users to generate a communication network topology graph and randomly initialize the flight states of each relay UAV and splice them into the initial state ; Step 4-3, in each time step of each round of training, traverse all agents, and each agent makes corresponding decisions After interacting with the environment, update the positions and flight states of each node to generate a new communication network topology graph The environmental state transfers to and obtain the reward . When any relay UAV crosses the boundary or the cumulative number of time steps reaches the preset upper limit, the MDP process of this round terminates. The interaction data group is stored in the replay area; Step 4-4: Randomly read several pieces of interaction data from the replay area. Traverse all agents. For each agent, according to the neural network update rule of the SAC algorithm, calculate the loss function of each network, and update the network parameters by backpropagation of the neural network in the way of gradient descent.
[0049] The following is a specific application case of the method of the present invention: training in a virtual simulation environment. The maximum number of steps allowed per round is 180, the time interval per step is 1 minute, and the simulation site is a square area. The communication network includes 1 empty base station node, 3 relay nodes, and 10 ground user nodes. Among them, the empty base station node is based on a stratospheric airship platform and flies at a constant altitude of 20 km; the relay nodes are based on UAV platforms and fly at a constant altitude of 2 km; the ground user nodes are all on the ground at an altitude of 0 km. The positions of the empty base station and the ground users will have random displacements to simulate the phenomenon of the airship being disturbed by the stratospheric wind field and the random movement of ground users. The relay UAVs are controlled by agents and satisfy a maximum speed of 25 m / s and a maximum acceleration The kinematic constraints are assumed to have an on-board energy source sufficient to support a 180-minute flight. The communication device has a transmission power of 1 W, a white noise power of -20 dBm, uses directional beam technology, has a signal gain of 20 dB, a noise gain of 9.5 dB, and a maximum transmission distance of 50 km. The data transmission rate per unit bandwidth, i.e., the spectral density, is used to measure the capacity.
[0050] The change curves of each performance index during the training process are as Figure 5 , Figure 6 , Figure 7 shown. Figure 5 is the minimum capacity curve graph during the training process; Figure 6 is the total network capacity curve graph during the training process; Figure 7 is the propulsion power consumption curve graph during the training process; the abscissa represents the number of training rounds, the ordinate represents the reward, and the ordinate value has no unit, and the dimension is 1. As the number of training rounds increases, the capacity of the ground user with the minimum capacity and the total capacity achieved by the entire network both increase, and the propulsion energy consumption of the UAV decreases, effectively improving the network efficiency and taking into account the quality of service of each ground user.
[0051] The trajectory planning algorithm obtained through sufficient training is subjected to simulation testing. Figure 8 , Figure 9 , Figure 10 , Figure 11 , Figure 12 shows the actual link connection diagram of the communication network at different times, which can show the positions of each node and the link connection situation. Among them, the blue five-pointed star represents the space-based base station, the red dot represents the relay UAV, and the green cross represents the ground user. The black arrow represents the actual link between each node , and the red arrow marks the multi-hop communication link corresponding to the ground user with the minimum capacity of the multi-hop communication link. MAXF represents the capacity of this multi-hop communication link, Agent_0 Reward represents the reward obtained by the first agent, Agent_1 Reward represents the reward obtained by the second agent, and Agent_2 Reward represents the reward obtained by the third agent. Actions represents the actions executed by the corresponding agent. The dimensions of Agent_0 Reward, Agent_1 Reward, Agent_2 Reward, Actions, and MAXF are all 1 and have no unit. Figures 8 to 12 The number next to the red dot in [] represents the UAV number, corresponding to its agent number.
[0052] Step 1, as Figure 8 shown: Agent_0 Reward: 28.93, Actions: -0.68, -0.44; Agent_1 Reward: 32.17, Actions: 1.0, 0.97; Agent_2 Reward: 32.78, Actions: -0.97, -1.0; MAXF: 0.01; Step 34, as Figure 9 shown: Agent_0 Reward: 76.64, Actions: -0.62, 0.11; Agent_1 Reward: 52.02, Actions: 1.0, -0.43; Agent_2 Reward: 27.9, Actions: -1.0, -0.98; MAXF: 0.014; Step 115, as Figure 10 shown: Agent_0 Reward: 30.39, Actions: 0.51, 0.37; Agent_1 Reward: 60.95, Actions: 0.99, 0.55; Agent_2 Reward: 34.59, Actions: -1.0, -0.8; MAXF: 0.017; Step 168, as Figure 11 shown: Agent_0 Reward: 43.55, Actions: 0.21, 0.24; Agent_1 Reward: 70.13, Actions: 0.98, -0.55; Agent_2 Reward: 55.38, Actions: -0.99, 0.37; MAXF: 0.021; Step 180, as Figure 12 shown: Agent_0 Reward: 46.04, Actions: -0.43, 0.39; Agent_1 Reward: 77.54, Actions: 0.97, -0.39; Agent_2 Reward: 59.01, Actions: -0.98, 0.0; MAXF: 0.026; In the initial state, the positions of the nodes are randomly distributed. It can be seen that as the simulation time progresses, each agent optimizes the network capacity in a dynamic scenario where the positions of the nodes and the link connection status change by real-time planning the movement trajectories of the relay UAVs.
[0053] Therefore, the present invention adopts an intelligent trajectory planning method for relay UAVs based on the topology graph of the near-space communication network, realizes the joint trajectory planning of each relay UAV node in the wireless communication network, optimizes the service quality of all ground users and the total capacity of the communication network, and optimizes the power consumption-capacity ratio, with high energy utilization efficiency.
[0054] The present invention provides an intelligent trajectory planning method for relay UAVs in a near-space communication network. There are many methods and ways to specifically implement this technical solution. The above is only the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by using the prior art.
Claims
1. A method for intelligent trajectory planning of a near-space communication network relay drone, characterized in that: The following steps are involved: Step 1: Design a method for generating a communication network topology graph: First, construct a directed graph called a fully connected graph. , where the node Including all nodes in the communication network, nodes It includes 1 air-based base station, M relay drones and N ground users. Add directed edges E from the base station node to each relay drone, between each relay drone, and from each relay drone to each ground user. The edges between each relay drone have two directions. Step 2: Convert the trajectory planning process of the stratospheric airship into a sequential decision-making process, and construct a multi-agent Markov decision process model to describe the trajectory planning task of the relay UAV. Discretize the time-continuous planning process. In each time slice, each agent observes the environmental state, makes decisions, and receives rewards in turn. At the beginning of the multi-agent Markov decision process, each relay UAV has a size of length A and width C. The multi-agent Markov decision process terminates if any relay drone flies out of the boundary or the total number of time slices reaches the upper limit. According to the task requirements, the state space, action space, and the state of two adjacent time slices are designed. and The reward function is used as the optimization objective; Step 3: Design the strategy network and value network. The strategy network and value network have the same structure, including graph attention network, gated recurrent unit network and multi-layer perceptron. The role of the strategy network is to output the corresponding strategy according to the input state; the role of the value network is to output the action value function value of the state and action combination according to the input state and action combination, and judge the reward effect of executing the output action in the current state; build a virtual simulation environment for relay drone trajectory planning to realize the generation of communication network topology map, random movement of air-based base stations and ground users, interaction between multiple agents and the environment, and state transfer; Step 4. In the virtual simulation environment of step 3, the agent is trained to find the optimal strategy based on the soft actor-critic algorithm, and a multi-agent reinforcement learning architecture with distributed training and distributed execution is adopted. Each agent includes an independent Actor network and Critic network, where Actor represents the actor and Critic represents the critic. The training and execution process is carried out within each agent.
2. The method according to claim 1, characterized in that Step 1 includes: Step 1-1, traverse all ground users, and for each ground user, use the depth-first search algorithm to search for the route from the airborne base station to the ground user n via the relay drone. multi-hop communication links, and A multi-hop communication link is abstracted as a directed acyclic graph with only a single source and a single sink. , v and e represent specific nodes and edges respectively, ; The number of nodes and edges included are and ; Step 1-2, traverse all multi-hop communication links of ground user n, and convert the multi-hop communication links from the i-th node To the jth node The link is abstracted as an edge ,side The characteristic of the corresponding link capacity is , the i-th node Features Including import The capacity of the edge and outflow The capacity of the edge ,Right now ; According to the line-of-sight transmission model, the channel capacity is calculated according to the following formula: , , in represents the channel bandwidth, Indicates that from the kth node To the jth node The channel gain, yes The intermediate parameters in the calculation formula are: Represents the kth node The spatial coordinates of Represents the kth node The transmission power; when When is the signal power at the receiving end, represents the variance of Gaussian white noise; when When Represents the noise power caused by co-channel interference; From the air-based base station to the ground user No. The capacity of a multi-hop communication link is equal to the minimum value of each link, expressed as: , Select the multi-hop communication link with the largest capacity The multi-hop communication link actually used is: , in Refers to returning the corresponding ; The multi-hop communication links actually used by each ground user are combined, and , get the actual link connection diagram; The node and edge features of The corresponding node features and edge features are spliced together; Steps 1-3, for ground users The multi-hop communication link graphs of the ground user n are combined to form the multi-hop communication link graph of all ground user n. ,make , Node features and edge features for Corresponding node features The mean and edge features of The mean of : , ; Steps 1-4, for all ground users Take and merge to form a communication network topology diagram , ; Node features and edge features Depend on The corresponding node features and edge features are spliced together: , , in represents the i-th node feature of the n-th ground user; Represents the edge feature of the edge from i to j of the nth ground user; Steps 1-5, if you As the input of the neural network, the following preprocessing is performed first: The nodes and edges of the directed graph Fill in the corresponding positions of node features and edge features with the value 0.
3. The method according to claim 2, characterized in that Step 2 includes: State space It is divided into two parts: ,in is the state of the graph, It is the relay drone status, communication network topology diagram Contains node features, edge features, and adjacency matrix.
4. The method according to claim 3, characterized in that Step 2 also includes: The status of all relay drones in a one-dimensional array structure is: , in Indicates the current time. , They represent the horizontal and vertical coordinates of the position of the mth relay drone, , They represent the x-direction speed and y-direction speed of the m-th relay drone, represents the propulsion power consumption of the mth relay UAV.
5. The method according to claim 4, characterized in that Step 2 also includes: the observation space of the agent is equal to the state space, and the action space is equal to ,in represent the magnitude and direction angle of the velocity increment of the mth relay UAV respectively; Reward Function ; in For global rewards, For local rewards: , , in, is the global minimum capacity reward item, which is a positive value and uses the optimal link with all ground users Capacity is proportional to the minimum value of: , ; is the global total capacity bonus, which is a positive value and corresponds to the actual link capacity of all ground users. Capacity in is proportional to: , ; is the local average capacity bonus, which is a positive value and the formula is: , ; It is the UAV propulsion power consumption bonus item, which is a negative value. The formula is: , ; This is a bonus item for the site boundary.
6. The method according to claim 5, characterized in that Step 3 includes: Step 3-1: Read the location information of airborne base stations, relay drones, and ground users to generate a communication network topology map Relay drone flight status , spliced into the current state at time t ; Step 3-2: The communication network topology diagram The graph attention network of the input policy network extracts new graph node features through message passing and aggregation; Step 3-3, traverse all agents, for each agent, organize the extracted graph node features into a one-dimensional vector, splice it with the flight status of the relay drone, pass it through the gated recurrent unit network and the fully connected layer, and output the decision action expressed in the form of the mean and standard deviation of the Gaussian distribution. The decision action corresponding to each agent is obtained by sampling the mean and standard deviation. ; Step 3-4, make the decision action Act on the corresponding relay drone to transfer the state of the relay drone itself; update the location information of each node at the next moment, generate a new communication network topology map, and splice it into the state at the next moment , calculate the reward obtained by each agent according to the reward function .
7. The method according to claim 6, characterized in that Step 4 includes: Step 4-1, initialization of agent neural network parameters; Step 4-2: Before each round of agent training begins, first read the location information of the airborne base station, relay drone, and ground user to generate a communication network topology map. , and randomly initialize the flight status of each relay drone , spliced into the initial state ; Step 4-3: In each time step of each round of training, all agents are traversed and each agent makes a corresponding decision action. , after interacting with the environment, update the location and flight status of each node and generate a new communication network topology , current status Transfer to the next state , and get rewards ; When any relay drone crosses the boundary, or the cumulative time step reaches the preset upper limit, the multi-agent Markov decision process terminates; the interaction data set Save to the playback area, Indicates the current state. Indicates decision-making action. is the decision action of the mth agent, It's a reward. It is the state of the next moment; Step 4-4: Randomly read interaction data from the playback area, traverse all agents, and for each agent, calculate the loss function of the actor-critic algorithm neural network of each agent according to the neural network update rule of the soft actor-critic algorithm, and update the network parameters by back-propagating the neural network in a gradient descent manner.
8. The method according to claim 7, characterized in that 4-4 includes: The soft actor-critic algorithm uses the loss function The gradient descent method is used to optimize the Critic network by adjusting the loss function The gradient descent method is used to optimize the Actor network. and The calculation formula is: , , in, represents the reward obtained in step t; Represents the discount factor, with a value between 0 and 1; is the target value function, and adopts the action value function learning mode of the offline strategy. In the learning mode, there are action value functions with the same structure and the target action value function , The output value of is used to update the action value function. The output value of is used to evaluate the expected reward after taking an action in the state, and is used by the following formula Internal parameters renew Internal parameters : , The update rate is a constant; is the action value function, which is used to evaluate the action in state Make the following move Expected reward after The term represents the policy entropy.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.
10. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 8 are executed.
Citation Information
Patent Citations
Method for optimizing network traffic scheduling based on near-end strategy optimization algorithm
CN115550268A
Distributed multi-unmanned aerial vehicle relay network coverage method
CN116017479A
Unmanned aerial vehicle safety path planning method based on maximum entropy multi-agent reinforcement learning
CN117908565A
Cited By
SDN routing method and architecture based on graph attention mechanism and dynamic priority playback
CN120321167A
Multi-unmanned aerial vehicle cooperative reasoning optimization method based on DAG topology generation and dynamic DNN model partitioning
CN120706632A
Municipal road surface intelligent inspection method and system based on unmanned aerial vehicle array
CN120707360A
A method and system for intelligent inspection of municipal roads based on unmanned aerial vehicle (UAV) arrays
CN120707360B