Implementation method of unmanned aerial vehicle relay communication system based on deep deterministic policy gradient algorithm

By optimizing the UAV relay communication system using the Deep Deterministic Policy Gradient (DDPG) algorithm, the problems of UAV flight trajectory and communication resource allocation were solved, and the maximum throughput of the ground terminal user link and the improvement of communication quality were achieved.

CN114980126BActive Publication Date: 2025-10-17NANJING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210544445.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-18
Publication Date
2025-10-17
Estimated Expiration
2042-05-18

AI Technical Summary

Technical Problem

Existing technologies make it difficult to maximize the throughput of ground terminal users and their links, and to achieve flight trajectory optimization of drones and reasonable allocation of communication resources.

Method used

The Deep Deterministic Policy Gradient (DDPG) algorithm is combined with the Actor-Critic network to optimize the mobility, energy consumption, interference and link scheduling problems in the UAV relay communication system. By constructing a DDPG network for parameter optimization, the flight trajectory of the UAV and the reasonable allocation of communication resources are achieved.

Benefits of technology

It maximizes the throughput of ground terminal user links, optimizes UAV flight trajectories, reduces the waste of communication resources, and improves communication quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114980126B_ABST
    Figure CN114980126B_ABST
Patent Text Reader

Abstract

The implementation method of the unmanned aerial vehicle relay communication system based on the deep deterministic policy gradient algorithm first constructs an unmanned aerial vehicle relay communication system model on a simulation software pycharm according to an application scene; then analyzes constraint problems in the multi-unmanned aerial vehicle relay communication system; then takes the position of a ground terminal user and the position of an unmanned aerial vehicle relay node as a state space, takes the speed, power and link scheduling set of the unmanned aerial vehicle relay node as an action space, and adopts the deep deterministic policy gradient algorithm to calculate an optimization problem; finally, a DDPG network is constructed, parameters are input into the DDPG network to optimize a target function, and the parameters of the DDPG network are acquired. The application can not only maximize the throughput of the ground terminal user and the link thereof, but also can realize flight trajectory optimization of the unmanned aerial vehicle and reasonable allocation of communication resources, while reducing the iteration number of the algorithm and accelerating the convergence process.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of unmanned aerial vehicle communication, and particularly relates to an implementation method of an unmanned aerial vehicle relay communication system based on a deep deterministic policy gradient algorithm. BACKGROUND

[0002] The rapid development of wireless mobile communication technology has promoted the generation of various new business scenarios. From the first generation of mobile communication (1G) to the current popular fifth generation of mobile communication (5G), the rapid development of mobile communication has greatly facilitated people's work and life, and gradually changed the production mode of society. However, the current development of mobile communication technology also faces many challenges, of which the most serious is the mass of terminal users, the differentiation and diversification of network business scenarios. According to the Cisco report, by 2023, there will be 5.3 billion users accessing the network, compared with 3.9 billion network users in 2018, with an annual network user growth rate of 6%. 2020 is considered the year of 5G commercialization, and new industries based on 5G communication technology, such as Internet of Things, Internet of Vehicles, video transmission, etc., have also developed rapidly. In addition, the combination of 5G and artificial intelligence technologies, such as unmanned driving technology, intelligent factories, intelligent logistics, etc., will be deeply integrated with industrial internet technology, further promoting the development of various fields towards intelligent and automated direction.

[0003] At present, although 5G commercialization is still in the process of popularization, the academic circles at home and abroad have already begun to study the potential key technologies of the sixth generation of mobile communication. From the demand for 6G key technologies, 6G not only needs to surpass the 5G standard in transmission rate, capacity and delay, but also requires an integrated network of air, space, land and sea to achieve seamless connection of different communication systems. In the process of 6G standardization evolution, some services for air users have been defined. Therefore, this will greatly promote the research on unmanned aerial vehicle communication technology.

[0004] Compared with traditional ground base station communication and satellite communication, the UAV communication has the following advantages: first, the UAV has the characteristics of high mobility, simple operation and complete controllability, and its dynamic scheduling and deployment are more convenient, so the UAV communication can realize rapid coverage and service diversion of traffic-intensive hot spot areas and reduce communication overhead; second, compared with communication satellites, the UAV is closer to the ground terminal, has short communication round-trip delay and less free space loss; third, the UAV communication system has less dependence on ground infrastructure and low construction cost; fourth, the UAV communication system is less affected by ground buildings and terrain shielding, and is usually a line-of-sight link, so the communication quality is good, and high-speed, high-reliability and low-latency communication can be realized. SUMMARY

[0005] The technical problem to be solved by the present application is to overcome the shortcomings of the prior art, and to provide an implementation method of a UAV relay communication system based on a deep deterministic policy gradient algorithm, which can maximize the throughput of ground terminal users and their links, and realize flight trajectory optimization and reasonable allocation of communication resources of the UAV.

[0006] The present application provides an implementation method of a UAV relay communication system based on a deep deterministic policy gradient algorithm, comprising the following steps,

[0007] Step S1. According to the application scenario, a UAV relay communication system model is constructed on the simulation software pycharm, including the representation of ground base stations, UAV relay nodes and ground terminal users;

[0008] Step S2. Analyze the constraint problem in the multi-UAV relay communication system, including the mobility problem and energy consumption problem of the UAV, the interference and link scheduling problem in the full-duplex mode, and the information rate problem, and convert the physical model into a mathematical optimization problem;

[0009] Step S3. Taking the position of the ground terminal user and the position of the UAV relay node as the state space, and taking the set of speed, power and link scheduling of the UAV relay node as the action space, the deep deterministic policy gradient algorithm is used to calculate the optimization problem;

[0010] Step S4. Constructing a DDPG network, inputting the above parameters into the DDPG network to optimize the objective function, and obtaining the parameters of the DDPG network.

[0011] As a further technical solution of the present application, in step S2, the mobility constraint formula of the UAV is

[0012]

[0013]

[0014] The energy consumption constraint formula of the UAV is:

[0015]

[0016] E trans [n] = p uav [n]·Δt, (4)

[0017] E[n]=E[n-1]-E fly [n]-E trans [n]; (5)

[0018] in, is the position coordinate of the UAV relay node m in time slot n, is the velocity vector of the UAV relay node m in time slot n, Δt is the time interval, D min is the minimum distance that should be satisfied between two UAV nodes, p uav [n] is the transmission power of the drone relay node, and m is the mass of the drone. Formula (1) represents the speed and position constraints of the drone in two adjacent time slots, formula (2) represents the minimum distance constraint that should be met between different drones, and formula (3) represents the flight energy consumption E of the drone in time slot n. fly [n], formula (4) represents the communication energy consumption E of the UAV in time slot n trans [n], formula (5) represents the total energy E[n] left by the drone at the end of time slot n.

[0019] Furthermore, in step S2, the constraint formula for interference and link scheduling between drones in full-duplex mode is:

[0020]

[0021]

[0022]

[0023] Where, formula (6) is the reachable link capacity between UAV i and UAV j, W is the bandwidth, is the transmission power of UAV i, is the path gain between UAV i and UAV j, η is the Gaussian noise power spectrum density, θ is the self-interference elimination coefficient, the smaller the θ value is, the smaller the self-interference of the UAV relay node is; Formula (7) and Formula (8) are the scheduling conditions of the UAV ground terminal i and UAV node j or UAV relay node i to UAV relay node j in session l in time slot n. Represents the scheduling of session l between ground terminal i and drone node j or drone relay node i to drone relay node j in time slot n.

[0024] Further, in step S3, the information rate constraint formula of the unmanned aerial vehicle relay node adopting the decode-and-forward relay strategy is

[0025]

[0026]

[0027]

[0028]

[0029] wherein, formula (9)-(11) are that in the multi-hop session, the information rate of the next hop is not greater than the information rate of the previous hop, is the reachable rate of the first hop of session l in time slot n, that is, the reachable rate of the source node to the first unmanned aerial vehicle relay node, is the reachable rate of the mth unmanned aerial vehicle relay node of session l in time slot n, is the reachable rate of the destination node of session l in time slot n; formula (12) is the end-to-end throughput of the source node to the destination node during session l.

[0030] Further, in step S3, the specific steps of the deep deterministic policy gradient algorithm are,

[0031] Step S31. According to the environment, the current state is initialized as s, and the feature vector is φ(s);

[0032] Step S32. In the current network of the actor, the action is selected according to the policy function a=πθ(φ(s))+N;

[0033] Step S33. In state s, the action a is executed, the next step state s' and the reward r are obtained, and whether the state isEnd is terminated;

[0034] Step S34. The four-tuple composed of {φ(s), a, r, φ(s')} is put into the experience replay pool D;

[0035] Step S35. State transition: s=s';

[0036] Step S36. Randomly sample m unrelated samples {φ(s j ),a j ,r j ,φ(s' j )} from the experience replay pool D, and calculate the target Q value y j : y j =r j +γQ'(φ(s' j ),π θ' (φ(s'j )),ω');

[0037] Step S37. Calculate the mean square error loss function Update the parameters ω of the Critic current network by gradient back propagation of the neural network;

[0038] Step S38. Calculate Update all parameters θ of the Actor current network by gradient back propagation of the neural network;

[0039] Step S39. Whether the algorithm meets the termination condition, if yes, end the iteration, otherwise go to step S32 to re-learn.

[0040] The advantages of the present application are that,

[0041] 1. By analyzing the mobility, energy consumption, interference, link scheduling and information rate constraint problems of the unmanned aerial vehicle node, the physical model is converted into a mathematical optimization problem, and the deep deterministic policy gradient algorithm is used to solve the optimization problem, so that the throughput of the ground terminal user and the link thereof can be maximized, and the flight trajectory of the unmanned aerial vehicle can be optimized.

[0042] 2. Under the premise of satisfying the maximum throughput of the session, the DDPG algorithm is optimized to realize the communication between the remote terminal users in a multi-hop manner, and the node selection of the multi-hop unmanned aerial vehicle and the reasonable allocation of the communication resources are realized.

[0043] 3. The deep deterministic policy gradient algorithm is adopted to fuse the Actor-Critic network and the deep learning network, so that the limitations of the Q-learning and DQN algorithms in the high-dimensional continuous state space can be broken through, the number of algorithm iterations can be reduced, and the convergence process can be accelerated. DETAILED DESCRIPTION

[0044] Figure 1 It is a scene diagram of the multi-hop unmanned aerial vehicle relay communication system of the present application;

[0045] Figure 2 It is a network architecture diagram of the DDPG algorithm of the present application;

[0046] Figure 3 It is a simulation scene diagram constructed according to the scene diagram in the pycharm software of the present application;

[0047] Figure 4 It is a change trend diagram of the system session average throughput optimized based on the DDPG algorithm of the present application;

[0048] Figure 5 It is an optimal running trajectory diagram of the unmanned aerial vehicle in session 1;

[0049] Figure 6 For the contrast chart of the trajectory of the unmanned aerial vehicle in session 1;

[0050] Figure 7 For the trajectory running chart of the unmanned aerial vehicle and the terminal user transmitting party in session 2;

[0051] Figure 8 For the trajectory running chart of the unmanned aerial vehicle and the terminal user receiving party in session 2;

[0052] Figure 9 For the time slot allocation chart of the terminal user in session 1;

[0053] Figure 10 For the schematic diagram of the multi-hop unmanned aerial vehicle relay in session 2;

[0054] Figure 11 For the unmanned aerial vehicle power control simulation chart when the exploration rate is 0.1;

[0055] Figure 12 For the unmanned aerial vehicle power control simulation chart when the exploration rate is 0.05;

[0056] Figure 13 For the throughput change trend simulation chart under different algorithms. DETAILED DESCRIPTION

[0057] The embodiment adopts pycharm software as a simulation platform, the programming language is python, and the TensorFlow framework is used to simulate the physical model. The problem of optimizing the trajectory of the unmanned aerial vehicle and allocating the communication resources in the unmanned aerial vehicle relay communication system is solved by using the reinforcement learning algorithm. The throughput of the ground terminal user and the link thereof is maximized. Moreover, the flight trajectory of the unmanned aerial vehicle is optimized and the communication resources are reasonably allocated. The communication quality of the terminal user is effectively ensured.

[0058] The embodiment adopts pycharm software as a simulation platform to construct and verify the trajectory optimization and the reasonable allocation of the communication resources of the unmanned aerial vehicle relay communication system based on the reinforcement learning. Please refer to Figure 1 and Figure 2 , the specific implementation includes the following,

[0059] (I) Construction of the unmanned aerial vehicle relay communication system

[0060] The embodiment constructs a physical model of a UAV relay communication system on a simulation platform according to an actual application scenario of UAV relay communication, which includes a ground base station, a UAV relay node, a ground terminal user, and an obstacle such as a high-rise building. According to actual conditions, the terminal user moves randomly on the ground and the position information is known, and there are two ways of communication between terminal users: end-to-end direct communication and UAV relay communication. When the distance between terminals is short and the channel condition is good, the terminal preferentially selects end-to-end direct communication. When the distance between terminal users is far or there is an obstacle without a direct path, the terminal can only communicate through the UAV relay node. In addition, the ground terminal device can also transmit information with the base station, and when the channel condition is poor, the terminal device preferentially selects to communicate with the base station through the UAV relay node. It is assumed that the system has L groups of sessions, which can be represented as The source node s(l) (l∈L) and the destination node d(l) (l∈L) of each group of sessions l cannot communicate end-to-end, and can only transmit data through the UAV relay node in a multi-hop manner.

[0061] (ii) Description of the UAV relay communication system model

[0062] For the UAV relay communication system model constructed in (i), the embodiment analyzes the mobility problem of the UAV, the energy consumption problem, the interference and link scheduling problem in the full-duplex mode, and the information rate problem.

[0063] (1) Mobility problem

[0064] In the embodiment, the entire data transmission period T of the system is divided into N equal time slots, and the length of each time slot is represented by Δt, i.e. Δt=T / N. It is assumed that the state of the UAV relay node and the ground terminal in the system does not change during the same time slot, and the coordinates of the source node s of session l in time slot n are represented as:

[0065]

[0066] The coordinates of the destination node d of session l in time slot n are represented as:

[0067]

[0068] The position coordinates of the UAV relay node m in time slot n can be represented as:

[0069]

[0070] The position coordinates of the UAV node m in the next time slot n+1 are represented as:

[0071]

[0072] then and The following conditions should be met:

[0073]

[0074]

[0075] wherein, denotes the velocity vector of the UAV relay node m at time slot n, D min denotes the minimum distance that should be met between two UAV nodes.

[0076] (2) Energy consumption problem

[0077] In this embodiment, the UAV is in flight during data transmission, and needs to reach the designated location before the energy is exhausted, so the energy consumption problem of the UAV relay node needs to be considered. Therefore, the energy consumption of the UAV during the entire flight process is mainly composed of two parts: energy consumption generated by communication and energy consumption generated by flight. The total energy of the UAV before starting flight is E, and E[n] represents the energy remaining in the UAV after flying the nth time slot.

[0078] E fly [n] represents the energy consumption generated by flight of the UAV in the nth time slot, and the flight speed of the UAV in the nth time slot is Therefore, the following relationship is obtained.

[0079]

[0080] wherein, m represents the mass of the UAV.

[0081] Then, E trans [n] represents the energy consumption generated by communication of the UAV in the nth time slot. The power allocation of the UAV in the nth time slot is represented as: Therefore, we have:

[0082] E trans [n] = p uav [n] · Δt, (4)

[0083] E[n] = E[n-1] - E fly [n] - E trans [n]; (5)

[0084] After the UAV reaches the destination (runs the last time slot), there should be E[N]≥0.

[0085] (3) Interference and link scheduling problem

[0086] In this embodiment, the UAV relay adopts full-duplex mode for information transmission. The UAV relay uses The scheduling of session l between the ground terminal i and the UAV node j or the UAV relay node i to the UAV relay node j in time slot n is represented. If there is data to be transmitted between the node i and the node j in time slot n, then Otherwise, There are the following constraints:

[0087]

[0088]

[0089] In this embodiment, let be the achievable link capacity of session l from the UAV relay node i to the relay node j in time slot t. In the full-duplex mode, the self-interference of the UAV relay node is not negligible, so the interference received by the UAV relay node j is composed of the mutual interference generated by other relay nodes in the system and the self-interference from the node j. The achievable link capacity between the node i and the node j can be calculated by the Shannon formula as shown in formula (8).

[0090]

[0091] where the first part represents the interference generated by other UAV relay nodes in the system to the node j (mutual interference), the second part represents the self-interference generated by the UAV relay node j, and the third part represents the noise power.

[0092] (4) Information rate constraint problem

[0093] In this embodiment, the UAV relay node adopts the Decode-and-Forward (DF) relay strategy, and the following constraints are considered without considering the delay:

[0094]

[0095]

[0096]

[0097]

[0098] where formulas (9)-(11) represent that in a multi-hop session, the information rate of the next hop should not be greater than that of the previous hop, represents the achievable rate of the first hop (from the source node to the first UAV relay node) of session l in time slot n, Rl(n, m) denotes the reachable rate of session l in the m-th UAV relay node in time slot n, Rl(n, d) denotes the reachable rate of session l in the destination node in time slot n; formula (12) represents the throughput of the end-to-end source node to the destination node during session l.

[0099] (III) Optimization of UAV relay communication system based on reinforcement learning

[0100] The embodiment uses an improved deep deterministic policy gradient algorithm to solve the optimization problem, and realizes the maximum throughput of the system session. First, the agent, state space, action space and reward mode of the model are defined. In this embodiment, the agent is a set of UAV relay nodes, the state space is composed of the positions of the ground terminal users and the positions of the UAV relay nodes, and is represented as:

[0101]

[0102] The action space is defined as the set of the speed of the UAV relay node, the power of the UAV relay node and the link scheduling, and is represented as:

[0103]

[0104] In this embodiment, the reward function is designed to consider two aspects: maximizing the throughput of the session within limited resources and reaching the destination before the fuel is exhausted. Therefore, the total reward function can be designed as:

[0105] r n =r(s n ,a n )=(1-κ end )(r c +r loc ),

[0106] where κ end is a binary variable representing whether the UAV is out of fuel. κ end = 1 indicates that the UAV is out of fuel, and the reward is 0, otherwise, the UAV is in normal state. r c represents the throughput of the system session, and r loc represents the reward brought by the change of the UAV position in different states.

[0107] The DDPG algorithm is divided into a training phase and an implementation phase. In each training, the UAV starts from the starting position with sufficient energy and ends when the energy is exhausted or the destination is reached. In the training phase, the specific implementation steps are as follows:

[0108] (1) The agent initializes the current state as s and the feature vector as φ(s) according to the environment;

[0109] (2) In the current network of the Actor, select the action a according to the policy function a = πθ (φ (s) ) + N;

[0110] (3) In the state s, perform the action a, obtain the next state s' and the reward r, and determine whether the state isEnd;

[0111] (4) Put the four-tuple {φ (s), a, r, φ (s')} into the experience replay pool D;

[0112] (5) State transition: s = s';

[0113] (6) Randomly sample m unrelated samples {φ (s j ), a j , r j , φ (s' j )} from the experience replay pool D, and calculate the target Q value y j :

[0114] y j = r j + γQ' (φ (s' j ), π θ' (φ (s' j )), ω') ;

[0115] (7) Calculate the mean square error loss function Update the parameters ω of the current network of the Critic using gradient backpropagation of the neural network;

[0116] (8) Calculate Update all parameters θ of the current network of the Actor using gradient backpropagation of the neural network;

[0117] (9) Whether the algorithm meets the termination condition, if yes, end the iteration, otherwise go back to step b) to learn again.

[0118] In the implementation phase, the unmanned vehicle will take appropriate action according to the current state through the trained Actor network.

[0119] The experimental verification is as follows:

[0120] (1) Experimental parameter setting, as shown in Table 1

[0121]

[0122] Table 1. Simulation parameter setting

[0123] (2) Experimental environment setting

[0124] In the present application, we simulate on pycharm software according to the actual application scene, and the simulation scene diagram is as followsFigure 3 The system is assumed to be composed of 20 UAV relay nodes, and the ground terminal users achieve communication with the ground base station or other users through multi-hop UAV relay nodes. Two groups of sessions are constructed in the relay system, session 1: four terminal users communicate with the BS through the UAV relay node 18, and the UAV relay node 18 operates in a square region with a side length of 2 km and a center at coordinates [-0.21, -14.25, 0.5]; session 2: one terminal communicates with another terminal through multi-hop UAV relay nodes. The ground base station is located at the coordinate origin [0, 0, 0.05], the starting position of the UAV relay node 18 is [-1.21, -14.25, 0.5], the terminal coordinates are [0.79, -14.25, 0.5], the flight height of the UAV is fixed at 500 m, and the ground terminals within the coverage of the UAV are randomly distributed in a square region of 1 km*1 km, and the ground terminals are in a random motion state.

[0125] (2) Experimental results verification

[0126] Figure 4 The average throughput trend of the session in the UAV relay communication system in the application is shown, and it can be seen that, after optimization by the DDPG algorithm, the throughput of the system session is obviously improved.

[0127] Figure 5 The optimal running trajectory graph of the UAV in session 1 is shown, Figure 6 The running trajectory comparison graph of the UAV in session 1 under different iteration numbers is shown, and with the continuous increase of the iteration number and the continuous updating of the DDPG network parameters, the learning behavior of the UAV is gradually optimized. When the iteration number is 8000, the DDPG network parameters tend to be stable, and the running trajectory of the UAV starts to be smooth. Figure 7 、 8 The trajectory running graphs of the UAV and the terminal user transmitting party and the terminal user receiving party in session 2 are shown respectively. In the process of learning and optimization of the agent, in order to maximize the throughput of the session, the UAV will run towards the user direction, at this time the distance between the UAV and the terminal user becomes close, and the communication rate will be continuously improved. At the end of the session period, the UAV will fly towards the set terminal point.

[0128] Figure 9 The terminal user time slot allocation graph is shown, and the number of available communication time slots for each terminal is evenly allocated, and the communication resources can be reasonably utilized. Figure 10 The UAV relay node routing graph in session 2 is shown, and through the multi-hop UAV relay node, two remote terminals can achieve high-rate communication. Figure 11 The running trajectory comparison graph of the UAV in session 1 under different iteration numbers is shown, and with the continuous increase of the iteration number and the continuous updating of the DDPG network parameters, the learning behavior of the UAV is gradually optimized. When the iteration number is 8000, the DDPG network parameters tend to be stable, and the running trajectory of the UAV starts to be smooth. Figure 10The total power consumption diagram of the UAV relay node in the case of session routing and scheduling. The power control of the UAV relay node is optimized under the premise of maximizing the communication rate between terminals. Through continuous learning iteration, the power consumption of the whole system is significantly reduced.

[0129] Figure 12 And Figure 13 The power consumption trend diagram of the UAV relay node when the exploration rate is 0.05 is shown. Compared with the exploration rate of 0.05, when the exploration rate is 0.1, the DDPG network has better convergence and the power control effect is better. By comparison, when the exploration rate is 0.05, the power consumption is much more than the DDPG network when the exploration rate is 0.1, and the DDPG algorithm falls into local optimum. When the exploration rate is 0.1, the average power consumption of the UAV relay node is 3.25W, when the exploration rate is 0.05, the average power consumption of the UAV relay node is 4.97W, and when the maximum power is used for transmission, the average power consumption of the UAV relay node is 20W. By comparison, it can be found that the power consumption of the session when the exploration rate is 0.1 is reduced by 34.6% compared with the exploration rate of 0.05, and the power consumption of the session when the exploration rate is 0.1 is reduced by 83.75% compared with the maximum power transmission.

[0130] (3) Summary of experimental results

[0131] This embodiment constructs a UAV relay communication system model on the simulation software pycharm according to the actual application scene, optimizes the flight trajectory, node selection and communication resources of the UAV relay node jointly, solves the problem by using the improved deep deterministic policy gradient algorithm, can realize the maximum throughput of the ground terminal user and its link, and can realize the flight trajectory optimization of the UAV and the reasonable allocation of the communication resources, can reduce the number of algorithm iterations, and can speed up the convergence process.

[0132] The above shows and describes the basic principles, main features and advantages of the embodiment. Those skilled in the art should understand that the embodiment is not limited by the above specific embodiments, and the above specific embodiments and the description in the specification are only to further illustrate the principles of the embodiment, and various changes and improvements can be made to the embodiment without departing from the spirit of the embodiment, and these changes and improvements all fall within the scope of the claimed embodiment. The scope of protection of the embodiment is defined by the claims and their equivalents.

Claims

1. The implementation method of the UAV relay communication system based on the deep deterministic policy gradient algorithm is characterized by: The following steps are included: Step S1. Build a UAV relay communication system model on the simulation software PyCharm according to the application scenario, including the representation of the ground base station, UAV relay node, and ground terminal users; Step S2. Analyze the constraints in the multi-UAV relay communication system, including UAV mobility and energy consumption, interference and link scheduling in full-duplex mode, and information rate, and transform the physical model into a mathematical optimization problem. Step S3. Using the location of the ground terminal user and the location of the UAV relay node as the state space, and the set of the UAV relay node's speed, power, and link scheduling as the action space, a deep deterministic policy gradient algorithm is used to calculate the optimization problem; Step S4. Build a DDPG network, input the above parameters into the DDPG network to optimize the objective function, and obtain the parameters of the DDPG network; In step S2, the mobility constraint formula of the UAV is: The energy consumption constraint formula of the UAV is: E trans [n]=p uav [n]·Δt,(4) E[n]=E[n-1]-E fly [n]-E trans [n];(5) in, is the position coordinate of the UAV relay node m in time slot n, is the velocity vector of the UAV relay node m in time slot n, Δt is the time interval, D min is the minimum distance that should be satisfied between two UAV nodes, p uav [n] is the transmission power of the drone relay node, and m is the mass of the drone. Formula (1) represents the speed and position constraints of the drone in two adjacent time slots, formula (2) represents the minimum distance constraint that should be met between different drones, and formula (3) represents the flight energy consumption E of the drone in time slot n. fly [n], formula (4) represents the communication energy consumption E of the UAV in time slot n trans [n], formula (5) represents the total energy E[n] left by the UAV at the end of time slot n; In step S2, the constraint formula for interference and link scheduling between drones in full-duplex mode is: Where, formula (6) is the reachable link capacity between UAV i and UAV j, W is the bandwidth, is the transmission power of UAV i, is the path gain between UAV i and UAV j, η is the Gaussian noise power spectrum density, θ is the self-interference elimination coefficient, the smaller the θ value is, the smaller the self-interference of the UAV relay node is; Formula (7) and Formula (8) are the scheduling conditions of the UAV ground terminal i and UAV node j or UAV relay node i to UAV relay node j in session l in time slot n. represents the scheduling of session l between ground terminal i and UAV node j or UAV relay node i to UAV relay node j in time slot n; In step S3, when the UAV relay node adopts the decoding and forwarding relay strategy, the information rate constraint formula is: Among them, formulas (9)-(11) mean that in a multi-hop session, the information rate of the next hop is not greater than the information rate of the previous hop. is the reachable rate of the first hop of session l in time slot n, that is, the reachable rate from the source node to the first UAV relay node, is the achievable rate of session l at the mth hop UAV relay node in time slot n, is the achievable rate of the destination node for session l in time slot n; formula (12) is the end-to-end throughput from the source node to the destination node during session l.

2. The method for implementing a UAV relay communication system based on a deep deterministic policy gradient algorithm according to claim 1, characterized in that: In step S3, the specific steps of the deep deterministic policy gradient algorithm are: Step S31. Initialize the current state to s and the feature vector to φ(s) according to the environment; Step S32. In the Actor current network, select an action according to the strategy function a = πθ(φ(s)) + N; Step S33. In state s, perform action a, obtain the next state s′ and reward r, and determine whether to terminate the state isEnd; Step S34: Put the four-tuple consisting of {φ(s), a, r, φ(s')} into the experience replay pool D; Step S35. Perform state transfer: s=s'; Step S36. Randomly sample m unrelated samples {φ(s j ),a j ,r j ,φ(s' j )}, calculate the target Q value y j :y j =r j +γQ'(φ(s' j ),π θ' (φ(s' j )),ω'); Step S37. Calculate the mean square error loss function Use the gradient back propagation of the neural network to update the parameters ω of the Critic current network; Step S38. Calculation Use the gradient back propagation of the neural network to update all parameters θ of the Actor's current network; Step S39: Check whether the algorithm meets the termination condition. If so, end the iteration; otherwise, go to step S32 to relearn.

Citation Information

Patent Citations

  • IRS-assisted unmanned aerial vehicle communication joint optimization method based on DDPG algorithm

    CN113162679A