Double-delay depth deterministic strategy gradient unmanned aerial vehicle relay trajectory optimization method

By optimizing the trajectory of fixed-wing UAVs using the TD3 algorithm and MDP model, the communication performance problem in a mixed probabilistic channel environment was solved, achieving efficient communication under energy and security constraints.

CN121619637APending Publication Date: 2026-03-06NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511926401.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Existing technologies in UAV relay systems have failed to effectively address the trajectory optimization problem of fixed-wing UAVs in complex environments, especially in mixed probabilistic channel environments. They cannot balance onboard energy constraints and flight safety constraints, resulting in poor communication performance.

Method used

The dual-delay deep deterministic policy gradient (TD3) algorithm, combined with the Markov decision process (MDP) model, is used to optimize the flight speed, elliptical trajectory parameters, and half-duplex communication time slot length of a fixed-wing UAV. A unified optimization framework is constructed to ensure that the UAV flies along a reasonable trajectory to maximize the channel capacity of the ground target node.

Benefits of technology

It significantly improves communication performance in mixed probabilistic channel environments, takes into account the energy constraints and flight safety of UAVs, and increases the average cumulative channel capacity of ground target nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121619637A_ABST
    Figure CN121619637A_ABST
Patent Text Reader

Abstract

The invention provides a double-delay depth deterministic strategy gradient unmanned aerial vehicle relay trajectory optimization method. The method comprises the following steps: step 1, establishing a three-node relay communication system; 2, calculating power consumption and energy consumption of the unmanned aerial vehicle; step 3, establishing an unmanned aerial vehicle relay trajectory optimization problem; and step 4, modeling the relay trajectory optimization problem of the unmanned aerial vehicle as a Markov decision process. According to the method, the flight speed of the unmanned aerial vehicle, the shape parameter of the elliptical trajectory and the time slot length of half-duplex communication are designed through a double-delay depth deterministic strategy gradient, so that the unmanned aerial vehicle flies along a reasonable closed elliptical trajectory, and relay service is provided for a ground source and a destination node in the process; and the average accumulated channel capacity of the ground destination node is maximized. A simulation experiment also shows that the system performance (the average accumulated channel capacity of a ground destination node) provided by the optimization method is obviously higher than that provided by a circular trajectory condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication technology, and more specifically to a dual-delay depth deterministic strategy gradient UAV relay trajectory optimization method. Background Technology

[0002] With the development of fifth-generation (5G) communication technology, unmanned aerial vehicles (UAVs), as a flexible aerial platform, have shown great potential in the field of auxiliary communication, especially in scenarios such as emergency communication, IoT data backhaul, and enhanced coverage in hotspot areas. UAVs acting as mobile relays can effectively solve the problem of poor communication link quality caused by terrain obstruction or long distances.

[0003] However, the performance of UAV relay systems is highly dependent on their flight trajectories. In complex urban or suburban environments, wireless channels often exhibit "mixed probability" characteristics, meaning that line-of-sight (LoS) links and non-line-of-sight (NLoS) links occur randomly, and their probability of occurrence is closely related to the UAV's position, altitude, and environment. This randomness of the channel presents a significant challenge to trajectory optimization. Furthermore, UAVs have limited onboard energy; when performing auxiliary communication tasks, they must complete the required tasks before their limited onboard energy is exhausted. This necessitates that the designed UAV trajectory be energy-efficient.

[0004] Most existing studies assume a deterministic channel model, failing to fully capture the randomness of the real environment. Furthermore, most optimization methods are computationally offline, making them ill-suited to real-time changes in channel conditions. Deep Reinforcement Learning (DRL) offers a novel approach to solving such sequential decision-making problems. Its "perception-learning-decision" paradigm is well-suited for handling high-dimensional, continuous state and action spaces and can adapt to environmental uncertainties. The Deep Deterministic Policy Gradient (DDPG) algorithm, a representative DRL algorithm for handling continuous action spaces, has been applied to UAV trajectory optimization. However, it suffers from overestimating action values. The Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm effectively improves upon DDPG's shortcomings by introducing a dual evaluation network, delayed updates, and target policy smoothing mechanisms, exhibiting better stability and convergence performance. Therefore, how to utilize the TD3 algorithm for trajectory optimization of fixed-wing UAV relays, maximizing relay service performance under limited onboard energy, is a pressing issue.

[0005] Furthermore, unlike rotary-wing UAVs, fixed-wing UAVs cannot hover or instantly change their flight direction; they can only continuously fly forward to obtain the necessary lift. To change course, the fuselage of a fixed-wing UAV needs to tilt (i.e., generate a roll angle). To ensure flight safety, fixed-wing UAVs must adhere to a maximum roll angle constraint during flight. Therefore, the roll angle constraint of a fixed-wing UAV limits its flight maneuverability. This safety constraint directly affects trajectory design and energy consumption; ignoring this constraint may prevent the designed flight trajectory from being practically implemented in engineering, or even create flight safety hazards. Most existing work treats fixed-wing UAVs as "ideal turning points," which is not consistent with reality. Therefore, how to balance energy constraints and flight safety constraints (primarily referring to roll angle constraints) to optimize the trajectory of fixed-wing UAV relays and maximize relay service capabilities is another urgent problem to be solved. Summary of the Invention

[0006] Purpose of the invention: This invention addresses the problem of overly simplistic solutions in existing technologies by providing a dual-delay depth deterministic strategy gradient UAV relay trajectory optimization method, comprising the following steps: Step 1: Establish a three-node relay communication system; Step 2: Calculate the drone's power consumption and energy consumption; Step 3: Establish the UAV relay trajectory optimization problem; Step 4: Model the UAV relay trajectory optimization problem as a Markov decision process.

[0007] Step 1 includes: the three-node relay communication system includes a fixed ground source node. A fixed ground destination node A mobile relay for an aerial drone; set up arrive The direction of the connecting line is axis; The axis rotates counterclockwise horizontally around the zero point. for Axis; perpendicular to shaft and The horizontal plane formed by the axes and the direction passing through the zero point is axis; The drone relay operates in a half-duplex amplification and forwarding mode. In the first time slot, the drone mobile relay R receives the signal transmitted by N1, and then amplifies the received signal from N1 before forwarding it to N2 in the second time slot.

[0008] In step 1, the drone's flight altitude is set to be constant. That is, in The coordinates on the axis are fixed as The flight path is constrained to an elliptical path; set up Located in a two-dimensional horizontal plane Place; Located in a two-dimensional horizontal plane Among them, for and The distance between them; The antenna height is ; The antenna height is The antenna height for drone relay is... .

[0009] In step 1, the coordinates of the center of the elliptical trajectory on the two-dimensional horizontal plane are set as follows: elliptical trajectory in The half-shaft on the shaft is ,and ,exist The half-shaft on the shaft is ,and The horizontal coordinates of the drone relay origin position are: The eccentricity of the elliptical trajectory is .

[0010] In step 1, the drone relay starts from the initial position at a constant speed. If it flies clockwise along an elliptical trajectory, then in At what moment, the angle of the drone's flight This considers the case where the flight around an elliptical trajectory does not exceed one revolution. Let be the perimeter of the elliptical trajectory, where Let be the eccentricity of the elliptical trajectory. Let be the parameter angle of the elliptical trajectory. Indicates the parameter angle The derivative of is used as the integration variable to integrate over the arc length of the ellipse to calculate the ellipse's circumference; in At what moment, the position of UAV relay R on the two-dimensional horizontal plane Given by equations (1) and (2): (1), (2), The drone travels along an elliptical trajectory at a constant speed. The flight takes place, and the time for one lap is... The location of the drone exhibits periodicity, with a period of... When determining the location of a drone relay, it is necessary to first determine the drone's flight time. according to Find the modulus, and substitute the remainder of the modulus operation into equations (1) and (2) to calculate the specific position on the two-dimensional horizontal plane; according to Calculate the antenna of the UAV relay R in Time and The distance of the antenna And the antenna of the drone relay R in Time and The distance of the antenna : (3), (4), When a drone flies horizontally at a fixed altitude, the drag force D it experiences is: (5), in, Zero lift-drag coefficient, air density, For the wing area of ​​the drone, For the flight speed of the drone, For the wingspan efficiency of drones, It is the wingspan of the drone. For the overall quality of the drone, It is the acceleration due to gravity. For drones The roll angle at any moment; the drone in Roll angle of time for ,in, for The radius of curvature of the elliptical trajectory of the UAV at any given time is given by equation (6): (6), in, For an elliptical trajectory in Half-shaft on the shaft, For an elliptical trajectory in Half-shaft on the shaft; according to and Then equation (5) can be rewritten as: (7).

[0011] Step 2 includes: the power consumption of the drone. Calculated from equation (8): (8), The drone relay starts at time 0, from the horizontal starting position. Starting from a point and flying along an elliptical trajectory, then... Energy consumption of drone relay for: (9), in, For the integral variable representing time, Represents a tiny increment in time; Indicates the drone at a certain time Instantaneous power consumption when flying along an elliptical trajectory.

[0012] Step 3 includes: In half-duplex amplification and repeater mode, any node in the system cannot receive a signal while transmitting it; any node in the system cannot transmit a signal while receiving it; due to the ground source node... With ground target node The direct link does not exist. The signal was transmitted to the air via a drone relay. Two time slots are required; a total of [number] slots are set. Each time slot If it is even, the first The first time slot and the first The lengths of the time slots are equal, for , It is an odd number. , satisfy ,in, and These are the minimum and maximum timeslot lengths, respectively. In the In each time slot, the ground source node transmits a unit power signal. The signal received by the drone relay R Written as: (10) in, for The transmission power, yes Large-scale channel gain between R and The Gaussian white noise received by R has an average power of , Written as: (11), in, It is the power per unit distance of the channel. To accumulate to the number The time of each time slot yes And R's antenna in The distance of time, It is the channel fading index. yes With R The probability of a line-of-sight link occurring at any given time. These are additional channel fading parameters for non-line-of-sight links. Calculated from equation (12): (12) in, and The terrain parameters that the drone flies over. yes The elevation angle between R and R; In the In the time slot, the drone relay will be in the first... Signals received from the ground source node in each time slot Scaled to unity power signal Specifically, as shown in equation (13): (13) Then, with power Will Launch to In the At the end of each time slot, Received signal for: (14) in, R is the transmit power. It is R and In the Large-scale channel gain per time slot yes Received white Gaussian noise, large-scale channel gain Written as: (15) in, To accumulate to the number The time of each time slot It is R and The antenna in The distance of time, Is R and exist The probability of a Loss channel appearing at any given time. Calculated from equation (16): (16) in, It is R and The angle of elevation between them; In the At the end of the time slot, the ground target node Received signal Rewritten as: (17) In the At the end of the time slot, the ground target node The signal-to-noise ratio (SNR) is the ratio of the received useful signal to the noise power. for: (18) Through the first and the Each time slot, destination node The obtained channel capacity is: ; Through all Transmission in time slots, destination node Cumulative channel capacity obtained with the assistance of aerial drones for: (19) First find Mean: (20), in This indicates calculating the mean. Through all Transmission in time slots, destination node Average cumulative channel capacity obtained with the assistance of aerial drone relay for: (twenty one), In the case of an elliptical flight trajectory, to maximize the target node If the average cumulative channel capacity is the objective, then the UAV relay trajectory optimization problem can be written as: (22a) (22b) (22c) (22d) (22e) (22f) (22g) Equation (22a) gives the objective of maximizing the average cumulative channel capacity of the destination node, where, , , and The speed of the drone The optimal solution, the elliptical trajectory in Half shaft on the shaft The optimal solution, the elliptical trajectory in Half shaft on the shaft The optimal solution and the first Length of each time slot The optimal solution; Equation (22b) is the energy constraint of the UAV, that is, the UAV at time... Total energy consumed The energy must be less than or equal to the energy carried by the drone. Equation (22c) gives the minimum speed constraint for the UAV. and maximum speed constraint Equation (22d) represents the length constraints of the two semi-axis of the elliptical trajectory; Equations (22e) and (22f) represent the roll angle constraints of the UAV, where... and They are The radius of curvature of the elliptical trajectory at time t and The radius of curvature of the elliptical trajectory at time step. It is the maximum roll angle of the drone; Equation (22g) gives the range of values ​​for the time slot identifier.

[0013] Step 4 includes: modeling the UAV relay trajectory optimization problem as a Markov Decision Process (MDP). During the MDP modeling process, the following steps are used: , and They represent the first The state space, action space, and reward function for each time slot; Indicates the current state. Indicates the current action. Indicates the current reward. Indicates the next state. Indicates the action to be performed in the next state. The reward for the next state is represented by equation (22a), (22b), (22c), (22d), (22e), (22f), and (22g). The optimization problem is transformed into an MDP. In the Each time slot, state space Set to: (twenty three), in, Indicates the first The two-dimensional horizontal plane position coordinates of a time-slot UAV Indicates the first Channel capacity obtained by the destination node in each time slot; when When it is an odd number, , ,in, Calculated from equation (20); In the Each time slot, action space Set to: (twenty four), The constraints are given by equation (25): (25) in, and They represent the time slot lengths respectively. Minimum and maximum constraints, elliptical trajectory in and Half shaft on the shaft and All are not greater than and Half the distance, that is and All less than or equal to ; In the Each time slot, the control strategy is written as , indicating the state is Execute action at time The probability of; In the Each time slot, reward function Written as: (26) The goal of MDP is to find an optimal strategy that maximizes cumulative reward. for: (27) in, It is a discount factor, with a value of [value missing]. , Indicates the first The reward function for each time slot; In control strategy Next, the state The expected cumulative reward is defined as follows: : (28) in, Indicates control strategy The expected value under, Represents the action space set, Action value function: (29) in, Represents the set of state spaces. Indicates from state Take action Transition to state And receive a reward The probability, For state The expected cumulative reward under the given conditions; A non-policy time difference (TD) control method is used: First, at the beginning of each traversal, the action value function is updated based on the state. ,action and rewards The observations, and the state obtained in the next time slot. The optimal action value function, i.e., the optimal Q function, is obtained through iteration according to formula (30): (30) in, For the target Q value, The current Q value, Indicates the learning rate; The learning rate is scalar. The target Q-value in equation (30) is obtained by employing the TD3 method based on Deep Neural Networks (DNN). The TD3 method consists of a policy network and two parallel evaluation networks. The policy network generates actions for the participants, and the evaluation networks evaluate the value of actions and states, guiding subsequent policy improvements. The policy network is represented as follows: The parameters are The two evaluation networks are represented as follows: , The corresponding network parameters are as follows: , The target policy network is represented as The parameters are The two target evaluation networks are respectively and The corresponding network parameters are as follows: , ; Will Called the current policy network, and This is called the current evaluation network; The entire training process of the TD3 method includes: First, the current policy network is trained according to the current state. Output a deterministic action By adding the mean, Standard deviation is Gaussian noise To construct exploratory actions and obtain the final action to be executed. Relaying actions in drones Then, the drone observed the next state. And calculate the current reward Subsequently, the empirical tuples Stored in the experience pool Next, set the batch size for extracting empirical tuples. and from Randomly select experience tuples from the pool; if the number of tuples in the experience pool is less than... Then all empirical tuples will be extracted, i.e., the number of empirical tuples extracted. , This represents the function that takes the minimum value; then, the algorithm will... A set of experience tuples is input into the policy network and the evaluation network. First, the target policy network... According to the selected number 1 empirical tuple The state in Output target action Then, the target action Adding truncated Gaussian noise for smoothing and regularization yields a new target action. ,in, The mean is 0 and the standard deviation is Gaussian noise in the range of values Internal truncation; Represents the target policy network The parameters, It is a positive number ranging from 0 to 1; the target Q value is approximated as , Determined by the minimum value of the outputs of the two objective evaluation networks: (31), in, , Indicates the first one drawn 1 empirical tuple The rewards in Indicates the first The objective evaluation network is in state and actions Q value under, ; This indicates finding the minimum value; Based on the approximation of the obtained target Q value Calculate the current evaluation network respectively loss function and current evaluation network loss function : (32), in, Indicates the first The current evaluation network is at the _th ... 1 empirical tuple The state in and actions The current Q value under the given condition; The loss function given by equation (32) is minimized, and the Adam optimizer is used to update the current evaluation network. and current evaluation network parameters and ; TD3 uses a delayed update mechanism, that is, every [time period]... Only update the current policy network and all target networks once per step, setting... If the number is even, when the update condition is met, the current policy network is updated by sampling the policy gradient. The policy gradient utilizes the first current evaluation network. Perform calculations based on randomly selected... The policy gradient for sampling is calculated using empirical tuples and equation (33). This updates the policy network parameters. : (33), in express right Find the gradient. Indicates the current policy network at the [number]th [time]. 1 empirical tuple The state in The deterministic action output; after delaying the update of the current policy network, the parameters of the target policy network are... and the parameters of the two objective evaluation networks. and Perform a soft update: (34), in, This is the soft update coefficient. .

[0014] The present invention also provides an electronic device, including a processor and a memory, the memory storing program code that, when executed by the processor, causes the processor to perform the steps of the method.

[0015] The present invention also provides a storage medium storing a computer program or instructions that, when the computer program or instructions are run on a computer, execute the steps of the method described.

[0016] This invention addresses the trajectory optimization problem of a fixed-wing UAV acting as a half-duplex mobile relay in a hybrid probabilistic channel environment. It constructs a unified optimization framework that aims to maximize the average cumulative channel capacity of the ground destination node while considering airborne energy constraints and flight safety constraints. Through a dual-delay, deep deterministic strategy gradient design, the UAV's flight speed, the shape parameters of its elliptical trajectory, and the time slot length for half-duplex communication are determined, enabling the UAV to fly along a reasonable closed elliptical trajectory. During this process, the UAV provides relay services to both the ground source and destination nodes, maximizing the average cumulative channel capacity of the ground destination node.

[0017] Compared with existing technologies, the beneficial effects of this invention are as follows: The method provided by this invention addresses the trajectory optimization problem of a fixed-wing UAV acting as a half-duplex mobile relay. In a hybrid probabilistic channel environment, it constructs a unified optimization framework that aims to maximize the average cumulative channel capacity of the ground destination node while considering airborne energy constraints and flight safety constraints (primarily roll angle constraints). Through a dual-delay, depth-deterministic strategy gradient design, it optimizes the UAV's flight speed, the shape parameters of the elliptical trajectory, and the time slot length for half-duplex communication, enabling the UAV to fly along a reasonable closed elliptical trajectory. During this process, the UAV provides relay services to both the ground source and destination nodes, maximizing the average cumulative channel capacity of the ground destination node. Simulation experiments also show that the system performance (average cumulative channel capacity of the ground destination node) provided by this optimization method is significantly higher than that of the circular trajectory case. Attached Figure Description

[0018] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, and the advantages of the present invention in the above and / or other aspects will become clearer.

[0019] Figure 1 This is a schematic diagram of the fixed-wing UAV relay communication system of the present invention.

[0020] Figure 2 This is a top view of an elliptical trajectory.

[0021] Figure 3 This is the convergence plot of the elliptical trajectory iteration.

[0022] Figure 4 This is the convergence plot of the circular trajectory iteration.

[0023] Figure 5 This is a schematic diagram of the optimized flight trajectory of a fixed-wing UAV. Detailed Implementation

[0024] Figure 1 The three-node relay communication system under consideration is presented, which includes a fixed ground source node. A fixed ground destination node And an aerial drone mobile relay R. Assume... arrive The direction of the connecting line is axis; The axis rotates counterclockwise horizontally around the zero point. for Axis; perpendicular to shaft and The horizontal plane formed by the axes and the direction passing through the zero point is The UAV relay operates in half-duplex amplification and forwarding mode. In the first time slot, the UAV mobile relay R receives the signal transmitted by N1, then amplifies the received signal from N1 and forwards it to N2 in the second time slot. Assume the UAV's flight altitude is constant. That is, in The coordinates on the axis are fixed as Its flight trajectory is constrained to an elliptical path. Furthermore, assume... Located in a two-dimensional horizontal plane Place; Located in a two-dimensional horizontal plane Among them, for and The distance between them; The antenna height is ; The antenna height is The antenna height for drone relay is the same as its flight speed, which is... .

[0025] Figure 2 A top view of the elliptical trajectory is given (here, the foci of the ellipse are at...). (Taking the axis as an example). Assume the center of the elliptical trajectory lies on a two-dimensional horizontal plane with coordinates of... elliptical trajectory in The half-shaft on the shaft is ,and ,exist The half-shaft on the shaft is ,and (That is, the focus is on) (In the case of the axis), the horizontal coordinate of the drone relay origin position is: Thus, the eccentricity of the elliptical trajectory is .

[0026] The drone relay starts at a constant speed from the originating position. ,according to Figure 2 The flight path, shown as clockwise along an elliptical trajectory, then... At that moment, the drone was flying at an angle of [missing information]. (At this point, assume the drone flies no more than one elliptical trajectory). Here, Let be the perimeter of the elliptical trajectory, where Let be the eccentricity of the elliptical trajectory. Thus, in At what moment, the position of UAV relay R on the two-dimensional horizontal plane Given by equations (1) and (2): (1), (2), It should be noted that equations (1) and (2) above only provide the two-dimensional horizontal coordinates of the UAV's position on the elliptical trajectory for a flight of no more than one revolution. In reality, the UAV may fly more than one revolution along the elliptical trajectory. The UAV travels along the elliptical trajectory at a constant speed. The flight takes approximately [time] to complete one orbit. Therefore, the location of the drone exhibits periodicity, with a period of... When determining the location of a drone relay, it is necessary to first determine the drone's flight time. according to Find the modulus, and substitute the remainder of the modulus operation into equations (1) and (2) to calculate the specific position on the two-dimensional horizontal plane.

[0027] Similarly, according to The antenna of the UAV relay R can be calculated. Time and and The distance of the antenna and The details are as follows: (3), (4), in, yes Antenna height, yes Antenna height, the height of the drone antenna and its flight altitude Consistent.

[0028] Existing research indicates that when a drone flies horizontally at a fixed altitude, it experiences significant drag. for: (5), in, Zero lift-drag coefficient, air density, For the wing area of ​​the drone, For the flight speed of the drone, For the wingspan efficiency of drones, It is the wingspan of the drone. For the overall quality of the drone, It is the acceleration due to gravity. For drones The roll angle at any given moment. The drone in... The roll angle at time t can be calculated as follows: ,in, for The radius of curvature of the elliptical trajectory of the UAV at any given time can be given by equation (6): (6), In equation (6), For an elliptical trajectory in Half-shaft on the shaft, For an elliptical trajectory in Half-shaft on the shaft, The angle at which the drone flies.

[0029] Next, according to and Then equation (5) can be rewritten as: (7), Since the UAV relay flies horizontally along an elliptical trajectory at a constant altitude and constant speed, the thrust and drag of the UAV are balanced, the vertical component of lift is balanced with gravity, and the horizontal component of lift provides the centripetal force for the motion. Therefore, the power consumption of the UAV can be calculated by equation (8): (8), As mentioned above, the drone relay starts at time 0, from the horizontal starting position. Starting from a point and flying along an elliptical trajectory, then... Energy consumption of drone relay for: (9), In half-duplex amplification and repeater mode, any node in the system cannot receive a signal while transmitting it; similarly, any node in the system cannot transmit a signal while receiving it. This is because the ground source node... With ground target node The direct link does not exist. The signal was transmitted to the air via a drone relay. Two time slots are required. A total of [number] time slots are set. Each time slot ( (even number), the Each time slot ( It is an odd number, that is, ) and the The lengths of the time slots are equal, for .here, satisfy ,in, and These are the minimum and maximum timeslot lengths, respectively.

[0030] In the In each time slot, the ground source node transmits a unit power signal. The signal received by the drone relay R Written as: (10) in, for The transmission power, yes Large-scale channel gain between R and The average power of the white Gaussian noise received by R is... .here, Written as: (11), in, It is the power per unit distance of the channel. To accumulate to the number The time of each time slot yes And R's antenna in The distance at any given time can be specifically expressed in equation (3). use Instead of calculation, It is the channel fading index. yes With R The probability of a line-of-sight (LoS) link appearing at any given moment. These are additional channel fading parameters for non-line-of-sight (NLoS) links. Calculated from equation (12): (12) in and R and Antenna height, and The terrain parameters that the drone flies over. yes The elevation angle between R and R.

[0031] In the Each time slot (here, It is an odd number, that is, The drone relay will be in the Signals received from the ground source node in each time slot Scaled to unity power signal Specifically, as shown in equation (13): (13) Then, with power Will Launch to In the first At the end of each time slot, The received signal is: (14) in, R is the transmit power. It is R and In the Large-scale channel gain per time slot yes The received Gaussian white noise has an average power of Large-scale channel gain here Written as: (15) in, It is the power per unit distance of the channel. To accumulate to the number The time of each time slot It is R and The antenna in The distance at time (specifically, the distance in equation (4)) Replace with (Calculated) It is the channel fading index. Is R and exist The probability of a Loss channel appearing at any given time. For NLoS cases, additional channel fading parameters, Calculated from equation (16): (16) in and R and Antenna height, and For the terrain parameters that the drone flies over, It is R and The angle of elevation between them.

[0032] In the At the end of the time slot, the ground target node Received signal Rewritten as: (17) In the At the end of the time slot, the ground target node The ratio of the received useful signal to the noise power, i.e., the signal-to-noise ratio (SNR), is: (18) Thus, through two time slots (the first) and the (time slots), destination node The obtained channel capacity is: Through all Transmission in time slots, destination node The cumulative channel capacity obtained with the assistance of aerial drones is: (19) in, Even number, Not greater than Odd numbers.

[0033] Observational formula (18) reveals that, The received signal-to-noise ratio is a random variable because, and They are random and statistically independent. To simplify the problem, we calculate... The average cumulative channel capacity obtained with the assistance of aerial drones. To do this, we first need to calculate... The mean is calculated using equation (20): (20), in This indicates calculating the mean. for time The probability of a Loss link between R and R can be calculated using equation (12); for Time R and The probability of a Loss link exists can be calculated using equation (16).

[0034] Thus, through all Transmission in time slots, destination node The average cumulative channel capacity obtained with the assistance of aerial drone relay is: (twenty one), In the case of an elliptical flight trajectory, to maximize the target node If the average cumulative channel capacity is the objective, then the trajectory design optimization problem can be written as: (22a) (22b) (22c) (22d) (22e) (22f) (22g) Equation (22a) gives the objective of maximizing the average cumulative channel capacity of the destination node, where, , , and The speed of the drone Elliptical trajectory in Half shaft on the shaft Elliptical trajectory in Half shaft on the shaft and the Length of each time slot The optimal solution; Equation (22b) is the energy constraint of the UAV, that is, the UAV at time... Total energy consumed The energy must be less than or equal to the energy carried by the drone. , The formula (9) can be used to... use Instead of calculation, equation (22c) gives the minimum speed constraint for the UAV. and maximum speed constraint Equation (22d) represents the length constraints of the two semi-axis of the elliptical trajectory; Equations (22e) and (22f) represent the roll angle constraints of the UAV, where... and They are and The radius of curvature of the elliptical trajectory at time moment can be specifically obtained from equation (6). Use respectively and Instead of calculation, This is the maximum roll angle of the drone; Equation (22g) gives the range of values ​​for the time slot identifier. It should be noted that: due to the total number of time slots... The number of timeslots is even, and the length of the odd-numbered timeslot is equal to the length of the following even-numbered timeslot. Therefore, in the above problem, The value range is 1 to All odd numbers.

[0035] Below, the UAV relay trajectory optimization problem is modeled as a Markov Decision Process (MDP). In the MDP modeling process, we use... , and They represent the first The state space, action space, and reward function for each time slot. Indicates the current state. Indicates the current action. Indicates the current reward. Indicates the next state. Indicates the action to be performed in the next state. This represents the reward in the next state. The optimization problem consisting of equations (22a), (22b), (22c), (22d), (22e), (22f), and (22g) can be transformed into an MDP.

[0036] In the Each time slot, state space Set to: (twenty three), in, Indicates the first The two-dimensional horizontal plane position coordinates of a time-slot UAV Indicates the first The channel capacity obtained by the destination node in each time slot. It should be noted that: when When it is an odd number, , ,in, It can be calculated using equation (20).

[0037] In the Each time slot, action space Set to: (twenty four), In equation (24) For the speed of drones, For the first The length of each time slot and For the shape parameters of the elliptical trajectory, Is it an elliptical trajectory in Half shaft on the shaft For an elliptical trajectory in Half-shaft on the shaft. , , and The value of is subject to constraints, which are given by equation (25): (25) In equation (25), and These represent the flight speeds of the drone. Minimum and maximum value constraints, and They represent the time slot lengths respectively. Minimum and maximum constraints, elliptical trajectory in and Half shaft on the shaft and All are not greater than and Half the distance, that is, and All less than or equal to .

[0038] In the For each time slot, the control strategy can be written as: , indicating the state is Execute action at time The probability of.

[0039] In the Each time slot, reward function Written as: (26) The goal of MDP is to find an optimal strategy that maximizes the cumulative reward, where the cumulative reward is: (27) in, It is a discount factor, with a value of [value missing]. , Indicates the first The reward function for each time slot.

[0040] In control strategy Next, the state The expected cumulative reward is defined as follows: Specifically, it is written as: (28) in, Indicates control strategy The expected value under, Represents the action space set, The action value function is given by equation (29).

[0041] In control strategy Next, the state Take action below The expected cumulative reward is defined as an action-value function: (29) in, Represents the set of state spaces. Indicates from state Take action Transition to state And receive a reward The probability of this (called the state transition probability). This is the discount factor. For state The expected cumulative reward is determined by the state transition probability. Since the state transition probability is unknown, the Twin Delayed Deep Deterministic policy gradient (TD3) method will be introduced to address this issue.

[0042] Due to the transition probability in equation (29) The problem is unknown, requiring a non-strategic temporal difference (TD) control method. First, the action value function (also known as the Q function) is updated at the beginning of each traversal. This is based on the state... ,action and rewards The observations, and the state obtained in the next time slot. The optimal action value function, i.e., the optimal Q function, is obtained through iteration according to formula (30): (30) in, For the target Q value, The current Q value, The learning rate is represented by Q. Here, Q represents the expected cumulative reward that the drone can obtain by taking a specific action in a given state. In fact, Q is estimated using the Q function to guide the drone relay in choosing the optimal action. The difference between the target Q value and the current Q value is the time difference (TD) error, which is calculated at the current time step and measures the gap between the target Q value and the actual reward.

[0043] Additionally, it should be noted that equation (30) gives the general update form of the action-value function based on time difference learning, where The scalar learning rate is used to iteratively correct the Q-value in the tabular time difference control method. However, the TD3 algorithm, which approximates the Q-function using a deep neural network, is used here, and the Q-value update is no longer directly applied to the scalar learning rate according to equation (30). Instead of explicit iteration, this is achieved by constructing a mean squared error loss function and using the Adam optimizer to perform gradient descent on the network parameters. The learning rate in equation (30) This can be considered as being replaced by the learning rate parameters of the evaluation network and policy network in the optimizer (learning rate of the evaluation network and learning rate of the policy network), therefore, they are no longer set separately in the actual implementation. Instead, it controls the network update step size by selecting appropriate evaluation network learning rate and policy network learning rate.

[0044] To obtain the target Q-value in equation (30), the TD3 method based on Deep Neural Networks (DNN) is adopted. This method is a model-free deterministic policy method used to learn policies in a continuous action space. The TD3 method consists of a policy network and two parallel evaluation networks. The policy network generates actions for the participant, and the evaluation networks evaluate the value of actions and states, guiding subsequent policy improvements. Specifically, the policy network is represented as follows: The parameters are The two evaluation networks are represented as follows: , The corresponding network parameters are as follows: , Furthermore, the policy network and both evaluation networks each have a target network with the same structure. The target policy network is represented as follows: The parameters are The two target evaluation networks are respectively and The corresponding network parameters are as follows: , Next, we will... Called the current policy network, and This is called the current evaluation network.

[0045] The entire training process of the TD3 method can be summarized as follows. First, the current policy network, based on the current state... Output a deterministic action To fully explore the continuous action space, a mean is added. Standard deviation is Gaussian noise To construct exploratory actions and obtain the final action to be executed. Relaying actions in drones Then, the drone can observe the next state. And calculate the current reward Subsequently, the experience tuples Stored in the experience pool Next, set the batch size for extracting empirical tuples. and from Randomly select experience tuples from the pool. If the number of tuples in the experience pool is less than... Then all empirical tuples will be extracted, that is, the number of empirical tuples extracted. , This represents the function that takes the minimum value. Then, the algorithm will... Each set of experience tuples is input into the policy network and the evaluation network. First, the target policy network... According to the selected number 1 empirical tuple The state in Output target action Then, truncated Gaussian noise is added to the target action for smoothing and regularization to obtain a new target action. ,in, The mean is 0 and the standard deviation is Gaussian noise in the range of values The truncation within. Here, Represents the target policy network The parameters, It is a positive number ranging from 0 to 1. Thus, the target Q value can be approximated as... , It is determined by the minimum value of the outputs of the two target evaluation networks.

[0046] The specific steps for solving the elliptical trajectory optimization design of UAVs are shown in Algorithm 1.

[0047] Algorithm 1: Step a1, obtain the parameters of the UAV relay system, including: ground source node Two-dimensional horizontal plane position Ground destination node Two-dimensional horizontal plane position , Antenna height , Antenna height The antenna altitude (flight altitude) of the drone R. , Transmission power The transmit power of R Channel power per unit distance Channel fading index Additional channel fading parameters in non-line-of-sight link scenarios Noise power Maximum speed limit of drone flight Minimum speed limit for drone flight Total energy of drones wing area of ​​drone Wingspan efficiency of drones Wingspan of drones Overall quality of drones Maximum roll angle of drone Zero lift drag coefficient air density Gravitational acceleration Terrain parameters of the drone flight and ; Step a2, initialize the current policy network. and two current evaluation networks and ; Step a3, initialize the target policy network and two target evaluation networks , The initialization parameters of the target network are: , , ; Step a4, Initialize the experience pool Set the maximum number of training iterations. The current number of training sessions is Set batch size ; Step a5, set the current number of training iterations. And initialize basic parameters, specifically: time slot identifier. Minimum constraint on time slot length Maximum value constraint of time slot length Accumulated to the number Time slot Gaussian noise added to enhance exploratory capabilities mean and standard deviation Gaussian noise added to the target action mean and standard deviation Parameters of the cutoff interval for Gaussian noise added to the target action Discount coefficient Update delay step interval for the current policy network and all target networks Drone flight speed Elliptical trajectory in Half shaft on the shaft Elliptical trajectory in Half shaft on the shaft Initial position of the UAV on the two-dimensional horizontal plane Remaining energy of drones , No. Channel capacity obtained by each time slot destination node Soft update coefficient ; Step a6, if , , If it is established, then the first Actions in a time slot Otherwise the first In the action space of each time slot, except , , Other variables according to Update; Step a7, based on the motion space In , set the The length of each time slot, i.e. Then, calculate the cumulative total up to the th. and the Time slot and ; Step a8, based on the motion space In and and in equation (6) Use respectively and Instead, calculations show that the drone is in and The radius of curvature of the elliptical trajectory at time t is and ; Step a9, if satisfied or If yes, return to step a5; otherwise, proceed to step a10. Step a10, according to Obtained from , , and Calculated according to formula (26) In equations (1) and (2) use Instead of calculation, we can obtain According to equation (20), the result is obtained. In this way, obtain ; Step a11, in Storage experience tuples ; Step a12, from Randomly selected from If the number of experience tuples in the experience pool is less than [number], then [the number of experience tuples in the experience pool is less than [number]; Then, during extraction, all empirical tuples will be extracted. Therefore, let , This represents a function that takes the minimum value. Step a13, based on the randomly selected Using a set of empirical tuples, the approximate value of the target Q is calculated using equation (31); Step a14, based on the randomly selected We use a set of empirical tuples to minimize the loss function (32), and use this to update the parameters of the two current evaluation networks. , ; Step a15: Determine if the current time step is satisfied. (in, If the condition is met (for the modulo operator), proceed to step a16; otherwise, proceed to step a17. Step a16, based on the randomly selected A set of empirical tuples are used to delay updating the network parameters of the current policy network using equation (33). Update the network parameters of all target networks: , , ; Step a17, in equation (9) use Instead, the energy consumption of the drone was calculated. ; Step a18, Update , ; Step a19, determine whether the condition is met. If the condition is met, proceed to step a20; otherwise, proceed to step a6. Step a20, determine Is it equal to If the result is equal to the given value, the algorithm ends; otherwise, proceed to step a5.

[0048] To address the optimization method proposed in this invention, simulation experiments were conducted. The simulation experiments solved the optimization problems of elliptical and circular trajectories based on the designed TD3 algorithm, simultaneously calculating the cumulative channel capacity received by the corresponding destination node. The results were compared with an algorithm based on Deep Deterministic Policy Gradient (DDPG) to verify the correctness and effectiveness of the proposed algorithm. In the simulation results, "Elliptical Trajectory TD3" represents the method proposed in this invention, "Circular Trajectory TD3" represents the optimal circular trajectory designed using TD3 under the same conditions, and "Elliptical Trajectory DDPG" and "Circular Trajectory DDPG" represent the optimal elliptical and optimal circular trajectories designed using DDPG, respectively.

[0049] The simulation experiments were conducted using PyCharm 2024.1, Python version 3.9, PyTorch framework version 2.2.2, and Gym version 0.26.1. A three-layer DNN neural network was used, with 128, 64, and 20 layers. The ReLU activation function was used for the hidden layers, and the Tanh activation function for the actions. The Adam optimizer was used to train the network. The learning rate for the policy network was 0.001, and the learning rate for the evaluation network was 0.003. Other parameter settings are shown in Table 1.

[0050] Table 1

[0051]

[0052] Figure 3 and Figure 4 The table shows the convergence of the reward function for the TD3 and DDPG algorithms during training on elliptical and circular trajectories, respectively. The results indicate that both trajectories converge under both TD3 and DDPG algorithms, with TD3 outperforming DDPG in both convergence performance and speed. Furthermore, the convergence effect of the elliptical trajectory is significantly better than that of the circular trajectory. The UAV flight results for both algorithms are shown in Table 2.

[0053] Table 2

[0054]

[0055] As shown in Table 2, under the same conditions, the channel capacity obtained by the TD3 algorithm is superior to that of DDPG, verifying that it effectively improves policy performance through a dual-evaluation network and a delayed update mechanism. Furthermore, the system performance under the elliptical trajectory is significantly higher than that under the circular trajectory. Notably, the optimal-performing "elliptical trajectory TD3" (i.e., the method proposed in this invention) exhibits characteristics of low speed and long flight time, indicating that the method has successfully learned a high-energy-efficiency flight mode. By optimizing energy consumption allocation, it extends the effective communication time, thereby maximizing the cumulative channel capacity.

[0056] Figure 5 Two flight trajectories of the UAV are presented, and the correctness of the flight trajectory designed by the method of the present invention can be seen from the shape of the trajectory.

[0057] This invention provides a dual-delay depth deterministic strategy gradient UAV relay trajectory optimization method. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.

Claims

1. A method for dual-delay deep-deterministic policy gradient UAV relay trajectory optimization, characterized in that, The method comprises the following steps: Step 1, establishing a three-node relay communication system; Step 2, calculating the power consumption and energy consumption of the UAV; Step 3, establishing a UAV relay trajectory optimization problem; Step 4, modeling the UAV relay trajectory optimization problem as a Markov decision process.

2. The method of claim 1, wherein, Step 1 includes: the three-node relay communication system contains a fixed ground source node , a fixed ground destination node and an aerial unmanned mobile relay R; Set To The direction of the line to The axis; The axis with zero point as the center counterclockwise horizontal rotation Is The axis; perpendicular to The horizontal plane formed by the axis and The direction of the axis passing through zero point is The axis; The UAV relay works in a half-duplex amplify-and-forward mode, in which the UAV mobile relay R receives the signal transmitted by N1 in the first time slot, and then transmits the received signal from N1 to N2 after scaling in the second time slot.

3. The method of claim 2, wherein, In step 1, the drone's flight altitude is set to be constant. That is, in The coordinates on the axis are fixed as The flight path is constrained to an elliptical path; Set located in a two-dimensional horizontal plane ; located in a two-dimensional horizontal plane , wherein is between ; the antenna height of ; the antenna height of ; the antenna height of the UAV relay is .

4. The method of claim 3, wherein, In step 1, the coordinates of the center of the elliptical trajectory on the two-dimensional horizontal plane are set as , the semi-axis of the elliptical trajectory on the axis is , and , the semi-axis on the axis is , and , the horizontal coordinate of the UAV relay originating position is , and the eccentricity of the elliptical trajectory is .

5. The method of claim 4, wherein, In step 1, the drone relay flies from the origination location at a constant speed , along an elliptical trajectory in a clockwise direction, then at time , the angle , the angle of flight of the drone relay is considered for no more than one revolution around the elliptical trajectory, is the circumference of the elliptical trajectory, where, is the eccentricity of the elliptical trajectory, is the parametric angle of the elliptical trajectory, denotes the differential of the parametric angle ; at time , the position of the drone relay R in the two-dimensional horizontal plane is given by equations (1) and (2): (1), (2), The UAV flies along an elliptical trajectory at a constant speed The UAV flies, and the time for one circle is The position of the UAV has periodicity, and the period is When the position of the UAV relay is calculated, the flight time of the UAV is taken modulo , and the remainder part of the modulo operation is substituted into formulas (1) and (2) to calculate the specific position in the two-dimensional horizontal plane; According to the distance of the antenna of the drone relay R at the time instant t from the antenna of the distance of the antenna of the drone relay R at the time instant t from the antenna of the distance of the antenna of the drone relay R at the time instant t from the antenna of : (3), (4), The drag experienced by a drone when flying horizontally at a fixed altitude. for: (5), in, Zero lift-drag coefficient, air density, For the wing area of ​​the drone, For the flight speed of the drone, For the wingspan efficiency of drones, It is the wingspan of the drone. For the overall quality of the drone, It is the acceleration due to gravity. For drones The roll angle at any moment; the drone in Roll angle of time for ,in, for The radius of curvature of the elliptical trajectory of the UAV at any given time is given by equation (6): (6), wherein is the semi-axis of the elliptic trajectory on the axis, is the semi-axis of the elliptic trajectory on the axis; According to and then equation (5) is rewritten as: (7)。 6. The method of claim 5, wherein, Step 2 includes power consumption of the drone Calculated from equation (8): (8), The UAV relay starts from time 0 from a horizontal launch location flies along an elliptical trajectory, then at time the energy consumption of the UAV relay is: (9), wherein, is an integral variable representing time, is a small increment of time; is the instantaneous power consumption of the UAV at time along an elliptical trajectory.

7. The method of claim 6, wherein, Step 3 includes: In half-duplex amplification and repeater mode, any node in the system cannot receive a signal while transmitting it; any node in the system cannot transmit a signal while receiving it; due to the ground source node... With ground target node The direct link does not exist. The signal was transmitted to the air via a drone relay. Two time slots are required; a total of [number] slots are set. Each time slot If it is even, the first The first time slot and the first The lengths of the time slots are equal, for , It is an odd number. , satisfy ,in, and These are the minimum and maximum timeslot lengths, respectively. In the first time slot, the ground source node transmits a unit power signal The signal received by the UAV relay R is then written as: (10), wherein is the transmit power of the R, is the large-scale channel gain between R and R, is the Gaussian white noise received by R with average power , is written as (11), in, It is the power per unit distance of the channel. To accumulate to the number The time of each time slot yes And R's antenna in The distance of time, It is the channel fading index. yes With R The probability of a line-of-sight link occurring at any given time. These are additional channel fading parameters for non-line-of-sight link scenarios. Calculated from equation (12): (12), wherein, and is a terrain parameter for the UAV to fly over, is is an elevation angle between R and In the In the time slot, the drone relay will be in the first... Signals received from the ground source node in each time slot Scaled to unity power signal Specifically, as shown in equation (13): (13), Then, the power The transmitted to , at the end of the time slot, the received signal is: (14), wherein is the transmit power of R, is R and is the large-scale channel gain of the th time slot, is is the received Gaussian white noise, and the large-scale channel gain is written as: (15), wherein is the time accumulated to the th time slot, is the distance of the antenna of R and at the th time instant, is the distance of R and at the th time instant, is the probability that R and have LoS channel at the th time instant, and is calculated by equation (16). (16), wherein is an elevation angle between R and is an elevation angle between R and At the end of the time slot, the ground destination node at the end of the time slot the received signal is rewritten as (17), In the At the end of the time slot, the ground target node The signal-to-noise ratio (SNR) is the ratio of the received useful signal to the noise power. for: (18), By the first and the second time slot, the destination node obtains the channel capacity as: ; Through transmission of all slots, the destination node cumulative channel capacity obtained through assistance of an aerial drone is: (19), First, find the mean of: (20), wherein denotes averaging; By transmission of all slots, the destination node obtains an average cumulative channel capacity with the assistance of aerial drone relays is: (21), In the case of an elliptical flight trajectory, to maximize the target node If the average cumulative channel capacity is the objective, then the UAV relay trajectory optimization problem can be written as: (22a), (22b), (22c), (22d), (22e), (22f), (22g), Equation (22a) gives the objective of maximizing the average cumulative channel capacity of the destination node, where, , , and The speed of the drone The optimal solution, the elliptical trajectory in Half shaft on the shaft The optimal solution, the elliptical trajectory in Half shaft on the shaft The optimal solution and the first Length of each time slot The optimal solution; Equation (22b) is the energy constraint of the UAV, that is, the UAV at time... Total energy consumed The energy must be less than or equal to the energy carried by the drone. Equation (22c) gives the minimum speed constraint for the UAV. and maximum speed constraint Equation (22d) represents the length constraints of the two semi-axis of the elliptical trajectory; Equations (22e) and (22f) represent the roll angle constraints of the UAV, where... and They are The radius of curvature of the elliptical trajectory at time t and The radius of curvature of the elliptical trajectory at time step. It is the maximum roll angle of the drone; Equation (22g) gives the range of values ​​for the time slot identifier.

8. The method of claim 7, wherein, Step 4 includes modeling the UAV relay trajectory optimization problem as a Markov Decision Process (MDP), in which , and represent the state space, action space, and reward function of the th time slot, respectively; represents the current state, represents the current action, represents the current reward, represents the next state, represents the action at the next state, represents the reward at the next state; and the optimization problem consisting of Equations (22a), (22b), (22c), (22d), (22e), (22f), and (22g) is converted to an MDP. In the first slot, the state space is set to: (23), wherein represents the two-dimensional horizontal plane position coordinate of the drone in the nth time slot, represents the channel capacity obtained by the destination node in the nth time slot; when is odd, , wherein is calculated from equation (20);​​ In the first time slot, the action space is set to: (24), The constraint condition is given by formula (25): (25), in, and They represent the time slot lengths respectively. Minimum and maximum constraints, elliptical trajectory in and Half shaft on the shaft and All are not greater than and Half the distance, that is and All less than or equal to ; At the first time slot, the control policy is written as , indicating the probability of performing action when the state is ; At the first time slot, the reward function is written as: (26), The goal of MDP is to find an optimal policy that maximizes the cumulative reward, the cumulative reward is given by: (27), wherein, is a discount factor, taking values , represents the reward function of the th time slot; In the control policy The expected cumulative reward under the state is defined as : (28), wherein, represents a control policy under the desired value, represents a set of action spaces, represents an action value function: (29), wherein, represents a set of state spaces, represents a probability of transitioning from state taking action to state and obtaining a reward , is the expected cumulative reward for state . Using a non-strategic temporal difference (TD) control method: first, at the beginning of each iteration, the action-value function is updated based on observations of states , actions , and rewards , and the resulting state at the next time slot, then the optimal action-value function, i.e., the optimal Q-function, is obtained iteratively according to equation (30): (30), wherein, is the target Q value, is the current Q value, denotes the learning rate; is the scalar learning rate; The target Q value in formula (30) is obtained by using a TD3 method based on a deep neural network DNN, the TD3 method comprising a policy network and two parallel evaluation networks, the policy network being used to generate actions for the participant, and the evaluation networks being used to evaluate the values of the actions and states and guide subsequent policy improvement, the policy network being represented as , the parameters being , the two evaluation networks being represented as , , the corresponding network parameters being , ; the target policy network being represented as , the parameters being , the two target evaluation networks being and , the corresponding network parameters being , ; Will Called the current policy network, and This is called the current evaluation network; The whole training process of the TD3 method includes: first, the current policy network outputs a deterministic action according to the current state ; then, an exploratory action is constructed by adding a Gaussian noise with a mean of and a standard deviation of , to obtain a final executed action ; after the UAV executes the action , the UAV observes the next state and calculates the current reward ; then, the experience tuple is stored in the experience pool ; then, the extraction experience tuple batch size is set, and the experience tuples are randomly extracted from ; if the number of tuples in the experience pool is less than , all experience tuples are extracted, that is, the number of extracted experience tuples , represents the minimum value function; then, the algorithm inputs experience tuples to the policy network and the evaluation network, first, the target policy network outputs a target action according to the state in the extracted experience tuple , then adds a truncated Gaussian noise to the target action for smoothing regularization, to obtain a new target action , wherein is the truncation of the Gaussian noise with a mean of 0 and a standard deviation of in the value interval ; represents the parameters of the target policy network , and is a positive number with a value of 0 to 1; the target Q value is approximated as , which is determined by the minimum value output by the two target evaluation networks: (31), wherein, , represents the reward in the extracted th experience tuple , represents the Q-value of the th target evaluation network at state and action ; ; represents to take the minimum value; According to the approximation of the target Q value obtained , loss functions of the current evaluation network and the current evaluation network are calculated respectively and : (32), in, Indicates the first The current evaluation network is at the _th ... 1 empirical tuple The state in and actions The current Q value under the given condition; By minimizing the loss function given by equation (32) and using the Adam optimizer while updating the parameters of the current critic network and the current critic network and ;​ TD3 adopts a delayed update mechanism, i.e., the current policy network and all target networks are updated only once every steps, and the update interval is set to be an even number When the update condition is met, the current policy network is updated by sampling the policy gradient, and the policy gradient is calculated by the first current evaluation network , according to the randomly extracted experience tuples, and the sampled policy gradient is calculated according to formula (33), so as to update the policy network parameters : (33), in express right Find the gradient. Indicates the current policy network at the [number]th [time]. 1 empirical tuple The state in The deterministic action output; after delaying the update of the current policy network, the parameters of the target policy network are... and the parameters of the two objective evaluation networks. and Perform a soft update: (34), wherein is a soft update coefficient, .

9. An electronic device, comprising: The device comprises a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.

10. A storage medium, characterized by The computer program or instructions are stored, and when the computer program or instructions are run on a computer, the steps of the method according to any one of claims 1 to 8 are executed.