Drone-Assisted Network Path Optimization Method and System Based on Deep Reinforcement Learning

Through a method based on deep reinforcement learning, the drone path is optimized, which solves the challenges of drone flight path planning in dynamic environments, and achieves efficient energy management and service quality satisfaction.

CN119094975BActive Publication Date: 2025-06-17NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411119870.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-15
Publication Date
2025-06-17
Estimated Expiration
2044-08-15

AI Technical Summary

Technical Problem

In a dynamically changing environment, how to reasonably plan the flight path of drones to meet the service quality needs of ground users, especially when energy limitations and users are dynamically moving.

Method used

Using a method based on deep reinforcement learning, the drone base station downlink scenario is constructed, the user dynamic movement is simulated, the joint optimization model is generated, and the network training is accelerated using dual competitive deep Q networks and imagination mechanisms to optimize the drone path.

Benefits of technology

It significantly reduces the energy consumption of drones, improves transmission efficiency in dynamic environments, and meets users' service quality needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119094975B_ABST
    Figure CN119094975B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of electromagnetic spectrum, and specifically discloses a method and system for optimizing the path of a drone-assisted network based on deep reinforcement learning. The method includes: constructing a typical downlink scenario of a drone base station, and simulating the behavior of users' dynamic movement according to the Gaussian Markov model; according to the downlink scenario of the drone base station, conducting system modeling and analysis, and generating a joint optimization model with the goal of minimizing energy consumption; solving the optimal solution of the joint optimization model through a deep reinforcement learning algorithm. This method combines a dual-competitive deep Q-network and an imagination mechanism (IM), uses a dual Q-network and a competitive structure to handle the problem of overestimation of the value function, and utilizes the IM to propagate information across rounds to accelerate network training, and can dynamically adapt to changing electromagnetic environments and user requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electromagnetic spectrum, and particularly relates to a method and system for optimizing the path of a drone-assisted network based on deep reinforcement learning. Background Art

[0002] In recent years, drone-assisted communication has received extensive attention due to its flexibility, economy, and potential to enhance coverage in various scenarios. Its excellent mobility has enabled drone-assisted communication systems to play an important role in emergency communication and broadband connection in remote areas. For example, in the case of natural disasters severely damaging ground communication infrastructure, drones can serve as temporary base stations to quickly restore the connection between ground users and the backbone network. In addition, compared with traditional ground communication systems, drones establish line-of-sight (LoS) links with ground users more frequently, providing higher-quality communication services. However, when serving users under energy constraints, ensuring the quality of service for ground users remains a major challenge. The dynamic movement of users leads to continuous changes in signal quality, and coupled with the different service requirements of different users, these factors pose higher requirements for the trajectory planning and service strategies of drones. Therefore, how to reasonably plan the flight path of drones in a dynamically changing environment to meet the needs of all users has become one of the key research issues.

[0003] Traditional drone trajectory planning problems can usually be abstracted as the traveling salesman problem and the pick-up problem. By solving these problems, the initial trajectory of the drone can be obtained. For example, some research focuses on the trajectory and resource allocation of drones to minimize the time consumption of specific tasks. This problem is represented by a set of line segments through discretization to represent the actual trajectory, and a block coordinate descent algorithm is used to decouple it into a resource allocation sub-problem and a trajectory optimization sub-problem. There is also research on jointly optimizing the three-dimensional drone trajectory and speed-time profile under the constraints of on-board energy, speed, acceleration, and completion time to maximize the total throughput of users. The original problem is decoupled into two sub-problems. First, the trajectory is optimized considering speed constraints, and in the second sub-problem, the speed and time for each time slot are optimized. Compared with traditional convex optimization algorithms, algorithms based on reinforcement learning may have higher performance and can adapt to more complex environments.

[0004] Reinforcement learning is a machine learning algorithm that continuously updates and iterates the existing model by using new data. Its characteristic lies in that during the optimization process, it can not only utilize the existing data but also obtain new data through exploration of the environment, thereby achieving the goal of real-time dynamic optimization. Some research has proposed a distributed trajectory optimization algorithm based on dual-stream attention multi-agent actor-critic to control multiple drones to achieve on-demand coverage of a wireless communication network. There is also research on using a MADDPG algorithm for drones to respond quickly to emergencies and improve the fairness of the covered area.

[0005] Although reinforcement learning has shown great potential in drone-assisted wireless networks, many current studies are still based on static user quality of service. In addition, when training algorithms, many existing studies rarely consider the problem of drone flight data sampling. However, the actual network environment is constantly changing, users' positions are moving in real time, and drones should utilize data samples as much as possible. This requires the drone path optimization to be reasonably adjusted according to the positions and service states of users. Therefore, there is an urgent need to more deeply explore how to efficiently implement reinforcement learning methods to achieve drone path optimization in a continuously changing environment. Summary of the Invention

[0006] To achieve the object of the present invention, the present application provides a method for optimizing the path of a drone-assisted network based on deep reinforcement learning, including:

[0007] Step S1: Construct a typical downlink scenario of a drone base station, and simulate the dynamic movement behavior of users according to the Gaussian Markov model;

[0008] Step S2: According to the downlink scenario of the drone base station, conduct system modeling analysis, and generate a joint optimization model with the goal of minimizing energy consumption;

[0009] Step S3: Solve the optimal solution of the joint optimization model through a deep reinforcement learning algorithm, fuse the dual-competition deep Q network and the imagination mechanism, and use IM cross-episode propagation of information to accelerate network training.

[0010] In some specific embodiments, the step S1 includes:

[0011] Step S11: The drone and the user mobile node select a direction and a speed to move from the current position to a new position;

[0012] Step S12: The user mobile node moves at a constant time interval each time, and calculates the new direction and speed of the user mobile node;

[0013] Step S13: If the drone or the user mobile node reaches the simulation boundary, it bounces back from the simulation boundary. The bounce-back angle is determined by the incident direction, and it continues to move along this path. Repeat the above steps until the end.

[0014] In some specific embodiments, in step S2, the joint optimization model is determined according to the following formula:

[0015]

[0016] Wherein, C1 to C3 represent the flyable area of the UAV; C4 represents that the UAV needs to fly to the end point in the last time slot; C5 represents the on-board energy limit of the UAV; C6 represents the minimum transmission data volume limit of the user; C7 represents the user signal-to-noise ratio limit; c k [t] represents the link state between the user and the UAV; t represents the time slot; k represents the number of users, represents the data volume transmitted from the UAV to the user, u[t] represents the user location, represents the remaining data volume to be transmitted, ω[t] represents the UAV location, ω end represents the UAV end point location, E[t] represents the remaining energy in the current time slot, E max represents the maximum on-board energy.

[0017] In some specific embodiments, step S3 includes:

[0018] Step S31: Determine the data transmission rate between the UAV and the user according to the transmission power of the UAV and the signal-to-noise ratio between the UAV and the ground user in the time slot, and determine the energy consumption of the UAV by using the closed analytical propulsion power consumption model;

[0019] Step S32: Set the state function, action function and reward function of the UAV;

[0020] Step S33: Determine the loss of the UAV data in the D3QN neural network according to the error between the predicted Q value and the target Q value, and update the neural network;

[0021] Step S34: Calculate the difference between the UAV state-action pairs through the IM mechanism according to the state function and action function of the UAV, and update the Critic and IM networks accordingly.

[0022] In some specific embodiments, in step S3, the state function of the UAV is:

[0023]

[0024] Wherein, u[t] represents the user location, represents the remaining data volume to be transmitted, ω[t] represents the UAV location, ω end represents the UAV end point location, E[t] represents the remaining energy in the current time slot, E max represents the maximum on-board energy.

[0025] In some specific embodiments, the action function of the user is: (+x d ,0,0), (-x d ,0,0), (0,+y d ,0), (0,-y d,0), (0,0, +z d ), (0,0, -z d ), and (0,0,0) respectively represent left, right, forward, backward, upward, downward, and hovering behaviors.

[0026] In some of the specific embodiments, the reward function includes:

[0027]

[0028] r2[t] = -κ 2a E[t] - κ 2b D sum

[0029]

[0030] where κ i , i = {1, 2a, 2b, 3, 4, 5} are reward parameters, and D sum is a normalization function with respect to the total transmitted data and is determined according to the following formula:

[0031]

[0032] where α and β are normalization parameters;

[0033] In some of the specific embodiments, the overall reward and punishment design determined according to the reward function is:

[0034]

[0035] In some of the specific embodiments, the closed - form propulsion power consumption model of the drone is expressed as follows:

[0036]

[0037] where P0, P i are two constants representing the profile power and induced power in the hovering state respectively, U tip is the tip speed of the rotor blade, v0 and ν are the rotor induced velocity and the fuselage drag ratio in the hovering state respectively, ρ and s are the air density and the rotor disk area respectively, A is the rotor disk area, and V is the speed of the drone at time slot t.

[0038] To achieve the same invention purpose, the present application also provides a drone - assisted network path optimization system based on deep reinforcement learning, including:

[0039] Behavior simulation module: used to construct a typical downlink scenario of the drone base station and simulate the behavior of the user's dynamic movement according to the Gaussian Markov model;

[0040] Optimization model generation module: used to perform system modeling and analysis according to the downlink scenario of the UAV base station, and generate a joint optimization model with the goal of minimizing energy consumption;

[0041] Model solution module: used to solve the optimal solution of the joint optimization model through a deep reinforcement learning algorithm, and train the UAV to learn the optimal solution through a DQN neural network, a double-competitive deep Q network, and an IM mechanism.

[0042] Beneficial effects of the above technical solutions:

[0043] In view of the problem of joint optimization of transmission efficiency and energy consumption of UAV-assisted ground mobile users, the present invention proposes a UAV path-assisted planning algorithm based on deep reinforcement learning. By integrating the imagination mechanism and the double-competitive deep Q network, the advantages of the IM mechanism for cross-round information propagation, the double Q network, and the competitive structure are effectively utilized, so as to efficiently adapt to the dynamically changing electromagnetic environment and user requirements. On the premise of meeting the user service quality, the energy consumption of the UAV is significantly reduced. Description of the drawings

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0045] Figure 1 Flowchart of a UAV-assisted network path optimization method based on deep reinforcement learning provided by an embodiment of the present invention;

[0046] Figure 2 UAV service ground user model diagram provided by an embodiment of the present invention;

[0047] Figure 3 Schematic diagram of the UAV network learning curve provided by an embodiment of the present invention;

[0048] Figure 4 Schematic diagram of the remaining energy curve of the UAV under different algorithms provided by an embodiment of the present invention;

[0049] Figure 5 Schematic diagram of the remaining energy of the UAV under different algorithms with different numbers of users provided by an embodiment of the present invention;

[0050] Figure 6 Schematic diagram of the remaining energy of the UAV under different amounts of transmitted data provided by an embodiment of the present invention;

[0051] Figure 7 Schematic diagram of the structure of a drone-assisted network path optimization system based on deep reinforcement learning provided by an embodiment of the present invention. Detailed implementation manners

[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0053] Examples of the embodiments are shown in the accompanying drawings, where the same or similar symbols represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as a limitation to the present invention.

[0054] Embodiment 1

[0055] An embodiment of the present invention provides a drone-assisted network path optimization method based on deep reinforcement learning. Referring to Figure 1 as shown, it includes:

[0056] Step S1: Construct a typical downlink scenario of a drone base station, and simulate the behavior of users' dynamic movement according to the Gaussian Markov model.

[0057] In a specific embodiment of the present invention, the step S1 includes:

[0058] Step S11: The drone and the user mobile node select a direction and a speed to move from the current position to a new position;

[0059] Step S12: The user mobile node moves at a constant time interval each time, and calculates the new direction and speed of the user mobile node;

[0060] Step S13: If the drone or the user mobile node reaches the simulation boundary, it bounces back from the simulation boundary. The bounce-back angle is determined by the incident direction, and it continues to move along this path. Repeat the above steps until the end.

[0061] Step S2: According to the downlink scenario of the drone base station, conduct system modeling analysis, and generate a joint optimization model with the goal of minimizing energy consumption.

[0062] In a specific embodiment of the present invention, in step S2, the joint optimization model is determined according to the following formula:

[0063]

[0064] Wherein, C1 to C3 represent the flyable area of the UAV; C4 represents that the UAV needs to fly to the end point in the last time slot; C5 represents the on-board energy limit of the UAV; C6 represents the minimum transmission data volume limit of the user; C7 represents the user signal-to-noise ratio limit; c k [t] represents the link state between the user and the UAV; t represents the time slot; k represents the number of users, represents the data volume transmitted from the UAV to the user, u[t] represents the user location, represents the remaining data volume to be transmitted, ω[t] represents the UAV location, ω end represents the UAV end point location, E[t] represents the remaining energy in the current time slot, E max represents the maximum on-board energy.

[0065] In a specific embodiment of the present invention, the data volume transmitted from the UAV to the user is expressed as:

[0066]

[0067] Wherein, c k [t] represents whether the user is linked to the UAV, R k [t] represents the transmission rate between the UAV and the user in time slot t.

[0068] In a specific embodiment of the present invention, the transmission rate between the UAV and the user in time slot t is determined according to the following formula:

[0069] R k [t]=Blog2(1 + γ k [t])

[0070] Wherein, γ k [t] is the user signal-to-noise ratio, and B is the user bandwidth.

[0071] Step S3: Solve the optimal solution of the joint optimization model through a deep reinforcement learning algorithm, and train the UAV to learn the optimal solution through a DQN neural network, a double-competitive deep Q network, and an IM mechanism.

[0072] Step S31: Determine the data transmission rate between the UAV and the user according to the transmission power of the UAV and the signal-to-noise ratio between the UAV and the ground user in the time slot, and determine the energy consumption of the UAV by using a closed-form analytical propulsion power consumption model;

[0073] Step S32: Set the state function, action function, and reward function of the UAV;

[0074] Step S33: Determine the loss of the UAV data in the D3QN neural network according to the error between the predicted Q value and the target Q value, and update the neural network;

[0075] Step S34: Calculate the difference between the UAV state-action pairs according to the state function and action function of the UAV through the IM mechanism, and update the Critic and IM networks accordingly.

[0076] In a specific embodiment of the present invention, the state function of the UAV is:

[0077]

[0078] where u[t] represents the user location, represents the remaining data volume to be transmitted, ω[t] represents the UAV location, ω end represents the UAV end location, E[t] represents the remaining energy in the current time slot, E max represents the maximum on-board energy.

[0079] In a specific embodiment of the present invention, the action function of the user is: (+x d ,0,0), (-x d ,0,0), (0,+y d ,0), (0,-y d ,0), (0,0,+z d ), (0,0,-z d ), and (0,0,0), representing left, right, forward, backward, up, down, and hover behaviors respectively.

[0080] In a specific embodiment of the present invention, the reward function includes:

[0081]

[0082] r2[t] = -κ 2a E[t] - κ 2b D sum

[0083]

[0084] where κ i , i = {1, 2a, 2b, 3, 4, 5} are reward parameters, D sum is a normalization function with respect to the total transmitted data and is determined according to the following formula:

[0085]

[0086] where α, β are normalization parameters;

[0087] In a specific embodiment of the present invention, the total reward and punishment determined according to the reward function is:

[0088]

[0089] In a specific embodiment of the present invention, the closed - analytical propulsion power consumption model of the unmanned aerial vehicle is expressed as follows:

[0090]

[0091] In the formula, P0, P i are respectively two constants representing the blade profile power and the induced power in the hover state, U tip is the tip speed of the rotor blade, v0 and ν are respectively the rotor induction speed and the fuselage drag ratio in the hover state, ρ and s are respectively the air density and the rotor disk area, A is the rotor disk area, and V is the speed of the unmanned aerial vehicle at time slot t.

[0092] Referring to Figure 2 as shown, in a downlink communication scenario, within a defined area R, one unmanned aerial vehicle with a total energy of E max serves as a base station to provide services for K ground users, where K = {1,..., K}. Assume that the target area R is a rectangle of [0, X max ×[0, Y max . E[t] represents the energy consumed by the unmanned aerial vehicle at time slot t, where E[0] = 0, and T t represents the total flight time slots of the unmanned aerial vehicle (each time slot t has the same duration δ t ), where t ∈ T, T = {0, 1,..., T t}, represents the minimum amount of data that the unmanned aerial vehicle needs to transmit to ground user k. Charging stations are respectively set near the starting point and the ending point of the unmanned aerial vehicle's flight. The unmanned aerial vehicle needs to complete the communication tasks of all ground users and fly to the destination ω max before consuming the maximum energy E end , otherwise it is regarded as a mission failure. This application takes minimizing the energy consumed by the unmanned aerial vehicle to complete the communication tasks and fly to the end point as the optimization goal.

[0093] The channel model between the unmanned aerial vehicle base station and the ground users is a probabilistic path loss model. The line - of - sight probability between the unmanned aerial vehicle and ground user k at time t is:

[0094]

[0095] In the formula, a and b are constants related to the type of propagation environment, and θ is the elevation angle, which can be specifically expressed as

[0096]

[0097] The LoS loss and the NLoS loss between the unmanned aerial vehicle and ground user k can be respectively expressed as:

[0098]

[0099]

[0100] Wherein, f is the carrier frequency, c represents the speed of light, and η LoS 、η NLoS are the average additional losses of the LoS and NLoS links respectively

[0101] The loss between the UAV and the ground user k can be expressed as:

[0102]

[0103] Given the UAV transmission power P, the signal-to-noise ratio (SNR) between the UAV and the ground user k at time slot t can be expressed as:

[0104]

[0105] Wherein, n0 is the noise power spectral density. To ensure the quality of service (QoS) of the user between the UAV and the user k, it is assumed in this paper that the SNR of the k-th ground at the t-th time slot should exceed a certain threshold γ m i n :

[0106] γ k [t]≥γ min

[0107] The closed analytical propulsion power consumption model of the UAV is expressed as follows:

[0108]

[0109] Wherein, P0, P i are two constants representing the blade power and the induced power in the hovering state respectively, U tip is the tip speed of the rotor blade, v0 and ν are the rotor induced speed and the fuselage drag ratio in the hovering state respectively, ρ and s are the air density and the rotor disk area respectively, A is the rotor disk area, and V is the speed of the UAV at time slot t

[0110] The total energy consumption E[T t is:

[0111]

[0112] The loss of the DQN neural network is calculated based on the error between the predicted Q value and the target Q value, and then the gradient descent method is used to update the parameters θ of the neural network. The predicted Q value is described as follows:

[0113]

[0114] where γ is the discount rate. The description of the target Q-value is as follows:

[0115]

[0116] The loss function is described as follows:

[0117]

[0118] In the double Q-network, two independent networks are used. The update of the target Q-value does not depend on the predicted Q-value of the same network, and the target Q-value is redefined in the following form:

[0119]

[0120] In the neural network with a competitive structure, the action-value function is decomposed into a state-value function and an advantage function, and its Q-value is redefined as follows:

[0121]

[0122] where φ, are the parameters for the advantage function and the state-value function respectively, and A is the number of all possible actions.

[0123] The IM mechanism consists of five steps: 1) First, randomly sample another set of experience samples (s', a') from the experience replay pool; 2) Then input both (s, a) and (s', a') into the feature encoder to obtain the features (q, q'); 3) Next, use the similarity function to calculate the features (q, q') to obtain the similarity vector v; 4) Pass the similarity vector v to the inference network to obtain the difference d(s, a, s', a') between the vectors; 5) Finally, use the mean squared error (MSE) of Q θ (s, a) + d and Q θ (s', a') as the loss function to update the Critic network and IM.

[0124] The simulation is built using the pytorch library in Python. D3QN contains two hidden layers, each hidden layer contains 256 neurons, and Relu is selected as the activation function. The learning rate is 0.0001, the discount rate is 0.99. The multi-head encoder of the IM structure and the hidden layer of the difference inference network have 256 and 64 neurons respectively. The similarity function uses cosine similarity, and the activation function is ReLU.

[0125] Figure 3The learning curve of the drone network is shown. The average reward value is used as a metric because it is a common method for evaluating reinforcement learning algorithms and can effectively reflect the overall performance of the algorithm. It can be seen that the agent trained by the IM-D3QN algorithm with the IM mechanism reached a platform area with good performance after about 1100 rounds, while the number of rounds required for the agent trained using the ordinary D3QN algorithm to reach the same reward value was about 1400 rounds, and the training efficiency was improved by about 21.4%. After the IM-D3QN with the IM mechanism continuously trains the agent, more information can be transferred to other states across rounds, and the algorithm can learn sample information more effectively and then have better performance.

[0126] In order to better demonstrate the performance of the algorithm, the proposed algorithm is compared with the Greedy algorithm and the KMeans algorithm. For the Greedy algorithm, the drone selects the nearest ground user who has not completed the communication task as the next target. When the drone completes the tasks of all users, it flies directly to the end point. For the KMeans algorithm, all ground users are clustered into several clusters. The drone first flies to the center of the nearest cluster until the tasks of all users in the cluster are completed, and then the drone selects the next nearest cluster. Similarly, when the tasks of all users are completed, it flies directly to the end point.

[0127] Figure 4 The curves of the remaining energy of the drone after completing the task under the three different algorithms as the number of iterations change are shown in detail. It is observed that in the early stage of IM-D3QN training, the drone attempts to complete the communication task in the flight area until the energy is exhausted, and its remaining energy is lower than that of the other two algorithms. However, as the number of training times increases, the performance of IM-D3QN gradually emerges, not only surpassing the greedy algorithm and KMeans algorithm, but also after about 1100 iterations, the remaining energy tends to be stable and close to the optimal state.

[0128] Figure 5 The comparison of the remaining energy of drones after completing tasks under different numbers of users using different algorithms is shown. It can be seen that the drones trained by the IM-D3QN algorithm have the largest remaining energy after completing tasks. Traditional greedy algorithms have difficulty coping with situations with a large number of users. D3QN and DQN based on deep reinforcement learning can better find suitable paths to reduce energy loss and complete communication tasks.

[0129] Figure 6It shows the situation of drones completing tasks with four algorithms under different amounts of transmitted data. It can be observed that the agents trained by the IM-D3QN algorithm have the most remaining energy overall under different data amounts. The drones controlled by the greedy algorithm and the KMeans algorithm need to stay at the relevant user or cluster center locations for a long time to complete data transmission, and the energy consumption increases significantly when the required amount of transmitted data is large.

[0130] The above test results show that the proposed scheme and method in this example have good performance. Especially in terms of energy efficiency using IM-D3QN, the method proposed in this application is significantly better than the traditional DQN algorithm and the other two traditional optimization algorithms.

[0131] Embodiment 2

[0132] An embodiment of the present invention provides a drone-assisted network path optimization system based on deep reinforcement learning. Referring to Figure 7 as shown, it includes:

[0133] Behavior simulation module 10: used to construct a typical downlink scenario of a drone base station and simulate the behavior of users' dynamic movement according to the Gaussian Markov model;

[0134] Optimization model generation module 20: used to perform system modeling and analysis according to the downlink scenario of the drone base station, and generate a joint optimization model with the goal of minimizing energy consumption;

[0135] Model solving module 30: used to solve the optimal solution of the joint optimization model through a deep reinforcement learning algorithm, and train the drone to learn the optimal solution through a DQN neural network and a double-competitive deep Q network.

[0136] In a specific embodiment of the present invention, the behavior simulation module 10 is used to perform the following steps:

[0137] Step S11: The drone and the user mobile node select a direction and speed to move from the current position to a new position;

[0138] Step S12: The user mobile node moves at a constant time interval each time, and calculates the new direction and speed of the user mobile node;

[0139] Step S13: If the drone or the user mobile node reaches the simulation boundary, it bounces back from the simulation boundary. The bounce-back angle is determined by the incident direction, and it continues to move along this path. Repeat the above steps until the end.

[0140] In a specific embodiment of the present invention, in step S2, the joint optimization model is determined according to the following formula:

[0141]

[0142] In the formula, C1 to C3 represent the flyable area of the UAV; C4 represents that the UAV needs to fly to the end point in the last time slot; C5 represents the on-board energy limit of the UAV; C6 represents the minimum transmission data volume limit of the user; C7 represents the signal-to-noise ratio limit of the user; c k [t] represents the link state between the user and the UAV; t represents the time slot; k represents the number of users, represents the data volume transmitted from the UAV to the user, u[t] represents the user position, represents the remaining data volume to be transmitted, ω[t] represents the UAV position, ω end represents the UAV end point position, E[t] represents the remaining energy in the current time slot, E max represents the maximum on-board energy.

[0143] In a specific embodiment of the present invention, the model solving module 30 is used to perform the following steps:

[0144] Step S31: Determine the data transmission rate between the UAV and the user according to the transmission power of the UAV and the signal-to-noise ratio between the UAV and the ground user in the time slot, and determine the energy consumption of the UAV by using the closed-form analytical propulsion power consumption model;

[0145] Step S32: Set the state function, action function and reward function of the UAV;

[0146] Step S33: Determine the loss of the UAV data in the D3QN neural network according to the error between the predicted Q value and the target Q value, and update the neural network;

[0147] Step S34: Calculate the difference between the UAV state-action pairs through the IM mechanism according to the state function and action function of the UAV, so as to update the Critic and IM networks.

[0148] In a specific embodiment of the present invention, in the model solving module 30, the state function of the UAV is:

[0149]

[0150] In the formula, u[t] represents the user position, represents the remaining data volume to be transmitted, ω[t] represents the UAV position, ω end represents the UAV end point position, E[t] represents the remaining energy in the current time slot, E max represents the maximum on-board energy.

[0151] In a specific embodiment of the present invention, the action function of the user is: (+x d ,0,0), (-xd , 0, 0), (0, +y d , 0), (0, -y d , 0), (0, 0, +z d ), (0, 0, -z d ), and (0, 0, 0) represent left, right, forward, backward, upward, downward, and hovering behaviors respectively.

[0152] In a specific embodiment of the present invention, the reward function includes:

[0153]

[0154] r2[t] = -κ 2a E[t] - κ 2b D sum

[0155]

[0156] where κ i , i = {1, 2a, 2b, 3, 4, 5} are reward parameters, and D sum is the normalization function with respect to the total transmitted data and is determined according to the following formula:

[0157]

[0158] where α, β are normalization parameters;

[0159] In a specific embodiment of the present invention, the total reward and punishment determined according to the reward function is:

[0160]

[0161] In a specific embodiment of the present invention, the amount of data transmitted by the UAV to the user is expressed as:

[0162]

[0163] where c k [t] represents whether the user is linked to the UAV, and R k [t] represents the transmission rate between the UAV and the user at time slot t.

[0164] In a specific embodiment of the present invention, the transmission rate between the UAV and the user at time slot t is determined according to the following formula:

[0165] R k [t] = Blog2(1 + γ k [t])

[0166] where γk [t] is the user signal-to-noise ratio, and B is the user bandwidth.

[0167] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

[0168] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general computer, a special computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct the computer or other programmable data processing terminal devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device realizes the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto the computer or other programmable data processing terminal devices, so that a series of operation steps are executed on the computer or other programmable terminal devices to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable terminal devices provide for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1Steps of the functions specified in one or more boxes. Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention. Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the said element.

[0169] The above has introduced the method and device provided by the present invention in detail. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

[0170] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", "one specific embodiment" or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0171] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of this application, and not to limit them; although the above application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A UAV-assisted network path optimization method based on deep reinforcement learning, characterized in that: include: Step S1: construct a typical drone base station downlink scenario and simulate the user's dynamic movement behavior based on the Gauss-Markov model; Step S2: Based on the UAV base station downlink scenario, a system modeling analysis is performed to generate a joint optimization model with the goal of minimizing energy consumption; Step S3: solving the optimal solution of the joint optimization model through a deep reinforcement learning algorithm, integrating the dual competitive deep Q network and the imagination mechanism and using IM to propagate information across rounds to accelerate network training; Step S3 includes: Step S31: determining the data transmission rate between the UAV and the user according to the transmission power of the UAV and the signal-to-noise ratio between the UAV and the ground user in the time slot, and determining the energy consumption of the UAV by using a closed analytical boost power consumption model; Step S32: setting the state function, action function and reward function of the drone; Step S33: determining the loss of the drone data in the D3QN neural network according to the error between the predicted Q value and the target Q value, and updating the neural network; Step S34: Calculate the difference between the drone state-action pair according to the drone state function and the drone action function through the IM mechanism, and update the Critic and IM networks accordingly; The state function of the drone is: Where u[t] represents the user location, represents the amount of data remaining to be transmitted, ω[t] represents the position of the drone, ω end represents the final position of the drone, E[t] represents the remaining energy in the current time slot, and E max Indicates the maximum onboard energy. The user's action function is: (+x d ,0,0),(-x d ,0,0),(0,+y d ,0),(0,-y d ,0),(0,0,+z d ), (0,0,-z d ) and (0,0,0), representing left, right, forward, backward, up, down, and hover behaviors respectively. The reward function includes: r2[t]=-κ 2a E[t]-k 2b D sum In the formula, κ i ,i={1,2a,2b,3,4,5} is the reward parameter, D sum Relative to the total amount of data transferred The normalization function of is determined according to the following formula: In the formula, α and β are normalization parameters; The overall reward and punishment design is determined according to the reward function: The closed analytical propulsion power consumption model of the UAV is expressed as follows: In the formula, P0, P i are two constants representing the blade power and induced power in the hovering state, U tip is the tip speed of the rotor blade, v0 and ν are the rotor induced speed and fuselage drag ratio in the hovering state, ρ and s are the air density and rotor disk area, A is the rotor disk area, and V is the speed of the UAV at time slot t.

2. The method for UAV-assisted network path optimization based on deep reinforcement learning according to claim 1 is characterized in that: The step S1 comprises: Step S11: The UAV and the user mobile node select a direction and speed to move from the current location to a new location; Step S12: the user mobile node moves at a constant time interval each time, and calculates a new direction and speed of the user mobile node; Step S13: If the UAV or user mobile node reaches the simulation boundary, it bounces back from the simulation boundary, the rebound angle is determined by the incident direction, and continues to move along this path, repeating the above steps until the end.

3. The method for UAV-assisted network path optimization based on deep reinforcement learning according to claim 1 is characterized in that: In step S2, the joint optimization model is determined according to the following formula: C4:ω[T t ]=ω end In the formula, C1~C3 represent the flight area of ​​the UAV; C4 represents that the UAV needs to fly to the destination in the last time slot; C5 represents the UAV onboard energy limit; C6 represents the user's minimum transmission data limit; C7 represents the user's signal-to-noise ratio limit; c k [t] represents the connection status between the user and the drone; t represents the time slot; k represents the number of users, represents the amount of data transmitted by the drone to the user, u[t] represents the user's location, represents the amount of data remaining to be transmitted, ω[t] represents the position of the drone, ω end represents the final position of the drone, E[t] represents the remaining energy in the current time slot, and E max Indicates the maximum onboard energy.

4. A UAV-assisted network path optimization system based on deep reinforcement learning, characterized in that: include: Behavior simulation module: used to build a typical drone base station downlink scenario and simulate the user's dynamic movement behavior based on the Gauss-Markov model; Optimization model generation module: used to perform system modeling analysis according to the UAV base station downlink scenario, and generate a joint optimization model with the goal of minimizing energy consumption; Model solving module: used to solve the optimal solution of the joint optimization model through a deep reinforcement learning algorithm, and train the drone to learn the optimal solution through a DQN neural network, a dual competitive deep Q network, and an IM mechanism; The model solving module 30 is used to perform the following steps: Step S31: determining the data transmission rate between the UAV and the user according to the transmission power of the UAV and the signal-to-noise ratio between the UAV and the ground user in the time slot, and determining the energy consumption of the UAV by using a closed analytical boost power consumption model; Step S32: setting the state function, action function and reward function of the drone; Step S33: determining the loss of the drone data in the D3QN neural network according to the error between the predicted Q value and the target Q value, and updating the neural network; Step S34: Calculate the difference between the drone state-action pair according to the drone state function and the drone action function through the IM mechanism, and update the Critic and IM networks accordingly; The state function of the drone is: Where u[t] represents the user location, represents the amount of data remaining to be transmitted, ω[t] represents the position of the drone, ω end represents the final position of the drone, E[t] represents the remaining energy in the current time slot, and E max Indicates the maximum onboard energy. The user's action function is: (+x d ,0,0),(-x d ,0,0),(0,+y d ,0),(0,-y d ,0),(0,0,+z d ), (0,0,-z d ) and (0,0,0), representing left, right, forward, backward, up, down, and hover behaviors respectively. The reward function includes: In the formula, κ i ,i={1,2a,2b,3,4,5} is the reward parameter, D sum Relative to the total amount of data transferred The normalization function of is determined according to the following formula: In the formula, α and β are normalization parameters; The overall reward and punishment design is determined according to the reward function: The closed analytical propulsion power consumption model of the UAV is expressed as follows: In the formula, P0, P i are two constants representing the blade power and induced power in the hovering state, U tip is the tip speed of the rotor blade, v0 and ν are the rotor induced speed and fuselage drag ratio in the hovering state, ρ and s are the air density and rotor disk area, A is the rotor disk area, and V is the speed of the UAV at time slot t.

Citation Information

Patent Citations

  • Model construction method and device, task allocation method and device, equipment and medium

    CN113032904A

  • Real-time spectrum resource optimization method and system based on deep reinforcement learning

    CN117768902A