Track and resource allocation joint optimization method in unmanned aerial vehicle edge computing network
By constructing a mobile edge computing network system model for UAVs and combining deep reinforcement learning and phase alignment methods, the UAV's 3D trajectory and resource allocation are optimized, solving the problems of unstable communication and low energy efficiency in UAV-MEC systems, and achieving efficient resource management and maximum energy efficiency.
Patent Information
- Application Number
- CN202511089379.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-21
AI Technical Summary
Unmanned aerial vehicle (UAV) MEC systems face challenges such as unstable communication links, low energy efficiency, security threats, and adaptability to dynamic environments. Traditional methods are insufficient to optimize the offloading ratio of UAV tasks, allocation of computing resources, and three-dimensional flight trajectory, resulting in the inability to meet the requirements of low power consumption and high real-time performance.
A mobile edge computing network system model for unmanned aerial vehicles (UAVs) is constructed. Combining deep reinforcement learning algorithms and phase alignment methods, the UAV's 3D trajectory, intelligent reflector phase shift matrix, and resource allocation are optimized through Markov decision processes. The R-DDQN algorithm is designed to solve the joint optimization problem.
It maximizes the energy efficiency of UAVs, optimizes the three-dimensional flight trajectory and resource allocation of UAVs, improves communication stability and energy efficiency, reduces energy consumption, and meets the requirements of low power consumption and high real-time performance.
Smart Images

Figure CN120994377A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of UAV trajectory optimization and resource allocation technology, specifically relating to a joint optimization method for trajectory and resource allocation in UAV edge computing networks. Background Technology
[0002] Due to their flexibility and wide coverage, unmanned aerial vehicles (UAVs) have been used to assist mobile edge computing (MEC) systems in performing computationally intensive tasks. By establishing line-of-sight (LOS) links with ground users, UAVs can act as "flying MEC servers," providing a large number of offloaded services with lower network overhead and execution latency. However, UAV-MEC systems still face challenges such as unstable communication links, low energy efficiency, security threats, and adaptability to dynamic environments.
[0003] To address these challenges, reconfigurable smart surface technology has been introduced into UAV-MEC systems. In complex urban areas or areas with dense obstacles, reconfigurable smart surfaces can bypass non-line-of-sight paths, enhance signal strength and stability between the UAV and ground terminals, reduce communication power consumption, and extend the UAV's endurance. However, jointly optimizing the UAV's task offloading ratio, computing resource allocation, 3D flight trajectory, and the phase shift matrix of the Reconfigurable Intelligence Surface (RIS) faces a highly complex non-convex optimization problem, making it difficult to meet the requirements of low power consumption and high real-time performance using traditional methods. Summary of the Invention
[0004] To address the problems existing in the prior art, this invention proposes a joint optimization method for trajectory and resource allocation in UAV edge computing networks. The method includes constructing a UAV mobile edge computing network system model, establishing a multivariate joint optimization model with the goal of maximizing UAV energy efficiency, modeling the optimization problem through a Markov decision process, solving the problem by combining deep reinforcement learning algorithms and phase alignment methods to obtain a joint optimization solution, and iteratively updating the UAV's three-dimensional trajectory, intelligent reflector phase shift matrix, and resource allocation results.
[0005] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0006] This invention proposes a joint optimization method for trajectory and resource allocation in UAV edge computing networks based on deep reinforcement learning. First, the UAV mobile edge computing network system is modeled, including its communication model, computation model, and energy consumption model. Then, a joint optimization problem is proposed, involving three-dimensional flight trajectory, computation offloading, resource allocation decision, and RIS phase shift, to maximize UAV energy efficiency. Compared with existing baseline methods, this invention uses UAVs as airborne base stations and airborne relay stations, and also considers applying intelligent reflective surfaces to the UAV network system to rationally schedule UAV resources and make real-time optimal three-dimensional flight trajectory and resource allocation schemes. Attached Figure Description
[0007] Figure 1 This is a flowchart of the joint optimization method for trajectory and resource allocation in the UAV edge computing network based on deep reinforcement learning in this invention;
[0008] Figure 2 This is a schematic diagram of the intelligent reflective surface-assisted UAV mobile edge computing network system model in this invention;
[0009] Figure 3 The three-dimensional flight trajectory of the drone;
[0010] Figure 4 A graph showing the relationship between the number of users and drone throughput;
[0011] Figure 5 This is a graph showing the relationship between the number of users and the energy efficiency of drones. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] This invention proposes a joint optimization method for trajectory and resource allocation in UAV edge computing networks, such as... Figure 1 This includes constructing a mobile edge computing network system model for UAVs, establishing a multivariate joint optimization model with the goal of maximizing UAV energy efficiency, modeling the optimization problem through a Markov decision process, solving the problem by combining deep reinforcement learning algorithms and phase alignment methods to obtain the joint optimization solution, and iteratively updating the UAV's 3D trajectory, intelligent reflector phase shift matrix, and resource allocation results.
[0014] like Figure 2The diagram illustrates the UAV mobile edge computing network system model constructed in this invention. Specifically, considering an 800m × 800m UAV flight area, the communication scenario includes a UAV equipped with an MEC server, intelligent reflectors fixedly mounted on the exterior surface of buildings, wireless access points (APs), and multiple ground users. This system uses time-division multiple access to avoid interference between different ground users during task offloading. The UAV flight time... Divided equally There are 1 time slot, and the duration of each time slot is 1. ,Right now In a three-dimensional Cartesian coordinate system, the horizontal position of the UAV in the t-th time slot is represented as: The drone's flight altitude is ,in High level, The number of high-level categories. This represents the horizontal difference for each altitude level; where the ground computing nodes are deployed in the coordinate system. The location of RIS is... RIS surface contains One reflector unit assists in UAV communication; in addition, ground users are randomly distributed in and Within the range, the drone's initial flight position is (0,0,80).
[0015] The construction of the UAV communication model includes: due to the altitude between the UAV and the smart reflector, the links between the ground user and the UAV (KU link), the UAV and the wireless access point (UA link), and the smart reflector and the wireless access point (RA link) are all line-of-sight links, and therefore all adopt the free-space path loss model; without loss of generality, the link between the ground user and the smart reflector (UR link) adopts the Rician fading channel model, which includes a Loss component and a non-line-of-sight (NLOS) component. The expression of the UAV communication model is:
[0016]
[0017]
[0018]
[0019]
[0020]
[0021]
[0022]
[0023]
[0024] in, Let be the channel gain from ground user k to UAV in the t-th time slot; For reference distance Path loss during the process; Indicates the location of ground user k. Let x be the x-coordinate of ground user k. Let k be the ordinate of the ground user. Represents the square of the Euclidean distance; This represents the flight altitude of the UAV in the t-th time slot;
[0025] This represents the channel gain from UAV to RIS in the t-th time slot; Rician factor, visible distance component , Indicates the angle of arrival of the signal in the t-th time slot UR link. The cosine value, Indicates the carrier wavelength. This indicates the distance between RIS reflective elements, the non-line-of-sight portion. It follows a complex Gaussian distribution with zero mean and unit variance;
[0026] Let represent the channel gain from RIS to AP in the t-th time slot; Indicates the location of the wireless access point. The x-coordinate of the wireless access point. The vertical coordinate of the wireless access point; Represents the height coordinates of the RIS; Indicates the angle of arrival of the RA link signal. The cosine value;
[0027] This represents the channel gain from UAV to AP in the t-th time slot;
[0028] Let be the diagonal phase shift matrix for the t-th time slot; This represents the phase shift of m RIS reflector units in the t-th time slot;
[0029] Calculate the relay transmission rate from the UAV to the AP for ground user k; For transmission bandwidth to ground users; Power for drone relay transmission; Noise power; Calculate the unloading ratio for drone missions. It is a decimal between 0 and 1; To calculate the length of the task.
[0030] Let be the transmission delay of the task of ground user k in the t-th time slot, which is relayed from the UAV to the AP.
[0031] The computational model includes: the computational latency and energy consumption of the computational task on the UAV; furthermore, the relay transmission process from the UAV to the AP and the computation process on the UAV are processed in parallel. The expression for the UAV computational model is:
[0032]
[0033]
[0034]
[0035] in, Let be the computational delay of the UAV terminal in the t-th time slot; The latency for processing tasks for ground users in the t-th time slot; Let be the computational energy consumption of the UAV terminal in the t-th time slot; The number of CPU cycles required to calculate a 1-bit task; For computing resources; The capacitance coefficient is determined by the chip structure; Let $t$ be the unloading delay from the ground user $k$ to the UAV within the $t$-th time slot. Let be the transmission delay of the task of ground user k in the t-th time slot, which is relayed from the UAV to the AP.
[0036] The construction of the UAV energy consumption model includes: the energy consumption for task relay from the UAV to the AP; simultaneously, the energy used by the UAV for flight also accounts for a large proportion. The expression for the UAV energy consumption model is:
[0037]
[0038]
[0039]
[0040]
[0041]
[0042] in, This indicates the instantaneous horizontal flight speed of the drone; Indicates the horizontal position of the drone; Indicates the duration of each time slot;
[0043] Let be the flight power of the UAV in the t-th time slot; This represents the drag power generated by the rotor of a drone when it is stationary due to air resistance. This indicates the tip rotation speed of the drone's rotor blades; This refers to the induced power required for a drone to overcome gravity when hovering or flying. This indicates the average speed of the induction rotor during flight; This indicates the drone's vertical flight power. This indicates the instantaneous vertical flight speed of the drone; Indicates the fuselage drag ratio; Indicates air density; Indicates rotor hardness; This represents the area of the circular region covered by the rotor when it rotates;
[0044] Let be the relay transmission energy consumption of the UAV terminal in the t-th time slot;
[0045] Let be the flight energy consumption of the UAV in the t-th time slot;
[0046] Let be the total energy consumption of the UAV terminal in the t-th time slot.
[0047] Based on the above, in this embodiment, the expression for maximizing the energy efficiency of the UAV is:
[0048]
[0049] Constraints: C1:
[0050] C2:
[0051] C3:
[0052] C4:
[0053] C5:
[0054] C6:
[0055] C7:
[0056] Among them, constraint C1 represents the task offloading ratio constraint; constraint C2 is the UAV computing resource constraint range constraint; constraint C3 is the RIS phase shift range constraint; constraint C4 is the task maximum latency limit constraint; constraints C5 and C6 are UAV position constraints; and constraint C7 is the UAV flight speed limit constraint.
[0057] This invention models optimization problems using Markov decision processes, transforming complex dynamic optimization problems into learnable sequential decision problems. This process rationally designs the following three elements:
[0058] (1) State-space design At the beginning of each time slot, the drone interacts with the environment and uses the observed environmental variables as its state, represented as follows: ;
[0059] (2) Action space design The set of actions that a drone can take, based on the observed state. Select Action , is represented as: ;
[0060] in, This refers to the horizontal movement distance of the drone; The vertical movement distance of the drone;
[0061] (3) Reward function design The quality of an action taken within a given time slot is measured by the following expression:
[0062]
[0063] in, Let represent the reward function for the t-th time slot; , The penalty coefficient is, and , It is a positive integer.
[0064] This embodiment provides a deep learning algorithm based on the R-DDQN algorithm, specifically including the following steps:
[0065] 101. Initialize the q-network weights ω and the target network weights Create an empty experience pool. ;
[0066] 102. The environmental state observed by the UAV in each time slot t is s. t ;
[0067] 103. Q-networks are based on the current state st Estimate Q(s) for each action t ,a t ,ω), using the ε-greedy strategy to select action a t ;
[0068] 104. Determine the optimal phase shift based on the current position of the UAV using the phase alignment method. ;
[0069] 105. Drone execution actions a t Afterwards, environmental feedback reward r t And continue to observe the next state s t+1 ;
[0070] 106. The obtained experience tuple (s) t ,a t ,r t , s t+1 Store in the experience pool This breaks the temporal correlation of data;
[0071] 107. From the experience pool Random sampling Based on this experience, the target network calculates the target Q-value;
[0072] 108. By calculating the loss function To update the q-network weights ω;
[0073] 109. Each Replace target network weights step by step ←ω;
[0074] 110. Repeat the steps of environment interaction, experience playback, network training, and target network synchronization until the predetermined number of training iterations is reached or the convergence condition is met.
[0075] The loss function in the R-DDQN algorithm acts on the parameters of the q-network. The q-network is updated via backpropagation, based on a set of empirical (s) i ,a i ,r i , s i+1 The process of calculating the loss function includes:
[0076]
[0077] Where Q(s) i ,a i ,ω) represents the state s when the network weight of the q-network is ω. i Choose action a i The Q value; γ is the discount factor; Indicates the target network weights are At that time, in state s i+1 Select action Q value; This indicates that when the network weights of the q-network are ω, in state s i+1 Choose the action that maximizes the Q value.
[0078] By controlling the RIS phase shift matrix, the reflection link and the direct link are phase-aligned to obtain the optimal reflection phase, as expressed by:
[0079]
[0080] in, Let m be the channel gain of the UR link for the m-th element. Let be the channel gain of the m-th element RA link; mod[ ] represents the modulo operation.
[0081] Figure 3 The three-dimensional flight trajectory of the UAV is shown when there are 6 ground users. The results indicate that in the method proposed in this invention, the UAV typically chooses the shortest flight path, striving to approach the ground user terminal and RIS to improve channel transmission conditions and reduce flight energy consumption. Compared to the method proposed in this invention, the other two benchmark methods result in a significant increase in the UAV's flight path, tending towards the access node.
[0082] Figure 4 The throughput of UAVs varied with the number of ground users under different benchmark methods. The results show that the throughput of UAVs gradually increases with the increase in the number of ground users. Meanwhile, the throughput of the method proposed in this invention is significantly better than other benchmark methods. This is because this method can effectively adjust the optimal position of the UAV, RIS phase, and the allocation of computing resources, and find a suitable task offloading ratio.
[0083] Figure 5 The energy efficiency of unmanned aerial vehicles (UAVs) was compared with the number of ground users under different benchmark methods. The results show that the energy efficiency of UAVs gradually increases with the increase of the number of ground users. Compared with other benchmark methods, the method proposed in this invention exhibits the best energy efficiency performance.
[0084] In summary, this invention models a mobile edge computing network system for unmanned aerial vehicles (UAVs), including its communication, computation, and energy consumption models. It proposes a joint optimization problem involving 3D flight trajectory, computational offloading, resource allocation decisions, and RIS phase shifting to maximize UAV energy efficiency. For the optimization problem, an R-DDQN algorithm is designed to effectively solve the objective optimization model by combining deep reinforcement learning algorithms with phase alignment methods. Finally, the proposed method is validated through various simulations, including UAV 3D flight trajectory, throughput, and energy efficiency.
[0085] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A joint optimization method for trajectory and resource allocation in an unmanned aerial vehicle (UAV) edge computing network, characterized in that, A mobile edge computing network system model for unmanned aerial vehicles (UAVs) is constructed. A multivariate joint optimization model is established with the goal of maximizing UAV energy efficiency. The optimization problem is modeled using a Markov decision process. The joint optimization solution is obtained by combining deep reinforcement learning algorithms and phase alignment methods. The UAV's 3D trajectory, intelligent reflector phase shift matrix, and resource allocation results are iteratively updated.
2. The method for joint optimization of trajectory and resource allocation in an unmanned aerial vehicle (UAV) edge computing network according to claim 1, characterized in that, The UAV mobile edge computing network system model includes a UAV equipped with an MEC server, a smart reflector fixedly installed on the exterior surface of a building, a wireless access point, and K ground users. The smart reflector includes M reflective elements to form a uniform linear array. The ground users offload all computing tasks to the UAV via line-of-sight links. Considering power consumption and computing resource limitations, the UAV chooses to offload some computing tasks to the wireless access point via line-of-sight links and to offload them via reflective links assisted by the smart reflector.
3. The method for joint optimization of trajectory and resource allocation in an unmanned aerial vehicle (UAV) edge computing network according to claim 1, characterized in that, The multivariate joint optimization model is expressed as: ; Constraints: ; ; ; ; ; ; ; in, Represents the optimization variable; This represents the horizontal position of the UAV in the t-th time slot; This represents the flight altitude of the UAV in the t-th time slot; This represents the offloading ratio of user k computational tasks handled by the UAV in the t-th time slot; This represents the computing resources used to process the user k-end unloading task in the t-th time slot; This represents the phase shift of the reflecting unit of the m-th intelligent reflecting surface in the t-th time slot; N represents the number of time slots in the UAV's flight. This represents the throughput of the UAV in the t-th time slot; This represents the total energy consumption of the UAV in the t-th time slot; This indicates the maximum computing resources available to the drone. This represents the processing latency of the user's task at time slot k in the t-th time slot; This represents the maximum user task processing latency in the t-th time slot; Indicates the maximum flight altitude of the drone; Indicates the minimum flight altitude of the drone; This represents the maximum flight coordinates of the drone along the horizontal axis. This represents the maximum flight coordinates of the drone along the horizontal vertical axis. This represents the instantaneous horizontal flight speed of the UAV in the t-th time slot; This indicates the maximum instantaneous flight speed of the drone in the horizontal direction; This represents the instantaneous vertical flight speed of the UAV in the t-th time slot; This indicates the maximum instantaneous flight speed of the drone in the vertical direction.
4. The method for joint optimization of trajectory and resource allocation in an unmanned aerial vehicle (UAV) edge computing network according to claim 3, characterized in that, Total energy consumption of the drone in the t-th time slot Represented as: ; ; ; ; Where K represents the number of ground users; This represents the computational energy consumption of the UAV in the t-th time slot; The capacitance coefficient of the UAV u is represented; This represents the number of CPU cycles required for the drone to compute 1 bit of a task for a ground user k terminal. This represents the relay transmission energy consumption of the UAV terminal in the t-th time slot; This indicates the energy consumption constraints for drone flight; This represents the flight energy consumption of the UAV in the t-th time slot; This represents the relay transmission power of the UAV in the t-th time slot; This represents the transmission delay of the task of ground user k in the t-th time slot, from the UAV relay to the wireless access point. This represents the offloading ratio of tasks handled by the drone for ground user k terminals; This represents the length of the computation task generated by ground user k in the t-th time slot; This represents the relay transmission rate from the UAV to the wireless access point for the ground user k's computation task in the t-th time slot; This represents the flight power of the UAV in the t-th time slot; This indicates the duration of each time slot.
5. The method for joint optimization of trajectory and resource allocation in an unmanned aerial vehicle (UAV) edge computing network according to claim 3, characterized in that, Flight power of the UAV in the t-th time slot express: ; in, This represents the drag power generated by the rotor of a drone when it is stationary due to air resistance. This refers to the induced power required for a drone to overcome gravity when hovering or flying. This indicates the drone's vertical flight power. This indicates the tip rotation speed of the drone's rotor blades; This indicates the average speed of the induction rotor during flight; This indicates the instantaneous vertical flight speed of the drone; Indicates the fuselage drag ratio; Indicates air density; Indicates rotor hardness; This indicates the area of the circular region covered by the rotor when it rotates.
6. The method for joint optimization of trajectory and resource allocation in an unmanned aerial vehicle (UAV) edge computing network according to claim 3, characterized in that, The throughput of the UAV in the t-th time slot Represented as: ; ; ; in, Let be the transmission delay of the task of ground user k in the t-th time slot, which is relayed from the UAV to the wireless access point. Calculate the relay transmission rate from the UAV to the wireless access point for ground user k. For transmission bandwidth to ground users; Power for drone relay transmission; This indicates the channel gain from the drone to the wireless access point; This represents the channel gain from the drone to the smart reflector; This represents the channel gain from the smart reflector to the UAV; Let be the diagonal phase shift matrix for the t-th time slot; Noise power; This represents the length of the computation task generated by ground user k in the t-th time slot.
7. The method for joint optimization of trajectory and resource allocation in an unmanned aerial vehicle (UAV) edge computing network according to claim 3, characterized in that, The latency of ground user k-end task processing in the t-th time slot Represented as: ; ; in, Let $t$ be the unloading delay from the ground user $k$ to the UAV within the $t$-th time slot. Let be the transmission delay of the task of ground user k in the t-th time slot, which is relayed from the UAV to the wireless access point. Let be the computational delay of the UAV terminal in the t-th time slot; This represents the offloading ratio of tasks handled by the drone for ground user k terminals; This represents the length of the computation task generated by ground user k in the t-th time slot; This represents the number of CPU cycles required to perform a task that does not compute 1 bit at the ground user k end; This represents the computing resources used to process the unloading task of ground user k in the t-th time slot.
8. The method for joint optimization of trajectory and resource allocation in an unmanned aerial vehicle (UAV) edge computing network according to claim 3, characterized in that, Modeling optimization problems using Markov decision processes includes: At the beginning of each time slot, the drone interacts with the environment and takes the observed environmental variables as its state. The state space is represented as follows: ; in, This represents the state of the t-th time slot; This indicates the position of the UAV in the t-th time slot; This represents the length of the computation task generated by ground user k in the t-th time slot; The set of actions that a drone can take, based on the observed state. Select Action The action space is represented as: ; in, This refers to the horizontal movement distance of the drone; The vertical movement distance of the drone; The reward function is used to measure the quality of the action taken in this time slot. The reward function is expressed as: in, Let represent the reward function for the t-th time slot; , The penalty coefficient is, and , It is a positive integer.
9. The method for joint optimization of trajectory and resource allocation in an unmanned aerial vehicle (UAV) edge computing network according to claim 7, characterized in that, By controlling the phase shift matrix of the reflecting unit in the smart reflector using a phase alignment method, the reflection link is phase-aligned with the direct reflection link to obtain the optimal reflection phase. The expression for the optimal reflection phase is: ; in, This represents the optimal phase of the m-th reflecting unit; This represents the channel gain from the drone u to the wireless access point; Let be the channel gain of the link between the ground user and the m-th reflection unit in the t-th time slot. Let be the channel gain of the link between the t-th time slot, the m-th reflection unit, and the wireless access point.