Unmanned aerial vehicle assisted method for joint task scheduling and motion trajectory optimization of mobile edge computing system

By combining task scheduling and trajectory optimization methods in UAV-assisted MEC systems, and utilizing a deep deterministic policy gradient algorithm to optimize the transmission power and flight parameters of UAVs, the problem of limited UAV computing power is solved, and the offloading efficiency of UAV-assisted edge computing systems is improved.

CN116257335BActive Publication Date: 2025-12-19BEIJING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211613821.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-15
Publication Date
2025-12-19
Estimated Expiration
2042-12-15

AI Technical Summary

Technical Problem

Existing drones have limited computing power and cannot meet the offloading requirements of terminal devices for computing tasks. Furthermore, fixed infrastructure is difficult to provide effective services in situations where communication facilities are sparse in the field or in the event of a sudden disaster.

Method used

This paper proposes a joint task scheduling and trajectory optimization method for UAV-assisted MEC systems. The method uses a deep deterministic policy gradient algorithm to jointly optimize the transmission power, task offloading, flight parameters and communication resources of the UAV. Considering the collaborative service of the UAV as an edge computing node and a communication relay node, the DDPG algorithm is used for optimization.

Benefits of technology

It effectively reduces the processing latency of computing tasks on ground terminal equipment and improves the offloading efficiency of UAV-assisted edge computing systems. In simulation experiments, the DDPG algorithm improves the processing latency performance by 20% compared to the DQN algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116257335B_ABST
    Figure CN116257335B_ABST
Patent Text Reader

Abstract

The application discloses a kind of unmanned plane auxiliary MEC system joint task scheduling and motion trajectory optimization method, in fully considering the position change of mobile device and the case of partial offloading of computing task in dynamic environment (ground terminal equipment and unmanned plane are under the condition of dynamic movement), by jointly optimizing user scheduling and computing task offloading mode selection, unmanned plane motion parameters (including flight angle, flight speed) and power distribution, communication, calculation and motion are jointly optimized, effectively reduce the computing task processing delay of terminal device, improve the offloading efficiency of unmanned plane auxiliary edge computing system.And compared with DQN and other baseline algorithms in simulation experiment, it is found that DDPG algorithm has significant improvement in processing delay.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of Internet of Vehicles service privacy protection, and particularly relates to a method for jointly scheduling tasks and optimizing motion trajectories of a UAV-assisted MEC system. BACKGROUND

[0002] With the development of the fifth generation mobile communication technology, new applications with intensive computing and delay sensitivity, such as automatic navigation, face recognition, network games, etc., are emerging rapidly. However, the computing capacity of terminal devices is often low and difficult to handle a large number of computing tasks; the traditional cloud computing mode concentrates computing resources in the cloud, and the intensive computing task access and the cloud away from the terminal device will cause large transmission delay. Mobile edge computing (MEC) can conveniently provide computing services to process terminal intensive computing tasks by deploying computing servers near the network edge close to the user side, thereby effectively reducing the computing processing delay and improving the user's service experience. Therefore, the emergence of the mobile edge computing mode greatly relieves the pressure on network bandwidth and cloud computing centers, and optimizes the response ability of computing and storage services.

[0003] The existing mobile edge computing server / computing center often adopts a fixed deployment mode to provide services for users. However, in the case of sparse communication facilities in the wild and sudden disasters, fixed infrastructure often has difficulty in providing effective services. Unmanned aerial vehicles (UAVs) have the characteristics of flexible maneuvering and easy deployment, and can be used to support ground communication relaying and edge computing services by establishing a line of sight (LOS) connection with ground terminals.

[0004] In the scheme of using UAVs to support ground communication, the UAVs often remain stationary and act as air base stations, which can be applied to scenarios such as infrastructure damage or large-scale event communication traffic burst. In addition, by utilizing the computing capacity of the UAVs, the UAVs can provide computing resources on demand for ground terminal devices, effectively improving the user's service experience.

[0005] Currently, academia and industry are researching on UAV-assisted MEC system. In the literature [1] (T. Ren et al., "Enabling Efficient Scheduling in Large-Scale UAV-Assisted Mobile-Edge Computing via Hierarchical Reinforcement Learning," in IEEE Internet of Things Journal, vol. 9, no. 10, pp. 7095-7109, 15 May 15, 2022, doi: 10.1109 / JIOT.2021.3071531.), for the problem of UAV motion planning and ground terminal device computing offloading scheduling, the authors decompose the scheduling problem into two sub-problems, and use a hierarchical reinforcement learning algorithm to optimize alternately to obtain a real-time scheduling strategy in a dynamic environment.

[0006] However, the prior art mainly considers that the UAV provides communication services as a communication relay node or provides computing services as an edge computing node, and less considers a scheme in which the two service modes are coordinated. Due to the limited computing capability of the UAV, the computing task offloading demand of the terminal device cannot be met. SUMMARY

[0007] The present application proposes a UAV-assisted MEC system joint task scheduling and motion trajectory optimization method to solve the problem that the existing UAV computing capability is limited and cannot meet the computing task offloading demand of the terminal device. The method not only considers the case that the UAV provides computing services as an edge computing node, but also considers the case that the UAV forwards the computing task to the base station side edge computing server as a communication relay node. Then, the user scheduling, computing task offloading selection, UAV flight parameters and communication resource allocation are jointly optimized, which effectively reduces the computing task processing delay of the ground terminal device and improves the offloading efficiency of the UAV-assisted edge computing system.

[0008] To achieve the above purpose, the present application provides the following technical solutions:

[0009] A UAV-assisted MEC system joint task scheduling and motion trajectory optimization method, taking minimizing the completion of all terminal computing task processing as the optimization objective, taking the UAV transmission power, task offloading variable, UAV flight speed, UAV flight angle, UAV self-energy, UAV moving area and terminal device moving area as constraints, and using a deep deterministic policy gradient algorithm to solve the optimization objective.

[0010] Further, the optimization objective and the constraints are expressed as follows:

[0011]

[0012] C1:

[0013] C2:

[0014] C3:

[0015] C4: 0≤β(n)≤2π

[0016] C5:

[0017] C6: {x u (n)∈[0,L], y u (n)∈[0,W]}

[0018] C7: {x m (n)∈[0,L], y m (n)∈[0,W]}

[0019] Wherein, t sum (n) is the total time delay of completing the computing task processing in the time slot T, which is expressed as:

[0020]

[0021] Wherein, is the transmission time delay of the ground terminal device completely offloading the computing task to the unmanned aerial vehicle, is the transmission time delay of the unmanned aerial vehicle partially offloading the computing task to the base station, t mu (n) is the computing task time delay of the unmanned aerial vehicle, t mb (n) is the time required for the base station to complete the computing task on the MEC server side;

[0022] C1 is the constraint on the transmission power P u (n) of the unmanned aerial vehicle, C2 is the constraint on the task offloading variable , C3 is the constraint on the flight speed v u (n) of the unmanned aerial vehicle, C4 is the constraint on the flight angle β(n) of the unmanned aerial vehicle, C5 is the constraint on the energy of the unmanned aerial vehicle itself, and C6 and C7 are respectively the restrictions on the moving areas of the unmanned aerial vehicle and the terminal device. fly = φ||v u (n)| 2 is the flight energy consumption of the unmanned aerial vehicle, wherein φ = 0.5M UAV t fly , M is the mass of the unmanned aerial vehicle, t fly is the flight time, E mu (n) = γ u (fu ) 3 t mu (n), E mu (n) is the local computing energy consumption of the UAV, f u is the computing power of the UAV, γ u is the impact factor of the chip structure on CPU processing, E U is the power of the UAV itself.

[0023] Further, the process of solving the optimization target by using the deep deterministic policy gradient algorithm is:

[0024] The state space is represented as where q(n) represents the position information of the UAV, p m (n) represents the position information of the ground terminal device, D m (n) represents the size of the remaining computing task data, represents the remaining power of the UAV;

[0025] The action space is represented as where v u (n) is the speed of the UAV, β(n) is the flight angle of the UAV, P u (n) is the transmission power of the UAV, is the task offloading variable;

[0026] The reward function is represented as r n = -t sum (n), where the reward function is negative total delay;

[0027] The Q-table is used to record and update the state-action value, i.e., Q(s, a), and the critic and actor networks are used for updating, which is represented as follows:

[0028] θ Q' ← ρθ Q + (1 - ρ)θ Q'

[0029] θ μ' ← ρθ μ + (1 - ρ)θ μ'

[0030] where θ Q is the parameter of the critic neural network, ρ is a constant, and θ μ is the parameter of the actor neural network.

[0031] Further, the critic network is updated as:

[0032] L(θ Q ) = E μ'[(y n -Q(s n ,a n |θ Q )) 2 ]

[0033] wherein E μ' is the mean function, y n is the Q target value, y n =r n +gamma Q(s n+1 , mu(s n+1 )|theta Q ), gamma is a discount factor, r n is a reward function.

[0034] Further, the Actor network is updated as:

[0035]

[0036] wherein, is the policy gradient function, mu is the policy network, and N is the size of random extraction from the experience pool.

[0037] Compared with the prior art, the present application has the following beneficial effects:

[0038] The unmanned aerial vehicle assisted MEC system joint task scheduling and motion trajectory optimization method provided by the present application fully considers the change of mobile device position and the partial offloading of computing tasks in a dynamic environment (the ground terminal device and the unmanned aerial vehicle are both in dynamic movement), jointly optimizes user scheduling and computing task offloading mode selection, unmanned aerial vehicle motion parameters (including flight angle, flight speed) and power allocation, and jointly optimizes communication, computation and motion, effectively reduces the computing task processing delay of the terminal device, and improves the offloading efficiency of the unmanned aerial vehicle assisted edge computing system. Compared with the baseline algorithm such as DQN in the simulation experiment, it is found that the DDPG algorithm has a significant improvement in processing delay. BRIEF DESCRIPTION OF DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can also be obtained by those skilled in the art based on these drawings.

[0040] Figure 1 is a model diagram of an unmanned aerial vehicle assisted edge system.

[0041] Figure 2 is a system architecture diagram of the unmanned aerial vehicle assisted MEC system joint task scheduling and motion trajectory optimization method provided by the embodiment of the present application.

[0042] Figure 3 Comparison results of DDPG and DQN in processing latency.

[0043] Figure 4 Comparison results of DDPG and Local baseline algorithm and Offload baseline algorithm in processing latency. DETAILED DESCRIPTION

[0044] In order to better understand the technical solutions, the method of the present application will be described in detail below with reference to the accompanying drawings.

[0045] 1) System model

[0046] Figure 1 The UAV-assisted MEC system diagram includes a UAV, multiple ground terminal devices (MDs) and a base station. It is considered that the time slot can be divided into T = {1, 2,..., N}. In time slot n, the position coordinates of terminal device m are p m (n) = [x m (n), y m (n), 0] T , and move randomly at a speed v m ; the position coordinates of the UAV are The coordinates of the base station b are l b (n) = [x b (n), y b (n), h] T .

[0047] It is considered that the UAV is deployed in the air, and the transmission between the UAV and the MDs is mainly line-of-sight. The channel gain between the UAV and the MDs can be expressed as:

[0048]

[0049] where g0represents the channel gain at a reference distance of 1 m, is the distance between the UAV and the MDs.

[0050] The channel gain between the UAV and the base station can be expressed as:

[0051]

[0052] where, is the distance between the UAV and the base station.

[0053] In time slot n, the computing task carried by MD m can be expressed as:

[0054] W m (n) = {D m (n), Cm (n),T m (n)}

[0055] where D m (n) is the data size of the computing task, C m (n) is the CPU cycles needed to process each bit of data, T m (n) is the maximum allowed time to process the computing task.

[0056] 2) Communication model

[0057] (1) Communication model between MDs and UAV

[0058] The transmission delay for MDs to offload all computing tasks to UAV is where D m (n) is the data size of the computing task carried by MDs, is the transmission rate between MDs and UAV.

[0059]

[0060] where B u is the available bandwidth for communication between UAV and MDs, P m is the transmission power of MDs, is the white Gaussian noise, P nlos is the non-line-of-sight transmission loss, and f(n) is a binary function, and f(n) e {0, 1}. (f(n) = 0 means that there is no obstruction between UAV and MDs, and f(n) = 1 means that there is an obstruction between UAV and MDs)

[0061] Therefore, the transmission energy consumption can be expressed as:

[0062]

[0063] (2) Communication model between UAV and base station

[0064] The transmission delay for UAV to offload part of the computing task to the base station is:

[0065]

[0066] where is the task offloading variable, is the transmission rate between UAV and base station.

[0067]

[0068] where B k is the available bandwidth for communication between UAV and base station, is the transmission power of UAV, is the maximum transmission power of the UAV, is the white Gaussian noise.

[0069] Therefore, after obtaining the transmission delay between the UAV and the base station, the transmission energy consumption can be expressed as:

[0070]

[0071] (3) UAV movement model

[0072] In time slot n, the UAV flies from q(n) to a new position coordinate, which can be expressed as:

[0073] q(n+1) = [x u (n) + l b (n) cos β(n), y u (n) + l b (n) sin β(n)]

[0074] where l b (n) = v u (n) t fly is the distance flown by the UAV, t fly is the flight time, and β(n) is the flight angle of the UAV.

[0075] Therefore, the flight energy consumption of the UAV can be expressed as:

[0076] E fly (n) = φ||v u (n) || 2

[0077] where φ = 0.5M UAV t fly , and M is the mass of the UAV.

[0078] 3) Calculation model

[0079] First, the UAV calculation task delay is:

[0080]

[0081] where f u is the calculation capability of the UAV, with the unit of cycles per second of CPU.

[0082] At this time, the energy consumption generated by the UAV calculation task is:

[0083] E mu (n) = γ u (f u ) 3 t mu (n)

[0084] wherein, gamma u is the impact factor of the chip structure on CPU processing.

[0085] In addition, the time required for the base station to complete the computing task on the MEC server side is:

[0086]

[0087] wherein, f b is the computing power of the base station, and the unit is the number of CPU cycles per second.

[0088] As Figure 2 shown, in the unmanned aerial vehicle assisted edge computing system, the terminal randomly generates a computing task and transmits it to the unmanned aerial vehicle, which is divided into two processing modes: the unmanned aerial vehicle acts as an edge computing node to complete the offloaded computing task processing, or the unmanned aerial vehicle is regarded as a relay forwarding and jointly completes the computing task processing with the edge node (base station).

[0089] The unmanned aerial vehicle assisted MEC system joint task scheduling and motion trajectory optimization method proposed by the application is as follows.

[0090] Optimization goal:

[0091] Considering completing the computing task processing within time slot T, the total delay can be represented as:

[0092]

[0093] wherein,

[0094] The application takes minimizing the completion of all terminal computing task processing as the goal, considers the joint optimization of ground terminal equipment and computing task offloading selection adjustment, UAV motion parameters (including flight angle, flight speed) and transmission power, and models the joint optimization problem as follows:

[0095]

[0096] C1:

[0097] C2:

[0098] C3:

[0099] C4: 0 <= beta (n) <= 2pi

[0100] C5:

[0101] C6: {x u (n) belongs to [0, L], y u(n)∈[0, W]

[0102] C7:{x m (n)∈[0, L], y m (n)∈[0, W]

[0103] Wherein, C1 is the constraint of the unmanned aerial vehicle transmission power, C2 is the constraint of the task offloading variable, C3 is the constraint of the unmanned aerial vehicle flight speed, C4 is the constraint of the unmanned aerial vehicle flight angle, C5 is the constraint of the unmanned aerial vehicle itself energy, C6 and C7 are the restrictions of the moving area of the unmanned aerial vehicle and the terminal device.

[0104] For the above optimization problem, the non-convex optimization problem is converted into Markov decision process (MDP). Considering that the channel condition and the device state of the system are dynamic and time-varying, further, the deep deterministic policy gradient (DDPG) algorithm is used to solve the problem, which can effectively solve the optimization problem with continuous action space.

[0105] The process of solving the optimization target is as follows:

[0106] The state space is represented as Wherein q(n) represents the position information of the unmanned aerial vehicle, p m (n) represents the position information of the ground terminal device, D m (n) represents the size of the remaining computing task data, represents the remaining power of the unmanned aerial vehicle;

[0107] The action space is represented as Wherein v u (n) is the speed of the unmanned aerial vehicle, β(n) is the flight angle of the unmanned aerial vehicle, P u (n) is the transmission power of the unmanned aerial vehicle, is the task offloading variable;

[0108] The reward function is represented as Wherein the reward function is the negative total delay.

[0109] In traditional reinforcement learning, Q-learning is widely used, which uses Q-table to record and update state-action value, that is, Q(s, a), however, Q-learning is difficult to extract and generalize features from previous experience, and the effect of processing large dimension problem is low. DDPG not only can predict Q value by constantly training from previous experience, but also uses critic and actor network for updating. In the DDPG algorithm, the critic network can be updated as:

[0110] L(θ Q )=E u' [(yn Q(s n , a n ) Q ) 2 ]

[0111] where E μ' is the mean function, y n is the Q target value, y n = r n + γQ(s n+1 , μ(s n+1 )|θ Q ), γ is the discount factor, and r n is the reward function.

[0112] The actor network is updated as:

[0113]

[0114] where is the policy gradient function, and μ is the policy network.

[0115] The DDPG soft update critic and actor target network process is represented as follows:

[0116] θ Q' ← ρθ Q + (1 - ρ)θ Q'

[0117] θ μ' ← ρθ μ + (1 - ρ)θ μ'

[0118] where θ Q is the parameter of the neural network, ρ is a constant, and θ μ is the parameter of the actor neural network.

[0119] The DDPG algorithm is as follows:

[0120] Randomly initialize the critic and actor networks;

[0121] Build the critic and actor target networks;

[0122] Initialize the experience replay pool;

[0123] For episode = 1, M do;

[0124] Initialize a random noise for exploration;

[0125] Initial state S1;

[0126] For t = 1, T do;

[0127] Randomly select an action a according to current policy t ;

[0128] After performing the action, the environment provides a timely reward r t+1 and a new state s t+1 ;

[0129] Add transitions (s t , a t , r t+1 , s t+1 ) to the experience pool

[0130] Randomly sample a mini-batch of transitions (s t , a t , r t+1 , s t+1 ) from the experience pool

[0131] Compute Q-values based on mean squared loss Update critic network parameters

[0132] Update actor network

[0133] Finally, soft update critic and actor target networks

[0134] The application effectively reduces the processing delay of the terminal device and effectively improves the offloading efficiency of the UAV-assisted edge computing system by jointly optimizing user scheduling and computing task offloading mode selection, UAV motion parameters (including flight angle, flight speed) and power allocation while fully considering the changes in the location of the mobile device and the partial offloading of the computing task in the dynamic environment. Compared with the baseline algorithm DQN in the simulation experiment, it is found that the DDPG algorithm has a significant improvement in processing delay. The comparison results of DDPG and DQN in processing delay (as shown in Figure 3 ) show that DDPG is improved by nearly 20% compared with DQN. As shown in Figure 4 , compared with the Local baseline algorithm (computing task only in local computing), DDPG is improved by nearly 60% in processing delay, and compared with the Offload baseline algorithm (computing task is completely offloaded to the base station for edge computing), DDPG is improved by nearly 34% in processing delay.

[0135] The above examples are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that the technical solutions recorded in the foregoing examples can be modified, or some technical features can be replaced by equivalent features, but these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1.A method for joint task scheduling and motion trajectory optimization of UAV-assisted MEC system, characterized in that, Taking minimizing completing all terminal computing task processing as an optimization target, taking the unmanned aerial vehicle transmission power, the task offloading variable, the unmanned aerial vehicle flight speed, the unmanned aerial vehicle flight angle, the unmanned aerial vehicle self energy, the unmanned aerial vehicle movement area and the terminal device movement area as constraints, and using a deep deterministic policy gradient algorithm to solve the optimization target; The optimization target and the constraint are expressed as follows: , where t sum (n) is the total delay of the computation task processing in the time slot is expressed as: ‘ wherein, ; a transmission delay for the ground terminal device to unload all computing tasks to the unmanned aerial vehicle, a transmission delay for the unmanned aerial vehicle to unload part of the computing task to the base station, a computing task delay for the unmanned aerial vehicle, a time required for the base station to complete the computing task on the MEC server side; represent the position information of the unmanned aerial vehicle; Transmission delay of the ground terminal device unloads all computing tasks to the unmanned aerial vehicle is represented as: ’ wherein is the size of the computing task data carried by the ground terminal device, is the transmission rate between the ground terminal device and the unmanned aerial vehicle, expressed as: ‘ wherein is the available bandwidth for the communication between the drone and the ground terminal device, is the transmission power of the ground terminal devices MDs, is the white Gaussian noise, is the non-line-of-sight transmission loss, is a binary function, and , denotes that there is no obstruction between the drone and the ground terminal device, denotes that there is an obstruction between the drone and the ground terminal device; the channel gain between the drone and the ground terminal devices MDs is denoted by ’ wherein g0represents a channel gain at a reference distance of 1 m, is a distance between the UAV and the ground terminal device MD, represents position information of the UAV, represents position information of the ground terminal device MD. Transmission delay of a drone offloading part of the computing task to a base station is represented as: ‘ wherein is a task offloading variable, is the transmission rate between the UAV and the base station, denoted as: ’ wherein is the available bandwidth for the communication between the UAV and the base station, is the transmission power of the UAV, is the maximum transmission power of the UAV, is the white Gaussian noise; the channel gain between the UAV and the base station is expressed as: ‘ wherein g0represents a channel gain at a reference distance of 1 m, is a distance between the UAV and the base station, represents position information of the UAV, and l(n) is a coordinate of the base station; Unmanned aerial vehicle computing task latency is represented as: ’ wherein C is the computing power of the drone in cycles per second of CPU; C m (n) is the CPU cycles required to process each bit of data; Time required for the base station to deploy the MEC server side to complete the computing task is represented as: ‘ wherein, is the computing power of the base station in cycles per second of CPU; C1 is the constraint on the transmission power of the UAV C2 is the constraint on the task offloading variable C3 is the constraint on the flight speed of the UAV C4 is the constraint on the flight angle of the UAV C5 is the constraint on the energy of the UAV itself, and C6 and C7 are the limits on the moving areas of the UAV and the terminal device, respectively; is the flight energy consumption of the UAV, wherein, , is the mass of the UAV, is the flight time, , is the local computing energy consumption of the UAV, is the computing capability of the UAV, is the impact factor of the chip structure on CPU processing, is the self-contained power of the UAV; The process of solving the optimization target by using the deep deterministic policy gradient algorithm is as follows: State space is represented as wherein represents position information of the UAV, represents position information of the ground terminal device D r (n) represents the size of the remaining computing task data amount, represents the remaining power of the UAV; The action space is represented as wherein is the speed magnitude of the UAV, is the flight angle of the UAV, is the transmission power of the UAV, is the task offloading variable; The reward function is represented as where the reward function is the negative total latency; The Q-table is used to record and update state-action values, i.e. with critic and actor networks updated simultaneously, represented as follows: ’ wherein, are parameters of the critic neural network, is a constant, are parameters of the actor neural network. 2.The method of claim 1, wherein, The critic network update is as follows: ‘ wherein, is the mean function, is the Q target value, , is the discount factor, is the reward function. 3.The method of claim 1, wherein, The actor network update is as follows: ’ where, For the policy gradient function, μ is the policy network, and N is the size of the random samples drawn from the experience pool.

Citation Information

Patent Citations

  • Scheduling optimization method and system for unmanned aerial vehicle assisted mobile edge computing

    CN114169234A