Unmanned aerial vehicle cooperative three-layer mobile edge computing deployment algorithm based on DDPG
By building a three-layer architecture mobile edge computing system and using DDPG algorithm to optimize drone deployment and task offloading, the problem of limited computing resources is solved, efficient utilization and low latency of computing resources is achieved, the base station service scope is expanded, and the low latency and high real-time requirements of VR/AR and other applications are met.
Patent Information
- Application Number
- CN202510085253.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-07-08
AI Technical Summary
The existing drone computing resources are limited, which is difficult to meet the low-latency needs of VR/AR and other applications. In addition, traditional scheduling algorithms have problems of uneven allocation of computing resources and backlog of tasks in multi-drone collaboration, which cannot effectively expand the base station service scope and cannot meet the needs of low-latency and high real-time.
Build a three-layer architecture mobile edge computing system, automatically learn drone deployment, task offload decisions and computing resource allocation through DDPG algorithm, combine drone collaborative offload mechanism, optimize computing resource allocation and task allocation, expand the base station service scope, and realize the drone flexibly adjusts its position in the air as a communication relay.
It effectively solves the problem of uneven allocation of computing resources, improves the utilization rate of computing resources, ensures low latency and high real-timeness, expands the service scope of the base station, and meets the needs of VR/AR and other applications.
Smart Images

Figure CN120282165A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of edge computing, UAV cooperative communication and intelligent optimization, and particularly relates to a UAV cooperative three-layer mobile edge computing deployment algorithm based on DDPG. Background Art
[0002] With the rapid development of 5G communication technology, the application and popularization of frontier technologies such as virtual reality (VR) and augmented reality (AR) have been greatly supported and guaranteed. Technologies such as VR / AR highly rely on low-latency communication to ensure real-time interaction and immersive experience. Therefore, how to optimize the time delay of these applications, especially in a mobile environment, is crucial for ensuring their efficient operation. However, existing user devices are usually limited by limited computing power and are difficult to meet the strict requirements of VR / AR applications for low latency. In this context, edge computing, by integrating computing resources to the network edge, close to the end user, can effectively reduce latency and provide real-time computing capabilities, which is a feasible solution. The computing resources of edge computing are generally concentrated in base stations, and the computing capabilities of these base stations are usually far beyond those of user devices. Therefore, when a user is within the service range of a base station, the user can offload computing tasks to the base station to meet the requirements of low latency. However, due to the fixed location of the base station and the geographical volatility of user demands, many users cannot always be within the service range of the base station, thus unable to obtain services in a timely manner and unable to meet the low-latency requirements.
[0003] However, in practical applications, the computing resources of UAVs are relatively limited compared with base stations. Especially when a single UAV needs to provide computing services for multiple users, problems such as task backlog and increased latency often occur, resulting in the inability to meet the low-latency requirements. At the same time, other UAVs may waste computing resources due to fewer task receptions or being in an idle state, thus leading to the problem of uneven load. Therefore, how to achieve cooperation and resource sharing among UAVs and optimize the allocation of computing resources has become the key to solving this problem.
[0004] In addition, multi-UAV cooperation faces a complex and changeable environment and requires rapid response, which poses a high challenge to traditional scheduling algorithms. Traditional scheduling algorithms are usually limited by factors such as stability, computational complexity, and real-time performance, and it is difficult to achieve efficient cooperation and optimization. Summary of the Invention
[0005] Aiming at the deficiencies in the existing technologies, the purpose of the present invention is to provide a UAV collaborative three-layer mobile edge computing deployment algorithm based on DDPG, which can effectively expand the service range of the base station, meet the requirements of low latency and high real-time performance for applications such as VR / AR; effectively solve the problem of uneven distribution of computing resources in the system and improve the utilization rate of computing resources; and optimize the task allocation and UAV position deployment in real time, thereby improving the collaborative efficiency and overall performance of the system. To achieve the above objects and other advantages according to the present invention, a UAV collaborative three-layer mobile edge computing deployment algorithm based on DDPG is provided, including:
[0006] Construct a three-layer architecture mobile edge computing system, and the three-layer architecture mobile edge computing system includes users, base stations and UAVs;
[0007] Regard the three-layer architecture mobile edge computing system as a Markov decision process, and automatically learn the optimal UAV deployment, task offloading decision and computing resource allocation through the DDPG algorithm;
[0008] Through the collaborative offloading mechanism between UAVs, when tasks can be offloaded between UAVs and there is a communication barrier between the user and the base station, tasks are offloaded through relay, further optimizing the latency and energy efficiency of the three-layer architecture mobile edge computing system.
[0009] Preferably, by flexibly adjusting the position of the UAV in the air to act as a communication relay between the user and the base station, helping the user offload tasks to the base station; the UAV can also directly provide computing services for the user, further improving the task processing speed and ensuring low latency and real-time performance. Combining the advantages of UAVs and edge computing can effectively expand the service range of the base station and meet the requirements of low latency and high real-time performance for applications such as VR / AR.
[0010] Preferably, through a three-layer mobile edge computing system based on UAV cooperation, which coordinates the task allocation between UAVs, optimizes the resource scheduling among users, UAVs and base stations, effectively solves the problem of uneven distribution of computing resources in the system, and improves the utilization rate of computing resources.
[0011] Preferably, based on the solution of deep reinforcement learning, combining the advantages of deep learning and reinforcement learning enables the UAV to automatically learn and optimize the decision-making strategy in a dynamic environment, and optimize the task allocation and UAV position deployment in real time, thereby improving the collaborative efficiency and overall performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a diagram of a UAV-assisted three-layer architecture mobile edge computing system for the UAV collaborative three-layer mobile edge computing deployment algorithm based on deep deterministic policy gradient according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0013] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0014] Referring to Figure 1 , a drone collaborative three-layer mobile edge computing deployment algorithm based on DDPG includes: constructing a three-layer architecture mobile edge computing system, and the three-layer architecture mobile edge computing system includes users, base stations and drones;
[0015] Regarding the three-layer architecture mobile edge computing system as a Markov decision process, and automatically learning the optimal drone deployment, task offloading decision and computing resource allocation through the DDPG algorithm;
[0016] Through the collaborative offloading mechanism between drones, when drones can offload tasks from each other and there are communication obstacles between users and base stations, tasks can be offloaded through relays, further optimizing the delay and energy efficiency of the three-layer architecture mobile edge computing system.
[0017] Embodiment 1
[0018] In the three-layer architecture mobile edge computing system, users generate computing tasks, which can be processed locally or offloaded to base stations or drones. Each drone is equipped with a computing server, and the base station is equipped with an MEC server to provide computing services for users. Assume that the number of users, drones and base stations in the system are U, A and C respectively. The user set is The drone set is The base station set is
[0019] The tasks generated by users vary in terms of data size, CPU cycles required for processing tasks, and maximum tolerable time delay. Therefore, we describe the task generated by user i as where S i represents the task size, C i represents the CPU cycles required to compute one bit of the task, and T i represents the maximum tolerable time delay of this user's task.
[0020] Adopt the binary offloading method to minimize the task processing delay. Thus, the computing tasks of users are regarded as indivisible entities, and the binary variables a ij ∈{0,1} and d ikLet \(x_{i}\in\{0,1\}\) represent the offloading processes of user \(i\) to the UAV and from user \(i\) to the base station, where \(x_{i} = 1\) means the user's task is offloaded, otherwise it is not offloaded.
[0021]
[0022] Among them, \(\mathcal{U}\) represents the set of UAVs, \(\mathcal{B}\) represents the set of base stations. For the computing tasks at the UAVs, they can be executed locally or offloaded to the base stations, and the binary offloading strategy is also adopted. Define a binary variable \(c_{jk}\) jk \(\in\{0,1\}\) to represent whether the computing task of UAV \(j\) is offloaded to base station \(k\), where \(c_{jk}=1\) means the task of the UAV is offloaded, otherwise it is not offloaded. In addition to offloading to the base station, the UAV can also offload the computing task to other UAVs. The binary variable \(b_{jj'}\) jj′ \(\in\{0,1\}\) is used to represent whether the computing task of UAV \(j\) is offloaded to UAV \(j'\), where \(b_{jj'}=1\) means the computing task of UAV \(j\) is offloaded to UAV \(j'\), otherwise it is not offloaded.
[0023] The latency of local computing for user \(i\) is:
[0024]
[0025] Where \(C\) i is the CPU cycles required to compute one bit of the user's task, \(S\) i is the size of the user's task, \(f\) i is the CPU frequency of the user's local computing. The energy consumed by user \(i\) in local computing is:
[0026]
[0027] Where \(k\) is a constant depending on the chip architecture of the user device.
[0028] The size of the task that the UAV (taking UAV \(j\) as an example) needs to execute includes the size of the task offloaded by the user to this UAV minus the size of the task offloaded by this UAV to other UAVs and then minus the size of the task offloaded by this UAV to the base station and plus the size of the task offloaded by other UAVs to this UAV Therefore, the size of the task that UAV \(j\) needs to execute is:
[0029]
[0030] Where the size of the tasks offloaded by all users to UAV \(j\) is:
[0031]
[0032] The task size that UAV j offloads to other UAVs is:
[0033]
[0034] The task size that UAV j offloads to the base station is:
[0035]
[0036] The task size that other UAVs offload to UAV j is:
[0037]
[0038] Define as the task size executed by user i on UAV j. Then, from the above calculations, we can obtain:
[0039]
[0040] And
[0041]
[0042] The computing delay of UAV j is:
[0043]
[0044] Where is the computing delay of user i on UAV j, as shown below:
[0045]
[0046] The computing energy consumption of UAV j is:
[0047]
[0048] Where is the computing energy consumption of user i on UAV j, as shown below:
[0049]
[0050] Where k is a constant, depending on the chip architecture of the user device, and f ij is the computing power allocated by UAV j to user i.
[0051] Based on the advantage that the BS has powerful computing resources, we can ignore the impact caused by the BS in terms of computing delay and energy consumption. Therefore, in our research, we assume that the computing power of the BS can meet the requirements of the computing tasks, without considering additional delay or energy consumption.
[0052] Assume that the communication link between the user and the UAV is controlled by a line-of-sight (LoS) channel, and the UAV serves the user through frequency division multiple access (FDMA). Therefore, the channel gain between user \(i\) and UAV \(j\) can be expressed as:
[0053]
[0054] where \(\eta_0\) represents the channel gain at a relative distance of 1 meter. Therefore, the transmission rate from user \(i\) to UAV \(j\) is
[0055]
[0056] where \(\sigma\) 2 is the communication channel noise power, \(B\) ij is the communication bandwidth allocated by UAV \(j\) to user \(i\), and \(P\) i is the transmission power of user \(i\). Therefore, the transmission delay from user \(i\) to UAV \(j\) is:
[0057]
[0058] The transmission energy consumption from user \(i\) to UAV \(j\) is:
[0059]
[0060] Assume that the communication link between the user and the base station is controlled by a LoS channel, and the base station serves the user through FDMA. Therefore, the channel gain between user \(i\) and base station \(k\) can be expressed as:
[0061]
[0062] Therefore, the transmission rate from user \(i\) to base station \(k\) is:
[0063]
[0064] where is the communication bandwidth allocated by base station \(k\) to user \(i\). Therefore, the transmission delay from user \(i\) to base station \(k\) is:
[0065]
[0066] The transmission energy consumption from user \(i\) to base station \(k\) is:
[0067]
[0068] Assume that the communication link between UAV \(j\) and UAV \(j'\) is controlled by a LoS channel, and the UAV provides services through FDMA. Therefore, the channel gain between the two UAVs can be expressed as:
[0069]
[0070] where L jj′ is the path loss between drone j and drone j′, as follows:
[0071]
[0072] where is the Euclidean distance between drone j and drone j′, f c is the carrier frequency, c is the speed of light, and η LoS is the additional attenuation factor added in the free space propagation model of the LoS link. Therefore, the transmission rate from drone j to drone j′ is:
[0073]
[0074] where is the communication bandwidth allocated by drone j′ to drone j, is the transmission power of drone j. Therefore, the transmission delay from drone j to drone j′ is:
[0075]
[0076] where S j is the task size unloaded by all users to drone j, and there is:
[0077]
[0078] Then the transmission energy consumption from drone j to drone j′ is:
[0079]
[0080] Assume that the communication link between the drone and the base station is controlled by the LoS channel, and the base station provides services to users through FDMA. Therefore, the channel gain between drone j and base station k can be expressed as:
[0081]
[0082] where is the Euclidean distance between drone j and base station k. Therefore, the transmission rate between drone j and base station k is:
[0083]
[0084] where is the communication bandwidth allocated by base station k to drone j, is the transmission power of drone j. Therefore, the transmission delay from drone j to base station k is:
[0085]
[0086] Then the transmission energy consumption from the UAV j to the base station k is as follows:
[0087]
[0088] Based on the above UAV-assisted three-tier architecture mobile edge computing system, considering the system delay (SystemDelay, SD) and the system energy consumption (System Energy Consumption, SEC) comprehensively, the objective function is set as follows:
[0089] Ψ = λ·SD + (1 - λ)·SEC.
[0090] Where λ is the weight parameter for measuring the ratio of SD and SEC.
[0091] The system delay includes the computing delay and the transmission delay, as follows:
[0092]
[0093] Among them, the computing delay is:
[0094]
[0095] The transmission delay is:
[0096]
[0097] Define the computing delay of user i as:
[0098]
[0099] The system energy consumption includes the computing energy consumption, the transmission energy consumption, and the UAV hovering energy consumption, which is expressed as follows
[0100]
[0101] Among them, the computing energy consumption is:
[0102]
[0103] The transmission energy consumption is:
[0104]
[0105] The UAV hovering energy consumption is:
[0106]
[0107] Among them, the UAV hovering power is:
[0108]
[0109] where ξ is the profile drag coefficient, ρ is the air density (kg / m 3 ), s is the rotor solidity, defined as the ratio of the total blade area to the disk area, S is the rotor disk area (m 2 ), Ω is the blade angular velocity (radians per second), R is the rotor radius (m), k′ is the correction coefficient of the induced power increment, and W is the aircraft weight (Newtons).
[0110] The hovering time of the UAV is the maximum of the following three: the delay t1 when the UAV receiving the user task directly performs local calculation without offloading, the delay t2 when the UAV offloads to another UAV for calculation, and the delay t3 when the UAV offloads to the base station for calculation, as follows:
[0111]
[0112] Therefore, the hovering time of the UAV is:
[0113]
[0114] Furthermore, by optimizing the UAV coordinates, the offloading decisions of the user and the UAV, and the resource allocation, the system cost is minimized and the system performance is maximized. The decision variables to be optimized include the coordinates of all UAVs, the task offloading variables, and the resource allocation variables.
[0115] The abscissa of the UAV is expressed as The ordinate is expressed as The task offloading variables include the user-UAV task offloading variable the user-BS offloading variable the UAV-UAV offloading variable and the UAV-BS offloading variable The above three offloading variables are all binary variables. The resource allocation variables include the computing resources allocated by the UAV to the user Then, the UAV deployment problem in the user-UAV-base station three-layer architecture is formulated as follows:
[0116]
[0117] s.t.
[0118]
[0119] Constraints C1 - C3 restrict that user tasks cannot be split and must be executed locally or can only be offloaded to one base station or one drone. Constraint C4 restricts the variable of drone cooperative offloading, that is, only when the user offloads a task to a drone can this drone offload the received task to another drone. Constraints C5 - C6 restrict that when the user offloads a task to a drone, this drone can only offload the received task to one base station or another drone. Constraint C7 is the drone capacity limit, that is, the limit on the number of users who can establish an offloading relationship with the drone. Constraint C8 is the base station capacity limit, that is, the limit on the number of drones that can establish an offloading relationship with the base station. Constraints C9 - C 10 restricts the range of computing resources. Constraint C 11 -C 14 is the coverage constraint of drones and base stations. A user can only establish an offloading relationship with a drone or a base station whose service range covers this user. Similarly, a drone can only establish an offloading relationship with a base station or other drones whose service range covers this drone. Constraint C 15 is the minimum distance constraint for drones to avoid collisions. Constraint C 16 is the maximum delay limit that user tasks can tolerate.
[0120] Furthermore, the three - layer architecture mobile edge computing system is regarded as a Markov decision process, and the DDPG algorithm is used to automatically learn the optimal drone deployment, task offloading decision, and computing resource allocation. The Markov decision process provides a systematic framework for describing state transitions, reward mechanisms, and control strategies in the decision - making process. Under this framework, the core goal of the problem is to optimize the drone deployment strategy through the interactive learning of agents, so as to reduce the total time delay and total energy consumption of the system.
[0121] Since both the state space and action space involved in this problem are continuous, this makes it impossible to directly apply traditional value iteration methods. Therefore, the DDPG algorithm is used for solution. DDPG is a reinforcement learning algorithm suitable for continuous action spaces. By training a deep neural network to approximate the policy function and value function, the optimal drone deployment strategy can be found. In this algorithm, the agent continuously adjusts its strategy through interaction with the environment to maximize the cumulative reward. The advantage of DDPG is that it can handle high - dimensional, continuous state and action spaces, and is suitable for complex drone deployment problems.
[0122] To accurately represent the state information of the system, the following three parts of information need to be added to the state space. The first part is the values of the decision variables in the original optimization problem P, namely the coordinates of the UAVs, the offloading variables, and the resource allocation variables. The second part is the values affected by the values of the decision variables, such as the channel gain, the distances between pairs of users, UAVs, and base stations, etc. The third part is other environmental information such as the numbers of users, UAVs, and base stations, and the flight altitude of the UAVs. Therefore, the state space can be expressed as
[0123]
[0124] H,D user2uav ,D uav2uav ,D uav2bs ;
[0125] user_count, uav_count, bs_count, uav_height];
[0126] where and are the horizontal and vertical coordinates of the UAV respectively, is the task offloading variable between the user and the UAV, is the offloading variable between UAVs, is the offloading variable between the UAV and the base station, is the offloading variable between the user and the base station, is the computing resource allocated by the UAV to the user, is the channel gain between the user and the UAV, D user2uav ,D uav2uav ,D uav2bs are the Euclidean distances between the user and the UAV, between UAVs, and between the UAV and the base station respectively. user_count, uav_count, bs_count, and uav_height are the numbers of users, UAVs, base stations, and the flight altitude of the UAV respectively. These state variables cover the spatial layout, resource allocation, and network environment characteristics of the system, and can comprehensively reflect the current state of the system and provide information for subsequent decision-making.
[0127] The goal of the problem is to minimize the total time delay and total energy consumption by determining the strategies for UAV deployment, task offloading, and resource allocation. The action space contains all the decision variables of the original optimization problem, namely the coordinates of all UAVs, the task offloading variables, and the computing resource allocation variables. The action space is defined as the set of actions that can be selected under all possible states, as follows:
[0128]
[0129] where and are the horizontal and vertical coordinates of the UAV respectively, is the task offloading variable between the user and the UAV, is the offloading variable between UAVs, is the offloading variable between the UAV and the base station, is the offloading variable between the user and the base station, is the computing resource allocated by the UAV to the user.
[0130] The reward function REWARD(STATE,ACTION) is a function that depends on the state STATE and the action ACTION, and is used to measure the advantages and disadvantages of taking the action ACTION in the state STATE. The designed reward function takes into account factors such as system delay and energy consumption, the satisfaction of system constraints, and the completion of tasks, as follows:
[0131]
[0132] where Q all is the reward for satisfying all constraints, Ψ is the objective function of the original optimization problem, N c is the total number of constraints, penalty i is the penalty value for violating a specific constraint. The reward function will adjust the reward through the optimization objectives of total time delay and total energy consumption. Lower delay and energy consumption will correspond to higher rewards to guide the agent towards minimizing these objectives. The reward function will encourage the agent to satisfy the constraints of the problem during the optimization process, such as resource allocation, physical location limitations of UAVs, constraints on task offloading, etc. Any behavior that violates the constraints will result in a penalty value for the reward, thus guiding the agent to avoid these behaviors that do not conform to the actual constraints. If all constraints are satisfied during the optimization process, the reward will be further increased.
[0133] The Markov decision process is represented by the five-tuple <STATE,ACTION,P,R,γ>, where STATE is the state space, ACTION is the action space, P is the state transition probability, R is the reward function, and γ is the discount factor, indicating the degree of influence of future rewards. The sum of the cumulative reward and the discounted reward can be expressed as:
[0134]
[0135] The state value function V π (a) is defined as follows:
[0136]
[0137] The action value function Q π (s,a) is then defined as follows:
[0138]
[0139] The goal of policy optimization can be transformed into optimizing the state-value function V π (a) or the action-value function Q π (s,a), by continuously adjusting the policy π to maximize the expected cumulative reward, thus achieving the optimal decision-making.
[0140] For a given quadruple <s t ,a t ,r t ,s t+1 >>, DDPG introduces the design of a dual deep neural network. This algorithm uses four deep neural networks, namely the Actor network μ(s,ω), the Critic network Q(s,a,λ), the target Actor network μ'(s,ω'), and the target Critic network Q'(s,a,λ'). The Actor network μ(s,ω) generates a continuous action a based on the current state s, and its goal is to maximize the state-action value function Q(s,a), that is, to select an optimal action to maximize the long-term reward. The Critic network Q(s,a,λ) is used to evaluate the pros and cons of executing action a in state s and calculate the corresponding value. To avoid the model falling into a local optimal solution during training, DDPG introduces exploration noise Therefore, given state s, the current action a can be expressed as:
[0141]
[0142] where represents the Gaussian noise added to the action.
[0143] The target action-value function can be expressed as:
[0144]
[0145] where R i represents the immediate reward, ψ is the discount factor, and Q′(s t+1 ,μ′(s t+1 ,ω′),λ′) is the evaluation of the next state by the target Critic network.
[0146] The loss function of the Critic network can be expressed as:
[0147]
[0148] The loss function of the Actor network can be expressed as:
[0149]
[0150] The DDPG algorithm updates the parameters of the Actor network by minimizing the above loss function. To update the parameters ω of the Actor network, the gradient of the loss function with respect to the parameters ω needs to be calculated as follows:
[0151]
[0152] The policy gradient estimates the expected value through sampling. So we have:
[0153]
[0154] where N is the batch size, is the state-action pair sampled from the experience replay pool. Using the gradient descent algorithm, the update rule for the parameters ω of the Actor network is:
[0155]
[0156] where α ω is the learning rate of the Actor network.
[0157] During the entire training process, the parameters of the Actor and Critic networks are continuously updated, and the parameters of the target network also need to be adjusted synchronously. The target network is updated using a soft update strategy, that is, through a smooth weighted average method, the parameters of the target network gradually approach the parameters of the current network. The update formula is as follows:
[0158] ω′←τω+(1-τ)ω′,λ′←τλ+(1-τ)λ′
[0159] where τ is the soft update coefficient, usually taking a small value to ensure the smoothness of the target network update process.
[0160] In this way, DDPG can effectively handle decision-making problems in continuous action spaces, thereby optimizing the deployment strategy of drones in the mobile edge computing environment. As the training progresses, the DDPG algorithm gradually learns how to select the optimal actions in various different environmental states, ensuring that the drone deployment plan achieves the optimal execution efficiency and robustness. The detailed process is shown in Algorithm 1.
[0161]
[0162] In summary, the present invention proposes a UAV collaborative three - layer MEC deployment algorithm based on DDPG, aiming to solve the problems of UAV position deployment and task assignment optimization in traditional edge computing architectures when facing dynamic environmental changes, uneven distribution of computing resources, and task offloading delays. The method includes the following steps: First, a three - layer architecture MEC system including users, base stations, and UAVs is established. Users generate computing tasks and choose to process them locally, on UAVs, or on base stations. Then, the system is regarded as a Markov decision process (MDP), and the DDPG algorithm is used to automatically learn the optimal UAV deployment, task offloading decisions, and computing resource allocation. The DDPG algorithm combines the three - layer MEC architecture with the UAV deployment problem, realizing a multi - level and multi - party collaborative optimization scheme, which can optimize the system's delay, energy consumption, and computing resource allocation through interaction with the environment. To further improve the system performance, a UAV - to - UAV collaborative offloading mechanism is also introduced. By enabling UAVs to offload tasks from each other, not only can the computing burden be effectively shared, but also when there are communication obstacles between users and base stations, tasks can be offloaded through relays, further optimizing the system's delay and energy efficiency.
[0163] The number of devices and the processing scale described here are used to simplify the description of the present invention. Applications, modifications, and variations of the present invention will be obvious to those skilled in the art.
[0164] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and the illustrated and described examples here.
Claims
1. A drone cooperative three - layer mobile edge computing deployment algorithm based on DDPG, characterized in that, Including: Construct a three-tier architecture mobile edge computing system, where the three-tier architecture mobile edge computing system includes users, base stations, and drones; Regard the three-tier architecture mobile edge computing system as a Markov decision process, and use the DDPG algorithm to automatically learn the optimal drone deployment, task offloading decision, and computing resource allocation; Through the cooperative offloading mechanism between drones, when drones can offload tasks from each other and there are communication obstacles between users and base stations, relay offloading tasks is used to further optimize the latency and energy efficiency of the three-tier architecture mobile edge computing system.
2. The drone cooperative three - layer mobile edge computing deployment algorithm based on DDPG according to claim 1, wherein, The user generates computing tasks and selects to process them locally, on a UAV, or at a base station; Each of the drones is equipped with a computing server; The base station is equipped with a mobile edge computing server to provide computing services for users.
3. The drone cooperative three - layer mobile edge computing deployment algorithm based on DDPG according to claim 1, characterized in that, The user-generated tasks vary in terms of data size, CPU cycles required to process the tasks, and maximum tolerated time delay. The task generated by user i is described as where S i represents the task size, C i represents the CPU cycles required to compute one bit of the task, and T i represents the maximum tolerated time delay for this user's task.
4. The UAV cooperative three-layer mobile edge computing deployment algorithm based on DDPG according to claim 3, characterized in that, Minimize the task processing delay through the binary offloading method, specifically: regard the user's computing task as an indivisible entity, and use the binary variables a ij ∈ {0, 1} and d ik ∈ {0, 1} to represent the offloading processes of the user with the UAV and the user with the base station respectively. A value of 1 indicates that the user task is offloaded, otherwise it is not offloaded.
5. The drone cooperative three - layer mobile edge computing deployment algorithm based on DDPG according to claim 4, characterized in that, Processing the computing tasks at the UAV through the binary offloading method, specifically: defining a binary variable c jk ∈ {0, 1} to represent whether the computing task of UAV j is offloaded to base station k. A value of 1 indicates that the task of the UAV is offloaded, otherwise it is not offloaded; In addition to offloading to the base station, the UAV can offload computing tasks to other UAVs, specifically: through the binary variable b jj′ ∈ {0, 1} is used to indicate whether the computing task of UAV j is offloaded to UAV j'. A value of 1 indicates that the computing task of UAV j is offloaded to UAV j', otherwise it is not offloaded.
6. The drone cooperative three-layer mobile edge computing deployment algorithm based on DDPG according to claim 1, characterized in that, The three-tier architecture mobile edge computing system also includes system latency and system energy consumption, where the system latency includes computing latency and transmission latency; The system energy consumption includes computing energy consumption, transmission energy consumption, and drone hovering energy consumption.
7. The drone collaborative three - layer mobile edge computing deployment algorithm based on DDPG according to claim 1, characterized in that, Further, by optimizing the drone coordinates, the offloading decisions of users and drones, and resource allocation, minimize the system cost and maximize the system performance, where the decision variables to be optimized include the coordinates of all drones, task offloading variables, and resource allocation variables.