Dynamic edge task unloading and unmanned aerial vehicle trajectory optimization adaptive method

Through deep reinforcement learning algorithms and AD3PG algorithms to optimize drone trajectory and task offload, the calculation delay and energy efficiency problems of drone-assisted mobile edge computing system in dynamic environments are solved, and efficient resource management and service coverage are achieved.

CN120343619APending Publication Date: 2025-07-18SHENYANG INSTITUTE OF CHEMICAL TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510282391.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Drone-assisted mobile edge computing systems face the problems of high computing delay and limited time sensitivity in dynamic environments. Traditional convex optimization methods cannot effectively solve the dynamic connection and energy efficiency optimization between the drone and user equipment.

Method used

Deep reinforcement learning algorithm is adopted to build a system model of drone-assisted mobile edge computing, and the optimization problem is transformed into Markov decision-making process. Combined with the AD3PG algorithm, the task offload ratio and drone trajectory are optimized, delayed update parameters, adjustment of neural network structure and learning rate are introduced, state space normalization is implemented, action output is constrained, boundary compliance, energy-aware termination and early completion of the protocol.

Benefits of technology

Adaptive task offloading and trajectory optimization of drones and user equipment in dynamic environments is realized, which reduces the total system delay, improves resource allocation efficiency, expands the coverage of computing resources, and enhances service accessibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343619A_ABST
    Figure CN120343619A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic edge task unloading and unmanned aerial vehicle trajectory optimization adaptive method, and relates to an unmanned aerial vehicle and user adaptive task technical method, and the method comprises the steps: building a system model of unmanned aerial vehicle assisted mobile edge calculation; establishing an optimization problem according to the system model; an optimization problem is converted into a Markov decision process for solving, a deep reinforcement learning algorithm is proposed to optimize a task unloading proportion and unmanned aerial vehicle trajectory optimization related parameters, and an optimal strategy is found. According to the method, the concept of low-altitude economic benefits of the unmanned aerial vehicle is integrated into the normal form of air mobile edge calculation, the ultra-low delay service is oriented, the delay challenge in cloud-based task unloading caused by geographical server user differences is solved, low-delay and faster service unloading is provided through the network edge, and the service unloading efficiency is improved. The time of resource allocation is saved to a great extent, the total time delay of the system is reduced, and the effectiveness of the dynamic method on the delay-sensitive application program in the unmanned aerial vehicle assisted mobile edge computing system is verified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an adaptive method for optimizing the trajectory of an unmanned aerial vehicle, specifically a method for task offloading with dynamic edges and adaptive optimization of the unmanned aerial vehicle trajectory. Background Art

[0002] The rapid development of mobile communication technology and the widespread deployment of Internet-connected devices have made mobile edge computing crucial for managing large amounts of data streams in modern networks. Different from traditional cloud computing with a centralized remote data center, the architecture of mobile edge computing creates a huge geographical gap between computing resources and terminal devices, resulting in low energy efficiency and latency. Mobile edge computing strategically places computing capabilities at the network edge near communication terminals. This architectural innovation improves the quality of service for latency-critical applications through localized data processing, reduces transmission latency by avoiding multi-hop journeys across wireless access networks, backhaul systems, and the Internet, and optimizes energy efficiency by minimizing long-distance data routing to reduce power consumption and extend the battery durability of mobile devices.

[0003] Given the rapid development of mobile edge computing, unmanned aerial vehicles (UAVs) have shown great potential as aerial mobile edge computing platforms, with high operational flexibility, efficient data collection capabilities, and excellent adaptability to harsh environments. These advantages make UAVs a key to 6G networks and an effective solution for handling complex computing tasks. The integration of UAVs and mobile edge computing effectively solves the key limitation of excessive energy consumption caused by concurrent data processing during UAV flight. By offloading computationally intensive tasks to the mobile edge computing system, this synergy significantly reduces the energy burden on the UAV platform while maintaining operational efficiency. In emergency situations, this paradigm shows particular efficacy: when a UAV transmits video of a disaster site to a mobile edge computing server, powerful processors and advanced algorithms can quickly analyze key details such as the location of victims and building damage. This server-based processing is much faster than on-board UAV computing, improving the early warning system through faster situation analysis.

[0004] The drone-assisted mobile edge computing system operates in static or dynamic environments. Although convex optimization is applicable to static scenarios, it encounters difficulties in dealing with dynamic scenarios due to high computational latency and limited time sensitivity. Deep reinforcement learning solves this problem by enabling drones to make adaptive decisions in a changing environment while maintaining real-time responsiveness. In drone data collection and mobile edge computing offloading, dynamic factors such as channel state, task load, and remaining battery of the drone require flexible strategies. Deep reinforcement learning allows drones to continuously adjust offloading strategies by monitoring the real-time resource usage and communication quality of the mobile edge computing server. By optimizing offloading decisions and drone positioning, a near-optimal solution that balances energy and computational efficiency can be achieved. Summary of the Invention

[0005] The object of the present invention is to propose an adaptive method for task offloading and drone trajectory optimization in a dynamic edge, which builds a system model of drone-assisted mobile edge computing; establishes an optimization problem according to the system model; transforms the optimization problem into a Markov decision process for solution, and proposes an algorithm of deep reinforcement learning to optimize the task offloading ratio and related parameters of drone trajectory optimization, and find the best strategy. The present invention saves the time of resource allocation, reduces the total system delay, and verifies the effectiveness of the dynamic method for delay-sensitive applications in the drone-assisted mobile edge computing system.

[0006] The object of the present invention is achieved by the following technical solutions:

[0007] An adaptive method for task offloading and drone trajectory optimization in a dynamic edge, comprising the following steps:

[0008] S1: Build a dynamic environment system model of drone-assisted mobile edge computing, a time delay model of drone-assisted mobile edge computing for task offloading, and an energy consumption model of drone-assisted mobile edge computing for task offloading. Among them, both the drone and the user equipment exhibit real-time mobility, and the dynamic changes of the line-of-sight (LoS) and non-line-of-sight (NLoS) noise power caused by the change of the occlusion condition between the drone and the user equipment are considered. It includes the initialization of the relevant parameters of the drone itself, the number of user equipment, the total tasks inside the system, the task volume randomly generated by the drone, the channel model parameters of the wireless transmission between the user equipment and the drone in the system, and the setting of the environmental site parameters;

[0009] S2: In the system model constructed according to S1, organize the three-dimensional coordinates of the drone in real-time flight, the location of the user device, the distance between the user device and the drone, the wireless transmission rate between the drone and the user device, the local latency of the user device for task offloading, the data transmission latency, the calculation latency of the drone for offloading tasks, the total latency of task offloading in the system, the data transmission energy consumption, the local computing energy consumption, the drone flight energy consumption, the drone working energy consumption, and the total energy consumption of task offloading in the system. Finally, under relevant constraints, establish an optimization problem with the goal of completing all tasks in the system and minimizing the system latency.

[0010] S3: The optimization problem constitutes an essentially non-convex mixed-integer programming task. Coupled with dynamic environmental changes and random data generation for each user device in each time slot, these characteristics cannot be solved by traditional convex optimization methods. Therefore, the formulated optimization problem is transformed into an equivalent Markov decision process, and the state space, action space, and reward function adapted to the system environment are set. Given the complex high-dimensional state and action spaces involved, we propose an Adaptive Delay Deep Deterministic Policy Gradient (AD3PG) algorithm method with four key improvements: introducing a delayed update parameter in the Actor-Critic network architecture; adjusting the neural network structure and learning rate for system optimization; implementing state space normalization to enhance training stability; for the mixed discrete-continuous action space, we use activation functions to constrain the action output and then scale it according to predefined action boundaries. In addition, to adapt to the system environment, this algorithm also follows three protocols: Boundary compliance: When the flight speed and direction selected by the drone exceed the operating boundary, the system will automatically switch the drone to the hover mode; Energy-aware termination: If the remaining battery capacity is insufficient to complete the flight operation and computing offloading in the current time slot, the offloading task will be terminated, and the user device will resume local computing; Early completion protocol: After completing the predefined task requirements before the scheduled time expires, the current cycle may terminate early.

[0011] Furthermore, the system model of drone-assisted mobile edge computing described in S1 includes the following steps:

[0012] S11: Include a drone and multiple user devices. The coverage area is limited to a bounded geographical area with a horizontal size of length (L) and width (W), and the drone maintains a specified flight height (H) within this area. The drone is equipped with a mobile edge computing server that can perform wireless data transmission and distributed edge computing services simultaneously. Within this defined area, N user devices are randomly distributed and exhibit real-time low-speed mobility, establishing a dynamic network topology that requires adaptive resource allocation.

[0013] S12: We assume that the system is in the total communication cycle T slotoperates within, divided into T discrete time slots T slot , in time slot t, the user equipment randomly generates a task of size D n (t), D n (t) ∈ [D n,min , D n,max . UE n can offload part of the task to the UAV through the shared wireless channel. For the UAV, each time slot consists of two phases, namely the flight phase and the communication phase, denoted as t fly and t comm , where, T slot = t fly + t comm . In each time slot, the UAV dynamically adjusts its trajectory according to the real-time position of the user equipment and selects a suitable user equipment for computing offloading. The UAV needs to complete a total data volume of D computing tasks within the entire communication cycle. The user equipment and the time slot set are respectively denoted as

[0014] S13: In time slot t, the position of its user equipment is defined as Q n (t) = [X n (t), Y n (t)] T , while the UAV maintains a fixed height H and a horizontal coordinate Q uav (t) = [X uav (t), Y uav (t)] T , and the UAV performs a hovering operation at this specified waypoint while establishing a communication link with the selected user equipment. The mobility pattern follows discrete-time dynamics, where the horizontal coordinate evolves as:

[0015] X uav (t + 1) = X uav (t) + t fly v(t)cosθ(t);

[0016] Y uav (t + 1) = Y uav (t) + t fly v(t)sinθ(t);

[0017] where, v(t) represents the flight speed and θ represents the rotation angle of the UAV within time slot t;

[0018] In time slot t, the distance between the user equipment and the UAV can be expressed as:

[0019]

[0020] Due to the dynamic position changes between the UAV and the user equipment, the blocking situation between wireless links exhibits time-related characteristics. The blocking state between devices is represented by N n,b (t) ∈ {0, 1}, where the presence of obstacles causes a decrease in signal quality by increasing the noise power, and the noise power is defined as P LOS ; when the user equipment operates in an unobstructed state, it is P NLOS , under the obstructed propagation condition, according to Shannon's theorem, the achievable transmission rate can be expressed as:

[0021]

[0022] where B represents the transmission bandwidth, P tx represents the uplink transmission power, and g n (t) represents the channel gain related to time slots between user equipment, and its magnitude is determined by d n (t);

[0023] S14: The delay model of the UAV-assisted mobile edge computing described by S1 includes the following: When the UAV-assisted mobile edge computing performs offloading task calculations, through the random offloading ratio w n (t) (0 ≤ w n (t) ≤ 1) to the UAV for calculation, C n (t) ∈ {0, 1} indicates whether a wireless connection is established with the UAV, where C n (t) = 1 indicates that the connection has been established, otherwise C n (t) = 0. If the connection is established, the local calculation delay can be expressed as:

[0024]

[0025] where, C ue represents the number of CPU cycles required for each unit of data, and f n represents the local computing power of the user equipment;

[0026] The transmission delay of the task, depending on the task data size and the transmission rate, is expressed as:

[0027]

[0028] When the UAV processes the offloaded task, an edge computing delay will be generated, which can be expressed as:

[0029]

[0030] where, f uav represents the edge computing power of the UAV;

[0031] In summary, the delay of task offloading in time slot t is as follows:

[0032]

[0033] S15: The energy consumption model of the UAV-assisted mobile edge computing described in S1 includes the following: During task offloading, both the UAV and the user equipment will generate energy consumption. For the UAV, the energy consumption mainly comes from flight and edge computing, while for the user equipment, the energy consumption comes from uplink data transmission and local computing. The data transmission energy consumption of the user equipment can be expressed as:

[0034] E n,trans (t) = P tx T n,trans (t);

[0035] The energy consumption generated by the user equipment for local computing can be expressed as:

[0036]

[0037] Among them, the parameter k = 10 -27 represents an architecture-related coefficient that quantifies the impact of chip design on the CPU processing efficiency;

[0038] The UAV flight energy consumption can be expressed as:

[0039]

[0040] where the UAV mass is expressed as M uav ;

[0041] The energy consumption of the UAV for edge computing can be expressed as:

[0042]

[0043] The total energy consumption of the system in time slot t is:

[0044] E(t) = E n,Local (t) + E n,trans (t) + E fly (t) + E n,uav (t).

[0045] Furthermore, the steps to establish an optimization problem according to the system model described in S2 are as follows:

[0046] S21: We define C = [C n (t)] N×T ∈ [0, 1] N×T as the strategy for the UAV to select offloading users, and W = [w n (t)] N×T ∈ [0, 1] N×TLet \(v = [v(t)]\) be the offloading ratio, where \(v\in[0,v_{max}]\). 1×T \(\in[0,v_{max}]\) max \) 1×T Let \(\theta = [\theta(t)]\) be the flight speed of the UAV, where \(\theta\in[0,2\pi]\). 1×T \(\in[0,2\pi]\) 1×T is the flight angle of the UAV. To minimize the total delay of the UAV to complete the task volume, we formulate the established problem as follows:

[0047]

[0048]

[0049] Furthermore, the constraint conditions of the optimization problem in S21 specifically include the following:

[0050] S211: Among them constraint and constraint indicate that within the t time slot, the user equipment is wirelessly connected to the UAV, and one user equipment can only be connected to one UAV;

[0051] S212: The constraint indicates whether there is an occlusion between the user equipment and the UAV in the t time slot, so as to select different noise powers \(P_{n1}\) LOS and \(P_{n2}\) NLOS ;

[0052] S213: The constraint indicates that the flight angle of the UAV is within \([0,2\pi]\);

[0053] S214: The constraint indicates that the flight speed of the UAV does not exceed its maximum speed limit \(v_{max}\) max ;

[0054] S215: Constraint, Constraint, Constraint and Constraint indicate that both the UAV and the user equipment move within a limited area and establish a connection with the user equipment to perform the offloading task;

[0055] S216: The constraint indicates that the task offloading ratio of the user equipment offloaded to the UAV is between \(\{0,1\}\);

[0056] S217: The constraint ensures that the energy consumed by the UAV during the entire communication cycle does not exceed the total battery capacity of the UAV;

[0057] S218: The constraint indicates that the UAV needs to complete a computing task with a data volume of D throughout the communication cycle.

[0058] Furthermore, the transformation of the optimization problem into a Markov decision process for solution described in S3 includes the following steps:

[0059] S31: List the state space in the Markov decision process of the optimization problem;

[0060] S32: List the action space in the Markov decision process of the optimization problem;

[0061] The reward function r describes the immediate reward obtained after executing action a from state s. To be consistent with the optimization goal of minimizing the total task completion delay, we define the reward for each time slot as the negative value of the instantaneous delay. Specifically, the reward for time slot t is expressed as:

[0062]

[0063] S34: So our goal is to maximize the cumulative system reward over the entire range:

[0064]

[0065] Furthermore, listing the state space in the Markov decision process of the optimization problem described in S31 includes the following steps:

[0066] S311: The remaining battery power of the UAV: E re ;

[0067] S312: The current position of the UAV: Q uav ;

[0068] S313: The remaining task volume of the UAV: D re ;

[0069] S314: The current positions of all user devices: Q n ;

[0070] S315: The randomly generated task volume of the user device within the current time slot t: D n ;

[0071] S316: The occlusion situation between the UAV and the user device: N n,b (1 ≤ n ≤ N);

[0072] The dimension of the system state is 4(N + 1), and all state variables are normalized to achieve numerical stability. At time slot t, the state s(t) can be formally expressed as:

[0073] s(t) = {E re(t), Q uav (t), D re (t), Q n (t), D n (t), N n,b (t), 1 ≤ n ≤ N}.

[0074] Furthermore, listing the action space in the Markov decision process of the optimization problem described in S32 includes the following steps:

[0075] S321: The user equipment offloading decision to determine whether to establish a connection with the drone: C n (t);

[0076] S322: The flight angle of the drone at the current time slot t: θ(t);

[0077] S323: The flight speed v(t) of the drone at the current time slot t;

[0078] S324: The offloading ratio w n (t) of the user equipment to offload the computing task to the drone;

[0079] S325: All action variables are normalized to achieve numerical stability. At time slot t, the state a(t) can be formally expressed as:

[0080] a(t) = {C n (t), θ(t), v(t), w n (t), 1 ≤ n ≤ N}.

[0081] Furthermore, a dynamic method for joint task offloading and drone trajectory optimization in an adaptive dynamic edge computing system, the execution process of the AD3PG algorithm described in S35 of S3, is as follows:

[0082] S351: The algorithm starts to execute, inputting and initializing the system environment area parameters; the mass, flight speed, battery power, and computing frequency of the drone; the number of user equipment and computing frequency, etc.;

[0083] S352: Initialize the deep reinforcement learning network parameters, including the discount rate, experience buffer capacity, experience pool size, learning rate, soft update parameter, action exploration noise, delayed update parameter ξ, etc.;

[0084] S353: Reset the environmental variables in the network, observe the current state space according to the occlusion situation, and record the samples into the experience pool;

[0085] S354: Output the exploratory action space through random policy perturbation;

[0086] S355: Execute actions, and update the remaining drone battery level, the total remaining system task volume, and the drone position at the next moment;

[0087] S356: Determine whether the drone position exceeds the boundary. If it exceeds the boundary, set the drone flight speed to zero and return to execute S353; otherwise, execute S358;

[0088] S357: Determine whether the remaining battery power of the drone is zero. If it is zero, set the offloading ratio to zero and return to execute S353; otherwise, execute S358;

[0089] S358: Obtain the position and occlusion situation of the user device at the next moment;

[0090] S359: Record the state space at the next moment and obtain the action space at the next moment from the network;

[0091] S3510: Output the drone flight speed and angle, the offloading ratio, and the offloading decision by adding the update delay parameter ξ;

[0092] S3511: Determine whether the total remaining system task volume is zero. If it is zero, end the algorithm process; otherwise, return to execute S353. Compared with the prior art methods, the present invention has the following advantages and effects:

[0093] 1. The model of the drone-assisted mobile edge computing constructed by the present invention is a pure dynamic environment, in which both the drone and the user device are moving in real time, and the occlusion situation between the drone and the user device is considered. Specifically, it is a method that adapts to the technologies of connection establishment between the drone and the user device, occlusion perception, dual noise power configuration, and adaptive task offloading ratio. Inside the method, the system is optimized by adjusting the neural network structure and learning rate, state space normalization is implemented to enhance training stability. For the mixed discrete-continuous action space, we use activation functions to constrain the action output and then scale it according to the predefined action boundaries. This can achieve joint task offloading and drone trajectory optimization in the edge computing system and find an approximate optimal solution by establishing a Markov decision process.

[0094] 2. In order to adapt to the system environment, the present invention also follows three protocols: Boundary compliance: When the flight speed and direction selected by the drone exceed the operation boundary, the system will automatically switch the drone to the hover mode; Energy-aware termination: If the remaining battery capacity is not sufficient to complete the flight operation and computing offloading in the current time slot, the offloading task will be terminated and the user device will resume local computing; Early completion protocol: After the predefined task requirements are completed before the scheduled time expires, the current cycle may be terminated early. This ensures the effective operation of the algorithm, greatly reduces the total system delay, and improves the efficiency of system resource allocation.

[0095] 3. In the context of the rapidly emerging low-altitude economy of unmanned aerial vehicles (UAVs), UAVs can serve as aerial mobile edge computing nodes, bringing computing resources and services directly to the predetermined locations. By overcoming the limitations of natural disasters, weak physical infrastructure, and extreme weather, UAVs can extend the coverage to areas that may not be reachable by traditional mobile edge computing deployments. Therefore, such flexible deployment ensures effective resource management and enhances service accessibility in areas far from the network or with insufficient services. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] Figure 1 is the flowchart of the implementation process of the method of the present invention;

[0097] Figure 2 is the model diagram of the UAV-assisted mobile edge computing system of the present invention;

[0098] Figure 3 is the algorithm framework diagram of the present invention;

[0099] Figure 4 is the algorithm execution flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0100] In order to enable those skilled in the art of this technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the scope of protection of the present invention.

[0101] The following describes the specific implementation of the present invention in detail with reference to specific embodiments.

[0102] As Figure 1 and Figure 2 shown, a dynamic edge task offloading and UAV trajectory optimization adaptive method according to an embodiment of the present invention includes the following steps:

[0103] S1: Build a dynamic environment system model for UAV-assisted mobile edge computing, a delay model for task offloading in UAV-assisted mobile edge computing, and an energy consumption model for task offloading in UAV-assisted mobile edge computing. Among them, both UAVs and user equipment exhibit real-time mobility, and the dynamic changes in the line-of-sight (LoS) and non-line-of-sight (NLoS) noise power caused by the change of the occlusion condition between UAVs and user equipment are considered. It includes the initialization of UAV-related parameters, the number of user equipment, the total tasks in the system, the task volume randomly generated by UAVs, the channel model parameters for wireless transmission between user equipment and UAVs in the system, and the setting of environmental site parameters;

[0104] Specifically, in the above S1, the UAV-assisted mobile edge computing system model is as follows:

[0105] S11: It includes a UAV and multiple user devices. The coverage area is limited to a bounded geographical area with a horizontal size of length (L) and width (W), and the UAV maintains a specified flight altitude (H) within this area. The UAV is equipped with a mobile edge computing server, capable of simultaneously performing wireless data transmission and distributed edge computing services. Within this defined area, N user devices are randomly distributed and exhibit real-time low-speed mobility, establishing a dynamic network topology that requires adaptive resource allocation;

[0106] S12: We assume that the system operates within a total communication cycle T slot and is divided into T discrete time slots T slot . In time slot t, the user device randomly generates a task of size D n (t), where D n (t) ∈ [D n,min , D n,max . The UE n can offload part of the task to the UAV through the shared wireless channel. For the UAV, each time slot consists of two phases, namely the flight phase and the communication phase, denoted as t fly and t comm , where T slot = t fly + t comm . In each time slot, the UAV dynamically adjusts its trajectory according to the real-time position of the user device and selects a suitable user device for computing offloading. The UAV needs to complete a total data volume of D computing tasks within the entire communication cycle. The user device and the time slot set are respectively denoted as

[0107] S13: Within time slot t, the position of its user device is defined as Q n (t) = [X n (t), Y n (t)] T , while the UAV maintains a fixed altitude H and a horizontal coordinate Q uav (t) = [X uav (t), Y uav (t)] T . The UAV performs a hover operation at this specified waypoint and simultaneously establishes a communication link with the selected user device. The mobility pattern follows discrete-time dynamics, where the horizontal coordinate evolves as:

[0108] X uav (t + 1) = X uav (t) + tfly v(t)cosθ(t);

[0109] Y uav (t + 1)= Y uav (t)+ t fly v(t)sinθ(t);

[0110] where v(t) represents the flight speed and θ represents the rotation angle of the UAV within time slot t;

[0111] In time slot t, the distance between the user equipment and the UAV can be expressed as:

[0112]

[0113] Due to the dynamic position changes between the UAV and the user equipment, the blocking situation between the wireless links exhibits time-related characteristics. The blocking state between the devices is represented by N n,b (t) ∈ {0, 1}, where the presence of obstacles will cause the signal quality to deteriorate by increasing the noise power, and the noise power is defined as P LOS ; it is P NLOS when the user equipment operates in an unobstructed state. Under the blocked propagation condition, according to Shannon's theorem, the achievable transmission rate can be expressed as:

[0114]

[0115] where B represents the transmission bandwidth, P tx represents the uplink transmission power, and g n (t) represents the channel gain related to the time slot between the user equipments, and its magnitude is determined by d n (t);

[0116] S14: The delay model of the UAV-assisted mobile edge computing described in S1 includes the following: When performing offloading task calculations in UAV-assisted mobile edge computing, through the random offloading ratio w n (t) (0 ≤ w n (t) ≤ 1) to the UAV for calculation, C n (t) ∈ {0, 1} represents whether a wireless connection is established with the UAV, where C n (t)= 1 indicates that the connection has been established, otherwise C n (t)= 0. If the connection is established, the local calculation delay can be expressed as:

[0117]

[0118] where C ue represents the number of CPU cycles required for each unit of data, f nRepresents the local computing power of the user equipment;

[0119] The transmission delay of the task, which depends on the task data size and the transmission rate, is expressed as:

[0120]

[0121] When the UAV processes the offloaded task, edge computing delay will be generated, which can be expressed as:

[0122]

[0123] where f uav represents the edge computing power of the UAV;

[0124] In summary, the delay of task offloading in time slot t is:

[0125]

[0126] S15: The energy consumption model of the UAV-assisted mobile edge computing described in S1 includes the following: During task offloading, both the UAV and the user equipment will generate energy consumption. For the UAV, the energy consumption mainly comes from flight and edge computing, while for the user equipment, the energy consumption comes from uplink data transmission and local computing. The data transmission energy consumption of the user equipment can be expressed as:

[0127] E n,trans (t) = P tx T n,trans (t);

[0128] The energy consumption generated by the user equipment for local computing can be expressed as:

[0129]

[0130] where the parameter k = 10 -27 represents an architecture-related coefficient that quantifies the impact of chip design on the CPU processing efficiency;

[0131] The flight energy consumption of the UAV can be expressed as:

[0132]

[0133] where the UAV mass is expressed as M uav ;

[0134] The energy consumption of the UAV for edge computing can be expressed as:

[0135]

[0136] The total energy consumption of the system in time slot t is:

[0137] E(t) = E n,Local (t) + E n,trans (t) + E fly (t) + E n,uav (t).

[0138] S2: In the system model constructed according to S1, organize the obtained three-dimensional coordinates of the UAV during real-time flight, the location of the user equipment, the distance between the user equipment and the UAV, the wireless transmission rate between the UAV and the user equipment, the local delay of the user equipment for task offloading, the data transmission delay, the calculation delay of the UAV for offloading tasks, the total delay of task offloading in the system, the data transmission energy consumption, the local computing energy consumption, the flight energy consumption of the UAV, the working energy consumption of the UAV, and the total energy consumption of task offloading in the system; finally, under relevant constraints, establish an optimization problem with the goal of completing all tasks in the system and minimizing the system delay.

[0139] Specifically, in S2, an optimization problem is established according to the system model, and the specific optimization problem is as follows:

[0140] S21: We define C = [C n (t)] N×T ∈[0, 1] N×T as the strategy for the UAV to select offloading users, W = [w n (t)] N×T ∈[0, 1] N×T as the offloading ratio, v = [v(t)] 1×T ∈[0, v max 1×T as the flight speed of the UAV, θ = [θ(t)] 1×T ∈[0, 2π] 1×T as the flight angle of the UAV. To minimize the total delay of the UAV to complete the task volume, we formulate the established problem as:

[0141]

[0142]

[0143] Specifically, in S21, an optimization problem is established according to the system model, and the specific constraint conditions are as follows:

[0144] S211: Among them The constraint and The constraint indicate that within the t time slot, the user equipment is wirelessly connected to the UAV, and one user equipment can only be connected to one UAV;

[0145] S212: The constraint is expressed as whether there is an occlusion between the user equipment and the UAV in the t time slot, so as to select different noise powers P​LOS and P NLOS ;

[0146] S213: The constraint indicates that the flight angle of the UAV is within [0, 2π].

[0147] S214: The constraint indicates that the flight speed of the UAV does not exceed its maximum speed limit v max ;

[0148] S215: Constraints, Constraints, Constraints and Constraints indicate that both the UAV and the user equipment move within a limited area and establish a connection with the user equipment to perform offloading tasks.

[0149] S216: The constraint indicates that the task offloading ratio of the user equipment unloaded onto the UAV is between {0, 1}.

[0150] S217: The constraint ensures that the energy consumed by the UAV throughout the communication cycle does not exceed the total battery capacity of the UAV.

[0151] S218: The constraint indicates that the UAV has to complete a computing task with a data volume of D throughout the communication cycle.

[0152] S3: The optimization problem constitutes an essentially non-convex mixed-integer programming task. Coupled with dynamic environmental changes and random data generation for each user equipment in each time slot, these characteristics cannot be solved by traditional convex optimization methods. Therefore, the formulated optimization problem is transformed into an equivalent Markov decision process, and the state space, action space, and reward function adapted to the system environment are set. Given the involvement of complex high-dimensional state and action spaces, we propose an Adaptive Delay Deep Deterministic Policy Gradient (AD3PG) algorithm method with four key improvements: introducing a delayed update parameter in the Actor-Critic network architecture; adjusting the neural network structure and learning rate for system optimization; implementing state space normalization to enhance training stability; for the mixed discrete-continuous action space, we use activation functions to constrain action outputs and then scale them according to predefined action boundaries. In addition, to adapt to the system environment, this algorithm also follows three protocols: Boundary compliance: When the flight speed and direction selected by the drone exceed the operation boundaries, the system will automatically switch the drone to the hover mode; Energy-aware termination: If the remaining battery capacity is insufficient to complete the flight operation and computing offloading within the current time slot, the offloading task will be terminated, and the user equipment will resume local computing; Early completion protocol: After completing the predefined task requirements before the expiration of the predetermined time, the current cycle may terminate early.

[0153] Specifically, in the above S3, according to the conversion of the problem into a Markov process, its state space is listed as follows:

[0154] S31: List the state space in the Markov decision process of the optimization problem. The specific state space includes the following:

[0155] S311: Remaining battery power of the drone: E re ;

[0156] S312: Current position of the drone: Q uav ;

[0157] S313: Size of the remaining task volume of the drone: D re ;

[0158] S314: Current positions of all user equipments: Q n ;

[0159] S315: Size of the task volume randomly generated by the user equipment within the current time slot t: D n ;

[0160] S316: Occlusion situation between the drone and the user equipment: N n,b (1 ≤ n ≤ N);

[0161] S317: The dimension of the system state is 4(N + 1), and all state variables are normalized to achieve numerical stability. At time slot t, the state s(t) can be formally expressed as:

[0162] s(t) = {E re (t), Q uav (t), D re (t), Q n (t), D n (t), N n,b (t), 1 ≤ n ≤ N}.

[0163] Specifically, in S3, according to the conversion of the problem into a Markov process, its action space is listed as follows:

[0164] S321: User equipment offloading decision, determining whether to establish a connection with the drone: C n (t);

[0165] S322: The flight angle of the drone at the current time slot t: θ(t);

[0166] S323: The flight speed v(t) of the drone at the current time slot t;

[0167] S324: The offloading ratio w n (t) of the user equipment to offload the computing task to the drone;

[0168] S325: All action variables are normalized to achieve numerical stability. At time slot t, the state a(t) can be formally expressed as:

[0169] a(t) = {C n (t), θ(t), v(t), w n (t), 1 ≤ n ≤ N}.

[0170] Specifically, in S3, according to the conversion of the problem into a Markov process, its reward function is listed as follows:

[0171] S33: The reward function r describes the immediate reward obtained after performing action a from state s. To be consistent with the optimization goal of minimizing the total task completion delay, we define the reward for each time slot as the negative value of the instantaneous delay. Specifically, the reward for time slot t is expressed as:

[0172]

[0173] S34: So our goal is to maximize the cumulative system reward over the entire range:

[0174]

[0175] The above is only the preferred embodiment of the present invention. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several deformations and improvements can be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicability of the patent.

[0176] Specifically, as Figure 3 and Figure 4 shown in the algorithm diagram, for the AD3PG algorithm proposed in S35 of S3, the steps are as follows:

[0177] S351: The algorithm starts to execute, and inputs and initializes the system environment area parameters; the mass of the drone, flight speed, battery power, and calculation frequency; the number of user devices and calculation frequency, etc.;

[0178] S352: Initialize the deep reinforcement learning network parameters, including the discount rate, experience buffer capacity, experience pool size, learning rate, soft update parameter, action exploration noise, delayed update parameter ξ, etc.;

[0179] S353: Reset the environmental variables in the network, and observe the current state space according to the occlusion situation and record the samples into the experience pool;

[0180] S354: Output the exploratory action space through random policy perturbation;

[0181] S355: Execute the action and update the remaining battery of the drone at the next moment, the remaining total task volume of the system, and the position of the drone;

[0182] S356: Determine whether the position of the drone exceeds the boundary. If it exceeds the boundary, set the flight speed of the drone to zero and return to execute S353; otherwise, execute S358;

[0183] S357: Determine whether the remaining battery of the drone is zero. If it is zero, set the offloading ratio to zero and return to execute S353; otherwise, execute S358;

[0184] S358: Obtain the positions of the user devices and the occlusion situation at the next moment;

[0185] S359: Record the state space at the next moment and obtain the action space at the next moment from the network;

[0186] S3510: Output the flight speed and angle of the drone, the offloading ratio, and the offloading decision by adding the update delay parameter ξ;

[0187] S3511: Determine whether the remaining task volume of the system is zero. If it is zero, end the algorithm process; otherwise, return to execute S353.

[0188] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention. Without departing from the concept of the present invention, several deformations and improvements can also be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect and practicality of the present invention.

Claims

1. An adaptive method for task offloading and UAV trajectory optimization of dynamic edges, characterized in that The method includes the following steps: S1: Establish a dynamic environment system model for unmanned aerial vehicle (UAV)-assisted mobile edge computing, a latency model for task offloading in UAV-assisted mobile edge computing, and an energy consumption model for task offloading in UAV-assisted mobile edge computing; where both the UAV and user equipment exhibit real-time mobility, and the dynamic changes in line-of-sight (LoS) and non-line-of-sight (NLoS) noise power caused by the change in the occlusion condition between the UAV and user equipment are considered; including the initialization of UAV-related parameters, the number of user equipment, the total tasks within the system, the task volume randomly generated by the UAV, the channel model parameters for wireless transmission between the user equipment and the UAV in the system, and the setting of environmental site parameters; S2: In the system model constructed according to S1, organize the three-dimensional coordinates of the UAV during real-time flight, the positions of user equipment, the distances between user equipment and the UAV, the wireless transmission rate between the UAV and user equipment, the local latency of task offloading by user equipment, data transmission latency, the computing latency of the UAV for offloading tasks, the total latency of task offloading in the system, data transmission energy consumption, local computing energy consumption, UAV flight energy consumption, UAV working energy consumption, and the total energy consumption of task offloading in the system; finally, under relevant constraints, establish an optimization problem with the goal of completing all tasks within the system and minimizing the system latency; S3: The optimization problem constitutes an essentially non-convex mixed integer programming task. Coupled with the dynamic environmental changes and the random data generation of each user equipment in each time slot, these characteristics cannot be solved by traditional convex optimization methods. Therefore, the formulated optimization problem is transformed into an equivalent Markov decision process, and a state space, an action space, and a reward function adapted to the system environment are set; given the complex high-dimensional state and action spaces involved, an Adaptive Delay Deep Deterministic Policy Gradient (AD3PG) algorithm method is proposed, with four key improvements: introducing a delayed update parameter in the Actor-Critic network architecture; adjusting the neural network structure and learning rate for system optimization; implementing state space normalization to enhance training stability; for the mixed discrete-continuous action space, we use an activation function to constrain the action output and then scale it according to predefined action boundaries; in addition, to adapt to the system environment, this algorithm also follows three protocols: boundary compliance: when the flight speed and direction selected by the UAV exceed the operation boundary, the system will automatically switch the UAV to the hover mode; energy-aware termination: if the remaining battery capacity is insufficient to complete the flight operation and computing offloading within the current time slot, the offloading task will be terminated and the user equipment will resume local computing; early completion protocol: after completing the predefined task requirements before the expiration of the predetermined time, the current cycle may be terminated early.

2. The adaptive method for task offloading and UAV trajectory optimization with dynamic edges according to claim 1, wherein, The system model of UAV-assisted mobile edge computing described in S1 includes the following steps: S11: It includes a drone and multiple user devices; the coverage area is limited to a bounded geographical area with a horizontal size of length (L) and width (W), and the drone maintains a specified flight altitude (H) within this area; the drone is equipped with a mobile edge computing server capable of simultaneous wireless data transmission and distributed edge computing services; within this defined area, N user devices are randomly distributed and exhibit real-time low-speed mobility, establishing a dynamic network topology that requires adaptive resource allocation. S12: Assume that the system runs within the total communication cycle T slot and is divided into T discrete time slots T slot . In time slot t, the user equipment randomly generates a task of size D n (t), where D n (t) ∈ [D n,min , D n,max ; the user equipment UE n can offload part of the task to the UAV through the shared wireless channel; for the UAV, each time slot consists of two phases, namely the flight phase and the communication phase, denoted as t fly and t comm , where T slot = t fly + t comm ; in each time slot, the UAV dynamically adjusts its trajectory according to the real-time position of the user equipment and selects a suitable user equipment for computing offloading; the UAV needs to complete a total data volume of D computing tasks within the entire communication cycle; the user equipment and the time slot set are respectively denoted as 3. The adaptive method for task offloading and UAV trajectory optimization with dynamic edges according to claim 1, characterized in that, The system model related parameters of the drone-assisted mobile edge computing described in S1 include the following steps: S13: During time slot t, the location of its user equipment is defined as Q n (T) = [X n (T), Y n (t)] T , while the drone maintains a fixed altitude H and horizontal coordinate Q uav (T) = [X uav (T), Y uav (T)] T , the drone hovers at this specified waypoint while establishing a communication link with the selected user equipment; the mobility pattern follows discrete-time dynamics, where the horizontal coordinate evolves as: X uav (t + 1)=X uav (t)+t fly v(t)cosθ(t); Y uav (t + 1)=Y uav (t)+t fly v(t)sinθ(t); Among them, v(t) represents the flight speed, and θ represents the rotation angle of the drone within time slot t. In time slot t, the distance between the user device and the drone can be expressed as: Due to the dynamic position changes between the UAV and the user equipment, the blocking situation between the wireless links exhibits time-related characteristics; the blocking state between devices is represented by N n,b (t) ∈ {0, 1}, where the presence of obstacles will cause the signal quality to degrade by increasing the noise power, and the noise power is defined as P LOS ; it is P NLOS when the user equipment operates in an unobstructed state. Under the blocked propagation condition, according to Shannon's theorem, the achievable transmission rate can be expressed as: where B represents the transmission bandwidth, and P tx represents the uplink transmission power, while g n (t) represents the channel gain related to the time slot between user equipments, and its magnitude is determined by d n (t); S14: The latency model of UAV-assisted mobile edge computing described in S1 is as follows: When performing offloading task calculation in UAV-assisted mobile edge computing, through the random offloading ratio w n (t) (0 ≤ w n (t) ≤ 1), the calculation is performed on the UAV side, and C n (t) ∈ {0, 1} indicates whether a wireless connection is established with the UAV, where C n (t) = 1 indicates that the connection has been established, otherwise C n (t) = 0. If the connection is established, the local calculation delay can be expressed as: Among them, C ue represents the number of CPU cycles required for each unit of data, and f n represents the local computing power of the user equipment; The transmission delay of the task, which depends on the task data size and transmission rate, is expressed as: When the drone processes the offloaded task, an edge computing delay will be generated, which can be expressed as: Among them, f uav represents the edge computing ability of the drone; To sum up, the delay of task offloading in time slot t is: S15: The energy consumption model of the drone-assisted mobile edge computing described in S1 includes the following: During task offloading, both the drone and the user device will generate energy consumption; for the drone, the energy consumption mainly comes from flight and edge computing, while for the user device, the energy consumption comes from uplink data transmission and local computing; the data transmission energy consumption of the user device can be expressed as: E n,trans P(t) = tx T n,trans (t); The energy consumption generated by the user device for local computing can be expressed as: Among them, the parameter k = 10 -27 represents an architecture-related coefficient that quantifies the impact of chip design on CPU processing efficiency; The flight energy consumption of the drone can be expressed as: where the mass of the drone is denoted as M uav ; The energy consumption of the drone for edge computing can be expressed as: The total energy consumption of the system within time slot t is: E(t) = E n,Local (t) + E n,trans (t) + E fly (t) + E n,uav (t).

4. An adaptive method for task offloading and UAV trajectory optimization with dynamic edges according to claim 1, characterized in that, The final optimization problem of the drone-assisted mobile edge computing described in S2 includes the following steps: S21: Define \(C = [C n (t)] N×T \in[0,1] N×T as the strategy for the UAV to select offloading users, \(W = [w n (t)] N×T \in[0,1] N×T as the offloading ratio, \(v = [v(t)] 1×T \in[0,v max 1×T as the flight speed of the UAV, \(\theta = [\theta(t)] 1×T \in[0,2\pi] 1×T as the flight angle of the UAV; To minimize the total delay of the UAV to complete the task volume, we formulate the established problem as:​ S211: Among them, C n (t) ∈ {0, 1}, Constraint and Constraint means that within the t time slot, the user equipment conducts a wireless communication connection with the drone, and one user equipment can only be connected to one drone; S212: N n,b (t) ∈ {0, 1}, The constraint is expressed as whether there is an occlusion between the user equipment and the UAV in the t time slot, so as to select different noise powers P LOS and P NLOS ; S213: 0 ≤ θ(t) ≤ 2π, The constraint indicates that the flight angle of the drone is within [0, 2π]; S214: 0 ≤ v(t) ≤ v max , The constraint indicates that the flight speed of the drone does not exceed its maximum speed limit v max ; S215: 0 ≤ X uav (t) + t fly v(t)cosθ(t) ≤ L, Constraint, 0 ≤ Y uav (t) + t fly v(t)sinθ(t) ≤ W, Constraint, 0 ≤ X n (t) ≤ L, Constraint and 0 ≤ Y n (t) ≤ L, The constraints indicate that both the drone and the user equipment move within a limited area and establish a connection with the user equipment to perform the offloading task; S216: 0 ≤ w n (t) ≤ 1, The constraint indicates that the task offloading ratio of the user equipment offloaded to the drone is between {0, 1}; S217: The constraint ensures that the energy consumed by the UAV during the entire communication cycle does not exceed the total battery capacity of the UAV; S218: The constraint indicates that the UAV needs to complete a computing task with a data volume of D within the entire communication cycle.

5. An adaptive method for task offloading and UAV trajectory optimization with dynamic edges according to claim 1, characterized in that, S3 transforms the optimization problem into a Markov decision process for solution, including the following steps: S31: List the state space in the Markov decision process of the optimization problem, including the following components: S311: Remaining battery power of the drone: E re ; S312: Current UAV position: Q uav ; S313: Size of the remaining tasks of the drone: D re ; S314: Current locations of all user devices: Q n ; S315: The amount of tasks randomly generated by the user equipment within the current time slot t: D n ; S316: Occlusion situation between the drone and the user equipment: N n,b (1 ≤ n ≤ N); S317: The dimension of the system state is 4(N + 1), and all state variables are normalized to achieve numerical stability. At time slot t, the state s(t) can be formally expressed as: s(t) = {E re (t), Q uav (t), D re (t), Q n (t), D n (t), N n,b (t), 1 ≤ n ≤ N}; S32: List the action space in the Markov decision process of the optimization problem, including the following components: S321: User equipment offloading decision to determine whether to establish a connection with the drone: C n (t); S322: The flight angle of the drone at the current time slot t: θ(t); S323: The flight speed v(t) of the drone at the current time slot t; S324: Offloading ratio of the user equipment to offload computing tasks to the drone: w n (t); S325: All action variables are normalized to achieve numerical stability. At time slot t, the state a(t) can be formally expressed as: a(t) = {C n (t), θ(t), v(t), w n (t), 1 ≤ n ≤ N}; S33: The reward function r describes the immediate reward obtained after performing action a from state s; to be consistent with the optimization goal of minimizing the total task completion delay, the reward for each time slot is defined as the negative value of the instantaneous delay; specifically, the reward for time slot t is expressed as: S34: So our goal is to maximize the cumulative system reward over the entire range:

6. The adaptive method for task offloading and UAV trajectory optimization with dynamic edges according to claim 1, wherein The execution process of the algorithm named AD3PG algorithm includes the following steps: S35: The execution flow of the AD3PG algorithm, characterized by including the following steps: S351: The algorithm starts to execute, inputting and initializing the parameters of the system environment area; the mass of the UAV, flight speed, battery power, and computing frequency; the number of user devices and computing frequency. S352: Initialize the parameters of the deep reinforcement learning network, including the discount rate, experience buffer capacity, experience pool size, learning rate, soft update parameter, action exploration noise, and delayed update parameter ξ. S353: Reset the environmental variables in the network, observe the current state space according to the occlusion situation, and record the samples into the experience pool. S354: Output the exploratory action space through random policy perturbation. S355: Execute the action and update the remaining battery of the UAV at the next moment, the remaining total task volume of the system, and the position of the UAV. S356: Determine whether the position of the UAV exceeds the boundary. If it exceeds the boundary, set the flight speed of the UAV to zero and return to execute S353; otherwise, execute S358. S357: Determine whether the remaining battery of the UAV is zero. If it is zero, set the offloading ratio to zero and return to execute S353; otherwise, execute S358. S358: Obtain the positions of the user devices and the occlusion situation at the next moment. S359: Record the state space at the next moment and obtain the action space at the next moment from the network. S3510: Output the flight speed and angle of the UAV, the offloading ratio, and the offloading decision by adding the update delay parameter ξ. S3511: Determine whether the remaining task volume of the system is zero. If it is zero, end the algorithm process; otherwise, return to execute S353.