Auxiliary edge computing task unloading method based on unmanned aerial vehicle

By constructing a drone-assisted edge computing system model and the CPER-MATD3 algorithm to optimize drone trajectories and task offloading strategies, the resource limitation and dynamic offloading problems in multi-drone collaborative assisted edge computing are solved, computing load balancing, communication delay minimization and energy consumption control are achieved, and task offloading efficiency is improved.

CN120704760APending Publication Date: 2025-09-26CHANGCHUN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510820443.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Edge computing assisted by multi-UAV collaboration faces problems such as limited resources, dynamic and changeable tasks, and complex offloading decisions, making it difficult to achieve computing load balancing, minimize communication delays, and control energy consumption.

Method used

A drone-assisted edge computing system model is constructed, and the CPER-MATD3 algorithm is used to optimize the drone trajectory and task offloading strategy. The weighted sum of latency and energy consumption is optimized through the discrete time model and Markov decision process combined with the user energy harvesting model.

Benefits of technology

It significantly reduces the total system cost, optimizes task offloading and improves efficiency, and effectively solves the resource conflicts and dynamic offloading challenges in multi-UAV collaborative assisted edge computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704760A_ABST
    Figure CN120704760A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of mobile communication, and particularly relates to a task unloading method based on unmanned aerial vehicle auxiliary edge computing, which comprises the following steps of: 1, constructing an unmanned aerial vehicle auxiliary edge computing system model; step 2, mathematical modeling is carried out on the current system model diagram, specifically, a trajectory model and an anti-collision model of the unmanned aerial vehicle, a communication model in the task unloading process and a calculation model in the task unloading process are included; 3, constructing a user energy collection model; 4, constructing an optimization problem by taking the weighted sum of the time delay and the energy consumption in the task processing process as a target; and 5, establishing an optimization problem as a Markov decision process, and solving an unmanned task unloading strategy in the unmanned aerial vehicle auxiliary edge computing system by adopting a CPER-MATD3 algorithm. According to the method, the overall effectiveness of the system can be remarkably improved, and optimization and efficiency improvement of task unloading are effectively realized in a multi-unmanned-aerial-vehicle-assisted mobile edge computing environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of mobile communication technology, and in particular to a method for offloading edge computing tasks assisted by drones. Background Art

[0002] With the rapid development of technologies such as 5G, the Internet of Things, and artificial intelligence (AI), the number of intelligent terminal devices is growing exponentially, giving rise to a large number of application scenarios that require extremely high computing resources and real-time responsiveness, such as high-definition video analysis, augmented reality, autonomous driving, and smart city management. These applications are characterized by being computationally intensive, data-intensive, and extremely sensitive to latency. While traditional cloud computing models offer powerful computing capabilities, their physical distance from data sources leads to high communication latency and network congestion during task transmission from the terminal to the cloud, making it difficult to meet the demand for low-latency, high-reliability services. To address this, mobile edge computing (MEC) has been proposed, deploying computing, storage, and network resources to edge nodes close to end users. This effectively reduces data transmission latency, improves task processing efficiency, and improves system responsiveness.

[0003] In this context, multi-UAV systems, as a new type of aerial assisted computing platform, show great potential in edge computing due to their high mobility, flexible deployment, and wide-area coverage. Multiple UAVs can not only serve as in-flight mobile edge servers, providing real-time computing services to ground users or devices, but can also complement the shortcomings of fixed infrastructure in special environments such as disaster relief, field patrols, and emergency communications, expanding the scope of MEC applications.

[0004] However, multi-UAV collaborative edge computing also faces numerous challenges. Firstly, UAVs are resource-constrained, such as limited computing power, battery capacity, and communication bandwidth, which can easily lead to resource conflicts and rapid energy depletion. Secondly, ground users' tasks are dynamic and ever-changing, and offloading demands are highly time-varying and spatially uneven, making the decision-making process for task offloading extremely complex. Therefore, designing efficient, intelligent, and dynamic offloading strategies that enable multiple UAVs to complete their tasks while balancing computational load, minimizing communication latency, and controlling energy consumption has become a key research issue. Summary of the Invention

[0005] (1) Technical problems solved

[0006] In response to the shortcomings of the existing technology, the present invention provides a method for offloading tasks based on drone-assisted edge computing, which solves the problems raised in the above background technology.

[0007] (2) Technical solution

[0008] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions:

[0009] A method for offloading tasks based on drone-assisted edge computing, comprising the following steps:

[0010] Step 1: Build a drone-assisted edge computing system model, which consists of M user UEs and N UAVs equipped with MEC servers. When user computing resources are insufficient, tasks can be offloaded to the UAVs equipped with MEC servers for computing.

[0011] Step 2: Mathematically model the current system model, including the trajectory model and collision avoidance model of the UAV, the communication model during task offloading, and the computational model during task offloading. The computational model during task offloading mainly includes two methods: local computation of the task by the user or offloading the task to a UAV equipped with an MEC server for computation.

[0012] Step 3: Build a user energy harvesting model;

[0013] Step 4: Construct an optimization problem with the goal of minimizing the weighted sum of latency and energy consumption during task processing.

[0014] Step 5: The optimization problem is formulated as a Markov decision process, and the CPER-MATD3 algorithm is used to solve the unmanned task offloading strategy in the UAV-assisted edge computing system.

[0015] Furthermore, the specific content of step 1 is as follows: the UAV-assisted edge computing system model adopts a discrete time model, which divides the communication time of the entire system into T time slots equally, and each time slot lasts for τ, including the service time τ com and flight time τ fly In this communication mode, only one user establishes contact with the drone in each time slot; in this system model, M drones fly on a plane with a fixed height of H. Each drone is equipped with a high-performance edge computing server. When the user's computing resources are insufficient, the task can be offloaded to the drone equipped with the MEC server for calculation.

[0016] Furthermore, the specific content of step 2 is: for the trajectory model and anti-collision model of the UAV, in the tth time slot, when the three-dimensional Cartesian coordinates of the user UEm are represented by p m (t) = [x m (t),y m (t),0], the three-dimensional Cartesian coordinates of the location of the UAVn are expressed as q n (t) = [x n (t),y n (t), H]; therefore, the Euclidean distance between user UEm and drone UAVn can be expressed as:

[0017]

[0018] At the tth time slot, the three-dimensional Cartesian coordinates of the UAVn’s location are expressed as q n (t) = [x n (t),y n (t),H], in a given area along the horizontal angle θ m (t)∈[0,2π) direction and velocity v n (t)∈[0,v max ] flight, the horizontal coordinate of the UAVn at the t+1th time slot can be expressed as:

[0019]

[0020] Because the drone needs to serve users within the user service area, its horizontal coordinates must meet the following requirements:

[0021]

[0022] where X max and Y max Indicates the boundaries of the service area;

[0023] Because there are multiple drones in the system, in order to avoid collisions between UAVs, they should maintain a minimum distance between them:q m (t)-q n (t)≥S min ; Among them, S min To avoid collisions, keep a safe distance;

[0024] For the communication model in the process of drone-assisted edge computing task offloading, the line-of-sight environment is an important factor that can affect the transmission speed between the drone and the terminal. Due to the obstruction of high-rise buildings and trees, the line-of-sight environment cannot always be guaranteed. Therefore, the line-of-sight environment between the drone and the user is random. The probability of the line-of-sight environment between the drone and the user can be expressed by the following function expression:

[0025]

[0026] Among them, C1 and C2 are constants determined by the environment. It represents the elevation angle between the user and the drone at the tth time slot, which can be expressed as follows:

[0027]

[0028] The probability of the drone and the user being in a non-line-of-sight environment can be expressed as:

[0029]

[0030] Depending on the LoS or NLoS connection between the drone and the user, the path loss of the received signal power of each user can be expressed as:

[0031]

[0032] Where f is the carrier frequency of the system, η Los and η NLos are the system constants in line-of-sight and non-line-of-sight environments, respectively, and c is the speed of light;

[0033] In order to simplify the path loss model of the UAV in the system and fully consider the path loss of the UAV in both line-of-sight and non-line-of-sight environments, the path probability is used to perform a weighted average on the above path loss formula to obtain the average path loss of the UAV, which can be expressed as:

[0034]

[0035] According to the average path loss between the UAVn and the user UEm, the channel gain between the UAVn and the user UEm can be obtained:

[0036]

[0037] Where α0 represents the power gain at a reference distance of 1 m;

[0038] According to the channel gain between the UAVn and the user UEm, the transmission rate between the UAVn and the user UEm can be:

[0039] Where B is the bandwidth of wireless communication, Uplink transmission power between UAVn and UEm, σ 2 is the noise power spectral density;

[0040] The computational model in the task offloading process mainly includes two methods. Specifically, in the tth time slot, part of the computational task of user UEm is offloaded to the UAVn. The proportion of this part of the task to the total task of the user is Indicates that the remaining task data is processed locally. The proportion of this part of the task volume to the total tasks of the user is expressed as express;

[0041] User UE k and UAV m The uninstall matching relationship is recorded as l m,n (t)={0,1}, when l m,n When (t) = 1, it indicates that the user UEk and UAV m Establish a connection, otherwise, l m,n (t) = 0; Assuming that there is only one drone providing services to the user in each time slot, it should satisfy:

[0042]

[0043] (1) Local computing model

[0044] In the tth time slot, the delay caused by local calculation can be expressed as:

[0045] Among them, D m (t) is the amount of tasks to be processed by user UEm in the tth time slot, s is the number of CPU cycles required to process each unit bit of data, and f m (t) represents the computing capability of user UEm;

[0046] In the tth time slot, the energy consumption generated by local computing can be expressed as:

[0047] in, It is the power factor determined by the CPU architecture;

[0048] (2) Model that offloads tasks to a drone equipped with a MEC server

[0049] The delay caused by offloading the task to the UAV equipped with the MEC server mainly consists of two parts: the transmission delay of user UEm offloading the task to UAVn and the computation delay of user UEm offloading the task to UAVn. Therefore, in the tth time slot, the transmission delay of user UEm offloading the task to UAVn can be expressed as:

[0050]

[0051] In the tth time slot, user UEm offloads the task to UAVn to calculate the delay, which can be expressed as:

[0052]

[0053] Among them, f n (t) represents the computing resources allocated to user UEm by UAVn in the tth time slot;

[0054] The energy consumption generated by offloading the task to the UAV equipped with the MEC server mainly includes three parts: the flight energy consumption of the UAV, the transmission energy consumption of the user UEm when offloading the task to the UAVn, and the computing energy consumption of the user UEm when offloading the task to the UAVn. Therefore, in the tth time slot, the transmission energy consumption of the user UEm when offloading the task to the UAVn can be expressed as:

[0055]

[0056] At the tth time slot, the energy consumption generated by the UAVn flight can be expressed as:

[0057] Among them, M иаv The quality of UAVn;

[0058] In the tth time slot, user UEm offloads the task to the UAVn to calculate the energy consumption, which can be expressed as:

[0059]

[0060] Among them, κ is the impact factor of hardware structure on the CPU processing of the UAV;

[0061] Because local computation and computation offloaded to the UAV are performed simultaneously, the total delay of the system at time slot t can be expressed as:

[0062] The total energy consumption of the system includes the energy consumption generated by local computing, the flight energy consumption of the UAV, the transmission energy consumption of the user UEm offloading the task to the UAVn, and the computing energy consumption of the user UEm offloading the task to the UAVn, which can be expressed as:

[0063] The cost of task processing includes the weighted sum of the total latency and total energy consumption in the multi-UAV assisted edge computing system at the tth time slot, which can be expressed as follows:

[0064] L(t)=λ1T(t)+λ2E(t).

[0065] Furthermore, the specific content of step 3 is to equip each user with an energy collection module, which can charge the battery through wireless signals. Assuming that energy collection starts at each time interval, in the initial state, assuming that the battery capacity of each user is full, the maximum battery capacity is Therefore, the energy of each user in the next time slot depends on the energy consumption and harvest of this time slot, which can be expressed by the following formula:

[0066]

[0067] Among them, e mn Indicates the energy collected by the user at the current moment.

[0068] Furthermore, the specific content of step 4 is: in the construction of a multi-UAV assisted mobile edge computing task offloading system, the weighted sum of the delay and energy consumption during task processing is minimized by jointly optimizing the position scheduling and task offloading strategy of the UAVs. Therefore, the optimization problem is constructed as follows:

[0069]

[0070] C2:θ n (t)∈[0,2π)

[0071] C3:v n (t)∈[0,v max ]

[0072] C4:0≤x n (t)≤X max

[0073] C5:0≤y n (t)≤Y max

[0074] C6:l m,n (t)={0,1}

[0075]

[0076] C9:q n (t)-q k (t)≥S min

[0077] C10:l m,n (t)R m,n (t)≤R max

[0078]

[0079] C13:T(t)≤T max (t)

[0080] Among them, constraint C1 represents the value range of the task offloading ratio; constraint C2 represents the constraint on the flight speed of UAVn; constraint C3 represents the constraint on the flight speed of UAVn; constraints C4 and C5 represent that UAVn moves within the area serving the user; constraint C6 represents whether user UEm establishes a connection with UAVn; constraint C7 indicates that in the tth time slot, UAVn can only provide service to one user; constraint C8 constrains the weight factors of the total system delay and total energy consumption; constraint C9 represents the constraint on the safe distance between UAVs; constraint C10 indicates that when user UEm decides to offload the task to UAVn, user UEm must be within the coverage range of UAVn; constraint C11 indicates that the battery power of user UEm does not exceed the minimum power threshold; constraint C12 indicates that the flight energy consumption and computing energy consumption of the UAV cannot exceed the maximum energy of its battery during the entire service cycle; constraint C13 indicates that the total delay of the system must be less than the maximum tolerable delay.

[0081] Furthermore, the specific content of step 5 is: establishing the optimization problem as a Markov decision process, and using the CPER-MATD3 algorithm to solve the unmanned task offloading policy in the UAV-assisted edge computing system. The Markov decision process consists of state, action and reward function;

[0082] State space S(t): includes the position information q of the UAVn n (t), user UEm location information p m (t), the size of the task amount generated by user UEM D n (t), the remaining power of the UAVn User UEm's battery status b m (t), can be defined as:

[0083]

[0084] Action space A(t): The action space is where the agent explores the state space and takes relevant actions, including task offloading strategies The unloading matching relationship between user UEm and drone UAVn is l m,n (t), the UAV flight deflection angle θ n (t), UAV flight speed ν n (t), can be expressed as:

[0085]

[0086] in, Indicates the task offloading strategy, l m,n(t) represents the unloading matching relationship between user UEm and UAV UAVn, θ n (t) represents the flight deflection angle of the UAVn, ν n (t) represents the flight speed of UAVn;

[0087] Reward function R(t): Since the goal of the optimization problem is to minimize the total system cost, the negative of the total system cost is used as the reward function, which can be defined as: R(t) = -[L(t) + Φ];

[0088] Where Φ is a penalty function, including penalties when the distance between drones is less than the safe distance, when the drone exceeds the minimum power threshold, and when the task completion time exceeds the maximum tolerable delay;

[0089] After modeling the Markov decision process of the optimization problem, the CPER-MATD3 algorithm is proposed to solve the unmanned task offloading strategy in the UAV-assisted edge computing system;

[0090] The MATD3 algorithm is an extension of the TD3 algorithm in the field of multi-agents. Its structure still adopts the form of centralized training and independent execution, that is, the input space of the value function of each agent includes not only its own observations and actions, but also the observations and actions of all other agents. In the MATD3 algorithm, each agent uses two Critic networks to estimate the action value function separately, which is similar to the idea of ​​Double DON. During the training process, the outputs of the two Critic networks are not directly used. Instead, the minimum value of the two is taken as the actual estimated value, which effectively reduces the volatility of the estimate and makes the estimate closer to the true value, thereby solving the overestimation problem to a certain extent. In addition, the MATD3 algorithm also adopts a policy delayed update method and a soft update strategy. Policy delayed update refers to the inconsistent update frequency of the network parameters of the Actor and Critic. Soft update means that the new parameters are equal to the weighted average of the old parameters and the new target parameters. This can reduce the fluctuation during parameter update, make the network training more stable, and further reduce the overestimation problem.

[0091] In the traditional MATD3 algorithm, the experience replay mechanism selects training samples through random sampling. This method fails to distinguish the importance of different experience samples, resulting in the failure to effectively discover and utilize high-value experience samples, thereby limiting the efficiency of network training. Although some studies have proposed using a prioritized experience replay mechanism to improve sampling efficiency, this method requires calculating and sorting the TD-error of all samples. This method ignores experiences with high immediate return values, which are also highly important. Therefore, this paper proposes the CPER-MATD3 algorithm to solve the problem of edge computing task offloading under the assistance of multiple drones. This algorithm not only considers experiences with high TD-error values, but also incorporates experiences with high immediate return values ​​as evaluation indicators. Through a composite priority sampling mechanism, it achieves efficient sample selection and significantly improves the utilization rate of experience samples.

[0092] The composite priority sampling mechanism mainly includes five steps:

[0093] The first step is to determine the TD-error value and immediate reward value of the agent in the current state;

[0094] The second step is to calculate the priority of the experience samples in the immediate reward standard and the TD-error standard respectively;

[0095] Y i =r t +ε

[0096] Y j =|δ t |+ε

[0097] The third step is to sort the priorities of the experience samples in the immediate return standard and the TD-error standard in ascending order to obtain rank(i) and rank(j), and then use the composite average sorting:

[0098]

[0099] The fourth step is to calculate the composite priority:

[0100]

[0101] Among them, α represents the relative weight of the evaluation priority;

[0102] The fifth step is to define the sampling probability:

[0103]

[0104] The steps of the CPER-MATD3 algorithm are as follows:

[0105] Step 1: Initialize the Actor network, Critic network, and target network, initialize the experience replay pool, and initialize the simulation parameters of the drone MEC system;

[0106] Step 2: The drone decides its actions based on the current strategy;

[0107] Step 3: After executing the action, the drone observes the next state and immediate reward, combines the state, action, reward, and next state into a four-tuple, and stores it in the experience replay pool;

[0108] Step 4: Use a composite priority sampling mechanism to extract small batches of samples from the experience replay pool to update the Actor and Critic networks;

[0109] Step 5: The parameters of the Critic network are updated by minimizing the loss function. The update rule of the Critic network is:

[0110]

[0111] Step 6: The Actor network is updated by sampling policy gradients. The update rule is:

[0112]

[0113] Step 7: Update the target network through soft update. The update rules are as follows:

[0114]

[0115] (3) Beneficial effects

[0116] Compared with the existing technology, the present invention provides a method for offloading tasks based on drone-assisted edge computing, which has the following beneficial effects:

[0117] The present invention constructs a UAV-assisted edge computing system model, a UAV trajectory model and an anti-collision model, a communication model between users and UAVs during task offloading, a computing model, and a user energy collection model. The weighted sum of time delay and energy consumption in the system is defined as a joint optimization problem, and the optimization problem is established as a Markov decision process. The CPER-MATD3 algorithm is used to solve the unmanned task offloading strategy in the UAV-assisted edge computing system, which can significantly reduce the total system cost and effectively achieve the optimization of task offloading and efficiency improvement. BRIEF DESCRIPTION OF THE DRAWINGS

[0118] Figure 1 is a flow chart of the method steps of the present invention;

[0119] Figure 2 This is a schematic diagram of the structure of the drone-assisted edge computing system model of the present invention;

[0120] Figure 3 This is a flowchart of the present invention using the CPER-MATD3 algorithm to optimize the total system cost. DETAILED DESCRIPTION

[0121] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0122] Example

[0123] like Figure 1-3 As shown, an embodiment of the present invention proposes a method for offloading tasks based on drone-assisted edge computing, including the following steps:

[0124] Step 1: Construct a drone-assisted edge computing system model, which consists of M users (UEs) and N drones (UAVs) equipped with MEC servers. When the user computing resources are insufficient, the task can be offloaded to the drone equipped with MEC servers for computing.

[0125] Step 2: Mathematically model the current system model diagram, including the trajectory model and anti-collision model of the drone, the communication model during the task offloading process, and the calculation model during the task offloading process. The calculation model during the task offloading process mainly includes two methods: the user calculates the task locally or the user offloads the task to the drone equipped with the MEC server for calculation.

[0126] Step 3: Build a user energy harvesting model;

[0127] Step 4: Construct an optimization problem with the goal of minimizing the weighted sum of latency and energy consumption during task processing.

[0128] Step 5: The optimization problem is formulated as a Markov decision process, and the CPER-MATD3 algorithm is used to solve the unmanned task offloading strategy in the UAV-assisted edge computing system.

[0129] Step 1: Construct a UAV-assisted edge computing system model. The system adopts a discrete time model and divides the communication time of the entire system into T time slots. Each time slot lasts for τ, including the service time τ. com and flight time τ flyIn this communication mode, only one user establishes contact with the drone in each time slot. In this system, M drones fly on a plane with a fixed height of H. Each drone is equipped with a high-performance edge computing server. When the user's computing resources are insufficient, the task can be offloaded to the drone equipped with the MEC server for computing.

[0130] Step 2: Mathematically model the current system model diagram. Specifically, for the trajectory model and anti-collision model of the UAV, at the tth time slot, when the three-dimensional Cartesian coordinates of the user UEm are represented by p m (t) = [x m (t),y m (t),0], the three-dimensional Cartesian coordinates of the location of the UAVn are expressed as q n (t) = [x n (t),y n (t), H]. Therefore, the Euclidean distance between user UEm and drone UAVn can be expressed as:

[0131]

[0132] At the tth time slot, the three-dimensional Cartesian coordinates of the UAVn’s location are expressed as q n (t) = [x n (t),y n (t),H], in a given area along the horizontal angle θ m (t)∈[0,2π) direction and velocity v n (t)∈[0,v max ] flight, the horizontal coordinate of the UAVn at the t+1th time slot can be expressed as:

[0133]

[0134] Because the drone needs to serve users within the user service area, its horizontal coordinates must meet the following requirements:

[0135]

[0136] where X max and Y max Indicates the boundaries of the service area;

[0137] Because there are multiple drones in the system, in order to avoid collisions between UAVs, they should maintain a minimum distance between them:q m (t)-q n (t)≥S min ; Among them, S min A safe distance to avoid collision.

[0138] For the communication model in the process of drone-assisted edge computing task offloading, line-of-sight environment is a key factor that can affect the transmission speed between the drone and the terminal. Due to obstruction by high-rise buildings and trees, line-of-sight environment cannot always be guaranteed. Therefore, the line-of-sight environment between the drone and the user is random. The probability of line-of-sight environment between the drone and the user can be expressed by the following function expression:

[0139]

[0140] Among them, C1 and C2 are constants determined by the environment. It represents the elevation angle between the user and the drone at the tth time slot, which can be expressed as follows:

[0141]

[0142] The probability of the drone and the user being in a non-line-of-sight environment can be expressed as:

[0143]

[0144] Depending on the LoS or NLoS connection between the drone and the user, the path loss of the received signal power of each user can be expressed as:

[0145]

[0146] Where f is the carrier frequency of the system, η Los and η NLos are the system constants in line-of-sight and non-line-of-sight environments, respectively, and c is the speed of light.

[0147] In order to simplify the path loss model of the UAV in the system and fully consider the path loss of the UAV in both line-of-sight and non-line-of-sight environments, the path probability is used to perform a weighted average on the above path loss formula to obtain the average path loss of the UAV, which can be expressed as:

[0148]

[0149] According to the average path loss between the UAVn and the user UEm, the channel gain between the UAVn and the user UEm can be obtained:

[0150]

[0151] Where α0 represents the power gain at a reference distance of 1 m.

[0152] According to the channel gain between the UAVn and the user UEm, the transmission rate between the UAVn and the user UEm can be:

[0153] Where B is the bandwidth of wireless communication, Uplink transmission power between UAVn and UEm, σ 2 is the noise power spectral density.

[0154] The computational model in the task offloading process mainly includes two methods. Specifically, in the tth time slot, part of the computational task of user UEm is offloaded to the UAVn. The proportion of this part of the task to the total task of the user is Indicates that the remaining task data is processed locally. The proportion of this part of the task volume to the total tasks of the user is expressed as express.

[0155] This system will user UE k and UAV m The uninstall matching relationship is recorded as l m,n (t)={0,1}, when l m,n When (t) = 1, it indicates that the user UE k and UAV m Establish a connection, otherwise, l m,n (t) = 0. Assuming that there is only one drone providing services to the user in each time slot, the following should be satisfied:

[0156]

[0157] (1) Local computing model

[0158] In the tth time slot, the delay caused by local calculation can be expressed as:

[0159] Among them, D m (t) is the amount of tasks to be processed by user UEm in the tth time slot, s is the number of CPU cycles required to process each unit bit of data, and f m (t) represents the computing capability of user UEm;

[0160] In the tth time slot, the energy consumption generated by local computing can be expressed as:

[0161] in, It is the power factor determined by the CPU architecture;

[0162] (2) Model that offloads tasks to a drone equipped with a MEC server

[0163] The delay in offloading the task to the UAV equipped with the MEC server mainly consists of two parts: the transmission delay of the task from user UEm to UAVn and the computation delay of the task from user UEm to UAVn. Therefore, in the tth time slot, the transmission delay of the task from user UEm to UAVn can be expressed as:

[0164]

[0165] In the tth time slot, user UEm offloads the task to UAVn to calculate the delay, which can be expressed as:

[0166]

[0167] Among them, f n (t) represents the computing resources allocated to user UEm by UAVn in the tth time slot;

[0168] The energy consumption generated by offloading tasks to the drone equipped with MEC server mainly includes three parts: the flight energy consumption of the drone, the transmission energy consumption of user UEm offloading tasks to drone UAVn, and the computing energy consumption of user UEm offloading tasks to drone UAVn. Therefore, in the tth time slot, the transmission energy consumption of user UEm offloading tasks to drone UAVn can be expressed as:

[0169]

[0170] At the tth time slot, the energy consumption generated by the UAVn flight can be expressed as:

[0171]

[0172] Among them, M иаv The quality of UAVn;

[0173] In the tth time slot, user UEm offloads the task to the UAVn to calculate the energy consumption, which can be expressed as:

[0174]

[0175] Among them, κ is the impact factor of hardware structure on the CPU processing of the UAV;

[0176] Because local computation and computation offloaded to the UAV are performed simultaneously, the total delay of the system at time slot t can be expressed as:

[0177] The total energy consumption of the system includes the energy consumption generated by local computing, the flight energy consumption of the UAV, the transmission energy consumption of the user UEm offloading the task to the UAVn, and the computing energy consumption of the user UEm offloading the task to the UAVn, which can be expressed as:

[0178] The cost of task processing includes the weighted sum of the total latency and total energy consumption in the multi-UAV assisted edge computing system at the tth time slot, which can be expressed as follows:

[0179] L(t)=λ1T(t)+λ2E(t)

[0180] Step 3: Construct a user energy collection model. Specifically, each user is equipped with an energy collection module, which can charge the battery through wireless signals. Assume that energy collection begins at each time interval. In the initial state, it is assumed that the battery capacity of each user is full. The maximum battery capacity is Therefore, the energy of each user in the next time slot depends on the energy consumption and harvest of this time slot, which can be expressed by the following formula:

[0181]

[0182] Among them, e mn Indicates the energy collected by the user at the current moment.

[0183] Step 4: With the goal of minimizing the weighted sum of latency and energy consumption during task processing, an optimization problem is constructed. Specifically, in the system for building multi-UAV-assisted mobile edge computing task offloading, the weighted sum of latency and energy consumption during task processing is minimized by jointly optimizing the UAV location scheduling and task offloading strategy. Therefore, the optimization problem is constructed as:

[0184]

[0185] C2:θ n (t)∈[0,2π)

[0186] C3:v n (t)∈[0,v max ]

[0187] C4:0≤x n (t)≤X max

[0188] C5:0≤y n (t)≤Y max

[0189] C6:l m,n (t)={0,1}

[0190]

[0191] C8:λ1+λ2=1,λ1∈[0,1],λ2∈[0,1],

[0192] C9:q n (t)-q k (t)≥S min

[0193] C10:l m,n (t)R m,n (t)≤R max

[0194]

[0195] C13:T(t)≤T max (t)

[0196] Among them, constraint C1 represents the value range of the task offloading ratio; constraint C2 represents the constraint on the flight speed of UAVn; constraint C3 represents the constraint on the flight speed of UAVn; constraints C4 and C5 represent that UAVn moves within the area serving the user; constraint C6 represents whether user UEm establishes a connection with UAVn; constraint C7 indicates that in the tth time slot, UAVn can only provide service to one user; constraint C8 constrains the weight factors of the total system delay and total energy consumption; constraint C9 represents the constraint on the safe distance between UAVs; constraint C10 indicates that when user UEm decides to offload the task to UAVn, user UEm must be within the coverage range of UAVn; constraint C11 indicates that the battery power of user UEm does not exceed the minimum power threshold; constraint C12 indicates that the flight energy consumption and computing energy consumption of the UAV cannot exceed the maximum energy of its battery during the entire service cycle; constraint C13 indicates that the total delay of the system must be less than the maximum tolerable delay.

[0197] Step 5: Establish the optimization problem as a Markov decision process and use the CPER-MATD3 algorithm to solve the unmanned task offloading strategy in the UAV-assisted edge computing system. The specific content is as follows: Establish the optimization problem as a Markov decision process and use the CPER-MATD3 algorithm to solve the unmanned task offloading strategy in the UAV-assisted edge computing system. The Markov decision process consists of state, action and reward function;

[0198] State space S(t): includes the position information q of the UAVn n (t), user UEm location information p m (t), the size of the task amount generated by user UEM D n(t), the remaining power of the UAVn User UEm's battery status b m (t), can be defined as:

[0199]

[0200] Action space A(t): The action space is where the agent explores the state space and takes relevant actions, including task offloading strategies The unloading matching relationship between user UEm and drone UAVn is l m,n (t), the UAV flight deflection angle θ n (t), UAV flight speed ν n (t), can be expressed as:

[0201]

[0202] in, Indicates the task offloading strategy, l m,n (t) represents the unloading matching relationship between user UEm and UAV UAVn, θ n (t) represents the flight deflection angle of the UAVn, ν n (t) represents the flight speed of UAVn;

[0203] Reward function R(t): Since the goal of the optimization problem is to minimize the total system cost, the negative of the total system cost is used as the reward function, which can be defined as: R(t) = -[L(t) + Φ];

[0204] Where Φ is a penalty function, including penalties when the distance between drones is less than the safe distance, when the drone exceeds the minimum power threshold, and when the task completion time exceeds the maximum tolerable delay;

[0205] After modeling the Markov decision process of the optimization problem, the CPER-MATD3 algorithm was proposed to solve the unmanned task offloading strategy in the UAV-assisted edge computing system.

[0206] The MATD3 algorithm is an extension of the TD3 algorithm for the multi-agent domain. Its architecture still utilizes centralized training and independent execution. Specifically, the input space of each agent's value function includes not only its own observations and actions, but also the observations and actions of all other agents. In the MATD3 algorithm, each agent uses two critic networks to independently estimate the action-value function, similar to the concept of Double DON. During training, the outputs of the two critic networks are not directly used. Instead, the minimum of the two is taken as the actual estimate. This effectively reduces the volatility of the estimate, bringing it closer to the true value and, to a certain extent, addressing the overestimation problem. Furthermore, the MATD3 algorithm employs a policy-delayed update method and a soft update strategy. Policy-delayed update refers to inconsistent update frequencies for the actor and critic network parameters. Soft update involves the new parameters being equal to the weighted average of the old parameters and the new target parameters. This reduces fluctuations during parameter updates, making network training more stable and further reducing the overestimation problem.

[0207] In the traditional MATD3 algorithm, the experience replay mechanism selects training samples through random sampling. This method fails to distinguish the importance of different experience samples, resulting in the ineffective discovery and utilization of high-value experience samples, thus limiting the efficiency of network training. Although some studies have proposed using a prioritized experience replay mechanism to improve sampling efficiency, this method requires calculating and ranking the TD-error of all samples, neglecting experiences with high immediate reward values, which are also highly important. Therefore, this paper proposes the CPER-MATD3 algorithm to address the problem of edge computing task offloading in multi-UAV-assisted environments. This algorithm not only considers experiences with high TD-error values, but also incorporates experiences with high immediate reward values ​​as evaluation metrics. Through a composite priority sampling mechanism, it achieves efficient sample selection and significantly improves the utilization of experience samples.

[0208] The composite priority sampling mechanism mainly includes five steps:

[0209] The first step is to determine the TD-error value and immediate reward value of the agent in the current state

[0210] The second step is to calculate the priority of the experience samples in the immediate return standard and the TD-error standard respectively.

[0211] Y i =r t +ε

[0212] Y j =|δ t |+ε

[0213] The third step is to sort the priorities of the experience samples in the immediate return standard and the TD-error standard in ascending order to obtain rank(i) and rank(j), and then use the composite average sorting:

[0214]

[0215] The fourth step is to calculate the composite priority:

[0216]

[0217] Among them, α represents the relative weight of the evaluation priority.

[0218] The fifth step is to define the sampling probability:

[0219]

[0220] The steps of the CPER-MATD3 algorithm are as follows:

[0221] Step 1: Initialize the Actor network, Critic network and its target network, initialize the experience replay pool, and initialize the simulation parameters of the drone MEC system.

[0222] Step 2: The drone decides its actions based on the current strategy

[0223] Step 3: After executing the action, the drone observes the next state and immediate reward, combines the state, action, reward and next state into a four-tuple and stores it in the experience replay pool.

[0224] Step 4: Use a compound priority sampling mechanism to extract small batches of samples from the experience replay pool to update the Actor and Critic networks.

[0225] Step 5: The parameters of the Critic network are updated by minimizing the loss function. The update rule of the Critic network is:

[0226]

[0227] Step 6: The Actor network is updated by sampling policy gradients. The update rule is:

[0228]

[0229] Step 7: Update the target network through soft update. The update rules are as follows:

[0230]

[0231] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A method for offloading tasks based on drone-assisted edge computing, characterized by: The steps include: Step 1: Build a drone-assisted edge computing system model, which consists of M user UEs and N UAVs equipped with MEC servers. When user computing resources are insufficient, tasks can be offloaded to the UAVs equipped with MEC servers for computing. Step 2: Mathematically model the current system model, including the trajectory model and collision avoidance model of the UAV, the communication model during task offloading, and the computational model during task offloading. The computational model during task offloading mainly includes two methods: local computation of the task by the user or offloading the task to a UAV equipped with an MEC server for computation. Step 3: Build a user energy harvesting model; Step 4: Construct an optimization problem with the goal of minimizing the weighted sum of latency and energy consumption during task processing. Step 5: The optimization problem is formulated as a Markov decision process, and the CPER-MATD3 algorithm is used to solve the unmanned task offloading strategy in the UAV-assisted edge computing system.

2. The method for offloading tasks based on drone-assisted edge computing according to claim 1, characterized in that: The specific content of step 1 is: the UAV-assisted edge computing system model adopts a discrete time model, which divides the communication time of the entire system into T time slots equally, and each time slot lasts for τ, including the service time τ com and flight time τ fly In this communication mode, only one user establishes contact with the drone in each time slot; in this system model, M drones fly on a plane with a fixed height of H. Each drone is equipped with a high-performance edge computing server. When the user's computing resources are insufficient, the task can be offloaded to the drone equipped with the MEC server for calculation.

3. The method for offloading tasks based on drone-assisted edge computing according to claim 1, characterized in that: The specific content of step 2 is: for the trajectory model and anti-collision model of the UAV, in the tth time slot, when the three-dimensional Cartesian coordinates of the user UEm are represented by p m (t) = [x m (t),y m (t),0], the three-dimensional Cartesian coordinates of the location of the UAVn are expressed as q n (t) = [x n (t),y n (t), H]; therefore, the Euclidean distance between user UEm and drone UAVn can be expressed as: At the tth time slot, the three-dimensional Cartesian coordinates of the UAVn’s location are expressed as q n (t) = [x n (t),y n (t),H], in a given area along the horizontal angle θ m (t)∈[0,2π) direction and velocity v n (t)∈[0,v max ] flight, the horizontal coordinate of the UAVn at the t+1th time slot can be expressed as: Because the drone needs to serve users within the user service area, its horizontal coordinates must meet the following requirements: where X max and Y max Indicates the boundaries of the service area; Because there are multiple drones in the system, in order to avoid collisions between UAVs, they should maintain a minimum distance between them:q m (t)-q n (t)≥S min ; Among them, S min To avoid collisions, keep a safe distance; For the communication model in the process of drone-assisted edge computing task offloading, the line-of-sight environment is an important factor that can affect the transmission speed between the drone and the terminal. Due to the obstruction of high-rise buildings and trees, the line-of-sight environment cannot always be guaranteed. Therefore, the line-of-sight environment between the drone and the user is random. The probability of the line-of-sight environment between the drone and the user can be expressed by the following function expression: Among them, C1 and C2 are constants determined by the environment. It represents the elevation angle between the user and the drone at the tth time slot, which can be expressed as follows: The probability of the drone and the user being in a non-line-of-sight environment can be expressed as: Depending on the LoS or NLoS connection between the drone and the user, the path loss of the received signal power of each user can be expressed as: Where f is the carrier frequency of the system, η Los and η NLos are the system constants in line-of-sight and non-line-of-sight environments, respectively, and c is the speed of light; In order to simplify the path loss model of the UAV in the system and fully consider the path loss of the UAV in both line-of-sight and non-line-of-sight environments, the path probability is used to perform a weighted average on the above path loss formula to obtain the average path loss of the UAV, which can be expressed as: According to the average path loss between the UAVn and the user UEm, the channel gain between the UAVn and the user UEm can be obtained: Where α0 represents the power gain at a reference distance of 1 m; According to the channel gain between the UAVn and the user UEm, the transmission rate between the UAVn and the user UEm can be: Where B is the bandwidth of wireless communication, Uplink transmission power between UAVn and UEm, σ 2 is the noise power spectral density; The computational model in the task offloading process mainly includes two methods. Specifically, in the tth time slot, part of the computational task of user UEm is offloaded to the UAVn. The proportion of this part of the task to the total task of the user is Indicates that the remaining task data is processed locally. The proportion of this part of the task volume to the total tasks of the user is expressed as express; User UE k and UAV m The uninstall matching relationship is recorded as l m,n (t)={0,1}, when l m,n When (t) = 1, it indicates that the user UE k and UAV m Establish a connection, otherwise, l m,n (t) = 0; Assuming that there is only one drone providing services to the user in each time slot, it should satisfy: (1) Local computing model In the tth time slot, the delay caused by local calculation can be expressed as: Among them, D m (t) is the amount of tasks to be processed by user UEm in the tth time slot, s is the number of CPU cycles required to process each unit bit of data, and f m (t) represents the computing capability of user UEm; In the tth time slot, the energy consumption generated by local computing can be expressed as: in, It is the power factor determined by the CPU architecture; (2) Model that offloads tasks to a drone equipped with a MEC server The delay caused by offloading the task to the UAV equipped with the MEC server mainly consists of two parts: the transmission delay of user UEm offloading the task to UAVn and the computation delay of user UEm offloading the task to UAVn. Therefore, in the tth time slot, the transmission delay of user UEm offloading the task to UAVn can be expressed as: In the tth time slot, user UEm offloads the task to UAVn to calculate the delay, which can be expressed as: Among them, f n (t) represents the computing resources allocated to user UEm by UAVn in the tth time slot; The energy consumption generated by offloading the task to the UAV equipped with the MEC server mainly includes three parts: the flight energy consumption of the UAV, the transmission energy consumption of the user UEm when offloading the task to the UAVn, and the computing energy consumption of the user UEm when offloading the task to the UAVn. Therefore, in the tth time slot, the transmission energy consumption of the user UEm when offloading the task to the UAVn can be expressed as: At the tth time slot, the energy consumption generated by the UAVn flight can be expressed as: Among them, M иаv The quality of UAVn; In the tth time slot, user UEm offloads the task to the UAVn to calculate the energy consumption, which can be expressed as: Among them, κ is the impact factor of hardware structure on the CPU processing of the UAV; Because local computation and computation offloaded to the UAV are performed simultaneously, the total delay of the system at time slot t can be expressed as: The total energy consumption of the system includes the energy consumption generated by local computing, the flight energy consumption of the UAV, the transmission energy consumption of the user UEm offloading the task to the UAVn, and the computing energy consumption of the user UEm offloading the task to the UAVn, which can be expressed as: The cost of task processing includes the weighted sum of the total latency and total energy consumption in the multi-UAV assisted edge computing system at the tth time slot, which can be expressed as follows: L(t)=λ1T(t)+λ2E(t).

4. The method for offloading tasks based on drone-assisted edge computing according to claim 1, characterized in that: The specific content of step 3 is to equip each user with an energy collection module, which can charge the battery through wireless signals. Assuming that energy collection begins at each time interval, in the initial state, assuming that the battery capacity of each user is full, the maximum battery capacity is Therefore, the energy of each user in the next time slot depends on the energy consumption and harvest of this time slot, which can be expressed by the following formula: Among them, e mn Indicates the energy collected by the user at the current moment.

5. The method for offloading tasks based on drone-assisted edge computing according to claim 1, characterized in that: The specific content of step 4 is: in the construction of a multi-UAV assisted mobile edge computing task offloading system, the weighted sum of the delay and energy consumption during task processing is minimized by jointly optimizing the position scheduling and task offloading strategy of the UAVs. Therefore, the optimization problem is constructed as follows: Among them, constraint C1 represents the value range of the task offloading ratio; constraint C2 represents the constraint on the flight speed of UAVn; constraint C3 represents the constraint on the flight speed of UAVn; constraints C4 and C5 represent that UAVn moves within the area serving the user; constraint C6 represents whether user UEm establishes a connection with UAVn; constraint C7 indicates that in the tth time slot, UAVn can only provide service to one user; constraint C8 constrains the weight factors of the total system delay and total energy consumption; constraint C9 represents the constraint on the safe distance between UAVs; constraint C10 indicates that when user UEm decides to offload the task to UAVn, user UEm must be within the coverage range of UAVn; constraint C11 indicates that the battery power of user UEm does not exceed the minimum power threshold; constraint C12 indicates that the flight energy consumption and computing energy consumption of the UAV cannot exceed the maximum energy of its battery during the entire service cycle; constraint C13 indicates that the total delay of the system must be less than the maximum tolerable delay.

6. The method for offloading tasks based on drone-assisted edge computing according to claim 1, characterized in that: The specific content of step 5 is: establishing the optimization problem as a Markov decision process, and using the CPER-MATD3 algorithm to solve the unmanned task offloading policy in the UAV-assisted edge computing system. The Markov decision process consists of state, action and reward function; State space S(t): includes the position information q of the UAVn n (t), user UEm location information p m (t), the size of the task amount generated by user UEM D n (t), the remaining power of the UAVn User UEm's battery status b m (t), can be defined as: Action space A(t): The action space is where the agent explores the state space and takes relevant actions, including task offloading strategies The unloading matching relationship between user UEm and drone UAVn is l m,n (t), the UAV flight deflection angle θ n (t), UAV flight speed ν n (t), can be expressed as: in, Indicates the task offloading strategy, l m,n (t) represents the unloading matching relationship between user UEm and UAV UAVn, θ n (t) represents the flight deflection angle of the UAVn, ν n (t) represents the flight speed of UAVn; Reward function R(t): Since the goal of the optimization problem is to minimize the total system cost, the negative of the total system cost is used as the reward function, which can be defined as: R(t) = -[L(t) + Φ]; Where Φ is a penalty function, including penalties when the distance between drones is less than the safe distance, when the drone exceeds the minimum power threshold, and when the task completion time exceeds the maximum tolerable delay; After modeling the Markov decision process of the optimization problem, the CPER-MATD3 algorithm is proposed to solve the unmanned task offloading strategy in the UAV-assisted edge computing system; The MATD3 algorithm is an extension of the TD3 algorithm in the field of multi-agents. Its structure still adopts the form of centralized training and independent execution, that is, the input space of the value function of each agent includes not only its own observations and actions, but also the observations and actions of all other agents. In the MATD3 algorithm, each agent uses two Critic networks to estimate the action value function separately, which is similar to the idea of ​​DoubleDON. During the training process, the outputs of the two Critic networks are not directly used. Instead, the minimum value of the two is taken as the actual estimated value, which effectively reduces the volatility of the estimate and makes the estimate closer to the true value, thereby solving the overestimation problem to a certain extent. In addition, the MATD3 algorithm also adopts a policy delayed update method and a soft update strategy. Policy delayed update refers to the inconsistent update frequency of the network parameters of the Actor and Critic. Soft update means that the new parameters are equal to the weighted average of the old parameters and the new target parameters. This can reduce the fluctuation during parameter update, make the network training more stable, and further reduce the overestimation problem. In the traditional MATD3 algorithm, the experience replay mechanism selects training samples through random sampling. This method fails to distinguish the importance of different experience samples, resulting in the failure to effectively discover and utilize high-value experience samples, thereby limiting the efficiency of network training. Although some studies have proposed using a prioritized experience replay mechanism to improve sampling efficiency, this method requires calculating and sorting the TD-error of all samples. This method ignores experiences with high immediate return values, which are also highly important. Therefore, this paper proposes the CPER-MATD3 algorithm to solve the problem of edge computing task offloading under the assistance of multiple drones. This algorithm not only considers experiences with high TD-error values, but also incorporates experiences with high immediate return values ​​as evaluation indicators. Through a composite priority sampling mechanism, it achieves efficient sample selection and significantly improves the utilization rate of experience samples. The composite priority sampling mechanism mainly includes five steps: The first step is to determine the TD-error value and immediate reward value of the agent in the current state; The second step is to calculate the priority of the experience samples in the immediate reward standard and the TD-error standard respectively; Y i =r t +e Y j =|δ t |+e The third step is to sort the priorities of the experience samples in the immediate return standard and the TD-error standard in ascending order to obtain rank(i) and rank(j), and then use the composite average sorting: The fourth step is to calculate the composite priority: Among them, α represents the relative weight of the evaluation priority; The fifth step is to define the sampling probability: The steps of the CPER-MATD3 algorithm are as follows: Step 1: Initialize the Actor network, Critic network, and target network, initialize the experience replay pool, and initialize the simulation parameters of the drone MEC system; Step 2: The drone decides its actions based on the current strategy; Step 3: After executing the action, the drone observes the next state and immediate reward, combines the state, action, reward, and next state into a four-tuple, and stores it in the experience replay pool; Step 4: Use a composite priority sampling mechanism to extract small batches of samples from the experience replay pool to update the Actor and Critic networks; Step 5: The parameters of the Critic network are updated by minimizing the loss function. The update rule of the Critic network is: Step 6: The Actor network is updated by sampling policy gradients. The update rule is: Step 7: Update the target network through soft update. The update rules are as follows: