Energy-saving low-delay task unloading method applied to mobile edge computing system
By adopting multi-UAV collaboration mechanism and deep reinforcement learning algorithm in the mobile edge computing system, the problems of high system power consumption, high delay and unstable training are solved, and more efficient and stable task offloading services are achieved.
Patent Information
- Application Number
- CN202510184087.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-27
AI Technical Summary
Existing mobile edge computing systems have problems such as high power consumption, high latency, and training instability due to falling into local optimality.
The multi-UAV collaboration mechanism is adopted, combined with a deep reinforcement learning algorithm, and optimized task offloading strategies are optimized by minimizing the weighted sum of energy consumption and delay by minimizing the weighted sum of energy consumption and delay.
It realizes more efficient task offloading services, improves the system's task processing efficiency and stability, and reduces energy consumption and time delay.
Smart Images

Figure CN120216115A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of mobile edge computing and deep reinforcement learning, and specifically to an energy-saving and low-latency task offloading method applied to a mobile edge computing system. Background Art
[0002] With the rapid development of mobile Internet technology, the limitations of high power consumption and high latency of traditional cloud computing architectures have gradually emerged. Especially in scenarios with high real-time requirements, the transmission of data from user devices to the central server not only consumes time but also consumes energy. This problem is very obvious in a large-scale Internet of Things environment, especially in places with a high density of people gathering (such as concerts, sports events, large public gatherings, etc.) or emergency rescue operations, where traditional communication facilities are difficult to meet the real-time processing requirements.
[0003] To solve this problem, mobile edge computing (MEC) has received extensive attention as an effective technology. By migrating computing resources and services to the network edge, MEC can provide low-latency and low-energy-consuming computing capabilities, reducing the energy consumption caused by long-distance data transmission while alleviating the pressure on the central server. Internet of Things devices offload real-time and complex computing tasks to nearby edge computing servers, enabling fast data processing and response without relying on traditional cloud data centers, significantly reducing latency and energy consumption.
[0004] At the same time, due to its super strong decision-making ability in a dynamically changing network environment, the deep reinforcement learning (DRL) algorithm has been widely applied to the mobile edge computing offloading field in recent years. Deep reinforcement learning can dynamically allocate resources according to the real-time state of edge devices (such as CPU utilization, memory usage, network bandwidth, etc.) and the characteristics of tasks (such as task size, latency requirements, etc.), maximizing resource utilization.
[0005] However, despite the significant theoretical advantages of deep reinforcement learning, traditional deep reinforcement learning algorithms still face some challenges in practical applications. In complex tasks, if there are multiple local optimal solutions in the environment, it may get stuck in a local optimum due to insufficient exploration and fail to find the global optimal strategy. It is also sensitive to hyperparameters during the training process, which may lead to inaccurate value estimates for policy updates, further making the training unstable. Summary of the Invention
[0006] The objective of the present invention is to provide an energy-saving and low-latency task offloading method for a mobile edge computing system, which can solve the technical problems of high power consumption, high latency in existing mobile edge computing systems, and unstable training due to falling into local optimality.
[0007] To achieve the above objective, the present invention adopts the following technical solutions.
[0008] An energy-saving and low-latency task offloading method for a mobile edge computing system includes the following steps.
[0009] S1. Establish a multi-UAV-assisted mobile edge computing (MEC) system model, including UAVs, users, and eavesdropping users. Each UAV is equipped with a micro mobile edge computing server to provide task offloading services for users.
[0010] S2. In the system model described in S1, calculate the weighted sum of the energy consumption and latency of the system.
[0011] S3. Establish a deep reinforcement learning model to minimize the weighted sum of energy consumption and latency, and select the optimal offloading action by the deep reinforcement learning model.
[0012] S4. Train the deep reinforcement learning model described in S3 until the deep reinforcement learning model achieves the preset goal.
[0013] Further, in S1, the established system model is specifically that within a rectangular flight area with side length L, there are U UAVs equipped with micro mobile edge computing servers, K Internet of Things wireless device users, and one eavesdropping user, where the flight altitude of the UAVs is H.
[0014] The set of unmanned aerial vehicles (UAVs) is represented as U = {1,..., u,..., U}, the set of wireless devices (WDs) is represented as K = {1,..., k,..., K}, and the eavesdropping user is represented as e.
[0015] Construct a three-dimensional coordinate system XYZ, where the X-axis and Y-axis represent the ground coordinates of the rectangular flight area, and the Z-axis represents the height from the ground. Then the three-dimensional coordinates q u of the u-th UAV u = [x u , y k , H], and the position of the k-th wireless device user WD k is represented as q k = [x k , y k , 0], and the position of the eavesdropping user e is represented as q e = [xe , y e , 0].
[0016] Furthermore, in step S2, the energy consumption of the system includes the wireless device user WD k local computing energy consumption, the wireless device user WD k transmission energy consumption for offloading tasks to the UAV, the computing energy consumption of the micro mobile edge computing server on the UAV to complete the offloading tasks, and the energy consumption of the UAV flight.
[0017] The latency of the system includes the wireless device user WD k delay in completing local computing tasks, the wireless device user WD k delay in offloading tasks to the UAV, and the delay of the micro mobile edge computing server on the UAV to complete the offloading tasks.
[0018] Furthermore, the channel gain between the wireless device user WD k and the UAV is: where u0 is the channel power gain at a reference distance of 1 m, is the actual signal propagation path.
[0019] The transmission rate at which the wireless device user WD k offloads computing tasks to the UAV is: where B represents the channel bandwidth, p k represents the maximum transmit power of WD k in the upload link, and σ 2 is the noise power of the channel.
[0020] The channel gain between the eavesdropping user e and the UAV is:
[0021] The transmission rate from the eavesdropping user e to the UAV is:
[0022] The secrecy rate of the system is: C m = max{(R k - R e ), 0}.
[0023] In the MEC system, the tasks of WD k in each time slot use a partial offloading strategy, and the offloading ratio is x k , then the delay for the wireless device user WD k to complete local computing tasks is: where s represents the CPU cycles required to process a unit byte, and f l,k represents the CPU computing frequency of the wireless device user WD k , and D k represents the total task volume of the user.
[0024] The local power consumption of the wireless device user WD k is: where γ is the computing efficiency parameter.
[0025] Then the local computing energy consumption of the wireless device user WD k is:
[0026] The delay for the wireless device user WD k to offload tasks to the UAV is:
[0027] Then the transmission energy consumption for the wireless device user WD k to offload tasks to the UAV is:
[0028] The delay for the micro mobile edge computing server on the UAV to complete the offloaded tasks is: where f u represents the computing frequency of the micro mobile edge computing server.
[0029] The power consumption of the micro mobile edge computing server is:
[0030] Then the computing energy consumption for the micro mobile edge computing server on the UAV to complete the offloaded tasks is:
[0031] When the UAV flies from the current hovering position to the next hovering position, the new hovering position is: q (u+1) = [x (u) + v (u) t fly cosβ (u) , y (u) + v (u) t fly sinβ (u) , H], where v(u) is the flight speed of the drone, and v (u) ∈[0, v max , t fly is the fixed flight time of the drone, cosβ (u) and sinβ (u) respectively represent the projection components of the flight direction of the drone in the x-axis direction and the y-axis direction.
[0032] Then the energy consumption of the drone flight is: E fly(u) = φ||v (u) || 2 , where φ = 0.5M UAV t fly , M UAV is the takeoff mass of the drone.
[0033] The total energy consumption of the system is:
[0034] The total delay of the system is:
[0035] Then the weighted sum of the energy consumption and delay of the system is: E k + ηT k , where η is the weight factor of the delay.
[0036] Furthermore, in step S3, an optimization problem is constructed to minimize the weighted sum of the energy consumption and delay: where, represents the total energy consumption of the system, represents the total delay of the system, η is the weight factor of the delay, f u ≥0, f l,k ≥0, indicating that the CPU of the micro mobile edge computing server and the wireless device user WD k is operating normally, represents that the execution delay of any WD is less than the maximum delay acceptable to the WD, represents that the drone is flying within a rectangular area with side length L.
[0037] Further, in step S3, the model-free Soft Actor-Critic (SAC) algorithm based on the deep reinforcement learning algorithm is adopted, and a regularization term of policy entropy is introduced into the objective function to encourage the policy network to maintain randomness during the learning process and explore the environment more comprehensively.
[0038] A double Q-network structure is adopted to alleviate the problem of value overestimation, and the policy is updated by the smaller value of the estimated values of the two Q-networks to improve the accuracy of value function estimation.
[0039] Construct a capability matrix between wireless device users and target drones to help the system quickly perform effective allocation and scheduling and improve system efficiency.
[0040] Further, the deep reinforcement learning model includes:
[0041] Construct the state space S of the system in the deep reinforcement algorithm, the task allocation information D of the device (ass) , the number of resources required for each task D (res) , the completion status D of the assigned tasks (sta) , the reachability and service capabilities D between tasks and drones or central servers (abi) ;
[0042] Construct the action space A of the system in the deep reinforcement algorithm. During time slot t, the drone and the offloading environment interact with each other, and the optimal offloading action A* is selected according to the current state and the observed environment. The action set includes the probability P of each task being assigned to a drone or a central server (d) and the offloading ratio x of the task (k) ;
[0043] Construct a reward function: Maximize the reward by minimizing the weighted sum of energy efficiency and delay on the premise of ensuring secure transmission.
[0044] Calculate the reward corresponding to each execution of the action, observe the next state S k+1 , and store the obtained (S (k) , A (k) , R (k) , S (k+1) ) in the experience storage pool B m .
[0045] In the Critic module, randomly extract a small batch (S m , A (k) , R (k) , S (k) , S (k+1) ) from B, and calculate the target Q value using the Bellman equation: y i = r i + γ(min(Q1(s i+1 , μ(s i+1 ))), Q2(s i+1 , μ(s i+1 )) - αlogπ(a i+1 | s i+1 )), where γ is the discount factor, y i is the target Q-value, μ(s i+1 ) is the action output by the target policy for a given state S k+1 , α is the temperature parameter used to balance exploration and exploitation of the policy, and π(a i+1 | s i+1 ) is the probability of the action generated by the policy network.
[0046] The SAC algorithm uses two Q-networks, Q1 and Q2, to reduce the overestimation problem of Q-values. The value of the state-action pair estimated by each Q-network is different, and the minimum value is taken as the target Q-value.
[0047] The current Q-value is obtained by predicting the inputs of the current state space S and action space A through two Q-networks, which are Q1(s, a) and Q2(s, a) respectively.
[0048] The mean squared error loss is calculated using the difference between the target Q-value and the current Q-value, and the Q-value network parameters are updated. The loss function is: L(θ Q ) = E[(Q1(s, a) - y i ) 2 + (Q2(s, a) - y i ) 2 .
[0049] The Adam optimizer is used to minimize the loss function, update the network parameters of Q1 and Q2, and perform soft updates on the policy network and Q-network. The update logic is: θ Q′ ← τθ Q + (1 ← τ)θ Q′ , θ π′ ← τθ π + (1 - τ)θ π′ , where τ is the soft update parameter.
[0050] Furthermore, in step S4, loop for interaction, storage, sampling, training, and updating until the deep reinforcement learning model reaches the expected goal.
[0051] After adopting the above technical solution, the present invention has the following beneficial effects: 1. The present invention adopts a multi-UAV cooperation mechanism, and through the cooperation among multiple UAVs equipped with micro mobile edge computing servers and IoT wireless device users, it can provide a more efficient task offloading service; 2. The system in the present invention uses multiple UAVs. Compared with a single-UAV system, it can handle task allocation and resource scheduling more flexibly, thereby improving the efficiency and stability of task processing; 3. The present invention adopts a model-free SAC algorithm based on deep reinforcement learning. Compared with traditional deep reinforcement learning algorithms, it can not only actively utilize policy entropy to balance exploration, but also maintain adaptability to the environment in a dynamic environment. According to the selection problem between multiple users and multiple UAVs, a capability matrix between users and UAVs is constructed as the state information of the SAC algorithm for effective allocation and scheduling. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is an overall distribution schematic diagram of the system model established by the present invention;
[0053] Figure 2 It is a schematic diagram of the principle of the method of the deep reinforcement learning algorithm in the present invention;
[0054] Figure 3 It is a comparison chart of the results of training the SAC algorithm adopted in the present invention for 1000 rounds with the DDPG (Deep Deterministic Policy Gradient) algorithm and the TD3 (Twin Delayed Deep Deterministic Policy Gradient) algorithm;
[0055] Figure 4 It is a comparison chart of the results of a wireless device user in the present invention choosing a UAV to offload tasks based on a capability matrix and the wireless device user evenly (average allocation) offloading tasks to UAVs;
[0056] Figure 5 It is a comparison chart of the results of training the SAC algorithm in the present invention for 1000 rounds under different temperature coefficients (alpha);
[0057] Figure 6 It is a comparison chart of the results of training the SAC algorithm in the present invention for 1000 rounds under different discount factors (gamma);
[0058] Figure 7 It is a comparison chart of the results of training the SAC algorithm in the present invention for 1000 rounds under different task lengths (task size). DETAILED DESCRIPTION OF THE INVENTION
[0059] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the features and performance of an energy-saving and low-latency task offloading method applied to a mobile edge computing system in the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0060] An energy-saving and low-latency task offloading method applied to a mobile edge computing system includes the following steps.
[0061] S1. Establish a multi-UAV-assisted mobile edge computing (MEC) system model, including UAVs, users, and eavesdropping users. Each UAV is equipped with a micro mobile edge computing server to provide task offloading services for users.
[0062] The established system model is specifically as follows: in a rectangular flight area with a side length of L, there are U UAVs equipped with micro mobile edge computing servers, K Internet of Things wireless device users, and one eavesdropping user, where the flight height of the UAVs is H.
[0063] The set of UAVs (Unmanned Aerial Vehicles) is represented as U = {1,..., u,..., U}, the set of wireless devices (Wireless Users) is represented as K = {1,..., k,..., K}, and the eavesdropping user is represented as e.
[0064] Construct a three-dimensional coordinate system XYZ, where the X-axis and Y-axis represent the ground coordinates of the rectangular flight area, and the Z-axis represents the height from the ground. Then, the three-dimensional coordinates q u of the u-th UAV u = [x u , y k , H], the position of the k-th wireless device user WD k is represented as q k = [x k , y e , 0], and the position of the eavesdropping user e is represented as q e = [x e , y
[0065] S2. In the system model described in S1, calculate the weighted sum of the energy consumption and latency of the system.
[0066] The energy consumption of the system includes the local computing energy consumption of the wireless device user WD k , the transmission energy consumption of the wireless device user WD k for offloading tasks to the UAV, the computing energy consumption of the micro mobile edge computing server on the UAV to complete the offloaded tasks, and the energy consumption of the UAV flight.
[0067] The time delay of the system includes the wireless device user WD k The delay in completing the local computing task, the wireless device user WD k The delay in offloading the task to the UAV, and the delay of the micro mobile edge computing server on the UAV in completing the offloading task.
[0068] The wireless device user WD k The channel gain between the wireless device user WD and the UAV is: where u0 is the channel power gain at a reference distance of 1m, is the actual signal propagation path.
[0069] The wireless device user WD k The transmission rate for offloading the computing task to the UAV by the wireless device user WD is: where B represents the channel bandwidth, p k represents the maximum transmit power of the WD k in the uplink, and σ 2 is the noise power of the channel.
[0070] The channel gain between the eavesdropping user e and the UAV is:
[0071] The transmission rate from the eavesdropping user e to the UAV is:
[0072] The secrecy rate of the system is: C m = max{(R k - R e ), 0}.
[0073] In the MEC system, the WD k uses a partial offloading strategy for the tasks in each time slot, and the offloading ratio is x k , then the delay of the wireless device user WD k in completing the local computing task is: where s represents the CPU cycles required to process a unit byte, and f l,k represents the CPU computing frequency of the wireless device user WD k , and D k represents the total task volume of the user.
[0074] The local power consumption of the wireless device user WD k is: where γ is the computing efficiency parameter.
[0075] Then, the local computing energy consumption of the wireless device user WD k is:
[0076] The wireless device user WD k The delay for offloading tasks to the UAV is:
[0077] Then, the transmission energy consumption of the wireless device user WD k for offloading tasks to the UAV is:
[0078] The delay for the micro mobile edge computing server on the UAV to complete the offloading task is: where f u represents the computing frequency of the micro mobile edge computing server.
[0079] The power consumption of the micro mobile edge computing server is:
[0080] Then, the computing energy consumption of the micro mobile edge computing server on the UAV to complete the offloading task is:
[0081] The UAV flies from the current hovering position to the next hovering position, and the new hovering position is: q (u+1) = [x (u) + v (u) t fly cosβ (u) , y (u) + v (u) t fly sinβ (u) , H], where v (u) is the flight speed of the UAV, and v (u) ∈[0, v max , t fly is the fixed flight time of the UAV, and they respectively represent the projection components of the flight direction of the UAV on the x-axis direction and the y-axis direction.
[0082] Then, the energy consumption of the UAV flight is: E fly(u) = φ||v (u) ||2 , where φ = 0.5M UAV t fly , M UAV is the takeoff mass of the UAV.
[0083] The total energy consumption of the system is:
[0084] The total delay of the system is:
[0085] Then the weighted sum of the energy consumption and delay of the system is: E k + ηT k , where η is the weight factor of the delay.
[0086] Furthermore, in step S3, an optimization problem is constructed by minimizing the weighted sum of the energy consumption and delay: where represents the total energy consumption of the system, represents the total delay of the system, η is the weight factor of the delay, f u ≥ 0, f l,k ≥ 0, indicating that the CPUs of the micro mobile edge computing server and the wireless device user WD k are operating normally, represents that the execution delay of any WD is less than the maximum delay acceptable to the WD, represents that the UAV is flying within a rectangular area with side length L.
[0087] S3. Establish a deep reinforcement learning model by minimizing the weighted sum of the energy consumption and delay, and select the optimal offloading action by the deep reinforcement learning model.
[0088] The model-free Soft Actor-Critic (SAC) algorithm based on the deep reinforcement learning algorithm is adopted, and a regularization term of policy entropy is introduced into the objective function to encourage the policy network to maintain randomness during the learning process and explore the environment more comprehensively.
[0089] The double Q-network structure is adopted to alleviate the problem of value overestimation, and the policy is updated by the smaller value among the estimated values of the two Q-networks to improve the accuracy of the value function estimation.
[0090] Construct a capability matrix between wireless device users and target drones to assist the system in quickly performing effective allocation and scheduling, thereby improving system efficiency.
[0091] Furthermore, the deep reinforcement learning model includes:
[0092] Construct the state space S of the system in the deep reinforcement algorithm, the task allocation information D of the device (ass) , the amount of resources required for each task D (res) , the completion status D of the assigned tasks (sta) , the reachability and service capabilities D between tasks and drones or the central server (abi) ;
[0093] Construct the action space A of the system in the deep reinforcement algorithm. During time slot t, the drones and the offloading environment interact with each other, and the optimal offloading action A* is selected according to the current state and the observed environment. The action set includes the probability P of each task being assigned to a drone or the central server (d) and the offloading ratio x of the task (k) ;
[0094] Construct the reward function: Maximize the reward by minimizing the weighted sum of energy efficiency and latency while ensuring secure transmission.
[0095] Calculate the reward corresponding to each execution of the action and observe the next state S k+1 , and store the obtained (S (k) , A (k) , R (k) , S (k+1) ) in the experience storage pool B m .
[0096] In the Critic module, randomly extract a small batch (S m , A (k) , R (k) , S (k) , S (k+1) ) from B, and calculate the target Q value using the Bellman equation: y i = r i + γ(min(Q1(s i+1 , μ(s i+1 )), Q2(s i+1 , μ(s i+1 ))) - αlogπ(a i+1 |s i+1 )), where γ is the discount factor and y i is the target Q value, and μ(si+1 ) is for the target policy to output the action at the given state S k+1 , α is the temperature parameter used to balance the exploration and exploitation of the policy, and π(a i+1 |s i+1 ) is the probability of the action generated by the policy network.
[0097] The SAC algorithm uses two Q-networks Q1 and Q2 to reduce the overestimation problem of Q-values. The value of each state-action pair estimated by each Q-network is different, and the minimum value is taken as the target Q-value.
[0098] The current Q-values are obtained by predicting the inputs of the current state space S and action space A through two Q-networks, which are Q1(s,a) and Q2(s,a) respectively.
[0099] The mean square error loss is calculated using the difference between the target Q-value and the current Q-value, and the Q-value network parameters are updated. The loss function is: L(θ Q ) = E[(Q1(s,a) - y i ) 2 +(Q2(s,a) - y i ) 2 .
[0100] The Adam optimizer is used to minimize the loss function, update the network parameters of Q1 and Q2, and perform a soft update on the policy network and Q-network. The update logic is: θ Q′ ← τθ Q +(1 - τ)θ Q′ , θ π′ ← τθ π +(1 - τ)θ π′ , where τ is the soft update parameter.
[0101] S4. Train the deep reinforcement learning model described in S3, and perform cyclic interaction, storage, sampling, training, and updating until the deep reinforcement learning model reaches the expected goal.
[0102] Specifically, the present invention provides an energy-saving and low-latency task offloading method applied to a mobile edge computing system, which uses the Deep Reinforcement Learning (DRL) algorithm to achieve the purpose of energy saving by minimizing the weighted sum of energy consumption and latency.
[0103] The system model is as Figure 1As shown in the figure, we designed U unmanned aerial vehicles (UAVs), and each UAV is equipped with a micro mobile edge computing server to provide computing power. All UAVs operate at a constant height H within a rectangular area of length L. In this area, there are K users using Internet of Things devices and one eavesdropping user. The set of UAVs is denoted as U = {1,..., u,..., U}, the set of wireless devices (WDs) is denoted as K = {1,..., k,..., K}, and the eavesdropping user is denoted as e.
[0104] In the whole system, WDs can offload the backlogged tasks to the UAVs, and after being assisted by the UAVs for computing, the results are transmitted to the users. We consider that the whole system operates in discrete time. Assuming that the computing power of a single WD is limited, in each time frame, it is selected to offload part of the tasks to the UAVs or a fixed central server for computing.
[0105] An Euclidean distance in a three-dimensional space is constructed to define the coordinate system. The X-axis and Y-axis represent the ground coordinates, and the Z-axis represents the height from the ground (the fixed flight height of the UAV is H). Then the three-dimensional coordinates of the u-th UAV are represented by q u = [x u , y u , H]. Since the ground height of the k-th user WD k is 0, its position is represented by q k = [x k , y k , 0]. The position of the eavesdropper is represented by q e = [x e , y e , 0]. Assuming that when the UAVs communicate with the WDs, the channels of the UAVs are line-of-sight links (LOS), and the fading effect of the channels is ignored. It can be obtained that the channel gain between the UAV and the WD k is: where u0 is the channel power gain when the reference distance is 1m, is the actual signal propagation path.
[0106] WD k The transmission rate of offloading tasks to the UAV can be obtained by the Shannon formula: where B represents the channel bandwidth, p k represents the maximum transmit power of WD k in the uplink, and σ 2 is the noise power of the channel.
[0107] Similarly, the channel gain between the UAV and the eavesdropping user can be obtained as follows:
[0108] The transmission rate from the eavesdropping user to the UAV can be obtained by the Shannon formula:
[0109] Therefore, the system secrecy rate is: C m =max{(R k -R e ),0}。
[0110] In our Mobile Edge Computing (MEC) system, the WD k uses a partial offloading strategy for the tasks in each time slot, and the offloading ratio is x k 。
[0111] Local computing mode: WD k offloads the task x k D k to the UAV for processing and executes (1 - x k )D k bits of the task locally. The total delay of the tasks completed by local computing is: where s represents the CPU cycles required to process one byte, f l,k represents the CPU computing frequency of the wireless device user WD k , D k represents the total task volume of the user,
[0112] The power consumption of local processing of the original data can be expressed as: where γ is the computing efficiency parameter.
[0113] The energy consumption of local computing is represented by the following equation:
[0114] Computing offloading mode: In the computing offloading mode, the time for WD k to offload the task to the UAV for transmission is:
[0115] When WD k offloads the task to the mobile edge computing system assisted by the UAV, the UAV is stationary relative to the ground. We assume that WD kHas a constant transmission power. Then the WD k 's transmission energy consumption is:
[0116] The UAV receives the WD k 's task and performs calculations. The time required for the UAV to execute the WD k task is expressed as: where f u represents the computing frequency of the micro mobile edge computing server.
[0117] When performing calculations on the mobile edge computing server, the power consumption can be expressed as:
[0118] The energy consumed by the UAV-assisted mobile edge computing system is:
[0119] When the UAV flies from the current hovering position to the next hovering position, the new hovering position can be expressed as: q (u+1) =[x (u) +v (u) t fly cosβ (u) ,y (u) +v (u) t fly sinβ (u) ,H], where v (u) belongs to [0,v max is the flight speed of the UAV, t fly is the fixed flight time of the UAV, cosβ (u) and respectively represent the projection components of the flight direction of the UAV on the x-axis direction and the y-axis direction.
[0120] At this time, the energy consumed by the UAV flight is: E fly(u) =φ||v (u) || 2 , where φ = 0.5M UAV t fly , M UAV is the takeoff mass of the UAV. Since the calculation results provided by the MEC server are very small, the data to be transmitted back can be ignored.
[0121] Construct the following optimization problem: f u ≥0, f l,k ≥0, wherein, is the total energy consumption of the system, is the total time delay of the system, η represents the weight factor of the relative time delay, f u ≥0, f l,k ≥0, indicating that the operation of the edge server CPU computing power and the local CPU computing power is normal, indicates that the execution delay of any WD is less than the maximum delay acceptable to the WD, indicates that the UAV is flying within a rectangular area with side length L.
[0122] The system model is solved using a deep reinforcement learning algorithm. The solution method is: generate the reachability and service capacity matrix between the generated tasks and the UAV or the central server.
[0123] Construct the state space S of the system in the deep reinforcement algorithm: the task allocation information D of the device (ass) , the number of resources required for each task D (res) , the completion status of the assigned tasks D (sta) , the reachability and service capacity D between the tasks and the UAV or the central server (abi) .
[0124] Construct the action space A of the system in the deep reinforcement algorithm: within time slot t, the UAV and the offloading environment interact, and the optimal offloading action A* is selected according to the current state and the observed environment. The action set includes the probability P of each task being assigned to the UAV or the central server (d) and the offloading ratio x of the task (k) .
[0125] Construct the reward function R. The goal of the system is to maximize the reward by minimizing the weighted sum of energy efficiency and time delay defined in the following problem while ensuring secure transmission:
[0126] Calculate the reward corresponding to each execution action, observe the next state S k+1 , and use the obtained (S (k) , A (k) , R (k) , S(k+1) )Stored in the experience storage pool B m .
[0127] In the Critic module, a small batch (S m , A (k) , R (k) , S (k) ) is randomly sampled from B (k+1) , and the target Q-value is calculated using the Bellman equation: y i = r i + γ(min(Q1(s i+1 , μ(s i+1 )), Q2(s i+1 , μ(s i+1 ))) - α log π(a i+1 |s i+1 ))), where γ is the discount factor, y i is the target Q-value, μ(s i+1 ) is the target policy used to output the action at the given state S k+1 , α is the temperature parameter used to balance the exploration and exploitation of the policy, and π(a i+1 |s i+1 ) is the probability of the action generated by the policy network.
[0128] In the SAC algorithm, two Q-networks Q1 and Q2 are used to reduce the overestimation problem of Q-values. The values of the state-action pairs estimated by each Q-network are different, and the minimum of them is taken as the target Q-value.
[0129] The current Q-values are obtained by predicting the inputs of the current state S and action A through the two Q-networks, which are Q1(s, a) and Q2(s, a) respectively.
[0130] The mean squared error loss is calculated using the difference between the above target Q-value and the current Q-value, and the Q-value network parameters are updated. The loss function is: L(θ Q ) = E[(Q1(s, a) - y i ). 2 +(Q2(s, a) - y i ). 2 .[[]END]]
[0131] The network parameters of Q1 and Q2 are updated by minimizing the loss function through the Adam optimizer, and the policy network and Q-network are softly updated: θ Q′ ← τθ Q +(1 - τ)θ Q′ , θ π′ ← τθπ +(1 - τ)θ π′ , where τ is the soft update parameter, usually taking a very small value.
[0132] Use Python to simulate the models used and evaluate the performance of the SAC algorithm. The specific steps are as follows:
[0133] DDPG is a deep reinforcement learning algorithm based on the Actor-Critic structure that can directly output deterministic actions. TD3 introduces delayed policy updates and a twin Q-network. When calculating the Q-value, noise is added to the target action, and different update frequencies are used in the updates of the policy network and the target Q-value to stabilize the training process. The SAC algorithm is an algorithm based on the maximum entropy reinforcement learning framework, with strong exploration capabilities, and can better handle complex environments. SAC also uses a twin Q-network to solve the problem of excessive Q-values, but its policy can not only maximize the expected reward but also maximize the entropy of the policy, encouraging more exploration. We use DDPG, TD3, and SAC to simulate the model respectively and obtain Figure 3 the results. The experimental results show that the performance of the SAC algorithm is better.
[0134] In Figure 4 it can be seen that compared with the average task distribution, using the ability matrix for dynamic task allocation has better performance. This is because dynamically allocating resources based on the capabilities of each agent can improve the task completion efficiency. The ability matrix adopts a more targeted method, optimizing the overall performance of the system. In contrast, the average task distribution may lead to suboptimal performance because it does not take into account the individual advantages of drones during task execution. Therefore, by using dynamic task allocation based on the ability matrix, the system can better handle the complexity of multi-objective optimization, thereby obtaining a faster convergence speed and better results.
[0135] In the SAC algorithm, the temperature coefficient can be used to control the balance between exploration and exploitation of the policy and plays a significant role in the entropy regularization term of the policy. It determines the entropy size of the policy distribution, thus affecting the randomness when choosing actions. An overly large temperature coefficient will increase the entropy of the policy, contributing to exploring more state spaces, while a smaller temperature coefficient makes the policy tend to deterministic action selection. From the simulation results Figure 5 it can be seen that when the temperature coefficient is 0.2, the system can achieve the optimal performance.
[0136] The discount factor affects the learning process of the policy by influencing the agent's attention to future rewards. In a dynamically changing environment, due to the fact that a low discount factor mainly relies on short-term feedback, it is difficult for the agent to adjust the policy when the environment changes, and this feedback may lose effectiveness when the environment changes. InFigure 6 It can be seen that when the discount factor is low, it takes a longer time to reach the optimal strategy, which shows that in a dynamically changing environment, choosing an appropriate discount factor is crucial for optimizing system performance.
[0137] In Figure 7 It can be seen that when the task sizes are different, the weighted sum of the required time delay and energy consumption is also different. The smaller the number of tasks, the smaller the value of the weighted sum of the required time delay and energy consumption, because processing larger tasks requires more time and resources.
[0138] The present invention designs a task offloading scheme to reduce the total cost of a multi-UAV-assisted mobile edge computing system. This scheme jointly optimizes factors such as the flight energy consumption, computing energy consumption, communication time delay, and user secure transmission of the UAVs. On this basis, an energy efficiency optimization model is established. To minimize the energy efficiency of the UAVs, we propose an optimization problem according to the established model. Considering the time-varying communication channels and dynamic task arrivals of the devices, an algorithm based on deep reinforcement learning is proposed to solve the problem of time-varying channels. It can be seen from the simulation results that after iterative training, the system energy efficiency and time delay are finally effectively reduced.
[0139] It should be noted that the parts not described in detail in this scheme are all prior arts. The above embodiments are only used to illustrate the present invention, but the present invention is not limited to the above embodiments. Any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention all fall within the protection scope of the present invention.
Claims
1. An energy-saving and low-latency task offloading method applied to a mobile edge computing system, characterized in that: The following steps are included: S1,establishes a multi-UAV-assisted mobile edge computing system model, including UAVs, IoT wireless device users, and eavesdropping users. Each UAV is equipped with a micro mobile edge computing server to provide task offloading services for users. S2, in the system model described in S1, calculates the weighted sum of the energy consumption and delay of the system, S3, establishes a deep reinforcement learning model to minimize the weighted sum of energy consumption and latency, and selects the optimal unloading action by the deep reinforcement learning model. S4, training the deep reinforcement learning model described in S3 until the deep reinforcement learning model achieves a preset goal.
2. The energy-saving and low-latency task offloading method applied to a mobile edge computing system according to claim 1, characterized in that: In S1, the established system model is as follows: in a rectangular flight area with a side length of L, there are U drones equipped with micro mobile edge computing servers, K IoT wireless device users, and one eavesdropping user. The drone’s flight altitude is H. The set of drones is denoted as U = {1, ..., u, ..., U}, the set of wireless devices is denoted as K = {1, ..., k, ..., K}, and the eavesdropping user is denoted as e. Construct a three-dimensional coordinate system XYZ, where the X-axis and Y-axis represent the ground coordinates of the rectangular flight area, and the Z-axis represents the height from the ground. Then the three-dimensional coordinate q of the u-th drone is u =[x u ,y u ,H], the kth wireless device user WD k The position is denoted as q k =[x k ,y k ,0], the location of the eavesdropping user e is denoted as q e =[x e ,y e ,0].
3. The energy-saving and low-latency task offloading method applied to a mobile edge computing system according to claim 2, characterized in that: In step S2, the energy consumption of the system includes the wireless device user WD k Local computing energy consumption, wireless device user WD k The transmission energy consumption of unloading tasks to drones, the computing energy consumption of the micro-mobile edge computing server on drones to complete the unloading tasks, and the energy consumption of drone flight. The system delay includes the wireless device user WD k Delay in completing local computing tasks, wireless device user WD k The delay in offloading tasks to the drone and the delay in the micro mobile edge computing server on the drone completing the offloading tasks.
4. The energy-saving and low-latency task offloading method applied to a mobile edge computing system according to claim 3, characterized in that: Wireless Device User WD k The channel gain between the UAV and Where u0 is the channel power gain when the reference distance is 1m, is the actual signal propagation path, the wireless device user WD k The transmission rate of offloading the computing task to the drone is Where B represents the channel bandwidth, p k Indicates WD k The maximum transmit power in the upload link, σ 2 is the noise power of the channel, and the channel gain between the eavesdropping user e and the drone is The transmission rate from eavesdropping user e to the drone is The confidentiality rate of the system is C m =max{(R k -R e ),0}, In the MEC system, WD k The tasks in each time slot use a partial offloading strategy, and the offloading ratio is x k , then the wireless device user WD k The delay to complete the local computation task is Among them, s represents the CPU cycle required to process a unit byte, and f l,k Indicates wireless device user WD k The CPU calculation frequency, D k Indicates the total amount of tasks for the user. Wireless Device User WD k The local power consumption is Among them, γ is the computational efficiency parameter, The wireless device user WD k The local computing energy consumption is Wireless Device User WD k The delay of unloading tasks to the drone is The wireless device user WD k The transmission energy consumption of unloading tasks to the UAV is The delay for the micro mobile edge computing server on the drone to complete the offloading task is Among them, f u represents the computing frequency of the micro mobile edge computing server, The power consumption of the micro mobile edge computing server is Then the computing energy consumption of the micro mobile edge computing server on the drone to complete the offloading task is The drone flies from the current hovering position to the next hovering position, and the new hovering position is q (u+1) =[x (u) +v (u) t fly cosβ (u) ,y (u) +v (u) t fly sinβ (u) ,H], Among them, v (u) is the flight speed of the drone, and v (u) ∈[0,v max ], t fly is the fixed flight time of the drone, cosβ (u) and sinβ (u) They represent the projection components of the UAV’s flight direction in the x-axis and y-axis directions respectively. The energy consumption of UAV flight is E fly(u) =φ||v (u) || 2 , Where, φ=0.5M UAV t fly , M UAV is the take-off mass of the drone, The total energy consumption of the system is The total delay of the system is Then the weighted sum of the system's energy consumption and delay is E k +ηT k , Where η is the weight factor of the delay.
5. The energy-saving and low-latency task offloading method applied to a mobile edge computing system according to claim 4, characterized in that: In step S3, the optimization problem is constructed by minimizing the weighted sum of energy consumption and delay in, represents the total energy consumption of the system, represents the total delay of the system, η is the weight factor of delay, f u ≥0,f l,k ≥0, indicating that the micro mobile edge computing server and wireless device user WD k The CPU is running normally, Indicates that the execution delay of any WD is less than the maximum delay that the WD can accept. It means that the drone is flying in a rectangular area with a side length of L.
6. The energy-saving and low-latency task offloading method applied to a mobile edge computing system according to claim 5, characterized in that: In step S3, a model-free Soft Actor-Critic algorithm based on a deep reinforcement learning algorithm is used. A regularization term of policy entropy is introduced into the objective function to encourage the policy network to maintain randomness during the learning process and explore the environment more comprehensively. A dual Q network structure is used to alleviate the problem of overestimation of value. The smaller value of the two Q network estimates is used to update the policy to improve the accuracy of the value function estimation. Build a capability matrix between wireless device users and target drones to help the system quickly and effectively allocate and schedule, thereby improving system efficiency.
7. The energy-saving and low-latency task offloading method applied to a mobile edge computing system according to claim 6, characterized in that: Deep reinforcement learning models include, Construct the state space S of the system in the deep reinforcement algorithm and the task allocation information D of the device (ass) , the number of resources required for each task D (res) , Completion of assigned tasks D (sta) , the reachability and service capabilities between the mission and the drone or central server (abi) , Construct the action space A of the system in the deep reinforcement algorithm. In the time slot t, the drone and the unloaded environment interact with each other, and the optimal unloading action A* is selected according to the current state and the observed environment. The action set includes the probability P of each task being assigned to the drone or the central server. (d) And the task offloading ratio x (k) , Building the Reward Function Under the premise of ensuring secure transmission, the reward is maximized by minimizing the weighted sum of energy efficiency and latency. Calculate the reward corresponding to each action and observe the next state S k+1 , the obtained (S (k) ,A (k) ,R (k) ,S (k+1) ) is stored in experience storage pool B m middle, In the Critic module, from B m Randomly draw a small batch (S (k) ,A (k) ,R (k) ,S (k+1) ), and use the Bellman equation to calculate the target Q value y i =r i +γ(min(Q1(s i+1 ,μ(s i+1 )),Q2(s i+1 ,μ(s i+1 )))-αlogπ(a i+1 |s i+1 )), Where γ is the discount factor, y i is the target Q value, μ(s i+1 ) is the target strategy used to output a given state S k+1 , α is the temperature parameter used for exploration and utilization of equilibrium strategies, π(a i+1 |s i+1 ) is the probability of the action generated by the policy network. The SoftActor-Critic algorithm uses two Q networks Q1 and Q2 to reduce the problem of overestimation of the Q value. Each Q network estimates a different value for the state-action pair, and the minimum value is taken as the target Q value. The current Q value is obtained by predicting the input of the current state space S and action space A through two Q networks, which are Q1(s,a) and Q2(s,a). The mean square error loss is calculated using the difference between the target Q value and the current Q value, and the Q value network parameters are updated. The loss function is: L(θ Q )E[(Q1(s,a)-y i ) 2 +(Q2(s,a)-y i ) 2 ], Use the Adam optimizer to minimize the loss function, update the network parameters of Q1 and Q2, and perform soft updates on the policy network and Q network. The update logic is θ Q′ ←τθ Q +(1-τ)θ Q′ , i π′ ←tth π +(1←τ)θ π′ , Among them, τ is the soft update parameter.
8. The energy-saving and low-latency task offloading method applied to a mobile edge computing system according to claim 7, characterized in that: In step S4, the interaction, storage, sampling, training, and updating are repeated until the deep reinforcement learning model reaches the desired goal.