An Edge Computing Task Offloading and Resource Allocation Method and Device
Through improved PPO algorithm and LSTM network optimization edge computing offload strategy, the problems of load balancing and insufficient scenario adaptability are solved, more efficient resource allocation and stability improvement are achieved, and equipment energy consumption and delay are reduced.
Patent Information
- Application Number
- CN202411225023.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-03
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2044-09-03
AI Technical Summary
The existing edge computing offload strategy lacks the ability to load balancing and adapt to multiple scenarios, resulting in low system stability and resource utilization, and the deep reinforcement learning algorithm has insufficient learning speed and stability.
The improved near-end strategy optimization (PPO) algorithm is adopted, combined with long and short-term memory network (LSTM) and crop function optimization, and a multi-layer deep learning network is built to optimize the objective functions of delay, energy consumption and load balancing. By introducing the load balancing degree of edge servers as the optimization goal, the stability and adaptability of the system are improved.
It significantly improves the system's resource utilization rate and adaptability to dynamic load changes, reduces equipment energy consumption and delay, and improves the system's stability and computing capabilities.
Smart Images

Figure CN119336484B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-access edge computing, and particularly relates to an edge computing offloading strategy, a resource allocation method, and a device improved based on the proximal policy optimization (PPO) algorithm. Background Art
[0002] With the rapid development of wireless communication technology and the sharp increase in the number of mobile devices, the amount of data generated at the network edge shows an explosive growth trend. Various computationally intensive and latency-sensitive tasks continue to emerge, posing unprecedented challenges to the computing power and battery capacity of mobile devices. Edge computing provides computing resources for mobile devices by deploying edge servers with strong computing capabilities near mobile devices, while significantly reducing data transmission latency and improving the quality of service (QoS). As one of the core technologies of MEC, computing offloading extends the computing power of mobile devices by transmitting computing tasks to edge servers close to users. This not only meets the requirements of applications for low latency and low energy consumption but also optimizes the user experience. During the computing offloading process, it is necessary to consider not only the location and proportion of task offloading but also environmental factors such as bandwidth, channels, and server computing resources.
[0003] Although edge computing can extend the computing power of mobile devices through computing offloading technology, it still faces problems such as slow response speed, high device energy consumption, and relatively single adaptable scenarios. Existing research mainly solves single objectives through heuristic algorithms or traditional deep reinforcement learning algorithms, pays insufficient attention to resource allocation in the computing offloading scenario, and does not pay attention to the load balancing of edge servers. Summary of the Invention
[0004] In view of the above problems, the present invention designs an improved deep reinforcement learning algorithm applicable to the scenarios of computing offloading and resource allocation, and realizes single-objective and multi-objective optimization by setting different latency and energy consumption coefficients to solve the deficiencies of existing research. The edge computing task offloading and resource allocation method of the present invention includes: an initialization step of obtaining task data of terminal devices and edge servers in an edge-terminal task scenario; the task parameters include the task offloading ratio of the terminal device; a data acquisition step of calculating task offloading parameters, the objective function, and constraint conditions of the offloading task for the edge-terminal task scenario based on the task data; a policy generation step of inputting the task offloading parameters, the objective function, and the conditional constraints into a proximal policy optimization model for training, and after the proximal policy optimization model completes policy convergence or reaches a predetermined number of training steps, outputting the task offloading policy and resource allocation policy of the edge-terminal task scenario by the proximal policy optimization model.
[0005] Furthermore, the proximal policy optimization model includes an Actor action network and a Critic evaluation network; the Actor action network is a deep learning network with a multi-layer architecture, sequentially including an input layer, a long short-term memory network layer, a fully connected layer composed of a Tanh activation layer, and an output layer.
[0006] Furthermore, the Critic evaluation network restricts the ratio r(θ) of the new and old policy probabilities during the training process through the Clip function;
[0007]
[0008] where ε is the truncation constant and v is the optimization constant, π θ (·) represents the new running policy, represents the old running policy, a t represents the running action, s t represents the environmental state.
[0009] Furthermore, the objective function and its constraints satisfy:
[0010]
[0011] where ω1 is the delay weight, ω2 is the energy consumption weight, ω3 is the load balancing weight, and the total task delay is the local execution delay of the task on the nth terminal device, is the edge delay of the task output to the edge server for execution, and the total task energy consumption is the energy consumption of the task processed locally on the nth terminal device, is the energy consumption of the task transmitted from the nth terminal device to the edge server, and LB is the load balancing degree of the edge server, L m is the CPU occupancy rate of the mth edge server, L avg represents the average value of the CPU utilization rate of the edge server, M is the number of edge servers in the edge-terminal task scenario, and the constraint condition C1 indicates that the task offloading ratio α is satisfied n The constraint condition C2 indicates that the bandwidth allocation ratio b is satisfied n , and the constraint condition C3 indicates that the computing resource allocation ratio k is satisfied n, the constraint C4 represents satisfying the limitation of the total bandwidth W on the bandwidth ratio allocated to the terminal device, the constraint C5 represents satisfying the limitation of the computing power of the edge server on the computing resource ratio allocated to the terminal device, and the constraint C6 represents satisfying the total delay of each task is subject to the maximum acceptable delay τ n of this task.
[0012] The present invention also proposes an edge computing task offloading and resource allocation device, including: an initialization module for obtaining task data of a terminal device and an edge server in an edge-terminal task scenario; the task parameters include the offloading ratio of the generated tasks of the terminal device; a data acquisition module for calculating task offloading parameters, the objective function and constraint conditions of the offloading tasks of the edge-terminal task scenario based on the task data; a policy generation module for inputting the task offloading parameters, the objective function and the conditional constraints into a proximal policy optimization model for training, and after the proximal policy optimization model completes policy convergence or reaches a predetermined number of training steps, outputting the task offloading policy and resource allocation policy of the edge-terminal task scenario by the proximal policy optimization model.
[0013] Furthermore, the proximal policy optimization model includes an Actor action network and a Critic evaluation network; the Actor action network is a deep learning network with a multi-layer architecture, sequentially including an input layer, a fully connected layer composed of a long short-term memory network layer and a Tanh activation layer, and an output layer.
[0014] Furthermore, the Critic evaluation network restricts the ratio r(θ) of the old and new policy probabilities during the training process through the Clip function:
[0015]
[0016] where ε is a truncation constant, v is an optimization constant, π θ (·) represents the new running policy, represents the old running policy, a t represents the running action, s t represents the environmental state.
[0017] Furthermore, the objective function and its constraint conditions satisfy:
[0018]
[0019]
[0020] where; ω1 is the delay weight, ω2 is the energy consumption weight, ω3 is the load balancing weight, and the total task delay is the local execution delay of the task on the nth terminal device, is the edge delay for the task to be executed by the edge server, and the total energy consumption of the task is the energy consumption for the task to be processed locally on the nth terminal device, is the energy consumption for the task to be transmitted from the nth terminal device to the edge server. LB is the load balancing degree of the edge server, L m is the CPU occupancy rate of the mth edge server, L avg represents the average value of the CPU utilization rate of the edge server, M is the number of edge servers in the edge-terminal task scenario. The constraint condition C1 represents that the task offloading ratio α is satisfied n , the constraint condition C2 represents that the bandwidth allocation ratio b is satisfied n , the constraint condition C3 represents that the computing resource allocation ratio k is satisfied n , the constraint condition C4 represents that the limit on the bandwidth ratio allocated to the terminal device by the total bandwidth W is satisfied. The constraint condition C5 represents that the computing power of the edge server on the ratio of the computing resources allocated to the terminal device is satisfied. The constraint condition C6 represents that the total delay of each task is limited by the maximum acceptable delay τ of the task n of.
[0021] The present invention also proposes a computer-readable storage medium storing computer-executable instructions, characterized in that when the computer-executable instructions are executed, the edge computing task offloading and resource allocation method as described above is implemented.
[0022] The present invention also proposes an electronic device including the edge computing task offloading and resource allocation device as described above.
[0023] The improved model based on the PPO algorithm of the present invention realizes the computing offloading and resource allocation method in the end-edge collaboration scenario, solves the problems of insufficient attention to the load balancing level of edge nodes in the existing computing offloading and the lack of attention to the time series characteristics of tasks in the actual scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a schematic diagram of the Actor network architecture introducing the LSTM network of the present invention.
[0025] Figure 2 is a schematic diagram of the IPPO algorithm framework of the present invention.
[0026] Figure 3 is a schematic diagram of the task offloading scenario system model under the end-edge collaboration architecture of the present invention.
[0027] Figure 4 It is the calculation offloading flow chart of the present invention.
[0028] Figure 5 It is the data calculation flow chart of the present invention.
[0029] Figure 6 It is the IPPO algorithm flow chart of the present invention.
[0030] Figure 7 It is the cumulative reward value trend chart of the IPPO algorithm of the present invention under different learning rate combinations.
[0031] Figure 8 It is the line chart of QoS values of each algorithm under different numbers of terminal devices.
[0032] Figure 9 It is the bar chart of the total delay of each algorithm under different numbers of terminal devices.
[0033] Figure 10 It is the bar chart of the total energy consumption of each algorithm under different numbers of terminal devices.
[0034] Figure 11 It is the schematic diagram of the edge computing task offloading and resource allocation device of the present invention.
[0035] Figure 12 It is the schematic diagram of an electronic device of the present invention.
[0036] Figure 13 It is the schematic diagram of the hardware structure of an electronic device of the present invention.
[0037] Among them, the reference numerals are:
[0038] 100: Electronic device 10: Preprocessing module
[0039] 20: Data processing module 30: Policy generation module
[0040] S1, S2, S3, S4, S2.1, S2.2, S2.3, S2.4, S2.5, S4.1, S4.2, S4.3, S4., S45, S4.6, S4.7, S4.8, S4.9, S4.10: Steps Detailed implementation manners
[0041] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific implementation methods described herein are only used to explain the present invention and are not used to limit the present invention.
[0042] The research work on computing offloading strategies in scenarios with limited existing resources pays less attention to the load balancing of edge nodes and is not comprehensive enough, which leads to uneven distribution of task loads among some edge nodes and affects the overall stability of the system. At the same time, the existing computing offloading strategies based on deep reinforcement learning are not optimized for the problems existing in the offloading scenarios, resulting in the execution speed and stability of the algorithms being difficult to meet the scenarios of a large number of mobile devices.
[0043] Aiming at the problems that the existing computing offloading pays insufficient attention to the load balancing level of edge nodes and does not attach importance to the time series characteristics of tasks in the actual scenario, the present invention proposes an improved model based on the PPO algorithm to implement a computing offloading and resource allocation method in the end-edge cooperation scenario.
[0044] The computing offloading and resource allocation method of the present invention takes the load balancing level of the edge server as the optimization goal of the offloading strategy. The existing computing offloading strategies mainly focus on latency and energy consumption and pay less attention to the load status level of the edge server. This may lead to a situation where a certain edge server receives multiple computing requests from mobile devices at the same time, while the servers with relatively less computing resources are often idle. This situation is not conducive to cooperative offloading between edge layers. The present invention analyzes the load balancing situation of the computing resources of the edge server through the CPU occupancy rate of the edge server, and improves the stability of the offloading system. By introducing the load balancing level of the edge server into the optimization goal, an objective function including latency, energy consumption, and load balancing degree is built, which improves the scalability and fault tolerance of the offloading system. Compared with a single optimization goal, it significantly improves the resource utilization rate of the system and the adaptability to dynamic load changes. It solves the problem that the existing computing offloading algorithms pay insufficient attention to the load of hardware devices.
[0045] The LSTM network is also adopted in the policy network to enhance the adaptation to tasks with time series characteristics and dynamic change states. In a multi-edge server system, especially in scenarios involving a large number of terminal devices, the environmental state is constantly changing dynamically. This dynamicity brings great difficulties to accurately obtaining this information, especially when considering that these data usually exhibit certain time series characteristics. LSTM can deeply mine the historical information in time series data and more accurately estimate the current state by capturing the dependencies in the sequence. In this way, by introducing LSTM, not only can the advantages of LSTM in processing time series data be fully utilized to improve the performance of the PPO algorithm in a dynamic environment, but it can also effectively handle partially observable data, solve the problem caused by incomplete data, and improve the stability and performance of the strategy.
[0046] In addition, the computing offloading and resource allocation method of the present invention expands the probability ratio of the algorithm clipping function (Clip) under the same hyperparameters. The original PPO algorithm has problems of slow learning speed and instability during the update process, and the limitation of the Clip function is one of the main reasons. In this paper, the Clip function of the algorithm is improved. By adjusting the update method of the probability ratio of the old and new policies under the same truncation constant ε, the range is expanded, the learning of the truncation critical range is enhanced, and the probability ratio of the old and new policies is limited to the interval, making it more conducive to exploring the policy probability ratio space compared to the original algorithm. Through the above improvements, not only the learning rate of the PPO algorithm is improved, but also its learning stability is enhanced. And on the premise of ensuring the correctness of the policy, faster convergence and better performance are achieved. The improved algorithm has a larger learning range under the same hyperparameters and speeds up the learning rate of the algorithm. It solves the problems of slow update and instability of the original algorithm.
[0047] In the first embodiment of the present invention, an edge computing task offloading and resource allocation architecture improved based on the PPO algorithm is provided, including:
[0048] 1. Construct an edge-end task offloading scenario including multiple terminal devices and multiple edge servers, and obtain the data and offloading models of the tasks generated by the terminal devices;
[0049] (1) Construct a system model of multiple edge servers:
[0050] The system includes M edge servers, represented by the set M = {1, 2, 3,..., M}; N terminal devices are represented by the set N = {1, 2, 3..., N}; the computing resources of the m-th edge server are represented by and the computing resources of the n-th terminal device are represented by .
[0051] (2) Data and offloading models of tasks generated by terminal devices:
[0052] Each terminal device generates a computing task at the beginning of the time slot represented by a triple {D n , C n , τ n}. Among them, D n represents the data size of the task on the mobile terminal device n; C n is the number of CPU cycles required for this task, and τ n is the maximum latency requirement of the task.
[0053] The offloading ratio of the task generated by the terminal device n is represented by α n = [0, 1], n ∈ N, and the remaining proportion of the task is executed locally on the device.
[0054] 2. Calculate data such as time delay, energy consumption, and system load balancing degree;
[0055] (1) Construct the time model of the task
[0056] According to Shannon's theorem, the transmission speeds of the mobile terminal device and the edge server can be expressed as:
[0057]
[0058] where b n represents the proportion of the bandwidth allocated to the terminal device in the total bandwidth, b n ∈[0,1], and satisfies W represents the total bandwidth of the uplink; P n is the power when the mobile terminal transmits the task; g n is the channel gain between the edge and the terminal device; N0 is the noise power spectral density.
[0059] The time for the terminal device to transmit the task data to the edge server can be expressed as:
[0060]
[0061] where α n is the proportion offloaded to the edge server.
[0062] The total time delays for local execution by the terminal device and execution by the edge node are respectively:
[0063]
[0064] where 1 - α n represents the proportion of local execution, represents the CPU frequency of the mobile terminal device n, C n is the total number of CPU cycles of this task, k n is the server CPU frequency allocated to this task, is the total computing frequency of the server.
[0065] The total time delay of the task is the maximum of the local execution time delay and the edge time delay, and can be expressed as:
[0066]
[0067] where, represents the partial execution time of the task, represents the time for the task to be processed at the edge (2) Construct the task energy consumption model
[0068] The energy consumption for local processing of tasks on the mobile device and the energy consumption for transmission to the edge node are respectively:
[0069]
[0070] Among them, μ is a constant related to the CPU structure of the edge device, representing the effective switching capacitance, and P n represents the power when the mobile device n transmits data.
[0071] The total energy consumption of the task is the sum of the energy consumption generated during the transmission process and the local execution energy consumption, and can be expressed as
[0072]
[0073] Among them, is the computing energy consumption of the mobile device, is the energy consumption for the task to be transmitted to the edge node.
[0074] (3) Construct a system load balancing model
[0075] Because there are multiple edge servers in the system, in order to effectively describe the load balancing situation of all devices in the edge layer, this model uses the standard deviation of the CPU occupancy rates of all edge servers to represent the system load balancing degree. The calculation method is as follows:
[0076]
[0077] Among them, the CPU occupancy rate of the m-th edge server can be expressed as L m , and L avg represents the average value of the CPU utilization rate of the edge server, and LB represents the system load balancing level.
[0078] 3. Describe the partial task offloading model in the scenario of multiple edge servers in each time slot, and set the objective function and conditional constraints of the system;
[0079] The system optimization goal is to minimize the weighted sum of the total delay, total energy consumption, and the edge layer load balancing value. The mathematical expression of this problem is as follows:
[0080]
[0081]
[0082] Among them, ω1, ω2, and ω3 are the weights of delay, energy consumption, and load balancing respectively. Considering that the load balancing value is relatively small, in order to balance the impact of load balancing on the total system cost, it is proven through experiments that setting its parameter ω3 between 1 and 2 can more effectively reflect the role of the load balancing value in the total system cost. Therefore, the load balancing weight parameter ω3 is set to (1, 2), and the sum of the weights of delay and energy consumption is still 1. The constraint conditions C1, C2, and C3 are the offloading ratio, bandwidth allocation ratio, and computing resource allocation ratio of the terminal device respectively, and they should all be within the range of 0 to 1. At the same time, C4 means that the bandwidth ratio allocated by the system to the terminal device should not exceed the limit of its total bandwidth W to ensure the reasonable utilization of bandwidth resources. C5 means that the computing resource ratio allocated by the server to the terminal device cannot exceed its own computing power to avoid overloading. C6 means that the total delay of each task cannot be greater than the maximum acceptable delay of the task. These constraint conditions jointly ensure the stability and efficiency of the system operation.
[0083] 4. Propose an improved algorithm based on the Proximal Policy Optimization algorithm, namely the IPPO algorithm, to minimize the weighted sum of delay, energy consumption, and load balancing level, and save the model parameters of the algorithm; including:
[0084] (1) Markov decision process
[0085] The Markov decision process is represented by a triple model {S, A, R} composed of three core elements: the state space (S), the action space (A), and the reward function (R).
[0086] The state space can be expressed as s t ={D n (t), C n (t), τ n (t), r n (t), L n (t)}, where D n (t) represents the data volume of the task generated by device n; C n (t) represents the computing volume of the task generated by device n; τ n (t) represents the maximum tolerance delay of the task; r n (t) represents the transmission speed; L n (t) is the CPU utilization rate of the edge server.
[0087] The action space includes: the offloading decision of the mobile terminal device, and the computing and bandwidth resources allocated by the edge server to the terminal device. It can be expressed as A = {α1, α2,..., α N ; G1, G2,..., G N ; b1, b2,..., b N ; k1, k2,..., kN}. αx represents the offloading ratio of the nth terminal device; G n represents the edge server object to which the task is to be offloaded; b n is the bandwidth ratio allocated by the edge server to the terminal device; k n represents the computing resource ratio obtained by the terminal device.
[0088] The reward function is to minimize the weighted sum of the system delay, energy consumption, and the load balancing level of the edge server.
[0089] It can be expressed as
[0090] (2) Improved Proximal Policy Optimization IPPO algorithm
[0091] In the conventional PPO algorithm, in each iteration, according to the samples collected by the current policy, the probability ratio is used to quantify the difference between the new policy and the old policy, and the policy parameters are updated accordingly. By finely controlling the update amplitude, the PPO algorithm ensures the stability of the training process, avoids drastic fluctuations in policy updates, and thus improves the convergence speed and performance of the algorithm. r(θ) = π θ (a t |s t ) / π θ' (a t |s t ) represents the ratio of the probabilities of the new and old policies taking action a t , and this ratio reflects the degree of change in action selection of the new policy relative to the old policy.
[0092] The core of such algorithms is to generate actions for the current state through the Actor action network, save a batch of samples to the experience replay pool, and then calculate the Q value of the current action through the Critic evaluation network. The two modules continuously update and optimize the network parameters using backpropagation of gradients.
[0093] Improvement 1: Introduce a Long Short-Term Memory network (LSTM) into the Actor network to make full use of its ability to process time series data. LSTM can deeply mine the historical information in time series data and more accurately estimate the current state by capturing the dependencies in the sequence. The state vector at time slot t is used as the input and passed to the LSTM network, and then the state vector processed by the LSTM network is further input into the network of the PPO algorithm. The Actor action network is a deep learning network composed of a 5-layer architecture, including an input layer, an LSTM layer, a fully connected layer composed of 2 Tanh activation layers, and a final output layer. Its network structure is as Figure 1 shown:
[0094] Improvement 2: The original PPO algorithm has problems of slow learning speed and instability during the update process, and the limitation of the Clip function is one of the main reasons. The Clip function of the algorithm is improved by adjusting the update method of the ratio of the new and old policy probabilities and expanding its range under the same truncation constant ε setting, thereby accelerating the learning rate of the algorithm. Specifically, the ratio of the new and old policy probabilities is restricted to the interval [1 / (1 + ε), 1 / (1 - ε)]. Secondly, in order to ensure that the policy can be optimized in the correct direction and enhance the learning of the truncation critical range, a constant v is introduced. When the ratio of the new and old policy probabilities exceeds the predetermined clipping range, it is re-restricted to a suitable interval. The calculation process of the improved clip function is as follows:
[0095]
[0096] where v ∈ (0, 1) and ε ∈ (0, 1).
[0097] The improved IPPO algorithm framework is as Figure 2 shown, including:
[0098] 4.1. Initialize the policy network π in the algorithm θ parameters θ, the old policy network parameters θ old and the value network parameters φ, and set hyperparameters including the learning rate, discount factor, and truncation coefficient, etc.;
[0099] 4.2. Input the environmental state space into the algorithm and learn through the long short-term memory network (LSTM) in the policy network to capture the system resource state with time series relationship and dynamic changes, and output the action probability π θ (a t |s t );
[0100] 4.3. Run the policy π in the environment θ to generate a batch of trajectories. Use the generated actions and environmental parameters to train the algorithm to obtain the reward value r t at this moment and the environmental state s t+1 at the next moment, and store (s t , a t , r t , s t+1 , π θ (a t |s t )) in the experience pool;
[0101] 4.4. Use the generalized advantage estimation advantage function
[0102] 4.5. Calculate the probability ratio of the new and old policies in the experience pool: Determine the final probability ratio through the improved clipping function, and use the hyperparameter v to control the probability ratio within
[0103] 4.6. Calculate the objective function value J(θ)
[0104]
[0105] 4.7. Update the policy network parameter θ through the gradient descent method to maximize the objective function J(θ);
[0106] 4.8. Calculate the loss value of the value function: loss = (V(φ) - v trace ) 2 , and update the value network parameter φ through the gradient descent method to minimize the loss value of the value function;
[0107] 4.9. Repeat the process of 4.2 - 4.8 until the policy converges or reaches the predetermined number of training steps;
[0108] 4.10. Output the final offloading decision and resource allocation plan.
[0109] In the second embodiment of the present invention, as Figure 3 shown, it shows an application scenario of the present invention, mainly applied to the task offloading scenario under the edge - cloud collaborative architecture. Due to the weak computing performance of mobile devices and the limitation of battery capacity, they are unable to process computationally intensive tasks. Edge nodes close to mobile devices have stronger computing resources and shorter data transmission times than cloud nodes. Therefore, through the computing offloading technology, the computing power of mobile devices can be extended, the task processing delay and mobile device energy consumption can be reduced, and thus the user service quality can be improved. The system includes N mobile devices, M edge servers, and a scheduler located at the base station. Each terminal device generates a task represented by a triple {D n , C n , τn }} at the beginning of each time instant, where D n (t) represents the data volume of the task generated by the mobile device at time t, C n (t) represents the number of CPU cycles required to process this task; τ n (t) represents the maximum tolerance delay of this task. The offloading ratio of the task generated by terminal device n is represented by α n = [0, 1], and the remaining proportion of the task is executed locally on the device. These two parts can be executed in parallel, thereby reducing the total delay and improving the service response time.
[0110] Based on the above system model architecture, in the third embodiment of the present invention, a method for edge computing task offloading and resource allocation is proposed, as Figure 4As shown in the figure, it is a schematic diagram of the overall process of the computing offloading strategy of the present invention. Specifically, it includes:
[0111] Step S1: Construct an edge-terminal task offloading scenario including multiple terminal devices and multiple edge servers, and obtain the data and offloading model of the tasks generated by the terminal devices;
[0112] In the system, the set M = {1, 2, 3,..., M} is used to represent M edge servers, and the set N = {1, 2, 3,..., N} is used to represent N terminal devices. Each terminal device generates a computing task at the beginning of a time slot. It is represented by a triple {D n , C n , τ n}. Among them, D n represents the data size of the task on the mobile terminal device n; C n is the number of CPU cycles required for this task, and τ n is the maximum latency requirement of the task. In this model, each task can be split into two parts and executed locally and at the edge simultaneously.
[0113] The offloading ratio of the task generated by the terminal device n is represented by α n = [0, 1], n ∈ N, and the remaining proportion of the task is executed locally on the device. These two parts can be executed in parallel, thereby reducing the total latency and improving the service response time.
[0114] Step S2: Calculate data such as time delay, energy consumption, and system load balance; as Figure 5 shown, it includes the following specific steps:
[0115] Step S2.1: Calculate the time for task transmission Divide the task data volume by the transmission speed, and the calculation formula is:
[0116]
[0117] Among them, α n is the ratio offloaded to the edge server, D n represents the total data volume of this task, and r n is the transmission speed of the data;
[0118] Step S2.2: Split the task according to the offloading ratio, and calculate the processing time of the task on the mobile device locally and at the edge node, the local latency and the edge latency The calculation formulas are:
[0119]
[0120] Among them, 1 - α nIndicates the proportion of local execution. Indicates the CPU frequency of the mobile terminal device n, C n Is the total number of CPU cycles of this task, k n Is the server CPU frequency allocated to this task. Is the total computing frequency of the server.
[0121] Step S2.3: Calculate the energy consumption of processing part of the task locally on the mobile device And the energy consumption during the task transmission process The specific calculation formulas are as follows:
[0122]
[0123] Among them, μ is a constant related to the CPU structure of the edge device representing the effective switching capacitance, with a value of 10 -26 , P n Represents the power when the mobile device n transmits data.
[0124] Step S2.4: Obtain the CPU occupancy rate of each edge server, and calculate its standard deviation as the system load balancing degree LB:
[0125]
[0126] Among them, M is the number of edge servers, L m Represents the CPU resource occupancy rate of the m-th edge node, L avg Represents the average value of all CPU utilization rates.
[0127] Step S2.5: Calculate the total delay and total energy consumption during the task offloading process. The total delay Is the maximum value of the local and edge delays. The total energy consumption Is the sum of the local execution and task transmission energy consumptions. The calculation formulas are as follows:
[0128]
[0129] Step S3: Describe the partial task offloading model in the scenario of multiple edge servers in each time slot, and set the objective function and conditional constraints of the system;
[0130] The objective function and its constraints are as follows:
[0131]
[0132] Among them, the constraint conditions C1, C2, and C3 are the offloading ratio, bandwidth allocation ratio, and computing resource allocation ratio of the terminal device respectively, and they should all be within the range of 0 to 1. At the same time, C4 means that the bandwidth ratio allocated by the system to the terminal device should not exceed the limit of its total bandwidth W to ensure the reasonable utilization of bandwidth resources. C5 means that the computing resource ratio allocated by the server to the terminal device cannot exceed its own computing power to avoid overloading. C6 means that the total delay of each task cannot be greater than the maximum acceptable delay of the task. These constraint conditions jointly ensure the stability and efficiency of the system operation.
[0133] Step S4: Input the data calculated in the above process into the IPPO algorithm, and obtain the algorithm parameter model and decision-making actions through iterative training. As Figure 6 shown, the specific steps of the algorithm model are as follows:
[0134] Step S4.1: Initialize the policy network πθ of the PPO algorithm 和 and the value network parameter φ, set hyperparameters such as the learning rate, discount factor, and truncation constant; the parameter settings of the IPPO algorithm are shown in Table 1:
[0135] Parameter Value Network type LSTM Number of neurons in the hidden layer of the Actor network 256,128 Number of neurons in the hidden layer of the Critic network 256,256,128 Activation function of the Actor network LeakyRelu, Tanh Activation function of the Critic network LeakyRelu Update frequency of the Actor network 10 <![CDATA[Actor network learning rate lr a > <![CDATA[10 -4 > <![CDATA[Critic network learning rate lr c > <![CDATA[10 -3 > Discount rate γ 0.6 Constant ε of the clip function 0.2 Constant ν of the clip function 0.05 Batch size 128
[0136] Table 1
[0137] Step S4.2: Input the state in the system environment into the policy network of the algorithm as the first state to start iterative training. After learning through the policy network containing LSTM, output the action probability π θ (a t |s t );
[0138] Step S4.3: Run the policy π in the environment θ to generate a batch of trajectories. Obtain the action a t at this moment through the action probability, the reward value r t and the environmental state s t+1 at the next moment, and store s t , a t , r t , s t+1 , π θ (a t |s t )) into the experience pool;
[0139] Step S4.4: Calculate the generalized advantage function to evaluate the quality of the action;
[0140] Step S4.5: Calculate the probability ratio of the new and old policies in the experience pool: Determine the final probability ratio through the improved clipping function, and use the hyperparameter v to control the probability ratio within
[0141] Step S4.6, calculate the objective function value:
[0142]
[0143] Step S4.7, update the policy network parameter θ through the gradient descent method to maximize the objective parameter J(θ);
[0144] Step S4.8, calculate the loss value of the value function: loss = (V(φ) - v trace ) 2 , and update the value network parameter φ through the gradient descent method to minimize the loss value of the value function;
[0145] Step S4.9, repeatedly execute S4.2 - S4.8 until the policy converges or reaches the predetermined number of training steps;
[0146] Step S4.10, output the final offloading decision and resource allocation plan.
[0147] The following details the technical effects obtained by the computing offloading and resource allocation method of the present invention through two experiments.
[0148] Experiment 1: IPPO algorithm learning rate combination experiment
[0149] In this experiment, four different learning rate combinations are selected to analyze the convergence and stability of the IPPO algorithm, and the number of training rounds is set to 1000. Figure 7 Shows the change of the reward value of the IPPO algorithm under different learning rate combinations. It can be observed that when the learning rate lr a of the Actor network and the learning rate lr c of the Critic network are both 10 -4 , the algorithm only converges stably at the 800th round of training, and the convergence speed is slow. When the learning rates are both 10 -3 , the algorithm cannot converge, and the change range of the reward value is too large to be shown in the figure. When the learning rate is set to lr a = 10 -3 , lr c = 10 -4 , the algorithm also cannot converge because the Actor network updates the parameters according to the unstable Q value, resulting in the algorithm being unable to learn effective information, thereby affecting the convergence of the algorithm. And when lr a = 10 -4 , lr c = 10 -3When the algorithm converges at the 600th round of training, it indicates that when lr c is greater than lr a , the algorithm can achieve better convergence. This is because the Critic network can learn the Q value faster and reach stability at this time, which is convenient for the Actor network to update parameters quickly in the correct direction. However, when lr a = 10 -3 , lr c = 10 -2 , the algorithm can never converge and is in a fluctuating state. Therefore, to ensure the convergence and performance of the IPPO algorithm, lr a and lr c are set to 10 -4 and 10 -3 respectively in the subsequent experiments.
[0150] Experiment 2: Experiment with Different Numbers of Terminal Devices
[0151] To comprehensively evaluate the effectiveness of the IPPO algorithm proposed in this chapter, the following algorithms are selected for comparison:
[0152] ① All Local (AL): All tasks generated by terminal devices are processed locally.
[0153] ② Random Offloading (RO): The tasks of each mobile terminal device are evenly distributed to all edge servers according to a random ratio, and the computing resources are also evenly distributed to the terminal devices.
[0154] ③ PPO algorithm offloading (PPO): The original PPO algorithm without improvement is used to evaluate the effectiveness of the optimization strategy proposed in this chapter.
[0155] The evaluation metrics mainly include the total system delay, the total system energy consumption, and the system quality of service value.
[0156] The total system delay is the total time consumed to process all terminal device tasks. The calculation formula is as follows:
[0157]
[0158] The total system energy consumption represents the total energy consumed during the local execution of all terminal devices and the data transmission to edge servers during the offloading process. The calculation formula is as follows:
[0159]
[0160] QoS is used to represent the optimization effect of each algorithm compared to the case where tasks are fully executed locally.
[0161]
[0162] where ω1 and ω2 are the delay weight and energy consumption weight respectively, and represent the delay and energy consumption when the task of device n is fully executed locally respectively.
[0163] When the number of fixed edge servers is 4 and the computing power is 4 Ghz, the number of terminal devices is set to {10, 15, 20, 25, 30, 35, 40} respectively. Analyze the changes in the total task processing delay, total device energy consumption, and QoS value of each algorithm under different numbers of terminal devices.
[0164] Figure 8 shows the QoS values of each algorithm under different numbers of devices. From Figure 8 it can be seen that the QoS value of the AL algorithm, as the baseline algorithm, always remains 0. As the number of terminal devices increases, the QoS value of the RO algorithm adopting the random offloading strategy shows a certain degree of fluctuation, which may be due to the fact that the random offloading strategy cannot effectively manage and allocate computing resources in the face of a large number of devices. In contrast, other algorithms show more stable performance in terms of QoS. Especially for the IPPO algorithm, its QoS value always remains at the highest level, with a maximum optimization of 6.62% compared to PPO. The reason for the small optimization amplitude is that the computing resources are still relatively sufficient under the current number of devices. Therefore, the QoS value has not increased significantly.
[0165] Figure 9 shows the changes in the total system delay under different numbers of terminal devices. From Figure 9 it can be found that the delay growth rates of the AL and RO algorithms are significantly higher than those of the other two algorithms, showing a linear growth trend. This indicates that with the continuous increase in the number of terminal devices, reasonable offloading decisions can significantly optimize the delay. After the number of terminal devices reaches 20, the optimization effect of the IPPO algorithm proposed in this paper is significantly enhanced compared to the PPO algorithm, and an average delay optimization of 9.23% can be achieved. This shows that as the number of terminal devices increases, the IPPO algorithm can better optimize the task processing delay and improve the system performance by making more reasonable task offloading decisions.
[0166] Figure 10 presents the impact of different numbers of terminal devices on the total system energy consumption. It can be observed that the IPPO algorithm always maintains the lowest energy consumption level and achieves a significant improvement in energy consumption optimization compared to the PPO algorithm, with an average optimization of 16.6%. This further proves that the IPPO algorithm can effectively reduce the device energy consumption during the task offloading process, thereby improving the utilization efficiency of the terminal device battery capacity.
[0167] It should be noted that in various embodiments of the present invention, the magnitudes of the serial numbers of the above steps do not mean the sequence of execution. The execution sequence of each step should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0168] The following is a device embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.
[0169] In the fourth embodiment of the present invention, an edge computing task offloading and resource allocation device is proposed, as Figure 11 shown, including:
[0170] An initialization module 10, configured to obtain task data of a terminal device and an edge server in an edge-terminal task scenario; the task data includes the offloading ratio α of the generated tasks of the nth terminal device n and the load balancing degree LB of the mth edge server;
[0171] A data acquisition module 12, configured to calculate task offloading parameters, an objective function, and constraint conditions of an offloading task for an edge-terminal task scenario based on the task data;
[0172] A policy generation module 13, configured to input the task offloading parameters, the objective function, and the conditional constraints into a proximal policy optimization model IPPO for training. When the IPPO completes policy convergence or reaches a predetermined number of training steps, the IPPO outputs a task offloading policy and a resource allocation policy for the edge-terminal task scenario.
[0173] The IPPO model includes an Actor action network and a Critic evaluation network; the Actor action network is a deep learning network with a multi-layer architecture, sequentially including an input layer, a long short-term memory network layer, a fully connected layer composed of 2 Tanh activation layers, and an output layer; the Critic evaluation network restricts the ratio r(θ) of the new and old policy probabilities during the training process through a Clip function:
[0174]
[0175] where ε is a truncation constant and v is an optimization constant, π θ (·) represents a new running policy, represents an old running policy, a t represents a running action, s t represents an environmental state.
[0176] In the fifth embodiment of the present invention, a computer-readable storage medium is proposed. For the edge computing task offloading and resource allocation device of the present invention, when its functions are implemented in the form of software functional units and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a computer-readable storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. Therefore, in the fifth embodiment of the present invention, a computer-readable storage medium is provided for storing a computer program for executing the edge computing task offloading and resource allocation method. It should be understood that the computer-readable storage medium in the embodiments of the present invention can be a volatile memory and / or a non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM).
[0177] Figure 12 is a schematic diagram of an electronic device of the present invention. As Figure 12As shown in the sixth embodiment of the present invention, an electronic device 100 is proposed, which includes the edge computing task offloading and resource allocation device as described above. Those of ordinary skill in the art can understand that all or part of the steps in the above method can be completed by a program to instruct related hardware (such as a processor, FPGA, ASIC, etc.). All or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, each module in the above embodiments can be implemented in the form of hardware, for example, by an integrated circuit to implement its corresponding function, or can be implemented in the form of a software functional module, for example, by a processor executing a program / instruction stored in a memory to implement its corresponding function. The embodiments of the present invention are not limited to any specific form of combination of hardware and software.
[0178] It should be noted that the structure of the electronic device shown in the drawings of the present invention does not constitute a limitation thereto. The actual knowledge structure recognition device may include more or fewer components than those shown in the drawings, or combine certain components, or have different component arrangements.
[0179] The electronic device of the present invention can be any device with data processing capabilities, and such a device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or by hardware, or by a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by a processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory and running them. Figure 13 It is a schematic diagram of the hardware structure of an electronic device of the present invention. As Figure 13 shown, from the hardware level, it is a hardware structure diagram of any device with data processing capabilities where the edge computing task offloading and resource allocation device of the present invention is located. In addition to Figure 13 the processor, memory, network interface, and non-volatile memory shown, the device of the embodiment is usually located in any device with data processing capabilities. According to the actual functions of the device with data processing capabilities, other hardware may also be included, which will not be elaborated here.
[0180] When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium may be a solid-state drive.
[0181] The research work on computing offloading strategies in existing resource-constrained scenarios pays less and less comprehensive attention to the load balancing of edge nodes, which results in uneven task load distribution among some edge nodes and affects the overall stability of the system. At the same time, the existing computing offloading strategies for deep reinforcement learning are not optimized for the problems existing in the offloading scenario, resulting in the execution speed and stability of the algorithm being difficult to meet the scenarios of a large number of mobile devices. In view of the above two problems, the present invention starts from the perspective of optimizing the offloading algorithm, proposes to use the load balancing degree of the edge server as the optimization target, and introduces a long short-term memory network (LSTM) model to learn the changes in the historical environment according to the characteristics of the offloading scenario, so as to achieve a precise response to the state changes, and optimizes the clipping function of the policy update to speed up the learning speed of the algorithm and improve its stability. On the one hand, the present invention constructs an objective function with delay, energy consumption, and load balancing as optimization targets, and realizes the degree of emphasis on a certain target by setting different weight coefficients, so that the offloading strategy has a more comprehensive adaptation scenario. On the other hand, by optimizing the PPO algorithm to adapt to the scenario of multiple edge servers and multiple mobile devices, the system can generate offloading decisions faster and more stably, and minimize the delay and instability brought by the running process of the algorithm to the offloading system.
[0182] The above embodiments are only used to illustrate the present invention and are not intended to limit the present invention. Those of ordinary skill in the relevant technical fields can make various changes and deformations without departing from the spirit and scope of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the present invention. The patent protection scope of the present invention shall be defined by the claims.
Claims
1. An edge computing task offloading and resource allocation method, characterized in that Including: An initialization step of obtaining task data of the terminal device and the edge server in the edge-terminal task scenario; The task parameters include the task offloading ratio of the terminal device; A data acquisition step of calculating, based on the task data, the task offloading parameters of the edge-terminal task scenario, and the objective function and constraints of the offloading task, where the objective function and its constraints satisfy: Wherein; ω1 is the delay weight, ω2 is the energy consumption weight, ω1 + ω2 = 1, ω3 is the load balancing weight, ω3 ∈ (1, 2), and the total task delay is the local execution delay of the task on the nth terminal device, is the edge delay for the task to be executed by the edge server, and the total task energy consumption is the energy consumption of the task processed locally on the nth terminal device, is the energy consumption for the task to be transmitted from the nth terminal device to the edge server, and LB is the load balancing degree of the edge server, L m is the CPU occupancy rate of the mth edge server, L avg represents the average value of the CPU utilization rate of the edge server, M is the number of edge servers in the edge-terminal task scenario, and the constraint condition C1 means that the task offloading ratio α is satisfied n The constraint condition C2 means that the bandwidth allocation ratio b is satisfied n , the constraint condition C3 means that the computing resource allocation ratio k is satisfied n , the constraint condition C4 means that the total bandwidth W restricts the bandwidth ratio allocated to the terminal device, the constraint condition C5 means that the computing power of the edge server restricts the computing resource ratio allocated to the terminal device, and the constraint condition C6 means that the total delay of each task is restricted by the maximum acceptable delay τ of the task n of; A policy generation step of inputting the task offloading parameters, the objective function, and the conditional constraints into a proximal policy optimization model for training. When the proximal policy optimization model completes policy convergence or reaches a predetermined number of training steps, the proximal policy optimization model outputs the task offloading policy and resource allocation policy of the edge-terminal task scenario.
2. The edge computing task offloading and resource allocation method according to claim 1, wherein The proximal policy optimization model includes an Actor action network and a Critic evaluation network; The Actor action network is a deep learning network with a multi-layer architecture, sequentially including an input layer, a fully connected layer composed of a long short-term memory network layer and a Tanh activation layer, and an output layer.
3. The edge computing task offloading and resource allocation method according to claim 2, characterized in that, The Critic evaluation network restricts the ratio r(θ) of the new and old policy probabilities during the training process through a Clip function where ε is a truncation constant and v is an optimization constant, π θ (·) represents the new running policy, represents the old running policy, a t represents the running action, s t represents the environmental state.
4. An edge computing task offloading and resource allocation device, characterized in that, Including: An initialization module for obtaining task data of the terminal device and the edge server in the edge-terminal task scenario; The task parameters include the offloading ratio of the generated tasks of the terminal device; A data acquisition module for calculating, based on the task data, the task offloading parameters of the edge-terminal task scenario, and the objective function and constraints of the offloading task, where the objective function and its constraints satisfy: Wherein; ω1 is the delay weight, ω2 is the energy consumption weight, ω1 + ω2 = 1, ω3 is the load balancing weight, ω3 ∈ (1, 2), and the total task delay is the local execution delay of the task on the nth terminal device, is the edge delay for the task to be executed by the edge server, and the total task energy consumption is the energy consumption for the task to be locally processed on the nth terminal device, is the energy consumption for the task to be transmitted from the nth terminal device to the edge server. LB is the load balancing degree of the edge server, L m is the CPU occupancy rate of the mth edge server, L avg represents the average value of the CPU utilization rate of the edge server, M is the number of edge servers in the edge-terminal task scenario. The constraint condition C1 means that the task offloading ratio α is satisfied n The constraint condition C2 means that the bandwidth allocation ratio b is satisfied n , the constraint condition C3 means that the computing resource allocation ratio k is satisfied n , the constraint condition C4 means that the total bandwidth W restricts the bandwidth ratio allocated to the terminal devices, the constraint condition C5 means that the computing power of the edge server restricts the computing resource ratio allocated to the terminal devices, and the constraint condition C6 means that the total delay of each task is restricted by the maximum acceptable delay τ of the task n ; A policy generation module for inputting the task offloading parameters, the objective function, and the conditional constraints into a proximal policy optimization model for training. When the proximal policy optimization model completes policy convergence or reaches a predetermined number of training steps, the proximal policy optimization model outputs the task offloading policy and resource allocation policy of the edge-terminal task scenario.
5. The edge computing task offloading and resource allocation device according to claim 4, wherein, The proximal policy optimization model includes an Actor action network and a Critic evaluation network; The Actor action network is a deep learning network with a multi-layer architecture, sequentially including an input layer, a fully connected layer composed of a long short-term memory network layer and a Tanh activation layer, and an output layer.
6. The edge computing task offloading and resource allocation device according to claim 5, characterized in that, The Critic evaluation network restricts the ratio r(θ) of the new and old policy probabilities during the training process through a Clip function where ε is a truncation constant and v is an optimization constant, π θ (·) represents the new running policy, represents the old running policy, a t represents the running action, s t represents the environmental state.
7. A computer-readable storage medium storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed, the edge computing task offloading and resource allocation method described in any one of claims 1 to 3 is implemented.
8. An electronic device, including the edge computing task offloading and resource allocation device described in any one of claims 4 to 6.
Citation Information
Patent Citations
Internet of Things edge task unloading method and device
CN113225377A
Resource scheduling method for optimizing edge energy consumption and load based on reinforcement learning
CN117194057A