Edge computing task unloading method based on deep reinforcement learning

By introducing a pre-sorted deep reinforcement learning model in edge computing task offloading, the task execution sequence is optimized, and the problems of high latency and energy consumption in the existing technology are solved, and faster convergence speed and lower computing costs are achieved.

CN120492047APending Publication Date: 2025-08-15GUANGDONG UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510425588.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-31
Filing Date
2025-04-07
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

When the existing edge computing task unloading method handles complex tasks with predecessor or successor tasks, it cannot further reduce latency and energy consumption through scheduling optimization, and the convergence speed is slow.

Method used

A pre-sorting model for computing unloading subtasks is introduced, a pre-sorted deep reinforcement learning task unloading model is built, and the execution order of edge computing tasks is optimized. Through deep reinforcement learning, the unloading decisions are optimized to reduce the delay and energy consumption of computing tasks.

Benefits of technology

The convergence speed of the deep reinforcement learning task offload model is improved, the decision time is reduced, and the delay and energy consumption of the calculation task are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492047A_ABST
    Figure CN120492047A_ABST
Patent Text Reader

Abstract

The invention provides an edge computing task unloading method based on deep reinforcement learning, and relates to the technical field of edge computing, and the method comprises the steps: firstly calculating the task unloading cost of user equipment under different unloading decisions, and constructing a cost optimization objective function; a calculation unloading sub-task pre-sorting model is introduced to construct a pre-sorting deep reinforcement learning task unloading model, the execution sequence of edge calculation tasks is optimized, the convergence speed of the deep reinforcement learning task unloading model is improved, and the decision time is shortened; a pre-sequencing deep reinforcement learning task unloading model is trained by taking minimization of the total cost of edge network terminal users as a target, and the trained pre-sequencing deep reinforcement learning task unloading model is utilized to output an unloading decision to perform task unloading, so that the time delay and energy consumption of a calculation task are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of edge computing, and more specifically, to an edge computing task offloading method based on deep reinforcement learning. Background Art

[0002] As the number of terminals accessing the edge continues to increase, traditional cloud computing architecture faces network transmission bottlenecks. Edge computing has emerged to alleviate computing and storage pressures in the cloud and improve the response speed of communication systems. Mobile Edge Computing (MEC), as an emerging computing model, reduces the burden on the core network and improves computing efficiency and response speed by distributing computing and storage resources to edge nodes close to data sources and user terminals. However, with the rapid development of the Internet and the Internet of Things, the number of smart mobile devices has shown an explosive growth trend, and the demand for data storage and computing resources has increased sharply.

[0003] To further improve the responsiveness of communication systems and reduce the demand for computing resources, a number of approaches have emerged in recent years using deep reinforcement learning (DRL) to optimize mobile edge computing. These methods combine the strengths of deep learning and reinforcement learning, demonstrating their exceptional ability to handle high-dimensional state spaces and complex policy optimization problems. Without explicit models or prior knowledge, these approaches achieve adaptation to dynamic computing environments and global optimization through interactive learning and policy sharing among intelligent agents.

[0004] Prior art has disclosed a method for offloading edge computing tasks in the Internet of Vehicles (IoV). This method constructs a three-layer IoV edge computing scenario model, obtains offloading tasks for mobile vehicles, and sets the total latency and total energy consumption costs during the task offloading process as a weighted total cost. The system total cost calculation problem is defined as a joint optimization problem of task offloading latency and energy consumption. This problem model is converted into a Markov decision process, and deep reinforcement learning is used to solve the joint optimization problem of task offloading latency and energy consumption, achieving effective IoV task offloading. However, this solution only treats edge computing tasks as independent units for local or edge server offloading, without considering inter-task dependencies or execution order. As a result, when processing complex tasks with predecessor or successor subtasks, it is unable to further reduce latency and energy consumption through scheduling optimization. Summary of the Invention

[0005] In order to solve the problems of slow convergence speed, high latency and energy consumption of computing tasks in current edge computing task offloading methods, the present invention proposes an edge computing task offloading method based on deep reinforcement learning, which reduces the latency and energy consumption of computing tasks and reduces costs. It introduces a pre-ordering model for computing offloading subtasks, improves the convergence speed of the deep reinforcement learning offloading model, and reduces decision-making time.

[0006] In order to achieve the above technical effects, the technical solutions of the present invention are as follows:

[0007] A method for offloading edge computing tasks based on deep reinforcement learning, the method comprising the following steps:

[0008] S1. Calculate the cost of task offloading by the user device under different offloading decisions;

[0009] S2. Based on the cost of task offloading by user devices under different offloading decisions, we construct a cost optimization objective function with the goal of minimizing the total cost for edge network end users, taking into account different offloading decisions and latency constraints.

[0010] S3. Construct a pre-sorted deep reinforcement learning task offloading model, wherein the pre-sorted deep reinforcement learning task offloading model includes a pre-sorting sub-model and a deep reinforcement learning sub-model. The pre-sorting sub-model is used to optimize the task offloading execution order in the user device, and the deep reinforcement learning sub-model includes a current action evaluation network and a target action evaluation network;

[0011] S4. Using the empirical data obtained from the interaction between the user device and the environment and the cost optimization objective function to train the pre-ordered deep reinforcement learning task offloading model, update the current action evaluation network and the target action evaluation network, and obtain the trained pre-ordered deep reinforcement learning task offloading model;

[0012] S5. Input the current state of the user device into the trained pre-sorted deep reinforcement learning task offloading model, optimize the execution order of task offloading based on the pre-sorted sub-model, output the offloading decision, and perform task offloading.

[0013] In this technical solution, the cost of task offloading for user devices under different offloading decisions is first calculated, and a cost optimization objective function is constructed. A pre-sorting model of computing offloading subtasks is introduced to construct a pre-sorted deep reinforcement learning task offloading model, which optimizes the execution order of edge computing tasks, improves the convergence speed of the deep reinforcement learning task offloading model, and reduces the decision-making time. The pre-sorted deep reinforcement learning task offloading model is trained with the goal of minimizing the total cost of edge network terminal users, and the trained pre-sorted deep reinforcement learning task offloading model is used to output the offloading decision for task offloading, thereby reducing the latency and energy consumption of computing tasks.

[0014] Preferably, the user equipment is a component of a mobile edge computing network, which includes: a cloud server, a base station and several user equipment, each base station is equipped with an edge server, the base station and the edge server are connected via a physical link, and the user equipment and the base station are connected via a wireless link;

[0015] The task offloading is edge computing task offloading, and the different offloading decisions include: retaining the edge computing task on the user device for execution, offloading the edge computing task to the edge server for execution, and offloading the edge computing task to the cloud server for execution.

[0016] Preferably, the mobile edge computing network time is discretized into T max time slots, forming a time slot set T, which is expressed as:

[0017] T={t|t∈1,2,...,T max}

[0018] The mobile edge computing network includes Z different types of tasks, and the task set L is expressed as:

[0019] L={Lz| z ∈1,2,...,Z}

[0020] The task set L generated by the user equipment of the mobile edge computing network j The expression is:

[0021] L j ={D j , B j , Z j , R j}

[0022] Among them, D j is the data size of the generated task, B j is the maximum tolerable delay, Z j is the task type, R j The size of the task result.

[0023] Preferably, the edge computing task is retained for execution on the user device, and the cost of execution on the user device is calculated, and the process is as follows:

[0024] Edge computing task L j Total uninstall time delay when uninstalling on local device Including calculation delay and waiting delay The expression is:

[0025]

[0026] Among them, Dj For edge computing tasks L j The size of the task data currently generated, pre(D j ) represents the edge computing task L j The data size of the predecessor task, is the CPU frequency required to execute the task, η is the CPU cycles required to execute a task of one megabyte in size;

[0027] Defining total execution energy consumption The expression is:

[0028]

[0029] Where ε is the effective capacitance coefficient per CPU cycle;

[0030] Edge computing task L j The total cost of executing on the local device is expressed as:

[0031]

[0032] Among them, γ is the weight coefficient.

[0033] Preferably, the process of offloading the edge computing task to the edge server for execution and calculating the execution cost is as follows:

[0034] Calculate the maximum information transmission rate from the user device to the edge server. The expression is:

[0035]

[0036] Among them, σ 2 is additive white Gaussian noise, From user devices to edge servers EV i The transmission bandwidth, From user devices to edge servers EV i The transmission power, h represents the channel gain, d j,i From user devices to edge servers EV i distance;

[0037] Edge computing task L j When offloading to edge servers, the total execution time is delayed Including calculation delay Waiting delay and transmission delay The expression is:

[0038]

[0039] in, Edge Server EV iAssigned to user equipment U j CPU frequency, pre(R j ) represents the edge computing task L j From user equipment U j To the edge server EV i The maximum information transmission rate, pre(D j ) represents the edge computing task L j The data size of the predecessor task, D j For edge computing tasks L j The size of the currently generated task data;

[0040] Define edge computing task L j Total energy consumption offloaded at edge servers The expression is:

[0041]

[0042] Among them, p j Indicates user equipment U j The transmission power, Edge Server EV i The transmission power, From user equipment U j To the edge server EV i One-way transmission time;

[0043] The edge computing task L j The total cost of offloading to the edge server is expressed as:

[0044]

[0045] Among them, γ is the weight coefficient.

[0046] Preferably, the process of offloading the edge computing task to the cloud server for execution and calculating the execution cost is as follows:

[0047] Calculate the maximum information transmission rate from the user device to the cloud server. The expression is:

[0048]

[0049] Among them, B j,cloud is the transmission bandwidth from user device to cloud server, p j,cloud is the transmission power from the user equipment to the cloud server, h represents the channel gain, d j,cloud is the distance between the user device and the cloud server;

[0050] Edge computing task L j When executing on a cloud server, the total execution time delay Including calculation delay Waiting delay and transmission delay The expression is:

[0051]

[0052] in, Assigned to user device U by cloud server j CPU frequency, pre(R j ) represents the edge computing task L j From user equipment U j The maximum information transmission rate to the cloud server, pre(D j ) represents the edge computing task L j The data size of the predecessor task, D j For edge computing tasks L j The size of the currently generated task data;

[0053] Define edge computing task L j Total energy consumption of cloud server execution The expression is:

[0054]

[0055] Among them, p cloud is the transmission power of the cloud server, From user equipment U j One-way transmission time to the cloud server;

[0056] The edge computing task L j The total cost of offloading to the cloud server is expressed as:

[0057]

[0058] Among them, γ is the weight coefficient.

[0059] Preferably, the expression of the construction cost optimization objective function described in step S2 is:

[0060]

[0061] The constraints are:

[0062]

[0063] Among them, j ={-1,0,1} is the constraint to limit the unloading position, B j The maximum tolerable delay for processing tasks.

[0064] Preferably, the pre-order sub-model in step S3 optimizes the user equipment U based on the priority attribute j The execution order of the tasks involved is as follows:

[0065] Defining a pre-ordered queue Used to store edge computing tasks L z Contains subtasks;

[0066] Defining parameters with priority attributes The expression is:

[0067]

[0068] Among them, I(l z ) is subtask l z The in-degree, O(l z ) is the out-degree of the subtask;

[0069] Calculate each subtask l z The priority attribute is in the pre-sorted queue in ascending order. Arrange the corresponding subtasks in, if the in-degree I(l z )=0, then the subtask l z As a queue The first subtask position of z ) is not equal to 0, and the out-degree O(l z )=0, then the subtask l z As a queue The last subtask.

[0070] Preferably, the current action evaluation network in step S3 includes the current Actor network μ and the current Critic network Q; the target action evaluation network includes the target Actor network μ′ and the target Critic network Q′; the deep reinforcement learning sub-model also includes the design of the state space, action space and reward function;

[0071] Among them, the expression of the state space in time slot t is:

[0072]

[0073] Among them, S t,j For user equipment U j The state at time slot t is expressed as:

[0074]

[0075] in, The user equipment U in the previous time slot j Get the size of computing resources allocated by the system, is the total cost of the system in the previous time slot, D t,j For user equipment U j The size of the task data generated in the current time slot, B t,j For user equipment U j The maximum tolerable delay of the task in the current time slot;

[0076] The expression of the action space in time slot t is:

[0077] a t ={a t,j |j∈[1,N]}

[0078] Among them, a t,j For user equipment U j The action in time slot t is expressed as:

[0079] a t,j =[a j,-1 , a j,0 , a j,1 ,...,a j,M}

[0080] Among them, a j,-1 Indicates that the action is to offload the computing task to the cloud server, a j,0 Indicates offloading computing tasks to local devices, a j,1 ,...,a j,M Indicates offloading computing tasks to an edge server;

[0081] The expression of the reward function in time slot t is:

[0082] r t ={r t,j |j∈[1,N]}where r t,j For user equipment U j The reward function at time slot t is expressed as:

[0083]

[0084] Among them, C j (ζ j ) is the user equipment U j In time slot t, ζ j Constrain the execution cost of offloading.

[0085] Preferably, in step S4, the process of using the empirical data obtained by the interaction between the user device and the environment and the cost optimization objective function to train the pre-ordered deep reinforcement learning task offloading model, updating the current network and the target network, and obtaining the trained pre-ordered deep reinforcement learning task offloading model is as follows:

[0086] Define user equipment U j The current Actor network parameters, current Critic network parameters, target Actor network parameters, and target Critic network parameters are:

[0087] Defining experience tuples t 、a t 、r t 、s t+1 >, the user equipment U j Interact with the environment, obtain the corresponding reward and next state according to the reward function, store them in the experience replay buffer pool, and randomly extract small batches of experience data for parameter update during subsequent model training;

[0088] Using loss function Loss t,j Update the current Critic network, loss function Loss t,j The expression is:

[0089]

[0090] Among them, b is the number of small batches of experience data randomly extracted from the experience return buffer pool, y t,j is the target estimate, Q t,j is the true action value; when the loss function Loss t,j When the minimum value is reached, the current Critic network is updated;

[0091] Use policy gradient ascent to update the current Actor network. The policy gradient ascent update expression is:

[0092]

[0093] Use soft update to update the target Critic network parameters and target Actor network parameters respectively. The expression is:

[0094]

[0095] Among them, τ∈[0,1] is the target network update parameter;

[0096] Current Actor network parameters Current Critic network parameters Target Actor Network Parameters Target Critic Network Parameters After all updates are completed, the trained pre-sorted deep reinforcement learning task offloading model is obtained.

[0097] Compared with the prior art, the beneficial effects of the technical solution of the present invention are: ​

[0098] The present invention proposes an edge computing task offloading method based on deep reinforcement learning. First, the cost of task offloading by user devices under different offloading decisions is calculated, and a cost optimization objective function is constructed. A pre-ordering model of computing offloading subtasks is introduced to construct a pre-ordered deep reinforcement learning task offloading model, which optimizes the execution order of edge computing tasks, improves the convergence speed of the deep reinforcement learning task offloading model, and reduces the decision-making time. The pre-ordered deep reinforcement learning task offloading model is trained with the goal of minimizing the total cost of edge network terminal users, and the trained pre-ordered deep reinforcement learning task offloading model is used to output the offloading decision for task offloading, thereby reducing the latency and energy consumption of computing tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0099] Figure 1 A schematic diagram showing a flow chart of a method for offloading edge computing tasks based on deep reinforcement learning proposed in Example 1 of the present invention;

[0100] Figure 2 A schematic diagram showing the architecture of a mobile edge computing network proposed in Example 2 of the present invention;

[0101] Figure 3 A flowchart showing the execution order of tasks included in optimizing user equipment proposed in Embodiment 2 of the present invention;

[0102] Figure 4 A structural diagram showing the pre-sorted deep reinforcement learning task offloading model proposed in Example 2 of the present invention;

[0103] Figure 5 A graph showing the total computational cost of executing edge computing tasks under different numbers of edge nodes under different offloading strategies proposed in Example 2 of the present invention. DETAILED DESCRIPTION

[0104] The accompanying drawings are for illustrative purposes only and are not to be construed as limiting this patent;

[0105] In order to better illustrate this embodiment, some parts of the drawings may be omitted, enlarged, or reduced, and do not represent the actual size;

[0106] It is understandable to those skilled in the art that descriptions of certain well-known contents may be omitted in the drawings.

[0107] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.

[0108] The positional relationships described in the drawings are for illustrative purposes only and should not be construed as limiting this patent;

[0109] Example 1

[0110] This embodiment proposes an edge computing task offloading method based on deep reinforcement learning. The flowchart of this method is shown in Figure 1 , including the following steps:

[0111] S1. Calculate the cost of task offloading by the user device under different offloading decisions;

[0112] S2. Based on the cost of task offloading by user devices under different offloading decisions, we construct a cost optimization objective function with the goal of minimizing the total cost for edge network end users, taking into account different offloading decisions and latency constraints.

[0113] S3. Construct a pre-sorted deep reinforcement learning task offloading model, wherein the pre-sorted deep reinforcement learning task offloading model includes a pre-sorting sub-model and a deep reinforcement learning sub-model. The pre-sorting sub-model is used to optimize the task offloading execution order in the user device, and the deep reinforcement learning sub-model includes a current action evaluation network and a target action evaluation network;

[0114] S4. Using the empirical data obtained from the interaction between the user device and the environment and the cost optimization objective function to train the pre-ordered deep reinforcement learning task offloading model, update the current action evaluation network and the target action evaluation network, and obtain the trained pre-ordered deep reinforcement learning task offloading model;

[0115] S5. Input the current state of the user device into the trained pre-sorted deep reinforcement learning task offloading model, optimize the execution order of task offloading based on the pre-sorted sub-model, output the offloading decision, and perform task offloading.

[0116] In this embodiment, the cost of task offloading by the user device under different offloading decisions is first calculated, a cost optimization objective function is constructed, and a pre-ordering model of computing offloading subtasks is introduced to construct a pre-ordered deep reinforcement learning task offloading model. The execution order of edge computing tasks is optimized, the convergence speed of the deep reinforcement learning task offloading model is improved, and the decision-making time is reduced. The pre-ordered deep reinforcement learning task offloading model is trained with the goal of minimizing the total cost of edge network terminal users, and the trained pre-ordered deep reinforcement learning task offloading model is used to output the offloading decision for task offloading, thereby reducing the latency and energy consumption of computing tasks.

[0117] Example 2

[0118] In this embodiment, the user equipment is a component of a mobile edge computing network, which includes: a cloud server, a base station, and several user equipment. Each base station is equipped with an edge server. The base station and the edge server are connected through a physical link, and the user equipment and the base station are connected through a wireless link. The architecture diagram of the mobile edge computing network is shown in FIG. Figure 2 As shown;

[0119] The task offloading is edge computing task offloading, and the different offloading decisions include: retaining the edge computing task on the user device for execution, offloading the edge computing task to the edge server for execution, and offloading the edge computing task to the cloud server for execution.

[0120] In this embodiment, the mobile edge computing network time is discretized into T max time slots, forming a time slot set T, which is expressed as:

[0121] T={t|t∈1,2,...,T max}

[0122] The mobile edge computing network includes Z different types of tasks, and the task set L is expressed as:

[0123] L={L z |z∈1,2,...,Z} the task set L generated by the user equipment of the mobile edge computing network j The expression is:

[0124] L j ={D j ,B j ,Z j ,R j}

[0125] Among them, D j is the data size of the generated task, B j is the maximum tolerable delay, Z j is the task type, R j The size of the task result.

[0126] In this embodiment, the edge computing task is retained for execution on the user device, and the cost of execution on the user device is calculated. The process is as follows:

[0127] Edge computing task L j Total uninstall time delay when uninstalling on local device Including calculation delay and waiting delay The expression is:

[0128]

[0129] Among them, D j For edge computing tasks L j The size of the task data currently generated, pre(D j ) represents the edge computing task L j The data size of the predecessor task, is the CPU frequency required to execute the task, η is the CPU cycles required to execute a task of one megabyte in size;

[0130] Defining total execution energy consumption The expression is:

[0131]

[0132] Where ε is the effective capacitance coefficient per CPU cycle;

[0133] Edge computing task L j The total cost of executing on the local device is expressed as:

[0134]

[0135] Among them, γ is the weight coefficient.

[0136] In this embodiment, the edge computing task is offloaded to the edge server for execution, and the execution cost is calculated. The process is as follows:

[0137] Calculate the maximum information transmission rate from the user device to the edge server. The expression is:

[0138]

[0139] Among them, σ 2 is additive white Gaussian noise, From user devices to edge servers EV i The transmission bandwidth, From user devices to edge servers EV i The transmission power, h represents the channel gain, d j,i From user devices to edge servers EV i distance;

[0140] Edge computing task L j When offloading to edge servers, the total execution time is delayed Including calculation delay Waiting delay and transmission delay The expression is:

[0141]

[0142] in, Edge Server EV i Assigned to user equipment U j CPU frequency, pre(R j ) represents the edge computing task L j From user equipment U j To the edge server EVi The maximum information transmission rate, pre(D j ) represents the edge computing task L j The data size of the predecessor task, D j For edge computing tasks L j The size of the currently generated task data;

[0143] Define edge computing task L j Total energy consumption offloaded at edge servers The expression is:

[0144]

[0145] Among them, p j Indicates user equipment U j The transmission power, Edge Server EV i The transmission power, From user equipment U j To the edge server EV i One-way transmission time;

[0146] The edge computing task L j The total cost of offloading to the edge server is expressed as:

[0147]

[0148] Among them, γ is the weight coefficient.

[0149] In this embodiment, the edge computing task is offloaded to the cloud server for execution, and the execution cost is calculated. The process is as follows:

[0150] Calculate the maximum information transmission rate from the user device to the cloud server. The expression is:

[0151]

[0152] Among them, B j,cloud is the transmission bandwidth from user device to cloud server, p j,cloud is the transmission power from the user equipment to the cloud server, h represents the channel gain, d j,cloud is the distance between the user device and the cloud server;

[0153] Edge computing task L j When executing on a cloud server, the total execution time delay Including calculation delay Waiting delay and transmission delay The expression is:

[0154]

[0155] in, Assigned to user device U by cloud server j CPU frequency, pre(R j ) represents the edge computing task L j From user equipment U j The maximum information transmission rate to the cloud server, pre(D j ) represents the edge computing task L j The data size of the predecessor task, D j For edge computing tasks L j The size of the currently generated task data;

[0156] Define edge computing task L j Total energy consumption of cloud server execution The expression is:

[0157]

[0158] Among them, p cloud is the transmission power of the cloud server, From user equipment U j One-way transmission time to the cloud server;

[0159] The edge computing task L j The total cost of offloading to the cloud server is expressed as:

[0160]

[0161] Among them, γ is the weight coefficient.

[0162] In this embodiment, the expression of the construction cost optimization objective function described in step S2 is:

[0163]

[0164] The constraints are:

[0165]

[0166] Among them, j ={-1,0,1} is the constraint to limit the unloading position, B j The maximum tolerable delay for processing tasks.

[0167] Specifically, the reward function is used to clarify the optimization goal, that is, to guide the system towards the desired optimization direction; the user device will continuously adjust the uninstallation strategy based on the feedback of the reward function to maximize the reward value and thus achieve the optimization goal; after the system performs a certain uninstallation action, the reward value calculated by the reward function can be used to determine whether the decision meets the optimization goal, thereby providing a quantitative evaluation standard for each uninstallation decision; the smaller the total execution cost, the higher the reward value, and the higher the reward value, the better the decision.

[0168] In this embodiment, the pre-order sub-model in step S3 optimizes the user equipment U based on the priority attribute. j The execution order of the tasks involved is as follows:

[0169] Defining a pre-ordered queue Used to store edge computing tasks L z Contains subtasks;

[0170] Defining parameters with priority attributes The expression is:

[0171]

[0172] Among them, I(l z ) is subtask l z The in-degree, O(l z ) is the out-degree of the subtask;

[0173] Calculate each subtask l z The priority attribute is in the pre-sorted queue in ascending order. Arrange the corresponding subtasks in, if the in-degree I(l z )=0, then the subtask l z As a queue The first subtask position of z ) is not equal to 0, and the out-degree O(l z )=0, then the subtask l z As a queue The last subtask of

[0174] Specifically, to optimize the user equipment U j The flowchart of the execution order of the included tasks can be found in Figure 3 First, initialize the sorting queue and check the in-degree I(l z ) is 0; if the in-degree is 0, the corresponding subtask is added to the sorting queue, and then the priority attribute parameters of the subsequent subtasks are calculated Then update the queue in descending order of the parameters, and return to the task queue to complete the sorting after the update is completed;

[0175] If the in-degree is not 0, it will be directly returned to the task queue to complete the sorting.

[0176] In this embodiment, the current action evaluation network in step S3 includes the current Actor network μ and the current Critic network Q; the target action evaluation network includes the target Actor network μ′ and the target Critic network Q′; the deep reinforcement learning sub-model also includes the design of the state space, action space and reward function;

[0177] Among them, the expression of the state space in time slot t is:

[0178] s t ={s t,j |j∈[1,N]}

[0179] Among them, S t,j For user equipment U j The state at time slot t is expressed as:

[0180]

[0181] in, The user equipment U in the previous time slot j Get the size of computing resources allocated by the system, is the total cost of the system in the previous time slot, D t,j For user equipment U j The size of the task data generated in the current time slot, B t,j For user equipment U j The maximum tolerable delay of the task in the current time slot;

[0182] The expression of the action space in time slot t is:

[0183] a t ={a t,j |j∈[1,N]}where a t,j For user equipment U j The action in time slot t is expressed as:

[0184] a t,j ={a j,-1 , a j,0 , a j,1 ,...,a j,M}

[0185] Among them, a j,-1 Indicates that the action is to offload the computing task to the cloud server, a j,0 Indicates offloading computing tasks to local devices, a j,1 ,...,a j,M Indicates offloading computing tasks to an edge server;

[0186] The expression of the reward function in time slot t is:

[0187] r t ={r t,j |j∈[1,N]}

[0188] Among them, r t,j For user equipment U j The reward function at time slot t is expressed as:

[0189] r t,j =-C j (ζ j )

[0190] Among them, C j (ζ j ) is the user equipment U j In time slot t, ζ j Constrain the execution cost of offloading.

[0191] In this embodiment, step S4 uses the empirical data obtained by the interaction between the user device and the environment and the cost optimization objective function to train the pre-ordered deep reinforcement learning task offloading model, updates the current network and the target network, and obtains the trained pre-ordered deep reinforcement learning task offloading model. The process is as follows:

[0192] Define user equipment U j The current Actor network parameters, current Critic network parameters, target Actor network parameters, and target Critic network parameters are:

[0193] Defining experience tuples t 、a t 、r t 、s t+1 >, the user equipment U j Interact with the environment, obtain the corresponding reward and next state according to the reward function, store them in the experience replay buffer pool, and randomly extract small batches of experience data for parameter update during subsequent model training;

[0194] Using loss function Loss t,j Update the current Critic network, loss function Loss t,j The expression is:

[0195]

[0196] Among them, b is the number of small batches of experience data randomly extracted from the experience return buffer pool, y t,j is the target estimate, Q t,j ​is the true action value; when the loss function Loss t,j When the minimum value reaches 0, the current Critic network is updated;

[0197] Use policy gradient ascent to update the current Actor network. The policy gradient ascent update expression is:

[0198]

[0199] Use soft update to update the target Critic network parameters and target Actor network parameters respectively. The expression is:

[0200]

[0201] Among them, τ∈[0,1] is the target network update parameter;

[0202] The structure diagram of the pre-sorted deep reinforcement learning task offloading model is as follows Figure 4 As shown;

[0203] Current Actor network parameters Current Critic network parameters Target Actor Network Parameters Target Critic Network Parameters After all updates are completed, the trained pre-sorted deep reinforcement learning task offloading model is obtained;

[0204] The total computational cost of executing edge computing tasks under different offloading strategies and different numbers of edge nodes is as follows: Figure 5 As shown in Figure 3, the total computational cost of executing edge computing tasks using the pre-sorted deep reinforcement learning task offloading model is less than the total computational cost of any offloading strategy.

[0205] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. A method for offloading edge computing tasks based on deep reinforcement learning, characterized in that: The following steps are involved: S1. Calculate the cost of task offloading by the user device under different offloading decisions; S2. Based on the cost of task offloading by user devices under different offloading decisions, we construct a cost optimization objective function with the goal of minimizing the total cost for edge network end users, taking into account different offloading decisions and latency constraints. S3. Construct a pre-sorted deep reinforcement learning task offloading model, wherein the pre-sorted deep reinforcement learning task offloading model includes a pre-sorting sub-model and a deep reinforcement learning sub-model. The pre-sorting sub-model is used to optimize the task offloading execution order in the user device, and the deep reinforcement learning sub-model includes a current action evaluation network and a target action evaluation network; S4. Using the empirical data obtained from the interaction between the user device and the environment and the cost optimization objective function to train the pre-ordered deep reinforcement learning task offloading model, update the current action evaluation network and the target action evaluation network, and obtain the trained pre-ordered deep reinforcement learning task offloading model; S5. Input the current state of the user device into the trained pre-sorted deep reinforcement learning task offloading model, optimize the execution order of task offloading based on the pre-sorted sub-model, output the offloading decision, and perform task offloading.

2. The edge computing task offloading method based on deep reinforcement learning according to claim 1 is characterized in that: The user equipment is a component of a mobile edge computing network, which includes: a cloud server, a base station and several user equipment, each base station is equipped with an edge server, the base station and the edge server are connected via a physical link, and the user equipment and the base station are connected via a wireless link; The task offloading is edge computing task offloading, and the different offloading decisions include: retaining the edge computing task on the user device for execution, offloading the edge computing task to the edge server for execution, and offloading the edge computing task to the cloud server for execution.

3. The edge computing task offloading method based on deep reinforcement learning according to claim 2 is characterized in that: Discretize the mobile edge computing network time into T max time slots, forming a time slot set T, which is expressed as: T={t|t∈1,2,...,T max } The mobile edge computing network includes Z different types of tasks, and the task set L is expressed as: L={L z |Z∈1,2,...,Z} The task set L generated by the user equipment of the mobile edge computing network j The expression is: L j ={D j ,B j ,Z j ,R j } Among them, D j is the data size of the generated task, B j is the maximum tolerable delay, Z j is the task type, R j The size of the task result.

4. The edge computing task offloading method based on deep reinforcement learning according to claim 2 is characterized in that: The process of retaining the edge computing task on the user device and calculating the cost of executing the task on the user device is as follows: Edge computing task L j Total uninstall time delay when uninstalling on local device Including calculation delay and waiting delay The expression is: Among them, D j For edge computing tasks L j The size of the task data currently generated, pre(D j ) represents the edge computing task L j The data size of the predecessor task, is the CPU frequency required to execute the task, η is the CPU cycles required to execute a task of one megabyte in size; Defining total execution energy consumption The expression is: Where ε is the effective capacitance coefficient per CPU cycle; Edge computing task L j The total cost of executing on the local device is expressed as: Among them, γ is the weight coefficient.

5. The edge computing task offloading method based on deep reinforcement learning according to claim 2 is characterized in that: The process of offloading edge computing tasks to edge servers and calculating execution costs is as follows: Calculate the maximum information transmission rate from the user device to the edge server. The expression is: Among them, σ 2 is additive white Gaussian noise, From user devices to edge servers EV i The transmission bandwidth, From user devices to edge servers EV i The transmission power, h represents the channel gain, d j,i From user devices to edge servers EV i distance; Edge computing task L j When offloading to edge servers, the total execution time is delayed Including calculation delay Waiting delay and transmission delay The expression is: in, Edge Server EV i Assigned to user equipment U j CPU frequency, pre(R j ) represents the edge computing task L j From user equipment U j To the edge server EV i The maximum information transmission rate, pre(D j ) represents the edge computing task L j The data size of the predecessor task, D j For edge computing tasks L j The size of the currently generated task data; Define edge computing task L j Total energy consumption offloaded at edge servers The expression is: Among them, p j Indicates user equipment U j The transmission power, Edge Server EV i The transmission power, From user equipment U j To the edge server EV i One-way transmission time; The edge computing task L j The total cost of offloading to the edge server is expressed as: Among them, γ is the weight coefficient.

6. The edge computing task offloading method based on deep reinforcement learning according to claim 2 is characterized in that: The process of offloading edge computing tasks to cloud servers and calculating execution costs is as follows: Calculate the maximum information transmission rate from the user device to the cloud server. The expression is: Among them, B j,cloud is the transmission bandwidth from user device to cloud server, p j,cloud is the transmission power from the user equipment to the cloud server, h represents the channel gain, d j,cloud is the distance between the user device and the cloud server; Edge computing task L j When executing on a cloud server, the total execution time delay Including calculation delay Waiting delay and transmission delay The expression is: in, Assigned to user device U by cloud server j CPU frequency, pre(R j ) represents the edge computing task L j From user equipment U j The maximum information transmission rate to the cloud server, pre(D j ) represents the edge computing task L j The data size of the predecessor task, D j For edge computing tasks L j The size of the currently generated task data; Define edge computing task L j Total energy consumption of cloud server execution The expression is: Among them, p cloud is the transmission power of the cloud server, From user equipment U j One-way transmission time to the cloud server; The edge computing task L j The total cost of offloading to the cloud server is expressed as: Among them, γ is the weight coefficient.

7. The edge computing task offloading method based on deep reinforcement learning according to claim 1 is characterized in that: The expression of the construction cost optimization objective function described in step S2 is: The constraints are: Among them, j ={-1,0,1} is the constraint to limit the unloading position, B j The maximum tolerable delay for processing tasks.

8. The edge computing task offloading method based on deep reinforcement learning according to claim 7 is characterized in that: Step S3: The pre-order sub-model optimizes the user equipment U based on the priority attribute. j The execution order of the tasks involved is as follows: Defining a pre-ordered queue Used to store edge computing tasks L z Contains subtasks; Defining parameters with priority attributes The expression is: Among them, I(l z ) is subtask l z The in-degree, O(l z ) is the out-degree of the subtask; Calculate each subtask l z The priority attribute is in the pre-sorted queue in ascending order. Arrange the corresponding subtasks in, if the in-degree I(l z )=0, then the subtask l z As a queue The first subtask position of z ) is not equal to 0, and the out-degree O(l z )=0, then the subtask l z As a queue The last subtask.

9. The edge computing task offloading method based on deep reinforcement learning according to claim 8 is characterized in that: In step S3, the current action evaluation network includes the current actor network μ and the current critic network Q; the target action evaluation network includes the target actor network μ′ and the target critic network Q′; the deep reinforcement learning sub-model also includes the design of the state space, action space and reward function; Among them, the expression of the state space in time slot t is: S t ={S t,j |j∈[1,N]} Among them, S t,j For user equipment U j The state at time slot t is expressed as: in, The user equipment U in the previous time slot j Get the size of computing resources allocated by the system, is the total cost of the system in the previous time slot, D t,j For user equipment U j The size of the task data generated in the current time slot, B t,j For user equipment U j The maximum tolerable delay of the task in the current time slot; The expression of the action space in time slot t is: a t ={a t,j |j∈[1,N]} Among them, a t,j For user equipment U j The action in time slot t is expressed as: a t,j ={a j,-1 ,a j,0 ,a j,1 ,...,a j,M } Among them, a j,-1 Indicates that the action is to offload the computing task to the cloud server, a j,0 Indicates offloading the computing task to the local device, a j,1 ,...,a j,M Indicates offloading computing tasks to an edge server; The expression of the reward function in time slot t is: r t ={r t,j |j∈[1,N]} Among them, r t,j For user equipment U j The reward function at time slot t is expressed as: r t,j =-C j (g) j ) Among them, C j (ζ j ) is the user equipment U j In time slot t, ζ j Constrain the execution cost of offloading.

10. The method for offloading edge computing tasks based on deep reinforcement learning according to claim 9, characterized in that: Step S4 uses the empirical data obtained by the interaction between the user device and the environment and the cost optimization objective function to train the pre-ordered deep reinforcement learning task offloading model, updates the current network and the target network, and obtains the trained pre-ordered deep reinforcement learning task offloading model. The process is as follows: Define user equipment U j The current Actor network parameters, current Critic network parameters, target Actor network parameters, and target Critic network parameters are: Defining experience tuples t 、a t 、r t 、s t+1 >, the user equipment U j Interact with the environment, obtain the corresponding reward and next state according to the reward function, store them in the experience replay buffer pool, and randomly extract small batches of experience data for parameter update during subsequent model training;​ Using loss function Loss t,j Update the current Critic network, loss function Loss t,j The expression is: Among them, b is the number of small batches of experience data randomly extracted from the experience return buffer pool, y t,j is the target estimate, Q t,j is the true action value; when the loss function Loss t,j When the minimum value is reached, the current Critic network is updated; Use policy gradient ascent to update the current Actor network. The policy gradient ascent update expression is: Use soft update to update the target Critic network parameters and target Actor network parameters respectively. The expression is: Among them, τ∈[0,1] is the target network update parameter; Current Actor network parameters Current Critic network parameters Target Actor Network Parameters Target Critic Network Parameters After all updates are completed, the trained pre-sorted deep reinforcement learning task offloading model is obtained.

Citation Information

Cited By

  • Edge computing task unloading method based on deep reinforcement learning

    CN122219999A