Depending task unloading decision algorithm based on hierarchical reinforcement learning

By dividing the dependent task offloading decision problem into two sub-problems: subtask offloading order and location decision, a hierarchical reinforcement learning algorithm is used to optimize the decision strategy, which solves the problem of unreasonable dependent task offloading strategy in edge computing, reduces task processing delay and improves task completion rate.

CN120631464APending Publication Date: 2025-09-12HENAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510731673.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies find it difficult to effectively solve the rationality problem of dependent task offloading strategies in edge computing, especially in the dependency relationship between subtasks and dynamic offloading environments, resulting in high task processing latency and unreasonable offloading strategies.

Method used

Abstract: In order to solve the problem of dependent task offloading and its application prospect, a hierarchical reinforcement learning-based dependent task offloading decision algorithm (HSOTO) is proposed. The dependent task offloading decision problem is divided into two sub-problems: sub-task offloading order decision and offloading location decision. These sub-problems are processed by OO module and OL module respectively. A reward feedback mechanism is designed to optimize the decision strategies of the two sub-problems.

Benefits of technology

Through the hierarchical reinforcement learning algorithm, the task processing delay is reduced, and the rationality of the dependent task offloading strategy and the task completion rate are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120631464A_ABST
    Figure CN120631464A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of edge computing, in particular to a dependent task unloading decision algorithm based on hierarchical reinforcement learning, which regards a dependent task unloading decision problem as a high coupling problem formed by two sub-problems, and proposes linkage solution of an unloading sequence decision module and an unloading position decision module. OO is proposed to solve a sub-task unloading sequence decision sub-problem, a sub-task sequence decision state conversion mechanism is constructed based on DAG, a reinforcement learning method is designed to solve the sub-problem, and the optimal unloading sequence of the sub-tasks is obtained. And proposing an OL to solve a sub-task unloading position decision sub-problem, constructing a sub-task unloading position decision model, and obtaining a final sub-task unloading decision. The OO transmits a subtask unloading sequence to the OL, the OL makes unloading position decisions in sequence according to the unloading sequence, two sub-problems are alternately optimized by taking unloading time delay as a linkage feedback factor between two modules, and the reasonability of setting depending on a task unloading strategy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of edge computing technology, and in particular to a dependent task offloading decision algorithm based on hierarchical reinforcement learning. Background Art

[0002] Edge computing (EC) is a distributed computing architecture that sinks computing resources to edge nodes close to users, thereby reducing the overall system latency. With the diversification of user needs and the improvement of hardware device performance, the tasks that users need to handle are becoming increasingly complex and usually consist of many interdependent subtasks. For example, in autonomous driving tasks, subtasks include environmental perception, target detection, environmental modeling, path planning, and motion control. In this case, in order to reduce the user's computing burden and obtain lower task processing latency, it is necessary to offload the above tasks to resource-rich edge servers for execution.

[0003] Dependent task offloading decision-making is one of the most important resource scheduling technologies in edge computing. It primarily addresses the order in which tasks should be offloaded and where they should be processed. The goal of this decision-making process is to reduce task processing latency by selecting a reasonable subtask offloading sequence and task offloading location, ensuring that tasks are efficiently processed within the latest completion time.

[0004] The complexity of offloading factors makes it difficult to determine the optimal offloading strategy for dependent tasks, primarily due to the following two factors: First, the dependencies between subtasks. Dependent tasks consist of multiple interdependent subtasks. Ignoring the constraints of these dependencies can lead to task processing failures and increased task processing latency. Existing research first topologically sorts the DAG and then allocates resources based on hierarchy or critical paths to ensure the legitimacy and basic fairness of the offloading order. However, this approach only ensures a reasonable order and cannot dynamically adapt to resource changes. Second, the offloading environment is highly dynamic. Edge server computing resources, network bandwidth, and latency are significantly affected by load fluctuations and network jitter. If the environment changes dramatically, static offloading strategies often fail, resulting in offloading failures and increased task processing latency. This approach utilizes online prediction of network computing resources and server load, and jointly optimizes offloading and resource allocation within the prediction window. However, this approach is highly dependent on prediction accuracy, and strategies fail when errors are large. In addition, the subtask offloading location decision is based on the subtask offloading order decision, and the subtask offloading order decision is affected by the subtask offloading location decision, resulting in a highly coupled problem, making it difficult to obtain the optimal dependent task offloading strategy, and thus resulting in poor rationality in the setting of the dependent task offloading strategy. Summary of the Invention

[0005] Dependent task offloading involves migrating tasks from the user's device to edge servers close to the user to reduce task processing latency. The dependencies between subtasks and the dynamic offloading environment create complex offloading factors, making it difficult to determine the optimal dependent task offloading strategy and leading to high task processing latency.

[0006] To obtain the optimal dependent task offloading strategy, the dependency relationship between subtasks and the impact of the high dynamics of the offloading environment should be comprehensively considered, and the focus should be on solving the high coupling problem between subtask offloading order decisions and subtask offloading location decisions.

[0007] To address the technical issue of poor rationality in setting dependent task offloading strategies, this paper proposes a hierarchical reinforcement learning-based dependent task offloading decision algorithm (HSOTO). This algorithm treats the dependent task offloading decision problem as a highly coupled problem consisting of two subproblems, and proposes a coordinated solution for the offloading order decision module and the offloading location decision module. First, an object-oriented (OO) module is proposed to solve the subtask offloading order decision subproblem. A state transition mechanism for the subtask order decision is constructed based on a distributed graph (DAG). A reinforcement learning approach is designed to solve this subproblem and obtain the optimal subtask offloading order. Then, an object-oriented (OL) module is proposed to solve the subtask offloading location decision subproblem. A time-based queue state transition mechanism is designed and a subtask offloading location decision model is constructed to obtain the final subtask offloading decision. Finally, the OO module transmits the subtask offloading order to the OL module, which makes offloading location decisions based on the offloading order. The OL module uses the offloading delay as a feedback factor between the two modules to alternately optimize the two subproblems. Experimental results show that the HSOTO algorithm outperforms the baseline algorithm in terms of both system latency and task completion rate.

[0008] The present invention provides a dependent task offloading decision algorithm based on hierarchical reinforcement learning, which includes:

[0009] Step 1: Use the OO module to build a task topology structure based on a directed acyclic graph, establish a subtask sequence decision state transition mechanism, and use a reinforcement learning algorithm to obtain the optimal offloading order of subtasks;

[0010] Step 2: Establish a queue state transition mechanism based on time flow through the OL module, build a subtask unloading location decision model, and use the reinforcement learning algorithm to obtain the optimal unloading location of the subtask;

[0011] Step three: establish a linkage feedback mechanism between the OO module and the OL module, pass the unloading order generated by the OO module to the OL module, and the OL module makes unloading location decisions in turn according to the unloading order, and uses the task unloading delay as a feedback link to alternately optimize the decision strategies of the OO module and the OL module.

[0012] Optionally, the OO module in step 1 constructs a subtask offloading sequence decision model and adopts the D3QN algorithm to obtain the optimal offloading sequence of the subtasks.

[0013] Optionally, the OL module in step 2 constructs a subtask unloading location decision model and adopts the D3QN algorithm to obtain the optimal unloading location of the subtask.

[0014] Optionally, in step three, the linkage feedback mechanism passes the unloading order generated by the OO module to the OL module. The OL module makes subtask unloading location decisions in turn according to the unloading order, and uses the task processing delay as the reward feedback of the OO module, alternately optimizing the decision strategies of the OO module and the OL module to obtain the optimal dependent task unloading strategy.

[0015] It should be noted that existing research on the highly coupled dependent task offloading decision-making problem typically treats the two sub-problems mentioned above as independent problems. For example, the scheduling priority of a task is determined by calculating the maximum distance from the task to the exit task, and computing resources are allocated to tasks according to priority order. Reinforcement learning has shown good performance in problems with high dynamic complexity. First, a multi-priority task sorting algorithm is proposed, and then an algorithm based on deep deterministic policy gradient is proposed to find the optimal offloading strategy. However, the high coupling between the two stages is ignored. In the face of the dynamically changing offloading environment, it is difficult to obtain the task priority at each moment, resulting in the inability to obtain the optimal offloading decision for the problem.

[0016] The present invention has the following beneficial effects:

[0017] For highly coupled problems, a "divide and conquer" approach is often adopted. This involves breaking the overall problem into several subproblems, solving each subproblem independently. An iterative feedback mechanism is introduced between the subproblems, with the lower layers verifying the strategies of the upper layers and providing feedback to the upper layers to achieve collaborative optimization. Hierarchical reinforcement learning (HRL) is an effective tool for solving coupled problems. It divides complex problems into multiple, mutually coupled subproblems. Using reinforcement learning algorithms, it considers specific influencing factors within each subproblem and formulates corresponding decision-making strategies. Then, through a collaborative feedback mechanism between the subproblems, it alternately optimizes the decision-making strategies for each subproblem, ultimately achieving the optimal solution for the overall problem. Based on this, the present invention divides the dependent task offloading decision problem into two sub-problems: sub-task offloading order decision and sub-task offloading location decision, and models the sub-task offloading order decision and sub-task offloading location decision respectively. A dependent task offloading decision algorithm based on hierarchical reinforcement learning is proposed, and the two sub-problems are respectively handed over to a sorting decision module based on sorting state transition and an offloading location decision module based on queue state transition for processing. A reward feedback mechanism is designed, and the delay gain is used as the feedback connection between the two sub-problems. The sub-task offloading order and the sub-task offloading location strategy are cyclically optimized, thereby reducing the task processing delay and improving the rationality of the dependent task offloading strategy setting. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 This is a flowchart of a hierarchical reinforcement learning-based dependent task offloading decision algorithm of the present invention;

[0020] Figure 2 This is a schematic diagram of a dependent task offloading decision scenario of the present invention;

[0021] Figure 3 This is a schematic diagram of the subtask sequence decision state transition of the present invention;

[0022] Figure 4 A schematic diagram of queue state transition according to the present invention;

[0023] Figure 5 A schematic diagram of the division of subtask processing stages of the present invention;

[0024] Figure 6 Schematic diagram of the HSOTO algorithm architecture of the present invention;

[0025] Figure 7 Schematic diagram of system delay comparison of the present invention;

[0026] Figure 8 Schematic diagram for comparing task completion rates of the present invention. DETAILED DESCRIPTION

[0027] To further illustrate the technical means and effects employed by the present invention to achieve its intended objectives, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementations, structures, features, and effects of the technical solutions proposed by the present invention. In the following description, references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.

[0028] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0029] Faced with the highly coupled problem of dependent task offloading decision-making, existing research on dependent task offloading decision-making usually treats the above two sub-problems as independent problems. For example, HEFT determines the scheduling priority of tasks by calculating the maximum distance from a task to the exit task, and allocates computing resources to tasks in order of priority. Reinforcement learning has shown good performance in problems with high dynamic complexity. First, a multi-priority task sorting algorithm is proposed, and then an algorithm based on deep deterministic policy gradient is proposed to find the optimal offloading strategy. However, the high coupling between the two stages is ignored. Faced with a dynamically changing offloading environment, it is difficult to obtain the task priority at each moment, resulting in the inability to obtain the optimal offloading decision for the problem.

[0030] For highly coupled problems, the idea of ​​"divide and conquer" is usually adopted. The whole problem is solved into several sub-problems, each sub-problem is solved independently, and an iterative feedback mechanism is introduced between the sub-problems. The lower layer verifies the upper layer's strategy and provides feedback to the upper layer to achieve collaborative optimization.

[0031] Hierarchical reinforcement learning (HRL) is an effective tool for solving coupled problems. It divides complex problems into multiple coupled subproblems. Using a reinforcement learning algorithm, it considers specific influencing factors within each subproblem and formulates corresponding decision-making strategies. Then, through a collaborative feedback mechanism between the subproblems, it alternately optimizes the decision-making strategies for each subproblem, ultimately achieving the optimal solution for the overall problem.

[0032] Based on this, the present invention divides the dependent task offloading decision problem into two sub-problems: sub-task offloading order decision and sub-task offloading location decision, and models the sub-task sorting decision and sub-task offloading location decision respectively. A dependent task offloading decision algorithm based on hierarchical reinforcement learning is proposed, and the two sub-problems are respectively handled by a sorting decision module based on sorting state transition and an offloading location decision module based on queue state transition. A reward feedback mechanism is designed, and the delay gain is used as the feedback connection between the two sub-problems. The sub-task offloading order and the sub-task offloading location strategy are cyclically optimized, thereby reducing the task processing delay.

[0033] refer to Figure 1 , illustrates the process of some embodiments of a hierarchical reinforcement learning-based dependent task offloading decision algorithm according to the present invention. This hierarchical reinforcement learning-based dependent task offloading decision algorithm is implemented through the collaborative work of an offloading order decision module (OO module) and an offloading location decision module (OL module), and specifically may include the following steps:

[0034] In step 1, a task topology structure based on a directed acyclic graph is constructed through the OO module, a subtask sequence decision state transition mechanism is established, and a reinforcement learning algorithm is used to obtain the optimal offloading sequence of subtasks.

[0035] In some embodiments, a task topology structure based on a directed acyclic graph (DAG) can be constructed through the OO module, a subtask sequence decision state transition mechanism can be established, and a reinforcement learning algorithm can be used to obtain the optimal unloading sequence of subtasks.

[0036] Among them, the OO module in step 1 constructs a subtask offloading sequence decision model and adopts the D3QN algorithm to obtain the optimal offloading sequence of subtasks.

[0037] In step 2, a queue state transition mechanism based on time flow is established through the OL module, a subtask unloading location decision model is constructed, and a reinforcement learning algorithm is used to obtain the optimal unloading location of the subtask.

[0038] Among them, the OL module in step 2 constructs a subtask unloading location decision model and adopts the D3QN algorithm to obtain the optimal unloading location of the subtask.

[0039] Step three: establish a linkage feedback mechanism between the OO module and the OL module, pass the unloading order generated by the OO module to the OL module, and the OL module makes unloading location decisions in turn according to the unloading order, and uses the task unloading delay as a feedback link to alternately optimize the decision strategies of the OO module and the OL module.

[0040] In some embodiments, a linkage feedback mechanism can be established between the two modules, and the unloading order generated by the OO module is passed to the OL module. The OL module makes unloading location decisions in turn according to the unloading order, and uses the task unloading delay as a feedback link to alternately optimize the decision strategies of the OO module and the OL module.

[0041] Among them, the linkage feedback mechanism in step three passes the unloading order generated by the OO module to the OL module. The OL module makes subtask unloading location decisions in turn according to the unloading order, and uses the task processing delay as the reward feedback of the OO module, alternately optimizing the decision-making strategies of the OO module and the OL module to obtain the optimal dependent task unloading strategy.

[0042] The technical solution of the present invention is described in detail as follows:

[0043] The present invention designs and implements a hierarchical reinforcement learning-based dependent task offloading decision algorithm HSOTO.

[0044] For multiple users with complex dependent tasks, the HSOTO method first divides the problem into two highly coupled subproblems: deciding the order of subtask offloading and deciding the location of subtask offloading. Then, using a hierarchical reinforcement learning algorithm, it solves both the order and location of subtask offloading, and jointly optimizes the two subproblems. Compared with baseline methods, the HSOTO method effectively reduces task processing latency and improves task completion rates.

[0045] Step 101: Problem description.

[0046] Figure 2 This is a schematic diagram of a dependent task offloading decision scenario. As shown in the figure, the network consists of two layers: the user layer and the edge server layer. The user layer includes numerous users with a certain level of computing power. They can process tasks independently or offload them to edge servers. The edge server layer, composed of multiple edge servers, provides offloading services to users. Assume that each user carries a time-sensitive dependent task, capable of making independent offloading decisions based on the current offloading environment. Edge servers are connected via optical fiber, enabling information sharing. Here, dependent tasks consist of multiple subtasks of different types. These subtasks are time-sensitive, meaning they must complete within a certain timeframe and are executed in a predecessor-successor order.

[0047] Therefore, dependent task offloading decisions refer to users making offloading decisions for each subtask based on the current offloading environment, such as the size and type of the subtask, as well as the load status of the user himself and the edge server, so as to minimize the overall system latency.

[0048] Step 102: Problem analysis.

[0049] like Figure 2 As shown in Figure 1, the dependent task offloading process is divided into two phases: the first is the subtask offloading order decision phase, in which the subtask offloading order is determined based on the task sorting status; the second is the subtask offloading location decision phase, in which the subtask offloading location decisions are made based on the subtask offloading order and the offloading environment status. Therefore, for the dependent task offloading decision problem, corresponding methods need to be proposed to solve these two sub-problems.

[0050] For sub-problem 1, the dependency between subtasks is the core constraint of the problem, which limits the execution order of tasks. This is mainly reflected in the fact that each subsequent subtask can only begin execution after all predecessors have completed and received their output results. Otherwise, additional waiting delays will be generated, affecting overall efficiency. In addition, in the subtask sorting decision problem, the state space and decision space are exploding. To minimize system latency while satisfying dependencies, the sorting module must not only efficiently track the sorting status of each subtask, but also take into account the global sorting effect. Therefore, the subtask sorting decision sub-problem can be regarded as a scheduling problem with dependency constraints. By designing the sorting decision period division, sorting state representation, sorting decision and state transition mechanism in the model, and using reinforcement learning algorithms, the sorting strategy is dynamically adjusted to cope with the ever-changing set of ready subtasks and dependencies.

[0051] For subproblem 2, the dynamically changing offloading environment is the primary constraint, introducing uncertainty into task scheduling and resource allocation. This is primarily manifested in the following: offloading decisions must be made based on subtask attributes, user, and edge server load. During the offloading process, resource occupation leads to uncertainty in the system's available resources, making the offloading environment dynamic. Therefore, the difficulty in deciding subtask offloading locations lies in how to obtain an efficient task offloading location strategy within a highly dynamic offloading environment. Furthermore, queue state transitions are a core modeling challenge, impacting resource availability at the time of migration and the waiting time for subsequent tasks. Therefore, subproblem 2 can be formulated as a dynamic resource scheduling and real-time optimization problem. By constructing an offloading location decision model that incorporates a queue state transition mechanism and combining it with the online interaction advantages of reinforcement learning, continuous awareness of environmental conditions and adaptive migration strategy optimization can be achieved.

[0052] Furthermore, the subtask offloading location decision is constrained by the subtask offloading order, which in turn is influenced by the subtask offloading location decision. Therefore, the two are highly coupled. The challenge lies in obtaining the globally optimal offloading decision for dependent tasks, ensuring that the subtask execution order satisfies the dependencies while dynamically adjusting the offloading order strategy based on the real-time status of the offloading decision. Therefore, the two subtasks can be treated as a joint optimization problem, with mutual feedback and iterative adjustments leading to the globally optimal subtask offloading order and location decisions.

[0053] Step 103: problem modeling.

[0054] The task-dependent task offloading decision problem is modeled, including offloading environment modeling, subtask offloading sequence decision, subtask offloading location decision, task processing delay modeling and problem solving path.

[0055] Step 1031, uninstall the environment.

[0056] For task-dependent task offloading decisions, the entities that need to be considered are tasks, users, and edge servers. The task is the object of the offloading decision, the user is the subject of the decision, and the edge server is the ultimate target of the offloading decision. These three entities are essential factors in the task-dependent task offloading decision. Below, we model the tasks, users, and edge servers separately.

[0057] Step 10311, task.

[0058] A task consists of multiple subtasks. This section will model subtasks and tasks separately.

[0059] Subtask attributes include data volume, latest completion time, task type, and processing density. Data volume refers to the size of the data that needs to be executed or transmitted during the subtask processing process; latest completion time refers to the maximum allowable delay during the subtask processing process; task type refers to the category to which the subtask belongs; and processing density refers to the ratio of the amount of computation required to execute the subtask to the amount of data, measured in CPU cycles.

[0060] Let v i,μ =(μ,λ i,μ ,τ i,μ ,s i,μ ,ρ i,μ ) represents the subtask μ of task i, where μ is the unique identifier of the subtask, and λ i,μ represents the data size of the subtask μ, τ i,μ represents the latest completion time of subtask μ, s i,μ represents the task type of subtask μ, ρ i,μ Indicates the processing density of the subtask.

[0061] A task has three attributes: subtasks, subtask dependencies, and inter-subtask data transfer. A subtask is the smallest execution unit, and it can be offloaded entirely. Subtask dependencies describe the order and data dependencies between subtasks, with predecessor and successor relationships existing between subtasks. Inter-subtask data transfer refers to the amount of data transferred from a subtask to its successor.

[0062] Let G i =(V i ,E i ,W i ) represents complex task i, where V i ={v i,μ |1≤μ≤N i} is the subtask set of task i, N i represents the number of subtasks in task i, v i,μ is the subtask μ of task i; E i ={e<v i,m ,v i,n >|v i,m ,v i,n ∈V i ,m≠n} is the dependency relationship between the subtasks of task i, e<v i,m ,v i,n > indicates subtask v i,n Depends on subtask v i,m , called v i,m It is v i,n The predecessor, v i,n It is v i,m The successor of W i ={w(e<v i,m ,v i,n >)|v i,m ,v i,n ∈V i ,m≠n}, where w(e<v i,m ,v i,n >) indicates subtask v i,m After processing is completed, it needs to be transferred to the subtask v i,n The amount of data. Let V = {V i |1≤i≤N u} represents the task set of all users, where N u Indicates the number of users in the network.

[0063] Step 10312, user.

[0064] A user has the following four attributes: task load, computing power, transmission power, and local queue. A task load is a dependent task consisting of multiple interdependent subtasks. Computing power is the processing performance required when executing a computing task locally. Transmission power is the wireless signal power provided by the user device during data transmission. The local queue is used to store tasks waiting for transmission and execution locally, and the local computing queue is used to store tasks waiting for execution locally.

[0065] Let u i =(G i ,f i ,P i ,Q i ) represents user i, where G i is the complex task i carried by user i, f i is the computing power of user i, measured by the time required to process one CPU cycle, P i is the transmission power of user i, is the queue set of user i, where represents the local transmission queue of user i, represents the kth subtask in the queue, Indicates the length of the local transmission queue, represents the local computing queue of user i, represents the kth subtask in the queue, Indicates the length of the local computing queue. Since the network contains multiple users, let U = {u i |1≤i≤N u} represents the user set, where N u Indicates the number of users in the network.

[0066] Step 10313, edge server.

[0067] Edge servers have three attributes: provided service type, computing capacity, and computing queue. The provided service type is the set of different subtasks that the edge server provides for execution, computing capacity is the performance of the edge server in processing data, and the computing queue is used to store tasks waiting to be executed.

[0068] Orders j =(Φ j ,F j ,Q j ) represents edge server j, where is the set of service types that edge server j can provide, represents the service type r of edge server j, N j F represents the number of service types that edge server j can provide; jrepresents the computing power of edge server j, measured by the time required to process one CPU cycle; is the computing queue set of the edge server, represents the computation queue maintained by edge server j for service type r, represents the kth subtask in the queue, = represents the length of the edge server computing queue. Therefore, all edge server computing queues and service type sets in the system can be represented as Q = {Q j |1≤j≤N e}, S={S j |1≤j≤N e Since the network includes multiple edge servers, let ES = {e j |1≤j≤N e} represents the edge server set, where N e Indicates the number of edge servers.

[0069] Step 1032: Subtask uninstallation order decision.

[0070] The relevant contents of subtask unloading sequence decision include sequence decision period division, sequence decision state and conversion process, sequence decision and unloading sequence.

[0071] Step 10321, sequentially decide time slot division.

[0072] Since the goal of subtask offloading sequence decision-making is to specify the offloading order for each subtask, we divide the sequence decision time slots according to the number of subtasks, and stipulate that each time slot only makes a sequence decision for one subtask. Let T1 denote the sequence decision cycle, and t1∈T1 denote the time slot t1 within it.

[0073] Step 10322, sequential decision state.

[0074] The state of a sequential decision is determined by all subtasks within a task. Therefore, when offloading sequential decisions, the state of all subtasks must be obtained. Since the state transition process is consistent across all subtasks, the state of a single subtask is discussed for this purpose.

[0075] The order decision status of each subtask is constrained by the following two factors: its own sorting status and its predecessor sorting status. The self-sorting status indicates whether the subtask is sorted; the predecessor sorting status indicates whether all predecessors are sorted.

[0076] Order s i,μ (t1)=(done i,μ (t1),state i,μ(t1)) represents the sequential decision status of task i’s subtask μ at the t1th decision, done i,μ (t1) indicates whether the subtask μ of task i is sorted at time slot t1. i,μ When (t1)=0, it means it is not sorted. i,μ When (t1)=1, it means it has been sorted; state i,μ (t1) indicates whether the predecessor subtasks of task i's subtask μ have been sorted at time slot t1. i,μ When (t1)=0, it means that the order is not complete. i,μ When (t1)=1, it means that the order is complete.

[0077] Step 10323, sequential decision making.

[0078] Sequential decision is to select a subtask for sequential decision and determine its priority, that is, the order of unloading the subtask. Let α i (t1) represents the unloading order decision of task i at time slot t1, where α i (t1)∈{1,2,...,N i}. When α i (t1) = μ, indicating that at time slot t1, subtask v i,μ is sorted. Let rank(v i,μ ) represents subtask v i,μ The sorting priority of subtask v i,μ The sorting is completed at time slot t1, then rank(v i,μ )=t1.

[0079] Step 10324, sequential decision state transition.

[0080] Sequential decision state transitions primarily describe the changes in the ordering state of subtasks due to sequential decision making. Depending on whether the subtask itself and its predecessors are sequenced, a subtask can be in one of three states: waiting, ready, and completed. The waiting state indicates that the subtask has not yet been sequenced and its predecessors have not yet been sequenced; the ready state indicates that the subtask has not yet been sequenced and its predecessors have all been sequenced; and the completed state indicates that the subtask has been sequenced and its predecessors have all been sequenced.

[0081] Order s i,μ (t1)=(0,0) represents subtask v i,μ The sequential decision state at time slot t1 is the waiting state, let s i,μ (t1)=(0,1) represents subtask v i,μ The sequential decision state at time slot t1 is the ready state, let s i,μ(t1)=(1,1) represents subtask v i,μ The sequential decision state at time slot t1 is the completed state.

[0082] Figure 3 This is a schematic diagram of the subtask sequence decision state transition. As shown in the figure, there are two changes in the subtask sorting state. One is from the waiting state to the ready state. When the subtask is in the waiting state, if all its predecessors have completed the sorting, then at the next sequence decision, the subtask will be converted from the waiting state to the ready state. The second is from the ready state to the completed state. When the subtask is in the ready state, if it is sorted in the current time slot, then at the next sequence decision, the subtask will be converted from the ready state to the completed state. In addition to the above changes, the subtask sequence decision state remains unchanged. Therefore, the subtask v i,μ The state transition formula is:

[0083]

[0084] Among them, pre(v i,μ ) represents the set of predecessor subtasks of subtask μ.

[0085] Step 10325, uninstall sequence.

[0086] After the subtasks are sorted, the subtask indexes are stored in the uninstallation order according to the subtask priority. i Represents the sorting sequence of tasks. Each order decision slot will obtain the unique identifier of the sorted subtask. Therefore, the update of the offloading order can be expressed as:

[0087] O i =O i ∪{α i (t1)} (2)

[0088] Among them, α i (t1) = μ, μ is the subtask v i,μ Unique identifier.

[0089] Step 1033: Subtask offloading location decision.

[0090] The subtask offloads the relevant content of location decision, including location decision time slot division, location decision state and its state transition, and location decision.

[0091] Step 10331, location decision time slot division.

[0092] The goal of subtask offloading location decision is to determine the processing location for each subtask. Based on the number of subtasks, the system time is divided into several equal time slots. In each time slot, a location decision is made for each subtask according to the order in which the subtasks are offloaded.

[0093] Let T2 represent the location decision cycle, t2∈T2 represent each time slot after division, and the length of each time slot is Δ. In each time slot t2, the subtask to be unloaded is obtained according to the subtask unloading order, and the subtask meets the following conditions: the subtask sorting priority is equal to the current time slot t2, that is, rank(v i,μ )=t2, and the decision completion time of the subtask satisfies t i,μ =t2.

[0094] Step 10332, location decision state.

[0095] The location decision state consists of the following six parts: subtask data size, subtask type, local computing queue state, local transmission queue state, edge server computing queue state, and the set of service types provided by the edge server.

[0096] make represents the location decision state of user i in time slot t2 when making location decision. i (t2) represents the unique identifier of user i making location decision at time slot t2, represents the amount of subtask data required for user i to make location decisions at time slot t2, represents the task type of location decision-making for user i in time slot t2, represents the queue state of user i’s local computing queue at time slot t2, represents the queue state of the local transmission queue of user i at time slot t2, Q(t2) represents the queue state of the computing queue of all edge servers at time slot t2, and Φ represents the set of service types that can be provided by all edge servers.

[0097] Step 10333, location decision.

[0098] The location decision is to choose where to offload the subtask according to the current location decision state. Let β i (t2)=j represents the location decision of time slot t2, where j∈{0,1,2,...,N e When j=0, it means that the subtask is processed locally, otherwise, it means that the subtask is offloaded to the corresponding edge server for processing.

[0099] Step 10334, position decision state transition.

[0100] Position decision state transitions primarily describe changes in the position decision state caused by position decisions. Position decision state transitions include the following two types: changes in pending offloaded subtasks and queue states. Queue state transitions are more complex. The following sections explain queue state changes in detail.

[0101] Figure 4 This is a diagram of queue state transitions. As shown in the figure, queue state transitions include the following three types: subtask arrival, subtask completion, and subtask discard. Subtask arrival indicates that a new subtask has been added to the queue at the current moment, subtask completion indicates that the subtask has been processed within the time interval, and subtask discard indicates that the subtask has been discarded due to a timeout within the current time interval. Therefore, the state of the task queue at the next moment is the current state plus the three aforementioned change factors. The queue state transition formula can be expressed as:

[0102] q(t2+1)=q(t2)+λ(t2)-λ fin (t2)-λ dis (t2) (3)

[0103] Among them, q(t2) represents the queue state of time slot t2, λ(t2) represents the impact of the arrival of the subtask in time slot t2, and λ fin (t2) represents the impact of subtask completion in time slot t2, λ dis (t2) represents the impact of subtask dropping in time slot t2.

[0104] For users, this includes the local computation queue and the local transmission queue. For edge servers, this includes the edge server computation queue. For each of these three queues, the impact of subtask completion and subtask abandonment on the queue is similar; only the impact of subtask arrival differs. The following describes the changes in each of these three queues.

[0105] The newly arrived subtasks in the local computation queue come from the subtasks processed locally. Therefore, the state of the local computation queue of user i in time slot t2+1 ​​can be expressed as:

[0106]

[0107] in, Indicates whether the subtask of time slot t2 is unloaded. When the subtask Uninstall processing, at this time β i (t2)∈{1,2,…,N e};when When the subtask Processed locally, β i (t2)=0.

[0108] The newly arrived subtasks in the local transmission queue come from the offloaded subtasks. Therefore, the state of the local transmission queue of user i in time slot t2+1 ​​can be expressed as:

[0109]

[0110] The newly arrived subtasks in the edge server computation queue come from the subtasks that have been transmitted in the local transmission queue. Therefore, the computation queue state corresponding to the service type r of edge server j in time slot t2+1 ​​can be expressed as:

[0111]

[0112] Therefore, according to the above-mentioned transition of the position decision state, the position decision state at time slot t2+1 ​​can be expressed as

[0113] Step 1034: Subtask processing delay.

[0114] First, the subtask completion time is calculated, then the subtask processing delay is calculated, and finally the total subtask processing delay is calculated.

[0115] Here, let φ i,μ To represent the subtask v i,μ The execution mode, when φ i,μ = 0, indicating local execution; when φ i,μ =1, indicating uninstall execution.

[0116] Step 10341, subtask completion time.

[0117] Figure 5 This is a diagram of the division of subtask processing stages. As shown in the figure, the subtask processing process can be divided into local processing and offload processing. The subtask completion time can be obtained based on the delay generated during the subtask processing and whether it exceeds the latest completion time. Therefore, there are two completion situations for subtasks: one is timeout completion, and the other is on-time completion. When completed on time, the subtask completion time is equal to the estimated completion time; when completed overtime, the subtask completion time is equal to the subtask decision completion time plus the subtask's latest completion time. At this point, the subtask will be discarded due to task processing failure, that is, execution failure. Therefore, subtask v i,μ The completion time can be expressed as:

[0118]

[0119] Among them, EFT i,μ Represents subtask v i,μ The estimated completion time is calculated in two ways: local processing and offload processing, depending on whether the process is offloaded or not.

[0120] When processing locally, the subtask processing process is divided into the following two stages: waiting for execution stage and execution stage.

[0121] make Represents subtask v i,μ The local start execution time of subtask vi,μ The decision completion time plus the local waiting delay; Represents subtask v i,μ The estimated completion time of the local processing is equal to the subtask v i,μ The start execution time plus the local execution delay.

[0122] When offloading, the subtask processing process is divided into the following four stages: the first is the waiting for transmission stage, the second is the transmission stage, the third is the waiting for execution stage, and the fourth is the execution stage.

[0123] make Represents subtask v i,μ The start transmission time of subtask v i,μ The decision completion time plus the transmission waiting delay; Represents subtask v i,μ At the transfer completion time, it is equal to subtask v i,μ The start transmission time plus the transmission delay; Represents subtask v i,μ The start time of the offloading process, which is equal to the subtask v i,μ The transmission completion time plus the transmission delay; Represents subtask v i,μ The estimated completion time of the edge server processing is equal to the subtask v i,μ The start execution time plus the execution delay.

[0124] Therefore, subtask v i,μ The estimated time to complete is:

[0125]

[0126] Step 10342, subtask processing delay.

[0127] Subtask processing delay is the delay caused by task execution, waiting and transmission during task processing. Figure 4 As shown in , the subtask processing delay consists of three parts: execution delay, waiting delay, and transmission delay. Among them, execution delay is the delay caused by the subtask calculation processing; waiting delay is the delay caused by the subtask waiting for calculation or transmission; transmission delay is the delay caused by the subtask transmission or the transmission of its output results. Assume that subtask v i,μ The time when the uninstallation decision is completed is t i,μ .

[0128] Step 103421, execution delay.

[0129] Depending on the execution location of the subtask, there are two ways to calculate its execution delay: local execution and offload execution.

[0130] When executed locally, Represents subtask v i,μ The local execution latency of Indicates the rounding up operation; when offloading to the edge server j for execution, let Represents subtask v i,μ The uninstall execution delay.

[0131] Step 103422, wait delay.

[0132] Depending on the location of the subtask processing, the waiting delay of the subtask includes the local waiting delay and the edge server waiting delay. Here, the present invention assumes that the subtask v i,μ is the kth subtask in the following queue.

[0133] When processing locally, a subtask's waiting latency only includes the time it spends waiting in the local computation queue. This waiting latency is governed by two factors: queue constraints and dependency constraints. Queue constraints mean that after a subtask arrives in the local computation queue, it must wait for the previous subtask in the queue to complete. Dependency constraints mean that the predecessor subtask cannot begin execution until its output has been transmitted locally.

[0134] make Represents subtask v i,μ The local waiting delay can be expressed as:

[0135]

[0136] in, Represents the waiting delay of the local computation queue, which is equal to the execution completion time of the previous subtask in the queue minus the subtask v i,μ The decision completion time, when the local calculation queue is not empty, otherwise represents the delay of waiting for the output of the predecessor subtask, which is equal to the transmission completion time of all predecessor subtask output results minus the subtask v i,μ The maximum decision completion time is

[0137] When offloading execution, the waiting delay of a subtask includes both the waiting delay in the transmission queue and the waiting delay in the edge server's computation queue. A subtask's transmission must wait for the previous subtask in the transmission queue to complete before it can begin. Offloading a subtask requires waiting for the previous subtask in the edge server's computation queue to complete and transmit its output to that location before it can begin.

[0138] make represents the transmission waiting delay, which is equal to the transmission completion time of the previous subtask in the transmission queue minus the subtask v i,μ The decision completion time, when the transmission queue is not empty otherwise

[0139] Similar to the waiting delay in the local queue, the waiting delay in the edge server computing queue is also constrained by the following two factors: queue and dependency. Represents subtask v i,μ The execution waiting delay can be expressed as:

[0140]

[0141] in, Indicates the waiting delay of the subtask in the edge server computing queue, which is equal to the execution completion time of the previous subtask in the queue minus the subtask v i,μ The transmission completion time, when the edge server calculation queue is not empty, otherwise, It represents the delay of waiting for the output result of the predecessor subtask, which is equal to the transmission completion time of all its predecessor subtask output results minus the subtask's v i,μ The transmission completion time is

[0142] Step 103423, transmission delay.

[0143] Depending on the task processing location, transmission latency consists of two components: subtask upload latency and result transmission latency. Subtask upload latency is the latency incurred when transmitting a subtask to the edge server via the wireless network; result transmission latency is the latency incurred when a subtask transmits its computational results to the location of its subsequent subtask.

[0144] When the subtask is uploaded, Represents subtask v i,μ The transmission delay is , where r(i, j) represents the transmission rate between user i and edge server j, using a wireless network model.

[0145] When the result is transmitted, the transmission delay can be calculated in the following two ways: two subtasks are processed locally and on the edge server respectively; and two subtasks are processed locally or on the edge server at the same time. i,q Process locally, subtask v i,μ When processed on edge server j, the transmission delay is:

[0146]

[0147] Among them, w(e<v i,q ,v i,μ >) is the subtask v i,q Transfer to subtask v i,μ When subtasks are processed locally or on edge servers simultaneously, the transmission delay within the same user and the transmission delay between edge servers can be ignored.

[0148] like Figure 5 As shown in the figure, the processing delay of a subtask can be divided into local processing delay and offloading processing delay. The local processing delay includes two parts: calculation waiting delay and execution delay. The offloading delay includes four parts: transmission waiting delay, transmission delay, calculation waiting delay and execution delay. Assume that the subtask μ of task i is offloaded at time t2, that is, O i (t2) = μ, let Represents subtask v i,μ The processing delay can be expressed as:

[0149]

[0150] Task processing cost is the cost incurred due to task delay. Processing cost calculation can be divided into the following two cases: First, if the subtask is completed within the latest completion time (i.e., the subtask processing delay is less than the latest completion time), then the subtask processing cost is equal to the subtask processing delay. Second, if the subtask is not completed within the latest completion time (i.e., the subtask processing delay is greater than the latest completion time), a larger penalty cost will be incurred due to subtask execution failure.

[0151] Let C i,μ (t2) represents subtask v i,μ Therefore, the processing cost of subtask v i,μ The processing cost is:

[0152]

[0153] Where κ represents the penalty cost.

[0154] Step 10343, total task processing delay.

[0155] In summary, the dependent task offloading decision problem P in the multi-user, multi-edge server scenario described in this paper can be modeled as an optimization problem P for minimizing edge computing system latency. That is, under the constraints of edge computing system resources, the goal of minimizing system latency is achieved by optimizing the task offloading strategy. Its mathematical formula can be expressed as:

[0156]

[0157] Among them, C1 to C6 are constraints. C1 is the subtask processing cost constraint, that is, the processing cost of completing the subtask within the latest completion time is equal to the subtask processing delay; otherwise, a penalty cost κ is incurred; C2 is the order decision state, that is, the order decision state of task i is affected by the offloading order state of all its subtasks; C3 is the subtask offloading order decision constraint, that is, during the offloading order decision process, only unsorted subtasks in the task and all predecessor subtasks have completed sorting are selected; C4 is the subtask offloading order constraint, that is, after the subtask sorting decision is completed, the subtask number is stored in the sorting sequence, and the subtask to be offloaded is obtained according to this sequence during the offloading location decision phase; C5 is the offloading environment state, that is, the subtask offloading decision is subject to the subtask attributes, user queue state, edge server queue state, and edge server service type constraints; C6 is the subtask order location constraint, that is, the subtask can only be offloaded to the edge server that can provide service for the subtask.

[0158] Step 104: problem-solving path.

[0159] According to formula (14), problem P is a mixed integer nonlinear optimization problem. In this problem, the subtask offloading order decision variable is a discrete variable, and the order decision problem itself is a nonlinear discrete optimization problem. At the same time, the subtask offloading location decision variable is also a discrete variable. Due to the dynamically changing network environment, the objective function has nonlinear characteristics, and the dependencies between subtasks and different sorting combinations further increase the nonlinearity of the problem. In addition, the two decision variables each have different optimization objectives and variable characteristics, and the subtask offloading location decision is based on the subtask offloading order decision, which in turn affects the subtask offloading location decision. In order to effectively solve this problem, it is necessary to treat it as a whole decision problem and reasonably divide the decision factors so as to find a solution that quickly obtains the optimal solution. Therefore, the present invention decomposes the dependent task offloading problem into two mutually coupled subproblems: the subtask offloading order decision problem (P1) and the subtask offloading location decision problem (P2).

[0160] The subtask offloading order decision problem (P1) can be expressed as:

[0161]

[0162] Among them, C2 to C3 are the constraints of the subtask offloading order decision problem, and the meaning of each constraint is consistent with the meaning of the constraint of problem P.

[0163] The subtask offloading location decision problem (P2) can be expressed as:

[0164]

[0165] Among them, C1, C4, C5 and C6 are the constraints of the subtask offloading location decision problem, and the meaning of each constraint is consistent with the meaning of the constraint of problem P.

[0166] Step 105: Dependent task offloading decision algorithm based on hierarchical reinforcement learning.

[0167] First, the HSOTO algorithm is described as a whole. Then, the unloading sequence decision module is designed to solve the subproblem P1 and the unloading location decision module is designed to solve the subproblem P2. Finally, the pseudo code of the algorithm is given.

[0168] Step 1051, algorithm idea.

[0169] Figure 6 This is the HSOTO algorithm architecture. As shown in the figure, the HSOTO algorithm consists of three modules: the environment module, the offload order decision module (OO), and the offload location decision module (OL). The environment module contains a collection of offload environment information, which is used to receive actions from the OO module and the OL module and simulate the environmental changes after the actions are executed. The OO module obtains the offload order of the subtask by obtaining the sequence decision status from the environment module. The OL module obtains the offload location decision of the subtask by obtaining the location decision status from the environment module.

[0170] The three modules above work together, interacting and feeding back to each other to jointly construct a solution mechanism for highly coupled problems. The coordinated data flow is as follows: First, the OO module obtains the sequential decision state from the environment module, makes an unloading location decision based on the sequential decision state, and interacts with the environment module to change the subtask unloading sequence state. Ultimately, it obtains the unloading order of the subtasks and inputs this unloading order into the OL module. Subsequently, the OL module obtains the subtasks to be unloaded in sequence according to the unloading order of the subtasks, obtains the current local queue and edge server queue state from the environment module, obtains the unloading location decision of the subtask, and interacts with the environment module to change the location decision state. Finally, the task processing delay generated during the interaction between the OL module and the environment module is input into the OO module as the module's delay reward to further optimize the subtask unloading sequence decision strategy.

[0171] Step 1052: Uninstall the sequence decision module.

[0172] Introduce the state space, action and reward of the OO module respectively.

[0173] Step 10521, state space.

[0174] The state space of the OO module includes the unloading order status of all subtasks. The sorting status of a subtask indicates whether the subtask has been sorted and its predecessor sorting status. Represents the state space of user i in time slot t1. After executing the sequential decision action in time slot t1, the state space of the subtask unloading sequential decision phase changes from becomes in, Indicates that after the action is executed, the sorting status of each subtask will be updated.

[0175] Step 10522, action.

[0176] The OO module's action is to select a subtask for sequencing in each time slot and determine its unloading order. Once the subtask is sequenced, its sequencing status will change, and the sequencing status of its successor subtasks will be updated at the same time. If all of its predecessor subtasks have been sequenced, the sequencing status of its predecessor subtasks will be updated. Otherwise, the sequencing status of the subtask will not change. Represents a sequential decision action, i.e. user i selects a subtask at time slot t1 Sort by,

[0177] Step 10522, reward.

[0178] The reward of the OO module consists of the following two parts: order decision reward and delay reward. Among them, the order decision reward is used to evaluate whether the selected subtask meets the dependency constraint conditions; the delay reward is used to evaluate the task processing delay corresponding to the offloading order. represents the sequential decision reward, which can be expressed as:

[0179]

[0180] in, Indicates that the sorted subtasks satisfy the dependency constraints and obtain a positive reward x; Indicates that the sorted subtasks do not meet the dependency constraints and receive a large negative reward. Denotes the delayed reward. Indicates that actions are executed during the subtask unloading order decision phase The reward obtained. Therefore, the reward function of the OO module can be expressed as:

[0181]

[0182] Here, γ represents the discount factor.

[0183] In order to accelerate convergence, the delay reward does not need to be calculated in every time slot. In the early stage of the sequence decision, since the algorithm is in the exploration stage, the obtained unloading order may not fully meet the constraints of the dependent task unloading. At this time, unloading according to the obtained unloading order will produce a large system delay, which has little impact on the exploration of the optimal unloading order strategy. Therefore, in the early stage of the subtask unloading order decision, only the sequence decision reward is used for optimization. When the sequence decision reward tends to be stable, it indicates that the current unloading order can meet the dependent constraints. At this time, the delay reward is added to the sequence decision reward to further optimize the unloading order decision strategy and obtain an unloading order that is suitable for the current environment. In order to judge whether the unloading order reward tends to be stable, the present invention uses a sliding reward variance to evaluate whether the unloading order reward tends to be stable, so as to determine whether to enable the task unloading decision module to optimize the unloading order strategy. Let ω represent the variable of whether the unloading order reward converges. When ω = 1, it means that it has converged, otherwise, ω = 0.

[0184] Step 1053: Uninstall the location decision module.

[0185] Introduce the state space, action and reward of the OL module respectively.

[0186] Step 10531, state space.

[0187] The state space of the OL module includes subtask attributes, local queue state and edge server state. Represents the state space of user i in time slot t2 in this module. After executing the unloading location decision action in time slot t2, the state space changes from becomes in, Indicates that an action is being performed After that, the subtasks to be offloaded, the user's local computing queue status, the local transmission queue status, and the edge server computing queue status will be updated.

[0188] Step 10532, action.

[0189] The action of the OL module is to choose whether to process the subtask locally or on the edge server in each time slot. represents the location decision action, indicating where user i offloads the subtask to be processed in time slot t2, where when , it means the subtask is processed locally, otherwise, Indicates that the subtask is processed on the edge server.

[0190] Step 10533, reward.

[0191] The reward of the OL module is used to evaluate the gain of the system delay caused by the user's interaction with the environment after performing the location decision action. Indicates that in the subtask unloading location decision stage based on the observation value User i performs an action The reward obtained. The optimization goal of this invention is to minimize the edge computing system latency, while the goal of reinforcement learning is to obtain the maximum reward, setting the reward to the inverse of the latency. Therefore, the reward function of user i in the subtask unloading location decision stage is defined as follows:

[0192]

[0193] Where γ represents the discount factor,

[0194] Step 1054, pseudo code.

[0195] Algorithm 1, shown in Table 1, is the pseudocode for the HSOTO algorithm. Algorithm 1 allows users to obtain the offloading order and location decisions of dependent subtasks based on the edge computing system environment, and also obtains the system latency resulting from the task offloading decision. The algorithm's inputs include the subtask offloading order decision status and the subtask offloading location decision status, and its outputs include the subtask offloading order, subtask offloading location decision, and task processing latency. The specific implementation of this algorithm is as follows: First, all users in the user set are traversed (line 1). Then, based on the task offloading order status, Algorithm 2 is called to obtain the subtask offloading order (line 2). Subsequently, the subtask offloading order convergence variable is obtained to determine whether the offloading order decision module has reached convergence, thereby determining whether to enable the location decision module to further optimize the subtask offloading order strategy (lines 3-10). Finally, if the offloading order decision module has reached convergence, Algorithm 2 is called, based on the offloading order decision and the task offloading decision status, to obtain the task offloading location decision and its processing latency (lines 11-14).

[0196] Table 1

[0197]

[0198]

[0199] Algorithm 2, shown in Table 2, is the pseudocode for the D3QN algorithm. Its primary function is to optimize a dependent task offloading decision strategy. This algorithm utilizes two networks to evaluate action values ​​and state values, respectively, to more accurately calculate the expected reward of the task offloading strategy. The specific implementation of this algorithm is as follows: First, at each time step in each round, an action is acquired and executed to obtain the reward and the state for the next time step, and the experience for the current time step is stored in the experience pool (lines 2-6). Then, a set of experiences is randomly sampled from the experience pool, and the target Q value for each experience is calculated (lines 7-12). Subsequently, the evaluation network parameters are updated by minimizing the target Q value (line 13). Finally, the evaluation network parameters are assigned to the target network parameters every round (lines 14-16).

[0200] Table 2

[0201]

[0202]

[0203] The specific implementation is as follows:

[0204] The following mainly introduces the experimental environment configuration, experimental results and discussion of the present invention.

[0205] Step 201, experimental environment configuration, mainly introduces the configuration environment, data set, baseline algorithm and evaluation indicators of the experiment.

[0206] Step 2011, configure the environment.

[0207] The experiments of the present invention were conducted in Python 3.6 and Tensorflow 1.4.0. The experimental scenarios include: the number of edge servers N = 5, the number of users M = 30, the users are evenly distributed throughout the entire area, each user generates a task, and the task consists of 5-25 subtasks with different task types. The number of task types is 3, and each edge server can handle a different number of task types. The neural network parameters are set as follows: the batch size is set to 32, the learning rate is 0.001, the discount factor is 0.9, the random exploration probability is initially 1 and gradually decreases to 0.01, and the RMSProp optimizer is used.

[0208] Step 2012, data set.

[0209] The experiments used the cluster-trace-v2018 cluster tracing dataset. This dataset, released in 2018, contains a production cluster system with 4,034 machines running jobs for eight days. The dataset provides DAG information for the actual workload. This work used the DAG information from the original dataset. Because the cluster tracing dataset lacked task type information, we added a task type attribute to each task. We then calculated the task's computational load by subtracting the start time from the end time and multiplying it by the planned CPU value, then performing appropriate scaling.

[0210] Step 2013, baseline algorithm.

[0211] In order to evaluate the performance of the HSOTO algorithm, the present invention selects the following baseline algorithms for comparison:

[0212] Local: Local execution means that all tasks are executed on the mobile device.

[0213] Random: Random offloading, that is, randomly selecting a subtask and offloading it to a randomly selected edge server for execution.

[0214] HEFT: A heuristic algorithm based on priority sorting that aims to minimize task completion time. The algorithm first calculates the priority of subtasks based on their average completion time. At each step, HEFT selects the subtask with the highest priority and then chooses the local or edge server processing task with the smallest estimated completion time.

[0215] Sort+DQN: A phased reinforcement learning algorithm. Sort first sorts subtasks according to their feasibility, and then DQN makes offloading decisions based on the sorting results.

[0216] Step 2014, evaluation indicators.

[0217] The experiment uses the following two indicators to evaluate the algorithm performance, namely system latency and task completion rate.

[0218] System latency is the sum of the processing latency of all subtasks of all user tasks in the edge computing system. Therefore, system latency can be expressed as:

[0219]

[0220] The task completion rate is the ratio of the number of subtasks that the edge computing system completes within the latest completion time to the total number of subtasks. Therefore, the task completion rate can be expressed as:

[0221]

[0222] Among them, N c It indicates the number of subtasks that are completed within the latest completion time, and N indicates the total number of subtasks.

[0223] Step 202: Results and discussion.

[0224] Step 2021, system delay.

[0225] Figure 7 This chart shows the changing trend of system latency under different influencing factors. When one of the influencing factors changes, the remaining factors take the default values. Figure 7 (a) is a trend chart showing how system latency changes with the number of users. As shown in the figure, as the number of users increases, the system latency gradually increases. This is because as the number of users increases, the number of tasks that need to be processed in the system also increases, which in turn leads to an increase in system latency. Figure 7 (b) Trend graph of system latency as the number of edge servers changes. As shown in the figure, as the number of edge servers increases, the system latency gradually decreases. This is because as the number of edge servers increases, the computing resources in the system increase, and tasks in the system can be processed faster. As shown in the figure, the system latency obtained by the HSOTO algorithm is better than other baseline algorithms. Among them, HSOTO is the best, followed by Sort+DQN, HEFT, Random, and Local. The system latency obtained by the Sort+DQN algorithm is also better than the other baseline algorithms. The HSOTO algorithm is better than the Sort+DQN algorithm. This shows that the HSOTO algorithm's hierarchical optimization method can further reduce system latency, further demonstrating the effectiveness of the subtask offloading order decision module.

[0226] Step 2022, task completion rate.

[0227] Figure 8 This is a graph showing the changing trend of task completion rate under different influencing factors. When one of the influencing factors changes, the remaining factors take the default value. Figure 8 (a) is a trend chart showing the task completion rate as the number of users changes. As shown in the figure, the task completion rate gradually decreases as the number of users increases. This is because as the number of users increases, the number of tasks that need to be processed in the system also increases, and the limited computing resources in the system lead to a decrease in the task completion rate. Figure 8(b) is a trend chart showing the task completion rate as the number of edge servers changes. As shown in the figure, the task completion rate gradually increases with the increase in the number of edge servers. This is because as the number of edge servers increases, the computing resources in the system increase, and more tasks can be processed within the latest completion time. As shown in the figure, the task completion rate achieved by the HSOTO algorithm is better than that of other baseline algorithms. Among them, HSOTO is the best, followed by Sort+DQN, HEFT, Random, and Local. The task completion rate achieved by the Sort+DQN algorithm is also better than that of the other baseline algorithms. The HSOTO algorithm is better than the Sort+DQN algorithm, which shows that the HSOTO algorithm's hierarchical optimization method can further improve the task completion rate, further demonstrating the effectiveness of the OO module.

[0228] In summary, the present invention obtains the optimal dependent task offloading strategy to minimize the task processing delay. The three main tasks are as follows: First, the divide-and-conquer strategy is adopted to divide the dependent task offloading decision problem into the subtask offloading order decision subproblem and the subtask offloading location decision subproblem; second, the two subproblems are modeled as Markov decision processes, and the unloading order decision module and the unloading location decision module based on reinforcement learning are designed to solve the two subproblems respectively; third, a reward feedback mechanism is adopted to establish feedback interaction between the two submodules, continuously optimize the dependent task offloading decision strategy, and accelerate the decision process. An efficient dependent task offloading strategy that minimizes task processing delay is obtained, and a system delay and task completion rate that are better than the baseline algorithm are obtained. The system delay of the HSOTO algorithm increases with the increase in the number of users, decreases with the increase in the number of edge servers, and is always better than other baseline algorithms. The task completion rate of the HSOTO algorithm decreases with the increase in the number of users, increases with the increase in the number of edge servers, and is always better than other baseline algorithms.

[0229] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A dependent task offloading decision algorithm based on hierarchical reinforcement learning, characterized in that: The following steps are involved: Step 1: Use the OO module to build a task topology structure based on a directed acyclic graph, establish a subtask sequence decision state transition mechanism, and use a reinforcement learning algorithm to obtain the optimal offloading order of subtasks; Step 2: Establish a queue state transition mechanism based on time flow through the OL module, build a subtask unloading location decision model, and use the reinforcement learning algorithm to obtain the optimal unloading location of the subtask; Step three: establish a linkage feedback mechanism between the OO module and the OL module, pass the unloading order generated by the OO module to the OL module, and the OL module makes unloading location decisions in turn according to the unloading order, and uses the task unloading delay as a feedback link to alternately optimize the decision strategies of the OO module and the OL module.

2. The dependent task offloading decision algorithm based on hierarchical reinforcement learning according to claim 1 is characterized in that: The OO module in step 1 constructs a subtask offloading order decision model and uses the D3QN algorithm to obtain the optimal offloading order of subtasks.

3. The dependent task offloading decision algorithm based on hierarchical reinforcement learning according to claim 1 is characterized in that: The OL module in step 2 constructs a subtask unloading location decision model and adopts the D3QN algorithm to obtain the optimal unloading location of the subtask.

4. The dependent task offloading decision algorithm based on hierarchical reinforcement learning according to claim 1 is characterized in that: In step three, the linkage feedback mechanism passes the unloading order generated by the OO module to the OL module. The OL module makes subtask unloading location decisions in turn according to the unloading order, and uses the task processing delay as the reward feedback of the OO module, alternately optimizing the decision strategies of the OO module and the OL module to obtain the optimal dependent task unloading strategy.