Dynamic Task Replication Method, Device and System in Edge Computing Environment
By applying the multi-arm gambling machine algorithm in an edge computing environment, dynamically selecting edge clusters for task replication, solving the problem of unpredictability of task completion in edge computing and significantly improving task replication performance.
Patent Information
- Application Number
- CN202111437730.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-29
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-11-29
AI Technical Summary
In an edge computing environment, existing task replication methods are difficult to effectively solve the unpredictability of task completion caused by unknown computing delays and uncertain network bandwidth.
Using a task replication decision algorithm based on a multi-arm gambling machine, by estimating the task calculation amount and calculating the confidence lower limit, the best edge cluster is dynamically selected for task replication to minimize the regret of the total completion time.
The performance of the task replication mechanism is improved, the average job completion time is reduced, and the completion time is reduced by 56.4% to 77.6% compared with other algorithms.
Smart Images

Figure CN114090218B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of edge computing, and particularly to a dynamic task replication method, device, and system in an edge computing environment. Background Art
[0002] With the development of edge computing, the data generated at the network edge has grown exponentially. It is expected that in the near future, the generation rate of edge cluster data will exceed the capacity of today's Internet. With the increase in the data aggregated at the edge and the rapid development of machine learning, machine learning tasks have become the main workload of edge systems. However, the limited resources of each edge cluster make it challenging to run machine learning tasks. As is well known, the completion of a job usually depends on its slowest task, i.e., the straggler. The traditional method to avoid stragglers is to offload tasks to a remote cloud, which results in huge wide area network latency and capital costs. Another promising alternative is to replicate tasks from overloaded edges to idle edges: when any one of the replicas completes, the task is completed. That is to say, the completion of the task depends on its fastest replica, which may reduce the task queue and computing latency.
[0003] However, there are the following challenges in achieving efficient task replication in edge clusters. First, to select the best replica location, it is necessary to know in advance the computing latency of the tasks running in the edge cluster, but such information cannot be known before making a replication decision and completing the replica. Second, the network resources between edges are usually time-varying, so the bandwidth is uncertain, which leads to uncertain transmission latency. These two intertwined challenges further make the completion of tasks unpredictable. Therefore, it is not easy to design an efficient replication algorithm that can continuously adapt to this dynamic and uncertain environment.
[0004] Existing replication methods cannot cope with these challenges. Detection-based algorithms require a large amount of time and cost to monitor and identify stragglers. Usually, such overhead is huge, so detection-based strategies have their inherent defects. Clone-based algorithms replicate a certain number of copies of tasks in advance and offload them to the corresponding edges. However, the latency is always unknown before executing the algorithm, so it is impossible to find the best edges to offload these replicas. Summary of the Invention
[0005] The object of the present invention is to propose a dynamic task replication method, device, and system in an edge computing environment to solve the problems existing in the existing task replication mechanism.
[0006] To achieve the above object of the invention, the following technical solutions are adopted in the present invention:
[0007] In a first aspect, a dynamic task replication method in an edge computing environment is proposed, including the following steps:
[0008] An optimization problem is established with the goal of minimizing the regret, which is the difference between the total completion time of operations in the edge environment and the total latency of operation completion under the ideal optimal replication decision.
[0009] The task replication decision algorithm based on the multi-armed bandit is used to solve the optimization problem, including:
[0010] At the beginning of the first time slot, estimate the task computation amount w according to the task type of the task and the size of the input data t ;
[0011] For each task t, calculate the lower confidence bound of the latency for replicating task t from edge cluster i to edge cluster j According to the lower confidence bound Determine all available edge clusters, and select r t ones smaller available edge clusters as the target edge clusters, and replicate the task to all target edge clusters for execution.
[0012] Furthermore, regret is calculated according to the following formula:
[0013]
[0014] delay a represents the completion time of job a, and ∑ a∈J delay a represents the total completion time of all jobs in the system. represents the theoretical best latency of job a. represents the theoretical best latency of all jobs in the system, and J represents the set composed of all jobs.
[0015] Furthermore, estimate the task computation amount w according to the task type of the task and the size of the input data t including: representing the amount of data to be processed by the dimension N of the input data vector, and obtaining the machine learning model structure type z of the task itself t , t and obtaining the task computation amount using the estimation function based on N and z.
[0016] Furthermore, the lower confidence bound of the latency for replicating task t from edge cluster i to edge cluster j is calculated as follows:
[0017]
[0018] x t represents the size of the input data of task t; y t represents the size of the output data of task t; Indicates the number of times the link from edge cluster i to edge cluster j is sampled when task t is completed; Indicates the number of times the link from edge cluster j to edge cluster i is sampled when task t is completed; Indicates the number of times edge cluster j is selected as the target edge cluster when task t is completed; b i,j Indicates the bandwidth coefficient from edge cluster i to edge cluster j, b j,i Indicates the bandwidth coefficient from edge cluster j to edge cluster i, f j Indicates the computing power coefficient of edge cluster j; Respectively indicate after the execution of task t b i,j 、b j,i And f j The lower confidence limit, and the calculation formula is as follows:
[0019]
[0020]
[0021]
[0022] Respectively indicate b i,j After being sampled The average value after times, b j,i After being sampled The average value after times, f j After being sampled The average value after times.
[0023] In the second aspect, a dynamic task replication device in an edge computing environment is proposed, including:
[0024] An optimization problem construction module, which is used to establish an optimization problem with the goal of minimizing the regret, which is the difference between the total completion time of the jobs in the edge environment and the total delay of the job completion under the ideal optimal replication decision;
[0025] An optimization problem solving module, which is used to solve the optimization problem by using a task replication decision algorithm based on the multi-armed bandit. The solution of the optimization problem includes:
[0026] At the beginning of the first time slot, estimate the task computation amount w according to the task type and the size of the input data of the task t ;
[0027] For each task t, calculate the lower confidence limit of the delay of copying task t from edge cluster i to edge cluster j According to the lower confidence limit Determine all available edge clusters, and select r t Individual The smaller available edge clusters are used as target edge clusters, and the tasks are copied to all target edge clusters for execution.
[0028] Furthermore, regret is calculated according to the following formula:
[0029]
[0030] delay a represents the completion time of job a, ∑ a∈J delay a represents the total completion time of all jobs in the system. represents the theoretical optimal latency of job a. represents the theoretical optimal latency of all jobs in the system, and J represents the set composed of all jobs.
[0031] Furthermore, the task computation amount w is estimated according to the task type of the task and the size of the input data. t It includes: representing the amount of data to be processed by the dimension N of the input data vector, and obtaining the machine learning model structure type z of the task itself. t , and using the estimation function based on N and z t to obtain the task computation amount.
[0032] Furthermore, the lower confidence limit of the latency for copying task t from edge cluster i to edge cluster j is calculated as follows:
[0033]
[0034] x t represents the size of the input data of task t; y t represents the size of the output data of task t; represents the number of times the link from edge cluster i to edge cluster j is sampled when task t is completed; represents the number of times the link from edge cluster j to edge cluster i is sampled when task t is completed; represents the number of times edge cluster j is selected as the target edge cluster when task t is completed; b i,j represents the bandwidth coefficient from edge cluster i to edge cluster j, b j,i represents the bandwidth coefficient from edge cluster j to edge cluster i, f j represents the computing power coefficient of edge cluster j; respectively represent the lower confidence limits of b i,j 、b j,i and f j after task t is executed, and the calculation formula is as follows:
[0035]
[0036]
[0037]
[0038] respectively represent b i,j after being sampled times, the average value of b j,i after being sampled times, the average value of b j after being sampled times, the average value.
[0039] In a third aspect, a computing device is proposed, including:
[0040] one or more processors;
[0041] a memory; and
[0042] one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and when the program is executed by the processor, it implements the dynamic task replication method in the edge environment as described in the first aspect of the present invention.
[0043] In a fourth aspect, a dynamic task replication system in an edge computing environment is proposed, including: at least one control node and several edge computing clusters. The control node is interconnected with the edge computing clusters and between each edge computing cluster via a network. The edge cluster feeds back its computing power and bandwidth status at the end of each time slot to the control node. The overloaded edge cluster timely transmits the relevant information of the task to be replicated to the control node. The control node makes a replication decision for the overloaded edge cluster by using the dynamic task replication method in the edge environment as described in the first aspect of the present invention and issues the decision to the edge cluster.
[0044] Compared with the prior art, the present invention has the following beneficial effects: For the first time, an algorithm based on the multi-armed bandit is applied to the task replication problem in an edge computing system. Previous work usually copies tasks from overloaded edges to idle edges to reduce queuing and computing delays by exchanging transmission delays. However, before making a replication decision, it is impossible to predict the completion delay of tasks replicated to different edges, which will affect the performance of the task replication mechanism. Therefore, the multi-armed bandit is applied to the task replication problem, and the bandwidth and edge computing capabilities between random variable edges are described. The proposed online task replication decision mechanism based on the multi-armed bandit model is superior to the existing technologies in terms of performance. The results show that the average job completion time according to the method of the present invention is reduced by 56.4% and 77.6% respectively compared with the "single offloading" and the "random algorithm". Description of the Drawings
[0045] Figure 1 Schematic diagram of the task replication model structure according to an embodiment of the present invention;
[0046] Figure 2 Schematic diagram of the task replication decision-making system structure according to an embodiment of the present invention;
[0047] Figure 3 Variation of regret after applying different task replication methods;
[0048] Figure 4 Variation of the average job completion delay after applying different task replication methods;
[0049] Figure 5 Average job completion delay at different task skewness levels after applying different task replication methods. Specific implementation manner
[0050] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0051] The biggest feature of executing dynamic task replication in the edge computing environment is randomness. Due to the continuous fluctuations of network bandwidth and edge cluster computing performance during task execution, it is impossible to predict the completion delay of tasks replicated to different edges before replication decision. Moreover, the total delay after decision is still a random quantity. The actual completion delay of each task can only be known after the task is actually completed. Therefore, the task completion delay is an unknown distribution. Specifically, in the task replication delay model, the bandwidth and edge computing capabilities satisfy unknown distributions and fluctuate over time, and cannot be predicted in advance. This uncertainty conforms to the multi-armed bandit model. Therefore, the multi-armed bandit model is used in the present invention to solve the randomness problem of replication.
[0052] The present invention regards the entire edge computing system as a multi-armed bandit, and each edge cluster is regarded as an arm, and an online target arm set (target edge cluster set) is selected for tasks on the hot edge cluster. The hot edge cluster is also the overloaded edge cluster. For task t on edge cluster i, r t copies of task t are replicated and transmitted to different edge clusters selected according to the method of the present invention. This process can be regarded as selecting r t arms from the current available edge cluster set of task t as the target edge clusters for task replication. Figure 1It shows how a job with three remaining tasks determines the optimal set of edges for executing replicas. Edge cluster 2 belongs to the hot edges, and the tasks on it (which may belong to certain jobs) need to be replicated. For the replication decision of task 3, the replication method of the present invention selects edge cluster 4 and edge cluster 7 as the target edge clusters, and replicates task 3 twice and sends them to these two edges for execution respectively. It can be seen that task 3 refuses to select edge 9 to execute the replicated replica. The replicated replica at edge 4 finishes first and returns the result, while the replicated replica at edge 7 has not yet returned its result. At this time, task 3 has been executed, and the replica of task 3 on edge 7 does not need to continue execution. It can also be seen from the figure that the computing power allocated to the replicas is uncertain, and the two-way bandwidth is also uncertain.
[0053] Task 3 makes a replication decision to select edge cluster 4 and edge cluster 7 as the target edge clusters, which is implemented based on the online task replication method of the multi-armed bandit. The implementation process of the dynamic task replication method in the edge computing environment proposed by the present invention is described in detail below.
[0054] First, an explanation of jobs and tasks is given. A job contains multiple tasks, and a task t is composed of a triple (x t , y t , z t ), which are x t representing the input data volume of the task, y t representing the output data volume of the task, and z t representing the type of the task. From the task type and the input data volume, the computing volume w t of the task can be obtained. In the embodiment of the present invention, J represents the set of all jobs, and K represents the set of all edge clusters, then K t represents the set of available edge clusters corresponding to task t.
[0055] For the convenience of description, in the following text, an edge cluster is sometimes also referred to as an edge, that is, edge i and edge cluster i have exactly the same meaning. In addition, edge group and edge cluster, and edge computing node can be used interchangeably.
[0056] The task replication delay from edge cluster i to j consists of the following parts: a) the delay d rep for sending the task from edge i to edge j; b) the computing delay d com of edge j; and c) the delay d ret for sending the result back from edge j to edge i. The d rep of the task depends on its input data size x t and the bandwidth trans i,j from edge cluster i to j. Therefore, there is Similarly, yt Indicates the size of the output data, trans j,i Indicates the bandwidth from edge cluster j to i. The computation delay depends on the amount of computation w required for the task t and the computing power com of edge j j . Thus, there is Therefore, the total replication delay for copying task t from edge cluster i to j is d t,i,j = d rep + d com + d ret . Denote as b i,j , and denote as b j,i , and denote as f j . From the sampling of the transmission delay and the input data volume, b i,j , b j,i and f j can be obtained. For example, and Similarly, it can be deduced that b i,j , b j,i and f j are sampled for each computation. Thus, the delay model for task replication is determined.
[0057] According to the task replication method of the present invention, its overall goal is to minimize the regret of the total completion time of all jobs in the edge system. Regret refers to the difference between the total delay of job completion and the total delay of job completion under the ideal optimal replication decision. For each task t, the set of target edge clusters π t is used as the set of replication decisions. Thus, there is |π t | = r t . Then, for each set of replication decisions π t generated by the algorithm, the actual delay is where i is the edge cluster that generates task t. A job consists of multiple different tasks, and the completion time of a job depends on the completion time of its slowest task. Therefore, the completion time of job a can be defined as delay a = max t∈a (d t ). Then, the total completion time of all jobs in the system is ∑ a∈J delay a .
[0058] Use the delay to represent the theoretical best delay of job a. Therefore, the theoretical optimal value of the total completion time of all jobs is Therefore, the regret based on the latency of the multi-armed bandit-based task replication system can be defined as:
[0059]
[0060] Furthermore, the following optimization problem is established:
[0061] Optimization objective:
[0062] Constraints:
[0063] (1) The set of target edge clusters for each replication decision belongs to the available target edge set of task t, and then belongs to the set composed of all edges:
[0064] (2) The size of the set of target edge clusters for each replication decision is equal to r t :
[0065] In the formula, J represents the set composed of all jobs, K represents the set composed of all edge clusters, and K t represents the available edge set of task t. The edge clusters determined to have no faults through heartbeat detection are the elements in K t . π t represents the set of target edge clusters for the replication decision of task t, which can be abbreviated as the replication decision set. The completion time of job a can be defined as delay a , and at the same time, the latency is used to represent the theoretical best latency of job a. Constraint (2) ensures that r t target edge clusters are selected from the currently available edge clusters to transmit r t copies of each machine learning task t.
[0066] The present invention uses a multi-armed bandit model to solve this optimization problem to address randomness. The specific solution process is given below.
[0067] First is the estimation of the task computation amount. To correctly estimate the completion delay of machine learning tasks, it is necessary to estimate the computation amount of each task. Generally speaking, the computation amount of machine learning tasks mainly depends on their model structures. Common machine learning model structures include linear regression models, clustering models, or probabilistic graphical models, etc. Common loss functions (measuring the difference between the inferred value and the actual value) include 2-norm loss, exponential loss, etc. Other complex models are usually derived from combinations, splicing, or modifications of these basic models. Therefore, the computation amount w can be estimated by analyzing the model type of the task and the size of the task input data t .
[0068] For inference tasks, the total amount of computation can be estimated based on the model structure by counting the number of atomic operations during the calculation process. Table 1 summarizes examples of loss functions for popular machine learning models. For example, for an N-dimensional linear regression model of the form y = w T x + b, the task calculation involves an N-dimensional vector multiplication and an addition. Since each dimension is a numerical value with a uniform precision, the dimension N of the input data vector actually represents the amount of data to be processed. The present invention constructs a function w t = A(N, z t ) to estimate the amount of computation for inference tasks. A has different specific implementations according to different model structures, and Table 1 summarizes the loss functions of some popular models. Based on the input dimension N and the model structure type z t , the amount of computation can be quickly estimated.
[0069] Table 1 Examples of loss functions for popular machine learning models
[0070]
[0071] For training tasks, if the optimal model parameter vector w * can be directly obtained during the training process. For example, w * = Θ(x, y, Φ) (where x is the input vector, y is the real value, and Φ is the loss function). Through the analysis of the closed-form expression, the relationship between the amount of computation and the size of the input data can be directly obtained. For most training tasks that require iterative updates, one iteration can be expressed as a closed-form expression w k+1 = Ω(w k , x, y, Φ) (k is the number of iteration rounds). Therefore, based on the dimension N of the input vector, the amount of computation for one iteration can be accurately estimated. The present invention adopts a method combining theory and experiment, and constructs a prediction function B(N, z t ) for the amount of computation of training tasks through the combination of the aforementioned closed-form expression and experiment. The specific expression of B varies according to the closed-form expressions derived from different models.
[0072] As described above, according to the task replication latency model, to estimate the latency of task t copied from edge i to edge j, in addition to the amount of computation w t of task t, it is also necessary to accurately estimate the random variables b i,j , b j,i and f jDue to insufficient bandwidth observation information and computing power at the target edge, the system faces a trade-off between exploration and exploitation. On the one hand, to accurately estimate the average latency of copying a task to each edge cluster, the system needs to try copying the task to different edge clusters; on the other hand, to minimize regret, the system tends to copy the task to the edge cluster with the minimum latency. For classical multi-armed bandits, a typical algorithm to solve the trade-off between exploration and exploitation is the Upper Confidence Bound (UCB) method. In the scenario of the present invention, the core idea is to maintain a lower confidence bound for the latency d of copying task t from each edge i to each edge j t,i,j Maintain a lower confidence bound and ensure that with a high probability, for example, higher than a specified probability value. Then, the algorithm makes a trade-off between exploration and exploitation by selecting multiple edges with the minimum lower confidence bound for task copying
[0073] To maintain the lower confidence bound of each parameter, the sample mean and the number of samplings of each parameter need to be maintained. Let denote the number of times the link from edge i to edge j has been sampled when the algorithm completes the t-th task. Define in the same way to denote the number of times edge j has been selected as the target edge when the t-th task is completed. For the sample average, after b i,j has been sampled times, use to denote the average value of b i,j . Define " in the same way to denote the average value of f j after it has been sampled times. Since the bandwidth and computing power of each edge are independent and identically distributed, the lower confidence bound can be constructed using concentration inequalities. Therefore, after task t is executed, the lower confidence bounds i,j of b j,i , b j and f and are as follows
[0074]
[0075]
[0076]
[0077] Therefore, according to the above three formulas and the relevant information x t , y t and w t of task t, the lower confidence bound of the latency of copying task t from edge i to edge j can be obtained as follows
[0078]
[0079] Furthermore, the sampling and calculation of
[0080]
[0081]
[0082]
[0083] After the task t is completed, a sampling is performed. Each sampling result is necessarily different. Among them, represents the value of b sampled by the system at the moment when the t-th task is completed; i,j value; Similarly. Since b i,j , b j,i and f j are random variables, these sampled samples are used to estimate the probability distribution functions of these three random samples.
[0084] Furthermore, the execution steps of the task replication method based on the multi-armed bandit include:
[0085] At the beginning of the time slot, determine the computational amount of the task according to the input data volume and task type of the task;
[0086] After obtaining the computational amounts of the subtasks of all jobs on each edge computing node within the current time slot, run the task replication algorithm based on the multi-armed bandit:
[0087] First, based on the optimistic initial value method, are all assigned the value of 0;
[0088] Then enter the continuous learning stage. For each task t, calculate the Select r t ones with smaller available edges as the target edge set, and then replicate the task to all the target edges for execution;
[0089] Subsequently, sample and calculate according to formulas (7), (8), and (9)
[0090] Repeat this process until all tasks of all jobs are completed.
[0091] At the beginning of the subsequent time slot, there is no need to assign the initial value again. Just run the algorithm in the continuous learning stage directly.
[0092] Refer to Figure 2, in one embodiment, the online task replication decision-making system based on multi-armed bandit is deployed in an edge computing system. The system includes: an edge computing cluster, a control node, and a network connecting each edge computing node. The task computing volume estimation module and the online decision-making module based on multi-armed bandit are both deployed on the control node. A series of jobs arrive at the edge computing system in each time slot. Each job consists of several tasks, and these tasks may arrive at different edge clusters. As Figure 2 shown, Task 3 submits the input data volume and task type to the control node. After a series of processes, the control node issues the replication decision made to Task 5 on Edge 2. Subsequently, Edge 2 replicates a copy for Task 5 according to the replication decision and transfers it to the corresponding target edge cluster for execution.
[0093] In this system, the control node interacts with each edge computing cluster periodically, and in real time feeds back the historical bandwidth status and computing power of the edge computing cluster to the control node. The control node makes a suitable replication decision for the task by processing the computing volume of the task and estimating the bandwidth and computing power, and issues it to the edge computing cluster where the task is located. After receiving the decision command sent by the control node, the edge cluster replicates r t copies of the task and distributes them to the target edge cluster for execution. During the execution process, the transmission delay and computing delay of this time are sampled to obtain the historical status of part of the bandwidth and computing power of the edge computing system. The specific execution process is as follows:
[0094] (S1) At the beginning of each time slot (the length of this time slot is fixed as the system configuration), a series of jobs arrive at the edge computing system. Each job consists of a series of tasks that randomly arrive at different edges. Among them, the hot edge replicates the tasks on it one by one. For the task t at the head of the queue, the input data volume and task type of the task t are sent to the control node;
[0095] (S2) The control node receives the relevant information of task t. The "task computing volume estimation module" estimates the computing volume of task t according to the input data volume and type of task t and sends it to the "online decision-making module";
[0096] (S3) The "edge system manager" collects the bandwidth information and computing power of the edge computing system every time it trains, and calculates the States (status) information of all available edges according to these historical bandwidth information and computing power information, that is
[0097] (S4) The "online decision-making module" receives the computing volume of task t, and according to the calculated by the "edge system manager" for all available edges t selects r The smallest edge cluster is used as the target edge cluster, and then the set of target edge clusters is sent as an Action to the "scheduler" of the control node;
[0098] (S5) After receiving the Action, the "scheduler" generates the edge cluster where the replication decision task t is located;
[0099] (S6) After the edge cluster where the task t is located receives the replication decision, according to the decision content, r copies of the task t are replicated and sent to the corresponding r t target edge clusters respectively; t
[0100] (S7) The target edge clusters that receive the copies execute the corresponding copies according to the first-come-first-served principle, and timely feedback the results to the edge cluster where the task t is located, and sample the corresponding bandwidth and computing power, and send the sampling information to the "edge system manager" of the control node;
[0101] (S8) After the edge where the task t is located receives the result returned by the first target edge cluster, it marks that the task t has been calculated;
[0102] (S9) Obviously, the completion delay of a job depends on the completion delay of its slowest task. This process repeats in turn until all tasks of all jobs are executed.
[0103] Among them, the overall goal of the control node scheduling is to minimize the regret of all jobs under the fluctuations of the edge computing cluster resources and the edge network bandwidth within a period of time (several time slots). The specific form of the established optimization problem can be seen in the above formula (2), which will not be elaborated here.
[0104] From the above embodiments and by comparing with other different algorithms, it can be further illustrated that the task replication method based on multi-armed bandit in the edge computing environment of the present invention is superior to other current advanced algorithms. The comparison methods include: 1) Local execution: For jobs arriving at the edge system, the tasks of these jobs randomly arrive at different edge clusters. Then the tasks are executed on the edge clusters where the tasks arrive without offloading or replication. 2) Random algorithm: A simple strategy in which the edge cluster randomly selects other edge clusters to replicate and offload copies in each time period. 3) Single offloading: A learning-based task offloading strategy that only selects one target edge cluster for offloading at a time.
[0105] The effect of the experiment is as Figures 3 to 5 shown Figure 3 Shows the regret of the online decision-making algorithm, local execution, random algorithm, and single offloading proposed by the present invention for the gambling machine. Obviously, the "online decision-making algorithm for the gambling machine" has the smallest regret due to its convergence. However, "local execution" has the largest regret due to its poor strategy. As Figure 3 shown, the "random algorithm" is not convergent because it is random when selecting the target edge cluster. Similar to the "online decision-making algorithm for the gambling machine", "single offloading" is also convergent. However, it has a higher regret than the "online decision-making algorithm for the gambling machine" because its learning speed is slow and it cannot fully utilize the resources of the edge system. As Figure 4 shown, by the 40th time slot, the average job completion times of "single offloading", "random algorithm", and "local execution" are 56.4%, 77.6%, and 128.1% higher than that of the "online decision-making algorithm for the gambling machine", respectively. It can be seen that the "online decision-making algorithm for the gambling machine" is far superior to other algorithms. As Figure 5 shown, as the skewness of the tasks increases, the latency of "local execution" becomes longer and longer. This is because the higher the skewness, more tasks are concentrated on one edge cluster, and "local execution" does not allow task offloading and replication, so the latency inevitably becomes longer and longer. However, as the skewness increases, the latencies of other algorithms become shorter and shorter. This is because when the skewness is high, tasks are concentrated in one edge cluster, and task offloading or replication can reduce the burden of local execution, thus greatly improving the performance. In any case, the "online decision-making algorithm for the gambling machine" is significantly superior to other algorithms.
[0106] With the rapid development of edge computing, edge clusters need to process a large number of tasks, causing some edge clusters to be overloaded, which is further translated into task completion lags. Previous work usually copies tasks from overloaded edges to idle edges to reduce task queuing and computing delays. However, before making the replication decision, it is impossible to predict the completion delays of tasks replicated to different edges, which will affect the overall task replication performance. The present invention first proposes an online task replication model and algorithm based on the multi-armed bandit. Through strict proof and measuring the gap between online decision-making and offline optimal decision-making, the regret of this bandit-based algorithm is ensured to be sublinear.
[0107] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, devices, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0108] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0109] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0110] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, the interaction manner between the control node and the edge computing cluster in the present invention, the method of collecting feedback information content and the online decision-making method of task replication are applicable in each system. Those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.
Claims
1. A dynamic task replication method in an edge computing environment, characterized in that, it includes the following steps: An optimization problem is established with the goal of minimizing the regret, which is the difference between the total completion time of jobs in the edge environment and the total job completion delay under the ideal optimal replication decision. The regret is defined by the following formula: delay a represents the completion time of job a, ∑ a∈J delay a represents the total completion time of all jobs in the system, represents the theoretically optimal delay of job a, represents the theoretically optimal delay of all jobs in the system, and J represents the set composed of all jobs; Use a task replication decision algorithm based on multi-armed bandits to solve the optimization problem, including: At the beginning of the first time slot, estimate the task computation amount w according to the task type of the task and the size of the input data t ; For each task t, calculate the lower confidence bound of the latency for copying task t from edge cluster i to edge cluster j According to the lower confidence bound Determine all available edge clusters and select r t ones of the smallest available edge clusters as the target edge clusters, and copy the task to all target edge clusters for execution.
2. The dynamic task replication method in an edge computing environment according to claim 1, characterized in that, The said delay a = max t∈a (d t ), where t ∈ a means that task t is included in job a; d t represents the actual delay of task t; The x t represents the input data size of task t, and y t represents the output data size of task t, and π t represents the set of replication target edge clusters included in the replication decision made for task t, and trans i,j represents the bandwidth from edge cluster i to edge cluster j, and trans j,i represents the bandwidth from edge cluster j to edge cluster i, and com j represents the computing power of edge cluster j; K t represents the set composed of all available edge clusters for task t.
3. The dynamic task replication method in an edge computing environment according to claim 1, characterized in that, Estimate the task computation amount w according to the task type of the task and the size of the input data t including: Use the dimension N of the input data vector to represent the amount of data to be processed, and obtain the machine learning model structure type z of the task itself t , and use an estimation function based on N and z t to obtain the computational workload of the task.
4. The dynamic task replication method in an edge computing environment according to claim 1, characterized in that, The lower confidence limit of the latency for copying task t from edge cluster i to edge cluster j is calculated as follows: x t represents the input data size of task t; y t represents the output data size of task t; represents the number of times the link from edge cluster i to edge cluster j is sampled when task t is completed; represents the number of times the link from edge cluster j to edge cluster i is sampled when task t is completed; represents the number of times edge cluster j is selected as the target edge cluster when task t is completed; b i,j represents the bandwidth coefficient from edge cluster i to edge cluster j, b j,i represents the bandwidth coefficient from edge cluster j to edge cluster i, f j represents the computing power coefficient of edge cluster j; respectively represent after task t is executed b i,j 、b j,i and f j 's lower confidence limit.
5. The dynamic task replication method in an edge computing environment according to claim 4, characterized in that, The calculation formula is as follows: respectively represent b i,j after being sampled the average value after being sampled j,i after being sampled the average value after being sampled j after being sampled the average value after being sampled; The calculation methods of 6. The dynamic task replication method in an edge computing environment according to claim 4, characterized in that, b i,j 、b j,i and f j are calculated as follows: trans i,j represents the bandwidth from edge cluster i to edge cluster j, trans j,i represents the bandwidth from edge cluster j to edge cluster i, com j represents the computing power of edge cluster j.
7. A dynamic task replication device in an edge computing environment, characterized in that, it includes: An optimization problem construction module for establishing an optimization problem with the goal of minimizing the regret, which is the difference between the total completion time of jobs in the edge environment and the total job completion delay under the ideal optimal replication decision. The regret is defined by the following formula: delay a represents the completion time of job a, ∑ a∈J delay a represents the total completion time of all jobs in the system, represents the theoretically optimal delay of job a, represents the theoretically optimal delay of all jobs in the system, and J represents the set composed of all jobs; An optimization problem solving module for using a task replication decision algorithm based on multi-armed bandits to solve the optimization problem. The solution of the optimization problem includes: At the beginning of the first time slot, estimate the task computation amount w according to the task type of the task and the size of the input data t ; For each task t, calculate the lower confidence bound of the latency for copying task t from edge cluster i to edge cluster j According to the lower confidence bound Determine all available edge clusters and select r t ones The smallest available edge clusters as the target edge clusters and copy the tasks to all the target edge clusters for execution.
8. A computing device, characterized in that, it includes: One or more processors; A memory; And One or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors. When the program is executed by the processor, it implements the dynamic task replication method in the edge environment as described in any one of claims 1-6.
9. A dynamic task replication system in an edge computing environment, characterized in that, it includes: At least one control node and several edge computing clusters. The control node is interconnected with the edge computing clusters and between the edge computing clusters via a network. The edge cluster feeds back its computing power and bandwidth status at the end of each time slot to the control node. The overloaded edge cluster timely transmits the relevant information of the tasks to be replicated to the control node. The control node makes a replication decision for the overloaded edge cluster using the dynamic task replication method in the edge environment as described in any one of claims 1-6 and issues the decision to the edge cluster.
Citation Information
Patent Citations
A decision-making method for task unloading and migration based on user mobility
CN109947545A
Content deployment and distribution method and system for mobile edge computing
CN111901392A