Cost-aware resource scheduling method and system

By establishing resource scheduling models and transformation optimization problems, combining task sorting algorithms and score calculation formulas, resource allocation is dynamically adjusted, and the problems of low resource utilization efficiency and task execution efficiency in the existing technology are solved, and the resource utilization efficiency and task execution efficiency are improved.

CN119987970AActive Publication Date: 2025-05-13SUN YAT SEN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510103288.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-13
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The prior art has low resource utilization efficiency and low task execution efficiency in resource scheduling, making it difficult to dynamically adjust resource allocation strategies to adapt to changing resource states.

Method used

By establishing a resource scheduling model, it is transformed into an optimization problem, and the tasks are sorted and allocated using task sorting algorithms and score calculation formulas, and resource allocation is dynamically adjusted to achieve the improvement of resource utilization efficiency and task execution efficiency.

Benefits of technology

It has achieved improvements in resource utilization efficiency and task execution efficiency, and can solve NP-Hard problems in polynomial time, and ensure that the distance between the solution of the approximate algorithm and the optimal solution exists in the upper bound.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987970A_ABST
    Figure CN119987970A_ABST
Patent Text Reader

Abstract

The invention discloses a cost-aware resource scheduling method and system. The method specifically comprises the following steps: establishing a resource scheduling model, and obtaining a first optimization problem according to the resource scheduling model; converting and establishing a second optimization problem according to the first optimization problem; obtaining a server set, a task set and an estimated completion time matrix of all tasks for all servers, and sorting the tasks in the task set by using a task sorting algorithm to obtain a task sequence; and according to the task sequence S and the estimated completion time matrix E of all tasks of the server set # imgabs0 # task set # imgabs1 # for all servers, solving the second optimization problem to obtain a final resource scheduling scheme. The method is high in resource utilization efficiency and task execution efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of resource scheduling, and more specifically, to a cost-aware resource scheduling method and system. Background Art

[0002] With the rapid development of artificial intelligence, AI tasks have brought huge computing challenges to traditional computing resources. In order to improve computing efficiency and reduce the cost of holding computing resources, a large number of task platforms have been generated to provide services to users. These platforms will purchase a large amount of computing resources and provide paid computing services to users. However, the task load is very complex. Too few computing resources will cause delays in user tasks, and too much computing resources will cause waste when idle. Therefore, task platforms usually rent computing resources from third-party service providers and dynamically scale the scale of resource rental when appropriate. In addition to the need to dynamically control the scale of computing resource rental, tasks in the task platform also need to be reasonably allocated to appropriate computing resources for execution. On the one hand, it can improve resource utilization, reduce average task completion time, and reduce maximum completion time on a fixed resource scale. On the other hand, it can provide better assistance for the reasonable scaling of computing resources.

[0003] Since third-party service providers can provide a wide variety of computing resources with different performance and prices, it is difficult to determine the optimal scheduling decision; different tasks have different characteristics and different execution time sensitivities, making it difficult to meet the diverse demands of task execution time as much as possible, and the task execution efficiency is low; due to the complexity of the scheduling problem, existing solutions mainly focus on heuristic solutions and reinforcement learning solutions. The scheduling solutions are time-consuming, cannot adapt well to the dynamic arrival of tasks, and have low resource utilization efficiency.

[0004] The prior art discloses a multi-rate data distribution method in a wireless network based on network coding, which adopts a dual coding graph to construct a coding graph by comprehensively considering the correlation between links, data packet lifetime and expected transmission time, so that the algorithm can be applied to a correlation network environment; a greedy algorithm is used to reduce the problem from an NP-hard problem to a linear level; before each round of distribution, vertices that are definitely timed out and vertices with low delay sensitivity are eliminated, so that vertices that can meet the delay requirements and have high delay sensitivity have a higher chance of being sent; and a mathematical method is used to describe the construction of the dual coding graph, and it is concluded that the complexity of the optimal algorithm is an NPhard problem; therefore, a link correlation-aware multi-rate dual coding algorithm (LMPC) is proposed, which determines the coding strategy and transmission rate by weighing the transmission delay and the life span of the data packet. This method does not dynamically adjust the allocation strategy for the dynamically changing resource status, but only allocates a specific number of resources, so the resource utilization efficiency is low. Summary of the invention

[0005] In view of the defects of low resource utilization efficiency and low task execution efficiency in the prior art, the present invention provides a cost-aware resource scheduling method and system, which has high resource utilization efficiency and high task execution efficiency.

[0006] The primary purpose of the present invention is to solve the above technical problems. The technical solutions of the present invention are as follows:

[0007] A cost-aware resource scheduling method, comprising:

[0008] S1: Establish a resource scheduling model, and obtain the first optimization problem according to the resource scheduling model;

[0009] S2: According to the first optimization problem, transform and establish a second optimization problem;

[0010] S3: Get a server collection Task Collection The estimated completion time matrix E of all tasks for all servers is used to sort the task set using the task sorting algorithm. Sort the tasks in to get the task sequence S;

[0011] S4: According to the task sequence S, the server set Task Collection The estimated completion time matrix E of all tasks for all servers is used to solve the second optimization problem and obtain the final resource scheduling solution.

[0012] Furthermore, the first optimization problem is as follows:

[0013]

[0014] X represents the allocation matrix, X i,j =1 means task j is executed in server i, Represents a server collection, represents the task set, j represents the task number, i represents the server number, P i represents the cost of renting the i-th server per unit time, e i,j represents the execution time of task j in server i, m j represents the resource requirement of the jth task, M i represents the total resources of the i-th server, w j represents the waiting time of the jth task, g i,j represents the time required for task j to be scheduled to server i, d j represents the time constraint of the jth task.

[0015] Furthermore, the second optimization problem is as follows:

[0016]

[0017]

[0018] λ represents the importance parameter of the total server cost, I j represents the intermediate computational effort of task j, X represents the allocation matrix, and X i,j =1 means task j is executed in server i, Represents a set of tasks, represents the server set, j represents the task number, i represents the server number, P i represents the cost of renting the i-th server per unit time, e i,j represents the execution time of task j in server i, m j represents the resource requirement of the jth task, M i represents the total resources of the i-th server, w j represents the waiting time of the jth task, g i,j represents the time required for task j to be scheduled to server i, d j represents the time constraint of the jth task.

[0019] Furthermore, the task sorting algorithm includes:

[0020] S301: In the task collection Select a task as the first task;

[0021] S302: Calculate the first task and server set The score of each server in , and the lowest score is used as the first score of the first task;

[0022] S303: putting the first score of the first task, the corresponding server, and the first task into a first set;

[0023] S304: In the task collection Select another task as the new first task, and repeat steps S302 to S303 until the task set All tasks in are selected;

[0024] S305: Sort the first tasks in the first set according to the first scores of the first tasks in ascending order to obtain a task sequence S.

[0025] Further, in step S4, solving the second optimization problem includes:

[0026] S401: In the task collection Select a task from the list as the second task;

[0027] S402: extracting a server corresponding to the second task from the task sequence S as the first server;

[0028] S403: Determine whether the first server is rented; if the first server is rented, execute step S404; if the first server is not rented, execute step S405;

[0029] S404: Determine the remaining resource conditions of the first server and the second task, and the time constraints of the first server and the second task; if both the remaining resource conditions and the time constraints are met, assign the second task to the first server; and execute step S407;

[0030] S405: Determine the time constraints between the first server and the second task; if the time constraints are met, rent the first server and assign the second task to the first server; execute step S407

[0031] S406: Execute the server unallocated algorithm;

[0032] S407: Select task set Another task in the task set is used as the new second task, and steps S402 to S406 are repeated until the task set All tasks in are selected and the final resource scheduling solution is obtained.

[0033] Furthermore, the server unallocated algorithm includes:

[0034] S40601: Collected on server Selecting the rented servers that meet the resource remaining condition and time constraint condition of the second task to form a second set;

[0035] S40602: Calculate the scores of the second task and each server in the second set, and use the server corresponding to the highest score as the second server;

[0036] S40603: Allocate the second task to the second server.

[0037] Furthermore, the time constraint conditions include:

[0038] g i,j +e i,j +w j ≤d j

[0039] w j represents the waiting time of the jth task, g i,j represents the time required for task j to be scheduled to server i, e i,j represents the execution time of task j in server i, dj represents the time constraint of the jth task.

[0040] Furthermore, the resource remaining condition includes:

[0041] m j ≤M i

[0042] m j represents the resource requirement of the jth task, M i Represents the total amount of resources of the i-th server.

[0043] Furthermore, in step S302 or step S40602, the score calculation formula is as follows:

[0044]

[0045] P i represents the cost of renting the i-th server per unit time, It indicates the estimated time for the jth task to be completed on the i-th server, C i,j represents the total cost of processing the jth task in the i-th server;

[0046]

[0047] C i,j represents the total cost of processing the jth task in the i-th server, m j represents the resource requirement of the jth task, S i,j represents the fraction of the j-th task processed in the i-th server.

[0048] A cost-aware resource scheduling system, comprising:

[0049] The first optimization problem module: establishes a resource scheduling model and obtains the first optimization problem according to the resource scheduling model;

[0050] Second optimization problem module: transform and establish a second optimization problem according to the first optimization problem;

[0051] Sorting module: Get server set Task Collection The estimated completion time matrix E of all tasks for all servers is used to sort the task set using the task sorting algorithm. Sort the tasks in to get the task sequence S;

[0052] Solution module: According to the task sequence S, server set Task Collection The estimated completion time matrix E of all tasks for all servers is used to solve the second optimization problem and obtain the final resource scheduling solution.

[0053] Compared with the prior art, the present invention has the following beneficial effects:

[0054] The present invention uses a sorting algorithm to sort tasks so that high-yield tasks can be scheduled as soon as possible. By transforming and establishing a second optimization problem, the problem to be solved is simplified, and the hard constraint of the task time is transformed into a soft constraint. By solving the second optimization problem, the solution of the NP-Hard problem can be completed in polynomial time, and there is an upper bound on the distance between the solution of the approximate algorithm and the optimal solution, ensuring that the algorithm can obtain a valuable solution under the worst input condition. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 A flow chart of a cost-aware resource scheduling method provided in Example 1.

[0056] Figure 2 This is a flowchart of the task sorting algorithm provided in Example 1.

[0057] Figure 3 This is a flowchart of the task sorting algorithm provided in Example 1.

[0058] Figure 4 This is a flowchart of the server unallocated algorithm provided in Example 1.

[0059] Figure 5 This is a bar chart of the rental cost in the online submission scenario of a real cluster task provided in Example 1.

[0060] Figure 6 A bar chart of the rental cost in the scenario where the tasks provided in Example 1 are uniformly submitted at the beginning of the scheduling period.

[0061] Figure 7 A bar chart of rental costs in a task submission scenario with a deadline requirement provided in Example 1.

[0062] Figure 8 This is a bar chart of the average task completion time in the real cluster task online submission scenario provided in Example 1.

[0063] Fig. 9 A histogram of average task completion time in a scenario where tasks provided in Example 1 are uniformly submitted at the beginning of a scheduling cycle.

[0064] Fig.10 A bar chart of average task completion time in a task submission scenario with a deadline requirement provided in Example 1.

[0065] Fig.11 A histogram of the maximum span of task completion in the real cluster task online submission scenario provided in Example 1.

[0066] Fig.12 A histogram of the maximum span of task completion in a scenario where the tasks provided in Example 1 are uniformly submitted at the beginning of a scheduling cycle.

[0067] Fig.13 A bar chart showing the maximum span of task completion in the task submission scenario with a deadline requirement provided in Example 1.

[0068] Fig.14 A line chart of the cluster deadline miss rate provided in Example 1.

[0069] Fig.15 This is a structural diagram of the AI ​​model training task platform provided in Example 2.

[0070] Fig.16 This is a structural diagram of the resource crowdsourcing platform provided in Example 3. DETAILED DESCRIPTION

[0071] The drawings are for illustrative purposes only and should not be construed as limiting the present patent;

[0072] In order to better illustrate the present embodiment, some parts in the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product;

[0073] It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0074] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0075] Example 1

[0076] like Figure 1 As shown, a cost-aware resource scheduling method includes:

[0077] S1: Establish a resource scheduling model, and obtain the first optimization problem according to the resource scheduling model;

[0078] S2: According to the first optimization problem, transform and establish a second optimization problem;

[0079] S3: Get a server collection Task Collection The estimated completion time matrix E of all tasks for all servers is used to sort the task set using the task sorting algorithm. Sort the tasks in to get the task sequence S;

[0080] S4: According to the task sequence S, the server set Task Collection The estimated completion time matrix E of all tasks for all servers is used to solve the second optimization problem and obtain the final resource scheduling solution.

[0081] It should be noted that the resource scheduling method mentioned in the present invention will be referred to as the CRS algorithm hereinafter.

[0082] Furthermore, the first optimization problem is as follows:

[0083]

[0084] X represents the allocation matrix, X i,j =1 means task j is executed in server i, Represents a server collection, represents the task set, j represents the task number, i represents the server number, P i represents the cost of renting the i-th server per unit time, e i,j represents the execution time of task j in server i, m j represents the resource requirement of the jth task, M i represents the total resources of the i-th server, w j represents the waiting time of the jth task, g i,j represents the time required for task j to be scheduled to server i, d j represents the time constraint of the jth task.

[0085] The total cost of renting server i on the platform is: C i (t) = P i (t)·T i ,in t represents time.

[0086] Constraint (1b) requires that each task can only be assigned to one server, constraint (1c) ensures that each task can be assigned according to the remaining resources, constraint (1d) ensures that the task can be completed before the time constraint, and constraint (1e) ensures that this problem is an integer linear programming problem.

[0087] It is difficult to directly solve the first optimization problem because: first, the total rental resources of the task platform are limited. The constraint (1d) is difficult to achieve the time constraint of each task. Second, when the task platform serves a large number of tasks and rents servers, the complexity of solving the integer linear programming problem is very expensive. This is because the integer linear programming problem is a recognized NP-Hard problem. Third, the problem Constraint (1d) requires the execution time of task j on server i, which cannot be accurately known in advance.

[0088] Furthermore, the second optimization problem is as follows:

[0089]

[0090] λ represents the importance parameter of the total server cost, I j represents the intermediate computational effort of task j, X represents the allocation matrix, and X i,j =1 means task j is executed in server i, Represents a set of tasks, represents the server set, j represents the task number, i represents the server number, P i represents the cost of renting the i-th server per unit time, e i,j represents the execution time of task j in server i, m j represents the resource requirement of the jth task, M i represents the total resources of the i-th server, w j represents the waiting time of the jth task, g i,j represents the time required for task j to be scheduled to server i, d j represents the time constraint of the jth task.

[0091] It should be noted that the following algorithms are all set to process tasks at a certain moment. You only need to execute them in a loop to process tasks in a time period.

[0092] It should be noted that the ← symbol in the present invention represents the meaning of assignment, which is equivalent to the = symbol.

[0093] Furthermore, if Figure 2 As shown, the task sorting algorithm includes:

[0094] S301: In the task collection Select a task as the first task;

[0095] S302: Calculate the first task and server set The score of each server in , and the lowest score is used as the first score of the first task;

[0096] S303: putting the first score of the first task, the corresponding server, and the first task into a first set;

[0097] S304: In the task collection Select another task as the new first task, and repeat steps S302 to S303 until the task set All tasks in are selected;

[0098] S305: Sort the first tasks in the first set according to the first scores of the first tasks in ascending order to obtain a task sequence S.

[0099] The pseudo code of the task sorting algorithm is as follows:

[0100] enter:

[0101] The estimated task completion time matrix E of all tasks currently waiting to be scheduled for all servers

[0102] The set of all tasks waiting to be scheduled

[0103] A collection of all rented servers

[0104] Output:

[0105] The sequence S after rescheduling all tasks waiting to be scheduled

[0106]

[0107]

[0108] Furthermore, if Figure 3 As shown, in step S4, solving the second optimization problem includes:

[0109] S401: In the task collection Select a task from the list as the second task;

[0110] S402: extracting a server corresponding to the second task from the task sequence S as the first server;

[0111] S403: Determine whether the first server is rented; if the first server is rented, execute step S404; if the first server is not rented, execute step S405;

[0112] S404: Determine the remaining resource conditions of the first server and the second task, and the time constraints of the first server and the second task; if both the remaining resource conditions and the time constraints are met, assign the second task to the first server; and execute step S407;

[0113] S405: Determine the time constraints between the first server and the second task; if the time constraints are met, rent the first server and assign the second task to the first server; execute step S407

[0114] S406: Execute the server unallocated algorithm;

[0115] S407: Select task set Another task in the task set is used as the new second task, and steps S402 to S406 are repeated until the task set All tasks in are selected and the final resource scheduling solution is obtained.

[0116] Furthermore, if Figure 4 As shown, the server unassigned algorithm includes:

[0117] S40601: Collected on server Selecting the rented servers that meet the resource remaining condition and time constraint condition of the second task to form a second set;

[0118] S40602: Calculate the scores of the second task and each server in the second set, and use the server corresponding to the highest score as the second server;

[0119] S40603: Allocate the second task to the second server.

[0120] This algorithm greedily selects appropriate resources for tasks based on the task sorting algorithm. When server resources do not meet the demand, this algorithm will automatically perform elastic scaling of the server.

[0121] The pseudo code for solving the second optimization problem is as follows:

[0122] enter:

[0123] The estimated task completion time matrix E of all tasks currently waiting to be scheduled for all servers

[0124] The set of all tasks waiting to be scheduled And the task sequence S obtained after the task sorting algorithm

[0125] A collection of all rented servers

[0126] Output:

[0127] The actual scheduling scheme is

[0128]

[0129]

[0130] The core idea of ​​this algorithm is to consider assigning the best server Best to each task j to be scheduled. j If Best jIf the resource request and remaining constraints (corresponding to constraint (2c)) and the execution time constraint (corresponding to constraint (2d)) can be satisfied, task j is assigned to Best first. j When Best j When it is not rented, try to keep it in a rented state. j If both constraints cannot be met, then select the Cand with the highest allocation value from the currently rented servers that meet both constraints. j If all servers cannot meet the service requirements, task j cannot be assigned at the current time t and will enter the next scheduling cycle for assignment.

[0131] Furthermore, the time constraint conditions include:

[0132] g i,j +e i,j +w j ≤d j

[0133] w j represents the waiting time of the jth task, g i,j represents the time required for task j to be scheduled to server i, e i,j represents the execution time of task j in server i, d j represents the time constraint of the jth task.

[0134] Furthermore, the resource remaining condition includes:

[0135] m j ≤M i

[0136] m j represents the resource requirement of the jth task, M i Represents the total amount of resources of the i-th server.

[0137] Furthermore, in step S302 or step S40602, the score calculation formula is as follows:

[0138]

[0139] P i represents the cost of renting the i-th server per unit time, It indicates the estimated time for the jth task to be completed on the i-th server, C i,j represents the total cost of processing the jth task in the i-th server;

[0140]

[0141] C i,jrepresents the total cost of processing the jth task in the i-th server, m j represents the resource demand of the jth task. All tasks currently waiting to be scheduled are for all servers. The estimated task completion time matrix can be expressed as E, S i,j represents the fraction of the j-th task processed in the i-th server.

[0142] A cost-aware resource scheduling system, comprising:

[0143] The first optimization problem module: establishes a resource scheduling model and obtains the first optimization problem according to the resource scheduling model;

[0144] Second optimization problem module: transform and establish a second optimization problem according to the first optimization problem;

[0145] Sorting module: Get server set Task Collection The estimated completion time matrix E of all tasks for all servers is used to sort the task set using the task sorting algorithm. Sort the tasks in to get the task sequence S;

[0146] Solution module: According to the task sequence S, server set Task Collection The estimated completion time matrix E of all tasks for all servers is used to solve the second optimization problem and obtain the final resource scheduling solution.

[0147] The results obtained by the algorithm for solving the second optimization problem have a technical effect with a theoretically guaranteed upper bound on performance, and the proof is given below:

[0148] There is a definition of approximation ratio in approximation algorithms. The approximation ratio is defined as the ratio between the approximate solution obtained by the approximate algorithm and the optimal solution that can be obtained in theory, that is, the gap between the performance upper bound of the approximate solution and the optimal solution. Let B represent an NP problem that maximizes the objective function, and A represent an approximate solution obtained by an approximate algorithm proposed to solve problem B. Then for each problem instance of NP problem B, we can get the formula:

[0149]

[0150] In the above formula, OPT(L) is the optimal solution of an instance L of NP problem B, and correspondingly, A(L) is the approximate solution obtained by approximate algorithm A to solve the instance. A It is defined as the approximation ratio of the approximation algorithm A, that is:

[0151]

[0152] Among them, if R A is a bounded constant greater than 1 and less than positive infinity, which means that the approximation algorithm A is an algorithm with a constant approximation ratio.

[0153] The elastic cloud resource AI task scheduling problem proposed in the present invention is a problem of minimizing the cost of renting a cloud server. Therefore, according to the definition of the above approximation ratio, the optimization objective of the present invention is transformed in the form of an inverse, that is, the problem to be solved is transformed into a target maximization problem, which is easier to derive the approximation ratio.

[0154] Definition C i,j is the cost of renting server i for task j. Therefore, the optimization goal is to minimize the total cost of all rented cloud servers. Now, in order to facilitate the derivation of the approximation ratio, define As the cost of renting server i for running task j increases, V i,j The value of will be reduced to meet the principle of maximizing the objective function, so the optimization goal at this time is It is necessary to maximize V as much as possible, that is, the approximate solution obtained by the algorithm can be close to the theoretical optimal solution.

[0155] For the sake of proof, the conditions are relaxed. First, assume that the total amount of resources of all cloud servers in the current task scheduling problem is the same, and the amount of resource requests required by the tasks submitted by the current users is also the same. In an actual production environment, this assumption is reasonable. For example, when users submit tasks of the same type in the same time period, such as DL reasoning tasks for image generation using large models, each task occupies the same amount of resources, and the video memory on medium-sized GPU cloud servers provided by cloud service providers is often a unified 32GB. Therefore, according to the assumption, the following formula can be obtained:

[0156]

[0157] Let A represent the problem raised by the present invention An approximation algorithm with polynomial time approximation solution after removing the deadline constraint, and the approximation ratio of the approximation algorithm is R A =c, where c is a constant. In addition, we define A′ as an approximate algorithm for solving the bin packing problem based on A. We define an instance of the bin packing problem I as follows: The volume of each item j is U j , the capacity of each box is O and satisfies the condition U j ≤O,

[0158] For the scenario of the present invention, the resource request amount of the task submitted by the user corresponds to mj =V j =U j , the maximum resource of each server is the same as the maximum capacity of each box, that is, M = O. Let: U = ∑ j U j . Use approximate algorithm A to solve problem instance I k , where k represents the number of servers in the platform, then when the first number k is obtained 1 Can satisfy Then, based on the idea of ​​First-Fit-Decreasing (FFD), the remaining items are loaded into k 2 Therefore, A′ uses a total of k boxes to solve the packing problem. 1 +k 2 boxes, and since A′ is an approximate algorithm obtained based on A in the packing problem, and the problem proposed by the present invention and the packing problem can also be matched according to the above content, A′ is obviously also an approximate algorithm with a polynomial time approximate solution.

[0159] Assume that the optimal solution obtained by solving Example I is to use only k * boxes can hold all the goods with a total volume of U. According to the above definition, we have So we have: Therefore, based on the above derivation, we can get k * ≥k 1 Since the total number of boxes in this problem is k, let k = k * +k ′ , then we can get k ′ ≤k 2 Therefore, the k used in the above FFD algorithm is 2 The total volume of the items in the boxes We set k 2 The minimum number of items in a box is l * boxes to pack, and it is easy to know k 2 > * , so we can get: Furthermore, According to the reference, we set it at l * The volume of the items packed in each box after Then the previous l * The volume of the items in each box will be greater than So we can get:

[0160]

[0161] Since in the scenario of the present invention, there will be at least 2 servers for rental scheduling, there must be Furthermore, if there is It will make That is, A′ is an approximate algorithm for solving the packing problem, and its approximation ratio can be obtained However, this is proved to be impossible by other works, unless P = NP. Therefore, in the scenario of the present invention, the cost of the solution obtained by the approximate algorithm CRS proposed by the present invention when the performance is the best will be:

[0162] Therefore, according to the above derivation, the problem to be solved in the present invention can be obtained The approximation ratio of the approximate solution obtained by the approximate algorithm, that is, the cost spent, is at least greater than or equal to times the optimal solution. Next, we prove that the upper bound of the cost required by the algorithm CRS proposed in this invention when solving the problem is the maximum multiple of the cost obtained by the theoretical optimal solution. First, we assume that F is the set of tasks to be scheduled obtained by a feasible solution obtained by the algorithm CRS proposed in this invention, and assume that Z is any subset of the task set submitted by the user. Then, we assume that: m(F)≥∑ i∈I M i This means that all the resources in the servers will be fully occupied by all the tasks in the set F. Otherwise, if this assumption is not true, it means that the server resources in the cluster are sufficient to run all the tasks submitted by users, which will result in the performance of all algorithms being very similar. F is the set of scheduled tasks, let F = {j 1 ,j 2 ,…,j l}, and k≤l, so the present invention sets F k ={j 1 ,j 2 ,…,j k} represents the first k tasks selected and scheduled by the algorithm CRS. Therefore, according to the greedy idea based on cost optimization in CRS, we can get:

[0163]

[0164] Among them, z∈Z\F k-1 Represents the removal of F from the set of tasks to be scheduled. k-1 The remaining task set after the task. According to the definition of cost performance in the CRS algorithm, we can draw the following conclusions:

[0165]

[0166] In addition, we assume that V(Z)>V(F l), otherwise, the CRS algorithm we proposed can produce optimal performance in any case, which is unreasonable. Under such an assumption, there must be m(j k ) < m(Z). Otherwise, the following inequality will occur:

[0167]

[0168] Therefore, it can be obtained that V(Z) ≤ V(F k ) ≤ V(F l ), which leads to a contradiction with the assumption made in the present invention.

[0169] After transforming the inequality, it can be obtained that Furthermore, from the above definitions of sets Z and F, it can be obtained that: Therefore, it can be further obtained that According to the previous assumption, the set of tasks to be scheduled obtained by CRS satisfies the condition Therefore, relative to the optimal solution, there must be m(F) ≥ m(OPT). If the set of solutions of tasks to be scheduled obtained according to CRS satisfies it means that all the submitted tasks are scheduled, that is At this time, the resource occupation of the tasks in the obtained solution set is the same as the video memory occupation of the tasks in the optimal solution set. Therefore, there is:

[0170]

[0171] Finally, since at the beginning of this section, for the convenience of proof, the reciprocal of the cost obtained by the CRS algorithm is defined as V, which can be transformed back at this time to obtain

[0172] Therefore, it can be obtained that the algorithm CRS can maintain the following performance bounds in any case:

[0173]

[0174] Compared with the existing technologies, the algorithm proposed by the present invention achieves optimal performance in terms of indicators such as the average task execution time and average task waiting time of all tasks. At the same time, the algorithm proposed by the present invention effectively reduces the cost of renting resources on the task platform and effectively improves the overall resource utilization rate of the rented resources on the task platform.

[0175] The dataset used for testing in the embodiments of the present invention is the data collected by Alibaba Cloud from its actual production clusters based on the released "Alibaba Cluster Trace Project". This publicly available dataset is widely used in the research field of the characteristics of modern Internet data centers and workloads.

[0176] The following indicators are used to evaluate the performance of the algorithm:

[0177] ① Economic cost: In the problem scenario of the present invention, the enterprise rents servers from the cloud service provider to run the tasks submitted by users, so the enterprise needs to pay the cloud service provider based on the pricing and rental time of the rented servers. Therefore, the most important optimization goal of the present invention is to minimize the economic cost of renting servers by the enterprise. An excellent scheduling strategy can minimize the economic cost of the enterprise while meeting the task requirements submitted by users as much as possible. Therefore, the enterprise cost of scheduling all tasks is used as an indicator for evaluation.

[0178] ② Average task completion time: In an actual production cluster, each user submits a task with specified relevant information at a certain time node. If there are sufficient computing resources in the cluster at this time, the task will be scheduled to run on a server. If there are no available resources, the task will be placed in a waiting queue for delayed scheduling. In addition, the execution time of a task is related to the performance of the server on which it actually runs. The better the performance, the shorter the running time. Therefore, the task completion time represents the time from when the task is submitted to when it is scheduled to be completed. The shorter the average task completion time, the better the user's needs can be met, so the average task completion time is used as an indicator for evaluation.

[0179] ③Maximum span of task completion (Makespan): The maximum span of task completion represents the sum of the time taken to complete all tasks submitted by users in a cluster, that is, the time from the time when the first task was submitted to the cluster to the time when the last task in the cluster was completed is the maximum span of task completion. When the resources in the cluster are reasonably used, the Makespan required to complete all tasks will be reduced, so the maximum span of task completion (Makespan) is used as an indicator for evaluation.

[0180] ④ Rate of unfinished tasks: When users specify a deadline for submitted tasks, it is necessary to reduce the number of tasks that are not completed before the deadline as much as possible, that is, to reduce the user's QoE loss. Therefore, when designing a scheduling algorithm, in addition to considering the economic cost of renting a cloud server and the task completion time, it is necessary to consider the user-specified deadline for each task. We define the rate of unfinished tasks as the ratio of the number of tasks in the cluster that are not completed before the deadline to the total number of tasks running in the cluster, and use it as an indicator for evaluation.

[0181] In this embodiment, the following algorithms are used as comparison algorithms:

[0182] ①Kubernetes default scheduling algorithm: In Kubernetes, which is widely used as the infrastructure of various cloud computing platforms, in order to reasonably allocate the tasks submitted to the Kubernetes cluster to the nodes with different resources in the cluster, researchers have designed a set of built-in default scheduling algorithms for Kubernetes. In the Kubernetes scheduler, whenever a task is submitted, a corresponding pod is created, and then according to the pre-selection strategy defined, all nodes in the cluster are traversed to obtain all nodes that meet the resource requirements specified by the pod; then, for these selected nodes, the resources on all nodes are scored and selected according to their preferred scoring mechanism. Usually, the nodes are scored based on the remaining resources and whether there are similar nodes scheduled recently. The priorities are sorted according to the scores, and finally the node with the highest priority is selected for scheduling.

[0183] ②Tetris algorithm: This algorithm is a resource scheduling strategy commonly used in cloud computing platforms, which aims to maximize the resource utilization and system performance of the cluster. Its design is inspired by the block organization mechanism in the Tetris game. The Tetris algorithm dynamically adjusts the location and number of tasks on the servers in the cluster according to the resource situation of the servers in the current cluster and the resource requirements of the submitted tasks to maximize the utilization of cluster resources. In addition, based on the idea of ​​load balancing, the algorithm will balance the resource utilization of different servers in the cluster as much as possible to avoid excessive concentration of tasks on certain servers, thereby improving the performance of the entire cluster.

[0184] ③Tiresias algorithm: This is a scheduling algorithm proposed by Microsoft Research in 2019 for clusters in its cloud computing platform Azure. The algorithm first proposes an offline heuristic scheduling strategy based on the workload of different servers in the cluster, resource utilization, and computing resource requirements of tasks, which can quickly handle most task scheduling scenarios in the cluster. In addition, it also uses the cluster's task operation history data and workload changes to prioritize different tasks and servers, and dynamically adjust the scheduling strategy to better adapt to different scenarios and needs.

[0185] ④Earliest Deadline First (EDF) algorithm: This is a scheduling algorithm commonly used when real-time tasks with deadline requirements exist in the cluster. EDF dynamically makes scheduling decisions based on the real-time status and deadline of the task. When a new task arrives or a task is completed, the system re-evaluates the task queue and re-orders it to ensure that it is always scheduled according to the earliest deadline. It can ensure that the task is completed before its deadline, so it is suitable for real-time systems with strict requirements on task response time and latency.

[0186] ⑤Moore's algorithm: This algorithm first designs a strategy based on the earliest deadline first strategy, and then checks whether the strategy is feasible. If the schedule can ensure that each task can meet its deadline, it is called a feasible schedule. If the scheduling strategy is not feasible, the largest task with a larger computational demand than the task that cannot meet the deadline will be repeatedly removed until the scheduling strategy becomes feasible. Finally, the algorithm schedules the previously removed tasks according to the idea of ​​random scheduling.

[0187] like Figures 5 to 14 As shown in the figure, the results of the algorithm are tested and analyzed based on three different task submission scenarios:

[0188] ① Scenario of online submission of real cluster tasks: In this scenario, test analysis is performed based on the data of tasks submitted for execution within a day in the dataset. During the day, tasks may be submitted for scheduling at every minute, i.e., time slot t, and the tasks in this scenario have no deadline constraints, which is similar to the tasks submitted in a normal cloud computing cluster.

[0189] ② All simulated tasks are submitted at the beginning of the cycle: According to experimental results, in the real online submission scenario of cluster tasks, the performance of the algorithm, especially the maximum span of task completion, will be affected by the last task submitted, namely the long tail effect. Therefore, all simulated tasks are submitted uniformly at the beginning of the scheduling cycle.

[0190] ③ Online submission scenario of tasks with deadline requirements: In this scenario, the deadline of each task is set based on the task QoS field of the data set, and the ratio of tasks completed on time obtained by each algorithm needs to be tested.

[0191] In addition, in real cloud server rental scenarios, the number of cloud servers available for rent in the same region by cloud service providers usually changes with changes in market demand, and sufficient cloud servers cannot be provided at all times. Therefore, the following three cloud server rental scenarios are simulated in the experiment:

[0192] ① 100%: This scenario corresponds to a situation where there are sufficient cloud servers available for rent in the resource pool of the cloud service provider. In the experimental scenario of the present invention, there are 4 cloud servers available for rent in each of the 5 types of cloud servers.

[0193] ②75%: This scenario corresponds to a situation where some other users in the cloud service provider's resource pool compete for leasing cloud servers. In the experimental scenario of the present invention, there are 3 cloud servers of each of the 5 types available for leasing;

[0194] ③50%: This scenario corresponds to a large number of users renting cloud servers in the resource pool of the cloud service provider, so the number of cloud servers currently available for rent is relatively small. In the experimental scenario of the present invention, there are 2 cloud servers available for rent for each of the 5 types of cloud servers.

[0195] Example 2

[0196] Based on the cost-aware resource scheduling method described in Example 1, this embodiment adopts the same cost-aware resource scheduling method as Example 1. Fig.15 As shown, it is an AI model training task platform that can be used to lease GPU computing power from a third-party cloud platform.

[0197] The platform needs to decide the number and time of renting GPU servers based on the load of tasks submitted by users. At this time, the platform can model the server of the present invention as a GPU server and modify the relevant data structure to meet the needs of the computing platform. The platform can model the tasks of the present invention as AI model training tasks entering the computing platform. When scheduling, the platform can access the cost-aware task allocation and resource dynamic scaling mechanism of the present invention to improve resource utilization and reduce rental costs.

[0198] Example 3

[0199] Based on the cost-aware resource scheduling method described in Example 1, this embodiment adopts the same cost-aware resource scheduling method as Example 1. Fig.16 As shown, it is a resource crowdsourcing platform that can be used for the idle computing power in society.

[0200] The platform needs to decide the amount and time of crowdsourcing computing power collection based on the load of tasks submitted by users, while considering whether to rent computing power from a third-party cloud platform. At this time, the platform can model the server of the present invention as a collected computing power resource, and modify the relevant data structure to meet the needs of the crowdsourcing platform. The platform can model the tasks of the present invention as tasks entering the computing platform. When scheduling, the platform can access the cost-aware task allocation and dynamic resource scaling mechanism of the present invention to improve resource utilization, reduce the occupation of unnecessary social computing power, and reduce invasiveness.

[0201] The same or similar reference numerals correspond to the same or similar components;

[0202] The terms used in the drawings to describe positional relationships are only used for illustrative purposes and should not be construed as limiting this patent;

[0203] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.

Claims

1. A cost-aware resource scheduling method, characterized in that: include: S1: Establish a resource scheduling model, and obtain the first optimization problem according to the resource scheduling model; S2: According to the first optimization problem, transform and establish a second optimization problem; S3: Get a server collection Task Collection The estimated completion time matrix E of all tasks for all servers is used to sort the task set using the task sorting algorithm. Sort the tasks in to get the task sequence S; S4: According to the task sequence S, the server set Task Collection The estimated completion time matrix E of all tasks for all servers is used to solve the second optimization problem and obtain the final resource scheduling solution.

2. A cost-aware resource scheduling method according to claim 1, characterized in that: The first optimization problem is as follows: X represents the allocation matrix, X i,j =1 means task j is executed in server i, Represents a server collection, represents the task set, j represents the task number, i represents the server number, P i represents the cost of renting the i-th server per unit time, e i,j represents the execution time of task j in server i, m j represents the resource requirement of the jth task, M i represents the total resources of the i-th server, w j represents the waiting time of the jth task, g i,j represents the time required for task j to be scheduled to server i, d j represents the time constraint of the jth task.

3. A cost-aware resource scheduling method according to claim 1, characterized in that: The second optimization problem is as follows: λ represents the importance parameter of the total server cost, I j represents the intermediate computational effort of task j, X represents the allocation matrix, and X i,j =1 means task j is executed in server i, Represents a set of tasks, represents the server set, j represents the task number, i represents the server number, P i represents the cost of renting the i-th server per unit time, e i,j represents the execution time of task j in server i, m j represents the resource requirement of the jth task, M i represents the total resources of the i-th server, w j represents the waiting time of the jth task, g i,j represents the time required for task j to be scheduled to server i, d j represents the time constraint of the jth task.

4. A cost-aware resource scheduling method according to claim 1, characterized in that: The task sorting algorithm includes: S301: In the task collection Select a task as the first task; S302: Calculate the first task and server set The score of each server in , and the lowest score is used as the first score of the first task; S303: putting the first score of the first task, the corresponding server, and the first task into a first set; S304: In the task collection Select another task as the new first task, and repeat steps S302 to S303 until the task set All tasks in are selected; S305: Sort the first tasks in the first set according to the first scores of the first tasks in ascending order to obtain a task sequence S.

5. A cost-aware resource scheduling method according to claim 1, characterized in that: In step S4, solving the second optimization problem includes: S401: In the task collection Select a task from the list as the second task; S402: extracting a server corresponding to the second task from the task sequence S as the first server; S403: Determine whether the first server is rented; if the first server is rented, execute step S404; if the first server is not rented, execute step S405; S404: Determine the remaining resource conditions of the first server and the second task, and the time constraints of the first server and the second task; if both the remaining resource conditions and the time constraints are met, assign the second task to the first server; and execute step S407; S405: Determine the time constraints between the first server and the second task; if the time constraints are met, rent the first server and assign the second task to the first server; execute step S407 S406: Execute the server unallocated algorithm; S407: Select task set Another task in the task set is used as the new second task, and steps S402 to S406 are repeated until the task set All tasks in are selected and the final resource scheduling solution is obtained.

6. A cost-aware resource scheduling method according to claim 5, characterized in that: The server unassigned algorithm includes: S40601: Collected on server Selecting the rented servers that meet the resource remaining condition and time constraint condition of the second task to form a second set; S40602: Calculate the scores of the second task and each server in the second set, and use the server corresponding to the highest score as the second server; S40603: Allocate the second task to the second server.

7. A cost-aware resource scheduling method according to claim 5 or 6, characterized in that: The time constraints include: g i,j +e i,j +w j ≤d j w j represents the waiting time of the jth task, g i,j represents the time required for task j to be scheduled to server i, e i,j represents the execution time of task j in server i, d j represents the time constraint of the jth task.

8. A cost-aware resource scheduling method according to claim 5 or 6, characterized in that: The resource remaining conditions include: m j ≤M i m j represents the resource requirement of the jth task, M i Represents the total amount of resources of the i-th server.

9. A cost-aware resource scheduling method according to claim 4 or 6, characterized in that: In step S302 or step S40602, the score calculation formula is as follows: P i represents the cost of renting the i-th server per unit time, It indicates the estimated time for the jth task to be completed on the i-th server, C i,j represents the total cost of processing the jth task in the i-th server; C i,j represents the total cost of processing the jth task in the i-th server, m j represents the resource requirement of the jth task, S i,j represents the fraction of the j-th task processed in the i-th server.

10. A cost-aware resource scheduling system, applied to the scheduling method according to any one of claims 1 to 9, characterized in that: include: The first optimization problem module: establishes a resource scheduling model and obtains the first optimization problem according to the resource scheduling model; Second optimization problem module: transform and establish a second optimization problem according to the first optimization problem; Sorting module: Get server set Task Collection The estimated completion time matrix E of all tasks for all servers is used to sort the task set using the task sorting algorithm. Sort the tasks in to get the task sequence S; Solution module: According to the task sequence S, server set Task Collection The estimated completion time matrix E of all tasks for all servers is used to solve the second optimization problem and obtain the final resource scheduling solution.

Citation Information

Patent Citations

  • Task scheduling algorithm for cost perception under cloud environment

    CN107992359A

  • Workflow execution time optimization method and device

    CN112308304A

  • Computing power scheduling method and system of AI server cluster

    CN118626224A

  • Cost-minimizing task scheduler

    WO2014124448A1