DAG job scheduling method and system

The state feature vectors of DAG jobs are extracted through graph convolutional neural network, predict the scheduling probability of tasks, and generate optimized scheduling strategies, which solves the problem of inefficiency of Kubernetes in DAG job scheduling, and achieves the improvement of resource utilization and the reduction of data transmission overhead.

WO2025152195A1PCT designated stage expired Publication Date: 2025-07-24HAINAN INST OF ZHEJIANG UNIV +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/073551
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-15
Filing Date
2024-01-23
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Kubernetes has problems of inefficiency when scheduling DAG jobs, especially when dealing with tasks with dependencies, resulting in low resource utilization and excessive overhead for intermediate data caches.

Method used

The DAG job scheduling method based on graph convolution neural network is adopted to obtain the native feature data of jobs and tasks, extract embedded feature vectors, generate state feature vectors, predict the scheduling probability of the task, and generate scheduling strategies based on optimization goals to optimize the allocation of tasks at nodes.

Benefits of technology

It improves the completion efficiency of DAG jobs, reduces the data transmission overhead between tasks, optimizes resource utilization, and reduces the cache overhead of intermediate data in the cluster.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024073551_24072025_PF_FP_ABST
    Figure CN2024073551_24072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention are a DAG job scheduling method and system. The method is as follows: when a progress status of a job is updated, the following steps are executed: acquiring native feature data of each job; extracting embedding feature vectors from the native feature data on the basis of a graph convolutional neural network, and generating a status feature vector on the basis of an extraction result; inputting the status feature vector into a first prediction network, predicting, by the first prediction network, the scheduling probability of each candidate task, and obtaining scheduling priorities of candidate tasks; and according to the scheduling priorities from highest to lowest, sequentially selecting and allocating nodes for the candidate tasks on the basis of the status feature vectors, and generating a scheduling strategy on the basis of a selection result. The scheduling strategy provided in the present invention can globally improve the efficiency and can reduce the cache overhead of intermediate data during job execution.
Need to check novelty before this filing date? Find Prior Art

Description

DAG job scheduling method and system Technical Field

[0001] The present invention relates to the field of cloud computing, and in particular to a DAG job scheduling technology based on Kubernetes in a cloud-native environment. Background Art

[0002] In recent years, cloud native has gradually become the standard paradigm in the Cloud Computing 2.0 era. Cloud native uses open source stacks to containerize workloads, leverages microservices architectures to improve flexibility and maintainability, leverages agile methodologies and DevOps to support continuous iteration and automated operations, and leverages cloud platform infrastructure to achieve elastic scaling, dynamic scheduling, and optimized resource utilization.

[0003] Kubernetes is an open-source container cluster management system that provides application deployment, maintenance, and expansion mechanisms. Submitting a job on Kubernetes requires uploading a YAML file (a configuration file) that clearly defines the execution content and resource requirements of each component in a specified format. Kubernetes will then begin scheduling and executing each component in the job according to the scheduling policies and rules.

[0004] However, Kubernetes's current scheduling capabilities for computing tasks in the form of a "Directed Acyclic Graph (DAG)" structure need to be improved.

[0005] Summary of the Invention

[0006] In response to the shortcomings of Kubernetes in the prior art in its poor scheduling capability for DAG jobs, the present invention provides a DAG job scheduling technology based on policy search and graph convolutional networks.

[0007] In order to solve the above technical problems, the present invention is solved by the following technical solutions:

[0008] A DAG job scheduling method performs the following steps when the progress status of a job is updated:

[0009] Obtain native feature data for each job;

[0010] The native feature data includes a job native feature vector and a task native feature vector corresponding to each task;

[0011] The job native feature vector includes node allocation information of the job;

[0012] The task native feature vector includes the node allocation information and the resource occupancy information of the corresponding task;

[0013] Extracting an embedded feature vector from the native feature data based on a graph convolutional neural network, and generating a state feature vector based on the extraction result;

[0014] Recording tasks of each job as candidate tasks, wherein the state feature vectors correspond to the candidate tasks in a one-to-one manner;

[0015] Inputting the state feature vector into a first prediction network, and having the first prediction network predict the probability of each candidate task being scheduled, to obtain the scheduling priority of each candidate task;

[0016] In descending order of the scheduling priorities, allocation nodes are selected for the candidate tasks in turn based on the state feature vectors, and a scheduling strategy is generated based on the selection results.

[0017] As an implementable approach:

[0018] The extraction results include:

[0019] A task embedding feature vector corresponding one-to-one to the candidate tasks;

[0020] A job embedding feature vector corresponding to each job;

[0021] A global embedding feature vector;

[0022] The state feature vector includes the global feature vector, and further includes a task feature vector and an operation feature vector corresponding to the candidate task.

[0023] As an implementable approach:

[0024] The graph convolutional neural network includes:

[0025] A first extraction network generates a task embedding feature vector for the candidate task based on the job native feature vector and the task native feature vector of the candidate task and the task embedding feature vector of the child task of the candidate task;

[0026] a second extraction network, generating a job embedding feature vector of the job based on the job native feature vector of the job and the task embedding feature vector corresponding to the job;

[0027] The third extraction network generates the global embedding feature vector based on the embedding feature vectors of each job.

[0028] As an implementable method, the specific steps of selecting an allocation node for a candidate task include:

[0029] The candidate task of the current node to be assigned is used as the target task;

[0030] Filter nodes that meet the target task execution requirements to obtain candidate nodes;

[0031] Obtaining a node feature vector of each candidate node, wherein the node feature vector includes resource utilization data of the corresponding node;

[0032] Based on the node feature vector and the state feature vector of the target task, the allocation priority of each candidate node is calculated, and the candidate node with the largest allocation priority is used as the target node of the target task.

[0033] As an implementable approach:

[0034] The second prediction network predicts the probability of each candidate node being assigned based on the node feature vector and the state feature vector of the target task, and obtains the assignment priority of each candidate node.

[0035] As an implementable embodiment, the scheduling strategy includes at least one scheduling action, and the scheduling action is used to indicate the tasks and nodes to be scheduled;

[0036] After generating the scheduling strategy, the rewards corresponding to each scheduling action are calculated based on the preset optimization goal, and the cumulative rewards are obtained by statistics;

[0037] The preset optimization goal is to advance the completion time of the last job and reduce the cache overhead occupied by intermediate data in the cluster.

[0038] The accumulated reward is used to update the graph convolutional neural network and the first prediction network.

[0039] As an implementation option, the reward r for the k-1th scheduling action is k-1 The calculation formula is: k-1 =-(t k -t k-1 )(|J k |+|ξ(J k )|);

[0040] in:

[0041] t k Indicates the global time when the kth scheduling action is triggered;

[0042] t k-1 Indicates the global time when the k-1th scheduling action is triggered;

[0043] J k is the time period [t k-1 ,t k ) is a collection of unfinished tasks;

[0044] ξ(J k ) is Jk The collection of parent tasks of all tasks in;

[0045] |J k |It's J k The number of inner elements;

[0046] |ξ(J k )|is ξ(J k ) The number of elements in it.

[0047] A DAG job scheduling method, comprising:

[0048] The monitoring module is used to obtain and transmit the native feature data of each job when the progress status of the job is updated;

[0049] The native feature data includes a job native feature vector and a task native feature vector corresponding to each task;

[0050] The job native feature vector includes node allocation information of the job;

[0051] The task native feature vector includes the node allocation information and the resource occupancy information of the corresponding task; the agent module is used to generate a corresponding scheduling strategy based on the native feature data;

[0052] The agent module includes:

[0053] a feature extraction unit configured to extract an embedded feature vector from the native feature data based on a graph convolutional neural network, and generate a state feature vector based on the extraction result; record the task of each job as a candidate task, and the state feature vector corresponds to the candidate task in a one-to-one manner;

[0054] a first strategy unit, configured to input the state feature vector into a first prediction network, and have the first prediction network predict the probability of each candidate task being scheduled, thereby obtaining a scheduling priority for each candidate task;

[0055] The second strategy unit is used to select allocation nodes for the candidate tasks in descending order of the scheduling priorities, and generate a scheduling strategy based on the selection results.

[0056] As an implementable approach, it also includes the Kubernetes platform;

[0057] The Kubernetes platform is used to generate a corresponding RPC request when the progress status of a job is updated, and the RPC request is for the native feature data of each job;

[0058] The monitoring module is configured to receive the RPC request and obtain the native feature data;

[0059] The second policy unit is further configured to generate a corresponding RPC response based on the scheduling policy, and feed the RPC response back to the Kubernetes platform.

[0060] A computer-readable storage medium stores a computer program, which implements the steps of any of the above methods when executed by a processor.

[0061] The present invention has significant technical effects due to the adoption of the above technical solutions:

[0062] The present invention can extract the embedded feature vector of the native feature data through a graph convolutional neural network when the progress of job completion in the system is updated, so as to obtain the state feature vector of each task, determine the task to be scheduled in the current round based on the state feature vector, and select a suitable node for the task based on the state feature vector. Since the state feature vector contains resource occupancy information and node allocation information, it can reflect the data location of tasks with dependencies, so as to guide the generation of a scheduling strategy that can reduce the data transmission overhead between tasks with dependencies. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0064] Figure 1 is a schematic diagram of a DAG job;

[0065] FIG2 is a schematic diagram of the life cycle of each task in the DAG job shown in FIG1 ;

[0066] FIG3 is a flow chart of a DAG job scheduling method according to the present invention;

[0067] FIG4 is a schematic diagram of the structure of a DAG job scheduling system according to the present invention. DETAILED DESCRIPTION

[0068] The present invention will be further described in detail below with reference to the examples. The following examples are intended to explain the present invention but the present invention is not limited to the following examples.

[0069] illustrate:

[0070] (1) Virtual machine nodes and clusters: A cluster consists of several virtual machine nodes (VM nodes), which are referred to as nodes in the following text.

[0071] In this manual, represents a cluster, where ci represents the i-th node. All virtual machine nodes can communicate with each other, forming a fully connected graph. At any time, the remaining amount of various resources of any node can be known in Represents node c i The remaining amount of the qth type of resources, It is the collection of all resource categories;

[0072] In actual use, resource categories such as CPU, memory, network bandwidth, hard disk capacity, etc. can be determined according to actual needs, and nodes can be heterogeneous, that is, they have completely different initial resource quantities.

[0073] (1) Assignments and tasks:

[0074] A job consists of a series of interdependent tasks, such as segmentation, transcoding, and merging. The scheduling method proposed in this invention is applicable to genetic computing jobs and video transcoding jobs.

[0075] When submitting a job, the user specifies the dependencies of each task within the job and the resource requests of each task in the YAML file.

[0076] A job can be modeled as a directed acyclic graph (DAG). The DAG can have multiple input tasks and multiple output tasks depending on the actual situation.

[0077] For a single job, we divide its tasks into input tasks and remaining tasks:

[0078] Inbound tasks require pulling a dataset from the Object Storage Service (OBS) as input. Pulling datasets consumes bandwidth and is a significant component of job completion time. Multiple inbound tasks within the same job may pull different datasets or the same dataset. If these inbound tasks are scheduled to the same virtual machine node and the required dataset is the same, then the dataset only needs to be pulled from OBS to that virtual machine node once. After the inbound task is processed, the results are sent to its own subtasks.

[0079] Remaining tasks only process data from the parent task and are unrelated to the dataset stored in OBS. Similarly, each task must send its processing results to its designated child task. If the task is an outbound task, it must also send the processing results back to OBS. The process of sending dataset processing results back consumes bandwidth and contributes significantly to job completion time.

[0080] (3) Scheduling strategy:

[0081] Scheduling refers to placing a task of a job on a virtual machine node at a certain moment.

[0082] A task can be placed on a node if and only if the VM node satisfies the resource request of the task.

[0083] At any time, the number of unfinished jobs in the system changes dynamically in real time, and the resource margin of each node in the cluster also changes in real time.

[0084] (4) Operation life cycle analysis:

[0085] Taking the DAG job shown in Figure 1 as an example, the following description is provided:

[0086] In Figure 1, n1 through n4 represent tasks. After the job is submitted, the first task, n1, becomes the incoming task. With no previous tasks to wait for, it is immediately placed in the pending queue, awaiting execution. Reasons for this waiting include the lack of a node that meets the resource requirements or the presence of higher-priority tasks (other jobs) within the current scheduling cycle. Therefore, n1's status is marked as pending. After this waiting period, task n1 is assigned the corresponding node, p(n1), and is scheduled.

[0087] Figure 2 shows the scheduling time and completion time of each task of the job shown in Figure 1, as well as the cache size and time period of the data on the node to which each task is bound. Before being executed, a task will go through two states: waiting (pending) and scheduling (scheduling). Tasks in the pending state cannot be scheduled. Tasks in the scheduling state can be scheduled. According to the functions registered in Kubernetes, they will perform actions such as enqueue, allocate (preferred), preempt (priority scheduling), reclaim (queue weight resource recovery), backfill (maximum scheduling), etc. in sequence, and then be bound to a suitable node.

[0088] The task execution process involves reading the required data into memory, processing it, and outputting it to a buffer (or disk). These three steps can occur simultaneously, allowing for simultaneous reading and processing (and output). Upon completion, the task's output data is cached on the hard drive mounted on the node to which it was scheduled. The resources occupied by the task (and the corresponding pod) are released, and the pod itself is destroyed. Upon completion of the parent task, the child task becomes available for scheduling.

[0089] Taking n1 as an example, its output data volume is s 12 、s 13 、s14 (The total is ), which are the inputs of n2 to n4 respectively. The child task will often read all the cached data output by the parent task, and then determine the part it really needs based on the execution content. Therefore, the cache duration of the parent task output data is the difference between the slowest child task completion time and its own completion time. See Figure 2, until n4 is destroyed. 12 and s 13 The cache occupied will be released.

[0090] Therefore, the cache overhead of the intermediate data of the job shown in Figure 1 can be calculated by multiplying the cache amount of each intermediate data by the sum of the corresponding cache existence time Reflection, where Δt i Represents task n i The slowest subtask completion time and n i The difference between its own completion time.

[0091] (5) The optimization goal of the scheduling method proposed in the present invention is:

[0092] Bring forward the last job's completion time:

[0093] The completion time of a job refers to the completion time of the last completed task in the job. During Kubernetes operation, multiple jobs may be submitted at the initial stage, and new jobs may be submitted at any time during job execution. Therefore, the optimization goal is to minimize the completion time of the last completed job.

[0094] Reduce the cache overhead of intermediate data in the cluster:

[0095] As can be seen above, the output data of the parent task may need to be temporarily cached locally, which consumes disk resources on the VM node where it is located. Even if a pair of dependent tasks are scheduled to the same VM node, data may still be cached (depending on the parent task's wait time), but the data transmission overhead is zero. This optimization goal is to minimize the disk resource usage of the cluster by intermediate data.

[0096] Example 1: A DAG job scheduling method, when the progress status of a job is updated, performs the following steps:

[0097] S100, obtaining native feature data of each operation;

[0098] The native feature data includes a job native feature vector and a task native feature vector corresponding to each task;

[0099] The job native feature vector includes node allocation information of the job;

[0100] The task native feature vector includes the node allocation information and the resource occupancy information of the corresponding task;

[0101] In this embodiment:

[0102] The native feature vector of the job corresponding to the i-th job is recorded as x i , the task native feature vector corresponding to the vth task in the i-th job is recorded as

[0103] The node allocation information includes the number of different nodes that have been allocated to the job (integer type variable) and whether the first node in the current node queue to be allocated has been allocated to a task of the job (boolean type variable);

[0104] The resource occupancy information includes the input data volume (float type variable), output data volume (float type variable), (estimated) remaining execution time (float type variable) of the corresponding task, and each output data Cache time (float type variable) where ξ i Represents task n i The collection of all subtasks.

[0105] Note: The jobs mentioned refer to jobs to be completed.

[0106] S200, extracting an embedded feature vector from the native feature data based on a graph convolutional neural network, and generating a state feature vector based on the extraction result;

[0107] Recording tasks of each job as candidate tasks, wherein the state feature vectors correspond to the candidate tasks in a one-to-one manner;

[0108] S300, inputting the state feature vector into a first prediction network, and having the first prediction network predict the probability of each candidate task being scheduled, to obtain the scheduling priority of each candidate task;

[0109] S400 , selecting allocation nodes for candidate tasks in descending order of the scheduling priorities based on state feature vectors, and generating a scheduling strategy based on the selection results.

[0110] This embodiment can extract the embedded feature vector of the native feature data through a graph convolutional neural network when the progress of job completion is updated in the system, so as to obtain the state feature vector of each task, determine the task to be scheduled in the current round based on the state feature vector, and select a suitable node for the task based on the state feature vector. Since the state feature vector contains resource occupancy information and node allocation information, it can reflect the data location of tasks with dependencies, so as to guide the generation of a scheduling strategy that can reduce the data transmission overhead between tasks with dependencies.

[0111] Furthermore, the extraction results in step S200 include:

[0112] A task embedding feature vector corresponding one-to-one to the candidate tasks;

[0113] A job embedding feature vector corresponding to each job;

[0114] A global embedding feature vector;

[0115] The state feature vector includes the global feature vector, and further includes a task feature vector and an operation feature vector corresponding to the candidate task.

[0116] The embedded feature vector is essentially a compressed perception of the original features and the structural information of the graph, so that the unstructured features are preserved in a numerical manner. This embodiment uses a graph convolutional neural network to extract the embedded feature vector;

[0117] 3 , the graph convolutional neural network in this embodiment includes:

[0118] First extraction network:

[0119] Generate a task embedding feature vector for the candidate task based on the job native feature vector and the task native feature vector of the candidate task, and the task embedding feature vector of the child task of the candidate task;

[0120] In this embodiment, the task embedding feature vector is extracted using the following formula:

[0121] in:

[0122] Represents the task embedding feature vector corresponding to the vth task of the i-th job;

[0123] ξ v Indicates the vth task n v The set of all subtasks of Represents the embedded feature vector corresponding to the u-th task of the i-th job, which is task n vsubtasks;

[0124] represents the task native feature vector of the vth task of the i-th job;

[0125] Both f1 and g1 represent a nonlinear mapping function, and both adopt a fully connected neural network with three hidden layers, that is, the first extraction network includes a first fully connected neural network for implementing the nonlinear mapping of f1 and a second fully connected neural network for implementing the nonlinear mapping of g1. The two fully connected neural networks here are shared by the feature extraction of all tasks of all jobs.

[0126] From the above, we can see that the task embedding feature vector of a task is the result of nonlinear transformation of the task embedding feature vectors of all its subtasks and adding the native feature vector of the task itself.

[0127] a second extraction network, generating a job embedding feature vector of the job based on the job native feature vector of the job and the task embedding feature vector corresponding to the job;

[0128] In this embodiment, the extraction of the task embedding feature vector is performed using the following formula:

[0129] in:

[0130] y i represents the job embedding feature vector corresponding to the i-th job;

[0131] m i represents the i-th job;

[0132] Represents the task embedding feature vector corresponding to the vth task of the i-th job;

[0133] x i represents the job-native feature vector of the i-th job;

[0134] Both f2 and g2 represent a nonlinear mapping function, and both adopt a fully connected neural network containing three hidden layers, that is, the second extraction network includes a third fully connected neural network for implementing the nonlinear mapping of f2 and a fourth fully connected neural network for implementing the nonlinear mapping of g2.

[0135] As can be seen from the above, the job embedding feature vector of a job is the result of nonlinear transformation of the task embedding feature vectors of all its tasks and adding its own job native feature vector.

[0136] The third extraction network generates the global embedding feature vector based on the embedding feature vectors of each job.

[0137] In this embodiment, the global embedding feature vector z is extracted using the following formula:

[0138] in:

[0139] m i represents the i-th job;

[0140] y i represents the job embedding feature vector corresponding to the i-th job;

[0141] Both f3 and g3 represent a nonlinear mapping function, and both adopt a fully connected neural network containing three hidden layers, that is, the third extraction network includes a fifth fully connected neural network for implementing the nonlinear mapping of f3 and a sixth fully connected neural network for implementing the nonlinear mapping of g3.

[0142] In this embodiment, the graph convolutional neural network includes 6 fully connected neural networks, each of which includes three hidden layers, and the number of neurons in the three hidden layers is set to 16, 32, and 16 respectively.

[0143] The extraction result of the graph convolutional neural network output is

[0144] For any candidate task, based on its associated task embedding feature vector Job embedding feature vector y i and the global embedding feature vector z constitute the state feature vector.

[0145] 3 , in this embodiment, the state feature vector is input into the first prediction network, and the first prediction network predicts the probability of each candidate task being scheduled. The specific implementation method is as follows:

[0146] First, the state feature vector of the candidate task (the vth task of the i-th job) is input into a nonlinear mapping q to obtain the output

[0147] where q is a fully connected neural network with three hidden layers.

[0148] Then all are fed into the SoftMax layer, which returns the probability of all tasks being selected.

[0149] That is, in this embodiment, the first prediction network includes a first feature transformation network and a first output network;

[0150] The input of the first feature network is all state feature vectors, and the first feature network is the fully connected neural network q mentioned above;

[0151] The input of the first output network is the output of the first feature network, and the output is the probability of each candidate task being selected. The first output network is a SoftMax layer.

[0152] Note that if a task is not schedulable, the corresponding mask variable will directly set the selection probability of the task to 0.

[0153] The specific steps of selecting an allocation node for a candidate task in this embodiment include:

[0154] S410, taking the candidate task of the current node to be assigned as the target task;

[0155] S420, screening nodes that meet the target task execution requirements to obtain candidate nodes;

[0156] That is, nodes that meet the resource request amount of the target task are selected as candidate nodes.

[0157] S430, obtaining a node feature vector of each candidate node, wherein the node feature vector includes resource utilization data of the corresponding node;

[0158] Resource utilization data includes characteristics such as the link bandwidth from the virtual node to the OBS and the CPU main frequency of the virtual machine node.

[0159] S440 , calculating the allocation priority of each candidate node based on the node feature vector and the state feature vector of the target task, and selecting the candidate node with the highest allocation priority as the target node of the target task.

[0160] That is, the second prediction network predicts the probability of each candidate node being assigned based on the node feature vector and the state feature vector of the target task, and obtains the assignment priority of each candidate node.

[0161] Referring to Figure 3, the specific implementation is as follows:

[0162] Performing nonlinear mapping on the node feature vector and the state feature vector of the target task;

[0163] in:

[0164] w l Represents the nonlinear mapping result of the lth candidate node corresponding to the target task;

[0165] h l Represents the node feature vector of the lth candidate node corresponding to the target task;

[0166] w is a fully connected neural network with three hidden layers.

[0167] Likewise, all They are input together into the SoftMax layer, which returns the probability of all candidate virtual machine nodes being selected.

[0168] That is, in this embodiment, the second prediction network includes a second feature transformation network and a second output network;

[0169] The input of the second feature network is the sum of the node feature vectors corresponding to the target task and all candidate nodes, and the first feature network is the fully connected neural network q mentioned above;

[0170] The input of the first output network is the output of the first feature network, and the output is the probability of each candidate task being selected. The first output network is a SoftMax layer.

[0171] The scheduling strategy in this embodiment includes at least one scheduling action, and the scheduling action is used to indicate the task and node to be scheduled. That is, a pair of target tasks and target nodes obtained through the above steps forms a scheduling action.

[0172] After generating the scheduling strategy, the rewards corresponding to each scheduling action are calculated based on the preset optimization goal, and the cumulative rewards are obtained by statistics;

[0173] The preset optimization goal is to advance the completion time of the last job and reduce the cache overhead occupied by intermediate data in the cluster.

[0174] The accumulated reward is used to update the graph convolutional neural network, the first prediction network, and the second prediction network, that is, to update the parameters of the above eight fully connected neural networks.

[0175] The reward r for the k-1th scheduling action k-1 The calculation formula is: k-1 =-(t k -t k-1 )(|J k |+|ξ(J k )|);

[0176] in:

[0177] t k Indicates the global time when the kth scheduling action is triggered;

[0178] t k-1 Indicates the global time when the k-1th scheduling action is triggered;

[0179] J k is the time period [t k-1 ,t k ) is a collection of unfinished tasks;

[0180] ξ(Jk ) is J k The collection of parent tasks of all tasks in;

[0181] |J k |It's J k The number of inner elements;

[0182] |ξ(J k )|is ξ(J k ) The number of elements in it.

[0183] The first penalty term in the above formula penalizes completion time, while the second penalty term penalizes situations where the cache space occupied by the output data of a completed parent task cannot be released due to unfinished child tasks. Given that it's impossible to detect the input and output volume of intermediate data, and given that the output data volume of a task is not optimizable (it doesn't change with the performance of the scheduling algorithm), we only penalize based on the "quantity" dimension.

[0184] In this embodiment, the Actor-Critic algorithm is used to update the parameters of the above 8 fully connected neural networks.

[0185] Let θ represent all the parameters of the agent (including all the weights and biases of the 8 fully connected neural networks), and s k Denotes the state input from the environment (i.e., the input of the first prediction network) that triggers the k-th scheduling action, and is represented by a k represents the kth scheduling action that is legal in the current state, then all (s k ,a k ,r k ,s k+1 )The parameter update formula corresponding to the quadruple is:

[0186] [Corrected 06.02.2024 according to Rule 91] where π θ (s k ,a k ) indicates the state is s k When the selected scheduling action is a k The probability of , which is P(task)·P(VM node) in Figure 3; γ is the rate of parameter update; The cumulative reward from the current step to the end of the algorithm; reward b k is the baseline function value of the kth round, and its significance is to reduce the excessive (s k ,a k ,r k ,s k+1 )The impact of quadruple variance on training efficiency.

[0187] Note: The training of the graph convolutional neural network, the first prediction network, and the second prediction network is carried out in a simulation environment, and the parameters are updated based on the Actor-Critic algorithm. This is the same as the above content, so it will not be repeated in this manual.

[0188] In this embodiment, the progress status of a job is updated when:

[0189] When any job is completed, the occupied resources of all nodes assigned to it are released;

[0190] When any task is completed, one or more of its subtasks can enter the schedulable state;

[0191] A new job is submitted.

[0192] The present invention integrates task dependencies, data locations, etc. to achieve optimal scheduling of task containers to virtual machine nodes, while minimizing job completion time and reducing data transmission overhead between tasks with dependencies. It also optimizes the priorities of schedulable tasks and candidate virtual machine nodes by establishing affinity relationships between tasks and data, thereby achieving optimal startup of jobs.

[0193] Example 2: A DAG job scheduling system includes an agent, wherein the agent includes:

[0194] The monitoring module is used to obtain and transmit the native feature data of each job when the progress status of the job is updated;

[0195] The native feature data includes a job native feature vector and a task native feature vector corresponding to each task;

[0196] The job native feature vector includes node allocation information of the job;

[0197] The task native feature vector includes the node allocation information and the resource occupancy information of the corresponding task;

[0198] An agent module, configured to generate a corresponding scheduling strategy based on the native feature data;

[0199] The agent module includes:

[0200] a feature extraction unit configured to extract an embedded feature vector from the native feature data based on a graph convolutional neural network, and generate a state feature vector based on the extraction result; record the task of each job as a candidate task, and the state feature vector corresponds to the candidate task in a one-to-one manner;

[0201] a first strategy unit, configured to input the state feature vector into a first prediction network, and have the first prediction network predict the probability of each candidate task being scheduled, thereby obtaining a scheduling priority for each candidate task;

[0202] The second strategy unit is used to select allocation nodes for the candidate tasks in descending order of the scheduling priorities, and generate a scheduling strategy based on the selection results.

[0203] Furthermore, it also includes the Kubernetes platform;

[0204] 4 , the Kubernetes platform is used to generate a corresponding RPC request when the progress status of a job is updated, and the RPC request is for the native feature data of each job;

[0205] The monitoring module is configured to receive the RPC request and obtain the native feature data;

[0206] The second policy unit is further configured to generate a corresponding RPC response based on the scheduling policy, and feed the RPC response back to the Kubernetes platform.

[0207] Embodiment 3: A computer-readable storage medium stores a computer program, which implements the steps of the method described in embodiment 1 when executed by a processor.

[0208] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0209] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0210] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatus, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0211] The present invention is described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device produce a device for implementing the functions specified in one or more processes in the flowcharts and / or one or more blocks in the block diagrams.

[0212] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0213] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce computer-implemented processing, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0214] It should be noted that:

[0215] References in this specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Therefore, appearances of the phrases "one embodiment" or "an embodiment" in various places throughout this specification do not necessarily refer to the same embodiment.

[0216] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0217] Furthermore, it should be noted that the specific embodiments described in this specification may vary in the shapes and names of their components. Any equivalent or simple variations based on the structure, features, and principles described in the patented concept of this invention are included within the scope of protection of this patent. Persons skilled in the art may make various modifications, additions, or substitutions to the described specific embodiments, and these modifications, as long as they do not deviate from the structure of the invention or exceed the scope defined by the claims, shall fall within the scope of protection of this invention.

Claims

1. A DAG job scheduling method, characterized in that, When the progress status of a job is updated, the following steps are performed: Obtain the native feature data of each job; The native feature data includes a job native feature vector and a task native feature vector corresponding to each task one by one; The job native feature vector contains the node allocation information of the job; The task native feature vector contains the node allocation information and also contains the resource occupancy information of the corresponding task; Based on a graph convolutional neural network, extract the embedded feature vectors from the native feature data, and generate a status feature vector based on the extraction results; Record the tasks of each job as candidate tasks, and the status feature vector corresponds to the candidate tasks one by one; Input the status feature vector into a first prediction network, and the first prediction network predicts the probability of each candidate task being scheduled to obtain the scheduling priority of each candidate task; In the order from largest to smallest of the scheduling priority, based on the status feature vector, select and allocate nodes for the candidate tasks in sequence, and generate a scheduling policy based on the selection results.

2. The DAG job scheduling method according to claim 1, wherein: The extraction results include: Task embedded feature vectors corresponding to the candidate tasks one by one; Job embedded feature vectors corresponding to the jobs one by one; A global embedded feature vector; The status feature vector contains the global feature vector, and also contains task feature vectors and job feature vectors corresponding to the candidate tasks.

3. The DAG job scheduling method according to claim 1, wherein: The graph convolutional neural network includes: A first extraction network that generates the task embedded feature vector of the candidate task based on the job native feature vector and task native feature vector of the candidate task, and the task embedded feature vector of the child tasks of the candidate task; A second extraction network that generates the job embedded feature vector of the job based on the job native feature vector of the job and the task embedded feature vector corresponding to the job; A third extraction network that generates the global embedded feature vector based on each job embedded feature vector.

4. The DAG job scheduling method according to claim 1, wherein The specific steps for selecting and allocating nodes for candidate tasks include: Take the candidate task of the currently to-be-allocated node as the target task; Screen the nodes that meet the execution of the target task to obtain candidate nodes; Obtain the node feature vectors of each candidate node, and the node feature vector contains the resource utilization data of the corresponding node; Based on the node feature vector and the status feature vector of the target task, calculate the allocation priority of each candidate node, and take the candidate node with the largest allocation priority as the target node of the target task.

5. The DAG job scheduling method according to claim 4, wherein: A second prediction network predicts the probability of each candidate node being allocated based on the node feature vector and the status feature vector of the target task to obtain the allocation priority of each candidate node.

6. The DAG job scheduling method according to claim 1, wherein The scheduling policy includes at least one scheduling action, and the scheduling action is used to indicate the task and node for performing the scheduling; After generating the scheduling policy, calculate the rewards corresponding to each scheduling action based on a preset optimization objective, and statistically obtain the cumulative reward; The preset optimization objective is to advance the completion time of the last job and reduce the cache overhead occupied by intermediate data in the cluster; The cumulative reward is used to update the graph convolutional neural network and the first prediction network.

7. The DAG job scheduling method according to claim 6, wherein The reward r of the (k - 1)-th scheduling action k-1 is calculated by the formula: r k-1 = -(t k - t k-1 )(|J k | + |ξ(Jk)|); Wherein: t k represents the global time when the k-th scheduling action is triggered; t k-1 represents the global time when the (k - 1)-th scheduling action is triggered; J k is the set of tasks that are not completed within the time period [t k-1 , t k ); ξ(J k ) is the set of parent tasks of all tasks in J k ; |J k |is J k The number of elements within; |ξ(J k )| is the number of elements in ξ(J k ).

8. A DAG job scheduling method, characterized in that, It includes: A monitoring module, configured to obtain and transmit the native feature data of each job when the progress status of the job is updated; The native feature data includes a job native feature vector and a task native feature vector corresponding to each task; The job native feature vector includes the node allocation information of the job; The task native feature vector includes the node allocation information and also includes the resource occupancy information of the corresponding task; An agent module, configured to generate a corresponding scheduling policy based on the native feature data; The agent module includes: A feature extraction unit, configured to extract an embedded feature vector from the native feature data based on a graph convolutional neural network, and generate a status feature vector based on the extraction result; each task of each job is recorded as a candidate task, and the status feature vector corresponds to the candidate task one by one; A first policy unit, configured to input the status feature vector into a first prediction network, and predict the probability of each candidate task being scheduled by the first prediction network to obtain the scheduling priority of each candidate task; A second policy unit, configured to sequentially select and allocate nodes for candidate tasks in descending order of the scheduling priority, and generate a scheduling policy based on the selection result.

9. The DAG job scheduling method according to claim 8, wherein It further includes a Kubernetes platform; The Kubernetes platform is configured to generate a corresponding RPC request when the progress status of a job is updated, and the RPC request is for the native feature data of each job; The monitoring module is configured to receive the RPC request and obtain the native feature data; The second policy unit is further configured to generate a corresponding RPC response based on the scheduling policy and feedback the RPC response to the Kubernetes platform.

10. A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Multi-coflow scheduling method based on graph neural network deep reinforcement learning

    CN111756653A

  • Workflow scheduling method and system based on graph convolutional neural network

    CN112711475A

  • DAG task scheduling method and device, equipment and storage medium

    CN114756358A

  • Task scheduling method, device and equipment for intelligent cloud platform

    CN114860398A

  • Deep reinforcement learning-based Yarn cluster workflow scheduling method

    CN116069473A