A Cloud Job Scheduling Method and System Based on an Exploration-Exploitation Separated Joint Neural Network
By using the exploration and use of separate joint neural networks for job scheduling in the cloud computing system, the efficiency of job scheduling and resource allocation in the cloud computing system is solved, and lower job delays and energy consumption are achieved.
Patent Information
- Application Number
- CN202210233352.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-10
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-03-10
AI Technical Summary
In cloud computing systems, how to efficiently schedule jobs and reasonably configure resources, reduce the time cost of job execution and the energy consumption of virtual machine clusters has become the core problem that cloud service providers need to solve.
The cloud job scheduling method based on exploring the use of separate joint neural networks is adopted. By decoupling the user load into sub-jobs, distributing it to multiple queues, and job attributes are composed of states input to multiple parallel neural networks to output action decisions. Use the return function to calculate the cost value of each action decision, select the action decision that obtains the minimum cost value as the best action decision, and schedule sub-jobs to multiple clusters.
By separate exploration and utilizing partial neural network training, avoiding falling into local optimization, improving the optimization effect of job scheduling, reducing job delay and energy consumption.
Smart Images

Figure CN114579281B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cloud computing, and more specifically, to a cloud job scheduling method and system based on exploring and utilizing a separated joint neural network. Background Art
[0002] The development of cloud computing has promoted the rapid development of the entire information industry. Cloud computing is the product of the integration of traditional computer technology and network technology such as distributed computing, parallel computing, utility computing, network storage, virtualization, load balancing, hot backup redundancy, etc. For a long time, due to the heterogeneity of resources, the diversity of jobs, the difference in service quality and the huge number of users in the cloud computing platform, the cloud computing system has to process a large number of jobs and data. In this case, deploying a virtual machine alone is likely to overload the virtual machine server, resulting in slow job response and increasing the risk of SLA violation. For this reason, cloud service providers and users tend to use a multi-queue multi-virtual machine cluster service model. Multi-queue refers to organizing jobs of different nature or requirements into multiple queues for submission. Multi-virtual machine cluster refers to a cluster of virtual machines belonging to multiple users that can communicate and coordinate with each other, and its overall performance is greatly improved through control technologies such as intra-cluster load balancing. In this service model, efficient job scheduling and reasonable resource allocation, reducing the time cost of job execution and the energy consumption of virtual machine clusters are technical issues that cloud service providers attach great importance to, which directly determines the service quality level and operating profit of the cloud platform. Job scheduling and resource allocation optimization under multiple queues and multiple clusters has become one of the core issues that need to be urgently solved in the field of cloud computing.
[0003] Many scholars and institutions have conducted multi-faceted research on the scheduling optimization problem of cloud computing. Some scholars have tried to use heuristic algorithms to solve this problem, but traditional heuristic algorithms require specific conditions to obtain the optimal solution. Faced with the complex and changeable cloud environment, its universality is not strong, and it is easy to fall into the local optimal solution in the process of solving multi-objective optimization problems, and it is impossible to obtain the global optimal solution. Reinforcement learning, as a model-free learning method, has strong decision-making ability and obtains the optimal solution to cloud scheduling optimization problems through a continuous trial and error mechanism. Therefore, some researchers have tried to use reinforcement learning methods to solve the job scheduling and resource management problems of cloud computing, but in the face of large-scale state space, reinforcement learning algorithms are prone to slow convergence or non-convergence. Deep neural networks have strong perception capabilities and can effectively cope with large-scale state spaces, which makes up for the shortcomings of reinforcement learning.
[0004] In practical applications, cloud scheduling is a complex and changeable problem. When batch submitting jobs to the cluster for execution in the cloud job system, how to perform effective job scheduling is still worth exploring. Summary of the invention
[0005] The present invention aims to overcome at least one defect (shortcoming) of the above-mentioned prior art, and provides a cloud job scheduling method and system based on an exploration-exploitation separated joint neural network, which is used to optimize and solve the problem of how to perform effective job scheduling during the batch job submission process in a cloud system.
[0006] The technical solution adopted by the present invention is a cloud job scheduling method based on an exploration-exploitation separated joint neural network, including:
[0007] Decouple the user load into sub-jobs and allocate the sub-jobs to multiple queues;
[0008] Form a state from the job attributes of multiple jobs in multiple queues;
[0009] Use the state as the input of multiple parallel neural networks, and the multiple parallel neural networks output corresponding action decisions;
[0010] Use the reward function to calculate the cost value of each action decision, and select the action decision with the minimum cost value as the best action decision for the jobs in this queue;
[0011] Schedule the sub-jobs in the queue to multiple clusters according to the selected best action decision;
[0012] Among them, the multiple parallel neural networks are divided into an exploration part and an exploitation part, and the neural network of the exploitation part is trained using a training set to obtain a trained neural network;
[0013] The specific operation of using the state as the input of multiple parallel neural networks is to input the state into the trained neural network of the exploitation part and the neural network of the exploration part.
[0014] Furthermore, the job attributes include the number of CPU cycles required by the job and the amount of data to be transmitted by the job.
[0015] Furthermore, using the state as the input of multiple parallel neural networks, and the multiple parallel neural networks output corresponding action decisions, specifically:
[0016] State , where N represents the number of queues, , M represents the number of jobs in each queue, ; represents job The number of CPU cycles required, represents job The amount of data to be transmitted; t represents the scheduling time slot;
[0017] The action decision output by each neural network Is expressed as: , where represents the function of the th neural network parameter, , represents the job in the queue scheduled to the cluster in, and K represents the number of computing clusters in the system. , that is:
[0018] .
[0019] Furthermore, the cost value of each action decision is calculated using the reward function, specifically:
[0020] The reward function is defined as:
[0021]
[0022] where s represents the job set, d represents the action decision, represents the weight of the job delay in the overall optimization goal, represents the optimization weight of the energy consumption;
[0023] Let D represent all the scheduling policies for the job set s, then the optimization goal is expressed as:
[0024]
[0025] where, represents the bandwidth that the job can occupy when assigned to the cluster ; represents the bandwidth allocated between the queue and the cluster ; represents the computing power of each job, represents the computing power of the cluster.
[0026] Furthermore, the action decision that obtains the minimum cost value is selected as the best action decision for the jobs in this queue, specifically obtained according to the following formula:
[0027] Best action decision .
[0028] A cloud job scheduling system based on a joint neural network with exploration-exploitation separation according to the present invention includes:
[0029] An allocation module for decoupling the user load into sub-jobs and allocating the sub-jobs to multiple queues;
[0030] A state construction module for forming a state from the multiple job attributes in the multiple queues;
[0031] An action decision-making module, configured to use the state as the input of multiple parallel neural networks, and the multiple parallel neural networks output corresponding action decisions;
[0032] An optimal action decision-making module, configured to calculate the cost value of each action decision by using a reward function, and select the action decision with the minimum cost value as the optimal action decision for the jobs in the queue;
[0033] A scheduling module, configured to schedule the sub-jobs in the queue to multiple clusters according to the selected optimal action decision;
[0034] Wherein, the multiple parallel neural networks are divided into an exploration part and a exploitation part, and the neural network in the exploitation part is trained by using a training set to obtain a trained neural network;
[0035] The step of using the state as the input of multiple parallel neural networks is specifically to input the state into the neural network in the trained exploitation part and the neural network in the exploration part.
[0036] Further, the job attributes include the number of CPU cycles required by the job and the amount of data to be transmitted by the job.
[0037] Further, the action decision-making module is specifically:
[0038] State , where N represents the number of queues, , M represents the number of jobs included in each queue, ; represents job the number of CPU cycles required, represents job the amount of data to be transmitted; t represents the scheduling time slot;
[0039] The action decision output by each neural network is expressed as: where represents the function of the th neural network parameter, , represents the job in queue scheduled to cluster in, K represents the number of computing clusters in the system, , that is: , namely:
[0040] .
[0041] Further, the optimal action decision-making module is specifically:
[0042] The reward function is defined as:
[0043]
[0044] where s represents the job set, d represents the action decision, represents the weight of job latency in the overall optimization objective, represents the optimization weight of energy consumption;
[0045] Let D represent all scheduling policies for the job set s, then the optimization objective is expressed as:
[0046]
[0047] where, represents the job is assigned to the cluster and the bandwidth it can occupy; represents the queue and the cluster the allocated bandwidth between them; represents the computing power of each job, represents the computing power of the cluster;
[0048] The best action decision is obtained according to the following formula .
[0049] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the above-mentioned cloud job scheduling method is implemented.
[0050] Compared with the prior art, the beneficial effects of the present invention are:
[0051] The present invention divides the neural network into two parts: an exploration part and a utilization part, separates the exploration and utilization parts, and executes different training methods respectively. During training, only the neural network of the utilization part is trained, while the exploration part is not trained to ensure the possibility of the model obtaining other scheduling strategies, thereby avoiding falling into local optimality and further improving the job scheduling problem of the present invention. Brief Description of the Drawings
[0052] Figure 1 is a model diagram of the cloud computing system in the present invention.
[0053] Figure 2 is a flowchart of a cloud job scheduling method based on a joint neural network with separated exploration and utilization in the present invention.
[0054] Figure 3 is a schematic diagram of the performance comparison before and after the improvement of the UDL algorithm in the specific experiment of the present invention under the condition of the same number of clusters with different numbers of queues.
[0055] Figure 4 This is a schematic diagram showing the performance comparison of the UDL algorithm before and after improvement under different clusters with the same queue number in the specific experiment of the present invention. Detailed implementation manners
[0056] The accompanying drawings of the present invention are only for illustrative purposes and should not be construed as limiting the present invention. To better illustrate the following embodiments, some components in the drawings will be omitted, enlarged or reduced, which do not represent the dimensions of the actual products; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0057] As Figure 1 shown, individual users, research institutions, etc. submit the jobs that need to be executed more in batches to the corresponding cloud service providers. After receiving the jobs, the cloud service providers input the jobs into a job scheduler composed of multiple trained neural networks for scheduling. The job scheduler generates corresponding scheduling policies according to the submitted jobs and schedules the jobs to the corresponding computing clusters for processing according to these policies. Before the user jobs are submitted, the cloud service providers need to train each neural network with the training job set aiming at reducing energy consumption and job completion time.
[0058] The present invention aims at Figure 1 the system model shown to schedule cloud computing jobs with the goal of minimizing job latency and energy consumption.
[0059] Embodiment 1
[0060] As Figure 1 shown, this embodiment provides a cloud job scheduling method based on a joint neural network with exploration and exploitation separation, including:
[0061] S1. Decouple the user load into sub-jobs and allocate the sub-jobs to multiple queues; a job decoupler can be used to decouple the user load with dependency relationships into sub-jobs, and then allocate them to multiple waiting queues, while ensuring the priority data transmission and execution of the parent jobs of the sub-jobs in the waiting queues, and each job in the queue has atomicity and can run independently. Each waiting queue has the same storage space, and the number of queues is dynamically adjusted according to actual needs.
[0062] A large number of basic devices form a huge data center. The servers in the neighborhood can be clustered into computing clusters according to geographical locations. Cloud system job scheduling is to schedule the sub-jobs in multiple queues to multiple clusters.
[0063] Assume that the number of computing clusters in the system is K, denoted as , . The number of job queues waiting to be scheduled is N, denoted as , ; The number of jobs in each queue is M, denoted as , . Therefore, the total number of jobs is M * N. The job represents the m-th job in the n-th queue.
[0064] S2. Combine the job attributes of multiple jobs in multiple queues into a state; the job 's job attributes include the number of CPU cycles required by the job and the amount of data to be transferred by the job, represented by a binary tuple ( ), represents the number of CPU cycles required by job , represents the amount of data to be transferred by job . is a random variable and follows a uniform distribution, that is , where and represent the minimum and maximum values of the job data volume respectively. Additionally, assume that the number of CPU cycles required by each job is linearly related to the data volume of the job, that is:
[0065]
[0066] where, represents the computing power to data volume ratio CDR, and its value depends on the type of job, and different types of jobs have different ratios CDR.
[0067] S3. Use the state as the input of multiple parallel neural networks, and the multiple parallel neural networks output corresponding action decisions;
[0068] Among them, the multiple parallel neural networks use a deep neural network DNN. The multiple parallel neural networks are divided into an exploration part and a exploitation part. The neural network in the exploitation part is trained using a training set to obtain a trained neural network;
[0069] The specific operation of using the state as the input of multiple parallel neural networks is to input the state into the trained neural network in the exploitation part and the neural network in the exploration part.
[0070] Combine the job attributes of multiple jobs in multiple queues into a state , the state , where N represents the number of queues, , M represents the number of jobs in each queue, ; t represents the scheduling time slot;
[0071] State As the input of multiple parallel deep neural networks, each deep neural network (DNN) outputs different action decisions, and the action decisions output by each DNN are expressed as: , where represents a function of the th neural network parameter, and the action decision is a binary sequence, expressed as , represents the job in the queue scheduled to the cluster in, where K represents the number of computing clusters in the system, , that is:
[0072] .
[0073] S4. Calculate the cost value of each action decision using the reward function, and select the action decision with the minimum cost value as the best action decision for the jobs in the queue;
[0074] S5. Schedule the sub-jobs in the queue to multiple clusters according to the selected best action decision.
[0075] In the specific implementation process, first build a communication model and a computing model. Specifically:
[0076] (1) Communication model
[0077] The attributes of the cluster are represented by a triple ( ), where represents the computing power of the cluster, that is, the number of CPU cycles, represents the communication power consumption of the cluster, represents the computing power consumption of the cluster.
[0078] The communication bandwidth between the queue and the cluster is expressed as , represents the bandwidth allocated between the queue and the cluster .
[0079] According to the definition of , at the scheduling time slot t, the number of jobs allocated from the queue to the cluster is:
[0080]
[0081] The communication model includes the transmission time and energy consumption required for transmitting job data. When multiple jobs in the same queue are scheduled to the same cluster simultaneously, the bandwidth from the queue to the machine is allocated to these jobs according to the principle of equal distribution. Therefore, if a job is assigned to the cluster , the bandwidth it can occupy is:
[0082] .
[0083] Communication latency is the time consumed for uploading job data to the server, specifically:
[0084]
[0085] Communication energy consumption is the energy consumed during the job transmission process, specifically:
[0086] .
[0087] Thus, the total communication energy consumption of all jobs in queue is:
[0088] .
[0089] (2) Computation model
[0090] The computation model includes the computation latency and computation energy consumption of jobs. According to the principle of equal distribution, the computing power of a certain cluster is evenly distributed to all jobs scheduled to that cluster. Similarly, the number of jobs scheduled to cluster is:
[0091]
[0092] Thus, the computing power obtained by each job is:
[0093]
[0094] Computation latency is the time consumed for a job to complete the computation, specifically:
[0095] .
[0096] Computation energy consumption is the energy consumed during the job computation process, specifically:
[0097]
[0098] The total computation energy consumption of all jobs in queue is:
[0099] .
[0100] (3)Optimization Objectives
[0101] In a certain scheduling time slot t, the scheduled job set is executed in parallel in the cluster. Therefore, the total delay required for this batch of jobs is as follows:
[0102]
[0103] The total energy consumption for executing this batch of jobs is the sum of the energy consumption of each job, that is:
[0104] .
[0105] In the specific implementation process, the optimization objective of the present invention is to minimize job delay and energy consumption. This is a multi-objective optimization problem. A weight factor is assigned to each objective to characterize its emphasis in the total optimization objective. Assume that the scheduling strategy adopted by the job set s is d, then the return function of the system is defined as follows:
[0106]
[0107] where s represents the job set, represents the weight of job delay in the total optimization objective, represents the optimization weight of energy consumption; The larger the value of, the greater the proportion of delay in the total optimization objective, and the smaller the energy consumption. If = 0, it means that the optimization objective only considers the energy consumption factor. If = 1, it means that the optimization objective only considers the delay factor.
[0108] The goal of the system is to obtain the optimal scheduling strategy, that is, to minimize job delay and energy consumption. Let D represent all scheduling strategies of the job set s, then the optimization objective is expressed as:
[0109]
[0110] Select the action decision that obtains the minimum cost value as the best action decision for the jobs in this queue. Specifically, it is obtained according to the following formula:
[0111] Best Action Decision .
[0112] In the specific implementation process, the experience replay mechanism of deep reinforcement learning can be adopted during the training process of multiple parallel deep neural networks. The samples generated by the multiple parallel deep neural networks are stored in the same experience sample pool and used as a common training sample set for each deep neural network to train, thereby increasing the scale and diversity of the training samples. In addition, small batches of samples are randomly drawn from the sample pool periodically for network model training to guide multiple agents to explore the optimal scheduling strategy. This not only improves the exploration ability of agents for the optimal strategy but also improves the utilization rate of training samples. Specifically, the current job set state space and the best action decision are used as samples and stored in the experience pool. After the number of samples in the experience pool reaches the expected threshold, a miniBatch number of samples are randomly drawn from it periodically for model training. The goal is to minimize the expected return value, and the gradient descent algorithm can be used in the training process to minimize the cross-entropy loss to optimize the parameter values of each DNN :
[0113] 。
[0114] Embodiment 2
[0115] The present invention also provides a cloud job scheduling system based on a joint neural network with exploration and exploitation separation, specifically including:
[0116] An allocation module, which is used to decouple the user load into sub-jobs and allocate the sub-jobs to multiple queues; specifically, a job decoupler can be used to decouple the user load with dependencies into sub-jobs and then allocate them to multiple waiting queues, while ensuring the priority data transmission and execution of the parent jobs of the sub-jobs in the waiting queues, and each job in the queue has atomicity and can run independently. Each waiting queue has the same storage space, and the number of queues is dynamically adjusted according to actual needs.
[0117] A large number of basic devices form a huge data center. The servers in the vicinity can be clustered into computing clusters according to geographical location. Cloud system job scheduling is to schedule the sub-jobs in multiple queues to multiple clusters.
[0118] Assume that the number of computing clusters in the system is K, denoted as , . The number of job queues waiting to be scheduled is N, denoted as , ; the number of jobs in each queue is M, denoted as , . Therefore, the total number of jobs is M*N. Jobs Denotes the m-th job in the n-th queue.
[0119] A status construction module for forming a status from the job attributes of multiple jobs in multiple queues; the job attributes of a job include the number of CPU cycles required by the job and the amount of data to be transferred by the job, represented by a binary tuple ( ). Denotes the job the number of CPU cycles required. Denotes the job the amount of data to be transferred. is a random variable and follows a uniform distribution, i.e., , where and represent the minimum and maximum values of the job data volume respectively. Additionally, assume that the number of CPU cycles required for each job is linearly correlated with the data volume of the job, i.e.:
[0120]
[0121] where, represents the computing power to data volume ratio CDR, and its value depends on the type of job, and different types of jobs have different ratios CDR.
[0122] An action decision module for using the status as the input of multiple parallel neural networks, and the multiple parallel neural networks output corresponding action decisions;
[0123] wherein, the multiple parallel neural networks adopt a deep neural network DNN, and the multiple parallel neural networks are divided into an exploration part and a exploitation part, and the neural network of the exploitation part is trained using a training set to obtain a trained neural network;
[0124] The specific operation of using the status as the input of multiple parallel neural networks is to input the status into the trained exploitation part neural network and the exploration part neural network.
[0125] Form a status from the job attributes of multiple jobs in multiple queues , the status , where N represents the number of queues, , M represents the number of jobs included in each queue, ; t represents the scheduling time slot;
[0126] The status is used as the input of multiple parallel deep neural networks, and each deep neural network DNN outputs different action decisions. The action decision output by each deep neural network DNN is represented as: , where represents the A function of neural network parameters, action decision is a binary sequence, expressed as , indicating the queue in the job scheduled to the cluster in, K represents the number of computing clusters in the system, , that is:
[0127] .
[0128] The optimal action decision module is used to calculate the cost value of each action decision by using the reward function, and select the action decision with the minimum cost value as the optimal action decision for the jobs in this queue;
[0129] The optimal action decision module needs to build a communication model and a computing model, specifically:
[0130] (1) Communication model
[0131] The cluster 's attributes are represented by a triple ( ), where represents the computing power of the cluster, that is, the number of CPU cycles, represents the communication power consumption of the cluster, represents the computing power consumption of the cluster.
[0132] The communication bandwidth between the queue and the cluster is expressed as , indicating the queue and the cluster is allocated the bandwidth between them.
[0133] According to 's definition, in the scheduling time slot t, the number of jobs allocated from the queue to the cluster is:
[0134]
[0135] The communication model includes the transmission time and energy consumption required to transmit job data. When multiple jobs in the same queue are scheduled to the same cluster at the same time, the bandwidth of this queue to this machine is allocated to these jobs according to the average allocation principle. Therefore, if the job is allocated to the cluster , then the bandwidth it can occupy is:
[0136] .
[0137] Communication delay is the time consumed for uploading job data to the server, specifically:
[0138]
[0139] Communication energy consumption is the energy consumed during the job transmission process, specifically:
[0140] 。
[0141] Thus, the new energy consumption of all jobs in the queue is:
[0142] 。
[0143] (2)Computing model
[0144] The computing model includes the computing delay and computing energy consumption of jobs. Adopting the principle of equal distribution, the computing power of a certain cluster is evenly distributed to all jobs scheduled to this cluster. Similarly, the number of jobs scheduled to cluster is:
[0145]
[0146] Thus, the computing power obtained by each job is:
[0147]
[0148] The computing delay is the time consumed for a job to complete computing, specifically:
[0149] 。
[0150] The computing energy consumption is the energy consumed during the job computing process, specifically:
[0151]
[0152] Queue The computing energy consumption of all jobs in is:
[0153] 。
[0154] (3)Optimization objective
[0155] In a certain scheduling time slot t, the scheduled job set is executed in parallel in the cluster. Therefore, the total delay required for this batch of jobs is:
[0156]
[0157] The total energy consumption for executing this batch of jobs is the sum of the energy consumption of each job, that is:
[0158] 。
[0159] In the specific implementation process, the optimization objective of the present invention is to minimize job latency and energy consumption. This is a multi-objective optimization problem. A weight factor is assigned to each objective to represent its emphasis in the total optimization objective. Assuming that the scheduling strategy adopted by the job set s is d, the reward function of the system is defined as follows:
[0160]
[0161] where s represents the job set, represents the weight of job latency in the total optimization objective, represents the optimization weight of energy consumption; The larger the value of, the greater the proportion of latency in the total optimization objective, and the smaller the energy consumption. If = 0, it means that the optimization objective only considers the energy consumption factor. If = 1, it means that the optimization objective only considers the latency factor.
[0162] The goal of the system is to obtain the optimal scheduling strategy, that is, to minimize job latency and energy consumption. Let D represent all scheduling strategies of the job set s, then the optimization objective is expressed as:
[0163]
[0164] Select the action decision that obtains the minimum cost value as the best action decision for the jobs in this queue, specifically obtained according to the following formula:
[0165] Best action decision 。
[0166] A scheduling module, configured to schedule sub-jobs in the queue to multiple clusters according to the selected best action decision;
[0167] Among them, the multiple parallel neural networks are divided into an exploration part and a exploitation part, and the neural network of the exploitation part is trained using a training set to obtain a trained neural network;
[0168] The inputting the state into the multiple parallel neural networks specifically refers to inputting the state into the trained neural network of the exploitation part and the neural network of the exploration part.
[0169] Further, the job attributes include the number of CPU cycles required by the job and the amount of data to be transmitted by the job.
[0170] Further, the action decision module is specifically:
[0171] State , where N represents the number of queues, , M represents the number of jobs in each queue, ; represents the number of CPU cycles required for a job ; represents the amount of data to be transferred for a job ; t represents the scheduling time slot;
[0172] The action decision output by each neural network is expressed as: where represents a function of the th neural network parameter, , represents the job in queue scheduled to cluster ; K represents the number of computing clusters in the system, , that is:
[0173] .
[0174] Furthermore, the optimal action decision module is specifically:
[0175] The reward function is defined as:
[0176]
[0177] where s represents the job set, d represents the action decision, represents the weight of the job delay in the overall optimization objective, represents the optimization weight of the energy consumption;
[0178] Let D represent all scheduling policies for the job set s, then the optimization objective is expressed as:
[0179]
[0180] where represents the bandwidth that the job can occupy when assigned to cluster ; represents the bandwidth allocated between queue and cluster ; represents the computing power of each job, represents the computing power of the cluster;
[0181] The optimal action decision is obtained according to the following formula .
[0182] In the specific implementation process, during the training process of multiple parallel deep neural networks, the experience replay mechanism of deep reinforcement learning can be adopted. The samples generated by the multiple parallel deep neural networks are stored in the same experience sample pool and used as a common training sample set for each deep neural network to train, thereby increasing the scale and diversity of the training samples. In addition, small batches of samples are randomly drawn from the sample pool periodically for network model training to guide multiple agents to explore the optimal scheduling strategy. This not only improves the exploration ability of the agents for the optimal strategy but also improves the utilization rate of the training samples. Specifically, the current job set state space and the optimal action decision are used as samples and stored in the experience pool. After the number of samples in the experience pool reaches the expected threshold, miniBatch-sized samples are randomly drawn from it periodically for model training. The goal is to minimize the expected return value, and the gradient descent algorithm can be used in the training process to minimize the cross-entropy loss to optimize the parameter values of each DNN :
[0183] 。
[0184] In the specific implementation process, the learning model based on multiple deep neural networks in this application is called the United Deep Learning (UDL) model, and the job scheduling algorithm based on the UDL model is called the UDL algorithm. In practical applications, when applying the UDL algorithm, all of the multiple parallel neural networks are trained, and this algorithm is called the basic UDL algorithm. When the algorithm is executed, the multiple parallel neural networks are divided into an exploration part and a utilization part. The neural networks in the utilization part are trained using the training set to obtain the trained neural networks, and this algorithm is called the improved UDL algorithm. This application conducts experiments on the basic UDL algorithm and the improved UDL algorithm. Under different numbers of job queues in the same cluster, the comparison results before and after the improvement of the UDL algorithm are as Figure 3 shown. In the experiment, the number of clusters is fixed at 5, takes a value of 0.9, and the number of queues gradually increases from 3 to 12. It can be seen from Figure 3 that as the number of job queues increases, the system load increases accordingly, and the return values of the algorithm models all show an upward trend. When the number of job queues is small, the optimization effect of the algorithm is not very obvious. However, when the number of job queues reaches 5 or more, the optimization effect of the improved UDL is better than that before the improvement. Because the number of computing clusters is fixed, when the number of job queues increases, the competition for limited computing resources among the queues will be more intense. In this case, the improved UDL shows better job scheduling performance than before the improvement.
[0185] In addition, in the case of different clusters with the same number of job queues, the comparison results before and after the improvement of the UDL algorithm are as follows Figure 4 shown. In the experiment, the number of job queues was fixed at 10, taking a value of 0.9, and the number of clusters gradually increased from 3 to 12. As can be seen from Figure 4 this, as the number of clusters increases, the available resources of the system increase, and the return value of the algorithm shows a downward trend. Similar to the previous experiment, when the number of job queues is fixed, the fewer the number of clusters, the more intense the competition for computing resources by jobs. At this time, the improved UDL performs better than before the improvement.
[0186] Based on the above experimental data, it can be seen that based on the UDL algorithm, the training method of the neural network is improved, so that multiple parallel neural networks are divided into an exploration part and an exploitation part. The neural network in the exploitation part is trained using the training set to obtain the trained neural network, and the trained neural network and the untrained neural network are used for target optimization, and better scheduling results can be obtained.
[0187] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, rather than limitations on the specific implementation manners of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the claims of the present invention shall be included within the protection scope of the claims of the present invention.
Claims
1. A cloud job scheduling method based on an exploration-exploitation separated joint neural network, characterized in that, it includes: Decouple the user load into sub-jobs and allocate the sub-jobs to multiple queues; Combine the multiple job attributes in the multiple queues into a state; Take the state as the input of multiple parallel neural networks, and the multiple parallel neural networks output corresponding action decisions; Use the reward function to calculate the cost value of each action decision, and select the action decision with the minimum cost value as the best action decision for the jobs in the queue; Schedule the sub-jobs in the queue to multiple clusters according to the selected best action decision; Among them, the multiple parallel neural networks are divided into an exploration part and an exploitation part, and the neural network in the exploitation part is trained using a training set to obtain a trained neural network; The step of taking the state as the input of multiple parallel neural networks specifically means inputting the state into the trained neural network in the exploitation part and the neural network in the exploration part.
2. The cloud job scheduling method based on an exploration-exploitation separated joint neural network according to claim 1, characterized in that, the job attributes include the number of CPU cycles required by the job and the amount of data to be transferred by the job.
3. The cloud job scheduling method based on an exploration-exploitation separated joint neural network according to claim 2, characterized in that, Taking the state as the input of multiple parallel neural networks, and the multiple parallel neural networks output corresponding action decisions, specifically: Status , where N represents the number of queues, , M represents the number of jobs in each queue, ; represents the job the number of CPU cycles required, represents the job the amount of data to be transferred; t represents the scheduling time slot; The action decision output by each neural network is expressed as: , where represents the function of the th neural network parameter, , represents the job in the queue scheduled to the cluster in, K represents the number of computing clusters in the system, , that is: 。 4. The cloud job scheduling method based on an exploration-exploitation separated joint neural network according to claim 3, characterized in that, Using the reward function to calculate the cost value of each action decision, specifically: The reward function is defined as: where \(s\) represents the job set and \(d\) represents the action decision, indicating the weight of job latency in the overall optimization objective, indicating the optimization weight of energy consumption; Let D represent all scheduling strategies for the job set s, then the optimization objective is expressed as: Among them, represents the bandwidth that a job can occupy when assigned to the cluster ; represents the bandwidth allocated between the queue and the cluster ; represents the computing power of each job, represents the computing power of the cluster.
5. The cloud job scheduling method based on an exploration-exploitation separated joint neural network according to claim 4, characterized in that, Selecting the action decision with the minimum cost value as the best action decision for the jobs in the queue is specifically obtained according to the following formula: Optimal Action Decision .
6. A cloud job scheduling system based on an exploration-exploitation separated joint neural network, characterized in that, it includes: An allocation module for decoupling the user load into sub-jobs and allocating the sub-jobs to multiple queues; A state construction module for combining the multiple job attributes in the multiple queues into a state; An action decision module for taking the state as the input of multiple parallel neural networks, and the multiple parallel neural networks output corresponding action decisions; A best action decision module for using the reward function to calculate the cost value of each action decision and selecting the action decision with the minimum cost value as the best action decision for the jobs in the queue; A scheduling module for scheduling the sub-jobs in the queue to multiple clusters according to the selected best action decision; Among them, the multiple parallel neural networks are divided into an exploration part and an exploitation part, and the neural network in the exploitation part is trained using a training set to obtain a trained neural network; The step of taking the state as the input of multiple parallel neural networks specifically means inputting the state into the trained neural network in the exploitation part and the neural network in the exploration part.
7. A cloud job scheduling system based on an exploration-exploitation separated joint neural network according to claim 6, characterized in that, the job attributes include the number of CPU cycles required by the job and the amount of data to be transmitted by the job.
8. A cloud job scheduling method based on an exploration-exploitation separated joint neural network according to claim 7, characterized in that, the action decision module is specifically: Status , where N represents the number of queues, , M represents the number of jobs in each queue, ; represents the number of CPU cycles required for the job , represents the amount of data to be transferred for the job ; t represents the scheduling time slot; The action decision output by each neural network is expressed as: , where represents the function of the -th neural network parameter, , represents the job in the queue scheduled to the cluster in which K represents the number of computing clusters in the system, , that is: 。 9. A cloud job scheduling method based on an exploration-exploitation separated joint neural network according to claim 2, characterized in that, the optimal action decision module is specifically: the reward function is defined as: where \(s\) represents the job set and \(d\) represents the action decision, represents the weight of the job delay in the overall optimization objective, represents the optimization weight of the energy consumption; let D represent all scheduling policies of the job set s, then the optimization objective is expressed as: Among them, represents the bandwidth that a job can occupy when assigned to a cluster ; represents the bandwidth assigned between a queue and a cluster ; represents the computing power of each job, represents the computing power of the cluster; Obtain the optimal action decision according to the following formula .
10. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by a processor, it implements the cloud job scheduling method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Exploit-explore on heterogeneous data streams
CN109313727A
Cloud job scheduling and resource allocation method
CN111722910A