FELSTM model-based cluster job waiting time prediction method
Through the cluster job waiting time prediction method based on the FELSTM model, the problem of insufficient prediction accuracy and flexibility in the prior art is solved, and more efficient and accurate cluster job waiting time prediction is achieved.
Patent Information
- Application Number
- CN202411551047.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-01
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-11-01
AI Technical Summary
The prior art lacks practicality and flexibility in predicting the waiting time of computing cluster jobs, making it difficult to effectively utilize high-performance computing resources.
The cluster job wait time prediction method based on the FELSTM model is adopted. By collecting and normalizing historical job data, the FELSTM model is constructed, and the parameter optimization is performed using the SmoothL1Loss loss function and Adam optimizer.
The prediction accuracy and efficiency of cluster job waiting time is improved, the data "overwhelming" problem caused by low data feature dimensions is avoided, and the efficiency of the training process and the accuracy of the prediction results are improved.
Smart Images

Figure CN120066670A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and particularly to a method for predicting the waiting time of cluster jobs based on the FELSTM model. Background Art
[0002] In a busy computing cluster, the scheduling efficiency of the job scheduling system will largely affect the utilization rate of the entire cluster. As the proportion of high-performance computing in scientific research is increasing, the situation that computing resources are difficult to match the growing computing demands is becoming increasingly prominent. Therefore, how to effectively and fully utilize the limited resources in the high-performance computing cluster for more intelligent job scheduling is a very necessary and worthy content for in-depth research.
[0003] The basis of cluster intelligent scheduling is to predict the waiting time required for computing cluster jobs to queue (abbreviated as cluster job waiting time). In the cluster management software LSF (Load Sharing Facility, LSF), although a function for predicting the job waiting time is provided, the implementation of this function requires the user to give the upper limit of the running time of each job in advance or estimate the running time of the job. Therefore, some scholars have begun to study the prediction of job running time to indirectly predict the waiting time of cluster jobs. However, since the running time can only be predicted for specific job tasks, indirectly predicting the job waiting time by predicting the job running time has problems of insufficient practicality and flexibility in actual use. The job waiting time is related to multiple factors such as the existing job status in the cluster, the user's job type and usage habits, etc. It is difficult to make effective predictions using traditional statistical analysis or machine learning methods. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for predicting the waiting time of cluster jobs based on the FELSTM model to solve the problems raised in the above background art.
[0005] To solve the above technical problems, the present invention provides the following technical solutions: A method for predicting the waiting time of cluster jobs based on the FELSTM model, the method comprising the following steps: Step S100. Collect historical job data in the target computer system and perform normalization processing on the historical job data; Step S200. Divide the normalized historical job data into an input sequence and an output sequence, where the input sequence is the data for predicting the output sequence, and the output sequence is the target sequence that the model expects to output; divide the divided input sequence and output sequence into a training data set, a validation set, and a test set; Step S300. Construct an FELSTM model through a deep learning framework, set the loss function as SmoothL1Loss, and use the Adam optimizer for parameter optimization; Step S400. Perform hyperparameter optimization on the FELSTM model to make the loss function of the model converge stably and reach the lowest value during the training process; after the training is completed, verify the model performance and evaluate the prediction accuracy on the test set.
[0006] The historical job data includes job ID, the number of resources requested by the user, the number of idle resources of the user, the number of resources the user is waiting for, the number of jobs the user is waiting for, the running time of the user's jobs, the CPU usage time of the user, the number of idle resources in the queue, the number of resources waiting in the queue, and the job waiting time; The job ID refers to the unique identifier assigned by the system when a certain job is submitted; the number of resources requested by the user refers to the number of CPU cores requested by the user when a certain job is submitted; the number of idle resources of the user refers to the upper limit of the number of cores used by the general computing queue minus the number of cores already used by the user when a certain job is submitted; the number of resources the user is waiting for refers to the total number of cores of the jobs the user is waiting to run when a certain job is submitted; the number of jobs the user is waiting for refers to the number of jobs the user is waiting to run when a certain job is submitted; the running time of the user's jobs refers to the total running time of the jobs the user has run when a certain job is submitted, in seconds; the CPU usage time of the user refers to the total CPU time of the jobs the user has run when a certain job is submitted, in seconds; the number of idle resources in the queue refers to the number of cores available for use in the queue when a certain job is submitted; the number of resources waiting in the queue refers to the total number of cores waiting to run in the queue when a certain job is submitted; the job waiting time refers to the time required for a certain job to be submitted and run, in seconds.
[0007] In step S100, the specific process of normalizing the historical job data is as follows: For the historical job data, use the Min - Max method to normalize the historical job data, and the calculation formula is as follows: y’=(yi - y_min) / (y_max - y_min), where y’ represents the normalized historical job data, yi represents the historical job data in the original sample, y_max represents the maximum value of the corresponding historical job data in the original sample, and y_min represents the minimum value of the corresponding historical job data in the original sample.
[0008] Step S200 includes: S201. Organize the normalized historical homework data into a historical homework dataset D, and D = {d1, d2,..., dn}, where d1 represents the feature vector corresponding to the 1st historical homework, d2 represents the feature vector corresponding to the 2nd historical homework, and so on. dn represents the feature vector corresponding to the nth historical homework, and n represents the data number of the historical homework, taking positive integers; extract the homework waiting time from the historical homework dataset D and form a homework waiting time set W, and W = {w1, w2,..., wn}. Similarly, w1 represents the homework waiting time corresponding to the 1st historical homework, w2 represents the homework waiting time corresponding to the 2nd historical homework, and wn represents the homework waiting time corresponding to the nth historical homework; S202. Define the input step size H, indicating the number of consecutive homework in each input sequence; divide the input sequence and the output sequence. For the input sequence, the following operations are performed: For each homework i, from H to n, generate the input sequence Xi, and Xi = [di - H + 1, di - H + 2,..., di]. Set the waiting time of the current homework to -1, and the corresponding input sequence Xi = [di - H + 1, di - H + 2,..., di - 1, di, -1]; For the output sequence, the following operations are performed: Extract the waiting time of the current homework from the homework waiting time set W as the output sequence Yi, and Yi = wi; Summarize all input sequences and output sequences to form an input sequence set X and an output sequence set Y respectively. The input sequence set X = [XH, XH + 1,..., Xn], and the output sequence set Y = [YH, YH + 1,..., Yn]; S203. Define the total number of samples of the input sequence set X and the output sequence set Y as N, and N = n - H + 1; divide the dataset according to the following ratio: the training set N_train is 80%, the validation set N_val is 10%, and the test set N_test is 10%; the specific division quantities are: N_train = ⌊0.8×N⌋, N_val = ⌊0.1×N⌋, N_test = N - N_train - N_val, where ⌊⌋ represents rounding down; according to the calculated division quantities, divide the training set, validation set, and test set from the input sequence set X and the output sequence set Y.
[0009] The FELSTM model structure in step S300 includes an input layer, a feature encoding layer, multiple LSTMs, a fully connected layer, and an output layer; The input layer receives the input sequence Xi, and the dimension of the input sequence Xi is (N, H, F), where F represents the feature dimension; the feature encoding layer encodes the input sequence Xi to expand its dimension; the multi-layer LSTM constructs a long short-term memory LSTM network, and the state update process of the LSTM is as follows: Input gate: i t = σ(W i · [h t-1 , x t + b i ), where i t represents the activation value of the input gate, which determines the influence of the current input on the cell state; σ represents the Sigmoid activation function, which maps the input to between 0 and 1; W i represents the weight matrix of the input gate, which is responsible for linearly transforming the previous hidden state and the current input; [h t-1 , x t represents the vector connecting the hidden state h t-1 at the previous time step and the input x t at the current time step, which is expressed as a combined input; b i represents the bias vector of the input gate, which is used to adjust the activation value of the input gate and increase the flexibility of the model; Forget gate: f t = σ(W f · [h t-1 , x t + b f ), where f t is the activation value of the forget gate, which controls the degree to which the cell state at the previous time step is retained in the current state, and the value range is between 0 and 1; W f represents the weight matrix of the forget gate, which is responsible for linearly transforming the previous hidden state and the current input; b f represents the bias vector of the forget gate, which is used to adjust the activation value of the forget gate and enhance the flexibility of the model; Output gate: o t = σ(W o · [h t-1 , x t + b o ), where o t represents the activation value of the output gate, which controls the degree of influence of the current cell state on the current hidden state, and the value range is between 0 and 1; W o represents the weight matrix of the output gate, which is responsible for linearly transforming the previous hidden state and the current input; b o represents the bias vector of the output gate, which is used to adjust the activation value of the output gate and enhance the flexibility of the model; Cell state: C ~ t = tanh(WC · [h t-1 , x t + b C ), C t = f t · C t-1 + i t · C ~ t ; where C ~ t represents the candidate cell state at the current time step; tanh represents the hyperbolic tangent activation function, which maps the input to between -1 and 1, providing a non-linear transformation; W C represents the weight matrix of the cell state, responsible for linearly transforming the previous hidden state and the current input; b C represents the bias vector of the cell state, used to adjust the calculation of the candidate cell state; C t represents the cell state at the current time step, combining the influence of the previous state and the current input; The fully connected layer passes the output of the multi-layer LSTM to the fully connected layer to generate the final output; the output layer directly uses the output of the fully connected layer as the final prediction value.
[0010] The feature encoding layer encodes the input sequence Xi to expand its dimension, specifically including: Z t = tanh(W z · x t + b z ), where Z t represents the output after feature encoding, W z represents the weight matrix, b z represents the bias term; The fully connected layer passes the output of the multi-layer LSTM to the fully connected layer to generate the final output, specifically including: y pred = W out · h T + b out , where y pred represents the final predicted output; W out represents the weight matrix of the output layer; h T represents the hidden state at the last time step, containing the information of the entire input sequence; b out represents the bias vector of the output layer.
[0011] The fully connected layer can meet the requirements of the feature encoding layer. However, there are also problems if only the fully connected layer is used. When the fully connected layer processes the input data, it multiplies the weight matrix by the input data matrix and then adds the bias term to obtain the output data. However, during the training process, the sizes of the weight matrix and the bias term change continuously, which results in the mean value of the output data not being zero and the value range being unrestricted. This situation will slow down the convergence speed of the optimization algorithm and is prone to the problem of gradient explosion. To solve this problem, the hyperbolic tangent (tanh) activation function is added after the fully connected layer to make the output data maintain a mean value of zero and the value range limited to [-1, 1], so as to accelerate the convergence speed of the optimization algorithm and prevent gradient explosion.
[0012] In step S300, the loss function is set to SmoothL1Loss, and the specific content includes: For each job i, calculate the predicted value y pred and the loss L between the true value i , where the true value is the output sequence Yi; the specific calculation formula is: , and the total loss is the average value L of the losses of all jobs total .
[0013] In step S300, the Adam optimizer is used for parameter optimization, and the specific content includes: In the implementation of the Adam optimizer, the update formulas for the first moment and the second moment are as follows: Update of the first moment: m t =β 1 m t-1 +(1 - β 1 )∇L total , where m t represents the first moment at the current time step; β 1 represents the decay rate of the first moment, usually taking a value of 0.9; m t-1 represents the first moment at the previous time step; ∇L total represents the gradient of the current loss function with respect to the parameters; Update of the second moment: v t =β 2 v t-1 +(1 - β 2 )(∇L total ) 2 , where v t represents the second moment at the current time step; β 2 represents the decay rate of the second moment, usually taking a value of 0.999; v t-1 represents the second moment at the previous time step; The parameter update formula is: θ new = θ old -(αm t ) / [(v t ) ^ (1 / 2) + ϵ], where θ new represents the updated parameter value, θ old represents the parameter value before update, α represents the learning rate, and ϵ represents a constant.
[0014] The specific content of hyperparameter optimization for the FELSTM model in step S400 includes: Define the hyperparameters as the input step size H, the learning rate α, the number of hidden layer nodes N h , the number of LSTM layers N L and the number of feature encoding nodes N f ; Select the grid search method for hyperparameter optimization. For each set of parameter combinations (H, α, N h , N L , N f ), perform the following operations: Construct and train the FELSTM model using different hyperparameter combinations, and record the value of the loss function L total during each training process; For each hyperparameter combination, calculate the average loss L avg on the validation set, and select the hyperparameter combination that minimizes the average loss L avg on the validation set as the final hyperparameters.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention introduces a feature encoding layer, which combines a fully connected layer and a hyperbolic tangent (tanh) activation function, aiming to perform efficient feature extraction on the input data. This design can not only accelerate the convergence process of the optimization algorithm but also effectively avoid the problem of gradient explosion; On this basis, the feature encoding layer is combined with the long short-term memory network (LSTM) to develop the FELSTM model; Using this model to predict the waiting time of cluster jobs not only avoids the problem of data "submergence" caused by the low dimensionality of cluster job data features but also improves the efficiency of the training process and the accuracy of the prediction results. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings: Figure 1 is a schematic flowchart of a method for predicting the waiting time of cluster jobs based on the FELSTM model of the present invention; Figure 2It is a schematic diagram of the data collection result of a method for predicting the waiting time of cluster jobs based on the FELSTM model of the present invention; Figure 3 It is a schematic diagram of data normalization of a method for predicting the waiting time of cluster jobs based on the FELSTM model of the present invention; Figure 4 It is a schematic diagram of the results of the input sequence and output sequence of a method for predicting the waiting time of cluster jobs based on the FELSTM model of the present invention; Figure 5 It is the overall structure diagram of the FELSTM model of a method for predicting the waiting time of cluster jobs based on the FELSTM model of the present invention; Figure 6 It is a diagram of the values of the model hyperparameters of a method for predicting the waiting time of cluster jobs based on the FELSTM model of the present invention. Specific embodiments
[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0018] Please refer to Figure 1 , the present invention provides a technical solution: A method for predicting the waiting time of cluster jobs based on the FELSTM model, the method includes the following steps: Step S100. Collect historical job data in the target computer system and perform normalization processing on the historical job data; Step S200. Divide the normalized historical job data into an input sequence and an output sequence, where the input sequence is the data for predicting the output sequence, and the output sequence is the target sequence that the model expects to output; divide the divided input sequence and output sequence into a training data set, a validation set, and a test set; Step S300. Build an FELSTM model through a deep learning framework, set the loss function to SmoothL1Loss, and use the Adam optimizer for parameter optimization; Step S400. Optimize the hyperparameters of the FELSTM model to make the loss function of the model converge stably and reach the lowest value during the training process; after the training is completed, verify the model performance and evaluate the prediction accuracy on the test set.
[0019] Historical job data includes job ID, the number of resources requested by the user, the number of idle resources of the user, the number of resources the user is waiting for, the number of jobs the user is waiting for, the running time of the user's jobs, the CPU usage time of the user, the number of idle resources in the queue, the number of resources waiting in the queue, and the job waiting time; The job ID refers to the unique identifier assigned by the system when a certain job is submitted; the number of resources requested by the user refers to the number of CPU cores requested by the user when a certain job is submitted; the number of idle resources of the user refers to the upper limit of the number of cores used in the general computing queue minus the number of cores already used by the user when a certain job is submitted; the number of resources the user is waiting for refers to the total number of cores of the jobs the user is waiting to run when a certain job is submitted; the number of jobs the user is waiting for refers to the number of jobs the user is waiting to run when a certain job is submitted; the running time of the user's jobs refers to the total running time of the jobs the user has run when a certain job is submitted, in seconds; the CPU usage time of the user refers to the total CPU time of the jobs the user has run when a certain job is submitted, in seconds; the number of idle resources in the queue refers to the number of cores available for use in the queue when a certain job is submitted; the number of resources waiting in the queue refers to the total number of cores waiting to run in the queue when a certain job is submitted; the job waiting time refers to the time required for a certain job to be submitted and run, in seconds.
[0020] In step S100, the specific process of normalizing the historical job data is as follows: For the historical job data, the Min - Max method is used to normalize the historical job data, and the calculation formula is as follows: y’=(yi - y_min) / (y_max - y_min), where y’ represents the normalized historical job data, yi represents the historical job data in the original sample, y_max represents the maximum value of the corresponding historical job data in the original sample, and y_min represents the minimum value of the corresponding historical job data in the original sample.
[0021] Step S200 includes: S201. Organize the normalized historical job data into a historical job data set D, and D = {d1, d2,..., dn}, where d1 represents the feature vector corresponding to the 1st historical job, d2 represents the feature vector corresponding to the 2nd historical job, and so on, dn represents the feature vector corresponding to the nth historical job, and n represents the data number of the historical job, taking a positive integer; extract the job waiting time from the historical job data set D and form a job waiting time set W, and W = {w1, w2,..., wn}, similarly, w1 represents the job waiting time corresponding to the 1st historical job, w2 represents the job waiting time corresponding to the 2nd historical job, and wn represents the job waiting time corresponding to the nth historical job; S202. Define the input step size H, which represents the number of consecutive jobs included in each input sequence; divide the input sequence and the output sequence. For the input sequence, the following operations are performed: For each job i, from H to n, generate the input sequence Xi, and Xi = [di - H + 1, di - H + 2,..., di]. Set the waiting time of the current job to -1, and the corresponding input sequence Xi = [di - H + 1, di - H + 2,..., di - 1, di, -1]; For the output sequence, the following operations are performed: Extract the waiting time of the current job from the job waiting time set W as the output sequence Yi, and Yi = wi; Summarize all input sequences and output sequences to form the input sequence set X and the output sequence set Y respectively. The input sequence set X = [XH, XH+1,..., Xn], and the output sequence set Y = [YH, YH+1,..., Yn]; S203. Define the total number of samples of the input sequence set X and the output sequence set Y as N, and N = n - H + 1; divide the data set according to the following ratio: the training set N_train is 80%, the validation set N_val is 10%, and the test set N_test is 10%; the specific division quantities are: N_train = ⌊0.8×N⌋, N_val = ⌊0.1×N⌋, N_test = N - N_train - N_val, where ⌊⌋ represents rounding down; according to the calculated division quantities, divide the training set, validation set, and test set from the input sequence set X and the output sequence set Y.
[0022] The FELSTM model structure in step S300 includes an input layer, a feature encoding layer, multiple LSTMs, a fully connected layer, and an output layer; The input layer receives the input sequence Xi, and the dimension of the input sequence Xi is (N, H, F), where F represents the feature dimension; the feature encoding layer encodes the input sequence Xi to expand its dimension; the multiple LSTMs construct a long short-term memory LSTM network. The state update process of the LSTM is as follows: Input gate: i t = σ(W i · [h t-1 , x t + b i ), where i t represents the activation value of the input gate, which determines the influence of the current input on the cell state; σ represents the Sigmoid activation function, which maps the input to between 0 and 1; W i represents the weight matrix of the input gate, which is responsible for linearly transforming the previous hidden state and the current input; [h t-1 , x trepresents the hidden state h connecting the previous time step t-1 and the input x at the current time step t as a vector, denoted as a combined input; b i represents the bias vector of the input gate, used to adjust the activation value of the input gate and increase the flexibility of the model; Forget gate: f t =σ(W f ·[h t-1 ,x t +b f ), where f t is the activation value of the forget gate, controlling the degree to which the cell state of the previous time step is retained in the current state, with a value range between 0 and 1; W f represents the weight matrix of the forget gate, responsible for linearly transforming the previous hidden state and the current input; b f represents the bias vector of the forget gate, used to adjust the activation value of the forget gate and enhance the flexibility of the model; Output gate: o t =σ(W o ·[h t-1 ,x t +b o ), where o t is the activation value of the output gate, controlling the degree to which the current cell state affects the current hidden state, with a value range between 0 and 1; W o represents the weight matrix of the output gate, responsible for linearly transforming the previous hidden state and the current input; b o represents the bias vector of the output gate, used to adjust the activation value of the output gate and enhance the flexibility of the model; Cell state: C ~ t =tanh(W C ·[h t-1 ,x t +b C ), C t =f t ·C t-1 +i t ·C ~ t ; where C ~ t represents the candidate cell state at the current time step; tanh represents the hyperbolic tangent activation function, mapping the input to between -1 and 1, providing a non-linear transformation; W C represents the weight matrix of the cell state, responsible for linearly transforming the previous hidden state and the current input; b C represents the bias vector of the cell state, used to adjust the calculation of the candidate cell state; C tRepresents the cell state at the current time step, combining the influence of the previous state and the current input; The fully connected layer passes the output of the multi-layer LSTM to the fully connected layer to generate the final output; the output layer directly uses the output of the fully connected layer as the final prediction value.
[0023] The feature encoding layer encodes the input sequence Xi to expand its dimension, specifically including: Z t =tanh(W z ·x t +b z ), where Z t represents the output after feature encoding, W z represents the weight matrix, b z represents the bias term; The fully connected layer passes the output of the multi-layer LSTM to the fully connected layer to generate the final output, specifically including: y pred =W out ·h T +b out , where y pred represents the output of the final prediction; W out represents the weight matrix of the output layer; h T represents the hidden state at the last time step, which contains the information of the entire input sequence; b out represents the bias vector of the output layer.
[0024] The fully connected layer can meet the requirements of the feature encoding layer, but there will also be problems if only the fully connected layer is used. When processing input data, the fully connected layer multiplies the weight matrix by the input data matrix and then adds the bias term to obtain the output data. However, during the training process, the sizes of the weight matrix and the bias term will change continuously, which will cause the mean value of the output data not to be zero and the value range to be unrestricted. This situation will slow down the convergence speed of the optimization algorithm and is prone to the problem of gradient explosion. To solve this problem, the hyperbolic tangent (tanh) activation function is added after the fully connected layer to make the output data maintain a mean value of zero and the value range is limited to [-1,1], so as to accelerate the convergence speed of the optimization algorithm and prevent gradient explosion.
[0025] The loss function set in step S300 is SmoothL1Loss, and the specific content includes: For each job i, calculate the loss L pred between the predicted value y i and the true value, and the true value is the output sequence Yi; the specific calculation formula is: , And the total loss is the average value L of all job losses total .
[0026] In step S300, the Adam optimizer is used for parameter optimization. The specific content includes: In the implementation of the Adam optimizer, the update formulas for the first moment and the second moment are as follows: Update of the first moment: m t =β 1 m t-1 +(1 - β 1 )∇L total , where m t represents the first moment at the current time step; β 1 represents the decay rate of the first moment, usually taking a value of 0.9; m t-1 represents the first moment at the previous time step; ∇L total represents the gradient of the current loss function with respect to the parameters; Update of the second moment: v t =β 2 v t-1 +(1 - β 2 )(∇L total ) 2 , where v t represents the second moment at the current time step; β 2 represents the decay rate of the second moment, usually taking a value of 0.999; v t-1 represents the second moment at the previous time step; The parameter update formula is: θ new =θ old -(αm t ) / [(v t )^(1 / 2)+ϵ], where θ new represents the updated parameter value, θ old represents the parameter value before update, α represents the learning rate, and ϵ represents a constant.
[0027] The specific content of hyperparameter optimization for the FELSTM model in step S400 includes: Define the hyperparameters as the input step size H, the learning rate α, the number of hidden layer nodes N h , the number of LSTM layers N L and the number of feature encoding nodes N f ; Select the grid search method for hyperparameter optimization. For each set of parameter combinations (H, α, N h , N L , N f ), perform the following operations: Construct and train the FELSTM model using different hyperparameter combinations, and record the loss function L during each training process totalValue; For each combination of hyperparameters, calculate the average loss L on the validation set avg , and select the combination of hyperparameters that minimizes the average loss L avg on the validation set as the final hyperparameters.
[0028] In this embodiment, job data in the general computing queue during the period from January 1, 2022 to April 30, 2023 was collected, and the original data set had a total of 2,987,315 job data records.
[0029] First, clean the 2,987,315 job data records, removing jobs that exited abnormally during waiting and jobs with conditional dependencies. Finally, 1,918,041 valid job data records are obtained. Then, for each job data record, collect the following 10-dimensional data: job ID, number of resources requested by the user (number of CPU cores), number of idle resources of the user, number of resources waiting for the user, number of jobs waiting for the user, running time of the user's job, CPU usage time of the user, number of idle resources in the queue, number of resources waiting in the queue, and job waiting time, as Figure 2 shown.
[0030] To analyze data with different dimensions, use the Min - Max method to normalize the data set, and the results are as Figure 3 shown.
[0031] Divide the divided input sequence and output sequence into training data set, validation data set and test data set, with their proportions being 80%, 10% and 10% respectively; when the parameter H is 7, the schematic diagrams of the input sequence and output sequence for data division are as Figure 4 shown, and each row in the schematic diagram of the input sequence results is used as the input of the FELSTM model, and its corresponding target output is the corresponding row in the schematic diagram of the output sequence results.
[0032] On the NVIDIA V100 GPU, use PyTorch to implement the construction of the FELSTM model, and perform model training and testing. The overall structure diagram of the FELSTM model is as Figure 5 shown. When constructing the FELSTM model, the loss function used is the SmoothL1Loss function, and the Adam optimizer is used to optimize the model parameters;
[0033] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0034] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A cluster job waiting time prediction method based on FELSTM model, characterized by: The method comprises the following steps: Step S100: Collect historical operation data in the target computer system and normalize the historical operation data; Step S200. Divide the normalized historical operation data into an input sequence and an output sequence, wherein the input sequence is data used to predict the output sequence, and the output sequence is a target sequence that the model expects to output; divide the divided input sequence and output sequence into a training data set, a validation set, and a test set; Step S300. Build a FELSTM model through a deep learning framework, set the loss function to SmoothL1Loss, and use the Adam optimizer for parameter optimization; Step S400. Optimize the hyperparameters of the FELSTM model so that the loss function of the model converges stably and reaches the minimum value during the training process; after the training is completed, verify the model performance and evaluate the prediction accuracy on the test set.
2. The cluster job waiting time prediction method based on the FELSTM model according to claim 1 is characterized in that: The historical job data includes job ID, number of resources applied by users, number of idle resources by users, number of resources waiting by users, number of jobs waiting by users, user job running time, user CPU usage time, number of idle resources in queues, number of resources waiting in queues, and job waiting time; The job ID refers to the unique identifier assigned by the system when a job is submitted; the number of user-applied resources refers to the number of CPU cores applied for by the user when a job is submitted; the number of user idle resources refers to the upper limit of the number of cores used by the general computing queue minus the number of cores already used by the user when a job is submitted; the number of user waiting resources refers to the total number of cores that the user is waiting to run jobs on when a job is submitted; the number of user waiting jobs refers to the number of jobs that the user is waiting to run on when a job is submitted; the user job running time refers to the total running time that the user has run jobs on when a job is submitted, in seconds; the user CPU usage time refers to the total CPU time that the user has run jobs on when a job is submitted, in seconds; the number of queue idle resources refers to the number of cores available for use in the queue when a job is submitted; the number of queue waiting resources refers to the total number of cores that the queue is waiting to run on when a job is submitted; the job waiting time refers to the time required for a job to be submitted to run, in seconds.
3. The cluster job waiting time prediction method based on the FELSTM model according to claim 1, characterized in that: In step S100, the specific process of normalizing the historical operation data is as follows: For historical operation data, the Min-Max method is used to normalize the historical operation data. The calculation formula is as follows: y'=(yi-y_min) / (y_max-y_min), Among them, y' represents the normalized historical operation data, yi represents the historical operation data in the original sample, y_max represents the maximum value of the corresponding historical operation data in the original sample, and y_min represents the minimum value of the corresponding historical operation data in the original sample.
4. The cluster job waiting time prediction method based on the FELSTM model according to claim 2 is characterized in that: The step S200 includes: S201. Arrange the normalized historical job data into a historical job data set D, where D={d1,d2,...,dn}, where d1 represents the feature vector corresponding to the first historical job, d2 represents the feature vector corresponding to the second historical job, and so on, dn represents the feature vector corresponding to the nth historical job, and n represents the data number of the historical job, which is a positive integer; extract the job waiting time from the historical job data set D, and form a job waiting time set W, where W={w1,w2,...,wn}, similarly, w1 represents the job waiting time corresponding to the first historical job, w2 represents the job waiting time corresponding to the second historical job, and wn represents the job waiting time corresponding to the nth historical job; S202. Define the input step length H, which represents the number of consecutive jobs contained in each input sequence; divide the input sequence and the output sequence, where the following operations are performed on the input sequence: For each job i, from H to n, generate an input sequence Xi, and Xi=[di-H+1,di-H+2,...,di], set the waiting time of the current job to -1, and the corresponding input sequence Xi=[di-H+1,di-H+2,...,di-1,di,-1]; The following operations are performed on the output sequence: extract the waiting time of the current job from the job waiting time set W as the output sequence Yi, and Yi=wi; Summarize all input sequences and output sequences to form input sequence set X and output sequence set Y respectively, where input sequence set X=[XH,XH+1,...,Xn] and output sequence set Y=[YH,YH+1,...,Yn]; S203. Define the total number of samples of the input sequence set X and the output sequence set Y as N, and N=n-H+1; divide the data set into the following proportions: the training set N_train is 80%, the validation set N_val is 10%, and the test set N_test is 10%; the specific number of divisions is: N_train=⌊0.8×N⌋, N_val=⌊0.1×N⌋, N_test=N-N_train-N_val, where ⌊⌋ represents rounding down; divide the input sequence set X and the output sequence set Y into a training set, a validation set, and a test set according to the calculated number of divisions.
5. The cluster job waiting time prediction method based on the FELSTM model according to claim 4 is characterized in that: The FELSTM model structure in step S300 includes an input layer, a feature encoding layer, a multi-layer LSTM, a fully connected layer, and an output layer; The input layer receives an input sequence Xi, and the dimension of the input sequence Xi is (N, H, F), where F represents the feature dimension; the feature encoding layer encodes the input sequence Xi to expand its dimension; the multi-layer LSTM constructs a long short-term memory LSTM network, and the state update process of the LSTM is as follows: Input gate: i t =σ(W i ·[h t-1 ,x t ]+b i ), where i t represents the activation value of the input gate; σ represents the Sigmoid activation function; W i represents the weight matrix of the input gate; [h t-1 ,x t ] represents the hidden state h of the previous time step t-1 and the input x at the current time step t A vector of , represented as a combined input; b i represents the bias vector of the input gate; Forget gate: f t =σ(W f ·[h t-1 ,x t ]+b f ), where f t The activation value of the forget gate; W f represents the weight matrix of the forget gate, which is responsible for linearly transforming the previous hidden state and the current input; b f Represents the bias vector of the forget gate; Output gate: o t =σ(W o ·[h t-1 ,x t ]+b o ), where o t Represents the activation value of the output gate; W o represents the weight matrix of the output gate, which is responsible for linearly transforming the previous hidden state and the current input; b o Represents the bias vector of the output gate, which is used to adjust the activation value of the output gate to enhance the flexibility of the model; Cell state: C ~ t =tanh(W C ·[h t-1 ,x t ]+b C ), C t =f t ·C t-1 +i t ·C ~ t ; where C ~ t represents the candidate cell state at the current time step; tanh represents the hyperbolic tangent activation function; W C The weight matrix representing the cell state; b C Bias vector representing the cell state; C t Represents the cell state at the current time step; Hidden state: h t =ot·tanh(C t ), where h t represents the hidden state of the current time step; The fully connected layer passes the output of the multi-layer LSTM to the fully connected layer to generate the final output; the output layer directly uses the output of the fully connected layer as the final prediction value.
6. The cluster job waiting time prediction method based on the FELSTM model according to claim 5 is characterized by: The feature encoding layer encodes the input sequence Xi to expand its dimension, specifically including: Z t =tanh(W z ·x t +b z ), where Z t represents the output after feature encoding, W z represents the weight matrix, b z represents the bias term; The fully connected layer passes the output of the multi-layer LSTM to the fully connected layer to generate the final output, specifically including: y pred =W out ·h T +b out , where y pred Represents the final predicted output; W out represents the weight matrix of the output layer; h T represents the hidden state of the last time step, which contains the information of the entire input sequence; b out Represents the bias vector of the output layer.
7. The cluster job waiting time prediction method based on the FELSTM model according to claim 6 is characterized by: The loss function in step S300 is set to SmoothL1Loss, and the specific contents include: For each job i, calculate the predicted value y pred The loss L between the actual value i , the true value is the output sequence Yi; the specific calculation formula is: , And the total loss is the average value of all operation losses L total .
8. The cluster job waiting time prediction method based on the FELSTM model according to claim 7 is characterized in that: In step S300, the Adam optimizer is used to perform parameter optimization, and the specific contents include: In the implementation of the Adam optimizer, the update formulas for the first-order moment and the second-order moment are as follows: First-order moment update: m t =β1m t-1 +(1-β1)∇L total , where m t represents the first-order moment of the current time step; β1 represents the decay rate of the first-order moment; m t-1 represents the first-order moment of the previous time step; ∇L total Represents the gradient of the current loss function with respect to the parameters; Second-order moment update: v t =β2v t-1 +(1-β2)(∇L total ) 2 , where v t represents the second-order moment of the current time step; β2 represents the decay rate of the second-order moment; v t-1 represents the second-order moment of the previous time step; The parameter update formula is: θ new =θ old -(αm t ) / [(v t )^(1 / 2)+ϵ], where θ new represents the updated parameter value, θ old represents the parameter value before update, α represents the learning rate, and ϵ represents a constant.
9. The cluster job waiting time prediction method based on the FELSTM model according to claim 8, characterized in that: The specific contents of the hyperparameter optimization of the FELSTM model in step S400 include: Define the hyperparameters as input step size H, learning rate α, number of hidden layer nodes N h , LSTM layer number N L And the number of feature encoding nodes N f ; Select the grid search method for hyperparameter optimization, for each set of parameter combinations (H, α, N h ,N L ,N f ), perform the following operations: construct and train the FELSTM model using different hyperparameter combinations, and record the loss function L during each training process total The value of For each hyperparameter combination, calculate the average loss L on the validation set avg , choose the average loss L on the validation set avg The smallest hyperparameter combination is taken as the final hyperparameter.
Citation Information
Patent Citations
Time sequence prediction method based on time convolution and LSTM
CN108764460A
BiLSTM voltage deviation prediction method based on Bayesian optimization
CN113554148A
Intelligent storm surge forecasting method based on LSTM-GM neural network model
CN113985496A
Time sequence prediction method fusing long short-term memory network and attention mechanism
CN116432697A
Time sequence prediction method and device based on TrAdaBoost-LSTM and medium
CN117371573A