Supercomputer operation duration prediction method based on semantics and time sequence

By improving the BERT architecture and GRU network to understand the semantics and timing information of the job path, the problem of low accuracy of job run time prediction in the prior art is solved, and the utilization rate of supercomputer resources is improved.

CN120372338AActive Publication Date: 2025-07-25CALCULATION AERODYNAMICS INST CHINA AERODYNAMICS RES & DEV CENT
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510884391.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-07-25
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

The prior art ignores semantic information in the job path in predicting the run-time duration of supercomputer jobs, making it difficult to improve the prediction accuracy and affect resource utilization.

Method used

Using an improved BERT-based architecture and GRU network, combining coarse-grained clustering and timing prediction, we understand the semantic information of the job path and capture the timing relationship between jobs, and build a prediction model of job run time.

Benefits of technology

It improves the prediction accuracy of job run time, enhances the resource utilization rate of supercomputer systems, and supports backfill scheduling optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372338A_ABST
    Figure CN120372338A_ABST
Patent Text Reader

Abstract

The invention provides a super computer operation duration prediction method based on semantics and a time sequence, relates to the field of super computers, and solves the problem that semantic information in an operation path and time sequence information between operations are neglected in super computer operation running duration prediction. The method comprises the following steps: firstly, acquiring job log data, distinguishing different user types through data grouping, storing the job log data of users in different data storage modes, and forming a model training set of the users in a coarse-grained clustering mode; then, a prediction model of the operation running time length is constructed in a mode of improving a BERT framework, model training is carried out, and in the training process, a time sequence prediction mode is combined, so that the prediction model predicts the time sequence of a new operation based on time sequence information; and after training is completed, the user submits operation path information of the new operation, and the prediction model determines the operation category of the new operation and outputs the operation running time. According to the method, the prediction accuracy of the operation running time is improved, and subsequent backfill scheduling is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of supercomputers and is applied to the job prediction process. Specifically, it relates to a method for predicting the running time of supercomputer jobs based on semantics and time series. Background Art

[0002] In recent years, with the continuous increase in the scale and complexity of supercomputers, their computing speed has reached billions or even trillions of calculations per second. Due to problems such as tight computing resources and waste of resources in the scheduling system, the improvement of system throughput has been restricted. How to improve the utilization efficiency of system resources has always been a research hotspot. A job refers to a task or program submitted by a user to a supercomputer system for execution, and these programs usually require a large amount of computing resources. In a supercomputer cluster, a user can submit jobs through command lines, script files, etc. Job scheduling can improve the utilization rate of hardware resources, optimize the throughput rate of jobs, and greatly reduce system overhead. Common job scheduling strategies include first-come-first-served, shortest-job-first, and round-robin. Currently, there are many studies on various scheduling methods. Supercomputer clusters usually adopt the first-come-first-served scheduling strategy to ensure service fairness. However, this may cause waste of computing resources.

[0003] A large number of jobs are submitted in a supercomputer, and the job management system will schedule these jobs and allocate computing resources. As a popular scheduling strategy, FCFS (first-come-first-served) makes jobs "queue up" and run in the order of submission. However, under the FCFS strategy, if the existing idle resources cannot meet the requirements of the job at the head of the waiting queue, they cannot be allocated to other jobs, resulting in waste of computing resources. As Figure 1 shown, the horizontal squares represent the running time Time of jobs, and the vertical squares represent system resources, such as the available processors Processors. Among them, jobs 1 and 2 are running jobs, and jobs 3, 4, and 5 are jobs waiting to run, that is, the waiting queue. Under the FCFS strategy, Figure 1 the remaining idle resources of job 2 cannot meet the requirements of job 3, and job 3 can only wait for job 2 to complete and release resources before being allocated to it, resulting in waste of idle resources during the running of job 2.

[0004] To optimize resource utilization, the following has been proposed, such as Figure 2The backfilling method shown allocates the reserved idle computing nodes to a small-scale, short-running non-head job without delaying the original head job, allowing the waiting jobs behind to run "in the gaps" to improve resource utilization. That is, when Job 2 is running and the idle resources are insufficient for Job 3, Job 5 is advanced to run "in the gaps" to improve resource utilization. The backfilling method requires obtaining the running time of the job in advance, while the running duration of the job in the traditional method is provided by the user. Usually, in order to prevent the job from timing out and being cancelled by the system, the user will provide an overestimated running duration of the job. When the running duration of the job is overestimated, it is not conducive to the system to perform job backfilling scheduling, still resulting in the phenomenon of low resource utilization, thus limiting the wide application of backfilling scheduling. Therefore, an algorithm that can accurately predict the running duration of the job is needed to solve this limitation problem.

[0005] The job path is the folder path where the code is located when the user submits the job. The user usually names the job path according to the user name, software name, project name, algorithm name, parameter information, etc. Therefore, the job path contains rich job semantic information. That is, the more similar the job paths are, the more similar the content expressed by the jobs is, and the more similar the running durations of the jobs usually are. Many studies extract features from the historical job logs and use machine learning algorithms to predict the running duration of the job. However, this type of method can only input various numerical information such as user number, submission time, and requested CPU number into the model, ignoring the semantic information contained in the job path, resulting in it being difficult to continue improving the accuracy of predicting the running duration of the job.

[0006] In the current specific operation, a large amount of historical log data is used to train the model. The model extracts different dimensional characteristics of the job based on the "general job log information" of the job, that is, the user number, submission time, and requested processor (CPU) number, etc., so as to predict the running time of the newly submitted job. The system then performs better job scheduling based on the predicted running duration of the job. The existing methods for predicting the running duration of the job mainly fall into two categories. One category is to analyze the code file or program compilation file of the job, and predict the running duration of the job by analyzing the running time of each function in the code and the input / output (I / O) time. However, usually the scheduling platform of the supercomputing center does not have the permission to obtain the user's code file. Therefore, this method is more commonly used in the user's compiler or system-level tools. The other category starts from historical jobs and looks for similar jobs in the historical jobs because similar jobs usually have similar running times. Therefore, the prediction of the running duration of the jobs in the supercomputing center is usually based on the analysis of historical jobs.

[0007] In the prior art, Zhou Longfang et al. published a paper titled "Hierarchical Clustering Algorithm for Job Names to Predict Job Running Time" in the Journal of National University of Defense Technology, Vol. 44, No. 5, October 2022; this paper proposed a hierarchical clustering algorithm for job names of letter-structure-number type, which better clusters jobs and thereby improves the accuracy of the prediction model. Specifically, the job path consists of three types: letters, special characters, and numbers. The author believes that the importance of the three gradually decreases; therefore, it first uses a clustering algorithm to cluster the letters, and then clusters the special characters in each category to obtain different "structures", and then clusters according to the number of digits in each category to obtain the final categories. After three rounds of clustering, the data is divided into thousands of categories. The data in the same category is considered to have similar job paths. The category is used as a new attribute and input into the model together with other attributes to predict the job running time. This technology converts the job path (string) that the model cannot understand into the category serial number (number) that the model can understand through clustering, so that the prediction model can understand the job path information.

[0008] In the prior art, the invention patent with the patent number CN202210132077.9 discloses "A Method for Proactively Predicting Supercomputer Job Failures Based on Application Similarity". This method adds the job path and job name to other features of the job log and also uses a machine learning algorithm to predict the job status; if the predicted job status is a non-success status (abnormal status), corresponding measures are taken for timely remedy to reduce resource waste. Specifically, this method calculates the similarity of the job name and job path respectively using the longest common subsequence similarity and the Levenshtein distance similarity, then performs clustering through a clustering algorithm, and then adds the category information as a new attribute to the data. After filtering out meaningless data, it uses a coarse-grained model and a fine-grained model to predict the job status respectively. The coarse-grained model inputs all the data into the model for training, and the fine-grained model is a model independently established for each user, and the data of each user is input into the corresponding user model respectively.

[0009] It can be seen from this that the prior art focuses on predicting the job running time around machine learning models and clustering technologies, and it is difficult to further improve the prediction accuracy. Therefore, it is necessary to explore new job prediction methods. Summary of the Invention

[0010] Based on the current situation in the background technology, the purpose of the present invention is to solve the problem of ignoring the semantic information in the job path in the prediction of the running duration of supercomputer jobs, and to explore the way to use pre-trained models in the field of natural language processing (especially the bidirectional encoder BERT based on the Transformer) to understand the semantic information in job logs. Therefore, a method for predicting the running duration of supercomputer jobs based on semantics and time series is proposed. First, the present invention separates different batch jobs through clustering and BERT, and uses the semantic understanding ability of BERT to obtain the running duration interval of jobs in a coarse-grained manner. Then, relevant technologies of the gated recurrent unit (GRU) are used to train new jobs using the time series of historical data, so as to realize the prediction of the running duration of jobs. The present invention further improves the prediction accuracy of the running duration of jobs, facilitates the subsequent backfill scheduling of jobs, and thus improves the data utilization rate of the supercomputer system.

[0011] The present invention adopts the following technical solutions to achieve the purpose: A method for predicting the running duration of supercomputer jobs based on semantics and time series, comprising the following steps: S1. Obtain the job log data of historical records and perform data preprocessing. After preprocessing, perform data grouping on the job log data, and distinguish different user types through data grouping; the job log data includes job path information; S2. Corresponding to different user types, adopt different data storage methods to store the job log data of this user; through a coarse-grained clustering method, screen and store different batches of batch jobs in the job log data of this user as the model training set; S3. Construct a prediction model for the running duration of jobs by improving the BERT architecture, and use the model training set to train the prediction model; S4. Combine the time series prediction method during the training of the prediction model, extract the time series information of the existing jobs in the model training set, and enable the prediction model to predict the time series of new jobs based on the time series information; S5. After the training of the prediction model is completed, the user submits the job path information of the new job to the prediction model. The prediction model determines the job category of the new job and outputs the corresponding running duration of the job to complete the job prediction process.

[0012] Specifically, in step S1, data cleaning is performed on the job log data of historical records; the job log data includes multiple job scripts with different job statuses, and each job script includes a script file and at least one job; take the job scripts with the job status of completed, delete the script files therein, and retain the job information therein; subsequently, according to the start time and end time recorded in the job information, calculate the job running duration of the job and store it in a list. The list also correspondingly stores the job ID, user ID, resource requirements, job path, and job name of the job, completing the preprocessing of the job log data.

[0013] Preferably, after the preprocessing of the job log data, different user types are distinguished according to the number of jobs corresponding to each user ID; a first threshold and a second threshold of the preset number of jobs are set, where the first threshold is less than the second threshold; the users with the number of jobs less than the first threshold are classified as the first type of users, the users with the number of jobs greater than or equal to the first threshold and less than the second threshold are classified as the second type of users, and the users with the number of jobs greater than or equal to the second threshold are classified as the third type of users, completing the distinction of user types.

[0014] Preferably, in step S2, for the first type of users, their job log data is deleted, and the job log data of the first type of users will not be used in the model training set; for the second type of users, the job log data of all second type of users is merged and stored as a whole, and the whole is used as a training part in the model training set; for the third type of users, the job log data of each specific user among the third type of users is stored independently, and each is used as a training part in the model training set.

[0015] Furthermore, in the coarse-grained clustering process, for each training part in the model training set, the K-Means clustering algorithm is used to cluster the job running duration column stored in the list, and the column is logarithmically scaled naturally during clustering; the absolute percentage error and further calculation of the weighted absolute percentage error are used to evaluate the clustering effect, and accordingly determine the selection of the evaluation index K value; after clustering, the clustering category of each job in each training part is obtained, which is the category of the batch processing jobs of different batches; the clustering category is stored together with the job path and job name corresponding to the job stored in the list to form the model training set.

[0016] Furthermore, in step S3, the constructed prediction model includes an input layer, a feature extraction layer, a position fusion layer, and an output layer; The input layer receives the input of the job path information in the model training set and generates an embedding vector composed of a word vector, a segment vector, and a position vector; The embedding vector is processed by the feature extraction layer, which extracts the vector encoding information of syntactic features and deep semantic features from the embedding vector by presetting the first encoder and the second encoder at different layers in the improved BERT architecture; The vector encoding information is input into the position fusion layer, and multi-dimensional semantic fusion is used to achieve feature enhancement and generate category probability distribution; The output layer selects the category with the highest confidence from the category probability distribution as the final classification result of semantic fusion, and obtains the job category corresponding to the job path information; based on the obtained job category, the corresponding job running time is output.

[0017] Specifically, in the position fusion layer, the variable-length sequence of vector encoding information is first converted into a fixed-dimensional global feature vector by the pooling layer, which then performs a random neuron masking operation through the Dropout layer, and then maps the vector features to the pre-clustered coarse-grained category space through the linear layer, which is output by the double encoding layer and performs mean fusion, and then generates the category probability distribution through the Softmax layer; The output layer uses the argmax method to select the category with the highest confidence from the category probability distribution.

[0018] Preferably, in step S4, the timing prediction method uses a GRU network to perform timing prediction. For batch jobs of the same batch in the same user, when the prediction model predicts the job running time of a new job, the job categories of a preset number of predicted jobs before the new job are input into the GRU network, the running time of the batch jobs of the batch is extracted, and the timing information of the new job is predicted as a reference for predicting the job category and job running time of the new job.

[0019] Preferably, in step S5, based on the job log data of the historical records, a corresponding model test set is constructed in the same manner as steps S1 and S2; the prediction model is tested using the model test set, and the accuracy deviation between the job running time output by the prediction model and the actual time is calculated; when the accuracy deviation meets the preset training standard, the training of the prediction model is completed.

[0020] Preferably, an online learning mechanism is added to the prediction model after training is completed. In actual applications, whenever a user submits the job path information of a new job, the prediction model adjusts some of its own model parameters to adapt to the job path information based on the data characteristics corresponding to the job path information of the new job, and predicts the corresponding job running time.

[0021] In summary, due to the adoption of this technical solution, the beneficial effects of the present invention are as follows: The present invention first uses large language models represented by BERT to understand the semantic information of job paths; existing technical papers use a clustering algorithm of "letter-structure-number" to obtain categories, and it cannot be shown that this clustering algorithm can understand the semantic information of job paths only through category numbers; existing patents use similarity for clustering, and the clustering results are added to other attributes and put into a prediction model for prediction. The job paths are ultimately converted into category numbers of clustering, and essentially cannot understand the semantic information of job paths. The present invention is based on the improvement of the model with the BERT architecture, and its hierarchical fusion method can be applied to job paths of different types and lengths, converting the job paths into encoded information, and then understanding the semantic information of job paths.

[0022] The present invention not only studies and understands the semantic information in job paths, but also introduces GRU technology to capture the temporal relationship between jobs in view of the phenomenon that the running times of batch jobs in the same batch are extremely similar; this enables the present invention to conduct a comprehensive evaluation from two dimensions of temporality and semantic information, further improving the prediction accuracy of job running duration. In contrast, when existing technologies use machine learning techniques after clustering, they will shuffle the data before training, ignoring the sequential relevance between jobs. Therefore, the present invention is more in-depth in extracting job semantics and also has more advantages in understanding job temporal information. Brief Description of the Drawings

[0023] Figure 1 It is a schematic diagram of resource waste caused by the FCFS strategy in the prior art; Figure 2 It is a schematic diagram of optimizing resource utilization by the backfilling method in the prior art; Figure 3 It is a schematic diagram for briefly describing the overall process of the method of the present invention; Figure 4 It is a schematic diagram for details of the operation process of the method of the present invention; Figure 5 It is a schematic diagram of the structure of the job running duration prediction model in the method of the present invention. Detailed Embodiment

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0025] Accordingly, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts fall within the scope of protection of the present invention.

[0026] Embodiment A method for predicting the job duration of a supercomputer based on semantics and time series Figure 3 A brief overview of the overall process of this method is shown and can be referred to synchronously; the general description of each key step of this method is as follows: S1. Obtain the job log data of the historical record and perform data preprocessing. After preprocessing, perform data grouping on the job log data, and distinguish different user types through data grouping; the job log data includes job path information. S2. Corresponding to different user types, use different data storage methods to store the job log data of this user; through a coarse-grained clustering method, screen and store the batch processing jobs of different batches in the job log data of this user as the model training set. S3. Construct a prediction model for the job running duration by improving the BERT architecture, and use the model training set to train this prediction model. S4. Combine the time series prediction method during the training process of the prediction model, extract the time series information of the existing jobs in the model training set, and enable the prediction model to predict the time series of new jobs based on the time series information. S5. After the training of the prediction model is completed, the user submits the job path information of the new job to the prediction model. The prediction model determines the job category of the new job and outputs the corresponding job running duration, completing the job prediction process.

[0027] In this embodiment, the prediction model for the job running duration constructed by improving the BERT architecture is more suitable for the job running duration prediction task. The BERT architecture itself adopts a multi-head self-attention mechanism, which can effectively capture the dependencies of the input sequence and improve the depth and accuracy of language understanding; among them, the self-attention mechanism generates new representations by calculating the relationships between each element in the input sequence, captures long-distance dependencies, and flexibly adjusts the context feature representations; and the multi-head self-attention mechanism, as an extension of the self-attention mechanism, is applied to each Transformer layer of the BERT architecture, using a multi-layer structure to converge semantic information at different levels and capture rich relationships and potential meanings in the text.

[0028] Next, this embodiment details the details and preferred methods of each step during the method execution through relatively specific numerical examples, etc. The detailed operation process during the method execution can be referred to Figure 4 for illustration.

[0029] First, in step S1, the job log data of the historical records is cleaned. The useless job script information in the job log data needs to be deleted, and only the information of each job at the finest granularity is retained; the job log data includes multiple job scripts with different job statuses, and each job script includes a script file and at least one job.

[0030] In this embodiment, the job scripts with the job status of completed (i.e., the status record is "Completed") are selected, and the script files ending with.batch are deleted from them, and only one or more job information is retained; then, according to the start time and end time recorded in the job information, the job running duration of the job is calculated and stored in a list. The job running duration can be saved in the column named "Seconds" in the list; the job ID, user ID, resource requirements (i.e., the required number of CPUs, etc.), job path, and job name of the job are also stored correspondingly in the list. After concatenating the job path and job name, they can be saved in the column named "Path". In this way, the column names in the list can be: JobID (job ID), UID (user ID), ReqCPUs (resource requirements), Seconds, and Path, thus completing the preprocessing of the job log data.

[0031] For the data grouping process, it aims to distinguish different user types: some users submit few jobs and with low frequency, and the training and research costs are high while the benefits are low; while some users submit many jobs and with high frequency, and the job log data of these jobs can be separated for separate model training, so as to better learn the job characteristics of these users. Therefore, in this embodiment, different user types are distinguished according to the number of jobs corresponding to each user ID; taking 311421 valid data records in the historical records during the training process as an example, the first threshold of 100 and the second threshold of 5000 for the number of jobs are preset. If the data volume is different in actual applications, the thresholds can also be adjusted accordingly to better distinguish user types.

[0032] In this embodiment, users with the number of jobs less than 100 are classified as "software development type" users, users with the number of jobs greater than or equal to 100 and less than 5000 are classified as "scientific exploration type" users, and users with the number of jobs greater than or equal to 5000 are classified as "engineering application type" users. These three types of users will correspond to different data processing measures, and this classification method combined with engineering practice has practical significance; after the classification is completed, the method can enter the step S2 part.

[0033] For "software development type" users, since their job data is very small, the prediction model is not sufficient to learn the relationship between the job path and the job running duration. Therefore, their job log data is deleted, and this part of the data will not be used in the model training set.

[0034] For "science exploration type" users, since their job data is relatively small, the job log data of all users of this type is merged and stored as a whole, which is used as a training part in the model training set.

[0035] For "engineering application type" users, the job log data of each specific user of this type is stored independently, and each is used as a training part in the model training set.

[0036] Subsequently, for the specific jobs of each user, the distribution of their job running durations is extremely uneven, which will lead to unsatisfactory results of traditional density-based clustering methods. Therefore, in the coarse-grained clustering process of this embodiment, for each training part in the model training set, the K-Means clustering algorithm is used to cluster the job running duration columns stored in the list. The coarse-grained clustering process will be used to screen different batches of batch jobs, where jobs in the same batch only differ in parameters and have very similar running times; while jobs in different batches usually have significant differences in running times due to data and test task differences; through the coarse-grained clustering process, clustering the job running durations can determine different job categories and their approximate time intervals, which is convenient for the subsequent prediction model to predict the job running duration of new jobs.

[0037] After the list corresponding to the job log data has been preprocessed and stored in step S1, K-Means clustering can be performed on the "Seconds" column, and the user and job information corresponding to the same row can be adjusted. Since the data distribution of the job running durations recorded in the "Seconds" column is extremely uneven, in this embodiment, natural logarithm scaling is performed on this column during clustering to better complete the clustering process.

[0038] Subsequently, the mean absolute percentage error is adopted and the weighted mean absolute percentage error is further calculated to evaluate the clustering effect, and based on this, the selection of the evaluation index K value is determined. The relevant formulas are as follows:

[0039]

[0040]

[0041] In the formula, represents the predicted value at the th time point, represents the actual value at the th time point, represents the number of jobs in the th category, represents the total number of user jobs in a training section, represents the weight of jobs in class. After clustering, the clustering categories of each job in each training section are obtained, that is, the categories of batch processing jobs represented as different batches, which can be saved in the column named "Class" in the list. Finally, the clustering category is stored together with the job path and job name corresponding to the job stored in the list to form a model training set, that is, including two columns "Path" and "Class", and is used to input into the prediction model of job running duration. In addition to using the K-Means clustering algorithm, other algorithms can also be selected according to actual needs to ensure the corresponding clustering effect.

[0042] In step S3 of this embodiment, the prediction model constructed by improving the BERT architecture extracts information of different dimensions through encoders at different levels. The advantage of its hierarchical fusion is that it can take into account the semantic extraction of job paths of different lengths. The prediction model finally outputs the category of the job, and the job running duration can be determined accordingly.

[0043] As Figure 5 shown, the prediction model constructed in this embodiment includes an input layer, a feature extraction layer, a position fusion layer, and an output layer.

[0044] The input layer receives the job path information in the model training set (which can be processed into example to input vectors) and generates an embedding vector composed of word vector Token Embeddings, segment vector Segment Embeddings, and position vector PositionEmbeddings ( to ).

[0045] The embedding vector is processed by the feature extraction layer. By presetting the first encoder and the second encoder at different layers in the improved BERT architecture, the vector encoding information of syntactic features and deep semantic features is respectively extracted from the embedding vector; in this embodiment, the encoder Encoder 8 at the 8th layer and the encoder Encoder12 at the 12th layer in the improved BERT architecture are respectively used for extraction here. The model layer information here can also be changed according to the naming method, length, etc. of the job path, so the layer position is not fixed; if there are a large number of long paths and medium paths in the training and actual application data, the encoders at the 8th layer and the 12th layer can be selected for processing according to the method of this embodiment.

[0046] In the vector encoding information input position fusion layer, multi-dimensional semantic fusion is performed to enhance features and generate a class probability distribution. In the position fusion layer of this embodiment, first, the pooling layer Pooler converts the variable-length sequence of vector encoding information into a fixed-dimensional global feature vector, which then performs a random neuron masking operation through the dropout layer to improve the generalization ability of the prediction model. Immediately afterwards, the linear layer Linear maps the vector features to the pre-clustering coarse-grained class space. After the output of the double-encoding layer and the execution of mean fusion, the syntactic features and the deep semantic features respectively form to and to vectors (the number of fused vectors changes from the original to ). Finally, through the normalized exponential function layer, that is, the Softmax layer, a class probability distribution is generated.

[0047] The output layer uses the maximum value indexing method (argmax method) to select the class with the highest confidence from the class probability distribution as the final classification result of semantic fusion, and obtains the job class corresponding to the job path information. According to the obtained job class, the corresponding job running duration is output.

[0048] As an optimization of this embodiment, in step S4, the time series prediction method uses a GRU network for time series prediction. For the batch processing jobs of the same batch of the same user, when the prediction model predicts the job running duration of a new job, it inputs the job classes of the predicted jobs with a preset number (for example, the first n pieces of job-related data) before the new job into the GRU network, extracts the running time series of the batch processing jobs of this batch, and predicts the time series information of the new job as a reference for predicting the job class and job running duration of the new job.

[0049] In the process of time series prediction, capturing time series data and long-term and short-term dependencies is the core point. The GRU network introduces gated recurrent units, which include an update gate and a reset gate, and can effectively capture long-term dependencies through the gate structure. After applying the GRU network in this embodiment, it can selectively retain and update historical information and determine what to retain and update for new information. This characteristic enables the GRU network to better process long sequence data, maintain stable information transmission, and improve the model performance and generalization ability.

[0050] In addition to using the GRU network for time series prediction, networks that can capture data time series information such as the long short-term memory network (LSTM) can also be used for time series prediction. However, compared with the LSTM network, the GRU network has a more concise structure, reduces the number of parameters, and improves the computational efficiency. Therefore, it is more suitable for processing large-scale data. In this embodiment, the GRU network is preferably used here. By improving the tight combination of the BERT architecture and the GRU network, without BERT-related methods, pure and accurate data of the same batch job types cannot be obtained, and data from different batches may be mixed together. Without the time series prediction of the GRU network, it is difficult to connect the time series information of the job type data of the same class (the same batch) before and after.

[0051] In step S5 of this embodiment, according to the historical job log data, a model test set is correspondingly constructed in the same manner as in steps S1 and S2; the model test set is used to test the prediction model, and the accuracy deviation between the job running duration output by the prediction model and the real time is calculated. , as shown in the following formula:

[0052]

[0053] In the formula, represents the accuracy deviation of the th user, represents the job running duration predicted by the model, represents the real job running duration of the job, represents the total number of users in the model test set. When the accuracy deviation meets the preset training standard, the training of the prediction model is completed.

[0054] As an optimization of this embodiment, an online learning mechanism is added to the prediction model after training. In actual application, whenever the job path information of a new job is submitted by a user, the prediction model adjusts some of its own model parameters according to the data characteristics corresponding to the job path information of the new job to adapt to the job path information, and predicts and outputs the corresponding job running duration.

[0055] The online learning mechanism enables the model to not only learn the rules from historical data during the initial training, but also automatically adjust and improve itself according to the newly emerging data after deployment. For the scenario of predicting the job running duration in this embodiment, when the job types are continuously updated, the prediction model will not become outdated or inaccurate due to long-term non-update. Since the online learning method is fine-tuned based on the latest submitted job data, it can help the model capture short-term change trends, such as changes in system resource usage within a specific time period, etc., so as to provide a more accurate prediction of the job running duration.

[0056] The online learning process can also be simply achieved through a feedback loop. After each prediction of the job running duration is completed, the corresponding actual job running duration is also recorded synchronously and compared with the prediction result. The prediction model can adjust its internal parameters by combining methods such as online gradient descent and passive-aggressive learning, so as to show stronger adaptability in a dynamic job environment.

Claims

1. A method for predicting the job duration of a supercomputer based on semantics and time series, characterized in that, It includes the following steps: S1. Obtain the job log data of historical records and perform data preprocessing. After preprocessing, perform data grouping on the job log data, and distinguish different user types through data grouping; The job log data includes job path information; S2. Corresponding to different user types, use different data storage methods to store the job log data of this user; through a coarse-grained clustering method, screen and store the batch processing jobs in different batches in the job log data of this user as the model training set; S3. Construct a prediction model for job running duration by improving the BERT architecture, and use the model training set to train this prediction model; S4. Combine the time series prediction method during the training of the prediction model, extract the time series information of the existing jobs in the model training set, and enable the prediction model to predict the time series of new jobs based on the time series information; S5. After the training of the prediction model is completed, the user submits the job path information of the new job to the prediction model. The prediction model determines the job category of the new job and outputs the corresponding job running duration, completing the job duration prediction process.

2. The supercomputer job duration prediction method according to claim 1, characterized in that: In step S1, perform data cleaning on the job log data of historical records; the job log data includes multiple job scripts with different job statuses, and each job script includes a script file and at least one job; select the job scripts with the job status of completed, delete the script files therein, and retain the job information; then, according to the start time and end time recorded in the job information, calculate the job running duration of this job and store it in a list. The job ID, user ID, resource requirements, job path, and job name of this job are also stored correspondingly in the list, completing the preprocessing of the job log data.

3. The supercomputer job duration prediction method according to claim 2, wherein: After the preprocessing of the job log data, distinguish different user types according to the number of jobs corresponding to each user ID; preset the first threshold and the second threshold of the number of jobs, where the first threshold is less than the second threshold; classify the users with the number of jobs less than the first threshold as the first type of users, classify the users with the number of jobs greater than or equal to the first threshold and less than the second threshold as the second type of users, and classify the users with the number of jobs greater than or equal to the second threshold as the third type of users, completing the distinction of user types.

4. The supercomputer job duration prediction method according to claim 3, wherein: In step S2, for the first type of users, delete their job log data, and the model training set will not use the job log data of the first type of users; for the second type of users, merge the job log data of all second type of users into a whole for storage, and the whole is used as a training part in the model training set; for the third type of users, store the job log data of each specific user in the third type of users independently, and each is used as a training part in the model training set.

5. The supercomputer job duration prediction method according to claim 4, characterized in that: In the coarse-grained clustering process, for each training part in the model training set, the K-Means clustering algorithm is used to cluster the job running time column stored in the list, and the column is scaled by natural logarithm during clustering; the absolute percentage error is used and the weighted absolute percentage error is further calculated to evaluate the clustering effect, and the selection of the evaluation indicator K value is determined accordingly; after clustering, the clustering category of each job in each training part is obtained, that is, the category of batch processing jobs represented by different batches; the clustering category is stored together with the job path and job name corresponding to the job stored in the list to form a model training set.

6. The supercomputer job duration prediction method according to claim 1, wherein: In step S3, the constructed prediction model includes an input layer, a feature extraction layer, a position fusion layer and an output layer; The input layer receives the job path information input in the model training set and generates an embedding vector consisting of word vectors, segment vectors, and position vectors; The embedding vector is processed by the feature extraction layer, which extracts the vector encoding information of syntactic features and deep semantic features from the embedding vector by presetting the first encoder and the second encoder at different layers in the improved BERT architecture; The vector encoding information is input into the position fusion layer, and multi-dimensional semantic fusion is used to achieve feature enhancement and generate category probability distribution; The output layer selects the category with the highest confidence from the category probability distribution as the final classification result of semantic fusion, and obtains the job category corresponding to the job path information; based on the obtained job category, the corresponding job running time is output.

7. The supercomputer job duration prediction method according to claim 6, wherein: In the position fusion layer, the variable-length sequence of vector encoding information is first converted into a fixed-dimensional global feature vector by the pooling layer, which then performs a random neuron masking operation through the Dropout layer, and then maps the vector features to the pre-clustered coarse-grained category space through the linear layer. After the double encoding layer outputs and performs mean fusion, the category probability distribution is generated through the Softmax layer; The output layer uses the argmax method to select the category with the highest confidence from the category probability distribution.

8. The supercomputer job duration prediction method according to claim 6, wherein: In step S4, the timing prediction method uses the GRU network to perform timing prediction. For batch jobs of the same batch in the same user, when the prediction model predicts the job running time of a new job, the job categories of a preset number of predicted jobs before the new job are input into the GRU network, the running time of the batch jobs of the batch is extracted, and the timing information of the new job is predicted as a reference for predicting the job category and job running time of the new job.

9. The supercomputer job duration prediction method according to claim 1, wherein: In step S5, based on the job log data in the historical records, a model test set is constructed in the same manner as steps S1 and S2; the prediction model is tested using the model test set, and the accuracy deviation between the job running time output by the prediction model and the actual time is calculated; when the accuracy deviation meets the preset training standard, the training of the prediction model is completed.

10. The supercomputer job duration prediction method according to claim 9, characterized in that: An online learning mechanism is added to the prediction model after training is completed. In actual applications, whenever the job path information of a new job is submitted by a user, the prediction model adjusts some of its own model parameters according to the data characteristics corresponding to the job path information of the new job to adapt to the job path information, and predicts and outputs the corresponding job running duration.

Citation Information

Patent Citations

  • A proactive prediction method for supercomputer job failures based on application similarity

    CN114169651B

  • Power failure maintenance operation duration prediction method and device

    CN117035146A

  • Method for predicting job running time based on job name hierarchical clustering algorithm

    CN117520118A

  • GPT-based directional label public opinion analysis method and corresponding device

    CN119740137A

  • Method and device for multi-pass human-machine conversation based on time sequence feature screening and encoding module

    KR102610897B1