Method and device for predicting operation time of computer, electronic equipment and storage medium
By obtaining the node marking information and actual running time in the job marking library, the predicted value of the computer job running to the key progress node is calculated, which solves the problems of strong data dependence and complex prediction in the prior art, and achieves efficient and accurate prediction of the job run time.
Patent Information
- Application Number
- CN202411963005.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art requires a large number of historical data samples for training when predicting the running time of computer jobs, and the process is complicated and not efficient enough.
By obtaining node marking information of the same type of job in the job marking library, and combining the actual running time, the running time prediction value when the job runs to the critical progress node is calculated. This method simplifies data dependency and improves prediction accuracy and efficiency.
It realizes the simple and efficient prediction of computer job running time without relying on a large amount of historical data, improving the accuracy and efficiency of prediction.
Smart Images

Figure CN120045427A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer job processing, and in particular, to a method, device, electronic device, and storage medium for predicting the running time of computer jobs. Background Art
[0002] With the increasing complexity of computer job processing requirements and the development of cluster computing technology, the prediction of the running time of computer jobs has become the key to improving resource scheduling efficiency, optimizing job scheduling, and reducing operating costs. Especially in large clusters or distributed computing environments, the running time of jobs is often affected by various factors. Therefore, accurately predicting the running time of jobs is of great significance for reasonably allocating computing resources, reducing job queuing waiting time, and improving computing efficiency.
[0003] In related technologies, some job running time prediction methods mainly rely on historical data and empirical models, and use prediction algorithms based on statistics or machine learning. However, these methods usually require a large number of historical data samples for training, and the process is relatively complex. Summary of the Invention
[0004] In view of this, embodiments of the present application provide a method, device, electronic device, and storage medium for predicting the running time of computer jobs, which are convenient for simply and efficiently predicting the running time of jobs.
[0005] To achieve the above invention objective, the following technical solutions are adopted: In a first aspect, an embodiment of the present application provides a method for predicting the running time of a computer job, including: obtaining node marking information of jobs of the same type as a first job in a job marking library, where the node marking information includes the proportion of the running time of jobs of the same type as the first job at a first processing progress node in the overall running time of the current time; obtaining the actual running time of the first processing progress node during the running process of the first job; and calculating and determining a predicted value of the overall running time of the current time when the first job runs to the first processing progress node according to the actual running time of the first processing progress node and the proportion of the running time at the first processing progress node in the overall running time of the current time.
[0006] In combination with the first aspect, in a first implementation manner of the first aspect, while obtaining the actual running time of the first processing progress node during the running process of the first job, the method further includes: obtaining the unique identity identifier of the first job, the configuration information identifier of the first job, and the first processing progress node identifier information; and forming marking information of the first processing progress node of the first job by using the unique identity identifier of the first job, the configuration information identifier of the first job, the first processing progress node identifier information, and the actual running time of the first processing progress node, and storing the marking information in the job marking library.
[0007] Combined with the first aspect and the first implementation manner of the first aspect, in the second implementation manner of the first aspect, after forming the marking information of the first processing progress node of the first job and storing it in the job marking library, the method further includes: after monitoring that the first job has finished running, determining whether the first job has run successfully; if it has not run successfully, canceling the marking information of the first processing progress node of the first job in the job marking library; if it has run successfully, storing the proportion of the running time of the first job at the first processing progress node in the actual value of the overall running time this time.
[0008] Combined with the first aspect and the first or second implementation manner of the first aspect, in the third implementation manner of the first aspect, while obtaining the actual running time of the first processing progress node during the running of the first job, the method further includes: obtaining the host identifier of the host running the first job and the host load information at the first processing progress; storing the host identifier and the host load information at the first processing progress in the job marking library and associating them with the marking information of the first processing progress node of the first job.
[0009] Combined with the first aspect and the first or second or third implementation manner of the first aspect, in the fourth implementation manner of the first aspect, the node marking information further includes the proportion of the running time of the job of the same type as the first job at the second processing progress node in the overall running time of that time; after calculating and determining the predicted value of the overall running time this time when the first job runs to the first processing progress node, the method further includes: obtaining the actual running time of the second processing progress node during the running of the first job; calculating and determining the predicted value of the overall running time this time when the first job runs to the second processing progress node according to the actual running time of the second processing progress node and the proportion of the running time at the second processing progress node in the overall running time of that time.
[0010] Combined with the first aspect and the first, second, third, or fourth implementation manner of the first aspect, in the fifth implementation manner of the first aspect, while obtaining the actual running time of the second processing progress node during the running of the first job, the method further includes: obtaining the unique identity identifier of the first job, the configuration information identifier of the first job, and the second processing progress node identifier information; and obtaining the host identifier of the host on which the first job runs and the host load information at the second processing progress; forming the marking information of the second processing progress node of the first job by using the unique identity identifier of the first job, the configuration information identifier of the first job, the second processing progress node identifier information, and the actual running time of the second processing progress node, and storing it in the job marking library; and storing the host identifier and the host load information at the second processing progress in the job marking library and associating them with the marking information of the second processing progress node of the first job.
[0011] Combined with the first aspect and the first, second, third, fourth, or fifth implementation manner of the first aspect, in the sixth implementation manner of the first aspect, the calculation of determining the predicted value of the current overall running time when the first job runs to the first processing progress node includes: dividing the actual running time of the first processing progress node by the proportion of the running time at the first processing progress node in the current overall running time, to obtain the predicted value of the current overall running time when the first job runs to the first processing progress node.
[0012] Combined with the first aspect and the first, second, third, fourth, fifth, or sixth implementation manner of the first aspect, in the seventh implementation manner of the first aspect, the method is applicable to the running time prediction of single-machine jobs and the running time prediction of computing cluster batch jobs; wherein, for the running time prediction of computing cluster batch jobs, after calculating and determining the predicted value of the current overall running time when each job runs to the first processing progress node, it further includes: comparing the magnitudes of the current overall running times among the jobs; and determining the maximum value among the predicted values of the current overall running times of the jobs as the predicted value of the current overall running time of the computing cluster batch job at the first processing progress node.
[0013] In a second aspect, an embodiment of the present application provides an apparatus for predicting the running time of a computer job, including: a first acquisition program unit configured to acquire node marking information of jobs of the same type as a first job in a job marking library, where the node marking information includes the proportion of the running time of a job of the same type as the first job at a first processing progress node in the overall running time of the current time; a second acquisition program unit configured to acquire the actual running time of the first processing progress node during the running of the first job; and a prediction program unit configured to calculate and determine a predicted value of the overall running time of the current time when the first job runs to the first processing progress node according to the actual running time of the first processing progress node and the proportion of the running time at the first processing progress node in the overall running time of the current time.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, including: one or more processors; a memory; one or more executable programs are stored in the memory, and the one or more processors read the executable program code stored in the memory to run a program corresponding to the executable program code for running any method described in the first aspect.
[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing one or more programs, and the one or more programs can be run by one or more processors to implement any method described in the first aspect. Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0017] Figure 1 It is a schematic flowchart of an embodiment of the method for predicting the running time of a computer job in the present application; Figure 2 It is a schematic flowchart of the collection or update process of node marking information in an embodiment of the present application Figure 3 It is a schematic diagram of an embodiment of the node marking information and host status information stored in the job marking library of the present application; Figure 4 It is a schematic diagram of the computing cluster and job organizational structure in an embodiment of the present application; Figure 5 It is a schematic block diagram of the overall prediction process of the running time of a job from the user side to the cluster management server side in a cluster environment in an embodiment of the present application; Figure 6 Schematic block diagram of an embodiment of the apparatus for predicting the running time of computer jobs in the present application; Figure 7 Schematic structural diagram of an embodiment of an electronic device in the present application. Detailed implementation manners
[0018] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0019] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.
[0020] As described in the background art, with the continuous growth of computer job processing requirements, the prediction of job running time has become an important means to improve resource scheduling and optimize job scheduling. Especially in the distributed computing and cluster job environment, accurately predicting the overall running time of jobs has important practical significance for improving computing resource utilization, reducing job latency, and enhancing system throughput.
[0021] The embodiments of the present application propose a method for predicting the running time of computer jobs, which is applicable to the running time prediction of single-machine jobs and cluster batch jobs. By marking the key progress nodes of the job and the actual running time, the overall running time of the job can be predicted, which can simply and efficiently provide relatively accurate prediction results to a certain extent, and thus provide a reliable basis for computing resource scheduling and job optimization processing.
[0022] Figure 1 Illustrates the method for predicting the running time of computer jobs provided by an embodiment of the present application. Refer to Figure 1 , in some embodiments of the present application, the method for predicting the running time of computer jobs includes: S110. Obtain the node marking information of the jobs of the same type as the first job in the job marking library.
[0023] A large amount of node marking information of jobs is stored in the job marking library, including job configuration information, identity identifiers of key processing progress nodes, running time, and the proportion of the running time at the key processing progress nodes in the overall running time of the current time, etc.
[0024] When predicting the running time of the first job during the operation, the node marking information of the job identical to the first job that needs to be obtained includes at least: the proportion of the running time of the job of the same type as the first job at the first processing progress node in the overall running time of the current run. For example, if the first job executed this time is a chip memory bandwidth performance test job, then it is necessary to obtain from the job marking library the actual running time and the corresponding proportion in the overall running time of the previous chip memory bandwidth performance test job of the same type at the same processing progress node, such as the first processing progress node, after the operation is completed, as the calculation parameters for predicting the overall running time of the first job this time.
[0025] In this embodiment, the formation of the node marking information in the job marking library can be that when the user submits a job, according to information such as the job type and job configuration, the key progress nodes of the job and their corresponding running time proportions are specified and stored in the job marking library in the form of key progress node marks. Or, the key progress nodes and the corresponding proportion of the running time of the current run are extracted from the historical job running data log and stored in the job marking library.
[0026] The key progress node marks formed by these two methods can both be used as the basis for estimating the running time of a new job. For the situation where there is no node marking information of the job of the same type as the first job in the job marking library, it is necessary to collect the node marking information first through the method specified by the user when submitting the job. Specifically, the job key progress node mark can be implemented by the user inserting a predetermined number of key progress node outputs in the script or application program of the submitted job.
[0027] See Figure 3 , the node marking information collected in the job marking library includes but is not limited to: the unique identity (ID) of the job, the job configuration information identifier, the key progress node identifier information, the key progress node running time (it should be noted that Figure 3 shows the estimated running time, which is for the current job. For the previous job, it is the actual running time of the current run of the previous job, in order to be different from Figure 3 the true running time of the current job at the first key processing progress node shown in the right table, and due to the limited writing space in the table, it is abbreviated in this way), the proportion of the running time of the job at the key progress node in the overall running time of the current run ( Figure 3 is expressed as the estimated running time percentage in, and the reason is the same as above), etc.
[0028] To help understand the technical solution provided by the embodiments of the present application, taking a big data processing job as an example, the description is as follows: Assume that for this big data processing job, when the user submits it, a unique identity identifier (ID) job001 is assigned to the job, and the job configuration information identifier is config3.1. The host configuration information in the server cluster required by this job can be marked, including the configuration settings of hardware resources and software resources.
[0029] In addition to the above two parameters, the marked key progress node identification information also includes at least one key progress node. The first progress node 1 is identified as data load (data loading), and this progress runs to load data from the storage system. For the running time of the key progress node, assume that the previous job of the same type, such as both being chip memory test jobs, the actual running time at the data load progress node was 3 minutes, and the proportion of the actual running time of the data load progress node in the overall running time of the current job after the previous job run was 22.7%, which is stored in the job marking library. When this job runs, the proportion of 22.7% of the previous key progress node in the overall running time of the current job is used as a parameter to estimate the overall running time of the current job at the corresponding progress node.
[0030] In some embodiments, the key progress nodes of a job may include the initialization stage, data processing stage, result output stage, etc. of the job. Specific quantified key progress nodes are such as 5%, 15%, 25%, 45%, 60%, 75%, 95%, etc. One or more key progress nodes can be set as needed.
[0031] S120. Obtain the actual running time of the first processing progress node during the running of the first job.
[0032] When the job starts running, the key progress nodes during the job running process are gradually triggered, and the actual running time can be output in the form of a log. Specifically, through the key progress node marks inserted by the user during the job running process, the monitoring module can capture the actual running time of the first job at the key progress nodes.
[0033] S130. Calculate and determine the predicted value of the overall running time of the first job when it runs to the first processing progress node according to the actual running time of the first processing progress node and the proportion of the running time.
[0034] After obtaining the actual running time of the first processing progress node, next, according to the actual running time of the first processing progress node and the proportion of the first processing progress node in the overall running time of the current job, calculate the overall running time of the job for this time.
[0035] In this embodiment, by obtaining the node marking information of jobs of the same type as the first job and combining the key progress node information during actual operation, such as the actual operation time, it is possible to estimate a predicted value of the overall operation time of the first job for this time when the job runs to the first processing progress node. Since the calculation and estimation in the prediction process are partially based on the actual operation time data, the dependence on a large amount of historical data is effectively avoided, thereby simply and efficiently improving the accuracy of job operation time prediction.
[0036] Specifically, calculating and determining the predicted value of the overall operation time of the first job for this time when it runs to the first processing progress node according to the actual operation time of the first processing progress node and the proportion of the operation time at the first processing progress node in the overall operation time of this time includes: dividing the actual operation time of the first processing progress node by the proportion of the operation time at the first processing progress node in the overall operation time of this time to obtain the predicted value of the overall operation time of the first job for this time when it runs to the first processing progress node. By performing a simple proportional calculation on the actual operation time and the proportion, it is possible to calculate a predicted value of the overall operation time of a job for this time when the first job runs to the first processing progress node, and since it is based on the actual operation time, it can accurately represent the overall operation time to a certain extent.
[0037] That is, the calculation formula for the overall operation time is: overall operation time = actual operation time of the first processing progress node / proportion of the estimated operation time of the first processing progress node. For example, assume that the actual operation time of the first processing progress node is 60 seconds, and the proportion of the estimated operation time of this node in the overall job operation time is 5%. According to the above formula, the predicted overall operation time of the job for this time will be: overall operation time = 60 / 0.05 = 1200s.
[0038] In this embodiment, by predicting the overall job operation time based on the proportion of the operation time in the key progress node marking and combining the actual operation duration of each key progress node, even in the case of lack of historical similar jobs or small historical data samples, it is still possible to predict the overall operation time of the job for this time based on the actual operation time of the key progress node. Moreover, since the prediction is based on the actual operation time, to a certain extent, the accuracy of the prediction can be improved.
[0039] In addition, before the first job runs, the running time of the previous (most recent) job of the same type as the first job at the first processing progress node can also be used as the estimated running time of the current first job at the corresponding key progress node to predict the overall running time. Specifically, the node marking information further includes the running time of a job of the same type as the first job at the first processing progress node; before the first job runs, the method further includes: calculating and determining the predicted value of the overall running time of the current run when the first job reaches the first processing progress node according to the running time of a job of the same type as the first job at the first processing progress node and the proportion of the running time at the first processing progress node in the overall running time of the current run. In this way, before the job runs, the predicted value of the overall running time of the current first job can be initially predicted through the estimated running time of the most recent job of the same type at the corresponding key progress node.
[0040] Similarly, before the first job runs, the predicted value of the overall running time of the first job = the running time of a job of the same type as the first job at the first processing progress node × the proportion of the running time at the first processing progress node in the overall running time of the current run.
[0041] During the running process of a new job, the data in the job marking library can also be updated by collecting the marking information output at the key progress node marks of the current job. For example, during the running process of the first job, in addition to predicting the overall running time, the output information of each key progress node of the first job can also be configured to be collected, and the data records in the job marking library can be updated. Therefore, referring to Figure 2 , in some embodiments, while obtaining the actual running time of the first processing progress node during the running process of the first job, the method further includes: obtaining the unique identity identifier of the first job, the configuration information identifier of the first job, and the identification information of the first processing progress node; forming the marking information of the first processing progress node of the first job with the unique identity identifier of the first job, the configuration information identifier of the first job, the identification information of the first processing progress node, and the actual running time of the first processing progress node, and storing it in the job marking library to realize the dynamic update of the job marking library.
[0042] Specifically, referring again to Figure 3, as described above, when submitting a job, the user will specify a unique job identity identifier for each job, such as a job ID and the corresponding job configuration information identifier, such as a job configuration ID. The job configuration ID is mainly used to indicate the requirements of the job for host resources, such as the host model, host load, the number of host slots (computing resources), the memory size requested by the job, etc. These identification information helps to distinguish different jobs and the specific configuration information of the jobs. At the same time, the first processing progress node identification information, such as a key progress node ID, can also be provided by the user to mark the key processing nodes in a specific job. These identification information are mainly used to distinguish between different jobs and different key progress nodes of the jobs, so as to facilitate the operation monitoring of the jobs and the accurate marking of key progress nodes.
[0043] In addition, in order to further improve the prediction accuracy, it is also necessary to consider the impact of host load on job running time. Therefore, continue to refer to Figure 3 , in some embodiments, while obtaining the actual running time of the first processing progress node in the running process of the first job, the method further includes: obtaining the host identifier of the host running the first job and the host load information at the first processing progress; storing the host identifier and the host load information at the first processing progress in a job marking library, and associating them with the marking information of the first processing progress node of the first job. In this embodiment, by associating the marking information of different key processing nodes with the host load information of the corresponding processing nodes, it is possible to dynamically predict the overall running time required to run the first job according to the host load condition, which is particularly suitable for predicting the running time of computing cluster jobs.
[0044] Since the impact of host load on job running time is non-linear, in the embodiments of the present invention, the host load can be divided into multiple intervals for processing. For example, the host load can be divided into the following levels: 40%, 50%, 60%, 70%, 80%, 90%, 100%, etc., and each level corresponds to a load range. For example, 40 means that the host load is less than or equal to 40%. During the running process of each job, when the host load changes, the prediction of the running time of the job will be adjusted according to the interval to which the current host load belongs.
[0045] After obtaining the above information, these information will be packaged into a marking record and stored in the job marking library. Among them, the job marking library is a relational database used to record the status and actual running duration of each key progress node in the job running process and other related host information.
[0046] In some embodiments, after obtaining the real-time marking information of the key progress nodes of the above-mentioned first job, when storing it in the job marking library, appropriate indexing and classification operations are performed on the real-time marking information of the key progress nodes to facilitate subsequent data acquisition and real-time update. Specifically, the update mechanism of the job marking library adopts a periodic update and rollback mechanism to ensure the accuracy and reliability of the data.
[0047] In this embodiment, by integrating the relevant identification information, actual running time of each key progress node of each job, and the corresponding host load information into an associated data of key progress marking information and host load information and storing it in the job marking library, the real-time tracking and accurate recording of the job running status can be achieved, and the availability and traceability of the job running data are improved.
[0048] In the embodiments of the present invention, by obtaining the node marking information of jobs of the same type as the first job and combining the progress node information during actual operation, the overall running time of the job can be calculated and corrected in real time. Further, since the key progress node information is associated with the corresponding host load information, not only the dependence on a large amount of historical data is effectively avoided, but also the predicted time can be dynamically adjusted according to the real-time feedback during the job running process, thereby significantly improving the accuracy and adaptability of the prediction.
[0049] See Figure 2 , after forming the marking information of the first processing progress node of the first job and storing it in the job marking library, the method further includes: after monitoring that the first job has finished running, determining whether the first job has run successfully; if it has not run successfully, canceling the marking information of the first processing progress node of the first job in the job marking library; if it has run successfully, storing the proportion of the running time of the first job at the first processing progress node in the actual value of the overall running time of this time.
[0050] Specifically, if a failure or error occurs during the running of the first job, all the marking information of the first job in the job marking library will be revoked to ensure the validity and accuracy of the data. In other words, whether the job runs successfully is the key factor for determining whether to update the corresponding marking information record in the database, ensuring that the information stored in the job marking library only includes accurate data of successful jobs, so that it has reference value. If the job runs successfully, the system will update the proportion of the actual running time of the key progress node in the job marking library. For example, assuming that the actual running time of the first processing progress node of the job is 60 seconds, the system will update the true proportion information of this first processing progress node in the overall running time, so that the data in the job marking library reflects the actual running situation of the first job, which helps to improve the accuracy of predicting the job running time at this progress node in the next job running process based on this proportion information.
[0051] In this embodiment, for the node marking information collected in the job marking library, by adopting a rollback mechanism, that is, during the operation of the job, when an error occurs or the operation is unsuccessful, the job marking library can restore to the previous state or abort the current operation, so as to ensure that during the operation process, especially when a critical operation or node fails, the job marking library only contains valid operation data, avoiding the information generated by failed jobs from affecting the prediction of the running time of subsequent jobs, thereby improving the accuracy and reliability of the job running time prediction.
[0052] As previously mentioned, in some embodiments, the overall running time of the job can be calculated by approaching point by point by setting multiple critical progress node marks, so as to improve the accuracy of prediction. For the case where the user sets at least two critical progress node marks when submitting a job, specifically, after calculating and determining the predicted value of the current overall running time when the first job runs to the first processing progress node, it further includes: obtaining the proportion of the running time of the job of the same type as the first job at the second processing progress node in the current overall running time from the job marking library; and obtaining the actual running time of the second processing progress node during the running of the first job; calculating and determining the predicted value of the current overall running time when the first job runs to the second processing progress node according to the actual running time of the second processing progress node and the proportion of the running time at the second processing progress node in the current overall running time.
[0053] In other words, by collecting running time data at multiple critical progress nodes, since the closer the critical progress node is to the end of the task running, the closer its predicted value is to the true running value, in this way, through multiple critical progress node marks, the predicted value can be dynamically corrected.
[0054] Specifically, the second processing progress node of the first job refers to the second critical processing progress node during the running of the job, and its specific position and actual running time in the entire job running process are different from those of the first processing progress node. According to the actual running time of the second processing progress node and the proportion of the running time of this node, extrapolate and calculate the predicted value of the current overall running time of the job corresponding to the second processing progress node, and moreover, the estimation of the predicted time can be corrected at each critical progress node, and the granularity of the running time prediction is finer.
[0055] Similarly, in the process of predicting the overall running time based on the actual running time and its time proportion of the second processing progress node, the marking information of the second processing progress node can also be collected and stored in the job marking library. Specifically, while obtaining the actual running time of the second processing progress node in the running process of the first job, it also includes: obtaining the unique identity identifier of the first job, the configuration information identifier of the first job, and the identifier information of the second processing progress node; forming the marking information of the second processing progress node of the first job with the unique identity identifier of the first job, the configuration information identifier of the first job, the identifier information of the second processing progress node, and the actual running time of the second processing progress node, and storing it in the job marking library.
[0056] Specifically, the unique identity identifier of the first job helps to uniquely identify the job in the job marking library and distinguish it from other jobs. The configuration information identifier of the first job provides detailed information about the job running environment, such as the required number of host slots and the host load requirements mentioned above. This is crucial for accurately predicting the job running time and scheduling jobs among multiple hosts.
[0057] In this embodiment, by collecting the information of each key progress node during the running process of the first job, the job marking library can be updated to the latest data record in real time, thereby providing more accurate basic data for subsequent running time prediction.
[0058] It can be understood that although this embodiment describes the technical concept of the technical solution of this embodiment with the first processing progress node and the second processing progress node, in fact, the specific number of processing progress nodes set in this application is not limited.
[0059] Similarly, in the process of collecting the second processing progress node, it can also include judging whether to revoke the data that has been saved in the job marking library according to whether the first job runs successfully through the aforementioned rollback mechanism, so as to ensure the validity of the data in the job marking library. For the sake of brief description, the specific detailed description of the collection process of the second processing progress node will not be given again. The relevant description of the first processing progress node part can be referred to. Generally speaking, the collection process of each progress node in each job can refer to the relevant description of the first processing progress node part.
[0060] In this embodiment, in a complex job running environment, introducing multiple key progress node markings and a point-by-point approximation calculation method can make the prediction result closer to the actual situation and the prediction granularity finer through point-by-point calculation without relying on a large amount of historical data.
[0061] The embodiments of the present invention are applicable to the running time prediction of single-machine jobs and the running time prediction of cluster batch jobs. In a cluster environment, for the running time prediction of cluster batch jobs, after calculating the predicted value of the overall running time of each job when it runs to the first processing progress node, it further includes: comparing the predicted values of the overall running time among various jobs, and taking the maximum running time among various jobs as the predicted value of the overall running time of the cluster batch job at the first processing progress node.
[0062] It should be noted that when there are multiple job progress node marks, for each job progress node mark, a corresponding predicted value of the overall running time at the corresponding progress node can be obtained according to the above method, realizing a more refined dynamic prediction with a finer granularity.
[0063] As mentioned above, the technical solution of the present invention is applicable to the running time prediction of single-machine jobs and the running time prediction of cluster batch jobs. In a cluster environment, multiple jobs run simultaneously, and their respective running times are often affected by cluster resources and the load of each host. Therefore, in the embodiments of the present invention, for the prediction of the running time of cluster jobs, it not only depends on the progress nodes and historical data of the jobs themselves, but also considers the influence of dynamic factors such as cluster status and host load, so as to improve the accuracy of prediction.
[0064] In the prediction process of cluster jobs, first, according to the progress node marks of each job, the overall running time of each job is calculated. For the running time prediction of cluster batch jobs, the further steps include: comparing the running times of each job to determine the maximum running time among them as the predicted running time of the batch job. This is because in a cluster environment, batch jobs often need to consider the slowest job as the standard for overall completion.
[0065] Specifically, during the job scheduling process, the monitoring system will obtain the load information of each host in real time and dynamically adjust the time prediction of the jobs running on that host according to the load level. If the host load is high (such as the load is between 80% and 100%), the jobs running on that host may be affected by resource bottlenecks, resulting in an extended running time. Therefore, the prediction system needs to dynamically adjust the estimated remaining running time of each job based on the current load range and historical running data.
[0066] For example, assume that in a cluster environment, a certain batch job contains several subtasks, and one of the subtasks is significantly affected by the host load during its execution. In this case, when the load of the host increases from 40% to 80%, the running time of the job may increase. At this time, the system adjusts according to the current load interval, extends the predicted time of the subtask, and thus affects the predicted running time of the entire batch job.
[0067] By comprehensively considering the host load interval and the proportion of the running time of the key progress nodes of the job, the present invention can achieve dynamic and accurate running time prediction, and effectively avoid prediction errors caused by changes in host load. This improvement makes the cluster job scheduling more efficient and accurate, helps to identify possible delays in advance, and optimizes the resource scheduling strategy.
[0068] In some embodiments, based on the prediction results of the host load interval, the job scheduling strategy in the cluster can be further optimized. For example, when the load of some nodes in the cluster is too high, the scheduling system can migrate tasks to nodes with lower load, reduce the impact of load imbalance on the running time of the job, and improve the utilization rate of cluster resources and the job completion efficiency.
[0069] Figure 4 Schematically shows the schematic diagram of the computing cluster and job organizational structure in an embodiment of the present application; see Figure 4 Specifically, for the prediction of the running time of the batch job in the computing cluster, in the actually deployed cluster, the jobs usually run in parallel. For a job submitted by a user, after it is submitted to the cluster management host, the cluster management host may split the job according to the situation (generally still called the job or task) and allocate it to multiple running hosts in the cluster. The running time of the job on each host can be calculated and the dynamic scheduling of cluster resources can be considered. Since multiple jobs in the cluster may have different running times, and the dependencies and resource consumptions between jobs are different. Therefore, when predicting the overall running time of the cluster job this time, the load information of the hosts in the entire cluster also needs to be considered.
[0070] In some embodiments, first, according to the historical data stored in the job tag library, obtain the proportion of the running time of the key progress nodes of each job, and gradually approximate and predict the overall running time of each job based on the historical running data. Then, based on the prediction results of the overall running time of each job, perform cluster job scheduling, compare the running times of each job, and use the job with the largest overall running time as the slowest task in the cluster operation. At this time, the total running time of the cluster is the maximum running time of this job. This is because, in the cluster job scheduling, the jobs run synchronously. If a job takes more time to complete, the overall job progress of the cluster will be restricted by it.
[0071] For example, in some batch jobs, the first job is divided into Job A and Job B, which are run in parallel in the same cluster. Job A needs to process less data and its predicted running time is 20 minutes; Job B needs to process a larger amount of data and its predicted running time is 45 minutes. According to the solution provided by the embodiments of the present invention, when calculating the predicted value of the overall running time of the cluster jobs this time, the result will be based on 45 minutes, representing the maximum running time of the cluster, thereby helping the scheduling system to optimize the job resource allocation and scheduling strategy.
[0072] In addition, in the embodiments of the present invention, it is also applicable to real-time correction in a dynamic environment. In other words, in a cluster environment, as the progress of job execution and the host load status change, the prediction result can be adjusted in real time. For example, when the host in the cluster has an excessive load at a certain processing progress node, the actual running time of the job may increase. At this time, the prediction value can be dynamically corrected according to the more refined association relationship between the key processing progress node marking information and the host load set according to the granularity. And through the way of approaching point by point, the running time prediction of the job can be dynamically adjusted in real time to ensure the efficient utilization of cluster resources.
[0073] In the embodiments of the present invention, through the dynamic adjustment prediction time method of approaching point by point, it is possible to provide a more refined running time prediction for cluster job scheduling. Especially in a dynamically changing and complex cluster environment, the dynamic real-time running time prediction helps to efficiently schedule and optimize cluster resources and improve the overall efficiency of cluster jobs.
[0074] In some embodiments, the method provided by the embodiments of the present invention can be applicable to the prediction of the running time of various job tasks. For example, the prediction of the running time of chip verification jobs, various jobs running on the cluster, including but not limited to circuit simulation, physical implementation, AI training and inference, high-performance numerical calculation, etc. According to different application scenarios, users can insert key progress node markings in combination with the characteristics of the jobs.
[0075] For example, for the training task of an AI model, since the time distribution in different training stages varies significantly, key progress node markings can be inserted in combination with the characteristics of the job itself. In AI training, key progress nodes can be inserted according to the training stages. For example, in the initial stage, such as 5% training completion, it is used to determine the model initialization time; in the middle stage, such as 50% training completion, it reflects the efficiency of data processing and parameter optimization during training; in the final stage, such as 90% training completion, it represents the model convergence speed and the time requirement for the final adjustment stage.
[0076] Suppose in a certain AI training task, the user discovers through analysis that data loading and model convergence are the main time bottlenecks. To address this, when submitting a job, key progress node markers can be inserted at the 5%, 50%, and 90% stages of the job's execution, and the estimated time for each key progress node can be input into a prediction model for dynamic adjustment. In the initial stage of training, the estimated value may be relatively high due to slow data loading. Therefore, based on no historical data, a default estimated running time can be set. However, as the training progresses and some marker information of key progress nodes has been collected, the actual estimated value in the marker information of the key progress nodes of the previous run on a certain host can be gradually approximated to the actual value, thereby achieving dynamic prediction of the overall running time.
[0077] Meanwhile, for batch processing jobs such as circuit simulation, the user can also insert key progress nodes according to the proportional relationship based on the characteristics of the job. For example, the entire simulation task can be divided into three nodes at 20%, 50%, and 80% according to the distribution of computing tasks. By monitoring the actual completion time of these key progress nodes, the running time prediction can be dynamically adjusted. This is especially applicable to situations where the running time of the task is greatly affected by dynamic factors such as host load.
[0078] According to the performance in some test cases, the more accurate the proportion of the estimated time of the key progress node, the greater the possibility that the predicted value approaches the true value. Therefore, it is recommended to insert at least two key progress nodes at 5% and 90% to ensure the reliability of the prediction results. It can be understood that the more key progress nodes are inserted, the higher the overall prediction accuracy.
[0079] See Figure 5 , which shows a schematic block diagram of the overall prediction process of the job running time from the user side to the cluster management server side in a cluster environment; Figure 6 shows a schematic block diagram of the structure of an embodiment of the device 200 for predicting the running time of computer jobs in the present application. See Figure 5 and Figure 6 , the present application also provides an embodiment of a device 200 for predicting the running time of computer jobs. The device 200 includes: A first acquisition program unit 210, configured to acquire the node marker information of jobs of the same type as the first job in the job marker library. The node marker information includes the proportion of the running time of jobs of the same type as the first job at the first processing progress node in the overall running time of the current time; A second acquisition program unit 220, configured to acquire the actual running time of the first processing progress node during the running of the first job; A prediction program unit 230 is configured to calculate and determine a predicted value of the overall running time of the first job when it runs to the first processing progress node according to the actual running time of the first processing progress node and the proportion of the running time at the first processing progress node in the overall running time of this time.
[0080] In some embodiments, the second acquisition program unit 220 is further configured to acquire the unique identity identifier of the first job, the configuration information identifier of the first job, and the identifier information of the first processing progress node. The apparatus further includes: a first storage program unit, configured to form marking information of the first processing progress node of the first job by using the unique identity identifier of the first job, the configuration information identifier of the first job, the identifier information of the first processing progress node, and the actual running time of the first processing progress node, and store the marking information in a job marking library.
[0081] In some embodiments, the apparatus further includes: a monitoring program unit, configured to, after monitoring that the first job has finished running, determine whether the first job has run successfully; if it has not run successfully, cancel the marking information of the first processing progress node of the first job in the job marking library; if it has run successfully, store the proportion of the running time of the first job at the first processing progress node in the actual value of the overall running time of this time.
[0082] In some embodiments, the second acquisition program unit 220 is further configured to acquire the host identifier of the host on which the first job runs and the host load information at the first processing progress; the first storage program unit is further configured to store the host identifier and the host load information at the first processing progress in the job marking library and associate them with the marking information of the first processing progress node of the first job.
[0083] In some embodiments, the node marking information further includes the proportion of the running time of a job of the same type as the first job at the second processing progress node in the overall running time of this time; the second acquisition program unit 220 is further configured to acquire the actual running time of the second processing progress node during the running process of the first job; the prediction program unit 230 is further configured to calculate and determine a predicted value of the overall running time of the first job when it runs to the second processing progress node according to the actual running time of the second processing progress node and the proportion of the running time at the second processing progress node in the overall running time of this time.
[0084] In some embodiments, the second acquisition program unit 220 is further configured to acquire the unique identity identifier of the first job, the configuration information identifier of the first job, and the identifier information of the second processing progress node; and acquire the host identifier of the host on which the first job runs and the host load information at the second processing progress. The device further includes: a second stored program unit, configured to form marking information of a second processing progress node of the first job by using a unique identity identifier of the first job, a configuration information identifier of the first job, second processing progress node identifier information, and an actual running time of the second processing progress node, and store the marking information in a job marking library; and store the host identifier and host load information at the second processing progress time in the job marking library, and associate the information with the marking information of the second processing progress node of the first job.
[0085] In some embodiments, the prediction program unit 230 is specifically configured to divide the actual running time of the first processing progress node by a proportion of the running time at the first processing progress node in the overall running time of the current time, so as to obtain a predicted value of the overall running time of the first job when running to the first processing progress node.
[0086] In some embodiments, the node marking information further includes a running time of a job of the same type as the first job at the first processing progress node; the prediction program unit 230 is further configured to, before running the first job, calculate and determine a predicted value of the overall running time of the first job when running to the first processing progress node according to the running time of the job of the same type as the first job at the first processing progress node and a proportion of the running time at the first processing progress node in the overall running time of the current time.
[0087] In some embodiments, the device is applicable to predicting the running time of a single-machine job and predicting the running time of a computing cluster batch job; wherein, for predicting the running time of a computing cluster batch job, the device further includes: a comparison program unit, configured to compare magnitudes of the overall running times of each job when running to the first processing progress node after calculating and determining the predicted value of the overall running time of each job when running to the first processing progress node; and a determination program unit, configured to determine a maximum value of the overall running times of each job as a predicted value of the overall running time predicted for the computing cluster batch job at the first processing progress node.
[0088] The device 200 for predicting the running time of a computer job in this embodiment has a similar implementation principle and technical effect to the foregoing method embodiment, and details are not described herein again, and reference can be made to each other.
[0089] Figure 7 FIG. is a schematic structural diagram of an embodiment of an electronic device for job processing according to the present application, which can implement any one of the methods in the embodiments of the present application. Refer to Figure 7, As an alternative embodiment, the above electronic device may include: a housing 41, a processor 42, a memory 43, a circuit board 44, and a power supply circuit 45. Among them, the circuit board 44 is disposed inside the space enclosed by the housing 41, and the processor 42 and the memory 43 are provided on the circuit board 44; the power supply circuit 45 is used to supply power to each circuit or device of the above electronic device; the memory 43 is used to store executable program codes; the processor 42 runs a program corresponding to the executable program code by reading the executable program code stored in the memory 43, and is used to run the method for predicting the running time of a computer job in any one of the foregoing embodiments.
[0090] For the specific running process of the above steps by the processor 42 and the steps further run by the processor 42 by running the executable program code, reference may be made to the description of Embodiment 1 of the method for predicting the running time of a computer job in this application, which will not be elaborated herein.
[0091] The electronic device exists in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones (such as iPhone), multimedia phones, functional phones, and low-end phones, etc. (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc., such as iPad. (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video playback modules (such as iPod), handheld game consoles, e-books, and intelligent toys and portable vehicle navigation devices. (4) Servers: Devices that provide computing services. The composition of a server includes a processor, a hard disk, a memory, a system bus, etc. A server is similar to a general computer architecture, but due to the need to provide highly reliable services, it has higher requirements in terms of processing power, stability, reliability, security, scalability, and manageability. (5) Other electronic devices with data interaction functions.
[0092] This application also provides an embodiment of a computer-readable storage medium. The computer-readable storage medium stores one or more programs, and the one or more programs can be run by one or more processors to implement the method for predicting the running time of a computer job in any one of the foregoing Embodiment 1.
[0093] In summary, according to the descriptions of the above embodiments, the method and device for predicting the running time of a computer job disclosed in this embodiment extrapolate the running time of the entire job based on the actual running times of some key nodes, and the estimated prediction time can be gradually corrected at each key node to approach the true overall running time of the job. To a certain extent, the accuracy of job running time prediction can be improved.
[0094] In addition, by allowing users to add key progress node markers to the job when submitting the job, the collected job running historical data can be collected at multiple key progress node levels, and the prediction data granularity is finer. Further, since the default running time value can be set by the user in the job key node marker, in the absence of similar historical data, a relatively reasonable overall running time prediction value for the current job can also be given based on this and the actual running time data.
[0095] It should be noted that in this article, except for qualifiers such as first priority and second priority indicating the first and second in queue priorities, the rest of the relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0096] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments.
[0097] For the convenience of description, the above device is described by dividing it into various units / modules according to functions. Of course, when implementing this application, the functions of the various units / modules can be implemented in the same or multiple software and / or hardware.
[0098] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program runs, it can include the processes of the embodiments of the above various methods. Among them, the storage medium can also be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0099] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for predicting the running time of a computer job, characterized in that: The method comprises: Obtaining node marking information of a job of the same type as the first job in a job marking library, wherein the node marking information includes a proportion of a running time of the job of the same type as the first job at a first processing progress node in the overall running time of the current time; Obtaining the actual running time at the first processing progress node during the running of the first job; According to the proportion of the actual running time and the running time at the first processing progress node in the current overall running time, a predicted value of the current overall running time when the first job runs to the first processing progress node is calculated and determined.
2. The method according to claim 1, characterized in that: While obtaining the actual running time of the first processing progress node during the running of the first job, the method further includes: obtaining the unique identification of the first job, the configuration information identification of the first job and the identification information of the first processing progress node; The unique identity of the first job, the configuration information of the first job, the identification information of the first processing progress node and the actual running time of the first processing progress node are used to form the marking information of the first processing progress node of the first job and stored in the job marking library.
3. The method according to claim 2, characterized in that After forming the mark information of the first processing progress node of the first job and storing it in the job mark library, the method further includes: after monitoring that the first job is completed, determining whether the first job is successfully executed; If the operation is unsuccessful, canceling the marking information of the first processing progress node of the first job in the job marking library; If the operation is successful, the proportion of the operation time of the first job at the first processing progress node in the actual value of the overall operation time is stored.
4. The method according to claim 2, characterized in that: While obtaining the actual running time of the first processing progress node in the running process of the first job, the method further includes: obtaining the host identifier of the host running the first job and the host load information at the first processing progress; The host identifier and the host load information at the first processing progress are stored in a job tag library, and associated with the tag information of the first processing progress node of the first job.
5. The method according to claim 1, characterized in that The node marking information also includes the proportion of the running time of the job of the same type as the first job at the second processing progress node in the overall running time of the current time; After calculating and determining the predicted value of the overall running time of the first job when it runs to the first processing progress node, the method further includes: Obtaining the actual running time of the second processing progress node during the running of the first job; According to the actual running time of the second processing progress node and the proportion of the running time at the second processing progress node in the current overall running time, a predicted value of the overall running time when the first job runs to the second processing progress node is calculated and determined.
6. The method according to claim 5, characterized in that While obtaining the actual running time of the second processing progress node during the running of the first job, the method further includes: obtaining the unique identification of the first job, the configuration information identification of the first job and the identification information of the second processing progress node; and obtaining the host identification of the host running the first job and the host load information during the second processing progress; The unique identity identifier of the first job, the configuration information identifier of the first job, the second processing progress node identifier information and the actual running time of the second processing progress node are used to form the tag information of the second processing progress node of the first job and stored in the job tag library; and the host identifier and the host load information during the second processing progress are stored in the job tag library and associated with the tag information of the second processing progress node of the first job.
7. The method according to claim 1, characterized in that The calculation to determine the predicted value of the overall running time of the first job when it runs to the first processing progress node includes: The actual running time of the first processing progress node is divided by the proportion of the running time at the first processing progress node in the current overall running time to obtain a predicted value of the overall running time when the first job runs to the first processing progress node.
8. The method according to claim 1, characterized in that The node marking information also includes the running time of a job of the same type as the first job at the first processing progress node; Before running the first job, the method also includes: calculating and determining a predicted value of the overall running time when the first job runs to the first processing progress node based on the running time of a job of the same type as the first job at the first processing progress node and the proportion of the running time at the first processing progress node in the overall running time.
9. The method according to claim 1, characterized in that: The method is applicable to the running time prediction of single-machine jobs and the running time prediction of batch processing jobs in computing clusters; wherein, For computing cluster batch processing job running time prediction, after calculating and determining the current overall running time prediction value when each job runs to the first processing progress node, it also includes: comparing the current overall running time prediction values between the various jobs; The maximum value of the current overall running time prediction values in each job is determined as the current overall running time prediction value of the computing cluster batch processing job at the first processing progress node.
10. A device for predicting the running time of a computer job, characterized in that: The device comprises: A first acquisition program unit is used to acquire node tag information of a job of the same type as the first job in a job tag library, wherein the node tag information includes a proportion of the running time of the job of the same type as the first job at the first processing progress node in the overall running time of the current time; A second acquisition program unit is used to obtain the actual running time of the first processing progress node during the running of the first job; The prediction program unit is used to calculate and determine the predicted value of the overall running time when the first job runs to the first processing progress node based on the actual running time of the first processing progress node and the proportion of the running time at the first processing progress node in the overall running time.
11. The device according to claim 10, characterized in that The second program acquisition unit is further used to acquire the unique identification of the first job, the configuration information identification of the first job and the first processing progress node identification information; The device also includes: a first storage program unit, which is used to form the marking information of the first processing progress node of the first job by using the unique identity identifier of the first job, the configuration information identifier of the first job, the first processing progress node identification information and the actual running time of the first processing progress node, and store it in the job marking library.
12. The device according to claim 11, characterized in that The device further comprises: a monitoring program unit, which is used to determine whether the first operation is successfully executed after monitoring that the first operation is completed; If the operation is unsuccessful, canceling the marking information of the first processing progress node of the first job in the job marking library; If the operation is successful, the proportion of the operation time of the first job at the first processing progress node in the actual value of the overall operation time is stored.
13. The device according to claim 11, characterized in that The second acquisition program unit is further used to acquire the host identifier of the host running the first job and the host load information during the first processing progress; The first storage program unit is further used to store the host identifier and the host load information at the first processing progress into a job tag library, and associate them with the tag information of the first processing progress node of the first job.
14. The device according to claim 10, characterized in that The node marking information also includes the proportion of the running time of the job of the same type as the first job at the second processing progress node in the overall running time of the current time; The second acquisition program unit is further used to obtain the actual running time of the second processing progress node during the running of the first job; The prediction program unit is further used to calculate and determine the overall running time of the first job according to the actual running time of the second processing progress node and the proportion of the running time at the second processing progress node in the overall running time.
15. The device according to claim 14, characterized in that The second program acquisition unit is further used to acquire the unique identification of the first job, the configuration information identification of the first job and the second processing progress node identification information; and obtaining a host identifier of a host running the first job and host load information during the second processing progress; The device also includes: a second storage program unit, which is used to store the unique identity identifier of the first job, the configuration information identifier of the first job, the second processing progress node identification information and the actual running time of the second processing progress node to form the marking information of the second processing progress node of the first job into the job marking library; and store the host identifier and the host load information at the second processing progress into the job marking library, and associate it with the marking information of the second processing progress node of the first job.
16. The device according to claim 10, characterized in that The prediction program unit is specifically configured to divide the actual running time of the first processing progress node by the proportion to obtain the overall running time of the first job.
17. The device according to claim 10, characterized in that The node marking information also includes the running time of a job of the same type as the first job at the first processing progress node; The prediction program unit is also used to calculate and determine, before running the first job, a predicted value of the overall running time of the first job running to the first processing progress node based on the running time of jobs of the same type as the first job at the first processing progress node and the proportion of the running time at the first processing progress node in the overall running time of the current time.
18. The device according to claim 10, characterized in that The device is suitable for predicting the running time of a single machine job and the running time of a batch processing job of a computing cluster; wherein, For computing cluster batch processing job running time prediction, the device also includes: a comparison program unit, which is used to compare the size of the current overall running time between each job after calculating and determining the current overall running time prediction value when each job runs to the first processing progress node; A determination program unit is used to determine the maximum value of the overall running time of each job as the predicted value of the overall running time of the computing cluster batch processing job running to the first processing progress node.
19. An electronic device, characterized in that: include: One or more processors; Memory; The memory stores one or more executable programs, and the one or more processors read the executable program codes stored in the memory to run the programs corresponding to the executable program codes, so as to run any method described in claims 1 to 9.
20. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method of any one of claims 1 to 9.