Job processing method and device, equipment, storage medium and program product

By obtaining and analyzing the operation data of the job, determining the standard operation curve and detecting abnormalities, the problem that the Slurm system cannot detect operation abnormalities in time is solved, and the reliability and efficiency of operation are improved.

CN119938375APending Publication Date: 2025-05-06DAWNING INFORMATION IND (BEIJING) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411813514.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The Slurm job scheduling system cannot detect an abnormality in time, resulting in a decrease in the operation efficiency of the job.

Method used

By obtaining the reference running data and real running data of the target job, the standard running curve is determined, and the job running status is saved when an exception is detected, the exception is handled and the job run is restored.

Benefits of technology

It improves the reliability and efficiency of job operation, promptly detects and handles job abnormalities, and reduces the number of job reruns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938375A_ABST
    Figure CN119938375A_ABST
Patent Text Reader

Abstract

The invention relates to a job processing method and device, equipment, a storage medium and a program product. The method comprises the steps that reference operation data of a target job in a reference time period and real operation data of the target job in a current time period are obtained, the reference time period is located before the current time period, a standard operation curve of the target job in a normal operation state in the current time period is determined according to the reference operation data, and real operation data of the target job in the current time period are obtained; under the condition of determining that the target operation runs abnormally at the target moment in the current time period according to the standard running curve and the real running data, storing the operation running state of the target operation at the target moment, carrying out exception processing on the target operation, and after the exception processing is finished, according to the operation running state, carrying out exception processing on the target operation. And performing operation recovery on the target job. By adopting the method, the operation efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a job processing method, device, equipment, storage medium and program product. Background Art

[0002] To ensure the reliability of job execution, the Slurm (Simple Linux Utility for Resource Management) job scheduling system classifies resources such as the central processing unit (CPU) and system memory, and runs jobs with different requirements on different computing nodes.

[0003] However, since the running process of a job involves multiple dimensions of running data, the Slurm job scheduling system cannot detect job abnormalities in a timely manner. In addition, after processing the abnormal job, the job needs to be rerun, which reduces the efficiency of job operation. Summary of the invention

[0004] Based on this, it is necessary to provide a job processing method, device, equipment, storage medium and program product that can improve job operation efficiency in response to the above technical problems.

[0005] In a first aspect, the present application provides a job processing method, comprising:

[0006] Obtain reference operation data of the target job in a reference period and actual operation data of the target job in a current period; wherein the reference period is before the current period;

[0007] Based on the reference operation data, determine the standard operation curve of the target operation under normal operation in the current period;

[0008] When it is determined based on the standard operation curve and the actual operation data that the target operation is operating abnormally at the target time in the current period, the operation status of the target operation at the target time is saved;

[0009] Perform exception handling on the target job, and after the exception handling is completed, resume the target job according to the job running status.

[0010] In the embodiments of the present application, on the one hand, the standard operation curve is used to promptly determine the operation anomalies of the target job, thereby improving the reliability of the job operation; on the other hand, when an abnormality occurs in the target job, the job operation status is immediately saved, and after the abnormality is handled, the target job is restored according to the job operation status, which can effectively improve the efficiency of job recovery and thereby improve the efficiency of job operation.

[0011] In one embodiment, obtaining reference operation data of a target job within a reference period includes:

[0012] Determine the target application type of the target application according to the command name of the target application for submitting the target job; determine the target monitoring indicator corresponding to the target application for submitting the target job according to the target application type and the correspondence between the pre-set candidate application types and the candidate monitoring indicators; obtain reference operating data of the target job under the target monitoring indicator within a reference time period.

[0013] In the embodiment of the present application, by acquiring the operating data of jobs under different application types under different monitoring indicators, the scope of data acquisition can be narrowed, while the rationality of the standard operating curve prediction is improved.

[0014] In one embodiment, determining a standard operation curve of a target operation in a normal operation state in a current period according to reference operation data includes:

[0015] According to the data processing method associated with the target application type, feature extraction is performed on the reference operation data to obtain reference operation features; based on the reference operation features, a standard operation curve of the target operation under normal operating conditions in the current period is determined.

[0016] In the embodiment of the present application, by using different data processing methods for reference operation data under different application types, the rationality of data processing can be guaranteed, thereby improving the accuracy of standard operation curve prediction.

[0017] In one embodiment, the real operation data includes real operation data at each time in the current period; determining that the target operation is abnormal at the target time in the current period according to the standard operation curve and the real operation data includes:

[0018] For each moment in the current time period, determine the operating deviation value between the actual operating data corresponding to the moment and the standard operating data corresponding to the moment in the standard operating curve; when the operating deviation value is greater than the deviation threshold, stop executing the operation of determining the operating deviation value between the actual operating data corresponding to other moments after the moment and the standard operating data corresponding to the standard operating curve; take the moment as the target moment, and determine the operating abnormality of the target job at the target moment in the current time period.

[0019] In the embodiment of the present application, by determining the target time when the abnormality occurs based on the operation deviation value between the actual operation data and the standard operation data corresponding to each moment in the current time period, the accuracy of the determination of the operation abnormality can be ensured.

[0020] In one embodiment, according to the running status of the job, the target job is resumed, including:

[0021] A target node is selected from each other node according to the node operation data of each other node; wherein each other node is a node in the cluster except the current node running the target job; and the operation of the target job is resumed on the target node according to the job operation status.

[0022] In the embodiment of the present application, by selecting a target node from each other node according to the node operation data of each other node, and resuming the operation of the target job on the target node, the reliability of the operation of the target job can be guaranteed.

[0023] In one of the embodiments, resuming the operation of the target job on the target node according to the operation status of the job includes:

[0024] Determine the job resource allocation strategy based on the job exception cause of the target job; and resume the operation of the target job on the target node based on the job running status and the job resource allocation strategy.

[0025] In an embodiment of the present application, by determining a job resource allocation strategy according to the job exception cause of the target job, and restoring the operation of the target job on the target node based on the job resource allocation strategy, the reliability of job recovery can be ensured.

[0026] In a second aspect, the present application further provides a job processing device, comprising:

[0027] A data acquisition module, used to acquire reference operation data of a target job in a reference period and real operation data of the target job in a current period; wherein the reference period is before the current period;

[0028] A curve determination module is used to determine a standard operation curve of a target operation under normal operation in a current period according to reference operation data;

[0029] The state saving module is used to save the operation state of the target job at the target time when it is determined that the target job is operating abnormally at the target time within the current period according to the standard operation curve and the actual operation data;

[0030] The job recovery module is used to handle exceptions of the target job and, after the exception handling is completed, to resume the operation of the target job according to the job running status.

[0031] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0032] Obtain reference operation data of the target job in a reference period and actual operation data of the target job in a current period; wherein the reference period is before the current period;

[0033] Based on the reference operation data, determine the standard operation curve of the target operation under normal operation in the current period;

[0034] When it is determined based on the standard operation curve and the actual operation data that the target operation is operating abnormally at the target time in the current period, the operation status of the target operation at the target time is saved;

[0035] Perform exception handling on the target job, and after the exception handling is completed, resume the target job according to the job running status.

[0036] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0037] Obtain reference operation data of the target job in a reference period and actual operation data of the target job in a current period; wherein the reference period is before the current period;

[0038] Based on the reference operation data, determine the standard operation curve of the target operation under normal operation in the current period;

[0039] When it is determined based on the standard operation curve and the actual operation data that the target operation is operating abnormally at the target time in the current period, the operation status of the target operation at the target time is saved;

[0040] Perform exception handling on the target job, and after the exception handling is completed, resume the target job according to the job running status.

[0041] In a fifth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:

[0042] Obtain reference operation data of the target job in a reference period and actual operation data of the target job in a current period; wherein the reference period is before the current period;

[0043] Based on the reference operation data, determine the standard operation curve of the target operation under normal operation in the current period;

[0044] When it is determined based on the standard operation curve and the actual operation data that the target operation is operating abnormally at the target time in the current period, the operation status of the target operation at the target time is saved;

[0045] Perform exception handling on the target job, and after the exception handling is completed, resume the target job according to the job running status.

[0046] The above-mentioned job processing method, device, equipment, storage medium and program product determine the standard operation curve of the target job in the normal operation state in the current period according to the reference operation data of the target job, and save the job operation state of the target job at the target moment when it is determined that the target job is abnormal at the target moment in the current period according to the standard operation curve and the real operation data; then, the target job is processed abnormally, and after the abnormal processing is completed, the target job is restored to operation according to the job operation state. Compared with the related art, in which the abnormality of the job cannot be detected in time, and the job needs to be re-run after the abnormal job is processed, the above-mentioned method, on the one hand, adopts the standard operation curve, can timely determine the abnormality of the target job, and improves the reliability of the job operation; on the other hand, in the case of abnormality of the target job, the job operation state is immediately saved, and after the abnormal processing is completed, the target job is restored to operation according to the job operation state, which can effectively improve the efficiency of job recovery, thereby improving the efficiency of job operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0048] Figure 1 A schematic diagram of a process flow of a job processing method in one embodiment;

[0049] Figure 2 A schematic diagram of a process for obtaining reference operation data in one embodiment;

[0050] Figure 3 A schematic diagram of a standard operating curve in one embodiment;

[0051] Figure 4 A schematic diagram of a process for determining an operational abnormality in one embodiment;

[0052] Figure 5 is a flowchart of a method for processing a job in another embodiment;

[0053] Figure 6 is a structural block diagram of a job processing device in one embodiment;

[0054] Figure 7 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0056] To ensure the reliability of job execution, the Slurm job scheduling system classifies resources such as the central processing unit (CPU) and system memory, and runs jobs with different requirements on different computing nodes.

[0057] However, since the running process of a job involves multiple dimensions of running data, the Slurm job scheduling system cannot detect job abnormalities in a timely manner. In addition, after processing the abnormal job, the job needs to be rerun, which reduces the efficiency of job operation.

[0058] Based on this, in an exemplary embodiment, Figure 1 As shown, a job processing method is provided, and the method is applied to a job processing device deployed in a Slurm scheduling system as an example for explanation, and specifically includes the following steps:

[0059] S101, obtaining reference operation data of a target job in a reference time period and actual operation data of the target job in a current time period.

[0060] Among them, the so-called job is the processing task submitted to the scheduling system through the application, the target job is the job with monitoring requirements; the reference period is the period in the historical period when the target job is in normal operation, and the reference period is before the current period; the real operation data is the real data generated by the target job during operation, as well as the node operation data of the current node where the target job is located. Furthermore, the reference operation data is the real operation data of the target job in the reference period.

[0061] In an embodiment of the present application, users can submit jobs to be processed to the Slurm scheduling system through various types of applications; then, the Slurm scheduling system schedules each job to a corresponding node for processing based on the computing resources required for each job.

[0062] In an optional implementation, a performance monitoring tool in the Slurm scheduling system can be used to collect and store real operating data of the target job in real time, wherein the real operating data includes but is not limited to CPU utilization, memory consumption, GPU utilization, and network traffic.

[0063] In order to detect abnormal situations of the target job in a timely manner, it is necessary to obtain the reference operation data of the target job in a reference period and the actual operation data of the target job in the current period according to the job identification information of the target job.

[0064] It is understandable that in order to ensure the reliability of the reference period, the previous period closest to the current period may be used as the reference period, that is, the reference operation data of the target job in the previous period is obtained.

[0065] S102, determining a standard operation curve of the target operation in a normal operating state within the current period according to the reference operation data.

[0066] The so-called standard operating curve is a curve that can characterize the operating data of the target operation under normal operating conditions.

[0067] In an optional implementation, the sample operation data and the sample operation curve corresponding to the sample operation data may be used to construct a first curve prediction model. The sample operation data may be only operation data when the operation is in a normal operation state, or may include both operation data in a normal operation state and operation data in an abnormal operation state, and this application does not limit this.

[0068] Exemplarily, the sample data features corresponding to the sample operation data may be determined, and the sample data features may be input into the initial prediction model, and the initial prediction model may output a prediction operation curve according to the sample data features; then, the initial prediction model may be optimized according to the deviation value between the prediction operation curve corresponding to the sample operation data and the sample operation curve, to obtain a first curve prediction model. The sample data features may include, but are not limited to, sample data features under the CPU frequency dimension, the number of instructions per cycle (IPC) dimension, the cache hit rate dimension, the context switch dimension, and the CPU migration dimension.

[0069] Based on this, the reference operating data can be input into the first curve prediction model being trained, and the first curve prediction model outputs a standard operating curve of the target operation under normal operating conditions in the current period based on the reference operating data and model parameters.

[0070] In another optional implementation, the operation change trend of the target operation within a reference time period can be determined based on the reference operation data; then, based on the operation change trend of the target operation within the reference time period, a standard operation curve of the target operation under normal operation in the current time period is determined.

[0071] S103, when it is determined based on the standard operation curve and the actual operation data that the target operation is operating abnormally at the target time in the current period, the operation operation status of the target operation at the target time is saved.

[0072] Among them, operation anomaly refers to the situation in which performance indicators exceed the normal range or abnormal behavior occurs during system operation. For example, CPU utilization, memory consumption, and network traffic fluctuate violently. These events usually indicate potential performance bottlenecks or resource overloads, which may affect the execution efficiency of computing tasks. The target time is the time when the target job has an abnormality.

[0073] It can be understood that since the standard operating curve can characterize the operating data of the target job under normal operating conditions within the target current time period, when there is a deviation between the actual operating data collected in real time and the standard operating curve, it can be determined that the target job is operating abnormally at the target time within the current time period.

[0074] In order to avoid the waste of resources caused by further expansion of the exception, when it is determined that the target job is running abnormally, it is necessary to immediately save the job running status of the target job at the target time. Exemplarily, the job running status information in the job memory dimension, register dimension and file handle dimension can be saved.

[0075] In addition, when no abnormal operation of the target job is detected, but the current node where the target job is located has abnormal conditions such as low CPU or GPU utilization or computing node crash, the job running status of the target job can also be saved immediately.

[0076] S104, performing exception processing on the target job, and after the exception processing is completed, resuming the operation of the target job according to the operation status.

[0077] After the job running status of the target job is saved, in an optional implementation, an alarm message including the job identifier of the target job and the node information of the current node where the target job is located can be sent to the operation and maintenance end associated with the Slurm scheduling system to prompt the operation and maintenance personnel to handle the exception of the target job in a timely manner.

[0078] In another optional implementation, in order to improve the efficiency of exception handling, corresponding candidate exception handling methods may be configured in advance for candidate abnormal operation forms with higher frequency. When it is determined that the target job is abnormal, based on the abnormal operation form of the target job, a query may be made from the correspondence between each candidate abnormal operation form and each candidate exception handling method to obtain the target exception handling method of the target job, and the target exception handling method may be used to handle the target job abnormally.

[0079] After the exception is handled, in an optional implementation, the job-related information about the target job in the current node where the target job is located can be cleared first, and then the target job can be restored in the current node based on the pre-saved job running status, so that the target job does not need to be rerun, but can continue to run directly from the location where the exception occurred.

[0080] In another optional implementation, in order to avoid recurrence of the same abnormal cause, a new target node may be selected from other nodes except the current node where the target job is located, and the target job may be restored on the target node according to the job running status.

[0081] In the above-mentioned job processing method, the standard operation curve of the target job in the normal operation state in the current period is determined according to the reference operation data of the target job, and when it is determined according to the standard operation curve and the real operation data that the target job is operating abnormally at the target moment in the current period, the operation state of the target job at the target moment is saved; then, the target job is processed abnormally, and after the abnormal processing is completed, the operation of the target job is restored according to the operation state. Compared with the related art, in which the abnormality of the job cannot be detected in time, and the job needs to be re-run after the abnormal job is processed, the above-mentioned method, on the one hand, adopts the standard operation curve, can timely determine the operation abnormality of the target job, and improves the reliability of the operation; on the other hand, in the case of an abnormality of the target job, the operation state is immediately saved, and after the abnormal processing is completed, the operation of the target job is restored according to the operation state, which can effectively improve the efficiency of job recovery, thereby improving the efficiency of job operation.

[0082] It is understandable that the execution process of a job involves data of multiple dimensions. If all the data are used as reference running data, the resource usage of data processing will be increased. Based on this, on the basis of the above embodiment, in this embodiment, an optional method for obtaining reference running data is provided, such as Figure 2 As shown, the specific steps include:

[0083] S201 : Determine a target application type of a target application according to a command name of a target application that submits a target job.

[0084] The target application is the application that submits the target job, the target application type is the type of the target application, and the command name is the name of the command initiated when operating the target job / target application. For example, the command name can be the command prompt cmd name.

[0085] In a high-performance computing environment, in order to effectively manage and optimize resources, the command name corresponding to the target application for submitting the target job can be obtained; then, the target application type of the target application can be determined by analyzing the form and content of the command name. For example, the application type can include but is not limited to compute-intensive, memory-intensive, and input / output (I / O)-intensive.

[0086] For example, after analyzing the form and content of the command name, if job A is determined to be a scientific computing program, numerical simulation, machine learning training, etc., then job A can be determined to be a compute-intensive job; if job B is determined to be a big data processing and database operation, then job B can be determined to be an I / O-intensive job.

[0087] S202 : Determine a target monitoring indicator corresponding to the target application for submitting the target job according to the target application type and the pre-set correspondence relationship between the candidate application types and the candidate monitoring indicators.

[0088] The candidate application types are used to characterize various possible application types; the candidate monitoring indicators are various data indicators existing in the operation data; and the target monitoring indicators are the monitoring indicators required to be obtained under the target application.

[0089] Since different types of applications have different requirements for system resources, for each candidate application type, the candidate monitoring indicators associated with the candidate application type can be determined according to the job running status of the job submitted by the application under the candidate application type; then, based on the candidate monitoring indicators associated with each candidate application type, the corresponding relationship between each candidate application type and each candidate monitoring indicator is constructed.

[0090] In one possible implementation, the type identifier of the target application type may be used as an index to query the constructed correspondence between each candidate application type and each candidate monitoring indicator to obtain the target monitoring indicator required for the target job.

[0091] For example, for computationally intensive applications, such as numerical simulation and scientific computing, such applications mainly rely on the computing power of the CPU, so the monitoring indicators may include CPU utilization, floating point operation performance FLOPS, instructions per cycle IPC, etc. By analyzing the operating data under the above indicators, it can help locate the CPU bottleneck and thus optimize the computing efficiency.

[0092] For memory-intensive applications, such as big data analysis or graphics processing tasks, the performance of such applications is often limited by the access speed and capacity of the memory. Therefore, monitoring indicators can include memory-related performance indicators, such as memory bandwidth, cache hit rate, page error rate, etc. By monitoring the operating data under the above indicators, the utilization of the memory subsystem can be effectively evaluated and a basis for memory optimization can be provided.

[0093] For I / O-intensive applications, such as database operations or file system-intensive access tasks, monitoring indicators may include I / O operation throughput, latency, disk read and write rates, etc. By analyzing the operating data under the above indicators, you can understand the pressure and potential bottlenecks of the disk subsystem and perform targeted I / O optimization.

[0094] S203, obtaining reference operation data of the target job under the target monitoring indicator within a reference period.

[0095] In one possible implementation, after the target monitoring indicator is determined, reference operating data of the target job under the target monitoring indicator within the reference period can be obtained from all operating data of the target job within the reference period based on the target monitoring indicator.

[0096] In addition, since only the reference operating data under the target monitoring indicators are used to predict the standard operating curve, in the subsequent detection process of job operation, in order to improve the accuracy of anomaly detection, it is also possible to only obtain the actual operating data of the target job under the target monitoring indicators in the current period and compare it with the standard operating curve.

[0097] In the embodiment of the present application, by acquiring the operating data of jobs under different application types under different monitoring indicators, the scope of data acquisition can be narrowed, while the rationality of the standard operating curve prediction is improved.

[0098] It is understandable that different types of applications have different requirements for system resources, which may cause noise and abnormal values ​​in the collected operating data under each monitoring indicator. Based on this, on the basis of the above embodiment, in this embodiment, an optional method for determining a standard operating curve is provided, such as Figure 3 As shown, the specific steps include:

[0099] S301 , extracting features from reference operation data according to a data processing method associated with a target application type to obtain reference operation features.

[0100] The so-called reference operation characteristics can characterize the characteristics of the reference operation data.

[0101] It is understandable that since the operating data corresponding to different application types have different characteristics, for each candidate application type, it is necessary to determine the data processing method corresponding to the candidate application type based on the possible abnormalities in the operating data corresponding to the candidate application type; then, based on the data processing method corresponding to each candidate application type, determine the association relationship between each candidate application type and each data processing method.

[0102] In one possible implementation, based on the type identifier of the target application type, a query can be performed in the association relationship between each candidate application type and each data processing method, thereby determining the data processing method associated with the target application type. Furthermore, the data processing method associated with the target application type can be used to perform operations such as data cleaning and feature extraction on the reference operation data to obtain reference operation features corresponding to the reference operation data.

[0103] For example, for computationally intensive applications, it is important to eliminate CPU utilization anomalies and errors caused by short-term fluctuations. Therefore, data cleaning and feature extraction can be performed on the reference running data through sliding window or median filtering technology to smooth the reference running data and reduce the impact of short-term spikes on the prediction.

[0104] For memory-intensive applications, it is necessary to detect and clean up extreme abnormal data caused by occasional memory failures, cache invalidation, or frequent page swapping. Therefore, the stability and accuracy of the reference running data can be ensured by removing the cache hit rate and page error rate that are too high or too low in the reference running data.

[0105] For I / O-intensive applications, special attention should be paid to abnormal data in the reference run data caused by I / O operation delays and sudden disk reads and writes. Therefore, the impact of abnormal data on prediction can be avoided by removing abnormal throughput fluctuation data or extreme latency data in the reference run data.

[0106] S302, determining a standard operation curve of the target operation in a normal operating state in the current period according to the reference operation characteristics.

[0107] In one possible implementation, the reference operating characteristics may be input into a trained second curve prediction model, and the second curve prediction model may output a standard operating curve of the target operation under normal operating conditions in the current period based on the reference operating characteristics and model parameters.

[0108] In another possible implementation, the data variation trend of the operation data can be analyzed based on the reference operation characteristics; then, based on the data variation trend, a standard operation curve of the target operation in the normal operation state in the current time period is determined.

[0109] In the embodiment of the present application, by using different data processing methods for reference operation data under different application types, the rationality of data processing can be guaranteed, thereby improving the accuracy of standard operation curve prediction.

[0110] In order to ensure the accuracy of abnormality determination, based on the above embodiments, in the embodiments of the present application, the real operation data includes the real operation data at each moment in the current period; further, an optional method for determining operation abnormality is provided, such as Figure 4 As shown, the specific steps include:

[0111] S401, for each moment in the current period, determining an operation deviation value between actual operation data corresponding to the moment and standard operation data corresponding to the moment in the standard operation curve.

[0112] Among them, the standard operating data is the operating data in the standard operating curve; the operating deviation value is used to characterize the deviation between the actual operating data and the standard operating data. Furthermore, the larger the operating deviation value, the more serious the deviation.

[0113] In one possible implementation, for each moment in the current time period, the operation deviation value at that moment may be determined according to the deviation between the standard operation data in the standard operation curve at that moment and the actual operation data at that moment.

[0114] Exemplarily, the standard operating data in the standard operating curve at that moment and the actual operating data at that moment can be simultaneously input into a trained deviation value determination model, and the deviation value determination model outputs the operating deviation value at that moment based on the standard operating data, the actual operating data and the model parameters.

[0115] It is understandable that, in order to ensure the high efficiency of abnormal monitoring, after obtaining the real operation data at the current moment, it is not necessary to start the abnormal detection process after obtaining the real operation data at all moments in the current period, but directly compare the real operation data with the standard operation data corresponding to the current moment in the standard operation curve to obtain the operation deviation value at the current moment. Further, the following steps S402-S403 can be continued based on the operation deviation value at the current moment.

[0116] S402, when the operation deviation value is greater than the deviation threshold, stop executing the operation of determining the operation deviation value between the actual operation data corresponding to other moments after the determination moment and the standard operation data corresponding to the standard operation curve.

[0117] The so-called deviation threshold is used to characterize a numerical value that can measure the degree of data deviation, which can be determined based on the experience of relevant technical personnel or based on a large number of experiments, and is not limited in this application.

[0118] In one possible implementation, the operation deviation value can be compared with the deviation threshold. If the operation deviation value is greater than the deviation threshold, it proves that the degree of deviation between the actual operation data of the target job at that moment and the standard operation data is large. Therefore, it can be determined that there is an abnormality in the operation of the target job at that moment.

[0119] At this time, there is no need to perform anomaly detection operations at subsequent moments, and the subsequent anomaly handling link for the target operation can be directly performed. That is, the operation of determining the operation deviation value between the actual operation data corresponding to other moments after the determination moment and the standard operation data corresponding to the standard operation curve is stopped.

[0120] If the operation deviation value is less than or equal to the deviation threshold, it proves that the deviation between the actual operation data and the standard operation data of the target job at this moment is small, that is, the target job is running normally at this moment, so the operation of the operation deviation value between the actual operation data corresponding to other moments after this moment and the standard operation data corresponding to the standard operation curve can continue to be executed.

[0121] S403, taking the time as the target time, and determining whether the target job runs abnormally at the target time in the current period.

[0122] When the running deviation value is greater than the deviation threshold, the moment when the abnormality is determined can be used as the target moment to save the job running status of the target job at the target moment, and determine that the target job runs abnormally at the target moment in the current period.

[0123] In the embodiment of the present application, by determining the target time when the abnormality occurs based on the operation deviation value between the actual operation data and the standard operation data corresponding to each moment in the current time period, the accuracy of the determination of the operation abnormality can be ensured.

[0124] In order to ensure the reliability of job recovery, based on the above embodiments, in an embodiment of the present application, an optional method for restoring the operation of a target job is provided, specifically, a target node is selected from each other node according to the node operation data of each other node; and the operation of the target job is restored on the target node according to the job operation status.

[0125] The other nodes are nodes in the cluster other than the current node running the target job.

[0126] In one feasible implementation, in order to avoid repeated job exceptions of the target job on the current node, a target node can be selected from each other node based on the node operation data of each other node; then, a target job in an interrupted operation state can be regenerated on the target node based on the pre-saved job operation status, and the operation of the target job can be resumed.

[0127] Exemplarily, other nodes with the lowest memory usage can be selected as target nodes based on the node operation data of other nodes; or other nodes whose resource types can be provided by other nodes and are consistent with those of the target node can be selected as target nodes.

[0128] It is understandable that after the target job is resumed on the target node, the job data of the target job on the current node may be released to ensure the memory usage of the current node.

[0129] In the embodiment of the present application, by selecting a target node from each other node according to the node operation data of each other node, and resuming the operation of the target job on the target node, the reliability of the operation of the target job can be guaranteed.

[0130] In order to further ensure the reliability of job recovery, based on the above embodiments, in the embodiments of the present application, another optional method for restoring the operation of the target job is provided, specifically, according to the cause of the job abnormality of the target job, the job resource allocation strategy is determined; according to the job running status and the job resource allocation strategy, the operation of the target job is restored on the target node.

[0131] The job exception reason is the reason why the target job runs abnormally; the job resource allocation strategy is used to represent the strategy for allocating resources for the target job.

[0132] In one feasible implementation, in order to avoid repeated exceptions due to the same reason in the target job, in the process of handling the exception of the target job, the cause of the job exception of the target job can be determined based on the actual operation data of the target job; then, based on the cause of the job exception, the original job resource allocation strategy of the target job can be adjusted to obtain the adjusted job resource allocation strategy.

[0133] For example, if the cause of the job abnormality of the target job is insufficient memory, the allocation amount of memory resources can be increased in the original job resource allocation strategy of the target job to obtain an adjusted job resource allocation strategy.

[0134] Furthermore, a target job in an interrupted state can be regenerated on the target node according to the pre-saved job running state, and the running resources of the target job can be allocated based on the adjusted job resource allocation strategy to resume the running of the target job.

[0135] In an embodiment of the present application, by determining a job resource allocation strategy according to the job exception cause of the target job, and restoring the operation of the target job on the target node based on the job resource allocation strategy, the reliability of job recovery can be ensured.

[0136] Figure 5 FIG. 2 is a flow chart of a method for processing a job in another embodiment. Based on the above embodiment, this embodiment provides an optional example of a method for processing a job. Figure 5 The specific implementation process is as follows:

[0137] S501 : Determine a target application type of a target application according to a command name of a target application that submits a target job.

[0138] S502 : Determine a target monitoring indicator corresponding to the target application for submitting the target job according to the target application type and the pre-set correspondence between the candidate application types and the candidate monitoring indicators.

[0139] S503, obtaining reference operation data of the target job under the target monitoring indicator in the reference time period, and actual operation data of the target job under the target monitoring indicator in the current time period.

[0140] Among them, the reference period is before the current period.

[0141] S504, extracting features from the reference operation data according to the data processing method associated with the target application type to obtain reference operation features, and determining a standard operation curve of the target operation in the current period under normal operation conditions according to the reference operation features.

[0142] S505, determining the operation deviation value between the actual operation data corresponding to each moment in the current period and the standard operation data at the corresponding moment in the standard operation curve.

[0143] S506, when the operation deviation value corresponding to any moment is greater than the deviation threshold, stop executing the operation of determining the operation deviation value between the actual operation data corresponding to other moments after this moment and the corresponding standard operation data in the standard operation curve, and use this moment as the target moment.

[0144] S507, saving the job running status of the target job at the target time, and performing exception processing on the target job.

[0145] S508: Select a target node from each other node according to the node operation data of each other node.

[0146] The other nodes are nodes in the cluster other than the current node running the target job.

[0147] S509: Determine a job resource allocation strategy according to the job exception cause of the target job.

[0148] S510, resuming the operation of the target job on the target node according to the job operation status and the job resource allocation strategy.

[0149] The specific process of the above S501-S510 can refer to the description of the above method embodiment, and its implementation principle and technical effect are similar, which will not be repeated here.

[0150] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0151] Based on the same inventive concept, the embodiment of the present application also provides a job processing device for implementing the job processing method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more job processing device embodiments provided below can refer to the limitations on the job processing method above, and will not be repeated here.

[0152] In an exemplary embodiment, Figure 6 As shown, a job processing device 1 is provided, comprising: a data acquisition module 10, a curve determination module 20, a state saving module 30 and a job recovery module 40, wherein:

[0153] The data acquisition module 10 is used to acquire reference operation data of the target job in a reference period and real operation data of the target job in a current period; wherein the reference period is before the current period;

[0154] The curve determination module 20 is used to determine the standard operation curve of the target operation in the current period under the normal operation state according to the reference operation data;

[0155] The state saving module 30 is used to save the operation state of the target operation at the target time when it is determined that the target operation is abnormal at the target time in the current period according to the standard operation curve and the actual operation data;

[0156] The job recovery module 40 is used to perform exception processing on the target job, and after the exception processing is completed, perform operation recovery on the target job according to the operation status of the job.

[0157] In an exemplary embodiment, the data acquisition module 10 is specifically used for:

[0158] Determine the target application type of the target application according to the command name of the target application for submitting the target job; determine the target monitoring indicator corresponding to the target application for submitting the target job according to the target application type and the correspondence between the pre-set candidate application types and the candidate monitoring indicators; obtain reference operating data of the target job under the target monitoring indicator within a reference time period.

[0159] In an exemplary embodiment, the curve determination module 20 is specifically used for:

[0160] According to the data processing method associated with the target application type, feature extraction is performed on the reference operation data to obtain reference operation features; based on the reference operation features, a standard operation curve of the target operation under normal operating conditions in the current period is determined.

[0161] In an exemplary embodiment, the real operation data includes the real operation data at each moment in the current period; the state saving module 30 is specifically used for:

[0162] For each moment in the current time period, determine the operating deviation value between the actual operating data corresponding to the moment and the standard operating data corresponding to the moment in the standard operating curve; when the operating deviation value is greater than the deviation threshold, stop executing the operation of determining the operating deviation value between the actual operating data corresponding to other moments after the moment and the standard operating data corresponding to the standard operating curve; take the moment as the target moment, and determine the operating abnormality of the target job at the target moment in the current time period.

[0163] In an exemplary embodiment, the job recovery module 40 includes:

[0164] A node determination unit, configured to select a target node from each other node according to the node operation data of each other node; wherein each other node is a node in the cluster other than the current node running the target job;

[0165] The job recovery unit is used to recover the operation of the target job on the target node according to the job running status.

[0166] In an exemplary embodiment, the job recovery unit is specifically used to:

[0167] Determine the job resource allocation strategy based on the job exception cause of the target job; and resume the operation of the target job on the target node based on the job running status and the job resource allocation strategy.

[0168] Each module in the above-mentioned job processing device can be implemented in whole or in part by software, hardware or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.

[0169] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 7 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store job operation data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a job processing method is implemented.

[0170] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0171] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0172] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0173] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0174] It should be noted that the data involved in this application (including but not limited to job operation data, etc.) are all data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0175] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0176] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0177] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for processing a job, characterized in that: The method comprises: Acquire reference operation data of a target job in a reference period and real operation data of the target job in a current period; wherein the reference period is before the current period; Determine, according to the reference operation data, a standard operation curve of the target operation in the current period under normal operation; When it is determined according to the standard operation curve and the actual operation data that the target operation is abnormally operating at the target time within the current period, the operation operation state of the target operation at the target time is saved; An exception process is performed on the target job, and after the exception process is completed, the target job is restored to operation according to the operation status of the job.

2. The method according to claim 1, characterized in that The obtaining of reference operation data of the target operation within a reference period includes: Determining a target application type of the target application according to a command name of a target application for submitting a target job; Determine the target monitoring indicator corresponding to the target application for submitting the target job according to the target application type and the pre-set correspondence between the candidate application type and the candidate monitoring indicator; Obtain reference operating data of the target job under the target monitoring indicator within a reference time period.

3. The method according to claim 2, characterized in that Determining, according to the reference operation data, a standard operation curve of the target operation in the current period under a normal operation state includes: Extracting features from the reference operation data according to a data processing method associated with the target application type to obtain reference operation features; A standard operation curve of the target operation in the current time period under normal operating conditions is determined according to the reference operation characteristics.

4. The method according to claim 1, characterized in that: The real operation data includes the real operation data at each moment in the current period; The determining, based on the standard operation curve and the actual operation data, that the target operation is operating abnormally at the target time within the current period includes: For each moment in the current period, determining an operation deviation value between the actual operation data corresponding to the moment and the standard operation data corresponding to the moment in the standard operation curve; When the operation deviation value is greater than the deviation threshold, stopping the operation of determining the operation deviation value between the real operation data corresponding to other moments after the moment and the standard operation data corresponding to the standard operation curve; The time is taken as the target time, and it is determined that the target job runs abnormally at the target time within the current period.

5. The method according to claim 1, characterized in that The step of restoring the target job according to the job running status includes: Selecting a target node from each other node according to the node operation data of each other node; wherein each other node is a node in the cluster except the current node running the target job; According to the job running status, the running of the target job is resumed on the target node.

6. The method according to claim 5, characterized in that The resuming the operation of the target job on the target node according to the operation status of the job includes: Determine a job resource allocation strategy according to the job abnormality cause of the target job; According to the job running state and the job resource allocation strategy, the running of the target job is resumed on the target node.

7. A job processing device, characterized in that: The device comprises: A data acquisition module, used to acquire reference operation data of a target job in a reference period and real operation data of the target job in a current period; wherein the reference period is before the current period; A curve determination module, used to determine a standard operation curve of the target operation in a normal operating state within the current period according to the reference operation data; A state saving module, configured to save the operation state of the target operation at the target time when it is determined that the target operation is abnormal at the target time within the current period according to the standard operation curve and the real operation data; The job recovery module is used to perform exception processing on the target job and, after the exception processing is completed, perform operation recovery on the target job according to the operation status of the job.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Electrical equipment state judgment and fault diagnosis method and device

    CN112782512A

  • Abnormal page recovery method and device, computer equipment and storage medium

    CN114064338A

  • Abnormal information determination method and device, computer equipment and storage medium thereof

    CN115421955A

  • Non-transitory computer-readable storage medium and printing system

    US20220317955A1