Task scheduling fault processing method and device, equipment, storage medium and program product

By evaluating the popularity of the scheduling model and the popularity of faults of the big data task, the UTRF classification model optimization processing strategy is adopted to solve the problem of task recovery when resources are limited, and the operation efficiency of the big data task is improved.

CN120492208APending Publication Date: 2025-08-15中移信息技术有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510670945.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the scheduling of big data tasks, the core tasks cannot be quickly restored when resources are limited, and the post-processing method is difficult to deal with a large number of failures, affecting the task operation efficiency.

Method used

By determining the popularity of the use and failure of the scheduling model of data tasks, the UTRF classification model is used for evaluation, and the pre-intervention processing strategy and post-processing strategy are formulated to optimize the failure rate and processing speed of the model.

Benefits of technology

The operation efficiency of data tasks is improved, and the failure rate is reduced through pre-prevention and post-optimization processing is carried out, which improves the stability and operation efficiency of tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492208A_ABST
    Figure CN120492208A_ABST
Patent Text Reader

Abstract

The invention discloses a task scheduling fault processing method and device, equipment, a storage medium and a program product, and relates to the technical field of task scheduling, the task scheduling fault processing method comprises the steps that the use heat and the fault heat of a scheduling model of a data task are determined, and the scheduling model is used for realizing data processing operation corresponding to the data task; and performing fault processing on the data task according to the use popularity and the fault popularity of the scheduling model of the data task. According to the scheme, the fault rate can be reduced through beforehand intervention of the data task, and resource allocation is optimized, so that the operation efficiency of the data task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of task scheduling technology, and in particular to a task scheduling fault handling method, apparatus, device, storage medium, and program product. Background Art

[0002] Existing scheduling for big data tasks mostly manages the task running status through the development tool side, and issues alarms based on indicators such as task failure and task operation duration to solve task failure problems. It is necessary to receive alarms and locate them before intervening to solve them.

[0003] However, with the substantial growth of big data tasks, the drawbacks of the existing task scheduling model are becoming increasingly apparent. On the one hand, when there are many tasks and limited resources, it is impossible to determine which task to restore first when a large number of alarms occur at the same time, and the core tasks cannot be quickly restored within limited time and resources. On the other hand, with an increasing number of failures, the current situation is difficult to meet the growing number of big data task failures through simple post-event manual inspections or simple post-processing of alarms, thus affecting the operating efficiency of big data tasks.

[0004] Therefore, how to reduce the failure rate of big data tasks and optimize the post-processing speed of failed tasks to improve the operating efficiency of big data tasks is a technical problem that needs to be solved urgently.

[0005] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention

[0006] The main purpose of this application is to provide a task scheduling fault handling method, device, equipment, storage medium and program product, aiming to solve the technical problem of how to reduce the failure rate of big data tasks and optimize the post-processing speed of faulty tasks to improve the operating efficiency of big data tasks.

[0007] To achieve the above objectives, the present application proposes a method for handling task scheduling failures, the method comprising:

[0008] Determining a usage heat and a failure heat of a scheduling model for a data task, the scheduling model being used to implement a data processing operation corresponding to the data task;

[0009] Fault processing is performed on the data task according to the usage heat and fault heat of the scheduling model of the data task.

[0010] In one embodiment, the step of determining the usage heat and the failure heat of the scheduling model of the data task includes:

[0011] Determining, based on scheduling metadata of a plurality of data tasks, a usage heat index and a fault heat index of a scheduling model of the data tasks;

[0012] Determining a usage heat category of the scheduling model according to the usage heat index;

[0013] Determining a fault heat category of the scheduling model according to the fault heat index;

[0014] The step of performing fault processing on the data task according to the usage heat and fault heat of the scheduling model of the data task includes:

[0015] Fault processing is performed on the data task according to the usage heat category and the fault heat category of the scheduling model of the data task.

[0016] In one embodiment, the step of performing fault processing on the data task according to the usage heat category and the fault heat category of the scheduling model of the data task includes:

[0017] Determining a fault handling priority of the data task according to a correspondence between a combination of a usage heat category and a fault heat category and a fault handling level, and the usage heat category and the fault heat category of a scheduling model of the data task;

[0018] Fault processing is performed on the data task according to the fault processing priority.

[0019] In one embodiment, the fault handling priority includes a fault prevention processing priority and a fault recovery processing priority, and the step of performing fault handling on the data task according to the fault handling priority includes:

[0020] According to the fault prevention processing priority, the scheduling model of the data task is optimized and resource tilted;

[0021] After a failure occurs in the data task, the scheduling model of the data task is restored according to the failure recovery processing priority.

[0022] In one embodiment, the task scheduling fault handling method further includes:

[0023] According to the fault recovery processing priority, an automatic re-run classification mechanism is implemented for the scheduling model of the data task, and the automatic re-run classification mechanism is classified based on the number of automatic re-runs and the automatic re-run time.

[0024] In one embodiment, the fault heat index includes the number of faults in the model cycle, the interval between the last model fault and the average execution time in the task cycle; and / or, the usage heat index includes the number of calls in the model cycle and the interval between the last model call.

[0025] In one embodiment, after the step of determining the usage heat index and the fault heat index of the scheduling model of the data tasks based on the scheduling metadata of the plurality of data tasks, the method further includes:

[0026] The number of failures in the model cycle is corrected according to the average execution time and scheduling operation information in the task cycle.

[0027] In one embodiment, before the step of determining the usage heat category of the scheduling model based on the usage heat index and before the step of determining the fault heat category of the scheduling model based on the fault heat index, the method further includes:

[0028] Collect usage heat index and failure heat index of historical scheduling models of several data tasks;

[0029] Based on expert prior knowledge, a cluster analysis algorithm is used to analyze the usage heat index and the fault heat index of the historical scheduling model, and the classification boundary values of different types of scheduling models are determined to obtain model classification parameters; the model classification parameters are used to classify the usage heat category and / or the fault heat category to obtain the usage heat category and / or the fault heat category.

[0030] In addition, to achieve the above-mentioned purpose, the present application also proposes a task scheduling fault processing device, which includes:

[0031] a determination module, configured to determine a usage heat and a failure heat of a scheduling model of a data task, wherein the scheduling model is used to implement a data processing operation corresponding to the data task;

[0032] The fault processing module is used to perform fault processing on the data task according to the usage heat and fault heat of the scheduling model of the data task.

[0033] In addition, to achieve the above-mentioned purpose, the present application also proposes a task scheduling fault handling device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the task scheduling fault handling method described above.

[0034] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the task scheduling fault handling method described above are implemented.

[0035] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the task scheduling fault handling method as described above.

[0036] One or more technical solutions proposed in this application have at least the following technical effects:

[0037] The task scheduling fault handling method, apparatus, device, storage medium, and program product proposed in the embodiments of the present application specifically determine the usage heat and fault heat of a scheduling model of a data task, where the scheduling model is used to implement the data processing operation corresponding to the data task; and perform fault handling on the data task based on the usage heat and fault heat of the scheduling model of the data task.

[0038] This application first determines the usage popularity and fault popularity of the scheduling model corresponding to the data task, and further determines the pre-intervention processing operation of the scheduling model based on the usage popularity and fault popularity of the scheduling model, optimizes the failure rate of the model, and at the same time determines the fault problem handling operation of the scheduling model, optimizes the post-processing speed of the fault task, thereby realizing fault handling of the data task and improving the operating efficiency of the data task. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0041] Figure 1 A flowchart of the first embodiment of the method for handling task scheduling failures provided in this application;

[0042] Figure 2 A flowchart of the second embodiment of the method for handling task scheduling failures provided in this application;

[0043] Figure 3 A flowchart of the third embodiment of the method for handling task scheduling failures provided in this application;

[0044] Figure 4 A flowchart of the fourth embodiment of the method for handling task scheduling failures provided in this application;

[0045] Figure 5 This is a schematic diagram of the module structure of the task scheduling fault processing device according to an embodiment of the present application;

[0046] Figure 6 Schematic diagram of the device structure of the hardware operating environment involved in the task scheduling fault handling method in the embodiment of the present application.

[0047] The purpose, features and advantages of this application will be further explained with reference to the accompanying drawings in conjunction with the embodiments. DETAILED DESCRIPTION

[0048] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0049] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0050] The main solution of the embodiment of the present application is: determining the usage heat and fault heat of the scheduling model of the data task, the scheduling model is used to implement the data processing operation corresponding to the data task; and performing fault processing on the data task according to the usage heat and fault heat of the scheduling model of the data task.

[0051] Technical terms involved in this application:

[0052] RFM Classification Model: RFM (Recency, Frequency, Monetary), also known as the Customer Value Classification Model, is a classic customer value analysis method widely used in marketing, customer relationship management, sales forecasting, and other fields. By quantitatively evaluating three key customer behavioral dimensions, the RFM classification model helps companies identify high-value customers, develop targeted marketing strategies, and optimize resource allocation. The RFM classification model calculates an RFM score based on the customer's most recent purchase date (R), purchase frequency (F), and purchase amount (M). These three dimensions are used to assess the customer's order activity value and are often used to segment or differentiate customer segments, such as high-value customers (high R, high F, high M), potential customers (low R, low F, low M), and churned customers (low R, low F, low M). The RFM classification model analyzes the model at a fixed point in time, so RFM results calculated at different times may vary.

[0053] In this embodiment, for ease of description, the following description is made with the task scheduling fault handling device as the execution subject.

[0054] Existing scheduling for big data tasks mostly manages the task running status through the development tool end, and performs manual inspections or abnormal alarms on the front-end page for each task's basic running status, such as the execution start time, end time, and task status. When problems are found, intervention is made to rerun or perform other recovery operations.

[0055] With the rapid growth of big data tasks, the drawbacks of the existing model are becoming increasingly apparent. First, with numerous tasks and limited resources, it's difficult to prioritize recovery when numerous alarms appear simultaneously, preventing the rapid recovery of core tasks within limited time and resources. Second, with the increasing number of failures, it's impossible to reduce the failure rate by optimizing specific models beforehand. The current approach of simply relying on post-event manual inspections or simply processing alarms after they are received is unable to meet the growing volume of big data task failures.

[0056] The present application provides a solution, which first determines the usage popularity and fault popularity of the scheduling model corresponding to the data task, and then further determines the pre-intervention processing operation of the scheduling model based on the usage popularity and fault popularity of the scheduling model, optimizes the failure rate of the model, and at the same time determines the fault problem processing operation of the scheduling model, optimizes the post-processing speed of the fault task, thereby realizing fault processing of the data task and improving the operating efficiency of the data task.

[0057] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device capable of implementing the above functions, a task scheduling fault handling device, etc. The following describes this embodiment and the following embodiments using the task scheduling fault handling device as an example.

[0058] Based on this, the embodiment of the present application provides a method for handling task scheduling failures, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the task scheduling fault handling method of this application.

[0059] In this embodiment, the task scheduling fault handling method includes steps S110 to S120:

[0060] Step S110, determining the usage heat and fault heat of a scheduling model of a data task, wherein the scheduling model is used to implement a data processing operation corresponding to the data task;

[0061] It should be noted that the present application is applied to fault handling of big data tasks, where big data tasks refer to tasks that process big data using a scheduling model, that is, data tasks, and the scheduling model refers to the data processing logic corresponding to the data task implemented through logical code. By calling the scheduling model of the data task, the data task can be run.

[0062] Understandably, existing data task scheduling often relies on development tools to manage task status. This involves manual inspections or exception alerts on the front-end page for each data task's basic status, such as its start time, end time, and task status. Once a problem is discovered, the task is re-run or undergoes other recovery operations. However, this post-processing approach to task scheduling is ineffective for handling large numbers of data tasks, effectively reducing the failure rate of data tasks, and impacting the efficiency of data tasks.

[0063] Therefore, in this embodiment, the task scheduling fault processing device first obtains the usage information and operation information of the scheduling model of the data task through a certain analysis algorithm (such as statistical analysis, machine learning model) through step S110, evaluates the usage heat and fault heat of the scheduling model, and determines the usage heat score or classification, as well as the fault heat score or classification of the scheduling model.

[0064] In a feasible implementation, the step of determining the usage heat and the fault heat of the scheduling model of the data task includes steps S1101 to S1103:

[0065] Step S1101, determining a usage heat index and a fault heat index of a scheduling model of a plurality of data tasks based on scheduling metadata of the data tasks;

[0066] It should be noted that the usage heat index and fault heat index of the scheduling model refer to data indicators used to evaluate the call heat and health status of the scheduling model of data tasks, such as indicators for health status evaluation from aspects such as call frequency, related task dependency, fault frequency, related business, and task timeliness, including but not limited to predicted failure probability, historical number of failures, historical number of calls, current operating status, number of related data tasks, and importance or level of related business.

[0067] Specifically, use log collection tools to capture the scheduling metadata of multiple data tasks, or obtain the scheduling metadata of multiple data tasks from the batch task development tool. Scheduling metadata refers to the multi-dimensional scheduling information related to the data task during its operation, such as the relevant scheduling model, upstream and downstream task call chains, data task dependencies, data task status, the start and end time of each task operation, the time when task failure occurs and the end time, etc.

[0068] Subsequently, in order to further improve the operating efficiency of multiple data tasks, the scheduling metadata is processed through certain algorithms (such as statistical analysis, machine learning models) to generate usage heat indicators and fault heat indicators of the scheduling model that reflect the health status of each data task, so that the usage heat indicators and fault heat indicators of the scheduling model can be used to perform task recovery or fault prevention on the data tasks.

[0069] Furthermore, the step S1101 further includes steps A01 to A02:

[0070] Step A01: Obtain the scheduling model, scheduling operation information, and scheduling fault information in the corresponding scheduling metadata according to the data task;

[0071] Step A02: Count the usage heat index and the fault heat index according to the scheduling operation information and the scheduling fault information and the type of the scheduling model.

[0072] Specifically, the task scheduling fault handling device first categorizes the collected scheduling metadata of multiple data tasks according to their names, obtaining the categorized scheduling metadata corresponding to each data task. It then retrieves specific field information from the categorized scheduling metadata, such as the scheduling model, scheduling operation information, scheduling fault information, associated data task information, and associated business attribute information. Each data task has a one-to-one correspondence with the model it calls, and each called model is used to process a single data task.

[0073] Subsequently, the above field information is processed using a preset statistical analysis algorithm, and classified and summarized according to different types of scheduling models, and finally the usage heat index and fault heat index data of the scheduling model corresponding to each data task are sorted out. Among them, the usage heat index of the scheduling model refers to the usage heat index related to the scheduling model that is customized by relevant personnel according to the scheduling requirements or experience of the data task, and the fault heat index of the scheduling model refers to the fault heat index related to the scheduling model that is customized by relevant personnel according to the scheduling requirements or experience of the data task. The usage heat index of the scheduling model includes but is not limited to the interval time of the model's most recent call, the number of calls within the model cycle, the importance of the model-related business, etc. The fault heat index of the scheduling model includes but is not limited to the interval time of the model's most recent failure, the number of failures within the model cycle, the average execution time within the task cycle, the probability of future failures, the average recovery time within the model cycle, the resource utilization rate within the cycle, etc. Among them, the cycle can be pre-set by relevant personnel based on actual needs.

[0074] The preset statistical analysis algorithm refers to the statistical analysis algorithm or program pre-set by relevant personnel according to the usage heat index and fault heat index of the selected scheduling model, such as time series analysis, descriptive analysis (such as mean, variance, etc.) or statistical algorithm of custom indicators.

[0075] In this embodiment, referring to Table 1, the task scheduling fault processing device obtains the scheduling model, scheduling operation information, scheduling fault information, and related data task information from the above-mentioned classified scheduling metadata and stores them in the database, wherein the scheduling operation information includes the scheduling start time and the scheduling end time; the scheduling fault information includes the scheduling fault time and the scheduling recovery time; and the related data task information includes the lower-level task scheduling name and the lower-level model name. Referring to Table 1, Table 1 obtains field information such as the task scheduling name, model name (i.e., scheduling model), lower-level task scheduling name, lower-level model name, scheduling start and end time, scheduling fault and recovery time for the data scheduling task of the scenic spot zipper daily details. The corresponding associated lower-level data scheduling task is the data scheduling task of the base station residence day details.

[0076] Table 1 Summary of scheduling metadata extraction field information

[0077]

[0078] Specifically, the usage heat index of the scheduling model selected by the task scheduling fault handling device is the model's most recent call interval UR and the number of calls within the model cycle UF. The fault heat index of the scheduling model is the model's most recent fault interval TR, the number of faults within the model cycle TF, and the average execution time within the task cycle. The following is a detailed explanation of the usage heat index and fault heat index of the above scheduling model:

[0079] UR (Use Recency): The time interval between the last time a scheduling model was used and the current time. UF (Use Frequency): The number of times a scheduling model was used within a cycle, which indicates the degree of dependence of the scheduling model corresponding to the lower-level data task on the scheduling model at this level. TR (Trouble Recency): The time interval between the last time a scheduling model failed and the current time. TF (Trouble Frequency): The number of times a scheduling model failed within a cycle, which indicates the failure frequency of the scheduling model. The average runtime within a task cycle is the total runtime of the data task within the cycle divided by the number of UF calls within the model cycle.

[0080] First, the type of the data task's scheduling model is determined using a specific field or identifier to determine whether the scheduling model is a daily call model or a monthly call model. The distinction between the daily call model and the monthly call model is based on the time dimension of the model's statistical data. In this embodiment, the scheduling model type can be more finely classified based on the time dimension of the model's statistical data, which is not limited here.

[0081] It should be understood that due to the different time dimensions of model statistical data, in order to ensure the applicability and effectiveness of the usage heat index and fault heat index data of the scheduling model, the usage heat index and fault heat index to be counted corresponding to different scheduling model types have different statistical periods. For example, the statistical period for the monthly call model is the last three months, and the statistical period for the daily call model is the last 7 days, that is, the last week.

[0082] Subsequently, based on the preset statistical analysis algorithm, the scheduling operation information and scheduling fault information of data tasks under different scheduling model types are calculated and summarized to obtain the usage heat index and fault heat index data of the scheduling model corresponding to each data task.

[0083] It should be clear that the acquired specific field information can be set by the relevant personnel based on the data required by the usage heat index and the fault heat index of the scheduling model, and is not limited to the above-mentioned specific field information.

[0084] In this embodiment, by summarizing and processing the scheduling operation information and scheduling fault information in the scheduling metadata, the usage heat index and fault heat index of the scheduling model of the data task can be further obtained, so as to optimize the scheduling processing of the data task according to the usage heat index and fault heat index of the scheduling model and improve the operation efficiency of the data task.

[0085] Step S1102: determining the usage popularity category of the scheduling model according to the usage popularity index;

[0086] Step S1103: Determine the fault heat category of the scheduling model according to the fault heat index.

[0087] The task scheduling fault handling device then classifies the scheduling model into different usage heat categories and fault heat categories based on preset usage heat and fault heat classification standards. The preset usage heat and fault heat classification standards can be classification rules or models pre-set by relevant personnel based on experience or actual needs. In this embodiment, the task scheduling fault handling device uses a heat classification model to perform classification processing based on the aforementioned usage heat index and fault heat index.

[0088] It should be noted that the heat classification model is a strategy for relevant personnel to distinguish between hot and cold data tasks, and is designed for data task fault handling with reference to the statistical analysis method of the customer value classification model. It is used to evaluate the call heat and fault heat of the data task scheduling model, and to classify the heat of the data task scheduling model based on the use heat and fault heat, namely the UTRF classification model. Among them, the customer value classification model can be an RFM classification model, a customer life cycle value model, a customer pyramid model, etc. In this embodiment, the customer value classification model is preferably an RFM classification model.

[0089] In this embodiment, the UTRF classification model constructs model fault indicators with reference to indicators such as the time from the last user consumption to the present and the frequency of user repeated consumption within a certain period of time in the RFM classification model.

[0090] It should be noted that model classification parameters refer to a series of parameters used to define and adjust the UTRF classification model. Model classification parameters determine how to classify data tasks based on input indicator data (such as number of calls, call interval time, number of failures, failure interval time, etc.), including but not limited to threshold settings, weight coefficients, classification rules, etc., which are usually pre-set by relevant personnel based on experience or actual needs.

[0091] Specifically, the task scheduling fault handling device uses the UTRF classification model to perform usage heat classification judgment on the number of calls UF within the model cycle and the interval time UR of the model's most recent call according to the model classification parameters, and obtains the usage heat classification result of the scheduling model, that is, the usage heat category, among which the usage heat of the scheduling model reflects the importance of model scheduling.

[0092] The task scheduling fault handling device uses the UTRF classification model to perform model fault heat classification judgment on the number of faults TF within the model cycle and the model's most recent fault interval TR indicators according to the model classification parameters, and obtains the fault heat classification result of the scheduling model, that is, the fault heat category. Among them, the fault heat of the scheduling model reflects the urgency of the model fault handling.

[0093] Among them, the fault heat classification results and the number of levels of the fault classification results can be pre-set by relevant personnel, such as the classification of model heat of high heat, medium heat, low heat, general heat, high coldness, medium coldness, and low coldness, as well as the classification of fault heat such as high fault, general fault, and low fault.

[0094] In this embodiment, the parameter indicators selected for the UTRF classification model are different according to different types of scheduling models. For example, the parameter indicators of the daily scheduling model are designed as follows, and the "interval time between the last call of the model within a week (days)", "number of calls of the model within a week", "interval time between the last failure of the model (days)", and "number of failures of the model within a week" are used as the parameters of the daily scheduling model; the parameter indicators of the monthly scheduling model are designed as follows, and the "interval time between the last call of the model (month)", "number of calls of the model within three months", "interval time between the last failure of the model (month)", and "number of failures of the model within three months" are used as the parameters of the monthly scheduling model. The embodiment of the present application can make a more detailed classification of the types of scheduling models based on the time dimension of the model statistical data, and then the parameter indicators selected for the UTRF classification model can be adjusted according to the actual situation. For example, the statistical period of each indicator can be reduced to the time dimension of days, hours, or minutes, which is not limited here.

[0095] Step S120 : performing fault processing on the data task according to the usage heat and fault heat of the scheduling model of the data task.

[0096] Specifically, certain analysis algorithms (such as statistical analysis and machine learning models) are used to evaluate the activity of the scheduling model in actual use and the activity of failures according to the usage popularity and failure popularity of the scheduling model of data tasks. Based on this, targeted fault handling strategies are formulated to ensure that data tasks can run efficiently and stably.

[0097] Furthermore, in combination with an implementation of the aforementioned step S110, the step S120 further includes step B01:

[0098] Step B01 : performing fault processing on the data task according to the usage heat category and the fault heat category of the scheduling model of the data task.

[0099] Specifically, the task scheduling fault handling device determines the usage heat category and fault heat category of the scheduling model based on the usage heat and fault heat classification results of the scheduling model of the data task according to the UTRF classification model, and thus determines the pre-fault prevention strategy and post-fault handling strategy of the data task based on the different categories of the usage heat category and the fault heat category, and performs pre-optimization processing and post-fault handling on the scheduling model of the data task according to the pre-fault prevention strategy and post-fault handling strategy to improve the operating efficiency of the data task.

[0100] Furthermore, in combination with an implementation of the aforementioned step S110, after step S1101, the following steps are further included:

[0101] Step C01 , correcting the number of faults in the model cycle according to the average execution time in the task cycle and the scheduling operation information.

[0102] Specifically, the average execution time within a task cycle usually reflects the performance of the data task under normal circumstances, and the number of failures within a model cycle is usually the number of times the scheduling model has frequent retries, abnormal exits, etc. within the statistical period. In order to improve the optimization effect of the fault heat index data of the scheduling model on the scheduling model, this embodiment needs to judge the overload operation of the data task, that is, the operation situation where the running time of the data task is greater than the average running time within the data task cycle, and regard this operation situation as a model failure. Therefore, this embodiment needs to judge the number of overloaded operations of the data task within the cycle based on the average execution time within the task cycle and the scheduling operation information, add the number of overloaded operations to the number of failures within the model cycle, and realize the correction and update of the number of failures within the model cycle, and use the corrected number of failures within the model cycle as the input parameter indicator of the UTRF classification model to improve the accuracy of the UTRF classification model in fault classification based on the corrected number of failures within the model cycle.

[0103] Through the above scheme, this embodiment first determines the usage popularity and fault popularity of the scheduling model corresponding to the data task, and further determines the pre-intervention processing operation of the scheduling model based on the usage popularity and fault popularity of the scheduling model to optimize the failure rate of the model. At the same time, it determines the fault problem processing operation of the scheduling model to optimize the post-processing speed of the fault task, thereby realizing fault processing of the data task and improving the operating efficiency of the data task.

[0104] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 2 , the step B01 includes steps S210 to S220:

[0105] Step S210: determining a fault handling priority of the data task based on a correspondence between a combination of a usage heat category and a fault heat category and a fault handling level, and the usage heat category and the fault heat category of the scheduling model of the data task;

[0106] Step S220: performing fault processing on the data task according to the fault processing priority.

[0107] In this embodiment, the task scheduling fault processing device combines the usage heat and fault heat category of the scheduling model of the data task according to the UTRF classification model to obtain a combined classification result. For example, the combined classification result of general heat + high fault is a more important optimization model.

[0108] Then, based on the combined classification results, combined with the correspondence between the pre-set combination of usage heat category and fault heat category and the fault handling level, the fault handling priority of the data task is determined. Subsequently, based on the fault handling priority, the pre-fault prevention strategy and post-fault handling strategy of the data task are determined, and the data task is optimized in advance and handled in post-fault handling according to the pre-fault prevention strategy and post-fault handling strategy.

[0109] The correspondence between the pre-set combination of usage heat category and fault heat category and the fault handling level refers to the fault handling priority corresponding to the classification result of the combination of model usage heat and model fault heat set by relevant personnel based on experience, actual needs, or custom rules.

[0110] This embodiment adopts the above scheme, specifically the UTRF classification model, and performs a model heat evaluation on the number of calls within the model cycle and the interval time of the model's most recent call according to the model classification parameters to obtain a usage heat category, and performs a fault heat evaluation on the number of failures within the model cycle and the interval time of the model's most recent failure to obtain a fault heat category, thereby understanding the usage heat and failure heat of the scheduling model. Then, the combined classification result of the scheduling model corresponding to the data task can be obtained by combining the usage heat category and the fault heat category, so as to further obtain the fault handling priority according to the combined classification result, so as to determine the core important scheduling model according to the fault handling priority, and optimize the operation efficiency of the data task by performing pre-fault prevention processing and post-fault priority processing on the core important scheduling model.

[0111] Based on the first or second embodiment of the present application, in the third embodiment of the present application, the same or similar contents as those in the first and / or second embodiment can be referred to the above introduction and will not be described in detail later. Figure 3 Before step S1102 and step S1103, steps S310 to S320 are also included:

[0112] Step S310, collecting usage heat indexes and fault heat indexes of historical scheduling models of several data tasks;

[0113] Step S320: Based on the expert prior knowledge, a cluster analysis algorithm is used to analyze the usage heat index and the fault heat index of the historical scheduling model, and the classification boundary values of different types of scheduling models are determined to obtain model classification parameters; the model classification parameters are used to classify the usage heat category and / or the fault heat category to obtain the usage heat category and / or the fault heat category.

[0114] Specifically, to obtain the model classification parameters of the UTRF classification model, the task scheduling fault processing device needs to collect the scheduling metadata of several previous data tasks and extract the usage heat index and fault heat index data of the historical scheduling models of data tasks under different model types from the previous scheduling metadata. The model classification parameters are used to classify the usage heat category and / or the fault heat category to obtain the usage heat category and / or the fault heat category.

[0115] Then, relying on cluster analysis algorithms, data analysis is performed on the usage heat index and fault heat index data of the historical scheduling models of data tasks under different model types, such as the K-means algorithm and the hierarchical clustering method, to obtain the natural grouping boundaries of the scheduling models of data tasks in terms of hot and cold characteristics, and further combined with the prior knowledge of experts familiar with the field of task scheduling fault handling, the natural grouping boundaries of the usage heat index and fault heat index of the scheduling models obtained by clustering are verified and adjusted, such as adjusting the weight coefficients of the UF and UR indicators used to measure the usage heat of the model, or adjusting the threshold limits of the UF and UR indicators of the model heat classification of different models. The task scheduling fault handling device can also apply the classification boundary values of the usage heat index and fault heat index of the scheduling models of the above-mentioned different types of scheduling models to the UTRF classification model, and perform usage heat classification and fault heat classification on other unseen data tasks to further optimize the classification boundary values of the usage heat index and fault heat index of the different types of scheduling models, and finally obtain the model classification parameters.

[0116] In this embodiment, using the daily and monthly call models as examples, model heat is categorized into hot, general, and cold models, and model fault heat (i.e., model fault degree) is categorized into high fault, general fault, and low fault. Tables 2 and 3 below provide exemplary classification thresholds for the usage heat index and fault heat index of the scheduling model under different hot and cold categories for the daily and monthly call models, respectively.

[0117] Table 2 Classification boundaries of the daily call model under different hot and cold categories

[0118]

[0119] Table 3 Classification boundaries of the call model in different cold and hot categories

[0120]

[0121] This embodiment adopts the above scheme, specifically by collecting the usage heat index and fault heat index data of the historical scheduling models of several data tasks; then, based on the cluster analysis algorithm and expert prior knowledge, the usage heat index and fault heat index data of the historical scheduling models are analyzed to determine the classification boundary values of different types of scheduling models to obtain model classification parameters, so as to facilitate the application of the model classification parameters to the UTRF classification model, thereby improving the accuracy of the usage heat classification and fault heat classification of the scheduling models of data tasks.

[0122] Based on the first embodiment, the second embodiment, or the third embodiment of the present application, in the fourth embodiment of the present application, the same or similar contents as those in the first embodiment, the second embodiment, and / or the third embodiment can be referred to above and will not be described in detail. Figure 4 The step S220 further includes steps S410 to S420:

[0123] Step S410: Optimizing the scheduling model of the data task and adjusting resource tilt according to the fault prevention processing priority;

[0124] Step S420 : After a failure occurs in the data task, the scheduling model of the data task is restored according to the failure recovery priority.

[0125] Specifically, referring to the classification results of usage heat and fault heat for model heat and fault heat in the third embodiment, and referring to Table 4 below, the three classifications of "hot model", "general model", and "cold model" obtained according to model heat and the three classifications of "high fault", "general fault", and "low fault" obtained according to model fault heat are combined to form 9 combination classification results of the scheduling model corresponding to the data task.

[0126] According to the ex ante fault prevention processing priority and ex post fault recovery processing level corresponding to the 9 combination classification results designed in Table 4 below, the ex ante fault prevention processing priority and ex post fault recovery processing level of the scheduling model corresponding to the data task are determined.

[0127] Specifically, in this application, the correspondence between the pre-fault prevention processing priority and the combined classification result depends on the expert prior knowledge, while the post-fault recovery processing priority is first sorted in descending order based on the model heat, and further sorted in descending order based on the fault heat in the same level model heat classification.

[0128] Table 4. Classification table of fault handling priorities of scheduling model

[0129]

[0130] Then, the task scheduling fault processing device refers to the fault prevention processing priority to determine the model optimization priority or importance of the scheduling model corresponding to the data task so that relevant personnel can perform model optimization processing, and adjust the response order of storage and computing queue resources according to the fault prevention processing priority, give priority to allocating resources to scheduling models with high fault prevention processing priority, or reserve resources for scheduling models with high fault prevention processing priority; at the same time, after one or more data tasks fail, the task scheduling fault processing device can determine the fault recovery processing priority corresponding to the scheduling model according to the UTRF classification model according to Table 4 above, and perform recovery processing on the scheduling model corresponding to one or more data tasks according to the fault recovery processing priority to ensure the priority recovery of the core model after the failure and improve the operating efficiency of the data task.

[0131] Furthermore, the task scheduling fault handling method further includes:

[0132] Step S430: According to the fault recovery processing priority, an automatic rerun classification mechanism is implemented for the scheduling model of the data task, wherein the automatic rerun classification mechanism is based on the number of automatic reruns and the automatic rerun time.

[0133] Specifically, in this embodiment, the task scheduling fault handling device can also refer to the fault prevention processing priority of the corresponding scheduling model for multiple data tasks, and implement an automatic re-running classification mechanism for the scheduling model corresponding to the data task. Specifically, the scheduling models with different fault prevention processing priorities are automatically re-run according to priority. For example, for the scheduling model with the highest fault prevention processing priority, the automatic re-run setting is preset to 10 re-run times or the automatic re-run time is set to recover immediately after the fault. The number of re-runs is reduced or the automatic re-run time is delayed for each level the fault prevention processing priority is reduced.

[0134] This embodiment adopts the above scheme, specifically by determining the usage heat index and fault heat index of the scheduling model of several data tasks according to the scheduling metadata of the data tasks; determining the usage heat category of the scheduling model according to the usage heat index; determining the fault heat category of the scheduling model according to the fault heat index; determining the fault handling priority of the data task according to the correspondence between the combination of the usage heat category and the fault heat category and the fault handling level, as well as the usage heat category and the fault heat category of the scheduling model of the data task; performing model optimization and resource tilt adjustment on the scheduling model of the data task according to the fault prevention processing priority; and recovering the scheduling model of the data task according to the fault recovery processing priority after a fault occurs in the data task.

[0135] This embodiment uses fault prevention processing priorities to sort the scheduling model for model optimization and tilt adjustment of resource reallocation; and adds a distinction in fault recovery processing priorities after a fault occurs, so that the scheduling model is recovered in sequence according to the fault recovery processing priorities, so that the core and important scheduling models can be effectively optimized in advance and handled after the fault occurs, thereby improving the operating efficiency of data tasks.

[0136] This application also provides a task scheduling fault handling device, please refer to Figure 5 , the task scheduling fault processing device includes:

[0137] A determination module 10 is configured to determine a usage heat and a failure heat of a scheduling model of a data task, wherein the scheduling model is configured to implement a data processing operation corresponding to the data task;

[0138] The fault processing module 20 is configured to perform fault processing on the data task according to the usage heat and fault heat of the scheduling model of the data task.

[0139] The task scheduling fault handling device provided in this application, which adopts the task scheduling fault handling method in the above-mentioned embodiment, can solve the technical problem of how to reduce the failure rate of big data tasks and optimize the post-processing speed of faulty tasks to improve the operating efficiency of big data tasks. Compared with the existing technology, the beneficial effects of the task scheduling fault handling device provided in this application are the same as the beneficial effects of the task scheduling fault handling method provided in the above-mentioned embodiment, and the other technical features of the task scheduling fault handling device are the same as the features disclosed in the above-mentioned embodiment method, which will not be repeated here.

[0140] The present application provides a task scheduling fault handling device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the task scheduling fault handling method in the above-mentioned embodiment one.

[0141] Reference below Figure 6, which shows a schematic diagram of the structure of a task scheduling fault handling device suitable for implementing an embodiment of the present application. The task scheduling fault handling device in the embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The task scheduling fault handling device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0142] like Figure 6 As shown, the task scheduling fault handling device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory 1002 or programs loaded from a storage device 1003 into a random access memory 1004. Random access memory 1004 also stores various programs and data required for the operation of the task scheduling fault handling device. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the task scheduling fault handling device to communicate wirelessly or wired with other devices to exchange data. Although the figure shows a task scheduling fault handling device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or provided instead.

[0143] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.

[0144] The task scheduling fault handling device provided in this application, which employs the task scheduling fault handling method of the above-mentioned embodiment, can solve the technical problem of how to reduce the failure rate of big data tasks and optimize the post-processing speed of faulty tasks to improve the operational efficiency of big data tasks. Compared with the prior art, the beneficial effects of the task scheduling fault handling device provided in this application are the same as those of the task scheduling fault handling method provided in the above-mentioned embodiment, and the other technical features of the task scheduling fault handling device are the same as those disclosed in the method of the previous embodiment, and are not further described here.

[0145] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0146] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0147] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer program) stored thereon, wherein the computer-readable program instructions are used to execute the task scheduling fault handling method in the above embodiment.

[0148] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0149] The computer-readable storage medium may be included in the task scheduling fault processing device; or may exist independently without being assembled into the task scheduling fault processing device.

[0150] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the task scheduling fault processing device, the task scheduling fault processing device: determines the usage heat and fault heat of the scheduling model of the data task, and the scheduling model is used to implement the data processing operation corresponding to the data task; and performs fault processing on the data task according to the usage heat and fault heat of the scheduling model of the data task.

[0151] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0152] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0153] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0154] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-mentioned task scheduling fault handling method. It can solve the technical problem of how to reduce the failure rate of big data tasks and optimize the post-processing speed of faulty tasks to improve the operating efficiency of big data tasks. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the task scheduling fault handling method provided in the above-mentioned embodiment, and will not be repeated here.

[0155] The present application also provides a computer program product, including a computer program, which implements the steps of the task scheduling fault handling method as described above when executed by a processor.

[0156] The computer program product provided in this application can solve the technical problem of reducing the failure rate of big data tasks and optimizing the post-processing speed of failed tasks, thereby improving the operational efficiency of big data tasks. Compared with the existing technology, the beneficial effects of the computer program product provided in this application are the same as those of the task scheduling fault handling method provided in the above embodiment, and will not be elaborated here.

[0157] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A task scheduling fault handling method, characterized in that: The task scheduling fault handling method includes: Determining a usage heat and a failure heat of a scheduling model for a data task, the scheduling model being used to implement a data processing operation corresponding to the data task; Fault processing is performed on the data task according to the usage heat and fault heat of the scheduling model of the data task.

2. The method according to claim 1, wherein The step of determining the usage heat and the failure heat of the scheduling model of the data task includes: Determining, based on scheduling metadata of a plurality of data tasks, a usage heat index and a fault heat index of a scheduling model of the data tasks; Determining a usage heat category of the scheduling model according to the usage heat index; Determining a fault heat category of the scheduling model according to the fault heat index; The step of performing fault processing on the data task according to the usage heat and fault heat of the scheduling model of the data task includes: Fault processing is performed on the data task according to the usage heat category and the fault heat category of the scheduling model of the data task.

3. The method according to claim 2, wherein The step of performing fault processing on the data task according to the usage heat category and the fault heat category of the scheduling model of the data task includes: Determining a fault handling priority of the data task according to a correspondence between a combination of a usage heat category and a fault heat category and a fault handling level, and the usage heat category and the fault heat category of a scheduling model of the data task; Fault processing is performed on the data task according to the fault processing priority.

4. The method according to claim 3, wherein The fault handling priority includes a fault prevention processing priority and a fault recovery processing priority. The step of performing fault handling on the data task according to the fault handling priority includes: According to the fault prevention processing priority, the scheduling model of the data task is optimized and resource tilted; After a failure occurs in the data task, the scheduling model of the data task is restored according to the failure recovery processing priority.

5. The method according to claim 4, wherein The task scheduling fault handling method further includes: According to the fault recovery processing priority, an automatic re-run classification mechanism is implemented for the scheduling model of the data task, and the automatic re-run classification mechanism is classified based on the number of automatic re-runs and the automatic re-run time.

6. The method according to claim 2, wherein The fault heat index includes the number of faults in the model cycle, the interval between the last fault of the model and the average execution time in the task cycle; and / or, the usage heat index includes the number of calls in the model cycle and the interval between the last call of the model.

7. The method according to claim 6, wherein After the step of determining the usage heat index and the fault heat index of the scheduling model of the data tasks based on the scheduling metadata of the plurality of data tasks, the method further includes: The number of failures in the model cycle is corrected according to the average execution time and scheduling operation information in the task cycle.

8. The method according to claim 2, wherein Before the step of determining the usage heat category of the scheduling model according to the usage heat index, and before the step of determining the fault heat category of the scheduling model according to the fault heat index, the method further includes: Collect usage heat index and failure heat index of historical scheduling models of several data tasks; Based on expert prior knowledge, a cluster analysis algorithm is used to analyze the usage heat index and the fault heat index of the historical scheduling model, and the classification boundary values of different types of scheduling models are determined to obtain model classification parameters; the model classification parameters are used to classify the usage heat category and / or the fault heat category to obtain the usage heat category and / or the fault heat category.

9. A task scheduling fault handling device, characterized in that: The task scheduling fault processing device includes: a determination module, configured to determine a usage heat and a failure heat of a scheduling model of a data task, wherein the scheduling model is used to implement a data processing operation corresponding to the data task; The fault processing module is used to perform fault processing on the data task according to the usage heat and fault heat of the scheduling model of the data task.

10. A task scheduling fault handling device, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the task scheduling fault processing method according to any one of claims 1 to 8.

11. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the task scheduling fault processing method according to any one of claims 1 to 8 are implemented.

12. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the task scheduling fault processing method according to any one of claims 1 to 8 are implemented.