Scheduling system task fault diagnosis alarm system, method and equipment based on log

By designing a log-based scheduling system task fault diagnosis and alarm system in the task scheduling system of the event-driven architecture, the problem of fault diagnosis and alarm caused by the dispersion and asynchronousness of the log data is solved, rapid diagnosis and efficient alarm are achieved, and the reliability and stability of the system are improved.

CN120216299APending Publication Date: 2025-06-27CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510256403.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the task scheduling system of the event-driven architecture, the dispersion and asynchronousness of log data make it difficult to diagnose and alert problems, and traditional log management systems are difficult to efficiently manage and analyze these log data distributed across multiple nodes.

Method used

A log-based scheduling system task fault diagnosis and alarm system is designed, including task manager, task executor, Logstash, Elasticsearch and Kafka. Through a unified log format and distributed storage mechanism, the full collection, storage and real-time analysis of logs are realized, and Kafka is listened to Kafka to obtain exception logs and alarm processing in real time.

Benefits of technology

It improves the efficiency of collecting and analyzing log data, realizes rapid diagnosis and alarm of task scheduling system failures, and improves the reliability and stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216299A_ABST
    Figure CN120216299A_ABST
Patent Text Reader

Abstract

The invention discloses a scheduling system task fault diagnosis and alarm system, method and device based on logs, and relates to the technical field of log management.The system comprises a task manager, a task executor, Logstash and Elasticsearch, the task manager is used for achieving creation, distribution, management and alarm of tasks, and the task executor is used for executing the tasks; and performing log generation according to a life cycle from creation of the task to completion of distribution, and the task executor is used for executing a service corresponding to the task, checking a log and generating the log according to a life cycle from the beginning of execution to the end of execution of the task. According to the invention, the reliability and stability of the task scheduling system can be improved, and the fault of the task scheduling system can be quickly diagnosed and alarmed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of log management, and particularly to a task fault diagnosis and warning system, method, and device for a log-based scheduling system. Background Art

[0002] In a task scheduling system based on the event-driven architecture style, the task manager, as the publisher of events, is responsible for publishing task scheduling events to the event channel; the executor, as the subscriber of events, subscribes to the task events it is interested in and executes the corresponding tasks when the events occur. The event channel, as the medium for transmitting events between the publisher and the subscriber, can be a message queue, an event bus, or other forms of communication channels. This architecture is particularly suitable for distributed and microservices environments and can effectively handle large-scale task scheduling and complex business scenarios. However, with the increase in the complexity of the task scheduling system, log-based fault diagnosis and early warning have become increasingly important and challenging. The main problems are as follows:

[0003] (1) The dispersion and inconsistency of log data: In an event-driven architecture, the task manager and the executor are usually distributed on different nodes, which results in the corresponding dispersion of log data on multiple nodes; in addition, due to the asynchronous processing characteristics of events, the logs of the same task may be scattered in the log files of multiple time points and components, increasing the difficulty of collecting log data and the cost of maintaining consistency;

[0004] (2) The breakage and difficulty in tracing of the event chain: In an event-driven task scheduling system, the processing of a task may involve the interaction of multiple events and components. Due to the asynchrony and uncertainty of events, the event chain may break due to network latency, component failures, or data processing errors, making it difficult to trace and reconstruct the entire process of task processing, increasing the complexity and uncertainty of fault diagnosis;

[0005] (3) The huge amount of log data: The amount of log data in an event-driven architecture is huge and is distributed in different event channels and components. How to efficiently manage and store this log data has become an urgent problem to be solved. Traditional log management systems usually store log data centrally on a centralized log server, but in an event-driven architecture, due to the loose coupling relationship between components, the distribution and transmission of log data become complex and it is difficult to achieve efficient centralized management;

[0006] (4) Difficulty in fault analysis and warning: In an event-driven architecture, the analysis of log data is crucial for fault diagnosis and warning. However, due to the asynchrony and distribution of log data, traditional log analysis methods are difficult to directly apply to an event-driven architecture. How to perform real-time analysis and processing of log data distributed in different event channels and components to quickly discover and locate faults is a major challenge currently faced.

[0007] In summary, the existing event-driven architecture has many defects and deficiencies in log collection, management, and analysis, seriously affecting the fault diagnosis and warning capabilities of the task scheduling system. Therefore, how to effectively implement fault diagnosis and warning of the task scheduling system to improve the reliability and stability of the task scheduling system has become an urgent problem to be solved currently. Summary of the Invention

[0008] This application provides a log-based task fault diagnosis and warning system, method, and device for a scheduling system, which can improve the reliability and stability of the task scheduling system and achieve rapid diagnosis and warning of faults in the task scheduling system.

[0009] In a first aspect, an embodiment of this application provides a log-based task fault diagnosis and warning system for a scheduling system. The log-based task fault diagnosis and warning system for a scheduling system includes:

[0010] A task manager, which is used to create, distribute, manage, and warn tasks, and generate logs according to the life cycle of the task from creation to distribution end;

[0011] A task executor, which is used to execute the business corresponding to the task, view logs, and generate logs according to the life cycle of the task from start of execution to end of execution;

[0012] Logstash, which is used to collect the logs generated by the task manager and task executor, obtain the full amount of logs and write them into the index of Elasticsearch, and write the abnormal logs into Kafka;

[0013] Elasticsearch, which is used to store and query the full amount of logs.

[0014] Combined with the first aspect, in an implementation manner,

[0015] The task manager includes a task creation module, a task distribution module, a task management module, a task warning module, and a unified log generator;

[0016] The task creation module is used to enter task information to create a task. The task information includes the name of the task, timing parameters, input and output, person in charge, and contact information of the person in charge;

[0017] The task distribution module is used to generate instances of each cycle of the task and perform load balancing, and distribute the instances of the task to the specified task executor, and the task executor performs an execution operation on the instances of the task;

[0018] The task alarm module is used to obtain the tasks corresponding to the exception logs from Kafka, and query the tasks corresponding to the specified logs from Elasticsearch for alarm processing;

[0019] The unified log generator of the task manager is used to generate logs for the life cycle of the task from creation to distribution end based on a specific log format, and write the generated logs into a log file.

[0020] Combined with the first aspect, in an implementation manner,

[0021] The task executor includes a task execution module, an instance Console, and a unified log generator;

[0022] The task execution module is used to execute the service corresponding to the task;

[0023] The instance Console is used to view the logs generated by the task manager in the form of a command window, and supports filtering the logs according to time and log format;

[0024] The unified log generator of the task executor is used to generate logs for the life cycle of the task from the start of execution to the end of execution based on a specific log format, and write the generated logs into a log file.

[0025] Combined with the first aspect, in an implementation manner, Logstash is specifically used for: collecting the specified logs in the logs generated by the task manager and the task executor, obtaining the full amount of logs and writing them into the index of Elasticsearch, filtering out the exception logs in the full amount of logs, and writing the task information, instance information, and time information of the tasks corresponding to the exception logs into Kafka.

[0026] In a second aspect, an embodiment of the present application provides a method for task fault diagnosis and alarm of a log-based scheduling system, which is carried out based on the above-mentioned scheduling system task fault diagnosis and alarm system. The method for task fault diagnosis and alarm of the log-based scheduling system includes:

[0027] Based on the unified log generators in the task manager and the task executor, generate logs and write the exception logs into Kafka;

[0028] Use the time as the index name, and store the generated logs in slices according to the date dimension to achieve distributed storage and query of the logs;

[0029] Monitor Kafka in real time to obtain the tasks corresponding to the exception logs, and obtain the contact information of the person in charge of the tasks for alarm and log analysis processing.

[0030] Combined with the second aspect, in one implementation, the writing the exception logs into Kafka specifically includes:

[0031] Filter the logs at the WARN and ERROR levels in the generated logs to obtain the exception logs;

[0032] Write the exception logs into Kafka.

[0033] Combined with the second aspect, in one implementation,

[0034] When storing the generated logs in slices according to the date dimension, it further includes: performing compression processing on the logs based on writing performance, log archiving, and log data expiration management;

[0035] For the writing performance, by increasing the threshold size of disk flushing and the data synchronization time interval, and disabling the relevant doc values and source fields to save disk space and improve the writing performance;

[0036] For log archiving, it is carried out through two dimensions of time granularity and access frequency. Among them, for the time granularity, it is to archive the index data for a certain period of time for the archiving operation, and for the access frequency, it is to archive the data that has not been accessed for the set time;

[0037] For log data expiration management, it is to clean up and delete the expired data.

[0038] Combined with the second aspect, in one implementation, the monitoring Kafka in real time to obtain the tasks corresponding to the exception logs, and obtaining the contact information of the person in charge of the tasks for alarm and log analysis processing, wherein, for the alarm, it specifically includes:

[0039] Monitor Kafka in real time to obtain the exception logs in Kafka, and parse to obtain the task ID, instance ID, and time of the task corresponding to the exception logs;

[0040] Confirm the es index where the log record is located through time, query the full - volume logs to generate a txt file through the instance ID, and query the contact information of the person in charge of the task through the task ID;

[0041] Perform alarm processing on the person in charge based on the contact information of the person in charge, and attach the exception log file and the link analysis data of the exception log file.

[0042] In combination with the second aspect, in one implementation, the Kafka is monitored in real time to obtain the tasks corresponding to the abnormal logs, and the contact information of the person in charge corresponding to the tasks is obtained for alarm and log analysis processing. For the log analysis, it specifically includes:

[0043] For the obtained abnormal logs, the key data included therein is obtained. The key data includes the instance ID of the instance of the task in a certain cycle, the event ID of each stage of the instance life cycle, the upstream event ID of the events of each stage of the instance life cycle, and the start time.

[0044] Based on the key data, the call link steps and link elapsed time of the task instance on the task executor node are parsed.

[0045] In the third aspect, an embodiment of the present application provides a task fault diagnosis and alarm device for a log-based scheduling system. The task fault diagnosis and alarm device for the log-based scheduling system includes a processor, a memory, and a log-based scheduling system task fault diagnosis and alarm program stored on the memory and executable by the processor. When the log-based scheduling system task fault diagnosis and alarm program is executed by the processor, the steps of the above-mentioned log-based scheduling system task fault diagnosis and alarm method are implemented.

[0046] The beneficial effects brought by the technical solution provided by the embodiment of the present application include:

[0047] (1) The collection efficiency and accuracy of log data are improved: By writing log data into a log file in a unified log format, the loss and omission of log data are avoided.

[0048] (2) The analysis efficiency and accuracy of log data are improved: Through custom rules, the rapid positioning of anomalies in log data is realized, improving the analysis efficiency and accuracy.

[0049] (3) The rapid diagnosis and alarm of faults in the task scheduling system are realized: Through the message queue, the abnormal information of the task is monitored in real time and pushed to relevant operation and maintenance personnel or responsible departments, improving the reliability and stability of the task scheduling system. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a schematic structural diagram of the task fault diagnosis and alarm system for the log-based scheduling system of the present application;

[0051] Figure 2 It is a schematic flowchart of the task fault diagnosis and alarm method for the log-based scheduling system of the present application;

[0052] Figure 3 It is a schematic hardware structure diagram of the task fault diagnosis and alarm device for the log-based scheduling system of the present application. DETAILED DESCRIPTION

[0053] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0054] First, some technical terms in the present application are explained to facilitate those skilled in the art to understand the present application.

[0055] Task scheduling system: A software system responsible for coordinating and managing the execution of various business tasks. It is used to receive task requests, assign tasks to appropriate executors for execution, and monitor the execution status of tasks to ensure that tasks can be completed smoothly according to the predetermined schedule and priority.

[0056] Task Manager: Responsible for receiving task requests, generating task events, and publishing these events to the event channel.

[0057] Task executor: subscribes to task events in the event channel and executes corresponding tasks based on the event content.

[0058] Event-driven architecture: A mainstream design that achieves loosely coupled communication between components in a task scheduling system through the generation, publication, subscription, and processing of events.

[0059] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.

[0060] On the first aspect, the embodiment of the present application provides a log-based scheduling system task fault diagnosis and alarm method, which improves the reliability and stability of the task scheduling system by improving the log collection, management and analysis mechanism, and realizes rapid diagnosis and alarm of task scheduling system faults.

[0061] In one embodiment, referring to Figure 1 , Figure 1 This is a schematic diagram of the structure of the log-based scheduling system task fault diagnosis and alarm system of this application. Figure 1As shown in the figure, the task fault diagnosis and warning system of the log-based scheduling system includes: a task manager, a task executor, Logstash, Elasticsearch, and Kafka. Logstash is an open-source data collection engine with real-time pipeline capabilities and is commonly used for log management and data analysis. Elasticsearch is a distributed, real-time search and analysis engine. Kafka is a distributed streaming processing platform.

[0062] The task manager is used to create, distribute, manage, and alarm tasks, and generate logs according to the life cycle of the task from creation to distribution; the task executor is used to execute the corresponding business of the task, view logs, and generate logs according to the life cycle of the task from the start of execution to the end of execution; Logstash is used to collect the logs generated by the task manager and the task executor, obtain the full amount of logs and write them into the index of Elasticsearch, and write the abnormal logs into Kafka; Elasticsearch is used to store and query the full amount of logs, and the index is partitioned and stored at the granularity of days.

[0063] Specifically, the task manager includes a task creation module, a task distribution module, a task management module, a task alarm module, and a unified log generator.

[0064] For the task creation module, the task creation module is used to enter task information to create a task. The task information includes the name of the task, timing parameters, input and output, person in charge, and contact information of the person in charge; the contact information of the person in charge includes information such as the contact phone number and email of the person in charge.

[0065] For the task distribution module, the task distribution module is used to generate and perform load balancing for each instance of the task, and distribute the instance of the task to the specified task executor, and the task executor performs an execution operation on the instance of the task.

[0066] For the task alarm module, the task alarm module is used to obtain the task corresponding to the abnormal log from Kafka, and query the task corresponding to the specified log from Elasticsearch for alarm processing; specifically, the task alarm module is responsible for consuming the information of the task corresponding to the abnormal log from the abnormal log topic of Kafka, and querying the detailed information of the task corresponding to the specified log from Elasticsearch, and performing an alarm by email or instant messaging.

[0067] For the unified log generator of the task manager, the unified log generator of the task manager is used to generate logs for the life cycle of a task from creation to distribution end based on a specific log format, and write the generated logs into a log file. Specifically, it is responsible for generating logs for the entire life cycle of a task from creation to distribution end according to the specific log format, and writing them into a log file. The specific log format is: [Time YYYY-MM-DD HH:MM:SS]-[Log level DEBUG / INFO / WARN / ERROR]-[TraceId task instance ID]-[ParentSpanId parent event Id]-[SpanID event ID]-[Log content message].

[0068] Specifically, the task executor includes a task execution module, an instance Console, and a unified log generator.

[0069] For the task execution module, the task execution module is used to execute the business corresponding to the task; that is, the task execution module is responsible for executing the actual business of the task, such as Sql, Shell, Spark, Mr, etc.

[0070] For the instance Console, the instance Console is used to view the logs generated by the task manager in the form of a command window, and supports filtering the logs according to time and log format.

[0071] For the unified log generator of the task executor, the unified log generator of the task executor is used to generate logs for the life cycle of a task from start of execution to end of execution based on a specific log format, and write the generated logs into a log file. Specifically, it is responsible for generating logs for the task from start of execution to end of execution according to the specific log format, and writing them into a log file. The specific log format is: [Time YYYY-MM-DD HH:MM:SS]-[Log level DEBUG / INFO / WARN / ERROR]-[TraceId task instance ID]-[ParentSpanId parent event Id]-[SpanID event ID]-[Log content message].

[0072] For Logstash, Logstash is specifically used to: collect the specified logs in the logs generated by the task manager and the task executor to obtain the full amount of logs and write them into the index of Elasticsearch, and filter out the abnormal logs in the full amount of logs (that is, filter out the logs of WARN (warning) and ERROR (error) levels), and write the task information, instance information, and time information of the tasks corresponding to the abnormal logs into Kafka, specifically into the abnormal log topic of Kafka.

[0073] Second aspect, the embodiments of the present application further provide a method for fault diagnosis and alarm of scheduling system tasks based on logs, which is carried out based on the above-mentioned scheduling system task fault diagnosis and alarm system.

[0074] In one embodiment, referring to Figure 2 , Figure 2 , which is a schematic flowchart of the method for fault diagnosis and alarm of scheduling system tasks based on logs in the present application. As Figure 2 shown, the method for fault diagnosis and alarm of scheduling system tasks based on logs includes:

[0075] S1: Based on the unified log generators in the task manager and the task executor, generate logs and write the abnormal logs into Kafka;

[0076] S2: Use time as the index name and store the generated logs in slices according to the date dimension to achieve distributed storage and query of logs;

[0077] S3: Listen to Kafka in real time to obtain the tasks corresponding to the abnormal logs, and obtain the contact information of the persons in charge corresponding to the tasks for alarm and log analysis processing.

[0078] Further, in one embodiment, writing the abnormal logs into Kafka specifically includes:

[0079] S101: Filter out the logs at the WARN and ERROR levels in the generated logs to obtain abnormal logs;

[0080] S102: Write the abnormal logs into Kafka.

[0081] Specifically, for the method for fault diagnosis and alarm of scheduling system tasks based on logs in the present application, log generation and collection are first performed. To solve the problems of scattered and asynchronous log data, the present application designs a unified log format, collection and storage mechanism. This mechanism integrates the corresponding log generation and collection modules in the task manager and the executor respectively, and at the same time filters out the logs at the WARN and ERROR levels and writes them into the abnormal log topic of Kafka for the task alarm module to listen and consume.

[0082] In the present application, for the Topic message format, it can be as follows:

[0083] {

[0084] "time":"2024:10:10 10:10:10",

[0085] "taskId":"fhgkhlkjfhdgfhgj975rtyutedcvbho",

[0086] "instanceId":"fhgkhlkjutyihjghgjhkjtedcvbho"

[0087] }。

[0088] To achieve Logstash writing logs to Elasticsearch while writing the data filtered by rules to Kafka, two output plugins need to be defined: one for Elasticsearch and the other for Kafka. The following is an example of configuring Logstash:

[0089]

[0090]

[0091]

[0092] Furthermore, in one embodiment, when storing the generated logs in slices according to the date dimension, it further includes: performing compression processing on the logs based on writing performance, log archiving, and log data invalidation management; for writing performance, by increasing the threshold size of disk flushing and the data synchronization time interval, and disabling relevant doc values and source fields to save disk space and improve writing performance; for log archiving, it is carried out through two dimensions of time granularity and access frequency. Among them, for time granularity, it is to archive the index data for a certain period of time for the archiving operation, and for access frequency, it is to archive the data that has not been accessed for a set time; for log data invalidation management, it is to clean up and delete the expired data.

[0093] Specifically, for the method of task fault diagnosis and warning of the log-based scheduling system of the present application, after log generation and collection, log management is carried out. In order to cope with the problem of the sharp increase in log data volume, using tasklogs-YYYY.MM.dd as the index name, the log data is stored in slices according to the date dimension, and the distributed storage technology is used to achieve the distributed storage and efficient query of the log data.

[0094] In addition, in order to reduce the storage cost of log data and improve query efficiency, compression processing can also be performed on the log data. When facing the pressure of a large amount of data writing, the three main aspects affecting log writing performance and storage management are writing performance, log archiving, and log data invalidation management.

[0095] For write performance, a trade - off can be made between data integrity and availability. If a certain probability of data loss can be tolerated, or if the data written unsuccessfully can be repaired through other compensation mechanisms, then measures such as increasing the disk flushing threshold size (flush_threshold_size) and the data synchronization time interval (sync_interval) can be taken. It is also possible to disable unnecessary doc values and _source fields to save disk space and improve write performance.

[0096] For log archiving, it can be carried out according to two dimensions: time granularity and access frequency. For time granularity, the index data of operations for a certain period (such as one week or half a month) can be archived (this strategy is most in line with the business of this patent). For access frequency, the data that has not been accessed for a long time (such as 15 days) can be indexed and archived.

[0097] For log data invalidation management, this application mainly involves cleaning and deleting expired data (such as data from three months ago), because normal historical task logs do not have practical value for problem location and analysis.

[0098] Further, in one embodiment, Kafka is listened to in real - time to obtain the tasks corresponding to the exception logs, and the contact information of the person in charge of the tasks is obtained for alarm and log analysis processing. Among them, for alarms, it specifically includes:

[0099] S301: Listen to Kafka in real - time to obtain the exception logs in Kafka, and parse to obtain the task ID, instance ID, and time of the task corresponding to the exception logs;

[0100] S302: Confirm the es index where the log record is located through time, query the full - volume logs to generate a txt file through the instance ID, and query the contact information of the person in charge of the task through the task ID;

[0101] S303: Carry out alarm processing to the person in charge based on the contact information of the person in charge, and attach the exception log file and the link analysis data of the exception log file.

[0102] Specifically, for the log-based scheduling system task fault diagnosis and alarm method of the present application, after log management, fault diagnosis and alarm are carried out. The task alarm module will listen to the abnormal log information of Kafka in real time, and parse out the task ID, instance ID, and time of the corresponding task. Confirm the es index where the log record is located through the time, query the full log to generate a txt file through the instance ID, and query the contact information of the person in charge of the task through the task ID. If it is an email alarm, an alarm email will be generated, attached with the abnormal log file and the link analysis data of the abnormal log file, and finally the alarm email will be sent. If it is an instant messaging method alarm, an alarm message containing task information, instance information, and abnormal events will be sent, and the detailed information of the log can be viewed through the Console of the instance. Among them, if a task instance runs across days, if you want to view the full log of the instance, you need to determine multiple es indexes through the start time to the end time of the instance for query. The person in charge of the task or the system operation and maintenance personnel can locate the specific abnormal point according to the log data and the call link information, and then stop or rerun the task.

[0103] Further, in one embodiment, Kafka is listened to in real time to obtain the task corresponding to the abnormal log, and the contact information of the person in charge of the task is obtained, and alarm and log analysis processing are carried out. Among them, for log analysis, it specifically includes:

[0104] S311: For the obtained abnormal log, obtain the key data it contains. The key data includes the instance ID of the instance of the task in a certain period, the event ID of each stage of the instance life cycle, the upstream event ID of the event of each stage of the instance life cycle, and the start time;

[0105] S312: Based on the key data, parse out the call link steps and link consumption time of the task instance at the task executor node.

[0106] Specifically, for the log-based scheduling system task fault diagnosis and alarm method of the present application, after fault diagnosis and alarm, log analysis is carried out. For the abnormal log file queried in the previous step, it contains four key data: the instance ID (TraceId) of the instance of the task in a certain period, the event ID (spanId) of each stage of the instance life cycle, the upstream event ID (ParentSpanID) of the event of each stage of the instance life cycle, and the start time. Through these four data, the call link steps of a certain task instance at each task executor node (the same TraceId indicates belonging to the same business process, and the link relationship between SpanId and ParentSpanId can connect the call link) and the specific link consumption time (the event difference between each step is the actual consumption of each processing link) can be parsed out.

[0107] In a third aspect, an embodiment of the present application provides a log-based scheduling system task fault diagnosis and alarm device. The log-based scheduling system task fault diagnosis and alarm device can be a device with data processing functions such as a personal computer (PC), a laptop, a server, etc.

[0108] Referring to Figure 3 , Figure 3 is a schematic hardware structure diagram of the log-based scheduling system task fault diagnosis and alarm device involved in the solution of the embodiment of the present application. In the embodiment of the present application, the log-based scheduling system task fault diagnosis and alarm device may include a processor, a memory, a communication interface, and a communication bus.

[0109] Among them, the communication bus can be of any type and is used to interconnect the processor, the memory, and the communication interface.

[0110] The communication interface includes interfaces such as input / output (I / O) interfaces, physical interfaces, and logical interfaces for implementing the interconnection of components inside the log-based scheduling system task fault diagnosis and alarm device, as well as interfaces for implementing the interconnection between the log-based scheduling system task fault diagnosis and alarm device and other devices (such as other computing devices or user devices). The physical interface can be an Ethernet interface, an optical fiber interface, an ATM interface, etc.; the user device can be a display, a keyboard, etc.

[0111] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical memory, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.

[0112] The processor may be a general - purpose processor, which can call the log - based scheduling system task fault diagnosis and alarm program stored in the memory and execute the log - based scheduling system task fault diagnosis and alarm method provided by the embodiments of the present application. For example, the general - purpose processor may be a central processing unit (CPU). Among them, the method executed when the log - based scheduling system task fault diagnosis and alarm program is called can refer to the various embodiments of the log - based scheduling system task fault diagnosis and alarm method of the present application, which will not be elaborated here.

[0113] Those skilled in the art can understand that Figure 3 the hardware structure shown in does not constitute a limitation to the present application, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0114] The terms "including" and "having" and any variations thereof in the specification, claims and drawings of the present application are intended to cover non - exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices. The descriptions such as "first", "second" and "third" are used to distinguish different objects, etc., and do not represent a sequence, nor do they limit that "first", "second" and "third" are of different types.

[0115] In the description of the embodiments of the present application, terms such as "exemplary", "for example" or "for instance" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary", "for example" or "for instance" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of terms such as "exemplary", "for example" or "for instance" is intended to present relevant concepts in a specific manner.

[0116] In the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B; the "and / or" in the text is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.

[0117] In some processes described in the embodiments of the present application, there are multiple operations or steps that appear in a specific order. However, it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of the present application or may be executed in parallel. The serial numbers of the operations are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed in sequence or in parallel, and these operations or steps may be combined.

[0118] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above and includes several instructions for causing a terminal device to execute the methods described in the various embodiments of the present application.

[0119] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A log-based scheduling system task fault diagnosis and alarm system, characterized in that: The log-based scheduling system task fault diagnosis and alarm system includes: The task manager is used to create, distribute, manage and warn tasks, and to generate logs according to the life cycle of tasks from creation to distribution. The task executor is used to execute the business corresponding to the task, view the log, and generate logs according to the life cycle from the start to the end of the task; Logstash, which is used to collect the logs generated by the task manager and task executor, obtain the full logs and write them into the Elasticsearch index, and write the exception logs into Kafka; Elasticsearch is used to store and query all logs.

2. A log-based scheduling system task fault diagnosis and alarm system as claimed in claim 1, characterized in that: The task manager includes a task creation module, a task distribution module, a task management module, a task alarm module, and a unified log generator; The task creation module is used to input task information to create a task. The task information includes the task name, timing parameters, input and output, person in charge, and contact information of the person in charge. The task distribution module is used to generate and load balance instances of each task cycle, and distribute the task instances to designated task executors, which execute the task instances. The task alarm module is used to obtain the task corresponding to the abnormal log from Kafka, and query the task corresponding to the specified log from Elasticsearch to perform alarm processing; The unified log generator of the task manager is used to realize log generation of the life cycle of a task from creation to distribution based on a specific log format, and write the generated log into a log file.

3. A log-based scheduling system task fault diagnosis and alarm system as claimed in claim 1, characterized in that: The task executor includes a task execution module, an instance console, and a unified log generator; The task execution module is used to execute the business corresponding to the task; The Console instance is used to view the logs generated by the task manager in the form of a command window, and supports filtering the logs according to time and log format; The unified log generator of the task executor is used to realize log generation of the life cycle of the task from the start of execution to the end of execution based on a specific log format, and write the generated log into a log file.

4. A log-based scheduling system task fault diagnosis and alarm system as claimed in claim 1, characterized in that: The Logstash is specifically used to collect the specified logs in the logs generated by the task manager and the task executor, obtain the full logs and write them into the index of Elasticsearch, filter out the abnormal logs in the full logs, and write the task information, instance information, and time information of the task corresponding to the abnormal log into Kafka.

5. A log-based scheduling system task fault diagnosis and alarm method, based on the scheduling system task fault diagnosis and alarm system according to any one of claims 1 to 4, characterized in that: The log-based scheduling system task fault diagnosis and alarm method includes: Based on the unified log generator in the task manager and task executor, log generation is realized and exception logs are written to Kafka. Use time as the index name and shard the generated logs according to the date dimension to achieve distributed storage and query of logs. Monitor Kafka in real time to obtain the tasks corresponding to the abnormal logs, and obtain the contact information of the person in charge of the tasks, and perform alarm and log analysis.

6. A method for diagnosing and warning task faults in a scheduling system based on logs as claimed in claim 5, characterized in that: The abnormal log is written into Kafka, specifically including: Filter the generated logs to get the WARN and ERROR level logs and get the exception logs; Write exception logs to Kafka.

7. A log-based scheduling system task fault diagnosis and alarm method as claimed in claim 5, characterized in that: When the generated logs are stored in shards according to the date dimension, the following steps are also included: based on write performance, log archiving, and log data expiration management, the logs are compressed; For write performance, we increase the threshold size for flushing to disk and the data synchronization interval, and disable related docvalues ​​and source fields to save disk space and improve write performance. Log archiving is done in two dimensions: time granularity and access frequency. Time granularity refers to archiving index data for a certain period of time, while access frequency refers to indexing and archiving data that has not been accessed within a set period of time. For log data expiration management, expired data is cleaned and deleted.

8. A method for diagnosing and warning task faults in a scheduling system based on logs as claimed in claim 5, characterized in that: The real-time monitoring of Kafka is to obtain the tasks corresponding to the abnormal logs, and obtain the contact information of the person in charge corresponding to the tasks, and perform alarm and log analysis processing, wherein the alarm specifically includes: Monitor Kafka in real time to obtain the exception logs in Kafka, and parse to obtain the task ID, instance ID, and time of the task corresponding to the exception log; Confirm the es index where the log record is located by time, query the full log to generate a txt file by instance ID, and query the contact information of the person in charge of the task by task ID; Based on the contact information of the person in charge, an alarm is issued to the person in charge, and the abnormal log file and the link analysis data of the abnormal log file are attached.

9. A log-based scheduling system task fault diagnosis and alarm method as claimed in claim 8, characterized in that: The real-time monitoring of Kafka is to obtain the tasks corresponding to the abnormal logs, and obtain the contact information of the person in charge corresponding to the tasks, and perform alarm and log analysis processing, wherein the log analysis specifically includes: For the acquired exception log, the key data contained therein is obtained, wherein the key data includes the instance ID of the instance of the task in a certain period, the event ID of each stage of the instance life cycle, the upstream event ID of the event at each stage of the instance life cycle, and the start time; Based on the key data, the call link steps and link duration of the task instance in the task executor node are analyzed and obtained.

10. A log-based scheduling system task fault diagnosis and alarm device, characterized in that: The log-based scheduling system task fault diagnosis and alarm device includes a processor, a memory, and a log-based scheduling system task fault diagnosis and alarm program stored in the memory and executable by the processor, wherein when the log-based scheduling system task fault diagnosis and alarm program is executed by the processor, the steps of the log-based scheduling system task fault diagnosis and alarm method as described in any one of claims 5 to 9 are implemented.