Task real-time evaluation and self-healing system and method based on big data platform
Through real-time evaluation and self-healing systems on the big data platform, the problem of task failures not being handled in a timely manner was solved, and efficient and stable data processing and operation and maintenance automation were achieved.
Patent Information
- Application Number
- CN202510829208.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-09
AI Technical Summary
The existing big data task platform is unable to evaluate task failures in real time and perform self-recovery, resulting in a lack of guarantee of work efficiency and stability.
A real-time task evaluation and self-healing system based on a big data platform is provided, which includes a content acquisition module, a fault assessment module and a fault self-healing module. By extracting task log information and measurement indicators in real time, the evaluation engine is used to judge faults and generate fault assessment reports, and self-healing processing is performed according to the reports.
It realizes the real-time fault perception and self-healing capabilities of the big data task platform, improves the timeliness and stability of data processing, and reduces operation and maintenance costs and manual intervention.
Smart Images

Figure CN120610833A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data service technology, and in particular to a real-time task evaluation and self-healing system, method, electronic device, computer storage medium, and computer program product based on a big data platform. Background Art
[0002] With the rapid development of technology, automated operations and maintenance have become an integral part of modern equipment and software. For example, by automatically sending diagnostic or evaluation data, equipment or software can send real-time information such as operating status and error reports to developers or manufacturers, enabling timely analysis and resolution of problems. Furthermore, with the rapid development of big data technology, the frequency and difficulty of handling task operation anomalies in ultra-large-scale cluster environments and complex data pipelines are increasing exponentially. For example, according to industry research data, in production environments with extremely large daily data processing volumes, data processing delays caused by task failures can reach millions of dollars per hour.
[0003] In today's complex big data environment, traditional manual assessment and diagnosis methods are inefficient and unable to meet the high business requirements for data processing timeliness and stability. This also leads to widespread problems for enterprise big data platforms, such as high operation and maintenance costs and low SLA (Service-Level Agreement) compliance rates. Therefore, there is an urgent need for a real-time assessment and self-healing system and method based on big data platforms that can assess and self-heal big data task failures in real time, thereby improving the timeliness and stability of platform data processing. Summary of the Invention
[0004] The main purpose of the present invention is to solve the technical problem that the big data task platform in the existing technology is unable to evaluate task failures in real time and perform self-recovery of failures, resulting in a lack of guarantee of efficiency and stability during work.
[0005] The first aspect of the present invention provides a real-time task evaluation and self-healing system based on a big data platform, comprising: The content acquisition module is used to extract the log information and measurement indicators of the task in real time through the big data task submission portal, and send the log information and measurement indicators to different message queues respectively; A fault assessment module is configured to call an assessment engine to extract the log information and the metric indicators from different message queues, determine whether a fault exists in the current task based on the log information and the metric indicators according to preset assessment rules, and if so, generate a fault assessment report for the current task; The fault self-healing module is used to search for a fault task whose fault type is a self-healable fault based on the content in the fault assessment report, and perform self-healing processing on the fault task when a task self-healing condition is met.
[0006] Optionally, in a first implementation of the first aspect of the present invention, the message queue includes a first message queue for storing the log information and a second message queue for storing the measurement indicator; The fault assessment module includes: An information reading unit, configured to call an evaluation engine to read the log information in the first message queue, read the metric indicators in the second message queue, and parse the read log information and metric indicators respectively; A fault analysis unit is configured to extract an evaluation rule from a preset evaluation rule library, and determine whether a current task has a fault based on the parsed log information and the measurement indicators according to the evaluation rule; The result generating unit is used to generate a fault assessment report based on the faulty task when a fault occurs in the current task, and store the report in the assessment result database.
[0007] Optionally, in a second implementation of the first aspect of the present invention, the fault assessment report includes at least a task fault type and a self-healing method corresponding to the faulty task; The fault self-healing module specifically includes: a method acquisition unit, configured to capture and parse the fault assessment report in the assessment result database, obtain a fault task whose fault type is a self-healable fault, and find a self-healing method corresponding to the fault task in the fault assessment report; The self-healing processing unit is used to determine whether the current system state and the faulty task meet all self-healing execution conditions, wherein the number of the self-healing execution conditions is at least one; if all the self-healing execution conditions are met, the self-healing method is called to perform self-healing processing on the faulty task.
[0008] Optionally, in a third implementation of the first aspect of the present invention, the self-healing processing unit is further configured to: Obtaining the fault task captured by the acquisition unit of the method; Determine whether the current task is in the self-healing time period. If so, continue the task. Otherwise, do not perform self-healing on the faulty task. Determine whether this is the first self-healing retry. If so, continue. Otherwise, do not self-heal the faulty task. Determine whether the faulty task is a memory overflow fault. If so, continue the task. Otherwise, do not perform self-healing on the faulty task. Determine whether the current faulty task is an SLA task. If so, modify the SQL node memory parameters of the faulty task and perform self-healing steps. Otherwise, continue. Determine whether the current cluster resources are idle. If so, modify the SQL node memory parameters of the faulty task and perform self-healing steps. If not, continue; Determine whether the current fault task has reached the maximum self-healing resource. If so, do not perform self-healing on the fault task. If not, modify the SQL node memory parameters of the fault task and execute the self-healing step.
[0009] Optionally, in a fourth implementation of the first aspect of the present invention, a rule maintenance module is further included for maintaining the evaluation rules used in the fault evaluation module; The rule maintenance module specifically includes: A code organization unit, configured to call a cache storage evaluation rule code snippet of an evaluation engine in the fault evaluation module, and insert the evaluation rule code snippet into the evaluation rule library; A rule updating unit is used to periodically call the evaluation engine to query whether there are new evaluation rule code snippets in the evaluation rule library of the fault evaluation module; if there are new evaluation rule code snippets, dynamic compilation is performed based on the new evaluation rule code snippets, and the evaluation rules used in the fault analysis unit are updated.
[0010] Optionally, in a fifth implementation of the first aspect of the present invention, the evaluation rule library includes an evaluation rule table and an evaluation constant table; The code organization unit is further specifically used for: Extracting evaluation rule information from the evaluation rule table, extracting an evaluation threshold corresponding to the evaluation rule information from the evaluation constant table, and generating an evaluation code snippet based on the evaluation rule information and the evaluation threshold.
[0011] A second aspect of the present invention provides a real-time task evaluation and self-healing method based on a big data platform, comprising: Obtain task log information and metrics in real time through the big data task submission portal, and send the log information and metrics to different message queues respectively; Calling the evaluation engine to extract the log information and the metric indicators from different message queues respectively, and judging whether there is a fault in the current task based on the log information and the metric indicators according to preset evaluation rules, and if so, generating a fault evaluation report for the current task; Based on the content in the fault assessment report, a fault task whose fault type is a self-healable fault is searched, and a self-healing process is performed on the fault task when a task self-healing condition is met.
[0012] Optionally, in a first implementation of the second aspect of the present invention, the message queue includes a first message queue for storing the log information and a second message queue for storing the measurement indicator; The call evaluation engine extracts the log information and the metric indicators from different message queues respectively, and determines whether there is a fault in the current task based on the log information and the metric indicators according to preset evaluation rules. If so, a fault evaluation report for the current task is generated, including: Calling an evaluation engine to read the log information in the first message queue, read the metric indicators in the second message queue, and parse the read log information and metric indicators respectively; Extracting evaluation rules from a preset evaluation rule library, and judging whether there is a fault in the current task based on the parsed log information and the measurement indicators according to the evaluation rules; When a fault occurs in the current task, a fault assessment report is generated based on the faulty task and stored in the assessment result database.
[0013] Optionally, in a second implementation of the second aspect of the present invention, the fault assessment report includes at least a task fault type and a self-healing method corresponding to the faulty task; The searching, based on the content in the fault assessment report, for a fault task whose fault type is a self-healable fault, and performing self-healing processing on the fault task when a task self-healing condition is met, includes: Capture and parse the fault assessment report in the assessment result database to obtain the fault task whose fault type is a self-healing fault, and find the self-healing method corresponding to the fault task in the fault assessment report; Determine whether the current system state and the fault task meet all self-healing execution conditions, where the number of the self-healing execution conditions is at least one; If all the self-healing execution conditions are met, the self-healing method is called to perform self-healing processing on the faulty task.
[0014] Optionally, in a third implementation of the second aspect of the present invention, determining whether the current system state and the faulty task meet all self-healing execution conditions, where the number of the self-healing execution conditions is at least one; and if all the self-healing execution conditions are met, calling the self-healing method to perform self-healing processing on the faulty task specifically includes: Obtaining the fault task captured by the acquisition unit of the method; Determine whether the current task is in the self-healing time period. If so, continue the task. Otherwise, do not perform self-healing on the faulty task. Determine whether this is the first self-healing retry. If so, continue. Otherwise, do not self-heal the faulty task. Determine whether the faulty task is a memory overflow fault. If so, continue the task. Otherwise, do not perform self-healing on the faulty task. Determine whether the current faulty task is an SLA task. If so, modify the SQL node memory parameters of the faulty task and perform self-healing steps. Otherwise, continue. Determine whether the current cluster resources are idle. If so, modify the SQL node memory parameters of the faulty task and perform self-healing steps. If not, continue; Determine whether the current fault task has reached the maximum self-healing resource. If so, do not perform self-healing on the fault task. If not, modify the SQL node memory parameters of the fault task and execute the self-healing step.
[0015] Optionally, in a fourth implementation of the second aspect of the present invention, before determining whether the current task has a fault based on the log information and the metric according to a preset evaluation rule, the method further includes: Calling the cache storage evaluation rule code snippet of the evaluation engine in the fault evaluation module, and inserting the evaluation rule code snippet into the evaluation rule library; The evaluation engine is periodically called to query whether there are new evaluation rule code snippets in the evaluation rule library of the fault evaluation module; if there are new evaluation rule code snippets, dynamic compilation is performed based on the new evaluation rule code snippets, and the evaluation rules used in the fault analysis unit are updated.
[0016] Optionally, in a fifth implementation of the second aspect of the present invention, the evaluation rule library includes an evaluation rule table and an evaluation constant table; Before calling the cache storage evaluation rule code snippet of the evaluation engine in the fault evaluation module and inserting the evaluation rule code snippet into the evaluation rule library, the method further includes: Extracting evaluation rule information from the evaluation rule table, extracting an evaluation threshold corresponding to the evaluation rule information from the evaluation constant table, and generating an evaluation code snippet based on the evaluation rule information and the evaluation threshold.
[0017] The third aspect of the present invention provides a real-time task evaluation and self-healing device based on a big data platform, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory to enable the real-time task evaluation and self-healing device based on the big data platform to perform the steps of the above-mentioned real-time task evaluation and self-healing method based on the big data platform.
[0018] A fourth aspect of the present invention provides a computer-readable storage medium, which stores instructions that, when executed on a computer, enable the computer to execute the steps of the above-mentioned real-time task evaluation and self-healing method based on a big data platform.
[0019] A fifth aspect of the present invention provides a computer program product, comprising a computer program / instruction, characterized in that when the computer program / instruction is executed by a processor, the steps of the above-mentioned real-time task evaluation and self-healing method based on the big data platform are implemented.
[0020] In the technical solution provided by the present invention, the real-time task evaluation and self-healing system based on the big data platform includes a content acquisition module, which is used to extract the log information and measurement indicators of the task in real time through the big data task submission portal, and send the log information and measurement indicators to different message queues respectively; a fault assessment module, which is used to call the evaluation engine to extract log information and measurement indicators from different message queues respectively, and judge whether the current task has a fault based on the log information and measurement indicators according to the preset evaluation rules. If so, a fault assessment report for the current task is generated; a fault self-healing module, which is used to find fault tasks whose fault type is a self-healable fault based on the content in the fault assessment report, and perform self-healing processing on the fault task when the task self-healing conditions are met. The technical solution provided by the present invention provides a task implementation evaluation and self-healing system based on the big data platform, which can perceive and evaluate the faults of big data tasks in real time and perform fault self-healing, which can improve the timeliness and stability of data processing of the big data business platform. In addition, the method, electronic device, computer-readable storage medium and computer program product provided by the present invention also solve the corresponding technical problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 This is a flowchart of a first embodiment of a method for real-time task evaluation and self-healing based on a big data platform according to an embodiment of the present invention; Figure 2This is a flow chart of a second embodiment of a method for real-time task evaluation and self-recovery based on a big data platform according to an embodiment of the present invention; Figure 3 1. This is a flow chart of the evaluation part in the second embodiment of the method for real-time task evaluation and self-recovery based on a big data platform according to an embodiment of the present invention; Figure 4 1. It is a flowchart of the self-healing part in the second embodiment of the method for real-time task evaluation and self-healing based on a big data platform in an embodiment of the present invention; Figure 5 A schematic diagram of an embodiment of a real-time task evaluation and self-healing system based on a big data platform in an embodiment of the present invention; Figure 6 A schematic diagram of an embodiment of a real-time task evaluation and self-healing device based on a big data platform in an embodiment of the present invention; Figure 7 The figure is a schematic diagram of the principle of a computer-readable medium in an embodiment of the present invention. DETAILED DESCRIPTION
[0022] Exemplary embodiments of the present invention will now be described more fully with reference to the accompanying drawings. However, exemplary embodiments can be implemented in various forms, and it should not be understood that the present invention is limited to the embodiments set forth herein. On the contrary, providing these exemplary embodiments enables the present invention to be more comprehensive and complete, making it easier to fully convey the inventive concept to those skilled in the art. In the figures, the same reference numerals represent the same or similar elements, components or parts, and thus their repeated description will be omitted.
[0023] Under the premise of being consistent with the technical concept of the present invention, the features, structures, characteristics or other details described in a specific embodiment do not exclude that they can be combined in one or more other embodiments in a suitable manner.
[0024] In the description of specific embodiments, the features, structures, characteristics, or other details of the present invention are described to enable those skilled in the art to fully understand the embodiments. However, this does not preclude those skilled in the art from practicing the technical solutions of the present invention without one or more of the specific features, structures, characteristics, or other details.
[0025] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.
[0026] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0027] The term "and / or" or "and / or" includes all combinations of any one or more of the associated listed items.
[0028] See also Figure 1 The first embodiment of the real-time task evaluation and self-healing method based on the big data platform in the embodiment of the present invention includes: S101, extracting the log information and metric indicators of the task in real time through the big data task submission portal, and sending the log information and metric indicators to different message queues respectively; It is understood that the execution subject of the present invention can be a real-time task evaluation and self-healing system based on a big data platform, or a terminal or server, and the specific implementation is not limited here. The embodiment of the present invention is described by taking the server as the execution subject as an example.
[0029] In this embodiment, during the working process of the big data task platform, the server can automatically monitor the working status of the task in real time. Specifically, the server supports multiple task submission entrances, including workflow, self-service query, task entrances from the client or web page, etc.; through these task entrances, the task log information and measurement indicators can be obtained, and the log information and the measurement indicators can be sent and saved in different message queues respectively.
[0030] In a specific embodiment, the log information may be a log file (log) during the execution of a big data platform task, which is mainly used to record various events and situations that occur when the program is running; the metric may be a metric (metric) in a Spark task, wherein the Spark task is a task generated in the open source big data processing Spark engine.
[0031] S102: Calling the evaluation engine to extract log information and metrics from different message queues; The evaluation engine is called to read the log information and metrics collected and saved in the previous step from different message queues.
[0032] S103. Determine whether the current task has a fault based on the log information and the metric according to the preset evaluation rules; S104. If yes, generate a fault assessment report for the current task; After obtaining the log information and metrics, the evaluation engine is called to determine whether a fault occurs during the processing of the big data task recorded in the log information based on pre-built evaluation rules, and a fault assessment report is generated based on the judgment result. The report is stored in the evaluation result database for subsequent fault self-recovery modules or manual queries.
[0033] Among them, the fault assessment report described in this embodiment includes at least the fault judgment result and the corresponding self-healing method; further, the fault judgment result specifically includes the fault type, fault description, target log information of the fault hit, target measurement indicators of the fault hit and fault conclusion.
[0034] After obtaining the fault assessment report, the contents contained in the fault assessment report can be stored in the assessment result database in the form of a data table. In addition, in this step, the pre-built assessment rules can be pre-set fault finding rules, including fault finding methods and fault determination threshold data.
[0035] S105 . Based on the content in the fault assessment report, search for faulty tasks whose fault type is a self-healable fault, and perform self-healing processing on the faulty tasks when the task self-healing conditions are met.
[0036] Based on the fault judgment results contained in the fault assessment report obtained in the above steps, determine whether the currently prone task faults include self-healing faults; if they do, determine whether the current system state and specific fault type meet the task self-healing conditions. There must be at least one task self-healing condition, such as whether the system time is within the self-healing time period and whether the maximum number of self-healing attempts has been reached. If all task self-healing conditions are met, the corresponding self-healing method in the fault assessment report for the faulty task requiring self-healing can be extracted, and the self-healing method can be called to perform self-healing processing on the faulty task.
[0037] The real-time task evaluation and self-healing method based on the big data platform provided by the embodiment of the present invention can perceive and evaluate the faults of big data tasks in real time and make intelligent decisions, realize automatic self-healing of tasks with faults, improve the timeliness and stability of data processing on the big data business platform, and improve the degree of automation of the operation and maintenance of the big data platform; meet the demand for real-time output of technical fault evaluation results.
[0038] Please see Figure 2-4 The second embodiment of the real-time task evaluation and self-recovery method based on the big data platform in the embodiment of the present invention includes: S201, extracting the log information and metric indicators of the task in real time through the big data task submission portal, and sending the log information and metric indicators to different message queues respectively; In this embodiment, the server supports multiple big data task submission entries, such as workflow, self-service query, and submission from Linux client or Jupyter, etc., where the Linux client refers to an application or tool running on the Linux operating system, such as a big data task processing platform built on Linux; the Jupyter refers to Jupyter Notebook, which is an application that can be used in data cleaning and conversion, numerical simulation, statistical modeling, machine learning and other tasks.
[0039] In a specific embodiment, the big data platform described in this embodiment is a management platform built based on the Apache DolphinScheduler system, which can be used to orchestrate and schedule big data tasks. The log information can include scheduled task information and self-service query logs, which can be obtained based on the big data platform system log files. The metrics are Spark metrics, which are obtained by collecting system data, such as the number of Spark shuffles and GC (garbage collection) time.
[0040] Please see the attached Figure 3 In this embodiment, the message queue includes a first message queue and a second message queue, both of which are Kafka queues. The first message queue and the second message queue are constructed to process log data and indicator data respectively. The first message queue is used to record the occurrence and processing status of real-time events through stored log information, and the second message queue is used to transmit performance indicators that require real-time monitoring through stored metrics. The two queues are relatively independent. After extracting log information, it is saved in the first message queue, and the metrics are sent to the second message queue for storage.
[0041] S202: Calling the evaluation engine to read log information from the first message queue, read metrics from the second message queue, and parse the read log information and metrics respectively; In this example, the Flink streaming framework is combined with the Janino dynamic rule engine to build the evaluation engine, enabling rapid loading and real-time validation of evaluation rules. Janino is a Java compiler that can compile Java source code files into bytecode files and can also compile Java expressions, blocks, classes, and source code files in memory.
[0042] For details, please refer to the attached Figure 3, determining whether the evaluation engine consumes data from the Kafka queue. If not, skip the process; if so, continue. The evaluation engine consumes data from the Kafka queue based on the Kafka connector built into the Flink streaming framework. This includes reading log information from the first message queue, reading metrics from the second message queue, and parsing the log information and metrics.
[0043] S203: Extracting evaluation rules from a preset evaluation rule library, and determining whether the current task has a fault based on the parsed log information and measurement indicators according to the evaluation rules; S204: When a fault occurs in the current task, a fault assessment report is generated based on the faulty task and stored in an assessment result database; The evaluation engine is invoked to extract evaluation rules from the evaluation rule library, analyze the parsed log information and metrics, and determine whether the current task has a fault. If a fault is present, a fault assessment report is generated and stored in the evaluation result database for subsequent query by the fault self-recovery module or manually. The fault assessment report includes at least the fault determination result and the corresponding self-recovery method; further, the fault determination result specifically includes the fault type, fault description, target log information hit by the fault, target metrics hit by the fault, and the fault conclusion.
[0044] In one specific implementation, the evaluation engine polls the evaluation rule repository at specified intervals to check for new rule code snippets. If so, the rule is added or replaced. This is accomplished by first using a cache to store code snippets within the evaluation engine's Flink code. Secondly, the new evaluation rule code snippets are inserted into the evaluation rule repository. Flink dynamically compiles all code snippets in the cache using Janino. Finally, the evaluation process is executed based on the compiled code to generate the evaluation results.
[0045] The evaluation rule library includes an evaluation type table, an evaluation rule table, and a constant table. The evaluation type table is used to store fault types and fault descriptions, such as "syntax error" or "data skew." The evaluation rule table is used to store evaluation rules (i.e., the Java code portion of the rules in the evaluation engine). The constant table stores thresholds required in the evaluation rules, such as the skew multiple threshold for determining data skew, the expansion multiple threshold for determining data expansion, and the average byte count threshold for determining and dividing file sizes. Considering that evaluation rules are variable, commonly used thresholds are placed in the constant table. This eliminates the need to modify the Java code of the rules in the evaluation engine; simply adjusting the thresholds adjusts the evaluation logic, reducing the number of operations required to change the evaluation logic.
[0046] The evaluation results database can also be a MySQL database. Based on the characteristics of MySQL databases, workflows and autonomous query systems can search, access, and even display information in the evaluation results database. The evaluation results database also includes an optimization suggestion table and an evaluation conclusion table. The optimization suggestion table is used to store general optimization rules for faults; the evaluation conclusion table is used to store conclusions obtained by running evaluation rules. This refers to the part where the evaluation engine analyzes metrics or log information when running Java code to determine the fault type, generate conclusions, and associate evaluation suggestions.
[0047] In one specific implementation, the system also includes using the evaluation engine's cache to store evaluation rule code snippets and inserting them into the evaluation rule library. The system then periodically calls the rule engine to query the evaluation rule library for new evaluation rule code snippets. If new evaluation rule code snippets are available, the Flink engine is called to dynamically compile the new evaluation rule code snippets using Janino and update the evaluation rules. This allows for adding or modifying evaluation rules without pausing the server, allowing for real-time adjustments to the specific evaluation content.
[0048] S205: Capture and parse the fault assessment report in the assessment result database to obtain the fault task whose fault type is a self-healable fault, and find the self-healing method corresponding to the fault task in the fault assessment report; The solution described in this embodiment can automatically capture the evaluation results of tasks with faults in the evaluation result database, and obtain faulty tasks whose fault type is a self-healing fault, parse the fault evaluation report in the evaluation result database, and find the self-healing method corresponding to the faulty task.
[0049] S206: Determine whether the current system status and fault task meet all self-recovery execution conditions; After obtaining the self-healable fault task and the corresponding self-healing method, the self-healable fault task is marked or stored in a queue to be self-healed. A determination is then made as to whether the current system state and the fault task meet all self-healing execution conditions. The self-healing execution condition can be at least one, such as whether the system time is within the self-healing time period and whether the maximum number of self-healing attempts has been reached.
[0050] S207: If all self-healing execution conditions are met, the self-healing method is called to perform self-healing processing on the faulty task.
[0051] For example, when the self-healing execution condition is "needs to be in the preset self-healing time period", it is determined whether the current time is in the self-healing time period; if the current time is in the self-healing time period, the fault self-healing method in the evaluation result is obtained to perform self-healing processing on the self-healable fault.
[0052] When there are multiple self-healing execution conditions, determine whether the current system state and the faulty task meet all the self-healing execution conditions. If all conditions are met, call the self-healing method to perform self-healing processing on the faulty task.
[0053] See also Figure 4 In a specific implementation, when performing specific fault self-recovery steps for a Spark task's memory overflow exception, the steps may be performed by calling a worker unit in the Apache Dolphin Scheduler. The specific steps include: Step 401: Call the worker unit to fetch the task currently evaluated as having a fault, and then execute step 402; Step 402: Determine whether the current time is the self-healing time period. If it is, proceed to step 403. If it is not, proceed to step 410, i.e., do not perform self-healing on the captured faulty task. Step 403: Determine whether the current self-healing operation is the first self-healing retry; if it is the first self-healing retry, proceed to step 404; if it is not the first retry, proceed to step 410, i.e., do not perform self-healing on the captured faulty task; Step 404: Determine whether the current fault is a memory overflow fault; if it is a memory overflow fault, continue to step 405; if it is not a memory overflow fault, execute step 410, that is, do not perform self-healing on the captured faulty task; Step 405: Determine whether the current faulty task is an SLA task; if the current faulty task is an SLA task, execute step 408 to modify the SQL node memory parameters of the faulty task and execute step 409, i.e., the self-healing step; if the current faulty task is not an SLA task, continue to execute step 406; Step 406: Determine whether the current cluster resources are idle. If so, execute step 408 to modify the SQL node memory parameters of the faulty task and execute step 409, i.e., the self-healing step; if not, continue to execute step 407; Step 407: Determine whether the current faulty task has reached the maximum self-healing resource. If so, execute step 410, i.e., do not perform self-healing on the faulty task. If not, execute step 408 to modify the SQL node memory parameters of the faulty task and execute step 409, i.e., the self-healing step.
[0054] In a specific implementation, modifying the SQL node memory parameters may be achieved by increasing the Executor memory configuration.
[0055] As for the detailed self-healing process, the Spark task is submitted through the big data platform. In a specific embodiment, the big data platform can be a scheduling platform; and based on the aforementioned content steps, the evaluation engine performs a real-time evaluation of the Spark running task and stores the evaluation results in the MySQL database. When the big data platform retries a failed Spark task, it will be based on whether it is a self-healing time period (for example, the early morning time period), whether it is the first retry, whether there is a memory overflow assessment report, and whether it is an SLA task. If all these conditions are met, when the big data platform submits the Spark task again, it will try to increase the Executor (executor) memory configuration to modify the memory parameters, and resubmit the task to try to resume the task operation so that the Spark task reduces the probability of memory overflow errors and reduces the alarms of the big data platform.
[0056] The real-time task evaluation and self-healing method based on the big data platform provided by the embodiment of the present invention can perceive task failures in real time and execute intelligent decisions, and has self-healing capabilities. It can automatically self-heal tasks with failures, and can efficiently realize automated operation and maintenance; this method can perform several attempts at self-healing operations in time periods such as the early morning of each day when manual duty is more difficult or there are fewer tasks, reducing the number of alarms of the big data platform in the early morning or other special time periods, improving the self-service query anomaly assessment diagnosis coverage and self-healing rate, further reducing the workload of problem investigation and labor costs, and improving the smoothness and satisfaction of user use.
[0057] The above describes the real-time task evaluation and self-healing method based on the big data platform in the embodiment of the present invention. The following describes the real-time task evaluation and self-healing system based on the big data platform in the embodiment of the present invention. Figure 5 In one embodiment of the present invention, a real-time task evaluation and self-healing system based on a big data platform includes: The content acquisition module 501 is used to extract the log information and measurement indicators of the task in real time through the big data task submission portal, and send the log information and measurement indicators to different message queues respectively; The fault assessment module 502 is configured to call an assessment engine to extract the log information and the metric indicators from different message queues, determine whether a fault exists in the current task based on the log information and the metric indicators according to preset assessment rules, and if so, generate a fault assessment report for the current task; The fault self-healing module 503 is configured to search for a fault task whose fault type is a self-healable fault based on the content in the fault assessment report, and perform self-healing processing on the fault task when a task self-healing condition is met.
[0058] The real-time task evaluation and self-healing system based on the big data platform provided by the embodiment of the present invention can perceive and evaluate the failure of big data tasks in real time and make intelligent decisions, realize automatic self-healing of tasks with failures, improve the timeliness and stability of data processing on the big data business platform, and improve the degree of automation of big data platform operation and maintenance.
[0059] In another embodiment of the present application, the message queue includes a first message queue for storing the log information and a second message queue for storing the metric; The fault assessment module 502 includes: An information reading unit, configured to call an evaluation engine to read the log information in the first message queue, read the metric indicators in the second message queue, and parse the read log information and metric indicators respectively; A fault analysis unit is configured to extract an evaluation rule from a preset evaluation rule library, and determine whether a current task has a fault based on the parsed log information and the measurement indicators according to the evaluation rule; The result generating unit is used to generate a fault assessment report based on the faulty task when a fault occurs in the current task, and store the report in the assessment result database.
[0060] In another embodiment of the present application, the fault assessment report includes at least a task fault type and a self-healing method corresponding to the faulty task; The fault self-recovery module 503 specifically includes: a method acquisition unit, configured to capture and parse the fault assessment report in the assessment result database, obtain a fault task whose fault type is a self-healable fault, and find a self-healing method corresponding to the fault task in the fault assessment report; The self-healing processing unit is used to determine whether the current system state and the faulty task meet all self-healing execution conditions, wherein the number of the self-healing execution conditions is at least one; if all the self-healing execution conditions are met, the self-healing method is called to perform self-healing processing on the faulty task.
[0061] In another embodiment of the present application, the self-healing processing unit is further configured to: Obtaining the fault task captured by the acquisition unit of the method; Determine whether the current task is in the self-healing time period. If so, continue the task. Otherwise, do not perform self-healing on the faulty task. Determine whether this is the first self-healing retry. If so, continue. Otherwise, do not self-heal the faulty task. Determine whether the faulty task is a memory overflow fault. If so, continue the task. Otherwise, do not perform self-healing on the faulty task. Determine whether the current faulty task is an SLA task. If so, modify the SQL node memory parameters of the faulty task and perform self-healing steps. Otherwise, continue. Determine whether the current cluster resources are idle. If so, modify the SQL node memory parameters of the faulty task and perform self-healing steps. If not, continue; Determine whether the current fault task has reached the maximum self-healing resource. If so, do not perform self-healing on the fault task. If not, modify the SQL node memory parameters of the fault task and execute the self-healing step.
[0062] In another embodiment of the present application, the real-time task evaluation and self-healing system based on the big data platform further includes a rule maintenance module for maintaining the evaluation rules used in the fault evaluation module; The rule maintenance module specifically includes: A code organization unit, configured to call a cache storage evaluation rule code snippet of an evaluation engine in the fault evaluation module, and insert the evaluation rule code snippet into the evaluation rule library; A rule updating unit is used to periodically call the evaluation engine to query whether there are new evaluation rule code snippets in the evaluation rule library of the fault evaluation module; if there are new evaluation rule code snippets, dynamic compilation is performed based on the new evaluation rule code snippets, and the evaluation rules used in the fault analysis unit are updated.
[0063] In another embodiment of the present application, the evaluation rule library includes an evaluation rule table and an evaluation constant table; The code organization unit is further specifically used for: Before calling the cache storage evaluation rule code snippet of the evaluation engine in the fault evaluation module, the evaluation rule information is extracted from the evaluation rule table, the evaluation threshold corresponding to the evaluation rule information is extracted from the evaluation constant table, and the evaluation code snippet is generated based on the evaluation rule information and the evaluation threshold.
[0064] In addition, the specific implementation scheme of the real-time task evaluation and self-healing system based on the big data platform described in the embodiments of the present application can be found in the contents of the aforementioned method embodiments, and will not be described in detail.
[0065] The real-time task evaluation and self-healing system based on the big data platform described in the embodiment of the present invention can perceive and evaluate the failure of big data tasks in real time and make intelligent decisions, realize automatic self-healing of tasks with failures, improve the timeliness and stability of data processing on the big data business platform, and improve the degree of automation of the operation and maintenance of the big data platform; further reduce the number of alarms on the big data platform, and improve the coverage rate of self-service query anomaly evaluation and diagnosis and the self-healing rate of system failures.
[0066] Based on the same inventive concept, an embodiment of this specification also provides an electronic device for real-time task evaluation and self-healing based on a big data platform. The following is a detailed description of the electronic device for real-time task evaluation and self-healing based on a big data platform in an embodiment of the present invention from the perspective of hardware processing.
[0067] Figure 6 This is a schematic diagram of the structure of an electronic device provided in the embodiment of this specification. Figure 6 The electronic device 600 according to this embodiment of the present invention will be described. Figure 6 The electronic device 600 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present invention.
[0068] like Figure 6 As shown, electronic device 600 is implemented as a general-purpose computing device. Components of electronic device 600 may include, but are not limited to, at least one processing unit 610, at least one storage unit 620, a bus 630 connecting various system components (including storage unit 620 and processing unit 610), and a display unit 640.
[0069] The storage unit stores program codes that can be executed by the processing unit 610, so that the processing unit 610 performs the steps according to various exemplary embodiments of the present invention described in the above processing method section of this specification. For example, the processing unit 610 can perform the following steps: Figure 1 or Figure 2 Steps shown.
[0070] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 6201 and / or a cache memory unit 6202 , and may further include a read-only memory unit (ROM) 6203 .
[0071] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205, such program modules 6205 including but not limited to: an operating system, one or more application programs, other program modules and program data, each of which or some combination may include an implementation of a network environment.
[0072] Bus 630 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0073] The electronic device 600 may also communicate with one or more external devices 100 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that enable a user to interact with the electronic device 600, and / or any device that enables the electronic device 600 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed through an input / output (I / O) interface 650. Furthermore, the electronic device 600 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 660. The network adapter 660 may communicate with other modules of the electronic device 600 through the bus 630. It should be understood that although Figure 6 Not shown, other hardware and / or software modules may be used in conjunction with electronic device 600, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0074] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the exemplary embodiments described in the present invention can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiment of the present invention can be embodied in the form of a software product, which can be stored in a computer-readable storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.) or on a network, and includes a number of instructions to enable a computing device (which can be a personal computer, server, or network device, etc.) to execute the above method according to the present invention. When the computer program is executed by a data processing device, the computer-readable medium is enabled to implement the above method of the present invention, that is: Figure 1 or Figure 2 The method shown.
[0075] In addition, the present invention also provides a computer program product and a computer-readable medium, wherein the computer program product includes a computer program / instructions, which, when executed by a processor, implements the real-time task evaluation and self-healing method based on a big data platform as described in any of the above embodiments.
[0076] in, Figure 7 A schematic diagram of a computer-readable medium provided in accordance with an embodiment of this specification.
[0077] accomplish Figure 1 or Figure 2 The computer program of the illustrated method can be stored on one or more computer-readable media. The computer-readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0078] The computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, wherein the readable program code is carried. The data signal propagated may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or component. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.
[0079] Program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0080] In summary, the present invention can be implemented in hardware, or as a software module running on one or more processors, or a combination thereof. Those skilled in the art will appreciate that, in practice, general-purpose data processing devices such as microprocessors or digital signal processors (DSPs) can be used to implement some or all of the functions of some or all of the components according to the embodiments of the present invention. The present invention can also be implemented as a device or apparatus program (e.g., a computer program or computer program product) for performing part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium or in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.
[0081] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the present invention is not inherently related to any specific computer, virtual device, or electronic device, and various general-purpose devices can also implement the present invention. The above description is only a specific embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.
[0082] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
[0083] If the technical solution of this application involves personal information, the product that applies the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing personal information. If the technical solution of this application involves sensitive personal information, the product that applies the technical solution of this application has obtained the individual's separate consent before processing sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, a clear and prominent sign is set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that they agree to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload their personal information; among which, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
[0084] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A real-time task evaluation and self-healing system based on a big data platform, characterized by: include: The content acquisition module is used to extract the log information and measurement indicators of the task in real time through the big data task submission portal, and send the log information and measurement indicators to different message queues respectively; A fault assessment module is configured to call an assessment engine to extract the log information and the metric indicators from different message queues, determine whether a fault exists in the current task based on the log information and the metric indicators according to preset assessment rules, and if so, generate a fault assessment report for the current task; The fault self-healing module is used to search for a fault task whose fault type is a self-healable fault based on the content in the fault assessment report, and perform self-healing processing on the fault task when a task self-healing condition is met.
2. The real-time task evaluation and self-healing system based on a big data platform according to claim 1 is characterized in that: The message queue includes a first message queue for storing the log information and a second message queue for storing the metric; The fault assessment module includes: An information reading unit, configured to call an evaluation engine to read the log information in the first message queue, read the metric indicators in the second message queue, and parse the read log information and metric indicators respectively; A fault analysis unit is configured to extract an evaluation rule from a preset evaluation rule library, and determine whether a current task has a fault based on the parsed log information and the measurement indicators according to the evaluation rule; The result generating unit is used to generate a fault assessment report based on the faulty task when a fault occurs in the current task, and store the report in the assessment result database.
3. The real-time task evaluation and self-healing system based on a big data platform according to claim 2 is characterized in that: The fault assessment report includes at least the task fault type and the self-recovery method corresponding to the fault task; The fault self-recovery module specifically includes: a method acquisition unit, configured to capture and parse the fault assessment report in the assessment result database, obtain a fault task whose fault type is a self-healable fault, and find a self-healing method corresponding to the fault task in the fault assessment report; The self-healing processing unit is used to determine whether the current system state and the faulty task meet all self-healing execution conditions, wherein the number of the self-healing execution conditions is at least one; if all the self-healing execution conditions are met, the self-healing method is called to perform self-healing processing on the faulty task.
4. The real-time task evaluation and self-healing system based on a big data platform according to claim 3 is characterized in that: The self-healing processing unit is further configured to: Obtaining the fault task captured by the acquisition unit of the method; Determine whether the current task is in the self-healing time period. If so, continue the task. Otherwise, do not perform self-healing on the faulty task. Determine whether this is the first self-healing retry. If so, continue. Otherwise, do not self-heal the faulty task. Determine whether the faulty task is a memory overflow fault. If so, continue the task. Otherwise, do not perform self-healing on the faulty task. Determine whether the current faulty task is an SLA task. If so, modify the SQL node memory parameters of the faulty task and perform self-healing steps. Otherwise, continue. Determine whether the current cluster resources are idle. If so, modify the SQL node memory parameters of the faulty task and perform self-healing steps. If not, continue; Determine whether the current fault task has reached the maximum self-healing resource. If so, do not perform self-healing on the fault task. If not, modify the SQL node memory parameters of the fault task and execute the self-healing step.
5. The real-time task evaluation and self-healing system based on a big data platform according to any one of claims 2 to 4, characterized in that: Also included is a rule maintenance module for maintaining the evaluation rules used in the fault evaluation module; The rule maintenance module specifically includes: A code organization unit, configured to call a cache storage evaluation rule code snippet of an evaluation engine in the fault evaluation module, and insert the evaluation rule code snippet into the evaluation rule library; A rule updating unit is used to periodically call the evaluation engine to query whether there are new evaluation rule code snippets in the evaluation rule library of the fault evaluation module; if there are new evaluation rule code snippets, dynamic compilation is performed based on the new evaluation rule code snippets, and the evaluation rules used in the fault analysis unit are updated.
6. The real-time task evaluation and self-healing system based on a big data platform according to claim 5 is characterized in that: The evaluation rule library includes an evaluation rule table and an evaluation constant table; The code organization unit is further specifically used for: Before calling the cache storage evaluation rule code snippet of the evaluation engine in the fault evaluation module, the evaluation rule information is extracted from the evaluation rule table, the evaluation threshold corresponding to the evaluation rule information is extracted from the evaluation constant table, and the evaluation code snippet is generated based on the evaluation rule information and the evaluation threshold.
7. A real-time task evaluation and self-healing method based on a big data platform, characterized in that: include: Obtain task log information and metrics in real time through the big data task submission portal, and send the log information and metrics to different message queues respectively; Calling the evaluation engine to extract the log information and the metric indicators from different message queues respectively, and judging whether there is a fault in the current task based on the log information and the metric indicators according to preset evaluation rules, and if so, generating a fault evaluation report for the current task; Based on the content in the fault assessment report, a fault task whose fault type is a self-healable fault is searched, and a self-healing process is performed on the fault task when a task self-healing condition is met.
8. A real-time task evaluation and self-healing device based on a big data platform, characterized in that: The real-time task evaluation and self-healing device based on the big data platform includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor calls the instructions in the memory to enable the big data platform-based task real-time evaluation and self-healing device to perform the steps of the big data platform-based task real-time evaluation and self-healing method as described in claim 7.
9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the program / instructions are executed by the processor, the steps of the real-time task evaluation and self-healing method based on the big data platform as described in claim 7 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by the processor, the steps of the real-time task evaluation and self-healing method based on the big data platform as claimed in claim 7 are implemented.