Error pattern diagnosis method, device and medium for data flow pipeline
Patent Information
- Application Number
- CN202611050509.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-15
- Publication Date
- 2026-10-09
AI Technical Summary
[0004]为了克服上述缺陷,提出了本申请,以提供解决或至少部分地解决现有技术中数据流水线错误诊断的效率较低且成本较高的技术问题
[0051]The error mode diagnosis method for data pipelines in this application includes: obtaining the log address of the failed task; obtaining the log content based on the log address; inputting the log content into a preset error mode library for pattern matching to obtain the matched error mode and the corresponding matching confidence level; if the matching confidence level is greater than a preset threshold, generating a first diagnostic result based on the matched error mode; otherwise, reading log-related data based on the log content and generating a second diagnostic result based on the log-related data. By obtaining the log addresses of failed tasks and retrieving their contents, the diagnostic process is automatically triggered, eliminating the need for manual log location. The log content is input into a pre-defined error pattern library for pattern matching, yielding the matched error patterns and their confidence levels, enabling rapid identification of known errors. When the matching confidence level is greater than a pre-defined threshold, a first diagnostic result is generated directly based on the matched error pattern, achieving rapid response and significantly improving diagnostic efficiency in high-confidence scenarios. Conversely, when the matching confidence level is less than or equal to the pre-defined threshold, related data is read from the log content, and a second diagnostic result is generated based on this data. This ensures that even in low-confidence or unmatched scenarios, accurate root cause analysis and treatment suggestions can still be obtained through in-depth diagnostics, thus forming a tiered diagnostic mechanism that balances efficiency and accuracy.
Smart Images

Figure CN122884731A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and specifically provides a method, device, and medium for error mode diagnosis of data pipelines. Background Technology
[0002] Existing data pipelines, task scheduling platforms, and log monitoring systems typically record task failure status, failure points, runtime parameters, log files, and alarm information. Some existing log analysis solutions can parse logs, perform keyword matching, or detect anomalies, and push anomalies to operations and maintenance personnel.
[0003] However, in actual production processes, troubleshooting failed tasks often involves multiple contextual information such as task configuration, data batches, code versions, and resource status. Most existing systems only provide scattered failure statuses and log entry points, still requiring maintenance personnel to perform a significant amount of manual work, such as downloading logs, searching keywords, consulting relevant personnel, and accumulating handling experience in personal notes or chat logs. This approach has the following main drawbacks: First, diagnostic methods relying on fixed rules or keyword matching can only cover known high-frequency errors. When dealing with changes in log representation, combined faults, or unknown long-tail errors, it is prone to missed or incorrect diagnoses, resulting in insufficient diagnostic accuracy. Second, the scattered information makes joint attribution difficult; manual troubleshooting requires repeatedly switching between multiple scattered information sources such as logs, task nodes, code versions, and responsible persons, leading to low efficiency. Finally, the conclusions of manual troubleshooting are difficult to form reusable and iterative structured knowledge, resulting in repetitive work, and valuable experience cannot be effectively accumulated and shared. Summary of the Invention
[0004] To overcome the aforementioned shortcomings, this application is proposed to provide a solution, or at least a partial solution, to the technical problems of low efficiency and high cost in data pipeline error diagnosis in the prior art. This application provides a method, apparatus, and medium for error mode diagnosis of data pipelines.
[0005] In a first aspect, this application provides an error mode diagnosis method for data pipelines, the method comprising:
[0006] Get the log address of failed tasks;
[0007] Obtain the log content based on the log address;
[0008] The log content is input into a preset error pattern library for pattern matching to obtain the matched error patterns and their corresponding matching confidence levels.
[0009] If the matching confidence level is greater than a preset threshold, a first diagnostic result is generated based on the matched error pattern.
[0010] Otherwise, log-related data is read based on the log content, and a second diagnostic result is generated based on the log-related data.
[0011] In one embodiment of the error mode diagnosis method for data pipelines in this application, obtaining the log content based on the log address includes:
[0012] Obtain the attribute information of the failed task;
[0013] Read the raw log data according to the log address;
[0014] The original log data is sliced according to the attribute information to obtain sliced logs;
[0015] Remove the security information from the slice log to obtain the log content.
[0016] In one embodiment of the error mode diagnosis method for data pipelines in this application, the step of inputting the log content into a preset error mode library for pattern matching includes:
[0017] The log content is matched against each error pattern in the preset error pattern library to obtain the rule matching result and the corresponding rule matching confidence.
[0018] The log content is matched with each error pattern in the preset error pattern library to obtain the template matching result and the corresponding template matching confidence.
[0019] The log content is vectorized to generate a log feature vector. The similarity between the log feature vector and the pattern feature vector of each error pattern in the preset error pattern library is calculated to obtain the vector matching result and the corresponding vector matching confidence.
[0020] The overall confidence level is determined based on the rule matching confidence level, template matching confidence level, and vector matching confidence level.
[0021] The final matched error pattern and match confidence are determined based on the overall confidence level.
[0022] In one embodiment of the error mode diagnosis method for data pipelines in this application, the attribute information includes at least one of task number, failed node, image version, code version, and trigger time; the log association data includes at least one of de-identified logs, node configuration, code snippets, historical cases, and runtime metrics.
[0023] The step of reading log-related data based on the log content includes:
[0024] Based on the task number and failure node in the log content, determine the target node to which the failed task belongs;
[0025] Obtain the de-identified logs and node configuration corresponding to the target node;
[0026] Obtain the code snippet that matches the image version or code version of the failed task;
[0027] Retrieve historical cases that match the log content;
[0028] Obtain the runtime metrics corresponding to the trigger time of the failed task.
[0029] In one embodiment of the error mode diagnosis method for data pipelines in this application, the second diagnosis result includes at least one of the following: fault cause, judgment basis, scope of impact, and fault handling suggestions;
[0030] The step of generating a second diagnostic result based on the log association data includes:
[0031] Root cause analysis is performed based on the log association data to generate the cause of failure for the failed task;
[0032] Based on the cause of the fault, data corresponding to the cause of the fault is obtained from the log-related data, which serves as the basis for determining the cause of the fault;
[0033] The scope of impact of the failed task is determined based on the associated log data.
[0034] Based on the cause of the fault, the judgment criteria, and the scope of impact, a fault handling opinion is generated.
[0035] In one embodiment of the error mode diagnosis method for data pipelines in this application, the second diagnosis result further includes candidate error modes; the method further includes:
[0036] Candidate responsible persons are generated based on the attribute information of failed tasks;
[0037] Based on the second diagnostic result and the candidate responsible person, obtain the manual confirmation result;
[0038] Based on the results of the manual verification, determine whether to store the candidate error pattern in the preset error pattern library.
[0039] In one embodiment of the error mode diagnosis method for data pipelines in this application, determining whether to store the candidate error mode in the preset error mode library includes:
[0040] If the manual confirmation result confirms that the candidate error pattern is valid, then the candidate error pattern is stored in the preset error pattern library;
[0041] If the manual confirmation result is to reject the candidate error pattern, then the candidate error pattern is not stored.
[0042] In one embodiment of the error mode diagnosis method for data pipelines in this application, the method further includes:
[0043] Set the lifecycle status of the candidate error modes stored in the preset error mode library to pending confirmation;
[0044] When the preset confirmation conditions are met, the lifecycle status of the candidate error mode is updated from pending confirmation to confirmed.
[0045] In a second aspect, an electronic device is provided, comprising:
[0046] At least one processor;
[0047] And, a memory communicatively connected to the at least one processor;
[0048] The memory stores a computer program, which, when executed by the at least one processor, is the aforementioned error mode diagnosis method for data pipelines.
[0049] In a third aspect, a computer-readable storage medium is provided, wherein a plurality of program codes are stored therein, the program codes being adapted to be loaded and run by a processor to perform the error mode diagnosis method for data pipelines as described in any of the preceding claims.
[0050] The above-described technical solutions of this application have at least one or more of the following beneficial effects:
[0051] The error mode diagnosis method for data pipelines in this application includes: obtaining the log address of the failed task; obtaining the log content based on the log address; inputting the log content into a preset error mode library for pattern matching to obtain the matched error mode and the corresponding matching confidence level; if the matching confidence level is greater than a preset threshold, generating a first diagnostic result based on the matched error mode; otherwise, reading log-related data based on the log content and generating a second diagnostic result based on the log-related data. By obtaining the log addresses of failed tasks and retrieving their contents, the diagnostic process is automatically triggered, eliminating the need for manual log location. The log content is input into a pre-defined error pattern library for pattern matching, yielding the matched error patterns and their confidence levels, enabling rapid identification of known errors. When the matching confidence level is greater than a pre-defined threshold, a first diagnostic result is generated directly based on the matched error pattern, achieving rapid response and significantly improving diagnostic efficiency in high-confidence scenarios. Conversely, when the matching confidence level is less than or equal to the pre-defined threshold, related data is read from the log content, and a second diagnostic result is generated based on this data. This ensures that even in low-confidence or unmatched scenarios, accurate root cause analysis and treatment suggestions can still be obtained through in-depth diagnostics, thus forming a tiered diagnostic mechanism that balances efficiency and accuracy. Attached Figure Description
[0052] The disclosure of this application will become more readily understood with reference to the accompanying drawings. It will be readily understood by those skilled in the art that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of this application. Furthermore, similar numbers in the drawings are used to denote similar components, wherein:
[0053] Figure 1 This is a schematic diagram of the main flow of an error mode diagnosis method for data pipelines in one embodiment of this application;
[0054] Figure 2 This is a schematic diagram of the main structure of an electronic device in one embodiment of this application. Detailed Implementation
[0055] Some embodiments of this application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of this application and are not intended to limit the scope of protection of this application.
[0056] In the description of this application, "module" and "processor" can include hardware, software, or a combination of both. A module can include hardware circuitry, various suitable sensors, communication ports, memory, and can also include software components, such as program code, or a combination of software and hardware. A processor can be a central processing unit, microprocessor, image processor, digital signal processor, or any other suitable processor. The processor has data and / or signal processing capabilities. The processor can be implemented in software, in hardware, or a combination of both. Non-transitory computer-readable storage media includes any suitable medium capable of storing program code, such as magnetic disks, hard disks, optical disks, flash memory, read-only memory, random access memory, etc. The term "A and / or B" means all possible combinations of A and B, such as only A, only B, or A and B. The terms "at least one A or B" or "at least one of A and B" have a similar meaning to "A and / or B" and can include only A, only B, or A and B. The singular terms "a" or "this" can also include plural forms.
[0057] Before providing a further detailed description of the embodiments of this application, some of the nouns and terms involved in the embodiments of this application will be explained. The nouns and terms involved in the embodiments of this application are subject to the following interpretations.
[0058] A data pipeline refers to an automated process consisting of one or more nodes such as data acquisition, cleaning, processing, labeling, quality inspection, training, or delivery. In this application, it specifically refers to an automated process that may generate failure logs during task execution, and whose operational status needs to be monitored and diagnosed.
[0059] The error pattern library is a pre-built and stored collection of structured knowledge in the system, used to provide matching basis for error identification of failed tasks in the data pipeline. Each error pattern includes at least the error category, log characteristics, applicable nodes, triggering conditions, handling suggestions, responsible person mapping, confidence threshold, and life cycle status. The life cycle status includes at least pending confirmation, confirmed, and obsolete.
[0060] Log-related data refers to the contextual information that needs to be analyzed in addition to the log content itself during in-depth diagnostics.
[0061] Attribute information refers to metadata information that describes failed tasks, and is used to accurately locate and slice logs.
[0062] See appendix Figure 1 , Figure 1 This is a schematic diagram of the main flow of an error mode diagnosis method for data pipelines according to an embodiment of this application. Figure 1As shown, the error mode diagnosis method for data pipelines in this application embodiment mainly includes the following steps S10-S50.
[0063] Step S10: Obtain the log address of the failed task.
[0064] Step S20: Obtain the log content based on the log address.
[0065] Step S30: Input the log content into the preset error pattern library for pattern matching, and obtain the matched error patterns and corresponding matching confidence levels.
[0066] Step S40: If the matching confidence is greater than the preset threshold, then generate the first diagnostic result based on the matched error pattern.
[0067] Step S50: Otherwise, read the log-related data based on the log content and generate a second diagnostic result based on the log-related data.
[0068] Based on steps S10-S50 above, by obtaining the log address of the failed task and retrieving the log content accordingly, the diagnostic process is automatically triggered without manual log location. The log content is input into a preset error pattern library for pattern matching, and the matched error patterns and their confidence levels are obtained, enabling known errors to be quickly identified. When the matching confidence level is greater than a preset threshold, a first diagnostic result is directly generated based on the matched error pattern, enabling rapid response and significantly improving diagnostic efficiency in high-confidence scenarios. When the matching confidence level is less than or equal to the preset threshold, related data is read based on the log content, and a second diagnostic result is generated based on the related data. This ensures that scenarios with low confidence or no matching can still obtain accurate root cause analysis and processing suggestions through deep diagnosis, thus forming a hierarchical diagnostic mechanism that balances efficiency and accuracy.
[0069] The following provides further explanation of steps S10 to S50.
[0070] Specifically, regarding step S10, during the data pipeline operation, when a task fails, the system automatically records its status and generates a log file. The storage location of this log file is the log address. Obtaining this address is the entry point to start the entire automated diagnostic process. Through this step, the starting point that originally required manual intervention is automated, solving the delays and omissions caused by manual monitoring and manual triggering of diagnostics.
[0071] The above is a further explanation of step S10. Step S20 will be further explained below.
[0072] In a specific embodiment, step S20 involves obtaining log content based on a log address, including: obtaining attribute information of the failed task; reading raw log data based on the log address; slicing the raw log data based on the attribute information to obtain sliced logs; and removing security information from the sliced logs to obtain the log content.
[0073] Specifically, the process begins by acquiring the attribute information of the failed task. This information may include the task number, failure node, project, data batch, runtime parameters, image version, code version, and trigger time. Next, the raw, unprocessed log data is read from the log address. Since the raw logs are typically very large, containing all records from the start to the end of the task, the raw log data is sliced based on the attribute information, extracting only key log segments relevant to the time before and after the failure, forming sliced logs. Log slicing significantly reduces the amount of data for subsequent analysis, improving processing efficiency. For example, raw log data can be sliced based on trigger time and failure node, or it can be sliced based on trigger time, failure node, error stack, resource metrics, and key anomaly segments to obtain sliced logs. Finally, to protect data security and user privacy, security information such as keys, accounts, customer data, and other sensitive information is further removed from the sliced logs, resulting in clean and secure log content.
[0074] The above is a further explanation of step S20. Step S30 will be further explained below.
[0075] Regarding step S30, in one specific embodiment, the log content is input into a preset error pattern library for pattern matching, including: performing rule matching between the log content and each error pattern in the preset error pattern library to obtain rule matching results and corresponding rule matching confidence; performing template matching between the log content and each error pattern in the preset error pattern library to obtain template matching results and corresponding template matching confidence; vectorizing the log content to generate log feature vectors, calculating the similarity between the log feature vectors and the pattern feature vectors of each error pattern in the preset error pattern library to obtain vector matching results and corresponding vector matching confidence; determining a comprehensive confidence based on rule matching confidence, template matching confidence, and vector matching confidence; and determining the final matched error pattern and matching confidence based on the comprehensive confidence.
[0076] Specifically, the process begins with rule matching, which involves screening log content based on predefined hard logic rules (such as specific error codes or keywords) to obtain rule matching results and corresponding rule matching confidence levels. Next, template matching is performed using predefined log templates (such as fixed log formats or structural patterns) to obtain template matching results and corresponding template matching confidence levels. Finally, vectorized similarity calculation is used, converting the log text into mathematical vectors and comparing their similarity with vectors of existing error patterns in an error pattern library to capture semantic similarity and obtain vector matching results and corresponding vector matching confidence levels. Then, through weighted summation, voting, or other predefined fusion strategies, a comprehensive confidence level that fully reflects the reliability of the matching is calculated. Finally, based on this comprehensive confidence level, the error pattern that best matches the current log and its final confidence level are selected from candidate error patterns, thus achieving high-accuracy classification and identification of log errors.
[0077] The above is a further explanation of step S30. Step S40 will be further explained below.
[0078] Specifically, when the matching confidence level exceeds a preset threshold, information such as handling suggestions, responsible person mapping, and error category pre-associated with the error pattern is extracted from a preset error pattern library. Based on this, according to the error category, log characteristics, and the trigger time and node information of the failed task, a first diagnostic result containing the error category, handling suggestions, and responsible person information is automatically generated. The handling suggestions can be directly output as processing steps in the diagnostic conclusion, and the responsible person mapping can be directly filled into the responsible person field of the diagnostic conclusion. Subsequently, the first diagnostic result can be pushed to the notification module for rapid response to failed tasks, thereby achieving efficient handling of known errors.
[0079] The above is a further explanation of step S40. Step S50 will be further explained below.
[0080] In one specific embodiment, the log-related data includes at least one of the following: anonymized logs, node configurations, code snippets, historical cases, and runtime metrics. Reading log-related data based on log content includes: determining the target node to which the failed task belongs based on the task number and failed node in the log content; obtaining the anonymized logs and node configurations corresponding to the target node; obtaining the code snippet that matches the image version or code version of the failed task; obtaining the historical cases that match the log content; and obtaining the runtime metrics corresponding to the trigger time of the failed task.
[0081] Specifically, the system first pinpoints the target node that failed based on the task number and failure node in the log content. Next, it retrieves the anonymized logs and node configuration files corresponding to that target node to understand its operating environment. Then, it automatically pulls relevant code snippets based on the image version or code version of the failed task, facilitating code-level analysis. Simultaneously, it can also obtain historical cases matching the current log content. Finally, it retrieves system performance metrics near the time the failed task was triggered, such as CPU and memory usage, to determine if the issue is resource-related. In this way, the system automatically aggregates information that was previously scattered across various locations, providing comprehensive context for in-depth diagnostics and solving the problem of inefficient manual cross-system information retrieval.
[0082] In one specific embodiment, the second diagnostic result includes at least one of the following: fault cause, judgment basis, scope of impact, and fault handling suggestions. Generating the second diagnostic result based on log association data includes: performing root cause analysis based on log association data to generate the fault cause of the failed task; obtaining data corresponding to the fault cause from the log association data based on the fault cause as the judgment basis for the fault cause; determining the scope of impact of the failed task based on the log association data; and generating fault handling suggestions based on the fault cause, judgment basis, and scope of impact to obtain the second diagnostic result.
[0083] Specifically, the process begins with root cause analysis based on comprehensive log correlation data. This can be achieved using methods such as large language models or expert systems to infer the cause of the failed task. Next, evidence directly related to this cause is automatically extracted from the log correlation data, such as specific log lines, abnormal configuration items, or lines of code, serving as the basis for judgment. Then, the potential impact of the failed task is analyzed and determined based on the log correlation data, including affected downstream tasks or data. Finally, a comprehensive analysis of the cause, judgment criteria, and impact scope generates specific troubleshooting recommendations, such as modifying a specific piece of code, rolling back a version, or adjusting a configuration. This approach transforms the output from vague error messages into a structured diagnostic report containing a complete analysis chain and action guidelines, improving the efficiency and quality of troubleshooting.
[0084] In one embodiment, the second diagnostic result further includes candidate error patterns; the method further includes: generating candidate responsible persons based on the attribute information of the failed task; obtaining manual confirmation results based on the second diagnostic result and the candidate responsible persons; and determining whether to store the candidate error patterns in a preset error pattern library based on the manual confirmation results.
[0085] Specifically, after generating the second diagnostic result, if it contains new candidate error patterns not yet included in the database, a confirmation process is initiated. First, one or more candidate responsible parties are generated based on the attribute information of the failed task. Then, the second diagnostic result, including the cause of the failure, the judgment criteria, the handling suggestions, and the candidate error patterns, along with the candidate responsible party information, is pushed to relevant personnel to obtain manual confirmation. Based on the received manual confirmation, it is determined whether to store the new candidate error pattern in the preset error pattern database. By verifying and confirming the machine diagnostic results through manual judgment, the potential for misjudgment and knowledge stagnation that may arise from relying solely on machines is resolved, enabling the diagnostic knowledge base to be dynamically updated.
[0086] In one embodiment, determining whether to store a candidate error pattern in a preset error pattern library includes: if the manual confirmation result is that the candidate error pattern is valid, then storing the candidate error pattern in the preset error pattern library; if the manual confirmation result is that the candidate error pattern is rejected, then abandoning the storage of the candidate error pattern.
[0087] Specifically, if the received manual verification result is positive, it means that the human expert agrees that the candidate error pattern is valid and valuable, and the candidate error pattern is stored in a pre-set error pattern library for future use when dealing with similar problems. Conversely, if the manual verification result is negative, it means that the expert believes that the candidate error pattern is incorrect or inaccurate, and the storage operation is abandoned. This storage mechanism based on manual verification helps ensure that all knowledge entering the error pattern library is verified and valid, avoiding the contamination and accumulation of erroneous knowledge.
[0088] Additionally, it includes: setting the lifecycle status of candidate error modes stored in the preset error mode library to pending confirmation; and updating the lifecycle status of candidate error modes from pending confirmation to confirmed when the preset confirmation conditions are met.
[0089] Specifically, to perform quality control and lifecycle management on newly added patterns, when a candidate error pattern is added to the preset error pattern library based on human verification, it is not immediately considered a fully reliable pattern. Instead, its lifecycle status is initially set to "pending verification." This indicates that although the pattern has undergone a single verification, its universality and accuracy still need to be tested by more cases. Only when a certain preset verification condition is met—for example, if the pattern is successfully matched multiple times in subsequent diagnoses and receives positive human feedback each time—will the candidate error pattern's lifecycle status be updated from "pending verification" to "verified." Furthermore, an error pattern may also change from "pending verification" or "verified" to "obsolete," for example, when its human verification is rejected, or when a pattern is not matched for a long time and is replaced by a new, better pattern. In such cases, its lifecycle status is set to "obsolete." By introducing a lifecycle management mechanism, the dynamic health and continuous iteration of the error pattern library are maintained, achieving a complete closed loop of knowledge accumulation, verification, application, and obsolescence.
[0090] It should be clarified that all the steps involved in this application, such as data processing, logical judgment, pattern matching, root cause analysis, and diagnostic conclusion generation, can be collaboratively executed by intelligent agents deployed in the data pipeline operation and maintenance system.
[0091] In addition, the time spent on each diagnosis, the hit rate, the results of manual confirmation, and the cost of model calls can be recorded for subsequent threshold and strategy optimization.
[0092] This application constructs a hierarchical diagnostic mechanism that prioritizes low-cost matching of known high-frequency errors using an error pattern library. Only when the matching confidence is below a preset threshold or a match cannot be found will log-related data be read and a second diagnostic result generated. This significantly reduces model call costs and system resource overhead while ensuring diagnostic accuracy, and shortens response time in high-frequency error scenarios.
[0093] This application performs joint attribution analysis on multi-dimensional information such as failed nodes, log fragments, task configurations, code versions, data batch attribution, and responsible person mapping. It can automatically generate candidate responsible persons and their corresponding confidence levels, reducing the time that operation and maintenance personnel need to repeatedly query and correlate information across multiple systems, and improving the efficiency of fault handling.
[0094] This application introduces a pattern evolution mechanism driven by manual confirmation results, which can automatically complete the entry of candidate error patterns into the database and the status change of existing error patterns (including pending confirmation, confirmed, and obsolete) based on the manual confirmation results. This allows each investigation experience to be precipitated into reusable diagnostic capabilities, enabling continuous optimization of the error pattern database and self-evolution of the system.
[0095] This application retains diagnostic conclusions, judgment criteria, responsible persons, handling suggestions, and manual confirmation records in a structured format, which facilitates subsequent fault review, audit tracking, and continuous governance of diagnostic quality, and provides traceable technical support for the stable operation of the data pipeline.
[0096] It should be noted that although the steps in the above embodiments are described in a specific order, those skilled in the art will understand that in order to achieve the effect of this application, different steps do not necessarily have to be executed in such an order. They can be executed simultaneously (in parallel) or in other orders, and these variations are all within the scope of protection of this application.
[0097] Those skilled in the art will understand that all or part of the processes in the method of the above-described embodiment can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable storage medium can include any entity or device capable of carrying the computer program code, a medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory, a random access memory, an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0098] Furthermore, this application also provides an electronic device, which may include at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program, which, when executed by the at least one processor, implements the error mode diagnosis method for data pipelines described in any of the above embodiments. See also Figure 2 As shown, Figure 2 The structure of an electronic device, including a processor 100 and a memory 200, is illustrated by way of example.
[0099] Furthermore, this application also provides a computer-readable storage medium. In one embodiment of the computer-readable storage medium according to this application, the computer-readable storage medium can be configured to store a program that performs the error mode diagnosis method for data pipelines described in the above-described method embodiments. This program can be loaded and run by a processor to implement the above-described error mode diagnosis method for data pipelines. For ease of explanation, only the parts related to the embodiments of this application are shown; for specific technical details not disclosed, please refer to the method section of the embodiments of this application. The computer-readable storage medium can be a memory device formed by various electronic devices. Optionally, in the embodiments of this application, the computer-readable storage medium is a non-transitory computer-readable storage medium.
[0100] The technical solution of this application has been described in conjunction with the specific embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of this application is obviously not limited to these specific embodiments. Without departing from the principles of this application, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of this application.
Claims
1. A method for error mode diagnosis in data pipelines, characterized in that, The method includes: Get the log address of failed tasks; Obtain the log content based on the log address; The log content is input into a preset error pattern library for pattern matching to obtain the matched error patterns and their corresponding matching confidence levels. If the matching confidence level is greater than a preset threshold, a first diagnostic result is generated based on the matched error pattern. Otherwise, log-related data is read based on the log content, and a second diagnostic result is generated based on the log-related data.
2. The error mode diagnosis method for data pipelines according to claim 1, characterized in that, The step of obtaining the log content based on the log address includes: Obtain the attribute information of the failed task; Read the raw log data according to the log address; The original log data is sliced according to the attribute information to obtain sliced logs; Remove the security information from the slice log to obtain the log content.
3. The error mode diagnosis method for data pipelines according to claim 1, characterized in that, The step of inputting the log content into a preset error pattern library for pattern matching includes: The log content is matched against each error pattern in the preset error pattern library to obtain the rule matching result and the corresponding rule matching confidence. The log content is matched with each error pattern in the preset error pattern library to obtain the template matching result and the corresponding template matching confidence. The log content is vectorized to generate a log feature vector. The similarity between the log feature vector and the pattern feature vector of each error pattern in the preset error pattern library is calculated to obtain the vector matching result and the corresponding vector matching confidence. The overall confidence level is determined based on the rule matching confidence level, template matching confidence level, and vector matching confidence level. The final matched error pattern and match confidence are determined based on the overall confidence level.
4. The error mode diagnosis method for data pipelines according to claim 2, characterized in that, The attribute information includes at least one of the following: task number, failed node, image version, code version, and trigger time; the log-related data includes at least one of the following: anonymized logs, node configuration, code snippets, historical cases, and runtime metrics. The step of reading log-related data based on the log content includes: Based on the task number and failure node in the log content, determine the target node to which the failed task belongs; Obtain the de-identified logs and node configuration corresponding to the target node; Obtain the code snippet that matches the image version or code version of the failed task; Retrieve historical cases that match the log content; Obtain the runtime metrics corresponding to the trigger time of the failed task.
5. The error mode diagnosis method for data pipelines according to claim 1, characterized in that, The second diagnostic result includes at least one of the following: cause of the fault, basis for judgment, scope of impact, and fault handling recommendations; The step of generating a second diagnostic result based on the log association data includes: Root cause analysis is performed based on the log association data to generate the cause of failure for the failed task; Based on the cause of the fault, data corresponding to the cause of the fault is obtained from the log-related data, which serves as the basis for determining the cause of the fault; The scope of impact of the failed task is determined based on the associated log data. Based on the cause of the fault, the judgment criteria, and the scope of impact, a fault handling opinion is generated.
6. The error mode diagnosis method for data pipelines according to claim 5, characterized in that, The second diagnostic result also includes candidate error patterns; the method further includes: Candidate responsible persons are generated based on the attribute information of failed tasks; Based on the second diagnostic result and the candidate responsible person, obtain the manual confirmation result; Based on the results of the manual verification, determine whether to store the candidate error pattern in the preset error pattern library.
7. The error mode diagnosis method for data pipelines according to claim 6, characterized in that, The step of determining whether to store the candidate error pattern in the preset error pattern library includes: If the manual confirmation result confirms that the candidate error pattern is valid, then the candidate error pattern is stored in the preset error pattern library; If the manual confirmation result is to reject the candidate error pattern, then the candidate error pattern is not stored.
8. The error mode diagnosis method for data pipelines according to claim 7, characterized in that, The method further includes: Set the lifecycle status of the candidate error modes stored in the preset error mode library to pending confirmation; When the preset confirmation conditions are met, the lifecycle status of the candidate error mode is updated from pending confirmation to confirmed.
9. An electronic device, characterized in that, include: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores a computer program that, when executed by the at least one processor, implements the error mode diagnosis method for data pipelines as described in any one of claims 1 to 8.
10. A computer-readable storage medium storing a plurality of program codes, characterized in that, The program code is adapted to be loaded and run by a processor to perform the error mode diagnosis method for data pipelines as described in any one of claims 1 to 8.