A Fault Log Processing Method, Device, Program Product and Medium

By standardizing and classifying the server log flow, accurately log details and summary data are generated, and the server diagnostic model cannot perceive the association between components and other factors is solved, achieving efficient fault diagnosis.

CN119537084BActive Publication Date: 2025-07-22INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510104236.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-07-22
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Existing server diagnostic models cannot effectively perceive the association between components and other factors, resulting in problems such as fault location interference and incomplete data coverage.

Method used

By analyzing the server log flow, standardized logs are generated, the status changes of the faulty components are recorded, and classification is carried out according to preset division rules to form different types of classification logs and summary data for diagnostic models to diagnose faults.

Benefits of technology

The data decoupling of the diagnostic model is achieved, allowing it to focus on diagnostic rules development or algorithm training without the need to deal with intricate server logs, improving the accuracy and efficiency of fault diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119537084B_ABST
    Figure CN119537084B_ABST
Patent Text Reader

Abstract

The present application discloses a fault log processing method, device, program product and medium, relating to the field of computer technology. The method includes: obtaining different types of fault logs and performing standardization processing on the fault logs to generate standardized logs; analyzing the log streams in the standardized logs to determine faulty components and recording the status change details of the faulty components to generate log details corresponding to the standardized logs; dividing the standardized logs according to a preset division rule to generate different types of classified logs, and respectively analyzing and processing the different types of classified logs to obtain respective corresponding target summary data; and performing fault diagnosis based on the log details and the target summary data. Through the technical solution of the present application, the fault diagnosis model can focus on the development of diagnosis rules or algorithm training without paying attention to the intricate data processing, realizing the data decoupling of the diagnosis model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to a method, device, program product and medium for processing fault logs. Background Art

[0002] As the core IT (Internet Technology) infrastructure, the stability of a server directly affects the IT service quality of a data center. Once a server fails, customers require rapid fault diagnosis, accurate problem location, and business recovery to ensure business continuity. Server diagnosis models can quickly analyze server logs and locate faults in an automated manner, improving the operation and maintenance efficiency of servers and becoming increasingly popular in data centers.

[0003] Current diagnosis models generally directly match fault keywords or fault rules based on the original server logs and dynamically maintain the keyword library or rules as needed, generally evolving and changing in a log stream manner. This method can only obtain the current state of the server and cannot perceive the association between components themselves and other factors. For diagnosis solutions in some specific scenarios, such as hard disk fault diagnosis, rule matching and diagnosis are performed on specific types of logs to locate problems related to hard disk components. However, there are numerous server fault logs, and the data structures and contents of various types of logs vary greatly. Directly performing rule matching and fault location on various types of data through a diagnosis model will result in situations such as positioning interference; and relying solely on the error analysis of a single log cannot cover all the associated data required for diagnosis.

[0004] Therefore, how to provide a solution to the above technical problems is an issue that those skilled in the art need to solve currently. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a method, device, program product and medium for processing fault logs. By analyzing the log stream of a server and summarizing and classifying various types of server logs, the diagnosis model can focus on diagnostic rules or algorithm training without having to concern itself with the intricate processing of server logs. The specific solutions are as follows:

[0006] In a first aspect, the present application discloses a method for processing fault logs, including:

[0007] Obtaining different types of fault logs and performing standardization processing on the fault logs to generate standardized logs;

[0008] Analyzing the log stream in the standardized logs to determine the faulty components and recording the status change details of the faulty components to generate log details corresponding to the standardized logs;

[0009] Divide the standardized logs according to the preset division rules to generate different types of classified logs, and analyze and process the different types of classified logs respectively to obtain the corresponding target summary data;

[0010] Perform fault diagnosis based on the log details and the target summary data.

[0011] Optionally, obtain different types of fault logs and standardize the fault logs to generate standardized logs, including:

[0012] Obtain different types of fault logs and perform fuzzy queries on the fault logs according to the log names to obtain multiple log classifications;

[0013] For each fault log in each log classification, perform structure recognition and parsing to obtain the corresponding target parameters, and standardize each line of data in the log stream corresponding to the fault log based on the target parameters to generate the standardized logs under each log classification.

[0014] Optionally, the target parameters include static parameters and dynamic parameters; the static parameters include component types and component slots, and the dynamic parameters include occurrence time, fault keywords, fault levels, problem status, and component information;

[0015] Among them, analyze the log stream in the standardized logs to determine the faulty components, and record the detailed status changes of the faulty components to generate the log details corresponding to the standardized logs, including:

[0016] Use the fault keywords to match in the log stream of the standardized logs to determine the faulty components, and record the detailed status changes of the faulty components according to the problem status to generate the log details corresponding to the standardized logs.

[0017] Optionally, record the detailed status changes of the faulty components to generate the log details corresponding to the standardized logs, including:

[0018] Obtain the self-status changes of the faulty components and the status changes of other factors related to the associated faulty components respectively; among them, the status changes of other factors are the status changes of the factors with a context relationship with the faulty components in the fault logs;

[0019] Combine the self-status changes and the status changes of other factors to generate composite status changes, and determine the detailed status changes of the faulty components according to the composite status changes to obtain the log details corresponding to the standardized logs.

[0020] Optionally, divide the standardized logs according to the preset division rules to generate different types of classified logs, including:

[0021] Divide the standardized logs into out-of-band general logs, storage specialty logs, and computing specialty logs according to the fault information recorded in the standardized logs. Among them, the out-of-band general logs contain the fault information of each component, the storage specialty logs contain the fault information of each device used for storage, and the computing specialty logs contain the fault information of each device used for computing.

[0022] Optionally, the out-of-band general logs include device fault diagnosis logs, system event logs, sensor logs, and black box logs.

[0023] Optionally, analyze and process the out-of-band general logs to obtain the corresponding target summary data, including:

[0024] Send the device fault diagnosis logs into the summary data to determine the first out-of-band general summary data;

[0025] Compare the system event logs with the first out-of-band general summary data to receive the target system event logs in the system event logs and obtain the second out-of-band general summary data. Among them, the target system event logs are the system event logs that do not have information with the same component slot and problem status after being compared with the first out-of-band general summary data;

[0026] Compare the black box logs with the second out-of-band general summary data to receive the target black box logs in the black box logs and obtain the third out-of-band general summary data. Among them, the target black box logs are the black box logs that do not have information with the same component slot and problem status after being compared with the second out-of-band general summary data;

[0027] Compare the sensor logs with the third out-of-band general summary data to receive the target sensor logs in the sensor logs and obtain the fourth out-of-band general summary data. Among them, the target sensor logs are the sensor logs that do not have information with the same component slot and problem status after being compared with the third out-of-band general summary data;

[0028] Perform status correction on the fourth out-of-band general summary data according to the current status of the components recorded in the sensor logs to obtain the target summary data corresponding to the out-of-band general logs.

[0029] Optionally, the storage specialty logs include: redundant array of independent disks logs, self-monitoring, analysis and reporting technology logs, physical drive information logs, and first message logs; the first message logs are the logs that record the message information of each device used for storage.

[0030] Optionally, analyze and process the storage specialty logs to obtain the corresponding target summary data, including:

[0031] Send the physical drive information logs into the summary data to determine the first storage specialty summary data;

[0032] Compare the redundant array of independent disks (RAID) log with the first storage specialty summary data to receive the target RAID log in the RAID log, and obtain the second storage specialty summary data; wherein, the target RAID log is the RAID log that does not have information with the same component slot and problem status after being compared with the first storage specialty summary data.

[0033] Compare the first message log with the second storage specialty summary data to receive the target first message log in the first message log, and obtain the third storage specialty summary data; wherein, the target first message log is the first message log that does not have information with the same component slot and problem status after being compared with the second storage specialty summary data.

[0034] Compare the self - monitoring, analysis, and reporting technology (SMART) log with the third storage specialty summary data to receive the target SMART log in the SMART log, and obtain the fourth storage specialty summary data; wherein, the target SMART log is the SMART log that does not have information with the same component slot and problem status after being compared with the third storage specialty summary data.

[0035] Perform status correction on the fourth storage specialty summary data according to the current status of the components recorded in the SMART log to obtain the target summary data corresponding to the storage specialty log.

[0036] Optionally, the computing specialty log includes: internal error log, maintenance log, process analysis log, and the second message log; the second message log is the log that records the message information of each device used for computing.

[0037] Optionally, analyzing and processing the computing specialty log to obtain the corresponding target summary data includes:

[0038] Send the second message log into the summary data to determine the first computing specialty summary data.

[0039] Compare the process analysis log with the first computing specialty summary data to receive the target process analysis log in the process analysis log, and obtain the second computing specialty summary data; wherein, the target process analysis log is the process analysis log that does not have information with the same component slot and problem status after being compared with the first computing specialty summary data.

[0040] Compare the internal error log with the second computing specialty summary data to receive the target internal error log in the internal error log and obtain the third computing specialty summary data; wherein, the target internal error log is the internal error log that does not have information with the same component slot and problem status after being compared with the second computing specialty summary data;

[0041] Compare the maintenance log with the third computing specialty summary data to receive the target maintenance log in the maintenance log and obtain the fourth computing specialty summary data; wherein, the target maintenance log is the maintenance log that does not have information with the same component slot and problem status after being compared with the third computing specialty summary data;

[0042] Perform status correction on the fourth computing specialty summary data according to the current status of the components recorded in the maintenance log to obtain the target summary data corresponding to the computing specialty log.

[0043] Optionally, perform fault diagnosis based on the log details and the target summary data, including:

[0044] Obtain the first diagnostic rule and the second diagnostic rule respectively, and combine the first diagnostic rule and the second diagnostic rule to obtain the target decision rule; wherein, the first diagnostic rule is the diagnostic rule obtained by sorting out and solidifying the rules of existing fault diagnosis instances, and the second diagnostic rule is the diagnostic rule obtained by performing rule matching on the fault data through an artificial intelligence algorithm;

[0045] Perform fault diagnosis on the log details and the target summary data by using the target decision rule, and output the root cause of the fault of the faulty component.

[0046] In a second aspect, the present application discloses an electronic device, including:

[0047] A memory for storing a computer program;

[0048] A processor for loading and executing the computer program to implement the foregoing fault log processing method.

[0049] In a third aspect, the present application discloses a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the foregoing fault log processing method are implemented.

[0050] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the foregoing fault log processing method is implemented.

[0051] The present application provides a method for processing fault logs, including: obtaining different types of fault logs, and performing standardization processing on the fault logs to generate standardized logs; analyzing the log streams in the standardized logs to determine faulty components, and recording the details of the status changes of the faulty components to generate log details corresponding to the standardized logs; dividing the standardized logs according to a preset division rule to generate different types of classified logs, and respectively analyzing and processing the different types of classified logs to obtain respective target summary data; performing fault diagnosis based on the log details and the target summary data.

[0052] The beneficial technical effects of the present application are as follows: By incorporating as many different types of fault logs into the analysis as possible, standardized logs with a standard structure are first generated; then, log stream analysis is performed on the standardized logs, not limited to screening for fault keywords in individual log points. The status of faulty components in the logs can be identified based on the continuous change characteristics of the logs, and the generated log details accurately reflect the true status of the components; further, the standardized logs are classified according to a preset division rule to obtain the target summary data corresponding to different types of classified logs. Therefore, each type of target summary data is not limited to a single type of fault log; finally, the generated log details and target summary data can be called by the diagnostic model for fault diagnosis, realizing data decoupling for the diagnostic model. In this way, the fault diagnosis model can focus on the development of diagnostic rules or algorithm training without having to concern itself with complex data processing.

[0053] In addition, a fault log processing device, program product, and medium provided by the present application correspond to the above-mentioned fault log processing method, and the effects are the same. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0055] Figure 1 It is a flowchart of a method for processing fault logs disclosed in the present application;

[0056] Figure 2 It is a schematic diagram of fault diagnosis disclosed in the present application;

[0057] Figure 3 It is a flowchart of the implementation details of a method for processing fault logs disclosed in the present application;

[0058] Figure 4 It is a schematic diagram of the data flow in a fault log disclosed in the present application;

[0059] Figure 5 A schematic diagram of a fault analysis process disclosed in this application;

[0060] Figure 6 A schematic diagram of the data flow in a fault log disclosed in this application;

[0061] Figure 7 A schematic diagram of a fault analysis process disclosed in this application;

[0062] Figure 8 A schematic diagram of the data flow in a fault log disclosed in this application;

[0063] Figure 9 A schematic diagram of a fault analysis process disclosed in this application;

[0064] Figure 10 A schematic diagram of platform diagnosis disclosed in this application;

[0065] Figure 11 A schematic diagram of the processing record of a fault log disclosed in this application;

[0066] Figure 12 A schematic diagram of the out-of-band general log summary process disclosed in this application;

[0067] Figure 13 A schematic diagram of the storage specialist log summary process disclosed in this application;

[0068] Figure 14 A schematic diagram of the computing specialist log summary process disclosed in this application;

[0069] Figure 15 A schematic diagram of the summary log disclosed in this application;

[0070] Figure 16 A schematic diagram of the out-of-band general data summary before processing disclosed in this application;

[0071] Figure 17 A schematic diagram of the out-of-band general data summary after processing disclosed in this application;

[0072] Figure 18 A schematic diagram of the storage specialist data summary before processing disclosed in this application;

[0073] Figure 19 A schematic diagram of the storage specialist data summary after processing disclosed in this application;

[0074] Figure 20 A schematic diagram of the overall fault log processing flow disclosed in this application;

[0075] Figure 21 Structural schematic diagram of a fault log processing device disclosed in this application;

[0076] Figure 22 Structural diagram of an electronic device disclosed in this application. Detailed implementation manners

[0077] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0078] Currently, server diagnosis models generally directly perform fault keyword or fault rule matching based on the original server logs, and dynamically maintain the keyword library or rules as needed. For example, the server SEL (System Event Log) as a general server log records most of the server component error reports, and currently many diagnosis solutions are based on SEL logs for server rule matching and diagnosis. However, the server's fault logs generally evolve and change in the form of a log stream, and the components, thresholds, etc. of the server involved are also in a continuous change process. Simply screening the corresponding fault states or keywords of a certain component or threshold can only obtain the current state and cannot perceive the association between the component itself and other factors.

[0079] For diagnosis solutions in some specific scenarios, such as hard disk fault diagnosis, rule matching and diagnosis are performed on specific types of logs, such as Smart (Self-Monitoring Analysis and Reporting Technology) logs, Raid (Redundant Array of Independent Disks) logs, etc. to locate hard disk component-related problems. However, there are numerous server fault logs, and the data structures and contents of various types of logs vary greatly. How to mine fault parameters from various non-standard data is a difficult point. Since there are correlations or independences between various types of logs and the relationships are relatively complex, directly performing rule matching and fault location on various types of data through a diagnosis model will result in situations such as location interference; and relying solely on the error analysis of a single log cannot cover all the associated data required for diagnosis.

[0080] To this end, the present application provides a fault log processing solution. By analyzing the log stream of the server and summarizing and classifying multiple types of server logs, the diagnostic model can focus on diagnostic rules or algorithm training without having to concern itself with the intricate server log processing.

[0081] An embodiment of the present invention discloses a fault log processing method. Refer to Figure 1 as shown, the method includes:

[0082] Step S11: Obtain different types of fault logs and perform standardization processing on the fault logs to generate standardized logs.

[0083] The reasons for server failures are intricate, and the logs recording error reports are also of various types. The log data for corresponding fault analysis and location is not limited to a single type of I fault log, such as IPMI (Intelligent Platform Management Interface), SMART, etc. logs. Therefore, in the embodiments of the present application, as many logs related to all server failures as possible are included in the analysis and converted into standardized error reporting data.

[0084] Specifically, obtain different types of fault logs, perform fuzzy queries on the fault logs according to the log names to obtain multiple log classifications; respectively perform structure recognition and parsing on the fault logs in each log classification to obtain corresponding target parameters, and based on the target parameters, perform standardization processing on each line of data in the log stream corresponding to the fault logs to generate standardized logs under each log classification.

[0085] In the embodiments of the present application, after obtaining different types of fault logs, the fault logs are searched and classified according to the log names. Taking the SEL log as an example, fuzzy queries are generally performed according to common naming methods such as sel.log, SEL.log, SEL_xx.log, etc. to match the SEL log. The IPMI logs, RAID logs, and other logs required for server diagnosis are obtained in the same way.

[0086] It can be understood that the types of fault logs under the same log classification are the same. If no valid log is matched during the query process, an error will be reported accordingly. After the query, according to the log classification, the log structures are respectively recognized and parsed to obtain corresponding target parameters. Among them, the target parameters include static parameters and dynamic parameters; the static parameters include component types and component slots, and the dynamic parameters include occurrence time, fault keywords, fault levels, problem status, and component information. Further, according to the positions and structures of key information such as target parameters, each line of data in the log stream is standardized to obtain standardized logs under different log classifications.

[0087] Step S12: Analyze the log stream in the standardized log to determine the faulty component, and record the details of the status change of the faulty component to generate log details corresponding to the standardized log.

[0088] In the embodiment of the present application, since the corresponding target parameters are obtained during the process of standardizing the fault log, the fault keywords can be used to match in the log stream of the standardized log to determine the faulty component.

[0089] Furthermore, the fault logs of the server, such as SEL, RAID, etc., are generally continuously recorded in the form of log streams, and the status of the corresponding components themselves and other factors also continuously change in the way of increasing time. Therefore, in the embodiment of the present application, the problem status of the faulty component is recorded. Specifically, the details of the status change of the faulty component are recorded according to the problem status to generate log details corresponding to the standardized log. It can be seen that the present invention is not limited to the screening of log point-like fault keywords, but performs stream analysis on the log based on the continuous change characteristics of the log, identifies the continuous change of the component status itself in the log and the association and influence of other factors on the component status, and accurately reflects the true status of the component.

[0090] Step S13: Divide the standardized log according to the preset division rules to generate different types of classified logs, and analyze and process the different types of classified logs respectively to obtain the corresponding target summary data.

[0091] After forming the standardized log from various logs and analyzing the log stream, there is another problem: multiple logs may record the fault point, and the fault data recorded by different logs, such as slot positions and statuses, are not the same. If direct diagnosis is performed without distinction, problems such as duplicate data diagnosis will occur. In order for the subsequent diagnosis model to diagnose correctly, it is necessary to integrate, deduplicate, and classify the logs.

[0092] In this process, the standardized log obtained by processing different types of fault logs is divided to generate different types of classified logs. Then, for each type of classified log, it is analyzed and processed to remove the duplicate data therein to obtain the final summary data.

[0093] In a specific embodiment, the different types of classified logs generated by partitioning include out-of-band general logs, storage specialist logs, and computing specialist logs; among them, the characteristic of out-of-band general logs is that they contain the fault information of each component, and the error reports are usually simple and direct, which is suitable for the direct judgment of simple faults; the characteristic of storage specialist logs is that they contain the fault information of each device used for storage (such as RAID cards, hard disks, etc.), and can be used to judge complex problems such as multiple hard disks and links, as well as the pre-inspection of potential fault types; the characteristic of computing specialist logs is that they contain the fault information of complex problems of each device used for computing (such as CPUs, memory, PCIE, etc.), and can be used to judge difficult problems of multiple components such as CPUs and memory.

[0094] Step S14: Perform fault diagnosis based on the log details and the target summary data.

[0095] Finally, perform fault diagnosis based on the log details and the target summary data, and output the root cause of the fault. It can be understood that by analyzing the log stream of the server and parsing out the log details of the continuous changes and multi-dimensional states of the components, it can provide accurate data for the diagnostic model to accurately judge the fault. After summarizing and classifying various logs of the server to form the standardized target summary data, since they all record the corresponding log details, therefore, during fault diagnosis, accurate and highly available classified data can be provided for the diagnostic model to call, enabling the diagnostic model or algorithm model to focus on diagnostic rules or algorithm training without having to worry about the complex processing of server log data, thus achieving data decoupling for the diagnostic model.

[0096] Specifically, when using the fault model for fault diagnosis: obtain the first diagnostic rule and the second diagnostic rule respectively, and combine the first diagnostic rule and the second diagnostic rule to obtain the target decision rule; based on the log details and the target summary data, use the target decision rule to perform fault diagnosis and output the root cause of the fault of the faulty component.

[0097] Among them, the first diagnostic rule is the diagnostic rule obtained by sorting out and solidifying the rules of existing fault diagnosis instances, and has the ability of large-scale and systematic analysis. Generally, the higher the quality of the data given, the more accurate the analysis conclusion; the second diagnostic rule is the diagnostic rule for rule matching of fault data through artificial intelligence (AI) algorithms. It usually involves processes such as feature engineering (word vector conversion word2vector, missing value balancing processing), model training (Xgboost model), and training and prediction using standardized classification data. Compared with the first diagnostic rule, the algorithm diagnosis corresponding to the second diagnostic rule requires higher data quality, and the more accurate the data, the more accurate the training and prediction effects.

[0098] Furthermore, the first diagnostic rule and the second diagnostic rule are summarized to form a target decision rule, and then the target summary data with attached log details is analyzed using the target decision rule to output the root cause, thereby realizing intelligent decision-making for fault diagnosis.

[0099] like Figure 2 The figure shows an exemplary method of using a diagnostic model to perform fault diagnosis to output the root cause of the fault. Among them, expert rule diagnosis corresponds to the first diagnostic rule, and AI algorithm diagnosis corresponds to the second diagnostic rule. Expert rule diagnosis sorts out and solidifies typical diagnostic examples based on experience, and AI algorithm diagnosis trains the algorithm model based on massive training data. By summarizing the results of expert rule diagnosis and AI algorithm diagnosis, the out-of-band general data obtained by classification, the stored specialist data, and the specialist data are calculated and the root cause is output. Since how to use the diagnostic model is not within the patent point of the present invention, it is mainly to complete the process. Therefore, the more detailed processing process can refer to the processing process of the existing diagnostic model, and the present invention will not repeat it.

[0100] The present application provides a fault log processing method, including: obtaining different types of fault logs, and standardizing the fault logs to generate standardized logs; analyzing the log stream in the standardized log to determine the faulty component, and recording the status change details of the faulty component to generate log details corresponding to the standardized log; dividing the standardized log according to preset division rules to generate different types of classified logs, and analyzing and processing the different types of classified logs respectively to obtain their respective corresponding target summary data; and performing fault diagnosis based on the log details and the target summary data.

[0101] The beneficial technical effects of the present application are: by incorporating as many different types of fault logs into the analysis as possible, a standardized log with a standard structure is first generated; then the standardized log is subjected to log stream analysis, which is not limited to the screening of log point fault keywords, and can identify the state of the faulty components in the log based on the continuous change characteristics of the log, and the generated log details accurately reflect the true state of the component; further, the standardized log is classified according to the preset division rules, and the target summary data corresponding to each different type of classified log is obtained, so each target summary data is not limited to a single type of fault log; finally, the generated log details and target summary data can be called by the diagnostic model for fault diagnosis, thereby realizing data decoupling of the diagnostic model. In this way, the fault diagnosis model can focus on diagnostic rule development or algorithm training without having to pay attention to complex data processing.

[0102] Based on the foregoing embodiments, this embodiment will specifically elaborate on S12 in the above embodiments. Fault logs of servers, such as SEL, RAID, etc., are generally continuously recorded in the form of log streams, and the status of corresponding components themselves and other factors also continuously change in an increasing order of time. Therefore, the process of recording the status change details of faulty components to generate log details corresponding to standardized logs may include the following steps:

[0103] Obtain the status change of the faulty component itself and the status change of other factors related to the faulty component respectively; wherein, the status change of other factors is the status change of the factors with a context relationship with the faulty component in the fault log;

[0104] Combine the status change of the component itself with the status change of other factors to generate a composite status change, and determine the status change details of the faulty component according to the composite status change to obtain log details corresponding to the standardized log.

[0105] In the embodiment of the present application, when a fault is triggered, the initialization of the target parameters is started. According to the foregoing content, the target parameters include static parameters and dynamic parameters. Recording the status change details of the faulty component, the core point of which is the problem status change.

[0106] In a specific implementation manner, regarding the recording of the status change of the component itself: Exemplarily, common single-point or multi-point statuses include: Issue Assert, indicating that a problem occurs; Issue Deassert, indicating that the problem is resolved; PresentAssert, indicating that the component is inserted; Present Deassert, indicating that the component is removed; Issue Assert>IssueDeassert, representing that the problem occurs and self-resolves; Present Assert>Issue Deassert, indicating the plugging and unplugging action of the component.

[0107] By combining the above several status changes, the current status can be determined. Exemplarily, the current status may include:

[0108] (1) Problem insertion: Only problem insertion, indicating that there is a problem currently;

[0109] (2) Problem self-repair: There is a problem and it is self-repaired. This scenario can be ignored in the inspection scenario

[0110] (3) Problem manual repair: There is a problem and it has been manually repaired. Adding this status helps the subsequent diagnosis module to determine multiple components or noise;

[0111] (4) Problem instability: The unstable status helps the diagnosis module to determine difficult judgments such as hard disk links;

[0112] (5) Fault reproduction: After the problem is manually repaired and occurs again, this problem helps the diagnostic module to make difficult judgments such as CPU problems and hard disk links.

[0113] The fault points are generally not isolated and point-like, but have a context. Therefore, in addition to the change in the state of the component itself, the context will also be considered. Thus, in another specific embodiment, regarding the recording of changes in the states of other factors of the component, for the server log, factors of automatic intervention and human intervention are mainly considered. Automatic intervention: Usually manifested as the system repairs the problem through self-intervention, such as OS (Operating System) restart, BMC (Baseboard Management Controller) restart, etc. Human intervention: Usually manifested as the problem is repaired by means of human intervention, and generally judged through the combination of composite states, typical combinations of several state changes. The following are several exemplary state changes of human intervention provided:

[0114] (1) System shutdown > BMC restart > System startup, generally indicating that the server has performed actions of manual power-off and power-on, and this state is mostly applicable to operations such as memory, CPU, and network card that require power-off;

[0115] (2) System shutdown > Chassis cover open > Chassis cover closed > System startup, generally indicating power-on and power-off of the server without power-off, and this state is mostly applicable to operations such as hard disks that do not require power-off;

[0116] (3) System startup > Chassis cover open > Chassis cover closed > System startup, generally indicating that the fault point is viewed and confirmed visually, and this state is generally applicable to replacement operations of special components such as fans.

[0117] Finally, combine the self-changes and other factor changes of the component to form a complete and composite component state change, which can truly reflect the actual state change of the component. The following are several exemplary composite state changes provided:

[0118] (1) Component causes downtime: Problem inserted > Automatic intervention > Problem resolved, indicating that it is corrected by means of system restart, and this problem affects the business and needs to be repaired;

[0119] (2) Component manually repaired: Problem inserted > Human intervention > Problem resolved, indicating that it is corrected by means of manual repair, and generally the problem is solved;

[0120] (3)Component problem reproduction: Problem insertion > Manual intervention > Problem resolution > Problem occurrence, which means fixing errors through manual repair. Generally, it indicates that the manual judgment is incorrect. This type of problem mostly presents as difficult and complex problems, resulting in multiple on-site visits and requires special attention.

[0121] Determine the status change details of the faulty component according to the composite state change to obtain the log details corresponding to the standardized log. That is, the log details record the change details of the component itself and other factors. It can be understood that the state change of the component itself can include problem insertion, problem resolution, in-place insertion, in-place resolution, etc.; the change of other factors can include server restart, power-on, power-off, chassis opening, model shutdown, BMC restart, BMC version change, etc.

[0122] It can be seen that this embodiment is not limited to the screening of log dot-like fault keywords, but needs to perform flow analysis on the log based on the continuous change characteristics of the log, identify the continuous change of the component's own state in the log and the association and influence of other factors on the component's state, and accurately reflect the true state of the component.

[0123] Such as Figure 3 Shown is an overall implementation detail flow block diagram exemplarily provided according to the above embodiment. For the fault diagnosis model, a log flow analysis processing module and a data classification module are added. The log flow analysis processing module performs log flow analysis on the server log, summarizes and standardizes more than a dozen types of fault logs related to the server, and then classifies them through the data classification module to form accurate and standardized classified summary data. Let the fault diagnosis model or AI algorithm model focus on the development of diagnostic rules or algorithm training without having to worry about intricate data processing. For the more specific processing process, reference can be made to the corresponding content disclosed in the foregoing embodiment, and details will not be elaborated here.

[0124] Exemplarily, the following uses 3 log implementation cases to illustrate the change of the component's own state and the assistance of other factors such as server power-on / off and chassis opening in judging the component state change.

[0125] Such as Figure 4 Shown is a data flow schematic diagram in a fault log for determining the component state through its own state change. The specific processing process is as follows:

[0126] 1. According to the fault keyword matching check, it is found that the front mylar of hard disk slot 0 has a fault at 04:25:24 on September 25, 2024 (keyword Front DISK0 Hard Disk Drive Detected Fault Assert), and the state of DISK0 is initialized as faulty;

[0127] 2. The status of hard disk DISK0 has changed. DISK0 was removed at 20:48:06 on October 10, 2024 (keyword: FrontDISK0 Hard Disk Drive Presence Deassert);

[0128] 3. The status of hard disk DISK0 has changed. DISK0 was reinserted at 20:49:42 on October 10, 2024 (keyword: FrontDISK0 Hard Disk Drive Presence Assert);

[0129] 4. The status of hard disk DISK0 has changed. The problem of DISK0 was fixed at 20:49:58 on October 10, 2024 (keyword: FrontDISK0 Hard Disk Drive Detected Fault Deassert);

[0130] In summary, it is recorded that the status of DISK0 itself has changed 4 times, undergoing 4 state transitions: hard disk failure > hard disk removal > hard disk insertion > hard disk normal. Based on this state transition, it can be concluded that although hard disk 0 had a fault, this hard disk has been repaired and its current status is normal. The process of analyzing the log of this component is as Figure 5 shown.

[0131] As Figure 6 shown is a schematic diagram of the data flow in a fault log for another method of determining the component status through its own status change. The specific processing process is as follows:

[0132] 1. According to the fault keyword matching, it is detected that DISK5 had a fault at 10:39:04 on April 22, 2023 (keyword: Front DISK5 Hard Disk Drive Detected Fault Assert), and the status of DISK5 was initialized as faulty;

[0133] 2. The status of hard disk DISK5 has changed. DISK5 was removed at 11:36:13 on April 23, 2024 (keyword: FrontDISK5 Hard Disk Drive Presence Deassert);

[0134] 3. The status of hard disk DISK5 has changed. DISK5 was reinserted at 11:38:42 on April 23, 2024 (keyword: FrontDISK5 Hard Disk Drive Presence Assert);

[0135] 4. The status of hard disk DISK5 has changed. The problem of DISK5 was fixed at 11:39:04 on April 23, 2024 (keyword: FrontDISK5 Hard Disk Drive Detected Fault Deassert);

[0136] 5. The status of hard disk DISK5 has changed. DISK5 had a fault at 07:17:55 on October 21, 2024 (keyword: FrontDISK5 Hard Disk Drive Detected Fault Assert);

[0137] In summary, the status bits of DISK5 have undergone 5 state transitions: hard disk failure > hard disk removed > hard disk inserted > hard disk normal > problem recurrence. Based on these state transitions, it can be concluded that there is a problem recurrence with hard disk 5, rather than just a simple fault. This type of problem usually means complex issues that require multiple on-site work orders and need to be closely monitored. The process of analyzing the log of this component is as Figure 7 shown. It can be seen that before using the present invention, only the fault of hard disk DISK5 could be diagnosed from this log. However, after using the present invention, it is possible to identify the fault of hard disk DISK5 and its recurrence after the fault is repaired. The state of hard disk recurrence here is extremely helpful for analyzing difficult problems with hard disks.

[0138] As Figure 8 shown is a schematic diagram of the data flow in a fault log for determining the status of a component through its own status change and other factor changes. The specific processing process is as follows:

[0139] 1. According to the matching of fault keywords, it is checked that the memory slot CPU0_C3D0 had a fault at 10:26:53 on October 27, 2024 (CPU0_C3D0 Uncorrectable ECC Assert), and the status of memory CPU0_C3D0 was initialized to faulty;

[0140] 2. For other relevant status changes, the operating system was shut down at 13:18:55 on October 27, 2024 (keyword: OSGraceful Shutdown);

[0141] 3. The status of memory CPU0_C3D has changed. The fault of memory CPU0_C3D0 was resolved at 13:18:56 on October 27, 2024 (keyword: CPU0_C3D0 Uncorrectable ECC Deassert);

[0142] 4. For other relevant status changes, the server performed a shutdown operation at 13:18:58 on October 31, 2024 (ACPI_PWRoff);

[0143] Summarizing the above, the recorded states of memory CPU0_C3D0 are in sequence: fault occurrence > system crash > memory fault resolved > server shutdown. The process of analyzing the logs of this component is as Figure 9 shown. It should be noted that the memory fault was resolved and the system was shut down at the same moment (with a 2-second interval). Considering this factor, the current state of the memory is faulty and has caused a system crash, rather than the fault being resolved.

[0144] As Figure 10 shown is an exemplary diagnostic screenshot generated after applying the present invention to a diagnostic model and then applying the diagnostic model to each operation and maintenance platform or service tool. Figure 11 shown is a schematic diagram of the log source corresponding to the diagnostic result.

[0145] Based on the above embodiments, this embodiment will specifically elaborate on S13 in the above embodiments. Among them, when the classified log obtained by partitioning is an out-of-band general log, in the first specific implementation manner, the out-of-band general log may include device fault diagnosis logs, system event logs, sensor logs, and black box logs. The process of analyzing and processing the out-of-band general log to obtain the corresponding target summary data may include the following steps:

[0146] Step 1: Send the device fault diagnosis log into the summary data to determine the first out-of-band general summary data;

[0147] Step 2: Compare the system event log with the first out-of-band general summary data to receive the target system event log in the system event log and obtain the second out-of-band general summary data; among them, the target system event log is the system event log that does not have information with the same component slot and problem status after being compared with the first out-of-band general summary data;

[0148] Step 3: Compare the black box log with the second out-of-band general summary data to receive the target black box log in the black box log and obtain the third out-of-band general summary data; among them, the target black box log is the black box log that does not have information with the same component slot and problem status after being compared with the second out-of-band general summary data;

[0149] Step 4: Compare the sensor log with the third out-of-band general summary data to receive the target sensor log in the sensor log and obtain the fourth out-of-band general summary data; among them, the target sensor log is the sensor log that does not have information with the same component slot and problem status after being compared with the third out-of-band general summary data;

[0150] Step 5: Perform status correction on the fourth out-of-band general summary data according to the current status of the components recorded in the sensor log to obtain the target summary data corresponding to the out-of-band general log.

[0151] It should be noted that, in a specific implementation manner, the device fault diagnosis log can be an IDL log. The IDL log is a server diagnosis log specific to a certain type of server. This log records the event history based on IPMI sensors and corresponds to the System Event Log (SEL), providing more comprehensive information. Each log contains processing suggestions and operation steps. Users can access these logs through the BMC Web interface and download or clear them as needed. The System Event Log (SEL) records server hardware status, error, and alert events, including hardware failures, system startup and shutdown, sensor status, etc. It is very important for server fault diagnosis and maintenance. The Sensor log records the real-time data of internal sensors of the server, such as temperature, voltage, etc., for monitoring the health status and performance of the server. The Blackbox log records the operation information of the server in the case of power failure, including power status, system configuration, etc., which helps to analyze the cause of problems when restoring the system after a power failure.

[0152] The following will be combined with Figure 12 to specifically illustrate the above embodiments. Traverse the standardized logs. First, send the standardized data parsed from the device fault diagnosis log into the summary data; then view the standardized data of the SEL log. If the slot or the type in the same state is not included in the summary data, receive and add this data to the summary data. Otherwise, discard this data; then view the standardized data of the Blackbox log. If the slot or the type in the same state is not included in the summary data, receive and add this data to the summary data. Otherwise, discard this data. Then view the standardized data of the sensor log. If the slot or the type in the same state is not included in the summary data, receive and add this data to the summary data. Otherwise, discard this data. It can be understood that since the components recorded in the sensor log are in the current state, the summary data is correspondingly corrected according to the current state of this component in the sensor log.

[0153] In the second specific implementation manner, when the classified log obtained by division is a storage specialty log, the storage specialty log can include independent redundant disk array logs, self-monitoring, analysis, and reporting technology logs, physical drive information logs, and first message logs; the first message log is a log that records the message information of each device used for storage. The process of analyzing and processing the storage specialty log to obtain the corresponding target summary data can include the following steps:

[0154] Step 1: Send the physical drive information log into the summary data to determine the first storage specialty summary data;

[0155] Step 2: Compare the redundant array of independent disks (RAID) log with the first storage specialty summary data to receive the target RAID log in the RAID log and obtain the second storage specialty summary data; among them, the target RAID log is the RAID log that does not have information with the same component slot and problem status after being compared with the first storage specialty summary data;

[0156] Step 3: Compare the first message log with the second storage specialty summary data to receive the target first message log in the first message log and obtain the third storage specialty summary data; among them, the target first message log is the first message log that does not have information with the same component slot and problem status after being compared with the second storage specialty summary data;

[0157] Step 4: Compare the self - monitoring, analysis, and reporting technology (SMART) log with the third storage specialty summary data to receive the target SMART log in the SMART log and obtain the fourth storage specialty summary data; among them, the target SMART log is the SMART log that does not have information with the same component slot and problem status after being compared with the third storage specialty summary data;

[0158] Step 5: Perform status correction on the fourth storage specialty summary data according to the current status of the components recorded in the SMART log to obtain the target summary data corresponding to the storage specialty log.

[0159] It should be noted that the redundant array of independent disks (RAID) log records data changes, events, and operations in the RAID system, such as data write, modification, deletion, etc. operations to ensure data consistency and integrity. The self - monitoring, analysis, and reporting technology (SMART) log records the health status information of the device, including key warning information such as available space, temperature threshold, reliability degradation, read - only status, and volatile memory device failure. The SMART log also records data read - write operations and host commands, which helps to analyze the operating status and potential failures of the device. The physical drive information log (Physical Drive Information, PdInfo) provides detailed information about physical devices (such as hard disks), including the version of the hard disk, RAID group, and hard disk information. These logs are usually used to help administrators diagnose hardware failures or optimize storage configurations. The first message log (Message log) usually records various messages in the system. The first message log in this embodiment is mainly the log related to the error reporting part of the RAID card and hard disk.

[0160] The following will be combined with Figure 13 to specifically describe the above embodiments. Traverse the standardized logs. First, send the standardized data after parsing the PdInfo log into the summary data; then view the RAID log standardized data. If the slot or the type under the same status is not included in the summary data, receive and add this piece of data to the summary data. Otherwise, discard this data; then view the Message log standardized data. If the slot or the type under the same status is not included in the summary data, receive and add this piece of data to the summary data. Otherwise, discard this data. Then view; then view the Smart log standardized data. If the slot or the type under the same status is not included in the summary data, receive and add this piece of data to the summary data. Otherwise, discard this data. Then view. It can be understood that since the components recorded in Smart are in the current state, the summary data is correspondingly corrected according to the current state of this component in the Smart log.

[0161] In the third specific implementation manner, when the classified log obtained by division is a computing specialty log, the computing specialty log may include an internal error log, a maintenance log, a process analysis log, and a second message log; the second message log is a log for recording the message information of each device used for computing. The process of analyzing and processing the computing specialty log to obtain the corresponding target summary data may include the following steps:

[0162] Step 1: Send the second message log into the summary data to determine the first computing specialty summary data;

[0163] Step 2: Compare the process analysis log with the first computing specialty summary data to receive the target process analysis log in the process analysis log and obtain the second computing specialty summary data; wherein, the target process analysis log is the process analysis log that does not have information with the same component slot and problem status after being compared with the first computing specialty summary data;

[0164] Step 3: Compare the internal error log with the second computing specialty summary data to receive the target internal error log in the internal error log and obtain the third computing specialty summary data; wherein, the target internal error log is the internal error log that does not have information with the same component slot and problem status after being compared with the second computing specialty summary data;

[0165] Step 4: Compare the maintenance log with the third computing specialty summary data to receive the target maintenance log in the maintenance log and obtain the fourth computing specialty summary data; wherein, the target maintenance log is the maintenance log that does not have information with the same component slot and problem status after being compared with the third computing specialty summary data;

[0166] Step 5: Perform status correction on the fourth calculation specialty summary data according to the current status of the components recorded in the maintenance log to obtain the target summary data corresponding to the calculation specialty log.

[0167] It should be noted that the internal error log (Internal Error, IERR) records error information that occurs during the operation of the system. This type of log is mainly used for troubleshooting and repair to help system administrators identify and solve problems during operation. The maintenance log (Maintenance) is mainly used to record operations and events related to system maintenance, and may also include information such as the maintenance plan of the device, the execution status, and the start and end times. The process analysis log (AnalyProcess) usually refers to the log that records process activities, such as user logins, application runs, system service startups, etc. This type of log helps monitor and analyze the running status of the system. The second message log (Message log) usually records various messages in the system. The second message log in this embodiment is mainly the log related to the error reporting parts of the CPU, memory, and PCIE.

[0168] The following will be combined with Figure 14 Specifically illustrate the above embodiments. Traverse the standardized log. First, send the standardized data parsed from the second message log into the summary data; then view the standardized data of the AnalyProcess log. If the slot or the type in the same state is not included in the summary data, receive and add this piece of data to the summary data, otherwise discard this data; then view the standardized data of the IERR log. If the slot or the type in the same state is not included in the summary data, receive and add this piece of data to the summary data, otherwise discard this data and then view; then view the standardized data of the Maintenance log. If the slot or the type in the same state is not included in the summary data, receive and add this piece of data to the summary data, otherwise discard this data and then view. It can be understood that since the components recorded in the Maintenance log are in the current state, the summary data is corrected accordingly according to the current state of this component in the Maintenance log.

[0169] The following uses a specific example of forming a classification log from the log for illustration. As Figure 15 shown, the sample therein includes four types of standardized logs: IDL, SEL, PdInfo, and RAID. When generating the corresponding summary data for the out-of-band general data in Figure 15 , before processing, as Figure 16As shown. The processing process is as follows: The memory CPU0_C0D0 data of IDL is recorded; The memory CPU0_C0D0 data of SEL is discarded; The hard disk DISK32 data of SEL is recorded; The hard disk DISK32 data of SDR is discarded, but the status of DISK32 is corrected to the current problem. The processed data finally obtained is as Figure 17 shown. For Figure 15 the storage specialist data in Figure 18 when generating the corresponding summary data, before processing it is as Figure 19 shown. The processing process is as follows: The hard disk slot8 data of PdInfo is recorded; The hard disk slot8 of the Raid card is discarded. The processed data finally obtained is as

[0170] It can be seen that before using the present invention, sending multiple problems to the diagnostic model will cause the diagnostic model to mistakenly think there are multiple hard disk and multiple memory problems, causing diagnostic troubles. After using the present invention, through classification and filtering, the diagnostic model can clearly judge a single hard disk failure of DISK32.

[0171] It should be noted that the above processing processes for out-of-band general logs, storage specialist logs, and computing specialist logs are described in a sequential processing manner. Similarly, parallel processing and multi-threaded technologies can also be used to process the currently classified fault logs to further improve the efficiency of log processing. During the parallel processing process, fault switching and load balancing functions can be provided to achieve parallel logging among multiple processors, thereby improving the processing speed. In addition, in a feasible implementation manner, incomplete or incorrect log data can also be processed through automatic recovery and cleaning strategies in parallel job scheduling. For example, a monitoring and debugging tool can be developed to perform fault troubleshooting and performance optimization according to monitoring metrics. The data cleaning strategy therein can remove the influence of abnormal behaviors and avoid abnormal data from contaminating the log quality.

[0172] As Figure 20The following is a schematic diagram of an overall implementation process provided by way of example in combination with the foregoing embodiments. After the logs are collected, the logs are integrated, decompressed, and classified to obtain different log classifications (such as out-of-band structured data such as equipment fault diagnosis and system event logs; out-of-band unstructured data such as black boxes and registers; in-band structured data such as GPUs and RAIDs; in-band unstructured data such as Messages and Smarts). Through fuzzy query, check whether the logs of the current log type can be matched. If no valid data is matched, give a reminder and record it. After the fault logs are standardized, record the abnormal change flow of the component status for the problem status of the dynamic parameters in the target parameters. Among them, for a more specific working process regarding the problem status change, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated here. Next, the standardized data is integrated, de-duplicated, and classified to obtain out-of-band general data, computing specialist data, storage specialist data, power supply specialist data, GPU specialist data, etc. Finally, based on the summary data obtained after classification, perform a fault decision to find the root cause of the fault.

[0173] Correspondingly, the embodiment of the present application also discloses a fault log processing device, see Figure 21 As shown, the device includes:

[0174] A standardization processing module 11, configured to obtain different types of fault logs and perform standardization processing on the fault logs to generate standardized logs;

[0175] A flow analysis module 12, configured to analyze the log flow in the standardized logs to determine the faulty components and record the status change details of the faulty components to generate log details corresponding to the standardized logs;

[0176] A data classification module 13, configured to divide the standardized logs according to a preset division rule to generate different types of classified logs, and respectively analyze and process the different types of classified logs to obtain their respective corresponding target summary data;

[0177] A fault diagnosis module 14, configured to perform fault diagnosis based on the log details and the target summary data.

[0178] Among them, for a more specific working process of each of the above modules, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated here.

[0179] It can be seen that through the above solution of this embodiment, it includes: obtaining different types of fault logs, and performing standardization processing on the fault logs to generate standardized logs; analyzing the log streams in the standardized logs to determine faulty components, and recording the detailed status changes of the faulty components to generate log details corresponding to the standardized logs; dividing the standardized logs according to preset division rules to generate different types of classified logs, and respectively analyzing and processing the different types of classified logs to obtain their respective corresponding target summary data; performing fault diagnosis based on the log details and the target summary data.

[0180] The beneficial technical effects of this application are as follows: by incorporating as many different types of fault logs as possible into the analysis, first generating standardized logs with a standard structure; then performing log stream analysis on the standardized logs, not limited to screening for log point-like fault keywords, being able to identify the status of faulty components in the logs based on the continuous change characteristics of the logs, and the generated log details accurately reflecting the true status of the components; further, classifying the standardized logs according to preset division rules and obtaining the target summary data corresponding to different types of classified logs. Therefore, each type of target summary data is not limited to a single type of fault log; finally, the generated log details and target summary data can be called by the diagnostic model for fault diagnosis, realizing data decoupling for the diagnostic model. In this way, the fault diagnosis model can focus on diagnostic rule development or algorithm training without having to concern itself with intricate data processing.

[0181] Furthermore, the embodiment of this application also discloses an electronic device Figure 22 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be considered as any limitation to the scope of use of this application.

[0182] Figure 22 It is a schematic structural diagram of an electronic device 20 provided by an embodiment of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the fault log processing method disclosed in any of the foregoing embodiments. Additionally, the electronic device 20 in this embodiment may specifically be a server.

[0183] In this embodiment, the power supply 23 is used to provide operating voltages for the various hardware devices on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and specific limitations thereof are not provided herein; the input / output interface 25 is used to obtain external input data or output data to the outside, and the specific interface type can be selected according to specific application requirements, and specific limitations thereof are not provided herein.

[0184] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, a random access memory, a magnetic disk, an optical disk, etc., and the resources stored thereon can include an operating system 221, a computer program 222, and data 223, etc. The data 223 can include various types of data. The storage method can be temporary storage or permanent storage.

[0185] Among them, the operating system 221 is used to manage and control the various hardware devices and the computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the fault log processing method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 can further include computer programs that can be used to complete other specific tasks.

[0186] Furthermore, the embodiment of the present application also discloses a computer-readable storage medium, and the computer-readable storage medium herein includes a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a magnetic disk, an optical disk, or any other form of storage medium known in the technical field. Among them, when the computer program is executed by a processor, the foregoing fault log processing method is implemented. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details are not described herein again.

[0187] Furthermore, the embodiment of the present application also provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, any implementation method of the foregoing fault log processing method is implemented.

[0188] The various embodiments in this specification are described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method part.

[0189] The steps of the fault log processing method or algorithm described in combination with the embodiments disclosed in this article can be implemented directly by hardware, software modules executed by a processor, or a combination of both. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0190] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0191] The above has introduced in detail a fault log processing method, device, program product and medium provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for processing fault logs, characterized in that, including: Obtain different types of fault logs, and perform standardization processing on the fault logs to generate standardized logs; Analyze the log stream in the standardized logs to determine the faulty components, and record the detailed status changes of the faulty components to generate log details corresponding to the standardized logs; wherein, the detailed status changes are: the problem status change records of the faulty components over time; the analysis process of the log stream is used to identify the continuous changes in the status of the components in the logs and the association and influence of other factors on the component status; Divide the standardized logs according to preset division rules to generate different types of classified logs, and respectively perform analysis processing on different types of the classified logs to obtain respective corresponding target summary data; wherein, the analysis processing process includes removing duplicate data in the classified logs; Perform fault diagnosis based on the log details and the target summary data.

2. The fault log processing method according to claim 1, wherein The obtaining different types of fault logs and performing standardization processing on the fault logs to generate standardized logs includes: Obtain different types of fault logs, and perform fuzzy query on the fault logs according to the log names to obtain multiple log classifications; Respectively perform structure recognition and parsing on the fault logs in each log classification to obtain corresponding target parameters, and based on the target parameters, perform standardization processing on each line of data in the log stream corresponding to the fault logs to generate standardized logs under each log classification.

3. The fault log processing method according to claim 2, wherein The target parameters include static parameters and dynamic parameters; the static parameters include component types and component slots, and the dynamic parameters include occurrence time, fault keywords, fault levels, problem status, and component information; Wherein, the analyzing the log stream in the standardized logs to determine the faulty components and recording the detailed status changes of the faulty components to generate log details corresponding to the standardized logs includes: Use the fault keywords to match in the log stream in the standardized logs to determine the faulty components, and record the detailed status changes of the faulty components according to the problem status to generate log details corresponding to the standardized logs.

4. The fault log processing method according to claim 1, wherein The recording the detailed status changes of the faulty components to generate log details corresponding to the standardized logs includes: Respectively obtain the self-status changes of the faulty components and the status changes of other factors associated with the faulty components; wherein, the status changes of other factors are the status changes of factors with a context relationship with the faulty components in the fault logs; Combine the self-status changes and the status changes of other factors to generate composite status changes, and determine the detailed status changes of the faulty components according to the composite status changes to obtain log details corresponding to the standardized logs.

5. The fault log processing method according to claim 1, characterized in that The dividing the standardized logs according to preset division rules to generate different types of classified logs includes: According to the fault information recorded in the standardized log, divide the standardized log into out-of-band general logs, storage specialty logs, and computing specialty logs; wherein, the out-of-band general logs contain the fault information of each component, the storage specialty logs contain the fault information of each device for storage, and the computing specialty logs contain the fault information of each device for computing.

6. The fault log processing method according to claim 5, wherein, The out-of-band general logs include device fault diagnosis logs, system event logs, sensor logs, and black box logs.

7. The fault log processing method according to claim 6, wherein Analyze and process the out-of-band general logs to obtain corresponding target summary data, including: Send the device fault diagnosis logs into the summary data to determine the first out-of-band general summary data; Compare the system event logs with the first out-of-band general summary data to receive the target system event logs in the system event logs and obtain the second out-of-band general summary data; wherein, the target system event logs are the system event logs that do not have information with the same component slot and problem status after being compared with the first out-of-band general summary data; Compare the black box logs with the second out-of-band general summary data to receive the target black box logs in the black box logs and obtain the third out-of-band general summary data; wherein, the target black box logs are the black box logs that do not have information with the same component slot and problem status after being compared with the second out-of-band general summary data; Compare the sensor logs with the third out-of-band general summary data to receive the target sensor logs in the sensor logs and obtain the fourth out-of-band general summary data; wherein, the target sensor logs are the sensor logs that do not have information with the same component slot and problem status after being compared with the third out-of-band general summary data; Perform status correction on the fourth out-of-band general summary data according to the current status of the components recorded in the sensor logs to obtain the target summary data corresponding to the out-of-band general logs.

8. The fault log processing method according to claim 5, wherein, The storage specialty logs include: redundant array of independent disks logs, self-monitoring, analysis, and reporting technology logs, physical drive information logs, and first message logs; the first message logs are the logs that record the message information of each device for storage.

9. The fault log processing method according to claim 8, wherein Analyze and process the storage specialty logs to obtain corresponding target summary data, including: Send the physical drive information logs into the summary data to determine the first storage specialty summary data; Compare the redundant array of independent disks logs with the first storage specialty summary data to receive the target redundant array of independent disks logs in the redundant array of independent disks logs and obtain the second storage specialty summary data; wherein, the target redundant array of independent disks logs are the redundant array of independent disks logs that do not have information with the same component slot and problem status after being compared with the first storage specialty summary data; Compare the first message log with the second storage specialist summary data to receive the target first message log in the first message log, and obtain the third storage specialist summary data; wherein, the target first message log is the first message log that has no information with the same component slot and problem status after being compared with the second storage specialist summary data; Compare the self-monitoring, analysis, and reporting technology log with the third storage specialist summary data to receive the target self-monitoring, analysis, and reporting technology log in the self-monitoring, analysis, and reporting technology log, and obtain the fourth storage specialist summary data; wherein, the target self-monitoring, analysis, and reporting technology log is the self-monitoring, analysis, and reporting technology log that has no information with the same component slot and problem status after being compared with the third storage specialist summary data; Perform status correction on the fourth storage specialist summary data according to the current status of the components recorded in the self-monitoring, analysis, and reporting technology log to obtain the target summary data corresponding to the storage specialist log.

10. The fault log processing method according to claim 5, wherein The computing specialist log includes: internal error log, maintenance log, process analysis log, and second message log; the second message log is a log that records the message information of each device used for computing.

11. The fault log processing method according to claim 10, characterized in that, Analyze and process the computing specialist log to obtain the corresponding target summary data, including: Send the second message log into the summary data to determine the first computing specialist summary data; Compare the process analysis log with the first computing specialist summary data to receive the target process analysis log in the process analysis log, and obtain the second computing specialist summary data; wherein, the target process analysis log is the process analysis log that has no information with the same component slot and problem status after being compared with the first computing specialist summary data; Compare the internal error log with the second computing specialist summary data to receive the target internal error log in the internal error log, and obtain the third computing specialist summary data; wherein, the target internal error log is the internal error log that has no information with the same component slot and problem status after being compared with the second computing specialist summary data; Compare the maintenance log with the third computing specialist summary data to receive the target maintenance log in the maintenance log, and obtain the fourth computing specialist summary data; wherein, the target maintenance log is the maintenance log that has no information with the same component slot and problem status after being compared with the third computing specialist summary data; Perform status correction on the fourth computing specialist summary data according to the current status of the components recorded in the maintenance log to obtain the target summary data corresponding to the computing specialist log.

12. The fault log processing method according to any one of claims 1 to 11, characterized in that, The fault diagnosis based on the log details and the target summary data includes: Obtain the first diagnostic rule and the second diagnostic rule respectively, and combine the first diagnostic rule and the second diagnostic rule to obtain a target decision rule; wherein, the first diagnostic rule is a diagnostic rule obtained by sorting out and solidifying the rules of existing fault diagnosis instances, and the second diagnostic rule is a diagnostic rule obtained by matching the rules of fault data through an artificial intelligence algorithm; Based on the log details and the target summary data, use the target decision rule to perform fault diagnosis and output the root cause of the fault of the faulty component.

13. An electronic device, characterized in that, It includes: A memory for storing a computer program; A processor for loading and executing the computer program to implement the fault log processing method according to any one of claims 1 to 12.

14. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the fault log processing method according to any one of claims 1 to 12 are implemented.

15. A computer-readable storage medium, characterized in that, For storing a computer program; wherein when the computer program is executed by the processor, the fault log processing method according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • System fault diagnosis method and device, equipment and storage medium

    CN116126574A

  • Fault diagnosis data labeling method and device

    CN119202708A