Server failure root cause analysis method, device, electronic device, and storage medium
By collecting and analyzing alarm information in the server, using level judgment and intellectual inference to build a causal relationship chain, combined with expert system review, the problem of root cause positioning of faults under multiple alarms on the server is solved, and fast and accurate fault source positioning and efficient operation and maintenance are achieved.
Patent Information
- Application Number
- CN202510750916.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-06-06
AI Technical Summary
In the prior art, it is difficult for servers to quickly and accurately locate the root cause of the failure in multiple interrelated alarm information scenarios, resulting in operation and maintenance difficulties and business interruptions.
By collecting server alarm information and log data regularly, using level judgment mechanisms and intellectual inference, an alarm causal relationship chain is automatically built, and reviewed in combination with historical case libraries to form a dual verification mechanism for AI analysis and expert systems, accurately locate the root cause of the fault and generate repair suggestions.
It realizes automated root cause analysis in multiple alarm information scenarios, quickly and accurately locates fault sources, improves fault diagnosis efficiency, reduces operation and maintenance complexity, and ensures operation and maintenance reliability and efficiency.
Smart Images

Figure CN120276907B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of server technology, and in particular to a method, device, electronic device, and storage medium for analyzing the root cause of server failures. Background Art
[0002] During daily operation, servers generate a large amount of diverse alarm information, including hardware failures, performance bottlenecks, network anomalies, and other types. The same server may generate multiple interrelated alarms in a short period of time, and these alarms may have complex causal relationships.
[0003] When encountering multiple interrelated alarm messages, there is a lack of an effective alarm correlation analysis mechanism, making it difficult to identify key clues from the complex alarm information. When a failure occurs, the operation and maintenance team is only busy dealing with the surface symptoms of the server, and it is difficult to quickly and accurately locate the root cause of the failure, resulting in serious problems such as server operation and business interruption. Summary of the Invention
[0004] The present application provides a server fault root cause analysis method, device, electronic device and storage medium to at least solve the problem in related technologies that it is difficult to quickly and accurately locate the root cause of the fault.
[0005] This application provides a server failure root cause analysis method, including:
[0006] Regularly collect server alarm information and server log data, and determine whether the alarm information belongs to the first level or the second level;
[0007] When the alarm information belongs to the first level, the root cause analysis and repair suggestions are directly provided based on the first level information;
[0008] When the alarm information belongs to the second level, intelligent agent reasoning is performed based on the alarm information and the second level information, and the root cause analysis and repair suggestions are provided based on the intelligent agent reasoning results. At the same time, the root cause analysis and repair suggestions are reviewed, and the final root cause analysis results and final repair suggestions are output.
[0009] Through this application, server log data and alarm information are collected regularly, and root cause analysis is performed in the case of multiple server alarms based on the level of the alarm information. A causal relationship chain between alarms is automatically established, the root cause fault component is accurately and quickly located, and repair suggestions are generated, which greatly improves the efficiency of fault diagnosis. The root cause analysis and repair suggestions are reviewed to form a dual verification mechanism of AI analysis and secondary review to avoid single AI diagnosis misjudgment. Therefore, it effectively solves the problem in related technologies that it is difficult to quickly and accurately locate the root cause of the fault, and achieves the effect of automated root cause analysis of the server in multiple alarm information scenarios and accurate location of the fault source.
[0010] In an optional embodiment, the first level is a single first fault segment dimension; the server log includes multiple component alarm logs;
[0011] When the alarm information is classified as level 1, the root cause analysis and repair suggestions are directly provided based on the level 1 information, including:
[0012] When the alarm information is a single first fault segment dimension, determining the root cause component corresponding to the single first fault segment dimension, and obtaining the first component alarm log corresponding to the root cause component from multiple component alarm logs;
[0013] Directly locate the root cause of the fault and provide repair suggestions based on the alarm log corresponding to the root cause component.
[0014] Through this application, when an alarm message of a single first fault segment dimension is received, the corresponding root cause component can be directly determined without the need for complex multi-source data association analysis or intelligent agent reasoning process. Therefore, this "direct mapping" processing method greatly shortens the fault diagnosis time, can start the repair process in the shortest time, quickly restore the normal operation of the server, and reduce the duration of business interruption caused by the fault.
[0015] In an optional embodiment, the second level includes multiple first fault segment dimensions, a single second fault segment dimension and multiple second fault segment dimensions, or mixed alarms of multiple first fault segment dimensions and at least one second fault segment dimension; the server log also includes a full log;
[0016] When the alarm information belongs to the second level, the intelligent agent performs reasoning based on the alarm information and the second level information, and performs fault root cause analysis and repair suggestions based on the intelligent agent reasoning results, including:
[0017] When the alarm information has multiple first fault segment dimensions, first agent reasoning is performed based on multiple component alarm logs, pre-acquired device serial numbers, and alarm times of multiple first fault segment dimensions to obtain first agent reasoning results. Based on the first agent reasoning results, a first fault root cause analysis and a first repair suggestion are performed.
[0018] When the alarm information is a single second fault segment dimension, determine the root cause component corresponding to the single second fault segment dimension, obtain server performance data and the second component alarm log corresponding to the root cause component, perform second agent reasoning based on the second component alarm log, server performance data, full log, and the pre-acquired device serial number and the alarm time of the single second fault segment dimension, and perform a second fault root cause analysis and provide a second repair suggestion based on the second agent reasoning result;
[0019] When the alarm information is a mixed alarm of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, server performance data is obtained, and third-agent reasoning is performed based on multiple component alarm logs, server performance data, full logs, and pre-acquired equipment serial numbers, multiple second fault segment dimensions or the alarm time of the mixed alarm, and a third fault root cause analysis and third repair suggestions are provided based on the third-agent reasoning results.
[0020] This application customizes differentiated agent reasoning strategies for different alarm scenarios (multiple first fault segment dimensions, a single second fault segment dimension, multiple second fault segment dimensions, or mixed alarms). For multiple first fault segment dimension alarms, the system focuses on component alarm logs and basic information reasoning to quickly resolve common faults. For single second fault segment dimension alarms, in-depth analysis is performed using performance data and full logs. In mixed alarm scenarios, comprehensive multi-source data reasoning is integrated to accurately locate the root cause of even complex faults, significantly improving adaptability and diagnostic accuracy across various fault scenarios.
[0021] In an optional embodiment, when the alarm information includes multiple first fault segment dimensions, first agent reasoning is performed based on multiple component alarm logs, pre-acquired device serial numbers, and alarm times of the multiple first fault segment dimensions to obtain a first agent reasoning result. Based on the first agent reasoning result, a first fault root cause analysis and a first repair suggestion are performed, including:
[0022] When the alarm information is multiple first fault segment dimensions, an improved constraint-based non-parametric algorithm is used to construct a first reasoning knowledge graph based on multiple component alarm logs, pre-acquired device serial numbers, and alarm times of multiple first fault segment dimensions, combined with a preset hardware dependency rule base;
[0023] Based on the first reasoning knowledge graph, the correlation and root causes of multiple first fault segment dimensions are searched, and the root causes of the faults are dynamically scored to obtain the first agent reasoning results of whether there is a unique root cause component or no root cause component;
[0024] When the first agent reasoning result is that there is a unique root cause component, performing a first fault root cause analysis on the unique root cause component and generating a first repair suggestion;
[0025] When the first intelligent agent reasoning result indicates that there is no root cause component, it is determined that multiple first fault segment dimensions are not correlated, corresponding first fault root cause analyses are performed on the multiple first fault segment dimensions respectively, and corresponding first repair suggestions are generated based on the corresponding first fault root cause analyses.
[0026] Through this application, for the alarm scenarios of multiple first fault segment dimensions, the preset hardware dependency rule base and the improved constraint-based non-parametric algorithm are combined to construct a first reasoning knowledge graph, which can quickly sort out the logical relationship between the alarm logs of multiple components. Through the existing hardware dependency rules in the rule base and the use of algorithm constraints, the scattered alarm information is efficiently converted into a structured knowledge graph, which intuitively presents the potential connections between multiple first fault segment dimension alarms, avoids the isolated analysis of each alarm, and greatly improves the efficiency of fault correlation mining. Based on the constructed knowledge graph, the root cause of the fault is found and dynamic scoring is implemented. Taking into account various factors such as the causal graph in-degree and historical matching, a quantitative evaluation is performed on each possible root cause of the fault. The multi-dimensional scoring mechanism can more comprehensively and accurately screen out the true root cause of the fault than a single standard judgment. Processing strategies are formulated for the two reasoning results: the existence of a unique root cause component and the absence of a root cause component. When there is a unique root cause component, it is directly analyzed in depth and repair suggestions are generated to quickly resolve the problem. When there is no root cause component, that is, multiple alarms are not related, each first fault segment dimension is analyzed independently to avoid mistakenly judging independent faults as cascading faults, ensuring that different fault scenarios can be handled reasonably and effectively.
[0027] In an optional embodiment, when the alarm information is a single second fault segment dimension, the root cause component corresponding to the single second fault segment dimension is determined, and the server performance data and the second component alarm log corresponding to the root cause component are obtained. Second agent reasoning is performed based on the second component alarm log, server performance data, full log, and the pre-acquired device serial number and the alarm time of the single second fault segment dimension. Based on the second agent reasoning result, a second fault root cause analysis and a second repair suggestion are performed, including:
[0028] When the alarm information is a single second fault segment dimension, determining the root cause component corresponding to the single second fault segment dimension based on the single second fault segment dimension, obtaining server performance data within a preset time period before the alarm information, and filtering out the second component alarm log corresponding to the root cause component from multiple component alarm logs;
[0029] Based on the second component alarm log, server performance data, full log, and pre-acquired device serial number and alarm time, and combined with the preset hardware dependency rule base, an improved constraint-based non-parametric algorithm is used to construct the second reasoning knowledge graph;
[0030] Based on the second reasoning knowledge graph, the direct root cause component and the correlation factor corresponding to the single second fault segment dimension are searched, and the second agent reasoning result of the presence of the direct root cause component or the presence of the correlation factor is obtained;
[0031] When the second agent reasoning result is for a direct root cause component, a second fault root cause analysis is performed on the direct root cause component and a second repair suggestion is generated; when the second agent reasoning result is for a correlation factor, a fault root cause analysis result and a second repair suggestion are generated based on the correlation factor.
[0032] Through this application, multi-source data such as server performance data within a preset time period before the alarm information, screened second component alarm logs, full logs, etc. are obtained, and combined with the device serial number and alarm time, the relevant information before and after the fault occurs is fully integrated. These data reflect the operating status of the server from different dimensions. By combining the preset hardware dependency rule base and the improved algorithm to build a knowledge graph, it is possible to deeply explore the potential connections between data, avoid missing key clues, and provide a rich and comprehensive data foundation for accurate fault analysis. For a single second fault segment dimension alarm, the constructed second reasoning knowledge graph is used to find the direct root cause components and correlation factors. When there is a direct root cause component, the source of the fault can be directly locked; if there are only correlation factors, it can also be comprehensively analyzed from aspects such as performance fluctuations and component associations to find the deep fault causes hidden under the surface phenomena, breaking the limitations of analyzing only a single alarm log and improving the diagnosis ability of complex faults.
[0033] In an optional embodiment, when the alarm information is a mixed alarm of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, server performance data is obtained, and third-agent reasoning is performed based on multiple component alarm logs, server performance data, full logs, and pre-acquired device serial numbers, multiple second fault segment dimensions, or the alarm time of the mixed alarm. Based on the third-agent reasoning results, a third fault root cause analysis and a third repair suggestion are performed, including:
[0034] When the alarm information is a mixed alarm of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, obtaining server performance data within a preset time period before the alarm information;
[0035] Based on multiple component alarm logs, server performance data, full logs, pre-acquired device serial numbers, multiple second fault segment dimensions, or alarm times of mixed alarms, and combined with a preset hardware dependency rule base, an improved constraint-based non-parametric algorithm is used to construct a third reasoning knowledge graph;
[0036] Based on the third reasoning knowledge graph, the correlation between multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension and the root cause of the fault are searched, and the third agent reasoning result of the existence of a unique root cause component, the existence of a root cause component, or the existence of a correlation factor is obtained;
[0037] When the third agent reasoning result is for a unique root cause component, performing a third fault root cause analysis on the unique root cause component and generating a third repair suggestion;
[0038] When the third agent reasoning result indicates that there is no root cause component, it is determined that the multiple second fault segment dimensions or the multiple first fault segment dimensions are not associated with the at least one second fault segment dimension, corresponding third fault root cause analyses are performed on the multiple second fault segment dimensions or the multiple first fault segment dimensions and the at least one second fault segment dimension, and corresponding third repair suggestions are generated based on the corresponding third fault root cause analyses;
[0039] When the third agent reasoning result has correlation factors, the correlation between multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension mixed alarm is found based on the correlation factors, and the correlation fault root cause analysis result and the third repair suggestion are generated based on the correlation.
[0040] Through this application, in the face of complex scenarios such as multiple second fault segment dimensions or mixed alarms, by integrating multiple component alarm logs, server performance data, full logs and other multi-source information, combined with the device serial number and alarm time, a third reasoning knowledge graph is constructed, which can effectively sort out the logical relationship between complex alarms. Based on the third reasoning knowledge graph, not only can the direct root cause components of the fault be found, but also the existing correlation factors can be excavated. By combining the preset hardware dependency rule base and the improved algorithm, the fault correlation is comprehensively analyzed from multiple dimensions, whether it is a single unique root cause component or a fault caused by the mutual correlation between multiple components, it can be accurately located. For situations where there is no obvious root cause component, the system can also analyze each alarm separately to prevent unrelated faults from being misjudged as cascading faults, greatly improving the accuracy of root cause location in complex fault scenarios.
[0041] In an optional implementation, the root cause analysis and repair suggestions are reviewed, and a final root cause analysis result and a final repair suggestion are output, including:
[0042] Based on the preset expert experience knowledge base and preset historical failure cases, the preset multi-dimensional check matrix and contradiction detection algorithm are used to review the root cause analysis and repair suggestions;
[0043] When a conflict occurs between the root cause analysis and the expert experience knowledge base during the review process, the preset priority level is used to determine whether to choose the root cause analysis or manual consultation method to resolve the conflict;
[0044] The confidence level of the fault root cause analysis is modified based on the conflict resolution result, the final fault root cause analysis result is output, and the final repair suggestion is generated based on the final fault root cause analysis result.
[0045] Through this application, based on the preset expert experience knowledge base and the preset structured fault diagnosis rule base, a preset multi-dimensional check matrix and contradiction detection algorithm are used to review the root cause analysis and repair suggestions of the fault, forming a dual verification mechanism of AI analysis + expert system review. It not only gives full play to the rapid analysis advantage of AI, but also ensures the accuracy of the diagnosis results through the expert system, effectively solving the possible misjudgment problem of a single AI system. This technology significantly reduces the complexity of operation and maintenance work, frees operation and maintenance personnel from tedious troubleshooting, and enables the data center operation and maintenance model to achieve a transformation from manual dominance to intelligent automation. While improving operation and maintenance reliability, it provides a strong guarantee for the efficient and stable operation of the data center.
[0046] The present application also provides a server failure root cause analysis device, comprising:
[0047] The alarm information collection and judgment module is used to regularly collect server alarm information and server log data, and judge whether the alarm information belongs to the first level or the second level;
[0048] The first fault root cause analysis module is used to directly perform fault root cause analysis and provide repair suggestions based on the first level information when the alarm information is the first level;
[0049] The second fault root cause analysis and review module is used to perform intelligent agent reasoning based on the alarm information and the second-level information when the alarm information is at the second level, and to perform fault root cause analysis and provide repair suggestions based on the intelligent agent reasoning results. At the same time, the fault root cause analysis and repair suggestions are reviewed, and the final fault root cause analysis results and final repair suggestions are output.
[0050] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned server failure root cause analysis methods when executing the computer program.
[0051] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned server failure root cause analysis methods are implemented.
[0052] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned server failure root cause analysis methods when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0054] Figure 1 A flowchart of a server failure root cause analysis method provided in an embodiment of the present application;
[0055] Figure 2 A flowchart of another server failure root cause analysis method provided in an embodiment of the present application;
[0056] Figure 3 A flowchart of another server failure root cause analysis method provided in an embodiment of the present application;
[0057] Figure 4 A flowchart of the server fault root cause analysis system provided in an embodiment of the present application;
[0058] Figure 5 A schematic diagram of conflict resolution by a fault diagnosis and review unit in a server fault root cause analysis system provided by an embodiment of the present application;
[0059] Figure 6 A structural block diagram of a server failure root cause analysis device provided in an embodiment of the present application;
[0060] Figure 7 Schematic diagram of the hardware structure of the electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0061] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0062] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0063] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0064] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the server failure root cause analysis method depends, the specific application environment architecture or specific hardware architecture is described herein.
[0065] An embodiment of the present application provides a server failure root cause analysis method, and the method is described in detail in conjunction with the execution process of the server failure root cause analysis method.
[0066] In this embodiment, a server fault root cause analysis method is provided, which can be used in a server fault root cause analysis system. The system includes an alarm analysis unit 1, an AI fault diagnosis agent unit 2, and a fault diagnosis review unit 3. The alarm analysis unit 1 is used to uniformly process server alarms; locate the root cause of the alarm fault and perform correlation analysis based on different alarm segment dimensions. The AI fault diagnosis agent unit 2 is responsible for in-depth analysis and intelligent reasoning of the alarm data. The fault diagnosis review unit 3 reviews the diagnosis results of the AI agent. Figure 1 FIG. 1 is a flow chart of a method for analyzing root causes of server failures according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0067] Step S101 : regularly collecting server alarm information and server log data, and determining whether the alarm information belongs to the first level or the second level.
[0068] Specifically, data collection tasks are automatically triggered at preset intervals (e.g., every minute, every 5 minutes). Leveraging standardized data collection interfaces and compatibility with multiple server monitoring protocols, this system enables real-time, stable collection of server alarm information and log data. Collected data includes, but is not limited to, hardware alarms (e.g., disk corruption, power failures), software alarms (e.g., process crashes, memory overflows), and comprehensive data such as server logs, application logs, and performance monitoring logs.
[0069] The alarm analysis unit automatically determines the severity level of each alarm using a pre-defined alarm classification rule base (e.g., categorizing alarms into the primary alarm segment (minor) and the secondary alarm segment (severe)), taking into account factors such as the alarm type, impact scope, and duration. It also counts the number of alarms received within the current cycle.
[0070] Step S102: When the alarm information belongs to the first level, the root cause analysis of the fault and the provision of repair suggestions are directly performed based on the first level information.
[0071] Specifically, for example, when the alarm analysis unit receives a single minor alarm, based on the principle of the most reasonable and economical use of resources, the fault diagnosis and review unit is directly called to locate the root cause of the fault and provide repair suggestions based on the alarm log of the component.
[0072] Step S103: When the alarm information belongs to the second level, intelligent agent reasoning is performed based on the alarm information and the second level information, and root cause analysis and repair suggestions are provided based on the intelligent agent reasoning results. At the same time, the root cause analysis and repair suggestions are reviewed, and the final root cause analysis results and final repair suggestions are output.
[0073] Specifically, when the alarm analysis unit receives multiple minor alarms, a single serious alarm, multiple serious alarms, or a mixture of minor and serious alarms, the server log data and the time of the alarm information, the corresponding equipment serial code and other information are submitted to the AI fault diagnosis agent unit for intelligent agent reasoning for correlation analysis, and based on the intelligent agent reasoning results, the fault diagnosis review unit is called to perform root cause analysis and provide repair suggestions; or the root cause analysis and repair suggestion processing operations are directly performed based on the intelligent agent results.
[0074] After the AI Fault Diagnosis Agent 2 infers the root cause analysis and remediation recommendations, the Fault Diagnosis Review Unit 3 verifies these results based on a knowledge base of expert experience accumulated through operational maintenance practices and a validated fault diagnosis rule system. This unit utilizes a built-in industry knowledge graph, a library of typical fault cases, and historical equipment operational data, along with a decision tree and verification algorithm developed through expert experience, to perform multi-dimensional verification of the root cause component and diagnostic conclusion initially determined by the AI Fault Diagnosis Agent 2. During the review process, the system focuses on assessing the feasibility (e.g., field operability), accuracy (match with historical cases), and rationality (compliance with operational and maintenance specifications) of the diagnostic results. For questionable diagnostic conclusions, the review unit initiates a correction mechanism or triggers a manual expert consultation process to ensure the final root cause analysis is highly reliable. This dual-security mechanism of "AI preliminary diagnosis + expert knowledge review" significantly improves the accuracy of complex equipment fault diagnosis, effectively avoiding potential misjudgments associated with AI diagnosis alone, and providing on-site operators with a more reliable basis for decision-making.
[0075] The server fault root cause analysis method provided in this embodiment regularly collects server log data and alarm information, implements root cause analysis in the case of multiple server alarms, automatically establishes a causal relationship chain between alarms, accurately and quickly locates the root cause fault component, and generates repair suggestions, greatly improving fault diagnosis efficiency. The root cause analysis and repair suggestions are reviewed, forming a dual verification mechanism of AI analysis and secondary review, avoiding single AI diagnosis misjudgments. Therefore, it effectively solves the problem in related technologies that it is difficult to quickly and accurately locate the root cause of the fault, and achieves the effect of automated root cause analysis of the server in multiple alarm information scenarios and accurately locating the source of the fault.
[0076] In this embodiment, a server fault root cause analysis method is provided, which can be used in a server fault root cause analysis system. The system includes an alarm analysis unit 1, an AI fault diagnosis agent unit 2, and a fault diagnosis review unit 3. The alarm analysis unit 1 is used to uniformly process server alarms and locate the root cause of alarm faults and conduct correlation analysis using different alarm segment dimensions. The AI fault diagnosis agent unit 2 is responsible for in-depth analysis and intelligent reasoning of alarm data. The fault diagnosis review unit 3 reviews the diagnosis results of the AI agent. Figure 2 FIG. 1 is a flow chart of a method for analyzing root causes of server failures according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:
[0077] Step S201: regularly collect server alarm information and server log data, and determine whether the alarm information belongs to the first level or the second level. Figure 1 Step S101 of the illustrated embodiment will not be described in detail here.
[0078] Step S202: When the alarm information belongs to the first level, the root cause analysis of the fault and the provision of repair suggestions are directly performed based on the first level information.
[0079] Specifically, if Figure 4 In the first workflow shown, the first level is a single first fault segment dimension; the server log includes multiple component alarm logs, and the above step S202 includes:
[0080] Step S2021: When the alarm information is a single first fault segment dimension, a root cause component corresponding to the single first fault segment dimension is determined, and a first component alarm log corresponding to the root cause component is obtained from multiple component alarm logs.
[0081] Specifically, a single first fault segment dimension is a single minor alarm.
[0082] When the alarm analysis unit 1 receives a single minor alarm, it determines the root cause component corresponding to the single minor alarm and the first component alarm log corresponding to the root cause component.
[0083] Step S2022: directly locate the root cause of the fault and provide repair suggestions based on the alarm log corresponding to the root cause component.
[0084] Specifically, based on the alarm log corresponding to the root cause component, in line with the principle of the most reasonable and economical use of resources, the reasoning logic of the AI fault diagnosis intelligent unit 2 is not enabled, and the fault diagnosis review unit 3 is directly called to locate the root cause of the fault and provide repair suggestions based on the alarm log of the root cause component.
[0085] In step S203, when the alarm information belongs to the second level, intelligent agent reasoning is performed based on the alarm information and the second level information, and root cause analysis of the fault and repair suggestions are provided based on the intelligent agent reasoning results. At the same time, the root cause analysis of the fault and the repair suggestions are reviewed, and the final root cause analysis results and final repair suggestions are output.
[0086] The server fault root cause analysis method provided in this embodiment customizes differentiated intelligent agent reasoning strategies for different alarm scenarios (multiple first fault segment dimensions, a single second fault segment dimension, multiple second fault segment dimensions, or mixed alarms). When multiple first fault segment dimension alarms are detected, the method focuses on component alarm logs and basic information reasoning to quickly resolve common faults. When a single second fault segment dimension alarm is detected, in-depth analysis is performed using performance data and full logs. In mixed alarm scenarios, multi-source data reasoning is fully integrated to accurately locate the root cause of even complex faults, significantly improving adaptability and diagnostic accuracy across various fault scenarios.
[0087] In this embodiment, a server fault root cause analysis method is provided, which can be used in a server fault root cause analysis system. The system includes an alarm analysis unit 1, an AI fault diagnosis agent unit 2, and a fault diagnosis review unit 3. The alarm analysis unit 1 is used to uniformly process server alarms; locate the root cause of the alarm fault and perform correlation analysis based on different alarm segment dimensions. The AI fault diagnosis agent unit 2 is responsible for in-depth analysis and intelligent reasoning of the alarm data. The fault diagnosis review unit 3 reviews the diagnosis results of the AI agent. Figure 3 FIG. 1 is a flow chart of a method for analyzing root causes of server failures according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:
[0088] Step S301: regularly collect server alarm information and server log data, and determine whether the alarm information belongs to the first level or the second level. Figure 2 Step S201 of the illustrated embodiment will not be described in detail here.
[0089] Step S302: When the alarm information is classified as level 1, the root cause analysis and repair suggestions are directly performed based on the level 1 information. Figure 2 Step S202 of the illustrated embodiment will not be described in detail here.
[0090] Step S303: When the alarm information belongs to the second level, intelligent agent reasoning is performed based on the alarm information and the second level information, and root cause analysis and repair suggestions are provided based on the intelligent agent reasoning results. At the same time, the root cause analysis and repair suggestions are reviewed, and the final root cause analysis results and final repair suggestions are output.
[0091] Specifically, if Figure 4 In the second to fourth workflow diagrams shown, the second level includes multiple first fault segment dimensions, a single second fault segment dimension and multiple second fault segment dimensions, or mixed alarms of multiple first fault segment dimensions and at least one second fault segment dimension; the server log also includes the full log; the above step S203 includes:
[0092] Step S3031, when the alarm information is multiple first fault segment dimensions, first intelligent agent reasoning is performed based on multiple component alarm logs and pre-acquired equipment serial numbers and alarm times of multiple first fault segment dimensions to obtain the first intelligent agent reasoning result, and based on the first intelligent agent reasoning result, the first fault root cause analysis and the first repair suggestion are provided.
[0093] Specifically, the AI fault diagnosis intelligent agent unit includes a data input interface, a multimodal correlation analysis engine root cause location module, and a repair suggestion generation module.
[0094] like Figure 4 In the third workflow diagram shown in Figure 1, the first fault segment dimension is a minor alarm. The AI fault diagnosis agent unit supports a multi-protocol adapter and receives structured data packets from the alarm analysis unit 1 through a data input interface. These packets contain: the device serial number (SN), which is used to associate device topology relationships in the CMDB; a sequence of alarm timestamps (T1...Tn) accurate to the millisecond level; the normalized encoding of component alarm logs (converting vendor-specific error codes according to the IPC-9592B standard); a server performance data matrix (5-second granularity time series metrics for CPU, memory, disk, and network); and the semantic feature vector of the full log (a 768-dimensional embedding representation pre-processed by the BERT model).
[0095] When N minor alarms are received, the device SN, the alarm time T1 corresponding to multiple minor alarms, the component alarm log WLG1,..., and the component alarm log WLGN data are submitted to the AI fault diagnosis intelligent unit 2 for correlation analysis, and the correlation between multiple minor alarms and the root cause of the fault are found, and repair suggestions are given.
[0096] In some optional implementations, step S3031 includes:
[0097] Step a1: When the alarm information is multiple first fault segment dimensions, based on multiple component alarm logs and pre-acquired equipment serial numbers and alarm times of multiple first fault segment dimensions, and combined with a preset hardware dependency rule base, an improved constraint-based non-parametric algorithm is used to construct a first reasoning knowledge graph.
[0098] In step a1, the multimodal correlation analysis engine uses a three-level analysis architecture to determine the relevance of alarms, including:
[0099] 1) Rule-level quick filtering: Preset hardware dependency rule base (e.g., "alarm triggering degradation rule for two disks in the same RAID group"). An example is as follows:
[0100] def detect RAID5 degradation (alarm information list):
[0101] # Filter out all disk error warnings
[0102] Disk error alarm = [alarm for alarm in alarm information list if alarm = 'disk error']
[0103] # Check if there are at least 2 disk errors and belong to the same RAID group
[0104] If len(disk error alarm)>= 2 and belongs to the same disk array group(disk error alarm):
[0105] return True, "RAID5 array is degraded. It is recommended to back up data immediately and replace the faulty disk."
[0106] else:
[0107] return False, "RAID5 degradation not detected".
[0108] 2) Temporal causal discovery layer, using the improved PCMCI+ algorithm (Partial Convergent Cross-Mapping, a constraint-based non-parametric algorithm):
[0109] By adding device topology constraints to the standard Peter-Clark (PC) algorithm, the dynamic adjustment formula for the time lag window is: τ = max(10s, 0.2×device physical distance coefficient).
[0110] Output is a fault propagation directed acyclic graph (DAG), where edge weights represent causal confidence.
[0111] 3) Knowledge graph inference layer: Build an operation and maintenance knowledge graph containing more than 3 million nodes. The implementation process is as follows:
[0112] Based on Neo4j (graph database) path query (such as: "power module failure → CPU frequency reduction → application timeout"), and using the Graph Attention Network (GAT) to predict potential propagation paths, we ultimately construct the first inference knowledge graph.
[0113] Step a2: Based on the first reasoning knowledge graph, the correlation and root cause of multiple first fault segment dimensions are searched, and the root cause of the fault is dynamically scored to obtain the first agent reasoning results of whether there is a unique root cause component or not.
[0114] Specifically, the root cause location module searches for the correlation and root causes of multiple first fault segment dimensions based on the first reasoning knowledge graph, and dynamically scores the root causes of the faults. The root cause scoring formula is as follows:
[0115] Root cause score = 0.4 × causal graph in-degree + 0.3 × historical matching + 0.2 × current health + 0.1 × topological centrality.
[0116] The in-degree of a causal graph refers to the number of other nodes pointing to a node in the fault causal relationship graph (i.e., the first-order reasoning knowledge graph). A higher in-degree indicates that the node is more influenced by other factors and is more likely to be a core point of fault propagation, thus being assigned a weight of 0.4.
[0117] The historical matching degree refers to the matching degree between the current fault characteristics and the historical known fault cases, and is assigned a weight of 0.3.
[0118] The current health refers to the degree to which the component's own status indicator deviates from the normal range when a fault occurs, and is assigned a weight of 02.
[0119] Topological centrality refers to the importance of a component in the system topology and is assigned a weight of 0.1.
[0120] The root cause location module also innovatively introduces a counterfactual verification mechanism, including: removing suspected root causes in the digital twin environment, verifying whether the remaining system alarms have disappeared, and finally outputting an explainable report to obtain the first-agent reasoning results of whether there is a unique root cause component or not.
[0121] Step a3: When the first intelligent agent reasoning result indicates that there is a unique root cause component, a first fault root cause analysis is performed on the unique root cause component and a first repair suggestion is generated; when the first intelligent agent reasoning result indicates that there is no root cause component, it is determined that multiple first fault segment dimensions are not correlated, corresponding first fault root cause analyses are performed on each of the multiple first fault segment dimensions, and corresponding first repair suggestions are generated based on the corresponding first fault root cause analyses.
[0122] Specifically, based on the analysis of the AI fault diagnosis intelligent unit 2, for a component with a unique root cause, the fault diagnosis review unit 3 is called to perform a root cause analysis on the unique root cause component and provide repair suggestions.
[0123] Based on the analysis of the AI fault diagnosis intelligent unit 2, if there is no root cause component, it means that there is no correlation between the minor alarms of multiple components. The fault diagnosis review unit 3 is called separately to perform fault root cause analysis and provide repair suggestions, that is, perform fault root cause analysis and provide repair suggestions for each component separately.
[0124] The repair suggestion provided in step a3 is implemented using a repair suggestion generation module, which generates a graded recommendation strategy based on the urgency of the alarm information, as shown in Table 1 below:
[0125] Table 1 Grading recommendation strategy
[0126]
[0127] In Table 1, in fault diagnosis and root cause analysis scenarios, confidence level is a key metric used to measure the reliability and trustworthiness of diagnostic results. It is typically expressed as a probability value (such as 95%). It reflects the diagnostic model or algorithm's confidence in the correctness of the currently determined root cause or conclusion.
[0128] Fault diagnosis systems (such as the reasoning model of an AI agent) analyze multiple sources of information, including alarm logs, performance data, and topological relationships, to output one or more possible root causes of the fault and assign a confidence score to each root cause. A higher confidence score indicates greater certainty in the system's judgment of the root cause. For example, when a server displays a RAID degradation alarm, the system, based on information such as disk error counts, RAID controller logs, and historical failure patterns, determines with 98% confidence that a "disk failure" is the root cause, indicating a very high degree of confidence in this conclusion.
[0129] Through this embodiment, for the alarm scenario of multiple minor alarms, the preset hardware dependency rule base and the improved constraint-based non-parametric algorithm are combined to construct a first reasoning knowledge graph, which can quickly sort out the logical relationship between the alarm logs of multiple components. Through the existing hardware dependency rules in the rule base and the use of algorithm constraints, the scattered alarm information is efficiently converted into a structured knowledge graph, which intuitively presents the potential connections between multiple minor alarms, avoids the isolated analysis of each alarm, and greatly improves the efficiency of fault correlation mining. Based on the constructed knowledge graph, the root cause of the fault is found and dynamic scoring is implemented. Taking into account various factors such as the causal graph in-degree and historical matching, a quantitative evaluation is performed on each possible root cause of the fault. The multi-dimensional scoring mechanism can more comprehensively and accurately screen out the true root cause of the fault than a single standard judgment. Processing strategies are formulated for the two reasoning results: the existence of a unique root cause component and the absence of a root cause component. When there is a unique root cause component, it is directly analyzed in depth and repair suggestions are generated to quickly resolve the problem. When there is no root cause component, that is, multiple alarms are not related, each first fault segment dimension is analyzed independently to avoid mistakenly judging independent faults as cascading faults, ensuring that different fault scenarios can be handled reasonably and effectively.
[0130] Step S3032, when the alarm information is a single second fault segment dimension, determine the root cause component corresponding to the single second fault segment dimension, and obtain the server performance data and the second component alarm log corresponding to the root cause component, perform second intelligent agent reasoning based on the second component alarm log, server performance data, full log, and the pre-acquired device serial number and the alarm time of the single second fault segment dimension, and perform a second fault root cause analysis and provide a second repair suggestion based on the second intelligent agent reasoning result.
[0131] In some optional embodiments, such as Figure 4 In the second workflow shown in FIG, step S3032 includes:
[0132] Step b1, when the alarm information is a single second fault segment dimension, determine the root cause component corresponding to the single second fault segment dimension based on the single second fault segment dimension, obtain the server performance data within the preset time period before the alarm information, and filter out the second component alarm log corresponding to the root cause component from multiple component alarm logs.
[0133] Specifically, the second fault segment dimension is recorded as a serious alarm.
[0134] When Alarm Analysis Unit 1 receives a single critical alarm, it identifies the root cause component based on the single critical alarm. Based on the alarm time, it searches for performance data and server log data within the past 48 hours and filters out the second component alarm log corresponding to the root cause component from the server log data. The device SN alarm time T1, the second component alarm log WLG1, the performance data PD, and the server's other log data LGD (full log) are submitted to AI Fault Diagnosis Agent Unit 2 for correlation analysis.
[0135] Step b2, based on the second component alarm log, server performance data, full log, and pre-acquired device serial number and alarm time, and combined with the preset hardware dependency rule base, an improved constraint-based non-parametric algorithm is used to construct a second reasoning knowledge graph.
[0136] Specifically, AI fault diagnosis agent unit 2 constructs a second reasoning knowledge graph using an improved constraint-based nonparametric algorithm based on the second component alarm log, server performance data, full log data, and pre-acquired device serial numbers and alarm times, combined with a preset hardware dependency rule base. The details of constructing the second reasoning knowledge graph are the same as those in step a1 above and are not repeated here.
[0137] Step b3: Search for direct root cause components and correlation factors corresponding to a single second fault segment dimension based on the second reasoning knowledge graph to obtain a second agent reasoning result indicating the existence of a direct root cause component or a correlation factor.
[0138] Specifically, correlation factors include the performance of other components, component anomalies that have not yet triggered alarms (only logs are recorded), business pressure, and other factors.
[0139] Step b4: When the second agent reasoning result indicates that there is a direct root cause component, a second fault root cause analysis is performed on the direct root cause component and a second repair suggestion is generated; when the second agent reasoning result indicates that there is a correlation factor, a fault root cause analysis result and a second repair suggestion are generated based on the correlation factor.
[0140] Specifically, for serious alarms of direct root cause components, the fault diagnosis review unit 3 is called to locate the root cause of the fault and provide repair suggestions. For serious alarms with related factors, such as the performance of other components, abnormalities of components that have not yet triggered alarms (only logs are recorded), business pressure and other factors, the AI fault diagnosis intelligent unit provides analysis results and repair suggestions.
[0141] Through this embodiment, multi-source data such as server performance data within a preset time period before the alarm information, screened second component alarm logs, and full logs are obtained, and combined with the device serial number and alarm time, the relevant information before and after the fault occurs is fully integrated. These data reflect the operating status of the server from different dimensions. By combining the preset hardware dependency rule base and the improved algorithm to build a knowledge graph, it is possible to deeply explore the potential connections between data, avoid missing key clues, and provide a rich and comprehensive data foundation for accurate fault analysis. For a single serious alarm, the constructed second reasoning knowledge graph is used to find the direct root cause components and correlation factors. When there is a direct root cause component, the source of the fault can be directly locked; if there are only correlation factors, it can also be comprehensively analyzed from aspects such as performance fluctuations and component associations to find the deep fault causes hidden under the surface phenomena, breaking the limitations of analyzing only a single alarm log and improving the diagnostic capabilities of complex faults.
[0142] Step S3033, when the alarm information is a mixed alarm of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, obtain server performance data, and perform third intelligent agent reasoning based on multiple component alarm logs, server performance data, full logs, and pre-acquired equipment serial numbers, multiple second fault segment dimensions or the alarm time of the mixed alarm, and perform a third fault root cause analysis and provide a third repair suggestion based on the third intelligent agent reasoning results.
[0143] In some optional embodiments, such as Figure 4 In the fourth workflow shown in FIG, step S3033 includes:
[0144] Step c1: when the alarm information is a mixed alarm of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, obtain server performance data within a preset time period before the alarm information.
[0145] Specifically, when N serious alarms or N mixed minor and serious alarms are received, based on the received alarm time, the performance data and server log data within 48 hours are retrieved, and the equipment SN, alarm time T1, component alarm log WLG1,..., component alarm log WLGN data, performance data PD and other server log data LGD (full log) are submitted to the AI fault diagnosis intelligent unit 2 for correlation analysis.
[0146] Step c2, based on multiple component alarm logs, server performance data, full logs, and pre-acquired device serial numbers, multiple second fault segment dimensions or alarm times of mixed alarms, and combined with the preset hardware dependency rule base, an improved constraint-based non-parametric algorithm is used to construct the third reasoning knowledge graph.
[0147] Specifically, AI fault diagnosis agent unit 2 constructs a third reasoning knowledge graph using an improved constraint-based nonparametric algorithm based on multiple component alarm logs, server performance data, full logs, pre-acquired device serial numbers, and the alarm times of multiple severe or mixed alarms, combined with a preset hardware dependency rule base. The details of constructing the third reasoning knowledge graph are the same as those in step a1 above and are not repeated here.
[0148] Step c3, based on the third reasoning knowledge graph, searches for the correlation and root cause of multiple second fault segment dimensions or multiple first fault segment dimensions with at least one second fault segment dimension, and obtains the third agent reasoning result of the existence of a unique root cause component, the existence of a root cause component, or the existence of a correlation factor.
[0149] Specifically, correlation factors include the performance of other components, component anomalies that have not yet triggered alarms (only logs are recorded), business pressure, and other factors.
[0150] Step c4, when the third-agent reasoning result is that there is a unique root cause component, a third fault root cause analysis is performed on the unique root cause component and a third repair suggestion is generated; when the third-agent reasoning result is that there is no root cause component, it is determined that multiple second fault segment dimensions or multiple first fault segment dimensions are not correlated with at least one second fault segment dimension, and corresponding third fault root cause analyses are performed on multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, and corresponding third repair suggestions are generated based on the corresponding third fault root cause analyses; when the third-agent reasoning result is that there is a correlation factor, the correlation between the mixed alarms of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension is searched based on the correlation factor, and the correlation fault root cause analysis result and the third repair suggestion are generated based on the correlation.
[0151] Specifically, for existing correlation factors, such as other component performance, component anomalies that haven't triggered alarms (only logging), and business pressure, the AI fault diagnosis agent directly searches for correlations between multiple alarms and the root cause of the fault. For a single, clearly identified root cause component, the fault diagnosis review unit 3 is invoked to analyze the root cause and provide repair recommendations. For alarms without a root cause component, indicating a lack of correlation between multiple component alarms, the fault diagnosis review unit 3 is invoked to analyze the root cause and provide repair recommendations.
[0152] In step S303, the AI fault diagnosis intelligent body unit 2 and the fault diagnosis review unit 3 collaborate in the workflow. First, the AI fault diagnosis intelligent body unit 2 infers the original diagnosis result, and then the fault diagnosis review unit 3 uses a composite marking system to perform a secondary review to obtain the final fault root cause analysis result and repair suggestions. When the review result is correct, the fault root cause analysis result and repair suggestions are added to the training set for continuous learning; when partially correct, partial result corrections are required to trigger the fine-tuning process; when the original diagnosis result inferred by the AI fault diagnosis intelligent body unit 2 is completely rejected, the expert intervention process is started.
[0153] Through this embodiment, in the face of complex scenarios such as multiple serious alarms or mixed alarms, by integrating multiple source information such as multiple component alarm logs, server performance data, full logs, and combining the device serial number and alarm time, a third reasoning knowledge graph is constructed, which can effectively sort out the logical relationship between complex alarms. Based on the third reasoning knowledge graph, not only can the direct root cause component of the fault be found, but also the existing correlation factors can be excavated. By combining the preset hardware dependency rule base and the improved algorithm, the fault correlation is comprehensively analyzed from multiple dimensions, whether it is a single unique root cause component or a fault caused by the mutual correlation between multiple components, it can be accurately located. For situations where there is no obvious root cause component, the system can also analyze each alarm separately to prevent unrelated faults from being misjudged as cascading faults, greatly improving the accuracy of root cause location in complex fault scenarios.
[0154] The fault diagnosis review unit 3 is used to review the root cause analysis and repair suggestions. The fault diagnosis review unit 3 is a key quality assurance link in the system fault diagnosis process. It relies on the expert experience knowledge base accumulated by the company in many years of operation and maintenance practice and the verified fault diagnosis rule system to authoritatively review the diagnosis results of the AI intelligent agent. This unit uses the built-in industry knowledge graph, typical fault case library and equipment operation and maintenance historical data, combined with the decision tree and verification algorithm formed by expert experience, to conduct multi-dimensional verification of the root cause components and diagnostic conclusions initially determined by the AI fault diagnosis intelligent agent unit 2. During the review process, the system will focus on evaluating the feasibility (such as on-site operability), accuracy (matching with historical cases) and rationality of the repair suggestions (compliance with operation and maintenance specifications) of the diagnostic results. For questionable diagnostic conclusions, the review unit will activate the correction mechanism or trigger the manual expert consultation process to ensure that the final output of the root cause analysis of the fault has a high degree of credibility. The above step S303 also includes:
[0155] Step S3034 , based on the preset expert experience knowledge base and preset historical fault cases, the fault root cause analysis and repair suggestions are reviewed using a preset multi-dimensional check matrix and contradiction detection algorithm.
[0156] Specifically, the preset expert experience knowledge base is a knowledge base built based on a multi-source knowledge fusion architecture, which also includes structured fault diagnosis rules, such as hardware diagnosis rules (such as disk bad sector patterns), software abnormality patterns (such as memory leak characteristics), and computer room physical topology constraints.
[0157] The Faiss vector database is used to store the features of historical fault cases (latitude = 256). The similarity calculation threshold is set to greater than 85%.
[0158] In order to update historical fault cases, a dynamic knowledge update method is used to obtain incremental knowledge from the following channels every month, as shown in Table 2:
[0159] Table 2 Incremental knowledge acquisition channels
[0160]
[0161] The default multi-dimensional check matrix is a five-layer check matrix, as shown in Table 3 below:
[0162] Table 3 Five-layer check matrix
[0163]
[0164] In step S3034, when the expert knowledge base is invoked, the review logic is monitored and rule matching statistics are performed to update the expert knowledge base. Monitoring includes the number and types of expert rules hit (e.g., hardware rules, software rules, topology rules), as well as the alarm signatures of missed rules (used to identify blind spots in the rule base). For example, when AI Fault Diagnosis Agent Unit 2 diagnoses a "memory fault," Fault Diagnosis Review Unit 3 invokes the hardware rule base in the expert knowledge base. If the rule "Prioritize controller for dual memory alarms in the same slot" is missed, this signature is recorded and a reminder to update the expert knowledge base is triggered.
[0165] By monitoring the alarm signatures of missed rules, scenarios missing or insufficiently covered in the expert rule base (such as new failure modes and alarm combinations under unique topologies) can be precisely identified. This avoids diagnostic bias caused by an outdated rule base, identifies rule blind spots, and improves the accuracy and completeness of the expert rule base. Triggering rule base update reminders prompts the operations team to add targeted rules (such as adding logic related to "cross-rack PSU cascading failures"), ensuring that the rule base evolves in sync with actual failure scenarios. By counting the number and type of matching rules (e.g., hardware rules account for 70% and topology rules account for 30%), the AI diagnostic results can be visually displayed to clearly demonstrate the alignment between expert experience and the user experience, providing a clear reference for manual review. This rule matching data allows prioritization of high-risk missed scenarios (such as multi-component alarms in complex topologies), preventing critical clues from being missed during manual review due to limited experience.
[0166] Step S3035: When a conflict occurs between the fault root cause analysis and the expert experience knowledge base during the review process, a preset priority level is used to determine whether to select the fault root cause analysis or manual consultation to resolve the conflict.
[0167] Specifically, the preset priority level is the priority level of the fault root cause analysis result.
[0168] When AI output conflicts with the rules in the expert experience knowledge base, Figure 5 The process shown resolves conflicts. Specifically, when a conflict is detected, it is determined whether the priority of the root cause analysis result is greater than or equal to level 3. If it is greater than or equal to level 3, a manual consultation method is selected to resolve the conflict. Otherwise, the root cause analysis result obtained by the AI fault diagnosis intelligent unit is used as the final analysis result and marked.
[0169] Step S3036: Modify the confidence level of the fault root cause analysis based on the conflict resolution result, output the final fault root cause analysis result, and generate a final repair suggestion based on the final fault root cause analysis result.
[0170] Specifically, the formula for adjusting the confidence level of the fault root cause analysis results output by AI is:
[0171] Final confidence = α × AI confidence + (1-α) × expert matching degree;
[0172] α is calculated dynamically, using the formula α = 1 / (1 + exp(-(number of similar cases in the expert database - 5)). The expert matching degree refers to the degree of match between the current fault characteristics and known fault cases in the expert knowledge base. It is typically expressed as a value between 0 and 1 (0 indicating a complete mismatch and 1 indicating a complete match). This value reflects the support provided by expert experience for the current fault diagnosis and serves as the basis for the fault diagnosis review unit (such as a human or rule engine) to revise the AI output results. The number of expert similar cases refers to the number of historical cases in the expert knowledge base that have similar characteristics to the current fault. "Similarity" is based on pre-set feature matching rules (e.g., identical alarm type, identical component, or similar topological location).
[0173] In order to better correct the confidence level of fault root cause analysis, a continuous learning interface is used for continuous learning. The feedback data format is as follows:
[0174] json
[0175] {
[0176] "original_ai": {"root_cause": "...", "confidence": 0.92}, / / Original AI diagnosis result, / / Root cause (e.g., power module voltage drop) / / Confidence (probability value between 0 and 1)
[0177] "adjusted_result": {"factor": "topology", "delta_confidence": -0.15}, / / Adjusted result / / Adjustment factors (such as topology constraints, expert rule conflicts, etc.) / / Confidence adjustment value (can be positive or negative, the cost is reduced by 15%)
[0178] "final_decision": {"action": "replace_psu", "executor": "human / AI"} / / Final decision / / Repair action (such as replacing the power module) / / Executor (human indicates manual execution, AI indicates automatic execution)
[0179] }.
[0180] This continuous learning interface uses a standardized feedback data format to transform empirical knowledge, such as manual review results and environmental constraints, into quantifiable model training data, achieving the following benefits:
[0181] 1) Continuously optimize AI models: By accumulating factor (adjustment factor) and delta_confidence (quantified adjustment range) data, we can identify model flaws (such as insufficient consideration of topological constraints) and optimize the algorithm accordingly.
[0182] 2) Improve diagnostic transparency: Record the complete logical chain from "AI initial judgment" to "final decision" to facilitate traceability and auditing;
[0183] 3) Support hybrid decision-making mode: clearly distinguish the boundaries between manual and automatic execution, and balance automation efficiency and manual reliability.
[0184] The server fault root cause analysis method provided in this embodiment is based on a preset expert experience knowledge base and a preset structured fault diagnosis rule base. It uses a preset multi-dimensional check matrix and contradiction detection algorithm to review the root cause analysis and repair suggestions, forming a dual verification mechanism of AI analysis + expert system review. This not only fully utilizes the rapid analysis advantages of AI, but also ensures the accuracy of the diagnosis results through the expert system, effectively solving the problem of misjudgment that may exist in a single AI system. This technology significantly reduces the complexity of operation and maintenance work, freeing operation and maintenance personnel from tedious troubleshooting, and enabling the transformation of the data center operation and maintenance model from manual control to intelligent automation. While improving operation and maintenance reliability, it provides a strong guarantee for the efficient and stable operation of the data center.
[0185] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0186] The embodiments of the present application also provide a server fault root cause analysis device, which is used to implement the above-mentioned embodiments and preferred embodiments. Details that have already been described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0187] This embodiment provides a server fault root cause analysis device, such as Figure 6 Shown, including:
[0188] The alarm information collection and judgment module 601 is used to regularly collect server alarm information and server log data, and judge whether the alarm information belongs to the first level or the second level;
[0189] The first fault root cause analysis module 602 is used to directly perform fault root cause analysis and provide repair suggestions based on the first level information when the alarm information is the first level;
[0190] The second fault root cause analysis and review module 603 is used to perform intelligent agent reasoning based on the alarm information and the second level information when the alarm information is at the second level, and to perform fault root cause analysis and provide repair suggestions based on the intelligent agent reasoning results, and at the same time review the fault root cause analysis and repair suggestions, and output the final fault root cause analysis results and final repair suggestions.
[0191] In some optional implementations, the first level is a single alarm severity level including a first fault segment dimension; the server log includes multiple component alarm logs; and the first fault root cause analysis module 602 includes:
[0192] A root cause component determination and alarm log acquisition unit, configured to, when the alarm information is a single first fault segment dimension, determine the root cause component corresponding to the single first fault segment dimension and acquire the first component alarm log corresponding to the root cause component;
[0193] The fault root cause location and repair suggestion providing unit is used to directly locate the fault root cause and provide repair suggestions based on the alarm log corresponding to the root cause component.
[0194] In some optional implementations, the second level includes multiple first fault segment dimensions, a single second fault segment dimension and multiple second fault segment dimensions, or mixed alarms of multiple first fault segment dimensions and at least one second fault segment dimension; the server log also includes a full log; and the second fault root cause analysis and review module 603 includes:
[0195] The first fault root cause analysis and first repair suggestion providing unit is used to perform first intelligent agent reasoning based on multiple component alarm logs and pre-acquired equipment serial numbers and alarm times of multiple first fault segment dimensions when the alarm information is multiple first fault segment dimensions, obtain the first intelligent agent reasoning result, and perform the first fault root cause analysis and provide the first repair suggestion based on the first intelligent agent reasoning result.
[0196] The second fault root cause analysis and second repair suggestion providing unit is used to determine the root cause component corresponding to the single second fault segment dimension when the alarm information is a single second fault segment dimension, and obtain the server performance data and the second component alarm log corresponding to the root cause component, perform second intelligent agent reasoning based on the second component alarm log, server performance data, full log, and the pre-acquired device serial number and the alarm time of the single second fault segment dimension, and perform second fault root cause analysis and second repair suggestion providing based on the second intelligent agent reasoning result.
[0197] The third fault root cause analysis and third repair suggestion providing unit is used to obtain server performance data when the alarm information is a mixed alarm of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, and perform third intelligent agent reasoning based on multiple component alarm logs, server performance data, full logs, and pre-acquired equipment serial numbers, multiple second fault segment dimensions or alarm times of mixed alarms, and perform third fault root cause analysis and third repair suggestion providing based on the third intelligent agent reasoning results.
[0198] The review unit is used to review the root cause analysis and repair suggestions of the fault based on the preset expert experience knowledge base and preset historical fault cases, using the preset multi-dimensional verification matrix and contradiction detection algorithm.
[0199] The conflict resolution unit is used to use a preset priority level to determine whether to select the fault root cause analysis or manual consultation method to resolve the conflict when a conflict occurs between the fault root cause analysis and the expert experience knowledge base during the review process.
[0200] The correction unit is used to correct the confidence of the fault root cause analysis based on the result after the conflict resolution, output the final fault root cause analysis result, and generate a final repair suggestion based on the final fault root cause analysis result.
[0201] In some optional implementations, the first fault root cause analysis and first repair suggestion providing unit includes:
[0202] A first reasoning knowledge graph construction subunit is configured to construct the first reasoning knowledge graph using an improved constraint-based non-parametric algorithm based on multiple component alarm logs, pre-acquired device serial numbers, and alarm times of multiple first fault segment dimensions, in combination with a preset hardware dependency rule base, when the alarm information is multiple first fault segment dimensions;
[0203] A first reasoning subunit is configured to search for correlations and root causes of multiple first fault segment dimensions based on the first reasoning knowledge graph, dynamically score the root causes of the faults, and obtain first agent reasoning results indicating the presence or absence of a unique root cause component;
[0204] The first fault root cause analysis and repair suggestion providing sub-unit is used to, when the first intelligent agent reasoning result is that there is a unique root cause component, perform the first fault root cause analysis on the unique root cause component and generate the first repair suggestion; when the first intelligent agent reasoning result is that there is no root cause component, determine that multiple first fault segment dimensions are not correlated, perform corresponding first fault root cause analysis on the multiple first fault segment dimensions respectively, and generate corresponding first repair suggestions based on the corresponding first fault root cause analysis.
[0205] In some optional implementations, the second fault root cause analysis and second repair suggestion providing unit includes:
[0206] The performance data and second component alarm log acquisition subunit is used to determine the root cause component corresponding to the single second fault segment dimension based on the single second fault segment dimension when the alarm information is a single second fault segment dimension, and to obtain the server performance data within a preset time period before the alarm information and to filter out the second component alarm log corresponding to the root cause component from multiple component alarm logs.
[0207] The second reasoning knowledge graph construction sub-unit is used to construct the second reasoning knowledge graph based on the second component alarm log, server performance data, full log, and pre-acquired device serial number and alarm time, and in combination with the preset hardware dependency rule base, using an improved constraint-based non-parametric algorithm.
[0208] The second reasoning subunit is used to search for the direct root cause component and correlation factors corresponding to the single second fault segment dimension based on the second reasoning knowledge graph, and obtain the second agent reasoning result that the direct root cause component or the correlation factor exists;
[0209] The second fault root cause analysis and repair suggestion sub-unit is provided, which is used to perform a second fault root cause analysis on the direct root cause component and generate a second repair suggestion when the second intelligent agent reasoning result has a direct root cause component; when the second intelligent agent reasoning result has a correlation factor, generate a fault root cause analysis result and a second repair suggestion based on the correlation factor.
[0210] In some optional implementations, the third fault root cause analysis and third repair suggestion providing unit includes:
[0211] A performance data acquisition subunit is configured to acquire server performance data within a preset time period before the alarm information is generated when the alarm information is a mixed alarm of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension;
[0212] The third reasoning knowledge graph construction subunit is used to construct the third reasoning knowledge graph based on multiple component alarm logs, server performance data, full logs, pre-acquired device serial numbers, multiple second fault segment dimensions, or alarm times of mixed alarms, and in combination with a preset hardware dependency rule base using an improved constraint-based non-parametric algorithm;
[0213] A third reasoning subunit is configured to search for the correlation between multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension and the root cause of the fault based on the third reasoning knowledge graph, and obtain a third agent reasoning result indicating the existence of a unique root cause component, the existence of a root cause component, or the existence of a correlation factor;
[0214] The third fault root cause analysis and repair suggestion providing sub-unit is used to, when the third intelligent agent reasoning result is that there is a unique root cause component, perform the third fault root cause analysis on the unique root cause component and generate the third repair suggestion; when the third intelligent agent reasoning result is that there is no root cause component, determine that multiple second fault segment dimensions or multiple first fault segment dimensions are not correlated with at least one second fault segment dimension, perform corresponding third fault root cause analysis on the multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, and generate corresponding third repair suggestions based on the corresponding third fault root cause analysis; when the third intelligent agent reasoning result is that there is a correlation factor, find the correlation between the mixed alarms of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension based on the correlation factor, and generate the correlation fault root cause analysis result and the third repair suggestion based on the correlation.
[0215] For the description of the features in the embodiment corresponding to the server failure root cause analysis device, reference can be made to the relevant description of the embodiment corresponding to the server failure root cause analysis method, which will not be repeated here.
[0216] The embodiment of the present application also provides an electronic device, such as Figure 7 As shown, it includes a memory 10 and a processor 20. The memory 10 stores a computer program, and the processor 20 is configured to run the computer program to execute the steps in any of the above server failure root cause analysis method embodiments.
[0217] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned server failure root cause analysis method embodiments when running.
[0218] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0219] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned server failure root cause analysis method embodiments are implemented.
[0220] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned server failure root cause analysis method embodiments are implemented.
[0221] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0222] The above is a detailed introduction to a server failure root cause analysis method, device, electronic device and storage medium provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A server failure root cause analysis method, characterized in that: include: Regularly collect server alarm information and server log data, and determine whether the alarm information belongs to the first level or the second level; When the alarm information belongs to the first level, the root cause analysis and repair suggestions are directly performed based on the first level information; When the alarm information belongs to the second level, the agent performs reasoning based on the alarm information and the second level information, and performs root cause analysis and repair suggestions based on the agent reasoning results. At the same time, the root cause analysis and repair suggestions are reviewed, and the final root cause analysis results and final repair suggestions are output; The second level includes multiple first fault segment dimensions, a single second fault segment dimension and multiple second fault segment dimensions, or mixed alarms of multiple first fault segment dimensions and at least one second fault segment dimension; the server log also includes a full log; When the alarm information belongs to the second level, performing intelligent agent reasoning based on the alarm information and the second level information, and performing fault root cause analysis and providing repair suggestions based on the intelligent agent reasoning results, including: When the alarm information is multiple first fault segment dimensions, an improved constraint-based non-parametric algorithm is used to construct a first reasoning knowledge graph based on multiple component alarm logs, pre-acquired device serial numbers, and alarm times of multiple first fault segment dimensions, combined with a preset hardware dependency rule base; Based on the first reasoning knowledge graph, the correlation and root cause of multiple first fault segment dimensions are searched, and the root cause of the fault is dynamically scored to obtain the first agent reasoning results of whether there is a unique root cause component or no root cause component. Based on the first agent reasoning results, the first fault root cause analysis and the first repair suggestion are provided; When the alarm information is a single second fault segment dimension, the root cause component corresponding to the single second fault segment dimension is determined based on the single second fault segment dimension, and the server performance data within a preset time period before the alarm information is obtained, and the second component alarm log corresponding to the root cause component is filtered out from the multiple component alarm logs; based on the second component alarm log, server performance data, full log, and pre-acquired device serial number and alarm time, and in combination with the preset hardware dependency rule base, an improved constraint-based non-parametric algorithm is used to construct a second reasoning knowledge graph; based on the second reasoning knowledge graph, the direct root cause component and correlation factors corresponding to the single second fault segment dimension are searched, and a second intelligent agent reasoning result is obtained for the existence of a direct root cause component or the existence of a correlation factor, and a second fault root cause analysis and a second repair suggestion are provided based on the second intelligent agent reasoning result; When the alarm information is a mixed alarm of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, obtain the server performance data within a preset time period before the alarm information; based on multiple component alarm logs, server performance data, full logs, and pre-acquired equipment serial numbers, multiple second fault segment dimensions or the alarm time of the mixed alarm, and combined with the preset hardware dependency rule base, an improved constraint-based non-parametric algorithm is used to construct a third reasoning knowledge graph; based on the third reasoning knowledge graph, the correlation and root cause of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension are searched, and the third intelligent agent reasoning results of the existence of a unique root cause component, the existence of a root cause component or the existence of a correlation factor are obtained, and based on the third intelligent agent reasoning results, a third fault root cause analysis and a third repair suggestion are provided.
2. The server failure root cause analysis method according to claim 1, characterized in that: The first level is a single first fault segment dimension; the server log includes multiple component alarm logs; When the alarm information is classified as level 1, the root cause analysis and repair suggestions are directly performed based on the level 1 information, including: When the alarm information is a single first fault segment dimension, determining a root cause component corresponding to the single first fault segment dimension, and obtaining a first component alarm log corresponding to the root cause component from the multiple component alarm logs; Directly locate the root cause of the fault and provide repair suggestions based on the alarm log corresponding to the root cause component.
3. The server failure root cause analysis method according to claim 1, characterized in that: Perform a first fault root cause analysis and provide a first repair suggestion based on the first agent's reasoning results, including: When the first agent reasoning result is that there is a unique root cause component, performing a first fault root cause analysis on the unique root cause component and generating a first repair suggestion; When the first intelligent agent reasoning result indicates that there is no root cause component, it is determined that multiple first fault segment dimensions are not correlated, corresponding first fault root cause analyses are performed on the multiple first fault segment dimensions respectively, and corresponding first repair suggestions are generated based on the corresponding first fault root cause analyses.
4. The server failure root cause analysis method according to claim 1, characterized in that: Based on the second agent's reasoning results, a second fault root cause analysis and second repair suggestions are provided, including: When the second agent reasoning result is for a direct root cause component, a second fault root cause analysis is performed on the direct root cause component and a second repair suggestion is generated; when the second agent reasoning result is for a correlation factor, a fault root cause analysis result and a second repair suggestion are generated based on the correlation factor.
5. The server failure root cause analysis method according to claim 1, characterized in that: Based on the third-party agent's reasoning results, the third-party fault root cause analysis and third-party repair suggestions are provided, including: When the third agent reasoning result is for a unique root cause component, performing a third fault root cause analysis on the unique root cause component and generating a third repair suggestion; When the third agent reasoning result indicates that there is no root cause component, it is determined that the multiple second fault segment dimensions or the multiple first fault segment dimensions are not associated with the at least one second fault segment dimension, corresponding third fault root cause analyses are performed on the multiple second fault segment dimensions or the multiple first fault segment dimensions and the at least one second fault segment dimension, and corresponding third repair suggestions are generated based on the corresponding third fault root cause analyses; When the third agent reasoning result has correlation factors, the correlation between multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension mixed alarm is found based on the correlation factors, and the correlation fault root cause analysis result and the third repair suggestion are generated based on the correlation.
6. The server failure root cause analysis method according to claim 1, characterized in that: Review the root cause analysis and repair suggestions, and output the final root cause analysis results and final repair suggestions, including: Based on the preset expert experience knowledge base and preset historical failure cases, the preset multi-dimensional check matrix and contradiction detection algorithm are used to review the fault root cause analysis and repair suggestions; When a conflict occurs between the root cause analysis and the expert experience knowledge base during the review process, the preset priority level is used to determine whether to choose the root cause analysis or manual consultation method to resolve the conflict; The confidence level of the fault root cause analysis is modified based on the conflict resolution result, the final fault root cause analysis result is output, and the final repair suggestion is generated based on the final fault root cause analysis result.
7. A server failure root cause analysis device, characterized in that: include: The alarm information collection and judgment module is used to regularly collect server alarm information and server log data, and judge whether the alarm information belongs to the first level or the second level; The first fault root cause analysis module is used to directly perform fault root cause analysis and provide repair suggestions based on the first level information when the alarm information is the first level; The second fault root cause analysis and review module is used to perform intelligent agent reasoning based on the alarm information and the second-level information when the alarm information is at the second level, and to perform fault root cause analysis and provide repair suggestions based on the intelligent agent reasoning results, while reviewing the fault root cause analysis and repair suggestions, and outputting the final fault root cause analysis results and final repair suggestions.
8. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the server failure root cause analysis method according to any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the server failure root cause analysis method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Server fault diagnosis method and device, storage medium and electronic equipment
CN118394561A
Adaptive operation and maintenance root cause positioning method and system based on deep learning
CN119691576A
Cited By
Abnormity analysis and intelligent diagnosis method and system based on multi-agent cooperation
CN121396745A