Server fault root cause analysis method and device, electronic equipment and storage medium

By building a server fault root cause analysis method, using level judgment and intellectual inference combined with knowledge graph, the problem of difficulty in fault location under multiple alarm information of the server is solved, and fast and accurate fault source positioning and diagnosis is achieved, reducing operation and maintenance complexity.

CN120276907AActive Publication Date: 2025-07-08INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510750916.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-08
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

In the prior art, it is difficult for servers to quickly and accurately locate the root cause of failure in multiple interrelated alarm information scenarios, resulting in operation and maintenance difficulties and business interruptions.

Method used

By collecting server alarm information and log data regularly, using level judgment mechanisms and agent reasoning, a causal relationship chain is built, combining preset hardware dependency rule bases and improved non-parametric algorithms to build a knowledge graph, perform multi-dimensional root cause analysis, and a dual-factor verification mechanism of AI analysis and expert system review is adopted.

Benefits of technology

It realizes fast and accurate fault source positioning in multiple alarm information scenarios, improves fault diagnosis efficiency, reduces operation and maintenance complexity, and ensures the accuracy and reliability of diagnostic results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276907A_ABST
    Figure CN120276907A_ABST
Patent Text Reader

Abstract

The invention discloses a server fault root cause analysis method and device, electronic equipment and a storage medium, and relates to the technical field of servers, server performance data, logs and alarm information are collected regularly, an alarm causal relationship chain is automatically constructed in a multi-alarm scene, root cause fault components are accurately identified, and the server fault root cause analysis efficiency is improved. A solution suggestion is intelligently generated in combination with a historical case library, the fault diagnosis efficiency is greatly improved, fault root cause analysis and repair suggestions are rechecked, a dual verification mechanism of agent analysis and secondary rechecking is formed, misjudgment of single agent diagnosis is avoided, and the fault diagnosis accuracy is improved. The technical problem that the root cause of the fault is difficult to quickly and accurately locate in the related technology is solved, and the technical effects of automatic root cause analysis and accurate location of the fault source of the server in a multi-alarm information scene are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of servers, and particularly to a method, device, electronic device, and storage medium for analyzing the root cause of server failures. Background Art

[0002] During the daily operation of a server, a large amount of diverse alarm information is generated, including various types such as hardware failures, performance bottlenecks, and network anomalies. The same server may generate multiple interrelated alarms within a short period, and there may be complex causal relationships between these alarms.

[0003] When faced with multiple interrelated alarm messages, there is a lack of an effective alarm correlation analysis mechanism, making it difficult to identify key clues from the numerous and complex alarm information. When a failure occurs, operations and maintenance personnel are only busy dealing with the surface symptoms of the server and find it difficult to quickly and accurately locate the root cause of the failure, resulting in the serious problem of server operation business interruption. Summary of the Invention

[0004] This application provides a method, device, electronic device, and storage medium for analyzing the root cause of server failures to at least solve the problem in related technologies of being difficult to quickly and accurately locate the root cause of a failure.

[0005] This application provides a method for analyzing the root cause of server failures, including: Regularly collecting server alarm information and server log data, and determining whether the alarm information belongs to the first level or the second level; When the alarm information belongs to the first level, directly perform root cause analysis of the failure and provide repair suggestions based on the first-level information; When the alarm information belongs to the second level, perform agent reasoning based on the alarm information and the second-level information, and perform root cause analysis of the failure and provide repair suggestions based on the agent reasoning result. At the same time, review the root cause analysis of the failure and the repair suggestions, and output the final root cause analysis result of the failure and the final repair suggestion.

[0006] Through this application, since the server log data and alarm information are regularly collected, and the root cause analysis in the case of multiple server alarms is realized according to the level of the alarm information, the causal relationship chain between alarms is automatically established, the root cause failure component is accurately and quickly located, and repair suggestions are generated, greatly improving the failure diagnosis efficiency. The root cause analysis of the failure and the repair suggestions are reviewed to form a double verification mechanism of AI analysis and secondary review, avoiding misjudgment of single AI diagnosis. Therefore, the problem in related technologies of being difficult to quickly and accurately locate the root cause of a failure is effectively solved, achieving the effect of automatic root cause analysis of the server in the scenario of multiple alarm information and accurately locating the failure source.

[0007] In an alternative embodiment, the first level is a single first fault segment dimension; the server log includes multiple component alarm logs; When the alarm information belongs to the first level, directly perform root cause analysis of the fault and provide repair suggestions based on the first level information, including: When the alarm information is a single first fault segment dimension, determine the root cause component corresponding to the single first fault segment dimension, and obtain the first component alarm log corresponding to the root cause component from multiple component alarm logs; Directly perform root cause location of the fault and provide repair suggestions based on the alarm log corresponding to the root cause component.

[0008] Through the present application, when receiving the alarm information of a single first fault segment dimension, the corresponding root cause component can be directly determined without performing complex multi-source data correlation analysis or agent reasoning process. Therefore, this "direct mapping" processing method greatly shortens the fault diagnosis time, can start the repair process in the shortest time, quickly restore the normal operation of the server, and reduce the service interruption duration caused by the fault.

[0009] In an alternative embodiment, the second level includes multiple first fault segment dimensions, a single second fault segment dimension and multiple second fault segment dimensions or a mixed alarm of multiple first fault segment dimensions and at least one second fault segment dimension; the server log further includes a full log; When the alarm information belongs to the second level, perform agent reasoning based on the alarm information and the second level information, and perform root cause analysis of the fault and provide repair suggestions based on the agent reasoning result, including: When the alarm information is multiple first fault segment dimensions, perform first agent reasoning based on multiple component alarm logs, the pre-acquired device serial number and the alarm time of multiple first fault segment dimensions to obtain a first agent reasoning result, and perform first fault root cause analysis and first repair suggestion providing based on the first agent reasoning result; When the alarm information is a single second fault segment dimension, determine the root cause component corresponding to the single second fault segment dimension, and obtain the server performance data and the second component alarm log corresponding to the root cause component. Perform second agent reasoning based on the second component alarm log, the server performance data, the full log, the pre-acquired device serial number and the alarm time of the single second fault segment dimension, and perform second fault root cause analysis and second repair suggestion providing based on the second agent reasoning result; When the alarm information is a mixed alarm of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, obtain the server performance data, and perform third-agent reasoning based on multiple component alarm logs, server performance data, full logs, and the device serial number, multiple second fault segment dimensions, or the alarm time of the mixed alarm obtained in advance, and perform third fault root cause analysis and provide third repair suggestions based on the third-agent reasoning result.

[0010] Through this application, customized differential agent reasoning strategies are developed for different alarm scenarios (multiple first fault segment dimensions, single second fault segment dimension, multiple second fault segment dimensions, or mixed alarms). When there are multiple first fault segment dimension alarms, focus on reasoning from component alarm logs and basic information to quickly handle common faults; when there is a single second fault segment dimension alarm, conduct in-depth analysis by combining performance data, full logs, etc.; in the mixed alarm scenario, comprehensively integrate multi-source data for reasoning to ensure that the root cause of complex faults can be accurately located, greatly improving the adaptability to various fault scenarios and the diagnostic accuracy.

[0011] In an optional implementation manner, when the alarm information is multiple first fault segment dimensions, perform first-agent reasoning based on multiple component alarm logs, the device serial number obtained in advance, and the alarm time of multiple first fault segment dimensions to obtain a first-agent reasoning result, and perform first fault root cause analysis and provide first repair suggestions based on the first-agent reasoning result, including: When the alarm information is multiple first fault segment dimensions, based on multiple component alarm logs, the device serial number obtained in advance, and the alarm time of multiple first fault segment dimensions, and in combination with a preset hardware dependency rule library, construct a first reasoning knowledge graph using an improved constraint-based non-parametric algorithm; Based on the first reasoning knowledge graph, find the relevance and fault root cause of multiple first fault segment dimensions, and perform dynamic scoring on the fault root cause to obtain a first-agent reasoning result of the component with a unique root cause and the component without a root cause; When it is the first-agent reasoning result of the component with a unique root cause, perform first fault root cause analysis on the unique root cause component and generate a first repair suggestion; When it is the first-agent reasoning result of the component without a root cause, determine that multiple first fault segment dimensions have no relevance, perform corresponding first fault root cause analysis on multiple first fault segment dimensions respectively, and generate corresponding first repair suggestions based on the corresponding first fault root cause analysis.

[0012] Through this application, for the alarm scenarios of multiple first fault segment dimensions, a first inference knowledge graph is constructed by combining a preset hardware dependency rule library and an improved constraint-based non-parametric algorithm, which can quickly sort out the logical relationships among the alarm logs of multiple components. By using the existing hardware dependency rules in the rule library and the algorithm constraint conditions, the scattered alarm information is efficiently converted into a structured knowledge graph, visually presenting the potential relationships among the alarms of multiple first fault segment dimensions, avoiding isolated analysis of each alarm, and greatly improving the efficiency of fault correlation mining. Based on the constructed knowledge graph, the root cause of the fault is found and dynamically scored. By comprehensively considering various factors such as the in-degree of the causal graph and the historical matching degree, each possible root cause of the fault is quantitatively evaluated. The multi-dimensional scoring mechanism can screen out the real root cause of the fault more comprehensively and accurately than a single standard judgment. For the two inference results of the existence of a unique root cause component and the non-existence of a root cause component, corresponding processing strategies are formulated respectively. When there is a unique root cause component, it is directly analyzed in depth and repair suggestions are generated to quickly solve the problem; when there is no root cause component, that is, when multiple alarms are not related, each first fault segment dimension is analyzed independently to avoid misjudging independent faults as cascaded faults, ensuring that different fault scenarios can be reasonably and effectively processed.

[0013] In an alternative embodiment, when the alarm information is a single second fault segment dimension, determine the root cause component corresponding to the single second fault segment dimension, and obtain the server performance data and the second component alarm log corresponding to the root cause component. Perform second-agent reasoning based on the second component alarm log, the server performance data, the full log, and the pre-obtained device serial number and the alarm time of the single second fault segment dimension, and provide second fault root cause analysis and second repair suggestions based on the second-agent reasoning result, including: When the alarm information is a single second fault segment dimension, based on the single second fault segment dimension, determine the root cause component corresponding to the single second fault segment dimension, and obtain the server performance data within a preset time period before the alarm information and the second component alarm log filtered from the multiple component alarm logs corresponding to the root cause component; Based on the second component alarm log, the server performance data, the full log, and the pre-obtained device serial number and the alarm time, and combined with the preset hardware dependency rule library, use an improved constraint-based non-parametric algorithm to construct a second inference knowledge graph; Based on the second inference knowledge graph, find the direct root cause component and relevant factors corresponding to the single second fault segment dimension, and obtain the second-agent reasoning result of the existence of a direct root cause component or the existence of relevant factors; When it is the second agent inference result with a direct root cause component, perform a second root cause analysis on the direct root cause component and generate a second repair suggestion; when it is the second agent inference result with correlation factors, generate a root cause analysis result and a second repair suggestion based on the correlation factors.

[0014] Through this application, multi-source data such as server performance data, filtered second component alarm logs, and full logs within a preset time period before obtaining the alarm information are obtained. Combining the device serial number and the alarm time, relevant information before and after the fault is comprehensively integrated. These data reflect the server operating status from different dimensions. By combining a preset hardware dependency rule library and an improved algorithm to construct a knowledge graph, the potential connections between data can be deeply mined, key clues can be avoided from being missed, and a rich and comprehensive data basis for accurate fault analysis can be provided. For a single second fault segment dimension alarm, use the constructed second inference knowledge graph to find the direct root cause component and correlation factors. When there is a direct root cause component, the fault source can be directly locked; if there are only correlation factors, it is also possible to comprehensively analyze from aspects such as performance fluctuations and component associations to find the deep fault cause hidden under the surface phenomenon, breaking the limitation of analyzing only from a single alarm log and improving the diagnostic ability for complex faults.

[0015] In an alternative implementation, when the alarm information is a mixed alarm of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, obtain the server performance data, and perform a third agent inference based on multiple component alarm logs, server performance data, full logs, and the previously obtained device serial number, the alarm time of multiple second fault segment dimensions or the mixed alarm, and provide a third root cause analysis and a third repair suggestion based on the third agent inference result, including: When the alarm information is a mixed alarm of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, obtain the server performance data within a preset time period before obtaining the alarm information; Based on multiple component alarm logs, server performance data, full logs, and the previously obtained device serial number, the alarm time of multiple second fault segment dimensions or the mixed alarm, and combine a preset hardware dependency rule library to use an improved constraint-based non-parametric algorithm to construct a third inference knowledge graph; Based on the third inference knowledge graph, find the relevance and root cause of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, and obtain a third agent inference result of having a unique root cause component, having a root cause component, or having correlation factors; When it is the third agent inference result of having a unique root cause component, perform a third root cause analysis on the unique root cause component and generate a third repair suggestion; When it is the third agent inference result without a root cause component, it is determined that multiple second fault segment dimensions or multiple first fault segment dimensions have no relevance to at least one second fault segment dimension. Corresponding third fault root cause analyses are respectively performed for multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, and corresponding third repair suggestions are generated based on the corresponding third fault root cause analyses; When it is the third agent inference result with relevant factors, the relevance of the mixed alarms of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension is searched based on the relevant factors, and a relevant fault root cause analysis result and a third repair suggestion are generated based on the relevance.

[0016] Through this application, in the face of complex scenarios such as multiple second fault segment dimensions or mixed alarms, by integrating multi-source information such as multiple component alarm logs, server performance data, and full-volume logs, and combining the device serial number and alarm time, a third inference knowledge graph is constructed, which can effectively sort out the logical relationships between complex alarms. Based on the third inference knowledge graph, not only can the direct root cause component of the fault be found, but also the existing relevant factors can be mined. By combining the preset hardware dependency relationship rule library and the improved algorithm, the fault association is comprehensively analyzed from multiple dimensions. Whether it is a single unique root cause component or a fault caused by the mutual association between multiple components, it can be accurately located. For the situation where there is no obvious root cause component, the system can also independently analyze each alarm to prevent misjudging unassociated faults as cascaded faults, greatly improving the accuracy of root cause location in complex fault scenarios.

[0017] In an alternative implementation manner, the fault root cause analysis and repair suggestions are reviewed, and the final fault root cause analysis result and the final repair suggestion are output, including: Based on the preset expert experience knowledge base and preset historical fault cases, the fault root cause analysis and repair suggestions are reviewed by using the preset multi-dimensional verification matrix and contradiction detection algorithm; When a conflict occurs between the fault root cause analysis and the expert experience knowledge base during the review process, the preset priority level determination is used to select the fault root cause analysis or select the method of manual consultation to resolve the conflict; Based on the result after conflict resolution, the confidence level of the fault root cause analysis is corrected, the final fault root cause analysis result is output, and the final repair suggestion is generated based on the final fault root cause analysis result.

[0018] Through this application, based on a preset expert experience knowledge base and a preset structured fault diagnosis rule base, a preset multi-dimensional verification matrix and a contradiction detection algorithm are used to review the root cause analysis of faults and repair suggestions, forming a dual verification mechanism of AI analysis + expert system review. This mechanism not only gives full play to the fast analysis advantage of AI but also ensures the accuracy of the diagnosis results through the expert system, effectively solving the misjudgment problem that may exist in a single AI system. This technology significantly reduces the complexity of operation and maintenance work, liberates operation and maintenance personnel from cumbersome fault troubleshooting, enables the operation and maintenance mode of the data center to transform from manual domination to intelligent automation, and while improving the reliability of operation and maintenance, provides a strong guarantee for the efficient and stable operation of the data center.

[0019] This application also provides a server fault root cause analysis device, including: An alarm information collection and judgment module, configured to regularly collect server alarm information and server log data, and judge whether the alarm information belongs to the first level or the second level; A first fault root cause analysis module, configured to directly perform root cause analysis of faults and provide repair suggestions based on the first-level information when the alarm information is at the first level; A second fault root cause analysis and review module, configured to perform agent reasoning based on the alarm information and the second-level information when the alarm information is at the second level, provide root cause analysis of faults and repair suggestions based on the agent reasoning results, and at the same time review the root cause analysis of faults and repair suggestions, and output the final root cause analysis result of faults and the final repair suggestions.

[0020] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above server fault root cause analysis methods when executing the computer program.

[0021] This application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above server fault root cause analysis methods are implemented.

[0022] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the above server fault root cause analysis methods are implemented. Description of the Drawings

[0023] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0024] Figure 1 It is a schematic flowchart of a method for analyzing the root cause of server failures provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of another method for analyzing the root cause of server failures provided by an embodiment of the present application; Figure 3 It is a schematic flowchart of yet another method for analyzing the root cause of server failures provided by an embodiment of the present application; Figure 4 It is a working flowchart of a system for analyzing the root cause of server failures provided by an embodiment of the present application; Figure 5 It is a schematic diagram of resolving conflicts by a fault diagnosis and review unit in a system for analyzing the root cause of server failures provided by an embodiment of the present application; Figure 6 It is a structural block diagram of a device for analyzing the root cause of server failures provided by an embodiment of the present application; Figure 7 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Specific embodiments

[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.

[0026] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0027] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0028] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the method for analyzing the root cause of server failures depends, the specific application environment architecture or specific hardware architecture will be described herein.

[0029] An embodiment of the present application provides a method for analyzing the root cause of server failures, and the method will be described in detail in combination with the execution process of the method for analyzing the root cause of server failures.

[0030] In this embodiment, a method for analyzing the root cause of server failures is provided, which can be used in a server failure root cause analysis system. The system includes an alarm analysis unit 1, an AI failure diagnosis agent unit 2, and a failure diagnosis review unit 3. Among them, the alarm analysis unit 1 is used for the unified processing of server alarms; the root cause positioning and correlation analysis of alarms are carried out in different alarm segment dimensions. The AI failure diagnosis agent unit 2 is responsible for the in-depth analysis and intelligent reasoning of alarm data. The failure diagnosis review unit 3 reviews the diagnosis results of the AI agent. Figure 1 It is a flowchart of the server failure root cause analysis method according to an embodiment of the present invention, as Figure 1 shown, and the process includes the following steps: Step S101, regularly collect server alarm information and server log data, and determine whether the alarm information belongs to the first level or the second level.

[0031] Specifically, the data collection task is automatically triggered at a preset time interval (such as every minute, every 5 minutes). Using a standardized data collection interface, compatible with multiple server monitoring protocols, real-time and stable collection of server alarm information and log data is achieved. The collected data includes but is not limited to hardware alarms (such as disk damage, power failure), software alarms (such as process crashes, memory overflows), and all data such as server logs, application logs, and performance monitoring logs.

[0032] The alarm analysis unit uses a preset alarm classification rule library (such as classifying alarms into the first alarm segment dimension, that is, the minor level, and the second alarm segment dimension, that is, the severe level), and combines factors such as the type, impact range, and duration of the alarm information to automatically determine the severity level of each alarm. At the same time, the number of alarm information in the current period is counted.

[0033] Step S102, when the alarm information belongs to the first level, directly perform root cause analysis of the failure and provide repair suggestions based on the first-level information.

[0034] Specifically, for example, when the alarm analysis unit receives a single minor alarm, in line with the principle of most reasonably saving resources, it directly calls the failure diagnosis review unit to locate the root cause of the failure and provide repair suggestions based on the alarm log of the component.

[0035] Step S103, when the alarm information belongs to the second level, perform agent reasoning based on the alarm information and the second-level information, and perform root cause analysis of the failure and provide repair suggestions based on the agent reasoning result. At the same time, review the root cause analysis of the failure and the repair suggestions, and output the final root cause analysis result of the failure and the final repair suggestion.

[0036] Specifically, when the alarm analysis unit receives multiple minor alarms, a single severe alarm, multiple severe alarms, or a mixed alarm of minor and severe alarms, it submits the server log data, the time of the alarm information, the corresponding device serial number, and other information to the AI fault diagnosis agent unit for correlation analysis and agent reasoning, and selects and invokes the fault diagnosis review unit based on the agent reasoning result to perform root cause analysis of the fault and provide repair suggestions; or directly performs the processing operation of root cause analysis of the fault and providing repair suggestions based on the agent result.

[0037] When the AI fault diagnosis agent unit 2 infers the root cause analysis result and repair suggestions of the fault, the fault diagnosis review unit 3 is used to review the root cause analysis result and repair suggestions of the fault based on the expert experience knowledge base accumulated in operation and maintenance practice and the verified fault diagnosis rule system. This unit conducts multi-dimensional verification on the root cause fault component and diagnosis conclusion initially determined by the AI fault diagnosis agent unit 2 through the built-in industry knowledge graph, typical fault case library, and device operation and maintenance historical data, combined with the decision tree and verification algorithm formed by expert experience. During the review process, the system will focus on evaluating the feasibility of the diagnosis result (such as on-site operability), accuracy (matching degree with historical cases), and rationality of the repair suggestions (meeting the requirements of operation and maintenance specifications). For doubtful diagnosis conclusions, the review unit will initiate a correction mechanism or trigger an artificial expert consultation process to ensure that the finally output root cause analysis of the fault has a high degree of credibility. This dual guarantee mechanism of "AI preliminary diagnosis + expert knowledge review" significantly improves the accuracy of complex device fault diagnosis, effectively avoids misjudgment situations that may occur in single AI diagnosis, and provides a more reliable decision-making basis for on-site operation and maintenance personnel.

[0038] The server root cause analysis method provided in this embodiment regularly collects server log data and alarm information, realizes root cause analysis in the case of multiple server alarms, automatically establishes a causal relationship chain between alarms, accurately and quickly locates the root cause fault component, and generates repair suggestions, greatly improving the fault diagnosis efficiency. It reviews the root cause analysis of the fault and repair suggestions to form a dual verification mechanism of AI analysis and secondary review, avoiding single AI diagnosis misjudgment. Therefore, it effectively solves the problem in the related technology of being difficult to quickly and accurately locate the root cause of the fault, achieving the effect of automatic root cause analysis of the server in the scenario of multiple alarm information and accurately locating the fault source.

[0039] In this embodiment, a method for analyzing the root cause of server failures is provided, which can be used in a server failure root cause analysis system. The system includes an alarm analysis unit 1, an AI fault diagnosis agent unit 2, and a fault diagnosis review unit 3. Among them, the alarm analysis unit 1 is used for unified processing of server alarms; fault root cause location and correlation analysis of alarms are carried out in different alarm segment dimensions. The AI fault diagnosis agent unit 2 is responsible for in-depth analysis and intelligent reasoning of alarm data. The fault diagnosis review unit 3 reviews the diagnosis results of the AI agent. Figure 2 is a flowchart of the method for analyzing the root cause of server failures according to an embodiment of the present invention, as Figure 2 shown. The process includes the following steps: Step S201, regularly collect server alarm information and server log data, and determine whether the alarm information belongs to the first level or the second level. For details, please refer to Figure 1 step S101 of the embodiment shown, which will not be elaborated here.

[0040] Step S202, when the alarm information belongs to the first level, directly perform root cause analysis of the failure and provide repair suggestions based on the first-level information.

[0041] Specifically, as Figure 4 shown in the first workflow, the first level is a single first fault segment dimension; the server log includes multiple component alarm logs. The above step S202 includes: Step S2021, when the alarm information is a single first fault segment dimension, determine the root cause component corresponding to the single first fault segment dimension, and obtain the first component alarm log corresponding to the root cause component from multiple component alarm logs.

[0042] Specifically, a single first fault segment dimension is a single minor alarm.

[0043] When the alarm analysis unit 1 receives a single minor alarm, determine the root cause component corresponding to the single minor alarm and the first component alarm log corresponding to the root cause component.

[0044] Step S2022, directly perform root cause location of the failure and provide repair suggestions based on the alarm log corresponding to the root cause component.

[0045] Specifically, based on the alarm log corresponding to the root cause component, following the principle of most reasonable and resource-saving use, do not enable the reasoning logic of the AI fault diagnosis agent unit 2, and directly call the fault diagnosis review unit 3 to perform root cause location of the failure and provide repair suggestions based on the alarm log of the root cause component.

[0046] Step S203: When the alarm information belongs to the second level, perform agent reasoning based on the alarm information and the second-level information, perform root cause analysis of the fault and provide repair suggestions based on the agent reasoning result, and at the same time review the root cause analysis of the fault and the repair suggestions, and output the final root cause analysis result of the fault and the final repair suggestions.

[0047] The method for root cause analysis of server faults provided in this embodiment customizes different agent reasoning strategies for different alarm scenarios (multiple first fault segment dimensions, single second fault segment dimension, multiple second fault segment dimensions, or mixed alarms). When there are alarms in multiple first fault segment dimensions, focus on component alarm logs and basic information reasoning to quickly handle common faults; when there is an alarm in a single second fault segment dimension, conduct in-depth analysis in combination with performance data, full logs, etc.; in the case of mixed alarm scenarios, comprehensively integrate multi-source data reasoning to ensure that the root cause of complex faults can be accurately located, greatly improving the adaptability to various fault scenarios and the diagnostic accuracy.

[0048] In this embodiment, a method for root cause analysis of server faults is provided, which can be used in a server fault root cause analysis system. The system includes an alarm analysis unit 1, an AI fault diagnosis agent unit 2, and a fault diagnosis review unit 3. Among them, the alarm analysis unit 1 is used for unified processing of server alarms; perform root cause location and correlation analysis of faults alarmed in different alarm segment dimensions. The AI fault diagnosis agent unit 2 is responsible for in-depth analysis and intelligent reasoning of alarm data. The fault diagnosis review unit 3 reviews the diagnosis results of the AI agent. Figure 3 is a flowchart of the method for root cause analysis of server faults according to an embodiment of the present invention, as Figure 3 shown, this process includes the following steps: Step S301: Regularly collect server alarm information and server log data, and determine whether the alarm information belongs to the first level or the second level. For details, please refer to Figure 2 Step S201 of the embodiment shown here, which will not be elaborated here.

[0049] Step S302: When the alarm information belongs to the first level, directly perform root cause analysis of the fault and provide repair suggestions based on the first-level information. For details, please refer to Figure 2 Step S202 of the embodiment shown here, which will not be elaborated here.

[0050] Step S303: When the alarm information belongs to the second level, perform agent reasoning based on the alarm information and the second-level information, perform root cause analysis of the fault and provide repair suggestions based on the agent reasoning result, and at the same time review the root cause analysis of the fault and the repair suggestions, and output the final root cause analysis result of the fault and the final repair suggestions.

[0051] Specifically, as Figure 4The second to fourth workflow diagrams shown. The second level includes multiple first fault segment dimensions, a single second fault segment dimension, and a mixture of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension; the server log also includes full logs; the above step S203 includes: Step S3031, when the alarm information is multiple first fault segment dimensions, perform first agent reasoning based on multiple component alarm logs, the pre-acquired device serial number, and the alarm times of multiple first fault segment dimensions to obtain a first agent reasoning result, and perform first fault root cause analysis and provide a first repair suggestion based on the first agent reasoning result.

[0052] Specifically, the AI fault diagnosis agent unit includes a data input interface, a multi-modal correlation analysis engine root cause location module, and a repair suggestion generation module.

[0053] As Figure 4 In the third workflow diagram shown, the first fault segment dimension is a minor alarm. The AI fault diagnosis agent unit supports multi-protocol adapters and receives structured data packets from the alarm analysis unit 1 through the data input interface, including: device SN code (i.e., device serial code, used to associate the device topology relationship in CMDB), an alarm timestamp sequence accurate to the millisecond level (T1...Tn), a normalized encoding of the component alarm log (converting the manufacturer-specific error code according to the IPC-9592B standard), a server performance data matrix (5-second granularity time series metrics of CPU / memory / disk / network), and a semantic feature vector of the full log (768-dimensional embedding representation pre-processed by the BERT model).

[0054] When receiving N minor alarms, submit the device SN, the alarm times T1 corresponding to multiple minor alarms, the component alarm logs WLG1,..., the component alarm log WLGN data to the AI fault diagnosis agent unit 2 for correlation analysis, and find the correlation and fault root cause between multiple minor alarms, and give a repair suggestion.

[0055] In some alternative embodiments, the above step S3031 includes: Step a1, when the alarm information is multiple first fault segment dimensions, based on multiple component alarm logs, the pre-acquired device serial number, and the alarm times of multiple first fault segment dimensions, and in combination with a preset hardware dependency rule library, construct a first inference knowledge graph using an improved constraint-based non-parametric algorithm.

[0056] In step a1, the multi-modal correlation analysis engine uses a three-level analysis architecture to implement alarm correlation determination, including: 1) Fast filtering at the rule layer: Preset a hardware dependency rule library (e.g., "The rule for triggering degradation when 2 disks in the same RAID group are alarmed"), and the example is as follows: def DetectRAID5Degradation(alarm message list): # Filter out all disk error alarms Disk error alarms = [alarm for alarm in alarm message list if alarm == 'Disk error'] # Check if there are at least 2 disk errors and they belong to the same RAID group if len(disk error alarms) >= 2 and BelongToTheSameDiskArrayGroup(disk error alarms): return True, "The RAID5 array has degraded. It is recommended to back up the data immediately and replace the faulty disks" else: return False, "No RAID5 degradation detected."

[0057] 2) Temporal causal discovery layer, applying the improved PCMCI+ algorithm (Partial Convergent Cross-Mapping, a constraint-based non-parametric algorithm): Based on the standard PC (Peter-Clark, PC) algorithm, add device topology constraint conditions, then the dynamic adjustment formula for the time lag window: τ = max(10s, 0.2 × device physical distance coefficient).

[0058] Output a directed acyclic graph (DAG) of fault propagation, and the edge weight represents the causal confidence.

[0059] 3) Knowledge graph reasoning layer, constructing an operation and maintenance knowledge graph containing more than 3 million nodes, and the implementation process is as follows: Based on path queries in Neo4j (a graph database) (e.g., "Power module failure → CPU frequency reduction → Application timeout"), and use a graph attention network (Graph Attention Network, GAT) to predict potential propagation paths, and finally construct the first inference knowledge graph.

[0060] Step a2, based on the first inference knowledge graph, find the relevance and fault root causes of multiple first fault segment dimensions, and dynamically score the fault root causes to obtain the first intelligent agent inference results of the existence of a unique root cause component and the non-existence of a root cause component.

[0061] Specifically, the root cause localization module searches for the correlations and root causes of multiple first fault segment dimensions based on the first inference knowledge graph, and dynamically scores the root causes of the faults. The root cause score formula is as follows: Root cause score = 0.4 × Causal graph in-degree + 0.3 × Historical matching degree + 0.2 × Current health degree + 0.1 × Topological centrality.

[0062] Among them, the causal graph in-degree refers to the in-degree of a node in the fault causal relationship graph (i.e., the first inference knowledge graph), which indicates how many other nodes point to it. The higher the in-degree, the more the node is affected by other factors, and the more likely it is to be the core point of fault propagation, with a weight of 0.4.

[0063] The historical matching degree refers to the matching degree between the current fault characteristics and historical known fault cases, with a weight of 0.3.

[0064] The current health degree refers to the degree to which the state indicators of the component itself deviate from the normal range when the fault occurs, with a weight of 0.2.

[0065] The topological centrality refers to the measure of the importance of a component in the system topological structure, with a weight of 0.1.

[0066] The root cause localization module also innovatively introduces a counterfactual verification mechanism, including: removing the suspected root cause in the digital twin environment, verifying whether the remaining system alarms disappear, and finally outputting an interpretability report to obtain the first agent inference results of the components with a unique root cause and the components without a root cause.

[0067] Step a3, when it is the first agent inference result of the component with a unique root cause, perform the first fault root cause analysis on the unique root cause component and generate the first repair suggestion; when it is the first agent inference result of the component without a root cause, determine that there is no correlation among multiple first fault segment dimensions, perform the corresponding first fault root cause analysis on multiple first fault segment dimensions respectively, and generate the corresponding first repair suggestions based on the corresponding first fault root cause analysis.

[0068] Specifically, based on the analysis of the AI fault diagnosis agent unit 2, for the component with a unique root cause, call the fault diagnosis review unit 3 to perform the fault root cause analysis and provide repair suggestions for the unique root cause component. Based on the analysis of the AI fault diagnosis agent unit 2, for the component without a root cause, it indicates that there is no correlation among the minor alarms of multiple components, and call the fault diagnosis review unit 3 to perform the fault root cause analysis and provide repair suggestions respectively, that is, perform the fault root cause analysis and provide repair suggestions for each component respectively.

[0069] The repair suggestions in step a3 are provided by the repair suggestion generation module, which generates a hierarchical recommendation strategy according to the urgency of the alarm information, as shown in Table 1 below: Table 1 Hierarchical Recommendation Strategy

[0070] In Table 1, in the scenario of fault diagnosis and root cause analysis, the confidence level is a key indicator used to measure the reliability and credibility of the diagnosis result, usually expressed as a probability value (such as 95%). It reflects the degree of confidence of the diagnosis model or algorithm in the "correctness of the current judged fault root cause or conclusion".

[0071] The fault diagnosis system (such as the inference model of the AI intelligent agent unit) analyzes multi-source information such as alarm logs, performance data, and topology relationships, outputs one or more possible fault root causes, and assigns a confidence score to each root cause. The higher the confidence level, the stronger the "certainty" of the system's judgment on the root cause. For example: when a certain server has a RAID Degradation alarm, the system judges that the root cause is "disk failure" with a confidence level of 98% based on information such as disk error count, RAID controller log, and historical fault patterns, indicating a very high credibility of this conclusion.

[0072] In this embodiment, for the alarm scenario of multiple minor alarms, by combining the preset hardware dependency rule library and the improved constraint-based non-parametric algorithm to construct the first inference knowledge graph, the logical relationship between the alarm logs of multiple components can be quickly sorted out. Through the existing hardware dependency rules in the rule library and using the algorithm constraint conditions, the scattered alarm information can be efficiently transformed into a structured knowledge graph, intuitively presenting the potential connections between multiple minor alarms, avoiding analyzing each alarm in isolation, and greatly improving the efficiency of fault correlation mining. Search for the fault root cause based on the constructed knowledge graph and implement dynamic scoring. Considering various factors such as the in-degree of the causal graph and historical matching degree, quantitatively evaluate each possible fault root cause. The multi-dimensional scoring mechanism can screen out the real fault root cause more comprehensively and accurately than the single standard judgment. For the two inference results of the component with a unique root cause and the component without a root cause, corresponding processing strategies are formulated respectively. When there is a component with a unique root cause, directly conduct in-depth analysis on it and generate repair suggestions to quickly solve the problem; when there is no component with a root cause, that is, when multiple alarms are not related, independently analyze each first fault segment dimension to avoid misjudging independent faults as cascaded faults, ensuring that different fault scenarios can be reasonably and effectively processed.

[0073] Step S3032: When the alarm information is a single second fault segment dimension, determine the root cause component corresponding to the single second fault segment dimension, obtain the server performance data and the second component alarm log corresponding to the root cause component, perform second agent reasoning based on the second component alarm log, the server performance data, the full volume of logs, and the device serial number and the alarm time of the single second fault segment dimension obtained in advance, and perform second fault root cause analysis and provide second repair suggestions based on the second agent reasoning result.

[0074] In some alternative embodiments, such as Figure 4 in the second workflow shown in, the above step S3032 includes: Step b1: When the alarm information is a single second fault segment dimension, based on the single second fault segment dimension, determine the root cause component corresponding to the single second fault segment dimension, and obtain the server performance data within a preset time period before the alarm information and the second component alarm log filtered out from multiple component alarm logs corresponding to the root cause component.

[0075] Specifically, the second fault segment dimension is recorded as a severe alarm.

[0076] When the alarm analysis unit 1 receives a single severe alarm, determine the corresponding root cause component based on the single severe alarm, retrieve the performance data and server log data within 48 hours forward based on the received alarm time, and filter out the second component alarm log corresponding to the root cause component from the server log data. Submit the device SN alarm time T1, the second component alarm log WLG1, the performance data PD, and the server other log data LGD (full volume of logs) to the AI fault diagnosis agent unit 2 for correlation analysis.

[0077] Step b2: Based on the second component alarm log, the server performance data, the full volume of logs, the device serial number and the alarm time obtained in advance, and in combination with the preset hardware dependency relationship rule library, construct a second inference knowledge graph using an improved constraint-based non-parametric algorithm.

[0078] Specifically, the AI fault diagnosis agent unit 2 constructs a second inference knowledge graph based on the second component alarm log, the server performance data, the full volume of logs, the device serial number and the alarm time obtained in advance, and in combination with the preset hardware dependency relationship rule library using an improved constraint-based non-parametric algorithm. The details of constructing the second inference knowledge graph are the same as the above step a1 and will not be elaborated here.

[0079] Step b3: Based on the second inference knowledge graph, find the direct root cause component and correlation factors corresponding to the single second fault segment dimension, and obtain the second agent inference result with a direct root cause component or correlation factors.

[0080] Specifically, the correlation factors include other component performances, component anomalies that have not triggered alarms (only logging), business pressure, and other factors.

[0081] Step b4, when it is the second agent inference result with a direct root cause component, perform a second root cause analysis on the direct root cause component and generate a second repair suggestion; when it is the second agent inference result with correlation factors, generate a root cause analysis result and a second repair suggestion based on the correlation factors.

[0082] Specifically, for a severe alarm of a direct root cause component, call the fault diagnosis review unit 3 to locate the root cause of the fault and provide a repair suggestion. For a severe alarm with correlation factors, such as other component performances, component anomalies that have not triggered alarms (only logging), business pressure, and other factors, the AI fault diagnosis agent unit provides an analysis result and a repair suggestion.

[0083] Through this embodiment, multi-source data such as server performance data, filtered second component alarm logs, and full logs within a preset time period before obtaining the alarm information are obtained. Combining the device serial number and the alarm time, relevant information before and after the fault is comprehensively integrated. These data reflect the server operation status from different dimensions. By combining the preset hardware dependency rule library and the improved algorithm to construct a knowledge graph, the potential connections between data can be deeply mined, avoiding missing key clues, and providing a rich and comprehensive data basis for accurately analyzing the fault. For a single severe alarm, use the constructed second inference knowledge graph to find the direct root cause component and correlation factors. When there is a direct root cause component, the fault source can be directly locked; if there are only correlation factors, it is also possible to comprehensively analyze from aspects such as performance fluctuations and component associations to find the deep fault cause hidden under the surface phenomenon, breaking the limitation of analyzing only from a single alarm log and improving the diagnostic ability for complex faults.

[0084] Step S3033, when the alarm information is a mixed alarm of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, obtain the server performance data, and perform a third agent inference based on multiple component alarm logs, server performance data, full logs, and the previously obtained device serial number, multiple second fault segment dimensions or the alarm time of the mixed alarm, and provide a third root cause analysis and a third repair suggestion based on the third agent inference result.

[0085] In some alternative embodiments, such as Figure 4 shown in the fourth workflow, the above step S3033 includes: Step c1, when the alarm information is a mixed alarm of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, obtain the server performance data within a preset time period before the alarm information.

[0086] Specifically, when N severe alarms or a mixture of N minor and severe alarms are received, based on the received alarm time, retrieve the performance data and server log data within 48 hours forward, and submit the device SN, alarm time T1, component alarm log WLG1, …, component alarm log WLGN data, performance data PD, and other server log data LGD (full logs) to the AI fault diagnosis agent unit 2 for correlation analysis.

[0087] Step c2: Based on multiple component alarm logs, server performance data, full logs, and the device serial number, multiple second fault segment dimensions, or the alarm time of a mixed alarm obtained in advance, and in combination with a preset hardware dependency rule library, construct a third inference knowledge graph using an improved constraint-based nonparametric algorithm.

[0088] Specifically, the AI fault diagnosis agent unit 2 constructs a third inference knowledge graph based on multiple component alarm logs, server performance data, full logs, and the device serial number, multiple severe alarms or the alarm time of a mixed alarm obtained in advance, and in combination with a preset hardware dependency rule library, using an improved constraint-based nonparametric algorithm. The details of constructing the third inference knowledge graph are the same as those in step a1 above and will not be elaborated here.

[0089] Step c3: Based on the third inference knowledge graph, find the relevance and fault root cause between multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, and obtain a third intelligent agent inference result of the existence of a unique root cause component, the existence of a root cause component, or the existence of a correlation factor.

[0090] Specifically, the correlation factors include factors such as the performance of other components, anomalies of components that have not triggered alarms (only logging), and business pressure.

[0091] Step c4: When the third intelligent agent inference result is the existence of a unique root cause component, conduct a third fault root cause analysis on the unique root cause component and generate a third repair suggestion; when the third intelligent agent inference result is the non-existence of a root cause component, determine that there is no relevance between multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, conduct corresponding third fault root cause analyses for multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension respectively, and generate corresponding third repair suggestions based on the corresponding third fault root cause analyses; when the third intelligent agent inference result is the existence of a correlation factor, find the relevance of the mixed alarm between multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension based on the correlation factor, and generate a relevance fault root cause analysis result and a third repair suggestion based on the relevance.

[0092] Specifically, for the existing correlation factors, such as the performance of other components, the anomalies of components that have not triggered alarms (only logging), business pressure, etc., the AI fault diagnosis agent unit directly searches for the correlation between multiple alarms and the root cause of the fault. For the component with a clear single root cause, the fault diagnosis review unit 3 is called to conduct root cause analysis of the root cause component and provide repair suggestions. For the alarm information without a root cause component, it indicates that there is no correlation between multiple component alarms, and the fault diagnosis review unit 3 is called respectively to conduct root cause analysis and provide repair suggestions.

[0093] In step S303, it involves the collaborative work process of the AI fault diagnosis agent unit 2 and the fault diagnosis review unit 3. First, the AI fault diagnosis agent unit 2 infers the original diagnosis result, and then the fault diagnosis review unit 3 uses the composite marking system for secondary review to obtain the final root cause analysis result of the fault and provide repair suggestions. When the review result is correct, the root cause analysis result of the fault and the repair suggestions are added to the training set for continuous learning; when it is partially correct, partial result correction is required to trigger the fine-tuning process; when the original diagnosis result inferred by the AI fault diagnosis agent unit 2 is completely rejected, the expert intervention process is started.

[0094] Through this embodiment, in the face of complex scenarios such as multiple severe alarms or mixed alarms, by integrating multi-source information such as the alarm logs of multiple components, server performance data, and full-scale logs, and combining the device serial number and alarm time, a third inference knowledge graph is constructed, which can effectively sort out the logical relationship between complex alarms. Based on the third inference knowledge graph, not only can the direct root cause component of the fault be found, but also the existing correlation factors can be mined. By combining the preset hardware dependency rule library and the improved algorithm, the fault correlation is comprehensively analyzed from multiple dimensions. Whether it is a single root cause component or a fault caused by the mutual correlation between multiple components, it can be accurately located. For the situation where there is no obvious root cause component, the system can also independently analyze each alarm to prevent misjudging uncorrelated faults as cascaded faults, greatly improving the accuracy of root cause location in complex fault scenarios.

[0095] The fault diagnosis review unit 3 is used to review the root cause analysis of faults and repair suggestions. The fault diagnosis review unit 3 is a key quality assurance link in the system fault diagnosis process. Relying on the expert experience knowledge base accumulated in the company's many years of operation and maintenance practice and the verified fault diagnosis rule system, it conducts an authoritative review of the diagnosis results of the AI intelligent agent. This unit conducts multi-dimensional verification on the fault root cause components and diagnosis conclusions initially determined by the AI fault diagnosis intelligent agent unit 2 through the built-in industry knowledge graph, typical fault case library, and equipment operation and maintenance historical data, combined with the decision tree and verification algorithm formed by expert experience. During the review process, the system will focus on evaluating the feasibility of the diagnosis results (such as on-site operability), accuracy (matching degree with historical cases), and the rationality of the repair suggestions (meeting the requirements of operation and maintenance specifications). For doubtful diagnosis conclusions, the review unit will start a correction mechanism or trigger a manual expert consultation process to ensure that the finally output root cause analysis of faults has a high degree of credibility. The above step S303 also includes: Step S3034, based on the preset expert experience knowledge base and preset historical fault cases, uses a preset multi-dimensional verification matrix and contradiction detection algorithm to review the root cause analysis of faults and repair suggestions.

[0096] Specifically, the preset expert experience knowledge base is a knowledge base constructed based on a multi-source knowledge fusion architecture. The multi-source knowledge fusion architecture also includes structured fault diagnosis rules, such as hardware diagnosis rules (such as disk bad track mode), software exception modes (such as memory leak characteristics), and physical topology constraints of the computer room.

[0097] The Faiss vector database is used to store the characteristics of historical fault cases (dimension = 256). The similarity calculation threshold is set to be greater than 85%.

[0098] In order to update historical fault cases, a dynamic knowledge update method is used to obtain incremental knowledge from the following channels every month, as shown in Table 2 below: Table 2 Incremental Knowledge Acquisition Channels

[0099] The preset multi-dimensional verification matrix is a five-layer verification matrix, as shown in Table 3 below: Table 3 Five-Layer Verification Matrix

[0100] In step S3034, when the expert experience knowledge base is called, the review logic is monitored, and the rule matching degree statistics operation is performed to update the expert experience knowledge base. The monitoring content includes the number and type of expert rules that are hit (such as hardware rules, software rules, topology rules) and the alarm characteristics of missed rules (used to identify blind spots in the rule base). For example, when the AI ​​fault diagnosis intelligent unit 2 diagnoses a "memory fault", the fault diagnosis review unit 3 calls the hardware rule base in the expert experience knowledge base. If the rule "the controller is given priority for dual memory alarms in the same slot" is not hit, the feature is recorded and a reminder to update the expert experience knowledge base is triggered.

[0101] By monitoring the alarm characteristics of missed rules, scenarios that are missing or insufficiently covered in the expert rule base (such as new failure modes and alarm combinations under special topological structures) can be accurately located to avoid diagnostic deviations caused by outdated rule bases. It can identify rule blind spots and improve the accuracy and completeness of the expert rule base. After triggering the rule base update reminder, it can promote the operation and maintenance team to supplement rules in a targeted manner (such as adding "cross-rack PSU cascading failure" related logic) to ensure that the rule base evolves synchronously with the actual fault scenario. Statistics on the number and type of hit rules (such as hardware rules accounting for 70% and topology rules accounting for 30%) can intuitively display the fit between AI diagnostic results and expert experience, and provide a clear reference basis for manual review. Through rule matching data, high-risk miss scenarios (such as multi-component alarms under complex topologies) can be prioritized to avoid missing key clues due to experience limitations during manual review.

[0102] Step S3035, when a conflict occurs between the fault root cause analysis and the expert experience knowledge base during the review process, a preset priority level is used to determine whether to select the fault root cause analysis or the manual consultation method to resolve the conflict.

[0103] Specifically, the preset priority level is the priority level of the fault root cause analysis result.

[0104] When AI output conflicts with the rules in the expert experience knowledge base, Figure 5 The process shown resolves conflicts. Specifically, when a conflict is detected, it is determined whether the priority of the root cause analysis result is greater than or equal to level 3. If it is greater than or equal to level 3, a manual consultation method is selected to resolve the conflict. Otherwise, the root cause analysis result obtained by the AI ​​fault diagnosis intelligent unit is used as the final analysis result and marked.

[0105] Step S3036, correcting the confidence of the fault root cause analysis based on the result after the conflict is resolved, outputting the final fault root cause analysis result, and generating a final repair suggestion based on the final fault root cause analysis result.

[0106] Specifically, the formula for adjusting the confidence level of the fault root cause analysis result output by AI is: Final confidence = α × AI confidence + (1-α) × expert matching degree; Among them, α is calculated dynamically, α= 1 / (1 + exp(-(number of similar cases in the expert database-5))), and the expert matching degree refers to the matching degree between the current fault characteristics and the known fault cases in the expert knowledge base, usually expressed as a value between 0 and 1 (0 means no match at all, 1 means a complete match), reflecting the support of expert experience for the current fault diagnosis, and is the basis for the fault diagnosis review unit (such as manual or rule engine) to correct the AI ​​output results. The number of expert similar cases refers to the number of historical cases in the expert knowledge base that are similar to the current fault characteristics, where "similar" must be based on preset feature matching rules (such as the same alarm type, the same component, and similar topological locations).

[0107] In order to better correct the confidence of the fault root cause analysis, the continuous learning interface is used for continuous learning. The feedback data format is as follows: json { "original_ai": {"root_cause": "...", "confidence": 0.92}, / / AI original diagnosis result, / / root cause (such as: power module voltage drop) / / confidence (probability value between 0-1) "adjusted_result": {"factor": "topology", "delta_confidence": -0.15}, / / Adjusted result / / Adjustment factors (such as topology constraints, expert rule conflicts, etc.) / / Confidence adjustment value (can be positive or negative, the principal and interest means a 15% reduction) "final_decision": {"action": "replace_psu", "executor": "human / AI"} / / Final decision / / Repair action (such as replacing the power module) / / Executor (human means manual execution, AI means automatic execution) }.

[0108] This continuous learning interface converts empirical knowledge such as manual review results and environmental constraints into quantifiable model training data through a standardized feedback data format, achieving the following value: 1) Continuously optimize AI models: By accumulating factor (adjustment factor) and delta_confidence (quantified adjustment range) data, identify model defects (such as insufficient consideration of topological constraints) and optimize the algorithm in a targeted manner; 2) Improve diagnostic transparency: Record the complete logical chain from "AI initial judgment" to "final decision" for easy traceability and auditing; 3) Support hybrid decision-making mode: Clearly distinguish the boundaries between manual and automatic execution to balance automation efficiency and manual reliability.

[0109] The server fault root cause analysis method provided in this embodiment is based on a preset expert experience knowledge base and a preset structured fault diagnosis rule base, and uses a preset multi-dimensional verification matrix and a contradiction detection algorithm to review the fault root cause analysis and repair suggestions, forming a dual verification mechanism of AI analysis + expert system review. It not only gives full play to the fast analysis advantage of AI, but also ensures the accuracy of the diagnostic results through the expert system, effectively solving the misjudgment problem that may exist in a single AI system. This technology significantly reduces the complexity of operation and maintenance work, liberates operation and maintenance personnel from tedious fault troubleshooting, enables the operation and maintenance mode of the data center to transform from manual dominance to intelligent automation, and provides a strong guarantee for the efficient and stable operation of the data center while improving operation and maintenance reliability.

[0110] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0111] The embodiment of the present application also provides a server fault root cause analysis device, which is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation by hardware, or a combination of software and hardware is also possible and contemplated.

[0112] This embodiment provides a server fault root cause analysis device, as Figure 6 shown, including: An alarm information collection and judgment module 601, configured to periodically collect server alarm information and server log data, and judge whether the alarm information belongs to the first level or the second level; A first fault root cause analysis module 602, configured to directly perform fault root cause analysis and provide repair suggestions based on the first level information when the alarm information is at the first level; A second fault root cause analysis and review module 603, configured to perform agent reasoning based on the alarm information and the second level information when the alarm information is at the second level, provide fault root cause analysis and repair suggestions based on the agent reasoning result, and at the same time review the fault root cause analysis and repair suggestions, and output the final fault root cause analysis result and the final repair suggestion.

[0113] In some alternative embodiments, the first level is a single alarm severity level including a first fault segment dimension; the server log includes multiple component alarm logs; the first root cause analysis module 602 includes: A root cause component determination and alarm log acquisition unit, configured to determine a root cause component corresponding to a single first fault segment dimension and acquire a first component alarm log corresponding to the root cause component when the alarm information is a single first fault segment dimension; A fault root cause location and repair suggestion providing unit, configured to directly perform fault root cause location and provide repair suggestions based on the alarm log corresponding to the root cause component.

[0114] In some alternative embodiments, the second level includes multiple first fault segment dimensions, a single second fault segment dimension, and multiple second fault segment dimensions or a mixed alarm of multiple first fault segment dimensions and at least one second fault segment dimension; the server log further includes a full log; the second root cause analysis and review module 603 includes: A first fault root cause analysis and first repair suggestion providing unit, configured to perform first agent reasoning based on multiple component alarm logs, a pre-acquired device serial number, and the alarm time of multiple first fault segment dimensions when the alarm information is multiple first fault segment dimensions, obtain a first agent reasoning result, and perform first fault root cause analysis and provide a first repair suggestion based on the first agent reasoning result.

[0115] A second fault root cause analysis and second repair suggestion providing unit, configured to determine a root cause component corresponding to a single second fault segment dimension and acquire server performance data and a second component alarm log corresponding to the root cause component when the alarm information is a single second fault segment dimension, perform second agent reasoning based on the second component alarm log, server performance data, full log, and a pre-acquired device serial number and the alarm time of the single second fault segment dimension, and perform second fault root cause analysis and provide a second repair suggestion based on the second agent reasoning result.

[0116] A third fault root cause analysis and third repair suggestion providing unit, configured to acquire server performance data when the alarm information is multiple second fault segment dimensions or a mixed alarm of multiple first fault segment dimensions and at least one second fault segment dimension, perform third agent reasoning based on multiple component alarm logs, server performance data, full log, and a pre-acquired device serial number, the alarm time of multiple second fault segment dimensions or the mixed alarm, and perform third fault root cause analysis and provide a third repair suggestion based on the third agent reasoning result.

[0117] A review unit, configured to review the root cause analysis and repair suggestions of faults by using a preset multi-dimensional verification matrix and a contradiction detection algorithm based on a preset expert experience knowledge base and preset historical fault cases.

[0118] A conflict resolution unit, configured to, when a conflict occurs between the root cause analysis of a fault and the expert experience knowledge base during the review process, use a preset priority level determination to select the root cause analysis of the fault or select a manual consultation method to resolve the conflict.

[0119] A correction unit, configured to correct the confidence level of the root cause analysis of the fault based on the result after conflict resolution, output the final root cause analysis result of the fault, and generate a final repair suggestion based on the final root cause analysis result of the fault.

[0120] In some alternative embodiments, the first root cause analysis and first repair suggestion providing unit includes: A first inference knowledge graph construction subunit, configured to, when the alarm information is of multiple first fault segment dimensions, construct a first inference knowledge graph by using an improved constraint-based non-parametric algorithm based on multiple component alarm logs, the device serial number obtained in advance, and the alarm times of multiple first fault segment dimensions, and in combination with a preset hardware dependency relationship rule base; A first inference subunit, configured to find the relevance and root cause of the fault of multiple first fault segment dimensions based on the first inference knowledge graph, and perform dynamic scoring on the root cause of the fault to obtain a first intelligent agent inference result of a component with a unique root cause and a component without a root cause; A first root cause analysis and repair suggestion providing subunit, configured to, when the first intelligent agent inference result is a component with a unique root cause, perform a first root cause analysis on the component with the unique root cause and generate a first repair suggestion; when the first intelligent agent inference result is a component without a root cause, determine that the multiple first fault segment dimensions have no relevance, perform corresponding first root cause analysis on the multiple first fault segment dimensions respectively, and generate corresponding first repair suggestions based on the corresponding first root cause analysis.

[0121] In some alternative embodiments, the second root cause analysis and second repair suggestion providing unit includes: A performance data and second component alarm log acquisition subunit, configured to, when the alarm information is of a single second fault segment dimension, determine a root cause component corresponding to the single second fault segment dimension based on the single second fault segment dimension, and acquire the server performance data within a preset time period before the alarm information and the second component alarm log corresponding to the root cause component screened from multiple component alarm logs.

[0122] The second inference knowledge graph construction subunit is used to construct a second inference knowledge graph based on the second component alarm log, server performance data, full volume log, and pre-acquired device serial number and alarm time, and by combining a preset hardware dependency relationship rule base and using an improved constraint-based non-parametric algorithm.

[0123] The second inference subunit is used to find the direct root cause component and correlation factors corresponding to a single second fault segment dimension based on the second inference knowledge graph, and obtain a second intelligent agent inference result with a direct root cause component or correlation factors existing. The second fault root cause analysis and repair suggestion providing subunit is used to, when the second intelligent agent inference result has a direct root cause component, perform a second fault root cause analysis on the direct root cause component and generate a second repair suggestion; when the second intelligent agent inference result has correlation factors, generate a fault root cause analysis result and a second repair suggestion based on the correlation factors.

[0124] In some optional implementation manners, the third fault root cause analysis and third repair suggestion providing unit includes: The performance data acquisition subunit is used to obtain the server performance data within a preset time period before the alarm information when the alarm information is a mixed alarm of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension. The third inference knowledge graph construction subunit is used to construct a third inference knowledge graph based on multiple component alarm logs, server performance data, full volume log, and pre-acquired device serial number, multiple second fault segment dimensions or the alarm time of the mixed alarm, and by combining a preset hardware dependency relationship rule base and using an improved constraint-based non-parametric algorithm. The third inference subunit is used to find the relevance and fault root cause of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension based on the third inference knowledge graph, and obtain a third intelligent agent inference result with a unique root cause component, a root cause component or correlation factors existing. The third failure root cause analysis and repair suggestion providing subunit is used to perform third failure root cause analysis on the unique root cause component and generate third repair suggestions when the third agent inference result has a unique root cause component; when the third agent inference result has no root cause component, it determines that multiple second failure segment dimensions or multiple first failure segment dimensions have no relevance to at least one second failure segment dimension, performs corresponding third failure root cause analysis on multiple second failure segment dimensions or multiple first failure segment dimensions and at least one second failure segment dimension respectively, and generates corresponding third repair suggestions based on the corresponding third failure root cause analysis; when the third agent inference result has relevant factors, it searches for the relevance of the mixed alarms of multiple second failure segment dimensions or multiple first failure segment dimensions and at least one second failure segment dimension based on the relevant factors, and generates a relevance failure root cause analysis result and third repair suggestions based on the relevance.

[0125] For the description of the features in the corresponding embodiment of the server failure root cause analysis device, reference can be made to the relevant description in the corresponding embodiment of the server failure root cause analysis method, which will not be elaborated here one by one.

[0126] An embodiment of the present application further provides an electronic device, as Figure 7 shown, including a memory 10 and a processor 20. A computer program is stored in the memory 10, and the processor 20 is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the server failure root cause analysis method.

[0127] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-mentioned embodiments of the server failure root cause analysis method when running.

[0128] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical disks and other various media that can store computer programs.

[0129] An embodiment of the present application further provides a computer program product. The above-mentioned computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-mentioned embodiments of the server failure root cause analysis method.

[0130] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and the computer program, when executed by a processor, implements the steps in any of the above-described embodiments of the server fault root cause analysis method.

[0131] Those skilled in the art can further realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0132] The above has introduced in detail a server fault root cause analysis method, device, electronic device, and storage medium provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for analyzing the root cause of server failures, characterized in that, Including: Regularly collect server alarm information and server log data, and determine whether the alarm information belongs to the first level or the second level; When the alarm information belongs to the first level, directly perform root cause analysis of the fault and provide repair suggestions based on the first-level information; When the alarm information belongs to the second level, perform agent reasoning based on the alarm information and the second-level information, and perform root cause analysis of the fault and provide repair suggestions based on the agent reasoning result. At the same time, review the root cause analysis of the fault and the repair suggestions, and output the final root cause analysis result of the fault and the final repair suggestions.

2. The server fault root cause analysis method according to claim 1, wherein The first level is a single first fault segment dimension; the server log includes multiple component alarm logs; When the alarm information belongs to the first level, directly perform root cause analysis of the fault and provide repair suggestions based on the first-level information, including: When the alarm information is a single first fault segment dimension, determine the root cause component corresponding to the single first fault segment dimension, and obtain the first component alarm log corresponding to the root cause component from the multiple component alarm logs; Directly perform fault root cause location and provide repair suggestions based on the alarm log corresponding to the root cause component.

3. The server fault root cause analysis method according to claim 2, wherein The second level includes multiple first fault segment dimensions, a single second fault segment dimension, and multiple second fault segment dimensions or a mixed alarm of multiple first fault segment dimensions and at least one second fault segment dimension; the server log also includes full logs; When the alarm information belongs to the second level, perform agent reasoning based on the alarm information and the second-level information, and perform root cause analysis of the fault and provide repair suggestions based on the agent reasoning result, including: When the alarm information is multiple first fault segment dimensions, perform first agent reasoning based on the multiple component alarm logs, the pre-obtained device serial number, and the alarm time of the multiple first fault segment dimensions to obtain a first agent reasoning result, and perform first fault root cause analysis and provide a first repair suggestion based on the first agent reasoning result; When the alarm information is a single second fault segment dimension, determine the root cause component corresponding to the single second fault segment dimension, and obtain server performance data and the second component alarm log corresponding to the root cause component. Perform second agent reasoning based on the second component alarm log, server performance data, full logs, the pre-obtained device serial number, and the alarm time of the single second fault segment dimension, and perform second fault root cause analysis and provide a second repair suggestion based on the second agent reasoning result; When the alarm information is multiple second fault segment dimensions or a mixed alarm of multiple first fault segment dimensions and at least one second fault segment dimension, obtain server performance data, and perform third agent reasoning based on the multiple component alarm logs, server performance data, full logs, the pre-obtained device serial number, and the alarm time of the multiple second fault segment dimensions or the mixed alarm, and perform third fault root cause analysis and provide a third repair suggestion based on the third agent reasoning result.

4. The server fault root cause analysis method according to claim 3, wherein, When the alarm information is multiple first fault segment dimensions, perform first agent reasoning based on the multiple component alarm logs, the device serial number obtained in advance, and the alarm times of the multiple first fault segment dimensions to obtain a first agent reasoning result, and perform first fault root cause analysis and provide a first repair suggestion based on the first agent reasoning result, including: When the alarm information is multiple first fault segment dimensions, based on the multiple component alarm logs, the device serial number obtained in advance, and the alarm times of the multiple first fault segment dimensions, and combined with a preset hardware dependency rule library, use an improved constraint-based non-parametric algorithm to construct a first inference knowledge graph; Based on the first inference knowledge graph, find the relevance and fault root cause of the multiple first fault segment dimensions, and perform dynamic scoring on the fault root cause to obtain a first agent reasoning result of the component with a unique root cause and the component without a root cause; When it is the first agent reasoning result of the component with a unique root cause, perform first fault root cause analysis on the unique root cause component and generate a first repair suggestion; When it is the first agent reasoning result of the component without a root cause, determine that the multiple first fault segment dimensions have no relevance, perform corresponding first fault root cause analysis on the multiple first fault segment dimensions respectively, and generate corresponding first repair suggestions based on the corresponding first fault root cause analysis.

5. The server fault root cause analysis method according to claim 3, wherein When the alarm information is a single second fault segment dimension, determine the root cause component corresponding to the single second fault segment dimension, and obtain the server performance data and the second component alarm log corresponding to the root cause component. Perform second agent reasoning based on the second component alarm log, the server performance data, the full amount of logs, the device serial number obtained in advance, and the alarm time of the single second fault segment dimension, and perform second fault root cause analysis and provide a second repair suggestion based on the second agent reasoning result, including: When the alarm information is a single second fault segment dimension, based on the single second fault segment dimension, determine the root cause component corresponding to the single second fault segment dimension, and obtain the server performance data within a preset time period before the alarm information and the second component alarm log filtered from the multiple component alarm logs corresponding to the root cause component; Based on the second component alarm log, the server performance data, the full amount of logs, the device serial number obtained in advance, and the alarm time, and combined with a preset hardware dependency rule library, use an improved constraint-based non-parametric algorithm to construct a second inference knowledge graph; Based on the second inference knowledge graph, find the direct root cause component and relevant factors corresponding to the single second fault segment dimension to obtain a second agent reasoning result of the existence of a direct root cause component or the existence of relevant factors; When it is the second agent reasoning result of the existence of a direct root cause component, perform second fault root cause analysis on the direct root cause component and generate a second repair suggestion; when it is the second agent reasoning result of the existence of relevant factors, generate a fault root cause analysis result and a second repair suggestion based on the relevant factors.

6. The server fault root cause analysis method according to claim 3, wherein When the alarm information is a mixed alarm of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, obtain server performance data, and perform third-agent reasoning based on multiple component alarm logs, server performance data, full logs, and the device serial number, multiple second fault segment dimensions or the alarm time of the mixed alarm obtained in advance, and perform third root cause analysis of the fault and provide a third repair suggestion, including: When the alarm information is a mixed alarm of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, obtain the server performance data within a preset time period before obtaining the alarm information; Based on multiple component alarm logs, server performance data, full logs, the device serial number obtained in advance, multiple second fault segment dimensions or the alarm time of the mixed alarm, and in combination with a preset hardware dependency rule library, construct a third inference knowledge graph by using an improved constraint-based non-parametric algorithm; Based on the third inference knowledge graph, find the relevance and fault root cause of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension, and obtain the third-agent reasoning result of the existence of a unique root cause component, the existence of a root cause component or the existence of a correlation factor; When the third-agent reasoning result is the existence of a unique root cause component, perform third root cause analysis on the unique root cause component and generate a third repair suggestion; When the third-agent reasoning result is the non-existence of a root cause component, determine that multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension have no relevance, perform corresponding third root cause analysis on multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension respectively, and generate corresponding third repair suggestions based on the corresponding third root cause analysis; When the third-agent reasoning result is the existence of a correlation factor, find the relevance of the mixed alarm of multiple second fault segment dimensions or multiple first fault segment dimensions and at least one second fault segment dimension based on the correlation factor, and generate a correlation fault root cause analysis result and a third repair suggestion based on the relevance; 7. The server fault root cause analysis method according to claim 1, wherein Review the fault root cause analysis and repair suggestion, and output the final fault root cause analysis result and the final repair suggestion, including: Based on a preset expert experience knowledge base and preset historical fault cases, review the fault root cause analysis and repair suggestion by using a preset multi-dimensional verification matrix and a contradiction detection algorithm; When a conflict occurs between the fault root cause analysis and the expert experience knowledge base during the review process, use a preset priority level determination to select the fault root cause analysis or select the method of manual consultation to resolve the conflict; Based on the result after conflict resolution, correct the confidence level of the fault root cause analysis, output the final fault root cause analysis result, and generate a final repair suggestion based on the final fault root cause analysis result.

8. A server fault root cause analysis device, characterized in that, Including: An alarm information collection and judgment module, which is used to regularly collect server alarm information and server log data, and judge whether the alarm information belongs to the first level or the second level; The first failure root cause analysis module is used to directly perform failure root cause analysis and provide repair suggestions based on the first-level information when the alarm information is at the first level; The second failure root cause analysis and review module is used to perform agent reasoning based on the alarm information and the second-level information when the alarm information is at the second level, perform failure root cause analysis and provide repair suggestions based on the agent reasoning result, and review the failure root cause analysis and repair suggestions at the same time, and output the final failure root cause analysis result and the final repair suggestion.

9. An electronic device, characterized in that, It includes: A memory for storing computer programs; A processor for implementing the steps of the server failure root cause analysis method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the server failure root cause analysis method according to any one of claims 1 to 7 when executed by a processor.

Citation Information

Patent Citations

  • Server fault diagnosis method and device, storage medium and electronic equipment

    CN118394561A

  • Adaptive operation and maintenance root cause positioning method and system based on deep learning

    CN119691576A

  • Intelligent operation and maintenance management and alarm system based on large model agent

    CN119847802A

  • System component failure diagnosis

    US20170097860A1

  • Root cause positioning method, and communication device and computer-readable storage medium

    WO2024012186A1

Cited By

  • Power distribution communication network fault positioning method, device, equipment and medium

    CN120811877A

  • A power distribution communication network fault locating method, device, equipment and medium

    CN120811877B

  • Abnormal problem checking method, checking system, computer system, storage medium and program product

    CN120821599A

  • Backup disaster recovery optimization method

    CN121037195A