Server fault diagnosis method and device

By establishing a binary tree regularized tree model and dynamic configuration of fault diagnosis rules, the problem of inaccurate fault reporting in server fault diagnosis is solved, and the accuracy and operation and maintenance efficiency of fault diagnosis are improved without upgrading the software.

CN115712538BActive Publication Date: 2025-08-26中移信息技术有限公司 +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110914115.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-10
Publication Date
2025-08-26
Estimated Expiration
2041-08-10

AI Technical Summary

Technical Problem

The existing server fault diagnosis method does not have high accuracy in reporting faults without upgrading the software, and may lead to improper handling of fault information by the operation and maintenance tools, affecting operation and maintenance efficiency.

Method used

By creating a fault diagnosis rule configuration file that is suitable for server faults, establish a binary tree regularized tree model, encode and sequence abnormal phenomena and log information, dynamically configure fault diagnosis rules, and improve the accuracy of fault information matching.

Benefits of technology

Without upgrading the software, the accuracy and maintainability of fault diagnosis are improved, the risk of misjudgment of fault information processing by operation and maintenance tools is reduced, and the operation and maintenance efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115712538B_ABST
    Figure CN115712538B_ABST
Patent Text Reader

Abstract

The present invention provides a server fault diagnosis method and device. The method comprises: obtaining a fault diagnosis rule configuration file adapted to the server fault; establishing a fault diagnosis rule model based on the fault diagnosis rule configuration file; and matching server abnormality information with the fault diagnosis rule model to determine fault information. The server fault diagnosis method and device provided by the present invention, by creating a fault diagnosis rule configuration file adapted to the server fault and dynamically adding or modifying specific fault types in the fault diagnosis rule configuration file, can update the fault diagnosis rule configuration file and change the parsing rules without upgrading the corresponding software, thereby improving the maintainability of the server and the accuracy of fault parsing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of wireless communications, and in particular to a method and device for diagnosing server faults. Background Art

[0002] When existing servers experience unadapted faults, the BMC (Baseboard Management Controller) can misjudge and report incorrect analysis results. Currently, there are short-term solutions for operations and maintenance personnel to use tools to adapt, mask, or handle misjudged reports, or delete related alarm codes. Long-term solutions involve equipment providers remediating these misreporting issues through software upgrades.

[0003] Among existing solutions, long-term solutions rely on software upgrades, but the release cycle is long and cannot meet short-term needs. Software upgrades may also require business interruption, inevitably causing losses and lacking flexibility. Short-term solutions rely on O&M tool adaptation, which places high demands on the software. Furthermore, by using command parameters to block relevant alarm codes, there is a risk that unadapted faults will not be correctly reported.

[0004] Therefore, it is of great significance to propose a method to improve the accuracy of fault reporting without upgrading the corresponding software. Summary of the Invention

[0005] The present invention provides a server fault diagnosis method and device, which are used to solve the technical problem in the prior art that the accuracy of fault reporting is low without upgrading software.

[0006] In a first aspect, the present invention provides a server fault diagnosis method, comprising:

[0007] Obtain the fault diagnosis rule configuration file adapted to the server fault;

[0008] Establishing a fault diagnosis rule model according to the fault diagnosis rule configuration file;

[0009] Match the server abnormal phenomenon information with the fault diagnosis rule model to determine the fault information.

[0010] In one embodiment, establishing a fault diagnosis rule model according to a fault diagnosis rule configuration file includes:

[0011] According to the fault diagnosis rule configuration file, a fault diagnosis rule model is established hierarchically using a regularized tree model of a binary tree;

[0012] Among them, in the fault diagnosis rule model, the abnormal phenomena of servers at each level are respectively regarded as sub-nodes of each layer of the root node.

[0013] In one embodiment, before matching the server abnormal phenomenon information with the fault diagnosis rule model, the method further includes:

[0014] Encode server abnormal phenomenon information according to each layer of sub-nodes in the fault diagnosis rule model;

[0015] Arrange the encoded server anomaly information in time series.

[0016] In one embodiment, matching server abnormality information with a fault diagnosis rule model to determine fault information includes:

[0017] Match the time-series abnormal phenomenon information with the sub-nodes of each layer of the fault diagnosis rule model in sequence according to the traversal order to determine the fault information;

[0018] The traversal order includes: starting from the top-level child node of the fault diagnosis rule model, and traversing to the lower-level child nodes in sequence.

[0019] In one embodiment, the server log information is encoded according to each layer of sub-nodes in the fault diagnosis rule model;

[0020] Arrange the encoded server log information in time series.

[0021] In one embodiment, when the fault information is not determined by matching the server abnormal phenomenon information with the fault diagnosis rule model, the fault information is determined by matching the time-series server log information with the fault diagnosis rule model.

[0022] In one embodiment, the fault diagnosis rule configuration file is a file in XML format.

[0023] In a second aspect, the present invention further provides a server fault diagnosis device, comprising:

[0024] Configuration generation module, used to obtain fault diagnosis rule configuration files adapted to server faults;

[0025] A model generation module is used to establish a fault diagnosis rule model according to a fault diagnosis rule configuration file;

[0026] The fault matching module is used to match the server abnormal phenomenon information with the fault diagnosis rule model to determine the fault information.

[0027] In a third aspect, the present invention further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-mentioned steps of the server fault diagnosis method when executing the computer program.

[0028] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements the steps of any of the above-mentioned server fault diagnosis methods when the computer program is executed by a processor.

[0029] The server fault diagnosis method, apparatus, electronic device, and storage medium provided by the present invention establish a fault diagnosis model by creating a fault diagnosis rule configuration file adapted to the server fault. Collected fault information is matched with the fault diagnosis model to determine the specific fault information, forming a set of fault diagnosis processes with dynamically configurable rules. By dynamically adding or modifying specific fault types in the fault diagnosis rule configuration file, the fault diagnosis rule configuration file can be updated and parsing rules can be changed without upgrading the corresponding software, thereby improving server maintainability and the accuracy of fault analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0031] Figure 1 A flow chart of a server fault diagnosis method provided by the present invention;

[0032] Figure 2 A schematic diagram of a binary tree model construction process according to an embodiment of the present invention;

[0033] Figure 3 A schematic diagram of a pre-order traversal process according to an embodiment of the present invention;

[0034] Figure 4 A schematic diagram of log format coding provided by one embodiment of the present invention;

[0035] Figure 5 A schematic diagram of an abnormal phenomenon processing flow provided by one embodiment of the present invention;

[0036] Figure 6 A schematic diagram of a model analysis provided for one embodiment of the present invention;

[0037] Figure 7A schematic diagram of a diagnostic flow chart of a server fault diagnosis method provided by the present invention;

[0038] Figure 8 A schematic structural diagram of a server fault diagnosis device provided by the present invention;

[0039] Figure 9 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0040] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0041] Figure 1 This is a flow chart of the server fault diagnosis method provided by the present invention. Figure 1 The server fault diagnosis method provided by the present invention may include:

[0042] S110, obtaining a fault diagnosis rule configuration file adapted to the server fault;

[0043] S120, establishing a fault diagnosis rule model according to the fault diagnosis rule configuration file;

[0044] S130: Match the server abnormal phenomenon information with the fault diagnosis rule model to determine the fault information.

[0045] The server fault diagnosis method provided by the present invention is performed by the server's baseboard management controller (BMC). As a management and control module of the server, the BMC has an important responsibility to monitor the current health status of the server.

[0046] It should be noted that the server in the embodiment of the present invention may be an Intel platform server, such as a Purley platform-based server, or a server of other platforms, and the present invention does not impose any limitation thereto.

[0047] In step S110 , the fault type of the server is obtained, a fault diagnosis rule configuration file adapted to the server fault type is established, and the fault diagnosis rule configuration file is imported into the BMC module of the server.

[0048] Optionally, the server failure includes host shutdown, host restart, fan abnormality, temperature abnormality, server disk loss, and BMC restart.

[0049] It is understandable that there are many types of server failure phenomena, each of which includes multiple levels of specific failure types. During server operation, if a new unidentified failure is discovered or an incorrectly identified failure is discovered, the server can be updated with the fault types it recognizes by modifying the fault diagnosis rule configuration file, adding server failure types to the configuration file, or modifying the original incorrectly identified failures, thereby dynamically adding and modifying the fault types it recognizes.

[0050] In step S120 , the fault information in the fault diagnosis rule configuration file is parsed, and a fault diagnosis rule model is established based on the hierarchical relationship corresponding to the parsed fault information.

[0051] It's understandable that many server failure phenomena have corresponding sub-phenomena. For example, the host shutdown phenomenon has corresponding sub-phenomena such as motherboard failure, MCA (Machine Check Architecture) failure, and overheating. The overheating phenomenon also has corresponding sub-phenomena, such as fan failure and sensor failure. Based on the correspondence between the failure phenomenon and its sub-phenomena, a corresponding fault diagnosis rule model can be established.

[0052] In step S130, the abnormal phenomenon information in the server is obtained, and the abnormal phenomenon information of the server is matched with the fault diagnosis rule model established in step S120. According to the pre-defined rules of the specific fault type in the model, the specific fault information represented by the abnormal phenomenon information is determined, thereby realizing the diagnosis of the server fault.

[0053] Optionally, the server exception information may be output information with Error or Warning printed by the server, or may be fault information presented in other forms.

[0054] It is understandable that server fault diagnosis can be achieved by inputting the collected server abnormal phenomenon information into the fault diagnosis model established corresponding to the fault diagnosis rule configuration file, and determining the specific fault cause based on the fault type already in the model.

[0055] The server fault diagnosis method provided by the present invention creates a fault diagnosis rule configuration file adapted to the server fault, establishes a fault diagnosis model, and matches collected fault information with the fault diagnosis model to determine the specific fault information, forming a set of fault diagnosis processes with dynamically configurable rules. By dynamically adding or modifying specific fault types in the fault diagnosis rule configuration file, the fault diagnosis rule configuration file can be updated and parsing rules can be changed without upgrading the corresponding software, thereby improving server maintainability and the accuracy of fault analysis.

[0056] In one embodiment, a fault diagnosis rule model is established based on a fault diagnosis rule configuration file, including: based on the fault diagnosis rule configuration file, a fault diagnosis rule model is established hierarchically using a regularized tree model of a binary tree; wherein, in the fault diagnosis rule model, abnormal phenomena of servers at each level are respectively used as child nodes of each layer of the root node.

[0057] Optionally, a binary tree regularized tree model is established based on the corresponding fault phenomena in the fault diagnosis rule configuration file and the corresponding specific fault information of the fault phenomena. The root node is used as the root node of the binary tree model, and the fault phenomena in the configuration file, such as server disk drop, BMC restart, host shutdown, fan abnormality, etc., are used as child nodes of the root node in the binary tree model. For each root node in the binary tree model, the fault phenomenon subnode also includes multiple child nodes. For example, the configuration file host shutdown includes motherboard failure, overtemperature and MCA failure as child nodes of the host shutdown node in the binary tree model. For the overtemperature of the child node of the host shutdown, the next layer of the configuration file includes fan failure and sensor failure, etc., which are used as child nodes of the overtemperature node in the binary tree model. According to the fault phenomenon information corresponding to the fault diagnosis rule configuration file and the abnormal phenomenon information at each level contained in the fault phenomenon information, the model is continuously refined downward with the regularized tree model of the binary tree at the top, and the fault diagnosis rule model is established in layers.

[0058] Specifically, each abnormal phenomenon in the fault diagnosis rule configuration file can be regarded as an independent node, and then each node is recursively read according to the model node in the fault diagnosis rule configuration file, and the read node information is used to construct a binary tree using the data structure model of the child-brother method. Figure 2The figure shows the process of constructing a binary tree based on the fault diagnosis rule configuration file using a recursive method and the child-sibling method. Generally speaking, the nodes in the fault diagnosis rule configuration file are recursively read. If the current node has child nodes, the first child node is inserted into the tree node as the left child, and the second child node is inserted into the tree node as the right child. If the current node has other child nodes, the current right child is used as the node, and the other child nodes are inserted into the right child of the node, and so on. If the current child node still has child nodes, the above node insertion process continues.

[0059] It can be understood that the child nodes under the root node in the fault diagnosis rule configuration file all store various fault phenomena, corresponding to the first layer of the tree structure. It can also be said that the type of abnormal phenomenon in the fault diagnosis rule configuration file determines the type of abnormal phenomenon that can be analyzed.

[0060] The server fault diagnosis method provided by the present invention establishes a fault diagnosis rule model in layers using a regularized tree model of a binary tree according to a fault diagnosis rule configuration file. By changing the specific fault type in the configuration file, the abnormal phenomenon type in each model can be flexibly configured, thereby changing the analysis process for specific phenomena, thereby achieving the purpose of dynamic configuration and improving the accuracy of fault diagnosis.

[0061] In one embodiment, before matching the server abnormal phenomenon information with the fault diagnosis rule model, the method further includes: encoding the server abnormal phenomenon information according to each layer of sub-nodes in the fault diagnosis rule model; and arranging the encoded server abnormal phenomenon information in a time series.

[0062] Optionally, the encoding of abnormal phenomenon data is to achieve matching of specific abnormal phenomena. 4 bytes (Byte1 to Byte4) can be used to represent different phenomena to encode the abnormal phenomenon data. Each subsequent additional phenomenon needs to be applied for according to the format. The 4 bytes can be agreed upon as follows: magic number 0xax (Byte4), first-level module (Byte3), second-level module (Byte2), extension or ACTION (Byte1). The magic number 0xax indicates that the current field is an abnormal phenomenon. Taking into account some special cases, the parsing of specific phenomena is simply blocked, and the lower 4 bits are agreed to be used to indicate the special purpose of this abnormal phenomenon. For example, when Bit3 is 1, it means that no parsing of the current abnormal phenomenon will be performed. When Bit3 is 0, parsing is performed by default. Other fields will be applied for later expansion. The first-level module describes the system-level modules and can be defined as follows: 0x01: Host, 0x02: BMC, 0x03: Memory, 0x04: PCIe, 0x05: CPU, 0x06: Fan, 0x07: Power Supply, 0x08: Motherboard, 0x09: Storage, 0x0a: Network, 0x0b: LCD, 0x0c: Upgrade, 0xe0: Other, etc. The second-level module is a refinement of the first-level module. For example, using the host module 0x01 as an example, the following scenarios exist: 0x01: Shutdown, 0x02: Hang, 0x03: Reboot, etc. The Byte1 field can be expanded later and remains 0x00 by default.

[0063] The encoded server anomaly information is arranged in a time series, that is, the collected server anomaly data is sorted in the order of occurrence. Based on experience, when a specific phenomenon is triggered, the current data has a strong correlation. Therefore, arranging each collected data in a time series and starting the matching model from the newest to the oldest on the time axis can improve matching efficiency.

[0064] The server fault diagnosis method provided by the present invention encodes server anomaly information based on each layer of sub-nodes in the fault diagnosis rule model. Compared to character strings, the use of encoded hexadecimal data not only facilitates machine recognition and processing during diagnosis, but also has excellent machine compatibility. Furthermore, the encoded hexadecimal data is implemented in ciphertext, making it difficult to decipher, highly secure, and not susceptible to malicious attacks or plagiarism.

[0065] In one embodiment, server abnormal phenomenon information is matched with a fault diagnosis rule model to determine fault information, including: matching the time-series abnormal phenomenon information with each layer of sub-nodes of the fault diagnosis rule model in a traversal order to determine the fault information; wherein the traversal order includes: starting from the top-level sub-node of the fault diagnosis rule model, and traversing to each level of lower-level sub-nodes in turn.

[0066] Optionally, a pre-order traversal method is used to match the time-series abnormal phenomenon information with the sub-nodes of each layer of the fault diagnosis rule model in sequence. That is, by first traversing the top-level model, matching the relevant abnormal phenomena, calling the corresponding callback function, and then traversing the left subtree first and then the right subtree to perform rule matching. Figure 3 Shown is the current flow through the model when an anomaly is triggered.

[0067] The server fault diagnosis method provided by the present invention determines the fault information by matching the time-series abnormal phenomenon information with the sub-nodes of each layer of the fault diagnosis rule model in the traversal order. By flexibly configuring the abnormal phenomenon types in each model, the analysis process for specific phenomena can be changed, thereby achieving the purpose of dynamic configuration and improving the accuracy of fault diagnosis.

[0068] In one embodiment, the method further includes encoding the server log information according to each layer of sub-nodes in the fault diagnosis rule model; and arranging the encoded server log information in a time series.

[0069] Optionally, the server's log information may include SEL (System Event Log) standard event log format event log, SDS (Smart Diagnosis System) log, normal data, and key file comparison information. Among them, the SEL event log is a standard event log defined in the IPMI (Intelligent Platform Management Interface) rules, and most key events are actually recorded in the SEL. The SDS log is a custom log format recorded by the intelligent diagnostic system H3C server, which is of great help in comprehensive analysis. Normal data is similar to environmental variables, such as BMC startup time, host configuration information, etc. Key file comparison information is, for example, the currently collected server boot log Bootlog information, and the comparison of historical Bootlog difference information can also be used as a relevant criterion.

[0070] Specifically, if Figure 4The following figure shows the server's log information data encoding format. The log encoding format consists of two parts totaling 16 bytes (Byte 16 to Byte 1). The first 8 bytes determine the specific format of the current log and are used to match a single log or log type. The second 8 bytes carry relevant slot information, the extension field. In the rule configuration file, only the first 8 bytes need to be stored according to the predefined format. Furthermore, according to the format definition, Byte 16 contains a control field. Therefore, when matching specific log types, the lower 4 bits of Byte 16 must be masked for comparison. Logs generated by each module must include the full 16 bytes. If 16 bytes are insufficient, the extension field can be used to expand the length of the extension field. Bytes 3 and 4 determine the length of the extension field.

[0071] Byte 16 is the control field, where Bit 2 represents the control field of the current model's log, facilitating matching and parsing of the corresponding logs during the final match. For example, if module A's logs contain only specific logs a, b, and c, then when designing model A, if you want it to parse all logs a, b, and c, you can simply set Bit 2. Bit 1 indicates whether to additionally parse logs from other modules. For example, if module A needs to additionally parse log d from module B, then Bit 1 needs to be set. Note that these fields are not inherited in nested models.

[0072] Byte 15 represents the log type included in the current model type. Currently, it can be divided into four main categories: SEL, SDS, bootlog, and configuration constant information. Using a single bit to represent a specific log type is convenient. For example, in the upper-level model, the corresponding SEL and SDS bits only need to be set, forming an OR relationship, which facilitates regularization and judgment. Furthermore, when the corresponding actual log is generated, the corresponding log must be the only bit in a specific bit.

[0073] Bytes 14 to 13 represent module types that can refer to the definitions in the exception phenomenon section, but the effect is not equivalent, as the module definition in the model is broader. Furthermore, a specific log might be defined in module A during model definition, and module B can reference that log if needed during model design. This provides a more flexible approach. In the overall design, the first-level module definition is broader, while the second-level module refines the first-level module.

[0074] Byte 12 to Byte 9 are customized log formats. In SEL logs, these four bytes are fixed event codes. However, in SDS logs, depending on the customization of each module, some of the secondary modules can be further refined to derive tertiary or quaternary modules, etc., and different related logs in the secondary modules can also be stored.

[0075] Byte 8 to Byte 5 are slot information, which stores the specific slot information of the current log.

[0076] Byte4 to Byte3: store the length of the extension field. The default value is 0, indicating no extension. If the current 16 bytes are insufficient to store the log content, this field is used to extend the log content.

[0077] Byte 2 stores the source of the log, such as BIOS, BMC itself, or storage card, ME, etc.

[0078] Byte1 is a reserved field.

[0079] The server fault diagnosis method provided by the present invention encodes server log information. Using the encoded hexadecimal data not only facilitates machine recognition and processing during diagnosis, but also has excellent machine compatibility. Furthermore, the encoded hexadecimal data is implemented in ciphertext, making it difficult to decipher, highly secure, and not susceptible to malicious attacks or plagiarism.

[0080] In one embodiment, the method further includes matching the server log information in a time-series arrangement with the fault diagnosis rule model to determine the fault information when the fault information cannot be determined by matching the server abnormal phenomenon information with the fault diagnosis rule model.

[0081] Optionally, the relevant log information of the server is stored in a time-series manner. When matching the server abnormal phenomenon information with the fault diagnosis rule model cannot directly determine the cause of the fault, the time-series stored server log information is matched with the fault diagnosis rule model to further determine the fault information. Figure 5 This is a flowchart of the abnormal phenomenon handling process, which shows how to further determine the fault information by traversing the log when an abnormal phenomenon occurs.

[0082] The server fault diagnosis method provided by the present invention encodes server log information and uses the encoded server log as a further basis for fault judgment, thereby increasing the basis for fault diagnosis and improving the fault diagnosis rate.

[0083] In one embodiment, the fault diagnosis rule configuration file is a file in XML format.

[0084] Optionally, the elements present in the XML file may include:<phenomenon_x>< / phenomenon_x> [M], fault phenomenon node, indicates the current fault phenomenon sequence number, x starts from 1, and this is a required item ([M] indicates a required item, [O] indicates an optional item).

[0085] <phenomenon_type> power_off< / phenomenon_type> [M]: An attribute under the fault phenomenon node, indicating the type of the fault object. "power_off" is the unique code corresponding to the fault object, for example, 0x10 represents "power_off", 0x11 represents "system_halt", and 0x12 represents "bmc_reset". This converts the highly readable fault phenomenon into a unique code, effectively encrypting it, facilitating code processing while maintaining relative confidentiality.

[0086] <quenue_depth> 4< / quenue_depth> [M]: Queue depth, indicating the depth of the subtree containing the current node. This attribute is used when constructing the tree structure and is also used to verify whether there are any problems with the current configuration.

[0087] <model_m_n>< / model_m_n> [M]: Model node, used to describe the model of the relevant criteria, where m represents the algorithm model of the current layer, and n represents the algorithm model of the current layer. Figure 6 As shown in the model analysis diagram, the various models are both mutually inclusive and parallel. This is done to achieve stratification. Because the human analysis process itself is inherently random, different people will produce different results. By organizing the chaotic, empirically based human analysis process into a structured code using machine language and outputting the analysis results in a specific format, we facilitate subsequent expansion based on this foundation.

[0088] <model_1_1_type> BIOS_State< / model_1_1_type> [M]: "BIOS_State" is a predefined code storage.

[0089] <time_ref>unlimited< / time_ref> [O]: This field describes the time validity of the time series log corresponding to the current model. For example, a time series log contains a memory UCE (Uncorrectable Error) event log from 20 minutes ago. The actual trigger is the most recent motherboard exception log. According to the current model, this memory UCE event log is likely to be misjudged. Therefore, a reference time for the model judgment is required. If the current model does not define a reference time, this time can be inherited from the previous model time.

[0090] <resultlevel> 0< / resultlevel> : This value indicates the reliability of the data obtained by the current model and whether to continue matching the next model after the current model is matched. 0 indicates to continue matching, and 1 indicates not to continue matching.

[0091] <operate> 1< / operate> [O]: Because models can nest models, this field indicates whether the model is added or deleted relative to the parent model. A value of 0 indicates addition, while a value of 1 indicates deletion. Absence of this field defaults to addition. If the parent model includes all child models by default, then child models do not need to be added individually. However, if a specific child model should be excluded from the parent model, this method can be used to remove it.

[0092] The server fault diagnosis method provided by the present invention selects a fault diagnosis rule configuration file in XML format. XML file format is an existing mature file format. XML file format is intuitive and has good scalability. BMC supports parsing of XML format files.

[0093] The following is an example of a diagnostic flow diagram of a server fault diagnosis method provided by the present invention to illustrate the technical solution provided by the present invention:

[0094] Diagnostic flow chart Figure 7 As shown, the steps in the figure can be divided into five parts as a whole, including rule configuration file, regularized tree structure, data collection, data timing and data parsing.

[0095] Specifically, the rule configuration file describes the overall parsing rules in XML format. The XML configuration file begins with a root node, followed by parallel definitions of various "phenomenon" nodes. These nodes represent different anomalies driving the parsing. Under the "phenomenon" node are multiple child nodes, which are both parallel and nested. These nodes are named "parsing models." These models are layered on top and further refined. Different model attributes are inserted as needed.

[0096] The regularized tree structure is to expand the above-mentioned rule configuration files in a binary tree structure in the project. When a specific phenomenon occurs, the phenomenon is matched in the tree structure. If the phenomenon matches, it is parsed step by step according to the predefined parsing model. Except for the configuration type, the data source for parsing should come from the time-series collected data.

[0097] The data collection module, which can be understood as various logistics points scattered across the country, collects raw data and encodes it according to established rules. The encoding rules can be simply understood as "log type_slot information" and then sends the data to the data analysis module.

[0098] Data time series processing involves arranging the collected data in chronological order. Experience shows that when a specific phenomenon is triggered, the current logs have a strong correlation. Therefore, each collected data is time-series processed. When matching rules, the logs are matched from the newest to the oldest on the timeline. The matching rules are compared based on the data type; if they are identical, a match is achieved.

[0099] The data parsing module serves as a bridge between the previous and next modules, connecting them. During the initialization phase, it verifies and constructs a tree-like parsing structure based on the current XML rule configuration file. It then begins monitoring data, including fault phenomenon data and log information data. To facilitate subsequent analysis, these data are encoded. Log information is placed into a time-series log information stream, and parsing analysis is initiated if a fault phenomenon is matched. The final data structure is then output in a specific format.

[0100] The present invention also provides a server fault diagnosis device, which can correspond to the server fault diagnosis method described above.

[0101] Figure 8 A schematic diagram of the structure of the server fault diagnosis device provided by the present invention is shown as follows: Figure 8 As shown, the device includes:

[0102] Configuration generation module 810, used to obtain a fault diagnosis rule configuration file adapted to the server fault;

[0103] A model generation module 820 is used to establish a fault diagnosis rule model according to the fault diagnosis rule configuration file;

[0104] The fault matching module 830 is used to match the server abnormal phenomenon information with the fault diagnosis rule model to determine the fault information.

[0105] The server fault diagnosis device provided by the present invention establishes a fault diagnosis model by creating a fault diagnosis rule configuration file adapted to the server fault. Collected fault information is matched with the fault diagnosis model to determine the specific fault information, forming a set of fault diagnosis processes with dynamically configurable rules. By dynamically adding or modifying specific fault types in the fault diagnosis rule configuration file, the fault diagnosis rule configuration file can be updated and parsing rules can be changed without upgrading the corresponding software, improving server maintainability and the accuracy of fault analysis.

[0106] In one embodiment, the configuration generation module 810 is specifically configured to:

[0107] Ensure that the fault diagnosis rule configuration file is in XML format.

[0108] In one embodiment, the model generation module 820 is specifically configured to:

[0109] Establish a fault diagnosis rule model based on the fault diagnosis rule configuration file, including:

[0110] According to the fault diagnosis rule configuration file, the fault diagnosis rule model is established hierarchically using a regularized tree model of a binary tree;

[0111] Among them, in the fault diagnosis rule model, the abnormal phenomena of servers at each level are respectively regarded as sub-nodes of each layer of the root node.

[0112] In one embodiment, the fault matching module 830 is specifically configured to:

[0113] Before matching server anomaly information with the fault diagnosis rule model, the following steps are also included:

[0114] Encode server abnormal phenomenon information according to each layer of sub-nodes in the fault diagnosis rule model;

[0115] Arrange the encoded server anomaly information in time series.

[0116] In one embodiment, the fault matching module 830 is further configured to:

[0117] Match server anomaly information with the fault diagnosis rule model to determine fault information, including:

[0118] Match the time-series abnormal phenomenon information with the sub-nodes of each layer of the fault diagnosis rule model in sequence according to the traversal order to determine the fault information;

[0119] The traversal order includes: starting from the top-level child node of the fault diagnosis rule model, and traversing to the lower-level child nodes in sequence.

[0120] In one embodiment, the fault matching module 830 is further configured to:

[0121] Encode the server log information according to each layer of sub-nodes in the fault diagnosis rule model;

[0122] Arrange the encoded server log information in time series.

[0123] In one embodiment, the fault matching module 830 is further configured to:

[0124] If the fault information cannot be determined by matching the server abnormal phenomenon information with the fault diagnosis rule model, the fault information is determined by matching the time-series server log information with the fault diagnosis rule model.

[0125] The present invention also provides an electronic device, such as Figure 9 As shown, the electronic device may include: a processor 910, a communication interface 920, a memory 930, and a communication bus 940, wherein the processor 910, the communication interface 920, and the memory 930 communicate with each other via the communication bus 940. The processor 910 may call the logic instructions in the memory 930 to execute the steps of the server fault diagnosis method, for example, including:

[0126] Obtain the fault diagnosis rule configuration file adapted to the server fault;

[0127] Establish a fault diagnosis rule model according to the fault diagnosis rule configuration file;

[0128] Match the server abnormal phenomenon information with the fault diagnosis rule model to determine the fault information.

[0129] In addition, the logic instructions in the above-mentioned memory 930 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0130] On the other hand, the present invention further provides a computer program product, comprising a computer program stored on a non-transitory computer-readable storage medium, wherein the computer program comprises program instructions. When the program instructions are executed by a computer, the computer can perform the steps of the server fault diagnosis method provided in each of the above method embodiments, for example, including:

[0131] Obtain the fault diagnosis rule configuration file adapted to the server fault;

[0132] Establish a fault diagnosis rule model according to the fault diagnosis rule configuration file;

[0133] Match the server abnormal phenomenon information with the fault diagnosis rule model to determine the fault information.

[0134] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the server fault diagnosis method provided in the above-mentioned method embodiments are implemented, for example, including:

[0135] Obtain the fault diagnosis rule configuration file adapted to the server fault;

[0136] Establish a fault diagnosis rule model according to the fault diagnosis rule configuration file;

[0137] Match the server abnormal phenomenon information with the fault diagnosis rule model to determine the fault information.

[0138] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0139] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A server fault diagnosis method, characterized in that: Applicable to baseboard management controllers, including: Obtain the fault diagnosis rule configuration file adapted to the server fault; Establishing a fault diagnosis rule model according to the fault diagnosis rule configuration file; Matching the server abnormal phenomenon information with the fault diagnosis rule model to determine the fault information; The establishing of a fault diagnosis rule model according to the fault diagnosis rule configuration file includes: According to the fault diagnosis rule configuration file, the fault diagnosis rule model is established hierarchically using a regularized tree model of a binary tree; Wherein, in the fault diagnosis rule model, the abnormal phenomena of servers at each level are respectively used as sub-nodes of each layer of the root node; Before matching the server abnormal phenomenon information with the fault diagnosis rule model, the method further includes: Encoding the server abnormal phenomenon information according to each layer of sub-nodes in the fault diagnosis rule model; Arrange the encoded server anomaly information in time series.

2. The server fault diagnosis method according to claim 1, characterized in that: The matching of the server abnormal phenomenon information with the fault diagnosis rule model to determine the fault information includes: Matching the time-series arranged abnormal phenomenon information with the sub-nodes of each layer of the fault diagnosis rule model in sequence according to the traversal order to determine the fault information; The traversal order includes: starting from the top-level child node of the fault diagnosis rule model, and traversing to the lower-level child nodes in sequence.

3. The server fault diagnosis method according to claim 1, characterized in that: Also includes: Encoding server log information according to each layer of sub-nodes in the fault diagnosis rule model; Arrange the encoded server log information in time series.

4. The server fault diagnosis method according to claim 3, characterized in that: Also includes: If the fault information cannot be determined by matching the server abnormal phenomenon information with the fault diagnosis rule model, the server log information arranged in time series is matched with the fault diagnosis rule model to determine the fault information.

5. The server fault diagnosis method according to claim 1, characterized in that: The fault diagnosis rule configuration file is a file in XML format.

6. A server fault diagnosis device, characterized in that: Applicable to baseboard management controllers, including: Configuration generation module, used to obtain fault diagnosis rule configuration files adapted to server faults; A model generation module, configured to establish a fault diagnosis rule model according to the fault diagnosis rule configuration file; A fault matching module is used to match the server abnormal phenomenon information with the fault diagnosis rule model to determine the fault information; The establishing of a fault diagnosis rule model according to the fault diagnosis rule configuration file includes: According to the fault diagnosis rule configuration file, the fault diagnosis rule model is established hierarchically using a regularized tree model of a binary tree; Wherein, in the fault diagnosis rule model, the abnormal phenomena of servers at each level are respectively used as sub-nodes of each layer of the root node; Before matching the server abnormal phenomenon information with the fault diagnosis rule model, the method further includes: Encoding the server abnormal phenomenon information according to each layer of sub-nodes in the fault diagnosis rule model; Arrange the encoded server anomaly information in time series.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the server fault diagnosis method according to any one of claims 1 to 5 are implemented.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the server fault diagnosis method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Intelligent fault diagnosis method and system, wind turbine generator set and storage medium

    CN109270458A