Server fault analysis method and device and medium
By performing event classification and correlation analysis on server log files, the problem that the BMC fault monitoring solution could not cover multiple fault types was solved, and more accurate fault analysis was achieved.
Patent Information
- Application Number
- CN202511181251.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-11-07
AI Technical Summary
Existing BMC-based server fault monitoring solutions cannot cover an increasing number of fault types, resulting in poor accuracy in fault analysis.
By obtaining the server's log files, processing the log files using log extraction template sets, extracting multiple log events, and classifying them based on event source and event type, the fault analysis results for the same event source are determined. Combined with correlation algorithms, frequent event item sets are identified, thereby improving the reliability and accuracy of fault analysis.
It improves the reliability and accuracy of server fault analysis, enabling accurate location of faulty components or component connections occurring at the same time, thus enhancing the flexibility and reliability of fault analysis.
Smart Images

Figure CN120915656A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and particularly relates to a server fault analysis method, device and medium. BACKGROUND
[0002] With the maturity of Internet technology, the services carried by the servers of the data center are more and more diverse, and the amount of data to be processed is also more and more large. In order to ensure that the servers can normally run in the intensive service state, it is usually necessary to monitor and process the faults of the servers.
[0003] Generally, the fault monitoring of the server can be performed with the help of the baseboard management controller (BMC) on the server.
[0004] However, in the scenario where the amount of services carried by the server is more and more large, the server fault analysis scheme based on the fault monitoring capability of the BMC cannot cover more and more fault types, resulting in poor accuracy of the server fault analysis. SUMMARY
[0005] The embodiments of the present application provide a server fault analysis method, device and medium, which improve the reliability and accuracy of the server fault analysis.
[0006] In a first aspect, a server fault analysis method is provided, comprising:
[0007] obtaining a log file of a server;
[0008] processing the log file based on a log extraction template set to obtain a plurality of log events of the server, and grouping the plurality of log events into a plurality of log event sets according to event attribute groups, wherein the event attribute groups include event sources and event types, the event sources include component identifiers or component connection relationship identifiers associated with the log events, and the event types include running information, alarm events or fault events;
[0009] if there are a first alarm event, a first fault event and first running information belonging to a same event source in the plurality of log event sets, the first alarm event, the first fault event and the first running information are determined as a fault analysis result of the same event source, wherein the event occurrence time of the first alarm event, the first fault event and the first running information is the same.
[0010] By adopting the technical solutions, the log files in the server can be processed based on the log extraction template to obtain multiple log events of the server, and the multiple log events can be classified based on the event attribute group composed of the event source and the event type, and then in the multiple log event sets, if an alarm event and a failure event about the same event source occur at the same time, the alarm event, the failure event and the running information about the same event source at the same time are determined as the analysis result of the same event source, which not only fully utilizes the log files to analyze the failure of the server components or the component connection relationship, improves the reliability of the server failure analysis, and further improves the accuracy of the server failure analysis by accurately determining the components or the component connection relationship that fail and alarm at the same time.
[0011] Optionally, the multiple remaining log events in the multiple log event sets are divided into multiple event item sets, wherein the event item sets include at least one log event belonging to the same event source, and the multiple remaining log events are log events in the multiple log event sets, except the first alarm event, the first failure event and the first running information belonging to the same event source.
[0012] The multiple event item sets are combined into multiple event item set groups two by two, wherein the two event item sets included in the event item set group are different.
[0013] The correlation degree parameters of the two event item sets in each event item set group are determined based on a correlation degree algorithm to obtain the correlation degree parameters corresponding to each event item set group.
[0014] The event item set group with the correlation degree parameter greater than a correlation degree parameter threshold is determined as a frequent event item set group, and the failure analysis result of the server is determined based on the frequent event item set group.
[0015] By adopting the technical solutions, in the case that the failure analysis cannot be performed from the same time dimension of the same event source, the multiple remaining log events in the multiple log event sets can be divided into multiple event item sets from the same event source dimension, the frequent event item set group with high frequency can be determined based on the correlation degree parameters between the two event item sets in the multiple event item sets, and the failure analysis result of the server is determined based on the frequent event item set group, which improves the reliability of the failure analysis.
[0016] Optionally, the event item set includes multiple log events, and the multiple remaining log events in the multiple log event sets are divided into multiple event item sets, including:
[0017] traverse the plurality of remaining log events to find at least one type of second alarm event, at least one type of second failure event and / or second running information belonging to a same event source, wherein a time difference between the second alarm event, the second failure event and / or the second running information is less than or equal to a time threshold;
[0018] combine the at least one type of second alarm event, the at least one type of second failure event and / or the second running information belonging to the same event source to obtain a first event item set associated with the same event source;
[0019] determine the first event item sets respectively associated with a plurality of different event sources as a plurality of event item sets.
[0020] By using the above technical solution, at least one type of second alarm event, at least one type of second failure event and / or second running information about a same event source and having a time difference less than a preset time length can be combined into a first event item set, so as to facilitate failure analysis on the state of the same event source within the preset time length, improve the diversity of failure analysis means for components in the server or the component association relationship, and improve the reliability of the failure analysis result.
[0021] Optionally, the event item set includes a plurality of log events, and the dividing the plurality of remaining log events in the plurality of log events into a plurality of event item sets includes:
[0022] traverse the plurality of remaining log events to find at least one type of second alarm event, at least one type of second failure event and / or second running information associated with each event source in an event source group, wherein the event source group includes target component identifiers of two target components having an association relationship in the server and a target component connection relationship identifier between the two target components;
[0023] combine the at least one type of second alarm event, the at least one type of second failure event and / or the second running information associated with each event source in the event source group to obtain a second event item set associated with the event source group;
[0024] determine the second event item sets respectively associated with a plurality of different event source groups as a plurality of event item sets.
[0025] By adopting the technical scheme, for two target components having an association relationship in the server, at least one type of alarm event, at least one type of fault event and / or operation information corresponding to each target component and the association relationship of the target components and having an event occurrence time less than a time length threshold can be divided into a second event item set, so that the state of the two components having an association relationship in the server and the association relationship between the two components within a preset time length is analyzed for fault, and the fault analysis reliability for the components having an association relationship in the server is improved.
[0026] Optionally, the event item set includes a plurality of log events, and the dividing of the plurality of remaining log events in the plurality of log event sets into a plurality of event item sets includes:
[0027] Iterating through the plurality of remaining log events, any two of the first alarm event, the first fault event and the first operation information associated with each event source in the event source group are searched;
[0028] Combining any two of the first alarm event, the first fault event and the first operation information associated with each event source in the event source group, a third event item set associated with the event source group is obtained;
[0029] The third event item sets respectively associated with a plurality of different event source groups are determined as a plurality of event item sets.
[0030] By adopting the technical scheme, for two target components having an association relationship in the server, at least one type of alarm event, at least one type of fault event and / or operation information corresponding to each target component and the association relationship of the target components and having an event occurrence time less than a time length threshold can be divided into a second event item set, so that the state of the two components having an association relationship in the server and the association relationship between the two components within a preset time length is analyzed for fault, and the fault analysis reliability for the components having an association relationship in the server is improved.
[0031] Optionally, the event item set includes a plurality of log events, and the dividing of the plurality of remaining log events in the plurality of log event sets into a plurality of event item sets includes:
[0032] Iterating through the plurality of remaining log events, any two of the first alarm event, the first fault event and the first operation information associated with each event source in the event source group are searched;
[0033] Combining any two of the first alarm event, the first fault event and the first operation information associated with each event source in the event source group, a third event item set associated with the event source group is obtained;
[0034] The fourth event item set respectively associated with a plurality of different event sources is determined as the plurality of event item sets.
[0035] By adopting the technical scheme, any two of the first alarm event, the first failure event and the first running information about the same event source and the same event occurrence time can be combined into one fourth event item set, so as to perform failure analysis on at least two types of log events based on the same event source at the same time, and further improve the flexibility and reliability of failure analysis on components in the server or the component association relationship.
[0036] Optionally, the log file includes a plurality of log information, and the log file is processed based on the log extraction template set to obtain a plurality of log events of the server, including:
[0037] For each log information, each log extraction template in the log extraction template set is traversed to perform keyword matching;
[0038] If there is a target string matching each keyword in the log extraction template in the log information, the target string matching each keyword is combined to obtain the log event associated with the log information;
[0039] The log events associated with each log information are combined to obtain the plurality of log events of the server.
[0040] Optionally, the log file of the server is obtained, including:
[0041] The log file of the server in the last failure analysis period is obtained;
[0042] The first alarm event, the first failure event and the first running information are determined as the failure analysis result of the same event source, including:
[0043] The first alarm event, the first failure event and the first running information are determined as the failure analysis result of the same event source of the server in the last failure analysis period.
[0044] In a second aspect, a server failure analysis apparatus is provided, including:
[0045] An obtaining module configured to obtain a log file of a server;
[0046] The extraction module is configured to process the log files based on a log extraction template set to obtain a plurality of log events of the server, and group the plurality of log events into a plurality of log event sets according to event attribute groups, wherein the event attribute groups include event sources and event types, the event sources include component identifiers or component connection relationship identifiers associated with the log events, and the event types include running information, alarm events or fault events.
[0047] The first determination module is configured to determine, if there are a first alarm event, a first fault event and first running information belonging to a same event source in the plurality of log event sets, the first alarm event, the first fault event and the first running information as a fault analysis result of the same event source, wherein the first alarm event, the first fault event and the first running information have the same event occurrence time.
[0048] In a third aspect, an electronic device is provided, including a memory, a processor and a computer program stored in the memory, and the processor executes the computer program to implement the method in the first aspect.
[0049] In a fourth aspect, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method in the first aspect.
[0050] In a fifth aspect, a computer program product is provided, and the computer program product includes a computer program, and the computer program is executed by a processor to implement the method in the first aspect.
[0051] It is to be understood that both the foregoing general description and the following detailed description are exemplary, and are intended to provide further explanation of the subject technology claimed. BRIEF DESCRIPTION OF DRAWINGS
[0052] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings in which:
[0053] Figure 1 FIG. 1 is a schematic diagram of a server component architecture according to an embodiment of the present application.
[0054] Figure 2 FIG. 2 is a schematic diagram of a server organization architecture according to an embodiment of the present application.
[0055] Figure 3is an execution flow schematic diagram of a server fault analysis method according to an embodiment of the present application.
[0056] Figure 4 is a schematic diagram of a BMC log file according to an embodiment of the present application.
[0057] Figure 5 is a schematic diagram of a plurality of event item set correlation according to an embodiment of the present application.
[0058] Figure 6 is an execution flow schematic diagram of another server fault analysis method according to an embodiment of the present application.
[0059] Figure 7 is a block diagram of a server fault analysis apparatus according to an embodiment of the present application.
[0060] Figure 8 is a schematic diagram of a computer program product according to an embodiment of the present application.
[0061] Figure 9 is a hardware block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0062] Embodiments of the present application will be described in more detail with reference to the drawings. While certain embodiments of the present application will be shown in the drawings, it should be understood that the present application can be embodied in various forms and should not be construed as being limited to the embodiments set forth in the drawings. Rather, these embodiments are provided so that the present application will be thorough and complete, and fully convey the scope of the present application to those skilled in the art.
[0063] It should be understood that each step recited in the method embodiments of the present application can be executed in different order and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present application is not limited in this respect.
[0064] The term "comprising" and variations thereof as used in the present application are open-ended, that is, "including but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions are given throughout the description. It should be noted that the concepts mentioned in the present application are merely used to distinguish different apparatuses, modules or units, and are not intended to limit the functions performed by these apparatuses, modules or units, or the order or interdependence of these functions.
[0065] It should be noted that the modification of "one", "multiple" mentioned in the present application is illustrative but not restrictive, and those skilled in the art should understand that unless otherwise explicitly indicated in the context, it should be understood as "one or more".
[0066] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present application are only for illustrative purposes, and are not used to limit the scope of the messages or information.
[0067] In order to solve the above problems, the server fault analysis method provided in the embodiments of the present application can be used to analyze the fault condition of the server of any organization architecture. Specifically, the actual architecture of the server can be determined, and the embodiments of the present application do not limit this. Optionally, the server fault analysis method provided in the present application can be used to analyze the fault condition of the components in the server, such as Figure 1 as shown, Figure 1 A server component architecture diagram is shown, wherein the server includes a first IO module (Input / Output Module) 101, a second IO module 102, a power module 103, a third IO module 104, a case 105, a built-in hard disk module 106, a super capacitor support 107, a wind deflector 108, a front hard disk backboard 109, a fan support 110, a fan module 111, a front hard disk 112, a flexible IO card 113, a mainboard 114, a redundant array of independent disks (Redundant Array of Independent Disks, RAID) control card 115, a trusted platform module (Trusted Platform Module, TPM) / trusted cryptography module (Trusted Cryptography Module, TCM) card 116, a memory 117, a processor 118 and a heat sink 119.
[0068] Optionally, in addition to the running status of the components of the server, the server fault analysis method provided in the embodiments of the present application can also be used to analyze the fault condition of the connection relationship of the components in the server; such as Figure 2 as shown, Figure 2A server organization architecture diagram provided by an embodiment of the present application is shown, wherein the server comprises a first central processing unit (CPU) 1, a second central processing unit (CPU2), 12 dual in-line memory modules (DIMM), a riser card 1, a riser card 2, a riser card 3, a 1*x8 Slimline interface, a 2*x8 Slimline interface, a flexible network interface controller (Flexible NIC), a redundant array of independent disks controller card (RAID controller card), a hard disk drive backplane (HDD backplane), a serial advanced technology attachment digital versatile disc-rewritable drive (SATA DVD-RW drive), a serial advanced technology attachment (SATA), a mini-serial attached SCSI high density interface (Mini-SAS HD), a universal serial bus key (USB key), an X557 type physical layer chip (PHY),2 RJ45, 2*10GE SFP+, integrating the X722 NIC, iBMC, GE, UART, VGA, Rear USB 3.0, Front USB 2.0, and 2*100Base-T.
[0069] The connection relationship between the components includes: the CPU1 and the CPU2 are connected through the UPI 0 (Ultra Path Interconnect 0) and the UPI 1 (Ultra Path Interconnect 1), the CPU1 and the Riser card1 are connected through the Peripheral Component Interconnect Express (PCIe) 3.0x24, the CPU1 is further connected with the 1*8Slimline and the RAID controller card through the PCIe 3.0x8; the CPU2 and the Riser card2 are connected through the PCIe 3.0x24, the CPU2 is further connected with the Flexible NIC and the 2*8Slimline through the PCIe 3.0x8 and the PCIe 3.0x16 respectively; at the same time, the RAID controller card is connected with the HDD backplane through two 4*serial connection Serial Attached Small Computer System Interface (SAS).
[0070] The Lewisburg LBG-2PCH (integrating the X722 NIC) is connected with the CPU1 through an 8-channel uplink (Uplink x8) and a direct media interface (DMI), meanwhile, the Lewisburg LBG-2PCH (integrating the X722 NIC) is connected with 3*Mini-SAS HD through a SATA interface, the Lewisburg LBG-2PCH (integrating the X722 NIC) is connected with the USB key through a universal serial bus (USB) 3.0, the Lewisburg LBG-2PCH (integrating the X722 NIC) is connected with the X557 type (PHY) through a serializer (Serializer / Deserializer, SerDes), the Lewisburg LBG-2PCH (integrating the X722 NIC) is connected with the Rear USB 3.0 through 2*USB 3.0, and is connected with the Front USB 2.0 through 2*USB 3.0, in addition, the Lewisburg LBG-2PCH (integrating the X722 NIC) is also connected with two 1*SATA and the SATA DVD-RW drive respectively, meanwhile, the X557 type (PHY) and the 2RJ45 are connected through a 2*10GBASE-T based on twisted pair, and a 1*SATA is also connected with the SATA DVD-RW drive.
[0071] The Lewisburg LBG-2PCH (integrating the X722 NIC) is also connected with the iBMC through 2*PCIe, a specific management link (SMLink) 0, a USB 3.0 and a low-pin count interface (LPC), meanwhile, the iBMC is also connected with the GE, the UART and the VGA respectively.
[0072] Figure 3 A flow chart of a server fault analysis method of an example embodiment of the present application is shown, the method can be applied to a terminal device, and the terminal device can be a computer, a notebook or a tablet computer, etc., as shown in the figure, the method of the embodiment of the present application can include: Figure 3 As shown in the figure, the method of the embodiment of the present application can include:
[0073] Step S301: Obtain the server's log file;
[0074] Step S302: Process the log file based on the log extraction template set to obtain multiple log events of the server, and divide the multiple log events into multiple log event sets according to event attributes;
[0075] The event attribute group includes an event source and an event type. The event source includes a component identifier or component connection relationship identifier associated with the log event. The event type includes operation information, alarm events, or fault events.
[0076] Step S303: If, in the multiple log event sets, there are a first alarm event, a first fault event, and a first operating information belonging to the same event source, then the first alarm event, the first fault event, and the first operating information are determined as the fault analysis results of the same event source.
[0077] The first alarm event, the first fault event, and the first operation information event occur at the same time.
[0078] In summary, the server fault analysis method provided in this application can process log files in the server based on log extraction templates to obtain multiple log events of the server. These multiple log events are then classified based on event attribute groups composed of event source and event type. Furthermore, within the multiple log event sets, using event source and event occurrence time as analysis dimensions, if alarm events and fault events related to the same event source occur simultaneously, the alarm events, fault events, and operational information related to the same event source at the same time are identified as the analysis results of the same event source. This not only fully utilizes log files for fault analysis of server components or component connections, improving the reliability of server fault analysis, but also further enhances the accuracy of server fault analysis by accurately identifying components or component connections that experience faults and alarms at the same time.
[0079] The following are Figure 3 The specific implementation methods of each step in the illustrated embodiment are described in detail below:
[0080] In step S301, the terminal device obtains the server's log file.
[0081] In this embodiment of the application, the server's log file is an operation and maintenance log file generated during the server's operation. The operation and maintenance log file is used to record the operation information of each component during the server's operation, as well as the connection operation information between components; wherein, the operation and maintenance log file can be a BMC log file or an OS log file.
[0082] As an example, Figure 4 As shown, Figure 4 A schematic diagram of a BMC log file is shown, which includes a 3rd dump folder for recording 3rd party related dump log information, an App dump folder for recording application related log information, a BMC log dump folder for recording specific module and component related dump log information, a Core dump folder for recording core log information generated by program crash, a Device dump folder for recording device related log data, a Log dump folder for recording log information during server running, an Opt performance folder for recording storage optimization and performance related log information, an OSDump folder for recording operating system related log information, a Register folder for storing register related log information, an RTOS dump folder for recording real-time operating system related log information, and a Sp log dump folder for recording specific service or module related log information.
[0083] The BMC log file further includes a Sp log-App log (splog_app.log) for recording application side SP related log events, and a dump log for recording various types of dump operation and process related logs.
[0084] In an optional embodiment, the terminal device can perform periodic fault analysis on the server, and the process of obtaining the log file of the server can include: obtaining the log file of the server in the last fault analysis period; so as to analyze the running status of the server in the last fault analysis period; wherein the fault analysis period can be determined based on actual needs, and the present application does not limit this. As an example, the fault analysis period can be analyzed once every 10 minutes, or once every 20 minutes; the computing power consumption in the server fault analysis process can be reduced, and the timeliness can be ensured.
[0085] In step S302, the terminal device processes the log file based on the log extraction template set, obtains a plurality of log events of the server, and divides the plurality of log events into a plurality of log event sets according to event attribute groups.
[0086] In the embodiments of the present application, the log extraction template set is a set of string extraction templates of keyword information written for different types of log information in the log file, wherein the string of key information extracted from the log information based on the log extraction template is used to represent the specific event content indicated by the log information; it should be noted that the log extraction template can be determined based on the type of log information, and the embodiments of the present application do not limit this; the type of log information is the running information, fault information or alarm information of the components in the server or the connection relationship between the components.
[0087] Among them, the event attribute group includes event source and event type, the event source includes component identification or component connection relationship identification associated with the log event, and the event type also includes running information, alarm event or fault event.
[0088] For example, the components can include CPU, hard disk (Disk), mainboard, memory, fan (Fan), power supply unit (Power Supply Unit, PSU / PS), hard disk backplane, network card (Network Adapter), RAID card, fiber channel host bus adapter (Fibre Channel Host Bus Adapter, FC HBA) card, etc.; the connection relationship between the components can include Pcie link, sas link, inter-integrated circuit (Inter-Integrated Circuit, I2C) link, etc.
[0089] The operation information of the components or the connection relationship between the components includes: operation information of the PSU power supply, such as input voltage, output voltage, input current, output current, input and output power, power manufacturer signal and the like; operation information of the RAID card, such as the specification of the RAID card, RAID card temperature information, RAID card SN information, RAID card link error code information and the like, wherein the specification of the RAID card can include RAID card model, RAID card firmware version information, RAID card cache and striping rule information, and battery backup unit (BBU) capacitor information and the like; hard disk backplane, such as the number of hard disks supported by the hard disk backplane, hardware backplane supporting expander (exp), hardware backplane supporting pass-through function, hardware backplane encoding information and hardware backplane logic version information and the like; operation information of the network card, such as network card slot information, network card model, network card identification (ID) information and network card number (number) information and the like; operation information of the mainboard, such as mainboard version information, working power supply voltage state, air inlet and outlet temperature, platform controller hub (PCH) temperature and voltage regulator (VR) power supply temperature and the like; wherein the mainboard version information includes BMC version number, complex programmable logic device (CPLD) version number and basic input / output system (BIOS) version number; operation information of the hard disk, such as hard disk specification, hard disk working temperature, hard disk smart information and hard disk type (logical hard disk or physical hard disk), wherein the hard disk specification includes: hard disk serial number (SN), hard disk firmware (FW), hard disk mode information, hard disk capacity and hard disk manufacturer information; operation information of the memory, such as memory specification, memory bank working temperature, memory bank slot information on the mainboard and the like, wherein the memory specification includes: memory bank SN, memory bank capacity, memory bank type, memory bank minimum voltage, memory bank rate and memory bank manufacturer; operation information of the CPU, such as CPU temperature and CPU specification, wherein the CPU specification includes: CPU model, CPU manufacturer, CPU SN and CPU working voltage and the like; operation information of the FC HBA card, such as FC HBA card slot information, FC HBA card model, FC HBA card ID information and FC HBA card number (number) information; operation information of the fan, such as fan speed, fan duty cycle and fan number.
[0090] It should be noted that the alarm information types in the server can generally include: fan alarms, such as fan rotation speed alarms and fan honor alarms; memory alarms, such as memory configuration error alarms, memory contact alarms, memory pre-time-out alarms, memory error correction code (ECC) alarms, memory uncorrectable error (UCE) alarms, and memory board alarms; storage alarms, such as CPU component related alarms, RAID card component related alarms, hard disk backplane component related alarms, Pcie card related alarms, network card related alarms, and IO module related alarms; storage alarms, such as hard disk failure alarms, hard disk pre-failure alarms, RAID group time-out, hard disk link failure, hard disk state alarms, hard disk has external configuration, hard disk loss, hard disk link error code alarms, PCIE hard disk alarms, self-diagnosis (SD) alarms, exp related alarms, and the like; management subsystem alarms; temperature alarms, such as CPU temperature alarms, CPU voltage for data queue (VDDQ) temperature alarms, CPU voltage regulator down (VRD) temperature alarms, CPU digital thermal sensor (DTS) temperature alarms, CPU core (CORE) temperature alarms, CPU VR power supply temperature alarms, memory temperature alarms, hard disk temperature alarms, hard disk group temperature alarms, power supply temperature alarms, power module primary / secondary temperature alarms, hard disk backplane temperature alarms, RAID card temperature alarms, Pcie card temperature alarms, BBU temperature alarms, optical module temperature alarms, Riser card temperature alarms, delay circuit temperature alarms, air inlet / outlet temperature alarms, IO module temperature alarms, PCH temperature alarms, and artificial intelligence (AI) module temperature alarms; power alarms, such as CPU voltage alarms, CPU voltage for core control plane (VCCP) alarms, CPU voltage for system agent (VSA) alarms, CPU voltage for input / output control (VCCIO) alarms, CPU voltage for memory controller power (VMCP) alarms, memory power supply alarms, PSU power supply alarms, Pcie card power supply alarms, BBU voltage alarms, mainboard power supply alarms, fan backplane power supply alarms, IO module power supply alarms, and AI module power supply alarms; and watchdog alarms.
[0091] Based on the above alarm information types, the alarm information of the components can include: alarm information of the PSU power supply, such as PSU temperature alarm, PSU redundancy failure alarm, PSU failure alarm, PSU input loss alarm, PSU output abnormality (output overvoltage, undervoltage, overcurrent) alarm, PSU mixed insertion alarm, and the like; RAID card alarm information, such as RAID card temperature alarm, RAID card BBU temperature alarm, RAID card group failure alarm, and other related alarms of the RAID card group; alarm information of the hard disk backplane, such as exp related alarm of the hard disk backplane, hard disk backplane temperature alarm, hard disk backplane power supply alarm, hard disk backplane component (such as CPLD self-check, clock or voltage, and the like) alarm; network card alarm information, such as network card failure or network card alarm; mainboard alarm information, such as mainboard wake-up circuit temperature alarm, mainboard power supply alarm, mainboard electronic tag related alarm, mainboard single board identification reading error alarm, mainboard real-time clock (RTC) battery alarm, mainboard CPLD self-check alarm, mainboard single board hardware address error alarm, mainboard clock alarm, mainboard management engine (ME) abnormal alarm, and other related alarms of the mainboard; hard disk alarm information, such as memory configuration misplacement alarm, memory UCE alarm, memory board alarm, memory contact alarm, memory ECC alarm, memory pre-failure alarm; CPU alarm information, such as CPU voltage alarm, CPU temperature alarm, CPU VCCP alarm, CPU VSA alarm, CPU VCCIO alarm, CPU VMCP alarm, and other related power supply alarms of the CPU, CPU VDDQ temperature alarm, CPU VRD temperature alarm, CPU DTS temperature alarm, CPU CORE temperature alarm, CPU VR power supply temperature alarm, and other related temperature alarms of the CPU, CPU component related alarm; FC HBA card alarm information.
[0092] Similarly, the fault information types in the server can generally include: runtime down faults and startup process down faults, wherein the runtime down faults can include: CPU faults such as CPU internal error (IERR) faults, CPU Uncore UCE faults, CPU Core UCE faults; memory UCE faults such as read / write UCE faults, check UCE faults; PCIE UCE faults such as Transaction Layer Packet (TLP) faults, suprise down errors, link exceptions and overflow errors; power supply faults such as CPU power supply faults, memory power supply faults, motherboard power supply faults, hard disk backplane power supply faults; cooling faults such as CPU overheating shutdown, PCH overheating shutdown and memory overheating shutdown, and operating system hang (OS hang); and the startup process down faults can include: motherboard power supply faults and motherboard firmware exceptions, wherein the motherboard firmware exceptions can include: microcode loading exceptions, ME self-check exceptions and Power On Self Test (POST) self-build exceptions.
[0093] It also needs to be explained that in the embodiments of the present application, the fault information types in the server can also generally include: historical alarms and current alarms, the historical alarms are alarm information that has been recovered after the alarm occurs, and the current server state is normal; the current alarms mean that the current server exists alarms that have not been recovered, that is, the server health state is abnormal.
[0094] Based on the above fault information types, the component fault information can include: PSU power supply fault information, RAID card fault information such as RAID card ECC error, RAID card fatal error and RAID card firmware error; hard disk fault information such as hard disk backplane link error fault record, hard disk backplane exp link exception record and hard disk backplane exp working exception; network card fault information; motherboard fault information, hard disk fault information such as hard disk medium error, hard disk sense code fault mark and abnormal record in hard disk smart information; memory fault information such as memory CE error, memory initialization exception error; CPU fault information and FC HBA card fault information.
[0095] In an alternative embodiment, the terminal device can process the log file based on the log extraction template set to obtain the plurality of log events of the server, which can include: for each piece of log information, performing keyword matching on each log extraction template in the log extraction template set; then, if there is a target string matching each keyword in the log extraction template in the log information, combining the target string matching each keyword to obtain the log event associated with the log information; further, combining the log event associated with each piece of log information to obtain the plurality of log events of the server. Each log extraction template in the log extraction template set can be traversed to determine the log extraction template for extracting the log event from the log information based on the target string matching each keyword in the log extraction template, and the target string matching each keyword in the determined log extraction template can be combined to obtain the log event associated with the log information. Traversing each log extraction template can improve the accuracy of the log event determined from the log information.
[0096] For example, as shown in Table 1, Table 1 shows the log extraction template determined for different types of log information, and the obtained log event.
[0097] Table 1
[0098] Table 1
[0099]
[0100] Table 1
[0101]
[0102] In Table 1, the first log information is hard disk bad block type log information, specifically, the first log information is: there is a bad block at the position of 10 (0x0d / s4) in the hard disk No. 75d20595, and the keywords in the log extraction template corresponding to the hard disk bad block type log information are: Puncturing, bad block, on, and the specific event content indicated by the log event PD*_Puncturing_bad_block is: there is a bad block on the hard disk.
[0103] The second log information is RAID controller memory correctable error type log information. Specifically, the second log information is that a correctable error occurs in the memory of the RAID controller on the PCIe slot 7. The keywords in the log extraction template corresponding to the RAID controller memory correctable error type log information are RAID, controller memory, and correctable errors. The extracted log event is RAID_ecc_errors, and the specific event content indicated is that a correctable error occurs in the RAID.
[0104] The third log information is hard disk slight connection abnormality type log information. Specifically, the third log information is that at 03:23:05 on November 14, 2023, a slight connection abnormality occurs in the hard disk No. 5. The keywords in the log extraction template corresponding to the hard disk slight connection abnormality type log information include time, Minor, Disk, link, and abnormal. The extracted log event is Disk*_minor_link_abnomal, and the specific event content indicated is that a slight connection abnormality occurs in the hard disk.
[0105] The fourth to sixth log information is hard disk medium error type log information. Specifically, the keywords in the log extraction template corresponding to the hard disk medium error type log information include PD, Sense, and 3 / * / *. The extracted log event is PD_s*_sense_medium, and the specific event content indicated is that a hard disk medium error occurs.
[0106] The seventh log information is hard disk reset type log information. Specifically, the seventh log information is that an exception occurs in the path 53473796b8748004 of the hard disk No. 7 located at the e0x41 / s4 position, and the system has performed a Type 03 type reset operation on the path to restore the connection. The keywords in the log extraction template corresponding to the hard disk reset type log information include PD, s, and reset. The extracted log event is PD_s*_reset, and the specific event content indicated is that a hard disk reset occurs.
[0107] The eighth log information is hard disk diagnosis failure type log information. Specifically, the eighth log information is that the diagnosis of the hard disk No. 7 located at the e0x41 / s4 position fails. The keywords in the log extraction template corresponding to the hard disk diagnosis failure type log information include PD, s, Diagnostics, and failed. The extracted log event is PD_s*_failure_diagnostics, and the specific event content indicated is that a hard disk diagnosis failure occurs.
[0108] The ninth log information is hard disk state abnormality class log information. Specifically, the ninth log information is as follows: at 15:41:56 on December 11, 2023, the state of the 15th disk with the serial number 69X0A04BFE4G is abnormal, and the system has recorded the minor warning (coded as 0x02000027); the keywords in the log extraction template corresponding to the hard disk state abnormality class log information include Disk, Disk state, abnormal, time, and Asserted<*>. The extracted log event Disk*_link_abnomal_asserted indicates that the specific event content is hard disk state abnormality.
[0109] The tenth log information is hard disk loss class log information. Specifically, the tenth log information is as follows: at 11:48:06 on December 10, 2023, it is detected that the 14th hard disk cannot be recognized (in a missing state), and a serious level of warning is triggered; the keywords in the log extraction template corresponding to the hard disk state abnormality class log information include Major, Disk, and missing. The extracted log event Disk*_major_missing indicates that the specific event content is hard disk loss.
[0110] The eleventh log information is hard disk interface state abnormality class log information. Specifically, the eleventh log information is as follows: the physical layer interface state of the hard disk bay corresponding to the 14th slot of the 41st hard disk is abnormal; the keywords in the log extraction template corresponding to the hard disk interface state abnormality class log information include PD, phy, and bad. The extracted log event PD_s*_phy_bad indicates that the specific event content is hard disk interface state abnormality.
[0111] The twelfth log information is hard disk health abnormality class log information. Specifically, the twelfth log information is as follows: the health state of the 44th hard disk (Intel 3.2TB, and the serial number is PHLN0312005M3P2BGN) is seriously abnormal; the keywords in the log extraction template corresponding to the hard disk health abnormality class log information include ID, Device Name, Manufacturer, Serial Number, Model, Firmware Version, and Health Status. The extracted log event indicates that the specific event content is that the health state of the hard disk with Intel 3.2TB and the serial number PHLN0312005M3P2BGN is seriously abnormal.
[0112] It can be understood that in the embodiments of the present application, the event attribute groups can be component identification-operation information groups, component identification-alarm event groups, component identification-failure event groups, component connection relationship identification-operation information groups, component connection relationship identification-alarm event groups, and component connection relationship identification-failure event groups.
[0113] As shown in Table 2, Table 2 shows a log event set based on a disk-alarm information group, including a disk minor link abnormality alarm (Disk*_minor_link_abnomal) and a disk missing (Disk*_major_missing).
[0114] Table 2
[0115]
[0116] As shown in Table 3, Table 3 shows a log event set based on a disk-operation information group, including a health state of a hard disk with an Intel 3.2 TB and a serial number of PHLN0312005M3P2BGN is severely abnormal, a hard disk reset (PD_s*_reset), and a hard disk state abnormality (Disk*_link_abnomal_asserted).
[0117] Table 3
[0118]
[0119] As shown in Table 4, Table 4 shows a log event set based on a disk-failure information group, including a hard disk medium error (PD_s*_sense_medium) failure, a hard disk diagnostic failure (PD_s*_failure_diagnostics) failure, a hard disk bad block (PD*_Puncturing_bad_block) failure, and a hard disk interface state abnormality (PD_s*_phy_bad) failure.
[0120] Table 4
[0121]
[0122] As shown in Table 5, Table 5 shows a log event set based on a RAID-failure information group, including a RAID correctable error (RAID_ecc_errors) failure.
[0123] Table 5
[0124] Log event Set of log events RAID_ecc_errors RAID_failure
[0125] In step S303, if the first alarm event, the first failure event and the first running information belonging to the same event source exist in the plurality of log event sets, the terminal device determines the first alarm event, the first failure event and the first running information as the failure analysis result of the same event source.
[0126] In the embodiment of the present application, the event occurrence time of the first alarm event, the first failure event and the first running information is the same.
[0127] In an optional implementation, the terminal device can traverse the plurality of log event sets, and in the case that the first alarm event, the first failure event and the first running information belonging to the same event source exist in the plurality of log event sets, the terminal device determines the first alarm event, the first failure event and the first running information as the failure analysis result of the same event source.
[0128] For example, if the first alarm event, the first failure event and the first running information of the hard disk at the same time exist in the plurality of log event sets, the terminal device determines the first alarm event, the first failure event and the first running information of the hard disk at the same time as the failure analysis result of the hard disk; if the first alarm event, the first failure event and the first running information of the Pcie link at the same time exist in the plurality of log event sets, the terminal device determines the first alarm event, the first failure event and the first running information of the Pcie link at the same time as the failure analysis result of the Pcie link.
[0129] In an optional implementation, the plurality of remaining log events in the plurality of log event sets are divided into a plurality of event item sets, wherein the event item sets include at least one log event from a same event source, and the plurality of remaining log events are log events in the plurality of log event sets other than the first alarm event, the first fault event, and the first running information from the same event source; then, the plurality of event item sets are combined into a plurality of event item set groups two by two, wherein the two event item sets in the event item set group are different; further; the correlation degree parameters of the two event item sets in each event item set group are determined based on a correlation degree algorithm, to obtain the correlation degree parameters corresponding to each event item set group; and the event item set group with a correlation degree parameter greater than a correlation degree parameter threshold is determined as a frequent event item set group, and the fault analysis result of the server is determined based on the frequent event item set group. In the case where fault analysis cannot be performed from the same time dimension of the same event source, the plurality of remaining log events in the plurality of log event sets are divided into a plurality of event item sets from the same event source dimension, and the frequent event item set group with high frequency is determined based on the correlation degree parameters between the two event item sets in the plurality of event item sets, and the fault analysis result of the server is determined based on the frequent event item set group, thereby improving the fault analysis reliability.
[0130] It should be noted that, in the embodiments of the present application, the event item set division strategy relied on in the process of dividing the plurality of remaining log events in the plurality of log event sets into a plurality of event item sets can be determined based on actual needs, and the embodiments of the present application do not limit this; for example, the event item set division strategy can be to divide at least one type of alarm event, at least one type of fault event, and / or running information with the same event source (component / component association relationship) and an event occurrence time difference less than a time threshold into one event item set; or, for two target components with an association relationship in the server, at least one type of alarm event, at least one type of fault event, and / or running information corresponding to each target component and the association relationship of the target component with an event occurrence time less than a time threshold are divided into one event item set.
[0131] The time threshold can be determined based on actual needs, and the embodiments of the present application do not limit this; for example, the time threshold can be 5 minutes, or 10 minutes.
[0132] In an optional implementation, the event item set includes a plurality of log events, and the process of dividing the plurality of remaining log events in the log event set into a plurality of event item sets can include: traversing the plurality of remaining log events, and searching for at least one type of second alarm event, at least one type of second fault event, and / or second running information belonging to a same event source, where the time difference between the second alarm event, the second fault event, and / or the second running information is less than or equal to a time threshold; then, combining the at least one type of second alarm event, the at least one type of second fault event, and / or the second running information belonging to the same event source to obtain a first event item set associated with the same event source; further, determining the first event item sets associated with a plurality of different event sources as the plurality of event item sets. The at least one type of second alarm event, the at least one type of second fault event, and / or the second running information about the same event source and having a time difference less than a preset time length can be combined into a first event item set, so as to perform fault analysis on the state of the same event source within the preset time length, improve the diversity of fault analysis means for components in the server or the component association relationship, and improve the reliability of the fault analysis result.
[0133] The time threshold can be determined based on actual needs, and embodiments of the present application do not limit this. For example, the time threshold can be 5 minutes, or 10 minutes.
[0134] For example, the terminal device can combine {component 1_second running information i1, component 1_second fault event f1, component 1_second alarm event a1, component 1_second fault event f2} 4 log events into a first event item set; and can also combine {component 2_second running information i2, component 2_second alarm event a2, component 2_second alarm event a3, component 2_second fault event f3} 5 log events into a first event item set.
[0135] In an optional implementation, the event item set includes a plurality of log events, and the process of dividing the plurality of remaining log events in the log event set into a plurality of event item sets includes: traversing the plurality of remaining log events, and searching for at least one type of second alarm event, at least one type of second fault event and / or second running information associated with each event source in an event source group, wherein the event source group includes target component identifiers of two target components having an association relationship in the server and a target component connection relationship identifier between the two target components; then, combining the at least one type of second alarm event, the at least one type of second fault event and / or the second running information associated with each event source in the event source group to obtain a second event item set associated with the event source group; and further determining the second event item sets respectively associated with a plurality of different event source groups as the plurality of event item sets. The at least one type of alarm event, the at least one type of fault event and / or the running information corresponding to each target component and the association relationship of the target component and having an event occurrence time less than a time threshold can be divided into a second event item set, so as to analyze the status of the two components having an association relationship in the server and the association relationship between the two components within a preset time length, and improve the fault analysis reliability of the components having an association relationship in the server.
[0136] For example, if the component 3 and the component 4 in the server have an association relationship c1, the terminal device can combine {component 3_second running information i3, component 3_second fault event f4, component 3_second fault event f5, component 4_second alarm event a4, component 4_second fault event f6, association relationship c1_second running information i4, association relationship c1_second fault event f7, association relationship c1_second fault event f8} 8 log events into a second event item set.
[0137] In an alternative embodiment, the event item set includes a plurality of log events, and the process of dividing a plurality of remaining log events in the plurality of log events into a plurality of event item sets includes: traversing the plurality of remaining log events to find any two of the first alarm event, the first failure event, and the first running information associated with each event source in the event source group; then, combining any two of the first alarm event, the first failure event, and the first running information associated with each event source in the event source group to obtain a third event item set associated with the event source group; and further determining the third event item sets respectively associated with a plurality of different event source groups as the plurality of event item sets. For two target components in the server that have an association relationship, any two of the first alarm event, the first failure event, and the first running information corresponding to each target component and the association relationship between the target components that occur at the same time can be divided into a third event item set, so that the status of the two components in the server and the association relationship between the two components at the same time can be analyzed for failure, and the failure analysis reliability for the components in the server that have an association relationship is further improved.
[0138] For example, if the component 5 and the component 6 in the server have an association relationship c2, the terminal device can combine the six log events {component 5_first running information i5, component 5_first failure event f2, component 5_first failure event f3, component 6_first alarm event a1, association relationship c2_first running information i6, association relationship c2_first failure event f7} into a second event item set.
[0139] In an alternative embodiment, the event item set includes a plurality of log events, and the process of dividing a plurality of remaining log events in the plurality of log events into a plurality of event item sets includes: traversing the plurality of remaining log events to find any two of the first alarm event, the first failure event, and the first running information associated with each event source in the event source group; then, combining any two of the first alarm event, the first failure event, and the first running information associated with each event source in the event source group to obtain a third event item set associated with the event source group; and further determining the third event item sets respectively associated with a plurality of different event source groups as the plurality of event item sets. For two target components in the server that have an association relationship, any two of the first alarm event, the first failure event, and the first running information corresponding to each target component and the association relationship between the target components that occur at the same time can be divided into a third event item set, so that the status of the two components in the server and the association relationship between the two components at the same time can be analyzed for failure, and the failure analysis reliability for the components in the server that have an association relationship is further improved.
[0140] For example, the terminal device can combine the two log events of {component 7_first running information i7, component 7_first failure event f3} into a fourth event item set; and can also combine the two log events of {component 8_first alarm event a4, component 8_first failure event f7} into a fourth event item set.
[0141] For example, as shown in Table 6, Table 6 shows an event item set table in the case of constructing an event item set from a component dimension according to an embodiment of the present application, wherein T1, T2, T3, T4, T5 and T6 represent different component identifiers, 1 represents that there is running information, alarm information or failure information corresponding to the component, and 0 represents that there is no running information, alarm information or failure information corresponding to the component.
[0142] Table 6
[0143] Component Component_operation_info Component_failure_event Component_alert_event T1 1 1 0 T2 1 1 1 T3 1 1 1 T4 1 1 1 T5 1 0 1 T6 1 0 1
[0144] For example, if the component is a hard disk, the event item set table in the case of constructing an event item set from a component dimension is shown in Table 7. In Table 7, the log event represented by hard disk_running information Disk_xx is that the running information of the hard disk includes xx; the log event represented by hard disk_running information Disk_SEAGATE_WRQ20P1 K0000E4065JB2_ST8000NM018B_E003 is that the running information of the hard disk includes a manufacturer of Seagate, a unique serial number of WRQ20P1 K0000E4065JB2, a model of 8TB enterprise nearline hard disk ST8000NM018B, and a version of E003 firmware of the hard disk; the log event represented by hard disk_failure event Disk_failure_predictive is that the hard disk has a pre-failure error fault; the log event represented by hard disk_failure event disk_failure_error_02 is that the hard disk has a fault with an error code of 02; the log event represented by hard disk_failure event disk_failure_error_f0 is that the hard disk has a fault with an error code of f0; the log event represented by hard disk_failure event disk_failure_error_fa is that the hard disk has a fault with an error code of fa; the log event represented by hard disk_failure event disk_media_error_count is that the hard disk has a storage medium (disk / flash) statistical fault; the log event represented by hard disk_failure event disk_failure_predictive_maintenance is that the hard disk has a predictive maintenance fault; and the log event represented by hard disk_alarm event disk_alarm_current is that the hard disk has an alarm.
[0145] Table 7
[0146]
[0147] Table 8
[0148]
[0149] As shown in Table 8, Table 8 shows an event item set table in a case that an event item set is constructed from a component association relationship dimension according to an embodiment of the present application, wherein C1, C2, C3, C4, C5 and C6 represent different component association relationship identifiers, 1 represents that there is running information, alarm information or fault information corresponding to a component, and 0 represents that there is no running information, alarm information or fault information corresponding to a component.
[0150] It should be noted that in the embodiment of the present application, the association degree algorithm can be an association rule (Apriori) algorithm or a frequent pattern growth (FP-Growth) algorithm, and the present application does not limit this; the association degree parameter can include support and confidence, wherein the association degree parameter threshold includes minimum support and minimum confidence, and the minimum support and the minimum confidence can be determined based on actual needs, and the present application does not limit this.
[0151] For example, for an event item set X and an event item set Y, the support of the event item set X is:
[0152] S (X) = O' (X) / N; (Formula 1)
[0153] In Formula 1, O' (X) is the number of log events containing all log events in the item set X, and N is the total number of log events.
[0154] The confidence of the event item set X and the event item set Y is:
[0155] C (X-Y) = O' (X∪Y) / O' ; (Formula 2)
[0156] Wherein, the support of the event item set X is used to represent the probability that the event item set X and the event item set Y appear at the same time; and the confidence of the event item set X and the event item set Y is used to represent the probability that the events in the event item set X reach the events in the event item set Y at the same time.
[0157] It should be further noted that in the embodiment of the present application, for an event item set, if the event item set is a frequent item set, all subsets of the event item set are also frequent item sets, and correspondingly, if the event item set is a non-frequent item set, all supersets of the event item set are also non-frequent item sets; wherein, as shown in Table 8, Figure 5 Figure 5 A multi-event item set association relationship schematic diagram provided by the embodiment of the present application is shown. If the item set {A, B} is non-frequent, then the superset of the item set {A, B}, {A, B, C}, {A, B, D}, etc. are also non-frequent, thereby reducing the calculation amount of determining frequent item sets and improving the fault analysis efficiency.
[0158] In an optional implementation, the process in which the terminal device determines the fault analysis result of the server based on the frequent event item set group can include: combining each log event included in each frequent event item set in the frequent event item set group into a frequent log event set, and determining the log event in the frequent log event set as the fault analysis result of the server.
[0159] For example, as shown in Table 9, Table 9 shows a frequent item set table provided by the embodiment of the present application, wherein the frequent log event set is determined based on the frequent event item set of the hard disk at the same time.
[0160] Table 9
[0161]
[0162] Table 9 (continued)
[0163]
[0164] It can be understood that when the time difference of the occurrence time of the multiple log events in the frequent log event set is less than or equal to the time length threshold value, the fault analysis result of the server indicates the component and / or the component association relationship, and at least one of the fault events, the alarm events and the running information appearing within the time length threshold value.
[0165] For example, as shown in Table 9, Table 9 shows a frequent item set table provided by the embodiment of the present application, wherein the frequent log event set is determined based on the frequent event item set of the hard disk at the same time. Figure 6 Figure 6 A flow chart of server fault analysis provided by the embodiment of the present application is shown, which includes:
[0166] In step S601, the log file of the server is acquired.
[0167] In step S602, the log file is processed based on the log extraction template set, the multiple log events of the server are obtained, and the multiple log events are grouped into multiple log event sets according to the event attribute groups.
[0168] The event attribute groups include the event source and the event type, the event source includes the component identifier or the component connection relationship identifier associated with the log event, and the event type includes the running information, the alarm event or the fault event.
[0169] Step S603, it is judged whether the first alarm event, the first failure event and the first running information belonging to the same event source exist in the multiple log event sets;
[0170] Step S604, if the first alarm event, the first failure event and the first running information belonging to the same event source exist in the multiple log event sets, the first alarm event, the first failure event and the first running information are determined as the failure analysis result of the same event source;
[0171] Step S605, if the first alarm event, the first failure event and the first running information belonging to the same event source do not exist in the multiple log event sets, the multiple log events in the multiple log event sets are divided into multiple event item sets;
[0172] Step S606, the multiple event item sets are combined into multiple event item set groups two by two;
[0173] Step S607, the correlation degree parameters of the two event item sets in each event item set group are determined based on the correlation degree algorithm, and the correlation degree parameters corresponding to each event item set group are obtained;
[0174] Step S608, the event item set group with the correlation degree parameter greater than the correlation degree parameter threshold is determined as a frequent event item set group, and the failure analysis result of the server is determined based on the frequent event item set group.
[0175] The present application exemplary embodiment provides a kind of server failure analysis device, Figure 7 The functional module schematic diagram of the server failure analysis device according to the present application exemplary embodiment is shown, the device is server, as Figure 7 As shown, the server failure analysis device 700 includes:
[0176] The acquisition module 701 is configured to acquire the log file of the server;
[0177] The extraction module 702 is configured to process the log file based on log extraction template set, obtain the multiple log events of the server, and divide the multiple log events into multiple log event sets according to event attribute group, wherein the event attribute group includes event source and event type, the event source includes component identifier or component connection relationship identifier associated with the log event, and the event type includes running information, alarm event or failure event;
[0178] The first determining module 703 is configured to determine, if there are a first alarm event, a first fault event and first running information belonging to a same event source in the plurality of log event sets, the first alarm event, the first fault event and the first running information as a fault analysis result of the same event source, wherein the first alarm event, the first fault event and the first running information have the same event occurrence time.
[0179] Optionally, the apparatus further comprises a second determining module 704 configured to:
[0180] divide a plurality of remaining log events in the plurality of log event sets into a plurality of event item sets, wherein the event item sets comprise at least one log event belonging to a same event source, and the plurality of remaining log events are log events in the plurality of log event sets other than the first alarm event, the first fault event and the first running information belonging to the same event source;
[0181] combine the plurality of event item sets into a plurality of event item set groups two by two, wherein the event item set groups comprise two different event item sets;
[0182] determine, based on a correlation degree algorithm, a correlation degree parameter of the two event item sets in each event item set group, to obtain a correlation degree parameter corresponding to each event item set group;
[0183] determine, as a frequent event item set group, an event item set group whose correlation degree parameter is greater than a correlation degree parameter threshold, and determine a fault analysis result of the server based on the frequent event item set group.
[0184] Optionally, the event item set comprises a plurality of log events, and the second determining module 704 is configured to:
[0185] traverse the plurality of remaining log events to find at least one type of second alarm event, at least one type of second fault event and / or second running information belonging to a same event source, wherein a difference between the event occurrence time of the second alarm event, the second fault event and / or the second running information is less than or equal to a time length threshold;
[0186] combine the at least one type of second alarm event, the at least one type of second fault event and / or the second running information belonging to the same event source to obtain a first event item set associated with the same event source;
[0187] determine, as the plurality of event item sets, first event item sets respectively associated with a plurality of different event sources.
[0188] Optionally, the event item set comprises a plurality of log events, and the second determining module 704 is configured to:
[0189] traversing the plurality of remaining log events to find at least one type of second alarm event, at least one type of second fault event and / or second running information associated with each event source in an event source group, wherein the event source group comprises target component identifiers of two target components existing in the server in an association relationship, and a target component connection relationship identifier between the two target components;
[0190] combining the at least one type of second alarm event, the at least one type of second fault event and / or the second running information associated with each event source in the event source group to obtain a second event item set associated with the event source group;
[0191] determining the second event item sets respectively associated with a plurality of different event source groups as the plurality of event item sets.
[0192] Optionally, the event item set comprises a plurality of log events, and the second determining module 704 is configured to:
[0193] traversing the plurality of remaining log events to find any two of the first alarm event, the first fault event and the first running information associated with each event source in an event source group;
[0194] combining the any two of the first alarm event, the first fault event and the first running information associated with each event source in the event source group to obtain a third event item set associated with the event source group;
[0195] determining the third event item sets respectively associated with a plurality of different event source groups as the plurality of event item sets.
[0196] Optionally, the event item set comprises a plurality of log events, and the second determining module 704 is configured to:
[0197] traversing the plurality of remaining log events to find any two of the first alarm event, the first fault event and the first running information associated with each event source in an event source group;
[0198] combining the any two of the first alarm event, the first fault event and the first running information associated with each event source in the event source group to obtain a third event item set associated with the event source group;
[0199] determining the third event item sets respectively associated with a plurality of different event source groups as the plurality of event item sets.
[0200] Optionally, the log file comprises a plurality of log information, and the second determining module 704 is configured to:
[0201] For each log entry, iterate through each log extraction template in the log extraction template set and perform keyword matching;
[0202] If there is a target string in the log information that matches each keyword in the log extraction template, then the target strings that match each keyword are combined to obtain the log event associated with the log information;
[0203] By combining the log events associated with each log message, multiple log events of the server can be obtained.
[0204] Optionally, the acquisition module 701 is configured to:
[0205] Obtain the server's log files from the previous fault analysis period;
[0206] The fault analysis result that identifies the first alarm event, the first fault event, and the first operational information as originating from the same event includes:
[0207] The first alarm event, the first fault event, and the first operational information are determined to be the fault analysis results of the same event source for the server in the previous fault analysis cycle.
[0208] An exemplary embodiment of this application also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the electronic device to perform a method according to an embodiment of this application.
[0209] An exemplary embodiment of this application also provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of this application.
[0210] like Figure 8 As shown, an exemplary embodiment of this application also provides a computer program product 800, including a computer program 801, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of this application.
[0211] refer to Figure 9The present invention describes a structural block diagram of an electronic device 900 that can serve as a terminal device of this application, which is an example of a hardware device that can be applied to various aspects of this application. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.
[0212] like Figure 9 As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded into a random access memory (RAM) 903 from a storage unit 908. The RAM 903 may also store various programs and data required for the operation of the electronic device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0213] Multiple components in electronic device 900 are connected to I / O interface 905, including: input unit 906, output unit 907, storage unit 908, and communication unit 909. Input unit 906 can be any type of device capable of inputting information to electronic device 900. Input unit 906 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 907 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 908 may include, but is not limited to, disk and optical disk. Communication unit 909 allows electronic device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0214] The computing unit 901 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs various methods and processes described above. For example, in some embodiments, the methods of the present embodiments can be implemented as a computer software program tangibly embodied in a machine readable medium, such as the storage unit 908. In some embodiments, portions or all of the computer program can be loaded and / or installed onto the electronic device 900 via the ROM 902 and / or the communication unit 909. In some embodiments, the computing unit 901 can be configured to perform the methods of the present embodiments by way of other any suitable means, such as by way of firmware.
[0215] Program code for carrying out methods of the present application can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be retrieved from a machine readable medium or from a propagated signal. The machine readable medium can be a storage medium, such as a magnetic, optical, or semiconductor storage medium. The program code can be downloaded from an on-line service or electronic bulletin board, over a communications network, such as the Internet, or via an electronic mail and / or fax.
[0216] In the context of the present application, a machine readable medium can be a tangible medium that can contain or store program for use by or in connection with an instruction execution system, apparatus, or device. The machine readable medium can be a machine readable signal medium or a machine readable storage medium. The machine readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0217] As used in this application, the terms "machine-readable medium" and "computer- readable medium" refer to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0218] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0219] In the embodiments described above, the functions of the flow or the functions described can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented in software, the functions can be embodied as one or more computer programs or instructions that are executed on or by one or more computers. The computer programs or instructions can be stored on or transferred from one computer-readable medium to another computer-readable medium, for example, a computer program product, an article of manufacture, a memory device, a floppy disk, a CD-ROM, one or more optical or other discs, a hard disk, a magnetic tape, a carrier wave, a transmitted, received or other electronically available medium, or a combination of these or any other medium from which a computer can read. The computer programs or instructions can be executed on or by one or more computers, servers, terminals, user devices, or other programmable devices to produce the functions described in the embodiments described above. The computer programs or instructions can be stored in any computer-readable medium, or transmitted from one computer-readable medium to another computer-readable medium, for example, a computer program product, an article of manufacture, a memory device, a floppy disk, a CD-ROM, one or more optical or other discs, a hard disk, a magnetic tape, a carrier wave, a transmitted, received or other electronically available medium, or a combination of these or any other medium from which a computer can read. The computer-readable medium can be a computer- accessible medium, a server, a data center, or other data storage device that is integrated with or includes one or more computer-readable media.
[0220] Although the present application has been described in connection with certain specific features and embodiments thereof, it is to be understood that it is intended to cover all modifications and variations of this application which are within the scope of the appended claims and their equivalents. Accordingly, the description and drawings are to be regarded as illustrative in nature and not as restrictive. It is intended that all such modifications and variations are included within the scope of the present application as defined by the following claims and their equivalents.
Claims
1. A method of server failure analysis, the method comprising: The method comprises: obtaining a log file of a server; processing the log file based on a log extraction template set to obtain a plurality of log events of the server, and grouping the plurality of log events into a plurality of log event sets according to event attribute groups, wherein the event attribute groups comprise event sources and event types, the event sources comprise component identifiers or component connection relationship identifiers associated with the log events, and the event types comprise running information, alarm events, or fault events; if a first alarm event, a first fault event, and first running information belonging to a same event source exist in the plurality of log event sets, the first alarm event, the first fault event, and the first running information are determined as a fault analysis result of the same event source, wherein the first alarm event, the first fault event, and the first running information have the same event occurrence time.
2. The server failure analysis method of claim 1, wherein, The method further comprises: dividing a plurality of remaining log events in the plurality of log event sets into a plurality of event item sets, wherein the event item sets comprise at least one log event belonging to a same event source, and the plurality of remaining log events are log events other than the first alarm event, the first fault event, and the first running information belonging to the same event source in the plurality of log event sets; combining the plurality of event item sets into a plurality of event item set groups two by two, wherein two event item sets in the event item set group are different; determining an association degree parameter of the two event item sets in each event item set group based on an association degree algorithm to obtain an association degree parameter corresponding to each event item set group; determining an event item set group with an association degree parameter greater than an association degree parameter threshold as a frequent event item set group, and determining a fault analysis result of the server based on the frequent event item set group.
3. The server failure analysis method of claim 2, wherein, The event item sets comprise a plurality of log events, and the dividing of the plurality of remaining log events in the plurality of log event sets into the plurality of event item sets comprises: traversing the plurality of remaining log events to find at least one type of second alarm event, at least one type of second fault event, and / or second running information belonging to a same event source, wherein a difference between event occurrence times of the second alarm event, the second fault event, and / or the second running information is less than or equal to a time length threshold; combining the at least one type of second alarm event, the at least one type of second fault event, and / or the second running information belonging to the same event source to obtain a first event item set associated with the same event source; determining first event item sets associated with a plurality of different event sources respectively as a plurality of event item sets.
4. The server failure analysis method of claim 2, wherein, The event item sets comprise a plurality of log events, and the dividing of the plurality of remaining log events in the plurality of log event sets into the plurality of event item sets comprises: traversing the plurality of remaining log events to find at least one type of second alarm event, at least one type of second failure event and / or second running information associated with each event source in an event source group, wherein the event source group comprises target component identifiers of two target components existing in the server in an associated relationship and a target component connection relationship identifier between the two target components; combining the at least one type of second alarm event, the at least one type of second failure event and / or the second running information associated with each event source in the event source group to obtain a second event item set associated with the event source group; determining the second event item sets respectively associated with a plurality of different event source groups as the plurality of event item sets.
5. The server failure analysis method of claim 2, wherein, The event item set comprises a plurality of log events, and the dividing the plurality of remaining log events in the plurality of log event sets into a plurality of event item sets comprises: traversing the plurality of remaining log events to find any two of the first alarm event, the first failure event and the first running information associated with each event source in an event source group; combining any two of the first alarm event, the first failure event and the first running information associated with each event source in the event source group to obtain a third event item set associated with the event source group; determining the third event item sets respectively associated with a plurality of different event source groups as the plurality of event item sets.
6. The server failure analysis method of claim 2, wherein, The event item set comprises a plurality of log events, and the dividing the plurality of remaining log events in the plurality of log event sets into a plurality of event item sets comprises: traversing the plurality of remaining log events to find any two of the first alarm event, the first failure event and the first running information associated with each event source in an event source group; combining any two of the first alarm event, the first failure event and the first running information associated with each event source in the event source group to obtain a third event item set associated with the event source group; determining the third event item sets respectively associated with a plurality of different event source groups as the plurality of event item sets.
7. The server failure analysis method of claim 1, wherein, The log file comprises a plurality of log information, and the processing the log file based on the log extraction template set to obtain a plurality of log events of the server comprises: traversing each log extraction template in the log extraction template set to perform keyword matching for each piece of log information; if there is a target string matching each keyword in the log extraction template in the log information, combining the target string matching each keyword to obtain the log event associated with the log information; combining the log event associated with each piece of log information to obtain the plurality of log events of the server.
8. The server fault analysis method of claim 1, wherein, The obtaining the log file of the server comprises: obtaining the log file of the server in the last failure analysis period; The determining the first alarm event, the first failure event and the first running information as the failure analysis result of the same event source comprises: The first alarm event, the first failure event, and first running information are determined as the failure analysis result of the same event source of the server in a previous failure analysis period.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory, wherein the computer program comprises instructions that, when executed by the processor, cause the electronic device to perform the method of any one of claims 1-8. The processor executes the computer program to implement the method of any one of claims 1 to 8.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method of any one of claims 1 to 8.