Server fault diagnosis method and device

By modifying the serial port configuration before the server operating system is started and using the SOL channel to transmit log data, the problem of server fault diagnosis that requires a physical serial port line in the existing technology is solved, and fast and convenient multi-server fault diagnosis and system downtime detection are achieved.

CN113835942BActive Publication Date: 2025-09-23NEW H3C TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111108372.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-22
Publication Date
2025-09-23
Estimated Expiration
2041-09-22

AI Technical Summary

Technical Problem

Existing server fault diagnosis methods require logging into the operating system and rely on physical serial ports. They are unable to perform batch diagnosis, are difficult to manage, and are time-consuming. They also make it difficult to diagnose problems that prevent the operating system from being accessed.

Method used

Before the server operating system starts, modify the serial port configuration through the IPMI interface, transfer the serial port log data to the virtual serial port, and then transmit it to the local computer through the SOL channel. Read the log data to determine the fault type.

Benefits of technology

It enables rapid diagnosis of multiple server failures without the need for physical serial lines, saving resources, facilitating management, reducing time costs, and being able to diagnose system downtime.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113835942B_ABST
    Figure CN113835942B_ABST
Patent Text Reader

Abstract

The present application provides a server fault diagnosis method and device. The method includes: connecting to at least one server; before the operating system of the server is started, notifying the server to modify the configuration of multiple serial ports on the server to transmit the log data of the modified serial port to the first virtual serial port on the server; transmitting the log data received by the first virtual serial port to the local second virtual serial port via the SOL channel; reading the log data of the second virtual serial port to determine the fault type. The technical solution of the present application can save the time cost of fault diagnosis and reduce the human resource consumption of fault diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to computer fault diagnosis, and in particular to a server fault diagnosis method and apparatus. Background Art

[0002] Various faults often occur during server deployment and operation. The existing fault diagnosis method requires logging into the OS (operating system), modifying the GRUB (boot loader) file, and redirecting the OS serial port to the BIOS (Basic Input Output System) serial port each time.

[0003] This method requires a physical serial port and an additional serial cable to retrieve logs. It also requires accessing the fault diagnosis service files within the operating system. This method is not suitable for diagnosing issues where the operating system cannot be accessed. Collecting logs for these issues requires repeated replication, and there is a risk of unsuccessful replication for low-probability issues. Furthermore, since logging into the operating system is required each time, this diagnostic method cannot diagnose multiple servers in batches. This makes management difficult, error-prone, and time-consuming. Summary of the Invention

[0004] The present invention provides a server fault diagnosis method, the method comprising:

[0005] Connecting to at least one server; before starting the operating system of the server, notifying the server to modify the configuration of multiple serial ports on the server so as to transmit log data of the modified serial ports to a first virtual serial port on the server; transmitting the log data received by the first virtual serial port to a local second virtual serial port via a SOL channel; and reading the log data of the second virtual serial port to determine the type of fault.

[0006] The present invention provides a server fault diagnosis device, the device comprising:

[0007] A server serial port modification module is used to connect to the server before the server operating system is started, notify the server to modify the configuration of multiple serial ports of the server, so as to transmit log data of the modified serial port to a first virtual serial port on the server, and transmit the log data received by the first virtual serial port to a local second virtual serial port via a SOL channel;

[0008] The fault determination module is configured to read log data of the second virtual serial port and determine a fault type according to the log data.

[0009] Based on the above technical solution, in an embodiment of the present invention, by connecting at least one server, the log data of multiple serial ports on a server can be sent to a local computer, thereby quickly finding the faulty server and determining the type of fault. In addition, before the server operating system is started, the serial port log data is transmitted via the SOL channel, which facilitates the detection of whether there is a system downtime. At the same time, it avoids the need to use an additional serial port line to connect each server, saving resources and facilitating management. Since the serial port log data of the server are all sent to a local computer, multiple faults of multiple servers can be diagnosed at once, saving time and cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the specification and, together with the description, serve to explain the principles of the specification.

[0011] Figure 1 This is a flow chart of a fault diagnosis method according to an embodiment of the present disclosure.

[0012] Figure 2 This is a flow chart of another fault diagnosis method according to an embodiment of the present specification.

[0013] Figure 3 This is a network architecture diagram of an embodiment of this specification.

[0014] Figure 4 This is a block diagram of a fault diagnosis device according to an embodiment of the present specification.

[0015] Figure 5 This is a block diagram of a computer device where a fault diagnosis device according to one embodiment of the present specification is located. DETAILED DESCRIPTION

[0016] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with this specification. Rather, they are merely examples of apparatus and methods consistent with certain aspects of this specification, as detailed in the appended claims.

[0017] The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this specification. As used in this specification and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0018] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, first information may also be referred to as second information, and similarly, second information may also be referred to as first information without departing from the scope of this specification. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to determining."

[0019] Next, the embodiments of this specification are described in detail.

[0020] like Figure 1 As shown, Figure 1 The present invention provides a flow chart of a server fault diagnosis method. The method can be used for servers with various operating systems, such as Linux, Windows, and Unix. The method includes:

[0021] Step 102: Connect to at least one server.

[0022] Step 104: Before the operating system of the server is started, the server is notified to modify the configuration of the multiple serial ports so as to transmit the log data of the modified serial ports to the first virtual serial port of the server.

[0023] Step 106: Transmit the log data received by the first virtual serial port to the local second virtual serial port via the SOL channel.

[0024] Step 108: Read the log data of the second virtual serial port to determine the fault type.

[0025] In the fault diagnosis method provided in the above-mentioned embodiment of the present application, step 102 connects to the server through a network transmission protocol to establish a communication channel.

[0026] In step 104, before the server operating system is started, an instruction is sent to notify the server to modify the serial port configuration.

[0027] The modified serial ports in this embodiment include: physical serial ports and OS serial ports.

[0028] The physical serial port includes a BIOS serial port and a BMC serial port.

[0029] Traditionally, accessing data from the BIOS or BMC serial ports requires connecting an additional serial cable. This not only requires the cost of purchasing the cable, but also requires manual connection. Furthermore, the cable's length is limited, making it impossible to send data to a remote terminal. This implementation uses the IPMI (Intelligent Platform Management Interface) interface to send commands to the server's BMC (Baseboard Management Controller), allowing the BMC to modify the serial port configuration based on the commands, thus eliminating the need for an additional serial cable.

[0030] As an example, a specific operation may be to bridge an idle UART serial port, connect the physical serial port to the UART serial port separately, and send the data of the physical serial port to the SOL (Serial Over Lan) virtual serial port on the server through the UART serial port.

[0031] By performing the above operations, data from the BIOS serial port and the BMC serial port are sent out through the SOL serial port, without the need for an additional serial port cable.

[0032] As an example, step 104 may also include replacing the first GRUB file on the server with a second GRUB file, where the second GRUB file has configuration information for redirecting the OS serial port to the first virtual serial port on the server, relative to the first GRUB file.

[0033] After the modification is complete, you can verify whether the modified configuration is effective.

[0034] The verification step includes determining whether the modified configuration is effective based on whether the server returns print information after restarting.

[0035] In step 106, the local sends a log data transmission instruction, and the SOL serial transmits the log data of the BIOS, BMC and OS serial ports of the server virtual serial port to the local virtual serial port via Ethernet.

[0036] In step 108, the local terminal reads the server log data through the local virtual serial port. Based on this log data, it is possible to diagnose which server has experienced a fault and determine the fault type based on the log content. Fault types are categorized as system downtime or non-system downtime. Because log data from the BIOS and BMC serial ports is included, this method can diagnose even a system downtime.

[0037] An embodiment is provided below to illustrate how to diagnose the fault type in practice.

[0038] In this embodiment, the local terminal initiates a connection to the server using the SSH protocol. After the connection is established, an operation instruction is sent to the server's BMC through the IPMI interface, so that the BMC modifies the serial port configuration based on the instruction.

[0039] The commands for modifying the BIOS serial port include:

[0040] Instructions for turning on the BIOS Debug mode switch;

[0041] Instructions for redirecting the BIOS serial port to the UART1 serial port.

[0042] The commands for modifying the BMC serial port include:

[0043] Command to open the BMC serial port;

[0044] Command to redirect the BMC serial port to the UART4 serial port;

[0045] BMC sets the command for UART1 serial port to bridge to UART4 serial port;

[0046] Instructions for transmitting data from the UART1 and UART4 serial ports via the SOL on the BMC.

[0047] The instructions for modifying the OS serial port include:

[0048] Instructions for replacing the GRUB file for your operating system.

[0049] The selected SSH connection protocol can also be replaced with HTTPS or Telnet (command line interface remote management protocol).

[0050] The UART1 and UART4 ports used in this example can be replaced by any two idle UART protocol ports, such as UART2 and UART3.

[0051] The above instructions can be executed in any order to achieve the same effect.

[0052] Regarding the instruction for replacing the GRUB file of the operating system, after the replacement is completed, the server is restarted to make the configuration of redirecting the OS serial port to the first virtual serial port of the server take effect.

[0053] After executing the above instructions, SOL's virtual serial port on the server will be able to receive log data from the BIOS serial port, BMC serial port, and OS serial port.

[0054] After sending a command to modify the serial port configuration, this embodiment further detects whether the serial port configuration has taken effect. The detection method includes the local terminal determining whether the modified configuration has taken effect based on whether it receives system power-on / off print information and BIOS power-on / off print information returned after the server is restarted. The system power-on / off print information is used to detect whether the modification to the GRUB file has taken effect, and the BIOS power-on / off print information is used to detect whether the modification to the BMC serial port and the BIOS serial port has taken effect. If the print information is received, the configuration has taken effect; if the print information is not received, the configuration has not taken effect.

[0055] Next, the local computer sends a log transfer instruction to the BMC for execution. The log transfer instruction transmits the data of the server virtual serial port to the local virtual serial port via SOL serial communication. The instruction can be as follows:

[0056] Run the command ipmitool.exe -U USERNAME -P PASSWORD -H BMCIP sol activate>os.log. USERNAME is the BMC username, PASSWORD is the password, and BMCIP is the server IP address. This command transfers data from the local virtual serial port to the local os.log file. If no os.log file exists locally, executing this command will automatically create one. This command is for illustrative purposes only and may not be the same as the preceding command.

[0057] After executing the above command, the local virtual serial port log data can be read by reading the os.log file. This log data will contain error information, which may include keywords or unique numerical identifiers. By searching the fault list, the fault type can be determined based on these keywords or unique numerical identifiers. The fault list can be compiled from publicly available information such as server manuals, operating system manuals, and documentation provided by software vendors. Fault types can be categorized as system downtime or non-system downtime, as needed.

[0058] In another embodiment of the present application, a method for performing fault diagnosis based on the determined fault type is also provided. Figure 2 As shown,

[0059] Step 204: Start the fault diagnosis service.

[0060] Step 206: Obtain the fault type.

[0061] Step 208: If the fault type is a system downtime fault, then as shown in step 210a, the downtime type in the log data is recorded; if it is not a system downtime fault, then as shown in step 210b, the faulty server is connected.

[0062] Step 212b: Check whether the fault diagnosis service is normal.

[0063] Step 214b: After connecting to the fault server, obtain the fault diagnosis file generated by the fault diagnosis service.

[0064] Step 216b: Determine the specific fault according to the fault diagnosis file.

[0065] In the above-mentioned method for further detecting specific faults, the fault diagnosis service can be kdump or sosreport, or kdump and sosreport are combined. In step 212b, the method for detecting whether the fault diagnosis service is normal includes setting a check digit for each fault diagnosis service, if the check digit is a preset value, then the service is normal.

[0066] An embodiment is provided below to illustrate how to further diagnose a fault.

[0067] In this embodiment, the fault diagnosis service is automatically started when the operating system is started. After reading the log data files of multiple servers, the fault type is determined based on the keywords in the log or the unique digital identifier, and the fault type is divided into system downtime fault and non-system downtime fault. If the read log data shows a system downtime fault, it is only necessary to record the fault type, or directly record the log data recording the fault keyword or unique digital identifier. If it is a non-system downtime fault, the fault server is connected in batches through the SSH protocol. During the connection stage, if the username and password verification method is used, the username and password need to be entered. If the key connection method is used, the public key and private key files need to be generated in advance, and the key authentication is set on the server. This operation is a one-time operation, and the username and password do not need to be entered for subsequent connections.

[0068] After the connection is successful, set a function check for each fault diagnosis service. When the check bit is 1, the current fault diagnosis service is functioning normally. If the fault diagnosis service is functioning normally, view the sosreport log file of the faulty server in the local SSH connection application window, debug directly in the sosreport debug module, and filter the sosreport log. Alternatively, use the SFTP module to directly download the sosreport log file and kdump core file.

[0069] Sosreport's log files and kdump's core files record the system status at the time of failure and the error reports generated by the application. With this information, staff can determine the specific failure.

[0070] The SSH connection protocol used in this example to connect to the faulty server can be replaced by other remote management protocols, such as RDP (remote desktop protocol), RFB (Remote Frame Buffer, graphical remote management protocol) and Telnet (command line interface remote management protocol). SFTP can be replaced by other file transfer protocols, such as FTP protocol. You can choose to establish an SFTP connection at the same time as the SSH connection is established, or you can establish the SFTP connection after determining that you want to download the file. After detecting the status of the fault diagnosis service in step 212b, you can decide whether to perform the next operation on the fault diagnosis service based on specific needs. For example, if the entire server system is in a production environment, the fault diagnosis service will not be restarted. If the server system is in a test environment, try to restart the fault diagnosis service. The operation strategies described above are just examples, and different strategies should be adopted according to different businesses.

[0071] like Figure 3 The figure shows a network architecture diagram of an embodiment. The network entities involved include server 310, server 320 and terminal 340. Server 310 and server 320 can be servers with different operating systems, and the server needs to be connected to Ethernet. Terminal 340 can be a personal computer, mobile phone, tablet computer, industrial console and other smart devices. Terminal 340 is connected to server 310 and server 320 respectively via Ethernet. In this example, the fault diagnosis method is implemented using fault diagnosis software running on terminal 340. The fault diagnosis software initiates a connection and notifies the server to modify the serial port configuration. After the serial port configuration is successfully modified, the server will send the serial port log data to terminal 340 through the SOL channel.

[0072] The SOL channel is established on the network 330 .

[0073] Serial port log data includes log data of the BIOS serial port, BMC serial port, and OS serial port.

[0074] In this example, the fault diagnosis software reads the received log data and determines the fault type. This determination can be made by searching a pre-stored fault list based on keywords or unique digital identifiers in the log data. The fault list can be compiled from publicly available information such as server manuals, operating system manuals, and documentation provided by software vendors.

[0075] It should be noted that although this example runs under Ethernet, if a forwarding device is added to the Ethernet, which is connected to the server and terminal 340, and forwards the server's log data to terminal 340, and forwards the instructions of terminal 340 to the server, then this example can be extended to a wide area network. Therefore, this solution cannot be limited to implementation under Ethernet.

[0076] Corresponding to the aforementioned method embodiments, this specification also provides embodiments of an apparatus and a terminal to which it is applied.

[0077] like Figure 4 As shown, Figure 4 FIG. 4 is a block diagram of a fault diagnosis device 400 according to an embodiment of the present disclosure, wherein the device includes:

[0078] The server serial port modification module 410 is used to connect to the server before the server operating system is started, and notify the server to modify the configuration of multiple serial ports of the server so as to transmit the log data of the modified serial port to the first virtual serial port on the server, and transmit the log data received by the first virtual serial port to the local second virtual serial port via the SOL channel.

[0079] The fault diagnosis service diagnosis module 420 is configured to read the log data of the second virtual serial port and determine the fault type according to the log data.

[0080] Optionally, the diagnostic module 420 may further perform fault diagnosis. Based on the fault type, if the fault type is a non-operating system downtime fault, the diagnostic module 420 may connect to the fault server to obtain a fault diagnosis file generated by the fault diagnosis service, determine a fault diagnosis result based on the fault diagnosis file, and if the fault type is a system downtime fault, record the downtime type based on the log data.

[0081] The specific details of the implementation process of the functions and effects of each module in the above-mentioned device can be found in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.

[0082] The embodiments of the fault diagnosis device of this specification can be applied to computer devices, such as servers or terminal devices. The device embodiments can be implemented through software, hardware, or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of the file processing in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for execution. From the hardware level, such as Figure 5 The figure is a hardware structure diagram of the computer device where the file processing device of the embodiment of this specification is located, except Figure 5 In addition to the processor 510, memory 530, network interface 520, and non-volatile memory 540 shown, the server or electronic device where the device 531 is located in the embodiment may also include other hardware according to the actual function of the computer device, which will not be described in detail.

[0083] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are only schematic, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this specification. Those of ordinary skill in the art can understand and implement it without paying any creative work.

[0084] Accordingly, an embodiment of this specification also provides an electronic device, which includes a memory, a processor, and a computer program stored and running on the memory, and when the processor executes the program, the fault diagnosis method in any of the above embodiments is implemented.

[0085] Accordingly, an embodiment of this specification further provides a computer storage medium on which a computer program is stored. When the program is executed by a processor, the fault diagnosis method in any of the above embodiments is implemented.

[0086] The present application may take the form of a computer program product implemented on one or more storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing program code. Computer-usable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. The information can be computer-readable instructions, data structures, modules of a program, or other data. Examples of computer storage media include but are not limited to: phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0087] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the instructions disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0088] It should be understood that the present description is not limited to the exact structure that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present description is limited only by the appended claims.

[0089] The above description is only a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this specification should be included in the scope of protection of this specification.

Claims

1. A method for diagnosing server faults, characterized in that: The method comprises: Connecting to at least one server; Before the operating system of the server is started, notifying the server to modify the configuration of multiple serial ports on the server so as to transmit log data of the modified serial ports to the first virtual serial port on the server; Transmitting the log data received by the first virtual serial port to the local second virtual serial port via the SOL channel; Reading log data of the second virtual serial port to determine the fault type; The serial ports on the server include: a physical serial port and an OS serial port; and modifying the configuration of multiple serial ports on the server includes: Bridging the idle UART serial ports, connecting the physical serial ports to the UART serial ports separately, and sending the data of the physical serial ports to the first virtual serial port through the UART serial ports; The first GRUB file on the server is replaced with a second GRUB file, where the second GRUB file carries configuration information for redirecting the OS serial port to the first virtual serial port.

2. The method according to claim 1, characterized in that The step of notifying the server to modify the configuration of multiple serial ports on the server includes: An instruction is sent to the BMC baseboard management controller of the server through the IPMI interface, so that the BMC baseboard management controller modifies the serial port configuration based on the instruction.

3. The method according to claim 1, characterized in that After modifying the configuration of multiple serial ports on the server, the following steps are also included: Verify that the modified configuration takes effect.

4. The method according to claim 1, wherein Also includes the steps: Perform fault diagnosis based on the determined fault type.

5. The method according to claim 4, characterized in that The steps of the fault diagnosis include: Starting a fault diagnosis service on the server; If the fault type is a non-operating system downtime fault, connecting to the fault server, obtaining a fault diagnosis file generated by the fault diagnosis service, and determining a fault diagnosis result according to the fault diagnosis file; If the fault type is a system downtime fault, the downtime type is recorded according to the log data.

6. The method according to claim 5, characterized in that The method further comprises the steps of: A check digit is set for each fault diagnosis service. If the check digit is a preset value, the service is normal.

7. A server fault diagnosis device, characterized in that: The device comprises: A server serial port modification module is configured to connect to the server before the server operating system is started, notify the server to modify the configuration of multiple serial ports of the server, transmit log data of the modified serial port to a first virtual serial port on the server, and transmit the log data received by the first virtual serial port to a local second virtual serial port via a SOL channel; wherein the serial ports on the server include: a physical serial port and an OS serial port; a fault determination module, configured to read log data of the second virtual serial port and determine a fault type according to the log data; The server serial port modification module is specifically used to: bridge the idle UART serial port, connect the physical serial port to the UART serial port separately, and send the data of the physical serial port to the first virtual serial port through the UART serial port; replace the first GRUB file on the server with a second GRUB file, wherein the second GRUB file carries configuration information for redirecting the OS serial port to the first virtual serial port.

Citation Information

Patent Citations

  • Method, device and system for binding virtual serial port and physical serial port

    CN102546840A

  • IPMI (intelligent platform management interface) based method for serial port redirection

    CN104363117A

  • Server fault discovery method and device, electronic equipment and storage medium

    CN112988439A