Hardware fault positioning method, server system and server

By using a CPLD to collect and store hardware signal data in real time when the BMC detects abnormal hardware functional modules, the problem of low efficiency in hardware fault location in the prior art is solved, and rapid fault location without disassembling the machine is achieved, thus reducing maintenance costs.

CN122019283APending Publication Date: 2026-05-12SHANGHAI EVEX INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI EVEX INFORMATION TECHNOLOGY CO LTD
Filing Date
2025-12-30
Publication Date
2026-05-12

Smart Images

  • Figure CN122019283A_ABST
    Figure CN122019283A_ABST
Patent Text Reader

Abstract

The invention provides a hardware fault positioning method, a server system and a server, and relates to the technical field of servers. According to the hardware fault positioning method, when a BMC monitors that an abnormal hardware function module exists in a server system, hardware signal data corresponding to the abnormal hardware function module are collected in real time through a CPLD, the hardware signal data are stored in a nonvolatile memory chip, and when the BMC monitors that the abnormal hardware function module is abnormally reproduced, the hardware signal data are stored in the nonvolatile memory chip. The hardware signal data corresponding to the abnormal hardware function module is read from the nonvolatile memory chip, the hardware signal data can be obtained without disassembling a machine, the root cause of the hardware fault is positioned based on the hardware signal data, rapid positioning of the hardware fault root cause is realized, and the positioning efficiency of the hardware fault is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of server technology, and in particular to a hardware fault location method, server system, and server. Background Technology

[0002] In the deployment and maintenance of server products, especially AI servers, rapid fault location and analysis are core requirements for ensuring the stable operation of server systems. Currently, AI servers commonly employ liquid cooling technology to cope with the heat dissipation pressure brought by high-density computing, resulting in a significantly more complex structure than traditional servers. For example, hardware components such as the motherboard, power module, and network interface card (wired network) of a liquid-cooled server need to be integrated into a closed liquid cooling pipeline system, and the acquisition of hardware signals and fault location rely on built-in monitoring modules (such as the Baseboard Management Controller (BMC)).

[0003] Currently, locating server hardware failures primarily relies on the BMC's one-click logging function. Specifically, the BMC generates logs by monitoring the real-time operating status of hardware functional modules (such as power supply status, core component temperature, voltage stability, etc.), and uses these logs to initially locate the faulty hardware functional module.

[0004] However, the above method of locating server hardware failures through BMC logs has the problem of low location efficiency. Summary of the Invention

[0005] This application provides a hardware fault location method, server system, and server to solve the problem of low location efficiency in related technologies that use BMC logs to locate server hardware faults.

[0006] In a first aspect, this application provides a hardware fault location method applied to a server system. The server system includes a BMC, a Complex Programmable Logic Device (CPLD), and a non-volatile memory chip connected in sequence. The method includes: in response to the BMC detecting an abnormal hardware functional module in the server system, sending a hardware signal acquisition command to the CPLD; in response to the hardware signal acquisition command, the CPLD acquires hardware signal data corresponding to the abnormal hardware functional module in real time, and transmits the hardware signal data to the non-volatile memory chip for storage via a first communication link; in response to the BMC detecting the abnormal hardware functional module recurring, the BMC reads the hardware signal data corresponding to the abnormal hardware functional module from the non-volatile memory chip via a second communication link, and the hardware signal data is used to locate the root cause of the hardware fault in the abnormal hardware functional module.

[0007] In one possible implementation, the hardware signal acquisition command carries communication link identification information, and the method further includes: the CPLD responding to the hardware signal acquisition command by switching the current communication link to a first communication link for connecting the CPLD and the non-volatile memory chip according to the communication link identification information.

[0008] In one possible implementation, in response to the BMC detecting the abnormal recurrence of an abnormal hardware functional module, hardware signal data corresponding to the abnormal hardware functional module is read from the non-volatile memory chip via a second communication link. This includes: in response to the BMC detecting the abnormal recurrence of an abnormal hardware functional module, the BMC sends a communication link switching command to the CPLD; the CPLD, in response to the communication link switching command, switches the current communication link to a second communication link for connecting the BMC and the non-volatile memory chip; and the BMC reads the hardware signal data corresponding to the abnormal hardware functional module from the non-volatile memory chip via the second communication link.

[0009] In one possible implementation, the hardware signal acquisition instruction also carries the hardware signal type and sampling frequency. The CPLD responds to the hardware signal acquisition instruction and, according to the hardware signal acquisition instruction, acquires the hardware signal data corresponding to the abnormal hardware function module in real time, including: the CPLD responds to the hardware signal acquisition instruction and, according to the hardware signal type and sampling frequency, acquires the hardware signal data corresponding to the abnormal hardware function module in real time.

[0010] In one possible implementation, the hardware signal data is transmitted to a non-volatile memory chip for storage via a first communication link, including: the CPLD transmitting the hardware signal data to the non-volatile memory chip via the first communication link; and the non-volatile memory chip classifying and storing the hardware signal data according to its spatial region.

[0011] In one possible implementation, in response to the BMC detecting the abnormal recurrence of an abnormal hardware functional module, the method further includes: the BMC sending a hardware signal stop acquisition command to the CPLD; and the CPLD responding to the hardware signal stop acquisition command by stopping the acquisition of hardware signal data corresponding to the abnormal hardware functional module.

[0012] Secondly, this application provides a server system, including: a BMC, a CPLD, and a non-volatile memory chip connected in sequence for communication;

[0013] The BMC is used to monitor the operating status of hardware functional modules in the server system in real time, and to send hardware signal acquisition instructions to the CPLD when an abnormal hardware functional module is detected.

[0014] The CPLD is used to, in response to a hardware signal acquisition command, acquire hardware signal data corresponding to the abnormal hardware function module in real time according to the hardware signal acquisition command, and transmit the hardware signal data to a non-volatile memory chip for storage based on the first communication link.

[0015] The BMC is also used to, in response to the detection of abnormal recurrence of an abnormal hardware functional module, read hardware signal data corresponding to the abnormal hardware functional module from the non-volatile memory chip based on the second communication link. The hardware signal data is used to locate the root cause of the hardware failure of the abnormal hardware functional module.

[0016] In one possible implementation, the CPLD includes a Serial Peripheral Interface (SPI) controller module and a communication switch switching module. The hardware signal acquisition command carries communication link identification information, hardware signal type, and sampling frequency.

[0017] The SPI controller module is used to collect hardware signal data corresponding to abnormal hardware function modules in real time according to the hardware signal type and sampling frequency, and transmit the hardware signal data to a non-volatile memory chip for storage based on the first communication link.

[0018] The communication switch switching module is used to switch the current communication link to the first communication link for connecting the SPI controller module and the non-volatile memory chip based on the communication link identification information.

[0019] In one possible implementation, the CPLD also includes a random access module, which is connected to the BMC via a third communication link;

[0020] The random access module is used to receive hardware signal acquisition commands, send communication link identification information to the communication switch switching module, and send hardware signal type and sampling frequency to the SPI controller module;

[0021] The random access module is also used to receive a communication link switching instruction sent by the BMC based on the third communication link, and send a communication link switching instruction to the communication switch switching module to enable the communication switch switching module to switch the current communication link to a second communication link for connecting the BMC and the non-volatile memory chip according to the communication link switching instruction.

[0022] Thirdly, this application provides a server, including: a processor and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the hardware fault location method provided in the first aspect above.

[0023] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the hardware fault location method provided in the first aspect above.

[0024] Fifthly, this application provides a computer program product, comprising: a computer program that, when executed by a processor, implements the hardware fault location method provided in the first aspect above.

[0025] The hardware fault location method, server system, and server provided in this application are applied to a server system comprising a BMC, a Complex Programmable Logic Device (CPLD), and a non-volatile memory chip connected in sequence. In response to the BMC detecting an abnormal hardware functional module in the server system, the hardware fault location method sends a hardware signal acquisition command to the CPLD. The CPLD then responds to the hardware signal acquisition command by acquiring hardware signal data corresponding to the abnormal hardware functional module in real time, and transmits the hardware signal data to the non-volatile memory chip for storage via a first communication link. Further, in response to the BMC detecting a recurrence of the abnormal hardware functional module, the BMC reads the hardware signal data corresponding to the abnormal hardware functional module from the non-volatile memory chip via a second communication link, which is used to locate the root cause of the hardware fault in the abnormal hardware functional module. In this application, when the BMC detects an abnormal hardware functional module in the server system, it uses a CPLD to collect the hardware signal data corresponding to the abnormal hardware functional module in real time and stores the hardware signal data in a non-volatile memory chip. Furthermore, when the BMC detects that the abnormal hardware functional module is recurring, it reads the hardware signal data corresponding to the abnormal hardware functional module from the non-volatile memory chip. This allows the hardware signal data to be obtained without disassembling the machine, and the root cause of the hardware failure can be located based on the hardware signal data, thus achieving rapid location of the root cause of the hardware failure and improving the efficiency of hardware failure location. Attached Figure Description

[0026] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0027] Figure 1 This is a schematic diagram of the server system provided in an embodiment of this application;

[0028] Figure 2 A flowchart illustrating the hardware fault location method provided in this application embodiment. Figure 1 ;

[0029] Figure 3 A schematic diagram of hardware signal waveforms provided in an embodiment of this application;

[0030] Figure 4 A flowchart illustrating the hardware fault location method provided in this application embodiment. Figure 2 ;

[0031] Figure 5 A schematic diagram of the address space of a non-volatile memory chip provided in an embodiment of this application;

[0032] Figure 6 This is a schematic diagram of the server structure provided in an embodiment of this application.

[0033] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0034] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0035] After the tested server is deployed in the customer's room, due to factors such as fluctuations in the power supply environment, changes in the pressure of the liquid cooling system, or hardware aging, the server may experience low-probability hardware abnormal events (such as alternating current loss (AC loss) events, inter-integrated circuit (I2C) communication interruptions, etc.).

[0036] In related technologies, server hardware fault location primarily relies on the BMC's one-click logging function. Specifically, the BMC generates logs by monitoring the real-time operating status of hardware functional modules (such as power supply status, core component temperature, voltage stability, etc.) and uses these logs to initially locate the faulty hardware functional module. However, the BMC logs can only pinpoint the malfunction of a specific hardware functional module. Data for locating the root cause of the hardware fault, such as the cause of the malfunction and the changing patterns of key signals affecting the malfunctioning module, can only be obtained through disassembling the machine and performing signal testing, resulting in low location efficiency. Furthermore, current AI servers have complex structures, especially liquid-cooled servers. Disassembling the machine requires pre-treatment operations such as draining and sealing pipelines, consuming significant manpower and time costs for each disassembly and reassembly, leading to high overall maintenance costs.

[0037] To address the problems existing in related technologies, this application embodiment, when the BMC detects an abnormal hardware functional module in the server system, uses a CPLD to collect the hardware signal data corresponding to the abnormal hardware functional module in real time and stores the hardware signal data in a non-volatile memory chip. Furthermore, when the BMC detects the abnormal hardware functional module recurring, it reads the hardware signal data corresponding to the abnormal hardware functional module from the non-volatile memory chip. This allows for the acquisition of hardware signal data without disassembling the machine, and the root cause of the hardware fault can be located based on the hardware signal data, achieving rapid location of the root cause of the hardware fault and improving the efficiency of hardware fault location.

[0038] The application scenarios of the embodiments of this application will be described below first.

[0039] The hardware fault location method provided in this application is applicable to fault location scenarios of complex hardware systems such as high-density AI servers and liquid-cooled servers.

[0040] The following is a combination of... Figure 1 The server system provided in the embodiments of this application will be described in detail.

[0041] Figure 1 This is a schematic diagram of the server system provided in an embodiment of this application. Figure 1 As shown, the server system includes a BMC, a CPLD, and a non-volatile memory chip that are connected in series.

[0042] Among them, BMC is used to monitor the operating status of hardware functional modules in the server system in real time, and send hardware signal acquisition instructions to CPLD when abnormal hardware functional modules are detected.

[0043] The CPLD is used to, in response to a hardware signal acquisition command, acquire hardware signal data corresponding to the abnormal hardware function module in real time according to the hardware signal acquisition command, and transmit the hardware signal data to a non-volatile memory chip for storage based on the first communication link.

[0044] The BMC is also used to, in response to the detection of abnormal recurrence of abnormal hardware functional modules, read hardware signal data corresponding to the abnormal hardware functional modules from the non-volatile memory chip based on the second communication link. The hardware signal data is used to locate the root cause of the hardware failure of the abnormal hardware functional modules.

[0045] For example, the CPLD can be an existing CPLD on the substrate of the server, or it can be a newly added CPLD. This application does not limit this, and the specific choice can be made according to the actual application requirements.

[0046] For example, such as Figure 1 As shown, the non-volatile memory chip is an external device independent of the CPLD, that is, the non-volatile memory chip is located outside the CPLD, and the CPLD and non-volatile memory chip are integrated on the same substrate (such as the motherboard and backplane of a server).

[0047] For example, the first communication link and the second communication link can be communication links based on the SPI protocol. In one possible implementation, the first communication link and the second communication link are based on, for example, Figure 1 The communication switch switching module shown uses a time-division multiplexing mechanism to share the SPI bus used to connect the communication switch switching module and the non-volatile memory chip.

[0048] like Figure 1 As shown, the first communication link is used to connect the CPLD and the non-volatile memory chip, enabling the writing of hardware signal data into the non-volatile memory chip; the second communication link is used to connect the BMC and the non-volatile memory chip, enabling the BMC to read hardware signal data from the non-volatile memory chip.

[0049] For example, a non-volatile memory chip can be a Flash chip.

[0050] like Figure 1As shown in the illustration, in the server system provided in this application embodiment, the BMC is communicatively connected to the terminal device. In one possible implementation, the BMC, based on a second communication link, reads the hardware signal data corresponding to the abnormal hardware functional module from the non-volatile memory chip, and then sends the hardware signal data to the terminal device. The terminal device then parses and processes the hardware signal data to obtain the signal waveform, and further uses the changing patterns of the signal waveform to quickly locate the root cause of the hardware fault. The terminal device contains an application program for parsing the hardware signal data.

[0051] For example, the terminal device can be a desktop computer, laptop computer, tablet computer, or all-in-one computer.

[0052] like Figure 1 As shown, optionally, the CPLD includes an SPI controller module and a communication switch switching module, and the hardware signal acquisition command carries communication link identification information, hardware signal type and sampling frequency;

[0053] The SPI controller module is used to collect hardware signal data corresponding to abnormal hardware function modules in real time according to the hardware signal type and sampling frequency, and transmit the hardware signal data to the non-volatile memory chip for storage based on the first communication link; the communication switch switching module is used to switch the current communication link to the first communication link used to connect the SPI controller module and the non-volatile memory chip according to the communication link identification information.

[0054] For example, such as Figure 1 As shown, the communication switch module includes a multiplexer (MUX). It is understood that the communication link switching function of the communication switch module can be implemented based on the MUX.

[0055] For example, the communication link identification information can be the identification information corresponding to the first communication link.

[0056] For example, hardware signal types include, but are not limited to, communication signals (such as I2C signals), power status monitoring data (such as power adapter AC power supply normal signal (Power Supply Unit AC OK, abbreviated as PSU_AC_OK), power enable signal (Enable, abbreviated as EN), and power status feedback signals (such as power PG (Power Good)).

[0057] For example, the hardware signal acquisition command may also carry the sampling duration. This application embodiment does not limit the sampling frequency and sampling duration carried in the hardware signal acquisition command; they can be determined according to actual application requirements.

[0058] like Figure 1 As shown, optionally, the CPLD also includes a random access module, which is connected to the BMC via a third communication link;

[0059] The random access module is used to receive hardware signal acquisition commands, send communication link identification information to the communication switch switching module, and send hardware signal type and sampling frequency to the SPI controller module. The random access module is also used to receive communication link switching commands sent by the BMC based on the third communication link, and send communication link switching commands to the communication switch switching module to enable the communication switch switching module to switch the current communication link to the second communication link for connecting the BMC and the non-volatile memory chip according to the communication link switching commands.

[0060] For example, the third communication link can be a communication link based on the I2C protocol.

[0061] like Figure 1 As shown, the random access module is equipped with random access memory (RAM). It is understandable that the functionality of the random access module can be implemented based on RAM.

[0062] The following is based on Figure 1 The server system shown is the execution entity. The specific implementation of the hardware fault location method provided in this application embodiment will be described in detail with reference to specific embodiments.

[0063] Figure 2 A flowchart illustrating the hardware fault location method provided in this application embodiment. Figure 1 .like Figure 2 As shown, a specific implementation of this hardware fault location method may include the following steps:

[0064] S201, in response to the BMC detecting an abnormal hardware function module in the server system, sends a hardware signal acquisition command to the CPLD.

[0065] In this step, one possible implementation is as follows: The BMC monitors the operating status of the hardware functional modules in the server system in real time (such as power supply status, core component temperature, voltage stability, etc.), and when it detects abnormal hardware events (such as AC Lost events, I2C communication interruptions, etc.), that is, when an abnormal hardware functional module occurs, it sends the hardware signal acquisition command to the random access module in the CPLD through the third communication link, namely the I2C communication link.

[0066] For example, when the BMC detects that the operating parameters of a hardware functional module exceed a preset threshold range, it indicates that the hardware functional module has malfunctioned. This embodiment of the application does not limit the preset threshold range; it can be determined based on actual application requirements.

[0067] For example, a hardware signal acquisition command may carry information such as sampling frequency, sampling duration, hardware signal type, and communication link identification information. The hardware signal type and communication link identification information are similar to those described above and will not be repeated here.

[0068] For example, the sampling frequency can be 1MHz (period 1µs), and the sampling duration can be the average sampling duration for each hardware signal, such as 2.86 minutes. This application does not specifically limit the sampling frequency and sampling duration; they can be determined according to actual application requirements.

[0069] For example, the hardware signal type may include one or more. This application embodiment does not specifically limit the number of hardware signal types, but can be determined according to the actual application requirements.

[0070] S202, the CPLD responds to the hardware signal acquisition command, and according to the hardware signal acquisition command, acquires the hardware signal data corresponding to the abnormal hardware function module in real time, and transmits the hardware signal data to the non-volatile memory chip for storage based on the first communication link.

[0071] In this step, one possible implementation is as follows: the random access module in the CPLD responds to the hardware signal acquisition command by sending the hardware signal type, sampling frequency, and sampling duration carried in the hardware signal acquisition command to the SPI controller module, so that the SPI controller module can acquire hardware signal data consistent with the hardware signal type from the abnormal hardware function module based on the sampling frequency and sampling duration; at the same time, the random access module sends the communication link identification information to the communication switch switching module, so that the communication switch switching module switches the current communication link to the first communication link consistent with the communication link identification information, which is used to connect the SPI controller module and the non-volatile memory chip, so that the SPI controller module can transmit the real-time acquired hardware signal data to the non-volatile memory chip for storage based on the first communication link.

[0072] It is understandable that in this step, the hardware signal data collected in real time by the SPI controller module includes normal hardware signal data and abnormal hardware signal data, or only abnormal hardware signal data.

[0073] It should be noted that in the hardware fault location method provided in this application embodiment, before the BMC detects the abnormal reproduction of the abnormal hardware functional module, if the amount of hardware signal data corresponding to the abnormal hardware functional module collected in real time by the CPLD is close to the upper limit of the storage space of the non-volatile memory chip, then the BMC or CPLD sends a data erase command to the non-volatile memory chip so that the non-volatile memory chip erases the corresponding hardware signal data in the order of the stored hardware signal data; or if the amount of hardware signal data corresponding to the abnormal hardware functional module collected in real time by the CPLD is close to the upper limit of the storage space of the non-volatile memory chip, then the CPLD automatically overwrites the earliest stored hardware signal data in the non-volatile memory chip according to the "first-in, first-out" rule, ensuring that the non-volatile memory chip always retains the latest hardware signal data.

[0074] S203, in response to the BMC detecting the abnormal reproduction of the abnormal hardware functional module, the BMC reads the hardware signal data corresponding to the abnormal hardware functional module from the non-volatile memory chip based on the second communication link. The hardware signal data is used to locate the root cause of the hardware failure of the abnormal hardware functional module.

[0075] For example, anomaly reproduction can be achieved by the BMC detecting anomalies in the operating parameters of a hardware functional module once again.

[0076] One possible implementation of this step is as follows: In response to the BMC detecting the abnormal recurrence of an abnormal hardware functional module, the BMC sends a communication link switching command to the random access module in the CPLD via the third communication link, namely the I2C communication link. This causes the random access module to send the communication link switching command to the communication switch switching module, which then switches the current communication link to a second communication link for connecting the BMC and the non-volatile memory chip. Based on the second communication link, the BMC reads the hardware signal data corresponding to the abnormal hardware functional module from the non-volatile memory chip and sends the read hardware signal data to the terminal device. The terminal device then parses and processes the hardware signal data to obtain the corresponding hardware signal waveform, and further uses this waveform to quickly locate the root cause of the hardware fault.

[0077] Figure 3 This is a schematic diagram of the hardware signal waveforms provided in an embodiment of this application. Figure 3 As shown, the hardware signal waveform can be a periodic square wave signal.

[0078] For example, if the non-volatile memory chip has a storage capacity of 1GB, capable of storing 8589.9s (143min) of data, a sampling frequency of 1MHz (period 1us), and only one hardware signal is collected, the hardware signal data read by the BMC from the non-volatile memory chip can fully encompass the hardware waveform data before, during, and after the hardware functional module malfunctions. In this scenario, a possible implementation for locating the root cause of hardware failure based on this hardware signal data could be: the BMC sends the read hardware signal data to the terminal device, allowing the terminal device to parse and process the hardware signal data to obtain the corresponding signal waveform, and then further quickly locate the root cause of the hardware failure based on the changing patterns of the signal waveform.

[0079] For example, if a non-volatile memory chip has a storage capacity of 1GB, capable of storing 8589.9s (143min) of data, a sampling frequency of 1MHz (period 1us), and 50 hardware signals are acquired, then if the average acquisition time for each hardware signal is 2.86min, it exceeds the acquisition and storage time of an oscilloscope. In this scenario, a possible implementation for locating the root cause of a hardware fault based on this hardware signal data could be: determining the time point corresponding to the abnormal recurrence of the abnormal hardware functional module based on the BMC logs, searching for the target hardware signal data corresponding to that time point in the hardware signal data based on that time point, and identifying the target hardware signal data and the hardware signal data following it as abnormal hardware signal data. The terminal device then parses this abnormal hardware signal data to obtain the corresponding abnormal hardware signal waveform, and quickly locates the root cause of the hardware fault based on the changing patterns of the abnormal hardware signal waveform.

[0080] It is understood that in the hardware fault location method provided in this application embodiment, abnormal hardware signal data of abnormal hardware functional modules in an abnormal state can be acquired in real time by a CPLD, and the abnormal hardware signal data can be analyzed by a terminal device connected to the BMC to reconstruct the abnormal hardware signal waveform, so that engineers can quickly locate the root cause of the hardware fault by combining the changing pattern of the abnormal hardware signal waveform. The abnormal hardware signal waveform obtained by analyzing the abnormal hardware signal data is consistent with the abnormal hardware signal waveform observed by an oscilloscope.

[0081] In this embodiment, when the BMC detects an abnormal hardware functional module in the server system, the CPLD collects the hardware signal data corresponding to the abnormal hardware functional module in real time and stores the hardware signal data in a non-volatile memory chip. Furthermore, when the BMC detects the abnormal hardware functional module recurring, the hardware signal data corresponding to the abnormal hardware functional module is read from the non-volatile memory chip. This allows hardware signal data to be obtained without disassembling the machine, reducing server maintenance costs. By locating the root cause of hardware failure based on this hardware signal data, the root cause of hardware failure can be quickly located, improving the efficiency of hardware failure location.

[0082] It is understood that the hardware fault location method provided in this application is beneficial for the rapid location and analysis of hardware faults in servers that have low probability of hardware abnormal events (such as AC Lost events, I2C communication interruptions, etc.).

[0083] Optionally, the hardware signal acquisition command carries communication link identification information. The hardware fault location method provided in this application embodiment further includes: the CPLD responds to the hardware signal acquisition command and switches the current communication link to a first communication link for connecting the CPLD and the non-volatile memory chip according to the communication link identification information.

[0084] The specific implementation method is similar to that described above, and will not be repeated here.

[0085] Optionally, in step S203, in response to the BMC detecting the abnormal recurrence of an abnormal hardware functional module, a possible implementation of the BMC reading the hardware signal data corresponding to the abnormal hardware functional module from the non-volatile memory chip based on the second communication link is as follows: In response to the BMC detecting the abnormal recurrence of an abnormal hardware functional module, the BMC sends a communication link switching command to the CPLD; in response to the communication link switching command, the CPLD switches the current communication link to the second communication link used to connect the BMC and the non-volatile memory chip; the BMC reads the hardware signal data corresponding to the abnormal hardware functional module from the non-volatile memory chip based on the second communication link.

[0086] The specific implementation method is similar to that described above, and will not be repeated here.

[0087] Optionally, the hardware signal acquisition instruction also carries the hardware signal type and sampling frequency. In step S202, the CPLD responds to the hardware signal acquisition instruction and acquires the hardware signal data corresponding to the abnormal hardware function module in real time according to the hardware signal acquisition instruction. One specific implementation method is that the CPLD responds to the hardware signal acquisition instruction and acquires the hardware signal data corresponding to the abnormal hardware function module in real time according to the hardware signal type and sampling frequency.

[0088] The specific implementation method is similar to that described above, and will not be repeated here.

[0089] The following is combined Figure 4 A detailed explanation is provided of a specific implementation method for transmitting hardware signal data to a non-volatile memory chip for storage based on the first communication link in step S202.

[0090] Figure 4 A flowchart illustrating the hardware fault location method provided in this application embodiment. Figure 2 .like Figure 4 As shown, a specific implementation of this hardware fault location method, which transmits hardware signal data to a non-volatile memory chip for storage based on a first communication link, may include the following steps:

[0091] S401, the CPLD transmits hardware signal data to the non-volatile memory chip based on the first communication link.

[0092] The specific implementation method is similar to that described above, and will not be repeated here.

[0093] The S402 non-volatile memory chip classifies and stores hardware signal data according to region space.

[0094] For example, the storage space of a non-volatile memory chip can be 256MB-1GB. This application does not limit the storage space size of the non-volatile memory chip; it can be determined based on actual application requirements.

[0095] One possible implementation of this step is as follows: A non-volatile memory chip stores hardware signal data in the corresponding preset address space based on the mapping relationship between the preset address space and the hardware signal type. The preset address space can be obtained by equally dividing the storage space of the non-volatile memory chip based on the sampling frequency, sampling duration, and number of bytes per sampling point of the hardware signal.

[0096] For example, the mapping relationship between the preset address space and the hardware signal type can be as follows: address space 1 to address space 10 are used to store the power EN signal, address space 11 to address space 100 are used to store the PSU_AC_OK signal, address space 101 to address space 1000 are used to store the I2C signal, and so on.

[0097] Another possible implementation of this step is as follows: Before the BMC sends the hardware signal acquisition instruction to the CPLD, it first sends an address space planning instruction to the CPLD to specify the memory space address range corresponding to different hardware signal types. When the CPLD acquires hardware signal data, it writes the corresponding hardware signal data into the specified address space of the non-volatile memory chip according to the hardware signal type identifier, thereby realizing the classified storage of different hardware signals.

[0098] Figure 5 This is a schematic diagram of the address space of a non-volatile memory chip provided in an embodiment of this application. Figure 5 As shown, the non-volatile memory chip may include multiple address spaces. Each address space is labeled with a start address (e.g., 0x0000) and an end address (e.g., 0x00FF).

[0099] For example, the range of the address space must match the amount of sampled data of the corresponding hardware signal (e.g., if a single signal requires 1MB of storage space, then the address range is 0x000000~0x00FFFF).

[0100] In this embodiment, the CPLD transmits hardware signal data to a non-volatile memory chip via a first communication link. The non-volatile memory chip then classifies and stores the hardware signal data according to a spatial region, enabling the BMC to quickly locate and read the hardware signal data in the corresponding spatial region based on the type of hardware signal, thereby improving the hardware fault location rate.

[0101] Optionally, in response to the BMC detecting the abnormal recurrence of the abnormal hardware functional module, the hardware fault location method provided in this application embodiment further includes: the BMC sending a hardware signal stop acquisition command to the CPLD; and the CPLD responding to the hardware signal stop acquisition command by stopping the acquisition of hardware signal data corresponding to the abnormal hardware functional module.

[0102] In this embodiment, one possible implementation is as follows: In response to the BMC detecting the abnormal reproduction of the abnormal hardware function module, the BMC sends a hardware signal stop acquisition command to the random access module in the CPLD through the third communication link. In response to the hardware signal stop acquisition command, the random access module sends the hardware signal stop acquisition command to the SPI controller module in the CPLD, so that the SPI controller module stops acquiring the hardware signal data corresponding to the abnormal hardware function module in response to the hardware signal stop acquisition command.

[0103] In this embodiment, in response to the BMC detecting the abnormal recurrence of an abnormal hardware functional module, the BMC sends a hardware signal stop acquisition command to the CPLD, causing the CPLD to stop acquiring the hardware signal data corresponding to the abnormal hardware functional module. This ensures that the hardware signal data stored in the non-volatile memory chip can completely cover the hardware signal data when the abnormal hardware functional module is in an abnormal state and before it is in an abnormal state, thereby improving the efficiency of hardware fault location.

[0104] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0105] Figure 6 This is a schematic diagram of the server structure provided in an embodiment of this application. Figure 6 As shown, the server 60 provided in this embodiment includes at least one processor 601 and a memory 602.

[0106] Optionally, the server 60 also includes a communication component 603. The processor 601, memory 602, and communication component 603 are connected via a bus 604.

[0107] In a specific implementation, at least one processor 601 executes computer execution instructions stored in memory 602, causing at least one processor 601 to perform the above-described method.

[0108] The specific implementation process of processor 601 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0109] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0110] The memory may include random access memory (RAM) and non-volatile memory (NVM), such as at least one disk storage device.

[0111] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0112] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0113] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.

[0114] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0115] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.

[0116] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0117] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0118] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0119] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0120] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0121] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A hardware fault location method, characterized in that, Applied to a server system, the server system comprising a baseboard management controller (BMC), a complex programmable logic device (CPLD), and a non-volatile memory chip connected in sequence, the method includes: In response to the BMC detecting an abnormal hardware functional module in the server system, a hardware signal acquisition command is sent to the CPLD; The CPLD responds to the hardware signal acquisition command, and according to the hardware signal acquisition command, acquires the hardware signal data corresponding to the abnormal hardware function module in real time, and transmits the hardware signal data to the non-volatile memory chip for storage based on the first communication link. In response to the BMC detecting the abnormal reproduction of the abnormal hardware functional module, the BMC reads the hardware signal data corresponding to the abnormal hardware functional module from the non-volatile memory chip based on the second communication link. The hardware signal data is used to locate the root cause of the hardware failure of the abnormal hardware functional module.

2. The hardware fault location method according to claim 1, characterized in that, The hardware signal acquisition command carries communication link identification information, and the method further includes: In response to the hardware signal acquisition command, the CPLD switches the current communication link to the first communication link used to connect the CPLD and the non-volatile memory chip according to the communication link identification information.

3. The hardware fault location method according to claim 1, characterized in that, In response to the BMC detecting the abnormal recurrence of the abnormal hardware functional module, the step of reading the hardware signal data corresponding to the abnormal hardware functional module from the non-volatile memory chip via the second communication link includes: In response to the BMC detecting the abnormal recurrence of the abnormal hardware functional module, the BMC sends a communication link switching command to the CPLD; In response to the communication link switching command, the CPLD switches the current communication link to the second communication link used to connect the BMC and the non-volatile memory chip. The BMC reads the hardware signal data corresponding to the abnormal hardware function module from the non-volatile memory chip based on the second communication link.

4. The hardware fault location method according to any one of claims 1 to 3, characterized in that, The hardware signal acquisition command also carries the hardware signal type and sampling frequency. The CPLD responds to the hardware signal acquisition command and, according to the command, acquires the hardware signal data corresponding to the abnormal hardware function module in real time, including: The CPLD responds to the hardware signal acquisition command and acquires the hardware signal data corresponding to the abnormal hardware function module in real time according to the hardware signal type and the sampling frequency.

5. The hardware fault location method according to any one of claims 1 to 3, characterized in that, The step of transmitting the hardware signal data to the non-volatile memory chip for storage based on the first communication link includes: The CPLD transmits the hardware signal data to the non-volatile memory chip based on the first communication link; The non-volatile memory chip classifies and stores the hardware signal data according to regional space.

6. The hardware fault location method according to any one of claims 1 to 3, characterized in that, In response to the BMC detecting the abnormal reproduction of the abnormal hardware functional module, the method further includes: The BMC sends a hardware signal to the CPLD to stop the acquisition command; In response to the hardware signal stop acquisition command, the CPLD stops acquiring hardware signal data corresponding to the abnormal hardware function module.

7. A server system, characterized in that, include: The substrate management controller (BMC), complex programmable logic device (CPLD), and non-volatile memory chip are connected in sequence via communication. The BMC is used to monitor the operating status of hardware functional modules in the server system in real time, and to send a hardware signal acquisition command to the CPLD when an abnormal hardware functional module is detected. The CPLD is used to, in response to the hardware signal acquisition command, acquire hardware signal data corresponding to the abnormal hardware function module in real time according to the hardware signal acquisition command, and transmit the hardware signal data to the non-volatile memory chip for storage based on the first communication link. The BMC is also used to, in response to the detection of abnormal recurrence of the abnormal hardware functional module, read hardware signal data corresponding to the abnormal hardware functional module from the non-volatile memory chip based on the second communication link, wherein the hardware signal data is used to locate the root cause of the hardware failure of the abnormal hardware functional module.

8. The server system according to claim 7, characterized in that, The CPLD includes a serial peripheral interface (SPI) controller module and a communication switch switching module. The hardware signal acquisition command carries communication link identification information, hardware signal type, and sampling frequency. The SPI controller module is used to collect hardware signal data corresponding to the abnormal hardware function module in real time according to the hardware signal type and the sampling frequency, and transmit the hardware signal data to the non-volatile memory chip for storage based on the first communication link. The communication switch switching module is used to switch the current communication link to the first communication link for connecting the SPI controller module and the non-volatile memory chip according to the communication link identification information.

9. The server system according to claim 8, characterized in that, The CPLD also includes a random access module, which is connected to the BMC via a third communication link; The random access module is used to receive the hardware signal acquisition command, send the communication link identification information to the communication switch switching module, and send the hardware signal type and the sampling frequency to the SPI controller module; The random access module is further configured to receive a communication link switching instruction sent by the BMC based on the third communication link, and send the communication link switching instruction to the communication switch switching module to enable the communication switch switching module to switch the current communication link to the second communication link for connecting the BMC and the non-volatile memory chip according to the communication link switching instruction.

10. A server, characterized in that, include: The processor, and the memory that is in communication with the processor; The memory stores instructions that the computer executes; The processor executes computer execution instructions stored in memory to implement the hardware fault location method as described in any one of claims 1 to 6.