Server Fault Chip Detection Method and Device

By setting up level detection and fault determination units in the server's substrate management controller, the problem of long-term failure positioning of servers is solved, and the effect of quickly positioning the fault chip and ensuring business reliability is achieved.

CN111949457BActive Publication Date: 2025-06-13CHINA GREATWALL TECH GRP CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010731524.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-27
Publication Date
2025-06-13
Estimated Expiration
2040-07-27

AI Technical Summary

Technical Problem

When a server fails, a lot of manpower and material resources are needed to determine the faulty chip, resulting in a long time to locate and recover.

Method used

By setting a level detection unit and a fault chip determination unit in the substrate management controller of the server, the interface level signal of the chip is detected, and whether there is a fault in the chip is determined based on these signals.

Benefits of technology

It realizes rapid positioning of faulty chips, saves manpower and material resources, shortens failure recovery time, and ensures the reliability of business services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111949457B_ABST
    Figure CN111949457B_ABST
Patent Text Reader

Abstract

This application is applicable to the field of server technology, and provides a method and device for detecting faulty chips of a server, which are applied to the baseboard management controller of the server. The baseboard management controller is connected to a preset number of chips through interfaces. The method includes: detecting the interface level signals of the connected chips; and determining whether there are faults in the corresponding chips according to the detected interface level signals. Thus, the specific faulty chips can be quickly located, enabling the server to quickly resume normal operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of servers, and particularly relates to a method and device for detecting faulty chips in a server. Background Art

[0002] With the rapid development of Internet applications, the computing volume and computing frequency of Internet applications have also increased. The increase in business computing volume has increased the operating pressure on the server, resulting in the core components of the server (such as processors, memory, etc.). In addition, as the server is used for a long time, some components of the server will also fail. As a type of machine that needs to run stably and reliably for a long time, once a failure occurs, it will have a great impact.

[0003] Currently, when a server fails, it often requires operation and maintenance personnel to analyze step by step according to the failure phenomenon, and it takes a lot of time to find out which specific component of the server has failed. For example, there are many functional chips on the motherboard of the server. Once a certain chip on the motherboard fails, it takes a lot of time to find the cause of the failure from top to bottom. For example, if the clock chip that provides the reference frequency for the CPU fails, from the perspective of the phenomenon, the server cannot start, and it is necessary to determine whether there is a problem with the operating system, whether there is a problem with the power supply, whether there is a problem with the hard disk controller, whether there is a problem with the power-on timing, whether there is a problem with the CPU itself, etc. Finally, it can be determined that the clock chip has a problem, which requires a lot of manpower and material resources. Summary of the Invention

[0004] In view of this, the embodiments of this application provide a method and device for detecting faulty chips in a server, so as to at least solve the problem that a large amount of manpower and material resources are required to determine the faulty chip when the server fails in the prior art.

[0005] The first aspect of the embodiments of this application provides a method for detecting faulty chips in a server, which is applied to the baseboard management controller of the server. The baseboard management controller is connected to a preset number of chips through interfaces. The method includes: detecting the interface level signals of the connected chips; and determining whether the corresponding chips are faulty according to the detected interface level signals.

[0006] The first aspect of the embodiments of this application provides a device for detecting faulty chips in a server, which is arranged in the baseboard management controller of the server. The baseboard management controller is connected to a preset number of chips through interfaces. The device includes: a level detection unit for detecting the interface level signals of the connected chips; and a faulty chip determination unit for determining whether the corresponding chips are faulty according to the detected interface level signals.

[0007] A third aspect of the embodiments of the present application provides a server, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.

[0008] A fourth aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0009] A fifth aspect of the embodiments of the present application provides a computer program product. When the computer program product runs on a server, the server is enabled to implement the steps of the above method.

[0010] The beneficial effects of the embodiments of the present application compared with the prior art are as follows:

[0011] The baseboard management controller in the server is connected to each chip through an interface. The baseboard management controller detects the interface level signal of the connected chip, and whether a corresponding chip has a fault can be determined through the interface level signal. Thus, when a certain chip on the motherboard fails, the specific faulty chip can be quickly located, without the need for operation and maintenance personnel to gradually analyze according to the fault phenomenon, which can save a large amount of manpower and material resources, enable the server to quickly resume normal operation, and ensure the reliability of business services. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0013] Figure 1 The flowchart of an example of the method for detecting a faulty chip of a server according to an embodiment of the present application is shown;

[0014] Figure 2 The schematic diagram of a response waveform when a chip connected to a BMC (Baseboard Management Controller) through an IIC (Inter-Integrated Circuit) interface has a fault is shown;

[0015] Figure 3 The schematic diagram of a signal waveform when a chip connected to a BMC through a fault indication interface has a fault is shown;

[0016] Figure 4A It shows a waveform schematic diagram of an example when the second chip connected to the BMC through the working signal interface is normal;

[0017] Figure 4B It shows a waveform schematic diagram of an example when the second chip connected to the BMC through the working signal interface fails;

[0018] Figure 5A It shows a waveform schematic diagram of an example when the third chip connected to the BMC through the working signal interface is normal;

[0019] Figure 5B It shows a waveform schematic diagram of an example when the third chip connected to the BMC through the working signal interface fails;

[0020] Figure 6 It shows a structural block diagram of an example of a server fault chip detection device according to an embodiment of the present application;

[0021] Figure 7 It is a schematic diagram of an example of a server according to an embodiment of the present application. Detailed implementation manners

[0022] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures, technologies, etc. are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0023] To illustrate the technical solutions described in the present application, the following will be described through specific embodiments.

[0024] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0025] It should also be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0026] It should be further understood that the term "and / or" as used in the specification and appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0027] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]" depending on the context.

[0028] In a specific implementation, the mobile terminal described in the embodiments of this application includes, but is not limited to, other portable devices such as mobile phones, laptop computers, or tablet computers having a touch-sensitive surface (e.g., a touch screen display and / or a touchpad). It should also be understood that in some embodiments, the above devices are not portable communication devices, but desktop computers having a touch-sensitive surface (e.g., a touch screen display and / or a touchpad).

[0029] In the following discussion, a mobile terminal including a display and a touch-sensitive surface is described. However, it should be understood that the mobile terminal may include one or more other physical user interface devices such as a physical keyboard, a mouse, and / or a joystick.

[0030] Various application programs that can be executed on the mobile terminal can use at least one common physical user interface device such as a touch-sensitive surface. One or more functions of the touch-sensitive surface and the corresponding information displayed on the terminal can be adjusted and / or changed between application programs and / or within the corresponding application programs. In this way, the common physical architecture of the terminal (e.g., the touch-sensitive surface) can support various application programs having a user interface that is intuitive and transparent to the user.

[0031] In addition, in the description of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0032] Figure 1 A flowchart showing an example of a server failure chip detection method according to an embodiment of this application is shown.

[0033] As Figure 1 shown, in step 110, the interface level signals of the connected chips are detected.

[0034] It should be noted that, in addition to the CPU (central processing unit) and the main bus controller, there are many chips on the server motherboard to implement various functions. For example, the server motherboard can be a server motherboard with servo chip management. BMC is usually used to monitor and manage the health status of the motherboard. For example, some important parameters on the motherboard such as voltage, temperature, power consumption, etc. can be monitored and recorded through BMC. In addition, the baseboard management controller is connected to the chip through an interface, and the functional types of the interfaces of the connected chips can be diversified. For example, they can be error interfaces for error signals or interrupt signals or general working signal interfaces, and there should be no restrictions here.

[0035] In step 120, according to the detected interface level signals, determine whether the corresponding chip has a fault.

[0036] In some examples of the embodiments of the present application, there are many IIC interfaces in BMC for management communication. Therefore, some chips (such as PCIE SWITCH / clock chip / SAS controller / power chip, etc.) need to communicate with BMC through IIC, as long as the IIC bus of the chip is interconnected with the BMC bus and there is no conflict in the device address. At this time, the fault detection process for the chip can be implemented by adding the corresponding function code configuration.

[0037] In some embodiments, for these chips with IIC interfaces that communicate with BMC, BMC can determine that all chips (or slave devices) without communication response are faulty. For example, assume that the IIC address of the clock chip Si52147-A01AGMR is 0X70, and there is no response when the BMC host accesses 0X70, then it can be considered that the Si52147-A01AGMR chip is faulty. Assume that the device address of a certain IIC slave device (or chip) is "1110A 2 A 1 A 0 ”, where A 2 A 1 A 0 is an address code that can be selected for the hardware. When the BMC (or master device) accesses it and the Si52147-A01AGMR chip (or slave device) cannot respond, it can be determined that the chip has a fault. Correspondingly, Figure 2 shows a schematic diagram of the response waveform in an example when a chip connected to BMC through the IIC interface has a fault. When BMC does not receive a response signal (or acknowledgment signal), it can be determined that the corresponding chip has a fault. Further, BMC can read the register value of the chip connected through the IIC interface, and when the register value is incorrect, it can be determined that the chip has a fault.

[0038] In addition, for a chip without an IIC interface, or when there is no master-slave device relationship between the chip and the BMC, a new interface connection relationship can be created between the BMC and the chip, so as to realize the identification process of the faulty chip.

[0039] In some examples of the embodiments of the present application, some chips (for example, the first chip) have a fault indication interface (for example, an error indicator light pin or an error interrupt pin) by themselves, then the fault indication interface can be connected to the interface of the BMC (for example, the GPIO (General-Purpose Input / Output Ports) of the BMC). Furthermore, when the BMC detects the first interface level signal from the fault indication interface, the BMC can determine that the first chip has a fault, that is, the BMC can determine whether the first chip has a fault by high and low levels (that is, the signal values 0 or 1).

[0040] For example, the chip Si5338N has an INTR (interrupt) function pin (that is, a fault indication interface), and the chip Si5338N and the BMC can be connected through an interface, and the BMC is configured, for example, the INTR function pin is valid at a low level. Figure 3 Fig. shows a waveform schematic diagram of an example when a chip connected to the BMC through a fault indication interface has a fault, where the waveform corresponding to signal 4 indicates that the fault indication interface is set from high to low during a short circuit.

[0041] In some examples of the embodiments of the present application, for chips (such as the second chip and the third chip) that have neither an IIC interface nor a fault indication interface, the working signal interfaces of the BMC and these chips can be connected, and the corresponding fault detection function can be realized through the configuration of the BMC.

[0042] In some embodiments, the BMC can realize the function of judging and identifying a faulty chip through a pin with a certain fixed function of a chip (for example, the second chip) (for example, the first working signal interface). Exemplarily, the second interface level signal from the first working signal interface can be detected within a set time period, and when the second interface level signal from the first working signal interface within the set time period meets the preset first faulty chip level condition, it is determined that the second chip has a fault. Here, the first faulty chip condition can be determined according to the level performance of the first working signal interface under normal working conditions, so as to identify the normal or faulty state of the chip.

[0043] Taking an example of the embodiments of the present application, for the cs (Chip Select) signal of SPI (Serial Peripheral Interface), the cs is set low to read the firmware information only after the chip initialization is completed, and the level will change between high and low continuously during the reading process. If the cs signal pin is always at a high level, it can be determined that this chip is faulty. Taking another example of the embodiments of the present application, the TCA9517DGKR chip is a simple level conversion chip. Under normal circumstances, when converting IIC signals, there should be a waveform change at a certain specific moment. If it is always at a high level, it can be determined that the TCA9517DGKR chip is faulty. Figure 4A shows a waveform schematic diagram of an example when the second chip connected to the BMC through the working signal interface is normal, and Figure 4B shows a waveform schematic diagram of an example when the second chip connected to the BMC through the working signal interface is faulty. As Figure 4A shown, when the W25Q128JVFIQ chip is working properly, the signals it outputs are intermittent high and low levels. As Figure 4B shown, when a chip fails and information cannot be read normally, the level waveform detected by the BMC will show a continuous high or low level, and at this time, it can be determined as a faulty chip.

[0044] In some embodiments, for a chip that has neither an IIC interface nor a fault indication interface (for example, the third chip), it can be achieved by connecting the BMC to a preset number (for example, multiple) of working signal interfaces (for example, GPIO pins) of the chip, and detecting whether the third interface level signals corresponding to the preset number of second working signal interfaces meet the preset second faulty chip level conditions, and when meeting the second faulty chip level conditions, it can be determined that the third chip is faulty.

[0045] For example, for the 88SE9230 chip, it has 8 GPIO pins that can be configured for related functions. Multiple of these pins can be selected and connected to the GPIO of the BMC, and then the BMC can identify whether the chip is faulty.

[0046] Figure 5A shows a waveform schematic diagram of an example when the third chip connected to the BMC through the working signal interface is normal, and Figure 5B shows a waveform schematic diagram of an example when the third chip connected to the BMC through the working signal interface is faulty. As Figure 5A and 5BAs shown, the GPIO0 and GPIO1 interfaces of the 88SE9230 chip are respectively connected to the BMC through interfaces. When the chip is normal, both the GPIO0 and GPIO1 interfaces output high level. When the BMC detects that the GPIO0 or GPIO1 outputs low level, the BMC can determine that the 88SE9230 chip has a fault, resulting in the inability to complete normal functions.

[0047] In some examples of the embodiments of the present application, when a faulty chip is detected, a fault prompt operation corresponding to the faulty chip can be executed. Here, corresponding fault prompt operation configurations are respectively set for a preset number of chips. Thus, by performing fault operations through personalized prompt operation configurations for different chips, users or operation and maintenance personnel can intuitively and quickly know the faulty chips.

[0048] In some embodiments, the number corresponding to the faulty chip can be displayed in the form of a list. For example, it is pre-agreed that the number for the clock chip Si52147 - A01AGMR is U23, and the corresponding number can be displayed when a fault occurs. In addition, the fault prompt content text corresponding to the number can also be directly displayed, such as "CPU reference clock chip fault".

[0049] Through the embodiments of the present application, it is realized to use the BMC hardware to detect the states of as many important chips as possible on the server motherboard, and all information can be collected and sorted out to achieve the goal of real-time detection and rapid location of faulty chips.

[0050] Figure 6 The structural block diagram of an example of the server faulty chip detection device according to the embodiments of the present application is shown. Here, the server faulty chip detection device 600 is disposed in the baseboard management controller (not shown) of the server, and the baseboard management controller is connected to a preset number of chips through interfaces.

[0051] As Figure 6 shown, the server faulty chip detection device 600 includes a level detection unit 610 and a faulty chip determination unit 620.

[0052] The level detection unit 610 is used to detect the interface level signals of each connected chip;

[0053] The faulty chip determination unit 620 is used to determine whether the corresponding chip has a fault according to the detected interface level signals of each.

[0054] In some embodiments, the baseboard management controller is connected to the fault indication interface of the first chip, and the faulty chip determination unit 620 includes a first faulty chip determination module (not shown), which is configured to determine that the first chip has a fault when there is a first interface level signal from the fault indication interface.

[0055] In some embodiments, the baseboard management controller is connected to the first working signal interface of the second chip. The faulty chip determination unit 620 includes a second faulty chip determination module (not shown), which is configured to determine that the second chip is faulty when the second interface level signal from the first working signal interface meets a preset first faulty chip level condition within a set time period.

[0056] It should be noted that for the information interaction, execution process, etc. between the above-mentioned devices / units, since they are based on the same concept as the method embodiments of the present application, their specific functions and the technical effects brought about can be specifically referred to in the method embodiment part, and will not be elaborated here.

[0057] Figure 7 is a schematic diagram of an example of the server in the embodiments of the present application. As Figure 7 shown, the server 700 of this embodiment includes: a processor 710, a memory 720, and a computer program 730 stored in the memory 720 and executable on the processor 710. When the processor 710 executes the computer program 730, it implements the steps in the method embodiment of the above-mentioned server faulty chip detection method, such as Figure 1 the steps 110 to 120 shown. Alternatively, when the processor 710 executes the computer program 730, it implements the functions of each module / unit in the above-mentioned device embodiments, such as Figure 6 the functions of the units 610 to 620 shown.

[0058] Exemplarily, the computer program 730 can be divided into one or more modules / units. The one or more modules / units are stored in the memory 720 and executed by the processor 710 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 730 in the server 700. For example, the computer program 730 can be divided into a level detection module and a faulty chip determination module, and the specific functions of each module are as follows:

[0059] The level detection module is used to detect the interface level signals of the connected chips.

[0060] The faulty chip determination module is used to determine whether the corresponding chip is faulty according to the detected interface level signals of each chip.

[0061] The server 700 can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The server may include, but is not limited to, a processor 710 and a memory 720. Those skilled in the art can understand,Figure 7 The server 700 is only an example and does not limit the server 700. It may include more or fewer components than those shown in the figure, or combine certain components, or have different components. For example, the server may also include input / output devices, network access devices, buses, etc.

[0062] The so-called processor 710 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0063] The memory 720 may be an internal storage unit of the server 700, such as the hard disk or memory of the server 700. The memory 720 may also be an external storage device of the server 700, such as a plug-in hard disk equipped on the server 700, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 720 may also include both the internal storage unit and the external storage device of the server 700. The memory 720 is used to store the computer program and other programs and data required by the server. The memory 720 may also be used to temporarily store the data that has been output or is to be output.

[0064] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment.

[0065] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0066] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0067] In the embodiments provided in this application, it should be understood that the disclosed device / server and method can be implemented in other ways. For example, the device / server embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0068] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0069] In addition, the functional units in the various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above units may be implemented in the form of hardware or in the form of software.

[0070] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, all or part of the processes in the above-described embodiment methods of the present application may also be completed by instructing relevant hardware through a computer program. The computer program may be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments may be implemented. Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0071] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for detecting faulty chips in a server, characterized in that, it is applied to the baseboard management controller of the server, and the baseboard management controller is connected to a preset number of chips through interfaces. The method includes: detecting the interface level signals of each connected chip; determining whether a corresponding chip has a fault according to the detected interface level signals of each; wherein, the chips are a first type of chip with a fault indication interface for itself, a second type of chip without a fault indication interface but capable of determining whether there is a fault through a certain fixed function pin, and a third type of chip without a fault indication interface but capable of determining whether there is a fault through multiple working signal interfaces. The corresponding interfaces include the fault indication interface of the first type of chip, the first working signal interface corresponding to the fixed function pin of the second type of chip, and the multiple preset working signal interfaces of the third type of chip. The first type of chip is a chip without an IIC interface but with a fault indication interface, and the second and third types of chips are chips without an IIC interface and a fault indication interface; wherein, the baseboard management controller is connected to the fault indication interface of the first type of chip. The determining whether a corresponding chip has a fault according to the detected interface level signals of each includes: when there is a first interface level signal from the fault indication interface, determining that the first type of chip has a fault; the baseboard management controller is connected to the first working signal interface of the second type of chip. The determining whether a corresponding chip has a fault according to the detected interface level signals of each includes: when the second interface level signal from the first working signal interface meets the preset first faulty chip level condition within a set time period, determining that the second type of chip has a fault; the baseboard management controller is connected to a preset number of second working signal interfaces of the third type of chip. The determining whether a corresponding chip has a fault according to the detected interface level signals of each includes: when the third interface level signals from each of the second working signal interfaces meet the preset second faulty chip level condition, determining that the third type of chip has a fault.

2. The method for detecting faulty chips in a server according to claim 1, characterized in that, after determining whether a corresponding chip has a fault according to the detected interface level signals of each, the method further includes: when there is a faulty chip, performing a fault prompt operation corresponding to the faulty chip, and corresponding fault prompt operation configurations are respectively set for the preset number of chips.

3. A device for detecting faulty chips in a server, characterized in that, it is arranged in the baseboard management controller of the server, and the baseboard management controller is connected to a preset number of chips through interfaces. The device includes: a level detection unit for detecting the interface level signals of each connected chip; a faulty chip determination unit for determining whether a corresponding chip has a fault according to the detected interface level signals of each; Among them, each of the chips is a first chip with a built-in fault indication interface, a second chip without a fault indication interface but capable of determining whether there is a fault through a certain fixed function pin, and a third chip without a fault indication interface but capable of determining whether there is a fault through multiple working signal interfaces. The corresponding interfaces include the fault indication interface of the first chip, the first working signal interface corresponding to the fixed function pin of the second chip, and the multiple preset working signal interfaces of the third chip. The first chip is a chip without an IIC interface but with a fault indication interface, and the second and third chips are chips without an IIC interface and a fault indication interface; Among them, the baseboard management controller is connected to the fault indication interface of the first chip, and the faulty chip determination unit includes: A first faulty chip determination module configured to determine that the first chip is faulty when there is a first interface level signal from the fault indication interface; The baseboard management controller is connected to the first working signal interface of the second chip, and the faulty chip determination unit includes: A second faulty chip determination module configured to determine that the second chip is faulty when the second interface level signal from the first working signal interface meets a preset first faulty chip level condition within a set time period; The baseboard management controller is connected to a preset number of second working signal interfaces of the third chip, and the faulty chip determination unit further includes: A third faulty chip determination module configured to determine that the third chip is faulty when the third interface level signals from each of the second working signal interfaces meet a preset second faulty chip level condition.

4. A server, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 2 are implemented.

5. A computer-readable storage medium storing a computer program, wherein, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 2 are implemented.

Citation Information

Patent Citations

  • Fault detection method and device

    CN104536855A

  • Method and device for detecting server failure

    CN106919490A

  • Server monitoring device, method and system thereof

    CN109508279A

  • Power supply fault type positioning method and device, equipment and medium

    CN110399029A