Fault location device and server system

By using the fault location device in the in-position indicator module and the control storage module, and by using LEDs to indicate faults in the graphics processor, the problem of difficult disassembly and location during power outages in existing technologies is solved. This enables rapid fault location and replacement even when the server power is on, reducing maintenance costs and the risk of equipment damage.

CN121579316BActive Publication Date: 2026-05-01INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610098403.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-23
Publication Date
2026-05-01
Estimated Expiration
2046-01-23

AI Technical Summary

Technical Problem

When multiple graphics processors fail, existing technologies require powering off and disassembling the server chassis, making it difficult to accurately locate and replace the faulty graphics processor, resulting in high repair costs and an increased risk of equipment damage.

Method used

A fault location device is provided, including a serial port connector, an on-site indicator module, and a control storage module. By interacting with the motherboard and storing fault information while the server power is on, the device uses light-emitting diodes to indicate the faulty graphics processor, enabling rapid location and replacement.

Benefits of technology

After a server power failure, the system can accurately locate and alert the user to the faulty graphics processor, reducing maintenance costs and the risk of incorrect equipment replacement, and improving the efficiency of fault location and replacement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121579316B_ABST
    Figure CN121579316B_ABST
Patent Text Reader

Abstract

The application provides a fault positioning device and a server system, which can be applied to the technical field of fault positioning. The fault positioning device comprises a first serial port connector; an in-place indication module configured to send an in-place signal to a first controller of a mainboard when the first serial port connector is electrically connected with a second serial port connector; and a control storage module configured to acquire and store fault information of at least one graphics processor according to the power supply of a server power supply and the level of the in-place signal. When the server power supply is switched to power off and the mainboard and the graphics processing card electrically connected with the control storage module are both powered by the power supply, the target fault information is determined from the fault information based on the fault identification of the mainboard and is sent to the first controller, so that the first controller determines at least one faulty graphics processor on the graphics processing card, and a first light-emitting diode corresponding to the at least one faulty graphics processor is lit or extinguished at a first frequency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of fault location technology, and more specifically to a fault location device and server system. Background Technology

[0002] To meet the current high computing power demands of servers, existing servers typically integrate multiple network interface cards (NICs), multiple central processing units (CPUs), multiple memory modules, and multiple graphics processing units (GPUs). However, with the increasing number of servers and their internal GPUs, when multiple GPUs fail, the server chassis needs to be powered off and disassembled first. This makes it difficult to accurately locate and replace the faulty GPU when the power is off. Summary of the Invention

[0003] In view of the above problems, this application provides a fault location device and a server system.

[0004] According to a first aspect of this application, a fault location device is provided, comprising: a first serial port connector; an on-premises indication module electrically connected to the first serial port connector, configured to send an on-premises signal to a first controller of the motherboard via the first serial port connector when the first serial port connector is electrically connected to a second serial port connector of the motherboard; and a control storage module electrically connected to a power supply within the fault location device, configured to acquire and store fault information of at least one graphics processor based on the power supply of the server power supply and the level of the on-premises signal, and, when the server power supply is switched off and both the motherboard and the graphics processing card electrically connected to the control storage module are powered by the power supply, determine target fault information from the fault information based on the fault identifier of the motherboard and send it to the first controller, so that the first controller determines at least one faulty graphics processor on the graphics processing card from the target fault information and controls a first light-emitting diode corresponding to the at least one faulty graphics processor to light up or turn off at a first frequency.

[0005] A second aspect of this application provides a server system comprising: a fault location device as described above; a motherboard electrically connected to the fault location device and a graphics processing card, including a baseboard management controller, a first controller, a second serial port connector, and a third serial port connector, wherein the baseboard management controller is configured to send at least one fault information of a predetermined format to the fault location device via the first controller and the second serial port connector, the first controller is configured to determine at least one faulty graphics processor from the target fault information, and send a control signal to the graphics processing card via the third serial port connector for controlling a first light-emitting diode corresponding to the at least one faulty graphics processor; and a graphics processing card including an input / output expander, at least one first light-emitting diode, at least one graphics processor, and a fourth serial port connector, wherein the input / output expander is electrically connected to the at least one first light-emitting diode and the at least one graphics processor, and the input / output expander is configured to, in response to the control signal transmitted via the fourth serial port connector, turn on or off the first light-emitting diode corresponding to the at least one faulty graphics processor, monitor the presence signal of the at least one graphics processor, and send it to the first controller.

[0006] According to an embodiment of this application, before the faulty graphics processor in the server needs to be replaced and the server power is not cut off, the fault location device is electrically connected to the motherboard so that the presence indication module and control storage module in the fault location device can interact and store information with the first controller in the motherboard, so as to provide retrievable information records for lighting the first light-emitting diode corresponding to the faulty graphics processor after the server power is cut off.

[0007] When the server power is off and the presence signal is valid, the power supply in the fault location device provides temporary, small-scale power to some components in the motherboard and graphics card, thereby supporting the flashing of the first LED corresponding to the faulty graphics processor.

[0008] Based on the current motherboard fault indicators, the system identifies whether a faulty graphics processor (GPU) exists within the server and retrieves relevant target fault information. This allows the system to control the first LED corresponding to the faulty GPU to blink. Consequently, even in the event of a server power outage, especially when multiple GPUs are faulty, the system can quickly locate and identify multiple faulty GPUs, enabling efficient replacement and reducing maintenance costs and equipment malfunctions and damage caused by incorrect replacements. Attached Figure Description

[0009] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0010] Figure 1A schematic diagram of a fault location device according to an embodiment of this application is shown;

[0011] Figure 2 A schematic diagram of a control storage module according to an embodiment of this application is shown;

[0012] Figure 3 A schematic diagram of an in-situ indication module according to an embodiment of this application is shown;

[0013] Figure 4 A schematic diagram of a fault location device according to another embodiment of this application is shown;

[0014] Figure 5 A schematic diagram of a fault location device according to yet another embodiment of this application is shown;

[0015] Figure 6 A flowchart of a fault location method according to an embodiment of this application is shown;

[0016] Figure 7 A schematic diagram of a server system according to an embodiment of this application is shown. Detailed Implementation

[0017] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.

[0018] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0019] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0020] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0021] To meet the current high computing power demands of servers, existing servers typically integrate multiple network interface cards (NICs), multiple central processing units (CPUs), multiple memory modules, and multiple graphics processing units (GPUs). However, with the increasing number of servers and their internal GPUs, when multiple GPUs fail, the server chassis needs to be powered off and disassembled first. This makes it difficult to accurately locate and replace the faulty GPU when the power is off.

[0022] This application provides a fault location device, a first serial port connector, an on-premises indicator module electrically connected to the first serial port connector, used to send an on-premises signal to a first controller on the motherboard via the first serial port connector when the first serial port connector is electrically connected to a second serial port connector on the motherboard, and a control storage module electrically connected to a power supply in the fault location device, used to acquire and store fault information of at least one graphics processor based on the power supply of the server power supply and the level of the on-premises signal, and when the server power supply is switched off and both the motherboard and the graphics processing card electrically connected to the control storage module are powered by the power supply, the target fault information is determined from the fault information based on the fault identifier of the motherboard and sent to the first controller, so that the first controller determines at least one faulty graphics processor on the graphics processing card from the target fault information and controls a first light-emitting diode corresponding to the at least one faulty graphics processor to light up or turn off at a first frequency.

[0023] Figure 1 A schematic diagram of a fault location device according to an embodiment of this application is shown.

[0024] like Figure 1 As shown, in the event of a failure of any one or more graphics processors (GPUs) 102 on the graphics processing card 101 (GPU card) within the server, a fault location device 103 can be inserted into the motherboard 104 within the server. This allows for an electrical connection between the motherboard 104 and the graphics processing card 101, enabling the location indication of the faulty graphics processor on the currently de-energized and disassembled graphics processing card 101. Specifically, the fault location device 103 may include a first serial port connector 105, a presence indication module 106, and a control storage module 107.

[0025] The first serial port connector 105 can be connected to the second serial port connector 108 on the motherboard 104, thereby electrically connecting the fault location device 103 to the motherboard 104. Both the first serial port connector 105 and the second serial port connector 108 may have corresponding power supply pin 1, clock pin 2, data pin 3, presence pin 4, and ground pin 5. The first serial port connector 105 can be a Micro USB (Micro Universal Serial Bus) or Mini USB (Mini Universal Serial Bus) interface, etc.

[0026] The presence indication module 106 can be electrically connected to the first serial port connector 105. It can be used to send a presence signal to the first controller 109 of the motherboard 104 through the first serial port connector 105 when the first serial port connector 105 is electrically connected to the second serial port connector 108 of the motherboard 104.

[0027] When the first serial port connector 105 and the second serial port connector 108 are electrically connected and the power supply U_3V3 provides normal power to the fault location device 103, the presence indicator module 106 can transmit a presence signal to the first controller 109, which is electrically connected to the presence pin of the second serial port connector 108, through the presence pin of the first serial port connector 105. This allows the first controller 109 in the motherboard 104 to obtain the current status of the external fault location device 103. Furthermore, it can control the level of the presence signal issued by the presence indicator module 106 according to the fault location process, thereby providing relevant progress prompts to the operator through the relevant indicator devices within the presence indicator module 106.

[0028] The control storage module 107 can be electrically connected to the power supply U_3V3 in the fault location device 103. It can be used to acquire and store fault information of at least one graphics processor 102 based on the power supply and presence signal level of the server power supply P3V3_STBY of the motherboard 104. When the server power supply P3V3_STBY is switched off and both the motherboard 104 and the graphics processing card 101, which are electrically connected to the control storage module 107, are powered by the power supply U_3V3, the control storage module 107 can determine the target fault information from the fault information based on the fault identifier of the motherboard 104 and send it to the first controller 109. This allows the first controller 109 to determine at least one faulty graphics processor on the graphics processing card 101 from the target fault information and control the first light-emitting diodes (LED1, LED2, LED3, LED4) corresponding to the at least one faulty graphics processor to light up or turn off at a first frequency.

[0029] The fault location device 103 may be equipped with a rechargeable 3.3V power supply U_3V3. Using the 3.3V power supply U_3V3, when the server power supply P3V3_STBY is normally powered and the fault location device 103 has received and stored at least one fault information, power is supplied only to the control storage module 107 and the presence indication module 106 within the fault location device 103. Simultaneously, the power supply U_3V3 within the fault location device 103 can also provide temporary power to some components on the motherboard 104 and the graphics processing card 101 when the power is cut off (server power supply P3V3_STBY is powered off) and the faulty graphics processor is disassembled and replaced. This allows for the illumination or extinguishing of the first light-emitting diode on the graphics processing card 101 to indicate the faulty graphics processor.

[0030] The fault information may include the identifier of the faulty graphics processor, the server serial number of the server on which the faulty graphics processor is located, the identifier of the motherboard 104 of the server on which the faulty graphics processor is located, the fault type of the faulty graphics processor, the fault model of the faulty graphics processor, etc., but is not limited to the above information.

[0031] Based on the power supply status of the server power supply P3V3_STBY and the level of the presence signal, the overall progress of fault location can be determined. When the server power supply P3V3_STBY can normally supply power to the motherboard 104 and graphics processing card 101, and the fault location device 103 is inserted into the motherboard 104 and sends a valid (high-level) presence signal to the first controller 109 on the motherboard 104, it can be confirmed that the current process is in the initial preparation stage of fault location. The control storage module 107 receives and stores the fault information corresponding to the faulty graphics processor sent by the motherboard 104, thereby storing the corresponding fault information in the fault location device 103. This allows the first LED on the graphics processing card 101 to light up or turn off during power-off replacement to locate and indicate the faulty graphics processor.

[0032] When the server power supply P3V3_STBY is powered off, and the fault location device 103, motherboard 104, and graphics processing card 101 are all temporarily powered by the power supply U_3V3 within the fault location device 103, and the fault location device 103 is inserted into the motherboard 104 and sends a valid (high-level) presence signal to the first controller 109 on the motherboard 104, it can be confirmed that the current fault location replacement process is underway. The control storage module 107 can determine the target fault information corresponding to the current motherboard 104 and graphics processing card 101 from at least one fault information stored internally, based on the fault identifier or server serial number and other identification information of the currently connected motherboard 104, and send it to the first controller 109 in the motherboard 104.

[0033] The first controller 109 can analyze the received target fault information, determine the faulty graphics processor in at least one graphics processor 102 on the current graphics processing card 101 and the first light-emitting diode corresponding to the faulty graphics processor, and then generate a corresponding control signal and send it to the input / output expander 110 on the graphics processing card 101, so that the input / output expander 110 controls the first light-emitting diode corresponding to at least one faulty graphics processor to light up or turn off at a first frequency according to the control signal, thereby accurately prompting relevant personnel that the graphics processor 102 at this location is a faulty graphics processor 102 and needs to be replaced.

[0034] According to an embodiment of this application, before the faulty graphics processor in the server needs to be replaced and the server power is not cut off, the fault location device is electrically connected to the motherboard so that the presence indication module and control storage module in the fault location device can interact and store information with the first controller in the motherboard, so as to provide retrievable information records for lighting the first light-emitting diode corresponding to the faulty graphics processor after the server power is cut off.

[0035] When the server power is off and the presence signal is valid, the power supply in the fault location device provides temporary, small-scale power to some components in the motherboard and graphics card, thereby supporting the flashing of the first LED corresponding to the faulty graphics processor.

[0036] Based on the current motherboard fault indicators, the system identifies whether a faulty graphics processor (GPU) exists within the server and retrieves relevant target fault information. This allows the system to control the first LED corresponding to the faulty GPU to blink. Consequently, even in the event of a server power outage, especially when multiple GPUs are faulty, the system can quickly locate and identify multiple faulty GPUs, enabling efficient replacement and reducing maintenance costs and equipment malfunctions and damage caused by incorrect replacements.

[0037] The aforementioned fault location device can also be applied to application scenarios such as PCIe external cards and memory that require disassembling and reassembling the server chassis while the server power is off for repair.

[0038] Figure 2 A schematic diagram of a control storage module according to an embodiment of this application is shown.

[0039] like Figure 2 As shown, the control storage module inside the fault location device 103 may include a memory 201 and a second controller 202.

[0040] Specifically, the memory 201 can be electrically connected to the first serial port connector 105. The memory 201 can be used to acquire and store fault information of at least one graphics processor 102 sent by the baseboard management controller 204 in the motherboard 104 when the server power supply P3V3_STBY is powered on and the bit signal is at an active level.

[0041] When the server power supply P3V3_STBY is powered on and sends a valid presence signal to the first controller 109 on the motherboard 104, the baseboard management controller 204 on the motherboard 104 can send at least one fault message to the first controller 109 in I2C format. In response to the presence of the fault location device 103, the first controller 109 sends a synchronized clock signal to the memory 201 and the second controller 202 via its clock pin, and switches the transmission link between its internal UART and I2C to write at least one fault message in I2C format to the memory 201 via its data pin, so that the memory 201 can store the received information. The memory 201 can be a hardware device with storage capabilities, such as an EEPROM. When the fault location device 103 is not electrically connected to the motherboard 104 or is not sending a presence signal to the first controller 109, the clock and data pins of the second serial port connector 108 can be in UART format to facilitate normal serial log transmission and reception with other devices.

[0042] The second controller 202 can be electrically connected to the memory 201. The second controller 202 can be used to determine target fault information from the fault information of at least one graphics processor 102 based on the fault identifier of the motherboard 104 from the first controller 109 when the server power supply P3V3_STBY is switched off and both the motherboard 104 and the graphics processing card 101 are powered by the power supply U_3V3; and when receiving an update signal sent by the first controller 109, to update the fault information stored in the memory 201 according to the update information corresponding to at least one faulty graphics processor in the update signal.

[0043] The second controller 202 can read information from the memory 201 or perform control operations such as monitoring, updating, and rewriting the information stored in the memory 201.

[0044] The update information may include, but is not limited to, the presence signal of each graphics processor 102, the processor identifier of each graphics processor 102, the processor identifier that occurred during the bit signal level change, and the fault identifier of the motherboard 104.

[0045] When the server power supply P3V3_STBY is powered off, a valid presence signal is detected by the presence indicator module 106, and the first controller 109 on the motherboard 104 is detected to be powered on normally, while the power-on signals of other hardware devices on the motherboard 104 are all invalid (low level). This confirms that the process of replacing the faulty graphics processor is underway. The second controller 202, based on the fault identifier or server serial number and other identification information of the current motherboard 104 sent by the first controller 109, reads and determines the target fault information from multiple fault information stored in the memory 201, and sends it to the first controller 109 on the motherboard 104 via the electrically connected first serial port connector 105 and second serial port connector 108.

[0046] After the first controller 109 controls the corresponding first light-emitting diode to flash at a first frequency according to the target fault information, the relevant faulty graphics processor is replaced. At the same time, the input / output expander 110 on the graphics processing card 101 can monitor the presence signals of multiple graphics processors 102. After the faulty graphics processor is replaced, the first controller 109 can generate a corresponding update signal through the presence signal monitored in real time by the input / output expander 110, and send the update signal to the second controller 202 in the fault location device 103 to notify the fault location device 103 of the current maintenance and replacement progress.

[0047] After receiving the update signal, the second controller 202 can combine the presence signal of each graphics processor 102 to extract the corresponding update information from the update signal, and perform processing such as erasing and marking on the corresponding fault information in the memory 201, thereby updating the fault information stored in the memory 201 and avoiding problems such as repeated replacement and maintenance.

[0048] According to embodiments of this application, the control storage module may specifically include a memory and a second controller. In response to the power supply status of the server and the level of the presence signal, the memory can write and store fault information of at least one graphics processor, so that even when the server power is off, the fault information can be retrieved by the first and second controllers. When the server power is completely off for unpacking and replacement maintenance, the second controller can quickly determine the target fault information from multiple fault messages and send it to the first controller via temporary power supply, so that the faulty graphics processor can be located and indicated by flashing of a first LED. Simultaneously, after the faulty graphics processor is replaced, the fault information stored in the memory can be updated, avoiding repeated replacement and maintenance, and improving the accuracy and efficiency of maintenance and replacement.

[0049] Please continue to refer to the above. Figure 2As shown, the control storage module 107 inside the fault location device may also include a multiplexer 203.

[0050] Specifically, the multiplexer 203 can be electrically connected to the memory 201, the second controller 202, and the first serial port connector 105. The multiplexer 203 can be used to establish a data transmission path between the first controller 109 and the memory 201 when the server power supply P3V3_STBY is powered on and the on-state signal is at an active level, so that at least one fault message is transmitted to the memory 201; and when the server power supply P3V3_STBY is powered off and both the motherboard 104 and the graphics processing card 101 are powered by the power supply U_3V3, it can establish a data transmission path between the first controller 109 and the second controller 202, so that the second controller 202 sends the target fault message and receives an update signal.

[0051] Multiplexer 203 can be configured on the data transmission channel between the memory 201, the second controller 202, and the data pins of the first serial port connector. In response to different fault repair and replacement processes, multiplexer 203 switches different data transmission links to enable data information exchange between the memory 201, the second controller 202, and the first controller 109.

[0052] According to an embodiment of this application, a multiplexer is provided on the data transmission channel between the memory, the second controller, and the data pins of the first serial port connector to facilitate switching between different data transmission links according to different demand processes, thereby improving the orderliness of data information control.

[0053] According to an embodiment of this application, the second controller 202 can also be used to: when the server power is switched to power-off and both the motherboard and the graphics processing card are powered by the power supply, read the fault information to be repaired from at least one fault information based on the fault identifier of the motherboard.

[0054] According to an embodiment of this application, when the fault information to be repaired includes fault information of a faulty graphics processor, the fault information to be repaired is identified as the target fault information and sent to the first controller through the first serial port connector.

[0055] According to an embodiment of this application, when the fault information to be repaired includes fault information of multiple faulty graphics processors, the fault type and fault model of the multiple faulty graphics processors are determined from the fault information to be repaired; based on the multiple fault types and multiple fault models, a fault prompting strategy is determined and target fault information is generated, so that the first controller controls the first light-emitting diodes corresponding to the multiple faulty graphics processors to light up or turn off at a first frequency according to the target fault information.

[0056] The memory 201 can store multiple fault information corresponding to multiple servers. Each fault information can contain a different number of faulty graphics processors and their corresponding fault information.

[0057] If the fault information to be repaired only includes the fault information of a faulty graphics processor, the second controller 202 can directly determine the fault information to be repaired as the target fault information, so that the first controller 109 can analyze the target fault information and control the corresponding first light-emitting diode to flash, etc.

[0058] When the fault information to be repaired includes fault information for multiple faulty graphics processors, the fault information to be repaired can be directly identified as the target fault information, allowing the first controller to parse the target fault information and control the corresponding first LED to blink. The fault type, fault model, and processor identifier of multiple faulty graphics processors can also be extracted from the fault information to be repaired. First, first control information is generated based on the multiple processor identifiers. Then, the lighting order of the multiple first LEDs is determined by multiplying the repair priority parameters corresponding to the fault type and the basic parameter of the number of faulty graphics processors pre-stored in memory. Next, based on the fault model of the same faulty graphics processor, the determined lighting order of the multiple first LEDs is reordered to obtain second control information. Based on the first and second control information, a fault indication strategy is determined to accurately locate the faulty graphics processor while allowing for orderly replacement of multiple faulty graphics processors in a zoned manner, improving the efficiency of repair and replacement. The first control information can be a predetermined period for controlling the first LEDs corresponding to all faulty graphics processors to blink simultaneously within a first time period. The second control information can determine the initial lighting sequence from largest to smallest based on the product of the maintenance priority parameter corresponding to the fault type and the basic parameter of the number of faulty graphics processors in the second time period. Then, from the faulty graphics processors lit in the same batch, they can be further lit in different time periods according to their model until all faulty graphics processors are lit and replaced.

[0059] The first controller 109 parses the target control information generated based on the fault indication strategy. According to the control sequence of the first light-emitting diodes in the target control information, the corresponding first light-emitting diodes can be turned on or off sequentially through the input / output expander in the graphics processing card.

[0060] According to an embodiment of this application, the second controller can determine whether a corresponding fault prompt strategy needs to be generated based on the number of faulty image processors on the current graphics processing card, thereby controlling at least one first light-emitting diode to light up or turn off (blink) according to the strategy in an orderly manner, thereby improving the fault location efficiency while ensuring the accuracy of the prompt.

[0061] According to an embodiment of this application, the second controller 202 can also be used to: upon receiving an update signal generated by the first controller based on the processor fault identifier in the target fault information, confirm the level state of the multiple in-situ signals based on the processor fault identifier in the target fault information, in the case of receiving an update signal generated by the first controller based on the in-situ signals and graphics processor identifiers of the multiple replaced graphics processors.

[0062] According to an embodiment of this application, when multiple in-situ signals are all at active levels, the fault status of at least one faulty graphics processor in the fault information is updated.

[0063] According to an embodiment of this application, when an invalid level exists among multiple in-situ signals, an alarm message is generated and sent to a first controller based on the graphics processor identifier corresponding to the in-situ signal with the invalid level, so that the first controller controls at least one first light-emitting diode to light up or turn off at a second frequency based on the alarm message.

[0064] The first controller 109 can monitor the status of the presence signals of multiple graphics processors in real time via the input / output expander 110 on the graphics processing card 101. After replacing at least one faulty graphics processor, the input / output expander 110 can send the presence signals of the current multiple graphics processors (both unreplaced and replaced) to the first controller 109. The first controller 109 can generate an update signal based on the multiple presence signals and the graphics processor identifiers of the multiple graphics processors, and send it to the second controller 202 in the fault location device.

[0065] After receiving the update signal, the second controller 202 determines the level state of the in-situ signal corresponding to the processor fault identifier from among the multiple in-situ signals in the update signal, based on the processor fault identifier in the fault information.

[0066] When the level of the presence signal corresponding to the processor fault identifier is active (high level), the second controller 202 can confirm that the replacement of the faulty graphics processor is correct, and then determine the presence signals of the graphics processors other than the faulty graphics processor among the multiple faulty graphics processors. When the levels of the presence signals of the graphics processors other than the faulty graphics processor are also active, the second controller 202 can update the fault status in the fault information stored in the memory.

[0067] When an invalid level (low level) is present in the presence signal corresponding to the processor fault identifier, the second controller 202 can confirm that the replaced faulty graphics processor is still faulty. Based on the processor identifier of the replaced faulty graphics processor, it generates an alarm message and sends it to the first controller 109. The first controller 109 then parses the alarm message and controls the corresponding first LED to flash. If the presence signal remains invalid after replacing the faulty graphics processor more than a predetermined number of times, the second controller 202 generates a pause maintenance message and sends it to the first controller 109. The first controller 109 then pulls down the presence pin of the first serial port connector based on the pause maintenance message, thereby controlling the second LED in the presence indicator module 106 to flash at an alarm frequency, facilitating system maintenance of the graphics processing card and server.

[0068] According to an embodiment of this application, upon receiving an update signal sent by the first controller, the system determines whether the replacement and repair of the faulty graphics processor is complete based on the level states of multiple in-situ signals. If the replacement and repair are confirmed to be complete, the fault information in the memory is updated to confirm that the replacement is complete. If the replacement and repair are not confirmed to be complete, the first light-emitting diode is lit up to provide multiple reminders, thereby assisting the operator in judging the current replacement status of the faulty graphics processor, improving the replacement completion rate, reducing repair costs, and mitigating equipment failures and damages caused by incorrect replacements.

[0069] Figure 3 A schematic diagram of an in-situ indication module according to an embodiment of this application is shown.

[0070] like Figure 3 As shown, the in-situ indication module of the fault location device may include an indication unit 301 and a first resistor R1.

[0071] Specifically, the indicator unit 301 can be electrically connected to the in-situ pin of the first serial port connector 105. The indicator unit 301 can be used to forward bias the second light-emitting diode LED5 in the indicator unit 301 and light it up at a third frequency when the in-situ pin of the first serial port connector 105 is pulled down to a low level.

[0072] The first end of the indicator unit 301 can be electrically connected to the power supply U_3V3 in the fault location device 103, and the second end of the indicator unit 301 can be electrically connected to the in-position pin of the first serial port connector 105. The first controller 109 pulls down the in-position pin of the second serial port connector 108, thereby pulling down the in-position pin of the first serial port connector 105 to a low level, thereby changing the voltage difference across the indicator unit 301, so that the second light-emitting diode LED5 in the indicator unit 301 can be forward biased and lit at a third frequency.

[0073] The first resistor R1 can be connected in parallel with the indicator unit 301 to pull up the presence pin of the first serial port connector 105 to a high level, thereby obtaining a valid presence signal that is sent to the first controller through the first serial port connector 105.

[0074] The first end of the first resistor R1, which is connected in parallel with the indicator unit 301, can be electrically connected to the power supply U_3V3, and the second end of the first resistor R1 can be electrically connected to the in-position pin of the first serial port connector 105. The first resistor R1 can be a 10kΩ resistor.

[0075] According to an embodiment of this application, the presence indication module may include an indication unit and a first resistor. The first resistor controls the presence pin of the first serial port connector to be in a high-level state, thereby sending a valid (high-level) presence signal to the first controller through the first serial port connector. By pulling the presence pin of the first serial port connector down to a low-level state, the second light-emitting diode in the indication unit is illuminated, indicating different states of the current process during different stages of server power-on and power-off, thus improving the efficiency of maintenance and replacement.

[0076] Please continue to refer to the above. Figure 3 As shown, the indicator unit 301 may include a second resistor R2 and a second light-emitting diode LED5.

[0077] Specifically, the first end of the second resistor R2 can be electrically connected to the power supply U_3V3, and the second end of the second resistor R2 can be electrically connected to the first end of the second light-emitting diode LED5.

[0078] The first terminal of the second light-emitting diode LED5 can be electrically connected to the second terminal of the second resistor R2, and the second terminal of the second light-emitting diode LED5 can be electrically connected to the in-situ pin of the first serial connector 105.

[0079] Responding to different stages of the maintenance process, the illumination of the second LED5 can have different meanings during the power-on and power-off phases of the server power supply P3V3_STBY. For example, during the power-on phase of the server power supply P3V3_STBY, the illumination of the second LED5 can indicate that at least one fault information has been written to the memory. During the power-off phase of the server power supply P3V3_STBY, the illumination of the second LED5 can indicate that at least one fault information in the memory has been updated, and the faulty graphics processor has been repaired and replaced.

[0080] According to an embodiment of this application, when at least one fault information has been written into the memory or the second controller has completed updating at least one fault information stored in the memory, the in-situ pin of the first serial connector 105 is pulled down to a low level, causing the second light-emitting diode LED5 to be forward biased and lit at a third frequency.

[0081] According to the embodiments of this application, by lighting up the second light-emitting diode at different stages of server power-on and power-off, the different states of the current process can be indicated. Thus, the different states can be obtained according to the different expressions such as the on / off state or flashing frequency of the second light-emitting diode, so as to promptly prompt the completion of storage or maintenance and replacement, thereby improving the efficiency of maintenance and replacement while quickly locating faults.

[0082] On the side opposite to the fault location device and the first serial port connector, a general interface such as a USB interface can also be provided for connecting external devices. This allows for monitoring and acquisition of relevant alarm information and fault information in the memory during the repair and replacement process, enabling targeted repair and maintenance of the graphics processor or faulty server.

[0083] Figure 4 A schematic diagram of a fault location device according to another embodiment of this application is shown.

[0084] like Figure 4 As shown, the fault location device 103 may include a power supply U_3V3, a first serial port connector 105, a presence indicator module, and a control storage module. The presence indicator module may include a first resistor R1, a second resistor R2, and a second light-emitting diode LED5. The control storage module may include a memory 201, a second controller 202, and a multiplexer 203. The specific connection methods between the components can be found in [reference needed]. Figures 1-3 The relevant description in the document.

[0085] When it is necessary to replace the faulty graphics processor in the server, the fault location device 103 can be electrically connected to the first controller 109 in the motherboard 104 through the first serial port connector 105 and the second serial port connector 108. When the server power supply P3V3_STBY is powered on and the presence signal is valid, the data transmission link between the first controller 109 and the memory 201 is turned on. The baseboard management controller 204 writes the fault information into the memory 201 through the first controller 109, and after the writing is completed, pulls down the presence pin of the first serial port connector 105 to light up the second light-emitting diode LED5 in the fault location device 103.

[0086] When the server power supply P3V3_STBY is powered off and the presence signal is valid, the data transmission link between the first controller 109 and the second controller 202 is established. The second controller 202 determines the target fault information from the memory 201 and sends it to the first controller 109. The first controller 109 parses the target fault information and controls the input / output expander 110 in the graphics processing card to light up the first LED corresponding to the faulty graphics processor. After the faulty graphics processor is replaced, the first controller 109 pulls down the presence pin of the first serial connector 105 to light up the second LED 5 in the fault location device 103.

[0087] Figure 5 A schematic diagram of a fault location device according to yet another embodiment of this application is shown.

[0088] like Figure 5 As shown, when the second controller 202 in the fault location device 103 malfunctions, the first controller 109 in the motherboard 104 can temporarily compensate for the operation process executed by the second controller 202. The specific connection methods between the devices and the writing of fault information into the memory 201 can be found in [reference needed]. Figure 4 The relevant description in the document.

[0089] When the server power supply P3V3_STBY is powered off and the presence signal is valid, the first controller 109 reads the fault information from the memory 201 and determines the target fault information. Then, based on the target fault information, it controls the input / output expander 110 in the graphics processing card to light up the first LED corresponding to the faulty graphics processor. After the faulty graphics processor is replaced, the first controller 109 pulls down the presence pin of the first serial connector 105 to light up the second LED 5 in the fault location device.

[0090] According to the embodiments of this application, the fault location device of this application can store fault information of multiple servers in a unified manner, and then power down and provide fault location prompts for each server in a unified manner. Alternatively, it can store fault information of only one server at a time, and then power down and provide fault location prompts for that server.

[0091] Figure 6 A flowchart of a fault location method according to an embodiment of this application is shown.

[0092] like Figure 6As shown, the fault location method of this embodiment may include operations S601 to S619. Starting S601, in response to the server power-on, the fault location device is inserted into the motherboard of the nth server. S602, the first controller determines whether the presence signal of the nth fault location device is valid. S603, if the presence signal is invalid, the clock and data pins of the second serial port connector are kept in UART (Universal Asynchronous Receiver-Transmitter) format.

[0093] When the presence signal is at an active level, the baseboard management controller writes at least one fault information into the memory via the first controller until the writing is complete, illuminating the second LED S605. A determination is made on whether fault information needs to be stored for N servers S606. If storage of fault information for N servers is required, a determination is made on whether all fault information for the N servers has been stored S607. If storage is not complete, the process returns to operations S602-S606. If storage is complete, the fault location device is unplugged and the power to the N servers is turned off S608. If storage of fault information for the N servers is not required, the fault location device is unplugged and the power to the nth server is turned off S609.

[0094] After the server power is off, the fault location device is inserted into the motherboard of the nth server (S610). The first controller determines whether the presence signal of the fault location device is valid (S611). If the presence signal is invalid, the clock and data pins of the second serial port connector are kept in UART format (S604). If the presence signal is valid, the second controller determines the target fault information from at least one fault information based on the fault identifier of the motherboard and sends it to the first controller (S612).

[0095] The first controller parses the received target fault information and controls the input / output expanders in the graphics processing card to light up the corresponding first LEDs, so as to replace the faulty graphics processor according to the lit first LEDs (S613). After replacing the faulty graphics processor, the second controller checks whether the presence signals of multiple graphics processors are valid (S614). If all presence signals are valid, the fault information in the memory is updated and the second LEDs are lit (S615). If there are invalid presence signals, the first LEDs are lit until the replacement is successful or the replacement is paused (S616). It is determined whether the faulty graphics processors in N servers have been replaced (S617). If it is not necessary to replace the faulty graphics processors in N servers, the process ends (S618). If it is necessary to replace the faulty graphics processors in N servers, it is determined whether the replacement of the faulty graphics processors in N servers has been completed (S619). If the replacement has not been completed, the process returns to operations S610 to S617. If the replacement has been completed, the process ends (S618).

[0096] Figure 7 A schematic diagram of a server system according to an embodiment of this application is shown.

[0097] like Figure 7 As shown, a server system involving fault diagnosis during server power outages may include the aforementioned fault location device, motherboard, and graphics processing card.

[0098] Specifically, the fault location device 103 can refer to the above-mentioned structure, function and effect, and will not be repeated here.

[0099] The motherboard 104 can be electrically connected to the fault location device 103 and the graphics processing card 101. It includes a baseboard management controller 204, a first controller 109, a second serial port connector 108, and a third serial port connector 701. The baseboard management controller 204 is used to send at least one fault information in a predetermined format to the fault location device 103 through the first controller 109 and the second serial port connector 108. The first controller 109 is used to determine at least one faulty graphics processor from the target fault information and send a control signal to the graphics processing card 101 through the third serial port connector 701 to control the first light-emitting diodes (LED1~LED4) set corresponding to the at least one faulty graphics processor.

[0100] The power supply terminal of the first controller 109 can be electrically connected to the power supply U_3V3 via the first diode D1, and simultaneously electrically connected to the server power supply P3V3_STBY via the second diode D2. The input and output terminals of the first controller 109 are grounded via the third resistor R3. The first diode D1 and the second diode D2 can be used to isolate the power supply U_3V3 from the server power supply P3V3_STBY. The third resistor R3 is used to pull down the in-position pin of the first serial connector 105 to a low level, causing the second light-emitting diode LED5 to blink. The third resistor can be a 100kΩ resistor.

[0101] There are two communication links between the baseboard management controller 204 and the first controller 109. One communication link can be a UART format link (TX and RX), and the other communication link can be an I2C format link (SCL and SDA).

[0102] The first controller 109 may contain a control multiplexing integration module. The control multiplexing integration module can perform control operations such as parsing and identifying target fault information, and can also switch the signal transmission format (UART or I2C) between the memory 201, the second controller 202 and the input / output expander 110 according to the level of the fault location device's on-site signal.

[0103] The graphics processing card 101 may include an input / output expander 110, at least one first light-emitting diode (LED1~LED4), at least one graphics processor (four graphics processors are shown in the figure), and a fourth serial port connector 702. The input / output expander 110 is electrically connected to at least one first light-emitting diode (LED1~LED4) and at least one graphics processor. The input / output expander 110 is used to turn on or off the first light-emitting diode corresponding to at least one faulty graphics processor in response to a control signal transmitted through the fourth serial port connector 702, monitor the presence signal of at least one graphics processor, and send it to the first controller 109.

[0104] The power supply terminal of the input / output expander 110 can be electrically connected to the power supply U_3V3 via the third diode D3, and the power supply terminal of the input / output expander can be electrically connected to the server power supply P3V3_STBY via the fourth diode D4. The third diode D3 and the fourth diode D4 are used to isolate the power supply U_3V3 from the server power supply P3V3_STBY.

[0105] The input / output expander 110 may include an in-situ signal detection interface and an LED connection interface corresponding to the number of graphics processors. The in-situ signal detection interface is electrically connected to the graphics processor, and the LED connection interface is electrically connected to the first LED. The input / output expander 110 may also include I2C format link interfaces (SCL and SDA), which are electrically connected to the first controller 109 through a third serial connector 701 and a fourth serial connector 702.

[0106] According to an embodiment of this application, based on the power supply status of the server power supply and the level status of the presence signal of the fault location device, the fault location device controls the first light-emitting diode corresponding to the faulty graphics processor to flash. This allows for rapid location and prompting of multiple faulty graphics processors when multiple graphics processors fail, enabling efficient replacement of the faulty graphics processors, reducing maintenance costs, and mitigating equipment failures and damage caused by incorrect replacements.

[0107] According to an embodiment of this application, the first controller 109 can also be used to: when a first light-emitting diode corresponding to at least one faulty graphics processor is turned off, the first controller generates an update signal and sends it to the fault location device based on the presence signal of at least one graphics processor monitored via the input / output expander and the graphics processor identifier of at least one graphics processor.

[0108] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0109] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0110] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.

[0111] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.

Claims

1. A fault location device, characterized in that, The fault location device includes: First serial port connector; The presence indicator module is electrically connected to the first serial port connector and is used to send a presence signal corresponding to the presence indicator module to the first controller of the motherboard through the first serial port connector when the first serial port connector is electrically connected to the second serial port connector of the motherboard. A control storage module is electrically connected to the power supply within the fault location device. The control storage module includes a memory electrically connected to the first serial port connector and a second controller electrically connected to the memory. The memory is used to acquire and store fault information of at least one graphics processor when the server power is on and the presence signal corresponding to the presence indication module is at a valid level. The second controller is used to determine target fault information from the fault information based on the fault identifier of the motherboard and send it to the first controller when the server power is switched off and both the motherboard and the graphics processing card electrically connected to the control storage module are powered by the power supply. This allows the first controller to... The target fault information identifies at least one faulty graphics processor on the graphics processing card, and controls a first light-emitting diode corresponding to the at least one faulty graphics processor to light up or turn off at a first frequency. When the at least one faulty graphics processor is replaced, the first controller generates an update signal based on the real-time monitored presence signals and graphics processor identifiers corresponding to multiple graphics processors. When the second controller receives the update signal sent by the first controller, it updates the fault information stored in the memory based on the processor fault identifier in the target fault information and the update information corresponding to the at least one faulty graphics processor in the update signal.

2. The fault location device according to claim 1, characterized in that, The control storage module further includes: A multiplexer, electrically connected to the memory, the second controller, and the first serial port connector, is used for... When the server power is on and the in-situ signal is at an active level, the data transmission path between the first controller and the memory is activated so that at least one fault information is transmitted to the memory. When the server power is off and both the motherboard and the graphics processing card are powered by the power supply, the data transmission path between the first controller and the second controller is activated, so that the second controller can send the target fault information and receive the update signal.

3. The fault location device according to claim 1, characterized in that, The second controller is also used for: When the server power is switched off and both the motherboard and the graphics processing card are powered by the power supply, the fault information to be repaired is read from at least one fault information based on the fault identifier of the motherboard. If the fault information to be repaired includes fault information of a faulty graphics processor, the fault information to be repaired is identified as the target fault information and sent to the first controller through the first serial port connector. When the fault information to be repaired includes fault information of multiple faulty graphics processors, the fault type and fault model of the multiple faulty graphics processors are determined from the fault information to be repaired. Based on multiple fault types and multiple fault models, a fault indication strategy is determined and the target fault information is generated, so that the first controller controls the first light-emitting diode corresponding to the multiple fault graphics processors to light up or turn off at the first frequency according to the target fault information.

4. The fault location device according to claim 1, characterized in that, The second controller is also used for: Upon receiving an update signal generated by the first controller based on the presence signals and graphics processor identifiers of the replaced multiple graphics processors, the level status of the multiple presence signals is confirmed based on the processor fault identifier in the target fault information. When all of the multiple in-situ signals are at an active level, the fault status of at least one faulty graphics processor in the fault information is updated; In the event that an invalid level exists among the plurality of in-situ signals, an alarm message is generated and sent to the first controller based on the graphics processor identifier corresponding to the invalid level in-situ signal, so that the first controller controls at least one first light-emitting diode to light up or turn off at a second frequency based on the alarm message.

5. The fault location device according to claim 1, characterized in that, The presence indication module includes: An indicator unit, which is electrically connected to the in-situ pin of the first serial port connector, is used to forward bias the second light-emitting diode in the indicator unit and light it up at a third frequency when the in-situ pin of the first serial port connector is pulled down to a low level. The first resistor, connected in parallel with the indicator unit, is used to pull up the presence pin of the first serial port connector to a high level, thereby obtaining a valid presence signal that is sent to the first controller through the first serial port connector.

6. The fault location device according to claim 5, characterized in that, The indicating unit includes: The second resistor has its first end electrically connected to the power supply and its second end electrically connected to the first end of the second light-emitting diode. The second light-emitting diode has its first end electrically connected to the second end of the second resistor, and its second end electrically connected to the in-situ pin of the first serial port connector.

7. The fault location device according to claim 5, characterized in that, When at least one fault information has been written to the memory or the second controller has completed updating the at least one fault information stored in the memory, the in-situ pin of the first serial connector is pulled down to a low level, causing the second light-emitting diode to be forward biased and lit at the third frequency.

8. A server system, characterized in that, The system includes: The fault location device as described in any one of claims 1-7; The motherboard, electrically connected to the fault location device and the graphics processing card, includes a baseboard management controller, a first controller, a second serial port connector, and a third serial port connector. The baseboard management controller is used to send at least one fault information in a predetermined format to the fault location device through the first controller and the second serial port connector. The first controller is used to determine at least one faulty graphics processor from the target fault information and send a control signal to the graphics processing card through the third serial port connector to control a first light-emitting diode corresponding to the at least one faulty graphics processor. A graphics processing card includes an input / output expander, at least one first light-emitting diode (LED), at least one graphics processor, and a fourth serial port connector. The input / output expander is electrically connected to the at least one first LED and the at least one graphics processor. The input / output expander is used to illuminate or extinguish the first LED corresponding to the at least one faulty graphics processor in response to a control signal transmitted through the fourth serial port connector, monitor the presence signal of the at least one graphics processor, and send it to the first controller.

9. The system according to claim 8, characterized in that, The first controller is further configured to: when the first light-emitting diode corresponding to the at least one faulty graphics processor is turned off, the first controller generates an update signal and sends it to the fault location device based on the presence signal of the at least one graphics processor monitored via the input / output expander and the graphics processor identifier of the at least one graphics processor.

Citation Information

Patent Citations

  • Hard disk fault detection method, system and device, computer equipment and storage medium

    CN116795610A