Power supply status monitoring system and server

By introducing a two-way signal monitoring overcurrent protection unit in the power supply unit, the problem of false alarm when the PSU fails is solved, and the operation and maintenance efficiency of the server is improved.

CN120523692BActive Publication Date: 2025-09-23INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511025012.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-09-23
Estimated Expiration
2045-07-24

AI Technical Summary

Technical Problem

When a PSU fails, the status of the overcurrent protection unit is prone to false alarms, affecting operation and maintenance efficiency.

Method used

Two signals from the power supply unit are introduced as the basis for the controller to monitor and determine whether the first overcurrent protection unit is in an abnormal state. The power distribution unit outputs the first indication signal and the second indication signal to update the control signal of the first overcurrent protection unit to prevent false alarms.

Benefits of technology

Effectively prevents false alarms of overcurrent protection unit status caused by PSU failure, improving operation and maintenance efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523692B_ABST
    Figure CN120523692B_ABST
Patent Text Reader

Abstract

The present application discloses a power supply status monitoring system and server, relating to the field of server technology. The system includes at least one power supply unit, a power distribution unit, a controller, and a first overcurrent protection unit. The power distribution unit is configured to output a first indication signal and a second indication signal, wherein the first indication signal is configured to indicate whether the input voltage of the power supply unit is within a preset input voltage range, and the second indication signal is configured to indicate whether the output voltage of the power supply unit is within a preset output voltage range. The controller is configured to update a control signal of the first overcurrent protection unit based on the first indication signal and the second indication signal, and determine whether to output a status abnormality signal based on the updated control signal, wherein the status abnormality signal is configured to indicate that the first overcurrent protection unit is in an abnormal state. The present application solves the technical problem in the related art that when a power supply unit fails, the status of the overcurrent protection unit is prone to false alarms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of servers, and in particular to a power supply status monitoring system and a server. Background Art

[0002] With the development of cloud computing and artificial intelligence (AI) applications, users' requirements for server computing performance are constantly increasing. For example, users require more graphics processing units (GPUs) to be deployed in servers to enhance their computing performance.

[0003] To enhance GPU power supply safety, some solutions require adding an overcurrent protection unit (OCP) to the GPU power input to provide timely protection against overcurrent or short circuits. However, in these solutions, the OCP status is prone to false alarms when the power supply unit (PSU) fails, impacting operational efficiency. Summary of the Invention

[0004] The present application provides a power supply status monitoring system and server to at least solve the technical problem in the related art that when a PSU fails, the status of the overcurrent protection unit is prone to false alarms.

[0005] The present application provides a power supply status monitoring system, comprising at least one power supply unit, a power distribution unit, a controller and a first overcurrent protection unit; the power distribution unit is connected in series between the power supply unit and the first overcurrent protection unit.

[0006] The power distribution unit is used to output a first indication signal and a second indication signal. The first indication signal is used to indicate whether the input voltage of the power supply unit is within a preset input voltage range. The second indication signal is used to indicate whether the output voltage of the power supply unit is within a preset output voltage range.

[0007] The controller is used to update the control signal of the first overcurrent protection unit based on the first indication signal and the second indication signal, and determine whether to output a state abnormality signal according to the updated control signal. The state abnormality signal is used to indicate that the first overcurrent protection unit is in an abnormal state.

[0008] The present application also provides a server, including the above-mentioned power supply status monitoring system.

[0009] The power supply status monitoring system and server provided in the embodiments of the present application can effectively prevent false alarms of the status of the first overcurrent protection unit due to PSU failure by introducing two signals from the power supply unit as the basis for the controller to monitor and determine whether the first overcurrent protection unit is in an abnormal state, thereby effectively improving operation and maintenance efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0011] Figure 1 This is a schematic diagram of a GPU power supply status monitoring system provided in an embodiment of the present application;

[0012] Figure 2 A schematic diagram of the structure of a power supply status monitoring system provided in an embodiment of the present application;

[0013] Figure 3 This is another structural diagram of a power supply status monitoring system provided in an embodiment of the present application;

[0014] Figure 4 A schematic diagram of the interconnection topology of a signal switching unit, a controller, and a baseboard management controller (BMC) provided in an embodiment of the present application;

[0015] Figure 5 This is a schematic diagram of another GPU power supply status monitoring system provided in an embodiment of the present application;

[0016] Figure 6 is a schematic diagram of the internal structure of a signal processing unit provided in an embodiment of the present application;

[0017] Figure 7 It is the logic truth table corresponding to each input signal and output signal in the signal processing unit;

[0018] Figure 8 This is a power-off sequence of an AC mains power failure system provided in the embodiment of the present application. Figure 1 ;

[0019] Figure 9 This is a power-off sequence of an AC mains power failure system provided in the embodiment of the present application. Figure 2 ;

[0020] Figure 10This is a power-off sequence of an AC mains power failure system provided in the embodiment of the present application. Figure 3 . DETAILED DESCRIPTION

[0021] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0022] To facilitate a clear description of the technical solutions of the embodiments of this application, the words "exemplary," "for example," and the like are used in the embodiments of this application to indicate examples, illustrations, or explanations. Any embodiment or design described in this application as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0023] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0024] In the embodiments of this application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects are in an "or" relationship.

[0025] The following is an explanation of some abbreviations and key terms involved in the application examples.

[0026] Power supply unit (PSU): used to convert AC power into DC power required by electronic equipment, providing stable power to various components of the equipment.

[0027] Power distribution board (PDB): used to distribute the power provided by the PSU to the various power-consuming units on the circuit board.

[0028] Overcurrent protection (OCP): When excessive current flows in a circuit, the protection device activates to cut off the circuit or limit the current to prevent equipment from being damaged by overcurrent.

[0029] Electronic fuse (EFUSE): An overcurrent protection component made of semiconductor materials. Compared with traditional fuses, it has the advantages of being reusable and having a fast response speed.

[0030] Voltage regulator module (VRM): This module converts the input voltage into the stable low voltage required by the processor, graphics card, etc., and provides dynamic response, efficient heat dissipation, and intelligent protection.

[0031] Baseboard Management Controller (BMC): The core controller in the server, used for remote monitoring, management, and maintenance of the server, enabling functions such as powering on and off the server and checking hardware status.

[0032] Complex programmable logic device (CPLD): A high-density programmable logic device with an integration density greater than 1000 gates and more input / output signals, product terms, and macrocells.

[0033] In today's digital age, cloud computing technology is evolving at an unprecedented rate, with its service models continuously expanding from basic storage and computing services to more complex areas such as analysis and decision support. Simultaneously, AI applications are experiencing explosive growth, encompassing cutting-edge fields such as natural language processing, computer vision, and reinforcement learning. These applications are expanding across a wide range of scenarios, from smart homes and transportation to medical diagnostics, profoundly changing the way people live and work.

[0034] Against this backdrop, user demands for server computing performance have dramatically increased. These demands are becoming increasingly stringent and escalating. In practical applications, this demand manifests itself in a number of specific requirements, one of which is the deployment of a greater number of graphics processing units (GPUs) in server systems.

[0035] Specifically, deep learning training requires complex matrix operations and neural network calculations on massive amounts of data. The parallel computing capabilities of GPUs can significantly shorten training time and improve training efficiency. In a cloud computing environment, multiple users simultaneously initiate computing requests, requiring the server to respond quickly and process them efficiently. Increasing the number of GPUs can enhance the server's parallel processing capabilities, ensuring that all user requests are processed promptly.

[0036] Therefore, in order to meet the ever-increasing demand for server computing performance from cloud computing and AI applications, deploying more GPUs in server systems has become an inevitable trend. By fully leveraging the parallel computing advantages of GPUs, the overall computing performance of the server can be effectively improved, providing users with better quality and more efficient services.

[0037] When GPUs are introduced on a large scale into server systems, the cumulative power consumption of multiple GPUs significantly impacts the current server power supply architecture. Existing server power supply systems are typically designed based on traditional load requirements, and their power capacity and power supply stability may not meet the power consumption requirements of a large number of GPUs running simultaneously. This can lead to problems such as voltage drops and increased current fluctuations, which in turn affect GPU performance and stability, and may even cause server system failures.

[0038] To effectively address GPU power supply safety issues and improve server system reliability and stability, some advanced solutions include adding an electronic fuse (EFUSE) protection circuit to the GPU power input. EFUSE, a semiconductor-based overcurrent protection component, offers advantages such as fast response, high accuracy, and reusability. Compared to traditional fuses, EFUSE can detect overcurrent or short-circuit faults much faster and quickly disconnect the circuit, effectively protecting the GPU and other server components from damage.

[0039] Specifically, when a GPU experiences an overcurrent or short-circuit fault, the EFUSE protection circuit immediately senses the abnormal current change and triggers a protective action within nanoseconds, disconnecting the faulty circuit from the power supply. This process not only prevents circuit board burnout caused by overcurrent or short-circuit, but also electrically isolates the faulty component from other normally functioning components, preventing the spread of the fault.

[0040] In addition, the EFUSE protection circuit also has a self-recovery function. After the fault is eliminated, it can automatically restore to normal working state without manual replacement of the fuse, thereby improving the maintainability and availability of the server system.

[0041] For example, in AI server applications, a single high-performance GPU running at full power can consume a high power consumption range of 600W to 800W. When a server needs to deploy eight such high-performance GPUs to meet large-scale AI computing tasks, the total power consumption will rise sharply to 4800W to 6400W.

[0042] From the PSU's perspective, if a 12V DC voltage is used to power the GPU, according to the basic power formula P=UI (where P is power, U is voltage, and I is current), the PSU needs to provide a current range of 400A-533A. This massive load current makes connector and cable selection a pressing challenge.

[0043] When it comes to connector selection, to ensure reliability and stability under high current conditions, the connector must have sufficient current-carrying and heat-dissipating capabilities. Based on current-carrying capacity standards, connector parameters such as the conductive cross-sectional area and contact resistance require careful design and optimization. This inevitably increases the size of the connector to increase the cross-sectional area of ​​the conductive material, reduce resistance, and minimize heat generation. However, within the compact structural space of AI servers, the introduction of large connectors presents a series of spatial layout issues. Servers typically utilize a highly integrated design with small spacing between hardware components. Large connectors may interfere with other components, impacting the overall assembly and maintenance of the server. For example, they may collide with components such as cooling fans and memory slots, preventing proper installation or affecting heat dissipation.

[0044] When selecting cables, high currents require thicker wires. According to the law of resistance, increasing wire diameter reduces cable resistance, thereby reducing heat generated under high current conditions. However, thicker cables have several impacts. When routing cables within the chassis, thicker cables occupy more space, further congesting already limited routing channels. This not only increases routing complexity but can also cause cables to cross and become tangled, impacting signal transmission stability and reliability. During cable management, thicker cables are less flexible, making them difficult to neatly organize and secure. This can cause cables to vibrate and loosen during server operation, leading to faults such as poor connection. Furthermore, thicker cables can negatively impact heat dissipation within the chassis. They block airflow, creating localized heat sinks and uneven temperature distribution within the server, impacting the performance and lifespan of hardware components.

[0045] At the same time, the conduction loss caused by high current in the system is also an issue that cannot be ignored. According to Joule's law, under high current conditions, even if the resistance of cables and connectors is low, a large amount of heat will be generated, leading to increased conduction loss. This not only wastes electricity and reduces the energy efficiency of the server, but also increases the electricity costs of the user's computer room. From an energy management perspective, increased conduction loss means that the server needs to consume more electricity to maintain its normal operation during operation, which runs counter to the current development trend of energy conservation, emission reduction, and green data centers. Therefore, how to reduce conduction loss under high current and improve the energy efficiency of the server is also one of the key issues that need to be considered in AI server design.

[0046] To summarize, in related technologies, the PSU power supply bus voltage in AI server systems usually uses high-voltage DC 54V, and then a power brick solution is used to convert the 54V power supply into a 12V power supply. The 12V power supply can power various components in the server system, such as the motherboard, hard disk, GPU, fan, etc.

[0047] Because AI servers are equipped with numerous GPUs, GPU power supply reliability is crucial to the continued stable operation of AI server services. Therefore, a system is required to monitor and manage the GPU power supply status in real time. Related technologies typically use a CPLD or BMC to obtain the EFUSE control signal EN and the power establishment completion monitoring signal PG corresponding to each GPU, and then determine whether the EFUSE is in an abnormal state.

[0048] For example, in some AI servers, N+N PSUs are usually used to provide 54V voltage. The 54V voltage is converted into 12V voltage through the power brick, and then converted into 12V standby voltage through 12V EFUSE. The 12V standby voltage can finally be converted into P3V3_STBY standby voltage to power the CPLD on the motherboard or GPU switch board.

[0049] like Figure 1 As shown, Figure 1 This is a schematic diagram of a GPU power supply status monitoring system provided in an embodiment of the present application.

[0050] For example, the following describes an 8-GPU configuration with 2+2 redundant PSU power supplies.

[0051] The GPU power supply status monitoring system includes: an alternating current (AC) mains power access unit, 2+2 redundant PSUs, a PDB, a P54V EFUSE, GPU EFUSEs (including EFUSE: P12V_GPU0, EFUSE: P12V_GPU1, EFUSE: P12V_GPU2, EFUSE: P12V_GPU3, EFUSE: P12V_GPU4, EFUSE: P12V_GPU5, EFUSE: P12V_GPU6, EFUSE: P12V_GPU7), GPU components (including GPU0, GPU1, GPU2, GPU3, GPU4, GPU5, GPU6, GPU7), a P12V_STBY EFUSE, a VR: P3V3_STBY, a CPLD, and a BMC.

[0052] The AC mains power access unit serves as the initial energy source for the entire system, responsible for introducing external AC mains power into the server. It must have standard interfaces and electrical characteristics, adaptable to regional power supply specifications, and provide a stable foundation for subsequent power conversion. This unit is also typically equipped with certain protection mechanisms, such as lightning protection and overvoltage protection, to prevent abnormal conditions in the external power grid from damaging the server's internal circuitry. Optionally, the AC mains power access unit can be used to provide a 220V voltage.

[0053] 2+2 redundant PSUs can be divided into two groups, with two PSUs in each group working in parallel. Under normal operating conditions, the two PSUs jointly provide power to the system, sharing the load and improving the power supply's output capacity. If any one PSU in a group fails, the remaining PSU can immediately assume the entire load, ensuring continuous server operation and avoiding business interruptions caused by power failures. Each PSU has an integrated power conversion circuit that can convert the input AC power into DC power suitable for internal use in the server and has comprehensive protection functions such as overcurrent protection, overvoltage protection, and short-circuit protection.

[0054] The PDB receives power from four PSUs and distributes it efficiently to various parts of the system. It features highly accurate power distribution and excellent electrical isolation, ensuring that power supply to different components does not interfere with each other. The PDB may also integrate power monitoring features, such as real-time voltage and current monitoring, to promptly detect abnormalities in the power supply process.

[0055] The P54V EFUSE plays a crucial role in overcurrent protection during power distribution. Positioned between the PDB and subsequent circuitry, it quickly fuses when an overcurrent condition occurs, shutting off the circuit and preventing damage to subsequent electronic components. Compared to traditional fuses, the P54V EFUSE offers advantages such as faster response and reusability, better meeting the power protection requirements of modern electronic devices.

[0056] Corresponding GPU EFUSEs are configured for each of the eight GPUs, from EFUSE: P12V_GPU0 to EFUSE: P12V_GPU7. These EFUSEs also function as overcurrent protection, providing independent power protection for each GPU. If a GPU experiences a short circuit or overcurrent fault, the corresponding GPU EFUSE will fuse, shutting off power to that GPU, preventing the fault from spreading to other components and ensuring the safe operation of the entire system.

[0057] P12V_STBY EFUSE: This EFUSE protects the standby voltage path. The standby voltage P12V_STBY provides power to critical management components in the server. These components require a stable power supply even when the server is in standby mode. The P12V_STBY EFUSE promptly cuts off power if an abnormality occurs in the standby voltage path, protecting related components from damage.

[0058] VR:P3V3_STBY can be used to convert the standby voltage P12V_STBY into a stable 3.3V voltage P3V3_STBY to power management components such as CPLD and BMC.

[0059] CPLDs can be used for server initialization, logic control, and signal processing. In power supply status monitoring systems, CPLDs can collect real-time status information from each power node, such as voltage, current, and EFUSE status, and perform judgments and processing based on pre-set logic rules. When a power anomaly is detected, the CPLD can promptly issue an alarm and implement appropriate protective measures, such as powering off or restarting components.

[0060] The BMC is responsible for comprehensive monitoring and management of the server's hardware status, including monitoring and alarming of parameters such as temperature, voltage, and fan speed. In the power status monitoring system, the BMC communicates with the CPLD and other power monitoring components to obtain detailed power status information and upload this information to the remote management terminal, allowing managers to understand the server's power status in real time. Furthermore, the BMC can remotely control and configure the server based on power status information, improving server manageability and maintainability.

[0061] In some embodiments, the GPU power supply status monitoring system operates as follows:

[0062] 220VAC AC mains power is drawn from the power distribution cabinet (PDC) and used as the PSU input. Four PSUs are plugged into the PDB to provide 2+2 redundant power. The output power is P54V_PSU, which then flows to the GPU switch board and outputs P54V power via EFUSE:P54V_PSU. This P54V power is converted by the VRM and output as P12V_IN. P12V_IN is split into multiple paths. One path is output as the standby voltage P12V_STBY via EFUSE:P12V_STBY, which is then output as P3V3_STBY via VR:P3V3_STBY to power the CPLD and BMC. The other paths sequentially pass through EFUSE:P12V_GPU0, EFUSE:P12V_GPU1, EFUSE:P12V_GPU2, EFUSE:P12V_GPU3, EFUSE:P12V_GPU4, EFUSE:P12V_GPU5, EFUSE:P12V_GPU6, and EFUSE:P12V_GPU7 to supply power to the corresponding components GPU0, GPU1, GPU2, GPU3, GPU4, GPU5, GPU6, and GPU7.

[0063] After P3V3_STBY is established and CPLD initialization is completed, the control signal EN_P12V_GPUn (n=0,1,2,3,4,5,6,7) will be output, and the PG signal PG_P12V_GPUn (n=0,1,2,3,4,5,6,7) of the corresponding GPU EFUSE will be monitored.

[0064] If the CPLD detects that EN_P12V_GPUn is high and the PG signal of the GPUn EFUSE changes from high to low, and the duration of the change is greater than the preset time (such as 25us), the CPLD determines that the GPUn EFUSE is in an abnormal state.

[0065] Once the GPUn EFUSE is determined to be in an abnormal state, the CPLD immediately outputs the abnormal state signal FAULT#_P12V_GPUn (n=0, 1, ..., 7). These abnormal state signals are then converted and transmitted by the signal switching unit and finally accurately transmitted to the BMC unit via the I2C bus.

[0066] After receiving an abnormal status signal from the CPLD, the BMC accurately records detailed fault information, including the faulty GPU number, fault type, and occurrence time, in a log file. These log files provide comprehensive and accurate fault analysis for maintenance personnel, helping them quickly locate the problem, troubleshoot the cause, and implement appropriate maintenance measures, thereby ensuring stable server system operation.

[0067] In some embodiments, if the PSU fails and loses power while the P3V3_STBY on the GPU switch board continues to supply power, the CPLD will remain in normal operation. At this point, the CPLD will record the operating status of the GPU EFUSE. At the moment the PSU loses power, the EN signal of the GPU EFUSE is high. Due to the GPU load, the EFUSE output voltage is quickly pulled down, and the PG signal of the GPU EFUSE also becomes low. At this point, based on the EN signal being high and the PG signal switching from high to low, the CPLD will mistakenly report an error indicating that the GPU EFUSE is in an abnormal state and report the error to the BMC, creating an abnormality log. This inevitably causes difficulties for computer room operations and maintenance personnel in analyzing and locating the problem.

[0068] Therefore, how to prevent the EFUSE status from being falsely reported due to PSU power failure is a technical problem that urgently needs to be solved.

[0069] In response to the above technical problems, an embodiment of the present application provides a power supply status monitoring system. The system introduces two signals from the PSU as the basis for the CPLD to monitor and determine whether the GPU EFUSE is in an abnormal state, which can effectively prevent false alarms of the EFUSE status caused by PSU power failure.

[0070] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0071] Reference Figure 2 , Figure 2 This is a structural diagram of a power supply status monitoring system provided in an embodiment of the present application.

[0072] In some embodiments, the power supply status monitoring system includes a power supply unit, a power distribution unit, a controller, and a plurality of first overcurrent protection units (such as Figure 2 EFUSE0, EFUSE1, EFUSE2, EFUSE3, EFUSE4, EFUSE5, EFUSE6, and EFUSE7); the power distribution unit is connected in series between the power supply unit and the plurality of first overcurrent protection units.

[0073] Optionally, the power supply unit may adopt a redundant design, such as an N+N redundant PSU design.

[0074] Exemplarily, the power supply unit may adopt 2+2 redundant PSUs.

[0075] The 2+2 redundant PSUs can be divided into two groups, each with two PSUs working in parallel. Under normal operating conditions, the two PSU groups jointly provide power to the system, sharing the load. If any one PSU in one group fails, the remaining PSU group immediately assumes the entire load, ensuring continuous server operation and avoiding business interruptions caused by power failures.

[0076] The 2+2 redundant PSU configuration ensures stable equipment operation while maintaining reasonable cost control. For example, a small or medium-sized data center with dozens of servers can use a 2+2 redundant PSU configuration to provide reliable power to the servers, preventing server downtime caused by a single PSU failure, which could impact normal business operations.

[0077] The power distribution unit (PDU) is the hub of power distribution, receiving power from multiple PSUs and distributing it efficiently to various parts of the system. It features highly accurate power distribution and excellent electrical isolation, ensuring that power supply to different components does not interfere with each other. It also integrates power monitoring features, such as real-time voltage and current monitoring, to promptly detect any anomalies in the power supply process.

[0078] In some implementations, the controller can be responsible for server initialization, logic control, and signal processing. In a power supply status monitoring system, the controller can collect real-time status information from each power node, such as voltage, current, and EFUSE status, and perform judgments and processing based on pre-set logic rules. When a power anomaly is detected, the controller can promptly issue an alarm and implement appropriate protective measures, such as shutting off the power supply or restarting the component.

[0079] Optionally, the controller may be a CPLD.

[0080] Optionally, the power supply status monitoring system further includes at least one processor (e.g. Figure 2 Processor 0, processor 1, processor 2, processor 3, processor 4, processor 5, processor 6, processor 7).

[0081] Optionally, the processor may be a GPU, or another type of processor or load.

[0082] In some embodiments, the system can equip each processor with a first overcurrent protection unit. These first overcurrent protection units can be used for overcurrent protection, providing an independent power protection mechanism for each processor. When a processor experiences a short circuit or overcurrent fault, the corresponding first overcurrent protection unit will fuse, cutting off the power supply to that processor, preventing the fault from spreading to other components and ensuring the safe operation of the entire system.

[0083] In some embodiments, the processors are connected to their corresponding first overcurrent protection units. For example, processor 0 is connected to EFUSE0, processor 1 is connected to EFUSE1, ..., processor 7 is connected to EFUSE7.

[0084] In some embodiments, the power distribution unit is used to output a first indication signal and a second indication signal, wherein the first indication signal is used to indicate whether the input voltage of the power supply unit is within a preset input voltage range, and the second indication signal is used to indicate whether the output voltage of the power supply unit is within a preset output voltage range.

[0085] Exemplarily, the first indication signal is used to reflect whether the input voltage of the above-mentioned power supply unit is within a normal fluctuation range. When the input voltage of the above-mentioned power supply unit is within a normal fluctuation range, the first indication signal may indicate that the input voltage of the above-mentioned power supply unit is normal. When the input voltage of the above-mentioned power supply unit is within an abnormal fluctuation range, the first indication signal may indicate that the input voltage of the above-mentioned power supply unit is abnormal.

[0086] The input voltage range is determined based on the design requirements and actual application scenarios of the power supply unit. It takes into account the reasonable variation range of the input voltage under various operating conditions to ensure that the power supply unit can operate normally under different input conditions. When the measured input voltage value is within the normal fluctuation range, the first indication signal outputs a specific level signal, indicating that the input voltage of the power supply unit is normal. Conversely, when the measured input voltage value exceeds the normal fluctuation range, the first indication signal rapidly changes its output level state, indicating that the input voltage of the power supply unit is abnormal.

[0087] Similarly, the second indication signal is used to reflect whether the output voltage of the above-mentioned power supply unit is within the normal fluctuation range. When the output voltage of the above-mentioned power supply unit is within the normal fluctuation range, the second indication signal can indicate that the output voltage of the above-mentioned power supply unit is normal. When the output voltage of the above-mentioned power supply unit is within the abnormal fluctuation range, the second indication signal can indicate that the output voltage of the above-mentioned power supply unit is abnormal.

[0088] In some embodiments, the controller is configured to update the control signal EN (eg, Figure 2 EN_0, EN_1, EN_2, EN_3, EN_4, EN_5, EN_6, EN_7), and according to the updated control signal and the monitoring signal PG corresponding to the first overcurrent protection unit (such as Figure 2 PG_0_new, PG_1_new, PG_2_new, PG_3_new, PG_4_new, PG_5_new, PG_6_new, PG_7_new), determine whether to output a state abnormality signal; the state abnormality signal is used to indicate that the above-mentioned first overcurrent protection unit is in an abnormal state.

[0089] In some embodiments, the first indication signal and the second indication signal can generate a new control signal after logical operation processing, and the control signal can be used as the latest control signal of the first overcurrent protection unit to control the conduction or shutdown of the first overcurrent protection unit.

[0090] The above implementation method of generating the first overcurrent protection unit control signal by processing the first indication signal and the second indication signal through logical operation can respond quickly and accurately according to the real-time changes in the power supply status, effectively ensuring the reliable operation of the system.

[0091] In some implementations, it may be determined whether the first overcurrent protection unit is in an abnormal state based on the updated control signal.

[0092] In some embodiments, the controller is further configured to obtain a monitoring signal corresponding to the first overcurrent protection unit, where the monitoring signal is configured to indicate whether the supply voltage of the first overcurrent protection unit is within a preset supply voltage range.

[0093] In some embodiments, the controller may determine whether to output a status abnormality signal based on the updated control signal and the monitoring signal.

[0094] It can be understood that when the PSU fails, the above-mentioned first indication signal and / or second indication signal will output a low level, thereby allowing the control signal of the above-mentioned first overcurrent protection unit to be set to a low level in advance, thereby ensuring that when the system is powered off, the control signal of the above-mentioned first overcurrent protection unit always becomes a low level before the above-mentioned monitoring signal, thereby avoiding false alarms and erroneous actions of the controller.

[0095] The power supply status monitoring system provided in the embodiment of the present application introduces two signals of the PSU (the first indication signal and the second indication signal mentioned above) as the basis for the controller to monitor and determine whether the first overcurrent protection unit is in an abnormal state. This can effectively prevent the occurrence of false alarms of the status of the first overcurrent protection unit due to a PSU failure, thereby effectively improving operation and maintenance efficiency.

[0096] Reference Figure 3 , Figure 3 This is another structural schematic diagram of a power supply status monitoring system provided in an embodiment of the present application.

[0097] In some embodiments, the controller includes a signal processing unit, which is used to perform logical operations on the first indication signal, the second indication signal and the control signal, and update the control signal based on the operation results. The control signal is used to control the first overcurrent protection unit to be turned on or off.

[0098] In some embodiments, the signal processing unit includes at least one first logic gate, a second logic gate, and a third logic gate; the first logic gate is used to perform a first operation on the first indication signal and the second indication signal, and output a first signal; the second logic gate is used to perform a second operation based on the first signal, and output a second signal; the third logic gate is used to perform a first operation on the second signal and the control signal, and update the control signal according to the operation result.

[0099] Optionally, the first operation includes a logical AND operation, and the second operation includes a logical OR operation.

[0100] The logical AND operation (AND) combines the logical states of multiple input signals. The output is true (a high level "1") only when all input signals are true (typically represented by a high level "1" in digital circuits); the output is false (a low level "0") if one or more input signals are false (a low level "0"). This operation enables the power management system to effectively perform comprehensive judgments on multiple conditions. For example, when determining whether a PSU is operating normally, the first and second indication signals can be simultaneously ANDed. Only when both the first and second indication signals are normal is the PSU considered to be operating normally.

[0101] Optionally, the first logic gate and the third logic gate may be implemented as complementary metal-oxide-semiconductor (CMOS) AND gate circuits. The CMOS AND gate circuits may be constructed by cross-coupling positive-channel metal oxide semiconductor field-effect transistors (PMOS) and N-type metal-oxide-semiconductor (NMOS) transistors, with input signals controlling the conduction states of the transistors to implement logic operations.

[0102] The logical OR operation (OR) is used to combine the logical states of multiple input signals. As long as one or more of all the input signals are true (usually represented by a high level "1" in digital circuits), the output result is judged to be true (high level "1"); only when all the input signals are false (low level "0"), the output result is false (low level "0").

[0103] Taking fault detection in a power management system as an example, the system may need to monitor the status of multiple PSUs. Each PSU will generate a first signal indicating normal or faulty status. By using these first signals as inputs for a logical OR operation, as long as one PSU fails, the output signal will become "1", thereby determining that at least one of the multiple PSUs is faulty.

[0104] Optionally, the second logic gate may adopt a CMOS OR gate structure, which may be composed of a PMOS in parallel and an NMOS in series, and the input signal realizes logic operation by controlling the conduction of the transistor.

[0105] In some embodiments, when the fluctuation amplitude of the input voltage of the above-mentioned power supply unit is less than or equal to the first amplitude threshold, the input voltage fluctuates within the allowable stable range. At this time, the above-mentioned first indication signal is a first level signal, indicating that the input voltage of the above-mentioned power supply unit is normal; when the fluctuation amplitude of the input voltage of the above-mentioned power supply unit is greater than the above-mentioned first amplitude threshold, the input voltage fluctuates beyond the normal range. The above-mentioned first indication signal is a second level signal, indicating that the input voltage of the above-mentioned power supply unit is abnormal.

[0106] Similarly, when the fluctuation amplitude of the output voltage of the above-mentioned power supply unit is less than or equal to the second amplitude threshold, the output voltage fluctuates within the allowable stable range. At this time, the above-mentioned second indication signal is a first level signal, indicating that the output voltage of the above-mentioned power supply unit is normal; when the fluctuation amplitude of the output voltage of the above-mentioned power supply unit is greater than the above-mentioned second amplitude threshold, the output voltage fluctuates beyond the normal range. The above-mentioned second indication signal is a second level signal, indicating that the output voltage of the above-mentioned power supply unit is abnormal.

[0107] Optionally, the first level signal may be a high level, and the second level signal may be a low level, which is not limited in the embodiment of the present application.

[0108] In some embodiments, the power supply status monitoring system further includes at least one logic gate buffer unit, which is connected in series between the power distribution unit and the controller.

[0109] The logic gate buffer unit may be used to perform isolation enhancement processing on the first indication signal and the second indication signal.

[0110] Among them, the above-mentioned logic gate buffer unit, as a core component for signal isolation and enhancement, can ensure the reliable transmission of the first indication signal and the second indication signal in a complex electromagnetic environment through level conversion, driving capability improvement and noise isolation.

[0111] In some embodiments, the power supply status monitoring system further includes a voltage adjustment unit, which is connected in series between the power distribution unit and the first overcurrent protection unit.

[0112] In some embodiments, the voltage adjustment unit may be configured to convert the voltage output by the power distribution unit into a power supply voltage for the processor, for example, converting the 54V voltage output by the power distribution unit into a 12V power supply voltage for the processor.

[0113] The voltage adjustment unit can reduce the input high voltage according to a predetermined ratio through an internal power conversion circuit, such as a switching power supply circuit or a linear voltage regulator circuit.

[0114] In some embodiments, the power supply status monitoring system further includes a second overcurrent protection unit, which is connected in series between the power distribution unit and the voltage adjustment unit.

[0115] The second overcurrent protection unit acts as a critical barrier between the power distribution unit and the voltage regulation unit, providing full overcurrent protection from the power input to the load. When an overcurrent condition occurs in the circuit, the second overcurrent protection unit quickly fuses, shutting off the circuit and preventing the excessive current from damaging subsequent electronic components.

[0116] Optionally, the second overcurrent protection unit may include an EFUSE. Compared with traditional fuses, EFUSE has advantages such as fast response speed and reusability, and can better meet the power protection requirements of the server.

[0117] In some embodiments, the power supply status monitoring system further includes a voltage conversion unit, which is connected in series between the voltage adjustment unit and the controller.

[0118] In some embodiments, the voltage conversion unit may be configured to convert the voltage output by the voltage adjustment unit into a power supply voltage for the controller, for example, converting the P12V_STBY voltage output by the voltage adjustment unit into the power supply voltage P3V3_STBY for the processor.

[0119] In some embodiments, the power supply status monitoring system further includes a third overcurrent protection unit, which is connected in series between the voltage adjustment unit and the voltage conversion unit.

[0120] The third overcurrent protection unit acts as a key protective layer between the voltage adjustment unit and the voltage conversion unit, providing full overcurrent protection from the voltage adjustment unit to the voltage conversion unit. When an overcurrent condition occurs in the circuit, the third overcurrent protection unit quickly fuses, shutting off the circuit and preventing the excessive current from damaging subsequent electronic components (such as the controller).

[0121] Optionally, the third overcurrent protection unit may include an EFUSE.

[0122] In some embodiments, the power supply status monitoring system further includes a BMC, and the controller is connected to the BMC.

[0123] In some implementations, the controller is configured to send a status abnormality signal to the BMC; and the BMC is configured to output a log file based on the status abnormality signal.

[0124] For example, when the controller detects that the system is in an abnormal state, it can send a state abnormality signal to the BMC. After receiving the state abnormality signal sent by the controller, the BMC can generate a detailed log file according to the log recording rules predefined by the system.

[0125] Optionally, the log file may include the time, cause, relevant parameter values, processor location, etc. of the exception.

[0126] Optionally, the storage location of the above log files can be a local storage device (such as flash memory, hard disk) or a remote server to ensure the security and reliability of the log data; access rights are managed through user authentication and authorization mechanisms, and only authorized users can access and view the log files.

[0127] In some embodiments, the power supply status monitoring system further includes a signal switching unit, which is connected in series between the BMC and the controller.

[0128] In some implementations, the controller is configured to send a status abnormality signal to the signal switching unit; and the signal switching unit is configured to store the status abnormality signal in a corresponding port register.

[0129] Upon receiving a status abnormality signal from the controller, the signal switching unit immediately activates the signal storage mechanism. It houses multiple port registers, each corresponding to a specific signal channel. Based on the characteristics of the status abnormality signal and pre-set storage rules, the signal switching unit accurately stores the signal in the corresponding port register. The port registers offer high-speed read and write speeds and maintain data stability, ensuring that the status abnormality signal is not lost or distorted during storage, providing a reliable data source for subsequent BMC read operations.

[0130] In some embodiments, the BMC is used to address the serial bus address of the signal switching unit through the serial bus to obtain the value in the port register, which can be used to indicate whether the first overcurrent protection unit is in an abnormal state.

[0131] The BMC can address the serial bus address of the signal switching unit through the serial bus. The serial bus address is the unique identifier of the signal switching unit in the system. The BMC can accurately locate the target signal switching unit by sending a specific address signal.

[0132] Once the BMC successfully addresses the signal switching unit, it sends a read command to the unit according to the serial bus communication protocol. Upon receiving the read command, the signal switching unit reads the abnormal status signal value from the corresponding port register and transmits this value back to the BMC via the serial bus. After obtaining the value from the port register, the BMC parses and analyzes it to determine whether each first overcurrent protection unit is in an abnormal state.

[0133] Optionally, the serial bus may be an I2C (inter-integrated circuit) bus.

[0134] For example, refer to Figure 4 , Figure 4This is a schematic diagram of the interconnection topology of a signal switching unit, a controller, and a BMC provided in an embodiment of the present application.

[0135] In some implementations, the controller outputs the FAULT signal for each processor's corresponding EFUSE (e.g., FAULT#_Processor0, FAULT#_Processor1, FAULT#_Processor2, FAULT#_Processor3, FAULT#_Processor4, FAULT#_Processor5, FAULT#_Processor6, and FAULT#_Processor7) to the signal switching unit. The signal state for FAULT#_Processor n (n = 0, 1, ..., 7) is stored in the corresponding single-port register of the signal switching unit. The BMC can then address the signal switching unit's I2C address (defined by the high and low level settings of address pins A2, A1, or A01) via the I2C bus and read the transmitted value of the signal switching unit's port register to obtain the corresponding EFUSE FAULT signal.

[0136] The power supply status monitoring system provided in the embodiment of the present application can effectively prevent the false alarm of the status of the first overcurrent protection unit due to a PSU failure by introducing two signals from the PSU as the basis for the controller to monitor and judge whether the first overcurrent protection unit is in an abnormal state, thereby effectively improving the operation and maintenance efficiency.

[0137] In order to better understand the power supply status monitoring system provided by this application, the following embodiments are illustrated using actual application scenarios.

[0138] Reference Figure 5 , Figure 5 This is a schematic diagram of another GPU power supply status monitoring system provided in an embodiment of the present application.

[0139] For example, the following describes an 8-GPU configuration with 2+2 redundant PSU power supplies.

[0140] The GPU power supply status monitoring system includes: an AC mains access unit, 2+2 redundant PSUs, a PDB, a second overcurrent protection unit (EFUSE: P54V), the first overcurrent protection unit GPU EFUSE (including EFUSE: P12V_GPU0, EFUSE: P12V_GPU1, EFUSE: P12V_GPU2, EFUSE: P12V_GPU3, EFUSE: P12V_GPU4, EFUSE: P12V_GPU5, EFUSE: P12V_GPU6, EFUSE: P12V_GPU7), processors (including GPU0, GPU1, GPU2, GPU3, GPU4, GPU5, GPU6, GPU7), a voltage regulation unit (VRM: P54V / P12V), a third overcurrent protection unit (EFUSE: P12V_STBY), a voltage conversion unit (VR: P3V3_STBY), a CPLD, and a BMC.

[0141] The 2+2 redundant PSU configuration ensures stable equipment operation while maintaining reasonable cost control. For example, a small or medium-sized data center with dozens of servers can use a 2+2 redundant PSU configuration to provide reliable power to the servers, preventing server downtime caused by a single PSU failure, which could impact normal business operations.

[0142] In some embodiments, the GPU power supply status monitoring system operates as follows:

[0143] 220VAC AC mains power is drawn from the power distribution cabinet (PDC) and used as the PSU input. Four PSUs are plugged into the PDB to provide 2+2 redundant power. The output power is P54V_PSU, which then flows to the GPU switch board and outputs P54V power via EFUSE:P54V_PSU. This P54V power is converted by the VRM and output as P12V_IN. P12V_IN is split into multiple paths. One path is output as the standby voltage P12V_STBY via EFUSE:P12V_STBY, which is then output as P3V3_STBY via VR:P3V3_STBY to power the CPLD and BMC. The other paths sequentially pass through EFUSE:P12V_GPU0, EFUSE:P12V_GPU1, EFUSE:P12V_GPU2, EFUSE:P12V_GPU3, EFUSE:P12V_GPU4, EFUSE:P12V_GPU5, EFUSE:P12V_GPU6, and EFUSE:P12V_GPU7 to supply power to the corresponding components GPU0, GPU1, GPU2, GPU3, GPU4, GPU5, GPU6, and GPU7.

[0144] After P3V3_STBY is established and CPLD initialization is completed, the control signal EN_P12V_GPUn (n=0,1,2,3,4,5,6,7) will be output, and the PG signal of the corresponding GPU EFUSE will be monitored at the same time.

[0145] In some implementations, two signals from the PSU, namely the first indication signal PSU_AC_OK and the second indication signal PSU_PWROK, are introduced as a basis for the CPLD to monitor and determine whether the GPU EFUSE is in an abnormal state.

[0146] Among them, PSU0_AC_OK of PSU0 is the first indication signal that the AC mains input corresponding to PSU0 is normal (this signal indicates that the PSU input mains power meets the normal fluctuation range), and the PSU0_PWROK signal is the second indication signal that the output voltage of PSU0 is established and the output voltage is normal.

[0147] Similarly, PSU1_AC_OK and PSU1_PWROK, PSU2_AC_OK and PSU2_PWROK, and PSU3_AC_OK and PSU3_PWROK are respectively first indication signals indicating that the AC mains input is normal and second indication signals indicating that the output voltage is normal corresponding to PSU1, PSU2, and PSU3.

[0148] In some embodiments, the two indication signals of each PSU are isolated and enhanced by the logic gate buffer unit to generate PSUi_AC_GOOD and PSUi_PG signals (i=0, 1, 2, 3) as inputs of the signal processing unit.

[0149] Furthermore, the signal processing unit performs logic operations on the PSUi_AC_GOOD and PSUi_PG signals to generate GPU EFUSE enable control signals EN_P12V_GPUn (n=0, 1, ..., 7) to control the corresponding GPU EFUSE to be turned on or off.

[0150] In some embodiments, the signal processing unit is specifically configured to:

[0151] The PSUi_AC_GOOD and PSUi_PG indication signals are incorporated into the GPU EFUSE enable signal generation logic. When a PSU fault occurs (i.e., AC power loss or abnormal PSU output), the GPUEFUSE EN signal is set low in advance. This ensures that the GPU EFUSE EN signal (EN_P12V_GPUn) always goes low before the GPU EFUSE PG signal when the system is powered off, preventing CPLD false alarms.

[0152] The power supply status monitoring system provided in the embodiment of the present application introduces the PSUi_AC_GOOD and PSUi_PG two-way indication signals of the PSU as signals for the CPLD to detect and determine whether the PSU is in normal working condition. This can avoid the problem of false alarm GPU EFUSE logging and improve the efficiency of fault location and operation and maintenance work.

[0153] Reference Figure 6 , Figure 6 It is a schematic diagram of the internal structure of a signal processing unit provided in an embodiment of the present application.

[0154] In some implementations, the signal processing unit may include 12 groups of “AND” logic gates and 1 group of “OR” logic gates.

[0155] Among them, four groups of "AND" logic gates implement the "AND" operation of PSU_i_AC_GOOD and PSU_i_PG of the four PSUs, and output corresponding first signals, including PSU_0_SHUT, PSU_1_SHUT, PSU_2_SHUT, and PSU_3_SHUT.

[0156] After the four first signals are subjected to an OR operation by the OR logic gate, a second signal PSU_ALL_SHUT is output.

[0157] When the second signal PSU_ALL_SHUT is at a low level, it indicates that the PSU is abnormally powered off (the power off reasons include abnormal AC mains input voltage and normal AC mains input voltage but abnormal PSU output voltage).

[0158] Furthermore, another eight groups of AND logic gates are used to perform AND operations on the second signal PSU_ALL_SHUT and the control signals EN (EN_P12V_GPUi) corresponding to the eight groups of GPUEFUSEs, and output corresponding control signals EN_P12V_GPUi_new, which can be used to control the corresponding GPU EFUSE to be turned on or off.

[0159] In order to better understand the embodiments of the present application, refer to Figure 7 , Figure 7 It is the logical truth table corresponding to each input signal and output signal in the signal processing unit.

[0160] Reference Figure 8 , Figure 8 This is a power-off sequence of an AC mains power failure system provided in the embodiment of the present application. Figure 1 .

[0161] in, Figure 8The timing diagram shown is a system power-off timing diagram when the AC mains power is lost before the above-mentioned signal processing unit is introduced.

[0162] After the AC power is lost, the PSU_AC_OK signal becomes low after the t1 time interval, PSU_PWROK becomes low after the t2 time interval after the PSU_AC_OK signal becomes low, P54V_PSU loses power after the t3 time interval after the PSU_AC_OK signal becomes low, and then P12V_STBY and P3V3_STBY lose power in sequence.

[0163] At this time, CPLD is still in normal working state. When EN_P12V_GPUn is detected as high level and PG_P12V_GPUn is detected as low level (such as Figure 8 In the timing diagram shown, the dotted box marks the position), the CPLD generates a status abnormality signal FAULT#_P12V_GPUn (the status abnormality signal is low, indicating that the EFUSE: P12V_GPUn power supply is abnormal), and transmits the GPUEFUSE abnormality signal to the BMC. Among them:

[0164] t1 represents the time interval from AC power failure to the PSU_AC_OK signal becoming low; t2 represents the time interval from the PSU_AC_OK signal changing from high to low to the PSU_PWROK signal changing from high to low; t3 represents the time interval from the PSU_PWROK signal changing from high to low to the P54V_PSU power-off; t4 represents the time interval from P54V_PSU power-off to P12V_STBY power-off; t5 represents the time interval from P12V_STBY power-off to P3V3_STBY power-off; t6 represents the time interval from EN_P12V_GPUn changing from high to low to PG_P12V_GPUn changing from high to low.

[0165] Reference Figure 9 , Figure 9 This is a power-off sequence of an AC mains power failure system provided in the embodiment of the present application. Figure 2 .

[0166] in, Figure 9 The timing diagram shown is a power-off timing diagram of the system after the AC mains power failure is introduced into the above signal processing unit. In this case, the AC input is abnormal, and PSU_AC_OK goes low before PSU_PWROK.

[0167] The difference from before the above processing unit is that when the AC power is lost (PSU input voltage is abnormal), the GPUEFUSE enable signal EN_P12V_GPUn is also pulled low synchronously, and the timing of other signals remains unchanged. At this time, the CPLD will detect that EN_P12V_GPUn becomes low before PG_P12V_GPUn (such as Figure 9 Therefore, the CPLD determines that EFUSE:P12V_GPUn is a normal power-off action and does not trigger the abnormal state signal indicating that the GPU EFUSE is in an abnormal state.

[0168] Reference Figure 10 , Figure 10 This is a power-off sequence of an AC mains power failure system provided in the embodiment of the present application. Figure 3 .

[0169] in, Figure 10 The timing diagram shown is a system power-off timing diagram when the AC mains power is lost after the above signal processing unit is introduced (the PSU output voltage is abnormal, and PSU_PWROK becomes low before PSU_AC_OK).

[0170] The difference from before the introduction of the above signal processing unit is that when the PSU output is abnormal, the enable signal EN_P12V_GPUn of GPU EFUSE is also synchronously pulled low, and the timing of other signals remains unchanged. At this time, the CPLD will detect that EN_P12V_GPUn becomes low before PG_P12V_GPUn (such as Figure 10 Therefore, the CPLD determines that EFUSE:P12V_GPUn is a normal power-off action and does not trigger the abnormal state signal indicating that the GPU EFUSE is in an abnormal state.

[0171] Based on the content described in the above embodiment, the following example illustrates the implementation of a power supply false alarm prevention solution for an AI server system equipped with eight 450W GPUs.

[0172] In some embodiments, the specific implementation process of this application includes the following steps:

[0173] 1) Deploy and install 8 GPUs on the system baseboard and divide the GPUs into two groups, each containing 4 GPUs.

[0174] 2) 54V equivalent current requirement = 8 * 450 / 54 / 0.97 + 800 / 54 / 0.97 = 84A. The OCP point setting for the 54V EFUSE upstream is 1.2 * 84 = 100.8A. Therefore, the sampling resistor (Rsense) values ​​for a single EFUSE include: 1mΩ, 3W, 1%, 2512 package, quantity 2.

[0175] Among them, 1mΩ is a low resistance design and is suitable for high current scenarios (such as server power supply, motor drive, etc.).

[0176] The maximum current supported by a 3W rated power is 173A. In actual operation, a safety margin must be reserved (usually 50%-70% of the rated power is used). That is, a single resistor can stably support a current of about 86A-121A.

[0177] 1% accuracy can meet most current monitoring needs (such as overcurrent protection and power calculation).

[0178] The 2512 package size helps to spread the heat and reduce the junction temperature.

[0179] 2) The system baseboard is powered using a high-performance power cable, drawing power from the PDB connected to the 2+2 redundant PSU end (P54V_PSU).

[0180] 3) P54V EFUSE can choose 54V solution.

[0181] 4) On the system baseboard, add a 12V EFUSE to the inputs of GPU0, GPU1, GPU2, GPU3, GPU4, GPU5, GPU6, and GPU7. Based on a single GPU power of 450W, the 12V supply current is 37.5A. Connect the EN_P12V_GPUn signal of each EFUSE to the CPLD, and the PG_P12V_GPUn signal is fed back to the CPLD (n = 0, 1, ..., 7).

[0182] 6) Based on Figure 7 The truth table shown in the figure is used to implement the truth table logic operation function through the CPLD, and a logical "AND" operation is performed on the obtained PSU_ALL_SHUT signal and the EN signal of each GPU EFUSE group. The output EN_P12V_GPUn_new is then fed back to the CPLD judgment logic unit (that is, if the CPLD detects that the PG_P12V_GPUn signal becomes low before EN_P12V_GPUn_new, it determines that the GPU EFUSE is in an abnormal state; if the CPLD detects that the PG_P12V_GPUn signal becomes low after EN_P12V_GPUn_new, it determines that the GPU EFUSE is normal), and the judgment signal FAULT#_P12V_GPUn is transmitted to the input port of the signal switching unit.

[0183] 7) Add a signal switching unit between the BMC and CPLD to implement signal switching from GPIO signals to the I2C bus. Write the PG_P12V_GPUn signal (n=0, 1, …, 7) to the port register. The BMC then accesses the port register through the I2C bus to obtain the register value and parse the corresponding FAULT signal.

[0184] 8) Finally, the BMC records the parsed FAULT signal in the fault log file.

[0185] The power supply status monitoring system provided in the embodiments of the present application can achieve the following beneficial effects:

[0186] 1) Introduce the PSU_AC_OK signal of the 54V PSU as the input signal for the CPLD to detect whether the GPU EFUSE is working abnormally when the AC power is lost.

[0187] 2) Introduce the PSU_PWROK signal of the 54V PSU as the input signal for the CPLD to detect whether GPUEFUSE is in an abnormal state when the PSU output power is abnormally cut off.

[0188] 3) Introducing a PSU signal processing unit into the CPLD allows for internal logic operations to determine whether the GPU EFUSE is in an abnormal state, thus avoiding the problem of false positive GPU EFUSE logging and improving the efficiency of fault location and operation and maintenance.

[0189] In some embodiments of the present application, a server is also provided, which includes a power supply status monitoring system; the power supply status monitoring system includes at least one power supply unit, a power distribution unit, a controller and a first overcurrent protection unit; the above-mentioned power distribution unit is connected in series between the power supply unit and the first overcurrent protection unit.

[0190] The power distribution unit is used to output a first indication signal and a second indication signal. The first indication signal is used to indicate whether the input voltage of the power supply unit is within a preset input voltage range. The second indication signal is used to indicate whether the output voltage of the power supply unit is within a preset output voltage range.

[0191] The above-mentioned controller is used to update the control signal of the first overcurrent protection unit based on the first indication signal and the second indication signal, and determine whether to output a state abnormality signal according to the updated control signal. The state abnormality signal is used to indicate that the first overcurrent protection unit is in an abnormal state.

[0192] In some embodiments, the controller is further configured to obtain a monitoring signal corresponding to the first overcurrent protection unit, where the monitoring signal is configured to indicate whether the supply voltage of the first overcurrent protection unit is within a preset supply voltage range.

[0193] In some embodiments, the controller is specifically configured to determine whether to output a status abnormality signal based on the updated control signal and the monitoring signal.

[0194] In some embodiments, the controller includes a signal processing unit, which is used to perform logical operations on the first indication signal, the second indication signal, and the control signal, and update the control signal based on the operation result. The control signal is used to control the first overcurrent protection unit to be turned on or off.

[0195] In some embodiments, the signal processing unit includes at least one first logic gate, a second logic gate, and a third logic gate; the first logic gate is used to perform a first operation on the first indication signal and the second indication signal, and output a first signal; the second logic gate is used to perform a second operation based on the first signal, and output a second signal; the third logic gate is used to perform a first operation on the second signal and the control signal, and update the control signal according to the operation result.

[0196] In some embodiments, the first operation includes a logical AND operation, and the second operation includes a logical OR operation.

[0197] In some embodiments, when the fluctuation amplitude of the input voltage of the above-mentioned power supply unit is less than or equal to the first amplitude threshold, the first indication signal is a first level signal; when the fluctuation amplitude of the input voltage of the power supply unit is greater than the first amplitude threshold, the first indication signal is a second level signal.

[0198] In some embodiments, when the fluctuation amplitude of the output voltage of the power supply unit is less than or equal to the second amplitude threshold, the second indication signal is a first level signal; when the fluctuation amplitude of the output voltage of the power supply unit is greater than the second amplitude threshold, the second indication signal is a second level signal.

[0199] In some implementations, the first level signal is a high level signal, and the second level signal is a low level signal.

[0200] In some embodiments, the system further includes at least one logic gate buffer unit, which is connected in series between the power distribution unit and the controller; the logic gate buffer unit is used to perform isolation enhancement processing on the first indication signal and the second indication signal.

[0201] In some embodiments, the system further includes at least one processor; the processor is connected to the first overcurrent protection unit.

[0202] In some embodiments, the system further includes a voltage adjustment unit, which is connected in series between the power distribution unit and the first overcurrent protection unit; the voltage adjustment unit is used to convert the voltage output by the power distribution unit into a power supply voltage for the processor.

[0203] In some embodiments, the system further includes a second overcurrent protection unit, which is connected in series between the power distribution unit and the voltage adjustment unit.

[0204] In some embodiments, the system further includes a voltage conversion unit connected in series between the voltage adjustment unit and the controller; the voltage conversion unit is configured to convert the voltage output by the voltage adjustment unit into a power supply voltage for the controller.

[0205] In some embodiments, the system further includes a third overcurrent protection unit, which is connected in series between the voltage adjustment unit and the voltage conversion unit.

[0206] In some embodiments, the system further includes a baseboard management controller connected to the baseboard management controller; the controller is configured to send an abnormal status signal to the baseboard management controller; and the baseboard management controller is configured to output a log file based on the abnormal status signal.

[0207] In some embodiments, the system further includes a signal switching unit, which is connected in series between the baseboard management controller and the controller; the controller is used to send a status abnormality signal to the signal switching unit; and the signal switching unit is used to save the status abnormality signal in a corresponding port register.

[0208] In some embodiments, the baseboard management controller is used to address the serial bus address of the signal switching unit via the serial bus to obtain the value in the port register, which is used to indicate whether the first overcurrent protection unit is in an abnormal state.

[0209] In some embodiments, the first overcurrent protection unit includes an electronic fuse.

[0210] It should be noted that the specific structure and implementation principle of the above power supply status monitoring system can refer to Figures 2 to 7 The embodiments shown are not described in detail here.

[0211] The server provided in an embodiment of the present application includes a power supply status monitoring system. The power supply status monitoring system introduces two signals from the power supply unit as the basis for the controller to monitor and determine whether the first overcurrent protection unit is working abnormally. This can effectively prevent the occurrence of false alarms of the status of the first overcurrent protection unit due to PSU abnormalities, thereby effectively improving the operation and maintenance efficiency of the server.

[0212] Optionally, the above-mentioned power supply status monitoring system can also be applied to switches, storage and other products with 54V power supply, which will not be described in detail in the embodiments of this application.

[0213] In the description of the embodiments of this application, it should be noted that, unless otherwise expressly specified or limited, the terms "coupled" and "connected" should be understood in a broad sense. For example, they may refer to a fixed connection, an indirect connection via an intermediate medium, internal communication between two components, or an interaction between two components. Those skilled in the art will understand the specific meanings of the above terms in the embodiments of this application based on specific circumstances.

[0214] It is understood that the division of modules within the computing device described above is merely a division of logical functions. Each function may correspond to a functional module, or two or more functions may be integrated into a single functional module. In actual implementation, all or some of the modules may be integrated into a single physical entity or distributed across different physical entities.

[0215] The above specific implementation methods further illustrate the purpose, technical solutions and beneficial effects of this application in detail. It should be understood that the above are only specific implementation methods of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of this application should be included in the scope of protection of this application.

[0216] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present application, rather than to limit them. Although the embodiments of the present application have been described in detail with reference to the above embodiments, ordinary technicians in this field should understand that they can still modify the technical solutions recorded in the above embodiments, or replace some or all of the technical features therein with equivalents. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the embodiments of the present application.

Claims

1. A power supply status monitoring system, characterized in that: The system comprises at least one power supply unit, a power distribution unit, a controller and a first overcurrent protection unit; the power distribution unit is connected in series between the power supply unit and the first overcurrent protection unit; The power distribution unit is configured to output a first indication signal and a second indication signal, wherein the first indication signal is configured to indicate whether the input voltage of the power supply unit is within a preset input voltage range, and the second indication signal is configured to indicate whether the output voltage of the power supply unit is within a preset output voltage range; The controller includes a signal processing unit, which is used to perform a logical operation on the first indication signal, the second indication signal, and the control signal of the first overcurrent protection unit, and update the control signal of the first overcurrent protection unit based on the operation result; wherein the signal processing unit includes at least one first logic gate, a second logic gate, and a third logic gate; the first logic gate is used to perform a first operation on the first indication signal and the second indication signal, and output a first signal; the second logic gate is used to perform a second operation based on the first signal, and output a second signal; the third logic gate is used to perform a first operation on the second signal and the control signal, and update the control signal according to the operation result; the first operation includes a logical AND operation, and the second operation includes a logical OR operation; The controller is used to determine whether to output a state abnormality signal according to the updated control signal; the state abnormality signal is used to indicate that the first overcurrent protection unit is in an abnormal state.

2. The system according to claim 1, wherein: The controller is also used for: A monitoring signal corresponding to the first overcurrent protection unit is obtained, where the monitoring signal is used to indicate whether the supply voltage of the first overcurrent protection unit is within a preset supply voltage range.

3. The system according to claim 2, characterized in that The controller is specifically used for: According to the updated control signal and the monitoring signal, it is determined whether to output the abnormal state signal.

4. The system according to claim 1, wherein: The control signal is used to control the first overcurrent protection unit to be turned on or off.

5. The system according to claim 4, characterized in that When the fluctuation amplitude of the input voltage of the power supply unit is less than or equal to the first amplitude threshold, the first indication signal is a first level signal; When the fluctuation amplitude of the input voltage of the power supply unit is greater than the first amplitude threshold, the first indication signal is a second level signal.

6. The system according to claim 4, characterized in that When the fluctuation amplitude of the output voltage of the power supply unit is less than or equal to the second amplitude threshold, the second indication signal is a first level signal; When the fluctuation amplitude of the output voltage of the power supply unit is greater than the second amplitude threshold, the second indication signal is a second level signal.

7. The system according to claim 5 or 6, characterized in that The first level signal is a high level signal, and the second level signal is a low level signal.

8. The system according to claim 1, wherein: The system further comprises at least one logic gate buffer unit, wherein the logic gate buffer unit is connected in series between the power distribution unit and the controller; The logic gate buffer unit is used to perform isolation enhancement processing on the first indication signal and the second indication signal.

9. The system according to claim 1, wherein: The system further includes at least one processor; the processor is connected to the first overcurrent protection unit.

10. The system according to claim 9, characterized in that The system further includes a voltage adjustment unit, wherein the voltage adjustment unit is connected in series between the power distribution unit and the first overcurrent protection unit; The voltage adjustment unit is used to convert the voltage output by the power distribution unit into the power supply voltage of the processor.

11. The system according to claim 10, wherein: The system further includes a second overcurrent protection unit connected in series between the power distribution unit and the voltage adjustment unit.

12. The system according to claim 10, wherein: The system further includes a voltage conversion unit, wherein the voltage conversion unit is connected in series between the voltage adjustment unit and the controller; The voltage conversion unit is used to convert the voltage output by the voltage adjustment unit into a power supply voltage for the controller.

13. The system according to claim 12, wherein: The system further includes a third overcurrent protection unit connected in series between the voltage adjustment unit and the voltage conversion unit.

14. The system according to claim 1, wherein: The system further comprises a baseboard management controller, the controller being connected to the baseboard management controller; The controller is used to send the abnormal status signal to the baseboard management controller; The baseboard management controller is configured to output a log file based on the abnormal status signal.

15. The system according to claim 14, wherein: The system further comprises a signal switching unit, wherein the signal switching unit is connected in series between the baseboard management controller and the controller; The controller is used to send the abnormal state signal to the signal switching unit; The signal switching unit is used to store the abnormal status signal in a corresponding port register.

16. The system according to claim 15, wherein: The baseboard management controller is used to address the serial bus address of the signal switching unit through a serial bus to obtain a value in the port register, where the value is used to indicate whether the first overcurrent protection unit is in an abnormal state.

17. The system according to claim 1, wherein: The first overcurrent protection unit includes an electronic fuse.

18. A server, characterized in that: The server includes the power supply status monitoring system according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • Over-voltage and over-current hardware protection circuit and DC power supply circuit

    CN101867177A

  • Intelligent power distributor control circuit

    CN114825908A