Server failure field data freezing system, method and server
By introducing an isolated energy storage unit, a data acquisition unit, and a non-volatile high-speed storage unit into the server, the system architecture solves the problem of data loss during sudden power failures of the server, and achieves reliable retention of critical data and improved accuracy of fault location.
Patent Information
- Application Number
- CN202611124165.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-28
- Publication Date
- 2026-08-25
AI Technical Summary
In existing technologies, when a server experiences a sudden power failure, the BMC and CPU lose power instantly, causing critical operational data to be lost in a timely manner, resulting in the loss of fault scene data and affecting the accuracy of fault location and analysis efficiency.
The system architecture adopts an isolated energy storage unit combined with a non-volatile high-speed storage unit and an independent logic control unit. When the main power supply is normal, data is buffered in a loop. When the power status indication signal is detected to be invalid or the bus voltage drops below the threshold, a freeze trigger signal is generated and the data is transferred to the non-volatile high-speed storage unit.
In the event of a sudden power outage, it can reliably retain critical operational status data, improve the accuracy of fault location and analysis efficiency, and prevent data loss.
Smart Images

Figure CN122633472A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of server hardware architecture and fault diagnosis technology, specifically to a server fault field data freezing system, method and server. Background Technology
[0002] With the rapid development of artificial intelligence, big data, and cloud computing technologies, data centers are placing increasingly higher demands on server reliability. During prolonged periods of high-load operation, a sudden hardware failure can lead to extended service interruptions if the root cause cannot be quickly identified, impacting the continuity and stability of data services. Therefore, post-fault analysis and precise fault location have become crucial for ensuring the stable operation of data centers, and a complete record of on-site operational data at the time of the fault is an essential prerequisite for accurate fault location.
[0003] In related technologies, the server industry commonly uses a Baseboard Management Controller (BMC) combined with a Central Processing Unit (CPU) logging system to monitor system status and record faults. During normal server operation, the BMC, powered by the motherboard power supply, collects telemetry data such as voltage, current, and temperature across various motherboard circuits at sampling periods of seconds or milliseconds, storing the logs in onboard non-volatile storage. Simultaneously, operating system-level error logs are written to the hard drive for later review by maintenance personnel. When a system anomaly occurs, the BMC or CPU attempts to record the current operating status information to non-volatile storage. However, when a sudden power failure occurs, the BMC and CPU may lose power instantaneously, preventing the timely saving of critical operational data and leading to data loss at the fault location. Summary of the Invention
[0004] This application provides a server fault scene data freezing system, method, and server to improve the problem in related technologies where, when a server experiences a sudden power failure, critical operating data cannot be saved in time due to the instantaneous power loss of the BMC and CPU, resulting in the loss of fault scene data.
[0005] In a first aspect, this application provides a server failure on-site data freezing system, comprising:
[0006] The isolated energy storage unit is connected to the server's main power bus through an isolation device. It is used to store electrical energy when the main power bus is powered normally and to release electrical energy to provide emergency working power when the main power bus fails to power.
[0007] The data acquisition unit is connected to the monitoring pins of the voltage regulator module (VRM) and / or power supply unit (PSU) on the server motherboard to collect real-time operating status data.
[0008] Non-volatile high-speed memory cell, the power supply terminal of the non-volatile high-speed memory cell is connected to the output terminal of the isolated energy storage cell;
[0009] The independent logic control unit is equipped with a power supply terminal, a detection input terminal, a data input terminal, and a data output terminal. The power supply terminal is connected to the output terminal of the isolated energy storage unit, the detection input terminal is connected to the power status indicator terminal and / or the bus voltage sampling terminal of the main power bus, the data input terminal is connected to the output terminal of the data acquisition unit, and the data output terminal is connected to the non-volatile high-speed storage unit. The independent logic control unit is used to cyclically write the operating status data into the internal buffer in normal mode. When the power status indicator signal is found to be invalid or the bus voltage drops beyond a preset threshold through the detection input terminal, a freeze trigger signal is generated to transfer the fault field data in the internal buffer to the non-volatile high-speed storage unit.
[0010] In one possible embodiment, the isolated energy storage unit includes: a supercapacitor array for storing electrical energy and releasing it when the main power bus fails to supply power; a charging control module, the input of which is connected to the main power bus and the output of which is connected to the supercapacitor array, for managing the charging of the supercapacitor array when the main power bus is supplying power normally; and an isolation device, connected in series between the supercapacitor array and the main power bus, for preventing the electrical energy of the supercapacitor array from flowing back to the main power bus when the main power bus fails to supply power, and for providing an independent power supply circuit for the independent logic control unit and the non-volatile high-speed storage unit.
[0011] In one possible embodiment, the isolation device is an ideal diode or a reverse-current protection diode.
[0012] In one possible embodiment, the independent logic control unit is a Complex Programmable Logic Device (CPLD) or a Field-Programmable Gate Array (FPGA).
[0013] In one possible embodiment, the non-volatile high-speed storage cell is a ferroelectric random access memory (FRAM).
[0014] In one possible embodiment, the data acquisition unit includes an analog-to-digital converter (ADC) and a digital interface. The analog input of the ADC is connected to the voltage sampling terminal, current sampling terminal, and temperature sampling terminal of the VRM and / or PSU for acquiring voltage, current, and temperature data of the VRM and / or PSU. The digital interface is connected to the status code output terminal of the VRM and / or PSU for reading the status code of the internal register of the VRM and / or PSU.
[0015] Secondly, this application provides a method for freezing server fault scene data, applied to an independent logic control unit in a server fault scene data freezing system provided above. The server fault scene data freezing method includes:
[0016] Under normal server conditions, acquire the operating status data of VRM and / or PSU on the server motherboard, and write the operating status data into the internal cache of the independent logic control unit in a circular cache manner;
[0017] Obtain the power status indication signal and / or bus voltage of the server's main power bus, and determine whether there is a power supply abnormality in the main power bus;
[0018] When a power supply anomaly occurs on the main power bus, a freeze trigger signal is generated to transfer the fault scene data stored in the internal cache to the non-volatile high-speed storage unit in the server fault scene data freeze system. In the event of a power supply failure on the main power bus, the isolated energy storage unit in the server fault scene data freeze system provides emergency working power to the independent logic control unit and the non-volatile high-speed storage unit.
[0019] In one possible embodiment, determining whether a power supply abnormality has occurred on the main power bus includes: determining that a power supply abnormality has occurred on the main power bus when the power status indication signal changes from an active level to an inactive level; and / or determining that a power supply abnormality has occurred on the main power bus when the voltage drop exceeds a preset threshold or the voltage drop slope exceeds a preset threshold.
[0020] In one possible embodiment, the server fault site data freezing method further includes: after the non-volatile high-speed storage unit completes data transfer, sending a power-off indication signal to the isolation energy storage unit to cause the isolation energy storage unit to stop supplying power to the independent logic control unit and the non-volatile high-speed storage unit.
[0021] Thirdly, this application provides a server, including a motherboard, and a server fault field data freezing system provided above, which is disposed on the motherboard.
[0022] The server fault field data freezing system, method, and server provided in this application include: an isolated energy storage unit connected to the server's main power bus via an isolation device, used to store electrical energy when the main power bus is powered normally and to release electrical energy to provide emergency working power when the main power bus fails; a data acquisition unit connected to the monitoring pins of the VRM and / or PSU on the server motherboard, used to acquire real-time operating status data; a non-volatile high-speed storage unit, the power supply terminal of which is connected to the output terminal of the isolated energy storage unit; and an independent logic control unit, equipped with a power supply terminal and detection input. The system includes a power supply terminal, a data input terminal, and a data output terminal. The power supply terminal is connected to the output terminal of the isolated energy storage unit. The detection input terminal is connected to the power status indicator terminal and / or the bus voltage sampling terminal of the main power bus. The data input terminal is connected to the output terminal of the data acquisition unit. The data output terminal is connected to the non-volatile high-speed storage unit. An independent logic control unit is used to cyclically write the operating status data to the internal buffer in normal mode. When the power status indicator signal is invalid or the bus voltage drops beyond a preset threshold, a freeze trigger signal is generated to transfer the fault field data in the internal buffer to the non-volatile high-speed storage unit.
[0023] This application establishes an isolated energy storage unit connected to the server's main power bus via an isolation device. This unit stores electrical energy when the main power bus is powered normally and provides emergency power to the independent logic control unit and non-volatile high-speed storage unit when the main power bus fails. Furthermore, a data acquisition unit collects real-time operating status data from the VRM and / or PSU on the server motherboard. The independent logic control unit cyclically caches this operating status data in normal mode and generates a freeze trigger signal when an invalid power status indication signal is detected or the bus voltage drops below a preset threshold. This transfers the internally cached fault data to the non-volatile high-speed storage unit, reliably preserving critical operating status data during sudden power outages, thereby improving the accuracy of server fault location and the efficiency of root cause analysis. Attached Figure Description
[0024] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0025] Figure 1 A schematic diagram of a server fault field data freezing system provided as an exemplary embodiment of this application;
[0026] Figure 2 A flowchart illustrating a method for freezing server fault scene data as provided in an exemplary embodiment of this application;
[0027] Figure 3A schematic diagram of the structure of a server provided for an exemplary embodiment of this application.
[0028] In the diagram, 10—Server fault field data freezing system; 11—Isolation energy storage unit; 111—Isolation device; 12—Data acquisition unit; 13—Non-volatile high-speed storage unit; 14—Independent logic control unit; 15—Main power bus; 16—VRM; 17—PSU.
[0029] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0030] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application.
[0031] The terms “first,” “second,” etc., used in the specification and claims of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, products, or apparatus.
[0032] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0033] In related technologies, when using a BMC combined with a CPU logging system to achieve system status monitoring and fault recording, a basic operating log can be generated under normal power supply conditions, which has a certain recording function for general alarms. However, when there is a sudden PSU failure, external power outage, or rapid drop in bus voltage, the data acquisition unit and write link still rely on the main power supply and often stop working before the voltage fails, resulting in the inability to retain key field data before and after the fault. In addition, the periodic sampling method is difficult to cover the rapid changes in voltage, current, and status signals in a short period of time, and the storage write time is also difficult to adapt to sudden power outage scenarios. As a result, the fault log often only reflects coarse-grained conclusions such as "power outage" or "power supply abnormality," making it difficult to reconstruct the field state before the fault and effectively distinguish whether it is a PSU-side failure, VRM abnormality, or a chain problem induced by the load side, thereby reducing the accuracy of fault root cause analysis and increasing troubleshooting time.
[0034] In view of this, how to reliably retain critical operational status data in the event of a sudden power failure has become an urgent technical problem to be solved. To address this issue, this application provides a server fault field data freezing system. This system comprises an isolated energy storage unit, a data acquisition unit, a non-volatile high-speed storage unit, and an independent logic control unit, forming a cooperative system architecture. When the main power supply is normal, operational status data is cyclically cached. When an invalid power status indication signal is detected or the bus voltage drops beyond a preset threshold, freezing is triggered, and the fault field data is transferred to the non-volatile high-speed storage unit. This improves data retention capability and fault location accuracy under sudden power failures.
[0035] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0036] Figure 1 A schematic diagram of a server fault field data freezing system provided as an exemplary embodiment of this application. Figure 1 As shown, the server fault site data freezing system 10 includes: an isolated energy storage unit 11, a data acquisition unit 12, a non-volatile high-speed storage unit 13, and an independent logic control unit 14, wherein:
[0037] The isolated energy storage unit 11 is connected to the main power bus 15 of the server through the isolation device 111. It is used to store electrical energy when the main power bus 15 is powered normally and to release electrical energy to provide emergency working power when the main power bus 15 fails to power.
[0038] The data acquisition unit 12 is connected to the monitoring pins of VRM16 and / or PSU17 on the server motherboard and is used to collect real-time operating status data.
[0039] The power supply terminal of the non-volatile high-speed storage cell 13 is connected to the output terminal of the isolated energy storage cell 11;
[0040] The independent logic control unit 14 is equipped with a power supply terminal, a detection input terminal, a data input terminal, and a data output terminal. The power supply terminal is connected to the output terminal of the isolated energy storage unit 11. The detection input terminal is connected to the power status indicator terminal and / or the bus voltage sampling terminal of the main power bus 15. The data input terminal is connected to the output terminal of the data acquisition unit 12. The data output terminal is connected to the non-volatile high-speed storage unit 13. The independent logic control unit 14 is used to cyclically write the operating status data into the internal cache in normal mode. When the power status indicator signal is invalid or the bus voltage drops beyond a preset threshold detected by the detection input terminal, a freeze trigger signal is generated to transfer the fault field data in the internal cache to the non-volatile high-speed storage unit 13.
[0041] The isolated energy storage unit 11 is connected to the server's main power bus 15 via an isolation device 111. The main power bus 15 is the DC power supply line (e.g., a 12V or 48V bus) output from the PSU 17 in the server, used to provide main power to the various power modules on the server motherboard. The isolated energy storage unit 11 stores electrical energy when the main power bus 15 is powered normally and releases electrical energy when the main power bus 15 fails to supply power, providing emergency working power. The isolation device 111 ensures that current flows unidirectionally from the main power bus 15 side to the isolated energy storage unit 11 side, preventing the electrical energy stored in the isolated energy storage unit 11 from flowing back to the main power bus 15 when the main power bus 15 fails to supply power.
[0042] The data acquisition unit 12 is connected to the monitoring pins of VRM16 and / or PSU17 on the server motherboard. VRM16 is a power supply module on the server motherboard that converts the main power bus voltage (e.g., 12V) into the low-voltage, high-current (e.g., 0.8V–1.8V) required by core chips such as the CPU or GPU. PSU17 is a hardware device that converts external AC power (e.g., 220V AC) into the DC voltage (e.g., 12V / 48V) required by the server. The data acquisition unit 12 is used to acquire real-time operating status data of VRM16 and / or PSU17, including analog telemetry data such as output voltage, output current, and temperature, as well as digital data such as internal register status codes.
[0043] The non-volatile high-speed storage unit 13 has its power supply terminal connected to the output terminal of the isolated energy storage unit 11. The non-volatile high-speed storage unit 13 is a storage device with non-volatility and high-speed write characteristics. It has the characteristics of data retention after power failure, fast write speed (nanosecond level), and can be written without erasure. It is used to permanently save the fault scene data in the event of a power failure.
[0044] The independent logic control unit 14 is equipped with a power supply terminal, a detection input terminal, a data input terminal, and a data output terminal. Its power supply terminal is connected to the output terminal of the isolated energy storage unit 11, and is independently powered by the isolated energy storage unit 11. The detection input terminal is connected to the power status indicator terminal and / or the bus voltage sampling terminal of the main power bus 15. The power status indicator terminal is the port where the main power bus 15 outputs a power status indicator signal (i.e., the Pwr_Good signal). When the main power bus 15 is powered normally, this signal is at an effective level (e.g., high level); when the main power bus 15 fails to supply power or the voltage drops below a preset threshold, this signal becomes at an ineffective level (e.g., low level). The bus voltage sampling terminal is used to collect data from the main power bus. The voltage value of line 15 is used by the independent logic control unit 14 to determine whether the bus voltage has dropped; the data input terminal is connected to the output terminal of the data acquisition unit 12 to receive operating status data; the data output terminal is connected to the data interface of the non-volatile high-speed storage unit 13 to transfer the fault field data to the non-volatile high-speed storage unit 13 when a fault occurs; the independent logic control unit 14 is a hardware logic controller implemented by a programmable logic device, which has the characteristics of hardware-level parallel processing and deterministic response delay, and can respond to power failures and complete data transfer operations within microseconds.
[0045] The independent logic control unit 14 is used to cyclically write operating status data to the internal cache in normal mode. Normal mode refers to the working mode when the main power bus 15 is powered normally, and the server is in normal operating state. The internal cache is the random access memory (RAM) or register array inside the independent logic control unit 14, which is used to temporarily store the operating status data within a recent period of time. Cyclic writing means continuously collecting operating status data according to a preset sampling period and updating the internal cache in a first-in-first-out manner, so that the internal cache always retains only the data within the most recent preset time period, such as 100ms.
[0046] The independent logic control unit 14 is also used to generate a freeze trigger signal when the power status indication signal is invalid or the bus voltage drops below a preset threshold is detected through the detection input terminal, so as to transfer the fault scene data in the internal cache to the non-volatile high-speed storage unit 13. Specifically, an invalid power status indication signal means that the Pwr_Good signal output by the main power bus 15 changes from an effective level to an invalid level, indicating that the power supply of the main power bus 15 has become abnormal; a bus voltage drop means that the voltage value of the main power bus 15 drops rapidly in a short period of time. The independent logic control unit 14 detects whether a bus voltage drop has occurred by judging whether the absolute value of the bus voltage is lower than a preset voltage threshold, or by judging whether the slope of the bus voltage drop (i.e., the rate of change of voltage per unit time) exceeds a preset slope threshold. When any of the above triggering conditions is detected, the independent logic control unit 14 immediately generates a freeze trigger signal, stops updating the internal cache, and writes the fault scene data (i.e., key operating status data and register snapshots in the most recent period before the power failure) stored in the internal cache to the non-volatile high-speed storage unit 13 with the highest priority via the Serial Peripheral Interface (SPI) bus or the Inter-Integrated Circuit (I2C) bus, thereby permanently "freezing" the fault scene data. It should be noted that the above-mentioned preset voltage threshold and preset slope threshold can be flexibly set by those skilled in the art according to the actual specifications and power supply capacity of the server main power bus. For example, the voltage threshold can be determined based on the nominal output voltage value of the PSU and its allowable fluctuation range, and the slope threshold can be determined based on the bus capacitance and load change characteristics. This application embodiment does not specifically limit this.
[0047] like Figure 1 As shown, the isolation energy storage unit 11, data acquisition unit 12, non-volatile high-speed storage unit 13, and independent logic control unit 14 together constitute an emergency data recording subsystem physically isolated from the main power domain. The "main power domain" refers to the area powered by the main power bus 15, including VRM16, PSU17, and other power modules on the server motherboard; the "isolated emergency domain" refers to the area independently powered by the isolation energy storage unit 11, including the independent logic control unit 14 and the non-volatile high-speed storage unit 13. When the main power bus 15 fails to supply power, the power modules in the main power domain (such as the CPU and BMC) stop working due to power loss, while the independent logic control unit 14 and the non-volatile high-speed storage unit 13 in the isolated emergency domain are still independently powered by the isolation energy storage unit 11, enabling them to continue collecting, freezing, and transferring fault site data.
[0048] The server fault field data freezing system provided in this application uses an isolated energy storage unit connected to the server's main power bus via an isolation device. This unit stores energy when the main power bus is supplying power normally and provides emergency power to the independent logic control unit and non-volatile high-speed storage unit when the main power bus fails. A data acquisition unit collects real-time operating status data from the VRM and / or PSU on the server motherboard. The independent logic control unit cyclically caches this operating status data in normal mode and generates a freeze trigger signal when it detects an invalid power status indicator signal or a bus voltage drop exceeding a preset threshold. This transfers the internally cached fault field data to the non-volatile high-speed storage unit, reliably preserving critical operating status data during sudden power anomalies, thereby improving the accuracy of server fault location and the efficiency of root cause analysis.
[0049] In some embodiments, the isolated energy storage unit includes: a supercapacitor array for storing electrical energy and releasing it when the main power bus fails to supply power; a charging control module, the input of which is connected to the main power bus and the output of which is connected to the supercapacitor array, for managing the charging of the supercapacitor array when the main power bus is supplying power normally; and an isolation device, connected in series between the supercapacitor array and the main power bus, for preventing the electrical energy of the supercapacitor array from flowing back to the main power bus when the main power bus fails to supply power, and for providing an independent power supply circuit for the independent logic control unit and the non-volatile high-speed storage unit.
[0050] The supercapacitor array, composed of multiple supercapacitors connected in series and / or parallel, is used to store electrical energy when the main power bus is supplying power normally and to release electrical energy when the main power bus fails, thus providing emergency working power. Supercapacitors possess characteristics such as high power density, fast charging and discharging speed, and long cycle life, enabling them to provide stable power output for a short period after the main power bus fails. Specific implementations of the supercapacitor array can utilize multiple individual supercapacitors connected in series to increase their voltage rating and in parallel to increase their capacity. Those skilled in the art can flexibly configure the array according to the required output voltage level and capacity of the isolated energy storage unit; this application does not impose specific limitations in this regard.
[0051] A charging control module, with its input connected to the main power bus and its output connected to the supercapacitor array, is used to manage the charging of the supercapacitor array when the main power bus is supplying power normally. The charging control module can be implemented using charging management chips or circuits known in the art, and its specific functions include, but are not limited to: constant current charging, constant voltage charging, or float charging management of the supercapacitor array; monitoring the voltage and current of the supercapacitor array to prevent overcharging or over-discharging; automatically stopping charging or switching to float charging mode after the supercapacitor array is fully charged; real-time monitoring of the health status of the supercapacitor array (such as capacitance decay, internal resistance changes, etc.), and outputting an alarm indication through a status signal terminal when the energy storage capacity drops below a preset threshold. In some embodiments, the charging control module also has a charging status output terminal, connected to the detection input terminal of an independent logic control unit, for outputting a charging status signal (such as charging, fully charged, fault, etc.) to the independent logic control unit so that the independent logic control unit can know the real-time status of the isolated energy storage unit.
[0052] An isolation device, connected in series between the supercapacitor array and the main power bus, allows current to flow from the main power bus side to the supercapacitor array side when the main power bus is supplying power normally, providing a path for the main power bus to charge the supercapacitor array; and prevents the power of the supercapacitor array from flowing back to the main power bus when the main power bus fails, thereby ensuring that the power stored in the supercapacitor array is supplied only to the independent logic control unit and the non-volatile high-speed memory unit, reducing the ineffective consumption of power on the main power bus side or causing abnormal impact on the main power domain after the failure.
[0053] Based on the above structure, when the main power bus is supplying power normally, the power supplied by the main power bus powers the various power modules (such as VRM, PSU, CPU, and BMC) within the main power domain in one path, and charges the supercapacitor array through the isolation device and charging control module in the other path. When the main power bus fails to supply power, the voltage on the main power bus side drops rapidly, and the isolation device immediately shuts off the discharge path of the supercapacitor array to the main power bus. The power stored in the supercapacitor array is supplied only to the independent logic control unit and the non-volatile high-speed storage unit through its output terminal, ensuring that the devices within the isolation emergency domain can continue to operate normally for a short period of time after the fault occurs (e.g., 500ms to 2s), and complete the collection, freezing, and transfer of fault site data. This dual-domain isolated power supply architecture, consisting of a "main power domain" and an "isolated emergency domain," ensures that even if the server's main power supply fails completely, or even if the power supply module (PSU) suffers physical damage (such as a system crash), the isolated energy storage unit can still independently provide stable and clean emergency power to the data freezing system at the fault site. This effectively improves the problem of critical data loss caused by the instantaneous shutdown of the BMC and CPU due to a main power failure in existing technologies.
[0054] In this embodiment, through the coordinated operation of a supercapacitor array, a charging control module, and isolation devices, the isolated energy storage unit automatically completes energy storage preparation when the server's main power bus is powered normally. When the main power bus fails, it can automatically provide emergency power within microseconds without any software intervention. This ensures that the independent logic control unit and the non-volatile high-speed storage unit still have sufficient power to complete the complete acquisition, freezing, and transfer of fault site data after the main power supply fails completely. This effectively improves the data loss problem caused by the BMC and CPU, which rely on the main power supply, stopping working the instant of power failure in related technologies, and thus provides complete data support for server fault diagnosis.
[0055] In some embodiments, the isolation device is an ideal diode or a reverse-current protection diode.
[0056] In one implementation, the isolation device is implemented using an anti-reverse current diode. This anti-reverse current diode utilizes the unidirectional conductivity of a PN junction, allowing current to flow only from the anode to the cathode, and is reverse-biased. Specifically, in this embodiment, the anode of the anti-reverse current diode is connected to the main power bus, and the cathode is connected to the supercapacitor array. When the main power bus is supplying power normally, the bus voltage is higher than the voltage of the supercapacitor array, and the anti-reverse current diode is forward biased, allowing current to flow from the main power bus side through the anti-reverse current diode to the supercapacitor array side, providing a path for the main power bus to charge the supercapacitor array. When the main power bus fails to supply power, the bus voltage rapidly drops below the voltage of the supercapacitor array, and the anti-reverse current diode is reverse biased and cut off. The energy stored in the supercapacitor array is prevented from flowing back to the main power bus side, thus ensuring that energy is supplied only to independent logic control units and non-volatile high-speed memory units. The anti-reverse current diode has advantages such as simple structure, high reliability, no need for control circuits, and low cost, making it suitable for scenarios where the on-state voltage drop requirement is not strict.
[0057] In another implementation, the isolation device is implemented using an ideal diode. The ideal diode refers to an active rectifier circuit composed of a power metal-oxide-semiconductor field-effect transistor (MOSFET) and an ideal diode controller. The power MOSFET is connected in series between the supercapacitor array and the main power bus. The ideal diode controller is connected across the drain and source of the power MOSFET to detect the voltage difference across the power MOSFET and control the switching on and off of the power MOSFET based on the detection result. Specifically, in this embodiment, the ideal diode controller continuously detects the voltage difference between the source and drain of the power MOSFET; when it detects current flowing from the main power bus side to the supercapacitor array side, the ideal diode controller controls the power MOSFET to turn on, providing a low-resistance path for the main power bus to charge the supercapacitor array; when it detects current flowing from the supercapacitor array side to the main power bus side (i.e., when the main power bus fails to supply power), the ideal diode controller immediately controls the power MOSFET to turn off, blocking the discharge path from the supercapacitor array to the main power bus. Compared to anti-reverse current diodes, ideal diodes offer advantages such as extremely low forward voltage drop (typically tens of millivolts, far lower than the 0.3V~0.7V of anti-reverse current diodes) and minimal reverse leakage current. This significantly reduces conduction losses when the main power bus is operating normally, preventing the anti-reverse current diodes from overheating due to excessive voltage drop under high current conditions, making them particularly suitable for high-current applications. Furthermore, the ideal diode controller can detect reverse current and turn off the power MOSFET within microseconds, effectively preventing energy from the supercapacitor array from flowing back into the main power bus.
[0058] Those skilled in the art can choose to use either a reverse-current protection diode or an ideal diode to implement the isolation device, depending on the specific requirements of the actual application scenario (such as system power level, efficiency requirements, cost control, size limitations, etc.). For example, in scenarios with lower power levels and less sensitivity to on-state voltage drop, a reverse-current protection diode can be used to reduce costs and simplify the design; in scenarios with high current and high efficiency requirements, an ideal diode can be used to reduce conduction losses and heat dissipation.
[0059] In this embodiment, by employing anti-backflow diodes or ideal diodes to implement isolation devices, the discharge path from the supercapacitor array to the main power bus can be quickly blocked when the main power bus fails. This ensures that the stored energy is supplied only to independent logic control units and non-volatile high-speed storage units, reducing ineffective energy consumption and abnormal impacts. It effectively achieves physical isolation between the main power domain and the isolated emergency domain, providing a guarantee for the isolated energy storage unit to independently and reliably provide emergency operating power. The use of anti-backflow diodes has the advantages of simple structure and low cost; the use of ideal diodes has the advantages of low forward voltage drop, low conduction loss, and low heat generation, making them suitable for high-current scenarios. In practical applications, the appropriate diode can be flexibly selected according to actual needs.
[0060] In some embodiments, the independent logic control unit is a CPLD or an FPGA.
[0061] In one implementation, the independent logic control unit is implemented using a CPLD. The CPLD features power-on operation, deterministic response delay, and moderate logic resources, making it suitable for executing real-time control logic with high deterministic requirements. In this embodiment, the CPLD integrates power status monitoring logic, data acquisition control logic, ring buffer management logic, and data transfer control logic. Specifically, the CPLD uses its general-purpose input / output (GPIO) ports to form detection input terminals, directly connected to the power status indicator terminal and / or bus voltage sampling terminal of the main power bus, for real-time monitoring of the Pwr_Good signal and bus voltage value. The CPLD internally includes hardware comparator logic or voltage detection logic, capable of determining within microseconds whether the Pwr_Good signal has changed from an effective level to an ineffective level, or whether the bus voltage has dropped below a preset threshold. The CPLD uses its GPIO ports or parallel bus interface to form data input terminals, connected to the output terminal of the data acquisition unit, for receiving the operating status data of the VRM and / or PSU acquired by the data acquisition unit. The CPLD integrates cache management logic (i.e., ring buffer control logic). In normal mode, the cache management logic continuously receives operating status data according to a preset sampling period and writes it to the CPLD's internal RAM in a first-in-first-out manner, retaining only data from the most recent preset time period. The CPLD also integrates fault triggering logic. When the power status monitoring logic detects an invalid Pwr_Good signal or a bus voltage drop exceeding a preset threshold, the fault triggering logic immediately generates a freeze trigger signal. The cache management logic responds to this freeze trigger signal by ceasing to update the internal RAM. Simultaneously, the data transfer control logic transfers the fault scene data stored in the internal RAM to a non-volatile high-speed memory unit via the SPI bus or I2C bus. Because the CPLD uses hardware logic to implement all the above functions, the entire detection, triggering, and transfer process is completely independent of the server's host operating system (OS) and BMC, and does not rely on the execution of any software programs. Its response latency is determined by the transmission delay of the hardware gate circuits, exhibiting determinism and predictability, and ensuring fault detection and triggering response are completed within microseconds.
[0062] In another implementation, the independent logic control unit is implemented using an FPGA. FPGAs are characterized by abundant logic resources, strong reconfigurability, and outstanding parallel processing capabilities, making them suitable for scenarios requiring complex logic processing or later upgrades and optimizations. Similar to CPLDs, FPGAs also integrate power status monitoring logic, data acquisition control logic, ring buffer management logic, and data transfer control logic. The connection relationships and working principles of their functional modules are basically the same as those of CPLD implementations. Specifically, the FPGA uses its GPIO ports to form detection input terminals and data input terminals, which are connected to the power status indicator terminal and / or bus voltage sampling terminal of the main power bus and the output terminal of the data acquisition unit, respectively. A ring buffer is implemented through its internal RAM or externally connected RAM. Fault trigger judgment and data transfer control are implemented through its internal logic. It is connected to a non-volatile high-speed memory unit via an SPI bus or I2C bus interface. Unlike CPLDs, FPGAs typically need to load configuration data from external non-volatile memory (such as Flash) upon power-up to complete logic configuration; the configuration loading time is typically in the millisecond range. However, once configured, the hardware logic inside the FPGA begins to run in parallel, and its response latency is also determined by the propagation delay of the hardware gates, exhibiting determinism and predictability. Therefore, during normal server operation, the FPGA, like the CPLD, can respond to power failures at hardware-level speed.
[0063] In practical applications, CPLDs or FPGAs can be selected to implement independent logic control units based on the specific application scenario requirements (such as logic complexity, resource requirements, upgrade flexibility, cost control, etc.). For example, for scenarios with relatively fixed logic functions and cost sensitivity, CPLDs can be used to reduce costs and simplify design; for scenarios with complex logic functions, which may require later upgrades and optimizations, and have high flexibility requirements, FPGAs can be used to meet more complex logic processing needs. In the embodiments of this application, both CPLDs and FPGAs are programmable logic devices, whose core feature is that specific logic functions can be implemented through hardware configuration. The control logic they execute exists in the form of hardware circuits, which is different from the scheme of implementing functions by executing software instructions through a processor. The selection of CPLDs and FPGAs makes the response speed, reliability, and determinism of independent logic control units significantly better than BMC or CPU schemes that rely on software operation.
[0064] In this embodiment, by employing a CPLD or FPGA to implement an independent logic control unit, control functions can be independently completed in a purely hardware logic manner without relying on the main operating system and BMC. Compared to BMC or CPU solutions that rely on software operation, CPLD or FPGA have significant advantages in terms of fast response speed, high reliability, and strong deterministic latency. Their response latency is determined by the transmission delay of the hardware gate circuits, enabling fault detection and trigger response to be completed within microseconds, reducing the uncertainty and delay caused by software scheduling and instruction execution. Furthermore, since all control logic of the CPLD or FPGA is fixed in hardware circuit form, it is unaffected by operating system crashes, BMC power failures, or software anomalies, ensuring reliable execution of fault field data freezing and transfer operations under extreme power failure scenarios.
[0065] In some embodiments, the non-volatile high-speed memory cell is FRAM.
[0066] FRAM is a non-volatile memory that utilizes the polarization effect of ferroelectric crystals to store data. Its core storage principle is that under the influence of an electric field, the polarization direction of the ferroelectric material changes with the direction of the electric field and remains polarized after the electric field is removed, thus achieving data storage. FRAM features fast write speeds (nanosecond level), direct writing without erasure, and extremely long read / write lifespan (approximately [missing information - likely related to read speeds and time).) Features include low power consumption and other characteristics.
[0067] In this embodiment, the FRAM communicates with an independent logic control unit via an SPI bus or an I2C bus, and its power supply terminal is connected to the output terminal of the isolated energy storage unit, which supplies power independently. The FRAM is internally divided into multiple storage areas, including but not limited to: a fault data storage area for storing fault scene data and a configuration parameter storage area for storing system configuration parameters (such as sampling frequency, trigger threshold, etc.).
[0068] When the main power bus is supplying power normally, the FRAM is in standby mode and does not perform write operations to reduce power consumption. In normal mode, the independent logic control unit continuously collects operating status data of the VRM and / or PSU at a preset sampling period and writes the data to the internal cache in a circular buffer manner. At this time, the FRAM does not need to participate in data writing and does not consume additional power. When the independent logic control unit detects an invalid power status indication signal or a bus voltage drop exceeding a preset threshold through its detection input, it immediately generates a freeze trigger signal, stops updating the internal cache, and writes the fault condition data stored in the internal cache to the FRAM with the highest priority via the SPI bus or I2C bus.
[0069] Because FRAM can be written to directly without erasure, and its write speed can reach nanosecond levels (typically 50ns~100ns), the entire data transfer process can be completed in microseconds to milliseconds, which is much shorter than the time window during which the isolated energy storage unit can maintain power (e.g., 500ms to 2s). Therefore, even in the extreme case of complete failure of the main power bus, the isolated energy storage unit still has sufficient power to support the FRAM to complete the writing operation of all fault field data, ensuring that the data is reliably and permanently preserved.
[0070] Furthermore, due to the high read / write lifespan of FRAM... This is significantly higher than that of traditional Flash memory. At a sub-scale, FRAM enables independent logic control units to perform high-frequency cyclic writing to the internal cache in normal mode without worrying about memory cell failure due to frequent writes. Simultaneously, FRAM's write power consumption is significantly lower than that of Flash memory, which is beneficial for extending emergency operating time under isolated energy storage unit power supply conditions. These characteristics make FRAM particularly suitable for the application scenarios described in this application, which require extremely fast data writing during power failures and frequent erase / write operations. In practical applications, those skilled in the art can select the specific model and specifications of FRAM according to specific needs (such as storage capacity, write speed, cost control, etc.), for example, choosing 2Mbit, 4Mbit, or higher capacity FRAM chips.
[0071] In this embodiment, by setting the non-volatile high-speed storage unit as FRAM, the characteristic of FRAM that it can be directly written to without erasure can be utilized, eliminating the erasure operation time required by traditional Flash memory. Combined with nanosecond-level write speeds, all fault site data can be transferred within microseconds to milliseconds, far less than the emergency power supply time window of the isolated energy storage unit, ensuring that fault site data can still be reliably and permanently preserved even after the main power bus has completely failed. Simultaneously, FRAM has high... With over 1000 read / write cycles, it can support long-term, high-frequency cyclic writing of internal cache to the independent logic control unit in normal mode. Moreover, its write power consumption is much lower than that of Flash memory, which is beneficial for extending emergency working time under the condition of independent power supply of isolated energy storage unit, thus providing sufficient power margin for complete transfer of fault field data. In addition, FRAM can retain stored data without any backup power supply after power failure, ensuring the long-term traceability of fault field data after the system is completely powered off, thus providing a complete and reliable data foundation for post-fault root cause analysis.
[0072] In some embodiments, the data acquisition unit includes an ADC and a digital interface. The analog input of the ADC is connected to the voltage sampling terminal, current sampling terminal, and temperature sampling terminal of the VRM and / or PSU for acquiring voltage, current, and temperature data of the VRM and / or PSU. The digital interface is connected to the status code output terminal of the VRM and / or PSU for reading the status code of the internal register of the VRM and / or PSU.
[0073] The data acquisition unit is connected to the voltage, current, and temperature sampling terminals of the VRM and / or PSU via the analog input terminals of the ADC, for acquiring voltage, current, and temperature data of the VRM and / or PSU. Specifically: the voltage sampling terminal acquires the analog output voltage signal of the VRM and / or PSU; the current sampling terminal acquires the analog output current signal of the VRM and / or PSU (typically obtained through a sampling resistor or current sense amplifier); and the temperature sampling terminal acquires the analog temperature signal of the VRM and / or PSU (typically obtained through an internal or external temperature sensor). Correspondingly, the ADC converts these analog signals into digital signals and transmits them to the independent logic control unit.
[0074] The data acquisition unit also connects to the status code output terminals of the VRM and / or PSU via a digital interface to read the status codes of the internal registers of the VRM and / or PSU. The VRM and PSU typically integrate digital communication interfaces such as a Power Management Bus (PMBus) or Serial Voltage Identification (SVID), and their internal registers store various status codes, such as overvoltage alarm flags, undervoltage alarm flags, overcurrent alarm flags, overtemperature alarm flags, fault codes, operating mode status, and enable status. Accordingly, the data acquisition unit directly reads these register status codes through the aforementioned digital interface without requiring ADC conversion.
[0075] The data acquisition unit summarizes the collected voltage, current, temperature data and internal register status codes, and transmits them to the data input terminal of the independent logic control unit through its output terminal; the independent logic control unit writes the above data into its internal cache in a cyclic buffering manner in normal mode.
[0076] The voltage sampling terminal is typically located at the output of the VRM and / or PSU to acquire the analog signal of the output voltage; the current sampling terminal is typically connected to the current sensing resistor or current sensing amplifier inside the VRM and / or PSU to acquire the analog signal of the output current; the temperature sampling terminal is typically connected to the internal or external temperature sensor of the VRM and / or PSU to acquire the analog signal of the temperature; the status code output terminal is a digital communication interface pin of the VRM and / or PSU (such as the DATA and CLK pins of PMBus, the DATA and CLK pins of SVID, etc.) to output the status codes of the internal registers.
[0077] In this embodiment, the VRM and PSU are two key components in the server power supply chain: the PSU converts external AC power into the DC voltage required by the server's main power supply (e.g., 12V / 48V), and the VRM further converts the main power bus voltage into low-voltage, high-current (e.g., 0.8V~1.8V) required by core chips such as the CPU or GPU. In actual fault diagnosis, if the acquired PSU output voltage drops abnormally first, followed by a drop in the VRM input voltage, the fault may originate from the PSU side; if the PSU output voltage is normal but the VRM output voltage is abnormal (e.g., voltage drop after current surge), the fault may originate from the load side (e.g., GPU / CPU short circuit) or the VRM itself. By simultaneously acquiring the operating status data of the VRM and / or PSU through the data acquisition unit, it is possible to accurately determine which stage of the power supply chain the fault occurred at after it occurred, thus providing precise data support for fault root cause analysis.
[0078] The ADC in the data acquisition unit can be implemented using a multi-channel high-precision ADC chip. Its sampling rate can be configured according to actual needs, such as 100kHz to 1MHz, to capture microsecond-level voltage / current transient changes. The ADC resolution can be selected according to actual accuracy requirements, such as 12-bit or 16-bit. The ADC sampling trigger can be controlled by an independent logic control unit to ensure that the sampling timing is synchronized with the buffer write operation of the independent logic control unit. The communication protocol of the digital interface depends on the interface type supported by the VRM and / or PSU, such as PMBus, SVID, or I2C. The data acquisition unit or independent logic control unit needs to support the corresponding communication protocol to complete the reading operation of the status code of the internal register of the VRM and / or PSU.
[0079] It should be noted that the specific selection of the above-mentioned ADC sampling rate, resolution, and digital interface protocol type depends on the specifications and performance requirements of the VRM and / or PSU in the actual application, and this application does not make specific limitations on this.
[0080] In this embodiment, by setting up an ADC and a digital interface for the data acquisition unit, the analog input of the ADC acquires analog data such as voltage, current, and temperature of the VRM and / or PSU, while the digital interface directly reads digital data such as status codes from the internal registers of the VRM and / or PSU. This achieves comprehensive and real-time acquisition of operational status data of key nodes in the server power supply link, providing a rich data source for fault diagnosis. Specifically, the ADC can capture transient changes in voltage / current with a microsecond-level sampling period, providing a data foundation for reconstructing the fine waveforms of the moment before power failure after a fault occurs. The digital interface can directly read alarm flags and fault codes such as overvoltage, undervoltage, overcurrent, and overtemperature from the internal registers of the VRM and / or PSU without ADC conversion, avoiding delays and errors that may be introduced by analog sampling, and ensuring that status codes are acquired accurately and promptly. Through this dual-channel acquisition architecture that combines analog and digital signals, key data from each link in the power supply link can be completely captured in the final moments before the main power bus fails, providing comprehensive, detailed, and accurate data support for post-fault root cause analysis (e.g., distinguishing whether the fault originates from the PSU side, VRM side, or load side).
[0081] The above embodiments illustrate the implementation of a server failure on-site data freezing system. Next, specific embodiments will be used to illustrate the application of this system.
[0082] Figure 2 This is a flowchart illustrating a server failure field data freezing method provided as an exemplary embodiment of this application. The server failure field data freezing method provided in this application embodiment is applied to an independent logic control unit in a server failure field data freezing system as described in any of the above embodiments. Figure 2 As shown, the method for freezing data at the server failure site includes:
[0083] S201. Under normal server operation, obtain the operating status data of VRM and / or PSU on the server motherboard, and write the operating status data into the internal cache of the independent logic control unit in a circular cache manner.
[0084] Among them, the server fault field data freezing system is used to continuously generate time window data before the fault when the server is under normal power supply, and to freeze the data after the power supply anomaly occurs; the independent logic control unit is the core of the method execution, which can be implemented by CPLD, FPGA or equivalent hardware logic devices. This unit works directly connected to the monitoring link on the server motherboard, without relying on the main operating system scheduling or the BMC to complete the final freeze control. The operational status data is acquired by the data acquisition unit from the monitoring link of the VRM and / or PSU. This operational status data is the original data source for subsequent fault location and includes at least one of voltage, current, power, temperature, and status information to comprehensively reflect the working status of the VRM and PSU. The status information can be a fault bit, alarm bit, enable status, power good status, or output regulation status. The internal cache is a temporary storage area configured inside the independent logic control unit, which can be on-chip RAM, register array, or static memory area directly connected to the logic unit at high speed. The circular cache refers to the internal cache being continuously written to at a fixed length and overwriting the oldest data after it is full, so that the cache always retains the continuous operational status data of the most recent period before the fault occurred, so as to realize the continuous formation of "fault pre-fault time window data".
[0085] In practical implementation, the independent logic control unit initiates data acquisition timing after the server powers on and enters normal operation to continuously receive operational status data from the data acquisition unit. The source channels for operational status data can include at least one of the following: VRM telemetry interface, PMBus bus, I2C bus, SPI interface, GPIO status pin, and analog-to-digital converter input port. The data acquisition unit uses these channels to sample and convert the raw data and transmits the results to the independent logic control unit. For numerical quantities such as voltage, current, power, and temperature, the data acquisition unit performs polling or parallel sampling according to a preset sampling period. The sampling period can be set to a fixed value within the range of 10 microseconds to 10 milliseconds. This sampling period, together with the internal buffer capacity, determines the length of the historical time window that can be retained. For example, if the internal buffer is organized by recording frames, with each frame containing a time stamp, VRM output voltage, VRM output current, PSU output voltage, PSU output current, temperature value, and status code, then several hundred or several thousand frames of data can be continuously recorded within a 1-millisecond sampling period. For status information, the data acquisition unit can use a parallel approach of interrupt latching and periodic sampling. Alarm pins with rapid transitions are directly latched by the hardware edge detection circuit, while stable state quantities are written into the record frame according to the sampling period to ensure that short-term state changes are not lost during the cache update process.
[0086] In terms of data writing organization, the internal cache is divided into a sequential address space and a write pointer is set. When the independent logic control unit receives a complete frame of running status data sent by the data acquisition unit, it writes the frame to the cache location currently pointed to by the write pointer, and appends an incrementing time sequence number or a local counter value. After writing is completed, the write pointer automatically moves forward. When the write pointer reaches the end of the cache, it wraps back to the beginning address of the cache and overwrites the earliest written content, thus forming a first-in-first-out circular cache. The writing process of this circular cache is independent of the scheduling of any external bus or processor and is completely completed autonomously by the hardware logic inside the independent logic control unit. In order to ensure the consistency of multi-source data at the same time, this application can first set up a sampling register group inside the independent logic control unit, and summarize the VRM and PSU related data obtained at the same sampling time into a single frame before writing it into the internal cache, so that each frame corresponds to a specific sampling time point. If the update rates of some monitored quantities are different, the previous sampling value is used for the unupdated fields and an update flag is set in the frame for subsequent restoration of the data evolution order.
[0087] Through the data acquisition and cyclic writing method in step S201 above, the internal cache always retains continuous fault scene data fragments from the most recent period before the main power bus power supply anomaly occurred, rather than only retaining coarse results after the fault occurred. This step uses an independent logic control unit to acquire data from the data acquisition unit and stores the operating status data in an overlay cache manner, so that subsequent freeze actions have a complete data foundation that can be transferred, and can reflect the dynamic changes on the VRM side and PSU side before and after the fault. It is understood that the specific examples above regarding sampling period, data frame format, bus type and cache capacity are only illustrative descriptions for the purpose of understanding the solution of this application, and are not intended to limit the scope of protection of this application.
[0088] S202. Obtain the power status indication signal and / or bus voltage of the server's main power bus, and determine whether the main power bus has experienced a power supply abnormality.
[0089] Among them, the main power bus is the input bus for the critical power supply of the server or motherboard, and is the direct target for power supply anomaly detection; the power status indication signal is a digital signal reflecting the power supply status of the main power bus, which can be composed of a power good signal, a power supply presence signal, or a board-level power effective indication signal output by the power module; the bus voltage is the actual voltage value of the main power bus, which can be sent to the independent logic control unit through voltage division sampling and analog-to-digital conversion; the power supply anomaly is the result of the trigger judgment of the freeze action, which can be manifested as an abnormal power supply status or voltage status of the main power bus; the independent logic control unit continuously monitors the above input quantities through hardware logic, so that anomaly identification does not depend on the software polling cycle.
[0090] In specific implementations, the power status indication signal can be connected to the digital input port of the independent logic control unit via a level conversion or isolation interface; the bus voltage can be acquired by using a resistor divider network to convert the bus voltage, which is higher than the input range of the logic device, into a sampleable level, and then quantized and sampled by an on-chip ADC or an external ADC; the independent logic control unit determines whether a power supply abnormality has occurred on the main power bus based on the acquired power status indication signal and / or bus voltage. Optionally, in some embodiments, anomaly identification can be performed based on changes in the power status indication signal, changes in the bus voltage, and / or voltage change trends.
[0091] By simultaneously introducing detection paths for digital status signals and analog voltage, the independent logic control unit can monitor the power supply status of the main power bus and generate a power supply anomaly determination result. This step establishes a hardware direct link from power supply status changes to freeze triggering, ensuring sufficient response speed for subsequent data freezes.
[0092] S203. When the main power bus experiences a power supply abnormality, a freeze trigger signal is generated to transfer the fault scene data stored in the internal cache to the non-volatile high-speed storage unit in the server fault scene data freeze system; wherein, when the main power bus fails to supply power, the isolated energy storage unit in the server fault scene data freeze system provides emergency working power to the independent logic control unit and the non-volatile high-speed storage unit.
[0093] The freeze trigger signal is a control signal output by the independent logic control unit after confirming a power supply anomaly. It is used to lock the internal cache contents and start the transfer sequence. The fault scene data is the key operational information that needs to be retained before and after the power supply anomaly. It includes at least the continuous operational status data frames already formed in the internal cache, and may include the marker frame at the time of the anomaly determination, the anomaly source identifier, and the last valid bus detection value. The non-volatile high-speed storage unit is used to permanently save the fault scene data in the event of a power outage. Specifically, it can use FRAM or other non-volatile memory with fast write capability before power failure and no need for erase preparation during the write process. The isolation energy storage unit is a temporary independent power source after the main power supply fails. Specifically, it can include supercapacitor banks, energy storage capacitor arrays, or backup energy storage modules with isolated outputs. Its output is connected to the emergency power input terminal of the independent logic control unit and the non-volatile high-speed storage unit, and maintains power isolation from the main power supply working path.
[0094] In specific implementation, the independent logic control unit generates a freeze trigger signal within the same control cycle after obtaining the power supply anomaly determination result in S202. This signal first acts on the internal cache write control logic, causing the circular cache write pointer to stop moving forward, and the data of each frame in the current cache area will no longer be overwritten by new sampled data. If there are still sampled values being received at the moment of the anomaly, the control logic can encapsulate the sampled value as a snapshot of the last frame at the moment of the anomaly and write it to the reserved area before closing the write operation. Subsequently, the independent logic control unit reads the valid data area in the internal cache, determines the order of the earliest and latest valid frames according to the current position of the write pointer, and reorganizes the circular storage order into a linear output order, and then writes it to the non-volatile high-speed memory unit via the SPI bus or I2C bus. To facilitate subsequent reading and parsing, the transferred data can be organized according to a fixed file header or a fixed record header, where the record header at least contains the total number of frames, the anomaly trigger type, the freeze time count value, the cache start index, and the checksum, and the record body stores the running status data of each frame in chronological order.
[0095] During normal power supply from the main power bus, the isolated energy storage unit is in a charging and holding state, and its output side is connected to the emergency power supply network via an isolation circuit or an ideal diode switching circuit. When the main power bus fails to supply power or the power drops below the switching threshold, the emergency power supply network is automatically taken over by the isolated energy storage unit, while the independent logic control unit and the non-volatile high-speed storage unit continue to operate until the freeze trigger, cache locking, and data transfer are all completed.
[0096] In one possible embodiment, the non-volatile high-speed storage unit uses FRAM devices, and the independent logic control unit writes continuously page by page or frame by frame without performing erase waiting, thereby completing the write of all fault site data to disk within the limited emergency power supply time. If the internal cache capacity is large, the control logic can first transfer the most recent frames and abnormal snapshots, and then continue to transfer the remaining historical frames until the voltage of the isolated energy storage unit drops to the termination threshold; however, in the complete implementation of this application, the isolated energy storage unit is configured according to the capacity required for full cache transfer. After the transfer is completed, the independent logic control unit can write an end identifier and verification result to the non-volatile high-speed storage unit to indicate that the fault site data has been completely sealed. Based on the above processing, in scenarios of PSU-side power supply failure, external power failure, or rapid bus drop, the continuous data before the fault and the fault triggering time information are stored together in the non-volatile high-speed storage unit, so that subsequent fault analysis can distinguish PSU anomalies, VRM anomalies, and load-side induced problems based on the time-continuous data chain. It should be understood that the above examples are only illustrative and not limiting.
[0097] Based on the above analysis, this application provides a method for freezing server fault scene data, including: acquiring the operating status data of the voltage regulation module and / or power module on the server motherboard during normal server operation, and writing the operating status data into the internal cache of the independent logic control unit in a circular cache manner; acquiring the power status indication signal and / or bus voltage of the server's main power bus, and determining whether the main power bus has experienced a power supply abnormality; generating a freeze trigger signal when the main power bus experiences a power supply abnormality, so as to transfer the fault scene data stored in the internal cache to the non-volatile high-speed storage unit in the server fault scene data freezing system; wherein, when the main power bus fails to supply power, the isolated energy storage unit in the server fault scene data freezing system provides emergency working power to the independent logic control unit and the non-volatile high-speed storage unit. This application uses an independent logic control unit to continuously collect VRM and power module operating status data at the hardware level, forms a continuous time window before the fault through internal caching, determines anomalies through main power bus status and voltage detection, and completes the transfer to a non-volatile high-speed storage unit with the support of an isolated energy storage unit. This allows key fault field data to be retained even when the main power supply suddenly fails. The fault record is no longer limited to the conclusion of power failure, but can output a time-series data chain that can be used for root cause location, thereby improving the accuracy of server fault location and the efficiency of fault root cause analysis.
[0098] In some embodiments, determining whether a power supply abnormality has occurred on the main power bus includes: determining that a power supply abnormality has occurred on the main power bus when the power status indication signal changes from an effective level to an ineffective level; and / or determining that a power supply abnormality has occurred on the main power bus when the voltage drop exceeds a preset threshold or the voltage drop slope exceeds a preset threshold.
[0099] In its implementation, the independent logic control unit receives a power status indication signal, such as the Pwr_Good signal, from the bus detection circuit and captures its edge changes. When the signal is detected to switch from an active level to an inactive level, an edge interrupt is triggered, and a power supply anomaly determination result is immediately output. Simultaneously, the bus voltage is continuously acquired by the voltage sampling circuit. The sampled values are converted from analog to digital and then input to the independent logic control unit. The independent logic control unit calculates the voltage drop amplitude and drop slope based on adjacent sampling points. When the voltage drop amplitude reaches or exceeds a preset amplitude threshold, or the drop slope reaches or exceeds a preset slope threshold, a power supply anomaly determination result is output. To adapt to different server platforms, the power status indication signal can be in the form of open-drain output, push-pull output, or logic level output. In practical applications, other models of this component can also be selected; this application does not limit this selection.
[0100] During operation, when the main power bus is supplying power normally, the power status indicator signal remains at a valid level, the bus voltage remains within the allowable range, and the independent logic control unit maintains normal monitoring. When an external power failure, PSU power supply failure, or rapid bus voltage drop occurs, the power status indicator signal will first or simultaneously undergo a level flip, or the bus voltage drop amplitude or drop slope will exceed the threshold. Based on this, the independent logic control unit will determine that the main power bus is experiencing a power supply abnormality and output trigger conditions to the subsequent freeze control logic.
[0101] By adopting the above judgment method, power supply anomalies can be identified by using both status signals and voltage quantization features. This enables anomaly judgment to cover scenarios such as power supply failure, rapid drops, and instability, and provides timely trigger signals to the fault site data freezing logic, thereby improving the reliability of power supply anomaly identification and the integrity of fault site data retention.
[0102] Based on the above embodiments, in some embodiments, the server fault site data freezing method further includes: after the non-volatile high-speed storage unit completes data transfer, sending a power-off indication signal to the isolation energy storage unit so that the isolation energy storage unit stops supplying power to the independent logic control unit and the non-volatile high-speed storage unit.
[0103] In one implementation, after data transfer is completed, the independent logic control unit generates a power-off indication signal based on a write confirmation flag or a completion response from the memory, and sends it to the enable terminal of the isolated energy storage unit via the control bus. Correspondingly, upon receiving the power-off indication signal, the isolated energy storage unit shuts down its output switch and stops supplying power to the independent logic control unit and the non-volatile high-speed memory unit, subsequently entering a recharge state. The power-off indication signal can be a high / low level toggle signal, a shutdown pulse, or a control frame with a status code, the specific signal form matching the control interface used. This control process establishes a clear timing correlation between the end of data transfer and the end of emergency power supply. After confirming that all cached data has been written to the non-volatile high-speed memory unit, the independent logic control unit promptly notifies the isolated energy storage unit to cut off the output, thereby switching the system from the data transfer power supply state to the power-off hold state. With this configuration, the isolated energy storage unit only discharges during the necessary data transfer period and stops supplying power after the transfer is completed. Combined with the recharging process after the main power is restored, this creates a closed-loop control system for fault site data preservation and energy storage management. By adopting the above processing mechanism, the non-volatile high-speed storage unit triggers power-off control as soon as it completes writing, which can reduce the ineffective consumption of emergency energy storage and enable the isolated energy storage unit to exit the discharge state in time after completing a fault recording task, so as to reserve sufficient power for subsequent re-entry into the fault freeze scenario, while maintaining the integrity and traceability of the saved fault data.
[0104] In another implementation, the power-off indication signal is generated by a timer module built into the independent logic control unit. When initiating a data transfer operation, the independent logic control unit synchronously starts the timer, setting the timing duration to a pre-configured transfer timeout protection period (this timeout protection period is greater than or equal to the maximum write duration of the non-volatile high-speed storage unit under the worst-case scenario). When the timer count reaches the preset timeout protection period, regardless of whether the write confirmation flag is set, the independent logic control unit forcibly sends a power-off indication signal to the isolated energy storage unit. This method is suitable for scenarios where the non-volatile high-speed storage unit fails to return a write completion response due to abnormal conditions. The timeout protection mechanism prevents the isolated energy storage unit from continuously discharging until its power is depleted while waiting for a response, ensuring that the system can exit the transfer power supply state within a controllable time under any circumstances. This helps reduce irreversible lifespan loss caused by deep discharge of the energy storage unit.
[0105] In another implementation, the power-down indication signal is generated by an independent logic control unit by monitoring the write enable signal of the non-volatile high-speed memory cell. During data transfer, the write enable signal of the non-volatile high-speed memory cell remains at a valid level (e.g., active high). Once the last frame of cached data is written, the memory cell automatically pulls the write enable signal back to an invalid level. The independent logic control unit captures the falling edge transition of this write enable signal in real time, delays the detection of the falling edge for a preset short holding time (e.g., 10 to 100 microseconds), and sends a power-down indication signal to the isolated energy storage unit after the internal state machine of the memory cell completes the full loop of the write operation. This method utilizes the memory cell's own write control signal as the basis for determining transfer completion, without waiting for the memory to return a completion response from the software layer, and without relying on a fixed duration set by a timer. It can trigger power-down control the instant the write operation is actually completed, achieving minimal delay between the end of transfer and the power-down indication, further reducing unnecessary discharge time of the isolated energy storage unit, and maximizing energy storage margin for subsequent fault freeze scenarios.
[0106] The above-mentioned implementation methods can be selected and configured or combined according to the actual application scenario: For scenarios with high write reliability requirements, the first method based on write confirmation flag can be used first; for scenarios that need to prevent abnormal unresponsiveness of storage units, a second timer timeout protection method can be added as a fallback mechanism; for scenarios that pursue the minimum transfer power consumption and the fastest power-off response, the third method based on write enable signal edge detection can be selected.
[0107] In summary, this application has at least the following advantages:
[0108] I. By introducing an independent logic control unit and a physically isolated emergency power supply domain into the server motherboard, a fault detection capability unavailable to traditional BMC management systems is achieved. Specifically, the independent logic control unit directly monitors the bus voltage based on hardware logic, without relying on the main operating system for scheduling or waiting for BMC polling responses. It can complete power supply anomaly detection and trigger data freeze within microseconds, significantly improving real-time performance compared to the millisecond to second-level sampling of traditional software logs. Simultaneously, the isolated energy storage unit (such as a supercapacitor) serves as an emergency power source, physically isolated from the main power supply through ideal diodes. Even if the main power supply module (PSU) completely fails or is physically damaged, this isolated domain can still maintain stable operation for at least 500ms to 2s, sufficient to support the complete transfer of data from the time window before the fault in the internal cache to the non-volatile high-speed storage unit, fundamentally improving the problem of log loss during power outages.
[0109] Second, a circular caching mechanism continuously records the operational status data of the VRM and / or PSU under normal server conditions, ensuring that the internal cache always retains continuous fault scene data fragments from the most recent period before the main power bus power supply anomaly occurred. Based on the recorded multi-source heterogeneous data (including voltage, current, power, temperature, and VRM internal status words, Pwr_Good timing signals, etc.), maintenance personnel can accurately reconstruct the dynamic evolution process before and after the fault occurred. For example, when the recorded current data shows a surge followed by a voltage drop, it can be determined that there is a sudden short circuit on the load side (GPU / CPU); when the voltage drops directly without any warning, it can be determined that there is a failure on the power supply side (PSU). This refined fault tracing capability based on waveform data cannot be provided by traditional simple logs that only record the fault results.
[0110] Third, in ultra-large-scale data center applications, the precise fault location capability provided by this application can effectively shorten the mean time to repair (MTBT). Because it can accurately locate specific damaged components (such as a particular phase VRM or a specific PSU), maintenance personnel do not need to blindly replace entire expensive motherboards or GPU accelerator cards using a "replacement method," significantly reducing spare parts waste and unnecessary downtime, and substantially lowering the maintenance cost per server. From a lifecycle perspective, this application provides significant practical value for data center operations and maintenance by shortening fault diagnosis time, reducing erroneous replacement rates, and improving maintenance efficiency.
[0111] Fourth, in this application, there is no software-level data interaction between the independent logic control unit and the main system. The data freezing process does not rely on the operating system's file system or the BMC's log service, fundamentally avoiding interference from software crashes, task blocking, or resource contention on the fault recording process, and achieving hardware-level deterministic automatic protection behavior. Furthermore, the non-volatile high-speed storage unit can use FRAM, with a write endurance of up to [missing information - likely a number]. Furthermore, unlike Flash memory, it does not require a large capacity capacitor to perform slow, long-term write / erase operations during the writing process. Even under extreme conditions (such as high temperature, severe vibration, and severe power supply fluctuations), it can still ensure the reliable preservation of fault data, greatly improving the data survival rate and providing complete and reliable original data for post-fault analysis. It has good social benefits and safety value.
[0112] Fifth, by using a control strategy that triggers power outage immediately after the transfer is completed, the isolated energy storage unit can promptly exit the discharge state and enter the preparation for recharging. This reduces the number of ineffective discharges that continue after the emergency energy storage has completed the freezing task, effectively prevents deep discharge of energy storage components, extends the cycle life of energy storage components such as supercapacitors, and ensures the sustainability of the freezing function in multiple consecutive abnormal events. This forms a complete closed-loop control from "emergency start-up - transfer power supply - power outage completion - recharging," achieving optimal management of emergency energy storage resources.
[0113] Figure 3 A schematic diagram of the structure of a server provided for an exemplary embodiment of this application. For example... Figure 3 As shown, the server 30 includes a motherboard 31 and a server fault field data freezing system 10, as described in any of the above embodiments, disposed on the motherboard 31.
[0114] This server integrates a fault scene data freezing and recording system onto the motherboard, enabling continuous acquisition and correlation processing of status information from the main power bus, PSU, VRM, and control links at the near-board level. It relies on the collaborative work of isolated energy storage units, data acquisition units, non-volatile high-speed storage units, and independent logic control units within the system. During normal power supply, it cyclically caches critical operational data. When an invalid power status indication signal is detected or the bus voltage drops beyond a preset threshold, it immediately triggers freezing and data transfer. This allows it to retain critical scene data before and after a fault even in the event of a sudden power outage or power anomaly. The server no longer only records coarse-grained power outage results but can more accurately distinguish between PSU-side failures, VRM anomalies, or load-side cascading problems, thereby improving fault location accuracy, shortening maintenance and troubleshooting time, and enhancing business continuity assurance capabilities.
[0115] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0116] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0117] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A server failure on-site data freezing system, characterized in that, include: An isolated energy storage unit is connected to the main power bus of the server through an isolation device. It is used to store electrical energy when the main power bus is powered normally and to release electrical energy to provide emergency working power when the main power bus fails to power. The data acquisition unit is connected to the monitoring pins of the voltage regulation module (VRM) and / or the power supply module (PSU) on the server motherboard. It is used to collect real-time operating status data, including the output voltage, output current, operating temperature, and internal register status codes of the VRM and / or the PSU. A non-volatile high-speed memory cell, wherein the power supply terminal of the non-volatile high-speed memory cell is connected to the output terminal of the isolated energy storage cell; An independent logic control unit is provided with a power supply terminal, a detection input terminal, a data input terminal, and a data output terminal. The power supply terminal is connected to the output terminal of the isolated energy storage unit. The detection input terminal is connected to the power status indicator terminal and / or the bus voltage sampling terminal of the main power bus. The data input terminal is connected to the output terminal of the data acquisition unit. The data output terminal is connected to the non-volatile high-speed storage unit. The independent logic control unit is used to cyclically write the operating status data into the internal cache in normal mode; When the power status indication signal is found to be invalid or the bus voltage drops below a preset threshold through the detection input terminal, a freeze trigger signal is generated to transfer the fault field data in the internal cache to the non-volatile high-speed storage unit.
2. The server fault on-site data freezing system according to claim 1, characterized in that, The isolated energy storage unit includes: A supercapacitor array is used to store electrical energy and release it when the main power bus fails to supply power. A charging control module, wherein the input terminal of the charging control module is connected to the main power bus and the output terminal of the charging control module is connected to the supercapacitor array, and is used to manage the charging of the supercapacitor array when the main power bus is supplying power normally; The isolation device is connected in series between the supercapacitor array and the main power bus to prevent the power of the supercapacitor array from flowing back to the main power bus when the main power bus fails to supply power, and to provide an independent power supply circuit for the independent logic control unit and the non-volatile high-speed memory unit.
3. The server fault field data freezing system according to claim 2, characterized in that, The isolation device is an ideal diode or a reverse-current protection diode.
4. The server fault field data freezing system according to any one of claims 1 to 3, characterized in that, The independent logic control unit is a complex programmable logic device (CPLD) or a field-programmable gate array (FPGA).
5. The server fault field data freezing system according to any one of claims 1 to 3, characterized in that, The non-volatile high-speed memory cell is a ferroelectric RAM.
6. The server fault field data freezing system according to any one of claims 1 to 3, characterized in that, The data acquisition unit includes an analog-to-digital converter (ADC) and a digital interface. The analog input terminal of the ADC is connected to the voltage sampling terminal, current sampling terminal, and temperature sampling terminal of the VRM and / or the PSU, and is used to acquire the voltage, current, and temperature data of the VRM and / or the PSU. The digital interface is connected to the status code output terminal of the VRM and / or the PSU, and is used to read the internal register status code of the VRM and / or the PSU.
7. A method for freezing on-site data during server failure, characterized in that, An independent logic control unit applied in the server fault field data freezing system as described in any one of claims 1 to 6, wherein the server fault field data freezing method comprises: Under normal server operation, acquire the operating status data of the voltage regulation module (VRM) and / or power supply module (PSU) on the server motherboard, and write the operating status data into the internal cache of the independent logic control unit in a circular cache manner; Obtain the power status indication signal and / or bus voltage of the main power bus of the server, and determine whether the main power bus has a power supply abnormality; When the main power bus experiences a power supply anomaly, a freeze trigger signal is generated to transfer the fault scene data stored in the internal cache to the non-volatile high-speed storage unit in the server fault scene data freeze system; wherein, when the main power bus fails to supply power, the isolated energy storage unit in the server fault scene data freeze system provides emergency working power to the independent logic control unit and the non-volatile high-speed storage unit.
8. The method for freezing on-site server fault data according to claim 7, characterized in that, The determination of whether the main power bus has experienced a power supply abnormality includes: When the power status indication signal changes from an active level to an inactive level, it is determined that the main power bus has experienced a power supply abnormality. And / or, when the voltage drop of the bus exceeds a preset voltage threshold or the drop slope exceeds a preset slope threshold, it is determined that the main power bus has a power supply abnormality.
9. The method for freezing on-site server fault data according to claim 7, characterized in that, Also includes: After the non-volatile high-speed storage unit completes the data transfer, a power-off indication signal is sent to the isolated energy storage unit so that the isolated energy storage unit stops supplying power to the independent logic control unit and the non-volatile high-speed storage unit.
10. A server, characterized in that, Includes a motherboard, and a server fault field data freezing system as described in any one of claims 1 to 6 disposed on the motherboard.