BMC and CPLD decoupling control system and server

CN122547720APending Publication Date: 2026-08-11SHANGHAI EVEX INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]本申请实施例提供的一种BMC与CPLD的解耦控制系统及服务器,用以解决现有BMC与CPLD耦合架构下存在的系统可维护性与运行可靠性不足的问题

Benefits of technology

[0020]本申请实施例提供的一种BMC与CPLD的解耦控制系统及服务器,通过设置供电控制模块,由逻辑控制单元获取节点在位识别信号与CPLD输出的BMC供电使能信号,并利用两路信号共同作为BMC供电控制依据,能够依据主板节点状态和CPLD控制需求对功率开关单元进行使能控制,并向BMC板卡稳定输出待机电压,使得BMC的供电来源不再单一依赖于CPLD,进而降低BMC与CPLD之间的耦合程度,提升开发维护效率和系统运行可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122547720A_ABST
    Figure CN122547720A_ABST
Patent Text Reader

Abstract

This application provides a decoupling control system and server for a BMC and CPLD. The system includes a power supply control module mounted on the server motherboard; wherein the power supply control module includes a logic control unit and a power switch unit; a first input terminal of the logic control unit is used to acquire a node presence identification signal on the server motherboard, and a second input terminal is used to acquire a power supply enable signal from the CPLD (Complex Programmable Logic Device) outputting the BMC power supply; the output terminal of the logic control unit is connected to the enable terminal of the power switch unit, and is used to control the power switch unit to conduct when the node presence identification signal is valid and / or the BMC power supply enable signal is valid; the input terminal of the power switch unit is used to connect to a standby power supply, and the output terminal of the power switch unit is used to output a standby voltage to the BMC board. This system is used to achieve power supply decoupling and fault isolation between the BMC and CPLD, improving system reliability and development and maintenance efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of server technology, and in particular to a decoupled control system and server for BMC and CPLD. Background Technology

[0002] With the rapid development of information technology, the demand for servers in data centers, cloud computing, and edge computing has exploded. As the core equipment for data storage and processing, the stability, maintainability, and scalability of servers directly affect the operational efficiency and reliability of the entire information system. In server hardware architecture design, the collaborative architecture between the Baseboard Management Controller (BMC) and the Complex Programmable Logic Device (CPLD) is a crucial element in ensuring stable server operation. The BMC is responsible for out-of-band management of the server, enabling important functions such as remote monitoring, firmware updates, and fault alarms; the CPLD undertakes hardware logic control tasks, such as power-on timing management, pin multiplexing, and hardware reset. The two work closely together in the server hardware architecture, and the effectiveness of their collaborative operation is crucial for servers of different form factors (such as rack-mount, blade, and edge servers) and configurations (such as different numbers of hard drives).

[0003] In existing technologies, the BMC and CPLD employ a deeply coupled design pattern. Functionally, they are tightly bound together, with the BMC's software logic and the CPLD's hardware logic interconnected through register definitions and communication protocols. For example, the BMC's power-on / power-down timing depends entirely on the control signals output by the CPLD, while the CPLD's hardware logic design must adapt to the BMC's communication interface requirements. In terms of development processes, BMC firmware development and CPLD hardware logic design must proceed synchronously; any change to one inevitably triggers joint testing of the other. Regarding hardware resource consumption, to meet the complex management needs of the BMC, the CPLD needs to integrate a large amount of logic resources, significantly increasing its device cost and power consumption. From a fault correlation perspective, faults in the BMC and CPLD can affect each other, creating single points of failure risk. In upgrade and maintenance, upgrading the BMC firmware requires simultaneous updates to the CPLD logic, and vice versa, leading to complex version management and requiring joint analysis of BMC software and CPLD hardware logs for fault diagnosis.

[0004] However, this existing deep coupling approach suffers from low development and maintenance efficiency and poor system reliability during server development, operation, and maintenance. Summary of the Invention

[0005] This application provides a decoupled control system and server for BMC and CPLD to solve the problems of insufficient system maintainability and operational reliability in the existing BMC and CPLD coupled architecture.

[0006] In a first aspect, embodiments of this application provide a decoupling control system for BMC and CPLD, including a power supply control module disposed on a server motherboard;

[0007] The power supply control module includes a logic control unit and a power switch unit.

[0008] The first input terminal of the logic control unit is used to obtain the node presence identification signal on the server motherboard, and the second input terminal is used to obtain the power supply enable signal of the baseboard management controller (BMC) output by the complex programmable logic device (CPLD).

[0009] The output of the logic control unit is connected to the enable terminal of the power switch unit, and is used to control the power switch unit to turn on when the node presence identification signal is valid and / or the BMC power supply enable signal is valid.

[0010] The input terminal of the power switch unit is used to connect to the standby power supply, and the output terminal of the power switch unit is used to output the standby voltage to the BMC board.

[0011] In one possible implementation, the power supply control module is configured such that when a CPLD malfunction causes the BMC power supply enable signal to become invalid, the logic control unit maintains the output of a drive signal to the power switch unit based on a continuously valid node presence identification signal, so that the BMC board continues to receive power when the CPLD malfunctions.

[0012] In one possible implementation, the system further includes an exception handling module integrated on the BMC board, which is used to acquire the working status information of the CPLD and, when it is determined from the working status information that the CPLD has abnormally hung up, to perform a reset process on the CPLD, wherein the reset process indicates that a hard reset or soft reset operation of the CPLD is triggered.

[0013] In one possible implementation, the exception handling module is further configured to: obtain the firmware status of the CPLD after the reset process; and when it is determined that the CPLD firmware is damaged based on the firmware status, perform firmware re-burning on the CPLD based on the CPLD firmware image stored on the BMC board.

[0014] In one possible implementation, the exception handling module is further configured to: determine whether the CPLD has returned to normal working state after the reset process; and generate and send a whole-machine power-off and power-on control command when the CPLD has not returned to normal working state, so as to trigger the whole-machine power-off and power-on operation or the server node hard reset operation.

[0015] In one possible implementation, the system also includes a status monitoring module integrated on the BMC board, which is used to periodically acquire the working status information of the CPLD; and when the working status information times out, has no response, or returns an abnormal value, it determines that the CPLD has abnormally hung up.

[0016] In one possible implementation, the system further includes a power-on control module mounted on the BMC board. The power-on control module includes a multi-stage DC-DC power supply chip, wherein the enable pins of the multi-stage DC power supply chips are cascaded sequentially. The power-on control module is configured to: after the first-stage DC power supply chip outputs a standby voltage and stabilizes, enable the second-stage DC power supply chip to output a first operating voltage after a first preset delay; after the first operating voltage stabilizes, enable the next-stage DC power supply chip to output a second operating voltage after a second preset delay, until the last-stage DC power supply chip outputs a target operating voltage; and after the target operating voltage stabilizes, generate and output a BMC reset signal after a third preset delay.

[0017] In one possible implementation, the system further includes a power-down control module, which includes a voltage detection chip and a target power timing controller. The voltage detection chip detects the main power supply voltage input to the entire machine and outputs an enable signal when the main power supply voltage drops to a preset threshold. The target power timing controller receives the enable signal and, based on the enable signal, outputs a first reset control signal on the first channel, generates a BMC reset signal after a delay, and resets the BMC chip. It also outputs a first power enable signal on the second channel to turn off the first standby power rail, and sequentially turns off the other low-voltage standby power rails after the first standby power rail has stabilized.

[0018] In one possible implementation, the voltage value of the first standby power rail is greater than that of the other low-voltage standby power rails.

[0019] Secondly, embodiments of this application provide a server, including a server motherboard, a CPLD, a BMC board, and a decoupled control system for the BMC and CPLD as described in the first aspect and / or various possible BMC and CPLDs in the first aspect; wherein the BMC board is a board in the form of an unbuffered dual in-line memory module (UDIMM) board, and is pluggably connected to the server motherboard through a UDIMM interface on the server motherboard.

[0020] This application provides a decoupling control system and server for BMC and CPLD. By setting up a power supply control module, the logic control unit obtains the node presence identification signal and the BMC power supply enable signal output by the CPLD. The two signals are used together as the basis for BMC power supply control. The power switching unit can be enabled and controlled according to the motherboard node status and CPLD control requirements, and the standby voltage is stably output to the BMC board. This makes the power supply of BMC no longer solely dependent on CPLD, thereby reducing the coupling between BMC and CPLD, improving development and maintenance efficiency and system operation reliability. Attached Figure Description

[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0022] Figure 1 A schematic diagram of the decoupling control system for BMC and CPLD provided in this application. Figure 1 ;

[0023] Figure 2 A schematic diagram of the specific structure of a decoupling control system for BMC and CPLD provided in this application;

[0024] Figure 3 A schematic diagram of the power-on control module provided in this application;

[0025] Figure 4 A schematic diagram of the power-off control module provided in this application;

[0026] Figure 5 A schematic diagram of the decoupling control system for BMC and CPLD provided in this application. Figure 2 .

[0027] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0028] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application.

[0029] It should be understood that the terms “comprising” and “having” and any variations thereof used in this application are intended to cover but not exclude inclusion. For example, a product or device that includes a series of components is not necessarily limited to those components that are explicitly listed, but may include other components that are not explicitly listed or that are inherent to such product or device.

[0030] As used in this application, the term "module" means any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code capable of performing the functions associated with that element.

[0031] In the server field, the BMC (Band Control Controller) handles out-of-band management, remote monitoring, and firmware updates, while the CPLD (Pin Locking Logic Controller) controls power-on sequencing, hardware reset, and pin multiplexing. Their coordinated operation is crucial for ensuring stable server operation. However, current technologies generally employ a deeply coupled design between the BMC and CPLD, which introduces problems in multiple dimensions.

[0032] In existing server solutions, the BMC and CPLD are designed with deep coupling. The CPLD outputs relevant control signals to the BMC, and the BMC then interacts with the CPLD according to preset register definitions and control flows to complete operations such as power supply control, status coordination, and startup management. Although this implementation can meet basic control requirements, the BMC software logic and the CPLD hardware logic are intertwined. When a function changes on one side, the other side often needs to be adjusted and re-verified synchronously.

[0033] Because the relevant control relationships are concentrated within the coupled architecture, practical applications are prone to problems such as long development iteration cycles, complex version matching, and unclear fault location boundaries. For example, when the BMC power supply control link depends on the CPLD output, if the CPLD logic is abnormal, misconfigured, or its state fails, it may further affect the normal power-on of the BMC board, thereby weakening its out-of-band management capabilities. At the same time, to adapt to different platforms and configurations, the coupled design may also lead to increased device resource consumption and maintenance costs, making it difficult to balance development efficiency and system reliability. It is evident that the existing deep coupling method between BMC and CPLD suffers from low development and maintenance efficiency and poor system reliability during development, operation, and maintenance.

[0034] To address the aforementioned issues, this application provides a decoupled control system and server for the BMC and CPLD. The system includes a power supply control module on the server motherboard, comprising a logic control unit and a power switch unit. The logic control unit has two independent input signal sources: a first input acquires a node presence identification signal from the server motherboard, generated by a hardware physical presence detection circuit, indicating whether a server node is present; the second input acquires a BMC power supply enable signal output from the CPLD. The logic control unit performs logical judgment on the two input signals. When at least one of the node presence identification signal and the BMC power supply enable signal is valid, it outputs a drive signal to control the power switch unit to conduct, allowing the standby power supply to be converted into standby voltage and output to the BMC board. This solution, by constructing a dual-input logic judgment mechanism with the node presence identification signal and the CPLD power supply enable signal running in parallel, prevents the BMC power supply from being solely controlled by the CPLD, thereby reducing the dependence of the BMC power supply process on a single control link and improving the system's operational reliability and ease of development and maintenance.

[0035] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0036] Figure 1 A schematic diagram of the decoupling control system for BMC and CPLD provided in this application. Figure 1 ,like Figure 1 As shown, the decoupling control system between the BMC and CPLD may include: a power supply control module mounted on the server motherboard; wherein the power supply control module includes a logic control unit and a power switch unit; the first input terminal of the logic control unit is used to acquire the node presence identification signal on the server motherboard, and the second input terminal is used to acquire the power supply enable signal of the baseboard management controller (BMC) output by the complex programmable logic device (CPLD); the output terminal of the logic control unit is connected to the enable terminal of the power switch unit, and is used to control the power switch unit to conduct when the node presence identification signal is valid and / or the BMC power supply enable signal is valid; the input terminal of the power switch unit is used to connect to the standby power supply, and the output terminal of the power switch unit is used to output the standby voltage to the BMC board.

[0037] In this embodiment, the power supply control module can be an independent functional circuit located on the server motherboard, used to replace the CPLD as the decision-making and execution core for power supply to the BMC. Its function is to aggregate enable signals from different physical sources and, according to preset hardware logic rules, ultimately control the supply of standby power to the BMC board. Furthermore, this module can be physically an independent and identifiable circuit unit, separated from the motherboard CPLD in both logical function and circuit wiring.

[0038] In some possible examples, the power control module can be an independent daughter card composed of discrete components, connected to the motherboard via headers and sockets, facilitating independent testing and replacement; it can also be an integrated circuit directly laid out and soldered onto a specific area of ​​the server motherboard; or it can be packaged as a dedicated power management chip to further simplify the motherboard layout.

[0039] Furthermore, the logic control unit is the decision-making core of the power supply control module, used to perform logical operations and judgments on multiple input enable signals and output the final power switch drive signal. The power switch unit is the execution core of the power supply control module, used to respond to the drive signal output by the logic control unit and execute the on and off actions of the power path. The two work together to form a complete "decision-execution" power supply control link, which is physically and functionally independent of the CPLD.

[0040] In some embodiments, the logic control unit can be implemented as: discrete logic gate circuits (such as AND gates and OR gate chips), small programmable logic devices, or combinational logic circuits built from diodes and transistors. The power switching unit can be implemented as: an integrated electronic fuse chip, discrete P-channel / N-channel power MOSFETs with driving circuits, or an integrated load switch chip. The specific selection can be configured according to the supply current, protection function requirements, and cost; this embodiment does not impose any restrictions.

[0041] The first input terminal can be a physical pin or electrical node on the logic control unit used to receive the first independent enable signal. This input terminal is dedicated to receiving a node presence identification signal. The node presence identification signal can be a physical status indication level generated by pure hardware circuitry, independent of any firmware and software, used to reliably characterize in real time whether the BMC board is correctly inserted and fixed on the motherboard connector. The validity of this signal is completely unaffected by the CPLD operating state and is one of the key signal sources for achieving power supply decoupling in this invention.

[0042] The second input is a physical pin or electrical node on the logic control unit used to receive a second independent enable signal. This input is dedicated to receiving the BMC power enable signal output by the CPLD. This signal is a control level generated by the internal logic of the CPLD on the motherboard based on the overall system operating status, representing the CPLD's software logic judgment result of the "allow BMC power-on" event. In this embodiment, this signal is downgraded from a single control source to one of the parallel input sources.

[0043] In some possible examples, the node presence identification signal can be generated in the following ways: a dedicated shorting pin is set on the BMC board, which shorts the pull-up detection line on the motherboard to ground when the board is inserted, generating a low-level active signal; an optical pair is set next to the connector, which blocks the light when the board is inserted, generating a state flip signal; or a micro switch is set at the end of the connector, which is mechanically triggered to turn on and off after the board is inserted, generating a signal.

[0044] The BMC power enable signal can be transmitted in two ways: directly outputting a standard level signal from the CPLD's general-purpose I / O port; or, when the CPLD output level does not match the logic control unit's input level, it can be converted by a dedicated level conversion chip before being connected.

[0045] The output terminal is the physical pin or electrical node on the logic control unit used to output the final drive signal. The enable terminal is a dedicated control pin on the power switching unit used to receive external control signals to turn the power path on or off. The two are in a drive-driven relationship.

[0046] In this embodiment, the logic control unit outputs a drive signal to turn on the power switch unit as long as the node presence identification signal is valid (indicating that the BMC board is physically present) or the BMC power enable signal is valid (indicating that the CPLD agrees to supply power). In one example, the logic control unit is preferably an AND gate logic device, which is configured to output a drive signal to control the power switch unit to turn on only when both the node presence identification signal and the BMC power enable signal are valid. This achieves dual hardware confirmation of BMC power supply, ensuring that the standby voltage is only output to the BMC board when both conditions are met simultaneously, providing higher power supply security for the system. When a CPLD malfunction causes the BMC power enable signal to become invalid, the node presence identification signal can still maintain the power switch unit's conduction, ensuring that the BMC does not lose power and continues to operate. This cuts off the path of CPLD failure affecting BMC power supply, allowing the BMC to continue operating without CPLD control, thus achieving hardware decoupling of the BMC and CPLD in the power supply link.

[0047] The input terminal is the physical pin or electrical node on the power switch unit that connects to the input power rail, used to connect to the standby power supply (such as P12V_STBY) provided by the server power module. The output terminal is the physical pin or electrical node on the power switch unit that connects to the load, used to output the standby voltage controlled by the switch to the BMC board. As the control switch in the entire power supply path, the power switch unit, when turned on, transfers the standby power from the input terminal to the output terminal with almost no loss, providing a stable standby voltage to the BMC board.

[0048] In some possible examples, standby power can come directly from the P12V_STBY copper strip next to the power connector on the motherboard, or from a separate standby voltage conversion circuit on the motherboard. Output to the BMC board can be achieved via motherboard wiring to the power pins of the BMC board connector, or via a wire-to-board connector in a cable connection scenario.

[0049] In some embodiments, the power supply control module described above also integrates an on-premises identification unit. This on-premises identification unit is a circuit that directly generates and transmits node on-premises identification signals, used to perform hardware detection of the physical on-premises status of the BMC board, and directly outputs the generated node on-premises identification signal to the first input terminal of the logic control unit.

[0050] In this embodiment, the presence identification unit, logic control unit, and power switch unit can be integrated into the same power supply control module. For example, they can be integrated on the same independent daughter card or located in the same adjacent area of ​​the motherboard. The presence identification unit generates a presence signal by detecting pin shorting, optical path obstruction, or switch triggering, and directly routes this signal to the first input terminal of the logic control unit. This integrated design reduces long-distance signal transmission between different boards or areas, lowers the risk of signal interference, and improves the modularity and maintainability of the entire power supply control function.

[0051] In one example, Figure 2 A schematic diagram of a decoupling control system for BMC and CPLD provided in this application is shown below. Figure 2As shown, the power supply control module on the motherboard side includes a logic control unit and a power switch unit. The logic control unit corresponds to the AND gate device in the figure, and the power switch unit corresponds to the EFUSE (Electronic Fuse) device in the figure. The logic control unit is configured with two independent input signals. One is the node presence identification signal output by the NODE0 presence identification unit, and the other is the BMC power supply enable signal output by the CPLD. The two signals are connected in parallel to the AND gate to complete the logic operation. The drive level output by the AND gate is directly connected to the enable terminal of the EFUSE. After the EFUSE is turned on, it outputs the P12V_STBY standby voltage to the back-end BMC board to complete the supply of the total standby power to the BMC board.

[0052] Furthermore, in a partially integrated implementation, the presence identification unit corresponding to the NODE0 presence identification can be integrated into the power supply control module. The presence identification unit determines the physical presence status of the BMC board by relying on hardware detection methods such as pin shorting, optical path obstruction, and switch triggering. It autonomously generates a node presence identification signal and directly transmits it to the first input terminal of the AND gate. The presence identification unit, AND gate, and EFUSE can be uniformly arranged in adjacent areas of the motherboard or integrated into the same daughter card, shortening the signal transmission path, reducing signal abnormality problems caused by electromagnetic interference, and realizing the modular integration of power supply control related hardware, which facilitates later assembly and fault diagnosis.

[0053] The entire hardware link has dual-channel signal redundancy drive capability. The AND gate device can rely on the continuous and effective node presence identification signal to output the drive level to maintain EFUSE conduction. Even if the CPLD is malfunctioning or the power supply enable signal fails, the hardware presence signal generated by the NODE0 presence identification unit can still independently ensure the continuous output of P12V_STBY, so that the BMC board can continue to power on and run without being affected by the CPLD failure. The isolation and decoupling of the BMC and CPLD operating states are achieved from the power supply head of the BMC board.

[0054] Furthermore, it should be noted that in some examples, the logic control unit and the power switching unit can be replaced with logic devices or power switching devices with equivalent control functions, without departing from the decoupled power supply control concept defined in this application.

[0055] The decoupled control system for BMC and CPLD provided in this application embodiment, during operation, sends the node presence identification signal on the server motherboard and the BMC power enable signal output by the CPLD to the logic control unit for logical combination. The control result drives the power switch unit to turn on or off the standby power path. When the node presence identification signal is valid, and / or the BMC power enable signal output by the CPLD is valid, the logic control unit outputs an enable control signal. Under the action of this signal, the power switch unit connects the standby power path to the BMC board, allowing standby voltage to be supplied to the BMC board, thus enabling the BMC to obtain standby power. Because the power supply control module combines the node presence identification signal and the CPLD enable state through the logic unit and implements power path control through the power switch unit, the BMC's power supply path no longer relies solely on the CPLD as a single control node. Therefore, it can change the rigid coupling relationship in the traditional architecture where the BMC's power supply must completely depend on the normal operation of the CPLD, achieving decoupled control between the BMC and CPLD.

[0056] Based on the above embodiments, the power supply control module is configured such that when the CPLD function malfunctions and the BMC power supply enable signal becomes invalid, the logic control unit maintains the output of the drive signal to the power switch unit based on the continuously valid node presence identification signal, so that the BMC board continues to receive power when the CPLD malfunctions.

[0057] This embodiment defines the fault-tolerant response behavior of the power supply control module in CPLD malfunction scenarios. CPLD malfunction refers to a state where the CPLD cannot perform its intended logic functions normally due to reasons such as firmware logic malfunction, clock failure, or power fluctuations. Examples include firmware corruption caused by unexpected power outages during CPLD firmware upgrades, internal state machine deadlock due to electromagnetic interference, or chip over-temperature triggering protection shutdown. In this case, the BMC power enable signal output by the CPLD is invalid due to a high impedance pin or being pulled low internally. Furthermore, an invalid state can mean that the level output by the CPLD to the second input terminal of the logic control unit no longer meets the valid determination condition. For example, if the system considers a high level as valid, a low level or high impedance output by the CPLD is considered invalid. This invalidity stems from a CPLD malfunction itself, rather than a normal system shutdown command.

[0058] Since the node presence identification signal is generated by a purely hardware presence detection circuit, it depends only on preset factors such as whether the BMC board is physically inserted, and is independent of the CPLD's operating state. Therefore, when a CPLD malfunction causes the second input signal to become invalid, the logic control unit can still rely on a continuously valid presence signal as the sole input source, keeping the output drive signal from flipping and ensuring the power switch unit remains seamlessly on. During CPLD malfunctions, the power switch unit remains on, the power supply path from standby power to the BMC board remains unobstructed, the BMC chip and its peripheral circuits do not lose power or restart, and the out-of-band management function is unaffected by a CPLD hangup.

[0059] In the above structure, once the node presence identification signal remains continuously valid, it indicates that the BMC board is still correctly plugged in and meets the power supply conditions. Even if the CPLD fails to output a valid BMC power supply enable signal due to an anomaly, the AND gate logic circuit control unit can still maintain the drive output of the power switch unit based on this continuously valid node presence identification signal, allowing the standby power supply to stably output standby voltage to the BMC board via the power switch unit. Therefore, the system will not cause the BMC to completely lose power due to a single control link interruption when the CPLD fails. The BMC can continue to maintain minimum management and monitoring capabilities, thereby improving the recoverability of the server motherboard under abnormal operating conditions and the continuity of remote maintenance. At the same time, since power supply maintenance mainly relies on the node presence signal, which is directly related to the board's plug-in status, a relatively independent relationship is formed between the BMC power supply path and the CPLD logic control. This improves power supply fault tolerance without increasing control complexity and reduces the risk of cascading failures caused by CPLD anomalies.

[0060] Furthermore, such as Figure 2 In the illustrated embodiment, the NODE0 presence identification signal and the CPLD control signal are input to the AND gate in parallel. The system can be further configured with a hardware bypass mechanism: when a CPLD anomaly is detected, the presence signal directly drives EFUSE to conduct through a pull-up resistor or a backup switch path, ensuring that as long as the BMC board is physically present, the power supply will not be interrupted, thereby cutting off the path of CPLD failure affecting the BMC power supply.

[0061] As can be seen from the above analysis, this implementation method can ensure continuous power supply when the BMC board is in normal position, while reducing the impact of CPLD failure on the BMC power-on link. That is, the BMC board does not lose power in the CPLD failure scenario, realizing complete decoupling between the power supply link and the CPLD operating state, ensuring that the out-of-band management function of the BMC can still operate normally during the CPLD hang, thereby improving the stability, maintenance convenience and out-of-band management continuity of the server hardware system.

[0062] Based on the above embodiments, the system also includes an exception handling module integrated on the BMC board, which is used to obtain the working status information of the CPLD and, when it is determined that the CPLD has abnormally hung up based on the working status information, to perform a reset process on the CPLD. The reset process indicates that a hard reset or soft reset operation of the CPLD is triggered.

[0063] In this embodiment, the anomaly handling module is the core of the software closed-loop mechanism for BMC reverse management and CPLD recovery, built upon the power supply decoupling described in the above embodiments. This mechanism empowers the BMC to proactively monitor, identify, and repair CPLD faults without being affected by the CPLD's power supply. Furthermore, the anomaly handling module can be a software functional unit running on the BMC board, manifested as an independent task thread, daemon process, or system service within the BMC firmware. Its functions cover a complete closed loop of status acquisition, anomaly determination, and reset repair. This module can operate using the BMC board's own processor, memory, and storage resources, with power supply guaranteed by the power supply control module, unaffected by the CPLD's state. Integration onto the BMC board means that the module runs locally on the server node, interacting directly with the CPLD via the board-level bus, resulting in low response latency and independence from network connectivity. In some embodiments, when the BMC board is a pluggable UDIMM (Unbuffered Dual In-line Memory Module) form, this module can be installed or replaced along with the BMC board.

[0064] Obtaining CPLD operational status information refers to collecting data reflecting whether the CPLD is operating normally, either through active polling or passive reception. Specific methods may include: the BMC periodically reading the CPLD's internal status register values ​​via the I2C or SPI bus; sending heartbeat requests and waiting for responses, using timeouts to determine if the CPLD is stuck; reading the CPLD firmware version number to verify its consistency; and collecting status feedback signals from peripherals controlled by the CPLD to indirectly infer whether the control logic is functioning correctly. This information can be obtained through existing management channels without requiring additional hardware.

[0065] An abnormal hangup refers to a frozen state in which the CPLD cannot respond to instructions or execute predetermined logic. The criteria for determining this state include: no response or invalid value returned for interactive instructions timeouts; inconsistent status register values; multiple consecutive failed heartbeat detections; and critical enable signals remaining at unexpected levels. Furthermore, the exception handling module can make judgments based on a combination of single or multiple rules.

[0066] Hard reset connects to the CPLD hardware reset pin via the BMC's GPIO pin, outputting a preset pulse level to trigger the chip's hardware reset logic. This restores all registers and state machines to their initial states and reloads the firmware, providing a thorough solution suitable for scenarios with complete logic deadlock. Soft reset writes a reset command word to a specific CPLD register via the communication bus, triggering an internal power-on reset or jumping to the firmware entry point for re-execution. This is a logic-level reset, requiring no external pin operation, and is faster. It is suitable for mild to moderate fault scenarios where the communication interface remains responsive but the internal logic is abnormal.

[0067] In some embodiments, a tiered reset strategy can be configured: first attempt a soft reset, and if the problem is not resolved, then perform a hard reset; and the number of automatic retries can be set (e.g., N automatic resets) to achieve automatic fault recovery without manual intervention.

[0068] By integrating an anomaly handling module onto the BMC board, the system proactively acquires CPLD operating status information without being affected by the CPLD's power supply. Upon determining an abnormal hang, it triggers a hard or soft reset for repair. Thus, the system constructs a closed-loop system for reverse monitoring and recovery of the CPLD by the BMC based on power supply decoupling. This achieves an architectural shift from the BMC being constrained by the CPLD to the BMC being able to proactively repair the CPLD, thereby improving the server motherboard's automatic recovery capability against CPLD failures and overall maintainability.

[0069] Based on the above embodiments, the exception handling module can also be used to: obtain the firmware status of the CPLD after the reset process; when it is determined that the CPLD firmware is damaged based on the firmware status, perform firmware re-burning on the CPLD based on the CPLD firmware image stored in the BMC board.

[0070] This embodiment further defines the exception handling module's deeper repair capabilities after a failed reset, namely firmware re-burning. In this embodiment, after the reset process, the CPLD undergoes a complete reboot process. At this time, the exception handling module re-establishes communication with the CPLD and reads its firmware-related information. The firmware status can refer to various indications reflecting whether the CPLD's current firmware is intact.

[0071] In some possible examples, the specific form of firmware status includes: CPLD power-on self-test results (CPLD verifies the firmware image in its internal flash memory during startup, and the verification result is reported through the status register); firmware version number integrity (if the read version number is all zeros or garbled characters, it indicates that the firmware is corrupted); and CPLD boot mode flag (if this flag indicates that the CPLD has entered safe mode or recovery mode, it means that normal firmware cannot be loaded). The above information can be obtained through the management bus between the BMC and the CPLD.

[0072] Understandably, since a reset operation can only repair temporary faults such as logic malfunctions or state machine deadlocks, it cannot restore the CPLD to normal function simply by resetting for permanent firmware damage caused by reasons such as firmware storage medium bit flipping, incomplete image due to unexpected power outages during firmware upgrades, or serious defects in the firmware logic itself. Therefore, the fault handling module analyzes the firmware status information. If it finds firmware verification failure, abnormal version number, or the CPLD entering recovery mode, it can determine that the root cause of the fault is firmware corruption rather than a temporary fault, thereby triggering the next level of repair measures.

[0073] In this embodiment, the BMC board pre-stores the official firmware image file of the CPLD, which can be stored in the BMC's onboard flash memory or an external storage chip. Firmware re-burning refers to the BMC using a communication bus to write the stored firmware image data into the CPLD's internal or external non-volatile storage medium according to the CPLD's burning protocol, replacing the corrupted firmware. In some possible examples, the burning interface can use bus protocols such as JTAG, SPI, or I2C; the burning process can follow the three-step process of erasing, programming, and verification specified by the CPLD manufacturer. Since the BMC's power supply is independently guaranteed by the power supply control module of claim 1, even if the CPLD is completely unresponsive during firmware corruption, the BMC can still operate normally and perform the re-burning operation.

[0074] In some embodiments, the BMC can also receive CPLD firmware files manually uploaded by the administrator through its out-of-band management interface (such as a web interface, IPMI, or Redfish) and trigger remote upgrade repairs without on-site operation, thereby expanding the flexibility and applicable scenarios of firmware repair.

[0075] By further acquiring the CPLD firmware status after the reset process and re-burning the CPLD using the official firmware image stored locally on the BMC board when firmware corruption is detected, a progressive fault recovery chain from reset repair to firmware re-burning is constructed. This improves the recovery success rate in firmware anomaly scenarios and reduces the risk of BMC power supply control and system management link interruption due to CPLD firmware corruption. As a result, the system can distinguish between temporary and permanent root causes of CPLD faults. For firmware corruption problems that cannot be resolved by reset, firmware recovery can be automatically completed under the guarantee of independent BMC power supply, avoiding situations where the entire machine cannot boot or requires on-site repair due to CPLD firmware anomalies, further improving the server's automatic repair capabilities and operational efficiency.

[0076] Based on the above embodiments, the exception handling module can also be used to: determine whether the CPLD has returned to normal working state after the reset process; when the CPLD has not returned to normal working state, generate and send a whole machine power-off and power-on control command to trigger the whole machine power-off and power-on operation or the server node hard reset operation.

[0077] This embodiment further defines the recovery strategy of the fault handling module when the reset operation fails to restore the CPLD, namely, a power restart at the whole machine or node level. In this embodiment, after the reset operation is completed, the fault handling module re-attempts to establish communication interaction with the CPLD, collects its working status information, and compares it with the expected value under normal conditions. In some examples, the criteria for restoring normal working conditions may include: the CPLD can respond normally to the BMC's read and write commands; the values ​​of each status register of the CPLD are restored to the expected default values; the peripherals controlled by the CPLD are restored to the correct control level; and the CPLD heartbeat signal or status reporting is restored to the normal cycle. If all the above indicators are met, the reset is considered successful, and the fault recovery process ends; if any one or more of the above indicators are continuously abnormal, the reset is considered invalid, and the repair method needs to be upgraded.

[0078] It is understood that this exception handling module can be implemented through a fault recovery process in the BMC firmware, an onboard state machine, a dedicated monitoring circuit, or a combination thereof. It can be connected to the system management bus, reset control interface, and power control interface connected to the BMC. This module is used to escalate to a higher-level recovery action after a CPLD partial reset failure, thereby clearing any residual abnormal register states, latch fault states, or timing mismatch states. In some examples, this exception handling module can be located in the motherboard power control area, the node reset control area, or the management interface area between the BMC and the CPLD.

[0079] Furthermore, the power-on / off control command for the entire server can be a hardware control signal issued by the BMC to control the switching on and off of the server's AC power supply. This command can act on the power management bus or the enable control pin of the power module on the server motherboard. In one example, this control command can be either a single pulse trigger signal or a management message containing an enable bit, duration parameters, and a target node identifier, so as to accurately execute the corresponding recovery operation on the entire server, server nodes, or board-level power domain.

[0080] The power-off and power-on operation, also known as an AC cycle, refers to a complete power-off and power-on process for all chips and modules in the server by disconnecting and reconnecting the AC input power. This method allows the CPLD to undergo a complete hardware power-off reset, clearing all volatile states and reloading the firmware from non-volatile memory. It has a thorough repair effect on latch-up effects or deep state machine jams caused by power fluctuations.

[0081] A server node hard reset refers to triggering a system-wide reset at the server node level via a hardware reset signal without disconnecting the AC power supply. This includes synchronous resets of the CPU, CPLD, and other peripheral chips, simulating a complete system restart. This method is faster than powering on and off the entire system and is suitable for scenarios where the CPLD can be recovered via a global reset signal.

[0082] In some possible implementations, the BMC can control the PS_ON signal of the power timing controller or power module on the server motherboard via GPIO pins to achieve power-on / off of the entire machine; alternatively, it can trigger a hard reset of the node by controlling the node reset logic circuit. Furthermore, the exception handling module can configure these two methods as a progressive strategy: first attempt a hard reset of the node; if that fails, then perform a power-on / off operation of the entire machine. In addition, the entire fault recovery process, from reset to re-burning and then to restarting the entire machine, can be executed automatically and recorded in the system event log without manual intervention. For example, the BMC records the time the CPLD hangs, the triggering scenario, and interactive exception information in the system event log, such as the SEL (System Event Log), to facilitate subsequent fault tracing. Some servers also save a snapshot of the CPLD's registers at the time of the hang to assist in locating the root cause.

[0083] By verifying whether the CPLD has returned to normal after the reset process, and actively triggering the whole machine power-off and power-on or node hard reset operation if it has not returned to normal, the fault state that cannot be eliminated by partial reset can be forcibly cleared, thereby improving the probability of successful recovery in CPLD abnormal scenarios and reducing the risk of BMC coordination failure, power supply link lock-up or system management unavailability caused by continuous CPLD abnormality. This also avoids long-term server downtime due to CPLD hang-up, further improving the system's operational reliability and automated operation and maintenance capabilities.

[0084] Based on the above embodiments, the system may also include a status monitoring module, integrated on the BMC board, for periodically acquiring the working status information of the CPLD; and when the working status information times out, has no response, or returns an abnormal value, it determines that the CPLD has abnormally hung up.

[0085] This embodiment further establishes an active monitoring mechanism for the CPLD's health status, providing a fault detection trigger source for the aforementioned anomaly handling closed loop. In this embodiment, the status monitoring module can be a software functional unit running on the BMC board, which can be manifested as a periodic scheduling task or background service in the BMC firmware. Its core function is to periodically initiate communication interactions with the CPLD, collect working status information, and perform health determination accordingly. The power supply for this module is guaranteed by the power supply control module, so even if the CPLD has entered an abnormal state, the monitoring function itself is not affected.

[0086] Periodically acquiring the CPLD's operational status information means that monitoring is performed cyclically at preset time intervals, which can be configured from several seconds to tens of seconds, to strike a balance between timely fault detection and communication overhead. Methods for acquiring operational status information include: the BMC sending status read commands to the CPLD via the I2C or SPI bus; sending heartbeat request packets and waiting for a response; and reading the real-time values ​​of specific registers within the CPLD. These operations reuse the existing management bus between the BMC and the CPLD, requiring no additional hardware connections. In some examples, operational status information may include various types of information such as the CPLD version, register status, and control status of various peripherals (e.g., fans, power supplies, buttons, indicator lights, etc.).

[0087] In this context, "timeout" refers to the BMC sending a communication request to the CPLD but receiving no response within a preset time threshold, indicating that the CPLD communication interface has stopped working. "No response" means that although the CPLD physically responds to the communication request, the returned data is an empty or invalid frame, indicating that its logical processing capability has been lost. "Returning an abnormal value" means that the data returned by the CPLD is formatted correctly, but its content deviates significantly from the expected value. For example, the status register value is all 0xFF or all 0x00, or the firmware version number read result does not match the known value, indicating that the CPLD's internal logic may be disordered. When any of the above conditions are triggered, the status monitoring module can determine that the CPLD is abnormally hung and generate an alarm event. This alarm event can serve as a direct input to trigger the fault handling module to perform a reset operation, or it can be independently recorded in the system event log and reported to the administrator through the out-of-band management interface, ensuring that the fault is detected and recorded as soon as possible.

[0088] In some embodiments, the status monitoring module can also continuously count the results of each interaction, and only confirm the hangup when an anomaly is determined multiple times in a row, so as to avoid misjudgment caused by momentary communication interference.

[0089] In practical implementation, this module can be implemented through software-timed tasks and firmware monitoring processes on the BMC board, or through independent watchdog logic, a dedicated status acquisition chip, or external auxiliary monitoring circuitry. The module is typically installed close to the BMC main control chip or the board's communication interface area to shorten signal paths and reduce the impact of bus impedance on sampling stability. In terms of form, this module can further manifest as a software monitoring function, a firmware background daemon thread, a hardware timeout comparator, or an integrated status monitoring circuit; the specific implementation can be selected based on server platform resources and reliability requirements.

[0090] By integrating a status monitoring module onto the BMC board, the operating status information of the CPLD is periodically acquired. When timeouts, no responses, or abnormal values ​​are returned, it is determined that the CPLD has abnormally hanged, thus establishing a continuous and proactive detection mechanism for the CPLD's operational status. With independent power supply to the BMC, this module can reliably and promptly detect CPLD faults, providing accurate trigger conditions for automatic reset and repair processes. This avoids the risk of cascading failures caused by prolonged undetected CPLD hangs, further improving the system's maintainability and operational reliability.

[0091] Based on the above embodiments, the system may further include a power-on control module, which is mounted on the BMC board. The power-on control module includes a multi-stage DC-DC power supply chip, wherein the enable pins of the multi-stage DC power supply chips are cascaded sequentially. The power-on control module is configured to: after the first-stage DC power supply chip outputs a standby voltage and stabilizes, enable the second-stage DC power supply chip to output a first operating voltage after a first preset delay; after the first operating voltage stabilizes, enable the next-stage DC power supply chip to output a second operating voltage after a second preset delay, until the last-stage DC power supply chip outputs a target operating voltage; and after the target operating voltage stabilizes, generate and output a BMC reset signal after a third preset delay.

[0092] In this embodiment, the power-on control module enables the BMC board to autonomously complete the step-by-step power-on and reset timing of multiple operating voltages after obtaining standby voltage, relying on onboard hardware circuitry, without the intervention of the CPLD or any external controller. Furthermore, this power-on control module is an independent hardware circuit unit on the BMC board responsible for generating and managing the power-on sequence and reset timing of each power rail of the BMC chip. Its location on the BMC board clarifies that this module resides on the same physical carrier as the BMC chip, achieving spatial and electrical separation from the motherboard CPLD. The module's power supply comes from the standby voltage output by the power switching unit, and its operation does not depend on any control signals from the CPLD. In a specific implementation, the power-on control module can be a power timing control unit used to generate multiple working voltages on the BMC board in a preset timing sequence and output a BMC reset signal after the target working voltage stabilizes. Its function is to convert the input standby power supply into multiple voltages suitable for the operation of each functional circuit of the BMC through a multi-stage DC-DC power supply chip, and provide reset control to the BMC after the voltage establishment meets the requirements, so that the BMC can complete the startup in a stable timing sequence during the power-on process.

[0093] In this context, a DC power supply chip can refer to an integrated circuit that converts an input DC voltage to another DC voltage level, such as converting a P3V3_STBY output to a P1V2_STBY output. Multi-stage refers to using at least two stages of DC power supply chips to meet the power supply requirements of the BMC chip, which needs multiple power rails with different voltages. In some embodiments, the DC power supply chip may take the form of an integrated synchronous buck converter, a low-dropout linear regulator, or a power module.

[0094] Furthermore, the enable pin can be a dedicated pin on the DC power supply chip used to receive external control signals to turn the output on or off. Cascading refers to the process where the power-good indicator output or output voltage of the preceding DC power supply chip, after being processed by a delay circuit, is connected to the enable pin of the following DC power supply chip, forming a hierarchical hardware enable chain. This cascading method pushes the power-on timing control logic down from traditional CPLD software programming to chip hardware pin interconnection, achieving pure hardware-based timing self-driving. The standby voltage can be the first voltage rail generated after the power-on control module receives power from the power switching unit, such as P0V8_STBY and P1V15_STBY. The first preset delay is used to ensure that the first-stage power supply output has reached a stable state; this delay can be achieved through an RC delay circuit, the chip's built-in soft-start time, or a dedicated delay device. In some examples, the first operating voltage can be P3V3_STBY and P1V8_STBY, the second operating voltage can be P1V2_STBY, and subsequent power rails such as P1V1_STBY can be generated. The second preset delay between each stage ensures that each power supply is fully stabilized before the next stage starts. The target operating voltage is the last power rail required by the BMC chip, and its voltage value and timing position are determined according to the specifications of the specific BMC chip model.

[0095] The BMC reset signal can be a hardware signal used to release the BMC chip from the reset state and enable it to begin executing firmware. A third preset delay ensures that the BMC reset state is released only after all power rails have stabilized, preventing abnormal BMC startup due to unprepared power. This reset signal can be generated by an independent delay circuit.

[0096] Based on the above analysis, it can be seen that by cascading the enable pins of multi-stage power chips, this module can divide the power-on process into several interconnected stable stages, thereby avoiding voltage drops, power-on jitter, or startup failures caused by instantaneous loading of a single-stage power supply.

[0097] In some embodiments, Figure 3 A schematic diagram of the power-on control module provided in this application is shown below. Figure 3As shown, the first-stage DC chip outputs P0V8_STBY and P1V15_STBY. After a stabilization delay of more than 1ms, the second-stage DC chip is enabled to output P3V3_STBY and P1V8_STBY. After the second stage stabilizes, a single DC chip is enabled to output P1V2_STBY. After the stabilization delay of this track exceeds 1.7ms, the next stage is enabled to output P1V1_STBY. After the final stage stabilizes, a further delay of more than 1ms triggers an independent delay line to generate a BMC_SRST reset signal.

[0098] By incorporating a power-on control module on the BMC board, which includes multi-stage DC power supply chips and cascaded enable pins, the purely hardware-driven, step-by-step power-on and reset timing of the BMC's multiple power rails is achieved. Consequently, after obtaining standby voltage, the entire power-on process of the BMC board is entirely self-driven by the onboard hardware circuitry, requiring no enable or timing control signals from the CPLD. This fundamentally eliminates the BMC's dependence on the CPLD for power-on timing, making the BMC board an independently power-on and operating unit. This further deepens the decoupling between the BMC and the CPLD, thereby improving the controllability of the BMC board's power-on timing, power supply stability, and overall system reliability.

[0099] Based on the above embodiments, the system further includes a power-down control module, which includes a voltage detection chip and a target power timing controller. The voltage detection chip is used to detect the main power supply voltage input to the whole machine and output an enable signal when the main power supply voltage drops to a preset threshold. The target power timing controller is used to receive the enable signal and, based on the enable signal, output a first reset control signal on the first channel, generate a BMC reset signal after delay processing to reset the BMC chip, and output a first power enable signal on the second channel to turn off the first standby power rail. After the first standby power rail stabilizes after power-down, it sequentially turns off the other low-voltage standby power rails.

[0100] This embodiment further defines the power-down timing control scheme for the BMC board and the entire system, enabling the system to autonomously complete the complete power-down timing sequence of BMC reset and graded shutdown of power rails based on dedicated hardware circuits after detecting a main power drop, without the need for CPLD to participate in any power-down detection, judgment, or control. Furthermore, this power-down control module can be a hardware circuit unit independent of the CPLD, specifically responsible for timing management in power-down scenarios. Its function is to completely take over all the power-down timing functions traditionally performed by the CPLD, including voltage drop detection, timing graded delay, and power rail enable control, thus completely decoupling the power-down link from the CPLD. In some embodiments, this module can be installed on the power input side or power management area of ​​the BMC board. The voltage detection chip is preferably located near the main power input path of the entire system to directly sample the main power bus voltage. The target power timing controller is connected to the voltage detection chip, the BMC reset network, and the power enable network through wires, PCB traces, or control ports, forming a unified output interface for reset control signals and power shutdown signals.

[0101] The main power supply voltage refers to the main power supply voltage of the server after AC-to-DC conversion. The voltage detection chip integrates a voltage comparator and a reference voltage source. When it detects that the main power supply voltage has dropped below a preset threshold, it indicates that the system is experiencing a power outage, and the chip immediately outputs an enable signal. The chip's detection and output actions are entirely performed in hardware, with a response time in the microsecond range. In some possible examples, the voltage detection chip can be a dedicated voltage monitoring chip or a reset chip, and the preset threshold can be configured through an external resistor divider network.

[0102] The target power timing controller is the timing execution unit in the power-down control module. It receives the enable signal from the voltage detection chip and outputs multiple control signals according to preset timing logic. This device is a hardware replacement for the original CPLD power-down timing function and can integrate multiple delay timers and output drivers internally. In some possible examples, the target power timing controller can be constructed using a dedicated power timing management chip, a programmable delay device, or a discrete timer circuit.

[0103] Furthermore, the first reset control signal can be a control output generated by the target power timing controller, processed by an RC delay circuit or an internal delay module of the chip, to form a reset pulse signal that meets the timing requirements of the BMC chip. In this reset mechanism, the BMC chip can be placed in a reset state before the power rail begins to shut down, stopping all read / write operations and logic processing, preventing the BMC from executing incorrect instructions or corrupting stored data due to voltage instability during power drops. This delay duration can be configured according to the BMC chip specifications; for example, a delay exceeding 1 millisecond can be used to generate the reset signal.

[0104] The first standby power rail can be a standby power rail with a higher voltage level in the system, such as P3V3_STBY. The target power sequencer outputs a first power enable signal, directly controlling the enable pin of the DC power chip or the power switch of this power rail to turn off its output. The low-voltage standby power rails are other power rails with lower voltage levels than the first standby power rail, such as P1V8_STBY, P1V2_STBY, P1V1_STBY, P0V8_STBY, and P1V15_STBY. In some examples, power-down stability can be guaranteed by detecting a good power signal from the first standby power rail or by setting a sufficient delay. Sequential shutdown can refer to shutting down power rails in a preset order, which can be implemented through cascaded delay logic within the target power sequencer. The design of shutting down the high-voltage power rails first, followed by the low-voltage power rails, avoids abnormal current paths when the low-voltage power rails fail first while the high-voltage power rails remain, protecting the BMC chip and other load devices.

[0105] In some embodiments, when the system starts up, the voltage detection chip continuously monitors the main power supply voltage input to the entire machine and compares the sampling result with a preset threshold. When the main power supply voltage drops below the threshold, the voltage detection chip immediately outputs an enable signal, and the target power timing controller enters a power-down control state accordingly. On one hand, the target power timing controller first outputs a first reset control signal on the first channel, and generates a BMC reset signal after a preset delay, causing the BMC chip to enter a reset state, thereby preventing the BMC from continuing to execute incomplete management logic during the power drop. On the other hand, the target power timing controller outputs a first power enable signal on the second channel to shut down the first standby power rail, causing the load on the power rail to lose power first and complete natural discharge or controlled discharge. After confirming that the first standby power rail has stabilized after power loss, the other low-voltage standby power rails are shut down in a predetermined order. This control process separates the BMC reset from the standby power rail shutdown timing, and ensures that each power rail is released in a stable sequence through threshold detection and delay control. This reduces power cross interference, reset timing disorder and residual voltage retention problems that occur when the main power supply drops abnormally, thereby improving the shutdown safety of the BMC board and the overall system reliability in power failure scenarios.

[0106] Furthermore, in one example, Figure 4 A schematic diagram of the power-down control module provided in this application is shown below. Figure 4As shown in the figure, when the voltage detection chip detects a main power supply drop, it outputs an enable signal to the target power supply timing controller. The first output of this controller generates a BMC_SRST reset signal after a delay to reset the BMC. The second output cuts off the 3V3 standby power supply, and after its power supply drops and stabilizes, the remaining five low-voltage standby power supplies (P1V8_STBY, P1V2_STBY, P1V1_STBY, P0V8_STBY, and P1V15_STBY) are cut off in sequence. The entire process is jointly completed by the voltage detection chip and the target power supply timing controller, and the CPLD does not participate in any link.

[0107] By constructing an independent power-off control module composed of a voltage detection chip and a target power supply timing controller, pure hardware timing control of BMC reset and power rail hierarchical sequential power-off in the whole machine power-off scenario is achieved. Thus, all functions of power-off detection, timing delay, and power supply cut-off are taken over by this dedicated hardware circuit, completely decoupling the power-off link from the CPLD, avoiding the risk of power-off timing chaos caused by CPLD abnormalities, ensuring the data security of the BMC chip during power-off, and further strengthening the integrity of the decoupling architecture between the BMC and the CPLD.

[0108] Based on the above embodiment, the voltage value of the first standby power rail is greater than that of other low-voltage standby power rails.

[0109] In this embodiment, the voltage value of the first standby power rail is greater than that of other low-voltage standby power rails, defining the relative voltage level relationship between power rails in the power-off timing. The first standby power rail is the first power rail to be cut off in the above power-off timing, and its nominal voltage value is relatively the highest among the standby power rails of the BMC board. The other low-voltage standby power rails are the remaining power rails that are cut off in sequence after the first standby power rail in the power-off timing, and their respective nominal voltage values are lower than that of the first standby power rail.

[0110] In a specific embodiment, referring to Figure 4 the shown structure, the first standby power rail corresponds to P3V3_STBY with a nominal voltage of 3.3V; the other low-voltage standby power rails correspond to P1V8_STBY (1.8V), P1V2_STBY (1.2V), P1V1_STBY (1.1V), P0V8_STBY (0.8V), and P1V15_STBY (1.15V) respectively. The voltage values of the above low-voltage standby power rails are all less than 3.3V.

[0111] The sequential design of first shutting down the high-voltage power rail and then the low-voltage power rail serves a clear electrical protection purpose. Within the BMC chip, there are ESD protection diodes or parasitic diode structures between different power domains. When the high-voltage power rail is still energized while the low-voltage power rail is de-energized, a forward conduction path may form from the high-voltage domain to the low-voltage domain through these diodes, generating uncontrolled leakage current or latch-up effects, potentially damaging the chip. By prioritizing the shutdown of the highest voltage power rail, the source of this abnormal current path across power domains can be cut off, ensuring the safe and reliable shutdown process of the subsequent low-voltage power rails.

[0112] In some possible implementations, the first standby power rail is not limited to 3.3V, as long as its voltage value is the highest relative to other standby power rails on the BMC board. For example, in some BMC chip power supply schemes, the first standby power rail may be 5V or 2.5V, and the voltage values ​​of other low-voltage standby power rails are correspondingly lower than that value.

[0113] By limiting the voltage of the first standby power rail to be greater than that of all other low-voltage standby power rails, the hardware protection sequence of step-by-step shutdown from high voltage to low voltage during power-down is clearly defined. Therefore, prioritizing the disconnection of the highest voltage power rail during system power-down effectively blocks abnormal leakage current paths caused by potential differences between high and low voltage domains, avoiding potential electrical damage to the BMC chip and other load devices. This further enhances the electrical safety and system reliability during power-down, and helps improve the operational reliability and maintenance convenience of the server in complex power supply scenarios.

[0114] As can be seen from the above embodiments, the decoupled control system of BMC and CPLD in this application embodiment achieves functional decoupling between BMC and CPLD through hardware timing control, independent power management, and modular architecture design, thereby improving the flexibility, reliability, and maintainability of the server system. Specifically, by separating the power-on / power-off timing control of BMC from CPLD, hardware-level power timing circuits (such as multi-level DC-DC chip cascade, voltage monitoring chips, and power timing controllers) are used to replace the logic control function of CPLD, and physical isolation between BMC board and CPLD is achieved through independent power supply paths (such as EFUSE control circuits), ultimately constructing a completely decoupled architecture between BMC and CPLD in terms of function, power, and fault domain. This solution separates the management function of BMC from the hardware logic control function of CPLD, and achieves independent operation and upgrades of both through hardware redundancy design and modular interfaces, thereby solving the core problems of low development efficiency, complex maintenance, and poor reliability in traditional coupled architectures. In practical applications, this system can be used in general-purpose servers in scenarios such as data centers, cloud computing, and edge computing. Through decoupling design, the BMC and CPLD can be developed, upgraded, and maintained independently, adapting to different hardware platforms and application scenarios to meet the needs of modern servers for flexibility, reliability, and scalability.

[0115] The following is based on Figure 5 Taking an example, the structure and working principle of the decoupling control system of BMC and CPLD in this application embodiment are further explained, such as... Figure 5 As shown, the system is divided into five functional units: power supply control module, power-on control module, power-off control module, anomaly handling module, and status monitoring module. Each module is divided into a motherboard-side power supply management unit and a BMC board-side timing and maintenance unit according to the hardware deployment location. The modules have independent division of labor and the fault domains are isolated from each other. The BMC and CPLD are decoupled from each other in the entire chain from power supply head, power-on and power-off timing, and anomaly maintenance.

[0116] The power supply control module, marked with a dashed box, is the core power supply isolation unit on the motherboard side of this system. It consists of a logic control unit and a power switch unit cascaded together, and is specifically responsible for managing the P12V_STBY standby main power supplied to the BMC board. The power-on control module and power-off control module are deployed inside the BMC board, and independently carry out the timing control logic for the BMC chip's step-by-step power-on and step-by-step power-off, without the CPLD participating in the timing scheduling. The exception handling module and status monitoring module are also integrated into the BMC board, and are responsible for collecting the CPLD's operating status in real time, identifying CPLD hang-up faults, and performing graded repair operations. The entire system severs the hardware and software binding relationship between the BMC and CPLD through a layered hardware architecture.

[0117] Furthermore, it can be understood that this system achieves decoupling design of the BMC power-on / off timing from the CPLD from the power timing perspective. The complete power-on and power-off timing of the BMC board is completely independent of the control of the CPLD and PMIC, relying on independent hardware circuits to autonomously meet the stringent timing constraints of the BMC chip. Figure 3 As shown, the power-on process relies on a power-on control module equipped with a multi-stage DC power supply chip. Through cascading chip enable pins and hardware delay circuitry, the voltage is output in a step-by-step manner, generating standby voltages according to a preset timing sequence and outputting a BMC reset signal after a delay. Figure 4 As shown, the power-down process relies on a power-down control module consisting of a voltage detection chip and a dedicated power timing controller. It independently completes voltage drop detection, BMC reset, and multi-rail power supply staged shutdown by hardware. The entire power-up and power-down sequence process does not require any control commands from the CPLD, eliminating the dependency between the two at the chip power supply timing level.

[0118] At the overall system architecture level, refer to Figures 1 to 5 As shown, the power supply control module relies on the node presence identification line and the CPLD dual-signal parallel control of the EFUSE power switch to achieve complete isolation of the BMC board and the motherboard CPLD operating environment, and independent fault domains. The node presence identification signal and the power enable signal output by the CPLD are jointly connected to the logic control unit. As long as the node presence identification signal remains valid, the power switch unit can be independently driven to continuously output the P12V_STBY standby voltage to supply the BMC board. Even if the CPLD experiences anomalies such as hang-up or logic failure, the BMC board can still continue to operate stably and will not lose power due to CPLD failure. In contrast, when the BMC itself fails, restarts, or undergoes firmware upgrades, the front-end EFUSE power supply path remains continuously conductive. The CPLD is not affected by BMC-side operations and can continue to perform basic server hardware management functions, stably maintaining core business operations such as overall machine power-on and basic fan speed adjustment, avoiding interruption of the main business process, and significantly improving the overall stability of the server operation.

[0119] Furthermore, the status monitoring module integrated into the BMC board undertakes the online health inspection function of the CPLD. The BMC will periodically interact with the CPLD registers to read the CPLD firmware version, hardware peripheral control status and other operating information in real time. Once there is an interaction timeout, no response, or abnormal parameter return, it can be determined that the CPLD has abnormally hung up, and then the abnormal handling module will be triggered to execute a layered progressive fault repair strategy.

[0120] The first-level repair strategy is to reset the CPLD. The BMC directly drives the CPLD hard reset or soft reset pin, causing the CPLD to reload its internal logic program. Most transient logic freezes and temporary interactive faults can be restored to normal through reset. At the same time, the system supports configuring the number of automatic resets, and no manual intervention is required throughout the process.

[0121] If the CPLD still fails to recover after a reset, the system executes a second-level repair strategy, performing a firmware re-burning repair. The BMC locally stores a complete official CPLD firmware image, which can be automatically downloaded to complete the firmware re-burning and repair CPLD failures caused by upgrade power outages or logical image corruption. Administrators can also remotely upload firmware images through the BMC out-of-band management interface to achieve remote, non-site upgrade repair.

[0122] If the CPLD fault persists after firmware re-burning, the system will execute the third-level repair strategy, whereby the BMC will issue a power-off or power-on command for the entire machine or a hard reset command for the server node. This will clear the underlying abnormal state and restore the CPLD's basic operational capabilities through a hardware reboot.

[0123] Furthermore, in terms of full lifecycle management of faults, the BMC fully records CPLD abnormal event information, writing the fault occurrence time, abnormal triggering scenario, and interactive error details into the system event log. For severe hang scenarios, a snapshot of the CPLD register at the time of the fault is also retained, making it easier for R&D and maintenance personnel to trace and locate the root cause of the fault. At the same time, the BMC synchronizes the CPLD health status to standardized out-of-band interfaces such as the Web management interface, IPMI, and Redfish in real time. Administrators can intuitively view CPLD health alarms without logging into the server operating system and identify potential hardware faults in advance.

[0124] In addition, this decoupled control system boasts flexible hardware architecture expansion capabilities. The BMC can be separated from the server motherboard using an independent UDIMM board form factor, equipped with a universal UDIMM interconnect interface. While optimizing the internal space layout of the chassis, it is compatible with multiple CPLDs with different logic configurations, and adapts to server products with multiple generations of CPU platforms, diverse hard drive configurations, and different rack configurations. Independent iteration and upgrades are achieved simultaneously at both the hardware and software levels. The BMC firmware and the internal logic of the CPLD are not version-bound, and both can independently complete version updates and function iterations without needing to simultaneously adapt to the other's program, significantly shortening the adaptation cycle for new hardware development and version upgrades.

[0125] The decoupling control system for BMC and CPLD provided in this application embodiment achieves deep decoupling of BMC and CPLD hardware and software from multiple dimensions, including power supply control module including logic control unit and power switch unit on the motherboard side, and power-on control module, power-down control module, status monitoring module and abnormal handling module independently deployed on BMC board. It has multiple technical advantages.

[0126] The system adds a node presence identification line to the motherboard power supply link, which, together with the CPLD output signal, completes the EFUSE conduction control. The two drive the P12V_STBY standby voltage to supply the BMC board in parallel. The power switching unit can be kept on by relying solely on a continuously effective node presence identification signal. Even if the CPLD hangs up, the BMC board can still obtain power stably and work normally. At the same time, in accordance with the inherent timing specifications of the BMC chip, an independent timing hardware circuit is built on the BMC board. The power-on link uses a multi-stage DC power chip cascaded with hardware delay circuits to achieve step-by-step power supply. The power-down link relies on a voltage detection chip and a dedicated power timing controller to complete voltage drop detection, BMC reset, and multi-power rail staged shutdown. The entire power-on and power-down process does not require the participation of CPLD and PMIC for control. The complete timing requirements of the BMC can be met by the hardware circuit alone. In terms of hardware structure, it adopts a UDIMM board-type independent BMC design, which is physically separated from the server motherboard. This optimizes the internal space layout of the chassis and can also match CPLDs with different logic configurations, making it compatible with multiple CPU platforms, hard drive configurations, and different types of server products.

[0127] Based on the aforementioned layered and decoupled architecture, the solution effectively improves product development and iteration efficiency. The BMC primarily handles out-of-band management software functions such as remote monitoring, firmware updates, and fault alarms, with a high frequency of iterations. The CPLD focuses on underlying hardware logic control, including power-on sequencing, hardware reset, and fan speed regulation. The development processes of the two are independent, allowing for parallel design iterations. Adjustments to a single module's functionality do not require a complete refactoring of the other module; only the corresponding module needs to be updated, effectively shortening the new product development and market launch cycle. Simultaneously, the architecture possesses strong design reusability, enabling one BMC hardware set to adapt to multiple versions of CPLD logic. This eliminates the need to repeatedly build entire control schemes for different server products, reducing redundant design workload and improving design flexibility.

[0128] This solution also optimizes overall system maintenance capabilities and reduces maintenance costs. The BMC firmware and CPLD logic support independent upgrades; updating BMC monitoring functions and fixing management vulnerabilities does not require modifying the CPLD logic, and adjusting CPLD hardware timing logic does not require reflashing the BMC firmware, resulting in lower upgrade risks. Hardware fault domains are isolated from each other, allowing for rapid identification of fault attribution when server anomalies occur, clearly determining whether the fault originates from the BMC out-of-band management module or the CPLD underlying hardware control module. This avoids the problem of mutual interference between faults in a coupled architecture, significantly shortening troubleshooting time.

[0129] Furthermore, the overall system reliability is significantly enhanced, mitigating the risk of single-point failures in a coupled architecture. The BMC and CPLD operate independently, and their failures do not affect each other. Even if the BMC experiences firmware anomalies or requires restart repair, the CPLD can still continuously perform low-level hardware management operations such as server power-on and basic fan speed control, ensuring uninterrupted server services and improving overall system stability. In addition, the solution reduces overall hardware resource consumption and material costs. Complex software functions such as remote monitoring and fault management are uniformly implemented by the BMC, which integrates processor, memory, and storage resources. The CPLD only retains basic hardware timing and control logic, allowing for the use of smaller, lower-cost devices with smaller logic resources, while also reducing CPLD power consumption.

[0130] This application also provides a server, including a server motherboard, a CPLD, a BMC board, and a decoupled control system for the BMC and CPLD in various possible implementations described above; wherein, the BMC board is a board in the form of an unbuffered dual in-line memory module (UDIMM) board, which is pluggably connected to the server motherboard through the UDIMM interface on the server motherboard.

[0131] In this embodiment, the server is a complete computing device comprising a server motherboard, CPLD, BMC board, and a decoupled control system. This server can be a rack server, blade server, edge server, or other form of data center computing node. It integrates the aforementioned decoupled control system, thus possessing all functional characteristics such as decoupled power supply between the BMC and CPLD, independent timing control, and closed-loop anomaly recovery.

[0132] The server motherboard is the main circuit board that houses the CPU, memory, chipset, CPLD, and various interface connectors. The CPLD, located on the server motherboard, is responsible for the overall hardware logic control, such as fan speed control, power timing, and reset management. The BMC board is a pluggable functional board independent of the server motherboard, housing the BMC chip and its peripheral circuitry, and is responsible for out-of-band remote management of the server.

[0133] The BMC board is a UDIMM form factor board, which limits the physical packaging standard it adopts. UDIMM is a physical interface standard for dual in-line memory modules widely used in the server field. Designing the BMC board as a UDIMM form factor means that its physical size, pin arrangement, and interface definition are compatible with the standard UDIMM slot specification. This form factor choice has several advantages: First, the UDIMM interface is a universal standard interface on server motherboards, eliminating the need to design a dedicated connector for the BMC and reducing motherboard design complexity; second, UDIMM slots are typically mounted vertically on the motherboard, and the BMC board is inserted vertically, which significantly saves motherboard space compared to the traditional planar layout; third, the modular nature of the UDIMM form factor makes the BMC board a field-replaceable unit that can be independently upgraded and replaced.

[0134] The BMC board's pluggable connection to the server motherboard via the UDIMM interface defines the electrical and mechanical connection between the two. Pluggable connection means the BMC board can be connected or disconnected from the motherboard by insertion or removal, eliminating the need for soldering and facilitating installation, removal, and replacement. The UDIMM interface contacts provide all electrical connection channels between the BMC board and the motherboard, including standby voltage power supply, reset signal, communication bus, and various GPIO signals. After the BMC board is inserted into the UDIMM slot, all functions, including power-on and communication with the CPLD and CPU, are performed through this interface.

[0135] In some embodiments, the UDIMM-type BMC board can be paired with different CPLD logic versions to adapt to servers with different CPU platforms or different hard drive configurations. When the BMC function needs to be upgraded, only the BMC board needs to be replaced, without modifying the server motherboard and CPLD; similarly, when it is necessary to adapt to different server models, the same BMC board can be reused in multiple motherboard designs, requiring only adjustments to the CPLD logic, thus achieving cross-model reuse of the BMC hardware platform.

[0136] In addition, since the BMC board adopts a pluggable UDIMM form, the administrator can directly plug and replace the BMC board during fault maintenance without disassembling the entire machine or motherboard, which significantly reduces the difficulty and time cost of on-site maintenance.

[0137] This embodiment defines the BMC board as a UDIMM board and enables pluggable connection via a UDIMM interface on the server motherboard, achieving physical separation and modular design of the BMC hardware from the motherboard at the server system level. This simplifies motherboard design and saves layout space by utilizing a universal standard interface. Furthermore, it makes the BMC board a field-replaceable unit that can be independently replaced and upgraded. The same BMC board can be paired with different CPLD logic to adapt to multiple server models, achieving cross-platform reuse of the BMC and improving the design flexibility, maintainability, and scalability of the server hardware architecture.

[0138] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. A decoupled control system of BMC and CPLD, characterized in that, This includes the power supply control module located on the server motherboard; The power supply control module includes a logic control unit and a power switch unit; The first input terminal of the logic control unit is used to obtain the node presence identification signal on the server motherboard, and the second input terminal is used to obtain the power supply enable signal of the substrate management controller (BMC) output by the complex programmable logic device (CPLD). The output terminal of the logic control unit is connected to the enable terminal of the power switch unit, and is used to control the power switch unit to turn on when the node presence identification signal is valid and / or the BMC power supply enable signal is valid. The input terminal of the power switch unit is used to connect to the standby power supply, and the output terminal of the power switch unit is used to output the standby voltage to the BMC board.

2. The system of claim 1, wherein, The power supply control module is configured as follows: When the CPLD malfunction causes the BMC power supply enable signal to become invalid, the logic control unit continues to output a drive signal to the power switch unit based on the continuously valid node presence identification signal, so that the BMC board continues to receive power when the CPLD malfunctions.

3. The system of claim 1, wherein, The system also includes an exception handling module integrated on the BMC board, which is used to obtain the working status information of the CPLD and, when it is determined that the CPLD has abnormally hung up based on the working status information, to perform a reset process on the CPLD. The reset process indicates that a hard reset or soft reset operation of the CPLD is triggered.

4. The system of claim 3, wherein, The exception handling module is also used for: After the reset process, the firmware status of the CPLD is obtained; When the CPLD firmware is determined to be corrupted based on the firmware status, the CPLD firmware is re-burned based on the CPLD firmware image stored on the BMC board.

5. The system of claim 3, wherein, The exception handling module is also used for: After the reset process, it is determined whether the CPLD has returned to normal working state; When the CPLD fails to return to normal operation, a power-off and power-on control command is generated and sent to trigger a power-off and power-on operation or a server node hard reset operation.

6. The system of claim 1, wherein, The system also includes a status monitoring module, which is integrated on the BMC board, for periodically acquiring the working status information of the CPLD; and when the working status information times out, has no response, or returns an abnormal value, it determines that the CPLD has abnormally hung up.

7. The system of any one of claims 1-6, wherein, The system also includes a power-on control module, which is set on the BMC board. The power-on control module includes a multi-stage DC-DC power supply chip, wherein the enable pins of the multi-stage DC power supply chip are cascaded in sequence. The power-on control module is configured to: after the first-stage DC power chip outputs a standby voltage and stabilizes, enable the second-stage DC power chip to output a first operating voltage after a first preset delay; after the first operating voltage stabilizes, enable the next-stage DC power chip to output a second operating voltage after a second preset delay, until the last-stage DC power chip outputs a target operating voltage; and after the target operating voltage stabilizes, generate and output a BMC reset signal after a third preset delay.

8. The system of any one of claims 1-6, wherein, The system also includes a power-down control module, which includes a voltage detection chip and a target power timing controller; The voltage detection chip is used to detect the main power supply voltage input to the whole machine, and outputs an enable signal when the main power supply voltage drops to a preset threshold. The target power timing controller is used to receive the enable signal and, according to the enable signal, output a first reset control signal on the first channel, generate a BMC reset signal after delay processing to reset the BMC chip; and output a first power enable signal on the second channel to turn off the first standby power rail, and after the first standby power rail has stabilized after power loss, turn off the other low-voltage standby power rails in sequence.

9. The system of claim 8, wherein, The voltage value of the first standby power rail is greater than that of the other low-voltage standby power rails.

10. A server, characterized by It includes a server motherboard, a CPLD, a BMC board, and a decoupling control system for the BMC and CPLD as described in any one of claims 1-9; wherein the BMC board is a board in the form of an unbuffered dual in-line memory module (UDIMM) board, and is pluggably connected to the server motherboard through a UDIMM interface on the server motherboard.