Power supply method and device for DPU of server and storage medium

By monitoring the DPU's in-situ status using CPLD and BMC, and dynamically adjusting power supply and heat dissipation strategies, the high power consumption management complexity of traditional server power supply systems is solved, achieving an efficient, flexible, and reliable power supply method, and optimizing the startup process and heat dissipation efficiency.

CN120994036APending Publication Date: 2025-11-21POWERLEADER COMPUTER SYST CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511508954.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Traditional server power supply systems struggle to meet the diverse needs of high-power hardware, resulting in complex power management, high hardware costs, large footprint, high heat dissipation costs, low efficiency, and an inability to intelligently identify the presence of the DPU.

Method used

By monitoring the presence status of the DPU through CPLD and BMC, the power distribution and heat dissipation strategies are dynamically adjusted, and power is supplied in stages to avoid resource waste and system complexity. The presence of the DPU is determined by the NCSI_PRSNT_N signal, enabling precise power supply and heat dissipation control.

Benefits of technology

It improves server startup reliability and efficiency, reduces hardware costs and system complexity, enhances system flexibility and fault tolerance, and optimizes startup process and heat dissipation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994036A_ABST
    Figure CN120994036A_ABST
Patent Text Reader

Abstract

The invention discloses a power supply method and device for a DPU of a server and a storage medium, and the power supply method comprises the steps: judging whether the DPU is in place or not when the server is detected to be powered on; if the DPU is in place, the CPLD controls the first main power supply to supply power to a mainboard, a fan and first hardware of the server, and the BMC obtains configuration information of the current DPU and sets a corresponding DPU heat dissipation control strategy; when the CPLD or the BMC receives a startup signal, the CPLD controls the second main power supply to supply power to the second hardware; if the DPU is not in place, when the CPLD or the BMC receives a startup signal, the CPLD simultaneously controls the first main power supply to supply power to the first hardware and the second main power supply to supply power to the second hardware; when the CPLD or the BMC receives the first main power supply state normal signal and the second main power supply state normal signal, the power-on starting time sequence of the server is completed. According to the invention, the power distribution and heat dissipation strategy is dynamically adjusted according to the in-place state of the DPU, meanwhile, the hardware cost and the system complexity are reduced, and a more intelligent, efficient and flexible server power supply method is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a power supply method, apparatus, and storage medium for a server's DPU. Background Technology

[0002] With the rapid development of information technology, servers play a crucial role in data centers, cloud computing, artificial intelligence, and other fields. Artificial intelligence servers, in particular, have seen their computing power demands increase dramatically, leading to a sharp rise in server power consumption. For example, some high-performance AI servers can consume over 10kW, meaning the total input current could reach 1000A or even higher. Against this backdrop of high power consumption, server power supply systems face significant challenges.

[0003] In related technologies, traditional server power supply systems typically use a single power module, which is insufficient to meet the diverse needs of high-power hardware (such as DPU cards and GPU cards). Power management becomes extremely complex when servers need to support multiple hardware configurations. Common AI servers often require one or two DPU cards, and the power consumption of a single DPU card is typically around 150W. A common practice among server manufacturers in DPU applications is to isolate the input power in blocks. To achieve this block isolation, traditional power supply systems usually require the addition of hardware components such as soft-start circuits. This not only increases hardware costs but also occupies a significant amount of PCB space, making it less feasible. Soft-start circuits require additional heat dissipation measures under high current conditions, further increasing system complexity and cooling costs. Furthermore, the use of soft-start circuits also leads to energy loss, reducing the overall system efficiency. Summary of the Invention

[0004] This invention provides a power supply method, device, and storage medium for a server's DPU, aiming to at least solve one of the technical problems existing in the prior art.

[0005] The technical solution of the present invention is a power supply method for a server's DPU, comprising: When the server is detected to be powered on, determine whether the DPU is in place; If the DPU is in place, the CPLD controls the first main power supply to provide power to the server's motherboard, fans, and first hardware. The BMC obtains the current configuration information of the DPU and sets the corresponding DPU heat dissipation control strategy. When the CPLD or BMC receives a power-on signal, the CPLD controls the second main power supply to power the second hardware. If the DPU is not in place, when the CPLD or BMC receives the power-on signal, the CPLD simultaneously controls the first main power supply to power the first hardware and the second main power supply to power the second hardware. When the CPLD or BMC receives the first main power status normal signal and the second main power status normal signal, it completes the power-on sequence for the server.

[0006] According to some embodiments of the present invention, the NCSI_PRSNT_N signal is monitored using a CPLD and a BMC, wherein the NCSI_PRSNT_N signal is a signal of the network component control interface; If the NCSI_PRSNT_N signal is a low-level signal, it indicates that the network controller is present, confirming that the DPU is in place; If the NCSI_PRSNT_N signal is high, it indicates that the network controller is not present, confirming that the DPU is not in place.

[0007] According to some embodiments of the present invention, if the DPU is in place, the step of the CPLD controlling the first main power supply to provide power to the server's motherboard, fan, and first hardware includes: If the DPU is in place, the CPLD will control the first main power supply to power on by default through the first main power supply start signal, so as to output the first voltage; The first voltage powers the server's motherboard, fan, and first hardware.

[0008] According to some embodiments of the present invention, when the CPLD or BMC receives a power-on signal, the CPLD controls the second main power supply to power the second hardware, including: Identify whether the first hardware device contains a DPU card; If a DPU card is present, the DPU card is powered on using the first voltage, and a power-on signal is sent. When the CPLD or BMC receives the power-on signal, the CPLD outputs a second main power supply turn-on signal to control the second main power supply to power on, so as to output the second voltage; The second hardware is powered by the second voltage; When the CPLD or BMC receives the second main power status normal signal, it completes the remaining power-on sequence of the server.

[0009] According to some embodiments of the present invention, the server includes N CRPS power modules and N PCIe slots, wherein the PCIe slots support the insertion and removal of DPU cards. The first hardware includes a first CRPS power module, a second CRPS power module, a first PCIe slot, and a second PCIe slot. The second hardware includes a third to Nth CRPS power modules, a third to Nth PCIe slots, a hard disk, and other electrical loads, wherein N is a natural number.

[0010] According to some embodiments of the present invention, it further includes: When the CPLD or BMC receives a power-off signal, it determines whether the DPU is in place; If the DPU is in place, the CPLD executes the preset first power-down sequence, the BMC determines that the current configuration is DPU and adjusts the DPU heat dissipation control strategy in standby mode; If the DPU is not in place, the CPLD executes the preset second power-down sequence.

[0011] According to some embodiments of the present invention, if the DPU is in place, the CPLD executes a preset first power-down sequence including: If the DPU is in place, the CPLD outputs a second main power off signal, which controls the second main power to shut down, so that the server is in standby mode. When the server is in standby mode, the CPLD continuously outputs a first main power-on signal, which controls the first main power supply to be in working state to output the first voltage, and the first voltage supplies power to the server's motherboard, fan and first hardware. BMC controls the fan to cool the DPU according to the preset DPU cooling control strategy.

[0012] According to some embodiments of the present invention, if the DPU is not in place, the CPLD executes a preset second power-down sequence including: If the DPU is not in place, the BMC determines that the current configuration is not DPU, and the CPLD outputs the first main power off signal and the second main power off signal simultaneously. The first main power supply is turned off by controlling the first main power supply to turn off, and the second main power supply is turned off by controlling the second main power supply to turn off, so that the server is in standby mode.

[0013] The present invention also relates to a computer device, including a memory and a processor, wherein the processor performs the above-described method when executing a computer program stored in the memory.

[0014] The present invention also relates to a computer-readable storage medium storing computer program instructions thereon, which, when executed by a processor, implement the above-described method.

[0015] The power supply method, device, and storage medium for the server's DPU provided in this invention embodiment have at least one of the following advantages or beneficial effects: When the server is detected to be powered on, it is determined whether the DPU is present, ensuring that power and heat dissipation resources are rationally allocated according to the DPU's presence during startup, avoiding resource waste and improving the overall reliability of the system. If the DPU is present, the CPLD precisely controls the allocation of power resources according to preset logic, ensuring that the motherboard, fan, and first hardware receive a stable power supply during the initial startup phase. The BMC dynamically adjusts the heat dissipation strategy according to the specific configuration information of the DPU, ensuring that the DPU remains within a safe temperature range during operation. Through precise heat dissipation control strategies, overheating or underheating of the DPU is avoided, thereby improving the efficiency of the heat dissipation system and extending the lifespan of the hardware.

[0016] When the CPLD or BMC receives a power-on signal, the CPLD controls the second main power supply to power the second hardware. Dividing power distribution into multiple stages avoids excessive impact on the power system during startup. Through staged power supply, the system can gradually load hardware, ensuring each piece of hardware operates under stable power, improving power utilization, and reducing power system losses. If the DPU is not present, when the CPLD or BMC receives a power-on signal, the system can simplify the power supply logic. The CPLD can directly and simultaneously control the first main power supply to power the first hardware and the second main power supply to power the second hardware, thus powering all hardware. This simplification reduces control complexity and improves system response speed.

[0017] When the CPLD or BMC receives the first and second main power status normal signals, the system will only complete the power-on sequence after confirming that all main power statuses are normal. This mechanism ensures that all hardware is in a normal state during startup, avoiding hardware failures or system instability caused by power supply issues. Through the power status detection and confirmation mechanism, the system can promptly detect and handle potential problems during startup, ensuring the orderliness and integrity of the startup process, avoiding startup failures or system anomalies caused by power supply problems, thereby optimizing the entire startup process.

[0018] Furthermore, additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the power supply system for a server in conventional technology; Figure 2 This is a general flowchart of the power supply method for the DPU of a server provided in an embodiment of the present invention; Figure 3This is a detailed flowchart of step S100 in the power supply method for the DPU of a server provided in an embodiment of the present invention; Figure 4 This is a detailed flowchart of step S300 in the power supply method for the DPU of a server provided in an embodiment of the present invention; Figure 5 This is a first block diagram of the power supply system for a server provided in an embodiment of the present invention; Figure 6 This is a second block diagram of the power supply system for the server provided in an embodiment of the present invention; Figure 7 This is a detailed flowchart of the power supply method for the DPU of a server provided in an embodiment of the present invention; Figure 8 This is a detailed flowchart of step S610 in the power supply method for the DPU of a server provided in an embodiment of the present invention; Figure 9 This is a flowchart of the server's transition from normal working state to standby state, provided in an embodiment of the present invention. Detailed Implementation

[0020] The following will provide a clear and complete description of the concept, specific structure, and technical effects of the present invention in conjunction with the embodiments and accompanying drawings, so as to fully understand the purpose, solution, and effects of the present invention.

[0021] It should be noted that, unless otherwise specified, when a feature is referred to as "fixed" or "connected" to another feature, it can be directly fixed or connected to the other feature, or indirectly fixed or connected to the other feature. The singular forms "a," "described," and "the" used herein are also intended to include the plural forms, unless the context clearly indicates otherwise. Furthermore, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this specification is for the purpose of describing particular embodiments only and not for limiting the invention. The term "and / or" as used herein includes any combination of one or more of the associated listed items.

[0022] It should be understood that although the terms first, second, third, etc., may be used to describe various elements in this invention, these elements should not be limited to these terms. These terms are only used to distinguish elements of the same type from each other. For example, a first element may also be referred to as a second element without departing from the scope of the invention, and similarly, a second element may also be referred to as a first element. Any and all instances or exemplary language (“e.g.,” “such as,” etc.) provided herein are intended only to better illustrate embodiments of the invention and, unless otherwise required, do not impose a limitation on the scope of the invention.

[0023] With the rapid development of big data and big data models, AI servers are increasingly occupying a major share of the server market. However, in AI servers, as the power consumption of a single GPU card or module increases, the total power consumption of the entire server is also increasing exponentially. This means that the total input current of the server is also increasing exponentially. For example, if the total power consumption of the entire server is above 10KW, it means that the total output current of a general-purpose 12V CRPS power supply reaches a staggering 1000A, or even more. Common AI servers typically require one or two DPU cards, and the power consumption of a single DPU card is usually around 150W. The common practice among server manufacturers in DPU applications is to isolate the input power supply in blocks. However, in high-power systems with a total current of 1000A, block isolation of the power supply becomes somewhat impractical.

[0024] In related technologies, traditional server power supply systems typically use a single power module, which is insufficient to meet the diverse needs of high-power hardware (such as DPU cards and GPU cards). Power management becomes extremely complex when servers need to support multiple hardware configurations. Common AI servers often require one or two DPU cards, and the power consumption of a single DPU card is typically around 150W. A common practice among server manufacturers in DPU applications is to isolate the input power in blocks. To achieve this block isolation, traditional power supply systems usually require the addition of hardware components such as soft-start circuits. This not only increases hardware costs but also occupies a significant amount of PCB space, making it less feasible. Soft-start circuits require additional heat dissipation measures under high current conditions, further increasing system complexity and cooling costs. Furthermore, the use of soft-start circuits also leads to energy loss, reducing the overall system efficiency.

[0025] Reference Figure 1 As shown, Figure 1 This is a schematic diagram of a server power supply system in conventional technology. In this technology, the main power from the power module passes through a soft-start circuit. The soft-start circuit is controlled by the motherboard CPLD, which controls EN1 and EN2 to control the power output or shutdown of the downstream devices. When a DPU card needs to be inserted in standby mode, the BMC sends a command to the CPLD to control EN2 to shut down the power supply of the soft-start 2 device, while the CPLD continues to control EN1 to turn on the power supply of the soft-start 1 device. This sequentially ensures that the DPU continues to be powered when inserted into card slots 1 and 2.

[0026] The above technical solution requires isolating all main power supplies using soft-start circuits. When the server's input current is low, such as within 200 to 300A, this is feasible, even with the increased cost of the soft-start circuits. However, when the AI ​​server's total power consumption is high, such as when the input current exceeds 1000A, many soft-start circuits are needed for isolation. Due to the high current, heat sinks are required within these circuits. This significantly increases the cost of hardware design and implementation, and also occupies considerable PCB space. Assuming a soft-start circuit costs 50 yuan and is designed to handle 100A, 10 such circuits would be needed, increasing the cost by 500 yuan, not including the cost of the heat sink. Therefore, its feasibility is low and does not align with the product development goals. Furthermore, when the input current is extremely high, the heat dissipation of the soft-start circuits causes unnecessary energy loss, not only failing to save energy but also increasing the difficulty and cost of server thermal design. Furthermore, the server cannot automatically identify whether a DPU card needs to be inserted, requiring manual issuance of BMC commands to the CPLD for execution. Moreover, it does not explain how to solve the heat dissipation problem when the DPU is powered.

[0027] Based on this, embodiments of the present invention provide a power supply method, device, and storage medium for a server's DPU, which dynamically adjusts power distribution and heat dissipation strategies according to the DPU's in-situ state, while reducing hardware costs and system complexity, and achieving a more intelligent, efficient, and flexible server power supply method.

[0028] Reference Figure 2 As shown, Figure 2 This is a general flowchart of the power supply method for the DPU of a server provided in this embodiment of the invention. The power supply method for the DPU of the server includes, but is not limited to, steps S100 to S500. Specifically, S100: When the server is detected to be powered on, determine whether the DPU is in place; S200: If the DPU is in place, the CPLD controls the first main power supply to provide power to the server's motherboard, fans, and first hardware. The BMC obtains the current DPU configuration information and sets the corresponding DPU heat dissipation control strategy. S300: When the CPLD or BMC receives a power-on signal, the CPLD controls the second main power supply to power the second hardware. S400: If the DPU is not in place, when the CPLD or BMC receives the power-on signal, the CPLD simultaneously controls the first main power supply to power the first hardware and the second main power supply to power the second hardware. S500: When the CPLD or BMC receives the first main power status normal signal and the second main power status normal signal, it completes the power-on sequence for the server.

[0029] In some embodiments of the present invention, the power supply method for the server's DPU includes: when the server is detected to be powered on, determining whether the DPU (Data Processing Unit) is present, and ensuring that power and heat dissipation resources are reasonably allocated according to the DPU's presence during startup to avoid resource waste and improve the overall reliability of the system. If the DPU is present, the CPLD (Complex Programmable Logic Device) precisely controls the allocation of power resources according to preset logic to ensure that the motherboard, fan, and first hardware receive a stable power supply during startup. The BMC (Baseboard Management Controller) dynamically adjusts the heat dissipation strategy according to the specific configuration information of the DPU (such as model, power consumption, etc.) to ensure that the DPU remains within a safe temperature range during operation. Through precise heat dissipation control strategies, excessive or insufficient heat dissipation of the DPU is avoided, thereby improving the efficiency of the heat dissipation system and extending the hardware lifespan.

[0030] Understandably, the motherboard and fans are fundamental components for server operation. Prioritizing power supply to them ensures the normal startup of basic server functions, such as motherboard initialization and fan cooling. Through CPLD logic control, power distribution chaos or errors can be avoided, thereby improving system stability and reducing system failures caused by power problems.

[0031] Subsequently, when the CPLD or BMC receives the power-on signal, the CPLD controls the second main power supply to power the second hardware. Dividing power distribution into multiple stages avoids excessive stress on the power system during startup. Through staged power supply, the system can gradually load hardware, ensuring each piece of hardware operates under stable power, improving power utilization, and reducing power system losses. This staged power supply mechanism can be flexibly adjusted according to different hardware configurations, enhancing the system's adaptability to startup under different power environments. Gradually supplying power based on the actual needs of the hardware (e.g., the presence of the DPU) avoids power waste and improves the system's adaptability and flexibility.

[0032] If the DPU is not present, when the CPLD or BMC receives a power-on signal, the system can simplify the power supply logic. The CPLD can directly and simultaneously control the first main power supply to power the first hardware and the second main power supply to power the second hardware, thus powering all hardware. This simplification reduces control complexity and improves system response speed. In this scenario, the system can centrally allocate resources to other hardware, avoiding resource waste caused by the absence of the DPU and optimizing overall resource utilization efficiency.

[0033] When the CPLD or BMC receives the first and second main power status normal signals, the system will only complete the power-on sequence after confirming that all main power statuses are normal. This mechanism ensures that all hardware is in a normal state during startup, avoiding hardware failures or system instability caused by power supply issues. Through the power status detection and confirmation mechanism, the system can promptly detect and handle potential problems during startup, ensuring the orderliness and integrity of the startup process, avoiding startup failures or system anomalies caused by power supply problems, thereby optimizing the entire startup process.

[0034] The power supply method for the DPU of a server provided in this invention enables refined management and control of the server's power-on process. By detecting hardware status, dynamically adjusting power distribution and heat dissipation strategies, and implementing phased power supply and power status detection, the system can achieve efficient, stable, and reliable startup under different hardware configurations and operating conditions. These techniques not only improve the server's operating efficiency but also enhance the system's fault tolerance and scalability. Furthermore, this invention eliminates the need for a soft-start circuit, reducing the material cost of soft-start circuits in hardware design, saving PCB layout space, making it suitable for high-power server design scenarios, reducing the difficulty of heat dissipation design, and making the server more energy-efficient. Simultaneously, it intelligently identifies the DPU's presence and dynamically adjusts the heat dissipation strategy based on the DPU's presence, avoiding overheating or underheating of the DPU, thereby improving the efficiency of the heat dissipation system and extending hardware lifespan.

[0035] Reference Figure 3 As shown, Figure 3 This is a detailed flowchart of step S100 in the power supply method for the DPU of a server provided in this embodiment of the invention. Step S100 includes, but is not limited to, steps S110 to S130. Specifically, S110: Use CPLD and BMC to monitor the NCSI_PRSNT_N signal, which is the signal of the network component control interface; S120: If the NCSI_PRSNT_N signal is low, it indicates that the network controller is present, confirming that the DPU is in place; S130: If the NCSI_PRSNT_N signal is high, it indicates that the network controller does not exist, confirming that the DPU is not in place.

[0036] In some embodiments of the present invention, the method for determining whether the DPU is present includes: monitoring the NCSI_PRSNT_N signal using a CPLD and a BMC. The NCSI_PRSNT_N signal is a signal of the Network Component Control Interface (NCSI) used to indicate the presence status of the network controller. A CPLD is a programmable logic device capable of rapidly responding to changes in hardware signals. It can monitor the level of the NCSI_PRSNT_N signal in real time through logic circuits. The BMC is the server's management system, responsible for monitoring hardware status and performing management tasks. It can obtain the status of the NCSI_PRSNT_N signal through communication with the CPLD and make corresponding decisions accordingly.

[0037] The CPLD can monitor the level of the NCSI_PRSNT_N signal in real time and transmit the signal status to the BMC. This real-time monitoring mechanism ensures that the system can quickly respond to changes in the DPU's presence status. For example, if the CPLD detects a low level for the NCSI_PRSNT_N signal, it indicates the presence of the network controller, thus inferring the DPU's presence; conversely, if the CPLD detects a high level for the NCSI_PRSNT_N signal, it indicates the absence of the network controller, thus inferring the DPU's absence. Using the high and low levels of the signal to determine the DPU's presence is a simple and reliable method. Low and high levels correspond to the presence and absence of the DPU, respectively. Determining the DPU's status through explicit signal levels reduces misjudgments caused by hardware failures or signal interference, avoids complex status detection logic, and improves the accuracy of the determination.

[0038] Based on the state of the NCSI_PRSNT_N signal, the CPLD can dynamically adjust the startup and operation process. For example, when the DPU is detected to be present, the CPLD starts the DPU-related hardware and cooling strategies; when the DPU is detected to be absent, the CPLD can skip the DPU-related initialization steps and directly start other hardware. By quickly and accurately determining the DPU status, the system can avoid unnecessary waiting and initialization operations, thereby improving startup and operation efficiency.

[0039] In some embodiments of the present invention, in step S200, if the DPU is in place, the CPLD controls the first main power supply to provide power to the server's motherboard, fan, and first hardware, specifically including steps S210 to S220. S210: If the DPU is in place, the CPLD will control the first main power supply to power on by default through the first main power supply start signal to output the first voltage; S220: Powers the server’s motherboard, fans and primary hardware via a primary voltage.

[0040] In some embodiments of the present invention, when the CPLD monitors the NCSI_PRSNT_N signal in real time and finds it to be a low-level signal, it indicates that the network controller is present and the DPU is in place. The CPLD outputs a first main power-on signal to control the power-on process of the first main power supply. After receiving the power-on signal, the first main power supply outputs a set first voltage to power the relevant hardware, such as the motherboard, fan and first hardware.

[0041] The motherboard is the core component of the server. Powering the motherboard enables it to power on and is responsible for handling and coordinating the operation of various hardware components; powering the fans, which are used for heat dissipation to ensure that the hardware operates within a normal temperature range; and powering the primary hardware, which is a specific hardware component related to the DPU, such as a network card, storage device, or DPU card.

[0042] Through the logic control of the CPLD, it is ensured that power is only supplied to the relevant hardware when the DPU is present, avoiding power failures or damage caused by hardware absence. After confirming the DPU's presence, the CPLD immediately controls the first main power supply to power on via the first main power-on signal, supplying power to critical hardware and thus accelerating system startup. The CPLD precisely controls the power-on process of the first main power supply via the first main power-on signal, ensuring that the motherboard, fans, and primary hardware receive a stable power supply.

[0043] Reference Figure 4 As shown, Figure 4 This is a detailed flowchart of step S300 in the power supply method for the DPU of a server provided in this embodiment of the invention. Step S300 includes, but is not limited to, steps S310 to S350. Specifically, S310: Identifies whether the first hardware component contains a DPU card; S320: If a DPU card is present, power on the DPU card using the first voltage and send a power-on signal; S330: When the CPLD or BMC receives a power-on signal, the CPLD outputs a second main power supply turn-on signal to control the second main power supply to power on, so as to output the second voltage; S340: Powers the second hardware via a second voltage; S350: When the CPLD or BMC receives the second main power status normal signal, it completes the remaining power-on sequence of the server.

[0044] In some embodiments of the present invention, when the CPLD or BMC receives a power-on signal, the method by which the CPLD controls the second main power supply to power the second hardware includes: first, identifying whether the first hardware has a DPU card; if the DPU card is detected, the CPLD will perform the corresponding power-on and boot operations; if the DPU card is not detected, the CPLD may skip operations related to the DPU card. For example, if the DPU card exists, the DPU card is powered on using a first voltage, and after the DPU card is powered on, the system will send a power-on signal to start the DPU card.

[0045] The CPLD or BMC monitors the power-on signal in real time. Upon receiving the power-on signal, the CPLD outputs a second main power-on signal. The CPLD controls the second main power supply to power on via this signal, outputting a second voltage. This second voltage powers the second hardware components, which are other hardware components related to the DPU card, such as storage devices and expansion cards. The CPLD or BMC monitors the status of the second main power supply. Upon receiving a normal status signal from the second main power supply, the system completes the remaining power-on sequence, such as starting the operating system and initializing all hardware components.

[0046] By identifying the presence of a DPU card in the primary hardware component through a detection mechanism, the system can avoid unnecessary power supply operations to non-existent hardware, thereby improving system reliability and efficiency. The system dynamically adjusts its power supply strategy based on the presence or absence of the DPU card, ensuring that power is only supplied to the DPU card when needed, thus avoiding resource waste. Furthermore, if the DPU card is absent, the system can skip operations related to the DPU card and continue booting other hardware, enhancing the system's fault tolerance.

[0047] By identifying the presence of a DPU card in the primary hardware and performing corresponding power-on and boot operations based on its presence, the system achieves precise hardware identification, phased power supply, and boot control. This technique not only improves system startup efficiency and reliability but also optimizes the utilization of power and hardware resources. Through the collaborative work of the CPLD and BMC, the system can achieve efficient, stable, and reliable operation in complex hardware environments, providing a solid foundation for the normal operation of the server.

[0048] In some embodiments of the present invention, the server includes N CRPS power modules and N PCIe slots, the PCIe slots supporting the insertion and removal of DPU cards, the first hardware including a first CRPS power module, a second CRPS power module, a first PCIe slot and a second PCIe slot; the second hardware including a third CRPS power module to the Nth CRPS power module, a third PCIe slot to the Nth NPCIE slot, a hard disk and other electrical loads, wherein N is a natural number.

[0049] Reference Figure 5 As shown, the server contains N CRPS (Common Redundant Power Supply) modules. These CRPS modules provide redundant power to ensure high system availability. The CRPS modules provide stable power to all server hardware components, support hot-swapping for easy maintenance and expansion. The server also contains N PCIe slots, supporting the insertion and removal of DPU cards. These PCIe slots are used to connect high-performance computing accelerator cards such as DPU cards, and support hot-swapping for dynamic hardware expansion and maintenance. By supporting the insertion and removal of DPU cards through N PCIe slots, the system can dynamically expand its computing power as needed, enhancing hardware scalability.

[0050] Redundant power is provided through N CRPS power modules, ensuring that the system can continue to operate normally even if one power module fails, thus improving system high availability. The CPLD can dynamically adjust the power distribution strategy based on the status of the power modules to ensure stable system operation.

[0051] In one embodiment, the power output from CRPS power module 1 (first CRPS power module) and CRPS power module 2 (second CRPS power module) is named Main Power Supply 1 (first Main Power Supply), and the power output from CRPS power modules 3 to CRPS power module N is named Main Power Supply 2 (second Main Power Supply), thus dividing the power supply. Main Power Supply 1 primarily supplies power to the motherboard (mainly the CPU and memory), fans, PCIe slot 1 (first PCIe slot), and PCIe slot 2 (second PCIe slot). PCIe slots 1 and 2 can support DPU cards. Main Power Supply 2 primarily supplies power to PCIe slots 3 to PCIe slot N, high-power GPU cards, network cards, hard drives, etc. This solution does not involve the application of a soft-start circuit.

[0052] The CPLD controls the power-on of the first and second CRPS power modules via a first main power-on signal, supplying power to the first hardware components. This includes the motherboard, fans, and the first and second PCIe slots. When the CPLD or BMC receives a power-on signal, the CPLD outputs a second main power-on signal, controlling the power-on of the third to Nth CRPS power modules to supply power to the second hardware components, including hard drives, other PCIe slots, and other electrical loads. Through staged power supply, the system can gradually load hardware, avoiding excessive stress on the power system during startup and improving startup efficiency and system stability.

[0053] Reference Figure 6As shown, the BMC and CPLD are responsible for power-on control of the power modules. The CPLD controls the output of the main power supply 1 through the PSU1_2_PSON_N signal (first main power-on signal) for power modules 1 and 2. The CPLD controls the output of the main power supply 2 through the PSU3_N_PSON_N signal (second main power-on signal) for power modules 3 to N.

[0054] Typically, scenarios using DPU cards utilize their NCSI (Network Controller Sideband Interface) function. NCSI is the communication interface protocol between the management controller (BMC) and the network card, allowing access to the BMC network via the network card. Therefore, the CPLD and BMC monitor the NCSI_PRSNT_N signal to intelligently identify whether a DPU card is currently configured on the machine, and then determine whether power supply is required. When using a DPU, the BMC sends commands to the CPLD to configure the relevant settings.

[0055] By clearly defining the server's hardware components and power supply strategy, particularly the specific division between the first and second hardware components, the system achieves high availability, phased power supply, hardware scalability, and enhanced reliability. This design not only optimizes power management but also improves system startup efficiency and operational stability. Through the collaborative work of the CPLD and BMC, the system can achieve efficient, stable, and reliable operation in complex hardware environments, providing a solid foundation for the normal operation of the server.

[0056] In one embodiment, the server is plugged into AC power and is in standby mode. The CPLD and BMC monitor the NCSI_PRSNT_N signal to determine whether the DPU is present. If the DPU is present: the internal logic of the CPLD is set up, the power-on control of the first power module PSU1 and the second power module PSU2 is separated from other power modules PSU, the BMC knows that the current configuration is DPU, and sets the predetermined fan control strategy. If the DPU is not present: the internal logic of the CPLD is set up, the power-on control of the first power module PSU1 and the second power module PSU2 is synchronized with other power modules, the BMC knows that the current configuration is non-DPU, and no fan control is required in standby mode.

[0057] If the DPU is in place, the CPLD will control the first power module PSU1 and the second power module PSU2 to power on via the first main power enable signal PSU1_2_PSON_N by default, and output the first main power 12V. It should be noted that although the motherboard will also receive the first main power at this time, since the system has not been powered on, the CPLD does not control the power-on sequence of the CPU and memory, so the CPU and memory will not start working and will not cause energy loss. The DPU card will power on and work after receiving the 12V input.

[0058] If the DPU is present, the BMC controls the fans to cool the DPU according to a predetermined strategy, preventing overheating during DPU operation. Subsequently, the CPLD and BMC receive a power-on signal. If the DPU is present, the CPLD outputs a second main power-on signal, PSU3_N_PSON_N, to power on power modules PSU3~N, initiating power supply to the GPU, hard drive, network card, and other major loads. If the DPU is absent, the CPLD simultaneously outputs a first main power-on signal, PSU1_2_PSON_N, and a second main power-on signal, PSU3_N_PSON_N, to power on all power modules, initiating power supply to the CPU, memory, GPU, hard drive, network card, and other major loads. When the CPLD or BMC receives the first main power status normal signal, PSU1_2_PWROK, and the second main power status normal signal, PSU3_N_PWROK, the system power-on sequence is complete. After power-on, the BMC monitors the system in real-time according to a predetermined cooling strategy, controlling the fans to cool the system, ensuring normal server operation.

[0059] Reference Figure 7 As shown, Figure 7 This is a detailed flowchart of the power supply method for the server's DPU provided in this embodiment of the invention. The power supply method for the server's DPU also includes, but is not limited to, steps S600 to S620. Specifically, S600: When the CPLD or BMC receives a power-off signal, it determines whether the DPU is in place; S610: If the DPU is in place, the CPLD executes the preset first power-down sequence, the BMC determines that the current configuration is DPU and adjusts the DPU heat dissipation control strategy in standby mode; S620: If the DPU is not in place, the CPLD executes the preset second power-down sequence.

[0060] In some embodiments of the present invention, when the CPLD or BMC receives a shutdown signal, it determines whether the DPU is in place. The shutdown signal can be triggered by user operation (e.g., through a management interface or physical button) or automatically by the system (e.g., for fault detection or maintenance needs). The CPLD or BMC is responsible for receiving and processing the shutdown signal to ensure that the system can correctly respond to the shutdown request.

[0061] The presence of the DPU is determined using hardware signals (such as the NCSI_PRSNT_N signal) or software detection mechanisms. If the DPU is present, the CPLD gradually shuts down the power module according to a preset first power-down sequence, ensuring the safe power-down of the DPU and other hardware components. The BMC determines that the current configuration is DPU and adjusts the DPU thermal management strategy in standby mode. If the DPU is absent, the CPLD gradually shuts down the power module according to a preset second power-down sequence, ensuring the safe power-down of other hardware components.

[0062] By using a preset power-down sequence, the system can gradually shut down the power module, avoiding damage to the hardware caused by sudden power outages and extending the hardware's lifespan. When the DPU is present, the BMC adjusts the DPU's thermal management strategy in standby mode to ensure the DPU remains within a safe temperature range, preventing hardware failures due to insufficient heat dissipation. When the DPU is absent, the system can skip DPU-related thermal operations, optimizing the allocation of thermal resources and improving overall system efficiency. The system can flexibly adjust operations based on the DPU's presence status. By dynamically adjusting the thermal management strategy, the system can optimize the allocation of thermal resources based on the DPU's presence status, avoiding resource waste.

[0063] Reference Figure 8 As shown, Figure 8 This is a detailed flowchart of step S610 in the power supply method for the DPU of a server provided in this embodiment of the invention. Step S610 includes, but is not limited to, steps S611 to S613. Specifically, S611: If the DPU is in place, the CPLD outputs a second main power off signal, which controls the second main power to shut down, so that the server is in standby mode. S612: When the server is in standby mode, the CPLD continuously outputs the first main power-on signal, and controls the first main power supply to be in working state to output the first voltage, and supplies power to the server's motherboard, fan and first hardware through the first voltage. S613: BMC controls the fan to cool the DPU according to the preset DPU heat dissipation control strategy.

[0064] In some embodiments of the present invention, if the DPU is present, the CPLD executes a preset first power-down sequence. The BMC determines that the current configuration is DPU and adjusts the DPU heat dissipation control strategy in standby mode. This includes: if the DPU is present, the CPLD outputs a second main power-off signal to control the second main power supply to shut down, which puts the server's second hardware (such as hard drives, other PCIe slots, etc.) into standby mode. The CPLD continuously outputs a first main power-on signal to control the first main power supply to be in working mode, outputting a first voltage to power the server's motherboard, fans, and first hardware (such as the DPU card). According to the preset DPU heat dissipation control strategy, the BMC controls the fans to dissipate heat for the DPU, dynamically adjusting the fan speed and operating mode to ensure that the DPU can still be kept within a safe temperature range in standby mode.

[0065] With the DPU in place, the server can gradually put non-critical hardware (such as hard drives and other expansion cards) into standby mode by shutting down the second main power supply while maintaining the operation of the first main power supply. This gradual power-down process prevents damage to these hardware components from sudden power outages. The CPLD and BMC work together to ensure that critical hardware (such as the motherboard, fans, and DPU card) continues to operate normally in standby mode, and a dynamic thermal management strategy ensures the safety of the DPU. This technique not only protects the hardware from damage caused by sudden power outages but also optimizes the utilization of power and cooling resources, enhancing system reliability and fault tolerance. Through the collaborative work of the CPLD and BMC, the system can achieve efficient, stable, and reliable operation in complex hardware environments, providing a solid foundation for the normal operation of the server.

[0066] In some embodiments of the present invention, in step S620, if the DPU is not in place, the CPLD executes a preset second power-down sequence, including but not limited to steps S621 to S622. Specifically, S621: If the DPU is not in place, the BMC determines that the current configuration is not DPU, and the CPLD outputs the first main power off signal and the second main power off signal at the same time. S622: Controls the first main power supply to shut down via the first main power supply shutdown signal, and controls the second main power supply to shut down via the second main power supply shutdown signal, so that the server is in standby mode.

[0067] In some embodiments of the present invention, if the DPU is not present, the method for the CPLD to execute a preset second power-down sequence includes: using a hardware signal (such as the NCSI_PRSNT_N signal) or other detection mechanism, the BMC determines whether the current configuration is a DPU configuration. If the DPU is not present, the BMC determines that the current configuration is not a DPU configuration and executes a non-DPU configuration power-down strategy. The CPLD outputs a first main power-off signal to control the first main power supply to shut down, and simultaneously outputs a second main power-off signal to control the second main power supply to shut down. By shutting down the first and second main power supplies, the server enters a standby state.

[0068] By simultaneously outputting the first and second main power-off signals from the CPLD, the system can gradually shut down all power modules, putting the server into standby mode and preventing hardware damage from sudden power outages. In the absence of the DPU, the CPLD simultaneously shuts down both the first and second main power supplies, simplifying the operation and reducing the complexity of the control logic. Furthermore, in the absence of the DPU, the system can completely shut down all power modules, avoiding unnecessary power consumption and improving power utilization efficiency.

[0069] In one embodiment, refer to Figure 9As shown, the process of a server transitioning from normal operation to standby includes: When the CPLD and BMC receive a power-off signal, they determine if the DPU is present. If the DPU is present, the current configuration is DPU-enabled. The BMC adjusts the DPU cooling fan strategy during standby, the CPLD executes the power-down sequence, and finally, the CPLD outputs the second main power-off signal PSU3_N_PSON_N to power down power modules PSU3~N, shutting down main power supply 2 (the second main power supply). It also outputs the first main power-on signal PSU1_2_PSON_N to power modules PSU1 and PSU2 to keep main power supply 1 powered on. If the DPU is not present, the current configuration is non-DPU-enabled. In standby mode, no fan control is needed. The CPLD executes the power-down sequence, and finally, the CPLD simultaneously outputs the first main power-off signal PSU1_2_PSON_N and the second main power-off signal PSU3_N_PSON_N to power down all power modules, shutting down main power supply 1 (the first main power supply) and main power supply 2 (the second main power supply). If the DPU is in place and the DPU card is powered by 12V normally, the BMC controls the fan to cool the DPU according to the established strategy to prevent the DPU card from overheating during operation; the server enters standby mode.

[0070] This paper details the implementation of the proposed solution for the entire power-on and power-off process of the server. Below, we provide an explanation of how to intelligently identify whether the DPU is present, along with some CPLD logic code: module top CSI_PRSNT_N, / / / NCSI cable presence monitoring input wire pwrbtn_n, / / / Power button input input wire PSU1_PWROK, / / / Power good signal input for PSU1 input wire PSU2_PWROK, / / / Power good signal input for PSU2 input wire PSU3_PWROK, / / / Power good signal input for PSU3 input wire PSUN_PWROK, / / / PSUN's Power Good signal input Output wire PSU1_2_PSON_N, / / / Power-on / off control signal output to PSU1 and PSU2 Output wire PSU3_N_PSON_N / / / Power-on / off control signal output to PSU3 and PSUN ); wire pson_n; / / / Control signal generated by the internal logic module to control the power-on and power-off of the PSU wire PSU1_2_PWROK; / / / Power Good signal obtained by logically ORing PSU1 and PSU2 wire PSU3_N_PWROK; / / / Power good signal obtained by logical ORing PSU3 to PSUN assign PSU1_2_PSON_N = (NCSI_PRSNT_N==1'b1) ? pson_n : 1'b0; / / / When the DPU is in position, the default output is low, causing PSU1 and 2 to output power. When the DPU is not in position, it follows the control power-on / off signals generated by the main system logic. assign PSU3_N_PSON_N = pson_n; / / / Other PSUs will default to following the control signals generated by the main system logic. It should be understood that the method steps in the embodiments of the present invention can be implemented or carried out by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable storage medium. The method can use standard programming techniques. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if necessary, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. Furthermore, for this purpose, the program can run on a programmed application-specific integrated circuit (ASIC).

[0071] Furthermore, the procedures described herein may be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by the context. The procedures described herein (or variations and / or combinations thereof) may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented by hardware or a combination thereof as code (e.g., executable instructions, one or more computer programs, or one or more applications) that commonly executes on one or more processors. The computer program comprises a plurality of instructions executable by one or more processors.

[0072] Furthermore, the method can be implemented in any suitable type of computing platform, including but not limited to personal computers, minicomputers, mainframes, workstations, networked or distributed computing environments, standalone or integrated computer platforms, or in communication with charged particle tools or other imaging devices, etc. Aspects of the invention can be implemented as machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, optical read and / or write storage medium, RAM, ROM, etc., such that it is readable by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the processes described herein. Furthermore, the machine-readable code, or portions thereof, can be transmitted via wired or wireless networks. The invention described herein includes these and other different types of non-transitory computer-readable storage media when such media comprises instructions or programs that implement the steps described above in conjunction with a microprocessor or other data processor. When programmed according to the methods and techniques described in the invention, the invention may also include the computer itself.

[0073] A computer program can be applied to input data to perform the functions described herein, thereby transforming the input data to generate output data stored in non-volatile memory. The output information can also be applied to one or more output devices, such as a display. In a preferred embodiment of the invention, the transformed data represents physical and tangible objects, including specific visual depictions of physical and tangible objects generated on the display.

[0074] The above description is merely a preferred embodiment of the present invention. The present invention is not limited to the above-described embodiments. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention, as long as they achieve the technical effects of the present invention by the same means, should be included within the scope of protection of the present invention. Within the scope of protection of the present invention, the technical solutions and / or implementation methods can have various modifications and variations.

Claims

1. A power supply method for a server's DPU, characterized in that, include: When the server is detected to be powered on, determine whether the DPU is in place; If the DPU is in place, the CPLD controls the first main power supply to provide power to the server's motherboard, fans, and first hardware. The BMC obtains the current configuration information of the DPU and sets the corresponding DPU heat dissipation control strategy. When the CPLD or BMC receives a power-on signal, the CPLD controls the second main power supply to power the second hardware. If the DPU is not in place, when the CPLD or BMC receives the power-on signal, the CPLD simultaneously controls the first main power supply to power the first hardware and the second main power supply to power the second hardware. When the CPLD or BMC receives the first main power status normal signal and the second main power status normal signal, it completes the power-on sequence for the server.

2. The power supply method for the server's DPU according to claim 1, characterized in that, The determination of whether the DPU is in place includes: The NCSI_PRSNT_N signal, which is a signal of the network component control interface, is monitored using a CPLD and a BMC. If the NCSI_PRSNT_N signal is a low-level signal, it indicates that the network controller is present, confirming that the DPU is in place; If the NCSI_PRSNT_N signal is high, it indicates that the network controller is not present, confirming that the DPU is not in place.

3. The power supply method for the server's DPU according to claim 1, characterized in that, If the DPU is in place, the CPLD controls the first main power supply to provide power to the server's motherboard, fans, and first hardware, including: If the DPU is in place, the CPLD will control the first main power supply to power on by default through the first main power supply start signal, so as to output the first voltage; The first voltage powers the server's motherboard, fan, and first hardware.

4. The power supply method for the server's DPU according to claim 3, characterized in that, When the CPLD or BMC receives a power-on signal, the CPLD controls the second main power supply to power the second hardware, including: Identify whether the first hardware device contains a DPU card; If a DPU card is present, the DPU card is powered on using the first voltage, and a power-on signal is sent. When the CPLD or BMC receives the power-on signal, the CPLD outputs a second main power supply turn-on signal to control the second main power supply to power on, so as to output the second voltage; The second hardware is powered by the second voltage; When the CPLD or BMC receives the second main power status normal signal, it completes the remaining power-on sequence of the server.

5. The power supply method for the server's DPU according to claim 1, characterized in that, The server includes N CRPS power modules and N PCIe slots. The PCIe slots support the insertion and removal of DPU cards. The first hardware includes a first CRPS power module, a second CRPS power module, a first PCIe slot, and a second PCIe slot. The second hardware includes a third to Nth CRPS power module, a third to Nth PCIe slot, a hard drive, and other electrical loads, where N is a natural number.

6. The power supply method for the server's DPU according to claim 3, characterized in that, Also includes: When the CPLD or BMC receives a power-off signal, it determines whether the DPU is in place; If the DPU is in place, the CPLD executes the preset first power-down sequence, the BMC determines that the current configuration is DPU and adjusts the DPU heat dissipation control strategy in standby mode; If the DPU is not in place, the CPLD executes the preset second power-down sequence.

7. The power supply method for the server's DPU according to claim 6, characterized in that, If the DPU is in place, the CPLD executes the preset first power-down sequence, including: If the DPU is in place, the CPLD outputs a second main power off signal, which controls the second main power to shut down, so that the server is in standby mode. When the server is in standby mode, the CPLD continuously outputs a first main power-on signal, which controls the first main power supply to be in working state to output the first voltage, and the first voltage supplies power to the server's motherboard, fan and first hardware. BMC controls the fan to cool the DPU according to the preset DPU cooling control strategy.

8. The power supply method for the server's DPU according to claim 6, characterized in that, If the DPU is not in place, the CPLD executes a preset second power-down sequence including: If the DPU is not in place, the BMC determines that the current configuration is not DPU, and the CPLD outputs the first main power off signal and the second main power off signal simultaneously. The first main power supply is turned off by controlling the first main power supply to turn off, and the second main power supply is turned off by controlling the second main power supply to turn off, so that the server is in standby mode.

9. A computer device comprising a memory and a processor, characterized in that, When the processor executes a computer program stored in the memory, it performs the method as described in any one of claims 1 to 8.

10. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they perform the method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Intelligent network card power-on control system and server

    CN114564095A

  • Method for setting server DPU network card to enter FS-5 state

    CN117075708A

  • Control method and device of server heat dissipation equipment, storage medium and electronic equipment

    CN117806438A