Server management system and server

By introducing interface modules between computing nodes and management nodes and using programmable logic devices and baseboard management controllers to achieve efficient interconnection and centralized management, the hardware maintenance complexity and cost issues caused by the computing units and management units sharing the same motherboard are solved, and the system scalability and management efficiency are improved.

CN120353303BActive Publication Date: 2025-09-12INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510837402.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-12
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

In the prior art, the computing unit and the management unit share a common motherboard, which means that when the motherboard is damaged, both need to be replaced, increasing the complexity and cost of hardware maintenance.

Method used

An interface module is introduced to connect multiple computing nodes through the programmable logic devices in the interface module, and these nodes are managed by the baseboard management controller to achieve efficient interconnection and centralized management.

Benefits of technology

Reduce management units, lower costs, improve system scalability and management efficiency, and enhance system reliability and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353303B_ABST
    Figure CN120353303B_ABST
Patent Text Reader

Abstract

The present application discloses a server management system and server, relating to the field of computer technology. The system includes a system that implements efficient interconnection and centralized management between multiple computing nodes and a management module by introducing an interface module between the computing module and the management module. Specifically, each computing node includes a computing unit and a programmable logic device. Multiple computing nodes are connected via the programmable logic device in the interface module. A first management module obtains operating data from each computing node through the interface module, enabling a single management module to manage multiple nodes. This reduces the number of management units and costs, thereby improving the scalability and management efficiency of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a server management system and a server. Background Art

[0002] The explosive growth of the internet has fueled the demand for high-concurrency processing capabilities. Clustering high-computing servers has become a key approach to improving performance. Decoupling and centralizing the management of server modules within clustered servers can reduce maintenance costs.

[0003] The servers in the related technology are all managed independently on a single node, that is, each node has its own separate management chip. Whether in a conventional server chassis or in an entire cabinet design, the server architecture usually manages each node independently. If you need to check the status of each node, you need to check it separately, or make another management chip for unified maintenance.

[0004] However, the management chip is mounted on the server motherboard. During actual execution, the computing unit or management unit shares the motherboard. Once the motherboard is broken, both need to be replaced at the same time, which not only affects the computing function, but also affects the normal operation of the entire management system, thereby increasing maintenance difficulty and replacement costs. Summary of the Invention

[0005] The present application provides a server management system and a server, which at least solves the problem in the related art that computing units or management units share a motherboard, and when the motherboard is damaged, both motherboards need to be replaced, which increases the complexity and cost of hardware maintenance.

[0006] The present application provides a server management system, comprising: a computing module, wherein the computing module comprises multiple computing nodes of the server, each computing node comprises a computing unit and a programmable logic device, and the computing unit in each computing node is connected to the programmable logic device; an interface module, wherein the interface module comprises multiple programmable logic devices, and the multiple programmable logic devices of the interface module are connected to the programmable logic devices of the multiple computing nodes; a first management module, wherein the first management module comprises at least one baseboard management controller, and the baseboard management controller is connected to the multiple programmable logic devices of the interface module, and the operating data of the multiple computing nodes is obtained through the interface module, and the multiple computing nodes are managed according to the operating data of the multiple computing nodes.

[0007] The present application also provides a server, including: a management system of the server of the above embodiment.

[0008] This application introduces an interface module between the computing module and the first management module to achieve efficient interconnection and centralized management between multiple computing nodes and the first management module. Specifically, each computing node includes a computing unit and a programmable logic device. Multiple computing nodes are connected via the programmable logic device in the interface module. The first management module obtains the operating data of each computing node through the interface module, achieving a design in which a single management module manages multiple nodes. This can reduce the number of management units and lower costs, thereby improving the scalability and management efficiency of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0010] Figure 1 A signal topology diagram of a server management system provided in an embodiment of the present application;

[0011] Figure 2 A multi-node management CPU debugging signal topology diagram provided by one embodiment of the present application;

[0012] Figure 3 A multi-node management PCIe signal topology diagram provided for one embodiment of the present application;

[0013] Figure 4 A multi-node management low-speed signal topology diagram provided in one embodiment of the present application;

[0014] Figure 5 A multi-node management SPI signal topology diagram provided for one embodiment of the present application;

[0015] Figure 6 A multi-node management USB signal topology diagram provided for one embodiment of the present application;

[0016] Figure 7 A topological diagram of computing nodes provided in one embodiment of the present application;

[0017] Figure 8 An example diagram of the association between the management module and the computing module is provided for one embodiment of the present application;

[0018] Figure 9 A design diagram of a dual management module provided in one embodiment of the present application;

[0019] Figure 10 This is an example diagram of the interface form of the management module provided in one embodiment of the present application;

[0020] Figure 11 A control flow chart of a dual management module provided in one embodiment of the present application;

[0021] Figure 12 A circuit flow diagram of an interface module and a management module provided in one embodiment of the present application;

[0022] Figure 13 A flowchart for powering on a computing module is provided for one embodiment of the present application. DETAILED DESCRIPTION

[0023] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0024] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0025] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0026] The embodiment of the present application provides a server management system, such as Figure 1 As shown, it includes: a computing module 101 , an interface module 102 and a first management module 103 .

[0027] Among them, the computing module 101 includes multiple computing nodes of the server, each computing node includes a CPU (Computing Unit) and a CPLD (Complex Programmable Logic Device), and the computing unit CPU in each computing node is connected to the programmable logic device CPLD; the interface module 102 includes multiple programmable logic devices CPLD, and the multiple programmable logic devices CPLD of the interface module are connected to the programmable logic devices CPLD of multiple computing nodes; the first management module 103 includes at least one BMC (Baseboard Management Controller), and the baseboard management controller BMC is connected to the multiple programmable logic devices CPLD of the interface module 102, obtains the operating data of multiple computing nodes through the interface module 102, and manages the multiple computing nodes according to the operating data of the multiple computing nodes.

[0028] Specifically, if Figure 1 As shown, the computing module 101 of the embodiment of the present application includes multiple computing nodes, each of which is composed of a computing unit (such as a CPU) and necessary supporting chips (such as a power controller, a clock chip, a memory, etc.), and each computing node is also equipped with a programmable logic device (CPLD). This design allows each computing node to operate independently, and can also communicate with the interface module through its internal CPLD. The interface module 102, as a key component connecting the computing module 101 and the first management module 103, includes multiple CPLDs for establishing connections with each computing node, and also provides a variety of interfaces (such as I2C, etc.) to realize data transmission and control signal transmission. The first management module 103 includes at least one baseboard management controller (BMC), which is responsible for monitoring the health of the entire system, including key parameters such as temperature and voltage, and can perform corresponding management operations based on the collected data. In particular, this management module supports hot-swappable function, which means that the management module can be replaced or maintained without shutting down the system, greatly improving the availability and flexibility of the system. In the actual implementation process, if Figure 1 As shown in the figure, the CPU records the internal temperature status, the voltage status on the motherboard, and other status in the motherboard CPLD register. After the CPLD on the IO board communicates with the motherboard CPLD through LTPI to obtain information, the BMC reads the relevant information from the CPLD register on the IO board through I2C.

[0029] In one embodiment of the present application, the first management module 103 includes at least one serial port, and the interface module 102 includes a serial port buffer, wherein the serial port buffer is connected to multiple programmable logic devices and the at least one serial port of the interface module 102 .

[0030] like Figure 2 As shown, the first management module 103 is configured with at least one serial port, and the interface module 102 includes a serial port buffer. The serial port buffer is connected to multiple programmable logic devices (CPLDs) in the interface module 102 and at least one serial port of the first management module 103. Specifically, the serial port buffer receives data from the serial port of the first management module 103 and temporarily stores the data until it is processed by an appropriate programmable logic device. In addition, the CPU can also connect the serial port to the mainboard CPLD, and the BMC can also debug the CPU through this path. However, in this case, debugging requires node selection. A switch button is provided on the mounting ear. By pressing this button, the CPLD on the IO board is triggered to switch, thereby selecting a node for serial port selection and debugging.

[0031] In one embodiment of the present application, the interface module 102 includes a Switch processor, wherein the Switch processor is respectively connected to multiple computing nodes and at least one baseboard management controller.

[0032] like Figure 3 As shown, the interface module 102 includes a Switch processor (PCIe Switch) for expanding PCIe so that the CPU can upgrade the BMC firmware through the operating system.

[0033] In one embodiment of the present application, the interface module 102 includes at least one communication protocol hub, wherein the at least one communication protocol hub is respectively connected to multiple computing nodes and at least one baseboard management controller.

[0034] The communication protocol hub may include an I3C hub and an I2C hub.

[0035] like Figure 4 As shown, the I3C hub and the I2C hub connect multiple computing nodes and at least one baseboard management controller (BMC), so that the BMC can simultaneously read memory information and temperature and voltage information on the motherboard.

[0036] In one embodiment of the present application, the interface module 102 includes at least one cache, wherein the at least one cache is used to store an upgrade program of at least one baseboard management controller, and the upgrade program is used for local upgrade and remote upgrade of the at least one baseboard management controller.

[0037] like Figure 5As shown, interface module 102 includes at least one cache (BIOS Flash). The cache acts as a high-speed storage unit for temporarily storing BMC firmware to be upgraded, avoiding frequent data loading from main memory or a remote server and improving upgrade efficiency. If the system has multiple BMCs, the cache can store their respective upgrade programs, enabling parallel upgrades of multiple BMCs and improving overall operational efficiency. Both local and remote upgrades are supported. Remote upgrades require the host computer software to specify which BMC to upgrade.

[0038] In addition, if Figure 6 As shown, the interface module 102 of this embodiment also includes a USB controller and a USB mux, which are used to aggregate USB and VGA signals and support earpiece key switching. The USB controller is a key component in computer hardware, responsible for managing communication with Universal Serial Bus (USB) devices. The USB mux is an electronic switch used to switch between multiple signal paths or share a single USB interface. It allows different USB devices or hosts to selectively connect to the same USB port as needed, or vice versa, allowing a device to selectively connect to different hosts or paths.

[0039] In one embodiment of the present application, each computing node includes a signal interface and a power supply unit, wherein multiple computing nodes are interconnected through the signal interface, and the power supply unit of each computing node controls the power supply on and off of the corresponding computing node.

[0040] like Figure 7 As shown, in the embodiment of the present application, each node is interconnected with the IO board in the rack using a cable tray, cables, or backplane. Each node can be independently maintained and powered by a separate power supply unit (PSU). Even if other nodes fail or require maintenance, the impact is limited to that single node, without affecting other normally functioning nodes. Each compute node is equipped with a signal interface for communicating with other nodes and the IO board. These interfaces support a variety of data transmission protocols (such as I2C and PCIe), ensuring efficient data exchange and system control. Through the signal interface, multiple compute nodes can be connected directly or indirectly, building a complex communication network. As a result, since each node has its own power supply unit and can independently control the power supply status, the system's fault tolerance is greatly improved. Even if a problem occurs in a node, the source of the problem can be quickly isolated, ensuring stable operation of the remaining components. The independent signal interface and power supply unit make operation of a single node simpler and faster, whether for routine maintenance or emergency repairs, without affecting other components of the system.

[0041] like Figure 7 As shown, the architecture of the embodiment of the present application can save the management interface space for each node, which can be used to place output interface components such as hard disks or network cards. Especially in the design of high-density node architecture, its configuration is often limited by the front and rear window space, and it is necessary to take a reduced configuration operation. Using this architecture can make it have a higher configuration.

[0042] In some embodiments, the relationship between the first management module 103 and the computing module 101 is as follows: Figure 8 As shown, computing module 101 houses only the CPU and necessary chips for CPU startup, such as a power controller, clock chip, memory, and a CPLD for controlling power-on timing. This simplifies the design of computing module 101 and facilitates replacement and maintenance. First management module 103 houses only the BMC and storage chips such as DRAM chips and EMMC required for BMC startup, as well as necessary interfaces such as USB, VGA, and serial ports. This minimizes the impact of plugging and unplugging the module for maintenance.

[0043] In one embodiment of the present application, the server management system further includes a second management module 104. The second management module 104 has the same structure as the first management module 103 and is redundant with the first management module 103. The baseboard management controller of the second management module 104 is connected to the multiple programmable logic devices of the interface module 102.

[0044] It is understandable that the embodiment of the present application needs to support the design of two management modules to manage all nodes at the same time. The two management modules supervise each other while being independent of each other. They can manage at the same time when both are in use, or they can manage separately when only one is in use. At the same time, when a management module fails, it can also be aligned to restart or update the firmware. The solution design is as follows Figure 9 shown.

[0045] Specifically, the second management module 104 shares the same hardware configuration and functional features as the first management module 103. This means both modules include at least one baseboard management controller (BMC) for monitoring the status of compute nodes and performing necessary management operations. A redundant relationship exists between the first management module 103 and the second management module 104. This means the two modules can operate simultaneously or back up each other, ensuring that if either module fails, the other can seamlessly take over its responsibilities, maintaining normal system operation. In practice, the baseboard management controller in the second management module 104 connects to the programmable logic devices (CPLDs) of multiple compute nodes via the interface module 102. This allows it to directly access operational data from each compute node and make appropriate management decisions based on this data. Even in extreme situations, such as when the first management module 103 becomes unavailable due to hardware failures or software errors, the second management module 104 can quickly take over all management tasks, ensuring uninterrupted service. This significantly improves the reliability and flexibility of the server management system.

[0046] It should be noted that the first management module 103 and the second management module 104 of the embodiment of the present application can be made into a gold finger shape and plugged into the IO adapter board. The IO adapter board selects a 4C+ connector. This interface is suitable for hot plugging and easy maintenance, and the number of pins meets the BMC chip selection. At the same time, the management module is designed to be in a standard gold finger shape for easier promotion and modularization, such as Figure 10 shown.

[0047] In one embodiment of the present application, the interface module 102 further includes a detection circuit, which is connected to the first management module 103 and the second management module 104 respectively, and is used to detect the insertion status of the first management module 103 and the second management module 104.

[0048] like Figure 9As shown, the detection circuit of the interface module 102 is connected to the first management module 103 and the second management module 104, respectively, to monitor the insertion status of these two management modules in real time. This means that the system can accurately determine whether each management module is correctly installed and respond accordingly. In actual implementation, when the first management module 103 or the second management module 104 is inserted, the detection circuit will identify this event through specific signals (such as a voltage level change). Once the insertion of a management module is detected, the embodiment of the present application automatically initializes and configures it according to preset logic, including but not limited to loading the necessary firmware and establishing a communication link. This helps quickly put the newly inserted management module into operation, reducing the need for human intervention. If the currently operating management module experiences an abnormality or is removed, the detection circuit can quickly detect it and trigger a redundancy switchover process. In this case, the backup management module will immediately take over all management tasks, ensuring uninterrupted system services. Therefore, the embodiment of the present application, through the integrated detection circuit, achieves real-time monitoring of the insertion status of the management modules, thereby providing more robust support at the hardware level. Even in the event of sudden hardware failures or accidental disconnections, the risk of data loss and service interruption can be effectively avoided.

[0049] In one embodiment of the present application, the programmable logic device of the interface module 102 is configured to perform the following steps: obtain the insertion status detected by the detection circuit; if the insertion status is the in-place state, obtain the management message of the first management module 103 and the second management module 104, and identify whether the management message carries fault information; if the management message carries fault information, record the fault information, disconnect the power supply of the corresponding management module, and switch to another management module; if the management message does not carry fault information, collect the operation data of the computing node.

[0050] like Figure 9 As shown, the programmable logic device (CPLD) of the interface module 102 of the embodiment of the present application is configured to perform a series of steps to ensure stable operation and fault management of the system. Specifically, the CPLD in the interface module 102 first obtains the insertion status information of the first management module 103 and the second management module 104 from the detection circuit to confirm which management module or modules are currently in the in-place state. If the insertion status of a management module (such as the first management module 103 or the second management module 104) is displayed as in-place, the CPLD will attempt to establish communication with the management module and obtain the management message sent by it. The management message typically contains information about the health of the system, such as the status of parameters such as temperature and voltage, as well as any possible fault information.

[0051] Furthermore, after receiving the management message, the CPLD analyzes the data to identify any fault information. If the management message contains fault information, it indicates a possible problem with the corresponding management module. The CPLD will then record this information so that subsequent maintenance personnel can troubleshoot the problem based on the log. The CPLD will also disconnect the power supply to the corresponding management module to prevent the fault from spreading and affecting the entire system, protecting other components from damage and allowing redundant modules to take over and maintain system stability. If no fault information is detected in the management message, the CPLD will continue its normal operation, communicating with each computing node through interface module 102, collecting their operating data (e.g., CPU temperature, motherboard voltage, etc.), and forwarding it to the management module for further processing or storage.

[0052] In one embodiment of the present application, interface module 102 further includes a multiplexer, which is connected to first management module 103 and second management module 104, respectively. The multiplexer is used to detect the insertion status of the external interface. The programmable logic device of interface module 102 is configured to perform the following steps: obtain the insertion status detected by the multiplexer; if the insertion status is in place, connect the external interface to the current management module.

[0053] like Figure 9 As shown, the interface module 102 of the embodiment of the present application not only includes the aforementioned detection circuit but also integrates a multiplexer (MUX). The addition of the multiplexer further enhances the flexibility and reliability of the system. Specifically, the multiplexer is connected to the first management module 103 and the second management module 104, respectively, allowing the system to dynamically switch between the two management modules as needed. In addition to connecting to the management module, the multiplexer is also used to detect the insertion status of external interfaces (such as USB, VGA, etc.). This enables the system to intelligently respond to user operations or the connection of external devices, thereby improving the user experience.

[0054] In actual implementation, the multiplexer first monitors the insertion status of each external interface. For example, when a USB device is plugged in, it sends a signal to the multiplexer notifying it of this event. Based on the external interface's insertion status, the multiplexer dynamically selects a data transmission path between the first management module 103 and the second management module 104. If the currently active management module (assuming it's the first management module 103) fails or requires maintenance, the system seamlessly switches to the backup second management module 104, ensuring uninterrupted service continuity.

[0055] In summary, the interface module 102 of the embodiment of the present application can intelligently manage and monitor the two management modules in the server system, thereby ensuring the continuous and stable operation of the system. The two management modules have equal status and there is no distinction between primary and secondary. They can collect data confidence from all nodes at the same time, but only one can manage all nodes at the same time. Figure 11 The management process of the embodiment of this application is described in detail:

[0056] 1) First, check whether the management module is in place. If not, terminate the process immediately.

[0057] 2) Check again to see if there is any abnormality reported. If there is an abnormality reported, it means that the module is faulty. You need to record the fault information first, then reset it and power it off to prevent information loss caused by direct power off.

[0058] 3) If no abnormality is detected, the management module can work normally, collect and record CPU-related information, and make it available for viewing by the host computer software.

[0059] 4) Then check whether there is access to the USB, VGA or serial port debugging interface. If there is access, it means that the user needs to operate and manage all nodes through this management module. At this time, the CPLD needs to switch the USB, VGA and serial ports to this management node for user viewing and operation until the USB, VGA and serial port devices are connected to another management module and then cut off to ensure the normal use of the external interface.

[0060] In one embodiment of the present application, the interface module 102 also includes an isolation circuit and a control circuit, and the management module includes a trigger circuit, a pre-charging circuit and a startup circuit. The programmable logic device of the interface module controls the insertion or removal of the management module through the isolation circuit, the control circuit, the trigger circuit, the pre-charging circuit and the startup circuit.

[0061] It is understandable that if Figure 12 As shown, a safe and reliable hot-swap operation can also be implemented between the interface module 102 and the management module (such as the first management module 103 or the second management module 104). Figure 12As shown, the interface module 102 also includes an isolation circuit and a control circuit. The isolation circuit is used to isolate the interface module 102 from other system circuits when the management module is not inserted or fails to function, thereby preventing current backflow or interference. The control circuit serves as a control bridge between the programmable logic device and the management module, executing corresponding actions based on the detected status. The management module includes a trigger circuit, a pre-charge circuit, and a startup circuit. The trigger circuit is used to send an "inserted" status signal to the interface module after the management module is inserted into place, notifying it that it can begin subsequent processes (such as pre-charging and startup). The pre-charge circuit is used to pre-charge the power line before the management module is officially powered on, avoiding inrush current caused by excessive capacitive loads, thereby protecting circuit safety. The startup circuit is used to control the main power path to open after the pre-charge is completed, so that the management module enters normal working state.

[0062] In one embodiment of the present application, the programmable logic device of the interface module 102 is configured to perform the following steps: obtaining the insertion status detected by the detection circuit; if the insertion status is the in-place state, sending a power supply signal to the control circuit, and the control circuit controls the start circuit to start the power supply of the pre-charging circuit based on the power supply signal, and the pre-charging circuit responds to charge; after identifying the start signal of the plug-in button of the management module, controlling the isolation circuit to access the signal of the management module.

[0063] It is understandable that the management module of the embodiment of the present application needs to support hot plugging. Due to the working requirements of the server, users need to perform maintenance without powering off or terminating the business. Conventional designs only design hot plugging and blind plugging for computing nodes, and do not consider hot plugging the management module in the case of decoupling computing and management. Figure 12 The hot insertion process of the embodiment of the present application is described in detail as follows:

[0064] The detection circuit in the interface module 102 first detects whether the management module has been correctly inserted into place. If the insertion status is detected to be "in place", that is, it is confirmed that the management module has been safely inserted, the CPLD sends a power supply signal to the control circuit, instructing the system to prepare to power the management module. After the control circuit receives the power supply signal, it controls the start circuit to activate the pre-charging circuit. The pre-charging circuit begins to pre-charge the management module to avoid damage to the device due to excessive transient current when power is directly turned on. The pre-charging circuit gradually increases the voltage according to the instructions of the start circuit to ensure that the management module is powered smoothly until the normal operating voltage is reached. After charging is completed, if the user presses the plug-in button on the management module (indicating that the user wants to connect or disconnect the management module), the CPLD will recognize this start signal. Once the start signal is recognized, the CPLD will control the isolation circuit to establish an electrical connection between the management module and other parts of the system, allowing signal transmission and data exchange.

[0065] During the actual execution process, the detection circuit is responsible for monitoring the physical insertion status of the management module to ensure that subsequent operations are triggered only when the management module is fully inserted. The control circuit acts as an intermediate layer, receiving instructions from the CPLD and controlling specific hardware actions, such as starting the pre-charging circuit. The startup circuit is specifically used to manage the charging process and ensure a smooth transition to full-power operation. The pre-charging circuit is responsible for the actual power supply and may include components such as current-limiting resistors and pre-charging MOSFETs to ensure safe power-on. The isolation circuit is used to cut off the connection between the management module and other system parts when not in use to prevent unnecessary interference or potential electrical problems.

[0066] In one embodiment of the present application, the programmable logic device of the interface module 102 is configured to perform the following steps: after identifying the shutdown signal of the plug-in button of the management module, control the isolation circuit to isolate the signal of the management module; send a power-off signal to the control circuit, and the control circuit controls the start-up circuit to disconnect the power supply of the pre-charging circuit based on the power-off signal; obtain the insertion status detected by the detection circuit, and if the insertion status is in the offline state, allow the management module to be unplugged.

[0067] It is understood that after the user presses the plug / unplug button, this application will safely disconnect power, isolate signals, and ultimately allow the management module to be removed in a predetermined sequence. Specifically, the user presses the plug / unplug button on the management module, indicating that they wish to safely remove the module. The programmable logic device detects the "off signal" emitted by the plug / unplug button, indicating that the removal operation is about to proceed and that a safe power-off process must be initiated. The isolation circuit electrically isolates the management module from the main system's communication channels and power paths, ensuring that signal lines are disconnected to prevent hardware anomalies caused by live signal transmission. The CPLD sends a "power-off signal" to the control circuit, instructing it to begin the power-off process. After receiving the power-off signal, the control circuit issues a command to the startup circuit, which then shuts off power to the pre-charge circuit. The pre-charge circuit stops functioning, cutting off all power to the management module. The detection circuit then re-queries the current insertion status to determine whether the management module has been physically removed (i.e., whether it is in the "offline" state). If the insertion status is "offline," removal of the management module is permitted. If the detection circuit confirms that the management module has been unplugged (the plug-in status is "offline"), it means that the entire power-off and isolation process has been completed, and the user can safely remove the module. In the actual implementation process, after the module is safely removed, the embodiment of the present application can prompt the user that "it is safe to unplug" through an LED indicator light, a buzzer, etc.

[0068] In summary, the hot swap process of the embodiment of the present application is as follows Figure 12 As shown in the figure, the hot insertion process is as follows: 1-2-3-4-5. The hot removal process is as follows: 4-5-2-3-1.

[0069] 1) The CPLD on the interface module 102 detects the presence signal of the management board;

[0070] 2) Power supply is enabled after the presence signal is detected. The management module is equipped with a slow start circuit to prevent the IO board power supply from being cut off due to excessive capacitive load of the inserted board and powering on too quickly.

[0071] 3) After the management module is powered on, it pre-charges the high-level signal to prevent the signal from affecting the IO board after it is connected;

[0072] 4) Wait for the hot swap button to be pressed;

[0073] 5) Press the button to open the isolation circuit on the interface module, allowing the signal to be connected.

[0074] In addition, the programmable logic device of the interface module 102 of the embodiment of the present application is configured to perform the following steps: obtain the insertion status detected by the detection circuit, if the insertion status is in an offline state, record the detection time, and calculate the actual duration between the current time and the detection time; determine whether the actual duration reaches the preset duration, if the actual duration reaches the preset duration, power on the computing module.

[0075] The preset duration, serving as the threshold for delayed startup, can be set based on actual circumstances and is not specifically limited. It should be understood that embodiments of the present application support powering on the computing module even when the management module is not present. The computing module supports operation without the management module, and after the management module is inserted, it can monitor the management module at any time. Specifically, embodiments of the present application use a detection circuit to determine whether the management module is fully inserted. If the insertion status is detected as "offline," meaning the management module is not present, the system records the time when this status was detected (referred to as the "detection moment"). The current time is monitored in real time, and the time difference between the "detection moment" and the "current moment" is calculated, referred to as the "actual duration." If the actual duration is less than the preset duration (e.g., 100ms), the power-on action is temporarily suspended. If the actual duration is greater than or equal to the preset duration, sufficient time has passed, and a power-on signal can be sent to the computing module to initiate operation. Thus, by setting a time threshold to delay power-on, power-on is only initiated after the management module has been "offline" for an extended period, preventing false triggering.

[0076] Furthermore, after the computing module is powered on, the embodiment of the present application can still power on the computing module without the management module. The computing module supports working without the management module, and the management node can be monitored at any time after the management module is inserted. Figure 13As shown, after the computing node is plugged into the power supply module (PSU), it is first powered by the backup power supply (STBY Power) to ensure that basic functions can operate. After the backup power supply is powered on, the computing unit CPU and the programmable logic device CPLD of the interface module 102 of the embodiment of the present application begin to work. At this time, it is possible to check whether the management module is in place. If not, the main power supply (Mail Power) is directly supplied to the computing module to start the CPU without the management module. The CPU startup does not need to be controlled by the management module, and the external pin configuration is directly completed by the CPLD.

[0077] 3) If the management module is in place, power on the management module directly. After the management module is powered on, it is necessary to check whether the main power supply of the CPU is powered on. If it is not powered on, it indicates that the system is ready to start up. The reset control of the management module is directly released to enable normal recognition. In the actual execution process, the user intervenes in the startup process by pressing the hot-swap function key. The embodiment of the present application sets a 5s delay waiting time. If the user does not press the hot-swap function key within these 5s, the startup process will automatically continue to avoid the situation where the user forgets to press the key. If the main power supply of the CPU is powered on, it means that the management module is hot-inserted in the power-on state. It is necessary to wait for the hot-swap button to be pressed before recognition. In the actual execution process, the embodiment of the present application sets a 1s self-test time. It checks whether the hot-swap button is pressed every 1s until the user presses it and then recognizes it. At this time, the management module is in the power-on recognition state and does not affect the system operation.

[0078] According to the server management system proposed in the embodiments of the present application, an interface module is introduced between the computing module and the management module to achieve efficient interconnection and centralized management between multiple computing nodes and the management module. Specifically, each computing node includes a computing unit and a programmable logic device. Multiple computing nodes are connected via the programmable logic device in the interface module. The first management module obtains the operating data of each computing node through the interface module, realizing a design in which a single management module manages multiple nodes. This can reduce the number of management units and lower costs, thereby improving the scalability and management efficiency of the system.

[0079] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0080] In addition, an embodiment of the present application also provides a server, including a management system of the server of the above embodiment.

[0081] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0082] The above is a detailed introduction to a server management system provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only intended to help understand the method and core ideas of the present application. It should be pointed out that, for those skilled in the art, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A server management system, characterized in that: include: A computing module, wherein the computing module includes multiple computing nodes of a server, each computing node includes a computing unit and a programmable logic device, and the computing unit in each computing node is connected to the programmable logic device; An interface module, the interface module comprising a plurality of programmable logic devices, the plurality of programmable logic devices of the interface module being connected to the programmable logic devices of the plurality of computing nodes; a first management module, the first management module including at least one baseboard management controller, the baseboard management controller being connected to the plurality of programmable logic devices of the interface module, obtaining operation data of the plurality of computing nodes through the interface module, and managing the plurality of computing nodes according to the operation data of the plurality of computing nodes; The interface module further includes a detection circuit connected to the first management module and configured to detect an insertion status of the first management module. The programmable logic device of the interface module is configured to perform the following steps: Obtaining the insertion state detected by the detection circuit; if the insertion state is offline, recording the detection time, and calculating the actual time between the current time and the detection time; Determine whether the actual duration reaches a preset duration, and if so, power on the computing module, wherein the computing module supports operation without a management module; The interface module also includes an isolation circuit and a control circuit, and the management module includes a trigger circuit, a pre-charging circuit and a startup circuit. The programmable logic device of the interface module controls the insertion or removal of the management module through the isolation circuit, the control circuit, the trigger circuit, the pre-charging circuit and the startup circuit.

2. The server management system according to claim 1, characterized in that: Each computing node includes a signal interface and a power supply unit, wherein the multiple computing nodes are connected to each other through the signal interface, and the power supply unit of each computing node controls the power supply on and off of the corresponding computing node.

3. The server management system according to claim 1, characterized in that: Also includes: The second management module has the same structure as the first management module, and the second management module and the first management module are redundant modules. The baseboard management controller of the second management module is connected to multiple programmable logic devices of the interface module.

4. The server management system according to claim 3, characterized in that: The detection circuit is connected to the first management module and the second management module respectively, and is used to detect the insertion status of the first management module and the second management module.

5. The server management system according to claim 4, characterized in that: The programmable logic device of the interface module is configured to perform the following steps: obtaining an insertion state detected by the detection circuit; If the insertion state is in place, obtaining management messages from the first management module and the second management module, and identifying whether the management messages carry fault information; If the management message carries fault information, the fault information is recorded and the power supply of the corresponding management module is disconnected; if the management message does not carry fault information, the operation data of the computing node is collected.

6. The server management system according to claim 3, characterized in that: The interface module further includes a multiplexer, which is connected to the first management module and the second management module respectively, and is used to detect the insertion status of the external interface.

7. The server management system according to claim 6, characterized in that: The programmable logic device of the interface module is configured to perform the following steps: Obtaining an insertion state detected by the multiplexer; If the insertion state is the in-place state, the external interface is connected to the current management module.

8. The server management system according to claim 1, characterized in that: The first management module includes at least one serial port, and the interface module includes a serial port buffer, wherein the serial port buffer is connected to a plurality of programmable logic devices of the interface module and the at least one serial port.

9. The server management system according to claim 1, characterized in that: The interface module includes a Switch processor, wherein the Switch processor is connected to the multiple computing nodes and the at least one baseboard management controller respectively.

10. The server management system according to claim 1, characterized in that: The interface module includes at least one communication protocol hub, wherein the at least one communication protocol hub is connected to the multiple computing nodes and the at least one baseboard management controller respectively.

11. The server management system according to claim 1, wherein: The interface module includes at least one cache, wherein the at least one cache is used to store an upgrade program of the at least one baseboard management controller, and the upgrade program is used for local upgrade and remote upgrade of the at least one baseboard management controller.

12. The server management system according to claim 1, characterized in that: The programmable logic device of the interface module is configured to perform the following steps: obtaining an insertion state detected by the detection circuit; If the insertion state is the in-position state, a power supply signal is sent to the control circuit, and the control circuit controls the start-up circuit to start the power supply of the pre-charging circuit based on the power supply signal, and the pre-charging circuit responds by charging; After identifying the start signal of the plug-in / plug-out button of the management module, the isolation circuit is controlled to access the signal of the management module.

13. The server management system according to claim 1, characterized in that: The programmable logic device of the interface module is configured to perform the following steps: After identifying a shutdown signal of the plug-in / plug-out button of the management module, controlling the isolation circuit to isolate the signal of the management module; Sending a power-off signal to the control circuit, wherein the control circuit controls the startup circuit to disconnect the power supply of the pre-charging circuit based on the power-off signal; The insertion state detected by the detection circuit is obtained, and if the insertion state is an offline state, the management module is allowed to be unplugged.

14. A server, characterized in that: include: A server management system according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Cabinet server and out-of-band management method

    CN117827731A

  • Server management system and method, computer equipment and storage medium

    CN118585383A