Server management system and server

By introducing an interface module between the computing module and the management module, and using programmable logic devices to connect multiple computing nodes, efficient interconnection and centralized management are achieved, the hardware maintenance complexity and cost problems caused by the sharing of the motherboard of the computing unit or management unit is solved, and the system's scalability and management efficiency are improved.

CN120353303AActive Publication Date: 2025-07-22INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510837402.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-22
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

In the prior art, the computing unit or management unit shares the motherboard, and the motherboard needs to be replaced at the same time due to damage, which increases the complexity and cost of hardware maintenance.

Method used

The interface module is introduced, multiple computing nodes are connected through programmable logic devices in the interface module, and multiple computing nodes are managed by the substrate management controller to achieve efficient interconnection and centralized management.

Benefits of technology

Reduce management units, reduce costs, improve system scalability and management efficiency, and improve system reliability and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353303A_ABST
    Figure CN120353303A_ABST
Patent Text Reader

Abstract

The invention discloses a management system of a server and the server, relates to the technical field of computers, and realizes efficient interconnection and centralized management between a plurality of computing nodes and a management module by introducing an interface module between the computing module and the management module. Specifically, each computing node comprises a computing unit and a programmable logic device, a plurality of computing nodes are connected through the programmable logic device in the interface module, and the first management module obtains the operation data of each computing node through the interface module, so that the design that one management module manages multiple nodes is realized, the management units can be reduced, and the management efficiency is improved. And the cost is reduced, so that the expandability and the management efficiency of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and particularly to a management system for a server and a server. Background Art

[0002] With the explosive growth of the Internet, the demand for high-concurrency processing capabilities has emerged. The clustered deployment of various high-computing-power servers has become a key means to improve performance. In the clustered deployment of servers, decoupling the design and unified management of each module in the server can reduce maintenance costs.

[0003] In the related art, servers are managed independently in a single node, that is, each node has its own separate management chip. Whether in a conventional server chassis or in a whole cabinet design, the server architecture usually manages each node independently. If you need to view the status of each node, you need to view it separately, or add another management chip for unified maintenance.

[0004] However, the management chip is board-mounted on the server motherboard. During the actual execution process, the computing unit or the management unit shares the motherboard. Once the motherboard fails, it needs to be replaced simultaneously, which not only affects the computing function but also the normal operation of the entire management system, thereby increasing the maintenance difficulty and replacement cost. Summary of the Invention

[0005] This application provides a management system for a server and a server, which at least solves the problem in the related art that the computing unit or the management unit shares the motherboard, and when the motherboard is damaged, it needs to be replaced simultaneously, increasing the complexity and cost of hardware maintenance.

[0006] This application provides a management system for a server, including: a computing module, the computing module includes multiple computing nodes of the server, each computing node includes a computing unit and a programmable logic device, and the computing unit in each computing node is connected to the programmable logic device; an interface module, the interface module includes multiple programmable logic devices, and the multiple programmable logic devices of the interface module are connected to the programmable logic devices of the multiple computing nodes; a first management module, the first management module includes at least one baseboard management controller, the baseboard management controller is connected to the multiple programmable logic devices of the interface module, obtains the operation data of the multiple computing nodes through the interface module, and manages the multiple computing nodes according to the operation data of the multiple computing nodes.

[0007] This application also provides a server, including: the management system for a server in the above embodiment. Through this application, by introducing an interface module between the computing module and the first management module, efficient interconnection and centralized management between multiple computing nodes and the first management module are achieved. Specifically, each computing node includes a computing unit and a programmable logic device. Multiple computing nodes are connected through the programmable logic devices in the interface module. The first management module obtains the operation data of each computing node through this interface module, realizing the design of managing multiple nodes with one management module, which can reduce management units, lower costs, and thus improve the scalability and management efficiency of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0009] Figure 1 It is a signal topology diagram of a server management system provided by an embodiment of this application; Figure 2 It is a multi-node management CPU debug signal topology diagram provided by an embodiment of this application; Figure 3 It is a multi-node management PCIe signal topology diagram provided by an embodiment of this application; Figure 4 It is a multi-node management low-speed signal topology diagram provided by an embodiment of this application; Figure 5 It is a multi-node management SPI signal topology diagram provided by an embodiment of this application; Figure 6 It is a multi-node management USB signal topology diagram provided by an embodiment of this application; Figure 7 It is a topology diagram of a computing node provided by an embodiment of this application; Figure 8 It is an association example diagram between a management module and a computing module provided by an embodiment of this application; Figure 9 It is a design diagram of a dual management module provided by an embodiment of this application; Figure 10 It is an interface form example diagram of a management module provided by an embodiment of this application; Figure 11 It is a control flow diagram of a dual management module provided by an embodiment of this application; Figure 12 It is a circuit flow diagram of an interface module and a management module provided by an embodiment of this application; Figure 13The power-on flowchart of a computing module provided by an embodiment of the present application. Detailed implementation manners

[0010] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0011] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0012] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0013] An embodiment of the present application provides a management system for a server, as Figure 1 shown, including: a computing module 101, an interface module 102, and a first management module 103.

[0014] Among them, the computing module 101 includes multiple computing nodes of the server, each computing node includes a CPU (Computing Unit) and a CPLD (Complex Programmable Logic Device), and the computing unit CPU in each computing node is connected to the programmable logic device CPLD; the interface module 102 includes multiple programmable logic devices CPLD, and the multiple programmable logic devices CPLD of the interface module are connected to the programmable logic devices CPLD of the multiple computing nodes; the first management module 103 includes at least one BMC (Baseboard Management Controller), and the baseboard management controller BMC is connected to the multiple programmable logic devices CPLD of the interface module 102, obtains the operation data of the multiple computing nodes through the interface module 102, and manages the multiple computing nodes according to the operation data of the multiple computing nodes.

[0015] Specifically, as Figure 1As shown in the figure, the computing module 101 of the embodiment of the present application includes a plurality of computing nodes. Each computing node consists of a computing unit (such as a CPU) and necessary support chips (such as a power controller, a clock chip, memory, etc.), and each computing node is also equipped with a programmable logic device (CPLD). This design enables each computing node to operate independently and also communicate with the interface module through the CPLD inside it. The interface module 102, as a key component connecting the computing module 101 and the first management module 103, includes multiple CPLDs for establishing connections with each computing node and also provides various interfaces (such as I2C, etc.) to achieve data transmission and the transfer of control signals. The first management module 103 includes at least one baseboard management controller (BMC). This baseboard management controller is responsible for monitoring the health status of the entire system, including key parameters such as temperature and voltage, and can perform corresponding management operations based on the collected data. In particular, this management module supports the hot-swap function, which means that the management module can be replaced or maintained without shutting down the system, greatly improving the availability and flexibility of the system. During the actual execution process, as Figure 1 shown, the CPU records the internal temperature status, the voltage status on the motherboard, and other statuses in the motherboard CPLD register. After the CPLD on the IO board obtains information through LTPI communication with the motherboard CPLD, the BMC reads the relevant information by reading the CPLD register on the IO board through I2C.

[0016] In an embodiment of the present application, the first management module 103 includes at least one serial port, and the interface module 102 includes a serial port buffer. Among them, the serial port buffer is connected to multiple programmable logic devices and at least one serial port of the interface module 102.

[0017] As Figure 2 shown, the first management module 103 is configured with at least one serial port, while the interface module 102 includes a serial port buffer. The serial port buffer is connected to multiple programmable logic devices CPLDs in the interface module 102 and at least one serial port of the first management module 103. Specifically, the serial port buffer receives data from the serial port of the first management module 103 and temporarily stores these data until they are processed by appropriate programmable logic devices. In addition, the CPU can also connect the serial port to the motherboard CPLD, and the BMC can also debug the CPU through this path. However, at this time, node selection is required for debugging. There is a switching button on the earpiece. By pressing this button, the CPLD on the IO board can be triggered to switch, so as to select the node for serial port selection and debugging.

[0018] In an embodiment of the present application, the interface module 102 includes a Switch processor. Among them, the Switch processor is respectively connected to multiple computing nodes and at least one baseboard management controller.

[0019] As Figure 3 shown, the interface module 102 includes a Switch processor (PCIe Switch) for expanding PCIe, enabling the CPU to upgrade the BMC firmware through the operating system.

[0020] In an embodiment of the present application, the interface module 102 includes at least one communication protocol hub, where the at least one communication protocol hub is respectively connected to a plurality of computing nodes and at least one baseboard management controller.

[0021] Among them, the communication protocol hub may include an I3C hub and an I2C hub.

[0022] As Figure 4 shown, the I3C hub and the I2C hub are connected to a plurality of computing nodes and at least one baseboard management controller (BMC), enabling the BMC to simultaneously read memory information and on-board temperature and voltage information on the motherboard.

[0023] In an embodiment of the present application, the interface module 102 includes at least one cache, where the at least one cache is used to store the upgrade program for at least one baseboard management controller, and the upgrade program is used for local upgrade and remote upgrade of at least one baseboard management controller.

[0024] As Figure 5 shown, the interface module 102 includes at least one cache (BIOS Flash). The cache serves as a high-speed storage unit for temporarily storing the BMC firmware program to be upgraded, avoiding frequent data loading from the main memory or remote server, and improving the upgrade efficiency. If there are multiple BMCs in the system, the cache can store their respective upgrade programs respectively to achieve parallel upgrade of multiple BMCs and improve the overall operation and maintenance efficiency. It supports local upgrade and remote upgrade. For remote upgrade, it is necessary to specify in the host computer software which BMC to use for the upgrade.

[0025] In addition, as Figure 6 shown, the interface module 102 in the embodiment of the present application further includes: a USB Controler and a USBMux for aggregating USB and VGA, and supporting earhook button switching at the same time. Among them, the USB Controler is a key component in computer hardware responsible for managing communication with Universal Serial Bus (USB) devices. The USB Mux is an electronic switch used to switch or share a single USB interface between multiple signal paths. It allows different USB devices or hosts to selectively connect to the same USB port as needed, or conversely, allows a device to selectively connect to different hosts or paths.

[0026] In one embodiment of the present application, each computing node includes a signal interface and a power supply unit. Among them, multiple computing nodes are interconnected through the signal interface, and the power supply unit of each computing node controls the power on and off of the corresponding computing node.

[0027] As Figure 7 shown, in the embodiment of the present application, each node is interconnected with the IO board in the form of a Cable Tray, cables, or a backplane in the rack. Each node can be independently maintained, and each node is powered by a separate power supply unit (PSU). Even when other nodes fail or need to be maintained, only the affected node is limited, and it will not affect other normally operating nodes. Each computing node is equipped with a signal interface for communicating with other nodes and the IO board. These interfaces support a variety of data transmission protocols (such as I2C, PCIe, etc.) to ensure efficient data exchange and system control. Through the signal interface, direct or indirect connections can be achieved between multiple computing nodes to build a complex communication network. Thus, since each node has its own power supply unit and can independently control the power supply state, the fault tolerance of the system is greatly improved. Even if a problem occurs in a certain node, the fault source can be quickly isolated to ensure the stable operation of the rest. And the independent signal interface and power supply unit make the operation of a single node more simple and fast. Whether it is for daily maintenance or emergency repair, it will not affect other parts of the system.

[0028] As Figure 7 shown, the architecture of the embodiment of the present application can save the management interface space for each node, which can be used to place out-interface components such as hard disks or network cards. Especially in the design of a high-density node architecture, its configuration is often limited by the space of the front and rear windows, and a reduction in configuration has to be adopted. With this architecture, it can have a higher configuration.

[0029] In some embodiments, the association between the first management module 103 and the computing module 101, as Figure 8 shown, the computing module 101 only places the CPU and the necessary chips required for CPU startup, such as a power controller, a clock chip, memory, a CPLD for controlling the power-on timing, etc., so as to simplify the design of the computing module 101 as much as possible for easy replacement and maintenance. The first management module 103 only places the BMC, DRAM particles required for BMC startup, storage chips such as EMMC, and necessary interfaces such as USB, VGA, and serial ports, so as to minimize its design so that the impact is minimized when it is plugged and unplugged for maintenance.

[0030] In an embodiment of the present application, the management system of the server further includes: a second management module 104. Among them, the second management module 104 has the same structure as the first management module 103. The second management module 104 and the first management module 103 are redundant modules to each other. The baseboard management controller of the second management module 104 is connected to multiple programmable logic devices of the interface module 102.

[0031] It can be understood that the embodiments of the present application need to support the design of simultaneously managing all nodes by two management modules. The two management modules monitor each other and are independent at the same time. They can be managed simultaneously when both are present, and can also be managed independently when only one is present. At the same time, when one management module fails, it can also be restarted or updated with firmware. The solution design is as Figure 9 shown.

[0032] Specifically, the second management module 104 has the same hardware configuration and functional characteristics as the first management module 103. This means that they both include at least one baseboard management controller (BMC) for monitoring the status of computing nodes and performing necessary management operations. A redundant relationship is formed between the first management module 103 and the second management module 104. This means that the two management modules can work simultaneously or back up each other, ensuring that in the event of the failure of any one module, the other can seamlessly take over its responsibilities and maintain the normal operation of the system. During actual execution, the baseboard management controller in the second management module 104 is connected to the programmable logic devices (CPLDs) of multiple computing nodes through the interface module 102, which enables it to directly obtain the operation data of each computing node and make corresponding management decisions accordingly. Even in extreme cases, such as when the first management module 103 becomes unavailable due to hardware failures, software errors, etc., the second management module 104 can quickly take over all management work to ensure that the service is not interrupted. Thereby, the reliability and flexibility of the server management system are greatly improved.

[0033] It should be noted that the first management module 103 and the second management module 104 in the embodiments of the present application can be made into the form of a gold finger and inserted on the IO adapter board. A 4C+ connector is selected on the IO adapter board. This interface is suitable for hot plugging, convenient for maintenance, and the number of pins meets the selection of the BMC chip. At the same time, designing the management module into a standard gold finger form is easier to promote and modularize, as Figure 10 shown.

[0034] In an embodiment of the present application, the interface module 102 further includes a detection circuit. The detection circuit is respectively connected to the first management module 103 and the second management module 104, and the detection circuit is used to detect the insertion states of the first management module 103 and the second management module 104.

[0035] As Figure 9As shown, the detection circuit of the interface module 102 is respectively connected to the first management module 103 and the second management module 104, and is used to monitor the insertion states of these two management modules in real time. This means that the system can accurately determine whether each management module is correctly installed in place and take corresponding response measures accordingly. During actual execution, when the first management module 103 or the second management module 104 is inserted, the detection circuit will identify this event through specific signals (such as level changes). Once the insertion of a certain management module is detected, the embodiment of the present application can automatically perform initialization configuration on it according to the preset logic, including but not limited to loading necessary firmware, establishing a communication link, etc. This helps the newly inserted management module enter the working state quickly and reduces the need for manual intervention. If the currently working management module has an abnormality or is removed, the detection circuit can quickly sense it and trigger the redundancy switching process. At this time, the standby management module will immediately take over all management work to ensure that the system service is not interrupted. Thus, the embodiment of the present application realizes real-time monitoring of the insertion states of the management modules through the integrated detection circuit, thereby being able to provide more stable support at the hardware level. Even in the face of sudden hardware failures or accidental disconnections, the risk of data loss and service interruption can be effectively avoided.

[0036] In an embodiment of the present application, the programmable logic device of the interface module 102 is configured to execute the following steps: obtain the insertion state detected by the detection circuit; if the insertion state is the in-place state, obtain the management messages of the first management module 103 and the second management module 104, and identify whether the management messages carry fault information; if the management messages carry fault information, record the fault information, then disconnect the power supply of the corresponding management module, and switch to another management module, if the management messages do not carry fault information, collect the operation data of the computing node.

[0037] As Figure 9 shown, the programmable logic device (CPLD) of the interface module 102 in the embodiment of the present application is configured to execute a series of steps to ensure the stable operation and fault management of the system. Specifically, the CPLD in the interface module 102 first obtains the insertion state information of the first management module 103 and the second management module 104 from the detection circuit to confirm which or which management modules are currently in the in-place state. If the insertion state of a certain management module (such as the first management module 103 or the second management module 104) shows the in-place state, the CPLD will attempt to establish communication with this management module and obtain the management messages sent by it. Among them, the management messages usually contain information about the system health status, such as the status of parameters such as temperature and voltage, as well as any possible fault information.

[0038] Further, after receiving the management message, the CPLD analyzes the data to identify whether there is fault information. If it is found that the management message carries fault information, it indicates that there may be a problem with the corresponding management module. Then these information will be recorded so that the subsequent maintenance personnel can troubleshoot the fault according to the log. At the same time, the CPLD will also perform the operation of disconnecting the power supply of the corresponding management module to prevent the fault from spreading and affecting the entire system, protect other components from damage, and allow the redundant module to take over the work to maintain the system stability. If no fault information is detected in the management message, the CPLD will continue its normal operation process, that is, communicate with each computing node through the interface module 102, collect their operation data (such as CPU temperature, motherboard voltage, etc.), and forward it to the management module for further processing or storage.

[0039] In an embodiment of the present application, the interface module 102 further includes a multiplexer, which is respectively connected to the first management module 103 and the second management module 104. The multiplexer is used to detect the insertion state of the external interface. The programmable logic device of the interface module 102 is configured to perform the following steps: obtain the insertion state detected by the multiplexer; if the insertion state is the in-position state, connect the external interface to the current management module.

[0040] As Figure 9 shown, the interface module 102 of the embodiment of the present application not only includes the above detection circuit, but also additionally integrates a multiplexer (MUX). The addition of the multiplexer further enhances the flexibility and reliability of the system. Specifically, the multiplexer is respectively connected to the first management module 103 and the second management module 104, allowing the system to dynamically switch between these two management modules as needed. In addition to connecting to the management module, the multiplexer is also used to detect the insertion state of external interfaces (such as USB, VGA, etc.). This enables the system to intelligently respond to user operations or the connection of external devices, thereby improving the user experience.

[0041] During the actual execution process, the multiplexer first monitors the insertion state of each external interface. For example, when a USB device is inserted, it sends a signal to the multiplexer to notify this event. According to the insertion state of the external interface, the multiplexer can dynamically select the data transmission path between the first management module 103 and the second management module 104. If the currently active management module (assuming it is the first management module 103) fails or needs maintenance, the system can seamlessly switch to the standby second management module 104 to ensure that the service continuity is not affected.

[0042] In summary, the interface module 102 of the present application embodiment can intelligently manage and monitor two management modules in the server system, thus ensuring the continuous and stable operation of the system. The two management modules are equal in status and there is no primary or secondary distinction. They can collect all node data information simultaneously, but only one can manage all nodes at the same time. The following combines Figure 11 to elaborate on the management process of the present application embodiment in detail: 1) First, detect whether the management module is in place. If it is not in place, directly end. 2) Then, detect whether there is an abnormal report. If there is an abnormal report, it indicates that the module has a fault. First, record the fault information, then reset and power off to prevent information loss caused by directly powering off.

[0043] 3) If no abnormality is detected, this management module can work normally, collect and record CPU-related information, and be available for the host computer software to view.

[0044] 4) Then, detect whether there is a USB, VGA, or serial port debugging interface connected. If there is a connection, it indicates that the user needs to operate and manage all nodes through this management module. At this time, the CPLD needs to switch the USB, VGA, and serial port to this management node for the user to view and operate until the USB, VGA, and serial port devices are connected to another management module and then switched away to ensure the normal use of the external interface.

[0045] In an embodiment of the present application, the interface module 102 further includes an isolation circuit and a control circuit, and the management module includes a trigger circuit, a pre-charge circuit, and a start circuit. The programmable logic device of the interface module controls the insertion or removal of the management module through the isolation circuit, the control circuit, the trigger circuit, the pre-charge circuit, and the start circuit.

[0046] It can be understood that, as Figure 12 shown, a safe and reliable hot-plug operation can also be achieved between the interface module 102 and the management module (such as the first management module 103 or the second management module 104). As Figure 12As shown, the interface module 102 further includes an isolation circuit and a control circuit. Among them, the isolation circuit is used to isolate the interface module 102 from other system circuits when the management module is not inserted or fails, preventing current backflow or interference. The control circuit serves as a control bridge between the programmable logic device and the management module, and performs corresponding actions according to the detected status. The management module includes a trigger circuit, a pre-charge circuit, and a startup circuit. Among them, the trigger circuit is used to send a "plugged in" status signal to the interface module after the management module is inserted in place, notifying it that the subsequent process (such as pre-charging and startup) can begin. The pre-charge circuit is used to pre-charge the power supply line before the management module is officially powered on, avoiding surge current caused by excessive capacitive load, thereby protecting the circuit safety. The startup circuit is used to control the main power path to be turned on after the pre-charging is completed, enabling the management module to enter the normal working state.

[0047] In an embodiment of the present application, the programmable logic device of the interface module 102 is configured to perform the following steps: obtain the insertion status detected by the detection circuit; if the insertion status is the in-place status, send a power supply signal to the control circuit, and the control circuit controls the startup circuit to start the power supply of the pre-charge circuit based on the power supply signal, and the pre-charge circuit responds to charge; after recognizing the startup signal of the plug-and-play button of the management module, control the isolation circuit to access the signal of the management module.

[0048] It can be understood that the management module of the embodiment of the present application needs to support hot plugging. Due to the working requirements of the server, users need to perform maintenance without power-off and without terminating the service. The conventional design only designs hot plugging and blind plugging for the computing node, and does not consider the hot plugging design of the management module in the case of decoupling of computing and management. The following combines Figure 12 The hot insertion process of the embodiment of the present application will be described in detail as follows: The detection circuit in the interface module 102 first detects whether the management module is correctly inserted in place. If the detected insertion status is the "in-place status", that is, it is confirmed that the management module has been safely inserted, the CPLD sends a power supply signal to the control circuit, indicating that the system is ready to supply power to the management module. After receiving the power supply signal, the control circuit controls the startup circuit to activate the pre-charge circuit. The pre-charge circuit starts to pre-charge the management module to avoid damaging the device due to excessive transient current during direct power-on. The pre-charge circuit gradually increases the voltage according to the instruction of the startup circuit to ensure a smooth power supply to the management module until the normal working voltage is reached. After the charging is completed, if the user presses the plug-and-play button on the management module (indicating that the user wishes to connect or disconnect the management module), the CPLD will recognize this startup signal. Once the startup signal is recognized, the CPLD will control the isolation circuit to establish an electrical connection between the management module and other parts of the system, allowing signal transmission and data exchange.

[0049] During the actual execution process, the detection circuit is responsible for monitoring the physical insertion state of the management module to ensure that subsequent operations are triggered only when the management module is fully inserted. The control circuit, as an intermediate layer, receives instructions from the CPLD and controls specific hardware actions, such as starting the pre-charge circuit. The start-up circuit is specifically used to manage the charging process to ensure a smooth transition to the full-power operation state. The pre-charge circuit is responsible for the actual power supply and may include components such as current-limiting resistors and pre-charge MOSFETs to ensure a safe power-on. The isolation circuit is used to disconnect the management module from other parts of the system when not in use to prevent unnecessary interference or potential electrical problems.

[0050] In an embodiment of the present application, the programmable logic device of the interface module 102 is configured to perform the following steps: after recognizing the closing signal of the plug-and-play button of the management module, control the isolation circuit to isolate the signal of the management module; send a power-off signal to the control circuit, and the control circuit controls the start-up circuit to disconnect the power supply of the pre-charge circuit based on the power-off signal; obtain the insertion state detected by the detection circuit, and if the insertion state is the offline state, allow the management module to be unplugged.

[0051] It can be understood that after the user presses the plug-and-play button, the present application can safely cut off the power, isolate the signal, and finally allow the management module to be unplugged in a predetermined order, as follows: The user presses the plug-and-play button on the management module, indicating the desire to safely unplug the module. The programmable logic device detects the "closing signal" sent by the plug-and-play button, indicating that the unplugging operation is about to be executed and the safe power-off process needs to be entered. The isolation circuit electrically isolates the communication channel and power path between the management module and the main system to ensure that the signal line is disconnected and avoid hardware abnormalities caused by the transmission of live signals. The CPLD sends a "power-off signal" to the control circuit to notify it to start the power supply shutdown process. After receiving the power-off signal, the control circuit issues an instruction to the start-up circuit, and the start-up circuit immediately shuts down the power supply to the pre-charge circuit. The pre-charge circuit stops working and cuts off all power supplies flowing to the management module. Then, it queries the current insertion state feedback by the detection circuit again to determine whether the management module has been completely unplugged physically (i.e., whether it is in the "offline state"). If the insertion state is the "offline state", the management module is allowed to be unplugged. If the detection circuit confirms that the management module has been unplugged (the insertion state is the "offline state"), it means that the entire power-off and isolation process has been completed, and the user can safely remove the module. During the actual execution process, after the module is safely removed, the embodiment of the present application can prompt the user "it is safe to unplug" through means such as LED indicators and buzzers.

[0052] In summary, the hot-swap process of the embodiment of the present application is as Figure 12 shown. The hot-insertion process proceeds according to 1-2-3-4-5. The hot-unplugging process proceeds according to 4-5-2-3-1.

[0053] 1) Detect the presence signal of the CPLD detection management board on the interface module 102; 2) Enable the power supply after detecting the presence signal. The management module is equipped with a soft-start circuit to prevent the capacitive load of the inserted board from being too large and the power of the IO board from being pulled dead due to too fast power-on; 3) After the management module is powered on, give a high-level signal for pre-charging to prevent the signal from affecting the IO board after being connected; 4) Wait for the hot-swap button to be pressed; 5) After pressing the button, turn on the isolation circuit on the interface module to allow the signal to be connected.

[0054] In addition, the programmable logic device of the interface module 102 in the embodiment of the present application is configured to perform the following steps: obtain the insertion state detected by the detection circuit. If the insertion state is the offline state, record the detection time, and calculate the actual duration between the current time and the detection time; determine whether the actual duration reaches a preset duration. If the actual duration reaches the preset duration, power on the calculation module.

[0055] Among them, the preset duration is used as the threshold for delayed start-up and can be set according to actual situations without specific limitations. It can be understood that the embodiment of the present application supports powering on the calculation module even when there is no management module. The calculation module supports working without a management module, and the management module can monitor the management module at any time after being inserted. Specifically, in the embodiment of the present application, the detection circuit is used to determine whether the management module has been inserted in place. If the detected insertion state is the "offline state", that is, the management module is not in place, the system will record the time point when this state is detected (referred to as the "detection time"). The current time is monitored in real time, and the time difference from the "detection time" to the "current time", that is, the "actual duration", is calculated. If the actual duration is less than the preset duration (such as 100 ms), the power-on action is not performed temporarily; if the actual duration is greater than or equal to the preset duration, it means that enough time has been waited, and a power-on signal can be sent to the calculation module to make it start working. Thus, by setting a time threshold for delayed power-on, it is ensured that power-on is only performed after the management module has been in the "offline state" for a long time, avoiding mis-triggering.

[0056] Further, after the calculation module is powered on, the embodiment of the present application supports powering on the calculation module even when there is no management module. The calculation module supports working without a management module, and the management module can monitor the management node at any time after being inserted. For example Figure 13As shown, after the power supply module (PSU) is inserted into the computing node, it is first powered by the standby power supply to ensure that the basic functions can run. After the standby power supply is powered on, the computing unit CPU of the interface module 102 and the programmable logic device CPLD of the interface module 102 start to work. At this time, it can be checked whether the management module is in place. If it is not in place, the main power supply (Mail Power) is directly supplied to the computing module to start the CPU without the management module. The startup of the CPU does not need to be controlled by the management module, and the external pin configuration is directly completed by the CPLD.

[0057] 3) If the management module is in place, power is directly supplied to the management module. After the management module is powered on, it is necessary to detect whether the main power supply of the CPU has been powered on. If it has not been powered on, it indicates that it is ready to start up, and the reset control of the management module is directly released to make it recognized normally. In the actual execution process, the user intervenes in the startup process by pressing the hot-swap function key. The embodiment of the present application sets a delay waiting time of 5s. If the user does not press the hot-swap function key within these 5s, the startup process will continue automatically to avoid the situation that the user forgets to press the key; if the main power supply of the CPU has been powered on, it means that the management module is hot-inserted in the powered-on state, and it needs to wait for the hot-swap button to be pressed before recognition. In the actual execution process, the embodiment of the present application sets a self-check time of 1s, and detects whether the hot-swap button is pressed every 1s until the user presses it before recognition. At this time, the management module is in the powered-on and unrecognized state, which does not affect the system operation.

[0058] According to the management system of the server proposed by the embodiment of the present application, by introducing an interface module between the computing module and the management module, efficient interconnection and centralized management between multiple computing nodes and the management module are realized. Specifically, each computing node includes a computing unit and a programmable logic device. Multiple computing nodes are connected through the programmable logic device in the interface module, and the first management module obtains the operation data of each computing node through this interface module, realizing the design of one management module managing multiple nodes, which can reduce the management unit and cost, thereby improving the scalability and management efficiency of the system.

[0059] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0060] In addition, the embodiment of the present application also provides a server, including the management system of the server in the above embodiment.

[0061] Those skilled in the art may further realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0062] The above has introduced in detail a server management system provided by this application. Specific examples are used herein to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A management system for a server, characterized in that, including: a computing module, the computing module includes a plurality of computing nodes of a server, each computing node includes a computing unit and a programmable logic device, and the computing unit in each computing node is connected to the programmable logic device; an interface module, the interface module includes a plurality of programmable logic devices, and connections are made between the plurality of programmable logic devices of the interface module and the programmable logic devices of the plurality of computing nodes; a first management module, the first management module includes at least one baseboard management controller, the baseboard management controller is connected to the plurality of programmable logic devices of the interface module, obtains the operation data of the plurality of computing nodes through the interface module, and manages the plurality of computing nodes according to the operation data of the plurality of computing nodes.

2. The management system of the server according to claim 1, characterized in that, Each of the computing nodes includes a signal interface and a power supply unit. Among them, the plurality of computing nodes are connected to each other through the signal interface, and the power supply unit of each computing node controls the power on / off of the corresponding computing node.

3. The management system of the server according to claim 1, wherein, It further includes: a second management module, the second management module has the same structure as the first management module, the second management module and the first management module are redundant modules to each other, and the baseboard management controller of the second management module is connected to the plurality of programmable logic devices of the interface module.

4. The management system of the server according to claim 3, characterized in that, The interface module further includes a detection circuit, the detection circuit is respectively connected to the first management module and the second management module, and the detection circuit is used to detect the insertion states of the first management module and the second management module.

5. The management system of the server according to claim 4, characterized in that, The programmable logic device of the interface module is configured to execute the following steps: obtain the insertion state detected by the detection circuit; if the insertion state is the in-place state, obtain the management messages of the first management module and the second management module, and identify whether the management messages carry fault information; if the management messages carry fault information, record the fault information, then cut off the power supply of the corresponding management module, and if the management messages do not carry fault information, collect the operation data of the computing nodes.

6. The management system of the server according to claim 3, characterized in that, The interface module further includes a multiplexer, the multiplexer is respectively connected to the first management module and the second management module, and the multiplexer is used to detect the insertion state of an external interface.

7. The management system of the server according to claim 6, characterized in that, The programmable logic device of the interface module is configured to execute the following steps: obtain the insertion state detected by the multiplexer; if the insertion state is the in-place state, connect the external interface to the current management module.

8. The management system of the server according to claim 1, characterized in that, The first management module includes at least one serial port, and the interface module includes a serial port buffer, where the serial port buffer is connected to the plurality of programmable logic devices of the interface module and the at least one serial port.

9. The management system of the server according to claim 1, characterized in that, The interface module includes a Switch processor, where the Switch processor is respectively connected to the plurality of computing nodes and the at least one baseboard management controller.

10. The management system of the server according to claim 1, characterized in that, The interface module includes at least one communication protocol hub, where the at least one communication protocol hub is respectively connected to the plurality of computing nodes and the at least one baseboard management controller.

11. The management system of the server according to claim 1, characterized in that, The interface module includes at least one cache, wherein the at least one cache is used to store upgrade programs for the at least one baseboard management controller, and the upgrade programs are used for local upgrade and remote upgrade of the at least one baseboard management controller.

12. The management system of the server according to claim 4, wherein The interface module further includes an isolation circuit and a control circuit, the management module includes a trigger circuit, a pre-charge circuit and a start-up circuit, and the programmable logic device of the interface module controls the insertion or removal of the management module through the isolation circuit, the control circuit, the trigger circuit, the pre-charge circuit and the start-up circuit.

13. The management system of the server according to claim 12, characterized in that, The programmable logic device of the interface module is configured to perform the following steps: Obtain the insertion state detected by the detection circuit; If the insertion state is the in-place state, send a power supply signal to the control circuit, and the control circuit controls the start-up circuit to start the power supply of the pre-charge circuit based on the power supply signal, and the pre-charge circuit responds to charge; After recognizing the start signal of the plug-and-play button of the management module, control the isolation circuit to access the signal of the management module.

14. The management system of the server according to claim 12, wherein The programmable logic device of the interface module is configured to perform the following steps: After recognizing the close signal of the plug-and-play button of the management module, control the isolation circuit to isolate the signal of the management module; Send a power-off signal to the control circuit, and the control circuit controls the start-up circuit to disconnect the power supply of the pre-charge circuit based on the power-off signal; Obtain the insertion state detected by the detection circuit, and if the insertion state is the offline state, allow the management module to be removed.

15. A server, characterized in that, Comprising: The management system of the server according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Equipment access detection device, PCIe routing card, system, control method and medium

    CN112783817A

  • Server backboard system and server operation control method

    CN115562942A

  • Cabinet server and out-of-band management method

    CN117827731A

  • Server management system and method, computer equipment and storage medium

    CN118585383A

Cited By

  • Multi-node server, multi-node server power failure control method, device and equipment

    CN120610614A

  • Server clock topology switching system and method

    CN122308556A

  • A server clock topology switching system and method

    CN122308556B