Server node management assembly, management system, server and management method

By using the switching switches and control modules in the distributed management architecture, the problems of server downtime and increased hardware costs caused by the centralized management architecture are solved, and high-reliability operation is achieved in the event of a failure.

CN121833353APending Publication Date: 2026-04-10XIAMEN YUANCHOU INTELLIGENT COMPUTING TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

The centralized management architecture of existing multi-node servers can cause server downtime in the event of a failure, increasing hardware costs and resulting in poor reliability.

Method used

A distributed management architecture is adopted, which switches the management controller of any server node to either the master or slave management controller by means of a switch and control module, avoiding the need to add an additional management controller and using the existing BMC module to achieve distributed master-slave management.

Benefits of technology

Save on hardware costs, improve server reliability, and ensure continued operation even in the event of a main management controller failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121833353A_ABST
    Figure CN121833353A_ABST
Patent Text Reader

Abstract

The invention discloses a server node management component, a management system, a server and a management method, relates to the technical field of servers, and constructs a server node management component, a control module is used for controlling a plurality of change-over switches, a management controller of any server node is switched to a main management controller, and a management module is used for controlling a plurality of change-over switches; the management controllers of other server nodes are switched to the slave management controllers, when the master management controller fails, any slave management controller is switched to the master management controller, and the functional modules of the server are controlled by using the management controllers in the server nodes. Compared with the prior art, no additional management controller is needed to control the function module of the server, the hardware cost is saved, when the master management controller breaks down, any slave management controller is switched to the master management controller, distributed master-slave management is formed, use of the function module of the server is not affected, and the reliability of the server is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of servers, and in particular to a management component of a server node, a management system, a server and a management method. BACKGROUND

[0002] In recent years, the performance of server processors continues to rise, the number of cores continues to increase, and the frequency is significantly improved. The performance of a single processor can be comparable to multiple CPUs (Central Processing Unit) in the past. This change promotes the evolution of server architecture from traditional dual CPU (two CPUs integrated on a single motherboard) to single CPU, so that a server of the same size can integrate multiple single-board motherboards to form a multi-node architecture. The multi-node server has similar performance to the traditional single-node server, but has stronger fault tolerance. When a single node fails, the remaining nodes can continue to work, but they need to share power supplies, cooling and other peripherals, which significantly increases the complexity of the architecture and the difficulty of management.

[0003] In related technologies, multi-node servers usually use centralized management architecture, which uses a separate BMC (Baseboard Management Controller) management board to uniformly manage and control all nodes and peripherals. This architecture can decouple the management module from other modules. However, since the server is often in a complex environment of high temperature, high electromagnetic interference and needs to run continuously for a long time, once the centralized BMC management board fails, it will directly cause the entire server to crash, seriously affecting the business continuity of the server, reducing the reliability of the server, and the additional BMC management board will increase the hardware cost of the server. SUMMARY

[0004] The present application provides a management component of a server node, a management system, a server and a management method to at least solve the problem of increasing hardware cost and poor reliability in related technologies based on an independent centralized management board to manage the server.

[0005] The present application provides a management component of a server node. The server includes a plurality of server nodes, and each server node includes a management controller. The management component includes: a plurality of switch switches, each switch switch controls at least one function module of the server; a control module, the control module controls the plurality of switch switches, switches the management controller of any server node to a master management controller, switches the management controllers of other server nodes to slave management controllers, and switches any slave management controller to a master management controller when the master management controller fails.

[0006] The application further provides a management system of a server node, comprising: at least one function module; a plurality of server nodes, each server node comprising a management controller; and the management assembly of the server node of the above embodiment, wherein the management assembly is connected with the function module and the management controller of each server node respectively.

[0007] The application further provides a server comprising the management system of the server node of the above embodiment.

[0008] The application further provides a management method of a server node, the method being applied to a control module in the management assembly of the server node of the above embodiment, and comprising: controlling a plurality of switching switches to switch the management controller of any server node into a master management controller and switch the management controllers of other server nodes into slave management controllers; and identifying whether the master management controller is faulty, and switching any slave management controller into the master management controller when the master management controller is faulty.

[0009] The application further provides a non-volatile computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the steps of the management method of the server node.

[0010] The application constructs a management assembly of a server node, comprising a plurality of switching switches and a control module, wherein the control module is used to control the plurality of switching switches to switch the management controller of any server node into a master management controller and switch the management controllers of other server nodes into slave management controllers, and to switch any slave management controller into the master management controller when the master management controller is faulty, and the switching switches control the function module of the server based on the master management controller. The management controller in the server node is used to control the function module of the server, and an additional management controller is not needed to control the function module of the server, thereby saving hardware cost. When the current master management controller is faulty, any slave management controller is switched into the master management controller, a distributed master-slave management is formed, the use of the function module of the server is not affected, and the reliability of the server is improved. Therefore, the technical problem of increasing hardware cost and poor reliability in the related art caused by managing the server based on an independent centralized management board can be solved, and the technical effects of saving hardware cost and improving the reliability of the server are achieved. BRIEF DESCRIPTION OF DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative effort.

[0012] Figure 1 A management architecture of a multi-node server of the related art is shown in FIG. 1.

[0013] Figure 2 A management component of a server node according to an embodiment of the present application is shown in FIG. 2.

[0014] Figure 3 A management system of a server node according to an embodiment of the present application is shown in FIG. 3.

[0015] Figure 4 A management architecture of a multi-node server according to a specific embodiment of the present application is shown in FIG. 4.

[0016] Figure 5 A hardware architecture of a multi-node server according to a specific embodiment of the present application is shown in FIG. 5.

[0017] Figure 6 A management flow of a multi-node server according to a specific embodiment of the present application is shown in FIG. 6.

[0018] Figure 7 A flowchart of a management method of a server node according to an embodiment of the present application is shown in FIG. 7. DETAILED DESCRIPTION

[0019] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, any other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the present application.

[0020] It should be noted that, in the description of the present application, the terms “include”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms “first”, “second” and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0021] Before describing the solutions of the present application, the related art of the present application is described to assist in understanding the solutions of the present application.

[0022] The general rack server in the prior art generally adopts a dual-CPU architecture, that is, one chassis adopts one mainboard, and two CPUs are designed on each mainboard to be interconnected. The CPU performance of the server is becoming more and more powerful, and more and more cores are integrated in the CPU, and the frequency is becoming higher and higher. Therefore, the current processor chip can be equivalent to the performance of multiple CPUs in the past. Therefore, based on the current powerful CPU performance, one CPU is designed on the mainboard, that is, a single-CPU architecture.

[0023] Therefore, in the server of the same size, two single-CPU mainboards can be integrated to form a multi-node server architecture instead of one dual-CPU mainboard. Taking a single-CPU dual-node server as an example, the dual-node server has two independent single-CPU nodes, and each node can work independently. Compared with the dual-CPU single-node server, the performance is similar, but the fault tolerance function is strong, that is, even if one node is broken, the other node can also work, thereby improving the overall reliability of the server. However, compared with the single-node server, the overall architecture of the multi-node server is more complex, because the single-node server exclusively occupies the peripheral devices (power supply, heat dissipation, storage, network, etc.) of the server, and the multi-node server needs to share these peripheral devices. Therefore, the multi-node server is more complex to manage. Therefore, a good management architecture affects the management efficiency and reliability of the multi-node server.

[0024] Therefore, the server in the related art generally adopts a centralized management architecture, that is, a separate BMC management board is designed in the server, and the management board manages all nodes and peripherals such as power supply, heat dissipation, and storage in the server. The advantage is that there is a special management module in the server, and the remaining modules are completely independent, and the decoupling is strong.

[0025] Figure 1 The management architecture of the multi-node server in the related art is shown in FIG. 1. Figure 1 As shown in FIG. 1, Figure 1 The management architecture of the multi-node server in the related art is shown in FIG. 1.

[0026] However, servers often operate 24 / 7 in environments with high temperatures, high noise, and high electromagnetic interference. If the management module of this centralized management architecture fails, the entire server will crash, affecting business operations. Furthermore, the separate design of a management board increases hardware costs.

[0027] To address these technical problems, this application proposes a server node management component.

[0028] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0029] The embodiments of this application provide a management component for a server node. The management component for a server node is described in detail below in conjunction with its constituent parts.

[0030] Figure 2 This is a schematic diagram of a server node management component provided according to an embodiment of this application.

[0031] It should be noted that the server in this embodiment includes multiple server nodes, and each server node includes a management controller.

[0032] like Figure 2 As shown, the management component 10 of the server includes: multiple toggle switches 11 and a control module 12.

[0033] Each switch 11 controls at least one functional module of the server; the control module 12 controls multiple switches 11 to switch the management controller of any server node to the master management controller, switch the management controller of other server nodes to the slave management controller, and switch any slave management controller to the master management controller when the master management controller fails.

[0034] It is understood that the embodiments of this application construct a server node management component 10, including multiple switching switches 11 and a control module 12. The control module 12 is used to control the multiple switching switches 11 to switch the management controller of any server node to the master management controller and switch the management controller of other server nodes to the slave management controller. When the master management controller fails, any slave management controller is switched to the master management controller. The switching switches 11 control the functional modules of the server based on the master management controller. By using the management controller in the server node to control the functional modules of the server, there is no need to add an additional management controller to control the functional modules of the server, saving hardware costs. Moreover, when the current master management controller fails, any slave management controller is switched to the master management controller, forming a distributed master-slave management, which does not affect the use of the server functional modules and improves the reliability of the server.

[0035] The server functional modules in this embodiment include a PSU (Power Supply Unit) power module, a cooling fan module, a hardware storage module, and other functional modules. The switch 11 is used to control the server's functional modules, such as controlling the fan to increase its speed for heat dissipation. The management controller can be a BMC. The control module 12 can be a module inside a CPLD (Complex Programmable Logic Device) in a relay board. The relay board is an intermediate component, and the management controller in the server node manages the server's functional modules through the relay board, which also enables communication between server nodes. The slave management controller in this application can also be understood as a backup management controller.

[0036] The BMC is a dedicated management controller on the server motherboard used to monitor and manage server hardware resources. It is independent of the main processor and operating system, and has its own processor, memory and operating system. It can still work normally even when the server is powered off. The BMC enables administrators to remotely monitor and manage the server's hardware status, and perform fault diagnosis and repair through remote access.

[0037] It should be noted that a server node is a complete functional unit of a server. This node is generally a motherboard, which integrates modules such as CPU, memory, and power supply. A server contains one or more nodes. In the embodiments of this application, the server node refers to the smallest functional unit of a server, including a motherboard. The motherboard is designed with functional modules such as CPU, BMC, and memory. The server node, when combined with peripheral power supply, heat dissipation, storage, and network functional modules, can become a complete server.

[0038] Furthermore, in some embodiments of this application, the input terminal of each switch 11 receives control signals from all server nodes, the output terminal of each switch is connected to the functional module, the control module receives the level signals and heartbeat signals from all server nodes, determines the main management controller based on the level signals and heartbeat signals from all server nodes, and uses the switching signal to control the output terminal of each switch to output the control signal of the main management controller.

[0039] Among them, the level signal is used to determine whether the server node is in place, and the level signal can be represented as NodeX_Present, where NodeX is the Xth server node; the heartbeat signal is used to determine whether the management controller of the server node is working properly, and the heartbeat signal can be represented as NodeX_Heartbeat.

[0040] It is understood that in this embodiment, the input terminal of each switch 11 receives control signals from all server nodes, and the output terminal of each switch 11 is connected to the functional module. The control module 12 receives the level signals and heartbeat signals from all server nodes, determines the main management controller based on the level signals and heartbeat signals from all server nodes, and uses the switching signal to control the output terminal of each switch 11 to output the control signal of the main management controller, so as to realize the functional module of determining the main management controller and using the main management controller to control the server.

[0041] The switching switch in this application embodiment can be a MUX (Multiplexer). A MUX is an electronic component or digital circuit whose main function is to select one signal from multiple input signals and transmit it to a unique output terminal. In this application, the MUX determines which input is selected by switching signals. That is, in this application embodiment, the control signal output by the main management controller can be transmitted to the functional module of the server. Moreover, there are multiple switching switches in this application, corresponding to the number of functional modules.

[0042] Specifically, for example, if there are three functional modules: a PSU power module, a cooling fan module, and a hardware storage module, then there are three switching switches, or three MUXs: one MUX corresponding to the PSU power module, one MUX corresponding to the cooling fan module, and one MUX corresponding to the hardware storage module.

[0043] For example, the input signal of a MUX includes the control signals of the fan cooling modules of two server nodes, and the output signal is the control signal of the fan cooling module of the server. In this embodiment, the control signal of the server node to be selected as the final control signal of the fan cooling module can be determined according to the switching signal.

[0044] In addition, it should be noted that the control signals output by the management controller are different for different functional modules. For example, NodeX_PSU_PMBUS is the control signal for the PSU power module, NodeX_FAN_Control is the control signal for the cooling fan module, and NodeX_BP_Sideband is the control signal for the storage module. Specifically, Node1_PSU_PMBUS is the control signal for the PSU power module output by server node 1, Node1_FAN_Control is the control signal for the cooling fan module output by server node 1, and Node1_BP_Sideband is the control signal for the storage module output by server node 1.

[0045] Furthermore, in some embodiments of this application, the control module 12 reads the level signal and heartbeat signal of the server nodes in sequence according to the sequence number of the multiple server nodes. If the server node is determined to be in the in-position state according to the level signal and to be in the normal working state according to the heartbeat signal, the management controller of the server node with the current sequence number is switched to the master management controller, and the management controllers of the server nodes with other sequence numbers are switched to the slave management controllers.

[0046] In this context, "in-place" refers to the state in which the server node is normally connected to the cluster and can be managed and scheduled. "Out-of-place" may be a state in which the server node is unable to be properly recognized and used by the cluster due to network disconnection, hardware failure, manual offline status (maintenance status), etc. The in-place state is determined based on the level signal and can be set according to specific circumstances. For example, a high level signal indicates an out-of-place state, while a low level signal indicates an in-place state.

[0047] It is understood that the control module 12 in this embodiment reads the level signal and heartbeat signal of the server nodes in sequence according to the sequence number of the multiple server nodes. If the server node is determined to be in the present state according to the level signal and to be in the normal working state according to the heartbeat signal, the management controller of the server node with the current sequence number is switched to the master management controller, and the management controller of the server nodes with other sequence numbers is switched to the slave management controller.

[0048] In short, the judgment conditions of the master management controller include several: (1) according to the preset order of nodes (e.g., according to the node number from smallest to largest); (2) the server node is in place; (3) the heartbeat signal of the server node is normal, and the management controller of the remaining server nodes is the slave management controller.

[0049] Assume there are two server nodes on the server: server node 1 and server node 2.

[0050] 1. First example: Read the level signal Node1_Present and heartbeat signal Node1_Heartbeat of server node 1. If Node1_Present is low, it means that it is in place. If Node1_Heartbeat is in normal state, then directly set the management controller of server node 1 as the master management controller.

[0051] 2. Second example: Read the level signal Node1_Present and heartbeat signal Node1_Heartbeat of server node 1. If Node1_Present is low, it indicates that it is in place. If Node1_Heartbeat is in an abnormal state, then continue to read the level signal Node2_Present and heartbeat signal Node2_Heartbeat of server node 2. If Node2_Present is low, it indicates that it is in place. If Node2_Heartbeat is in a normal state, then the management controller of server node 2 is set as the master management controller.

[0052] In addition, it should be noted that if the management controller of server node 2 is currently the master management controller, the management controller of server node 1 is inserted, and the heartbeat signal is in normal condition and the level signal is in position, then the management controller of server node 2 will still be regarded as the master management controller and the management controller of server node 2 will be regarded as the slave management controller.

[0053] Furthermore, in some embodiments of this application, the control module 12 sets the read and write permissions of the management controller according to the serial number of the server node, wherein the main management controller is set with read and write permissions, and the slave management controller is set with read permissions and write permissions are prohibited.

[0054] It is understood that the permissions of the master and slave management controllers in this application embodiment are distinguished. The master management controller is set to have read and write permissions, while the slave management controller is set to have read permissions and is prohibited from having write permissions. This allows the master management controller to read the data of the server node based on read permissions and to manage the functional modules of the server based on write permissions, thereby ensuring that only the master management controller can perform management, while the slave management controller only has read permissions.

[0055] Furthermore, in some embodiments of this application, the management component 10 of the server node is provided with a first register and a plurality of second registers, wherein the first register stores the operating data of the functional modules, and the second registers are configured corresponding to the server node and store the operating data of the server node.

[0056] The first register, also known as the peripheral register, is used to store the running data of the functional module and can be named Peripheral_Register. The number of the second registers corresponds to the number of server nodes and can be called mirror registers. They are used to store the running data of the corresponding server nodes and can be named Node_Mirror_Register.

[0057] It is understood that the server node management component 10 in this application embodiment also includes a first register and a plurality of second registers, so that the main management controller can uniformly manage the server based on the operation data of the functional modules read from the first register and the operation data of the server node read from the second registers.

[0058] In addition, it should be noted that if the server node where the main management controller is located is damaged or removed, the management controller that was previously selected from the management controllers with the level signal in the in-place state and the heartbeat signal in the normal state will be selected as the main management controller according to the sequence number. The server will be managed according to the data in its corresponding second register, and the data in the first register will be updated after scanning the peripherals of the server and the status of each server node.

[0059] Furthermore, in some embodiments of this application, the management component 10 of the server node further includes a bus module.

[0060] The bus module is connected to the bus of each management controller, and each management controller communicates with each other through the bus module. The bus can be an I2C (Inter-Integrated Circuit) bus.

[0061] It is understood that the server node management component 10 in this application embodiment also includes a bus module, which communicates with the management controller of multiple server nodes through the bus module, so that the main management controller manages the functional modules of the server based on the data of each server node.

[0062] Furthermore, in some embodiments of this application, the functional module includes a cooling fan. After the management component is initialized, the control module controls the cooling fan to output a preset target speed. When the main management controller fails, during the process of switching between the main management controller and the slave management controller, the control module controls the cooling fan to output the maximum speed.

[0063] The target speed is an initial target speed set according to the specific situation, and no specific limit is made on it.

[0064] It is understood that in this embodiment of the application, after the server node management component 10 is initialized and before the master management controller and slave management controller are determined, the control module 12 first takes over the control of the server's cooling fan, cools the server with a default target speed to prevent the server from overheating, and controls the cooling fan to output the maximum speed during the process of switching the master management controller and slave management controller when the master manager fails, so as to avoid the server overheating and crashing.

[0065] The following specific examples illustrate the process by which the management component of a server node determines the management controller and the process of switching the management controller.

[0066] Example 1: The example is a server with two nodes (node ​​1 and node 2).

[0067] Both Node 1 and Node 2 are equipped with independent management controllers (BMC1 and BMC2, respectively). The management components include two switches (MUX1 and MUX2) and one control module. MUX1 controls the server's power module, and MUX2 controls the cooling fan module. The inputs of both switches are connected to the control signals of BMC1 and BMC2, and their outputs are connected to the power module and the fan module, respectively.

[0068] Initially, the control module prioritizes reading the level signal (low level indicates presence) and heartbeat signal (periodic pulses indicate normal operation) of node 1 according to its sequence number. Since node 1 is present and functioning normally, the control module sets BMC1 as the master controller and BMC2 as the slave controller, and sends a switching signal to MUX1 and MUX2. At this time, MUX1 outputs power control signals for BMC1 (such as adjusting the supply voltage), and MUX2 outputs fan speed control signals for BMC1, achieving dominant control of the functional modules. Simultaneously, the control module grants read / write permissions to BMC1 (allowing modification of power thresholds and fan speed parameters), but only grants read permissions to BMC2 (allowing access only to the current device status).

[0069] If BMC1 experiences a sudden failure (such as interruption of the heartbeat signal or abnormal level signal), the control module will immediately trigger a switching mechanism upon detecting the anomaly: switching BMC1 to the master management controller, synchronously updating the switching signals of MUX1 and MUX2 so that their outputs switch to the control signals of BMC2; simultaneously adjusting the permission configuration to grant read and write permissions to BMC2, while BMC1 (in a fault state) is restricted to read-only. During this process, the bus module ensures real-time communication between BMC1 and BMC2. The first register stores functional module data such as power supply voltage and fan speed in real time, while the second registers corresponding to nodes 1 and 2 record their respective operating data such as CPU load and memory usage, ensuring that status information is not lost during master-slave switching and that the server continues to operate stably.

[0070] Example 2: The following description is based on a 3-node server with 3 server nodes (node ​​1, node 2, and node 3).

[0071] Each of the three nodes is equipped with a management controller, BMC1, BMC2, and BMC3. The management components include three switches (MUX1, MUX2, and MUX3) and one control module. MUX1 controls the power supply module, MUX2 controls the cooling fan module, and MUX3 controls the hardware storage module. The input of each switch receives control signals from BMC1, BMC2, and BMC3, and the output is connected to the corresponding functional module.

[0072] Initially, the control module reads the level signal and heartbeat signal sequentially according to the node number. It first checks node 1; if its level signal is low (indicating it's present) and its heartbeat signal is normal (stable periodic pulses), the control module sets BMC1 as the master controller and BMC2 and BMC3 as slave controllers. Subsequently, the control module sends switching signals to the three switches, causing MUX1 to output the power control signal for BMC1 (e.g., setting the power supply), MUX2 to output the fan speed control signal for BMC1, and MUX3 to output the hardware memory allocation signal for BMC1. Simultaneously, the control module grants read / write permissions to BMC1 (allowing modification of power thresholds, fan speed limits, etc.), while only granting read permissions to BMC2 and BMC3 (allowing them to view the current status of each module).

[0073] When BMC1 malfunctions (e.g., the heartbeat signal disappears, or the level signal goes high), the control module detects the anomaly and immediately checks the next node – node 2. If node 2 is present and functioning normally, BMC2 is switched to the master controller, and BMC1 (in a fault state) and BMC3 are set as slave controllers. At this time, the control module updates the switching signal, switching the outputs of MUX1, MUX2, and MUX3 to the control signal of BMC2. Simultaneously, permissions are adjusted, granting read and write permissions to BMC2, while restricting BMC1 and BMC3 to read-only.

[0074] If BMC2 also fails, the control module will continue to monitor node 3, switch BMC3 to the master management controller, and repeat the switching process. During this period, the bus module ensures real-time communication between BMC1, BMC2, and BMC3. The first register stores operating data of functional modules such as power consumption, fan speed, and network bandwidth. The second registers corresponding to nodes 1, 2, and 3 record information such as CPU utilization and memory usage. When the master management controller switches, the data in these registers ensures that the new master management controller can quickly obtain the device status, guaranteeing uninterrupted server operation.

[0075] According to the server node management component proposed in this application embodiment, the control module is used to control multiple switching switches to switch the management controller of any server node to the master management controller and switch the management controllers of other server nodes to the slave management controllers. When the master management controller fails, any slave management controller is switched to the master management controller. The switching switches control the server's functional modules based on the master management controller. By using the management controller in the server node to control the server's functional modules, there is no need to add an additional management controller to control the server's functional modules, saving hardware costs. Furthermore, when the current master management controller fails, any slave management controller is switched to the master management controller, forming a distributed master-slave management system that does not affect the use of the server's functional modules and improves the server's reliability.

[0076] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0077] Embodiments of this application also provide a management system for server nodes.

[0078] Figure 3 This is a schematic diagram of a server node management system provided according to an embodiment of this application.

[0079] like Figure 3 As shown, the management system 20 of the server node includes: at least one functional module 21, multiple server nodes 22, and a management component 10 for the server nodes.

[0080] Each server node includes a management controller; the management component 10 is connected to the functional module 21 and the management controller of each server node.

[0081] It is understood that the embodiments of this application also construct a server node management system 20, including a functional module 21, a server node 22 and a server node management component 10. The management component 10 is linked to the functional module and the management controller of each server node 22. The management controller manages the functional module 22 to realize the management of the server node and the server.

[0082] It should be noted that the description of the features in the embodiment corresponding to the management system of the server node can be found in the relevant description of the embodiment corresponding to the management component of the server node, and will not be repeated here.

[0083] According to the server node management system proposed in the embodiments of this application, the management components and functional modules are linked to the management controller of each server node, and the management controller manages the functional modules to realize the management of server nodes and service requests.

[0084] Embodiments of this application also provide a server, including a management system for the server nodes described in the above embodiments.

[0085] The following describes the server, server management system and management components of this application embodiment through a specific embodiment, taking the server containing two server nodes (hereinafter referred to as nodes) as an example.

[0086] This application establishes a distributed master-slave management system. Figure 4 This is a schematic diagram of the management architecture of a multi-node server provided according to a specific embodiment of this application. The overall server architecture is as follows: Figure 4 As shown, these two nodes represent the minimum possible computer system and can operate independently. Each node contains a BMC management unit (equivalent to the management controller mentioned above) for managing its own node and the server (some multi-node server nodes do not have BMC management, but this solution requires each node to have a BMC management module). The server also includes a relay board, through which nodes manage server peripherals (i.e., functional modules) and communication between nodes. Finally, the server needs to perform distributed management and master / slave switching according to a specific process algorithm.

[0087] Figure 5 This is a schematic diagram of the hardware architecture of a multi-node server according to a specific embodiment of this application. The hardware scheme of the server architecture is as follows: Figure 5 As shown, the control signals for Node1 are: Node1_PSU_PMBUS, Node1_FAN_Control, and Node1_BP_Sideband; the control signals for Node2 are: Node2_PSU_PMBUS, Node2_FAN_Control, and Node2_BP_Sideband. These three control signals correspond to different functional modules.

[0088] These control signals are connected in pairs to the CPLD on the relay board, enabling signal switching from one to many. In actual designs, there may be more control signals, which are also processed according to this logic. The goal is to ensure that each node's control signals can control the server's peripheral functional modules through controllable signal switching.

[0089] The presence signals of Node 1 and Node 2: The Node1_Present and Node2_Present signals are connected to the CPLD. The CPLD determines whether the current node is present by the level change of these two signals. It is usually designed that a high level means the node is not present, and a low level means the node is present.

[0090] The heartbeat signals of the BMC modules on both nodes, Node1_Heartbeat and Node2_Heartbeat, are connected to the CPLD, which is the control module. This allows the CPLD to determine whether the BMC modules on both nodes are functioning correctly. The I2C buses of the two BMC modules are also connected to the CPLD, enabling communication between the two BMC modules and with the CPLD on the relay board.

[0091] The CPLD internally includes a control module for controlling the operation of signal switching switches. It also contains a peripheral information register: the master node's `_Peripheral_Register`. The master node (i.e., the master node's management controller, as mentioned above) stores acquired peripheral information data in this register. The BMC (Browser Controller) uses this information to analyze the server's peripheral operating status for management. The master node's BMC can read and write (update data) this register, but the standby node's (i.e., the standby node's management controller, as mentioned above) BMC can only read and not write.

[0092] The CPLD internally has multiple node mirror registers: Node_Mirror_Register, with one register for each node. This register stores management information data collected by the node's BMC. Only the node itself can write (i.e., update data) the data in this register; other nodes can only read it.

[0093] Figure 6 This is a schematic diagram illustrating the management process of a multi-node server according to a specific embodiment of this application. The workflow of the entire server node management architecture is as follows: Figure 6 As shown, the details are as follows:

[0094] 1. CPLD initialization.

[0095] After the server is powered on, the CPLD on the relay board starts working first to complete the initialization. At this time, the CPLD takes over the server's fan control and cools the server at a default speed to prevent the server from overheating.

[0096] 2. Detect the level signal of the detection node to determine whether it is in place and whether the heartbeat signal is normal.

[0097] The CPLD sequentially checks the presence signals of the nodes to confirm their physical presence. After the BMC on each node completes initialization, the heartbeat signal begins operation. Each node's BMC scans its own device's operational status and writes its management information data to the Node_Mirror_Register of each child node, updating the data periodically. Upon receiving the heartbeat signals from each node, the CPLD considers that the node's BMC module is functioning correctly.

[0098] 3. Determine the primary node and backup node based on the test results.

[0099] If node 1 is in place and its heartbeat signal is normal, the CPLD will set node 1 as the primary node and the other nodes as backup nodes. If node 1 is not in place, or if node 1 is in place but its heartbeat signal is not received, node 2 will be set as the primary node and the other nodes as backup nodes.

[0100] Therefore, there are three conditions for determining the master node: (1) it follows the preset order of nodes (e.g., from smallest to largest node number); (2) the node is in place; and (3) the node's heartbeat signal is normal. The remaining nodes are backup nodes.

[0101] 4. Manage the server.

[0102] Once the master node is determined, the CPLD control signal switching switch connects the master node's control signals to the corresponding server peripherals. After the signal is completed, the CPLD notifies the master node BMC that it can manage all modules of the server.

[0103] After receiving the notification, the master node (BMC) gains control and begins scanning the shared peripherals of the server, placing the scanned peripheral information data for management into the master node's `_Peripheral_Register`. The master node then integrates the peripheral management information data with the management data read from each node to manage the server in a unified manner.

[0104] In addition, it should be noted that if the master node is damaged or removed, the CPLD will take over the system fan control and set the fan to maximum to prevent the system from overheating and crashing. The second node will be set as the master node. Node 2 will then manage the server based on the management information data in the original master node's _Peripheral_Register, updating the peripheral status register information after scanning the server peripherals and the status of each node.

[0105] If node 2 is currently the master node managing the server, and node 1 is inserted, node 2 will still be the master node, and node 1 will be the backup node.

[0106] In summary, this application embodiment utilizes the existing BMC modules on each node of the server to replace the additional BMC management board in the original solution for system management, saving costs. Moreover, management authority is distributed among each node, so even if the main management node fails, other nodes can take over the management of the server as backups, greatly improving the reliability of the server.

[0107] Embodiments of this application also provide a method for managing server nodes.

[0108] Figure 7 This is a flowchart of a server node management method provided according to an embodiment of this application.

[0109] like Figure 7 As shown, the server node management method, applied to the control module of the server node management component in the above embodiment, includes the following steps:

[0110] In step S101, multiple switching switches are controlled to switch the management controller of any server node to the master management controller and switch the management controller of other server nodes to the slave management controller.

[0111] In step S102, it is determined whether the main management controller is faulty. If the main management controller is faulty, any management controller will be switched to the main management controller.

[0112] For a description of the features in the corresponding embodiment of the management method for this server node, please refer to the relevant description of the corresponding embodiment of the management component of the server node, which will not be repeated here.

[0113] According to the server node management method proposed in the embodiments of this application, when a failure of the main management controller is detected, any slave management controller can be switched to the main management controller, forming a distributed master-slave management, which does not affect the normal use of the server and improves the reliability of the server.

[0114] Embodiments of this application also provide a non-volatile computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described server node management method embodiments when running.

[0115] In one exemplary embodiment, the aforementioned non-volatile computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0116] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0117] The foregoing has provided a detailed description of a server node management component, management system, server, and management method provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A management component for a server node, characterized in that, The server includes multiple server nodes, each server node including a management controller, wherein the management component includes: Multiple toggle switches, each of which controls at least one functional module of the server; The control module controls the plurality of switching switches to switch the management controller of any server node to the master management controller and switch the management controller of other server nodes to the slave management controller. When the master management controller fails, any of the slave management controllers is switched to the master management controller.

2. The server node management component according to claim 1, characterized in that, Each of the switching switches receives control signals from all the server nodes at its input, and the output of each of the switching switches is connected to the functional module. The control module receives the level signals and heartbeat signals from all the server nodes, determines the main management controller based on the level signals and heartbeat signals from all the server nodes, and uses the switching signals to control the output of each of the switching switches to output the control signals of the main management controller.

3. The server node management component according to claim 2, characterized in that, The control module reads the level signal and heartbeat signal of the server nodes sequentially according to their serial numbers. If the level signal indicates that the server node is in an in-position state and the heartbeat signal indicates that the server node is in a normal working state, the management controller of the server node with the current serial number is switched to the master management controller, and the management controllers of the other server nodes are switched to the slave management controllers.

4. The server node management component according to claim 3, characterized in that, The control module sets the read and write permissions of the management controller according to the serial number of the server node. The main management controller is set with the read and write permissions, and the slave management controller is set with the read permission and the write permission is disabled.

5. The server node management component according to claim 1, characterized in that, The management component is provided with a first register and a plurality of second registers, wherein the first register stores the operating data of the functional module, and the second registers are configured corresponding to the server node and store the operating data of the server node.

6. The server node management component according to claim 1, characterized in that, Also includes: A bus module is provided, which is connected to the bus of each of the management controllers, and each of the management controllers communicates with each other through the bus module.

7. A management system for server nodes, characterized in that, include: At least one functional module; Multiple server nodes, each of which includes a management controller; The management component of the server node according to any one of claims 1 to 6, wherein the management component is connected to the functional module and the management controller of each server node respectively.

8. A server, characterized in that, The management system includes the server node as described in claim 7.

9. A method for managing server nodes, characterized in that, The method is applied to the control module of the management component of the server node according to any one of claims 1 to 6, wherein the method includes: Control multiple switching switches to switch the management controller of any server node to the master management controller and switch the management controller of other server nodes to the slave management controller; If the main management controller is faulty, any of the slave management controllers will be switched to become the main management controller if the main management controller is faulty.

10. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the server node management method as described in claim 9.

Citation Information

Cited By

  • Server hardware circuits and their mode control methods, server

    CN122309451A