A control system, method, server cluster system, and medium of a server

By adopting a centralized power supply and management approach, the problems of excessive power consumption and low management efficiency of server clusters have been solved, achieving efficient resource utilization and stable system operation.

CN119739275BActive Publication Date: 2025-12-12INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202412000456.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-12-12
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

Server clusters consume excessive power and are inefficient to manage. Existing technologies that require separate power supply modules for each server result in resource waste and difficulty in unified management.

Method used

A centralized power supply method is adopted, which obtains the power supply requirements of each server through the control module and adjusts the output power of the power supply module to achieve unified power supply and management of the server cluster.

Benefits of technology

It reduced the power consumption of the server cluster, improved management efficiency and system reliability, reduced resource waste, and ensured business continuity and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119739275B_ABST
    Figure CN119739275B_ABST
Patent Text Reader

Abstract

The application discloses a kind of server's control system, method, server cluster system and medium, it is related to server technical field, power supply module is set to power supply for all servers in entire server cluster, control module can effectively obtain the power supply demand of each server, and according to the actual demand of each server, the output power of power supply module output to each server is regulated, so as to realize the centralized power supply of all servers in server cluster, reduce power consumption, control module can be aimed at all servers unified power consumption management and server system state management, improve management efficiency.Through centralized management and real-time monitoring, operating personnel can directly understand the running condition of all servers in server cluster through control module, quickly respond to fault condition, improve the reliability and stability of entire server cluster system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of servers, and in particular to a server control system and method, a server cluster system and a medium. BACKGROUND

[0002] With the rapid development of information technology, the importance and complexity of servers as the core components of data centers continue to increase. In order to meet the changing business needs and technical challenges, in complex application scenarios such as data centers, a server cluster composed of multiple servers is required to work cooperatively. In order to realize power supply to each server, a power supply module is separately arranged in each server in the server cluster, and the power supply module supplies power to the corresponding server under the control of the server. In this case, the power consumption of the entire server cluster is very large, and the operating personnel can only monitor and manage the power consumption of each server one by one, which is inefficient.

[0003] Therefore, how to reduce the power consumption of the server cluster and improve the management efficiency of the server cluster is a problem to be solved by those skilled in the art. SUMMARY

[0004] The purpose of the embodiments of the present application is to provide a server control system and method, a server cluster system and a medium, which can solve the problem of excessive power consumption caused by a server cluster.

[0005] To solve the above technical problems, the embodiments of the present application provide a server control system applied to a server cluster, which comprises:

[0006] a power supply module, the power supply module comprising a plurality of output terminals connected one-to-one with a plurality of servers in the server cluster, the output terminal of the power supply module being connected with a power supply terminal of the corresponding server;

[0007] a control module, a first end of the control module being connected with data terminals of the plurality of servers in the server cluster through a first communication bus, and a second end of the control module being connected with a control terminal of the power supply module through a second communication bus, the control module being configured to acquire power supply requirements of the servers through the first communication bus, determine a power supply strategy of the power supply module according to the power supply requirements, and control the power supply module to work based on the power supply strategy through the second communication bus; the power supply strategy of the power supply module comprising output powers of the output terminals of the power supply module.

[0008] In some embodiments, all the servers in the server cluster are divided into N groups of servers, and each group of servers comprises at least two servers; the control module comprises N control sub-modules connected one-to-one with the N groups of servers, and the power supply module comprises N power supply sub-modules connected one-to-one with the N groups of servers; N is a positive integer greater than 1.

[0009] The power supply sub-module is configured to supply power to the corresponding group of servers under the control of the corresponding control sub-module.

[0010] In some embodiments, the control sub-module comprises:

[0011] The first switch module comprises a fixed end and a plurality of switch ends connected one-to-one with a plurality of servers in the corresponding group of servers, and the switch end of the first switch module is connected with the data end of the corresponding server;

[0012] The first multiplexer has a first end connected with the fixed end of the first switch module;

[0013] The processor has a first end connected with the second end of the first multiplexer, and is configured to control the first switch module to switch through the first multiplexer to obtain the power supply demand of each server in the corresponding group of servers.

[0014] In some embodiments, the control sub-module further comprises:

[0015] The second multiplexer has a first end connected with the second end of the processor;

[0016] The pin expander comprises a bus end and M pin ends, the pin end of the pin expander is connected with the signal end of the corresponding power supply sub-module, and the bus end is connected with the second end of the second multiplexer;

[0017] The processor is further configured to control the power supply sub-module to start through the pin expander, and obtain the working state of the power supply sub-module through the pin expander; the working state of the power supply sub-module comprises an insertion state of the power supply sub-module, a power supply state of the power supply sub-module, and a fault state of the power supply sub-module.

[0018] In some embodiments, the processor comprises:

[0019] The management controller has a first end connected with the second end of the first multiplexer, a second end connected with the first end of the second multiplexer, and a third end connected with the control end of the corresponding power supply sub-module through the second communication bus, and is configured to obtain the power supply demand of each server in the corresponding group of servers through the first multiplexer, and control the working of the corresponding power supply sub-module according to the power supply demand;

[0020] The programmable logic device has a first end connected with the detection end of the management controller, and is configured to detect the fault state of the management controller;

[0021] The control sub-module further comprises:

[0022] a second switch module, a first end of which is connected with the third end of the first multiplexer, a second end of which is connected with the third end of the second multiplexer, a third end of which is connected with the second end of the programmable logic device, and a fourth end of which is connected with the fourth end of the management controller of another control submodule of the control module except for itself;

[0023] The programmable logic device is further configured to send a takeover signal to the management controller of another control submodule connected with the fourth end of the second switch module when detecting that the management controller fails, so that the management controller of another control submodule takes over the work of the failed management controller through the second switch module.

[0024] In some embodiments, the specific process of detecting the failure state of the management controller comprises:

[0025] sending a heartbeat signal to the management controller and starting a timer;

[0026] if the response signal returned by the management controller is not received when the count value of the timer reaches the preset value, it is determined that the management controller fails.

[0027] In some embodiments, the server control system further comprises:

[0028] a temperature sensor arranged in the corresponding group of servers, an output end of which is connected with the first input end of the processor, for detecting the working temperature of the corresponding group of servers, so that the processor determines whether the working temperature of the server is abnormal;

[0029] and / or,

[0030] a liquid leakage sensor arranged in the corresponding group of servers, an output end of which is connected with the second input end of the processor, for detecting the liquid leakage condition of the corresponding group of servers, so that the processor stops the power supply for the server with the liquid leakage condition when the liquid leakage condition exists in the server.

[0031] To solve the above technical problems, the embodiments of the present application further provide a server cluster system comprising a server cluster and a server control system as described above, and the output end of the server control system is connected with the power supply end of each server in the server cluster, and the input end is connected with the data end of each server.

[0032] To solve the above technical problems, the embodiments of the present application further provide a server control method applied to the server control system as described above; the server control method comprises:

[0033] obtaining the power supply demand of each server in the server cluster;

[0034] determining the power supply strategy of the power supply module according to the power supply demand;

[0035] The control power supply module works based on a power supply strategy; the power supply strategy of the power supply module includes output power of each output terminal of the power supply module.

[0036] To solve the above technical problems, the embodiment of the present application further provides a computer readable storage medium, and the computer readable storage medium stores a computer program, and the computer program is executed by a processor to realize the steps of the control method of the server.

[0037] From the above technical solution, it can be seen that the power supply module is arranged to supply power to all servers in the entire server cluster, the control module can effectively obtain the power supply demand of each server, and the output power of the power supply module output to each server is regulated according to the actual demand of each server, thereby realizing centralized power supply of all servers in the server cluster. The beneficial effects of the present application are that the centralized power supply mode is adopted to reduce power consumption, the control module can uniformly manage power consumption and server system state for all servers, and the management efficiency is improved. Through centralized management and real-time monitoring, the operating personnel can directly understand the running status of all servers in the server cluster through the control module, quickly respond to fault conditions, and improve the reliability and stability of the entire server cluster system. BRIEF DESCRIPTION OF DRAWINGS

[0038] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0039] Figure 1 A structural schematic diagram of a server control system provided by the embodiment of the present application is shown in the figure.

[0040] Figure 2 A structural schematic diagram of a server whole cabinet provided by the embodiment of the present application is shown in the figure.

[0041] Figure 3 Another structural schematic diagram of a server whole cabinet provided by the embodiment of the present application is shown in the figure.

[0042] Figure 4 An internal structural schematic diagram of a server control system provided by the embodiment of the present application when one control submodule is arranged is shown in the figure.

[0043] Figure 5 An internal structural schematic diagram of a server control system provided by the embodiment of the present application when two control submodules are arranged is shown in the figure.

[0044] Figure 6 A flowchart of a control method of a server provided by an embodiment of the present application is shown in FIG. 1. DETAILED DESCRIPTION

[0045] The technical solutions in the embodiments of the present application will be clearly and completely described with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, but not all embodiments of the present application. Based on the embodiments in the present application, any other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0046] The terms “include” and “have” and any variations thereof in the specification and the above drawings of the present application are intended to cover the non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but can include steps or units not listed.

[0047] In order for those skilled in the art to better understand the technical solutions of the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0048] Next, a control system of a server provided by an embodiment of the present application is described in detail, which is applied to a server cluster; referring to FIG. 2, Figure 1 as shown, Figure 1 A structural diagram of a control system of a server provided by an embodiment of the present application is shown in FIG. 2; the control system of the server includes:

[0049] A power supply module 1 includes a plurality of output terminals corresponding to a plurality of servers in the server cluster one by one, and the output terminal of the power supply module 1 is connected to the power supply terminal of the corresponding server;

[0050] A control module 2, a first end is connected to the data terminal of each server in the server cluster through a first communication bus, and a second end is connected to the control terminal of the power supply module 1 through a second communication bus, for obtaining the power supply demand of each server through the first communication bus, determining the power supply strategy of the power supply module 1 according to the power supply demand, and controlling the power supply module 1 to work based on the power supply strategy through the second communication bus; the power supply strategy of the power supply module 1 includes the output power of each output terminal of the power supply module 1.

[0051] It is not difficult to understand that the power supply module 1 needs to supply power to all servers in the server cluster, so it needs to establish a one-to-one corresponding power supply circuit with each server, so when the server cluster includes A servers, the power supply module 1 will set up A output ends corresponding to the connection of A servers, respectively realizing the power supply of each server. The power supply demand of the server refers to the demand power of the server work, so the corresponding power supply strategy of the power supply module 1 needs to be controlled, that is, the output power of each output end of the power supply module 1 needs to be configured. In this centralized power supply case, only a single power supply module 1 needs to be redundantly designed, reducing the number of PSUs that need to be redundantly set, reducing the power consumption of the entire server cluster, and avoiding excessive resource waste.

[0052] It is not difficult to understand that the power supply module 1 needs to supply power to all servers in the server cluster, so it needs to establish a one-to-one corresponding power supply circuit with each server, so when the server cluster includes A servers, the power supply module 1 will set up A output ends corresponding to the connection of A servers, respectively realizing the power supply of each server. The power supply demand of the server refers to the demand power of the server work, so the corresponding power supply strategy of the power supply module 1 needs to be controlled, that is, the output power of each output end of the power supply module 1 needs to be configured. In this centralized power supply case, only a single power supply module 1 needs to be redundantly designed, reducing the number of PSUs that need to be redundantly set, reducing the power consumption of the entire server cluster, and avoiding excessive resource waste.

[0053] It should be noted that the specific implementation of the server cluster is not particularly limited herein, and the number of servers, the specific type and implementation of the server are not particularly limited herein. The specific type and implementation of the control module 2 and the power supply module 1 are not particularly limited herein, and the control module 2 can be implemented by a processor or a controller, and the power supply module 1 can be implemented by a power supply system including multiple PSUs. The specific type and implementation of the first communication bus and the second communication bus are not particularly limited herein, the first communication bus is mainly used for data transmission and can be implemented by an I2C bus (Inter-Integrated Circuit, two-wire serial bus), and the control module 2 controls the power supply module 1 through the second communication bus, so it can be implemented by a PMBUS bus (Power Management Bus, power management bus).

[0054] It can be understood that the control module 2 can be implemented by a PMC (Power Management Controller, power management controller), and the PMC can be set as a separate board card connected to the server and other modules such as the power supply module 1 through a connector. When the control module 2, i.e., the PMC board fails, the maintenance personnel only needs to replace the PMC board instead of the entire control system, and there is no need to maintain the power supply module 1, which greatly saves the maintenance cost and time. Through this modular design, the fault repair is more rapid and convenient, the downtime of the server cluster system is reduced, the maintainability, reliability and management efficiency of the entire server cluster system are improved, and the user is provided with higher system availability, lower business interruption risk, more efficient resource utilization and lower operation and maintenance cost. The control system provided by the application can also be applied to other application scenarios that require multiple servers or multiple computers, can realize centralized power supply of multiple devices, and is not limited to the application scenario of the server cluster.

[0055] As a specific embodiment, in order to facilitate application, the server cluster and the control system of the server are arranged in a server cabinet, and each server in the server cluster is a Node (computing node) of the control system of the server. Referring to FIG. 1, Figure 2 Figure 2 ​The structural schematic diagram of a server whole cabinet provided by the embodiment of the present application; the control system designed by the present application can be applied to a whole cabinet with 8 servers, and each server can be used as a computing node. Two control systems, Power Management Module 0 and Power Management Module 1, are designed in the server whole cabinet in the embodiment, Power Management Module 0 is used to manage and control the power supply of the four computing nodes Node 0-3, and Power Management Module 1 is used to manage and control the power supply of the four computing nodes Node 4-7. The control system is connected with the four computing nodes corresponding to itself respectively, and is used to monitor and manage the running condition of each computing node. In addition, a communication link is designed between Power Management Module 0 and Power Management Module 1 to realize the master-backup redundancy design. The two control systems are the backup system of each other, when normally operating, Power Management Module 0 is only responsible for the monitoring and management of Node 0-3, and Power Management Module 1 is responsible for Node 4-7. When Power Management Module 1 fails, Power Management Module 0 can be used as an alternative system to take over the monitoring and management of Node 4-7, and at this time, Power Management Module 0 can still manage Node 0-3, ensuring that the monitoring and management functions of the whole cabinet are not affected.

[0056] Further, referring to Figure 3 shown, Figure 3 The structural schematic diagram of another server whole cabinet provided by the embodiment of the present application; two control systems are arranged in the whole cabinet in the embodiment to realize the power supply control of 16 computing nodes (16 servers), Power Management Module 0 is used to manage and control the power supply of the eight computing nodes Node 0-7, and Power Management Module 1 is used to manage and control the power supply of the eight computing nodes Node 8-16. If 16 or even more computing nodes are placed in the server whole cabinet, the control system designed by the present application can still effectively monitor and manage these computing nodes. In theory, as long as A computing nodes can be placed in the whole cabinet, the control system designed by the present application can effectively monitor and manage them.

[0057] It is not difficult to understand that in addition to being able to control the power supply module 1 to supply power to each server in the whole cabinet, since it establishes a communication connection with the server through the first communication bus, the control system can also be responsible for monitoring and managing the system state of the server, including temperature, fan speed, network connection and other key parameters. Directly multiplexing the control system can monitor the server state in real time, covering the hardware state, temperature, fan speed, network connection and other aspects of the server, realize the monitoring and management of the server, and further realize the centralized management of the server. Through centralized management and real-time monitoring, the administrator can fully understand the running status of all servers in the whole cabinet, quickly respond to potential problems, and improve the reliability and stability of the system. The control system can further support remote management function through setting wireless communication module and other ways, so that the operator can not be limited by geographical location, remotely access and manage the whole cabinet through network connection, and improve the flexibility and convenience of management.

[0058] Specifically, the BMC (Baseboard Management Controller) and CPLD (Complex Programmable Logic Device) in the server cooperate to realize real-time detection of various working states and state parameters of the server, and the control module 2 establishes a communication connection with the BMC and CPLD in the server through the first communication bus and receives the detection results related to the server, so as to realize the detection of the working temperature, the speed of the cooling fan and other server states in the server by multiplexing the control system. The control system can also realize monitoring of the power output by the power supply module 1 to each server, and the power monitoring system can timely discover power abnormalities such as overvoltage, undervoltage and overcurrent, and issue an alarm to avoid system downtime caused by power failure and ensure business continuity. The control module 2 can be provided with a prompt module to realize prompt and alarm of server failure and power failure and other conditions, so as to realize timely response to the failure conditions of the whole cabinet. In order to distinguish the failure conditions of each server, the prompt module can also be provided with a plurality of prompt sub-modules corresponding to each server one by one, and when a server or its input power exists an abnormality, the control module 2 can control the corresponding prompt sub-module to alarm, so as to enable the operator to timely locate the failure. The power monitoring and server monitoring and management of the server whole cabinet can not only improve the operation efficiency and reliability of the data center, but also provide a powerful guarantee for energy saving and emission reduction and business continuity. The power monitoring and server monitoring and management of the server whole cabinet play a crucial role in the operation of the data center. Through real-time monitoring and management of the power usage in the cabinet, power monitoring can optimize energy use, reduce energy consumption and improve energy efficiency. This not only helps to reduce operating costs, but also improves the environmental friendliness of the system.

[0059] It should be noted that on the basis of realizing the power supply function of the device, the control module 2 can also be further connected to the switch, so that the operator can obtain the state information of each server in the whole cabinet through the switch and the control module 2, and can also issue an instruction through the switch, which is forwarded to the CPLD and BMC of the corresponding server by the control module 2, so that the unified management of multiple servers corresponding to a control system can be realized by using one switch. The switch can be further integrated in the server whole cabinet. For a server whole cabinet, only one switch can be set, and the unified management of all servers in a server whole cabinet can be realized by using the communication connection between the control modules 2 of each control system and the communication connection between the control module 2 and the server. A single operation interface or management interface is realized for the whole cabinet, which is convenient to apply.

[0060] The server whole cabinet can integrate servers, storage devices, network devices, cooling systems, control systems including power supply modules 1 and control modules 2, etc. Through efficient management and monitoring of the control system, the stability and reliability of the whole cabinet are ensured. This integrated design simplifies the deployment and management of the data center, and the administrator can monitor and manage the running state of the whole cabinet through a single interface, improving the management efficiency. At the same time, the control system can monitor the temperature, humidity, power usage and other key parameters of the server and the power supply module 1 in real time, optimize energy use through intelligent algorithms, reduce energy consumption, and improve energy efficiency. In addition, the monitoring system can timely discover potential faults and issue alarms to avoid server downtime and ensure business continuity. The server whole cabinet can be more applied to edge computing and other scenarios to provide low-latency, high-reliability computing and storage services to meet changing business needs and technical challenges.

[0061] As a specific embodiment, the server whole cabinet designed based on the control system provided by the present application can realize the monitoring and management of the whole server whole cabinet. In the centralized power supply whole cabinet, each computing node is not powered by a separate PSU, but is powered by the power supply module in the cabinet. Taking the power supply module composed of a Power Management Module and a PSU as an example, as shown in Figure 4 , Figure 4 is a schematic diagram of the internal structure of a control system of a server when one control submodule is provided for an embodiment of the present application; a Power Management Module manages 4 computing nodes. The control module in the control system is composed of a PMC board card, which includes BMC, CPLD and related MUX, I2C SWITCH and GPIO Expander devices on the board card. As Figure 4As shown, considering that the power supply module needs to supply power to four servers, four PSUs are arranged in the power supply module for the convenience of application, and two PSUs are additionally arranged as a redundant design. When the four servers all work at the maximum demand power, the control module can control the four PSUs in the power supply module to supply power to the four servers respectively for the convenience of control, but when the demand power of the servers is small, for example, the demand power of NODE 0 is small, after the first PSU supplies power to NODE 0, there will be a margin, and the margin can be provided to NODE 1, if the demands of NODE 0 and NODE 1 are both small, one PSU can be used to supply power to both of them, if the demand power of NODE 0 or NODE 1 is large, the margin of the first PSU after supplying power to NODE 0 and the second PSU can be used to supply power to NODE 1 at the same time, so as to improve the power supply efficiency of the power supply module, and this power supply mode can also try to reduce the situation that all the PSUs work at the same time, greatly reducing the power consumption of the whole server cluster system and reducing resource waste.

[0062] Taking the first communication bus as an example of the I2C bus, the BMC on the PMC board realizes efficient monitoring and management of the computing nodes through a series of I2C connections and switching mechanisms. Specifically, the I2C0 of the BMC is connected to the first multiplexer MUX0, and then connected to the first switching module I2C SWITCH0 through I2C_MUX0. The downstream of the first switching module I2C SWITCH0 is designed with 4 groups of I2C_NODE 0-3 signals, respectively representing the server data of the four computing nodes. The BMC can connect to the BMC of the computing nodes 0-3 through this link to obtain the IP (Internet Protocol) of the BMC, so that the control module 2 determines which computing node the signal corresponds to and monitors the running state of the computing node, such as temperature, fan speed, and power supply. Taking the computing node Node0 as an example, the BMC in Node0 is used to manage the running state of the node itself. When running normally, the communication link is connected to the CPLD of Node0 through I2C, and then connected to the PMC board through I2C_NODE0. This design ensures that the PMC can monitor and manage the running parameters of Node0 in real time. If the CPLD of Node0 fails, the BMC of Node0 can directly connect to the PMC through I2C to ensure real-time monitoring of the computing node. This redundant design improves the reliability and stability of the system, ensuring that the monitoring and management functions of the system can be seamlessly switched in the case of single-point failure, ensuring the continuity of business and the normal operation of the system. Secondly, the BMC can monitor the temperature, liquid cooling state, power state and other key parameters of each computing node in the whole cabinet in real time through the I2C link and GPIO signal. This real-time monitoring function helps to discover and handle potential problems in a timely manner, preventing system failures and ensuring efficient operation of the whole cabinet. For company customers, real-time monitoring and early warning functions can discover and solve potential problems in advance, reduce unexpected downtime and maintenance costs, and improve the overall performance and reliability of the system.

[0063] The application provides a control system of a server, based on which centralized power supply control of multiple servers can be realized, and state monitoring and management of the multiple servers can be effectively realized, a complete power supply monitoring and management module is integrated by designing a control module and a communication link thereof. Unified power consumption management and server system state management of all servers in an entire cabinet can be effectively realized, the control module 2 can be responsible for monitoring and managing power supply use of the entire cabinet, real-time monitoring of power consumption of each server, optimization of energy use, reduction of energy consumption, and provision of detailed power consumption reports. Meanwhile, the control module 2 can detect power supply abnormalities such as overvoltage, undervoltage and overcurrent and timely issue an alarm to avoid system shutdown caused by power supply failure. The reliability and management efficiency of the entire server system are improved, and the control system has high flexibility and scalability and can support different sizes and types of entire cabinets. The control system can be effectively applied to small edge computing environments and large data centers, and meet the needs of different scenarios. Users can flexibly configure and expand the system according to their own business needs, avoid excessive investment and resource waste, and improve resource utilization and economic benefits.

[0064] In some embodiments, all servers in the server cluster are divided into N groups of servers, each group of servers including at least two servers; the control module 2 includes N control submodules connected one-to-one with the N groups of servers, and the power supply module 1 includes N power supply submodules connected one-to-one with the N groups of servers; N is a positive integer greater than 1.

[0065] The power supply submodule is used to supply power for the corresponding group of servers under the control of the corresponding control submodule.

[0066] In order to improve the reliability and stability of the system, master-standby redundancy design of the control module 2 can also be performed, multiple control submodules are arranged in the control module 2 to be redundant to each other, and automatic switching to another control submodule can be flexibly performed when any one control submodule has a problem. In order to improve the device utilization rate and prolong the service life of the device, when initially setting, multiple servers in the same server cluster or in the same entire cabinet are divided into N groups, and each control submodule manages and controls the corresponding group of servers, so that the task amount of a single control submodule is avoided to be too large, and the utilization rate of all control submodules is improved. At this time, for any group of servers, the control submodule is the control module 2 of the group of servers, that is, one control system is divided into several control subsystems, and multiple servers are divided into the control subsystems for control and management, and one control subsystem includes one control submodule and a power supply submodule. The specific grouping mode of the servers and the implementation mode of the corresponding control submodule and power supply submodule are not particularly limited in the application.

[0067] It can be understood that the control module 2 includes two control sub-modules, that is, the control system designs two control modules 2, and each control module 2 is responsible for managing a part of the computing nodes. Once one of the control modules 2 fails, the other control module 2 can automatically take over the management authority of the failed control module 2, ensuring the continuous operation and high availability of the entire server system. This redundant design eliminates the risk of single point failure and avoids the whole cabinet downtime caused by the failure of a single module, thereby ensuring the continuity and stability of the business. Through the redundant design, the reliability and stability of the entire control system are significantly improved, higher system availability and lower business interruption risk are achieved, and the user experience is improved.

[0068] As a specific embodiment, refer to Figure 5 As shown in the figure, Figure 5 is an internal structure diagram of a control system of a server when two control sub-modules are provided; the control system is designed with two control sub-systems (PMC 0 and PMC 1) and corresponding two control sub-modules (Power Management Module 0 and Power Management Module 1) in a whole cabinet for monitoring and management of the server. In normal operation, Power Management Module 0 manages Node0-3, and Power Management Module 1 manages Node4-7. Once one of the control sub-modules fails, for example, the BMC of Power Management Module 0 fails and is dead, cannot continue to operate and communicate, the control system will automatically switch to the other control sub-module Power Management Module 1 to take over the work of Power Management Module 0. As Figure 5As shown, taking BMC failure of PMC 1 as an example, the BMC and the CPLD are connected through the WDT timing module to ensure the response of the BMC. Once the BMC does not return a response to the CPLD for a period of time, it can be considered that the BMC has failed and hung up, the BMC of the PMC fails, and the control system automatically switches to PMC 0. At this time, the BMC of the PMC 0 board is connected to the I2C SWITCH1 of the PMC 1 board through I2C3, and the I2C SWITCH1 is connected to the CPLD, MUX0 and MUX1 of the PMC 1 board respectively. The BMC of the PMC 0 board can be connected to the I2C SWITCH0 of the PMC 1 board through the I2C SWITCH1 and MUX0 of the PMC 1 board, so as to access the Node 4-7. In addition, the I2C SWITCH1 of the PMC 1 board can also enable the BMC of the PMC 0 board to access the MUX1 of the PMC 1 board through switching, so as to obtain the real-time data of the PSU 1. Therefore, in the design of the present scheme, even if the BMC of the PMC 1 hangs up, through a series of I2C link switching designs on the hardware, the PMC 0 board can monitor and manage all the computing nodes and PSUs hung under the PMC 1. There is also a I2C_CPLD link connected to the CPLD at the downstream of the I2C SWITCH1 of the PMC 1. Therefore, the BMC of the PMC 0 board can also obtain the alarm information of the temperature sensor and the liquid leakage sensor originally managed by the PMC 1 board, so as to ensure that all important information in the entire cabinet can be seamlessly switched to the PMC 0 system, and the operation of the entire cabinet is stable. The PMC is designed as a separate board card connected with other modules through connectors. When the BMC of the PMC board fails, maintenance personnel can replace the PMC board as needed, rather than the entire control system, which greatly saves maintenance cost and time. The modular design not only improves the maintainability of the system, but also ensures that the system function can be quickly restored after failure.

[0069] It should be noted that the specific redundancy design manner between the control subsystems in the control system, i.e., between the control submodules, is not particularly limited in the present application. When the number of control submodules is even, the control submodules can be set in a two-by-two redundancy manner, i.e., every two control submodules are a group, and the two control submodules in a group are redundant to each other, and when one control submodule in a group fails, the other control submodule takes over. The latter control submodule can also be in sequence as the redundancy of the former control submodule, i.e., the second control submodule is the redundancy of the first control submodule, the Nth control submodule is the redundancy of the (N-1)th control submodule, and finally the first control submodule is the redundancy of the Nth control submodule. When the ith control submodule fails, the (i+1)th control submodule takes over.

[0070] Specifically, through the control system, real-time monitoring and management of power supply and server system status of the whole cabinet can be realized, so that the administrator can directly use the control system module to comprehensively master the running status of all servers in the whole cabinet, including key parameters such as power consumption, temperature condition of centralized power supply, fan speed of the server, network connection, etc. In addition, the whole control system also adopts a redundant design, which can automatically switch to another control subsystem flexibly when any control subsystem has a problem. This redundant mechanism greatly enhances the fault tolerance and high availability of the control system. Even if a control sub-module fails, it will not affect the management and monitoring functions of the whole cabinet, ensuring the continuity of business and the stability of the system. The whole cabinet server and power supply monitoring and management system designed by the application has high flexibility and scalability, can support the state monitoring and management of N servers in the whole cabinet, and can be flexibly applied to different types and sizes of whole cabinets according to needs, with high reliability, high efficiency management capability and easy maintenance.

[0071] In some embodiments, the control sub-module comprises:

[0072] The first switching module comprises a fixed end and a plurality of switching ends connected one-to-one with a plurality of servers in a corresponding group of servers, and the switching end of the first switching module is connected with the data end of the corresponding server.

[0073] The first multiplexer MUX0 is connected with the fixed end of the first switching module.

[0074] The processor is connected with the second end of the first multiplexer MUX0 through the first end, and is used to control the first switching module to switch through the first multiplexer MUX0, so as to obtain the power supply demand of each server in the corresponding group of servers.

[0075] It can be understood that the use of pins of the processor can be reduced by the cooperation of the switching module and the multiplexer in the control sub-module, and the pin resources can be saved. The first switching module can transmit the data of the connected multiple servers to the processor through switching, and only one pin of the processor is needed to realize the acquisition of the data of the multiple servers through the cooperation of the multiplexer, so as to realize the monitoring and management of the multiple servers. The multiplexer is arranged to facilitate the redundant connection between the multiple control sub-modules. The specific types and implementation manners of the first switching module, the first multiplexer MUX0 and the processor are not particularly limited in the application. Figure 4 As shown in the figure, the first switching module is implemented by I2C Switch 0 (I2C link switching switch).

[0076] Specifically, by configuring the switching module and the multiplexer for the control submodule, the control submodule can conveniently monitor and manage multiple servers, save resources, improve the utilization rate of the communication bus, avoid address conflicts, and improve the efficiency of power management.

[0077] In some embodiments, the control submodule further comprises:

[0078] a second multiplexer MUX1, a first end of which is connected to a second end of the processor;

[0079] a pin expander, the pin expander comprising a bus end and M pin ends, the pin ends of the pin expander being connected to the signal ends of the corresponding power supply submodule, and the bus end being connected to a second end of the second multiplexer MUX1;

[0080] The processor is further configured to control the power supply submodule to start through the pin expander, and to obtain the working state of the power supply submodule through the pin expander; the working state of the power supply submodule comprises an insertion state of the power supply submodule, a power supply state of the power supply submodule, and a fault state of the power supply submodule.

[0081] Considering that the control submodule needs to consider the working state of the power supply module 1 when controlling the power supply module 1, the working state of the power supply module 1 is usually represented by a pin signal, and therefore a pin expander can be further provided in the control submodule to expand multiple pins to transmit the pin signals that need to be communicated between the control submodule and the power supply submodule, so as to ensure accurate control between the control submodule and the power supply submodule. Meanwhile, the multiplexer is provided to facilitate redundant connection between multiple control submodules. The specific types and implementation manners of the second multiplexer MUX1 and the pin expander are not particularly limited herein.

[0082] As a specific embodiment, as shown in FIG. 1, the control submodule comprises a processor 1, a switching module 2, a first multiplexer MUX2, a second multiplexer MUX1, and a pin expander 3. Figure 4 and Figure 5As shown, when the power supply sub-module is implemented by multiple PSUs, the management signals of the PSUs include ALERT_N, AC_OK_N, PS_ON_N and PRSNT_N, which are connected to the GPIO Expander chip through four pins respectively, and then transmitted to the BMC on the PMC board through the second multiplexer MUX1. The GPIO Expander chip can be flexibly selected according to the number of PSUs to adapt to different application scenarios. The I2C2 of the BMC is connected to the GPIO Expander through the second multiplexer MUX1, so that the BMC as the core of the management system can monitor and obtain the state of the PSU in real time, such as power-on / off, in-place state or alarm signal. Among them, ALERT_N is used to report abnormal conditions or failures in the PSU; AC_OK_N is used to indicate the presence or absence of the input voltage of the PSU; PS_ON_N is used to control the start of the PSU; PRSNT_N is used to indicate whether the PSU has been correctly installed or inserted into the system. The control submodule can determine whether the PSU is installed only after receiving the PRSNT_N signal, and issue the PS_ON_N signal to control the PSU to start after the installation is completed, and determine whether the PSU can normally output voltage after receiving the AC_OK_N signal, and control the PSU to output the corresponding power to each server in the case of normal voltage output of the PSU; At the same time, the control submodule obtains ALERT_N to judge whether the PSU has failed, and controls it to stop when the PSU fails, controls other PSUs to replace its power supply or alarms to inform the operator to timely eliminate the fault.

[0083] Specifically, by configuring the control submodule with a pin expander and a multiplexer, the control submodule can conveniently monitor and manage multiple PSUs, save resources, improve the efficiency of power management, and ensure that the power supply submodule can provide accurate and reliable power supply for the server under normal working conditions.

[0084] In some embodiments, the processor comprises:

[0085] The management controller is connected with the second end of the first multiplexer MUX0 at the first end, connected with the first end of the second multiplexer MUX1 at the second end, and connected with the control end of the corresponding power supply submodule through the second communication bus at the third end, for obtaining the power supply demand of each server in the corresponding group of servers through the first multiplexer MUX0, and controlling the work of the corresponding power supply submodule according to the power supply demand;

[0086] The programmable logic device is connected with the detection end of the management controller at the first end, for detecting the fault state of the management controller;

[0087] The control submodule further comprises:

[0088] a second switch module, a first end of which is connected to a third end of the first multiplexer MUX0, a second end of which is connected to a third end of the second multiplexer MUX1, a third end of which is connected to the second end of the programmable logic device, and a fourth end of which is connected to a fourth end of the management controller of another control submodule of the control module 2 except for itself;

[0089] The programmable logic device is further configured to send a takeover signal to the management controller of the other control submodule connected to the fourth end of the second switch module when detecting that the management controller fails, so that the management controller of the other control submodule takes over the work of the failed management controller through the second switch module.

[0090] It is not difficult to understand that, in order to realize the redundancy between the plurality of control submodules, the second switch module is also needed to be arranged in the control submodule to realize the communication link between the two control submodules, so as to help the two control submodules to realize the data takeover in case of failure, etc. Meanwhile, in order to facilitate the failure detection, the processor includes the management controller and the programmable logic device, wherein the management controller undertakes the monitoring and control of the server and the power supply submodule by the control submodule, and the programmable logic device is particularly used for detecting the failure of the management controller and notifying another management controller to take over the work in case of failure of the management controller. The specific types and implementation manners of the management controller, the programmable logic device and the second switch module are not particularly limited in the present application. The management controller can be implemented by a BMC, and the programmable logic device can be implemented by a CPLD, as shown in FIG. 2. Figure 4 The second switch module is implemented by an I2C Switch 1 (I2C link switch).

[0091] Specifically, the processor specifically includes the management controller and the programmable logic device, the programmable logic device is arranged to detect the failure of the management controller, and the management controller is used as the control core to realize the entire monitoring and management function of the processor, so as to ensure the reliable work of the control module 2 in the control system and improve the reliability and safety of the entire control system. Meanwhile, the second switch module is arranged to realize the takeover between the control submodules, the second switch module is used to realize the effective acquisition of all data of the failed control submodule by the takeover control submodule in case of failure, and the key communication link connection is realized. This design ensures that the control system can seamlessly switch the management authority in case of failure, and continues to monitor and manage the computing node and the power supply module 1.

[0092] In some embodiments, the specific process of detecting the failure state of the management controller by the programmable logic device includes:

[0093] sending a heartbeat signal to the management controller and starting a timer;

[0094] If the response signal returned by the management controller has not been received when the count value of the timer reaches the preset value, it is determined that the management controller has failed.

[0095] It can be understood that a WDT (Watchdog Timer) timing module can be specifically used for fault detection. The heartbeat detection is performed between the BMC and the CPLD through the WDT timing module. If no feedback response signal is received within a predetermined time, the WDT timing module will time out, triggering the fault detection mechanism. This mechanism can timely discover the failure of the BMC and start the fault switching process, ensuring the stable operation of the control system. The specific types and implementation modes of the heartbeat signal and the counter are not particularly limited in the present application. The specific values of the preset values can be set according to the normal response time of the heartbeat signal, which are not particularly limited in the present application.

[0096] Specifically, through the WDT timing module and the I2C link switching design, the control system can automatically switch to the standby PMC when the BMC failure is detected, ensuring seamless connection of monitoring and management functions, and guaranteeing the operation stability and business continuity of the entire cabinet. This automatic switching mechanism reduces the need for manual intervention, improves the automation management level of the control system, and reduces the operation and maintenance cost and complexity. The entire cabinet server and power monitoring and management system designed based on the present application significantly improves the reliability and management efficiency of the entire server system through multiple redundancies and intelligent management, bringing users higher system availability, lower business interruption risk, more efficient resource utilization, and lower operation and maintenance cost.

[0097] In some embodiments, further comprising:

[0098] a temperature sensor arranged in the corresponding group of servers, and having an output end connected to a first input end of the processor, for detecting the working temperature of the corresponding group of servers, so that the processor determines whether the working temperature of the server is abnormal;

[0099] and / or,

[0100] a liquid leakage sensor arranged in the corresponding group of servers, and having an output end connected to a second input end of the processor, for detecting the liquid leakage condition of the corresponding group of servers, so that the processor stops power supply to the server with the liquid leakage condition when the liquid leakage condition exists in the server.

[0101] It is also important to consider the temperature management of the whole server cabinet and the management of the liquid cooling system. High temperature can cause the machine to run slowly and even cause downtime. Abnormal liquid leakage can cause equipment damage and other safety problems. Therefore, multiple temperature sensors can be designed in the control system to monitor the temperature at different positions in the whole cabinet. As shown in Figure 4 The temperature sensors are directly connected to the CPLD on the PMC board through the GPIO signal, and the data is stored in the CPLD register. The BMC reads the register information in the CPLD through I2C1 to obtain the temperature at different positions in the whole cabinet in real time. The management of the liquid cooling system in the liquid cooling whole cabinet or liquid cooling server is also supported. Multiple liquid cooling leakage detection sensors can be designed in the control system and placed at key positions of the liquid cooling system in the whole cabinet. This allows the PMC to monitor the temperature and liquid leakage failure of the liquid cooling system in real time, ensuring the normal operation of the liquid cooling system and avoiding server damage caused by liquid cooling failure. Once a liquid leakage problem occurs in a computing node, the leakage detection sensor transmits the status to the CPLD through the GPIO signal, and the BMC can immediately obtain the liquid leakage alarm. According to the liquid leakage situation at different positions, the BMC can appropriately select to turn off the power supply of the corresponding computing node to prevent short circuit and burning caused by liquid leakage. The specific setting method of the type, number and position of the temperature sensor and the leakage sensor is not particularly limited in this application and can be adjusted and set according to the actual setting of the server.

[0102] Specifically, by setting the temperature sensor and / or the liquid leakage sensor, the control system can timely detect the temperature and / or liquid leakage in the liquid cooling system and take corresponding measures to avoid server failure caused by high temperature or liquid leakage, further realizing comprehensive monitoring of the server and further improving the safety and reliability of the server. The comprehensive monitoring and management capability not only improves the operation efficiency of the data center, but also provides a strong guarantee for the stable operation of the server. The whole cabinet monitoring and management architecture designed by the present application can effectively solve many problems in the management of the traditional server whole cabinet, realize efficient and reliable power consumption management and server system state monitoring, and ensure the stable operation and business continuity of the data center.

[0103] To solve the above technical problems, the embodiment of the present application also provides a server cluster system, which comprises a server cluster and a server control system as described above. The output end of the server control system is connected to the power supply end of each server in the server cluster, and the input end is connected to the data end of each server.

[0104] The features of the server cluster system provided by the embodiment of the present application can be referred to the related description of the embodiment of the server control system, which will not be repeated here.

[0105] Referring to Figure 6 as shown, Figure 6 A flowchart of a server control method is provided in an embodiment of the present application. To solve the above technical problem, an embodiment of the present application further provides a server control method applied to a server control system as described above. The server control method comprises the following steps.

[0106] S11: obtaining power supply requirements of each server in a server cluster;

[0107] S12: determining a power supply strategy of a power supply module according to the power supply requirements;

[0108] S13: controlling the power supply module to work based on the power supply strategy; the power supply strategy of the power supply module comprises output power of each output terminal of the power supply module.

[0109] The features of the server control method provided in the embodiment of the present application can be referred to the related description of the embodiment of the server control system, which will not be repeated here.

[0110] It can be understood that if the server control method in the above embodiment is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and performs all or part of the steps of the method of each embodiment of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), electrically erasable programmable ROM, register, hard disk, removable magnetic disk, CD-ROM, magnetic disk or optical disk and various program code storage media.

[0111] Based on this, an embodiment of the present application further provides a computer readable storage medium, and the computer readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the server control method as described above are implemented.

[0112] The features of the computer readable storage medium provided in the embodiment of the present application can be referred to the related description of the embodiment of the server control method, which will not be repeated here.

[0113] An embodiment of the present application further provides a computer program product, which comprises computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the server control method described in the above embodiment are implemented.

[0114] The description of the features of the computer program product provided by the embodiments of the present application can refer to the description of the embodiments of the control method of the server, which will not be repeated here.

[0115] The above describes in detail the server control system, method, server cluster system and medium provided by the embodiments of the present application. Each embodiment in the specification is described in a progressive manner, and each embodiment mainly describes the difference from other embodiments. The same or similar parts of each embodiment can be referred to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the related parts can be referred to the method part.

[0116] The skilled person can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware, computer software or a combination of both. In order to clearly show the interchangeability of hardware and software, the components and steps of each example have been described in the above description. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0117] The above describes in detail the server control system, method, server cluster system and medium provided by the embodiments of the present application. The principle and implementation of the present application is described by specific examples. The above description of the embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary skilled person in the art, without departing from the principle of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the present application.

Claims

1. A control system of a server, characterized by, The application is applied to a server cluster; a control system of the server comprises: a power supply module, the power supply module comprises a plurality of output terminals corresponding to a plurality of servers in the server cluster, and the output terminal of the power supply module is connected with the power supply terminal of the corresponding server; a control module, a first end of the control module is connected with the data terminal of a plurality of servers in the server cluster through a first communication bus, a second end of the control module is connected with the control terminal of the power supply module through a second communication bus, the control module is used for acquiring the power supply demand of each server through the first communication bus, determining the power supply strategy of the power supply module according to the power supply demand, and controlling the power supply module to work based on the power supply strategy through the second communication bus; the power supply strategy of the power supply module comprises the output power of each output terminal of the power supply module; all servers in the server cluster are divided into N groups of servers, and each group of servers comprises at least two servers; the control module comprises N control sub-modules corresponding to the N groups of servers, and N is a positive integer greater than 1; the control sub-module comprises: a first switching module, the first switching module comprises a fixed end and a plurality of switching terminals corresponding to a plurality of servers in a corresponding group of servers, and the switching terminal of the first switching module is connected with the data terminal of the corresponding server; a first multiplexer, a first end of the first multiplexer is connected with the fixed end of the first switching module; a processor, a first end of the processor is connected with a second end of the first multiplexer, and the processor is used for controlling the first switching module to switch through the first multiplexer to acquire the power supply demand of each server in the corresponding group of servers.

2. The control system of a server according to claim 1, wherein, the power supply module comprises N power supply sub-modules corresponding to the N groups of servers; the power supply sub-module is used for supplying power for the corresponding group of servers under the control of the corresponding control sub-module.

3. The control system of a server according to claim 2, wherein, the control sub-module further comprises: a second multiplexer, a first end of the second multiplexer is connected with a second end of the processor; a pin expander, the pin expander comprises a bus end and M pin ends, the pin end of the pin expander is connected with the signal end of the corresponding power supply sub-module, and the bus end is connected with a second end of the second multiplexer; the processor is further used for controlling the power supply sub-module to start through the pin expander, and acquiring the working state of the power supply sub-module through the pin expander; the working state of the power supply sub-module comprises the insertion state of the power supply sub-module, the power supply state of the power supply sub-module and the fault state of the power supply sub-module.

4. The control system of the server according to claim 3, wherein the processor comprises: a management controller, a first end of the management controller is connected with a second end of the first multiplexer, a second end of the management controller is connected with a first end of the second multiplexer, and a third end of the management controller is connected with the control terminal of the corresponding power supply sub-module through a second communication bus, the management controller is used for acquiring the power supply demand of each server in the corresponding group of servers through the first multiplexer, and controlling the working of the corresponding power supply sub-module according to the power supply demand; A programmable logic device, a first end of which is connected to a detection end of the management controller, for detecting a fault state of the management controller; The control sub-module further comprises: A second switching module, a first end of which is connected to a third end of the first multiplexer, a second end of which is connected to a third end of the second multiplexer, a third end of which is connected to a second end of the programmable logic device, and a fourth end of which is connected to a fourth end of the management controller of another control sub-module in the control module except itself; The programmable logic device is further configured to send a takeover signal to the management controller of another control sub-module connected to the fourth end of the second switching module when detecting that the management controller has failed, so that the management controller of another control sub-module takes over the work of the failed management controller through the second switching module.

5. The control system of the server according to claim 4, wherein, The specific process of detecting the fault state of the management controller by the programmable logic device comprises: sending a heartbeat signal to the management controller and starting a timer; if a response signal returned by the management controller has not been received when the count value of the timer reaches a preset value, determining that the management controller has failed.

6. The control system of a server according to any one of claims 1 to 5, wherein, Further comprising: a temperature sensor arranged in a corresponding group of servers, an output end of which is connected to a first input end of the processor, for detecting the working temperature of the corresponding group of servers, so that the processor determines whether the working temperature of the server is abnormal; and / or, a liquid leakage sensor arranged in a corresponding group of servers, an output end of which is connected to a second input end of the processor, for detecting the liquid leakage condition of the corresponding group of servers, so that the processor stops power supply for the server with the liquid leakage condition when the liquid leakage condition exists in the server.

7. A server cluster system, characterized by A control system of a server as claimed in any one of claims 1 to 6, an output end of the control system of the server being connected to a power supply end of each server in the server cluster, and an input end of the control system of the server being connected to a data end of each server.

8. A control method of a server, characterized by, The server control method comprises: obtaining power supply requirements of each server in the server cluster; determining a power supply strategy of the power supply module according to the power supply requirements; controlling the power supply module to work based on the power supply strategy; and the power supply strategy of the power supply module comprises output power of each output end of the power supply module.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to realize the steps of the server control method of claim 8. The computer readable storage medium stores a computer program, and the computer program is executed by the processor to realize the steps of the server control method of claim 8.

Citation Information

Patent Citations

  • Operation control method and device for server power supply

    CN118778790A

  • Power detecting circuit board, power detecting system, and immersed liquid cooling tank

    US20240223099A1