Intelligent Server System with Multiple Heterogeneous Nodes
Through the intelligent server system of multiple heterogeneous nodes, the node manager is used to centrally manage the status information of the GPU BOX server and general server, which solves the problem of high material cost of communication cables in traditional systems and achieves more efficient connection and management.
Patent Information
- Application Number
- CN202211110775.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-13
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-09-13
AI Technical Summary
The material cost of communication cables in traditional intelligent server systems is relatively high.
An intelligent server system with multiple heterogeneous nodes is used to interact with the status information of the GPU BOX server and the general server, and a node manager is used to realize centralized management of status information and fault analysis, reducing the use of communication cables.
It reduces the material cost of communication cables, improves the system's connection reliability and management convenience, reduces the difficulty of wire management, and facilitates monitoring and troubleshooting.
Smart Images

Figure CN115454633B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of server design, and particularly to an intelligent server system with multiple heterogeneous nodes. Background Art
[0002] With the continuous rise and development of cloud computing technology and related derivative technologies and products, the business volume of the Internet industry has gradually shown an explosive growth, thus promoting the development of intelligent server systems.
[0003] The intelligent server system abandons the traditional server solution that uses a general-purpose server with a CPU architecture as the computing power core, but includes a heterogeneous server system with a GPU BOX server using a CPU architecture and a general-purpose server using a CPU architecture, and uses a GPU Module that supports parallel computing to provide sufficient computing power for the system. However, the intelligent server system in the traditional technology has a relatively high material cost of communication cables. Summary of the Invention
[0004] Based on this, it is necessary to provide an intelligent server system with multiple heterogeneous nodes that can reduce the material cost of communication cables in the intelligent server system for the above technical problems.
[0005] In one embodiment, a cabinet server system with multiple heterogeneous nodes is provided. The intelligent server system includes:
[0006] A first preset number of GPU BOX servers for outputting first status information; the first status information is used to characterize the working status of the GPU BOX server; the GPU BOX server adopts a GPU architecture;
[0007] A first preset number of general-purpose servers, communicatively connected to the corresponding GPU BOX servers, for receiving the first status information output by the corresponding GPU BOX servers and outputting the corresponding first status information and second status information; the second status information is used to characterize the working status of the general-purpose server; the general-purpose server adopts a CPU architecture;
[0008] A node manager, communicatively connected to each general-purpose server, for communicatively connecting to the management node of the cabinet management server; the node manager is further configured to receive the corresponding first status information and second status information output by each general-purpose server and output each first status information and second status information to the management node, so that the cabinet management server records each first status information and second status information.
[0009] In one embodiment, the node manager includes: a network switching module, the network switching module includes a second preset number of network interfaces; the network switching module is communicatively connected to corresponding general servers through each network interface; the network switching module is configured to receive corresponding first status information and second status information output by each general server, and output each first status information and second status information;
[0010] The second preset number is greater than the first preset number; a network interface control module, communicatively connected to the network switching module, is configured to communicatively connect to the management node of the rack-mounted management server, and is further configured to receive each first status information and second status information output by the network switching module, and output each first status information and second status information to the management node, so that the rack-mounted management server records each first status information and second status information.
[0011] In one embodiment, the node manager further includes: a baseboard management control module, communicatively connected to the network switching module, is configured to receive each first status information and second status information output by the network switching module; the baseboard management control module is further configured to perform fault analysis according to each first status information and second status information, and output a first target instruction; the first target instruction is used to instruct the corresponding GPU BOX server or general server to avoid first-priority faults; wherein, the first-priority faults include abnormal power-off faults and liquid leakage faults.
[0012] In one embodiment, the node manager further includes a field programmable gate array module; the field programmable gate array module is communicatively connected to the baseboard management control module; the baseboard management control module is further configured to output each first status information and second status information to the field programmable gate array module; the field programmable gate array module is configured to perform fault analysis according to each first status information and second status information, and output a second target instruction; the second target instruction is used to instruct the corresponding GPU BOX server or general server to avoid second-priority faults; wherein, the second-priority faults include downclocking faults and overclocking faults; the priority of the second-priority faults is lower than the priority of the first-priority faults.
[0013] In one embodiment, the field programmable gate array module is communicatively connected to the network interface control module; the field programmable gate array module is further configured to perform traffic control analysis according to each first status information and second status information, and output a third target instruction; wherein, the third target instruction is used to instruct the network interface control module to output a fourth target instruction to the network switching module; the fourth target instruction is used to instruct the network switching module to control the bandwidth of each general server.
[0014] In one embodiment, the GPU BOX server includes a first complex programmable logic module; the general server includes a second complex programmable logic module; wherein, the second complex programmable logic module is communicatively connected to the corresponding first complex programmable logic module; the network switching module is communicatively connected to the second complex programmable logic module; the first complex programmable logic module is configured to output first status sub-information to the corresponding second complex programmable logic module when the corresponding GPU BOX server is in the standby mode; wherein, the first status information includes the first status sub-information; the first status sub-information is used to represent that the GPU BOX server is in the standby mode; the second complex programmable logic module is configured to receive the first status sub-information, and is further configured to output the first status sub-information and second status sub-information to the network switching module when the corresponding general server is in the standby mode; wherein, the second status information includes the second status sub-information; the second status sub-information is used to represent that the corresponding general server is in the standby mode; the network switching module is further configured to receive the first status sub-information and the second status sub-information, and output the first status sub-information and the second status sub-information to the baseboard management control module; the baseboard management control module is further configured to receive the first status sub-information and the second status sub-information; the baseboard management control module is further configured to output a fifth target instruction when receiving the first status sub-information and the second status sub-information; wherein, the fifth target instruction is used to instruct the corresponding second complex programmable logic module to power on, and output a sixth target instruction to the corresponding first complex programmable logic module when the corresponding second complex programmable logic module completes power-on; the sixth target instruction is used to instruct the corresponding first complex programmable logic module to power on.
[0015] In one embodiment, the first complex programmable logic module is further configured to output a target feedback instruction to the corresponding second complex programmable logic module when completing power-on; the second complex programmable logic module is further configured to receive the target feedback instruction; the second complex programmable logic module is further configured to output a global reset instruction when receiving the target feedback instruction; wherein, the global reset instruction is used to instruct the corresponding GPU BOX server and general server to perform a global reset.
[0016] In one embodiment, the intelligent server system further includes: a cabinet management server, the cabinet management server includes a management node, and the management node is communicatively connected to a node manager; the cabinet management server is configured to receive and record each first status information and second status information.
[0017] In one embodiment, the intelligent server system further includes a switch; wherein, the management node is communicatively connected to the node manager through the switch.
[0018] In one embodiment, the node manager is an intelligent network card.
[0019] Based on this, the above intelligent management server system outputs, through a first preset number of GPU BOX servers, the first status information used to characterize the working status of the GPU BOX servers; where the GPU BOX servers adopt a GPU architecture; then, through a first preset number of general servers, the first status information output by the corresponding GPU BOX servers is received, and the corresponding first status information and the second status information used to characterize the working status of the general servers are output; where the general servers adopt a CPU architecture; then, the node manager is used to communicatively connect to the management node of the cabinet-style management server, receive the corresponding first status information and second status information output by each general server, and output each first status information and second status information to the management node of the cabinet-style management server, so that the cabinet-style management server records each first status information and second status information, thereby realizing the cascading of the cabinet-style management server, the node manager, the corresponding general servers, and the corresponding GPU BOX servers, reducing the communication cables required for communication connection, thus reducing the material cost of the communication cables of the intelligent server system, improving the overall connection reliability of the intelligent server system, and reducing the cable management difficulty of the intelligent server system. In addition, it is convenient for the cabinet-style management server to perform monitoring operations or issue instructions on the corresponding general servers and GPU BOX servers according to each second status information and the first status information corresponding to the second status information. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 FIG. 6 is a first internal structure diagram of an intelligent server system with multiple heterogeneous nodes in an embodiment;
[0021] Figure 2 FIG. 10 is a schematic connection diagram of a GPU BOX server and a general server 200 in a specific example;
[0022] Figure 3 FIG. 14 is an internal structure diagram of a GPU BOX server and a general server 200 in a specific example;
[0023] Figure 4 FIG. 18 is a second internal structure diagram of an intelligent server system with multiple heterogeneous nodes in an embodiment;
[0024] Figure 5 FIG. 22 is an internal structure diagram of a node manager in an embodiment;
[0025] Figure 6 FIG. 26 is a third internal structure diagram of an intelligent server system with multiple heterogeneous nodes in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] To facilitate the understanding of the present application, the present application will be described more comprehensively below with reference to the relevant drawings. Embodiments of the present application are shown in the drawings. However, the present application can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present application more thorough and comprehensive.
[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used in the specification of this application herein are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0028] It can be understood that the terms "first", "second", etc. used in this application may be used herein to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, without departing from the scope of this application, a first resistor may be referred to as a second resistor, and similarly, a second resistor may be referred to as a first resistor. Both the first resistor and the second resistor are resistors, but they are not the same resistor.
[0029] It can be understood that in the following embodiments, "connection", if there is a transfer of electrical signals or data between the connected circuits, modules, units, etc., should be understood as "electrical connection", "communication connection", etc.
[0030] As used herein, the singular forms "a", "an" and "the" may also include the plural forms unless the context clearly dictates otherwise. It should also be understood that the terms "comprise / include" or "have" etc. specify the presence of the stated features, wholes, steps, operations, components, parts, or combinations thereof, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, components, parts, or combinations thereof.
[0031] In one embodiment, as Figure 1 shown, there is provided an intelligent server system with multiple heterogeneous nodes. The intelligent server system with multiple heterogeneous nodes includes a first preset number of GPU BOX servers 100, a first preset number of general servers 200, and a node manager 300. Among them, the general server 200 is communicatively connected to the corresponding GPU BOX server 100, and the node manager 300 is communicatively connected to the general server 200.
[0032] It can be understood that in an intelligent server system with multiple heterogeneous nodes, the number of GPU BOX servers 100 and general servers 200 is the same and there is a one-to-one correspondence. When the number of GPU BOX servers 100 is the first preset number, the number of general servers 200 is also the first preset number.
[0033] In a specific example, the first preset number can be determined according to the model of the node manager 300 selected by the user. Specifically, the number of intelligent nodes included in different models of the node manager 300 is different, and the first preset number is determined according to the number of intelligent nodes of the node manager 300. The node manager 300 is communicatively connected to the corresponding general server 200 through each intelligent node, and the general server 200 is communicatively connected to the corresponding GPU BOX server 100 at the same time. The node manager 300 is used to communicatively connect to the management node of the rack-mounted management server, thereby realizing the cascading of the rack-mounted management server, the node manager 300, the corresponding general server 200, and the corresponding GPU BOX server 100, and reducing the material cost of the communication cables of the intelligent server system. In addition, the intelligent node can be but is not limited to an AI node. The above is only a specific example, and it is flexibly set according to user needs in actual applications and is not limited here.
[0034] The GPU BOX server 100 is a server device adopting a GPU architecture. The GPU BOX server 100 can output first status information. Among them, the first status information is used to characterize the working status of the GPU BOX server 100.
[0035] In a specific example, the first status information includes the working voltage, working power consumption, working environment temperature, and working frequency of the GPU BOX server 100, and also includes the working status of the PCIe devices in the GPU BOX server 100. Among them, the working status of the PCIe devices includes the fan speed, hard disk working power consumption, and network card working power consumption, etc. The above is only a specific example, and it is flexibly set according to user needs in actual applications and is not limited here.
[0036] The general server 200 is a server device adopting a CPU architecture. The general server 200 can receive the first status information output by the corresponding GPU BOX server 100 and output the second status information and the first status information output by the corresponding GPU BOX server 100. Among them, the second status information is used to characterize the working status of the general server.
[0037] In a specific example, the general server 200 can package the second status information and the first status information output by the corresponding GPU BOX server 100, and output a status information packet to the node manager 300. The above is only a specific example, and in actual applications, it can be flexibly set according to user requirements and will not be limited here.
[0038] In a specific example, the second status information includes the operating voltage, operating power consumption, operating ambient temperature, and operating frequency of the general server 200, and also includes the operating status of the PCIe devices in the general server 200. Among them, the operating status of the PCIe devices includes fan speed, hard disk operating power consumption, network card operating power consumption, etc. The above is only a specific example, and in actual applications, it can be flexibly set according to user requirements and will not be limited here.
[0039] In a specific example, such as Figure 2 and Figure 3 as shown, the GPU BOX server includes an RJ45 network port and a BMC chip, and the general server includes two RJ45 network ports, a BMC chip, and a Lanswitch chip. Among them, the RJ45 network port of the GPU BOX server is connected to one of the RJ45 network ports of the general server through a communication cable, and the other RJ45 network port of the general server is connected to the node manager through a communication cable. The Lanswitch chip in the general server realizes network port cascading and network packet isolation, and transmits management information through the BMC chip. In addition, PCIe signals and sideband signals for interactive management are transmitted between the GPU BOX server and the general server through a general cable. The above sideband signals include the presence signal and the global reset signal output by the general server, etc. The Lanswitch chip in the general server can be but is not limited to the Marvell 88E6321 chip. The above is only a specific example, and in actual applications, it can be flexibly set according to user requirements and will not be limited here.
[0040] The node manager 300 is a device that can be used to communicate with the management nodes of cabinet-style management servers. Among them, the node manager 300 can receive the second status information output by each general server 200 and the first status information corresponding to the second status information, and output each second status information and the first status information corresponding to the second status information to the management nodes of the cabinet-style management server, so that the cabinet-style management server can receive and record each second status information and the first status information corresponding to the second status information, which is convenient for the cabinet-style management server to perform monitoring operations or issue instructions for the corresponding general server 200 and GPU BOX server 100 according to each second status information and the first status information corresponding to the second status information.
[0041] In one embodiment, the node manager 300 may be, but is not limited to, a smart network card.
[0042] Based on this, the above intelligent management server system outputs, through a first preset number of GPU BOX servers 100, first status information for characterizing the working status of the GPU BOX servers 100; wherein, the GPU BOX servers 100 adopt a GPU architecture; then, through a first preset number of general servers 200, the first status information output by the corresponding GPU BOX servers 100 is received, and the corresponding first status information and second status information for characterizing the working status of the general servers 200 are output; wherein, the general servers 200 adopt a CPU architecture; then, the node manager 300 is used to communicatively connect to the management node of the rack-mounted management server, receive the corresponding first status information and second status information output by each general server 200, and output each first status information and second status information to the management node of the rack-mounted management server, so that the rack-mounted management server records each first status information and second status information, thereby realizing the cascading of the rack-mounted management server, the node manager 300, the corresponding general servers 200, and the corresponding GPU BOX servers 100, reducing the communication cables required for communication connection, thus reducing the material cost of the communication cables of the intelligent server system, improving the overall connection reliability of the intelligent server system, and reducing the cable management difficulty of the intelligent server system. In addition, it is convenient for the rack-mounted management server to perform monitoring operations or issue instructions on the corresponding general servers 200 and GPU BOX servers 100 according to each second status information and the first status information corresponding to the second status information.
[0043] In one embodiment, as Figure 4 shown, the node manager 300 includes a network switching module 310 and a network interface control module 320.
[0044] The network switching module 310 includes a second preset number of network interfaces; the network switching module 310 communicatively connects to the corresponding general servers 200 through each network interface; the network switching module 310 can receive the corresponding first status information and second status information output by each general server 200, and output each first status information and second status information. Wherein, the second preset number is greater than the first preset number.
[0045] In a specific example, as Figure 5As shown in the figure, the network switching module 310 can be, but is not limited to, a Lanswitch chip, and the Lanswitch chip can be, but is not limited to, a Marvell 88E6321 chip. The Marvell 88E6321 chip can have P0 to P4 connected downstream as five network interfaces, corresponding to connecting five groups of intelligent nodes, and the five groups of intelligent nodes are respectively communicatively connected to the corresponding general servers 200. The above are only specific examples, and in actual applications, it can be flexibly set according to user needs, and no limitation is imposed here.
[0046] The network interface control module 320 is communicatively connected to the network switching module 310, can be communicatively connected to the management node of the cabinet-style management server, and can also receive the first status information and the second status information output by the network switching module 310, and output the first status information and the second status information to the management node, so that the cabinet-style management server records the first status information and the second status information.
[0047] In a specific example, as Figure 5 shown, the network interface control module 320 can be, but is not limited to, a NIC chip, and the NIC chip can be, but is not limited to, an Inter E810 chip. The external expansion port SPF28 is directly connected to the management node or connected to the management node through a switch, and is connected to the TX / RX port of the network switching module 310, that is, the Lanswitch chip, through the SerDes protocol. The above are only specific examples, and in actual applications, it can be flexibly set according to user needs, and no limitation is imposed here.
[0048] In this embodiment, the node manager 300 configured with the network switching module 310 and the network interface control module 320 receives the corresponding first status information and second status information output by each general server 200, and outputs the first status information and the second status information to the management node, so that the cabinet-style management server records the first status information and the second status information, improving the convenience of the intelligent server system and facilitating the cabinet-style management server to perform monitoring operations or issue instructions for the corresponding general server 200 and the GPU BOX server 100 according to the second status information and the first status information corresponding to the second status information.
[0049] In one of the embodiments, as Figure 4 shown, the node manager further includes a baseboard management control module 330. Among them, the baseboard management control module is communicatively connected to the network switching module 310.
[0050] The baseboard management control module 330 can receive the first status information and the second status information output by the network switching module 310, and can also perform fault analysis according to the first status information and the second status information, and output a first target instruction.
[0051] Among them, the first target instruction is used to instruct the corresponding GPU BOX server 100 or general server 200 to avoid first-priority faults. It can be understood that the first-priority faults include abnormal power-off faults and liquid leakage faults.
[0052] In a specific example, such as Figure 5 shown, the baseboard management control module 330 can be but is not limited to a BMC chip, and the BMC chip can be but is not limited to an Aspeed AST2600 chip; at the same time, the BMC chip is communicatively connected to the network switching module 310, i.e., the Lanswitch chip, through MDC / MDIO to configure the working mode of the network switching module 310 and isolate network packets. In addition, P5 of the network switching module 310, i.e., the Lanswitch chip, is connected to MAC4 of the baseboard management module, i.e., the BMC chip, in the node manager 300 through RGMII. The MAC3 of the baseboard management module, i.e., the BMC chip, provides an NCSI signal. The above is only a specific example, and it can be flexibly set according to user requirements in actual applications, and is not limited here.
[0053] In a specific example, the baseboard management control module 330 can output the first target instruction obtained after fault analysis based on each first status information and second status information to the network switching module 310, so as to be forwarded by the network switching module 310 to the corresponding general server 200 or the network switching module 310 and the corresponding general server 200 to the corresponding GPU BOX server, thereby instructing the corresponding GPU BOX server 100 or general server 200 to avoid first-priority faults. In addition, the network switching module 310 is further configured to forward the first target instruction to the network interface control module 320, and output the first target instruction to the management node of the cabinet management server communicatively connected to the cabinet through the network interface control module 320, so that the cabinet management server records the first target instruction, which also enables the operation and maintenance personnel to timely understand whether first-priority faults occur in the intelligent server system through the first target instruction in the cabinet management server. The above is only a specific example, and it can be flexibly set according to user requirements in actual applications, and is not limited here.
[0054] In this embodiment, by configuring the baseboard management control module 330 in the node manager 300, and outputting the first target instruction after fault analysis of each first status information and second status information through the baseboard management control module 330 to instruct the corresponding GPU BOX server or general server to avoid first-priority faults including abnormal power-off faults and liquid leakage faults, the convenience and reliability of the intelligent server system are improved.
[0055] In one of the embodiments, such as Figure 4As shown, the node manager further includes a field programmable gate array module 340. Among them, the field programmable gate array module 340 is communicatively connected to the baseboard management control module 330.
[0056] The field programmable gate array module 340 is used to perform fault analysis based on each first status information and second status information, and output a second target instruction. It can be understood that the second target instruction is used to instruct the corresponding GPU BOX server 100 or general server 200 to avoid second-priority faults. Among them, the second-priority faults include downclocking operation faults and overclocking operation faults, and the priority of the second-priority faults is lower than that of the first-priority faults. In addition, the baseboard management control module 330 can also output each first status information and second status information to the field programmable gate array module 340.
[0057] In a specific example, as Figure 5 shown, the field programmable gate array module 340 can be, but is not limited to, an FPGA chip. The FPGA chip can also give the PCIe resources from the motherboard to the network interface control module 320, that is, the NIC chip. The above is only a specific example, and it can be flexibly set according to user needs in actual applications, and no limitation is made here.
[0058] In a specific example, the field programmable gate array module 340 can output the second target instruction obtained by performing fault analysis based on each first status information and second status information to the baseboard management control module 330, so as to be forwarded to the corresponding general server 200 through the baseboard management control module 330 and the network switching module 310 in sequence, or forwarded to the corresponding GPU BOX server through the network switching module 310 and the corresponding general server 200, so as to instruct the corresponding GPU BOX server 100 or general server 200 to avoid second-priority faults. In addition, the network switching module 310 is also used to forward the second target instruction to the network interface control module 320, and output the second target instruction to the management node communicatively connected to the cabinet management server through the network interface control module 320, so that the cabinet management server records the second target instruction, which also realizes that the operation and maintenance personnel can timely understand whether there are second-priority faults in the intelligent server system through the second target instruction in the cabinet management server. The above is only a specific example, and it can be flexibly set according to user needs in actual applications, and no limitation is made here.
[0059] In this embodiment, by configuring the field programmable gate array module 340 in the node manager 300, and performing fault analysis on each first status information and second status information through the field programmable gate array module 340, and then outputting a second target instruction to instruct the corresponding GPU BOX server or general server to avoid the second-priority faults including downclocking faults and overclocking faults, the convenience and reliability of the intelligent server system are improved.
[0060] In one embodiment, as Figure 4 shown, the field programmable gate array module 340 is communicatively connected to the network interface control module 320.
[0061] Among them, the field programmable gate array module 340 can also perform traffic control analysis based on each first status information and second status information, and output a third target instruction. It can be understood that the third target instruction is used to instruct the network interface control module to output a fourth target instruction to the network switching module; the fourth target instruction is used to instruct the network switching module to control the bandwidth of each general server.
[0062] In a specific example, as Figure 5 shown, the field programmable gate array module 340 can be but is not limited to an FPGA chip, the network interface control module 320 can be but is not limited to a NIC chip, and the network switching module 310 can be but is not limited to a Lanswitch chip. The FPGA chip can perform traffic control analysis based on each first status information and second status information and then output the third target instruction to the NIC chip. The NIC chip will output a first target instruction to the Lanswitch chip according to the indication of the third target instruction. The NIC chip can control the bandwidth of each general server according to the indication of the first target instruction, so as to adjust the network resources and computing resources of the server system, and complete the real-time monitoring of the network resources and computing resources of the server system, avoiding problems such as network congestion caused by limited bandwidth of the network uplink resources of the server system, thereby improving the data processing efficiency of the server system. The above is only a specific example, and it is flexibly set according to user needs in actual applications, and is not limited here.
[0063] In this embodiment, the field programmable gate array module 340 performs traffic control analysis based on each first status information and second status information and then outputs a third target instruction to instruct the network interface control module 320 to output a fourth target instruction to the network switching module; at the same time, the network switching module 310 controls the bandwidth of each general server according to the fourth target instruction, avoiding problems such as network congestion caused by limited bandwidth of the network uplink resources of the server system, thereby improving the data processing efficiency of the server system.
[0064] In one embodiment, asFigure 6 As shown, the GPU BOX server 100 includes a first complex programmable logic module 110; the general server 200 includes a second complex programmable logic module 210. Among them, the second complex programmable logic module 210 is communicatively connected to the corresponding first complex programmable logic module 110, and the network switching module 310 is communicatively connected to the second complex programmable logic module 210.
[0065] It can be understood that when the corresponding GPU BOX server 100 is in the standby mode, the first complex programmable logic module 110 can output first status sub-information to the corresponding second complex programmable logic module 210. Among them, the first status information includes the first status sub-information; the first status sub-information is used to represent that the GPU BOX server 100 is in the standby mode.
[0066] The second complex programmable logic module 210 can receive the first status sub-information output by the first complex programmable logic module 110, and can also output the first status sub-information and the second status sub-information to the network switching module 310 when the corresponding general server 200 is in the standby mode. Among them, the second status information includes the second status sub-information; the second status sub-information is used to represent that the corresponding general server 200 is in the standby mode.
[0067] The network switching module 310 can also receive the first status sub-information and the second status sub-information, and output the first status sub-information and the second status sub-information to the baseboard management control module 330.
[0068] The baseboard management control module 330 can also receive the first status sub-information and the second status sub-information output by the network switching module 310, and output a fifth target instruction when receiving the first status sub-information and the second status sub-information. Among them, the fifth target instruction is used to instruct the corresponding second complex programmable logic module 210 to power on, and output a sixth target instruction to the corresponding first complex programmable logic module 110 when the corresponding second complex programmable logic module 210 completes power-on; the sixth target instruction is used to instruct the corresponding first complex programmable logic module 110 to power on.
[0069] In a specific example, such as Figure 3As shown, the first complex programmable logic module 110 can be, but is not limited to, a CPLD chip, and the second complex programmable logic module 210 can be, but is not limited to, a CPLD chip. When each GPU BOX server 100 and general server 200 start to power on, the management node of the rack management server will start recording and monitoring each second status information and the information related to the second status information. When the corresponding GPU BOX server 100 is in the standby mode, that is, the corresponding GPU BOX server 100 is in the Stand by Ready state, it outputs the first status sub-information indicating that the GPU BOX server 100 is in the standby mode to the corresponding second complex programmable logic module 210.
[0070] Subsequently, the second complex programmable logic module 210 can receive the first status sub-information output by the first complex programmable logic module 110; when the corresponding general server 200 is in the standby mode, that is, the corresponding general server 200 is in the Stand by Ready state, so at this time both the corresponding GPU BOX server 100 and general server 200 are in the Stand by Ready state, then it outputs the first status sub-information and the second status sub-information indicating that the corresponding general server 200 is in the standby mode to the network switching module 310.
[0071] Next, the network switching module 310 can also receive the first status sub-information and the second status sub-information, and output the first status sub-information and the second status sub-information to the baseboard management control module 330. The baseboard management control module 330 receives the first status sub-information and the second status sub-information output by the network switching module 310, and when receiving the first status sub-information and the second status sub-information, it outputs the fifth target instruction to the second complex programmable logic module 210 through the network switching module 310.
[0072] Finally, the second complex programmable logic module 210 receives the fifth target instruction transmitted by the network switching module 310, and automatically powers on itself when receiving the fifth target instruction. When the second complex programmable logic module 210 completes its power-on, it outputs the sixth target instruction to the corresponding first complex programmable logic module 110 to enable the corresponding first complex programmable logic module 110 to power on. In addition, the sixth target instruction can be, but is not limited to, a 12V power-on signal provided by the VR chip in the general server 200. The above is only a specific example, and it is flexibly set according to user requirements in actual applications, and is not limited here.
[0073] In this embodiment, when the baseboard management control module 330 receives the first status sub-information and the second status sub-information, it outputs a fifth target instruction to instruct the corresponding second complex programmable logic module 210 to power on, and when the corresponding second complex programmable logic module 210 completes power-on, it outputs a sixth target instruction to the corresponding first complex programmable logic module 110. At the same time, the corresponding first complex programmable logic module 110 is instructed to power on through the sixth target instruction, so as to realize the automatic and orderly power-on of each GPU BOX server 100 and general server 200, and avoid the phenomenon of device loss during the power-on process of each GPU BOX server 100 and general server 200.
[0074] In one embodiment, as Figure 6 shown, the first complex programmable logic module 110 can also output a target feedback instruction to the corresponding second complex programmable logic module 210 when it completes its own power-on. Then, the second complex programmable logic module 210 is further configured to output a global reset instruction when it receives the target feedback instruction output by the first complex programmable logic module 110, and can instruct the corresponding GPU BOX server 100 and general server 200 to perform a global reset according to the global reset instruction.
[0075] In a specific example, as Figure 3 shown, the first complex programmable logic module 110 can be but is not limited to a CPLD chip, and the second complex programmable logic module 210 can be but is not limited to a CPLD chip. The first complex programmable logic module 110 can also output a target feedback instruction to the corresponding second complex programmable logic module 210 when it completes its own power-on, that is, it can output a Power Good signal to the corresponding second complex programmable logic module 210. Then, the second complex programmable logic module 210 is further configured to output a global reset instruction when it receives the Power Good signal output by the first complex programmable logic module 110, that is, to release the global reset instruction given by the PCH, so as to realize the global reset of the corresponding GPU BOX server 100 and general server 200. The above is only a specific example, and it is flexibly set according to user needs in actual applications, and is not limited here.
[0076] In this embodiment, when the first complex programmable logic module 110 completes its startup, it outputs a target feedback instruction to the corresponding second complex programmable logic module 210. Subsequently, the second complex programmable logic module 210 is further configured to output a global reset instruction when receiving the target feedback instruction output by the first complex programmable logic module 110, thereby implementing global reset of the corresponding GPU BOX server 100 and general server 200 according to the global reset instruction, increasing the management of the automatic power-on and power-off mechanism for each GPU BOX server 100 and general server 200, and facilitating maintenance by operation and maintenance personnel.
[0077] In one embodiment, the intelligent server system further includes a cabinet management server. Wherein, the cabinet management server includes a management node, and the management node is communicatively connected to the node manager 300; the cabinet management server is configured to receive and record each first status information and second status information. Therefore, the convenience of the intelligent server system is improved.
[0078] In one embodiment, the intelligent server system further includes a switch. Wherein, the management node is communicatively connected to the node manager 300 through the switch, that is, the switch can forward each first status information and second status information received from the node manager to the management node of the cabinet management server, so as to facilitate the cabinet management server to receive and record each first status information and second status information.
[0079] In the description of this specification, the descriptions referring to terms such as "some embodiments", "other embodiments", "ideal embodiments", etc. mean that the specific features, structures, materials, or features described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example.
[0080] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0081] The above-described embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent shall be subject to the appended claims.
Claims
1. An intelligent server system with multiple heterogeneous nodes includes: A first preset number of GPU BOX servers for outputting first status information; The first status information is used to characterize the working status of the GPU BOX servers; The GPU BOX servers adopt a GPU architecture; The first preset number of general servers are communicatively connected to the corresponding GPU BOX servers, and are used to receive the first status information output by the corresponding GPU BOX servers and output the corresponding first status information and second status information; the second status information is used to characterize the working status of the general servers; the general servers adopt a CPU architecture; A node manager communicatively connected to each of the general servers and used to communicatively connect to the management node of a rack-mounted management server; the node manager is further used to receive the corresponding first status information and the second status information output by each of the general servers and output each of the first status information and the second status information to the management node, so that the rack-mounted management server records each of the first status information and the second status information; The node manager includes: A network switching module, the network switching module includes a second preset number of network interfaces; the network switching module is communicatively connected to the corresponding general servers through each of the network interfaces; the network switching module is used to receive the corresponding first status information and the second status information output by each of the general servers and output each of the first status information and the second status information; The second preset number is greater than the first preset number; A network interface control module, which is used to communicatively connect to the network switching module and the management node respectively, and is further used to receive each of the first status information and the second status information output by the network switching module and output each of the first status information and the second status information to the management node.
2. The intelligent server system according to claim 1, characterized in that The node manager further includes: A baseboard management control module communicatively connected to the network switching module and used to receive each of the first status information and the second status information output by the network switching module; the baseboard management control module is further used to perform fault analysis according to each of the first status information and the second status information and output a first target instruction; the first target instruction is used to instruct the corresponding GPU BOX server or general server to avoid first-priority faults; wherein, the first-priority faults include abnormal power-off faults and liquid leakage faults.
3. The intelligent server system according to claim 2, wherein The node manager further includes a field programmable gate array module; the field programmable gate array module is communicatively connected to the baseboard management control module; The baseboard management control module is further used to output each of the first status information and the second status information to the field programmable gate array module; The field programmable gate array module is used to perform fault analysis based on each of the first state information and the second state information, and output a second target instruction; the second target instruction is used to instruct the corresponding GPU BOX server or the general server to avoid second-priority faults; wherein, the second-priority faults include downclocking faults and overclocking faults; the priority of the second-priority faults is lower than the priority of the first-priority faults.
4. The intelligent server system according to claim 3, characterized in that The field programmable gate array module is communicatively connected to the network interface control module; The field programmable gate array module is further used to perform traffic control analysis based on each of the first state information and the second state information, and output a third target instruction; wherein, the third target instruction is used to instruct the network interface control module to output a fourth target instruction to the network switching module; the fourth target instruction is used to instruct the network switching module to control the bandwidth of each of the general servers.
5. The intelligent server system according to claim 2, wherein The GPU BOX server includes a first complex programmable logic module; the general server includes a second complex programmable logic module; wherein, the second complex programmable logic module is communicatively connected to the corresponding first complex programmable logic module; the network switching module is communicatively connected to the second complex programmable logic module; The first complex programmable logic module is used to output first state sub-information to the corresponding second complex programmable logic module when the corresponding GPU BOX server is in the standby mode; wherein, the first state information includes the first state sub-information; the first state sub-information is used to represent that the GPU BOX server is in the standby mode; The second complex programmable logic module is used to receive the first state sub-information, and is further used to output the first state sub-information and second state sub-information to the network switching module when the corresponding general server is in the standby mode; wherein, the second state information includes the second state sub-information; the second state sub-information is used to represent that the corresponding general server is in the standby mode; The network switching module is further used to receive the first state sub-information and the second state sub-information, and output the first state sub-information and the second state sub-information to the baseboard management control module; The baseboard management control module is further used to receive the first state sub-information and the second state sub-information; the baseboard management control module is further used to output a fifth target instruction when receiving the first state sub-information and the second state sub-information; wherein, the fifth target instruction is used to instruct the corresponding second complex programmable logic module to power on, and output a sixth target instruction to the corresponding first complex programmable logic module when the corresponding second complex programmable logic module completes power-on; the sixth target instruction is used to instruct the corresponding first complex programmable logic module to power on.
6. The intelligent server system according to claim 5, characterized in that, The first complex programmable logic module is further used to output a target feedback instruction to the corresponding second complex programmable logic module when power-on is completed; The second complex programmable logic module is further configured to receive the target feedback instruction; the second complex programmable logic module is further configured to output a global reset instruction when receiving the target feedback instruction; wherein, the global reset instruction is used to instruct the corresponding GPU BOX server and the general server to perform a global reset.
7. The intelligent server system according to claim 1, wherein The intelligent server system further includes: The cabinet-type management server, the cabinet-type management server includes the management node, and the management node is communicatively connected to the node manager; the cabinet-type management server is configured to receive and record each of the first status information and the second status information.
8. The intelligent server system according to claim 7, wherein The intelligent server system further includes a switch; wherein, the management node is communicatively connected to the node manager through the switch.
9. The intelligent server system according to claim 1, wherein The node manager is an intelligent network card.
Citation Information
Patent Citations
Cloud service architecture based on hybrid heterogeneous processor
CN113609068A
Heterogeneous interconnection system and cluster
CN114968895A