Server management method, computer program product and server cabinet
By defining the target management interface in the server management controller, the server information of each server, including the target U-bit location, is solved, and the cable structure and management cost problems added by additional connection to the U-bit asset management module in the prior art are solved, thereby achieving efficient management and cost reduction of the server.
Patent Information
- Application Number
- CN202510667795.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The prior art monitors the server information inside the server cabinet by additionally connecting the cabinet management controller to the U-bit asset management module, resulting in an increase in cable structure and management costs in the cabinet.
By defining the target management interface in the server management controller, the server information of each server, including the target U-bit location, is obtained, and the server management is realized based on the PIN determination of the PCIe resource bus unit.
The cable structure in the cabinet is reduced, and the identification and unified management of U-bit assets, network addresses and other information of each server node are realized, reducing management costs.
Smart Images

Figure CN120196513A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of servers, and particularly to a server management method, a computer program product, and a server cabinet. Background Art
[0002] For a converged architecture server whole cabinet (referred to as a server cabinet), by modularizing the cable interfaces of the servers to form blind plug connectors, the centralized power supply bus, network bus, and liquid cooling bus in the cabinet form a bus interface module, greatly simplifying the cable connection in the cabinet. However, due to business requirements, the server cabinet also needs to add a PCIe resource bus unit to form a multi-bus converged architecture cabinet with a centralized power supply bus, network bus, liquid cooling bus, and PCIe resource bus.
[0003] In related technologies, by additionally connecting a U-position asset management module to the cabinet management controller to monitor the server information inside the server cabinet, the cable structure and management cost in the cabinet are increased. Summary of the Invention
[0004] The present application provides a server management method, a computer program product, a server cabinet, an electronic device, and a computer-readable storage medium, so as to at least solve the problem in related technologies that by using a U-position asset management module to monitor the server information inside the server cabinet, the cable structure and management cost in the cabinet are increased.
[0005] The present application provides a server management method, which is applied to a cabinet management controller of a server cabinet; the server cabinet includes multiple servers; each server has multiple target buses with a unified standard; the multiple target buses of each server converge on the PCIe resource bus unit of the server; the method includes: When powered on, periodically access the server management controllers of each server; the server management controller defines a target management interface for providing server information; Obtain the server information of each server from the target management interface of each server management controller; the server information includes the target U-position of the corresponding server; the target U-position is determined by each server management controller according to each PIN of the corresponding PCIe resource bus unit of the corresponding server; each PIN is used to indicate the partition identifier and U-position identifier of the server; Manage each server based on the server information of each server.
[0006] The present application also provides a server management method, which is applied to the server management controllers of each server in a server cabinet; each server included in the server cabinet has a variety of target buses with a unified standard; the various target buses of each server converge on the PCIe resource bus unit of the server; the method includes: In the case of being powered on, define a target management interface for providing server information; In the case of being accessed by the cabinet management controller of the server cabinet, provide server information to the cabinet management controller based on the target management interface, so that the cabinet management controller manages the server based on the server information obtained from the target management interface; The server information includes the target U-position; the target U-position is determined based on the following method: Determine the target U-position of the corresponding server according to each PIN possessed by the PCIe resource bus unit of the corresponding server; each PIN is used to indicate the partition identifier and U-position identifier of the server.
[0007] The present application also provides a computer program product, which is applied to the cabinet management controller of a server cabinet; the server cabinet includes multiple servers; each server has a variety of target buses with a unified standard; the various target buses of each server converge on the PCIe resource bus unit of the server; the computer program product includes: A first communication module, which is used to periodically access the server management controllers of each server included in the server cabinet in the case of being powered on; the server management controller defines a target management interface for providing server information; An acquisition module, which is used to obtain the server information of each server from the target management interface of each server management controller; the server information includes the target U-position of the corresponding server; the target U-position is determined by each server management controller according to each PIN possessed by the corresponding PCIe resource bus unit of the corresponding server; each PIN is used to indicate the partition identifier and U-position identifier of the server; A management module, which is used to manage each server based on the server information of each server.
[0008] The present application also provides a computer program product, which is applied to the server management controllers of each server in a server cabinet; each server included in the server cabinet has a variety of target buses with a unified standard; the various target buses of each server converge on the PCIe resource bus unit of the server; the computer program product includes: A definition module, which is used to define a target management interface for providing server information in the case of being powered on; A second communication module, configured to, when being accessed by a cabinet management controller of a server cabinet, provide server information to the cabinet management controller based on a target management interface, so that the cabinet management controller manages the server based on the server information obtained from the target management interface; The server information includes a target U-position; the target U-position is determined based on the following method: Determine the target U-position of the corresponding server according to each PIN of the PCIe resource bus unit of the corresponding server; each PIN is used to indicate the partition identifier and U-position identifier of the server.
[0009] This application also provides a server cabinet, which includes a cabinet management controller and multiple servers; each server has multiple target buses and a server management controller with a unified standard; the multiple target buses of each server converge on the PCIe resource bus unit of the server; When powered on, the server management controller defines a target management interface for providing server information; the server information includes the target U-position of the corresponding server; the target U-position is determined by each server management controller according to each PIN of the corresponding PCIe resource bus unit of the corresponding server; each PIN is used to indicate the partition identifier and U-position identifier of the server; When powered on, the cabinet management controller periodically accesses the server management controllers of each server; obtains the server information of each server from the target management interfaces of each server management controller, and manages each server based on the server information of each server.
[0010] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above server management methods when executing the computer program.
[0011] This application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above server management methods are implemented.
[0012] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the above server management methods are implemented.
[0013] With this application, since the server has multiple target buses with unified standards, and the multiple target buses converge on the PCIe resource bus unit of the server, each server management controller determines the target U-bit position of the server based on each PIN of the corresponding PCIe resource bus unit of the corresponding server. This target U-bit position belongs to the server information. The server management controller defines an interface for providing server information to the target management interface, and the cabinet management controller obtains the server information of each server by accessing the target management interface. Therefore, it can solve the technical problem in the related art that by additionally connecting the cabinet management controller to the U-bit asset management module to monitor the server information inside the server cabinet, the cable structure and management cost inside the cabinet are increased. This application can further reduce the cable structure inside the cabinet, and can realize the identification and unified management of the U-bit assets, network addresses, and other server information of each server node, reducing the management cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0015] Figure 1 FIG. 1 is a schematic flowchart of a method for managing a server provided by an embodiment of the present application; Figure 2 FIG. 2 is a schematic diagram of the cabinet structure of a server cabinet provided by an embodiment of the present application; Figure 3 FIG. 3 is a schematic diagram of the architecture of partitions in a server cabinet provided by an embodiment of the present application; Figure 4 FIG. 4 is a schematic flowchart of a method for managing a server provided by an embodiment of the present application; Figure 5 FIG. 5 is a schematic diagram of the connection between a server and a PCIe bus connector provided by an embodiment of the present application; Figure 6 FIG. 6 is a schematic block diagram of a computer program product provided by an embodiment of the present application; Figure 7 FIG. 7 is a schematic block diagram of a computer program product provided by an embodiment of the present application; Figure 8 FIG. 8 is a schematic diagram of the interaction process between a cabinet management controller and the server management modules of each server provided by an embodiment of the present application; Figure 9 FIG. 9 is a schematic diagram of the interaction process between a cabinet management controller and the server management modules of each server provided by an embodiment of the present application. Detailed implementation manners
[0016] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0017] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0018] To meet the requirements of flexible resource allocation in different application scenarios and different needs such as artificial intelligence, machine learning, and intelligent computing, data centers are accelerating the transformation from a computing-centric architecture to a data-centric integrated architecture. The integrated architecture system centralizes similar resources at the physical level to form a general computing resource pooling server, a PCIe resource switching server or unit, a heterogeneous acceleration resource pooling server, etc. respectively. Through decoupling, the resources are decomposed into different independent pooling server chassis to facilitate centralized management and expansion of resources. However, each pooling server node still needs to be connected to the PCIe resource switching server through cables to enable flexible resource allocation. This makes the scale of the interconnection cables at the cabinet level huge, the wiring complex, and the structure redundant, becoming an invisible resistance to deployment and operation.
[0019] In some open data center cabinet design standards, such as OCP / Open19 / ODCC / OCTC, by modularizing the cable interfaces of the servers to form blind plug connectors, the centralized power supply bus, network bus, and liquid cooling bus in the cabinet form bus interface modules, greatly simplifying the cable connection in the cabinet. Such blind plug servers have formed a deployment scale and a customized ecosystem, which can significantly improve cost and energy efficiency. For the integrated architecture server whole cabinet, it is also necessary to add a PCIe resource bus module so that the independent pooling resource servers are aggregated to the PCIe resource switching server or unit through high-speed blind plug connectors and the PCIe bus module to form a four-bus integrated architecture cabinet of the centralized power supply bus, network bus, liquid cooling bus, and PCIe resource bus (hereinafter referred to as the server cabinet). Centered on the PCIe resource switching unit, a flexible configurable integrated architecture server system is constructed.
[0020] Although the pooled server nodes inside the server cabinet are highly decoupled, they need to be able to operate like an integrated server system and be conveniently fault-located and operation and maintenance managed. Therefore, an efficient solution is needed to identify and uniformly manage the positions, server types, management addresses, and basic information of each pooled server node in the cabinet. Inside a traditional server cabinet, since each server node is relatively independent and lacks a unified interconnection interface, to achieve functions such as automatically and real-time monitoring the cabinet space status, automatically and real-time monitoring the model of the server nodes in the cabinet, the number of occupied U positions, and the specific positions in the rack, etc., an RMC is usually used to additionally connect to a U-position asset management module to manage server information, which increases the cable structure and management cost inside the cabinet.
[0021] To enable those skilled in the art of this technical field to better understand the solution of this application, the following further elaborates on this application in conjunction with the accompanying drawings and specific implementation manners.
[0022] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the management method of the server depends, the specific application environment architecture or specific hardware architecture is described herein.
[0023] The embodiments of this application provide a management method for a server. In combination with the execution process of the management method of the server, the method is described in detail.
[0024] The management method for the server provided by the embodiments of this application can be executed by a Rack Management Controller (RMC) of a server cabinet.
[0025] In practical applications, an RMC is usually used in data centers, computer rooms, or server cabinets, and is responsible for monitoring and managing the hardware devices inside the cabinet, including servers, storage devices, network devices, etc. Its core functions are remote management, monitoring, fault troubleshooting, and resource allocation of the hardware. An RMC usually integrates some advanced management functions, such as power management, temperature control, and hardware health monitoring.
[0026] The server cabinet includes multiple servers; each server has multiple target buses with a unified standard; The embodiments of this application can specifically be an integrated architecture server whole cabinet. This integrated architecture server whole cabinet includes multiple target buses, including but not limited to a power supply bus, a network bus, a liquid cooling bus, a resource bus, etc., and may also include other buses, such as an address bus, an expansion bus, a system bus, a serial bus, and a parallel bus, etc.
[0027] The multiple target buses of each server converge on the PCIe resource bus unit of the server.
[0028] The PCIe resource bus unit is a key component in the PCIe system responsible for managing and allocating bus resources. By providing functions such as resource allocation, device addressing, bandwidth management, and configuration management, it ensures that multiple devices can work efficiently and stably while sharing the bus.
[0029] Multiple target buses converge on the PCIe resource bus unit of the server. Each server service can manage the buses of the server cabinet without the need to connect to the U-bit asset management module through the server management controller additionally, which can effectively improve the efficiency of resource management and simplify the architecture and management of the server cabinet.
[0030] See Figure 1 , one of the flow schematic diagrams of a server management method provided by an embodiment of the present application includes the following steps 110, 120, and 130: Step 110: Periodically access the server management controllers of each server when powered on. Step 120: Obtain the server information of each server from the target management interface of each server management controller; the server management controller defines the target management interface for providing server information; the server information includes the target U-bit position of the corresponding server; the target U-bit position is determined by each PIN of the corresponding PCIe resource bus unit of the corresponding server; each PIN is used to indicate the partition identifier and U-bit identifier of the server. Step 130: Manage each server based on the server information of each server.
[0031] The server cabinet in the embodiment of the present application includes multiple partitions, and each partition has multiple servers. Each partition includes multiple (for example, 4) computing servers, one memory pooling server, and multiple (for example, 4) heterogeneous acceleration pooling servers.
[0032] The cabinet management controller in the server cabinet needs to understand the server information of each server in detail to achieve overall scheduling of the acceleration calculators (such as GPUs) in multiple heterogeneous acceleration pooling servers. The server information includes at least some of the server type, model, serial number, U-bit number, target U-bit address, U-bit identifier (starting U-bit address), and the first network address.
[0033] The cabinet server of the server can periodically obtain the server information of each server from the server management controller of each server.
[0034] The server management controller is a key tool for ensuring the healthy operation of the server. Through functions such as remote control, hardware monitoring, and logging, it helps administrators maintain full control of the server at any time. It is especially crucial for the operation of data centers, which can reduce server downtime and improve server availability and maintenance efficiency.
[0035] The operating status of the server management controller has little correlation with the corresponding server and can be regarded as independent. The server management controller itself has an IP address and can directly communicate with the cabinet management controller.
[0036] In practical applications, the server management controller can specifically be the Baseboard Management Controller (BMC). The BMC is a dedicated hardware management controller in the server, which is usually integrated on the motherboard of the server and is used to provide hardware monitoring, management, and remote control functions.
[0037] Each server has a PCIe resource bus unit. To save the management cost of the server, in the embodiment of the present application, each PIN is set in the PCIe resource bus unit, and each PIN is used to indicate the partition identifier and U-bit identifier of the server. The partition identifier is used to indicate the partition where the server is located, and the U-bit identifier is used to indicate the starting U-bit position of the server.
[0038] The server management controller reads each PIN in each PCIe resource bus unit to obtain the partition identifier and U-bit identifier of its corresponding server, and can determine the target U-bit position of the server based on the partition identifier and U-bit identifier of the server, as well as the number of U-bits of the server. The target U-bit position is used to indicate the highest U-bit position of the server.
[0039] Then, each server management controller configures a first network address for each server according to the target U-bit position of the corresponding server. Based on the first network addresses of each server, a local area network inside the server can be constructed, and the servers inside the server cabinet can communicate with each other within this local area network.
[0040] To enable the cabinet management controller to obtain the server information including the target U-bit position, the server management controller defines a target management interface based on the server information of the corresponding server. The target management interface is a software interface, and the BMC can expose the server information of the server through this target management interface. The server information belongs to hardware resources, that is, the BMC exposes hardware resources through a software interface.
[0041] The cabinet management controller can directly access the target management interfaces of each server management controller to obtain the server information of each server.
[0042] The target management interface can be a Redfish interface, which is a standard interface based on RESTful API and aims to provide more modern, easily extensible, and automated management functions for server management.
[0043] After obtaining the server information of each server, the cabinet management controller can manage each server, generate a management topology diagram, and provide computing services, memory services, etc. externally based on the management diagram containing the server information of each server.
[0044] Through the server management method provided by this application, since the server has multiple target buses with a unified standard, and the multiple target buses converge on the PCIe resource bus unit of the server, each server management controller determines the target U-bit position of the server according to each PIN of the corresponding PCIe resource bus unit of the corresponding server. This target U-bit position belongs to the server information. The server management controller defines to provide the server information to the target management interface, and the cabinet management controller obtains the server information of each server by accessing the target management interface. Therefore, it can solve the technical problem in the related art that by additionally connecting a U-bit asset management module to the cabinet management controller to monitor the server information inside the server cabinet, the cable structure and management cost inside the cabinet are increased. This application can further reduce the cable structure inside the cabinet, and can realize the identification and unified management of the U-bit assets, network addresses, and other server information of each server node, reducing the management cost.
[0045] In some embodiments, the server information further includes the first network address of the server; the first network address is determined by the server management controller according to the target U-bit position of the corresponding server; after obtaining the server information of each server from the target management interface of each server management controller, the method further includes: Controlling an automated network building tool to build a local area network inside the server cabinet based on the first network address in the server information, so that each server conducts network communication under the local area network.
[0046] After obtaining the first network address of each server in the embodiment of this application, the local area network inside the server can be built through the automated network building tool based on the first network address of each server and the network addresses of each switch in the server, so that each server conducts network communication under the local area network.
[0047] The automated network building tool includes at least one of, but is not limited to, Ansible, Chef, NetBox, etc.
[0048] The embodiments of the present application build a local area network inside the server cabinet based on the first network addresses of each server, which can significantly improve the performance, security, manageability, and scalability of the network inside the server cabinet, and has high cost-effectiveness.
[0049] In some embodiments, multiple target buses include PCIe buses; the server cabinet includes multiple partitions; each server in each partition includes a memory pooling server, multiple computing servers, and multiple heterogeneous acceleration pooling servers; Multiple computing servers are connected to the memory pooling server through the PCIe bus to share the memory resources of the memory pooling server; Multiple computing servers are connected to multiple heterogeneous acceleration pooling servers through the PCIe bus to share the acceleration computing resources in the multiple heterogeneous acceleration pooling servers.
[0050] The server cabinet in the embodiments of the present application includes multiple partitions, and each partition includes a memory pooling server, multiple computing servers, and multiple heterogeneous acceleration pooling servers.
[0051] Among them, a memory pooling server (Memory Pooling Server) is a server system that uses memory pooling technology to optimize memory management. Memory pooling technology pre-allocates a large memory area (memory pool), and then allocates and reclaims memory blocks from it, avoiding frequent memory allocation and release operations, thereby reducing memory fragmentation and improving memory utilization efficiency, especially suitable for scenarios with high concurrency and high real-time performance requirements.
[0052] The memory pooling server can specifically be a CXL (Compute Express Link) memory pooling server, and the CXL memory pooling server is a new type of server architecture that uses CXL technology to implement memory pooling.
[0053] A computing server refers to a server specifically designed to handle compute-intensive tasks. The computing server mainly focuses on providing efficient computing power for executing a large number of data processing, computing tasks, virtualization, simulation, analysis, and other tasks.
[0054] A heterogeneous acceleration pooling server refers to a server that integrates multiple computing accelerators (such as CPUs, GPUs, FPGAs, TPUs, etc.) and manages and schedules resources through a pooling method. It uses different types of hardware accelerators to process different types of workloads and improves computing efficiency and flexibility through resource pooling.
[0055] See Figure 2, The embodiment of the present application provides a schematic diagram of the cabinet structure of a server cabinet. There are two partitions in this server cabinet, namely the first partition zone0 and the second partition zone1. Each partition includes 1 CXL memory pooling server (abbreviated as CXL Memory Box in this embodiment), 4 general computing servers (abbreviated as Host in this embodiment), and 4 heterogeneous acceleration pooling servers (abbreviated as GPU Box in this embodiment).
[0056] Among them, the 4 computing servers included in the first partition zone0 are respectively Host 00, Host 01, Host 02, and Host 03, 1 CXL memory pooling server is CXL Memory Box 00, and the 4 heterogeneous acceleration pooling servers are respectively GPU Box 00, GPU Box 01, GPU Box 02, and GPU Box 03.
[0057] Among them, the 4 computing servers included in the second partition zone1 are respectively Host 10, Host 11, Host 12, and Host 13, 1 CXL memory pooling server is CXL Memory Box 10, and the 4 heterogeneous acceleration pooling servers are respectively GPU Box 10, GPU Box 11, GPU Box 12, and GPU Box 13.
[0058] The PCIe bus in the embodiment of the present application is used for the interconnection of PCIe memory pooling resources within the partition. That is, each computing server is connected to the memory pooling server through the PCIe bus, so that multiple computing servers can share the memory resources of the memory pooling server. In addition, multiple computing servers are connected to multiple heterogeneous acceleration pooling servers through the PCIe bus, so as to share the acceleration computing resources in the multiple heterogeneous acceleration pooling servers.
[0059] In some embodiments, the memory pooling server includes a memory pool, a memory exchange chip, and pooling management software; The acceleration pooling server includes a PCIe exchange chip and a computing accelerator; The pooling management software is used to control the memory exchange chip in the memory pooling server to allocate the memory resources in the memory pool for each computing server according to the memory allocation requests of each computing server; The pooling management software is used to control the PCIe exchange chip in each acceleration pooling server to allocate the acceleration resources in the acceleration calculator for each computing server according to the acceleration computing requests of each computing server.
[0060] The memory allocation request in the embodiment of this application is used to request the allocation of memory for a computing server. After the pooling management software obtains the memory allocation request of the computing server, it controls the memory exchange chips in each memory pooling server to allocate the memory resources in the memory pool for each computing server.
[0061] The memory exchange chip (CXL Switch) can effectively manage the data interaction between the memory in the memory pooling server and external devices (such as computing servers), transfer infrequently used data from the memory to the hard disk, and load the required data back from the hard disk to the memory.
[0062] The memory exchange chip can specifically be a CXL memory exchange chip.
[0063] The acceleration computing request in the embodiment of this application is used to request the allocation of acceleration computing resources for a computing server. After the pooling management software obtains the acceleration computing request of the computing server, it can control the PCIe switch chips in each acceleration pooling server to allocate an acceleration calculator for the computing server, so as to allocate acceleration resources for it.
[0064] The PCIe switch chip (PCIe Switch) is a hardware device used to expand the PCI Express (PCIe) bus, mainly used to manage and optimize the data transmission between PCIe devices of multiple servers, expand the scalability and parallel processing capabilities of the system, improve the data transmission efficiency, and manage the traffic between different devices.
[0065] The acceleration calculator can be at least one of a GPU acceleration calculator, an FPGA acceleration calculator, an ASIC acceleration calculator, and a CPU acceleration calculator, etc.
[0066] The pooling management software in the embodiment of this application combines the memory exchange chip and the PCIe switch chip, and can realize the adjustment of the pooled memory resource ratio and acceleration computing resources, so that the configuration of each computing server in the partition can be flexibly adjusted, thereby providing different computing power services according to the actual business scenario, avoiding resource redundancy, and improving resource utilization efficiency.
[0067] See Figure 3 , the embodiment of this application provides a schematic diagram of the architecture of a partition in a server cabinet, including a cabinet management controller RMC, 4 computing servers (Host*4), a CXL memory pooling server (CXL Memory Box*1), and 4 heterogeneous acceleration pooling servers (CPU Box*4). Each server has a corresponding server management controller BMC, and the RMC and BMC communicate through a network bus.
[0068] The CXL memory pooling server (CXL Memory Box*1) is equipped with a pooling management software (MCPU), a CXL memory switching chip (CXL Switch), a memory expansion controller MXC, and pooled memory (CXL Memory); The heterogeneous acceleration pooling server (CPU Box*4) includes a PCIe switching chip (PCIe Switch) and a computing accelerator (such as a CPU Device).
[0069] In addition, the computing server (Host*4), the CXL memory pooling server (CXL Memory Box*1), and the heterogeneous acceleration pooling server (CPU Box*4) are also connected through a PCIe bus. Specifically, the computing server (Host*4) can be connected to the CXL memory switching chip (CXL Switch) of the acceleration pooling server (CPU Box*4) and the PCIe switching chip (PCIe Switch) of the heterogeneous acceleration pooling server (CPU Box*4) through the PCIe bus.
[0070] In some embodiments, multiple target buses include a network bus; the server cabinet includes a first switch, a second switch, and a third switch; The network bus includes a service data network bus, a remote direct memory access RDMA data network bus, and an out-of-band management network bus; Multiple computing servers are connected to the first switch through the service data network bus to provide computing resources to the external devices of the multiple computing servers through the first switch; Multiple heterogeneous acceleration pooling servers are connected to the second switch through the RDMA data network bus to provide acceleration resources to the external devices of the multiple heterogeneous acceleration pooling servers through the second switch; The cabinet management controller and each server management controller are connected to the third switch through the out-of-band management network bus to centrally control and manage each server through the third switch.
[0071] Continue Figure 2 Furthermore, the server cabinet in the embodiment of the present application further includes a first switch, a second switch, and a third switch. These three switches can be TOR switches. Specifically, in Figure 2 the first switch can specifically refer to a TORHost service network switch, the second switch can specifically refer to a TOR RDMA switch; the third switch can specifically refer to a TOR management network switch.
[0072] Multiple computing servers can be connected to the first switch via the service data network bus in the network bus, so as to provide computing resources to the external devices of the multiple computing servers through the first switch; multiple heterogeneous acceleration pooling servers are connected to the second switch via the RDMA data network bus in the network bus, so as to provide acceleration resources to the external devices of the multiple heterogeneous acceleration pooling servers through the second switch; the cabinet management controller and each server management controller are connected to the third switch via the out-of-band management network bus in the network bus, so as to centrally control and manage each server through the third switch.
[0073] In some embodiments, the multiple target buses include a water-cooling bus; each server includes a liquid-cooling pipeline; The water-cooling bus is respectively connected to each liquid-cooling pipeline and an external water-cooling unit; the water-cooling bus is used to convey the coolant of the external water-cooling unit to the liquid-cooling pipeline to dissipate heat from each server.
[0074] The water-cooling bus provides cooling water and circulating power by an external water-cooling unit (external means outside the cabinet), is connected to each server through the liquid-cooling pipeline, thereby dissipating heat from the server and controlling the temperature of server components within the working range.
[0075] Continue Figure 2 , Figure 2 The server cabinet is connected to the external water-cooling unit, and the external water-cooling unit can be connected to each server, switch, etc., so as to dissipate heat from each part inside the server cabinet.
[0076] In some embodiments, the multiple target buses include a power supply bus; the server cabinet includes a power management device; Each server in the server cabinet is connected to the power management device via the power supply bus to obtain the power allocated by the power management device through the power supply bus.
[0077] The servers and switches inside the server cabinet can be connected to the power management device via the power supply bus. The power management device is mainly used to provide stable power supply in data centers and large enterprises. The power management device can distribute the power of the power supply to each server through the power supply common line to ensure the stable power supply of the server and guarantee the normal operation of the server cabinet.
[0078] In some embodiments, the power management device includes two power racks; The cabinet management controller is located in any one of the power racks.
[0079] The power management device in the embodiments of the present application can specifically be two power racks. A power rack is a device architecture used in places such as data centers and computer rooms, and is specifically designed to support the integration and management of various IT devices and power systems. A power rack usually combines functions such as power supply, heat dissipation, equipment installation, and monitoring to ensure that devices operate in an efficient and safe environment.
[0080] A rack management controller RMC is deployed in one of the power racks.
[0081] The power rack in the embodiments of the present application can specifically be a PowerShelf. The PowerShelf not only provides basic mounting bracket functions, but also additionally provides power supply (through a power distribution unit PDU) and heat dissipation management (such as built-in fans or a cooling system), enabling the device to operate under more optimized conditions.
[0082] Continue Figure 2 , inside the server mechanism, there are a power rack PowerShelf 00 and a power rack PowerShelf01. Among them, PowerShelf 00 includes a rack management controller, which can be characterized as PowerShelf 00 (BMC) in Figure 2 .
[0083] The BMC in the embodiments of the present application is set at any power rack node. On the one hand, it is used for the management and control of the two power racks. On the other hand, it is responsible for the identification and centralized management of the two converged architecture partitions in the entire cabinet, enabling each partition of the server mechanism to operate as flexibly as a single server system.
[0084] In some embodiments, the method further includes: Based on a preset power-on sequence, determine the target server to be powered on currently, and send a power-on request to the target management interface of the server management controller of the target server to power on the target server through the server management controller; At a preset duration after sending the power-on request, obtain the power-on status of the target server from the target management interface of the target server. When the power-on status is the powered-on status and there are unpowered servers, determine the next target server to be powered on based on the power-on sequence.
[0085] After the BMC obtains the server information of each server in the server cabinet, it can establish the management ability for the entire server cabinet and can perform management operations based on Figure 3 the provided architecture schematic diagram.
[0086] The server nodes in the partition have power-on timing requirements. When powering on, a certain power-on sequence is required. For example, the power-on sequence from first to last is: the heterogeneous acceleration resource pool server (GPU Box) nodes and the CXL memory pooling server (CXL Memory Box) nodes are powered on simultaneously, and the general computing server (Host) nodes are powered on later. When powering off, the power-off sequence needs to be reversed.
[0087] Through the above management architecture, a coordinated power-on and power-off operation interface can be implemented in the RMC. When powering on, first send a Redfish interface request to the GPU Box nodes and the CXL Memory Box BMC for power-on operations, and poll the Redfish interface for the power-on status of the GPU Box and the CXL Memory Box. When the power-on status is the powered-on status, it indicates that the GPU Box nodes and the CXL Memory Box nodes have completed power-on. Thus, a power-on request can be sent to the next target server, the Host node, and finally the coordinated power-on operation is completed. By establishing this management architecture, each partition can operate like a single complete server system, greatly reducing the management complexity and enabling the system to be flexibly adjusted and easily managed.
[0088] See Figure 4 , the second flowchart of the server management method provided by the embodiment of the present application can be executed by the server management controller of each server in the server cabinet. Each server included in the server cabinet has a variety of target buses with unified standards; the various target buses of each server converge on the PCIe resource bus unit of the server. The method includes: steps 410 to 420: Step 410, when powered on, define a target management interface for providing server information; Step 420, when accessed by the cabinet management controller of the server cabinet, provide server information to the cabinet management controller based on the target management interface, so that the cabinet management controller can manage the server based on the server information obtained from the target management interface; The server information includes the target U-bit position; the target U-bit position is determined based on the following method: Determine the target U-bit position of the corresponding server according to each PIN of the PCIe resource bus unit of the corresponding server; each PIN is used to indicate the partition identifier and the U-bit identifier of the server.
[0089] The server cabinet includes multiple servers; each server has a variety of target buses with unified standards; An embodiment of the present application may specifically be an integrated architecture server rack, and the integrated architecture server rack includes a variety of target buses, including but not limited to a power supply bus, a network bus, a liquid cooling bus, a resource bus, etc., and may also include other buses, such as an address bus, an expansion bus, a system bus, a serial bus, and a parallel bus, etc.
[0090] The various target buses of each server converge on the PCIe resource bus unit of the server.
[0091] The PCIe resource bus unit is a key component in the PCIe system responsible for managing and allocating bus resources. It ensures that multiple devices can work efficiently and stably while sharing the bus by providing functions such as resource allocation, device addressing, bandwidth management, and configuration management.
[0092] The various target buses converge on the PCIe resource bus unit of the server. Each server service can manage the buses of the server cabinet without the need to connect to the U-bit asset management module through the server management controller additionally, which can effectively improve the efficiency of resource management and simplify the architecture and management of the server cabinet.
[0093] The server management controller is a key tool for ensuring the healthy operation of the server.
[0094] The server management controller may specifically be a (Baseboard Management Controller, BMC). The BMC is a dedicated hardware management controller in the server, which is usually integrated on the motherboard of the server and is used to provide hardware monitoring, management, and remote control functions.
[0095] Each server has a PCIe resource bus unit. To save the management cost of the server, in the embodiment of the present application, each PIN is set in the PCIe resource bus unit. Each PIN is used to indicate the partition identifier and the U-bit identifier of the server. The partition identifier is used to indicate the partition where the server is located, and the U-bit identifier is used to indicate the starting U-bit position of the server.
[0096] The server management controller reads each PIN in each PCIe resource bus unit, thereby obtaining the partition identifier and the U-bit identifier of its corresponding server. Based on the partition identifier and the U-bit identifier of the server, and the number of U-bits of the server, the target U-bit position of the server can be determined. The target U-bit position is used to indicate the highest U-bit position of the server.
[0097] The target U-bit position can be subsequently used to configure the first network address for each server, and then build a local area network inside the server. The servers inside the server cabinet can perform network communication within this local area network.
[0098] The server management controller can store the target U-bit position in its own management memory.
[0099] To enable the rack management controller to obtain server information including the target U-bit position, the first network address, etc., the server management controller defines a target management interface based on the server information of the corresponding server. The rack management controller can directly access the target management interfaces of each server management controller to obtain the server information of each server.
[0100] The target management interface can be a Redfish interface. The Redfish interface is a standard interface based on RESTful API, aiming to provide more modern, easily extensible and automated management functions for server management.
[0101] After obtaining the server information of each server, the rack management controller can manage each server, generate a management topology map, and provide computing services, memory services, etc. based on the management map containing the server information of each server.
[0102] Through the server management method provided by this application, the multiple target buses of each server with a unified standard converge on the PCIe resource bus unit of the server. Each server management controller determines the target U-bit position of the server according to each PIN of the corresponding PCIe resource bus unit of the corresponding server. The target U-bit position belongs to the server information. The server management controller defines a target management interface for providing server information. The rack management controller obtains the server information of each server by accessing the target management interface, further reducing the cable structure in the rack, and can realize the identification and unified management of the U-bit assets, network addresses and other information of each server node, reducing the management cost.
[0103] In some embodiments, the PCIe resource bus unit includes a PCIe bus connector; the PCIe bus connector includes an I / O expander and each PIN, and each server management controller is connected to the I / O expander of the corresponding PCIe bus connector; Determining the target U-bit position of the corresponding server according to each PIN of the corresponding PCIe resource bus unit of the corresponding server includes: Obtaining and parsing each PIN in the PCIe bus connector based on the connected I / O expander to obtain the U-bit identifier of the server; the U-bit identifier is used to indicate the starting U-bit position; Determining the target U-bit position of the belonging server according to the starting U-bit position and the number of U-bits.
[0104] The U-bit identifier and partition identifier of each server node are encoded by an I2C IO expansion device located on the PCIe data bus connector of the cabinet. When the server node is inserted into the cabinet and the BMC in the server is powered on and starts up, the I2C is used to access the IO expansion device to summarize the levels of each PIN, and the U-bit identifier and partition identifier are parsed based on the levels of each PIN.
[0105] See Figure 5 , the embodiment of the present application provides a connection schematic diagram between the server and the PCIe bus connector. The server includes a BMC and a PCIe Device (PCIe device), and the PCIe bus connector includes an IO (input / output) expander (specifically, an inter-integrated circuit communication protocol I2C IO expander) and PCIe PINS (8 PINs). Each PIN corresponds to one bit, and the 8 PINs are b0, b1, b2, b3, b4, b5, b6, and b7 respectively. In practical applications, 1 byte can be used to record the U-bit identifier and partition identifier. It is defined that bit7-bit6 of 1 byte is the partition identifier of the server, and bit5-bit0 is the U-bit identifier of the server in the cabinet.
[0106] Figure 5 In, the BMC in the server is connected to the I2C IO expander on the PCIe bus connector through the clock line SDC and the data line SDA, and the 8 IO PINs after the I2C IO expander are obtained through the I2C protocol and parsed according to the above table to obtain the U-bit identifier and partition identifier of the server itself, while the PCIe device is interconnected and communicates with the remaining servers through other PINs of the connector.
[0107] The foregoing embodiments have illustrated that each partition of the server cabinet includes multiple servers. The partition identifier of 00 indicates belonging to the first identifier, and the partition identifier of 01 indicates belonging to the second partition. Computing servers (represented by host), memory pooling servers (CXL Memory Box), heterogeneous acceleration pooling servers (GPU Box), etc. Each device in the partition has its target U-bit identifier and partition identifier, and its target U-bit position can be determined according to the U-bit identifier and the U-bit number.
[0108] Of course, in addition to each server, the server cabinet also has various switches (such as the aforementioned TOR Host service switch, TOR RDMA switch, TOR management network switch), power management devices (PowerShelf), etc., all of which have corresponding target U-bit positions.
[0109] Referring to Table 1, the embodiment of the present application provides the U positions occupied by each node inside the server cabinet (the highest position of the occupied U position is the target U position of the node), the node number, the partition identifier and the U position identifier:
[0110] Table 1 BMC can obtain and parse the U-bit identification of the corresponding server based on the connected IO expander, and can determine the target U-bit position of the server by comparing the starting U-bit position and the number of U bits. For example, in Table 1, the U-bit identification is 000001, and the U-bit identification (U-bit starting position) corresponding to the computing server Host 13 in partition 1 is 00 1101, and the corresponding decimal value is 13, which occupies 1 U bit, and the occupied U bit is 13, and its target U bit position is 13+1-1=13.
[0111] Similarly, the U-bit identifier (U-bit starting position) corresponding to the memory pooling server GPU Box 10 in partition 1 is 00 1010, and the corresponding decimal value is 10. It occupies 3 U-bits, and the occupied U-bits are 10, 11, and 12 respectively. The target U-bit position is 10+3-1=12.
[0112] The above table is converted into the following JSON code and stored in the BMC file system of each server node, so that when the server node BMC reads each PIN, it can obtain each target U-bit position according to the partition identifier and U-bit identifier obtained by parsing each PIN.
[0113] The embodiment of the present application obtains and parses the U-bit identifier of the corresponding server based on the connected IO expander; the U-bit identifier is used to indicate the starting U-bit position; the starting U-bit position and the number of U bits are processed to determine the target U-bit position of the server to which it belongs, without the need to use an additional U-bit asset management module, thereby reducing resource management costs.
[0114] In some embodiments, the server information further includes a first network address of the server; after determining the target U-bit position of the corresponding server according to each PIN possessed by the PCIe resource bus unit corresponding to the corresponding server, the method further includes: The first network address is determined based on: Obtaining a second network address of the cabinet management controller; The data of the preset byte position of the second network address of the cabinet management controller is replaced with the target U position, and the first network address of the server is obtained and configured.
[0115] RMC can centrally manage the server nodes in the cabinet, and a network link needs to be established from RMC to the BMC of each node server. After the server BMC reads the U-bit identification label, it automatically configures its own management network address according to the U-bit, and establishes a management local area network in the cabinet.
[0116] When RMC starts up, it automatically configures its second network address (also known as the management network address) in the cabinet management local area network. The BMC of each server node in the cabinet identifies its partition information according to Bit7-Bit6 of the U-bit label, and identifies the U-bit label (denoted as Umin) according to Bit5-Bit0. By reading the number of U-bits occupied by the chassis (denoted as Un) set in the motherboard read-only memory (EEPROM) when the server leaves the factory, it obtains the target U-bit position (denoted as Umax) occupied in the cabinet, and uses this Umax as a variable to configure the management network address of the server node in the cabinet management local area network.
[0117] Specifically, the data of the preset byte bit of the second network address of the cabinet management controller can be replaced with the target U-bit position to obtain the first network address of the affiliated server. For example, the preset byte is the last byte. Assuming that the second network address of the BMC is 169.254.10.100, the first network address and the second network address can be IPv4 addresses or IPv6 addresses, and there is no restriction on this.
[0118] Continuing the embodiment of Table 1, after the first network address is configured for each node, the first network addresses of each node are shown in Table 2 below:
[0119] Table 2 The embodiment of the present application configures the first network address for each server based on the target U-bit position of each server to achieve the purpose of binding the network address to the U-bit position. Subsequently, it can facilitate the administrator to more reasonably plan the network topology, avoid network bottlenecks or traffic conflicts, help improve the network management efficiency of the data center, enhance the fault troubleshooting ability, optimize resource utilization, and improve the overall network performance and reliability.
[0120] In some embodiments, the server cabinet includes at least two partitions; the partition identifier is used to indicate the partition where the server is located; The server information further includes the node number of the server; the mapping relationship between the U-bit label and the partition identifier and the node number of each server is stored in the controller file system of each server; Before defining the target management interface for providing server information, the method further includes: The node number of the server is determined based on the following method: Parse each PIN of the PCIe resource bus unit of the server to obtain the U-bit identifier and partition identifier of the server; Query the node number corresponding to the U-bit identifier and partition identifier of the corresponding server according to the mapping relationship.
[0121] In each embodiment of the present application, each server has a corresponding partition identifier. There is a mapping relationship between the U-bit identifier and partition identifier of the server and the node number. The node number is also called the node identifier. For a server, the node identifier is the server identifier and can uniquely indicate the server.
[0122] This mapping relationship can be stored in the form of Table 1 above. The node code mapped to the U-bit identifier and partition identifier of the server can be queried from Table 1 above, making the server information provided by the target management interface more complete, helping the RMC to more comprehensively understand the information of each server node, and improving network manageability.
[0123] In some embodiments, the server information further includes operation status information and / or chassis asset information; before defining the target management interface for providing server information, the method further includes: Monitor the operation status information of the server; the operation status information includes at least one of the power-on state, power-on / off state, and health state of the server; and / or, Obtain the chassis asset information from the read-only memory of the server. The chassis asset information includes at least one of the server type, model, serial number, and U-bit number.
[0124] It can be understood that the server information in the embodiments of the present application further includes operation status information and / or chassis asset information.
[0125] The operation status information is obtained by the server management controller monitoring the server and is stored in the management memory of the server management controller.
[0126] The chassis asset information belongs to information that cannot be modified and is set in the motherboard read-only memory (EEPROM) when the server leaves the factory. The chassis asset information includes at least one of the server type, model, serial number, and U-bit number (representing the height of the occupied U-bit).
[0127] Of course, other information can also be included, such as the production batch number, etc.
[0128] Referring to Table 3 below, the embodiments of the present application show various server information corresponding to each field of the target management interface. Each field represents a type of server information, and each field has a corresponding interpretation and source:
[0129] Table 3 The target management interface of the server management controller also provides the operating status information of the server and / or the chassis asset information, thereby enabling the target management interface to provide more abundant server information and enabling the cabinet management controller to obtain server information more comprehensively.
[0130] As can be seen from the above, the BMC can obtain server information including the target U-position of the server, the first network address, the node number, the operating status information, and the chassis asset management information, etc. The BMC can summarize this information in the target management interface (such as redfish). Thus, the RMC can obtain more comprehensive server information from the target management interface of the BMC, greatly reducing the complexity of management and enabling the system to be flexibly adjusted and easily managed.
[0131] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation manner.
[0132] Figure 6 One of the block diagrams of the computer program product provided according to the embodiments of the present application. This computer program product can be applied to the cabinet management controller of a server cabinet; the server cabinet includes multiple servers; each server has a variety of target buses with a unified standard; the various target buses of each server converge on the PCIe resource bus unit of the server; Such as Figure 6 shown, this computer program product includes: a first communication module 610, an acquisition module 620, and a management module 630; The first communication module 610 is used to periodically access the server management controller of each server included in the server cabinet when powered on; the server management controller defines a target management interface for providing server information; The acquisition module 620 is used to obtain the server information of each server from the target management interface of each server management controller; the server information includes the target U-position of the corresponding server; the target U-position is determined by each server management controller according to each PIN of the corresponding PCIe resource bus unit of the corresponding server; each PIN is used to indicate the partition identifier and the U-position identifier of the server; The management module 630 is used to manage each server based on the server information of each server.
[0133] In some embodiments, this computer program product further includes: A building module, configured to control an automated network building tool to build a local area network inside a server cabinet based on a first network address in each server information, so that each server conducts network communication under the local area network.
[0134] In some embodiments, multiple target buses include a PCIe bus; the server cabinet includes multiple partitions; each server in each partition includes a memory pooling server, multiple computing servers, and multiple heterogeneous acceleration pooling servers; Multiple computing servers are connected to the memory pooling server through the PCIe bus to share the memory resources of the memory pooling server; Multiple computing servers are connected to multiple heterogeneous acceleration pooling servers through the PCIe bus to share the acceleration computing resources in the multiple heterogeneous acceleration pooling servers.
[0135] In some embodiments, multiple target buses include a network bus; the server cabinet includes a first switch, a second switch, and a third switch; The network bus includes a service data network bus, a Remote Direct Memory Access (RDMA) data network bus, and an out-of-band management network bus; Multiple computing servers are connected to the first switch through the service data network bus to provide computing resources to external devices of the multiple computing servers through the first switch; Multiple heterogeneous acceleration pooling servers are connected to the second switch through the RDMA data network bus to provide acceleration resources to external devices of the multiple heterogeneous acceleration pooling servers through the second switch; The cabinet management controller and each server management controller are connected to the third switch through the out-of-band management network bus to centrally control and manage each server through the third switch.
[0136] In some embodiments, multiple target buses include a water cooling bus; each server includes a liquid cooling pipeline; The water cooling bus is respectively connected to each liquid cooling pipeline and an external water cooling unit; the water cooling bus is used to transport the coolant of the external water cooling unit to the liquid cooling pipeline to dissipate heat from each server.
[0137] In some embodiments, multiple target buses include a power supply bus; the server cabinet includes a power management device; Each server in the server cabinet is connected through the power supply bus and the power management device to obtain the power allocated by the power management device through the power supply bus.
[0138] In some embodiments, the power management device includes two power racks; The cabinet management controller is located in any one of the power racks.
[0139] In some embodiments, the computer program product further includes: A control module, configured to: Based on a preset power-on sequence, determine a target server to be powered on currently, and send a power-on request to a target management interface of a server management controller of the target server, so as to power on the target server through the server management controller; After a preset duration from sending the power-on request, obtain the power-on status of the target server from the target management interface of the target server. When the power-on status is the powered-on status and there are unpowered servers, determine the next target server to be powered on based on the power-on sequence.
[0140] Figure 7 It is the second block diagram of the computer program product provided according to the embodiments of the present application. This computer program product can be applied to the server management controllers of each server in a server cabinet; each server included in the server cabinet has a variety of target buses with a unified standard; the variety of target buses of each server converge on the PCIe resource bus unit of the server; As Figure 7 shown, this computer program product includes: a definition module 710 and a second communication module 720; The definition module 710 is configured to define a target management interface for providing server information when powered on; The second communication module 720 is configured to, when accessed by a cabinet management controller of the server cabinet, provide server information to the cabinet management controller based on the target management interface, so that the cabinet management controller manages the server based on the server information obtained from the target management interface; The server information includes a target U-bit position; the target U-bit position is determined based on the following method: Determine the target U-bit position of the corresponding server according to each PIN of the PCIe resource bus unit of the corresponding server; each PIN is used to indicate the partition identifier and the U-bit identifier of the server.
[0141] In some embodiments, the server information further includes operation status information and / or chassis asset information; this computer program product further includes: A monitoring module, configured to monitor the operation status information of the server; the operation status information includes at least one of the server power-on status, the power-on / off status, and the health status; and / or A chassis asset acquisition module, configured to acquire chassis asset information from the read-only memory of the server, and the chassis asset information includes at least one of the server type, model, serial number, and number of U-bits.
[0142] In some embodiments, the PCIe resource bus unit includes a PCIe bus connector; the PCIe bus connector includes an IO expander and each PIN, and each server management controller is connected to the IO expander corresponding to the PCIe bus connector; The definition module 710 is specifically used for: Based on the connected IO expander, each PIN in the PCIe bus connector is obtained and parsed to obtain the U-bit identifier of the server; the U-bit identifier is used to indicate the starting U-bit position; According to the starting U-bit position and the number of U bits, the target U-bit position of the server is determined.
[0143] In some embodiments, the computer program product comprises: The first processing module is used for: The first network address is determined based on: Obtaining a second network address of the cabinet management controller; The data of the preset byte position of the second network address of the cabinet management controller is replaced with the target U position, and the first network address of the server is obtained and configured.
[0144] In some embodiments, the server cabinet includes at least two partitions; the partition identifier is used to indicate the partition where the server is located; the server information also includes the node number of the server; the controller file system of each server stores the mapping relationship between the U-bit identifier of each server and the partition identifier and the node number; The computer program product also includes: The second processing module is used for: The node ID of the server is determined based on: Parse each PIN of the PCIe resource bus unit of the server to obtain the U bit identifier and partition identifier of the server; According to the mapping relationship, query the node number corresponding to the U bit identifier and the partition identifier of the corresponding server.
[0145] For the description of the features in the embodiments corresponding to the computer program product, reference can be made to the relevant description of the embodiments corresponding to the server management method, which will not be repeated here.
[0146] Figure 8 This is one of the schematic diagrams of the interaction process between the cabinet management controller and the server management modules of each server provided according to an embodiment of the present application.
[0147] The server cabinet includes a cabinet management controller and multiple servers; each server has multiple target buses of unified standards and a server management controller; the multiple target buses of each server converge on the PCIe resource bus unit of the server; see Figure 8, the interaction process includes the following steps 810 to 830: Step 810, when the server management controller is powered on, define a target management interface for providing server information; the server information includes the target U-bit position of the corresponding server; the target U-bit position is determined by each server management controller according to each PIN of the corresponding PCIe resource bus unit of the corresponding server; each PIN is used to indicate the partition identifier and U-bit identifier of the server; Step 820, when the cabinet management controller is powered on, periodically access the server management controllers of each server; Step 830, the cabinet management controller obtains the server information of each server from the target management interface of each server management controller, and manages each server based on the server information of each server.
[0148] For the detailed implementation of the above process, see the foregoing embodiments and will not be elaborated here.
[0149] Through the server management method provided by this application, multiple target buses of each server with a unified standard converge on the PCIe resource bus unit of the server. Each server management controller determines the target U-bit position of the server according to each PIN of the corresponding PCIe resource bus unit of the corresponding server. The target U-bit position belongs to important server information. The server management controller defines a target management interface for providing server information. The cabinet management controller obtains the information of each server by accessing the target management interface, further reducing the cable structure in the cabinet, and can realize the identification and unified management of the U-bit assets, network addresses and other information of each server node, reducing the management cost.
[0150] See Figure 9 , the embodiment of this application provides a schematic diagram of the interaction process between the cabinet management controller in a server cabinet and the server management controllers of each server, including the following steps: For any server management controller: The server management controller completes its own power-on startup; The server management controller reads each PIN of the PCIe resource bus unit to determine the partition identifier and U-bit identifier of the server; The server management controller determines the target U-bit position of the corresponding server according to the U-bit identifier and the number of U-bits; The server management controller configures the first network address of the server according to the target U-bit position of the corresponding server; The server management controller determines, from the BMC file system, that the node number having a mapping relationship with the partition identifier and the U-bit identifier of the server is the node number of the server; The server management controller obtains the operating status information; the operating status information includes at least one of the server power-on status, the power-on / off status, and the health status; The server management controller reads the read-only memory to obtain the chassis asset information, and the chassis asset information includes at least one of the server type, model, serial number, and U-bit number; The server management controller externally provides, through the redfish interface, server information including at least some of the target U-bit position, the first network address, the operating status information, the node number, and the chassis asset information; The cabinet management controller completes its own power-on startup; The cabinet management controller obtains the server information of each server through the target management interface of each server management controller via the network bus, and aggregates the server information of each server; The cabinet management controller automatically incorporates each server node into management according to the server information of each server; The cabinet management controller generates a centralized management topology according to the server information of each server.
[0151] For the detailed implementation process of the above process, refer to the foregoing embodiments, and details will not be elaborated herein.
[0152] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the foregoing embodiments of the server management method.
[0153] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the foregoing embodiments of the server management method when running.
[0154] In an exemplary embodiment, the foregoing computer-readable storage medium may include, but is not limited to: a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc, and other various media that can store computer programs.
[0155] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the steps in any one of the foregoing embodiments of the server management method.
[0156] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered as exceeding the scope of this application.
[0157] The above has introduced in detail a server management method, a computer program product, a server cabinet, an electronic device, and a computer-readable storage medium provided by this application. Specific examples are used herein to elaborate on the principles and implementation manners of this application. The descriptions of the above embodiments are only applicable to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A management method for a server, characterized in that, Cabinet management controller applied to a server cabinet; the server cabinet includes a plurality of servers; Each server has a variety of target buses with unified standards; The variety of target buses of each server converge on the PCIe resource bus unit of the server; the method includes: When powered on, periodically access the server management controller of each server; the server management controller defines a target management interface for providing server information; Obtain the server information of each server from the target management interface of each server management controller; the server information includes the target U-bit position of the corresponding server; the target U-bit position is determined by each server management controller according to each PIN of the corresponding PCIe resource bus unit of the corresponding server; each of the PINs is used to indicate the partition identifier and the U-bit identifier of the server; Manage each server based on the server information of each server.
2. The management method of the server according to claim 1, wherein The server information further includes the first network address of the server; the first network address is determined by the server management controller according to the target U-bit position of the corresponding server; after obtaining the server information of each server from the target management interface of each server management controller, the method further includes: Control the automated network construction tool to construct the local area network inside the server cabinet based on the first network address in the server information of each server, so that each server can perform network communication under the local area network.
3. The management method of the server according to claim 1, characterized in that, The variety of target buses includes PCIe buses; The server cabinet includes a plurality of partitions; each server in each partition includes a memory pooling server, a plurality of computing servers, and a plurality of heterogeneous acceleration pooling servers; The plurality of computing servers are connected to the memory pooling server through the PCIe bus to share the memory resources of the memory pooling server; The plurality of computing servers are connected to the plurality of heterogeneous acceleration pooling servers through the PCIe bus to share the acceleration computing resources in the plurality of heterogeneous acceleration pooling servers.
4. The management method of the server according to claim 3, characterized in that, The variety of target buses includes network buses; the server cabinet includes a first switch, a second switch, and a third switch; The network buses include a service data network bus, a remote direct memory access (RDMA) data network bus, and an out-of-band management network bus; The plurality of computing servers are connected to the first switch through the service data network bus to provide computing resources to the external devices of the plurality of computing servers through the first switch; The plurality of heterogeneous acceleration pooling servers are connected to the second switch through the RDMA data network bus to provide acceleration resources to the external devices of the plurality of heterogeneous acceleration pooling servers through the second switch; The cabinet management controller and each server management controller are connected to the third switch through the out-of-band management network bus to centrally control and manage each server through the third switch.
5. The management method of the server according to claim 1, characterized in that, The variety of target buses includes a water cooling bus; each server includes a liquid cooling pipeline; The water-cooling bus is respectively connected to each of the liquid-cooling pipelines and an external water-cooling unit; the water-cooling bus is used to transport the coolant of the external water-cooling unit to the liquid-cooling pipelines to dissipate heat from each server.
6. The management method of the server according to claim 1, wherein, The multiple target buses include a power supply bus; the server cabinet includes a power management device; Each server in the server cabinet is connected through the power supply bus and the power management device to obtain the power allocated by the power management device through the power supply bus.
7. The management method of the server according to claim 6, wherein The power management device includes two power racks; The cabinet management controller is located in any one of the power racks.
8. The management method of the server according to any one of claims 1-7, characterized in that, The method further includes: Based on a preset power-on sequence, determining a target server to be powered on currently, and sending a power-on request to a target management interface of the server management controller of the target server, so as to power on the target server through the server management controller; After a preset time period after sending the power-on request, obtaining the power-on status of the target server from the target management interface of the target server, and when the power-on status is the powered-on status and there are unpowered servers, determining the next target server to be powered on based on the power-on sequence.
9. A management method for a server, characterized in that, Applied to the server management controllers of each server in a server cabinet; each server included in the server cabinet has multiple target buses with a unified standard; The multiple target buses of each server converge on the PCIe resource bus unit of the server; the method includes: When powered on, defining a target management interface for providing server information; When accessed by the cabinet management controller of the server cabinet, providing the server information to the cabinet management controller based on the target management interface, so that the cabinet management controller manages the server based on the server information obtained from the target management interface; The server information includes a target U-position; the target U-position is determined based on the following method: Determining the target U-position of the corresponding server according to each PIN included in the PCIe resource bus unit of the corresponding server; each PIN is used to indicate the partition identifier and the U-position identifier of the server.
10. The management method of the server according to claim 9, characterized in that, The server information further includes operation status information and / or chassis asset information; Before defining the target management interface for providing server information, the method further includes: Monitoring the operation status information of the server; the operation status information includes at least one of the server power-on status, the power-on / off status, and the health status; and / or, Obtaining chassis asset information from the read-only memory of the server, where the chassis asset information includes at least one of the server type, model, serial number, and the number of U-units.
11. The management method of the server according to claim 10, wherein The PCIe resource bus unit includes a PCIe bus connector; the PCIe bus connector includes an IO expander and each PIN, and each server management controller is connected to the IO expander of the corresponding PCIe bus connector; The determining the target U-position of the corresponding server according to each PIN included in the corresponding PCIe resource bus unit of the corresponding server includes: Obtain and parse each PIN in the PCIe bus connector based on the connected IO expander to obtain the U-bit identifier of the server; the U-bit identifier is used to indicate the starting U-bit position; Determine the target U-bit position of the corresponding server according to the starting U-bit position and the number of U-bits.
12. The management method of the server according to claim 9, wherein The server information further includes the first network address of the server; after determining the target U-bit position of the corresponding server according to each PIN of the corresponding PCIe resource bus unit of the corresponding server, the method further includes: Determine the first network address based on the following method: Obtain the second network address of the chassis management controller; Replace the data of the preset byte position of the second network address of the chassis management controller with the target U-bit position to obtain and configure the first network address of the server.
13. The management method of the server according to claim 9, characterized in that, The server chassis includes at least two partitions; the partition identifier is used to indicate the partition where the server is located; the server information further includes the node number of the server; the mapping relationship between the U-bit identifier and the partition identifier and the node number of each server is stored in the controller file system of each server; Before defining the target management interface for providing server information, the method further includes: Determine the node number of the server based on the following method: Parse each PIN of the PCIe resource bus unit of the server to obtain the U-bit identifier and the partition identifier of the server; Query the node number corresponding to the U-bit identifier and the partition identifier of the corresponding server according to the mapping relationship.
14. A computer program product, characterized in that, Applied to the chassis management controller of the server chassis; the server chassis includes multiple servers; Each server has a variety of target buses with a unified standard; The various target buses of each server converge on the PCIe resource bus unit of the server; The computer program product includes: A first communication module, configured to periodically access the server management controllers of each server included in the server chassis when powered on; the server management controller defines a target management interface for providing server information; An acquisition module, configured to obtain the server information of each server from the target management interface of each server management controller; the server information includes the target U-bit position of the corresponding server; the target U-bit position is determined by each server management controller according to each PIN of the corresponding PCIe resource bus unit of the corresponding server; the respective PINs are used to indicate the partition identifier and the U-bit identifier of the server; A management module, configured to manage each server based on the server information of each server.
15. A computer program product, characterized in that, Applied to the server management controller of each server in the server chassis; each server included in the server chassis has a variety of target buses with a unified standard; The various target buses of each server converge on the PCIe resource bus unit of the server; The computer program product includes: A definition module, configured to define a target management interface for providing server information when powered on; A second communication module, configured to, when accessed by a cabinet management controller of the server cabinet, provide the server information to the cabinet management controller based on the target management interface, so that the cabinet management controller manages the server based on the server information obtained from the target management interface; The server information includes a target U-position; the target U-position is determined based on the following method: Determine the target U-position of the corresponding server according to each PIN of the PCIe resource bus unit of the corresponding server; each PIN is used to indicate the partition identifier and the U-position identifier of the server.
16. A server cabinet, characterized in that, The server cabinet includes a cabinet management controller and multiple servers; each server has multiple target buses and a server management controller with a unified standard; The multiple target buses of each server converge on the PCIe resource bus unit of the server; When powered on, the server management controller defines a target management interface for providing server information; The server information includes the target U-position of the corresponding server; the target U-position is determined by each server management controller according to each PIN of the corresponding PCIe resource bus unit of the corresponding server; each PIN is used to indicate the partition identifier and the U-position identifier of the server; When powered on, the cabinet management controller periodically accesses the server management controllers of each server; Obtain the server information of each server from the target management interface of each server management controller, and manage each server based on the server information of each server.
17. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to implement the steps of the server management method according to any one of claims 1-8 or 9-13 when executing the computer program.
18. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the server management method according to any one of claims 1-8 or 9-13 when executed by a processor.
Citation Information
Patent Citations
Data center cabinet U bit management device and method
CN116644771A
Server asset information reporting method, device, equipment and medium
CN119356975A
Server cabinet and server system
CN119917442A
Server, server asset information acquisition method and apparatus, and server asset information providing method and apparatus
US20250156364A1
Cited By
System, server and method for sharing memory resource pool, medium and program product
CN120448141A