Server management method, computer program product and server cabinet

By setting the PIN pin in the PCIe resource bus unit of the server and defining the target management interface, the increased cable and management cost problems in the prior art are solved, and unified management and cost reduction of server information are achieved.

CN120196513BActive Publication Date: 2025-08-08INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510667795.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-08
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

The prior art increases the cable structure and management costs in the cabinet by additionally connecting the U-bit asset management module to monitor the server information inside the server cabinet.

Method used

By setting the PIN pin in the PCIe resource bus unit of the server, it is used to indicate the partition identifier and the U-bit identifier. The server management controller defines the target management interface, and the cabinet management controller accesses the interface to obtain server information, realizing unified management of the server.

Benefits of technology

It reduces the cable structure in the cabinet, reduces management costs, and realizes the identification and unified management of U-bit assets, network addresses and other information of server nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196513B_ABST
    Figure CN120196513B_ABST
Patent Text Reader

Abstract

The present application discloses a server management method, a computer program product, and a server cabinet, which relate to the field of server technology, including: since the server has multiple target buses of unified standards, and the multiple target buses converge on the PCIe resource bus unit of the server, each server management controller determines the target U-bit position of the server according to each PIN possessed by the PCIe resource bus unit of the corresponding server, and the target U-bit position belongs to the server information. The server management controller defines a target management interface for providing the server information, and the cabinet management controller obtains the information of each server by accessing the target management interface. Therefore, it can solve the technical problem of related technologies that monitor server information by additionally connecting a U-bit asset management module, which increases the cable structure and management cost in the cabinet, and can further reduce the cable structure in the cabinet, realize the identification and unified management of each server, and reduce the management cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of server technology, and in particular to a server management method, computer program product, and server cabinet. Background Art

[0002] The converged architecture server cabinet (referred to as the server cabinet) modularizes the server cable interfaces into blind-plug connectors, allowing the cabinet's centralized power bus, network bus, and liquid cooling bus to form bus interface modules, greatly simplifying cabling within the cabinet. However, due to business requirements, the server cabinet also needed to add a PCIe resource bus unit, forming a multi-bus converged architecture cabinet with a centralized power bus, network bus, liquid cooling bus, and PCIe resource bus.

[0003] Related technologies monitor server information inside a server cabinet by additionally connecting a cabinet management controller to a U-position asset management module, which increases the cable structure and management costs inside the cabinet. Summary of the Invention

[0004] The present application provides a server management method, a computer program product, a server cabinet, an electronic device, and a computer-readable storage medium to at least solve the problem in the related art of monitoring server information inside a server cabinet through a U-position asset management module, which increases the cable structure and management costs within the cabinet.

[0005] The present application provides a server management method, which is applied to a cabinet management controller of a server cabinet; the server cabinet includes multiple servers; each server has multiple target buses of a unified standard; the multiple target buses of each server converge on a PCIe resource bus unit of the server; the method includes:

[0006] When powered on, the server management controller of each server is periodically accessed; the server management controller defines a target management interface for providing server information;

[0007] Obtaining server information of each server from a target management interface of each server management controller; the server information includes a target U-bit position of the corresponding server; the target U-bit position is determined by each server management controller based on each PIN possessed by a PCIe resource bus unit corresponding to the corresponding server; each PIN is used to indicate a partition identifier and a U-bit identifier of the server;

[0008] Manage each server based on its server information.

[0009] The present application also provides a server management method, which is applied to a server management controller of each server in a server cabinet; each server contained in the server cabinet has multiple target buses of a unified standard; the multiple target buses of each server converge on a PCIe resource bus unit of the server; the method comprises:

[0010] In the case of power on, define the target management interface for providing server information;

[0011] When accessed by a cabinet management controller of a server cabinet, providing server information to the cabinet management controller based on a target management interface, so that the cabinet management controller manages the server based on the server information obtained from the target management interface;

[0012] The server information includes the target U-bit location; the target U-bit location is determined based on the following:

[0013] The target U-bit position of the corresponding server is determined according to each PIN possessed by the PCIe resource bus unit of the corresponding server; each PIN is used to indicate the partition identifier and U-bit identifier of the server.

[0014] The present application also provides a computer program product, which is applied to a rack management controller of a server rack; the server rack includes multiple servers; each server has multiple target buses of a unified standard; the multiple target buses of each server converge on a PCIe resource bus unit of the server; the computer program product includes:

[0015] A first communication module is configured to periodically access a server management controller of each server contained in the server cabinet when the server cabinet is powered on; the server management controller defines a target management interface for providing server information;

[0016] An acquisition module is configured to acquire server information of each server from a target management interface of each server management controller; the server information includes a target U-bit position of the corresponding server; the target U-bit position is determined by each server management controller based on each PIN possessed by a PCIe resource bus unit corresponding to the corresponding server; each PIN is used to indicate a partition identifier and a U-bit identifier of the server;

[0017] The management module is used to manage each server based on the server information of each server.

[0018] The present application also provides a computer program product, which is applied to a server management controller of each server in a server cabinet; each server contained in the server cabinet has multiple target buses of a unified standard; the multiple target buses of each server converge on a PCIe resource bus unit of the server; the computer program product includes:

[0019] A definition module, used for defining a target management interface for providing server information when the power is on;

[0020] a second communication module, configured to, when accessed by a cabinet management controller of the server cabinet, provide server information to the cabinet management controller based on a target management interface, so that the cabinet management controller manages the server based on the server information obtained from the target management interface;

[0021] The server information includes the target U-bit location; the target U-bit location is determined based on the following:

[0022] The target U-bit position of the corresponding server is determined according to each PIN possessed by the PCIe resource bus unit of the corresponding server; each PIN is used to indicate the partition identifier and U-bit identifier of the server.

[0023] The present application also provides a server cabinet, comprising a cabinet management controller and multiple servers; each server has multiple target buses of a unified standard and a server management controller; the multiple target buses of each server converge on a PCIe resource bus unit of the server;

[0024] When powered on, the server management controller defines a target management interface for providing server information; the server information includes a target U-bit position for the corresponding server; the target U-bit position is determined by each server management controller based on the respective PINs of the PCIe resource bus unit corresponding to the corresponding server; each PIN is used to indicate the partition identifier and U-bit identifier of the server;

[0025] When powered on, the rack management controller periodically accesses the server management controller of each server, obtains server information of each server from the target management interface of each server management controller, and manages each server based on the server information of each server.

[0026] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned server management methods when executing the computer program.

[0027] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned server management methods are implemented.

[0028] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned server management methods when executed by a processor.

[0029] Through this application, since the server has multiple target buses with unified standards, and the multiple target buses converge on the PCIe resource bus unit of the server, each server management controller determines the target U-bit position of the server according to the various PINs possessed by the corresponding PCIe resource bus unit of the corresponding server. The target U-bit position belongs to the server information. The server management controller defines a target management interface for providing server information. The cabinet management controller obtains the information of each server by accessing the target management interface. Therefore, the technical problem of the related technology that the server information inside the server cabinet is monitored by additionally connecting the cabinet management controller to the U-bit asset management module, which increases the cable structure and management cost in the cabinet, can be solved. This application can further reduce the cable structure in the cabinet, and can realize the identification and unified management of the U-bit assets, network addresses and other server information of each server node, thereby reducing management costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0031] Figure 1 One of the flow charts of a server management method provided in an embodiment of the present application;

[0032] Figure 2 A schematic diagram of the structure of a server cabinet provided in an embodiment of the present application;

[0033] Figure 3 A schematic diagram of the partition architecture in a server cabinet provided in an embodiment of the present application;

[0034] Figure 4 A second flow chart of a server management method provided in an embodiment of the present application;

[0035] Figure 5 A schematic diagram of the connection between the server and the PCIe bus connector provided in an embodiment of the present application;

[0036] Figure 6 One of the block diagrams of the computer program product provided in the embodiment of the present application;

[0037] Figure 7 A second block diagram of a computer program product provided in an embodiment of the present application;

[0038] Figure 8This is a schematic diagram of the interaction flow between the rack management controller and the server management modules of each server provided in an embodiment of the present application;

[0039] Figure 9 This is a second schematic diagram of the interaction flow between the cabinet management controller and the server management modules of each server provided in an embodiment of the present application. DETAILED DESCRIPTION

[0040] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0041] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0042] To meet the flexible resource allocation needs of diverse application scenarios and requirements, such as artificial intelligence, machine learning, and intelligent computing, data centers are accelerating their transition from a compute-centric architecture to a data-centric converged architecture. Converged architecture systems physically centralize similar resources, forming various types of servers, such as general-purpose computing resource pools, PCIe resource exchange servers or units, and heterogeneous acceleration resource pools. Decoupling these resources into separate, independent pooled server chassis facilitates centralized resource management and expansion. However, each pooled server node still needs to be connected to the PCIe resource exchange server via cables to enable flexible resource allocation. This results in large-scale interconnect cables at the cabinet level, complex wiring, and redundant structures, which pose an invisible obstacle to deployment and operation.

[0043] Some open data center cabinet design standards, such as OCP / Open19 / ODCC / OCTC, modularize the server cable interfaces into blind-plug connectors, allowing the centralized power bus, network bus, and liquid cooling bus within the cabinet to form bus interface modules. This greatly simplifies the cabling within the cabinet. Such blind-plug servers have already achieved deployment scale and a customized ecosystem, significantly improving cost and energy efficiency. For the entire converged architecture server cabinet, it is also necessary to add a PCIe resource bus module. This allows independent pooled resource servers to converge to a PCIe resource exchange server or unit via high-speed blind-plug connectors and PCIe bus modules, forming a four-bus converged architecture cabinet (hereinafter referred to as the server cabinet) with a centralized power bus, network bus, liquid cooling bus, and PCIe resource bus. With the PCIe resource exchange unit as the center, a flexible and configurable converged architecture server system is constructed.

[0044] Although the pooled server nodes within the server cabinet are highly decoupled, they need to operate as a single server system and facilitate fault location and operation and maintenance management. Therefore, an efficient solution is needed to identify and uniformly manage the location, server type, management address, and basic information of each pooled server node within the cabinet. In traditional server cabinets, since each server node is relatively independent and lacks a unified interconnection interface, automatic real-time monitoring of the cabinet space status, the model of the server node within the cabinet, the number of U-bits occupied, and the specific location in the rack are usually achieved by using RMC to additionally connect the U-bit asset management module to manage server information, which increases the cabling structure and management costs within the cabinet.

[0045] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0046] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the server management method depends, the specific application environment architecture or specific hardware architecture is described herein.

[0047] An embodiment of the present application provides a server management method, and the method is described in detail in conjunction with the execution process of the server management method.

[0048] The server management method provided in the embodiment of the present application may be executed by a rack management controller (RMC) of a server rack.

[0049] In practice, RMC is typically used in data centers, computer rooms, or server cabinets to monitor and manage the hardware within, including servers, storage devices, and network equipment. Its core functions include remote hardware management, monitoring, troubleshooting, and resource allocation. RMC often integrates advanced management features, such as power management, temperature control, and hardware health monitoring.

[0050] The server cabinet includes multiple servers; each server has multiple target buses with unified standards;

[0051] The embodiment of the present application can specifically be a converged architecture server cabinet, which includes multiple target buses, including but not limited to a power supply bus, a network bus, a liquid cooling bus, a resource bus, etc., and can also include other buses, such as an address bus, an expansion bus, a system bus, a serial bus, and a parallel bus, etc.

[0052] The various target buses of each server converge into the PCIe resource bus unit of the server.

[0053] The PCIe resource bus unit (RBU) is a key component in PCIe systems responsible for managing and allocating bus resources. It provides functions such as resource allocation, device addressing, bandwidth management, and configuration management, ensuring that multiple devices can operate efficiently and stably while sharing the bus.

[0054] Multiple target buses converge on the server's PCIe resource bus unit. Each server service can manage each bus of the server cabinet without having to connect to the U-bit asset management module through the server management controller. This can effectively improve the efficiency of resource management and simplify the architecture and management of the server cabinet.

[0055] See also Figure 1 , an embodiment of the present application provides a flow chart of a server management method, including the following steps 110, 120 and 130:

[0056] Step 110: When the system is powered on, periodically access the server management controller of each server;

[0057] Step 120: Obtain server information of each server from a target management interface of each server management controller. The server management controller defines a target management interface for providing server information. The server information includes a target U-bit position of the corresponding server. The target U-bit position is determined by each server management controller based on respective PINs of a PCIe resource bus unit corresponding to the corresponding server. Each PIN indicates a partition identifier and a U-bit identifier of the server.

[0058] Step 130: Manage each server based on the server information of each server.

[0059] The server cabinet in the embodiment of the present application includes multiple partitions, each partition has multiple servers, and each partition includes multiple (for example, 4) computing servers, one memory pooling server and multiple (for example, 4) heterogeneous acceleration pooling servers.

[0060] In order to achieve overall scheduling of acceleration computers (such as GPUs) in multiple heterogeneous acceleration pooled servers, the cabinet management controller in the server cabinet needs to have detailed knowledge of the server information of each server. The server information includes at least part of the server type, model, serial number, U-bit number, target U-bit address, U-bit identifier (starting U-bit address), and the first network address.

[0061] The cabinet server of the server may periodically obtain server information of each server from the server management controller of each server.

[0062] The Server Management Controller (SMC) is a critical tool for ensuring healthy server operations. Through remote control, hardware monitoring, and logging, it helps administrators maintain complete control of their servers at all times. It is particularly crucial for data center operations, reducing server downtime and improving server availability and maintenance efficiency.

[0063] The operating status of the server management controller is not highly correlated with the corresponding server and can be considered independent of each other. The server management controller itself has an IP address and can communicate directly with the cabinet management controller.

[0064] In practical applications, the server management controller may be specifically a Baseboard Management Controller (BMC). The BMC is a dedicated hardware management controller in the server, which is usually integrated on the server's motherboard and is used to provide hardware monitoring, management, and remote control functions.

[0065] Each server has a PCIe resource bus unit. To reduce server management costs, this embodiment of the application provides various PIN pins in the PCIe resource bus unit. Each PIN is used to indicate the server's partition identifier and U-bit identifier. The partition identifier indicates the partition in which the server is located, and the U-bit identifier indicates the server's starting U-bit position.

[0066] The server management controller reads each PIN in each PCIe resource bus unit to obtain the partition identifier and U-bit identifier of its corresponding server. Based on the partition identifier and U-bit identifier of the server and the number of U bits of the server, the target U-bit position of the server can be determined. The target U-bit position is used to indicate the highest U-bit position of the server.

[0067] Next, each server management controller configures a first network address for each server according to the target U-bit position of the corresponding server. Based on the first network address of each server, a local area network inside the server can be constructed, and the servers inside the server cabinet can perform network communication within the local area network.

[0068] In order to enable the cabinet management controller to obtain server information including the target U-bit location, the server management controller defines a target management interface based on the server information of the corresponding server. The target management interface is a software interface. The BMC can expose the server information of the server through the target management interface. The server information belongs to the hardware resources, that is, the BMC exposes the hardware resources through the software interface.

[0069] The cabinet management controller can obtain the server information of each server by directly accessing the target management interface of each server management controller.

[0070] The target management interface may be a Redfish interface, which is a standard interface based on a RESTful API and is intended to provide more modern, easily extensible, and automated management functions for server management.

[0071] After obtaining the server information of each server, the cabinet management controller can manage each server, generate a management topology map, and provide computing services, memory services, etc. based on the management map containing the information of each server.

[0072] Through the server management method provided by the present application, since the server has multiple target buses of unified standards, and the multiple target buses converge on the PCIe resource bus unit of the server, each server management controller determines the target U-bit position of the server according to the various PINs possessed by the corresponding PCIe resource bus unit of the corresponding server. The target U-bit position belongs to the server information. The server management controller defines a target management interface for providing server information. The cabinet management controller obtains the information of each server by accessing the target management interface. Therefore, it can solve the technical problem of the related technology that the server information inside the server cabinet is monitored by additionally connecting the cabinet management controller to the U-bit asset management module, which increases the cable structure and management cost in the cabinet. The present application can further reduce the cable structure in the cabinet, and can realize the identification and unified management of the U-bit assets, network addresses and other server information of each server node, thereby reducing management costs.

[0073] In some embodiments, the server information further includes a first network address of the server; the first network address is determined by the server management controller according to a target U-bit position of the corresponding server; after obtaining the server information of each server from the target management interface of each server management controller, the method further includes:

[0074] The control automation network building tool builds a local area network inside the server cabinet based on the first network address in the information of each server, so that each server can perform network communication under the local area network.

[0075] After obtaining the first network address of each server, the embodiment of the present application can use an automated network construction tool to build a local area network within the server based on the first network address of each server and the network address of each switch in the server, so that each server can communicate over the local area network.

[0076] The automated network construction tool includes but is not limited to at least one of Ansible, Chef, NetBox, etc.

[0077] The embodiment of the present application builds a local area network inside the server cabinet based on the first network address of each server, which can significantly improve the performance, security, manageability, and scalability of the network inside the server cabinet and has high cost-effectiveness.

[0078] In some embodiments, the multiple target buses include a PCIe bus; the server cabinet includes multiple partitions; each server in each partition includes a memory pooling server, multiple computing servers, and multiple heterogeneous acceleration pooling servers;

[0079] Multiple computing servers are connected to the memory pooling server via the PCIe bus to share the memory resources of the memory pooling server;

[0080] Multiple computing servers are connected to multiple heterogeneous acceleration pool servers through a PCIe bus to share accelerated computing resources in the multiple heterogeneous acceleration pool servers.

[0081] The server cabinet of the embodiment of the present application includes multiple partitions, each of which includes a memory pooling server, multiple computing servers, and multiple heterogeneous acceleration pooling servers.

[0082] A memory pooling server is a server system that uses memory pooling technology to optimize memory management. By pre-allocating a large memory area (a memory pool), and then allocating and reclaiming memory blocks from it, memory pooling avoids frequent memory allocation and release operations, thereby reducing memory fragmentation and improving memory efficiency. This technology is particularly suitable for scenarios with high concurrency and real-time performance requirements.

[0083] The memory pooling server may specifically be a CXL (Compute Express Link) memory pooling server. The CXL memory pooling server is a new server architecture that uses CXL technology to implement memory pooling.

[0084] Computing servers refer to servers specially designed to handle computing-intensive tasks. Computing servers mainly focus on providing efficient computing capabilities for performing large-scale data processing, computing tasks, virtualization, simulation, analysis and other tasks.

[0085] A heterogeneous acceleration pooled server is one that integrates multiple computing accelerators (such as CPUs, GPUs, FPGAs, and TPUs) and manages and schedules resources through a pooled approach. It utilizes different types of hardware accelerators to handle different workloads and improves computing efficiency and flexibility through resource pooling.

[0086] See also Figure 2 The embodiment of the present application provides a schematic diagram of the cabinet structure of a server cabinet. The server cabinet has two partitions, namely the first partition zone0 and the second partition zone1. Each partition includes one CXL memory pooling server (referred to as CXL Memory Box in this embodiment), four general-purpose computing servers (referred to as Host in this embodiment), and four heterogeneous acceleration pooling servers (referred to as GPU Box in this embodiment).

[0087] The first zone, zone0, includes four computing servers: Host 00, Host 01, Host 02, and Host 03; one CXL memory pooling server: CXL Memory Box 00; and four heterogeneous acceleration pooling servers: GPU Box 00, GPU Box 01, GPU Box 02, and GPU Box 03.

[0088] The second zone, zone 1, includes four computing servers: Host 10, Host 11, Host 12, and Host 13; one CXL memory pooling server: CXL Memory Box 10; and four heterogeneous acceleration pooling servers: GPU Box 10, GPU Box 11, GPU Box 12, and GPU Box 13.

[0089] In this embodiment of the application, the PCIe bus is used to interconnect PCIe memory pooling resources within a partition. That is, each compute server is connected to a memory pooling server via the PCIe bus, thereby allowing multiple compute servers to share the memory resources of the memory pooling server. In addition, multiple compute servers are connected to multiple heterogeneous acceleration pooling servers via the PCIe bus, thereby sharing the accelerated computing resources of multiple heterogeneous acceleration pooling servers.

[0090] In some embodiments, the memory pooling server includes a memory pool, a memory switch chip, and pooling management software;

[0091] The acceleration pooled server includes PCIe switching chips and computing accelerators;

[0092] The pooling management software is used to control the memory switching chip in the memory pooling server to allocate memory resources in the memory pool to each computing server based on the memory allocation request of each computing server;

[0093] The pooling management software is used to control the PCIe switching chip in each acceleration pool server to allocate acceleration resources in the acceleration calculator to each computing server based on the acceleration computing request of each computing server.

[0094] In the embodiment of the present application, the memory allocation request is used to request memory allocation for the computing server. After obtaining the memory allocation request from the computing server, the pooling management software controls the memory switching chip in each memory pooling server to allocate memory resources in the memory pool to each computing server.

[0095] The memory switch chip (CXL Switch) can effectively manage data interaction between the memory in the memory pooled server and external devices (such as computing servers). It can transfer infrequently used data from the memory to the hard disk and load required data from the hard disk back to the memory.

[0096] The memory switching chip may specifically be a CXL memory switching chip.

[0097] In the embodiment of the present application, the accelerated computing request is used to request the allocation of accelerated computing resources to the computing server. After obtaining the accelerated computing request from the computing server, the pooled management software can control the PCIe switching chip in each accelerated pooled server to allocate an accelerated calculator to the computing server, thereby allocating accelerated resources to it.

[0098] A PCIe switch is a hardware device used to extend the PCI Express (PCIe) bus. It is primarily used to manage and optimize data transmission between PCIe devices on multiple servers, expand system scalability and parallel processing capabilities, improve data transmission efficiency, and manage traffic between different devices.

[0099] The accelerated calculator may be at least one of a GPU accelerated calculator, an FPGA accelerated calculator, an ASIC accelerated calculator, and a CPU accelerated calculator.

[0100] The pooled management software of the embodiment of the present application is combined with a memory switching chip and a PCIe switching chip to realize the adjustment of pooled memory resource allocation and accelerated computing resources, so that the configuration of each computing server in the partition can be flexibly adjusted, thereby providing different computing power services according to actual business scenarios, avoiding resource redundancy, and improving resource utilization efficiency.

[0101] See also Figure 3 An embodiment of the present application provides an architectural diagram of partitions in a server cabinet, including a cabinet management controller (RMC), four computing servers (Host*4), a CXL memory pooling server (CXL Memory Box*1), and four heterogeneous acceleration pooling servers (CPU Box*4). Each server has a corresponding server management controller (BMC), and the RMC and BMC communicate through a network bus.

[0102] The CXL Memory Box*1 includes pooling management software (MCPU), CXL memory switch chips (CXL Switch), memory expansion controller MXC, and pooled memory (CXL Memory).

[0103] The heterogeneous acceleration pooled server (CPU Box*4) includes a PCIe switch chip (PCIe Switch) and a computing accelerator (such as a CPU Device).

[0104] In addition, the computing server (Host*4), CXL memory pooling server (CXL Memory Box*1) and heterogeneous acceleration pooling server (CPU Box*4) are also connected through the PCIe bus. The computing server (Host*4) can be connected to the CXL memory switch chip (CXL Switch) of the acceleration pooling server (CPU Box*4) and the PCIe switch chip (PCIe Switch) of the heterogeneous acceleration pooling server (CPU Box*4) through the PCIe bus.

[0105] In some embodiments, the plurality of target buses include a network bus; the server cabinet includes a first switch, a second switch, and a third switch;

[0106] The network bus includes the business data network bus, the remote direct memory access (RDMA) data network bus, and the out-of-band management network bus;

[0107] The plurality of computing servers are connected to the first switch via a service data network bus to provide computing resources to external devices of the plurality of computing servers via the first switch;

[0108] The plurality of heterogeneous acceleration pooled servers are connected to the second switch via an RDMA data network bus, so as to provide acceleration resources to external devices of the plurality of heterogeneous acceleration pooled servers via the second switch;

[0109] The cabinet management controller and each server management controller are connected to the third switch via an out-of-band management network bus, so as to centrally control and manage each server via the third switch.

[0110] continue Figure 2 The server cabinet of the embodiment of the present application further includes a first switch, a second switch and a third switch. These three switches can be TOR switches. Specifically, Figure 2 In the figure, the first switch may specifically refer to the TORHost business network switch, the second switch may specifically refer to the TOR RDMA switch; and the third switch may specifically refer to the TOR management network switch.

[0111] Multiple computing servers can be connected to the first switch through the business data network bus in the network bus, so that computing resources can be provided to the external devices of the multiple computing servers through the first switch; multiple heterogeneous acceleration pooled servers are connected to the second switch through the RDMA data network bus in the network bus, so that acceleration resources can be provided to the external devices of the multiple heterogeneous acceleration pooled servers through the second switch; the cabinet management controller and each server management controller are connected to the third switch through the out-of-band management network bus in the network bus, so that each server can be centrally controlled and managed through the third switch.

[0112] In some embodiments, the multiple target buses include a water-cooled bus; each server includes a liquid cooling circuit;

[0113] The water cooling bus is connected to each liquid cooling pipeline and the external water cooling unit respectively; the water cooling bus is used to transport the cooling liquid of the external water cooling unit to the liquid cooling pipeline to dissipate heat for each server.

[0114] The water cooling bus is provided with cooling water and circulation power by an external water cooling unit (external refers to the outside of the cabinet), and is connected to each server through liquid cooling pipes, thereby dissipating the heat of the server and controlling the temperature of the server components within the operating range.

[0115] continue Figure 2 , Figure 2 The server cabinet is connected to an external water cooling unit, which can be connected to various servers, switches, etc., so as to dissipate heat for various parts inside the server cabinet.

[0116] In some embodiments, the plurality of target buses includes a power bus; the server cabinet includes a power management device;

[0117] Each server in the server cabinet is connected to the power management device via a power supply bus, so as to obtain power allocated to it by the power management device via the power supply bus.

[0118] The servers and switches inside the server cabinet can be connected to the power management device through the power supply bus. The power management device is mainly used to provide a stable power supply in data centers and large enterprises. The power management device can distribute the power of the power supply to each server through the power supply line, ensuring stable power supply to the server and the normal operation of the server cabinet.

[0119] In some embodiments, the power management device includes two power racks;

[0120] The cabinet management controller is located in any power rack.

[0121] The power management devices in the embodiments of the present application can specifically be two power racks. A power rack is an equipment architecture used in data centers, computer rooms, and other locations, specifically designed to support the integration and management of various IT equipment and power systems. A power rack typically combines power supply, cooling, equipment installation, and monitoring functions to ensure efficient and secure equipment operation.

[0122] A rack management controller (RMC) is deployed in one of the power racks.

[0123] The power rack in the embodiment of the present application may specifically be a PowerShell. The PowerShell not only provides basic mounting bracket functionality but also additionally provides power supply (via a power distribution unit (PDU)) and heat management (e.g., a built-in fan or cooling system), enabling the equipment to operate under more optimized conditions.

[0124] continue Figure 2 The server structure includes power rack PowerShell 00 and power rack PowerShell 01, wherein PowerShell 00 includes a cabinet management controller. Figure 2 It can be represented as PowerShell 00 (BMC).

[0125] In the embodiment of the present application, the BMC is set at any power rack node. On the one hand, it is used for the management and control of the two power racks. On the other hand, it is responsible for the identification and centralized management of the two fusion architecture partitions in the entire cabinet, so that each partition of the server organization can operate as flexibly as a single server system.

[0126] In some embodiments, the method further comprises:

[0127] Based on a preset power-on sequence, determining a target server to be powered on, and sending a power-on request to a target management interface of a server management controller of the target server, so as to power on the target server through the server management controller;

[0128] After a preset time has passed since the power-on request was sent, the power-on status of the target server is obtained from the target management interface of the target server. If the power-on status is powered on and there are unpowered servers, the next target server to be powered on is determined based on the power-on sequence.

[0129] After obtaining the server information of each server in the server cabinet, BMC can establish the management capability of the entire server cabinet. Figure 3 Provided architectural diagram for management operations.

[0130] The server nodes in the partition have power-on sequence requirements. When powering on, they must follow a certain power-on sequence. For example, the power-on sequence is as follows: heterogeneous acceleration resource pool server (GPU Box) nodes and CXL memory pool server (CXLMemory Box) nodes are powered on at the same time, and general computing server (Host) nodes are powered on in sequence. When powering off, they need to be powered off in the reverse sequence.

[0131] The aforementioned management architecture enables coordinated power-on and power-off interfaces to be implemented in the RMC. During power-on, a Redfish interface request is first sent to the GPUBox node and CXL Memory Box BMC to initiate the power-on operation. The Redfish interface for the GPU Box and CXL Memory Box power-on status is then polled. If the power-on status indicates "Powered On," indicating that the GPU Box and CXL Memory Box nodes have been powered on, a power-on request can be sent to the next target server, the Host node, completing the coordinated power-on operation. This management architecture enables each partition to operate as a single, complete server system, significantly reducing management complexity and enabling flexible system adjustments and ease of management.

[0132] See also Figure 4 The present invention provides a second flow chart of a server management method, which can be executed by a server management controller of each server in a server cabinet. Each server contained in the server cabinet has multiple target buses of a unified standard. The multiple target buses of each server converge on a PCIe resource bus unit of the server. The method includes steps 410 to 420:

[0133] Step 410: When the system is powered on, define a target management interface for providing server information.

[0134] Step 420: When being accessed by a rack management controller of the server rack, provide server information to the rack management controller based on a target management interface, so that the rack management controller manages the server based on the server information obtained from the target management interface;

[0135] The server information includes the target U-bit location; the target U-bit location is determined based on the following:

[0136] The target U-bit position of the corresponding server is determined according to each PIN possessed by the PCIe resource bus unit of the corresponding server; each PIN is used to indicate the partition identifier and U-bit identifier of the server.

[0137] The server cabinet includes multiple servers; each server has multiple target buses with unified standards;

[0138] The embodiment of the present application can specifically be a converged architecture server cabinet, which includes multiple target buses, including but not limited to a power supply bus, a network bus, a liquid cooling bus, a resource bus, etc., and can also include other buses, such as an address bus, an expansion bus, a system bus, a serial bus, and a parallel bus, etc.

[0139] The various target buses of each server converge into the PCIe resource bus unit of the server.

[0140] The PCIe resource bus unit (RBU) is a key component in PCIe systems responsible for managing and allocating bus resources. It provides functions such as resource allocation, device addressing, bandwidth management, and configuration management, ensuring that multiple devices can operate efficiently and stably while sharing the bus.

[0141] Multiple target buses converge on the server's PCIe resource bus unit. Each server service can manage each bus of the server cabinet without having to connect to the U-bit asset management module through the server management controller. This can effectively improve the efficiency of resource management and simplify the architecture and management of the server cabinet.

[0142] The Server Management Controller is a critical tool for ensuring the health of your servers.

[0143] The server management controller can be specifically a Baseboard Management Controller (BMC). A BMC is a dedicated hardware management controller in a server. It is usually integrated on the server's motherboard and is used to provide hardware monitoring, management, and remote control functions.

[0144] Each server has a PCIe resource bus unit. To reduce server management costs, the present embodiment sets up various PINs in the PCIe resource bus unit. Each PIN is used to indicate the server's partition identifier and U-bit identifier. The partition identifier is used to indicate the partition where the server is located, and the U-bit identifier is used to indicate the server's starting U-bit position.

[0145] The server management controller reads each PIN in each PCIe resource bus unit to obtain the partition identifier and U-bit identifier of its corresponding server. Based on the partition identifier and U-bit identifier of the server and the number of U bits of the server, the target U-bit position of the server can be determined. The target U-bit position is used to indicate the highest U-bit position of the server.

[0146] The target U-bit position can be subsequently used to configure a first network address for each server, thereby building a local area network within the server, and the servers inside the server cabinet can perform network communication within the local area network.

[0147] The server management controller may store the target U-bit location in its own management memory.

[0148] In order to enable the cabinet management controller to obtain server information including the target U-bit position, the first network address, etc., the server management controller defines a target management interface based on the server information of the corresponding server. The cabinet management controller can obtain the server information of each server by directly accessing the target management interface of each server management controller.

[0149] The target management interface may be a Redfish interface, which is a standard interface based on a RESTful API and is intended to provide more modern, easily extensible, and automated management functions for server management.

[0150] After obtaining the server information of each server, the cabinet management controller can manage each server, generate a management topology map, and provide computing services, memory services, etc. based on the management map containing the information of each server.

[0151] Through the server management method provided by the present application, multiple target buses of each server with unified standard multiple target buses converge on the PCIe resource bus unit of the server. Each server management controller determines the target U-bit position of the server according to the various PINs possessed by the corresponding PCIe resource bus unit of the corresponding server. The target U-bit position belongs to the server information. The server management controller defines a target management interface for providing server information. The cabinet management controller obtains the information of each server by accessing the target management interface, further reducing the cable structure in the cabinet, and can realize the identification and unified management of the U-bit assets, network addresses and other information of each server node, thereby reducing management costs.

[0152] In some embodiments, the PCIe resource bus unit includes a PCIe bus connector; the PCIe bus connector includes an IO expander and various PINs, and each server management controller is connected to the IO expander corresponding to the PCIe bus connector;

[0153] Determining the target U-bit position of the corresponding server according to each PIN of the PCIe resource bus unit of the corresponding server includes:

[0154] Based on the connected IO expander, each PIN in the PCIe bus connector is obtained and parsed to obtain the U-bit identifier of the server; the U-bit identifier is used to indicate the starting U-bit position;

[0155] Determine the target U-bit position of the server based on the starting U-bit position and the number of U bits.

[0156] The U-bit identifier and partition identifier of each server node are encoded through the I2CIO expansion device located on the cabinet PCIe data bus connector. After the server node is inserted into the cabinet and the BMC in the server is powered on and started, the I2C access IO expansion device is used to summarize the voltage levels of each PIN pin and parse the U-bit identifier and partition identifier based on the voltage levels of each PIN pin.

[0157] See also Figure 5 , an embodiment of the present application provides a connection diagram between a server and a PCIe bus connector, where the server includes a BMC and a PCIe Device (PCIe device), and the PCIe bus connector includes an IO (input / output) expander (specifically, an I2C IO expander for inter-integrated circuit communication protocol) and PCIe PINS (8 PIN pins), each PIN corresponds to one bit, and the 8 PINs are b0, b1 b2, b3, b4, b5, b6, and b7, respectively. In actual applications, 1 byte can be used to record the U bit identifier and partition identifier, and bit7-bit6 of the 1 byte are defined as the partition identifier of the server, and bit5-bit0 are defined as the U bit identifier of the server in the cabinet.

[0158] Figure 5 In the figure, the BMC in the server is connected to the I2CIO expander on the PCIe bus connector through the clock line SDC and the data line SDA. The 8 IO PINs behind the I2C IO expander are obtained through the I2C protocol and compared with the above table for parsing to obtain the U bit identifier and partition identifier of the server itself. The PCIe device communicates with other servers through the other PINs of the connector.

[0159] The preceding embodiment illustrates that each partition of a server cabinet includes multiple servers, with partition ID 00 representing the first partition and 01 representing the second partition. These servers include compute servers (represented by host), memory pooling servers (CXL Memory Box), and heterogeneous acceleration pooling servers (GPU Box). Each device in a partition has a target U-bit ID and partition ID. The target U-bit position can be determined based on the U-bit ID and the number of U bits.

[0160] Of course, in addition to the servers, the server cabinet also has various switches (such as the aforementioned TOR Host business switch, TOR RDMA switch, TOR management network switch), power management equipment (PowerShelf), etc., all of which have corresponding target U-positions.

[0161] Referring to Table 1, the embodiment of the present application provides the U-bits occupied by each node inside the server cabinet (the highest occupied U-bit is the target U-bit position of the node), the node number, the partition identifier, and the U-bit identifier:

[0162]

[0163] Table 1

[0164] The BMC can obtain and parse the U-bit identifier of the corresponding server based on the connected I / O expander. It can determine the target U-bit position of the server by comparing the starting U-bit position and the number of U bits. For example, in Table 1, the U-bit identifier is 000001, and the corresponding U-bit identifier (U-bit starting position) for compute server Host 13 in partition 1 is 00 1101. The corresponding decimal value is 13, which occupies one U bit. Therefore, the target U-bit position is 13 + 1 - 1 = 13.

[0165] Similarly, the U-bit identifier (U-bit starting position) corresponding to the memory pooling server GPU Box 10 in partition 1 is 00 1010, and the corresponding decimal value is 10. It occupies 3 U bits, and the occupied U bits are 10, 11, and 12 respectively. Its target U-bit position is 10+3-1=12.

[0166] The above table is converted into the following JSON code and stored in the BMC file system of each server node, so that when the server node BMC reads each PIN, it can obtain each target U-bit position according to the partition identifier and U-bit identifier obtained by parsing each PIN.

[0167] The embodiment of the present application obtains and parses the U-bit identifier of the corresponding server based on the connected IO expander; the U-bit identifier is used to indicate the starting U-bit position; the starting U-bit position and the number of U bits are processed to determine the target U-bit position of the server to which it belongs, without the need for an additional U-bit asset management module, thereby reducing resource management costs.

[0168] In some embodiments, the server information further includes a first network address of the server; after determining the target U-bit position of the corresponding server according to each PIN possessed by the PCIe resource bus unit corresponding to the corresponding server, the method further includes:

[0169] The first network address is determined based on:

[0170] Obtaining a second network address of a cabinet management controller;

[0171] The data of the preset byte position of the second network address of the cabinet management controller is replaced with the target U bit position to obtain and configure the first network address of the server.

[0172] RMC centrally manages server nodes within a cabinet. This requires establishing a network link from RMC to the BMC of each server node. After reading the U-bit identification identifier, the server BMC automatically configures its own management network address based on the U bit, establishing a management local area network within the cabinet.

[0173] When RMC starts, it automatically configures its second network address (also called the management network address) in the cabinet management local area network. The BMC of each server node in the cabinet identifies its partition information based on Bit7-Bit6 of the U-bit identifier, and identifies the U-bit identifier (denoted as Umin) based on Bit5-Bit0. By reading the number of U bits occupied by the chassis (denoted as Un) set in the motherboard read-only memory (EEPROM) when the server leaves the factory, it obtains the target U-bit position occupied by itself in the cabinet (denoted as Umax), and uses Umax as a variable to configure the management network address of the server node in the cabinet management local area network.

[0174] Specifically, the data of the preset byte position of the second network address of the cabinet management controller can be replaced with the target U bit position to obtain the first network address of the server to which it belongs. For example, the preset byte is the last byte. Assuming that the second network address of the BMC is 169.254.10.100, the first network address and the second network address can be IPv4 addresses or IPv6 addresses, without any restriction.

[0175] Continuing with the embodiment of Table 1, after configuring the first network address of each node, the first network address of each node is shown in the following Table 2:

[0176]

[0177] Table 2

[0178] The embodiment of the present application configures a first network address for each server based on the target U-bit position of each server to achieve the purpose of binding the network address to the U-bit position. Subsequently, it will be convenient for administrators to plan the network topology more reasonably, avoid network bottlenecks or traffic conflicts, and help improve the network management efficiency of the data center, enhance troubleshooting capabilities, optimize resource utilization, and improve overall network performance and reliability.

[0179] In some embodiments, the server cabinet includes at least two partitions; the partition identifier is used to indicate the partition where the server is located;

[0180] The server information also includes the node number of the server; the mapping relationship between the U bit identifier and partition identifier of each server and the node number is stored in the controller file system of each server;

[0181] Before defining a target management interface for providing server information, the method further includes:

[0182] The node ID of the server is determined based on:

[0183] Parse the PINs of the server's PCIe resource bus unit to obtain the server's U-bit identifier and partition identifier;

[0184] According to the mapping relationship, query the node number corresponding to the U bit identifier of the corresponding server and the partition identifier.

[0185] In the embodiment of the present application, each server has a corresponding partition identifier. There is a mapping relationship between the server's U-bit identifier, the partition identifier, and the node number. The node number is also called the node identifier. For the server, the node identifier is the server identifier and can uniquely indicate the server.

[0186] This mapping relationship can be stored in the form of the above Table 1. The node code that has a mapping relationship with the server's U-bit identifier and partition identifier can be queried from the above Table 1, so that the server information provided by the target management interface is more complete, which helps RMC to understand the information of each server node more comprehensively and improve network manageability.

[0187] In some embodiments, the server information further includes operating status information and / or chassis asset information; before defining a target management interface for providing the server information, the method further includes:

[0188] Monitor the server's operating status information; the operating status information includes at least one of the server's power-on status, power-on status, and health status; and / or,

[0189] The chassis asset information is obtained from a read-only memory of the server, where the chassis asset information includes at least one of a server type, a model, a serial number, and a U-bit number.

[0190] It is understandable that the server information in the embodiment of the present application also includes operating status information and / or chassis asset information.

[0191] The operating status information is obtained by the server management controller through monitoring of the server and is stored in the management memory of the server management controller.

[0192] Chassis asset information is non-modifiable and is set in the motherboard read-only memory (EEPROM) when the server leaves the factory. Chassis asset information includes at least one of the server type, model, serial number, and U-bit number (indicating the height of the U-bit occupied).

[0193] Of course, other information may also be included, such as production batch number and so on.

[0194] Referring to Table 3 below, the embodiment of the present application shows various server information corresponding to each field of the target management interface. Each field represents a type of server information, and each field has a corresponding interpretation and source:

[0195]

[0196] Table 3

[0197] The target management interface of the server management controller also provides server operation status information and / or chassis asset information, thereby enabling the target management interface to provide richer server information and enabling the cabinet management controller to obtain server information more comprehensively.

[0198] From the above, we can see that BMC can obtain server information including the server's target U-bit position, first network address, node number, operating status information and chassis asset management information. BMC can summarize this information in the target management interface (such as redfish). As a result, RMC can obtain more comprehensive server information from the BMC's target management interface, greatly reducing the complexity of management and enabling the system to be flexibly adjusted and easy to manage.

[0199] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0200] Figure 6 This is a block diagram of a computer program product according to an embodiment of the present application. The computer program product can be applied to a rack management controller of a server rack; the server rack includes multiple servers; each server has multiple target buses of a unified standard; the multiple target buses of each server converge on a PCIe resource bus unit of the server;

[0201] like Figure 6 As shown, the computer program product includes: a first communication module 610, an acquisition module 620 and a management module 630;

[0202] The first communication module 610 is used to periodically access the server management controller of each server contained in the server cabinet when the server cabinet is powered on; the server management controller defines a target management interface for providing server information;

[0203] An acquisition module 620 is configured to acquire server information of each server from a target management interface of each server management controller; the server information includes a target U-bit position of the corresponding server; the target U-bit position is determined by each server management controller based on respective PINs of a PCIe resource bus unit corresponding to the corresponding server; each PIN is used to indicate a partition identifier and a U-bit identifier of the server;

[0204] The management module 630 is used to manage each server based on the server information of each server.

[0205] In some embodiments, the computer program product further comprises:

[0206] The building module is used to control the automatic network building tool to build a local area network inside the server cabinet based on the first network address in the information of each server, so that each server can perform network communication under the local area network.

[0207] In some embodiments, the multiple target buses include a PCIe bus; the server cabinet includes multiple partitions; each server in each partition includes a memory pooling server, multiple computing servers, and multiple heterogeneous acceleration pooling servers;

[0208] Multiple computing servers are connected to the memory pooling server via the PCIe bus to share the memory resources of the memory pooling server;

[0209] Multiple computing servers are connected to multiple heterogeneous acceleration pool servers through a PCIe bus to share accelerated computing resources in the multiple heterogeneous acceleration pool servers.

[0210] In some embodiments, the plurality of target buses include a network bus; the server cabinet includes a first switch, a second switch, and a third switch;

[0211] The network bus includes the business data network bus, the remote direct memory access (RDMA) data network bus, and the out-of-band management network bus;

[0212] The plurality of computing servers are connected to the first switch via a service data network bus to provide computing resources to external devices of the plurality of computing servers via the first switch;

[0213] The plurality of heterogeneous acceleration pooled servers are connected to the second switch via an RDMA data network bus, so as to provide acceleration resources to external devices of the plurality of heterogeneous acceleration pooled servers via the second switch;

[0214] The cabinet management controller and each server management controller are connected to the third switch via an out-of-band management network bus, so as to centrally control and manage each server via the third switch.

[0215] In some embodiments, the multiple target buses include a water-cooled bus; each server includes a liquid cooling circuit;

[0216] The water cooling bus is connected to each liquid cooling pipeline and the external water cooling unit respectively; the water cooling bus is used to transport the cooling liquid of the external water cooling unit to the liquid cooling pipeline to dissipate heat for each server.

[0217] In some embodiments, the plurality of target buses includes a power bus; the server cabinet includes a power management device;

[0218] Each server in the server cabinet is connected to the power management device via a power supply bus, so as to obtain power allocated to it by the power management device via the power supply bus.

[0219] In some embodiments, the power management device includes two power racks;

[0220] The cabinet management controller is located in any power rack.

[0221] In some embodiments, the computer program product further comprises:

[0222] Control module for:

[0223] Based on a preset power-on sequence, determining a target server to be powered on, and sending a power-on request to a target management interface of a server management controller of the target server, so as to power on the target server through the server management controller;

[0224] After a preset time has passed since the power-on request was sent, the power-on status of the target server is obtained from the target management interface of the target server. If the power-on status is powered on and there are unpowered servers, the next target server to be powered on is determined based on the power-on sequence.

[0225] Figure 7 This is a second block diagram of a computer program product according to an embodiment of the present application. The computer program product can be applied to a server management controller of each server in a server cabinet; each server contained in the server cabinet has multiple target buses of a unified standard; the multiple target buses of each server converge on a PCIe resource bus unit of the server;

[0226] like Figure 7 As shown, the computer program product includes: a definition module 710 and a second communication module 720;

[0227] A definition module 710 is used to define a target management interface for providing server information when the power is on;

[0228] The second communication module 720 is configured to provide server information to the cabinet management controller based on a target management interface when accessed by the cabinet management controller of the server cabinet, so that the cabinet management controller manages the server based on the server information obtained from the target management interface;

[0229] The server information includes the target U-bit location; the target U-bit location is determined based on the following:

[0230] The target U-bit position of the corresponding server is determined according to each PIN possessed by the PCIe resource bus unit of the corresponding server; each PIN is used to indicate the partition identifier and U-bit identifier of the server.

[0231] In some embodiments, the server information further includes operating status information and / or chassis asset information; and the computer program product further includes:

[0232] A monitoring module for monitoring the server's operating status information; the operating status information includes at least one of the server's power-on status, power-on status, and health status; and / or

[0233] The chassis asset acquisition module is used to obtain chassis asset information from the read-only memory of the server, where the chassis asset information includes at least one of the server type, model, serial number, and U-bit number.

[0234] In some embodiments, the PCIe resource bus unit includes a PCIe bus connector; the PCIe bus connector includes an IO expander and various PINs, and each server management controller is connected to the IO expander corresponding to the PCIe bus connector;

[0235] The definition module 710 is specifically used to:

[0236] Based on the connected IO expander, each PIN in the PCIe bus connector is obtained and parsed to obtain the U-bit identifier of the server; the U-bit identifier is used to indicate the starting U-bit position;

[0237] Determine the target U-bit position of the server based on the starting U-bit position and the number of U bits.

[0238] In some embodiments, the computer program product comprises:

[0239] The first processing module is configured to:

[0240] The first network address is determined based on:

[0241] Obtaining a second network address of a cabinet management controller;

[0242] The data of the preset byte position of the second network address of the cabinet management controller is replaced with the target U bit position to obtain and configure the first network address of the server.

[0243] In some embodiments, the server cabinet includes at least two partitions; the partition identifier is used to indicate the partition where the server is located; the server information also includes the node number of the server; the controller file system of each server stores the mapping relationship between the U-bit identifier of each server and the partition identifier and the node number;

[0244] The computer program product further comprises:

[0245] The second processing module is configured to:

[0246] The node ID of the server is determined based on:

[0247] Parse the PINs of the server's PCIe resource bus unit to obtain the server's U-bit identifier and partition identifier;

[0248] According to the mapping relationship, query the node number corresponding to the U bit identifier of the corresponding server and the partition identifier.

[0249] For descriptions of features in the embodiments corresponding to the computer program product, reference can be made to the relevant descriptions of the embodiments corresponding to the server management method, which will not be repeated here.

[0250] Figure 8 This is one of the flow diagrams of interaction between the cabinet management controller and the server management modules of each server provided according to an embodiment of the present application.

[0251] The server cabinet includes a cabinet management controller and multiple servers; each server has a unified standard multiple target bus and a server management controller; the multiple target buses of each server converge on the PCIe resource bus unit of the server; see Figure 8 The interaction process includes the following steps 810 to 830:

[0252] Step 810: When powered on, the server management controller defines a target management interface for providing server information; the server information includes a target U-bit position of the corresponding server; the target U-bit position is determined by each server management controller based on the respective PINs of the PCIe resource bus unit corresponding to the corresponding server; each PIN is used to indicate the partition identifier and U-bit identifier of the server;

[0253] Step 820: When the rack management controller is powered on, it periodically accesses the server management controller of each server.

[0254] Step 830: The rack management controller obtains server information of each server from the target management interface of each server management controller, and manages each server based on the server information of each server.

[0255] The detailed implementation of the above process can be found in the above embodiments and will not be described again here.

[0256] Through the server management method provided by the present application, multiple target buses of each server with unified standard multiple target buses converge on the PCIe resource bus unit of the server. Each server management controller determines the target U-bit position of the server according to the various PINs possessed by the corresponding PCIe resource bus unit of the corresponding server. The target U-bit position is important server information. The server management controller defines a target management interface for providing server information. The cabinet management controller obtains information of each server by accessing the target management interface, further reducing the cable structure in the cabinet, and can realize the identification and unified management of the U-bit assets, network addresses and other information of each server node, thereby reducing management costs.

[0257] See also Figure 9 The embodiment of the present application provides a schematic diagram of an interaction process between a cabinet management controller in a server cabinet and a server management controller of each server, including the following steps:

[0258] For any server management controller:

[0259] The server management controller itself is powered on and started up;

[0260] The server management controller reads each PIN of the PCIe resource bus unit to determine the partition identifier and U-bit identifier of the server;

[0261] The server management controller determines the target U-bit position of the corresponding server according to the U-bit identifier and the number of U bits;

[0262] The server management controller configures the first network address of the server according to the target U-bit position of the corresponding server;

[0263] The server management controller determines from the BMC file system that a node ID that is mapped to the partition ID and the U-bit ID of the server is the node ID of the server;

[0264] The server management controller obtains operation status information; the operation status information includes at least one of the server power-on status, power-on status and health status;

[0265] The server management controller reads the read-only memory to obtain chassis asset information, where the chassis asset information includes at least one of a server type, a model, a serial number, and a U-bit number;

[0266] The server management controller provides the server information including at least part of the target U-bit position, the first network address, the operation status information, the node number and the chassis asset information to the outside through the redfish interface;

[0267] The cabinet management controller itself is powered on and started;

[0268] The rack management controller obtains the server information of each server from the target management interface of each server management controller through the network bus, and summarizes the server information;

[0269] The cabinet management controller automatically manages each server node based on the server information;

[0270] The rack management controller generates a centralized management topology based on the information of each server.

[0271] The detailed implementation process of the above process can be found in the above embodiment and will not be described again here.

[0272] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps of any of the above-mentioned server management method embodiments.

[0273] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned server management method embodiments when running.

[0274] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0275] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned server management method embodiments are implemented.

[0276] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0277] The above is a detailed introduction to a server management method, computer program product, server cabinet, electronic device and computer-readable storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only intended to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A server management method, characterized in that: A cabinet management controller applied to a server cabinet; the server cabinet includes a plurality of servers; Each server has multiple target buses with unified standards; Multiple target buses of each server converge into a PCIe resource bus unit of the server; the method includes: When powered on, the server management controller of each server is periodically accessed; the server management controller defines a target management interface for providing server information; Obtaining server information of each server from a target management interface of each server management controller; the server information includes a target U-bit position of the corresponding server; the target U-bit position is determined by each server management controller based on respective PINs possessed by a PCIe resource bus unit corresponding to the corresponding server; the respective PINs are used to indicate a partition identifier and a U-bit identifier of the server; Manage each server based on its server information.

2. The server management method according to claim 1, characterized in that: The server information further includes a first network address of the server; the first network address is determined by the server management controller according to a target U-bit position of the corresponding server; after obtaining the server information of each server from the target management interface of each server management controller, the method further includes: The control automation network building tool builds a local area network inside the server cabinet based on the first network address in the information of each server, so that each server can perform network communication under the local area network.

3. The server management method according to claim 1, wherein: The multiple target buses include a PCIe bus; The server cabinet includes multiple partitions; each server in each partition includes a memory pooling server, multiple computing servers, and multiple heterogeneous acceleration pooling servers; The multiple computing servers are connected to the memory pooling server via the PCIe bus to share memory resources of the memory pooling server; The multiple computing servers are connected to the multiple heterogeneous acceleration pooling servers through the PCIe bus to share the accelerated computing resources in the multiple heterogeneous acceleration pooling servers.

4. The server management method according to claim 3, characterized in that: The multiple target buses include a network bus; the server cabinet includes a first switch, a second switch, and a third switch; The network bus includes a business data network bus, a remote direct memory access (RDMA) data network bus, and an out-of-band management network bus; The plurality of computing servers are connected to the first switch via the service data network bus to provide computing resources to external devices of the plurality of computing servers via the first switch; The plurality of heterogeneous acceleration pooled servers are connected to the second switch via the RDMA data network bus, so as to provide acceleration resources to external devices of the plurality of heterogeneous acceleration pooled servers via the second switch; The cabinet management controller and each server management controller are connected to the third switch via the out-of-band management network bus, so as to centrally control and manage each server via the third switch.

5. The server management method according to claim 1, characterized in that: The multiple target buses include a water-cooled bus; each server includes a liquid cooling pipeline; The water cooling bus is connected to each of the liquid cooling pipelines and the external water cooling unit respectively; the water cooling bus is used to transport the cooling liquid of the external water cooling unit to the liquid cooling pipeline to dissipate heat for each server.

6. The server management method according to claim 1, wherein: The plurality of target buses include a power supply bus; the server cabinet includes a power management device; Each server in the server cabinet is connected to the power management device via the power supply bus, so as to obtain power allocated thereto by the power management device via the power supply bus.

7. The server management method according to claim 6, characterized in that: The power management device includes two power racks; The cabinet management controller is located in any power rack.

8. The server management method according to any one of claims 1 to 7, characterized in that: The method further comprises: Determine a target server to be powered on based on a preset power-on sequence, and send a power-on request to a target management interface of a server management controller of the target server, so as to power on the target server through the server management controller; Within a preset time after sending the power-on request, the power-on status of the target server is obtained from the target management interface of the target server. If the power-on status is the powered-on state and there is a server that is not powered on, the next target server to be powered on is determined based on the power-on sequence.

9. A server management method, characterized in that: A server management controller applied to each server in a server cabinet; each server contained in the server cabinet has multiple target buses with unified standards; Multiple target buses of each server converge into a PCIe resource bus unit of the server; the method includes: In the case of power on, define the target management interface for providing server information; When accessed by a rack management controller of the server rack, providing the server information to the rack management controller based on the target management interface, so that the rack management controller manages the server based on the server information obtained from the target management interface; The server information includes a target U-bit position; the target U-bit position is determined based on the following method: The target U-bit position of the corresponding server is determined according to each PIN possessed by the PCIe resource bus unit of the corresponding server; the each PIN is used to indicate the partition identifier and U-bit identifier of the server.

10. The server management method according to claim 9, characterized in that: The server information also includes operating status information and / or chassis asset information; Before defining a target management interface for providing server information, the method further includes: Monitoring the server's operating status information; the operating status information includes at least one of the server's power-on status, power-on status, and health status; and / or, The chassis asset information is obtained from a read-only memory of the server, where the chassis asset information includes at least one of a server type, a model, a serial number, and a U-bit number.

11. The server management method according to claim 10, characterized in that: The PCIe resource bus unit includes a PCIe bus connector; the PCIe bus connector includes an IO expander and various PINs, and each server management controller is connected to the IO expander corresponding to the PCIe bus connector; The determining of the target U-bit position of the corresponding server according to each PIN possessed by the PCIe resource bus unit corresponding to the corresponding server includes: Obtaining and parsing each PIN in the PCIe bus connector based on the connected IO expander to obtain a U-bit identifier of the server; the U-bit identifier is used to indicate the starting U-bit position; The target U-bit position of the server is determined according to the starting U-bit position and the number of U bits.

12. The server management method according to claim 9, characterized in that: The server information further includes a first network address of the server; after determining the target U-bit position of the corresponding server according to each PIN possessed by the PCIe resource bus unit corresponding to the corresponding server, the method further includes: The first network address is determined based on the following method: Obtaining a second network address of the cabinet management controller; The data of the preset byte position of the second network address of the cabinet management controller is replaced with the target U bit position to obtain and configure the first network address of the server.

13. The server management method according to claim 9, characterized in that: The server cabinet includes at least two partitions; the partition identifier is used to indicate the partition where the server is located; the server information also includes the node number of the server; the controller file system of each server stores the mapping relationship between the U-bit identifier of each server and the partition identifier and the node number; Before defining a target management interface for providing server information, the method further includes: The node number of the server is determined based on: Parse the PINs of the server's PCIe resource bus unit to obtain the server's U-bit identifier and partition identifier; The node number corresponding to the U-bit identifier and the partition identifier of the corresponding server is queried according to the mapping relationship.

14. A computer program product, characterized in that A cabinet management controller applied to a server cabinet; the server cabinet includes a plurality of servers; Each server has multiple target buses with unified standards; The various target buses of each server converge on the server's PCIe resource bus unit; The computer program product comprises: A first communication module is configured to periodically access a server management controller of each server contained in the server cabinet when the server cabinet is powered on; the server management controller defines a target management interface for providing server information; an acquisition module configured to acquire server information of each server from a target management interface of each server management controller; the server information includes a target U-bit position of the corresponding server; the target U-bit position is determined by each server management controller based on respective PINs possessed by a PCIe resource bus unit corresponding to the corresponding server; the respective PINs are used to indicate a partition identifier and a U-bit identifier of the server; The management module is used to manage each server based on the server information of each server.

15. A computer program product, characterized in that A server management controller applied to each server in a server cabinet; each server contained in the server cabinet has multiple target buses with unified standards; The various target buses of each server converge on the server's PCIe resource bus unit; The computer program product comprises: A definition module, used for defining a target management interface for providing server information when the power is on; a second communication module, configured to, when accessed by a cabinet management controller of the server cabinet, provide the server information to the cabinet management controller based on the target management interface, so that the cabinet management controller manages the server based on the server information obtained from the target management interface; The server information includes a target U-bit position; the target U-bit position is determined based on the following method: The target U-bit position of the corresponding server is determined according to each PIN possessed by the PCIe resource bus unit of the corresponding server; the each PIN is used to indicate the partition identifier and U-bit identifier of the server.

16. A server cabinet, characterized in that: The server cabinet includes a cabinet management controller and multiple servers; each server has a unified standard multiple target bus and server management controller; The various target buses of each server converge on the server's PCIe resource bus unit; The server management controller defines a target management interface for providing server information when powered on; The server information includes a target U-bit position of the corresponding server; the target U-bit position is determined by each server management controller according to each PIN possessed by the corresponding PCIe resource bus unit of the corresponding server; each PIN is used to indicate the partition identifier and U-bit identifier of the server; When powered on, the cabinet management controller periodically accesses the server management controller of each server; The server information of each server is acquired from the target management interface of each server management controller, and each server is managed based on the server information of each server.

17. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the server management method according to any one of claims 1 to 8 or 9 to 13 when executing the computer program.

18. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the server management method according to any one of claims 1 to 8 or 9 to 13 are implemented.

Citation Information

Patent Citations

  • Data center cabinet U bit management device and method

    CN116644771A

  • Server asset information reporting method, device, equipment and medium

    CN119356975A