Heterogeneous accelerated pooling server system, pooling resource allocation method, and media

By using a heterogeneous accelerated pooled server system, dynamically configuring I/O devices and establishing heterogeneous computing resource allocation paths, the scalability and operational efficiency issues of traditional heterogeneous accelerated server systems in large-scale and diversified applications are solved, thereby achieving improved high-efficiency computing and resource utilization.

CN119883590BActive Publication Date: 2026-05-29INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INSPUR SUZHOU INTELLIGENT TECH CO LTD
Filing Date
2024-09-30
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Traditional heterogeneous accelerated server systems suffer from limited scalability, closed ecosystems, energy waste, high upgrade and expansion costs, and low operating efficiency when facing large-scale and diversified applications, thus failing to meet the demands of high-performance computing.

Method used

The heterogeneous accelerated pooling server system includes a general computing resource pool, an internal bus switching module, an Ethernet switching module, and multiple heterogeneous computing accelerated resource pools. It dynamically configures I/O devices through globally unique codes and internal bus switching modules, establishes heterogeneous computing resource allocation paths, supports direct connection between accelerated processors and Ethernet controllers, and realizes flexible allocation and dynamic adjustment of resources.

Benefits of technology

It improves system scalability and operating efficiency, reduces communication latency, enhances computing performance and resource utilization, and supports efficient computing for diverse application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119883590B_ABST
    Figure CN119883590B_ABST
Patent Text Reader

Abstract

The application provides a heterogeneous acceleration pooling server system, a pooling resource allocation method and a medium, and relates to the technical field of computers.The heterogeneous acceleration pooling server system comprises a general computing resource pool, an internal bus exchange module, an Ethernet exchange module and a plurality of heterogeneous computing acceleration resource pools.The internal bus exchange module comprises a plurality of high-speed data interconnection modules connected to each other, a plurality of IO devices are arranged on each module, the IO devices of the plurality of high-speed data interconnection modules are allocated a globally unique code, and the general computing resource pool configures the required IO devices according to resource requirements after obtaining the resource requirements, establishes a heterogeneous computing resource allocation path to match the general computing resource pool with one or more heterogeneous computing acceleration resource pools corresponding to the resource requirements, can realize that each general computing unit can access any heterogeneous acceleration device, improves the system expansion capability, and improves the operation efficiency of the heterogeneous acceleration server when coping with large-scale and diversified applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a heterogeneous accelerated pooling server system, a pooling resource allocation method, and a medium. Background Technology

[0002] With the development of computer applications such as artificial intelligence and big data, the demand for computing power has grown rapidly. Simultaneously, with the diversification of computing application scenarios and the emergence of various application areas such as machine learning, artificial intelligence, autonomous driving, and the industrial internet, traditional general-purpose processors are encountering increasing performance bottlenecks when processing massive amounts of computing and data / images, such as low computational parallelism, low bandwidth, and excessive latency. Traditional server architectures based on single-core processors can no longer meet today's demands. To efficiently handle various computing tasks, different types of heterogeneous acceleration processors have emerged. Traditional heterogeneous acceleration services simply integrate existing processors of different specifications and models. The traditional heterogeneous acceleration service architecture uses an I / O tree structure. During startup, the host unit enumerates I / O bridges or I / O devices in the I / O tree structure using a depth-first enumeration method, assigning a unique ID number to each I / O tree. Routing paths can only be allocated within each I / O tree, limiting the diversified allocation and use of I / O device resources. Traditional heterogeneous acceleration servers suffer from limited system scalability, closed ecosystems, energy waste, and high upgrade and expansion costs, resulting in low operating efficiency when dealing with large-scale and diversified applications. Summary of the Invention

[0003] This invention provides a heterogeneous accelerated pooling server system, a pooling resource allocation method, and a medium to address the shortcomings of traditional heterogeneous accelerated pooling server systems, such as low operating efficiency, limited performance, and restrictions on the commercial application of high-performance computing technology.

[0004] This invention provides a heterogeneous accelerated pooling server system, comprising:

[0005] General-purpose computing resource pool, internal bus switching module, Ethernet switching module, and multiple heterogeneous computing acceleration resource pools;

[0006] The general computing resource pool is communicatively connected to the internal bus switching module via the first system interconnect bus;

[0007] Each heterogeneous computing acceleration resource pool communicates with the internal bus switching module via a second system interconnect bus.

[0008] The internal bus switching module includes multiple interconnected high-speed data interconnect modules. Each high-speed data interconnect module is equipped with multiple I / O devices. The I / O devices of the multiple high-speed data interconnect modules are assigned globally unique codes in a unified manner, so as to access any I / O device in any high-speed data interconnect module through the globally unique codes. After the general computing resource pool obtains the resource requirements, it configures the required I / O devices according to the resource requirements through the internal bus switching module, and establishes a heterogeneous computing resource allocation path so that the general computing resource pool matches one or more heterogeneous computing acceleration resource pools corresponding to the resource requirements.

[0009] According to the heterogeneous acceleration pooled server system provided by the present invention, the heterogeneous computing acceleration resource pool includes:

[0010] At least one accelerator card, wherein the accelerator card integrates an accelerator processor and an Ethernet controller;

[0011] The acceleration processor is directly connected to the Ethernet controller;

[0012] The Ethernet controller is connected to the Ethernet switching module.

[0013] The heterogeneous accelerated pooling server system provided by the present invention further includes:

[0014] A high-speed serial computer expansion bus standard link is used for accelerated processor interconnection between multiple servers.

[0015] According to the heterogeneous accelerated pooling server system provided by the present invention, the system further includes a remote direct data access network interface card (RDBNIC), which communicates with the kernel driver of the accelerated processor through a kernel driver standard interface.

[0016] The remote direct data access network card is connected to the Ethernet switching module, so that multiple accelerator processors connected to the multiple remote direct data access network cards can complete direct data transmission and remote direct access to video memory.

[0017] According to the heterogeneous accelerated pooling server system provided by the present invention, the internal bus switching module includes an I / O switching chip controller. The I / O switching chip controller is used to enumerate all I / O devices, define network topology partitioning according to the resource requirements, dynamically adjust the matching between the accelerated processor and the I / O devices, and establish resource allocation paths.

[0018] According to the heterogeneous accelerated pooling server system provided by the present invention, the globally unique code is obtained by encoding the high-speed data interconnect module number and its internal IO device number using a segmented encoding method.

[0019] According to the heterogeneous accelerated pooled server system provided by the present invention, the general computing resource pool includes:

[0020] Multiple general-purpose computing units, each including a general-purpose processor and memory, are combined using virtualization software to form computing resource containers of arbitrary granularity.

[0021] According to the heterogeneous accelerated pooling server system provided by the present invention, the system further includes a pooling management controller;

[0022] The general computing power unit in the general computing resource pool and the acceleration card in the heterogeneous computing acceleration resource pool are respectively equipped with a pooled node management controller. The pooled management controller manages multiple pooled node management controllers in a unified manner to perform at least one of the following on the general computing resource pool and the heterogeneous computing acceleration resource pool: node topology identification, centralized display of asset information, and coordinated power-on / off.

[0023] According to the heterogeneous accelerated pooled server system provided by the present invention, the pooling management controller is also used to perform at least one of the following functions: integrated monitoring, fault early warning, and visual management of the system through a standardized service interface.

[0024] According to the heterogeneous accelerated pooling server system provided by the present invention, the system further includes a pooling management engine;

[0025] The pooling management engine is used to manage the internal bus switching module. By calling the physical channels and software interfaces in the internal bus switching module, it discovers the current network topology to obtain the current asset allocation status, and coordinates the overall computing power resources to switch the network topology to meet the resource requirements based on the physical channels and software interfaces.

[0026] According to the heterogeneous accelerated pooled server system provided by the present invention, the pooling management engine is further used for:

[0027] Hot removal, hot insertion, and hot reset processes are performed on any accelerator card in the heterogeneous computing acceleration resource pool to control whether the accelerator card is put into use.

[0028] According to the heterogeneous accelerated pooled server system provided by the present invention, the pooling management engine is further used for:

[0029] Based on a load balancing-based optimized scheduling algorithm, the accelerator cards in the heterogeneous computing acceleration resource pool are dynamically allocated and adjusted.

[0030] According to the heterogeneous accelerated pooled server system provided by the present invention, the pooling management engine is further used for:

[0031] Establish expert templates for different application scenarios;

[0032] When switching application scenarios, different expert templates are switched by calling the API according to the current application scenario; the expert templates that meet the requirements of the application scenario are visualized through a visual web interface.

[0033] According to the heterogeneous acceleration pooling server system provided by the present invention, the heterogeneous acceleration pooling server system is a single-machine heterogeneous acceleration pooling server system, wherein a single host in the single-machine heterogeneous acceleration pooling server system supports expansion of 8 acceleration cards, 16 acceleration cards, and 32 acceleration cards.

[0034] According to the heterogeneous acceleration pooling server system provided by the present invention, the heterogeneous acceleration pooling server system is a multi-host heterogeneous acceleration pooling server system, in which each host is connected to a high-speed data interconnect module, and 32 acceleration processor cards are shared through multiple high-speed data interconnect modules.

[0035] According to the heterogeneous acceleration pooling server system provided by the present invention, the system further includes a computing power virtualization subsystem, which aggregates the heterogeneous computing acceleration resource pools into a unified virtual computing power resource pool. The computing power virtualization subsystem includes:

[0036] Pooled servers, which are computing nodes in the system cluster, are used for task execution;

[0037] A management server, which is used to run management components in the system cluster;

[0038] An application server, which runs applications within a system cluster, can be a container, a virtual machine, or a physical machine.

[0039] The present invention also provides a pooling resource allocation method, applicable to the heterogeneous accelerated pooling server system described in any of the above claims, comprising:

[0040] A general computing unit in the general computing resource pool initiates a request for dynamic adjustment of heterogeneous computing resources to the pooling management controller based on resource requirements.

[0041] The pooling management engine determines the IO devices to be accessed in the internal bus switching module based on the resource requirements, accesses the IO devices in one or more high-speed data interconnect modules through a globally unique code, reconstructs the heterogeneous computing resource allocation path, and matches the heterogeneous computing resources corresponding to the resource requirements for the general computing unit.

[0042] The heterogeneous computing resources include resources in a single acceleration resource pool, resources across acceleration resource pools, and resources across pooled servers.

[0043] According to the pooled resource allocation method provided by the present invention, the reconstruction of the heterogeneous computing resource allocation path includes releasing the accelerator card in a certain general computing unit and allocating it to other general computing units. Specifically, the pooled management engine performs hot removal of the accelerator card and sends a request to the internal bus switching module to obtain the physical location of the accelerator card.

[0044] The accelerator card is located based on its physical location, and then reset.

[0045] The reset accelerator cards are reassigned to the general-purpose computing units to be assigned.

[0046] According to the pooled resource allocation method provided by the present invention, the step of releasing the accelerator card in a certain general-purpose computing unit and allocating it to other general-purpose computing units further includes:

[0047] The accelerator card is trained according to the application scenario of the general computing unit to be allocated;

[0048] The trained accelerator cards are reassigned to the general computing units to be assigned.

[0049] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the pooled resource allocation method described in any of the preceding claims.

[0050] The heterogeneous accelerated pooling server system, pooling resource allocation method, and medium provided by this invention include a general-purpose computing resource pool, an internal bus switching module, an Ethernet switching module, and multiple heterogeneous computing accelerated resource pools. The general-purpose computing resource pool is communicatively connected to the internal bus switching module via a first system interconnect bus. Each heterogeneous computing accelerated resource pool is communicatively connected to the internal bus switching module via a second system interconnect bus. The internal bus switching module includes multiple interconnected high-speed data interconnect modules, each high-speed data interconnect module having multiple I / O devices, and the I / O devices of the multiple high-speed data interconnect modules are allocated in a coordinated manner. A globally unique code is used to access any IO device in any high-speed data interconnect module. After the general computing resource pool obtains the resource requirements, it configures the required IO devices according to the resource requirements through the internal bus switching module, and establishes a heterogeneous computing resource allocation path so that the general computing resource pool matches one or more heterogeneous computing acceleration resource pools corresponding to the resource requirements. By assigning a globally unique code to multiple IO devices on each high-speed data interconnect module, each general computing unit can access any heterogeneous acceleration device, improving the system's scalability and enhancing the operating efficiency of the heterogeneous acceleration server when dealing with large-scale and diversified applications. Attached Figure Description

[0051] To more clearly illustrate the technical solutions in this invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0052] Figure 1 This is a functional structure diagram of the heterogeneous acceleration pooling server system provided in an embodiment of the present invention;

[0053] Figure 2 This is a schematic diagram of the data flow of a heterogeneous accelerator card provided in an embodiment of the present invention;

[0054] Figure 3 This is a schematic diagram of the RDMA communication data flow of the heterogeneous accelerator card provided in an embodiment of the present invention;

[0055] Figure 4 This is a schematic diagram of end-to-end communication between distributed heterogeneous acceleration servers provided in an embodiment of the present invention;

[0056] Figure 5 This is a schematic diagram of the expansion of a distributed heterogeneous acceleration server cluster provided in an embodiment of the present invention;

[0057] Figure 6This is a schematic diagram of the heterogeneous computing power resource pool management architecture provided in an embodiment of the present invention;

[0058] Figure 7 This is a schematic diagram of the heterogeneous computing power resource pooling management engine architecture provided in an embodiment of the present invention;

[0059] Figure 8 This is a schematic diagram of the expert template mechanism for heterogeneous computing power resource pools provided in an embodiment of the present invention.

[0060] Figure 9 This is a schematic diagram of the existing heterogeneous acceleration server system structure provided in the embodiments of the present invention;

[0061] Figure 10 This is a schematic diagram of the single-machine 32-card heterogeneous acceleration server system structure provided in an embodiment of the present invention;

[0062] Figure 11 This is a schematic diagram of a single-rack 64-card heterogeneous acceleration server structure provided in an embodiment of the present invention;

[0063] Figure 12 This is a schematic diagram of a multi-host shared 32-card heterogeneous acceleration server structure provided in an embodiment of the present invention;

[0064] Figure 13 This is a schematic diagram of a multi-host shared 32-card heterogeneous acceleration server structure provided in an embodiment of the present invention;

[0065] Figure 14 This is a diagram showing the relationship between heterogeneous computing power software-defined virtualization and pooled servers provided in this embodiment of the invention;

[0066] Figure 15 This is a schematic diagram of the dynamic allocation process of heterogeneous computing power resources provided in an embodiment of the present invention;

[0067] Figure 16 This is a schematic diagram of the heterogeneous computing power software-defined virtualization deployment process provided in an embodiment of the present invention. Detailed Implementation

[0068] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0069] Figure 1 The functional structure diagram of the heterogeneous accelerated pooling server system provided in the embodiments of the present invention is as follows: Figure 1 As shown, the heterogeneous accelerated pooling server system provided in this embodiment of the invention includes:

[0070] General-purpose computing resource pool, internal bus switching module, Ethernet switching module, and multiple heterogeneous computing acceleration resource pools;

[0071] The general computing resource pool is communicatively connected to the internal bus switching module via the first system interconnect bus;

[0072] Each heterogeneous computing acceleration resource pool communicates with the internal bus switching module via a second system interconnect bus.

[0073] Each heterogeneous computing acceleration resource pool is communicatively connected to the Ethernet switching module;

[0074] The internal bus switching module includes multiple interconnected high-speed data interconnect modules. Each high-speed data interconnect module is equipped with multiple I / O devices. The I / O devices of the multiple high-speed data interconnect modules are assigned globally unique codes in a unified manner, so as to access any I / O device in any high-speed data interconnect module through the globally unique codes. After the general computing resource pool obtains the resource requirements, it configures the required I / O devices according to the resource requirements through the internal bus switching module, and establishes a heterogeneous computing resource allocation path so that the general computing resource pool matches one or more heterogeneous computing acceleration resource pools corresponding to the resource requirements.

[0075] In this embodiment of the invention, the globally unique code is obtained by encoding the high-speed data interconnect module number and its internal IO device number using a segmented encoding method.

[0076] In this embodiment of the invention, the general-purpose computing resource pool is a resource pool composed of general-purpose processors such as CPUs, and the heterogeneous computing acceleration resource pool is a resource pool composed of heterogeneous processors such as GPUs. The internal bus switching module contains multiple high-speed data interconnect modules, for example, four high-performance switching boards, each of which extends 10 high-speed interfaces (IO devices). The general-purpose computing resource pool is composed of general-purpose computing units, primarily providing the system's computing functions, while the heterogeneous computing acceleration resource pool is composed of heterogeneous computing units, primarily providing accelerated computing functions for the entire system, achieving higher computing performance, parallelism, and energy efficiency. The internal bus switching module is a high-performance switching unit, including high-speed data interconnect modules, tightly integrated with the basic software to achieve physical decoupling between general-purpose computing and heterogeneous accelerated computing. It also enables efficient configuration management of heterogeneous acceleration cards and direct resource sharing across nodes, improving the efficiency of heterogeneous resource access and providing a highly reliable and high-performance data interconnect network for the entire system.

[0077] In this embodiment of the invention, the heterogeneous accelerated pooling server system adopts a distributed architecture, catering to diverse application needs. It departs from the traditional CPU-centric design, employing key technologies such as hardware decoupling and consistent high-speed interconnection. Centered on an internal bus switching module, it decouples and reconstructs the traditional server architecture, enabling collaborative computing between multiple general-purpose processor platforms and various heterogeneous computing acceleration units. This achieves hardware decoupling, pooling, and reconstruction of large-scale computing resources, memory resources, heterogeneous acceleration resources, and storage resources. Through software-defined system design, it achieves dynamic scheduling of resources. The system resource pool is integrated through unified management, heat dissipation, and power supply, forming a high-performance host with high heterogeneous computing power and flexible resource allocation. The entire system can flexibly provide diverse computing power for various application scenarios, improving computing performance even under conditions of limited key components and devices, driven by application needs, and through system-level architectural key technology innovation to integrate diverse computing power.

[0078] In this embodiment of the invention, the acceleration processor is a graphics processing unit (GPU), a field programmable gate array (FPGA), etc. Through the interconnection of heterogeneous computing acceleration resource pools and software-defined system design, the system topology can be dynamically adjusted to achieve on-demand allocation of heterogeneous resources.

[0079] The distributed heterogeneous accelerated pooled server system achieves device resource pooling through high-performance internal bus switching and software-defined system design. This decouples devices from the physical link layer of the CPU, allowing for flexible adjustment of port configurations and resource allocation paths in the switching network. Fine-grained partitioning of shared resource pools enables on-demand elastic allocation of device resources and multi-host sharing. Addressing current data center host system performance limitations such as system interconnect bandwidth constraints, performance mismatches between different storage levels, and low I / O resource utilization, the pooled system design flexibly achieves dynamic resource allocation and load balancing for diverse scenario requirements. It also enables the fusion of multi-platform processor computing power and collaborative scheduling of heterogeneous accelerated resources, alleviating data center performance expansion bottlenecks.

[0080] Traditional heterogeneous acceleration services simply integrate existing processors of different specifications and models. The traditional heterogeneous acceleration service architecture uses an I / O tree structure. During startup, the host unit enumerates I / O bridges or I / O devices in the I / O tree structure using a depth-first enumeration method, assigning a unique ID number to each I / O tree. Routing paths can only be allocated within each I / O tree, limiting the diversified allocation and use of I / O device resources. This results in limited system scalability, a closed ecosystem, energy waste, and high upgrade and expansion costs. When dealing with large-scale, diversified applications, traditional heterogeneous acceleration servers are inefficient.

[0081] The heterogeneous accelerated pooling server system provided in this invention includes a general-purpose computing resource pool, an internal bus switching module, an Ethernet switching module, and multiple heterogeneous computing accelerated resource pools. The general-purpose computing resource pool is communicatively connected to the internal bus switching module via a first system interconnect bus. Each heterogeneous computing accelerated resource pool is communicatively connected to the internal bus switching module via a second system interconnect bus. The internal bus switching module includes multiple interconnected high-speed data interconnect modules. Each high-speed data interconnect module has multiple I / O devices. The I / O devices of the multiple high-speed data interconnect modules are assigned globally unique codes to access any I / O device in any high-speed data interconnect module. After obtaining resource requirements, the general-purpose computing resource pool configures the required I / O devices according to the resource requirements through the internal bus switching module, establishing a heterogeneous computing resource allocation path so that the general-purpose computing resource pool matches one or more heterogeneous computing accelerated resource pools corresponding to the resource requirements. By assigning globally unique codes to the multiple I / O devices on each high-speed data interconnect module, each general-purpose computing unit can access any heterogeneous accelerated device, improving system scalability and increasing the operating efficiency of the heterogeneous accelerated server when dealing with large-scale, diversified applications.

[0082] Based on any of the above embodiments, the heterogeneous computing acceleration resource pool includes:

[0083] At least one accelerator card, wherein the accelerator card integrates an accelerator processor and an Ethernet controller;

[0084] The acceleration processor is directly connected to the Ethernet controller;

[0085] The Ethernet controller is connected to the Ethernet switching module.

[0086] In this embodiment of the invention, an innovative design is implemented for the heterogeneous accelerator card, directly connecting the accelerator processor and the Ethernet controller to address the computational power challenges of large-scale models. This direct connection between the accelerator processor and the Ethernet controller via a switching chip differs from traditional methods where the accelerator processor connects to a switching chip via a Network Interface Card (NIC), then to the CPU via the switching chip, and finally to the NIC via a CPU-to-NIC interconnection. The heterogeneous computing accelerator card integrates a GPU chip and an Ethernet controller, enabling direct data stream reading, such as... Figure 2As shown, the RDMA communication library needs to call an interface to initiate a network transmission request. The driver in the software stack supports registering the relevant interfaces of the RDMA network card. The RDMA network card will communicate with the GPU kernel driver through the kernel driver standard interface to complete the direct transmission of data and realize direct remote access to the video memory. Only the GPU and the RDMA network card are involved in the entire data path, avoiding redundant jumps of data to system memory and effectively reducing communication latency.

[0087] In this embodiment of the invention, the RDMA network card will communicate with the GPU kernel driver through the kernel driver standard interface to complete the direct transmission of data, reducing the number of data hops by 50% compared with the traditional distributed typical path; the accelerator processor integrates remote direct memory access function, the accelerator card supports dual-port 100G Ethernet and remote direct memory access (RDMA) function, and supports GPU chip and Ethernet through software stack and driver.

[0088] In this embodiment of the invention, the accelerator card further includes a remote direct data access (RDA) network card. The RDA network card communicates with the kernel driver of the accelerator processor via a kernel driver standard interface to complete direct data transmission and remote direct access to video memory. The system's remote direct data access communication software stack is as follows: Figure 3 As shown, this can reduce data transfer latency in the accelerator processor and improve computing performance.

[0089] In this embodiment of the invention, the distributed acceleration server supports two end-to-end communication methods compared to traditional AI servers, such as... Figure 4 As shown. The first end-to-end communication method is high-speed serial computer extended bus standard (peripheral component interconnect express, PCIe) link communication, and the second end-to-end communication method is direct connection communication technology between the switching chip and the Ethernet controller. Through these two end-to-end communication methods, the entire system can achieve horizontal scaling to realize interconnection between systems, thereby achieving larger-scale heterogeneous computing acceleration pooling.

[0090] Based on this heterogeneous accelerator card, a distributed heterogeneous acceleration server is developed. The entire machine can be equipped with a switching chip and a PCIe-compatible heterogeneous acceleration processor module. Servers are interconnected via heterogeneous accelerator cards and switches, forming a hardware interconnection of multiple training servers to achieve cluster computing power.

[0091] Based on any of the above embodiments, the heterogeneous acceleration pooling server system further includes:

[0092] A high-speed serial computer expansion bus standard link is used for interconnecting accelerated processors between multiple servers. The high-speed serial computer expansion bus standard link also includes parameter plane switches, service plane switches, storage plane switches, storage nodes, and service nodes.

[0093] The parameter plane switch is used to cluster and expand the acceleration processors in the heterogeneous computing acceleration resource pool.

[0094] The storage plane switch is used for cluster expansion of the storage nodes;

[0095] The service plane switch is used to cluster the service nodes.

[0096] Heterogeneous acceleration server cluster expansion methods based on high-speed serial computer expansion bus standard links, such as... Figure 5 As shown: Distributed acceleration server GPU nodes achieve cluster expansion through parameter plane switches, with 1:1 non-blocking; the service plane and storage plane cluster networks are connected to separate switches, supporting 1:1 non-blocking or 2:1 convergence; the cluster out-of-band management network is connected to the TOR (Top of Rack) out-of-band management switch. Management data for the cluster is transmitted to the TOR switch via independent physical ports or networks to achieve cluster management and control. Out-of-band management is a management method independent of user service data transmission; it centrally integrates and manages network devices through a dedicated management channel. In a cluster environment, the out-of-band management network is connected to the TOR out-of-band management switch, and cluster management data is transmitted to the TOR switch through a dedicated connection, thereby achieving management and control of the entire cluster. This management method is typically used to ensure high availability, manageability, and security of the cluster. Specifically, the ways in which the out-of-band management network connects to the TOR out-of-band management switch can include:

[0097] Using dedicated physical ports or networks ensures that management data and user business data are transmitted on separate links, preventing management data from being interfered with by business data. Through the management interface provided by the TOR switch, management and monitoring of each node in the cluster are achieved, including configuration, status monitoring, and fault diagnosis. Remote management and maintenance of network devices are possible without interfering with normal business operations. Furthermore, centralized management through the TOR switch significantly improves management efficiency and accuracy, reducing the risk of service interruptions due to manual intervention.

[0098] It should be noted that the cluster expansion topology scheme can be adjusted according to the actual scheme, and this invention does not limit it.

[0099] Based on any of the above embodiments, the general computing resource pool includes:

[0100] Multiple general-purpose computing units, each including a general-purpose processor and memory, are combined using virtualization software to form computing resource containers of arbitrary granularity.

[0101] In this embodiment of the invention, the general computing resource pool is based on I / O centralization, further decoupling CPU and memory to form a centralized computing resource pool. With the addition of software-defined capabilities, massive amounts of CPUs and memory can be virtualized into computing resource containers of arbitrary granularity through virtualization software, forming computing nodes of various sizes and configurations.

[0102] Based on any of the above embodiments, the internal bus switching module includes an I / O switching chip controller, which is used to enumerate all I / O devices, define network topology partitioning according to the resource requirements, dynamically adjust the pairing between the acceleration processor and I / O devices, and establish resource allocation paths.

[0103] This invention fully adapts to the complex characteristics of distributed heterogeneous accelerated pooling servers, providing reliable support for overall machine management and control as well as pooled resource management. It develops a heterogeneous computing power resource pool overall machine management system to improve the availability and manageability of the server system; and develops a heterogeneous computing power resource pooling management engine to achieve the reconfiguration management and dynamic allocation of heterogeneous computing power resources, improving the usability and manageability of pooled resources. It meets the need for unified monitoring and management of ultra-large-scale distributed heterogeneous accelerated pooling servers. It breaks through key technologies such as automatic discovery of complex topologies, centralized asset management, and automatic switching of heterogeneous resources in heterogeneous computing power pooling systems, enabling on-demand rapid deployment, automatic switching, and centralized management of physical resources in heterogeneous computing power systems. The overall architecture of the heterogeneous computing power resource pool management is as follows: Figure 6 As shown.

[0104] In this embodiment of the invention, the system includes:

[0105] A pooling management controller is provided in the general computing power unit in the general computing resource pool and the acceleration card in the heterogeneous computing acceleration resource pool. The pooling management controller manages multiple pooling node management controllers in a unified manner to perform at least one of the following on the general computing resource pool and the heterogeneous computing acceleration resource pool: node topology identification, centralized display of asset information, and coordinated power-on / off.

[0106] The pooled management controller is used for at least one of node topology identification, centralized display of asset information, and coordinated power-on / off.

[0107] In this embodiment of the invention, the pooling management controller is also used to perform at least one of the following functions: integrated monitoring, fault warning, and visual management of the system through a standardized service interface.

[0108] Traditional servers typically use a general-purpose computing unit (GPCU) as the core, coupled with a fixed number of accelerators. At the management level, the accelerator modules are managed by the GPCU's baseboard controller. However, for distributed heterogeneous accelerator pooling servers, GPCU and heterogeneous computing acceleration units are decoupled, pooled, and reconfigured through a high-performance switching unit. This requires that the allocation of GPCU and heterogeneous computing resources be based on demand. Traditional GPCU-centric management architectures cannot perceive and coordinate overall computing resources, and cannot meet requirements such as automatic discovery of complex topologies, centralized asset management, and automatic switching of heterogeneous resources. Therefore, this invention designs a hierarchical heterogeneous computing resource pool management architecture, centered on the Pooled System Management Controller (PSMC) within the high-performance switching unit. Heterogeneous computing acceleration units and GPCU are each paired with a pooled node management controller. The pooled node management controller acts as the central management node, and distributed nodes are uniformly managed by the central management node, achieving full lifecycle management of the heterogeneous computing resource pool.

[0109] The pooling management controller is crucial in the management of distributed heterogeneous accelerated pooled servers. It acts as a bridge for information communication and is responsible for unified management. Its main functions include node topology identification, centralized display of asset information, and coordinated power-on / off. Through standardized service interfaces, it provides operational capabilities, enabling integrated monitoring, fault warnings, and visualized management. The heterogeneous computing resource pool system management system interconnects the pooling node management controllers at all levels via a network, building an independent engine management platform.

[0110] In this embodiment of the invention, the pooling management engine is used to manage the internal bus switching module. By calling the physical channels and software interfaces in the internal bus switching module, it discovers the current network topology to obtain the current asset allocation status, and coordinates the overall computing power resources to switch the network topology to meet the resource requirements based on the physical channels and software interfaces.

[0111] In a heterogeneous computing resource pool system, the discovery, management, and elastic adjustment of heterogeneous computing resources are crucial. This project designs a pooling management engine as the core management unit for dynamic resource adjustment. The pooling management engine will control the high-performance switching unit to achieve automatic resource topology discovery and flexible automatic resource switching, and provide a standard interface for the data center monitoring and management platform to achieve centralized management of massive data center-level resources.

[0112] As a centralized management layer, the pooling management engine enables on-demand allocation and elastic scaling of general-purpose and heterogeneous computing resources by managing and controlling high-performance I / O switching chips.

[0113] Based on any of the above embodiments, for a heterogeneous computing resource system, the internal bus switching module, as a high-performance switching unit, is the infrastructure connecting general-purpose computing units and heterogeneous computing units, and constructing general-purpose resource pools and heterogeneous acceleration resource pools. The pooling management engine acts as the glue between the data center monitoring and management platform and the infrastructure, enabling resource management and switching scheduling. The pooling management engine controls the high-performance switching unit to achieve automatic resource topology discovery and flexible automatic resource switching, and provides a standard interface to the data center monitoring and management platform for centralized resource scheduling and management. The overall hardware and software solution of the pooling management engine is as follows: Figure 7 As shown, the pooling management engine includes:

[0114] Hardware layer, driver layer, functional layer, and interface layer;

[0115] The hardware layer is used to connect the general computing resource pool and the multiple heterogeneous computing acceleration resource pools, and to allocate resources to the general computing resource pool and the multiple heterogeneous computing acceleration resource pools;

[0116] The driver layer is used for data transmission between the hardware layer and the functional layer;

[0117] The functional layer is used to implement the functional requirements of the user;

[0118] The interface layer is used to obtain user demand information and provide users with visual information.

[0119] In embodiments of the present invention, the functional layer is used for, but is not limited to:

[0120] (1) Information interaction is performed between the general asynchronous transceiver and the pooling management controller, and a network interface is provided to visualize the resource list and topology information. The resource list and topology information include at least one of heterogeneous device resource pool information, general computing unit information and I / O port information.

[0121] (2) Establish expert templates for different performance indicators, and switch between different expert templates according to the performance indicator requirements of the application scenario when switching application scenarios;

[0122] In addition, a callable API is provided to display the resource logic topology of the visualization application of the expert template based on the callable API.

[0123] In this embodiment of the invention, the physical link layer binding between heterogeneous computing devices and CPUs is removed, enabling on-demand elastic allocation of heterogeneous computing resources by general-purpose computing units. For different application scenarios, different expert templates for device allocation relationships can be established in the pooling management engine. When switching application scenarios, the topology is automatically switched based on different templates, simplifying the application scenario switching process. Based on the network API interface provided by the pooling management engine, a system interface standard is researched, and a system framework standard is established to implement expert template descriptions and dynamic switching interfaces for heterogeneous computing resources, allowing applications to request resources based on expert templates. Simultaneously, callable APIs are provided to visualize the application of expert templates and to display the logical topology of allocated resources, such as... Figure 8 As shown.

[0124] In this embodiment of the invention, based on parameters such as business requirements, resource status, and performance indicators, an expert template description and dynamic switching interface for heterogeneous computing power resources are implemented, enabling applications to apply for resources based on the expert templates. At the same time, a network interface and a visual WEB interface are provided to realize the visual application of the expert templates.

[0125] (3) Dynamic allocation and adjustment of heterogeneous computing units are achieved through hot removal, hot insertion and hot reset of devices. In this embodiment of the invention, the pooling management engine achieves dynamic allocation and adjustment of heterogeneous computing units at the second level through key technologies such as hot removal, hot insertion and hot reset of devices.

[0126] (4) Dynamically balance and allocate physical resources according to user needs, and release the computing power of heterogeneous computing units. In this embodiment of the invention, the pooling management engine dynamically balances and allocates physical resources based on the load balancing optimization scheduling algorithm, realizes fine-grained on-demand allocation of storage resources, and maximizes the release of heterogeneous computing power.

[0127] In a distributed heterogeneous accelerated pooling server, the general-purpose host resource pool and the heterogeneous accelerated resource pool are interconnected through high-performance switching units. Each high-performance switching unit consists of eight high-performance I / O switching chips. These chips have embedded controllers and provide dedicated physical channels and software interfaces for the pooling management engine to access. The pooling management engine uses these programmable interfaces to achieve unified centralized management and control of all high-performance I / O switching chips. The pooling management engine provides expert templates for different scenarios (such as deep learning and numerical computation), flexibly enables topology switching, and provides a visual UI layer with standard network interfaces.

[0128] Within the heterogeneous computing resource pool, users can utilize the interfaces provided by the pooling management engine to perform operations such as on-demand allocation, dynamic scaling, and resource release of general-purpose and heterogeneous computing units. Taking AI applications as an example, users can match different numbers of heterogeneous computing accelerator cards with general-purpose computing resources according to load requirements. When the AI ​​application stops, users can release the heterogeneous computing resources back to the entire heterogeneous computing resource pool to facilitate efficient resource flow and full utilization.

[0129] Based on any of the above embodiments, in the related technologies, the existing heterogeneous acceleration pooling server system topology is as follows: Figure 9 As shown, most of them use heterogeneous chips to support eight NVIDIA A100 Tensor Core GPUs and two AMD Milan CPUs interconnected by NVIDIA NVLink in a 4U space, while supporting both liquid cooling and air cooling technologies. The NF5688M6 is a GPU server with extreme scalability optimized for large-scale data centers, supporting eight NVIDIA A100 Tensor Core GPUs interconnected by third-generation NVLink and two Intel Ice Lake CPUs, and supporting up to 13 PCIe Gen4 I / O cards.

[0130] In this embodiment of the invention, the heterogeneous acceleration pooling server system is a single-machine 32-card heterogeneous acceleration pooling server system, which includes:

[0131] Two fully interconnected topologies, each containing four switching chips, support expansion to 32 general-purpose processor chip cards;

[0132] The internal interconnect bandwidth of the single-machine 32-card heterogeneous acceleration pooled server system reaches 384GB / s, and the inter-card interconnect bandwidth reaches 128GB / s.

[0133] A single-machine 32-card heterogeneous acceleration server system solution, such as Figure 10 , 11 As shown: The rack can accommodate two 32-card systems, supporting a maximum of 64 GPU cards for expansion; the internal interconnect bandwidth of 16 cards reaches 384GB / s, and the cross-CPU interconnect bandwidth between 16 cards is 128GB / s.

[0134] Distributed heterogeneous accelerated resource pooling servers differ from traditional servers. Centered around multiple data exchange networks and I / O interconnection networks, the system architecture can be divided into general-purpose computing resource pools, high-performance switching units, heterogeneous accelerated computing resource pools, and network switching units according to their logical functions. The general-purpose computing resource pool, based on centralized I / O, further decouples CPU and memory to form a centralized computing resource pool. With software-defined capabilities, it can virtualize massive amounts of CPUs and memory into computing resource containers of arbitrary granularity, forming computing nodes of various sizes and configurations. The high-performance switching unit primarily implements high-performance data exchange functions. By adopting a distributed data exchange architecture, it can achieve flexible network topology partitioning through software definition, quickly and dynamically adjusting the combination between computing and I / O modules, effectively improving the scalability and flexibility of the entire system and ensuring hardware reconfiguration. The heterogeneous accelerated resource pool mainly provides heterogeneous acceleration services, dynamically reconfiguring CPU and GPU resources to form heterogeneous computing server clusters to meet the computing power requirements of high-performance applications.

[0135] Based on any of the above embodiments, the heterogeneous acceleration pooling server system is an 8-host shared 32-card heterogeneous acceleration pooling server system, which includes:

[0136] Two fully interconnected topologies, each containing four switching chips, with eight hosts sharing 32 accelerator processor cards.

[0137] An 8-host, 32-card heterogeneous acceleration server system solution, such as... Figure 12 , 13 As shown, it supports dynamic allocation and rapid deployment of expert templates, enabling multiple hosts to share a heterogeneous computing acceleration resource pool.

[0138] In this embodiment of the invention, a single system supports expansion of 8 / 16 / 32 heterogeneous computing accelerator cards, supports no fewer than 8 domestically produced hosts sharing a heterogeneous computing acceleration resource pool, and supports dynamic resource allocation at the second level.

[0139] Based on any of the above embodiments, the system further includes a computing power virtualization subsystem, which aggregates the heterogeneous computing acceleration resource pool into a unified virtual computing power resource pool. The computing power virtualization subsystem includes:

[0140] Pooled servers, which are computing nodes in the system cluster, are used for task execution;

[0141] A management server, which is used to run management components in the system cluster;

[0142] An application server, which runs applications within a system cluster, can be a container, a virtual machine, or a physical machine.

[0143] In this embodiment of the invention, the distributed heterogeneous accelerated pooling server is a new type of high-performance computing infrastructure. It differs greatly from traditional servers in terms of system architecture, hardware logic, and infrastructure design, and requires a more efficient and flexible management architecture to achieve lifecycle management of pooled resources.

[0144] In this embodiment of the invention, the logical relationship between the heterogeneous computing power software-defined virtualization management system and its subsystems and the pooled server system is as follows: Figure 14 As shown: By aggregating computing resources through a unified virtual computing resource pool, it is possible to support computing tasks to occupy resources during execution and not occupy resources during idle periods; for users, this reduces resource usage costs. For the cluster, it can improve the turnover frequency of cluster resources, providing limited resources to more users. It supports remote access to computing resources, breaking spatial limitations; allowing users to access computing resources on any node in the cluster from any node. It supports virtualization of computing resources, dividing a single computing card into multiple computing cards for use. It matches computing resources with computing tasks, improving resource utilization. It supports the integration of various types of computing resources across nodes for a single computing task. This allows computing tasks to be provided with computing resources far exceeding the physical machine resources. It supports dynamic adjustment of virtual resources. During the execution of computing tasks, the virtual resources used by the computing tasks can be dynamically adjusted, thus supporting functions such as mixed deployment and time-sharing. Mixed deployment schedules different types of tasks to the same physical resources. Through scheduling and resource isolation control measures, while ensuring the service level agreement, it fully utilizes resource capabilities and reduces operating costs. This technology dynamically adjusts resource allocation based on online load by co-deploying online services and computing tasks. For example, when online load is high, computing tasks will relinquish resources to ensure smooth operation of online services; conversely, when online load is low, computing tasks will utilize idle resources to improve resource utilization. Time-sharing is a time-based resource allocation strategy that distributes resources to different applications at different times to adapt to changing demands. Through coordinated time-sharing scheduling in operations and maintenance, resources can be allocated to different applications at different times to reduce resource procurement during peak periods, further saving costs. By employing features such as co-deployment and time-sharing, data center resource utilization can be significantly improved, operating costs reduced, and stable business operations ensured.

[0145] The heterogeneous acceleration pooling server system provided in this invention features a novel heterogeneous computing accelerator card, a distributed heterogeneous acceleration pooling server system, and a management and scheduling method. It supports direct remote access to memory, reducing data hops by 50% compared to typical traditional distributed paths. Based on a heterogeneous computing accelerator card hardware and software prototype verification platform, it optimizes the compatibility of various components such as processors, accelerator cards, networks, and operating systems, ensuring hardware and software compatibility. Breakthroughs have been achieved in the development of the distributed heterogeneous acceleration pooling server system and in hardware and software co-optimization. Research has been conducted on the full lifecycle management and scheduling technology of heterogeneous computing power resource pools, enabling dynamic resource scheduling that adapts to business needs. It supports dynamic adjustment of processor / accelerator ratios, topology discovery, and on-demand switching to meet diverse application load requirements. Remote invocation of heterogeneous computing power resources based on Ethernet is implemented, supporting TCP and RDMA data transmission protocols. Furthermore, it achieves virtualization and dynamic scaling management of heterogeneous computing power resources.

[0146] This invention, application-oriented and centered on system architecture design, utilizes a multi-layered collaborative design encompassing "system-component-chip-software" to research and develop a high-performance, high-efficiency distributed heterogeneous acceleration pooling server system. This system achieves unified management and scheduling of heterogeneous computing resources, fully leveraging the innovative advantages of diversified computing power, making computing power user-friendly and easy to use. This is crucial for promoting the widespread application of artificial intelligence technology and the development of my country's computing power industry. Mastering the core technologies of distributed heterogeneous acceleration pooling servers will lead to the formation of an independent technology system. Heterogeneous computing is the primary computing architecture for future data center computing infrastructure. Current lags in research on key heterogeneous computing technologies result in dependence on external systems for crucial computing infrastructures such as intelligent computing data centers and supercomputing data centers, limiting their ability to expand and manage computing power. This invention aims to drive research into the core technologies of distributed heterogeneous acceleration pooling servers between academia and enterprises, forming an independent technology system. Breakthroughs will be achieved in key technologies for heterogeneous computing accelerator cards, distributed heterogeneous acceleration pooling servers, heterogeneous computing resource pool lifecycle management and scheduling technologies, and heterogeneous computing resource virtualization and remote scheduling technologies, constructing an independent and controllable industrial chain, value chain, and ecosystem. Developing independently controllable heterogeneous acceleration chips, heterogeneous acceleration cards, distributed heterogeneous acceleration pooled servers, and related supporting software, and conducting demonstration applications in large-scale data centers, will help promote the development of the computing power industry. It will also drive the coordinated development of upstream and downstream industrial chains, creating massive economic benefits. Furthermore, it will encourage more enterprises and research institutions to increase investment in core components and equipment for heterogeneous computing, providing strong support for the development of next-generation information technology.

[0147] The pooled resource allocation method provided by the present invention is described below. The pooled resource allocation method described below can be referred to in correspondence with the heterogeneous accelerated pooled server system described above.

[0148] The pooling resource allocation method provided by this invention is applicable to the heterogeneous accelerated pooling server system described in any of the above claims, and includes:

[0149] Step 101: A general computing unit in the general computing resource pool initiates a request for dynamic adjustment of heterogeneous computing resources to the pooling management controller based on resource requirements;

[0150] Step 102: The pooling management engine determines the IO devices to be accessed in the internal bus switching module according to the resource requirements, accesses the IO devices in one or more high-speed data interconnect modules through a globally unique code, reconstructs the heterogeneous computing resource allocation path, and matches the heterogeneous computing resources corresponding to the resource requirements for the general computing unit.

[0151] The heterogeneous computing resources include resources in a single acceleration resource pool, resources across acceleration resource pools, and resources across pooled servers.

[0152] In this embodiment of the invention, for a heterogeneous computing resource system, the internal bus switching module, as a high-performance switching unit, serves as the infrastructure connecting general-purpose computing units and heterogeneous computing units, and constructing general-purpose resource pools and heterogeneous acceleration resource pools. The pooling management engine acts as the glue between the data center monitoring and management platform and the infrastructure, enabling resource management and switching scheduling. The pooling management engine controls the high-performance switching unit to achieve automatic resource topology discovery and flexible automatic resource switching, and provides a standard interface to the data center monitoring and management platform for centralized resource scheduling and management.

[0153] In this embodiment of the invention, the distribution of resources is divided into three cases from the perspective of resource distribution, which can effectively utilize resources. Based on this, a scheduling system based on heterogeneous distance is developed. According to the scheduling strategy, it supports the heterogeneous resource scheduling capability of a single acceleration resource pool, across acceleration resource pools, and across pooled servers, thereby meeting the computing power requirements of different applications for pooled servers.

[0154] This invention fully adapts to the complex characteristics of distributed heterogeneous accelerated pooling servers, providing reliable support for overall machine management and control as well as pooled resource management. It develops a heterogeneous computing power resource pool overall machine management system to improve the availability and manageability of the server system; and it develops a heterogeneous computing power resource pooling management engine to achieve the reconstruction management and dynamic allocation of heterogeneous computing power resources, improving the usability and manageability of pooled resources. It meets the need for unified monitoring and management of ultra-large-scale distributed heterogeneous accelerated pooling servers. It breaks through key technologies such as automatic discovery of complex topologies, centralized asset management, and automatic switching of heterogeneous resources in heterogeneous computing power pooling systems, enabling on-demand rapid deployment, automatic switching, and centralized management of physical resources in heterogeneous computing power systems.

[0155] Figure 15A timing diagram of the pooled resource allocation method provided in the embodiments of the present invention is shown below. Figure 15 As shown, the pooled resource allocation method provided in this embodiment of the invention includes:

[0156] (0) Taking the release of a heterogeneous accelerator card device and its allocation to a general computing unit as an example, before the user initiates a request for dynamic adjustment of heterogeneous computing power resources through the pooling management engine, it is necessary to ensure that the application layer process related to the heterogeneous computing power device has ended, so as to avoid abnormal access of the device by the application layer and cause program abnormality.

[0157] (1) The pooling management engine completes the hot removal of the device;

[0158] (2) The pooling management engine sends a request to the high-performance switching unit to obtain the physical location of the device;

[0159] (3) The pooling management engine confirms the physical location of the device;

[0160] (4) The pooling management engine resets the resources of the heterogeneous computing power device and retrains it, so that the running state is restored to the default value;

[0161] (5) The pooling management engine confirms that the device has been reset.

[0162] (6) The pooling management engine reassigns the device to another general computing unit. The computing unit will see the new device without the business being aware of it, thus completing the dynamic switching of heterogeneous computing resources.

[0163] This invention enables cross-node, multi-host sharing, on-demand resource allocation and elastic application, maximizing the release of heterogeneous computing power and achieving performance optimization in different application scenarios.

[0164] In this embodiment of the invention, the pooled resource allocation method further includes:

[0165] The management and control service receives a resource request initiated by the heterogeneous resource management communication component in the application server, and sends the configuration to the switching network according to the resource request;

[0166] The configuration is distributed to the heterogeneous virtual computing power management node component in the pooling server through the search and exchange network, so that the pooling server loads the configuration to complete the virtualization deployment of computing power resources.

[0167] In this embodiment of the invention, the heterogeneous intelligent computing power virtualization subsystem aggregates the accelerated computing resource pools on the pooled servers into a unified virtual computing power resource pool. Based on this unified virtual computing power resource pool, the high-speed network heterogeneous computing power resource remote invocation subsystem, the heterogeneous virtual computing power management subsystem, and the heterogeneous computing cluster resource allocation and system performance optimization subsystem are responsible for managing and using the unified virtual computing power. The deployment of the heterogeneous computing power software-defined virtualization management system requires pooled servers, control servers, and application servers. The specific deployment locations and invocation chains of the subsystem modules are as follows: Figure 16 As shown.

[0168] In this embodiment of the invention, for the cluster, the turnover frequency of cluster resources can be increased, providing limited resources to more users. It supports remote access to computing resources, breaking spatial limitations; allowing users to access computing resources on any node in the cluster from any node. It supports virtualization of computing resources, dividing a single computing card into multiple computing cards. It matches computing resources to computing tasks, improving resource utilization. It supports integrating various types of computing resources across nodes for a single computing task. Thus, it can provide computing resources to computing tasks far exceeding the physical machine resources.

[0169] Distributed acceleration server GPU nodes achieve cluster expansion through parametric plane switches, achieving 1:1 non-blocking. The service plane and storage plane cluster networks are connected to separate switches, supporting 1:1 non-blocking or 2:1 convergence. The cluster out-of-band management network is connected to a TOR (Top of Rack) out-of-band management switch. Management data for the cluster is transmitted to the TOR switch via independent physical ports or networks to achieve cluster management and control. Out-of-band management is a management method independent of user service data transmission; it centrally integrates and manages network devices through a dedicated management channel. In a cluster environment, the out-of-band management network is connected to the TOR out-of-band management switch, and cluster management data is transmitted to the TOR switch through a dedicated connection, thereby achieving management and control of the entire cluster. This management method is typically used to ensure high availability, manageability, and security of the cluster.

[0170] During computational tasks, the virtual resources used can be dynamically adjusted to support features such as mixed deployment and time-sharing. Mixed deployment schedules different types of tasks onto the same physical resources. Through scheduling and resource isolation control measures, it fully utilizes resource capabilities and reduces operating costs while ensuring service level agreements (SLAs). This technology dynamically adjusts resource allocation based on online load. For example, when online load is high, computational tasks will relinquish their resource usage to ensure smooth operation of online services; conversely, when online load is low, computational tasks will utilize idle resources to improve resource utilization. Time-sharing is a time-based resource allocation strategy that distributes resources to different applications at different times to adapt to changing demands. With coordinated time-sharing scheduling by operations and maintenance, resources can be allocated to different applications at different times to reduce resource procurement during peak periods, further saving costs. Mixed deployment and time-sharing can significantly improve data center resource utilization, reduce operating costs, and ensure stable business operations.

[0171] The pooled resource allocation method provided in this invention interconnects servers between systems through a Ethernet switching module; it adjusts network port configuration and heterogeneous computing resource allocation paths according to resource requirements so that the general computing resource pool matches one or more heterogeneous computing acceleration resource pools corresponding to the resource requirements; it supports dynamic adjustment of processor / accelerator ratio, topology discovery, and on-demand switching; it can achieve adaptive resource dynamic scheduling to meet diverse application load requirements; and it improves the operating efficiency and performance of the heterogeneous acceleration pooled server system.

[0172] On the other hand, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the pooled resource allocation method provided by the above methods. The method includes: a general computing power unit in a general computing resource pool initiates a dynamic adjustment request for heterogeneous computing power resources to a pooling management controller based on resource requirements; adjusting network port configuration and heterogeneous computing resource allocation path through a pooling management engine to match heterogeneous computing power resources corresponding to the resource requirements for the general computing power unit; resetting and retraining the matched heterogeneous computing power resources; and allocating the trained acceleration processor to the general computing power unit to complete the dynamic switching of heterogeneous computing power resources; wherein the heterogeneous computing power resources include resources in a single acceleration resource pool, resources across acceleration resource pools, and resources across pooling servers.

[0173] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0174] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0175] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A heterogeneous accelerated pooling server system, characterized in that, include: General-purpose computing resource pool, internal bus switching module, Ethernet switching module, and multiple heterogeneous computing acceleration resource pools; The general computing resource pool is connected to the internal bus switching module via a first system interconnect bus; each heterogeneous computing acceleration resource pool is connected to the internal bus switching module via a second system interconnect bus; each heterogeneous computing acceleration resource pool is connected to the Ethernet switching module, which is used for interconnection between servers in the system. The internal bus switching module includes multiple interconnected high-speed data interconnect modules. Each high-speed data interconnect module is equipped with multiple I / O devices. The I / O devices of the multiple high-speed data interconnect modules are uniformly assigned globally unique codes, so that any I / O device in any high-speed data interconnect module can be accessed through the globally unique codes. After the general computing resource pool obtains the resource requirements, it configures the required I / O devices according to the resource requirements through the internal bus switching module, and establishes a heterogeneous computing resource allocation path so that the general computing resource pool matches one or more heterogeneous computing acceleration resource pools corresponding to the resource requirements. The heterogeneous computing acceleration resource pool includes: At least one accelerator card, wherein the accelerator card integrates an accelerator processor and an Ethernet controller; The acceleration processor is directly connected to the Ethernet controller; The Ethernet controller is connected to the Ethernet switching module; The system also includes a remote direct data access network card (RDBNIC), which communicates with the kernel driver of the accelerator processor through a kernel driver standard interface. The RDBNIC is connected to the Ethernet switching module so that multiple accelerator processors connected to multiple RDBNICs can complete direct data transmission and remote direct access to video memory.

2. The heterogeneous accelerated pooling server system according to claim 1, characterized in that, The system also includes: A high-speed serial computer expansion bus standard link is used for accelerated processor interconnection between multiple servers.

3. The heterogeneous accelerated pooling server system according to claim 1, characterized in that, The internal bus switching module includes an I / O switching chip controller, which is used to enumerate all I / O devices, define network topology partitioning according to the resource requirements, dynamically adjust the pairing between the acceleration processor and I / O devices, and establish resource allocation paths.

4. The heterogeneous accelerated pooling server system according to claim 1, characterized in that, The globally unique code is obtained by encoding the high-speed data interconnect module number and its internal IO device number using a segmented encoding method.

5. The heterogeneous accelerated pooling server system according to claim 1, characterized in that, The general computing resource pool includes: Multiple general-purpose computing units, each including a general-purpose processor and memory, are combined using virtualization software to form computing resource containers of arbitrary granularity.

6. The heterogeneous accelerated pooling server system according to claim 1, characterized in that, The system also includes a pooling management controller; The general computing power unit in the general computing resource pool and the acceleration card in the heterogeneous computing acceleration resource pool are respectively equipped with a pooled node management controller. The pooled management controller manages multiple pooled node management controllers in a unified manner to perform at least one of the following on the general computing resource pool and the heterogeneous computing acceleration resource pool: node topology identification, centralized display of asset information, and coordinated power-on / off.

7. The heterogeneous accelerated pooling server system according to claim 6, characterized in that, The pooled management controller is also used for at least one of the following: integrated monitoring, fault warning, and visual management of the system through a standardized service interface.

8. The heterogeneous accelerated pooling server system according to claim 1, characterized in that, The system also includes a pooling management engine; The pooling management engine is used to manage the internal bus switching module. By calling the physical channels and software interfaces in the internal bus switching module, it discovers the current network topology to obtain the current asset allocation status, and coordinates the overall computing power resources to switch the network topology to meet the resource requirements based on the physical channels and software interfaces.

9. The heterogeneous accelerated pooling server system according to claim 8, characterized in that, The pooling management engine is also used for: Hot removal, hot insertion, and hot reset processes are performed on any accelerator card in the heterogeneous computing acceleration resource pool to control whether the accelerator card is put into use.

10. The heterogeneous accelerated pooling server system according to claim 8, characterized in that, The pooling management engine is also used for: Based on a load balancing-based optimized scheduling algorithm, the accelerator cards in the heterogeneous computing acceleration resource pool are dynamically allocated and adjusted.

11. The heterogeneous accelerated pooling server system according to claim 8, characterized in that, The pooling management engine is also used for: Establish expert templates for different application scenarios; When switching application scenarios, different expert templates are switched by calling the API according to the current application scenario; the expert templates that meet the requirements of the application scenario are displayed visually through a visual web interface.

12. The heterogeneous accelerated pooling server system according to claim 1, characterized in that, The heterogeneous acceleration pooling server system is a single-machine heterogeneous acceleration pooling server system. In the single-machine heterogeneous acceleration pooling server system, a single host supports expansion with 8 acceleration cards, 16 acceleration cards, and 32 acceleration cards.

13. The heterogeneous accelerated pooling server system according to claim 1, characterized in that, The heterogeneous acceleration pooling server system is a multi-host heterogeneous acceleration pooling server system. In the multi-host heterogeneous acceleration pooling server system, each host is connected to a high-speed data interconnect module, and 32 acceleration processor cards are shared through multiple high-speed data interconnect modules.

14. The heterogeneous accelerated pooling server system according to claim 1, characterized in that, The system further includes a computing power virtualization subsystem, which aggregates the heterogeneous computing acceleration resource pools into a unified virtual computing power resource pool. The computing power virtualization subsystem includes: Pooled servers, which are computing nodes in the system cluster, are used for task execution; A management server, which is used to run management components in the system cluster; An application server, which runs applications within a system cluster, can be a container, a virtual machine, or a physical machine.

15. A pooled resource allocation method, applicable to the heterogeneous accelerated pooled server system according to any one of claims 1 to 14, characterized in that, include: A general computing unit in the general computing resource pool initiates a request for dynamic adjustment of heterogeneous computing resources to the pooling management controller based on resource requirements; The pooling management engine determines the IO devices to be accessed in the internal bus switching module based on the resource requirements, accesses the IO devices in one or more high-speed data interconnect modules through a globally unique code, reconstructs the heterogeneous computing resource allocation path, and matches the heterogeneous computing resources corresponding to the resource requirements for the general computing unit. The heterogeneous computing resources include resources in a single acceleration resource pool, resources across acceleration resource pools, and resources across pooled servers.

16. The pooled resource allocation method according to claim 15, characterized in that, The process of reconstructing the heterogeneous computing resource allocation path includes releasing the accelerator card in a general-purpose computing unit and allocating it to other general-purpose computing units. Specifically, the pooling management engine performs hot removal of the accelerator card and sends a request to the internal bus switching module to obtain the physical location of the accelerator card. The accelerator card is located based on its physical location, and then reset. The reset accelerator cards are reassigned to the general-purpose computing units to be assigned.

17. The pooled resource allocation method according to claim 16, characterized in that, The step of releasing the accelerator card in a general-purpose computing unit and allocating it to other general-purpose computing units also includes: The accelerator card is trained according to the application scenario of the general computing unit to be allocated; The trained accelerator cards are reassigned to the general computing units to be assigned.

18. A non-transitory readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the pooled resource allocation method as described in any one of claims 15 to 17.