Heterogeneous accelerated pooling server system, pooling resource allocation method, and medium

By using a heterogeneous accelerated pooled server system, I/O devices and resource allocation paths are dynamically configured, achieving high-efficiency computing performance of heterogeneous accelerated servers in large-scale and diversified applications. This solves the problem of limited scalability of traditional heterogeneous accelerated server systems and improves operational efficiency and resource management efficiency.

WO2026065965A1PCT designated stage Publication Date: 2026-04-02INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Traditional heterogeneous accelerated server systems suffer from limited scalability, closed ecosystems, energy waste, high upgrade and expansion costs, and low operating efficiency when facing large-scale and diversified applications, thus failing to meet the demands of high-performance computing.

Method used

The system employs a heterogeneous accelerated pooled server system, including a general computing resource pool, an internal bus switching module, an Ethernet switching module, and multiple heterogeneous computing accelerated resource pools. It dynamically configures I/O devices through globally unique codes and internal bus switching modules, establishes heterogeneous computing resource allocation paths, supports direct connection between accelerated processors and Ethernet controllers to achieve direct data transmission, and uses high-speed serial computer extended bus standard links and remote direct data access network cards for interconnection. It combines a pooled management controller and engine for unified resource management and dynamic allocation.

Benefits of technology

It improves system scalability and operating efficiency, enables high-efficiency computing performance of heterogeneous acceleration servers in large-scale and diversified applications, reduces communication latency and energy consumption, supports the integration of multi-platform processor computing power and the collaborative scheduling of heterogeneous resources, and simplifies resource management and expansion processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025082427_02042026_PF_FP_ABST
    Figure CN2025082427_02042026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computers and provides a heterogeneous accelerated pooling server system, a pooling resource allocation method, and a medium. The heterogeneous accelerated pooling server system comprises a general-purpose computing resource pool, an internal bus switching module, an Ethernet switching module, and a plurality of heterogeneous accelerated computing resource pools. The internal bus switching module comprises a plurality of interconnected high-speed data interconnection modules, each module is provided with a plurality of IO devices, and the IO devices of the plurality of high-speed data interconnection modules are collectively allocated globally unique codes. After the general-purpose computing resource pool acquires a resource requirement, a required IO device is configured on the basis of the resource requirement, and a heterogeneous computing resource allocation path is established, so that the general-purpose computing resource pool matches one or more heterogeneous accelerated computing resource pools corresponding to the resource requirement, thereby enabling each general-purpose computing power unit to access any heterogeneous acceleration device, improving the system expansion capability, and improving the operating efficiency of a heterogeneous acceleration server when dealing with large-scale and diversified applications.
Need to check novelty before this filing date? Find Prior Art

Description

Heterogeneous acceleration pooling server system, pooling resource allocation method and medium

[0001] Cross-reference to Related Applications

[0002] This application claims priority to the Chinese patent application No. 202411390541.X, filed on September 30, 2024, and entitled "Heterogeneous acceleration pooling server system, pooling resource allocation method and medium", the entire content of which is incorporated herein by reference. TECHNICAL FIELD

[0003] The present application relates to the field of computer technology, and in particular to a heterogeneous acceleration pooling server system, a pooling resource allocation method and a medium. BACKGROUND

[0004] With the development of computer applications such as artificial intelligence and big data, the demand for computing power is growing rapidly. At the same time, with the gradual diversification of computing application scenarios, the emergence of different application fields such as machine learning, artificial intelligence, unmanned driving and industrial internet, traditional general-purpose processors encounter more and more performance bottlenecks in mass processing and computing, mass data / pictures, such as low computing parallelism, low bandwidth, long latency, etc. The traditional server architecture with single-core processors as the computing core cannot meet today's needs. In order to efficiently process various computing tasks, different types of heterogeneous acceleration processors have emerged. Traditional heterogeneous acceleration services are achieved by simply integrating different specifications and models of existing processors. The traditional heterogeneous acceleration service architecture adopts an I / O (Input / Output) tree structure. The host unit enumerates the I / O bridges or I / O devices in the I / O tree structure based on a depth-first enumeration method during the startup process, and assigns a unique ID (Identity Document) number to each I / O tree. The routing path can only be allocated within each I / O tree, which limits the diversified allocation and use of I / O device resources. The system expansion capability is limited, the ecology is closed, the energy consumption is wasted, and the upgrade and expansion cost is high. When dealing with large-scale and diversified applications, the traditional heterogeneous acceleration server has low running efficiency. SUMMARY

[0005] The present application provides a heterogeneous acceleration pooling server system, a pooling resource allocation method and a medium to solve the defects of low running efficiency, performance limitation and restriction of commercial application of high-performance computing technology of the traditional heterogeneous acceleration pooling server system.

[0006] The present application provides a heterogeneous acceleration pooling server system, comprising:

[0007] a general computing resource pool, an internal bus switching module, an Ethernet switching module, and a plurality of heterogeneous computing acceleration resource pools;

[0008] The general computing resource pool is communicatively connected with the internal bus switching module through a first system internal interconnection bus.

[0009] Each of the plurality of heterogeneous computing acceleration resource pools is communicatively connected with the internal bus switching module through a second system internal interconnection bus.

[0010] The internal bus switching module comprises a plurality of interconnected high-speed data interconnection modules, each of which is provided with a plurality of IO (Input Output) devices, and the IO devices of the plurality of high-speed data interconnection modules are collectively assigned a globally unique code for accessing any IO device in any high-speed data interconnection module through the globally unique code; after the general computing resource pool obtains resource requirements, the internal bus switching module configures the required IO devices according to the resource requirements to establish a heterogeneous computing resource allocation path to match the general computing resource pool with one or more of the plurality of heterogeneous computing acceleration resource pools corresponding to the resource requirements.

[0011] According to the heterogeneous acceleration pooled server system provided by the present application, the heterogeneous computing acceleration resource pool comprises:

[0012] At least one acceleration card, an acceleration processor and an Ethernet controller integrated in the acceleration card;

[0013] The acceleration processor is directly connected with the Ethernet controller.

[0014] The Ethernet controller is connected with the Ethernet switching module.

[0015] According to the heterogeneous acceleration pooled server system provided by the present application, the system further comprises:

[0016] A high-speed serial computer expansion bus standard link, which is used for interconnection of acceleration processors among the plurality of servers.

[0017] According to the heterogeneous acceleration pooled server system provided by the present application, the system further comprises a remote direct data access network card, which intercommunicates with the kernel driver of the acceleration processor through a kernel driver standard interface,

[0018] The remote direct data access network card is connected with the Ethernet switching module to enable the plurality of acceleration processors connected by the plurality of remote direct data access network cards to complete direct transmission of data and remote direct access of display memory.

[0019] According to the heterogeneous acceleration pooled server system provided in the application, the internal bus exchange module comprises an I / O exchange controller, the I / O exchange controller is used for enumerating all I / O devices, and the network topology division is defined according to resource requirements, the matching between the acceleration processors and the I / O devices is dynamically adjusted, and the resource allocation path is established.

[0020] According to the heterogeneous acceleration pooled server system provided in the application, the global unique code is numbered by the high-speed data interconnection module and the IO device number in the high-speed data interconnection module, and is obtained by using a segmented coding mode.

[0021] According to the heterogeneous acceleration pooled server system provided in the application, the general computing resource pool comprises:

[0022] A plurality of general computing units, each general computing unit comprising a general processor and a memory, and the general processor and the memory are combined in any number by virtualization software to form a computing resource container of any granularity.

[0023] According to the heterogeneous acceleration pooled server system provided in the application, the system further comprises a pooled management controller;

[0024] The general computing unit in the general computing resource pool and the acceleration card in the heterogeneous computing acceleration resource pool are respectively provided with a pooled node management controller, the pooled management controller uniformly manages a plurality of pooled node management controllers, and at least one of node topology identification, asset information centralized display and cooperative power-on and power-off of the general computing resource pool and the heterogeneous computing acceleration resource pool is performed.

[0025] According to the heterogeneous acceleration pooled server system provided in the application, the pooled management controller is further used for at least one of integrated monitoring, fault early warning and visual management of the system through a standardized service interface.

[0026] According to the heterogeneous acceleration pooled server system provided in the application, the system further comprises a pooled management engine;

[0027] The pooled management engine is used for managing the internal bus exchange module, discovering the current network topology by calling the physical channel and the software interface in the internal bus exchange module, obtaining the current asset allocation, and adjusting the integrated computing resource switching network topology according to the physical channel and the software interface to meet the resource requirements.

[0028] According to the heterogeneous acceleration pooled server system provided in the application, the pooled management engine is further used for:

[0029] The hot removal, hot insertion and hot reset processing of any acceleration card in the heterogeneous computing acceleration resource pool are performed to control whether the acceleration card is put into use.

[0030] According to the heterogeneous acceleration pooled server system provided in the application, the pooling management engine is further used for:

[0031] The optimization scheduling algorithm based on load balancing is used for dynamically allocating and adjusting the acceleration cards in the heterogeneous computing acceleration resource pool.

[0032] According to the heterogeneous acceleration pooled server system provided in the application, the pooling management engine is further used for:

[0033] Expert templates in different application scenarios are established.

[0034] Different expert templates are switched according to the current application scenario by calling an API (Application Programming Interface) when the application scenario is switched; and the expert templates meeting the application scenario requirements are visually displayed through a visual WEB (World Wide Web) interface.

[0035] According to the heterogeneous acceleration pooled server system provided in the application, the heterogeneous acceleration pooled server system is a single-machine heterogeneous acceleration pooled server system, and a single host in the single-machine heterogeneous acceleration pooled server system supports 8 acceleration card extensions, 16 acceleration card extensions and 32 acceleration card extensions.

[0036] According to the heterogeneous acceleration pooled server system provided in the application, the heterogeneous acceleration pooled server system is a multi-host heterogeneous acceleration pooled server system, and each host in the multi-host heterogeneous acceleration pooled server system is connected with a high-speed data interconnection module, and 32 acceleration processor cards are shared through the multiple high-speed data interconnection modules.

[0037] According to the heterogeneous acceleration pooled server system provided in the application, the system further comprises a computing power virtualization subsystem, the computing power virtualization subsystem aggregates the heterogeneous computing acceleration resource pool into a unified virtual computing power resource pool, and the computing power virtualization subsystem comprises:

[0038] The pooled server is a computing node in the system cluster and is used for task running.

[0039] The management and control server is used for running a management and control component in the system cluster.

[0040] The application server is used for running an application program in the system cluster, and the application server is a container, a virtual machine or a physical machine.

[0041] The application further provides a pooled resource allocation method, which is suitable for the heterogeneous acceleration pooled server system in any of the above aspects, and comprises the following steps:

[0042] A general computing resource pool initiates a heterogeneous computing resource dynamic adjustment request to the pool management controller according to resource requirements;

[0043] The pool management engine determines an IO device to be accessed in the internal bus switching module according to resource requirements, accesses the IO device in one or more high-speed data interconnection modules through a globally unique code, reconstructs a heterogeneous computing resource allocation path, and matches the general computing unit with the heterogeneous computing resource corresponding to the resource requirements.

[0044] The heterogeneous computing resource includes a resource in a single acceleration resource pool, a resource across acceleration resource pools, and a resource across pool servers.

[0045] According to the pool resource allocation method provided in the application, when reconstructing the heterogeneous computing resource allocation path, the acceleration card in a certain general computing unit is released and distributed to other general computing units, specifically including: the pool management engine sends a request to the internal bus switching module after hot removing the acceleration card, and obtains the physical location of the acceleration card.

[0046] According to the physical location, the acceleration card is found and reset,

[0047] The reset acceleration card is redistributed to the general computing unit to be distributed.

[0048] According to the pool resource allocation method provided in the application, the acceleration card in a certain general computing unit is released and distributed to other general computing units, and further includes:

[0049] The reset acceleration card is trained according to the application scenario of the general computing unit to be distributed;

[0050] The trained acceleration card is redistributed to the general computing unit to be distributed.

[0051] The application also provides a computer non-volatile readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the pool resource allocation method of any one of the above.

[0052] The heterogeneous acceleration pooling server system, the pooling resource allocation method and the medium provided by the application, the heterogeneous acceleration pooling server system comprises a general computing resource pool, an internal bus exchange module, an Ethernet exchange module and a plurality of heterogeneous computing acceleration resource pools; the general computing resource pool is in communication connection with the internal bus exchange module through a first system internal interconnection bus; each heterogeneous computing acceleration resource pool is in communication connection with the internal bus exchange module through a second system internal interconnection bus; the internal bus exchange module comprises a plurality of interconnected high-speed data interconnection modules, a plurality of IO devices are arranged on each high-speed data interconnection module, the IO devices of the plurality of high-speed data interconnection modules are allocated a globally unique code in an overall plan, so as to access any IO device in any high-speed data interconnection module through the globally unique code; after the general computing resource pool obtains resource requirements, the required IO devices are configured by the internal bus exchange module according to the resource requirements, the heterogeneous computing resource allocation path is established to make the general computing resource pool match one or more heterogeneous computing acceleration resource pools corresponding to the resource requirements, by allocating a globally unique code for the plurality of IO devices on each high-speed data interconnection module, each general computing unit can access any heterogeneous acceleration device, the system expansion capability is improved, and when large-scale and diversified applications are dealt with, the running efficiency of the heterogeneous acceleration server is improved. BRIEF DESCRIPTION OF DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the application or the related art, the drawings needed to be used in the embodiments or the related art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0054] Fig. 1 is a functional structure schematic diagram of the heterogeneous acceleration pooling server system provided by the embodiment of the application;

[0055] Fig. 2 is a heterogeneous acceleration card data flow schematic diagram provided by the embodiment of the application;

[0056] Fig. 3 is a heterogeneous acceleration card RDMA communication data flow schematic diagram provided by the embodiment of the application;

[0057] Fig. 4 is a distributed heterogeneous acceleration server end-to-end communication schematic diagram provided by the embodiment of the application;

[0058] Fig. 5 is a distributed heterogeneous acceleration server cluster expansion schematic diagram provided by the embodiment of the application;

[0059] Fig. 6 is a heterogeneous computing resource pool management architecture schematic diagram provided by the embodiment of the application;

[0060] Fig. 7 is a heterogeneous computing resource pooling management engine architecture schematic diagram provided by the embodiment of the application;

[0061] FIG. 8 is a schematic diagram of a heterogeneous computing resource pool expert template mechanism according to an embodiment of the present application;

[0062] FIG. 9 is a schematic diagram of a prior heterogeneous acceleration server system structure according to an embodiment of the present application;

[0063] FIG. 10 is a schematic diagram of a single-machine 32-card heterogeneous acceleration server system structure according to an embodiment of the present application;

[0064] FIG. 11 is a schematic diagram of a single-cabinet 64-card heterogeneous acceleration server structure according to an embodiment of the present application;

[0065] FIG. 12 is a schematic diagram of a multi-machine shared 32-card heterogeneous acceleration server structure according to an embodiment of the present application;

[0066] FIG. 13 is a schematic diagram of a multi-machine shared 32-card heterogeneous acceleration server structure according to an embodiment of the present application;

[0067] FIG. 14 is a relationship diagram of heterogeneous computing power software-defined virtualization and pooling servers according to an embodiment of the present application;

[0068] FIG. 15 is a schematic diagram of a heterogeneous computing power resource device dynamic allocation process according to an embodiment of the present application;

[0069] FIG. 16 is a schematic diagram of a heterogeneous computing power software-defined virtualization deployment process according to an embodiment of the present application. DETAILED DESCRIPTION

[0070] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions in the present application will be described below in conjunction with the accompanying drawings in the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.

[0071] FIG. 1 is a functional structure diagram of a heterogeneous acceleration pooling server system according to an embodiment of the present application. As shown in FIG. 1, the heterogeneous acceleration pooling server system according to an embodiment of the present application includes:

[0072] a general computing resource pool, an internal bus switching module, an Ethernet switching module, and a plurality of heterogeneous computing acceleration resource pools;

[0073] The general computing resource pool is in communication connection with the internal bus switching module through a first system internal interconnection bus;

[0074] Each of the heterogeneous computing acceleration resource pools is in communication connection with the internal bus switching module through a second system internal interconnection bus;

[0075] Each of the heterogeneous computing acceleration resource pools is in communication connection with the Ethernet switch module;

[0076] The internal bus switch module comprises a plurality of interconnected high-speed data interconnection modules, each of which is provided with a plurality of IO devices, and the IO devices of the plurality of high-speed data interconnection modules are collectively assigned a globally unique code for accessing any IO device in any high-speed data interconnection module through the globally unique code; after the general computing resource pool obtains the resource requirement, the required IO device is configured by the internal bus switch module according to the resource requirement, and a heterogeneous computing resource allocation path is established to enable the general computing resource pool to match one or more heterogeneous computing acceleration resource pools corresponding to the resource requirement.

[0077] In the embodiments of the present application, the globally unique code is obtained by using a segmented coding mode to code the number of the high-speed data interconnection module and the number of the IO device inside the high-speed data interconnection module.

[0078] In the embodiments of the present application, the general computing resource pool is a resource pool composed of general processors such as CPUs (Central Processing Unit, central processor), the heterogeneous computing acceleration resource pool is a resource pool composed of heterogeneous processors such as GPUs (Graphics Processing Unit, graphics processor), the plurality of high-speed data interconnection modules in the internal bus switch module are, for example, four high-performance switch boards, each of which expands 10 high-speed interfaces (IO devices) externally, the general computing resource pool is composed of general computing units, mainly providing the computing function of the system, the heterogeneous computing acceleration resource pool is composed of heterogeneous computing units, mainly providing the acceleration computing function of the entire system, achieving higher computing performance, parallelism, and energy efficiency ratio; the internal bus switch module is a high-performance switching unit, comprising high-speed data interconnection modules, which are closely combined with the basic software, realizing physical decoupling of general computing and heterogeneous acceleration computing, and at the same time realizing efficient configuration management and resource cross-node direct sharing of the heterogeneous acceleration card, improving the heterogeneous resource access efficiency, and providing a high-reliability, high-performance data interconnection network for the entire system.

[0079] In the embodiments of the present application, the heterogeneous acceleration pooling server system adopts a distributed architecture, faces diversified application requirements, changes the traditional design centered on CPU, and through key technologies such as hardware decoupling and consistent high-speed interconnection, takes the internal bus exchange module as the core, decouples and restructures the traditional server architecture, realizes collaborative computing of multiple general-purpose processor platforms and multiple heterogeneous computing acceleration units, realizes hardware decoupling, pooling and restructuring of large-scale computing resources, memory resources, heterogeneous acceleration resources and storage resources, and realizes dynamic scheduling of resources through software-defined system design. The system resource pool is integrated through unified management, heat dissipation and power supply to form a high-performance host with high-heterogeneous computing power and flexible scheduling and allocation of resources. The whole system can flexibly provide multiple computing power for multiple application scenarios and realize the fusion of multiple computing power through system-level architecture key technology innovation under the condition of limited key components and devices, and improve computing performance.

[0080] In the embodiments of the present application, the acceleration processor is a graphics processing unit (GPU), a field programmable gate array (FPGA) or the like, and through the heterogeneous computing acceleration resource pool interconnection and software-defined system design, the system topology can be dynamically adjusted to realize on-demand allocation of heterogeneous resources.

[0081] The distributed heterogeneous acceleration pooling server system realizes device resource pooling through high-performance exchange of the internal bus and software-defined system design, breaks the binding relationship between the device and the CPU at the physical link layer, can flexibly adjust the port configuration and resource allocation path of the exchange network, finely divides the shared resource pool, and realizes on-demand elastic allocation of device resources and multi-host sharing. In view of the problems that the performance expansion of the current data center host system is limited by the system interconnection bandwidth, the performance between the storage at different levels is not matched, and the I / O resource utilization is low, through the pooling system design, the dynamic allocation and load balancing of resources are flexibly realized for diversified scene requirements, the fusion of multi-platform processor computing power and the collaborative scheduling of heterogeneous acceleration resources are realized, and the performance expansion bottleneck problem of the data center is alleviated.

[0082] The traditional heterogeneous acceleration service is to simply integrate different specifications and different models of existing processors, the traditional heterogeneous acceleration service architecture adopts an I / O tree structure, the host unit enumerates the I / O bridge or I / O device in the I / O tree structure based on a depth-first enumeration method during the startup process, and allocates a unique ID number for each I / O tree, and the routing path can only be allocated within each I / O tree, which limits the diversified allocation and use of I / O device resources. The system expansion capability is limited, the ecology is closed, the energy consumption is wasted, the upgrade and expansion cost is high, and when dealing with large-scale and diversified applications, the traditional heterogeneous acceleration server has low running efficiency.

[0083] The heterogeneous acceleration pool server system provided by the embodiments of the present application comprises a general computing resource pool, an internal bus switching module, an Ethernet switching module and a plurality of heterogeneous computing acceleration resource pools; the general computing resource pool is in communication connection with the internal bus switching module through a first system internal interconnection bus; each heterogeneous computing acceleration resource pool is in communication connection with the internal bus switching module through a second system internal interconnection bus; the internal bus switching module comprises a plurality of interconnected high-speed data interconnection modules, each high-speed data interconnection module is provided with a plurality of IO devices, the IO devices of the plurality of high-speed data interconnection modules are collectively assigned a globally unique code, so as to access any IO device in any high-speed data interconnection module through the globally unique code; after the general computing resource pool obtains resource requirements, the required IO devices are configured by the internal bus switching module according to the resource requirements, a heterogeneous computing resource allocation path is established to make the general computing resource pool match one or more heterogeneous computing acceleration resource pools corresponding to the resource requirements, by assigning a globally unique code to the plurality of IO devices on each high-speed data interconnection module, each general computing unit can access any heterogeneous acceleration device, the system expansion capability is improved, and when large-scale and diversified applications are dealt with, the running efficiency of the heterogeneous acceleration server is improved.

[0084] Based on any of the above embodiments, the heterogeneous computing acceleration resource pool comprises:

[0085] At least one acceleration card, an acceleration processor and an Ethernet controller are integrated in the acceleration card;

[0086] The acceleration processor is directly connected with the Ethernet controller;

[0087] The Ethernet controller is connected with the Ethernet switching module.

[0088] In the embodiments of the present application, the heterogeneous acceleration card is innovatively designed, and the acceleration processor is directly connected with the Ethernet controller to cope with the computing power challenge of large models. The acceleration processor is directly connected with the Ethernet controller through the switching processor, which is different from the traditional acceleration processor connected to the switching processor through the network interface card (NIC), then connected to the CPU through the switching processor, and then interconnected with the network card through the CPU. The heterogeneous computing acceleration card integrates GPU and Ethernet controller, realizes direct data stream reading, as shown in FIG. 2. The RDMA (Remote Direct Memory Access) communication library needs to call the interface to initiate a network transmission request. The driver in the software stack supports the registration of the RDMA network card related interface. The RDMA network card will interwork with the GPU kernel driver through the kernel driver standard interface, complete the direct transmission of data, and realize the direct access of the video memory far end. The entire data path only involves the GPU and the RDMA network card, avoiding the redundant jump of data to the system memory, and effectively reducing the communication delay.

[0089] In the embodiments of the present application, the RDMA network card interworks with the GPU kernel driver through the kernel driver standard interface, completes the direct transmission of data, and the data hop count is reduced by 50% compared with the traditional distributed typical path. The acceleration processor integrates the remote direct memory access function. The acceleration card supports dual-port 100G Ethernet and remote direct memory access (RDMA) function, and supports GPU and Ethernet through software stack and driver.

[0090] In the embodiments of the present application, the acceleration card further includes a remote direct memory access network card, which interworks with the kernel driver of the acceleration processor through the kernel driver standard interface, completes the direct transmission of data and the direct access of the video memory far end. The communication software stack of the system remote direct memory access is shown in FIG. 3. The data transmission delay of the acceleration processor can be reduced, and the computing performance can be improved.

[0091] In the embodiments of the present application, the distributed acceleration server supports two kinds of end-to-end communication modes compared with the traditional AI server, as shown in FIG. 4. The first end-to-end communication mode is high-speed serial computer expansion bus standard (peripheral component interconnect express, PCIe) link communication. The first end-to-end communication mode is a direct communication technology of the switching processor and the Ethernet controller. Through the two kinds of end-to-end communication modes, the entire system can realize horizontal expansion to realize interconnection between systems, and further realize larger-scale heterogeneous computing acceleration pooling.

[0092] Based on the heterogeneous acceleration card, a distributed heterogeneous acceleration server is developed, which can carry a switching processor and a PCIe compatible heterogeneous acceleration processor module. The servers are interconnected by the heterogeneous acceleration cards and switches, forming a hardware interconnection of the multi-training server whole machine, realizing the cluster computing power.

[0093] Based on any of the above embodiments, the heterogeneous acceleration pooling server system further comprises:

[0094] A high-speed serial computer expansion bus standard link is used for interconnection of acceleration processors between multiple servers. The high-speed serial computer expansion bus standard link is also provided with a parameter plane switch, a service plane switch, a storage plane switch, a storage node and a service node;

[0095] The parameter plane switch is used for cluster expansion of the acceleration processors in the heterogeneous computing acceleration resource pool;

[0096] The storage plane switch is used for cluster expansion of the storage nodes;

[0097] The service plane switch is used for cluster expansion of the service nodes.

[0098] The heterogeneous acceleration server cluster expansion mode based on the high-speed serial computer expansion bus standard link is shown in Figure 5: the distributed acceleration server GPU node realizes cluster expansion through the parameter plane switch, 1:1 non-blocking; the service plane and the storage plane cluster network are connected to separate switches, which can support 1:1 non-blocking or 2:1 convergence; the cluster out-of-band management network is connected to the TOR (Top of Rack) out-of-band management switch. Through independent physical ports or networks, the management data of the cluster is transmitted to the TOR switch to realize the management and control of the cluster. Out-of-band management is a management method independent of user service data transmission, which integrates and manages network devices through a dedicated management channel. In a cluster environment, the out-of-band management network is connected to the TOR out-of-band management switch, and the management data of the cluster is transmitted to the TOR switch through a special connection method, thereby realizing the management and control of the entire cluster. This management method is usually used to ensure the high availability, manageability and security of the cluster. Specifically, the out-of-band management network connected to the TOR out-of-band management switch can include:

[0099] The use of independent physical ports or networks ensures that the management data and user service data are transmitted on different links, avoiding interference of the management data by the service data. Through the management interface provided by the TOR switch, the management and monitoring of each node in the cluster are realized, including configuration, state monitoring, fault diagnosis, etc. In the case of not interfering with the normal operation of the service, the network equipment is remotely managed and maintained. At the same time, through the centralized management of the TOR switch, the management efficiency and accuracy can be greatly improved, and the risk of service interruption caused by manual intervention can be reduced.

[0100] It should be noted that the cluster expansion topology scheme can be adjusted according to the actual scheme, and the present application is not limited.

[0101] Based on any of the above embodiments, the general computing resource pool comprises:

[0102] A plurality of general computing units, each general computing unit comprising a general processor and a memory, and the general processor and the memory are combined in any number by virtualization software to form a computing resource container of any granularity.

[0103] In the embodiments of the present application, the general computing resource pool is further decoupled from the CPU and the memory on the basis of I / O centralization to form a centralized computing resource pool, and is assisted by software-defined capabilities. A large number of CPUs and memories of computers can be combined by virtualization software to form a computing resource container of any granularity, and computing nodes of various scales and configurations can be formed.

[0104] Based on any of the above embodiments, the internal bus exchange module comprises an I / O exchange controller, which is used for enumerating all I / O devices, and defining network topology division according to resource requirements, dynamically adjusting the matching between the acceleration processor and the I / O device, and establishing a resource allocation path.

[0105] The embodiments of the present application fully adapt to the complexity characteristics of the distributed heterogeneous acceleration pooling server, and provide reliable support for overall management control and pooling resource management; develop an overall management system for the heterogeneous computing resource pool, improve the availability and manageability of the server system; develop a heterogeneous computing resource pooling management engine, realize the reconstruction management and dynamic allocation of the heterogeneous computing resource, and improve the ease of use and manageability of the pooled resource. Meet the needs of unified monitoring and management of super-large-scale distributed heterogeneous acceleration pooling servers. Breakthrough key technologies such as automatic discovery of complex topology, centralized asset management, and automatic switching of heterogeneous resources in the heterogeneous computing resource pooling system, realize the on-demand rapid deployment, automatic switching, and centralized management of the physical resources of the heterogeneous computing system. The overall architecture of the heterogeneous computing resource pool management is shown in FIG. 6.

[0106] In the embodiments of the present application, the system comprises:

[0107] The pooled management controller is used for at least one of node topology identification, centralized display of asset information, and cooperative power-on and power-off.

[0108] The pooled management controller is used for at least one of node topology identification, centralized display of asset information, and cooperative power-on and power-off.

[0109] In the embodiment of the present application, the pooled management controller is also used for at least one of integrated monitoring, fault early warning, and visual management of the system through a standardized service interface.

[0110] The traditional server is usually centered on a general computing unit with a fixed number of accelerators. In the management hierarchy, the general computing unit manages the accelerator modules hung below through the baseboard controller. For a distributed heterogeneous acceleration pooled server, the general computing unit and the heterogeneous computing acceleration unit are decoupled, pooled, and reconstructed through a high-performance switching unit. The general computing and heterogeneous computing resource collocation relationship needs to be allocated on demand according to the computing power demand. The traditional management architecture centered on the general computing unit cannot perceive and coordinate the integrated computing power resources, and cannot meet the needs of complex topology automatic discovery, centralized asset management, and automatic switching of heterogeneous resources. Therefore, the embodiment of the present application designs a hierarchical heterogeneous computing resource pool management architecture, taking the pooled management controller (PSMC) in the high-performance switching unit as the center, and the heterogeneous computing acceleration unit and the general computing unit collocating the pooled node management controller. The pooled management controller is the center management node, and the distributed nodes are uniformly managed by the center management node, realizing the whole life cycle management of the heterogeneous computing resource pool.

[0111] The pooled management controller is crucial in the management of the distributed heterogeneous acceleration pooled server, and is a bridge for information communication, responsible for unified management. The main functions include node topology identification, centralized display of asset information, and cooperative power-on and power-off. Through a standardized service interface, external operation and maintenance capabilities are provided to realize integrated monitoring, fault early warning, and visual management. The heterogeneous computing resource pool whole machine management system realizes the interconnection of pooled node management controllers at all levels through a network, and builds a management platform with independent engines.

[0112] In the embodiment of the present application, the pooled management engine is used for managing the internal bus switching module, discovering the current network topology through the physical channel and software interface in the internal bus switching module to obtain the current asset allocation, and switching the network topology of the integrated computing power resources according to the physical channel and software interface to meet the resource demand.

[0113] In the heterogeneous computing resource pool system, the discovery, management and elastic adjustment of the heterogeneous computing resources are crucial. The pool management engine is designed as a core management unit of dynamic resource adjustment. The pool management engine controls the high-performance I / O exchange processor to realize automatic discovery of resource topology, flexible automatic switching of resources, and provides a standard interface for the data center monitoring and management platform to realize centralized management of massive resources at the data center level.

[0114] The pool management controller is a centralized management layer. The pool management engine controls the high-performance I / O exchange processor to realize on-demand allocation and elastic expansion of general computing resources and heterogeneous computing resources.

[0115] Based on any of the above embodiments, for the heterogeneous computing resource system, the internal bus exchange module as a high-performance exchange unit is the infrastructure for connecting general computing units and heterogeneous computing units, and building general resource pools and heterogeneous acceleration resource pools. The pool management engine is the adhesive between the data center monitoring and management platform and the infrastructure for realizing resource management and switching scheduling. The pool management engine controls the high-performance exchange unit to realize automatic discovery of resource topology, flexible automatic switching of resources, and provides a standard interface for the data center monitoring and management platform to realize centralized scheduling and management of resources. The overall software and hardware technical solution of the pool management engine is shown in FIG. 7. The pool management engine includes:

[0116] a hardware layer, a driver layer, a function layer and an interface layer;

[0117] The hardware layer is used to connect the general computing resource pool and the plurality of heterogeneous computing acceleration resource pools, and to allocate resources to the general computing resource pool and the plurality of heterogeneous computing acceleration resource pools.

[0118] The driver layer is used for data transmission between the hardware layer and the function layer.

[0119] The function layer is used to realize the functional requirements required by the user.

[0120] The interface layer is used to obtain user demand information and provide visual information to the user.

[0121] In the embodiments of the present application, the function layer is used for, but not limited to:

[0122] (1) information interaction with the pool management controller through the universal asynchronous receiver-transmitter, providing a network interface to visually display the resource list and topology information, the resource list and topology information including at least one of the heterogeneous device resource pool information, the general computing unit information and the I / O port information.

[0123] (2) establishing expert templates with different performance indicators, and switching different expert templates according to the performance indicator requirements of the application scenario when the application scenario is switched.

[0124] and providing a callable API to display the logical topology of the resources according to the visual application of the expert template of the callable API.

[0125] In the embodiments of the present application, the binding relationship between the heterogeneous computing device and the CPU at the physical link level is released, and the general computing unit is implemented to flexibly allocate the heterogeneous computing resources on demand. For different application scenarios, different device allocation relationship expert templates can be established in the pool management engine, and the topology is automatically switched according to different templates when the application scenario is switched, so as to simplify the application scenario switching process. Based on the network API interface provided by the pool management engine, the system interface standard is researched, the system framework standard is established, the expert template description and dynamic switching interface of the heterogeneous computing resource are realized, and the application can apply for resources according to the expert template. At the same time, the callable API is provided to realize the visual application of the expert template, and the logical topology of the allocated resources is displayed, as shown in FIG. 8.

[0126] In the embodiments of the present application, according to the parameters such as business requirements, resource states and performance indicators, the expert template description and dynamic switching interface of the heterogeneous computing resource are realized, so that the application can apply for resources according to the expert template, and the network interface and the visual WEB interface are provided to realize the visual application of the expert template.

[0127] (3) Dynamically allocating and adjusting the heterogeneous computing unit through hot removal, hot insertion and hot reset of the device. In the embodiments of the present application, the pool management engine realizes the dynamic allocation and adjustment of the heterogeneous computing unit in seconds through the key technologies such as hot removal, hot insertion and hot reset of the device.

[0128] (4) Dynamically balancing and deploying the physical resources according to the user requirements, and releasing the computing power of the heterogeneous computing unit. In the embodiments of the present application, the pool management engine dynamically balances and deploys the physical resources based on the optimization scheduling algorithm of load balancing, realizes the on-demand deployment of fine-grained storage resources, and maximizes the release of the heterogeneous computing power.

[0129] In the distributed heterogeneous acceleration pool server, the general host resource pool and the heterogeneous acceleration resource pool are connected through a high-performance exchange unit. The high-performance exchange unit is composed of eight high-performance I / O exchange processors. The high-performance I / O exchange processor is embedded with an embedded controller, which provides a dedicated physical channel and a dedicated software interface for the pool management engine to call. The pool management engine realizes the unified centralized management and control of all high-performance I / O exchange processors through these programmable interaction interfaces. The pool management engine provides expert templates in different scenarios (such as deep learning, numerical calculation, etc.), flexibly realizes topology switching, and realizes a visual UI layer to provide a standard network interface.

[0130] In the heterogeneous computing resource pool, users can realize the operation of on-demand allocation, dynamic scaling and resource release of general-purpose, heterogeneous and other computing units through the interface provided by the pool management engine. Taking AI application as an example, users can call different numbers of heterogeneous computing accelerator cards and general-purpose computing resources according to the load demand, and when the AI application stops, users can release the heterogeneous computing resources back to the entire heterogeneous computing resource pool, so as to realize efficient flow and full utilization of resources.

[0131] Based on any of the above embodiments, in the related art, the topology of the existing technical scheme of the heterogeneous acceleration pooling server system is shown in FIG. 9, and most of them support 8 NVIDIA A100 Tensor Core GPUs interconnected by NVIDIA NVLink and 2 AMD Milan CPUs in a 4U space with heterogeneous processors, and support liquid cooling and air cooling technologies. NF5688M6 is a GPU server designed for large-scale data centers with extreme expansion capability, supporting 8 third-generation NVLink interconnected NVIDIA A100 Tensor Core GPUs and two Intel Ice Lake CPUs, and supporting up to 13 PCIe Gen4 I / O cards.

[0132] In the embodiments of the present application, the heterogeneous acceleration pooling server system is a single-machine 32-card heterogeneous acceleration pooling server system, which includes:

[0133] Two groups of full interconnection topologies containing four switching processors support 32 general-purpose processor card extensions;

[0134] The internal interconnection bandwidth of the single-machine 32-card heterogeneous acceleration pooling server system reaches 384GB / s, and the inter-card interconnection bandwidth reaches 128GB / s.

[0135] The single-machine 32-card heterogeneous acceleration server system scheme is shown in FIGS. 10 and 11: the whole cabinet can place two groups of 32-card systems, and can support a maximum of 64 GPU card extensions; the internal interconnection bandwidth of 16 cards reaches 384GB / s, and the cross-CPU interconnection bandwidth between 16 cards is 128GB / s.

[0136] The distributed heterogeneous acceleration resource pooling server is different from the traditional server, and is centered on multiple data exchange networks and I / O interconnection networks. The system architecture can be divided into a general computing resource pool, a high-performance exchange unit, a heterogeneous acceleration computing resource pool and a network exchange unit according to different logical functions. The general computing resource pool is further decoupled from the CPU and the memory on the basis of I / O centralization to form a centralized computing resource pool, and is assisted by software-defined capabilities. A large number of CPUs and memory computers can form computing resource containers of any size through virtualization software to form computing nodes of various scales and configurations. The high-performance exchange unit mainly implements high-performance data exchange functions. Through the use of a distributed data exchange architecture, flexible network topology division can be achieved through software definition, and the matching between computing and I / O modules can be quickly and dynamically adjusted to realize the dynamic combination between the two and effectively improve the scalability and flexibility of the entire system to ensure the realization of hardware reconstruction. The heterogeneous acceleration resource pool mainly provides heterogeneous acceleration services, dynamically reconstructs CPU and GPU resources to form a heterogeneous computing server cluster, and meets the demand of high-performance applications for computing power.

[0137] Based on any of the above embodiments, the heterogeneous acceleration pooling server system is an 8-host shared 32-card heterogeneous acceleration pooling server system, which includes:

[0138] Two groups of full interconnection topologies containing four exchange processors, 8 hosts sharing 32 acceleration processor cards.

[0139] The 8-host shared 32-card heterogeneous acceleration server system scheme is shown in FIGS. 12 and 13, supports dynamic allocation and expert template rapid deployment, and realizes multi-host shared heterogeneous computing acceleration resource pool.

[0140] In the embodiments of the present application, a single system supports 8 / 16 / 32 heterogeneous computing acceleration card expansion, supports not less than 8 domestic host sharing heterogeneous computing acceleration resource pool, and supports second-level resource dynamic allocation.

[0141] Based on any of the above embodiments, the system further includes a computing power virtualization subsystem, which aggregates the heterogeneous computing acceleration resource pool into a unified virtual computing power resource pool. The computing power virtualization subsystem includes:

[0142] A pooling server, which is a computing node in the system cluster and is used for task running;

[0143] A management and control server, which is used for running a management and control component in the system cluster;

[0144] An application server, which is used for running an application program in the system cluster, and the application server is a container, a virtual machine or a physical machine.

[0145] In the embodiments of the present application, the distributed heterogeneous acceleration pooling server is a new form of high-performance computing power infrastructure, which has great differences from traditional servers in system architecture, hardware logic, infrastructure design, etc., and needs more efficient and flexible management architecture to realize the life cycle management of pooled resources.

[0146] In the embodiments of the present application, the logical relationship between the heterogeneous computing power software-defined virtualization management system and its subsystem and the pooling server system is shown in FIG. 14: by aggregating computing power resources through a unified virtual computing power resource pool, it can be realized that the computing task occupies resources during execution and does not occupy resources during idle time; for users, it reduces resource usage costs. For the cluster, it can improve the turnover frequency of cluster resources and provide limited resources for more people to use. Support remote calling of computing power resources, breaking the space limit; allow users to call computing power resources on any node on the cluster. Support virtualization of computing power resources, dividing a computing card into multiple computing cards for use. Make computing power resources match computing tasks to improve resource utilization. Support integrating various types of computing power resources to a computing task across nodes. Thus, it can provide computing tasks with much more computing resources than physical machine resources. Support dynamic adjustment of virtual resources. Computing tasks can dynamically adjust the virtual resources used by computing tasks during operation, thereby supporting functions such as mixed deployment and time sharing. Mixed deployment means scheduling different types of tasks to the same physical resources, and through scheduling and resource isolation control means, on the basis of guaranteeing the service level agreement, fully utilizing the resource capacity and reducing the operating cost. This technology mixes online services and computing tasks through mixed deployment, and dynamically adjusts the allocation of resources according to the size of online pressure. For example, when the online pressure is large, the computing task will exit the occupied resources to ensure the smooth operation of online services; when the online pressure is small, the computing task will occupy idle resources to improve resource utilization. Time sharing is a time-based resource allocation strategy, which allocates resources to different applications in different periods to adapt to changes in demand. Through the time-sharing scheduling of operation and maintenance, resources are allocated to different applications according to different periods to reduce resource procurement during the big promotion period and further save costs. Through functions such as mixed deployment and time sharing, the resource utilization rate of the data center can be significantly improved, the operating cost can be reduced, and the stable operation of the business can be guaranteed.

[0147] The heterogeneous acceleration pooling server system provided by the embodiments of the present application designs a new type of heterogeneous computing acceleration card, a distributed heterogeneous acceleration pooling server system and a management and scheduling method, supports remote direct access of memory, and the data hop count is reduced by 50% compared with the traditional distributed typical path; based on the hardware and software prototype verification platform of the heterogeneous computing acceleration card, the adaptation and optimization of various components such as processors, acceleration cards, networks and operating systems are carried out, and the compatibility of the acceleration card hardware and software is ensured. Breakthrough in the development of distributed heterogeneous acceleration pooling server system, hardware and software collaborative optimization, research on the whole life cycle management and scheduling technology of heterogeneous computing resource pool, realize the dynamic scheduling of business adaptive resources, support the dynamic adjustment of processor / accelerator ratio, topology discovery and on-demand switching, meet the diversified application load demand; realize the remote calling of heterogeneous computing resources based on Ethernet, support TCP (Transmission Control Protocol), RDMA data transmission protocol; realize the virtualization and dynamic scaling management of heterogeneous computing resources.

[0148] The embodiments of the present application take application as the guide and take system architecture design as the core, through the multi-level collaborative design of "system-component-processor-software", research and develop a high-performance and high-efficiency distributed heterogeneous acceleration pooling server system, realize the unified management and scheduling of heterogeneous computing resources, fully exert the system innovation advantage of diversified computing power, make the computing power easy to use and easy to use, which is very important for promoting the popularization and application of artificial intelligence technology and the development of computing power industry in China. Mastering the core technology of distributed heterogeneous acceleration pooling server, forming a self-contained technology system. Heterogeneous computing is the main computing architecture adopted by future data center computing infrastructure. The lag in current research on key technologies of heterogeneous computing leads to the dependence of important computing infrastructure such as intelligent computing data center and supercomputing data center on the outside, which is limited in computing power expansion and computing power management. Through the embodiments of the present application, the academia and enterprises are driven to carry out research on the core technology of distributed heterogeneous acceleration pooling server, and a self-contained technology system is formed. Breakthrough in key technologies of heterogeneous computing acceleration card, key technologies of distributed heterogeneous acceleration pooling server, life cycle management and scheduling technology of heterogeneous computing resource pool, and virtualization and remote scheduling technology of heterogeneous computing resource, build a self-contained industrial chain, value chain and ecological system. Develop self-contained heterogeneous acceleration processors, heterogeneous acceleration cards, distributed heterogeneous acceleration pooling servers and related supporting software, carry out demonstration application in large-scale data centers, which is conducive to the development of computing power industry. Drive the coordinated development of upstream and downstream industrial chains, create huge economic benefits. Drive more enterprises and research institutions to increase investment in heterogeneous computing core components and equipment, and provide strong support for the development of new generation information technology.

[0149] The pooling resource allocation method provided by the present application is described below, and the pooling resource allocation method described below can be mutually corresponding to the heterogeneous acceleration pooling server system described above.

[0150] The pool resource allocation method provided in the application is applicable to the heterogeneous acceleration pool server system of any one of the above, comprising:

[0151] Step 101, a general computing resource pool unit initiates a heterogeneous computing resource dynamic adjustment request to the pool management controller according to resource requirements;

[0152] Step 102, the pool management engine determines the IO device to be accessed in the internal bus switching module according to the resource requirements, accesses the IO device in one or more high-speed data interconnection modules through a globally unique code, reconstructs the heterogeneous computing resource allocation path, and matches the heterogeneous computing resource corresponding to the resource requirements for the general computing unit;

[0153] The heterogeneous computing resource includes the resource in a single acceleration resource pool, the resource across the acceleration resource pool, and the resource across the pool server.

[0154] In the embodiment of the application, for the heterogeneous computing resource system, the internal bus switching module as a high-performance switching unit is the infrastructure connecting the general computing unit and the heterogeneous computing unit, and constructing the general resource pool and the heterogeneous acceleration resource pool, and the pool management engine is the adhesive between the data center monitoring management platform and the infrastructure to realize resource management and switching scheduling. The pool management engine controls the high-performance switching unit to realize resource topology automatic discovery, resource flexible automatic switching, and provides a standard interface for the data center monitoring management platform to realize resource centralized scheduling and management.

[0155] In the embodiment of the application, the distribution of resources is divided into three cases from the dimension of resource distribution, which can effectively utilize resources, and on this basis, a scheduling system based on heterogeneous distance is developed to support the heterogeneous resource scheduling capability of the single acceleration resource pool, the cross acceleration resource pool, and the cross pool server according to the scheduling strategy, thereby meeting the computing power requirements of the pool server for different applications.

[0156] The embodiment of the application fully adapts to the complexity characteristics of the distributed heterogeneous acceleration pool server, provides reliable support for overall management control and pool resource management, develops the overall management system of the heterogeneous computing resource pool to improve the availability and manageability of the server system, develops the pool management engine of the heterogeneous computing resource to realize the reconstruction management and dynamic allocation of the heterogeneous computing resource, and improves the ease of use and manageability of the pool resource. Meet the needs of unified monitoring and management of super-large-scale distributed heterogeneous acceleration pool servers. Breakthrough key technologies such as automatic discovery of complex topology, centralized asset management, and automatic switching of heterogeneous computing pool systems, realize the on-demand rapid deployment, automatic switching, and centralized management of physical resources of the heterogeneous computing system.

[0157] FIG. 15 is a timing diagram of the pool resource allocation method provided by the embodiments of the present application. As shown in FIG. 15, the pool resource allocation method provided by the embodiments of the present application includes the following steps:

[0158] (0) Taking releasing a certain heterogeneous acceleration card device and allocating it to a certain general-purpose computing unit as an example, before the user initiates a request for dynamic adjustment of heterogeneous computing power resources through the pool management engine, it is ensured that the application layer process related to the heterogeneous computing power device has ended to avoid program exceptions caused by abnormal access of the application layer to the device;

[0159] (1) The pool management engine completes the hot removal of the device;

[0160] (2) The pool management engine sends a request to the high-performance exchange unit to obtain the physical location of the device;

[0161] (3) The pool management engine confirms the physical location of the device;

[0162] (4) The pool management engine resets the heterogeneous computing power device resource and re-performs Training to restore the running state to the default value;

[0163] (5) The pool management engine confirms that the device has completed the reset;

[0164] (6) The pool management engine reallocates the device to another general-purpose computing unit, which will see the newly added device without business awareness, that is, the dynamic switching of the heterogeneous computing power resource is completed.

[0165] The embodiments of the present application can realize cross-node, multi-host sharing, on-demand resource allocation and elastic application, maximize the release of heterogeneous computing power, and realize performance optimization in different application scenarios.

[0166] In the embodiments of the present application, the pool resource allocation method further includes:

[0167] The management and control service receives a resource application initiated by the heterogeneous resource management communication component in the application server, and issues a configuration to the exchange network according to the resource application;

[0168] The configuration is issued to the heterogeneous virtual computing power management node component in the pool server through the search exchange network, so that the pool server loads the configuration to complete the computing power resource virtualization deployment.

[0169] In the embodiments of the present application, the heterogeneous intelligent computing power virtualization subsystem aggregates the pool of accelerated computing resources on the pool server into a unified virtual computing power resource pool. On the basis of the unified virtual computing power resource pool, the heterogeneous computing power resource remote invocation subsystem, the heterogeneous virtual computing power management subsystem, the heterogeneous computing cluster resource allocation and system performance optimization subsystem are responsible for managing and using the unified virtual computing power. The deployment of the heterogeneous computing power software-defined virtualization management system requires a pool server, a management and control server and an application server. The specific subsystem module deployment location and call link are shown in FIG. 16.

[0170] In the embodiments of the present application, for the cluster, the turnover frequency of the cluster resources can be improved, and limited resources can be provided for more people to use. Remote invocation of computing power resources is supported, breaking the spatial limit; users are allowed to invoke computing power resources on any node on the cluster. Virtualization of computing power resources is supported, and a computing card can be divided into multiple computing cards for use. Computing power resources are matched with computing tasks to improve resource utilization. Cross-node integration of various types of computing power resources for a computing task is supported. Thus, computing resources much larger than physical machine resources can be provided for a computing task.

[0171] The distributed acceleration server GPU node realizes cluster expansion through a parameter face switch, 1:1 non-blocking; the service plane and the storage plane cluster network are connected to separate switches, which can support 1:1 non-blocking or 2:1 convergence; the cluster out-of-band management network is connected to a TOR (Top of Rack) out-of-band management switch. Through independent physical ports or networks, the management data of the cluster is transmitted to the TOR switch to realize the management and control of the cluster. Out-of-band management is a management method independent of user service data transmission, which integrates and manages network devices through a dedicated management channel. In the cluster environment, the out-of-band management network is connected to the TOR out-of-band management switch, and the management data of the cluster is transmitted to the TOR switch through a special connection method, thereby realizing the management and control of the entire cluster. This management method is usually used to ensure the high availability, manageability and security of the cluster.

[0172] During the running of the computing task, the virtual resources used by the computing task can be dynamically adjusted, thereby supporting functions such as mixed deployment and time sharing. Mixed deployment refers to scheduling different types of tasks to the same physical resource, and through scheduling and resource isolation and other control means, on the basis of guaranteeing a service level agreement, the resource capacity is fully utilized, and the operation cost is reduced. This technology dynamically adjusts the allocation of resources according to the size of online pressure by mixing the deployment of online services and computing tasks. For example, when the online pressure is large, the computing task exits the occupied resource, to ensure the smooth running of the online business; and when the online pressure is small, the computing task occupies the idle resource, to improve the resource utilization. Time sharing is a time-based resource allocation strategy, and through the allocation of resources to different applications in different periods, the change in demand is adapted. Through the time sharing scheduling cooperation of operation and maintenance, the resources are allocated to different applications according to different periods, to reduce the resource procurement during the big promotion period, and further save the cost. Through functions such as mixed deployment and time sharing, the resource utilization rate of the data center can be significantly improved, the operation cost is reduced, and the stable operation of the business is ensured.

[0173] The pool resource allocation method provided by the embodiments of the present application realizes the interconnection of each server between systems through a network switching module; the network port configuration and the heterogeneous computing resource allocation path are adjusted according to resource requirements, so that the general computing resource pool matches one or more heterogeneous computing acceleration resource pools corresponding to the resource requirements, supports dynamic adjustment of the processor / accelerator ratio, topology discovery and on-demand switching, can realize adaptive resource dynamic scheduling of the business, meets the diversified application load requirements, and improves the operation efficiency and performance of the heterogeneous acceleration pooled server system.

[0174] In another aspect, the present application also provides a computer nonvolatile readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the pool resource allocation method provided by each method, and the method comprises: a certain general computing power unit in a general computing resource pool initiates a heterogeneous computing power resource dynamic adjustment request to a pool management controller according to resource requirements; the network port configuration and the heterogeneous computing resource allocation path are adjusted by a pool management engine, to match the general computing power unit with the heterogeneous computing power resource corresponding to the resource requirements; the matched heterogeneous computing power resource is reset and retrained, the trained acceleration processor is allocated to the general computing power unit, and the dynamic switching of the heterogeneous computing power resource is completed; wherein the heterogeneous computing power resource comprises resources in a single acceleration resource pool, resources across acceleration resource pools, and resources across pooled servers.

[0175] The device embodiments described above are merely illustrative, wherein the units illustrated as separate components can or can not be physically separate, and the components illustrated as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0176] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software products, and the computer software products can be stored in a computer non-volatile readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods of various embodiments or some parts of the embodiments.

[0177] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A heterogeneous accelerated pooling server system, comprising: The system comprises: a general computing resource pool, an internal bus switching module, an Ethernet switching module, and a plurality of heterogeneous computing acceleration resource pools; the general computing resource pool is communicatively connected to the internal bus switching module through a first system internal interconnection bus; each heterogeneous computing acceleration resource pool is communicatively connected to the internal bus switching module through a second system internal interconnection bus; each heterogeneous computing acceleration resource pool is communicatively connected to the Ethernet switching module, and the Ethernet switching module is configured to interconnect servers between systems; the internal bus switching module comprises a plurality of interconnected high-speed data interconnection modules, each high-speed data interconnection module is provided with a plurality of IO devices, the IO devices of the plurality of high-speed data interconnection modules are collectively assigned a globally unique code, so as to access any IO device in any high-speed data interconnection module through the globally unique code; after the general computing resource pool obtains a resource requirement, the required IO devices are configured by the internal bus switching module according to the resource requirement, and a heterogeneous computing resource allocation path is established to match one or more heterogeneous computing acceleration resource pools corresponding to the resource requirement.

2. The heterogeneous accelerated pooling server system of claim 1, wherein, The heterogeneous computing acceleration resource pool comprises: at least one acceleration card, wherein an acceleration processor and an Ethernet controller are integrated in the acceleration card; the acceleration processor is directly connected to the Ethernet controller; the Ethernet controller is connected to the Ethernet switching module.

3. The heterogeneous accelerated pooling server system of claim 2, wherein, The system further comprises: a high-speed serial computer expansion bus standard link, which is configured to interconnect acceleration processors between a plurality of servers.

4. The heterogeneous accelerated pooling server system of claim 2, wherein, The system further comprises a remote direct data access network card, which communicates with a kernel driver of the acceleration processor through a kernel driver standard interface, the remote direct data access network card is connected to the Ethernet switching module, so that a plurality of acceleration processors connected by a plurality of remote direct data access network cards can complete direct transmission of data and remote direct access of display memory.

5. The heterogeneous accelerated pooling server system of claim 2, wherein, The internal bus switching module comprises an I / O switching controller, which is configured to enumerate all I / O devices, and dynamically adjust the matching between the acceleration processor and the I / O devices according to the network topology division defined by the resource requirement, to establish a resource allocation path.

6. The heterogeneous accelerated pooling server system of claim 1, wherein, The globally unique code is obtained by coding the number of the high-speed data interconnection module and the number of the IO device inside the high-speed data interconnection module in a segmented coding manner.

7. The heterogeneous accelerated pooling server system of claim 1, wherein, The general computing resource pool comprises: a plurality of general computing units, each general computing unit comprising a general processor and a memory, and any number of general processors and memories are combined by virtualization software to form a computing resource container of any granularity.

8. The heterogeneous accelerated pooling server system of claim 1, wherein, The system further comprises a pooling management controller; The general computing resource pool and the accelerator card of the heterogeneous computing acceleration resource pool are respectively provided with a pool management controller, the pool management controller uniformly manages a plurality of pool management controllers, and at least one of node topology identification, asset information centralized display, and cooperative power-on and power-off of the general computing resource pool and the heterogeneous computing acceleration resource pool is performed.

9. The heterogeneous accelerated pooling server system of claim 8, wherein, The pool management controller is further configured to at least one of integrally monitor, failure early warning, and visual management of the system through a standardized service interface.

10. The heterogeneous accelerated pooling server system of claim 1, wherein, The system further comprises a pool management engine; The pool management engine is configured to manage the internal bus exchange module, discover the current network topology by calling the physical channel and the software interface in the internal bus exchange module, obtain the current asset allocation, and coordinate the integrated computing resource to switch the network topology to meet the resource demand according to the physical channel and the software interface.

11. The heterogeneous accelerated pooling server system of claim 10, wherein, The pool management engine is further configured to: The pool management engine is further configured to:

12. The heterogeneous accelerated pooling server system of claim 10, wherein, Based on the load balancing optimization scheduling algorithm, the pool management engine is further configured to: The pool management engine is further configured to:

13. The heterogeneous accelerated pooling server system of claim 10, wherein, Establish expert templates under different application scenarios; Switch different expert templates according to the current application scenario when the application scenario is switched; and visually display the expert templates meeting the application scenario demand through a visual WEB interface. The heterogeneous acceleration pooling server system is a single-machine heterogeneous acceleration pooling server system, and a single host in the single-machine heterogeneous acceleration pooling server system supports 8, 16, and 32 extended accelerator cards.

14. The heterogeneous accelerated pooling server system of claim 1, wherein, The heterogeneous acceleration pooling server system is a multi-host heterogeneous acceleration pooling server system, each host in the multi-host heterogeneous acceleration pooling server system is connected with a high-speed data interconnection module, and 32 accelerator processor cards are shared through a plurality of high-speed data interconnection modules.

15. The heterogeneous accelerated pooling server system of claim 1, wherein, The system further comprises a computing power virtualization subsystem, the computing power virtualization subsystem aggregates the heterogeneous computing acceleration resource pool into a unified virtual computing power resource pool, and the computing power virtualization subsystem comprises:

16. The heterogeneous accelerated pooling server system of claim 1, wherein, A pool server in the system cluster is configured as a computing node and is configured to run a task; A management and control server in the system cluster is configured to run a management and control component; An application server in the system cluster is configured to run an application program, and the application server is a container, a virtual machine, or a physical machine. The general computing resource pool and the accelerator card of the heterogeneous computing acceleration resource pool are respectively provided with a pool management controller, the pool management controller uniformly manages a plurality of pool management controllers, and at least one of node topology identification, asset information centralized display, and cooperative power-on and power-off of the general computing resource pool and the heterogeneous computing acceleration resource pool is performed.

17. A method of pooling resource allocation, adapted for use in a heterogeneous accelerated pooling server system as claimed in any one of claims 1 to 16, characterized in that, ​ ​ determining, by the pooling management engine, an IO device to be accessed in the internal bus switching module according to the resource requirement, accessing the IO device in one or more high-speed data interconnection modules through a globally unique code, reconstructing a heterogeneous computing resource allocation path, and matching the general computing unit with a heterogeneous computing resource corresponding to the resource requirement; wherein the heterogeneous computing resource includes a resource in a single acceleration resource pool, a resource across acceleration resource pools, and a resource across pooling servers.

18. The method of claim 17, wherein, The reconstructing of the heterogeneous computing resource allocation path includes releasing an acceleration card in a certain general computing unit and allocating the acceleration card to another general computing unit, and specifically includes: sending, by the pooling management engine, a request to the internal bus switching module after hot removing the acceleration card, and obtaining a physical location of the acceleration card; finding the acceleration card according to the physical location, resetting the acceleration card, and reallocating the reset acceleration card to a general computing unit to be allocated.

19. The method of claim 18, wherein, The releasing of the acceleration card in the certain general computing unit and the allocating of the acceleration card to the other general computing unit further include: training the reset acceleration card according to an application scenario of the general computing unit to be allocated; and reallocating the trained acceleration card to the general computing unit to be allocated.

20. A non-transitory readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by a processor to implement the pooling resource allocation method according to any one of claims 17-19.

Citation Information

Patent Citations

  • Multi-accelerator card heterogeneous server and resource link reconstruction method

    CN117687956A

  • Server system, resource scheduling method of server system, chip and chip grain

    CN118210634A

  • Cloud-based framework for analysis using accelerators

    US20240095090A1