Computing system and method, device, medium, and program product

By constructing heterogeneous acceleration pools and resource pools in the computing system, and utilizing Ethernet switching modules and remote direct data access technology, the problem of high bandwidth and low latency communication in large-scale model training was solved, enabling the rapid and stable operation of model training tasks.

WO2026065985A1PCT designated stage Publication Date: 2026-04-02LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Large-scale, long-term model training tasks face challenges such as high communication volume and high time costs under high bandwidth and low latency requirements, which are difficult to solve effectively with existing technologies.

Method used

A computing system is constructed, including a computing resource pool and a heterogeneous acceleration pool, which are connected through an Ethernet switching module. Remote direct data access is achieved using an Ethernet controller and computing cores to realize high-bandwidth and low-latency data transmission. The management core is used for task distribution and binding relationship management.

Benefits of technology

It enables fast and stable operation of model training tasks, provides sufficient computing power support, reduces communication latency and improves data transmission efficiency, and meets the high bandwidth and low latency requirements of large-scale model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025084423_02042026_PF_FP_ABST
    Figure CN2025084423_02042026_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses a computing system and method, a device, a medium, and a program product in the technical field of computers. In the present application, heterogeneous accelerator cards in which Ethernet controllers are built form a heterogeneous acceleration pool, the heterogeneous acceleration pool is connected to an Ethernet switching module by means of the Ethernet or remote direct memory access, and the Ethernet switching module is connected to a computing resource pool; and a management core in the computing resource pool can manage accelerator cards in at least one heterogeneous acceleration pool and a binding relationship between computing cores and the accelerator cards, and distribute a processing task to the accelerator cards in the at least one heterogeneous acceleration pool, so that the accelerator cards in the at least one heterogeneous acceleration pool collaboratively complete the processing task. Data transmission is accelerated by means of the Ethernet or remote direct memory access, which not only can provide sufficient computing power support for operation of tasks, but also can provide high bandwidth and low latency for high-speed communication and large-data-volume transmission of the tasks, thereby facilitating the rapid and stable operation of model training tasks.
Need to check novelty before this filing date? Find Prior Art

Description

A computing system, method, device, medium and program product

[0001] Cross-reference to Related Applications

[0002] This application claims priority to the Chinese patent application No. 202411388137.9, filed on September 30, 2024, and entitled “A computing system, method, device, medium and program product”, the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0003] The present application relates to the technical field of computer, and in particular, relates to a computing system, method, device, medium and program product. BACKGROUND

[0004] At present, model training needs massive computing power support, usually needs a large number of servers as nodes, forms a cluster through high-speed network, and interconnects and intercommunicates between servers to complete the task by mutual cooperation. However, for large-scale and long-time model training task, the communication amount required by only single computing iteration reaches the order of hundreds of GB, in addition, there are various communication needs of parallel mode. If the network bandwidth is not large enough and the delay is long, not only the marginal computing power decreases, but also the time cost of model training increases. In addition, the large model training also has high requirements for delay and packet loss.

[0005] Therefore, how to construct a corresponding computing power system for the model training task is a problem to be solved by those skilled in the art. SUMMARY

[0006] In a first aspect, the present application provides a computing system, comprising: a computing resource pool, at least one heterogeneous acceleration pool, and an Ethernet switching module connected between the computing resource pool and the at least one heterogeneous acceleration pool;

[0007] The at least one heterogeneous acceleration pool comprises: a plurality of acceleration cards; each acceleration card comprises: an Ethernet controller and at least one operation core, and the Ethernet controller communicates with the operation core in a remote direct memory access mode;

[0008] Each acceleration card in the at least one heterogeneous acceleration pool communicates with the Ethernet switching module in an Ethernet mode or a remote direct memory access mode;

[0009] The computing resource pool comprises: a plurality of computing cores, and a same computing core has a binding relationship with at least one acceleration card in the at least one heterogeneous acceleration pool;

[0010] The plurality of computing cores comprises a management core, the management core is used for managing each acceleration card in the at least one heterogeneous acceleration pool and the binding relationship, and distributing a processing task to each acceleration card in the at least one heterogeneous acceleration pool.

[0011] Optionally, each accelerator card in the at least one heterogeneous acceleration pool comprises a golden finger and a power connector; the golden finger and the power connector are connected to each device in the corresponding accelerator card and supply power to each device in the corresponding accelerator card.

[0012] Optionally, the golden finger and the power connector in each accelerator card in the at least one heterogeneous acceleration pool supply power at the same time in the corresponding accelerator card startup process.

[0013] Optionally, the Ethernet controller in each accelerator card in the at least one heterogeneous acceleration pool is connected to the computing core in the corresponding accelerator card through a switch chip.

[0014] Correspondingly, the Ethernet controller in each accelerator card in the at least one heterogeneous acceleration pool and the computing core in the corresponding accelerator card realize communication in a remote direct memory access mode through the switch chip.

[0015] Optionally, each accelerator card in the at least one heterogeneous acceleration pool comprises a common clock device; the common clock device is connected to the computing core, the switch chip and the Ethernet controller in the corresponding accelerator card.

[0016] Optionally, each accelerator card in the at least one heterogeneous acceleration pool comprises a clock generator; the clock generator is connected between the common clock device and the computing core, between the common clock device and the switch chip, and between the common clock device and the Ethernet controller in the corresponding accelerator card.

[0017] Correspondingly, the clock generator is used to select a clock source for the computing core, the switch chip and the Ethernet controller in the corresponding accelerator card; the clock source is the common clock device in the corresponding accelerator card or a clock device in the computing core bound to the corresponding accelerator card.

[0018] Optionally, each accelerator card in the at least one heterogeneous acceleration pool comprises a control unit; the control unit is used to realize timing control of information transmission in the corresponding accelerator card.

[0019] Optionally, each accelerator card in the at least one heterogeneous acceleration pool comprises a link selector; the link selector is connected to the computing core, the switch chip, the Ethernet controller, the common clock device, the power device and the control unit in the corresponding accelerator card.

[0020] Optionally, the Ethernet switching module comprises at least one switch; the at least one switch is connected to the computing resource pool and the at least one heterogeneous acceleration pool.

[0021] Optionally, each accelerator card in the at least one heterogeneous acceleration pool comprises a first node controller; the first node controller is used to collect device information and running information in the corresponding accelerator card.

[0022] Correspondingly, each computing core in the computing resource pool comprises a second node controller; the second node controller is configured to collect device information and running information in the corresponding computing core;

[0023] Correspondingly, the Ethernet switching module comprises a central controller; the central controller is configured to collect device information and running information in the acceleration cards collected by the first node controllers, and device information and running information in the computing cores collected by the second node controllers.

[0024] Optionally, the first node controller is configured to collect state information of the acceleration chips and sensor data of the board cards in the corresponding acceleration cards.

[0025] Optionally, the second node controller is configured to collect core running information and sensor data in the corresponding computing cores.

[0026] Optionally, the central controller is configured to construct a topology map comprising each computing core in the computing resource pool and each acceleration card in the at least one heterogeneous acceleration pool according to the collected information.

[0027] Optionally, the central controller is configured to synchronize the collected information to a management core in the computing resource pool.

[0028] Correspondingly, the management core manages each acceleration card in the at least one heterogeneous acceleration pool and the binding relationship according to the received synchronization information, and distributes processing tasks to each acceleration card in the at least one heterogeneous acceleration pool.

[0029] Optionally, the management core is configured to formulate a corresponding task allocation strategy and a binding relationship adjustment strategy according to the received synchronization information.

[0030] Optionally, the central controller is configured to generate log data according to the collected information, and analyze the log data to perform fault diagnosis.

[0031] Optionally, the computing system further comprises a bus switching module connected to the at least one heterogeneous acceleration pool; the bus switching module is further connected to another computing resource pool.

[0032] In a second aspect, the present application provides a computing method applied to a management core, comprising:

[0033] receiving a processing task;

[0034] sending the processing task to an Ethernet switching module in a computing system, so that the Ethernet switching module distributes the processing task to at least one heterogeneous acceleration pool in the computing system by an Ethernet mode or a remote direct data access mode;

[0035] The computing system comprises a computing resource pool, at least one heterogeneous acceleration pool, and an Ethernet switching module connected between the computing resource pool and the at least one heterogeneous acceleration pool.

[0036] The at least one heterogeneous acceleration pool comprises: a plurality of acceleration cards; each acceleration card comprises: an Ethernet controller and at least one operation core, the Ethernet controller and the operation core communicate in a remote direct memory access mode;

[0037] Each acceleration card in the at least one heterogeneous acceleration pool communicates with the Ethernet switch module in an Ethernet mode or a remote direct memory access mode;

[0038] The computing resource pool comprises: a plurality of computing cores, and each computing core has a binding relationship with at least one acceleration card in the at least one heterogeneous acceleration pool; and a management core for any one of the plurality of computing cores.

[0039] In a third aspect, the present application provides an electronic device, comprising:

[0040] A memory for storing computer readable instructions;

[0041] A processor for executing the computer readable instructions to implement the computing method disclosed above.

[0042] In a fourth aspect, the present application provides one or more non-volatile computer readable storage media storing computer readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the computing method disclosed above.

[0043] In a fifth aspect, the present application provides a computer program product comprising computer readable instructions, which, when executed by a processor, implement the steps of the computing method disclosed above. BRIEF DESCRIPTION OF DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of the provided drawings.

[0045] FIG. 1 is a schematic diagram of a computing system disclosed by one or more embodiments of the present application;

[0046] FIG. 2 is a schematic diagram of a GPU acceleration card disclosed by one or more embodiments of the present application;

[0047] FIG. 3 is a power supply circuit diagram of a GPU acceleration card disclosed by one or more embodiments of the present application;

[0048] FIG. 4 is a clock circuit diagram of a GPU acceleration card disclosed by one or more embodiments of the present application;

[0049] FIG. 5 is a schematic diagram of communication between a GPU and an Ethernet controller according to one or more embodiments of the present disclosure;

[0050] FIG. 6 is a schematic diagram of a RDMA communication software stack according to one or more embodiments of the present disclosure;

[0051] FIG. 7 is an architecture diagram of a computing system according to one or more embodiments of the present disclosure;

[0052] FIG. 8 is an architecture diagram of another computing system according to one or more embodiments of the present disclosure;

[0053] FIG. 9 is an architecture diagram of yet another computing system according to one or more embodiments of the present disclosure;

[0054] FIG. 10 is a schematic diagram of a communication method according to one or more embodiments of the present disclosure;

[0055] FIG. 11 is a schematic diagram of a topology according to one or more embodiments of the present disclosure;

[0056] FIG. 12 is a schematic diagram of accelerator card management according to one or more embodiments of the present disclosure;

[0057] FIG. 13 is another schematic diagram of accelerator card management according to one or more embodiments of the present disclosure;

[0058] FIG. 14 is a schematic diagram of kernel communication of a general computing node according to one or more embodiments of the present disclosure;

[0059] FIG. 15 is a schematic diagram of a heterogeneous virtual computing power management subsystem according to one or more embodiments of the present disclosure;

[0060] FIG. 16 is a schematic diagram of a heterogeneous computing power resource pool management architecture according to one or more embodiments of the present disclosure;

[0061] FIG. 17 is a server structure diagram according to one or more embodiments of the present disclosure;

[0062] FIG. 18 is a terminal structure diagram according to one or more embodiments of the present disclosure. DETAILED DESCRIPTION

[0063] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other examples obtained by those of ordinary skill in the art without creative work fall within the scope of the present disclosure.

[0064] Currently, model training needs massive computing power support, usually needs a large number of servers as nodes, forms a cluster through high-speed network, interconnects and intercommunicates between servers, and cooperates with each other to complete the task. However, for large-scale and long-time model training tasks, the communication amount required by only a single computing iteration reaches the order of hundreds of GB, and there are also various communication needs of parallel modes. If the network bandwidth is not large enough and the delay is long, not only the marginal computing power decreases, but also the time cost of model training increases. In addition, the large model training also has high requirements for delay and packet loss. Therefore, the application provides a computing scheme, which can construct a corresponding computing power system for the model training task, accelerate data transmission through Ethernet or remote direct data access mode, not only can provide sufficient computing power support for the running of the task, but also can provide high bandwidth and low delay for high-speed communication and large data transmission of the task, which is beneficial to the rapid and stable operation of the model training task.

[0065] Referring to FIG. 1, the embodiment of the application discloses a computing system, comprising: a computing resource pool, at least one heterogeneous acceleration pool, and an Ethernet switching module connected between the computing resource pool and the at least one heterogeneous acceleration pool.

[0066] Among them, the at least one heterogeneous acceleration pool comprises: a plurality of acceleration cards; each acceleration card comprises: an Ethernet controller and at least one operation core, and the Ethernet controller and the operation core communicate in a remote direct data access mode. Each acceleration card in the at least one heterogeneous acceleration pool communicates with the Ethernet switching module in an Ethernet mode or a remote direct data access mode. The computing resource pool comprises: a plurality of computing cores, and a same computing core has a binding relationship with at least one acceleration card in the at least one heterogeneous acceleration pool. The plurality of computing cores include a management core, the management core is used for managing each acceleration card in the at least one heterogeneous acceleration pool and the binding relationship, and distributing processing tasks to each acceleration card in the at least one heterogeneous acceleration pool.

[0067] It should be noted that each acceleration card includes a circuit structure for realizing related data operation, which can be specifically: GPU, FPGA, etc., which are collectively referred to as operation cores in the embodiment. The computing resource pool is composed of a large number of processor cores, and a processor core is: a computing core; which includes a processor core specially used as a management end, which is referred to as a management core in the embodiment.

[0068] In an implementation manner, each acceleration card in the at least one heterogeneous acceleration pool comprises: a golden finger and a power connector; the golden finger and the power connector connect each device in the corresponding acceleration card and supply power for each device in the corresponding acceleration card. Among them, the golden finger and the power connector in each acceleration card in the at least one heterogeneous acceleration pool supply power at the same time in the starting process of the corresponding acceleration card.

[0069] In an embodiment, the Ethernet controller in each of the accelerator cards in the at least one heterogeneous acceleration pool is connected with the computing core in the corresponding accelerator card via a switch chip; accordingly, the Ethernet controller in each of the accelerator cards in the at least one heterogeneous acceleration pool and the computing core in the corresponding accelerator card communicate with each other in a remote direct memory access mode via the switch chip.

[0070] In an embodiment, each of the accelerator cards in the at least one heterogeneous acceleration pool comprises a homogenous clock device connected with the computing core, the switch chip and the Ethernet controller in the corresponding accelerator card. In the at least one heterogeneous acceleration pool, each of the accelerator cards comprises a clock generator connected between the homogenous clock device and the computing core, between the homogenous clock device and the switch chip, and between the homogenous clock device and the Ethernet controller in the corresponding accelerator card; accordingly, the clock generator is configured to select a clock source for the computing core, the switch chip and the Ethernet controller in the corresponding accelerator card; the clock source is the homogenous clock device in the corresponding accelerator card or a clock device in the computing core bound to the corresponding accelerator card.

[0071] In an embodiment, each of the accelerator cards in the at least one heterogeneous acceleration pool comprises a control unit; the control unit is configured to implement timing control of information transmission in the corresponding accelerator card. The control unit can be a CPLD, an MCU (Microcontroller Unit) or the like.

[0072] In an embodiment, each of the accelerator cards in the at least one heterogeneous acceleration pool comprises a link selector; the link selector is connected with the computing core, the switch chip, the Ethernet controller, the homogenous clock device, the power device and the control unit in the corresponding accelerator card.

[0073] In an embodiment, the Ethernet switching module comprises at least one switch; the at least one switch is connected with the computing resource pool and the at least one heterogeneous acceleration pool.

[0074] In an embodiment, each of the accelerator cards in the at least one heterogeneous acceleration pool comprises a first node controller; the first node controller is configured to collect device information and running information in the corresponding accelerator card; accordingly, each of the computing cores in the computing resource pool comprises a second node controller; the second node controller is configured to collect device information and running information in the corresponding computing core; accordingly, the Ethernet switching module comprises a central controller; the central controller is configured to collect the device information and running information in the accelerator cards collected by the first node controllers and the device information and running information in the computing cores collected by the second node controllers.

[0075] The first node controller is configured to collect state information of the acceleration chips in the corresponding acceleration card and sensor data of the card. The second node controller is configured to collect core running information in the corresponding computing core and sensor data. The central controller is configured to construct a topology graph including the computing cores in the computing resource pool and the acceleration cards in the at least one heterogeneous acceleration pool according to the collected information.

[0076] In an embodiment, the central controller is configured to synchronize the collected information to the management core in the computing resource pool.

[0077] In an embodiment, the management core is configured to formulate a corresponding task allocation strategy and a binding relationship adjustment strategy according to the received synchronization information.

[0078] In an embodiment, the central controller is configured to generate log data according to the collected information, and analyze the log data to perform fault diagnosis.

[0079] In an embodiment, the computing system further includes a bus exchange module connected to the at least one heterogeneous acceleration pool.

[0080] It can be seen that in the embodiment, the heterogeneous acceleration card with the built-in Ethernet controller forms the heterogeneous acceleration pool, the heterogeneous acceleration pool is connected to the Ethernet switch module in an Ethernet mode or a remote direct memory access mode, and the Ethernet switch module is connected to the computing resource pool. The management core in the computing resource pool can manage the acceleration cards in the at least one heterogeneous acceleration pool, the binding relationship between the computing cores and the acceleration cards, and distribute processing tasks to the acceleration cards in the at least one heterogeneous acceleration pool, so that the acceleration cards in the at least one heterogeneous acceleration pool cooperatively complete the processing tasks. The computing system for model training task constructed in the embodiment accelerates data transmission in an Ethernet mode or a remote direct memory access mode, can not only provide sufficient computing power support for the running of the task, but also provide high bandwidth and low latency for high-speed communication and large data transmission of the task, and is beneficial to the rapid and stable operation of the model training task.

[0081] The application can directly encapsulate GPU data into Ethernet frames to realize direct communication between Ethernet and GPU for the heterogeneous acceleration card. The heterogeneous acceleration card supports dual-port 100G Ethernet and RDMA (Remote Direct Memory Access) functions, supports GPU chip, Ethernet controller integration and GPU memory remote direct access functions through software stack and driver, and thus realizes direct data flow from GPU to Ethernet.

[0082] The design core of the heterogeneous acceleration card is an Ethernet controller, which is mainly responsible for data transmission and network interaction, can provide high-bandwidth, ultra-low-latency Ethernet service, and meets the demand of high-speed transmission of the entire system network. In an example, the Ethernet controller adopts two SFP28 type input / output interfaces for external connection, transmits the external computing task to the inside of the acceleration card through network transmission; adopts PCIe Gen4 X8 connection for internal connection, and transmits the data transmitted from the outside to the GPU chip based on the Peer-to-Peer communication mechanism, and the GPU chip processes the data. When the GPU chip and the Ethernet controller interact with each other, based on the remote direct access function of the GPU memory, the point-to-point transmission of the GPU memory data between different GPUs is realized. The acceleration card integrated with the network card (i.e. the Ethernet controller) can realize ultra-low latency and support PCIe Gen4 full link.

[0083] Referring to FIG. 2, a GPU acceleration card principle block diagram is shown in FIG. 2. In FIG. 2, the GPU acceleration card is a PCIE (Peripheral Component Interconnect Express, a high-speed serial computer expansion bus standard) form card, including a PCIE gold finger connector and a power connector, and externally connects two 100G networks and connects to a switch through an optical module and an optical fiber line, and can perform data transmission with other GPUs. The uplink PCIE Gen4x16 link of the GPU acceleration card is connected to the gold finger, and two PCIE Gen4x16 links are connected to the network control chip and the GPU acceleration chip by the switch. The card form of the GPU acceleration card is a PCIe card form, full height 111.15mm, length 267.7mm, double width 39.04m). The key chips of the card include a GPU acceleration chip (i.e. operation core), a switch chip (i.e. switching chip), a network chip (i.e. network card) and a CPLD (Complex Programmable Logic Device, Complex Programmable Logic Device). The PCIE Gen4x16 of the gold finger accesses the PCIE_SW, and the PCIE_SW outputs another two PCIE X16 respectively connected to the GPU chip and the NIC (Network Interface Card) chip. The GPU chip realizes high-performance data processing and computing functions, the switch realizes port expansion of the PCIE, the NIC can realize 100G network optical port expansion, and completes network virtualization, offload acceleration, data flow forwarding and other work, wherein the CPLD chip is used for interrupt timing control and information interaction. The heterogeneous computing acceleration card fuses the hardware heterogeneous computing environment through software, so that the heterogeneous processors can communicate and transmit data through the PCIE high-speed bus and high-performance network Ethernet, realize cross-platform operation of computing tasks, and complete collaborative computing through higher-level system division and task scheduling, realize higher computing performance and better computing efficiency. The acceleration card can be deployed in various forms such as single host 4 cards, 8 cards, 16 cards or 32 cards, to adapt to the landing needs of various industries, and can be widely applied to fields such as intelligent finance, intelligent recommendation, fast search, content review, artificial intelligence generation, intelligent customer service and the like.

[0084] As shown in FIG. 3, the default power supply mode of the GPU acceleration card is that the power connector input P12V is used for whole board power supply, and the gold finger provides P3V3_STBY for CPLD and FRU power supply. Considering that the power supply in the whole machine may cause power supply shortage when the power supply transient impact occurs, the P12V_GF of the PCIE gold finger is reserved for the power supply of the SW and the NIC, which can share about 50W of power consumption. The board card 12V power supply is obtained through the board cable connection of the host mainboard, and the maximum power is 270W. The board card 3V3 power supply is converted to the whole board power supply through the P12V on the board card. The board card 3V3_STBY power supply is obtained through the gold finger connection of the HOST mainboard, and the power meets the PCIe 5.0 standard. At the same time, for the power supply transient impact application scene (such as: GPU startup, dynamic hot plug scene, etc.), the board card reserves a selected soldering resistor, which connects the gold finger and the Power Connector 12V power supply, and the gold finger simultaneously provides 3V3_STBY and 12V (60W Max).

[0085] As shown in FIG. 4, the GPU card accesses the 100MHz homologous clock signal from the PCIe gold finger, and accesses the Clock Buffer, which expands 3 paths of 100MHz clock to the GPU, the switch chip and the NIC chip. At the same time, the Clock Generator non-homologous clock scheme is reserved to avoid clock delay and jitter caused by long line. If the clock comes from the host, the clock line is long, but the embodiment can select the clock: the host or the local.

[0086] It should be noted that the GPU and the Ethernet controller NIC are directly connected through the switch chip PCIe SW, which can cope with the computing power challenge of large models. The acceleration card is connected with the Ethernet controller through the switch chip, which is different from the traditional NIC card connection switch chip connection CPU, and then the CPU and the network card are interconnected. The heterogeneous computing acceleration card integrates the GPU chip and the Ethernet controller, realizes the direct data flow from the GPU to the RDMA, as shown in FIG. 5, the GPU and the Ethernet controller NIC are connected through the switch chip Switch, and at the same time, the communication is realized in the memory remote direct access mode. Moreover, the heterogeneous computing acceleration card supports double-port 100G Ethernet and RDMA function.

[0087] In an example, the system RDMA communication software stack is shown in FIG. 6. Generally, communication between cluster nodes is completed through a network. For an RDMA-enabled network card device, an interface needs to be called to initiate a network transmission request. The driver in the software stack supports registration of RDMA network card related interfaces. The RDMA network card will interwork with the GPU kernel driver through the standard interface of the kernel driver to complete direct data transmission and realize remote direct memory access. The entire data path only involves the GPU and the RDMA network card, avoiding redundant jumps of data to the system memory and effectively reducing communication delay.

[0088] As can be seen in this example, GPUs can be directly connected; GPUs are directly connected to Ethernet controllers through the switch chip; host CPUs are interconnected through Ethernet controllers and switches. Among them, the heterogeneous computing accelerator card realizes the combination of the GPU chip and the Ethernet chip under the same root port through the direct connection of the switch chip and the Ethernet controller; through the high-bandwidth and low-latency interconnection RDMA technology, effective aggregation between chips is realized.

[0089] In an example, the architecture diagram of the computing system can refer to FIG. 7. The general computing resource pool (corresponding to another computing resource pool) is interconnected with the GPU and FPGA and other heterogeneous computing acceleration resource pools through the internal bus and the internal bus exchange module. In addition, the heterogeneous acceleration card realizes direct remote access to the memory in the heterogeneous acceleration card through the Ethernet to realize the RDMA function, reduces the data transmission delay of the acceleration card, and improves the computing performance. The whole system can be horizontally expanded to realize interconnection between systems, and further realize larger-scale heterogeneous computing acceleration pooling. Among them, through the high-performance exchange of the system internal bus and the software-defined system design, the device resource pooling is realized, the binding relationship between the device and the CPU at the physical link layer is released, the port configuration and resource allocation path of the exchange network can be flexibly adjusted, the shared resource pool can be finely divided, and the on-demand elastic allocation and multi-host sharing of device resources are realized. In view of the problems that the performance expansion of the current data center host system is limited by the system interconnection bandwidth, the performance between the storage levels is not matched, the I / O resource utilization is low, and the like, through the pooling system design, the dynamic allocation and load balancing of resources are flexibly realized for diversified scene requirements, the fusion of multi-platform processor computing power and the cooperative scheduling of heterogeneous acceleration resources are realized, and the performance expansion bottleneck problem of the data center is relieved.

[0090] The heterogeneous computing acceleration resource pool (i.e., heterogeneous acceleration pool) shown in FIG. 7 can maximize computing capability and efficiency. According to the differences between resource categories such as general computing, heterogeneous computing, and the like, the same type of computing resources are integrated to form a resource pool, and resources among different devices are recombined on demand. Through hardware reconstruction, fine-grained segmentation and pooling of resources are achieved, and general computing resources and heterogeneous computing resources are more closely combined. The super-high-speed internal and external interconnection technology of full interconnection is used to connect each resource pool, and the fusion of heterogeneous computing power is achieved. At the same time, the computing resources can realize task scheduling and load balancing according to the business scenario. Through software definition, a business-aware resource reconstruction decision system is established to complete intelligent reconstruction of hardware resources, realize pooling and centralized management of resources, and thus complete dynamic adjustment, flexible combination, and intelligent allocation, improve the computing efficiency and response rate of the whole system, and realize the intelligentization and high efficiency of heterogeneous computing.

[0091] In an example, the topology inside the heterogeneous computing acceleration resource pool and the internal bus exchange module is refined, and FIG. 8 corresponding to FIG. 7 can be obtained. As shown in FIG. 8, the internal bus exchange module (high-performance exchange chip in FIG. 8) includes a plurality of exchange chips, and the Ethernet exchange module (network exchange unit in FIG. 8) includes a plurality of network switches NET.

[0092] It should be noted that the internal bus exchange module includes not only a plurality of exchange chips, but also CPLD, management controller, baseboard controller, network exchange chip, and UART (Universal Asynchronous Receiver / Transmitter) chip; one heterogeneous computing acceleration resource pool can include not only a plurality of GPUs, but also baseboard controller and CPLD, and the like, as shown in FIG. 9. That is, each resource pool can realize flexible configuration of resources through reconstruction and decoupling, and the local multi-channel monitoring and diagnosis network is interconnected with the processor, which focuses on monitoring and management and fault detection at the hardware layer, collects the running state information of the host system in real time and performs monitoring and management, can also be connected to many sensors to read environmental conditions and control temperature through fans, and also supports other system management functions, including remote power control, serial local area network, monitoring and error recording of server host and memory. At the same time, when the system controller diagnoses a fault, it can be displayed in a comprehensive and intuitive manner through light diagnosis of the panel indicator light, and the host system fault and early warning can be indicated in a clear form.

[0093] Compared with traditional AI servers, the heterogeneous computing acceleration resource pool supports two P2P communication modes, as shown in FIG. 10. In addition to the PCIe link P2P communication, direct connection technology with an Ethernet controller through a switch chip can also be used. The RDMA communication library needs to call the interface to initiate a network transmission request, and the driver in the software stack supports registering the RDMA network card related interface. The RDMA network card will interwork with the GPU kernel driver through the standard interface of the kernel driver to complete the direct transmission of data and realize the direct access of the video memory at a remote end. The entire data path only involves the GPU and the RDMA network card, avoiding redundant jumps of data to the system memory and effectively reducing communication delay. Pearl represents an internal bus exchange module.

[0094] In an example, the servers are interconnected through heterogeneous acceleration cards and switches, and the hardware interconnection of the multi-training server whole machine is formed to realize the cluster computing power. A feasible topology is shown in FIG. 11. In FIG. 11, the distributed acceleration server GPU node P2P realizes cluster expansion through a parameter plane switch, and 1:1 non-blocking; the service plane and storage plane cluster networks are connected to separate switches, which can support 1:1 non-blocking or 2:1 convergence; the cluster out-of-band management network is connected to a TOR out-of-band management switch.

[0095] Since the management of a single acceleration card component is very important in a large-scale cluster system, the embodiment also realizes the overall out-of-band management of the system for the heterogeneous computing acceleration card. The main functions include device monitoring, log management, fault diagnosis, configuration management, and the like, as shown in FIG. 12.

[0096] Hardware device monitoring: The server monitors the hardware devices of the acceleration card through the out-of-band path by the management chip BMC through the I2C bus protocol to monitor the components in the acceleration card system, including the state of the acceleration chip, the temperature and voltage of the board card, and the like. The state of each chip in the acceleration card is collected through the boundary scan chain to realize fault positioning; control codes are sent to each chip through the boundary scan chain to realize the functions of turn-on test, hardware configuration, module reset, fault isolation, and the like. Temperature monitoring: The management chip dynamically obtains the heat dissipation requirements and temperature information required by the working environment of the acceleration card. For example, the temperature of the heterogeneous acceleration chip, the temperature of the memory of the heterogeneous acceleration chip, the temperature at the air inlet and outlet, and the like. Voltage monitoring: The power chip is accessed to monitor the dynamic board-level voltage state.

[0097] Fault diagnosis and log management: The system BMC can obtain error information such as overvoltage, undervoltage and overtemperature from the acceleration chip through the I2C bus protocol, and can also obtain the serial number, firmware version and driver version of the acceleration chip, as well as the working voltage of the chip, the power consumption of the board and other information. When the acceleration card system fails, the BMC will record the log generated by the real-time monitoring of the machine state in the local file system. The on-site technical personnel can export the log through the BMC web page or other tools, analyze the log, locate the problem and solve the problem.

[0098] Configuration management: Support for configuration of network, user, alarm and other software, and support for import and export functions of configuration files.

[0099] In terms of acceleration card management, the computing unit and the management unit are decoupled and standardized, and the common management, security and control functions are separated from the computing unit. This way, different computing power platforms and management platforms can be compatible, supporting multiple interfaces and unified management of multiple cores, meeting the needs of different application scenarios. The management unit needs to implement computing power allocation and management, multi-core module voltage regulation and power management, ensure system efficiency design; realize the monitoring of resource utilization, I / O throughput and resource health status, realize collaborative power-on and power-off, centralized management of resources and topology, and guarantee system availability; realize on-demand rapid deployment and automatic management of hardware resources, monitoring and fault management of key resource information, intelligent positioning and recovery of faults, and guarantee the reliability of the computing system. The system management module is responsible for unified management, providing operation and maintenance capabilities through standardized service interfaces, realizing integrated monitoring, fault warning, visual management, etc.

[0100] The system power-on and power-off collaborative control process is as follows:

[0101] 1. After the PSU in each unit is powered on, the VR outputs each STBY power, and sends the PG of the last stby power to the CPLD.

[0102] 2. When the STBY power in each chassis is turned on, the CPLD and BMC work normally, and the host and the CPLD in the resource pool send the STBY PG signal to the BMC through I2C / UART.

[0103] 3. The host and resource pool BMC send the STBY PG completion signal to the BMC of the core switching unit through the network. At this time, each unit waits for the key-on signal of the switching unit to perform the next operation.

[0104] 4. After pressing the power-on key of the core switching unit, the power-on signal is sent to the CPLD, and the CPLD controls the EN signal of the Main Power.

[0105] 5. At the same time, the Power on signal is sent to the BMC, and the BMC sends the power-on signal to the BMC of each chassis through the network.

[0106] 6. The host and the BMC of the resource pool send the power-on signal to the respective CPLD through I2C / UART, the CPLD controls the EN of each VR, and after receiving the last power-on signal, the BMC sends the power-on completion signal to the BMC of the core switching unit and the mCPU through the network.

[0107] The system power-on reset workflow is as follows:

[0108] 1. After the STBY power-on is completed, the BMC of the host scans the ID of the local interconnection cable CDFP interface and the corresponding CDFP ID of the core switching unit to establish a mapping table.

[0109] 2. Similarly, the BMC of the core switching unit scans the ID of the local CDFP interface and the corresponding CDFP ID of the resource pool to establish a mapping table.

[0110] 3. After the power-on of each unit is completed, the PERST signal is waited for, the host sends the PERST to each module by the CPU and the CPLD, and the BMC of the host sends the PERST information to the BMC of the core switching unit through the TOR switch.

[0111] 4. The PERST of the SW in the core switching unit is triggered by the PLTRST of the mCPU, the PERST of the SW is informed to the CPLD by the BMC through I2C, the CPLD sends the PERST to the SW and the CDFP, and the resource pool sends the PERST from the interconnection cable CDFP to the CPLD, which is sent to each GPU device by the CPLD.

[0112] Specifically, please refer to FIG. 13, and the implemented acceleration card management functions are as follows:

[0113] Topology identification: support for viewing the topology view, and can view node summary information such as node type, power-on state, overall health status, etc.

[0114] Asset information management: support for viewing device information level by level, and can view detailed device information such as device asset information and high-speed interface connection state through the Web / Redfish page.

[0115] Collaborative control: support for centralized power-on and power-off control, each unit is powered on and powered off in sequence; support for system reset function, automatically detects the reset signal for reset during the power-on process and the host restart process; support for reset control after resource reallocation, etc.

[0116] In system design, the discovery, management and flexible adjustment of system resources are critical. The control module is the core management unit of dynamic resource adjustment, which controls the high-performance switching unit to realize automatic discovery of resource topology and flexible automatic switching of resources, including: resource identification and display: providing network interface and visual WEB interface for resource list and topology display, including heterogeneous computing resource unit information, I / O port information, etc.; resource dynamic allocation: through key technologies such as hot removal, hot insertion and hot reset of devices, the dynamic allocation and adjustment of heterogeneous computing units are realized within seconds; load balancing: based on the optimization scheduling algorithm of load balancing, the physical resources are dynamically balanced and allocated to maximize the release of heterogeneous computing power; expert template: according to business requirements, resource state and performance indicators, etc. Parameters, realize the expert template description and dynamic switching interface of heterogeneous computing power resources, so that applications can apply for resources according to the expert template, and at the same time provide network interface and visual WEB interface to realize the visualization of the expert template application.

[0117] In order to fully exert the performance of heterogeneous multi-core and adapt to various application scenarios, it is very important for applications to be able to freely migrate between multiple cores. However, migration between different instruction set heterogeneous multi-cores has always been a difficult problem in the industry. Load migration and resource scheduling are applied in the following scenarios: the operating system performs load balancing, and the process is migrated to a static or lightly loaded core; according to different load types, the process is migrated to a core with a different instruction set architecture; when the power state changes, some processes need to be migrated to achieve power control; when a core overheats, the process is migrated to another core, etc.

[0118] The operating system of a general-purpose computing node (i.e. host) contains multiple kernels, each of which is compiled and run under a specific ISA instruction set. The entire system implements balanced scheduling of computing power on demand through unified system software for applications. At the operating system level, inter-kernel communication and collaboration are implemented, so that the entire system still has a global state under different ISA core instances. On this basis, through a unified scheduling core and advanced scheduling techniques, the application program is unaware of the underlying hardware, and through the cooperation of the operating system and the scheduling system, the application can automatically select the optimal mapping to achieve performance improvement of the multi-instruction set heterogeneous multi-core system, as shown in FIG. 14.

[0119] For a general-purpose computing node, the embodiment provides a heterogeneous virtual computing power management subsystem. The system combines heterogeneous resource computing power virtualization and remote scheduling technology at the technical level, and combines the use and management needs of users and administrators for heterogeneous virtualization resources at the application level. Please refer to FIG. 15 for details. The system ensures that different users have the following basic capabilities when using virtualization resources in different scenarios:

[0120] Transparency: The system enables users to invoke heterogeneous virtual computing power execution applications locally or remotely without modifying system component code and user code, just like executing on physical resources.

[0121] Low overhead: The system enables applications running on it to perform as close to physical resources as possible when using heterogeneous virtual computing power.

[0122] Isolation: The system manages the allocation and release of heterogeneous virtual computing power for each container, enabling containers to be completely isolated from each other while sharing physical resources.

[0123] To achieve the above goals, the heterogeneous virtual computing power management subsystem mainly includes three components. The heterogeneous virtual computing power management platform is responsible for the front-end interaction logic for end users and administrators. The heterogeneous virtual computing power management center component is responsible for resource scheduling and management of the entire system. The heterogeneous virtual computing power management node component is responsible for specific operation execution on each computing node.

[0124] Heterogeneous virtual computing power management platform: a user and administrator-oriented heterogeneous resource scheduling and management platform responsible for receiving user job requests and creating virtualized computing environments based on specific business needs. It also receives and aggregates node device information from node components and distributes the device aggregation information to the center component. Specifically, the platform includes the following functions:

[0125] User management: The platform manages users and roles by storing user and related group information, permissions, etc.

[0126] Resource management: The platform manages all heterogeneous computing power physical devices and virtual devices by providing different views of these devices, including node resource summary, node state information, virtual device disable / enable / move, etc.

[0127] Quota management: The platform manages and allocates virtual heterogeneous computing power resources by configuring specific virtual heterogeneous computing power resource quotas, including virtual total computing power limit, virtual total video memory limit, single virtual device computing power limit, and video memory limit, etc.

[0128] Strategy management: The platform configures and manages various scheduling strategies based on administrator's specific needs by abstracting different strategies of the scheduler in the center component. See the center component section for specific scheduling strategies.

[0129] Please refer to FIG. 16, a heterogeneous computing resource pool management architecture diagram, a Pooled System Management Controller (PSMC) in a high-performance switching unit is taken as a center, a heterogeneous computing acceleration unit and a general-purpose computing unit are respectively matched with a Pooled Node Management Controller (PNMC), the PSMC is taken as a central management node, the PNMC is taken as a distributed node and is uniformly managed by the PSMC, and the whole life cycle management of the heterogeneous computing resource pool is realized. In the heterogeneous computing resource pool system, the discovery, management and elastic adjustment of the heterogeneous computing resources are crucial, the Pooled Management Engine is designed as a core management unit of resource dynamic adjustment in the project, the Pooled Management Engine controls the high-performance switching unit to realize resource topology automatic discovery, resource flexible automatic switching, and provides a standard interface for a data center monitoring and management platform to realize massive resource centralized management at the data center level.

[0130] The PSMC is crucial in the management of the distributed heterogeneous acceleration pooled server and is a bridge for information communication, is responsible for unified management work, and mainly includes node topology identification, asset information centralized display and cooperative power-on and power-off. Through a standardized service interface, an operation and maintenance capability is provided to the outside, integrated monitoring, fault early warning and visual management are realized. The heterogeneous computing resource pool whole machine management system realizes interconnection of the pooled node management controllers at all levels through a network, and builds a management platform of independent engines.

[0131] A computing method provided by an embodiment of the present application is introduced below, and the computing method described below can be mutually referred to with other embodiments described herein.

[0132] The embodiment of the present application discloses a computing method applied to a management core, including: receiving a processing task; and sending the processing task to an Ethernet switching module in a computing system, so that the Ethernet switching module distributes the processing task to at least one heterogeneous acceleration pool in the computing system in an Ethernet mode or a remote direct memory access mode.

[0133] The computing system includes: a computing resource pool, at least one heterogeneous acceleration pool and an Ethernet switching module connected between the computing resource pool and the at least one heterogeneous acceleration pool; the at least one heterogeneous acceleration pool includes: a plurality of acceleration cards; each acceleration card includes: an Ethernet controller and at least one operation core, the Ethernet controller and the operation core communicate in a remote direct memory access mode; each acceleration card in the at least one heterogeneous acceleration pool communicates with the Ethernet switching module in an Ethernet mode or a remote direct memory access mode; the computing resource pool includes: a plurality of computing cores, and a same computing core has a binding relationship with at least one acceleration card in the at least one heterogeneous acceleration pool; and the management core is any one of the plurality of computing cores.

[0134] It can be seen that the embodiment provides a computing method, which accelerates data transmission in an Ethernet mode or a remote direct data access mode, can not only provide sufficient computing power support for task running, but also provide high bandwidth and low latency for high-speed communication and large data transmission of the task, and is beneficial to rapid and stable operation of model training tasks.

[0135] Next, an electronic device provided by an embodiment of the present application is introduced, and the electronic device described below can be referred to with other embodiments described herein.

[0136] An electronic device is disclosed by an embodiment of the present application, comprising:

[0137] a memory for saving computer readable instructions;

[0138] a processor for executing the computer readable instructions to implement the method disclosed by any of the above embodiments.

[0139] Further, an electronic device is also provided by an embodiment of the present application. The electronic device can be a server as shown in FIG. 17 or a terminal as shown in FIG. 18. FIG. 17 and FIG. 18 are structural diagrams of electronic devices according to an exemplary embodiment, and the contents in the figures should not be considered as any limitation on the use range of the present application.

[0140] FIG. 17 is a structural schematic diagram of a server provided by an embodiment of the present application. The server can specifically include at least one processor, at least one memory, a power supply, a communication interface, an input / output interface and a communication bus. The memory is used to store computer readable instructions, which are loaded and executed by the processor to implement the related steps in the calculation disclosed by any of the preceding embodiments.

[0141] In the embodiment, the power supply is used to provide working voltage for each hardware device on the server; the communication interface can create a data transmission channel between the server and external devices, and the communication protocol followed by the communication interface is any communication protocol applicable to the technical solution of the present application, which is not specifically limited here; the input / output interface is used to obtain external input data or output data to the outside world, and the specific interface type can be selected according to the specific application needs, which is not specifically limited here.

[0142] In addition, the memory as a carrier for resource storage can be a read-only memory, a random access memory, a magnetic disk or an optical disk, etc., and the resources stored thereon include an operating system, computer readable instructions and data, etc., and the storage mode can be temporary storage or permanent storage.

[0143] The operating system is used to manage and control various hardware devices on the server and computer readable instructions to enable the processor to operate and process data in the memory, which can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer readable instructions that can be used to complete the computing method disclosed in any of the preceding embodiments, the computer readable instructions can further include computer readable instructions that can be used to complete other specific work. In addition to the data including application update information and other data, the data can also include application developer information and other data.

[0144] FIG. 18 is a structural schematic diagram of a terminal according to an embodiment of the present application. The terminal can specifically include, but is not limited to, a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.

[0145] Generally, the terminal in the embodiment includes a processor and a memory.

[0146] The processor can include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor can be implemented in at least one of a hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), and a PLA (Programmable Logic Array). The processor can also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor can be integrated with a GPU (Graphics Processing Unit) that is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor can also include an AI (Artificial Intelligence) processor for processing machine learning-related computing operations.

[0147] The memory can include one or more non-transitory computer-readable storage media that stores computer-readable instructions. The memory can also include high-speed random access memory and non-volatile memory such as one or more disk storage devices, flash memory devices. In this embodiment, the memory is used to store at least the following computer-readable instructions, wherein the computer-readable instructions are loaded and executed by the processor, and can realize the related steps in the computing method executed by the terminal side disclosed in any of the preceding embodiments. In addition, the resources stored in the memory can also include operating systems, data, etc., and the storage mode can be temporary storage or permanent storage. The operating system can include Windows, Unix, Linux, etc. The data can include but is not limited to application update information.

[0148] In some embodiments, the terminal can also include a display screen, an input / output interface, a communication interface, a sensor, a power supply, and a communication bus.

[0149] Those skilled in the art can understand that the structure shown in FIG. 18 does not constitute a limitation on the terminal, and can include more or fewer components than shown.

[0150] A non-volatile storage medium provided by an embodiment of the present application is introduced below, and the non-volatile storage medium described below can be referred to with other embodiments described herein.

[0151] One or more non-volatile computer-readable storage media storing computer-readable instructions are used to save computer-readable instructions, wherein the computer-readable instructions are executed by the processor to realize the computing method disclosed in the preceding embodiments. The non-volatile computer-readable storage medium is a computer-readable non-volatile storage medium, which is a carrier for storing resources, and can be a read-only memory, a random access memory, a magnetic disk or an optical disk, etc. The resources stored thereon include an operating system, computer-readable instructions and data, etc., and the storage mode can be temporary storage or permanent storage.

[0152] A computer program product provided by an embodiment of the present application is introduced below, and the computer program product described below can be referred to with other embodiments described herein.

[0153] A computer program product includes computer-readable instructions, which are executed by the processor to realize the steps of the computing method disclosed above.

[0154] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other.

[0155] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in random access memory (RAM), non-volatile memory (ROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disk, a removable disk, a CD-ROM, or any other form of non-transitory storage medium known in the art. The processor can be configured to execute the software module.

[0156] The principles and implementations of the present application have been described in relation to specific examples, which are presented only by way of illustration and for purposes of description and should not be construed as limiting the present application.

Claims

1. A computing system, comprising: The application relates to a computing resource pool, at least one heterogeneous acceleration pool and an Ethernet switch module connected between the computing resource pool and the at least one heterogeneous acceleration pool. The at least one heterogeneous acceleration pool comprises a plurality of acceleration cards; each acceleration card comprises an Ethernet controller and at least one operation core, and the Ethernet controller communicates with the operation core in a remote direct memory access mode. Each acceleration card in the at least one heterogeneous acceleration pool communicates with the Ethernet switch module in an Ethernet mode or a remote direct memory access mode. The computing resource pool comprises a plurality of computing cores, and each computing core has a binding relationship with at least one acceleration card in the at least one heterogeneous acceleration pool. The plurality of computing cores comprise a management core, which is used for managing each acceleration card in the at least one heterogeneous acceleration pool and the binding relationship, and distributing processing tasks to each acceleration card in the at least one heterogeneous acceleration pool. Each acceleration card in the at least one heterogeneous acceleration pool comprises a golden finger and a power connector; the golden finger and the power connector are connected with each device in the corresponding acceleration card and supply power for each device in the corresponding acceleration card.

2. The computing system of claim 1, wherein, The golden finger and the power connector in each acceleration card in the at least one heterogeneous acceleration pool supply power simultaneously in the process of starting up the corresponding acceleration card.

3. The computing system of claim 2, wherein, An exchange chip is connected between the Ethernet controller in each acceleration card in the at least one heterogeneous acceleration pool and the operation core in the corresponding acceleration card.

4. The computing system of claim 1, wherein, Correspondingly, the Ethernet controller in each acceleration card in the at least one heterogeneous acceleration pool and the operation core in the corresponding acceleration card realize communication in a remote direct memory access mode through the exchange chip. Each acceleration card in the at least one heterogeneous acceleration pool comprises a homologous clock device, which is connected with the operation core, the exchange chip and the Ethernet controller in the corresponding acceleration card.

5. The computing system of claim 1, wherein, Each acceleration card in the at least one heterogeneous acceleration pool comprises a clock generator, which is connected between the homologous clock device and the operation core, between the homologous clock device and the exchange chip and between the homologous clock device and the Ethernet controller in the corresponding acceleration card.

6. The computing system of claim 5, wherein, Correspondingly, the clock generator is used for selecting a clock source for the operation core, the exchange chip and the Ethernet controller in the corresponding acceleration card; the clock source is the homologous clock device in the corresponding acceleration card or a clock device in the bound computing core. Each acceleration card in the at least one heterogeneous acceleration pool comprises a control unit, which is used for realizing time sequence control of information transmission in the corresponding acceleration card.

7. The computing system of claim 1, wherein, Each acceleration card in the at least one heterogeneous acceleration pool comprises a link selector, which is connected with the operation core, the exchange chip, the Ethernet controller, the homologous clock device, a power device and the control unit in the corresponding acceleration card.

8. The computing system of claim 1, wherein, The Ethernet switch module comprises at least one switch; the at least one switch is connected with the computing resource pool and the at least one heterogeneous acceleration pool.

9. The computing system of claim 1, wherein, Each acceleration card in the at least one heterogeneous acceleration pool comprises a first node controller, which is used for collecting device information and running information in the corresponding acceleration card.

10. The computing system of claim 1, wherein, ​ Correspondingly, each computing core in the computing resource pool comprises a second node controller; the second node controller is configured to collect device information and running information in the corresponding computing core; Correspondingly, the Ethernet switch module comprises a central controller; the central controller is configured to collect the device information and running information in the acceleration card collected by the first node controller, and the device information and running information in the computing core collected by the second node controller.

11. The computing system of claim 10, wherein, The first node controller is configured to collect state information of the acceleration chip and sensor data of the board card in the corresponding acceleration card.

12. The computing system of claim 10, wherein, The second node controller is configured to collect core running information and sensor data in the corresponding computing core.

13. The computing system of claim 10, wherein, The central controller is configured to construct a topology graph comprising each computing core in the computing resource pool and each acceleration card in the at least one heterogeneous acceleration pool according to the collected information.

14. The computing system of claim 10, wherein, The central controller is configured to synchronize the collected information to the management core in the computing resource pool. Correspondingly, the management core manages each acceleration card in the at least one heterogeneous acceleration pool and the binding relationship according to the received synchronization information, and distributes processing tasks to each acceleration card in the at least one heterogeneous acceleration pool.

15. The computing system of claim 10, wherein, The management core is configured to formulate a corresponding task allocation strategy and a binding relationship adjustment strategy according to the received synchronization information.

16. The computing system of claim 10, wherein, The central controller is configured to generate log data according to the collected information, and analyze the log data to perform fault diagnosis.

17. The computing system of any one of claims 1 to 16, wherein, The computing system further comprises a bus switch module connected to the at least one heterogeneous acceleration pool; the bus switch module is further connected to another computing resource pool.

18. A computing method, comprising: Applied to the management core, comprising: receiving a processing task; and sending the processing task to an Ethernet switch module in a computing system, so that the Ethernet switch module distributes the processing task to at least one heterogeneous acceleration pool in the computing system by Ethernet or remote direct memory access; The computing system comprises a computing resource pool, the at least one heterogeneous acceleration pool, and the Ethernet switch module connected between the computing resource pool and the at least one heterogeneous acceleration pool; The at least one heterogeneous acceleration pool comprises a plurality of acceleration cards; each acceleration card comprises an Ethernet controller and at least one computing core, and the Ethernet controller communicates with the computing core by remote direct memory access; Each acceleration card in the at least one heterogeneous acceleration pool communicates with the Ethernet switch module by Ethernet or remote direct memory access; The computing resource pool comprises a plurality of computing cores, and a same computing core has a binding relationship with at least one acceleration card in the at least one heterogeneous acceleration pool; the management core is any one of the plurality of computing cores.

19. An electronic device, comprising: Comprising: a memory for storing computer readable instructions; a processor for executing the computer readable instructions to implement the method of claim 18.

20. One or more non-transitory computer-readable storage media storing computer- readable instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of the method of claim 18.

21. A computer program product comprising computer readable instructions, characterized in that, The computer-readable instructions, when executed by a processor, implement the method of claim 18.

Citation Information

Patent Citations

  • Computing engine communication method and device

    CN116028238A

  • Heterogeneous acceleration board card calculation method and device, equipment and medium

    CN116192849A

  • Multi-accelerator card heterogeneous server and resource link reconstruction method

    CN117687956A

  • Computing system, method, device, medium and program product

    CN119201469A