A computing system, method, apparatus, medium, and program product

By constructing a computing system and using Ethernet switching modules to connect the computing resource pool and the heterogeneous acceleration pool, high-bandwidth and low-latency data transmission is achieved, solving the computing power requirements of large-scale model training tasks and improving the efficiency and stability of model training.

CN119201469BActive Publication Date: 2026-04-03LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Large-scale, long-term model training tasks require collaboration among numerous servers, resulting in high communication volumes and stringent requirements for bandwidth and latency. Existing networks cannot meet the computational power demands of model training, leading to increased time costs.

Method used

A computing system is constructed, including a computing resource pool and a heterogeneous acceleration pool, which are connected through an Ethernet switching module. The system utilizes an Ethernet controller and computing cores for remote direct data access, manages cores to distribute tasks, and achieves high-bandwidth, low-latency data transmission.

Benefits of technology

Provide sufficient computing power for model training tasks, enable high-speed communication and large data volume transmission, and improve the speed and stability of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119201469B_ABST
    Figure CN119201469B_ABST
Patent Text Reader

Abstract

This application discloses a computing system, method, device, medium, and program product in the field of computer technology. This application enables heterogeneous accelerator cards with built-in Ethernet controllers to form a heterogeneous acceleration pool. The heterogeneous acceleration pool is connected to an Ethernet switching module via Ethernet or remote direct data access, and the Ethernet switching module is connected to a computing resource pool. A management core in the computing resource pool can manage the binding relationships between each accelerator card in at least one heterogeneous acceleration pool, the computing core, and the accelerator cards, and distribute processing tasks to each accelerator card in at least one heterogeneous acceleration pool, enabling the accelerator cards in at least one heterogeneous acceleration pool to collaboratively complete the processing tasks. Accelerating data transmission via Ethernet or remote direct data access not only provides sufficient computing power for task execution but also provides high bandwidth and low latency for high-speed communication and large-volume data transmission, which is beneficial for the rapid and stable operation of model training tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a computing system, method, apparatus, medium and program product. Background Technology

[0002] Currently, model training requires massive computing power, typically necessitating a large number of servers acting as nodes, forming a cluster via a high-speed network. These servers interconnect and collaborate to complete the task. However, large-scale, long-duration model training tasks can require hundreds of gigabytes of communication just for a single computation iteration, not to mention the communication needs of various parallel modes. Insufficient network bandwidth and high latency not only lead to diminishing marginal returns on computing power but also increase the time cost of model training. Furthermore, large model training places high demands on latency and packet loss control.

[0003] Therefore, how to build a corresponding computing power system for model training tasks is a problem that needs to be solved by those skilled in the art. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a computing system, method, device, medium, and program product to build a corresponding computing power system for model training tasks. The specific solution is as follows:

[0005] In a first aspect, this application provides a computing system, including: a computing resource pool, at least one heterogeneous acceleration pool, and an Ethernet switching module connected between the computing resource pool and the at least one heterogeneous acceleration pool;

[0006] The at least one heterogeneous acceleration pool includes: multiple acceleration cards; each acceleration card includes: an Ethernet controller and at least one computing core, wherein the Ethernet controller and the computing core communicate via remote direct data access.

[0007] Each accelerator card in the at least one heterogeneous acceleration pool communicates with the Ethernet switching module via Ethernet or remote direct data access.

[0008] The computing resource pool includes: multiple computing cores, and the same computing core is bound to at least one accelerator card in the at least one heterogeneous acceleration pool;

[0009] The plurality of computing cores includes a management core, which is used to manage each accelerator card in the at least one heterogeneous acceleration pool and the binding relationship, and to distribute processing tasks to each accelerator card in the at least one heterogeneous acceleration pool.

[0010] Optionally, each accelerator card in the at least one heterogeneous accelerator pool includes: a gold finger and a power connector; the gold finger and the power connector connect to each device in the corresponding accelerator card and supply power to each device in the corresponding accelerator card.

[0011] Optionally, the gold fingers and power connectors of each accelerator card in the at least one heterogeneous accelerator pool are simultaneously powered during the power-on process of the corresponding accelerator card.

[0012] Optionally, a switching chip is connected between the Ethernet controller in each accelerator card and the computing core in the corresponding accelerator card in the at least one heterogeneous acceleration pool;

[0013] Accordingly, the Ethernet controller in each accelerator card in the at least one heterogeneous acceleration pool communicates with the computing core in the corresponding accelerator card through the switching chip via remote direct data access.

[0014] Optionally, each accelerator card in the at least one heterogeneous acceleration pool includes: a co-source clock device, which is connected to the computing core, switching chip and Ethernet controller in the corresponding accelerator card.

[0015] Optionally, each acceleration card in the at least one heterogeneous acceleration pool includes: a clock generator, which is connected between the same source clock device and the computing core, the same source clock device and the switching chip, and the same source clock device and the Ethernet controller in the corresponding acceleration card;

[0016] Accordingly, the clock generator is used to select a clock source for the computing core, switching chip and Ethernet controller in the corresponding accelerator card; the clock source is: a clock device of the same source in the corresponding accelerator card or a clock device in the computing core bound to the corresponding accelerator card.

[0017] Optionally, each accelerator card in the at least one heterogeneous accelerator pool includes: a control unit; the control unit is used to implement timing control of information transmission in the corresponding accelerator card.

[0018] Optionally, each acceleration card in the at least one heterogeneous acceleration pool includes: a link selector; the link selector connects to the computing core, switching chip, Ethernet controller, co-source clock device, power supply device and control unit in the corresponding acceleration card.

[0019] Optionally, the Ethernet switching module includes: at least one switch; the at least one switch connects the computing resource pool and the at least one heterogeneous acceleration pool.

[0020] Optionally, each accelerator card in the at least one heterogeneous acceleration pool includes: a first node controller; the first node controller is used to collect device information and operating information in the corresponding accelerator card;

[0021] Accordingly, each computing core in the computing resource pool includes: a second node controller; the second node controller is used to collect device information and operating information in the corresponding computing core;

[0022] Accordingly, the Ethernet switching module includes: a central controller; the central controller is used to collect device information and operating information from the accelerator card collected by the first node controller, and device information and operating information from the computing core collected by the second node controller.

[0023] Optionally, the first node controller is used to collect the status information of the acceleration chip in the corresponding acceleration card and the sensor data of the board.

[0024] Optionally, the second node controller is used to collect core operation information and sensor data in the corresponding computing core.

[0025] Optionally, the central controller is used to construct a topology map including each computing core in the computing resource pool and each accelerator card in the at least one heterogeneous acceleration pool based on the collected information.

[0026] Optionally, the central controller is used to synchronize the collected information to the management core in the computing resource pool;

[0027] Accordingly, the management core manages each accelerator card in the at least one heterogeneous acceleration pool and the binding relationship based on the received synchronization information, and distributes processing tasks to each accelerator card in the at least one heterogeneous acceleration pool.

[0028] Optionally, the management core is used to formulate corresponding task allocation strategies and binding relationship adjustment strategies based on the received synchronization information.

[0029] Optionally, the central controller is used to generate log data based on the collected information and analyze the log data for fault diagnosis.

[0030] Optionally, the computing system further includes a bus switching module connected to the at least one heterogeneous acceleration pool; the bus switching module is also connected to another computing resource pool.

[0031] Secondly, this application provides a calculation method applied to a management kernel, including:

[0032] Receive and process tasks;

[0033] The processing task is sent to the Ethernet switching module in the computing system, so that the Ethernet switching module distributes the processing task to at least one heterogeneous acceleration pool in the computing system via Ethernet or remote direct data access.

[0034] The computing system includes: a computing resource pool, the at least one heterogeneous acceleration pool, and the Ethernet switching module connected between the computing resource pool and the at least one heterogeneous acceleration pool;

[0035] The at least one heterogeneous acceleration pool includes: multiple acceleration cards; each acceleration card includes: an Ethernet controller and at least one computing core, wherein the Ethernet controller and the computing core communicate via remote direct data access.

[0036] Each accelerator card in the at least one heterogeneous acceleration pool communicates with the Ethernet switching module via Ethernet or remote direct data access.

[0037] The computing resource pool includes: multiple computing cores, each computing core being bound to at least one accelerator card in the at least one heterogeneous acceleration pool; the management core is any one of the multiple computing cores.

[0038] Thirdly, this application provides an electronic device, comprising:

[0039] Memory, used to store computer programs;

[0040] A processor for executing the computer program to implement the aforementioned disclosed computation method.

[0041] Fourthly, this application provides a non-volatile storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the aforementioned disclosed calculation method.

[0042] Fifthly, this application provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the aforementioned disclosed computation method.

[0043] As can be seen from the above scheme, this application provides a computing system, including: a computing resource pool, at least one heterogeneous acceleration pool, and an Ethernet switching module connected between the computing resource pool and the at least one heterogeneous acceleration pool; the at least one heterogeneous acceleration pool includes: multiple acceleration cards; each acceleration card includes: an Ethernet controller and at least one computing core, the Ethernet controller and the computing core communicate via remote direct data access; each acceleration card in the at least one heterogeneous acceleration pool communicates with the Ethernet switching module via Ethernet or remote direct data access; the computing resource pool includes: multiple computing cores, the same computing core having a binding relationship with at least one acceleration card in the at least one heterogeneous acceleration pool; the multiple computing cores include a management core, the management core being used to manage each acceleration card in the at least one heterogeneous acceleration pool and the binding relationship, and to distribute processing tasks to each acceleration card in the at least one heterogeneous acceleration pool.

[0044] As can be seen, the technical effects of this application are as follows: heterogeneous accelerator cards with built-in Ethernet controllers form a heterogeneous acceleration pool. The heterogeneous acceleration pool is connected to an Ethernet switching module via Ethernet or remote direct data access, and the Ethernet switching module is connected to a computing resource pool. The management core in the computing resource pool can manage the binding relationship between each accelerator card in at least one heterogeneous acceleration pool, the computing core, and the accelerator cards, and distribute processing tasks to each accelerator card in at least one heterogeneous acceleration pool, enabling each accelerator card in at least one heterogeneous acceleration pool to collaboratively complete the processing task. This application provides a computing power system for model training tasks. By accelerating data transmission via Ethernet or remote direct data access, it not only provides sufficient computing power support for task operation but also provides high bandwidth and low latency for high-speed communication and large data volume transmission, which is beneficial for the rapid and stable operation of model training tasks.

[0045] Correspondingly, the computing device, equipment, medium, and program product provided in this application also have the above-mentioned technical effects. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0047] Figure 1 This is a schematic diagram of a computing system disclosed in this application;

[0048] Figure 2 This is a block diagram of a GPU accelerator card disclosed in this application;

[0049] Figure 3 This is a power supply circuit diagram of a GPU accelerator card disclosed in this application;

[0050] Figure 4 This is a clock circuit diagram of a GPU accelerator card disclosed in this application;

[0051] Figure 5 This is a schematic diagram of communication between a GPU and an Ethernet controller disclosed in this application;

[0052] Figure 6 This is a schematic diagram of an RDMA communication software stack disclosed in this application;

[0053] Figure 7 This is an architecture diagram of a computing system disclosed in this application;

[0054] Figure 8 This is an architecture diagram of another computing system disclosed in this application;

[0055] Figure 9 This is an architecture diagram of another computing system disclosed in this application;

[0056] Figure 10 This is a schematic diagram of a communication method disclosed in this application;

[0057] Figure 11 This is a schematic diagram of a topology disclosed in this application;

[0058] Figure 12 This is a schematic diagram of an accelerator card management system disclosed in this application;

[0059] Figure 13 This is another schematic diagram of accelerator card management disclosed in this application;

[0060] Figure 14 This is a schematic diagram of kernel communication for a general-purpose computing node disclosed in this application;

[0061] Figure 15 This is a schematic diagram of a heterogeneous virtual computing power management subsystem disclosed in this application;

[0062] Figure 16 This is a schematic diagram of a heterogeneous computing power resource pool management architecture disclosed in this application;

[0063] Figure 17 A server architecture diagram provided for this application;

[0064] Figure 18 A terminal structure diagram provided for this application. Detailed Implementation

[0065] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other instances obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0066] Currently, model training requires massive computing power, typically necessitating a large number of servers as nodes forming a cluster via a high-speed network. These servers interconnect and collaborate to complete the task. However, large-scale, long-duration model training tasks can require hundreds of gigabytes of communication data for a single computation iteration, not to mention the communication needs of various parallel modes. Insufficient network bandwidth and high latency not only lead to diminishing marginal returns on computing power but also increase the time cost of model training. Furthermore, large model training places high demands on latency and packet loss. To address these issues, this application provides a computing scheme that constructs a corresponding computing power system for model training tasks. By accelerating data transmission via Ethernet or remote direct data access, it not only provides sufficient computing power for task execution but also offers high bandwidth and low latency for high-speed communication and large-volume data transmission, facilitating the rapid and stable operation of model training tasks.

[0067] See Figure 1 As shown in the figure, this application discloses a computing system, including: a computing resource pool, at least one heterogeneous acceleration pool, and an Ethernet switching module connected between the computing resource pool and the at least one heterogeneous acceleration pool.

[0068] The system includes at least one heterogeneous acceleration pool comprising multiple accelerator cards. Each accelerator card includes an Ethernet controller and at least one computing core, with the Ethernet controller and computing core communicating via Remote Direct Data Access (RDA). Each accelerator card in the at least one heterogeneous acceleration pool communicates with an Ethernet switching module via Ethernet or RDA. The computing resource pool includes multiple computing cores, each computing core having a binding relationship with at least one accelerator card in the at least one heterogeneous acceleration pool. Among the multiple computing cores is a management core, which manages the accelerator cards and their binding relationships in the at least one heterogeneous acceleration pool and distributes processing tasks to the accelerator cards in the at least one heterogeneous acceleration pool.

[0069] It should be noted that each accelerator card includes a circuit structure for implementing relevant data operations, which may specifically be a GPU, FPGA, etc., and in this embodiment, they are collectively referred to as computing cores. The computing resource pool consists of numerous processor cores, each processor core being a computing core; this includes a dedicated processor core for management, which in this embodiment is referred to as the management core.

[0070] In one embodiment, each accelerator card in the at least one heterogeneous accelerator pool includes a gold finger and a power connector; the gold finger and the power connector connect to and supply power to each device in the corresponding accelerator card. The gold finger and power connector in each accelerator card in the at least one heterogeneous accelerator pool are powered simultaneously during the power-on startup process of the corresponding accelerator card.

[0071] In one embodiment, a switching chip is connected between the Ethernet controller in each accelerator card and the corresponding computing core in the at least one heterogeneous acceleration pool; correspondingly, the Ethernet controller in each accelerator card and the corresponding computing core in the at least one heterogeneous acceleration pool communicate remotely via the switching chip.

[0072] In one embodiment, each accelerator card in the at least one heterogeneous acceleration pool includes: a homogeneous clock device, which is connected to the computing core, switching chip, and Ethernet controller in the corresponding accelerator card. Each accelerator card in the at least one heterogeneous acceleration pool includes: a clock generator, which is connected between the homogeneous clock device and the computing core, the homogeneous clock device and the switching chip, and the homogeneous clock device and the Ethernet controller in the corresponding accelerator card; correspondingly, the clock generator is used to select a clock source for the computing core, switching chip, and Ethernet controller in the corresponding accelerator card; the clock source is: the homogeneous clock device in the corresponding accelerator card or the clock device in the computing core bound to the corresponding accelerator card.

[0073] In one embodiment, each accelerator card in the at least one heterogeneous accelerator pool includes a control unit; the control unit is used to implement timing control of information transmission in the corresponding accelerator card. The control unit may specifically be a CPLD, MCU (Microcontroller Unit), etc.

[0074] In one embodiment, each acceleration card in the at least one heterogeneous acceleration pool includes: a link selector; the link selector connects to the computing core, switching chip, Ethernet controller, co-source clock device, power supply device and control unit in the corresponding acceleration card.

[0075] In one implementation, the Ethernet switching module includes: at least one switch; the at least one switch connects a computing resource pool and at least one heterogeneous acceleration pool.

[0076] In one embodiment, each accelerator card in at least one heterogeneous acceleration pool includes: a first node controller; the first node controller is used to collect device information and operating information in the corresponding accelerator card; correspondingly, each computing core in the computing resource pool includes: a second node controller; the second node controller is used to collect device information and operating information in the corresponding computing core; correspondingly, the Ethernet switching module includes: a central controller; the central controller is used to collect the device information and operating information in the accelerator card collected by the first node controller, and the device information and operating information in the computing core collected by the second node controller.

[0077] The first node controller is used to collect status information of the acceleration chips in the corresponding accelerator cards and sensor data from the boards. The second node controller is used to collect core operation information and sensor data from the corresponding computing cores. The central controller is used to construct a topology map including each computing core in the computing resource pool and each accelerator card in at least one heterogeneous acceleration pool based on the collected information.

[0078] In one implementation, a central controller is used to synchronize the collected information to a management core in a computing resource pool; accordingly, the management core manages the accelerator cards and binding relationships in at least one heterogeneous acceleration pool based on the received synchronization information, and distributes processing tasks to the accelerator cards in at least one heterogeneous acceleration pool.

[0079] In one implementation, the management core is used to formulate corresponding task allocation strategies and binding relationship adjustment strategies based on the received synchronization information.

[0080] In one implementation, the central controller is used to generate log data based on the collected information and analyze the log data for fault diagnosis.

[0081] In one embodiment, the computing system further includes a bus switching module connected to at least one heterogeneous acceleration pool; the bus switching module is also connected to another computing resource pool.

[0082] As can be seen, in this embodiment, heterogeneous accelerator cards with built-in Ethernet controllers form a heterogeneous acceleration pool. This pool connects to an Ethernet switching module via Ethernet or remote direct data access, and the Ethernet switching module connects to a computing resource pool. The management core in the computing resource pool manages the binding relationships between each accelerator card in at least one heterogeneous acceleration pool, the computing core, and the accelerator cards, and distributes processing tasks to each accelerator card in the pool, enabling them to collaboratively complete the processing tasks. This embodiment provides a computing power system for model training tasks. By accelerating data transmission via Ethernet or remote direct data access, it not only provides sufficient computing power to support task execution but also offers high bandwidth and low latency for high-speed communication and large-volume data transmission, which is beneficial for the rapid and stable operation of model training tasks.

[0083] This application enables heterogeneous accelerator cards to directly encapsulate GPU data into Ethernet frames, achieving direct communication between Ethernet and the GPU. The heterogeneous accelerator card supports dual-port 100G Ethernet and RDMA (Remote Direct Memory Access) functionality. Through software stacks and drivers, it supports the integration of GPU chips and Ethernet controllers, as well as remote direct access to GPU memory, thereby enabling data flow directly from the GPU to the Ethernet.

[0084] The core of a heterogeneous accelerator card design lies in its Ethernet controller, which is primarily responsible for data transmission and network interaction. It provides high-bandwidth, ultra-low-latency Ethernet services to meet the high-speed transmission requirements of the entire system. In one example, the Ethernet controller uses two SFP28 input / output interfaces to connect externally, transmitting external computing tasks to the accelerator card via network transmission. Internally, it uses a PCIe Gen4 X8 connection, transmitting externally transmitted data to the GPU chip based on a peer-to-peer communication mechanism for data processing. When the GPU chip interacts with the Ethernet controller, it utilizes the GPU memory remote direct access function to achieve point-to-point data transmission between different GPUs. The integrated network card (i.e., Ethernet controller) in the accelerator card achieves ultra-low latency and supports the entire PCIe Gen4 link.

[0085] Please see Figure 2 A block diagram of a GPU accelerator card is shown below. Figure 2 As shown. Figure 2The GPU accelerator card is a PCIe (Peripheral Component Interconnect Express) form factor card, including a PCIe gold finger connector and a power connector. It outputs two 100G network links and connects to a switch via optical modules and fiber optic cables, enabling data transfer with other GPUs. The GPU accelerator card's uplink PCIe Gen4 x16 link connects to the gold finger, and the switch outputs two PCIe Gen4 x16 links to the network controller chip and the GPU accelerator chip respectively. The GPU accelerator card is a PCIe card, with a total height of 111.15mm, a length of 267.7mm, and a width of 39.04mm. Key chips on the board include the GPU accelerator chip (i.e., the computing core), the switch chip, the network chip (i.e., the network interface card), and the CPLD (Complex Programmable Logic Device). The PCIe Gen4 x16 connector on the Goldfinger interface connects to the PCIe_SW. The PCIe_SW outputs two additional PCIe x16 connectors, which connect to the GPU chip and the NIC (Network Interface Card) chip, respectively. The GPU chip performs high-performance data processing and computation, the Switch expands the PCIe ports, and the NIC expands the network optical ports to 100G, handling network virtualization, offloading acceleration, and data stream forwarding. A CPLD chip is used for interrupt timing control and information exchange. This heterogeneous computing accelerator card integrates software and hardware heterogeneous computing environments, enabling communication and data transmission between heterogeneous processors via the high-speed PCIe bus and high-performance Ethernet, achieving cross-platform operation of computing tasks. It also achieves collaborative computing through higher-level system partitioning and task scheduling, resulting in higher computing performance and better efficiency. The accelerator card can be deployed in various configurations, including 4, 8, 16, or 32 cards per host, to meet diverse industry needs and can be widely used in smart finance, intelligent recommendation, fast search, content moderation, AI generation, and intelligent customer service.

[0086] like Figure 3As shown, the default power supply method for the GPU accelerator card is to power the entire board with a P12V input from the power connector, and to power the CPLD and FRU with a P3V3_STBY input from the gold fingers. Considering that the power supply within the entire system may be insufficient during power transients, a P12V_GF input from the PCIe gold fingers is reserved to power the SW and NIC, which can share approximately 50W of power consumption. The 12V power supply of the card is obtained by connecting to the host motherboard through the card cable, with a maximum power of 270W. The 3V3 power supply of the card is all converted from the P12V on the card to power the entire board. The 3V3_STBY power supply of the card is obtained by connecting to the HOST motherboard through the gold fingers on the card, and the power meets the PCIe 5.0 standard. At the same time, for power transient surge application scenarios (such as GPU startup, dynamic hot-swapping, etc.), optional soldering resistors are reserved on the card to connect the gold fingers and the Power Connector 12V power supply, with the gold fingers providing both 3V3_STBY and 12V (60W Max).

[0087] like Figure 4 As shown, the GPU card receives a 100MHz clock signal from the PCIe gold fingers, which is then connected to a ClockBuffer. The ClockBuffer expands to provide three 100MHz clock signals to the GPU, Switch chip, and NIC chip. A ClockGenerator non-homogeneous clock scheme is also reserved to avoid clock delays and jitter caused by excessively long lines. If the clock comes from the host, the clock line will be longer, but in this embodiment, the clock can be selected: either host or local.

[0088] It should be noted that the GPU and Ethernet controller NIC are directly connected via a PCIe switch chip (SW), which can handle the computational challenges of large-scale models. This direct connection technology between the accelerator card and the Ethernet controller via the switch chip differs from the traditional method where the NIC card connects to the switch chip, which in turn connects to the CPU, and then the CPU interconnects with the network card. Heterogeneous computing accelerator cards integrate GPU chips and Ethernet controllers, enabling data flow directly from the GPU to RDMA, such as... Figure 5 As shown, the GPU and Ethernet controller NIC are connected via a switch chip and communicate via remote direct memory access. Furthermore, the heterogeneous computing accelerator card supports dual-port 100G Ethernet and RDMA functionality.

[0089] In one example, the system RDMA communication software stack is as follows: Figure 6As shown. Typically, communication between cluster nodes is accomplished via the network. For network interface cards (NICs) supporting RDMA, an interface call is required to initiate a network transmission request. The driver in the software stack supports registering the relevant interfaces for the RDMA NIC. The RDMA NIC will then communicate with the GPU kernel driver through the kernel driver's standard interface to complete direct data transmission, enabling direct remote memory access. The entire data path involves only the GPU and the RDMA NIC, avoiding redundant jumps in data transfer to system memory and effectively reducing communication latency.

[0090] As can be seen in this example, GPUs can be directly connected; GPUs are directly connected to the Ethernet controller via a switching chip; host CPUs are interconnected via Ethernet controllers and switches. Specifically, the heterogeneous computing accelerator card achieves the integration of GPU chips and Ethernet chips under the same root port through the direct connection between the switching chip and the Ethernet controller; and effective aggregation between chips is achieved through High Bandwidth Low Latency Interconnect (RDMA) technology.

[0091] In one example, the architecture diagram of the computing system can be referenced. Figure 7 The general-purpose computing resource pool (corresponding to another computing resource pool) interconnects with heterogeneous computing acceleration resource pools such as GPUs and FPGAs via an internal bus and an internal bus switching module. Furthermore, heterogeneous acceleration cards directly access each other's memory remotely via Ethernet using RDMA functionality, reducing data transmission latency and improving computing performance. The entire system can be horizontally scaled to achieve inter-system interconnection, enabling larger-scale heterogeneous computing acceleration pooling. Specifically, high-performance internal bus switching and software-defined system design achieve device resource pooling, decoupling the physical link layer between devices and CPUs. This allows for flexible adjustment of port configurations and resource allocation paths in the switching network, fine-grained partitioning of shared resource pools, and on-demand elastic allocation and multi-host sharing of device resources. Addressing the current limitations of data center host system performance expansion due to system interconnect bandwidth constraints, performance mismatches between different levels of storage, and low I / O resource utilization, the pooled system design flexibly achieves dynamic resource allocation and load balancing for diverse scenario requirements. It also enables the fusion of multi-platform processor computing power and collaborative scheduling of heterogeneous acceleration resources, alleviating the bottleneck problem of data center performance expansion.

[0092] Figure 7The heterogeneous computing acceleration resource pool (i.e., the heterogeneous acceleration pool) shown maximizes computing power and efficiency. Based on the differences in resource categories such as general-purpose computing and heterogeneous computing, similar computing resources are integrated to form a resource pool, enabling on-demand resource reconfiguration between different devices. Fine-grained resource segmentation and pooling are achieved through hardware reconfiguration, more tightly integrating general-purpose and heterogeneous computing resources. Ultra-high-speed internal and external interconnection technology connects the various resource pools, achieving the fusion of heterogeneous computing power. Simultaneously, computing resources can be scheduled and load-balanced according to business scenarios. A business-aware resource reconfiguration decision system is established through software definition, completing intelligent hardware resource reconfiguration, achieving resource pooling and centralized management, thereby enabling dynamic adjustment, flexible combination, and intelligent allocation, improving the overall system's computing efficiency and response rate, and realizing intelligent and efficient heterogeneous computing.

[0093] In one example, refining the topology within the heterogeneous computing acceleration resource pool and the internal bus switching module yields the following: Figure 7 corresponding Figure 8 .like Figure 8 As shown, the internal bus switching module ( Figure 8 The high-performance switching chip in the middle contains numerous switching chips, including Ethernet switching modules. Figure 8 The network switching unit (NET) contains numerous network switches.

[0094] It should be noted that the internal bus switching module includes not only numerous switching chips, but also: CPLD, management controller, baseboard controller, network switching chip, and UART (Universal Asynchronous Receiver / Transmitter) chip; a heterogeneous computing acceleration resource pool may include not only numerous GPUs, but also baseboard controllers and CPLDs, etc. Please refer to [link to relevant documentation] for details. Figure 9 In other words, each resource pool can achieve flexible resource configuration through reconfiguration and decoupling. It interconnects with the processor via a local multi-channel monitoring and diagnostic network, focusing on hardware-level monitoring, management, and fault detection. It collects and manages the host system's operational status information in real time, and can also connect to numerous sensors to read environmental conditions and control temperature via fans. It also supports other system management functions, including remote power control, serial LAN, server host and memory monitoring, and error logging. Furthermore, when the system controller diagnoses a fault, it provides a comprehensive and intuitive display via panel indicator lights using optical path diagnostics, clearly indicating host system faults and issuing warnings.

[0095] Compared to traditional AI servers, heterogeneous computing acceleration resource pools support two P2P communication methods, such as... Figure 10As shown. Besides PCIe P2P communication, direct connection technology between the switching chip and the Ethernet controller is also possible. The RDMA communication library needs to call the interface to initiate a network transmission request. The driver in the software stack supports registering the relevant interfaces of the RDMA network card. The RDMA network card will communicate with the GPU kernel driver through the kernel driver standard interface to complete direct data transmission and realize direct remote access to video memory. Only the GPU and the RDMA network card participate in the entire data path, avoiding redundant jumps in data to system memory and effectively reducing communication latency. Pearl represents the internal bus switching module.

[0096] In one example, servers are interconnected via heterogeneous accelerator cards and switches, forming a hardware interconnection of multiple training servers to achieve cluster computing power. A feasible topology would be as follows: Figure 11 As shown, in Figure 11 In the distributed acceleration server, GPU nodes achieve cluster expansion through parameter plane switches in a 1:1 non-blocking manner; the service plane and storage plane cluster networks are connected to separate switches, which can support 1:1 non-blocking or 2:1 convergence; the cluster out-of-band management network is connected to the TOR out-of-band management switch.

[0097] In large-scale cluster systems, the management of individual accelerator card components is crucial. Therefore, this embodiment also implements overall out-of-band management for heterogeneous computing accelerator cards, with main functions including device monitoring, log management, fault diagnosis, and configuration management. Figure 12 As shown.

[0098] Hardware Monitoring: The server monitors the accelerator card hardware via an out-of-band path through the management chip (BMC) using the I2C bus protocol. This monitoring includes the status of the accelerator chips, board temperature, voltage, and other sensor data. Boundary scan chain data is used to collect the status of each chip within the accelerator card for fault location; control codes are also sent to each chip via the boundary scan chain to perform continuity testing, hardware configuration, module reset, and fault isolation. Temperature Monitoring: The management chip dynamically acquires the heat dissipation requirements and temperature information needed for the accelerator card's operating environment. Examples include the temperature of heterogeneous accelerator chips, their memory temperature, and inlet / outlet air vent temperatures. Voltage Monitoring: The power supply chip monitors and dynamically acquires relevant board-level voltage statuses.

[0099] Fault Diagnosis and Log Management: The system's BMC can obtain error information such as overvoltage, undervoltage, and overtemperature from the accelerator chip via the I2C bus protocol. It can also obtain the accelerator chip's serial number, firmware version, driver version, chip operating voltage, and board power consumption. When an accelerator card system malfunctions, the BMC records the logs generated by real-time monitoring of the machine status to the local file system. Field technicians can export the logs through the BMC web page or other tools to troubleshoot, analyze, locate, and resolve problems.

[0100] Configuration Management: Supports configuration of software such as network, users, and alarms, and supports importing and exporting configuration files.

[0101] In terms of accelerator card management, the computing unit and management unit are decoupled and standardized, separating common management, security, and control functions from the computing unit. This ensures compatibility with different computing and management platforms, supports unified management of multiple interfaces and chips, and meets the needs of different application scenarios. The management unit needs to implement computing power allocation and management, multi-chip module voltage regulation and power consumption management to ensure high-efficiency system design; monitor resource utilization, I / O throughput, and resource health status; achieve coordinated power-on / off, centralized management of resources and topology to ensure system availability; and enable on-demand rapid deployment and automatic management of hardware resources, monitoring of key resource information and fault management, intelligent fault location and recovery to ensure computing system reliability. The system management module is responsible for unified management, providing operation and maintenance capabilities through standardized service interfaces, and achieving integrated monitoring, fault early warning, and visualized management.

[0102] The system power-on / off coordination control process is as follows:

[0103] 1. After the PSU in each unit is powered, the VR switches out the STBY power and sends the PG of the last STBY power to the CPLD.

[0104] 2. After the STBY power-on in each chassis is complete and the CPLD and BMC are working normally, the host and the CPLD in the resource pool will send the STBY power-on PG signal to the BMC via I2C / UART.

[0105] 3. The host and resource pool BMC send the STBY PG completion signal to the core switching unit BMC via the network. At this time, each unit waits for the switching unit's power-on signal to proceed with the next operation.

[0106] 4. After pressing the power button of the core switching unit, the power-on signal is sent to the CPLD, and the CPLD controls the EN signal of the Main Power.

[0107] 5. At the same time, the Power on signal is sent to the BMC, and the BMC sends the power-on signal to the BMC of each chassis via the network.

[0108] 6. The host and resource pool BMC send the power-on signal to their respective CPLDs via I2C / UART. The CPLDs control the EN of each VR and, after receiving the last power PG, send the power-on completion signal to the BMC and mCPU of the core switching unit via the network and BMC.

[0109] The system power-on reset process is as follows:

[0110] 1. After STBY is powered on, the host's BMC will scan the ID of the CDFP interface of the local interconnect cable and the corresponding CDFP ID of the core switching unit to establish a mapping table.

[0111] 2. Similarly, the BMC of the core switching unit will scan the ID of the local CDFP interface and the CDFPID of the corresponding resource pool to establish a mapping table.

[0112] 3. After each unit is powered on, it waits for the PERST signal. The host sends PERST to each module via the CPU and CPLD. The host's BMC sends the PERST information to the core switching unit BMC through the TOR switch.

[0113] 4. Within the core switching unit, PERST events except for SW are triggered by mCPU's PLTRST. PERST events for SW are communicated to the CPLD via I2C by the BMC, and the CPLD then sends PERST events to the SW and CDFP. PERST events for the resource pool originating from the interconnect cable CDFP are sent to the CPLD, which then forwards them to each GPU device.

[0114] For details, please refer to Figure 13 The implemented accelerator card management functions are as follows:

[0115] Topology identification: Supports viewing the topology view, which allows you to view node summary information, such as node type, power-on status, and overall health status.

[0116] Asset Information Management: Supports viewing device information at each level. Detailed device information, such as device asset information and high-speed interface connection status, can be viewed through the Web / Redfish page.

[0117] Collaborative control: Supports centralized power-on / off control, with each unit powered on and off in sequence; supports system reset function, automatically detecting and resetting the system during power-on and host restart; supports reset control after resource reallocation, etc.

[0118] In system design, the discovery, management, and elastic adjustment of system resources are crucial. The control module, as the core management unit for dynamic resource adjustment, controls the high-performance switching unit to achieve automatic resource topology discovery and flexible automatic resource switching. This includes: Resource identification and display: providing network interfaces and a visual web interface to display resource lists and topologies, including heterogeneous computing resource unit information and I / O port information; Dynamic resource allocation: achieving second-level dynamic allocation and adjustment of heterogeneous computing units through key technologies such as hot removal, hot insertion, and hot reset of devices; Load balancing: dynamically allocating physical resources based on optimized load balancing scheduling algorithms, achieving fine-grained on-demand allocation of storage resources, and maximizing the release of heterogeneous computing power; Expert templates: providing expert template descriptions and dynamic switching interfaces for heterogeneous computing resources based on business needs, resource status, and performance indicators, allowing applications to request resources based on expert templates, while also providing network interfaces and a visual web interface for the visual application of expert templates.

[0119] To fully leverage the performance of heterogeneous multi-core processors and adapt to diverse application scenarios, the ability for applications to migrate freely between cores is crucial. However, migration between heterogeneous multi-core processors with different instruction sets has long been a challenge in the industry. Load migration and resource scheduling are applied in the following scenarios: operating system load balancing, where processes are migrated to static or lightly loaded cores; migration to cores with different instruction set architectures based on different load types; power consumption changes, requiring the migration of some processes to achieve power control; and migrating processes to other cores when a core overheats.

[0120] The operating system of a general-purpose computing node (i.e., the host) contains multiple kernels, each compiled and running under a specific ISA instruction set. The entire system uses unified system software to achieve on-demand balanced scheduling of computing power for applications. At the operating system level, inter-kernel communication and collaboration are implemented, ensuring the system maintains a global state even when running instances of different ISA cores. Building upon this, a unified scheduling core employs advanced scheduling techniques to achieve application-level hardware detachment. Through collaboration between the operating system and the scheduling system, applications can automatically select the optimal mapping, thereby improving the performance of heterogeneous multi-core systems with multiple instruction sets. Figure 14 As shown.

[0121] For general-purpose computing nodes, this embodiment provides a heterogeneous virtual computing power management subsystem. At the technical level, this system combines heterogeneous resource computing power virtualization and remote scheduling technologies; at the application level, it addresses the usage and management needs of users and administrators for heterogeneous virtualized resources. For details, please refer to [link to relevant documentation]. Figure 15 This system ensures that different users have the following basic capabilities when using virtualized resources in different scenarios:

[0122] Transparency: Without modifying system component code or user code, the system allows users to invoke local or remote heterogeneous virtual computing power to execute applications as if they were running on physical resources.

[0123] Low overhead: The system enables applications running on it to perform as close as possible to the performance of physical resources when using heterogeneous virtual computing power.

[0124] Isolation: The system manages the allocation and release of heterogeneous virtual computing power for each container, enabling containers to be completely isolated from each other while sharing physical resources.

[0125] To achieve the above objectives, the heterogeneous virtual computing power management subsystem mainly comprises three components. The heterogeneous virtual computing power management platform is responsible for the front-end interaction logic for end users and administrators. The heterogeneous virtual computing power management center component is responsible for resource scheduling and management of the entire system. The heterogeneous virtual computing power management node component is responsible for the specific operation execution on each computing node.

[0126] Heterogeneous Virtual Computing Power Management Platform: A heterogeneous resource scheduling and management platform for users and administrators, responsible for receiving user job requests and creating virtualized computing environments based on specific business needs. It also receives and aggregates node device information reported by node components, and then distributes the aggregated device information to the central component. Specifically, the platform includes the following functions:

[0127] User Management: The platform enables basic management of users and roles by storing information such as users, related groups, and permissions.

[0128] Resource Management: The platform manages all heterogeneous computing power physical and virtual devices by statistically analyzing them and providing different views, including node resource overview, node status information, and operations such as disabling, enabling, and removing virtual devices.

[0129] Quota management: The platform manages and allocates these computing resources by configuring specific quotas for virtual heterogeneous computing resources. This includes virtual total computing power limits, virtual total video memory limits, and computing power and video memory limits for individual virtual devices.

[0130] Policy Management: The platform configures and manages various scheduling policies based on the administrator's specific needs by abstracting different policies of the scheduler in the central component. Specific scheduling policies can be found in the description of the central component.

[0131] Please see Figure 16This diagram illustrates a heterogeneous computing resource pool management architecture. Centered on the Pooled System Management Controller (PSMC) within the high-performance switching unit, heterogeneous computing acceleration units and general-purpose computing units are each paired with a Pooled Node Management Controller (PNMC). The PSMC acts as the central management node, while the PNMC serves as a distributed node, managed uniformly by the PSMC to achieve full lifecycle management of the heterogeneous computing resource pool. In this system, the discovery, management, and elastic adjustment of heterogeneous computing resources are crucial. This project designs a pooled management engine as the core management unit for dynamic resource adjustment. The pooled management engine controls the high-performance switching unit to achieve automatic resource topology discovery and flexible automatic resource switching, and provides a standard interface for the data center monitoring and management platform to achieve centralized management of massive data center-level resources.

[0132] The Power Supply Management System (PSMC) plays a crucial role in the management of distributed heterogeneous accelerated pooled servers. It serves as a bridge for information communication and is responsible for unified management. Its main functions include node topology identification, centralized display of asset information, and coordinated power-on / off. Through standardized service interfaces, it provides operational capabilities, enabling integrated monitoring, fault warnings, and visualized management. The heterogeneous computing resource pool system interconnects the management controllers of pooled nodes at all levels via a network, building an independent engine management platform.

[0133] The following describes a calculation method provided by an embodiment of this application. The calculation method described below can be referred to in conjunction with other embodiments described herein.

[0134] This application discloses a computing method applied to a management core, comprising: receiving a processing task; sending the processing task to an Ethernet switching module in a computing system, so that the Ethernet switching module distributes the processing task to at least one heterogeneous acceleration pool in the computing system via Ethernet or remote direct data access.

[0135] The computing system includes: a computing resource pool, at least one heterogeneous acceleration pool, and an Ethernet switching module connecting the computing resource pool and the at least one heterogeneous acceleration pool; the at least one heterogeneous acceleration pool includes: multiple acceleration cards; each acceleration card includes: an Ethernet controller and at least one computing core, the Ethernet controller and the computing core communicate via remote direct data access; each acceleration card in the at least one heterogeneous acceleration pool communicates with the Ethernet switching module via Ethernet or remote direct data access; the computing resource pool includes: multiple computing cores, the same computing core being bound to at least one acceleration card in the at least one heterogeneous acceleration pool; the management core is any one of the multiple computing cores.

[0136] As can be seen, this embodiment provides a computing method that accelerates data transmission via Ethernet or remote direct data access. This not only provides sufficient computing power to support the operation of the task, but also provides high bandwidth and low latency for high-speed communication and large data transmission, which is beneficial for the rapid and stable operation of the model training task.

[0137] The following describes an electronic device provided by an embodiment of this application. The electronic device described below can be referred to in conjunction with other embodiments described herein.

[0138] This application discloses an electronic device, including:

[0139] Memory, used to store computer programs;

[0140] A processor is configured to execute the computer program to implement the methods disclosed in any of the above embodiments.

[0141] Furthermore, embodiments of this application also provide an electronic device. The aforementioned electronic device can be, for example,... Figure 17 The server shown can also be as follows: Figure 18 The terminal shown. Figure 17 and Figure 18 These are all diagrams illustrating the structure of an electronic device according to an exemplary embodiment. The content in the diagrams should not be considered as any limitation on the scope of this application.

[0142] Figure 17 This is a schematic diagram of a server structure provided in an embodiment of this application. The server may specifically include: at least one processor, at least one memory, a power supply, a communication interface, an input / output interface, and a communication bus. The memory stores a computer program, which is loaded and executed by the processor to implement the relevant steps in the calculations disclosed in any of the foregoing embodiments.

[0143] In this embodiment, the power supply is used to provide operating voltage for each hardware device on the server; the communication interface can create a data transmission channel between the server and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0144] In addition, the memory, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored on it include operating system, computer programs and data, etc., and the storage method can be temporary storage or permanent storage.

[0145] The operating system manages and controls the various hardware devices and computer programs on the server to enable the processor to perform operations and processes on the data in the memory. It can be Windows Server, Netware, Unix, Linux, etc. In addition to computer programs capable of performing the calculation methods disclosed in any of the foregoing embodiments, the computer programs may further include computer programs capable of performing other specific tasks. The data may include application update information and application developer information.

[0146] Figure 18 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application. The terminal may include, but is not limited to, a smartphone, tablet computer, laptop computer, or desktop computer.

[0147] Typically, the terminal in this embodiment includes a processor and a memory.

[0148] The processor may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor can be implemented using at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor may also include a main processor and coprocessors. The main processor, also known as the CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor may also include an AI (Artificial Intelligence) processor, which handles computational operations related to machine learning.

[0149] The memory may include one or more computer non-volatile storage media, which may be non-transitory. The memory may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In this embodiment, the memory is used to store at least the following computer programs, which, after being loaded and executed by the processor, are capable of implementing the relevant steps in the computational methods executed on the terminal side as disclosed in any of the foregoing embodiments. Furthermore, the resources stored in the memory may also include operating systems and data, and the storage method may be temporary or permanent storage. The operating system may include Windows, Unix, Linux, etc. The data may include, but is not limited to, application update information.

[0150] In some embodiments, the terminal may further include a display screen, an input / output interface, a communication interface, a sensor, a power supply, and a communication bus.

[0151] Those skilled in the art will understand that Figure 18 The structure shown does not constitute a limitation on the terminal and may include more or fewer components than illustrated.

[0152] The following describes a non-volatile storage medium provided in an embodiment of this application. The non-volatile storage medium described below can be referred to in conjunction with other embodiments described herein.

[0153] A non-volatile storage medium is provided for storing a computer program, wherein the computer program, when executed by a processor, implements the computation method disclosed in the foregoing embodiments. The non-volatile storage medium is a computer-readable non-volatile storage medium, which, as a carrier for resource storage, can be a read-only memory, random access memory, disk, or optical disk, etc. The resources stored thereon include an operating system, computer programs, and data, and the storage method can be temporary storage or permanent storage.

[0154] The following describes a computer program product provided by an embodiment of this application. The computer program product described below can be referred to in conjunction with other embodiments described herein.

[0155] A computer program product includes a computer program / instructions that, when executed by a processor, implement the steps of the aforementioned disclosed computation method.

[0156] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0157] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of non-volatile storage medium known in the art.

[0158] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A computing system, characterized in that, include: A computing resource pool, at least one heterogeneous acceleration pool, and an Ethernet switching module connected between the computing resource pool and the at least one heterogeneous acceleration pool; The at least one heterogeneous acceleration pool includes: multiple acceleration cards; each acceleration card includes: an Ethernet controller and at least one computing core, the Ethernet controller and the computing core communicate via remote direct data access; a switching chip is connected between the Ethernet controller in each acceleration card and the computing core in the corresponding acceleration card in the at least one heterogeneous acceleration pool. Each accelerator card in the at least one heterogeneous acceleration pool communicates with the Ethernet switching module via Ethernet or remote direct data access. The computing resource pool includes: multiple computing cores, and the same computing core is bound to at least one accelerator card in the at least one heterogeneous acceleration pool; The plurality of computing cores includes a management core, which is used to manage each accelerator card in the at least one heterogeneous acceleration pool and the binding relationship, and to distribute processing tasks to each accelerator card in the at least one heterogeneous acceleration pool; Each accelerator card in the at least one heterogeneous acceleration pool includes: a homogeneous clock device, which is connected to the computing core, switching chip and Ethernet controller in the corresponding accelerator card; the Ethernet controller in each accelerator card in the at least one heterogeneous acceleration pool communicates with the computing core in the corresponding accelerator card through the switching chip to realize remote direct data access, thereby realizing point-to-point transmission of video memory data between different accelerator cards. Wherein, each acceleration card in the at least one heterogeneous acceleration pool includes: a clock generator, wherein the clock generator is connected between the same source clock device and the computing core, the same source clock device and the switching chip, and the same source clock device and the Ethernet controller in the corresponding acceleration card; Accordingly, the clock generator is used to select a clock source for the computing core, switching chip and Ethernet controller in the corresponding accelerator card; the clock source is: a clock device of the same source in the corresponding accelerator card or a clock device in the computing core bound to the corresponding accelerator card.

2. The computing system according to claim 1, characterized in that, Each accelerator card in the at least one heterogeneous accelerator pool includes: a gold finger and a power connector; the gold finger and the power connector connect to each device in the corresponding accelerator card and supply power to each device in the corresponding accelerator card.

3. The computing system according to claim 2, characterized in that, The gold fingers and power connectors of each accelerator card in the at least one heterogeneous accelerator pool are simultaneously powered during the power-on process of the corresponding accelerator card.

4. The computing system according to claim 1, characterized in that, Each accelerator card in the at least one heterogeneous accelerator pool includes a control unit; the control unit is used to implement timing control of information transmission in the corresponding accelerator card.

5. The computing system according to claim 1, characterized in that, Each acceleration card in the at least one heterogeneous acceleration pool includes: a link selector; the link selector connects to the computing core, switching chip, Ethernet controller, co-source clock device, power supply device and control unit in the corresponding acceleration card.

6. The computing system according to claim 1, characterized in that, The Ethernet switching module includes: at least one switch; the at least one switch connects the computing resource pool and the at least one heterogeneous acceleration pool.

7. The computing system according to claim 1, characterized in that, Each accelerator card in the at least one heterogeneous acceleration pool includes: a first node controller; the first node controller is used to collect device information and operating information from the corresponding accelerator card; Accordingly, each computing core in the computing resource pool includes: a second node controller; the second node controller is used to collect device information and operating information in the corresponding computing core; Accordingly, the Ethernet switching module includes: a central controller; the central controller is used to collect device information and operating information from the accelerator card collected by the first node controller, and device information and operating information from the computing core collected by the second node controller.

8. The computing system according to claim 7, characterized in that, The first node controller is used to collect the status information of the computing cores in the corresponding accelerator card and the sensor data of the board.

9. The computing system according to claim 7, characterized in that, The second node controller is used to collect core operation information and sensor data from the corresponding computing core.

10. The computing system according to claim 7, characterized in that, The central controller is used to construct a topology map including each computing core in the computing resource pool and each accelerator card in the at least one heterogeneous acceleration pool based on the collected information.

11. The computing system according to claim 7, characterized in that, The central controller is used to synchronize the collected information to the management core in the computing resource pool; Accordingly, the management core manages each accelerator card in the at least one heterogeneous acceleration pool and the binding relationship based on the received synchronization information, and distributes processing tasks to each accelerator card in the at least one heterogeneous acceleration pool.

12. The computing system according to claim 7, characterized in that, The management core is used to formulate corresponding task allocation strategies and binding relationship adjustment strategies based on the received synchronization information.

13. The computing system according to claim 7, characterized in that, The central controller is used to generate log data based on the collected information and analyze the log data for fault diagnosis.

14. The computing system according to any one of claims 1 to 13, characterized in that, The computing system further includes a bus switching module connected to the at least one heterogeneous acceleration pool; the bus switching module is also connected to another computing resource pool.

15. A calculation method, characterized in that, Applied to the management core, including: Receive and process tasks; The processing task is sent to the Ethernet switching module in the computing system, so that the Ethernet switching module distributes the processing task to at least one heterogeneous acceleration pool in the computing system via Ethernet or remote direct data access. The computing system includes: a computing resource pool, the at least one heterogeneous acceleration pool, and the Ethernet switching module connected between the computing resource pool and the at least one heterogeneous acceleration pool; The at least one heterogeneous acceleration pool includes: multiple acceleration cards; each acceleration card includes: an Ethernet controller and at least one computing core, the Ethernet controller and the computing core communicate via remote direct data access; a switching chip is connected between the Ethernet controller in each acceleration card and the computing core in the corresponding acceleration card in the at least one heterogeneous acceleration pool. Each accelerator card in the at least one heterogeneous acceleration pool communicates with the Ethernet switching module via Ethernet or remote direct data access. The computing resource pool includes: multiple computing cores, each computing core being bound to at least one accelerator card in the at least one heterogeneous acceleration pool; the management core is any one of the multiple computing cores. Each accelerator card in the at least one heterogeneous acceleration pool includes: a homogeneous clock device, which is connected to the computing core, switching chip and Ethernet controller in the corresponding accelerator card; the Ethernet controller in each accelerator card in the at least one heterogeneous acceleration pool communicates with the computing core in the corresponding accelerator card through the switching chip to realize remote direct data access, thereby realizing point-to-point transmission of video memory data between different accelerator cards. Wherein, each acceleration card in the at least one heterogeneous acceleration pool includes: a clock generator, wherein the clock generator is connected between the same source clock device and the computing core, the same source clock device and the switching chip, and the same source clock device and the Ethernet controller in the corresponding acceleration card; Accordingly, the clock generator is used to select a clock source for the computing core, switching chip and Ethernet controller in the corresponding accelerator card; the clock source is: a clock device of the same source in the corresponding accelerator card or a clock device in the computing core bound to the corresponding accelerator card.

16. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the method as described in claim 15.

17. A non-volatile storage medium, characterized in that, Used to store a computer program, wherein the computer program, when executed by a processor, implements the method as described in claim 15.

18. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the method as described in claim 15.

Citation Information

Patent Citations

  • Memory management and use method and device, equipment and medium

    CN112286688A

  • Hardware calculation module, device and method, electronic device and storage medium

    CN116627888A

  • Multi-accelerator card heterogeneous server and resource link reconstruction method

    CN117687956A