Computing system, control method and apparatus therefor, device, non-volatile storage medium, and switch

By introducing a high-speed serial computer expansion bus and switch into the computing system, and obtaining accelerator card address information to configure the address mapping table, the problem of low cross-server communication efficiency was solved, realizing a high-bandwidth multi-machine, multi-card computing cluster, and improving communication efficiency and latency performance.

WO2026157350A1PCT designated stage Publication Date: 2026-07-30LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
LANGCHAO ELECTRONIC INFORMATION IND CO LTD
Filing Date
2025-10-13
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

The low efficiency of cross-server communication in computing systems prevents acceleration even with increased computing resources, making communication bottlenecks a limiting factor.

Method used

By introducing a first switch into the computing system, connecting it to the server's accelerator card via a high-speed serial computer expansion bus, obtaining local device address information and configuring an address mapping table, data forwarding between accelerator cards is achieved. Cross-server data forwarding is performed using a non-transparent bridge or PCIe architecture.

Benefits of technology

It improves cross-server communication bandwidth, reduces latency, and enables high-bandwidth multi-machine, multi-GPU computing clusters, solving the problems of limited computing power in a single server and insufficient bandwidth in multiple servers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025127373_30072026_PF_FP_ABST
    Figure CN2025127373_30072026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computers, and discloses a computing system, a control method and apparatus therefor, a device, a non-volatile storage medium, and a switch. First switches use high-speed serial computer expansion buses to implement interconnection between accelerator cards located on multiple servers, acquire local device addresses of the accelerator cards on the servers, configure address mapping tables of the servers on the basis of the local device addresses of the accelerator cards, and perform data forwarding between the accelerator cards on different servers on the basis of the address mapping tables. Thus, the present invention effectively solves the problems of limited computing power of a single server and insufficient bandwidth and high latency in Ethernet-based interconnection between multiple servers, implements cross-server vertical scaling based on a high-speed serial expansion bus, increases the bandwidth of cross-server communication, and reduces the latency of cross-server communication, thereby laying a foundation for a computing cluster that enables high-bandwidth domain communication across multiple hosts and multiple accelerator cards.
Need to check novelty before this filing date? Find Prior Art

Description

Computing systems and their control methods, devices, equipment, non-volatile storage media and switches

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese Patent Application No. 202510124950.3, filed on January 26, 2025, entitled “Computing System and Control Method, Apparatus, Device, Medium and Switch Thereof”, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of computer technology, and in particular to computing systems and their control methods, apparatus, devices, non-volatile storage media and switches. Background Technology

[0004] With the development of artificial intelligence technology, the demand for computing power has increased dramatically. This necessitates not only deploying accelerator cards on individual servers to improve single-machine computing power, but also interconnecting a large number of servers to form multi-machine, multi-card clusters. However, when computing systems expand to a certain scale, communication bottlenecks prevent further acceleration even with increased computing resources.

[0005] Therefore, there is a technical problem in the related technologies where communication efficiency across servers in computing systems is low. Summary of the Invention

[0006] The purpose of this application is to provide a computing system and its control method, apparatus, device, non-volatile storage medium and switch for improving the communication efficiency across servers in a computing system.

[0007] To address the aforementioned technical problems, according to a first aspect, this application provides a computing system comprising: multiple servers and a first switch;

[0008] The first switch is connected to the server's acceleration card via a high-speed serial computer expansion bus.

[0009] The first switch is configured to obtain the local device address information of the accelerator cards on the server, configure the address mapping table of each server according to the local device address information, and perform data forwarding between accelerator cards on different servers according to the address mapping table.

[0010] In some embodiments, the first switch forwards data between accelerator cards on different servers according to an address mapping table, including:

[0011] The first switch receives data packets sent by the accelerator card through the first non-transparent bridge port, performs address translation and bus identifier translation on the data packets according to the address mapping table, and then forwards the data packets to the destination accelerator card through the second non-transparent bridge port.

[0012] In some embodiments, the address mapping table includes the address mapping relationship between the virtual device address allocated by the server to the remote device and the corresponding local device address information;

[0013] The first switch forwards data between acceleration cards on different servers based on the address mapping table, including:

[0014] After receiving the data packet sent by the accelerator card, the first switch forwards the data packet to the destination accelerator card according to the address mapping relationship.

[0015] In some embodiments, the southbound interface of the accelerator card is connected to the first switch via a high-speed serial computer expansion bus.

[0016] In some embodiments, the number of first switches is multiple, and the computing system further includes a first controller;

[0017] The first controller is configured to generate an address mapping table using the local device address information collected by the first switch, and send the address mapping table to the first switch to configure the address mapping table on the first switch.

[0018] In some embodiments, different first switches are interconnected via a high-speed serial computer expansion bus.

[0019] In some embodiments, the system further includes a first out-of-band management controller and a second controller located on the board where the first controller is located, and a third controller located on the board where the first switch is located.

[0020] The first out-of-band management controller is configured to implement out-of-band management of the first switch; the second controller is configured to control the power-on and power-off sequence of the first switch; and the third controller is configured to configure the high-speed serial computer expansion bus interface of the first switch.

[0021] In some embodiments, the first controller is further configured to receive configuration requirement parameters for the target resource pool, obtain status parameters of the accelerator cards through the first switch, and perform resource scheduling of the accelerator cards according to the configuration requirement parameters and status parameters to construct the target resource pool.

[0022] In some embodiments, the first switch is further configured to receive configuration requirement parameters for the target resource pool, obtain status parameters of the accelerator cards, and perform resource scheduling of the accelerator cards according to the configuration requirement parameters and status parameters to construct the target resource pool.

[0023] In some embodiments, different accelerator cards located on the same server are interconnected via a high-speed serial computer expansion bus.

[0024] In some embodiments, a second switch disposed on the server is also included;

[0025] Different accelerator cards located on the same server are interconnected via a high-speed serial computer expansion bus and a second switch.

[0026] In some embodiments, a second switch disposed on the server is also included;

[0027] The first switch is connected to the server's acceleration card via a high-speed serial computer expansion bus, and includes:

[0028] The first switch is connected to the accelerator card via a high-speed serial computer expansion bus and the second switch.

[0029] In some embodiments, the first switch is connected to the server's acceleration card via a high-speed serial computer expansion bus, including:

[0030] The first switch and the accelerator card are connected via a copper cable based on the high-speed serial computer extended bus protocol.

[0031] In some embodiments, the first switch is connected to the server's acceleration card via a high-speed serial computer expansion bus, including:

[0032] The first switch and the accelerator card are connected via a first fiber optic link based on a high-speed serial computer extended bus protocol.

[0033] In some embodiments, the first optical fiber link includes a first optical interface disposed on a server, a second optical interface disposed on a first switch, and a first optical fiber;

[0034] Both the first optical interface and the second optical interface include: a first relay module, a second relay module, and a linear direct-drive optoelectronic conversion device;

[0035] The first relay module is configured to relay data signals from a high-speed serial computer extended bus, and the first channel of the linear direct-drive optoelectronic conversion device is located between the first relay module and the first optical fiber.

[0036] The second relay module is configured to relay auxiliary signals of the high-speed serial computer expansion bus; the second channel of the linear direct-drive optoelectronic conversion device is located between the second relay module and the first optical fiber.

[0037] In some embodiments, the first switch includes a third switch that is directly connected to the server and a fourth switch configured to enable interconnection between different third switches.

[0038] To solve the above-mentioned technical problems, according to the second aspect, this application also provides a switch, which is connected to an accelerator card located on different servers via a high-speed serial computer expansion bus;

[0039] The switch is configured to obtain the local device address information of the accelerator cards on the server, configure the address mapping table of each server according to the local device address information, and perform data forwarding between accelerator cards on different servers according to the address mapping table.

[0040] To address the aforementioned technical problems, according to a third aspect, this application also provides a control method for a computing system, comprising:

[0041] Obtain the local device address information of the server's accelerator card through the high-speed serial computer expansion bus;

[0042] Configure the address mapping table of each server according to the address information of each local device;

[0043] Data forwarding between accelerator cards on different servers is performed based on the address mapping table.

[0044] To address the aforementioned technical problems, according to a fourth aspect, this application also provides a control device for a computing system, comprising:

[0045] The acquisition unit is configured to acquire the local device address information of the server's accelerator card via a high-speed serial computer expansion bus;

[0046] The configuration unit is set to configure the address mapping table of each server based on the address information of each local device;

[0047] The control unit is configured to forward data between accelerator cards on different servers based on an address mapping table.

[0048] To address the aforementioned technical problems, according to a fifth aspect, this application also provides a control device for a computing system, comprising:

[0049] The memory is configured to store computer programs;

[0050] The processor is configured to execute computer programs, and when the computer programs are executed by the processor, they implement the steps of the control method of the computing system described above.

[0051] To solve the above-mentioned technical problems, according to the sixth aspect, this application also provides a non-volatile storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the control method of the computing system described above.

[0052] To address the aforementioned technical problems, according to the seventh aspect, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the control method for the aforementioned computing system.

[0053] The computing system provided in this application has the advantage of interconnecting accelerator cards located on multiple servers through a high-speed serial computer expansion bus using a first switch. The first switch obtains the local device address of the accelerator card on the server, configures the address mapping table of each server according to the local device address of each accelerator card, and forwards data between accelerator cards on different servers according to the address mapping table. This effectively solves the problems of limited computing power of a single server and insufficient bandwidth and large latency in interconnecting multiple servers via Ethernet. It realizes cross-server vertical expansion based on a high-speed serial computer expansion bus, improves the bandwidth of cross-server communication, reduces the latency of cross-server communication, and lays the foundation for realizing high-bandwidth domain communication computing clusters with multiple machines and multiple cards.

[0054] The control method, apparatus, equipment, non-volatile storage medium, and switch for the computing system provided in this application have the aforementioned beneficial effects, which will not be elaborated further here. Attached Figure Description

[0055] To more clearly illustrate the technical solutions of the embodiments or related technologies of this application, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0056] Figure 1 is a schematic diagram of the structure of a computing system provided in an embodiment of this application;

[0057] Figure 2 is a schematic diagram of a cross-domain interconnection architecture for accelerator cards provided in an embodiment of this application;

[0058] Figure 3 is a schematic diagram of another cross-domain interconnection architecture for accelerator cards provided in an embodiment of this application;

[0059] Figure 4 is a schematic diagram of the structure of an accelerator card provided in an embodiment of this application;

[0060] Figure 5 is a schematic diagram of a PCIe southbound interconnection cluster topology provided in an embodiment of this application;

[0061] Figure 6 is a schematic diagram of the structure of a first switch provided in an embodiment of this application;

[0062] Figure 7 is a schematic diagram of the connection between a first controller and a first switch provided in an embodiment of this application;

[0063] Figure 8 is a schematic diagram of the control structure of a control base plate provided in an embodiment of this application;

[0064] Figure 9 is a schematic diagram of a centralized collaborative management structure for cluster interconnection based on PCIe bus provided in an embodiment of this application;

[0065] Figure 10 is a schematic diagram of a PCIe optical interconnect link provided in an embodiment of this application;

[0066] Figure 11 is a schematic diagram of another PCIe optical interconnect link provided in an embodiment of this application. Detailed Implementation

[0067] The core of this application is to provide a computing system and its control method, apparatus, equipment, non-volatile storage medium and switch, for improving the communication efficiency across servers in the computing system.

[0068] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0069] To facilitate understanding of the technical solutions provided in the embodiments of this application, some key terms used in the embodiments of this application will be explained here first.

[0070] With the rapid development of artificial intelligence (AI) technology, the demand for computing power is growing exponentially. Especially in typical AI applications such as generative AI, large models, and deep learning, heterogeneous acceleration server systems also face many challenges such as improving computing efficiency and system performance. Continuous research and innovation are needed to support the sustainable development of high-performance computing infrastructure.

[0071] Artificial intelligence model computation mainly includes model training and model inference.

[0072] Traditional AI server architectures are insufficient to meet the ever-increasing computing power demands of AI, necessitating system-level innovation by integrating more heterogeneous accelerator cards within the server to enhance system computing power. Practical experience has shown that the computational speed of AI models is affected by factors such as single-card speed, the number of accelerator cards, and the multi-card speedup ratio.

[0073] The single-card speed is affected by the computing speed of the accelerator card and the speed of data input / output (IO). The model computing performance of a single card can be accelerated through precision training, operator fusion, and gradient accumulation.

[0074] Theoretically, the more accelerator cards there are, the faster the model computation speed. However, as the scale of computation increases, such as with the further growth of the dataset, the parallel computation of the model will have limitations. When the computing system expands to a certain scale, due to the existence of communication bottlenecks, the marginal effect of increasing computing resources becomes obvious, and even increasing computing resources may not achieve acceleration. At this time, it is necessary to optimize the communication topology to optimize the overall performance of the computing system.

[0075] The speedup of multi-GPU systems is determined by computation and communication efficiency, and needs to be optimized in conjunction with the algorithm and the network topology in the cluster. A multi-dimensional hybrid parallel strategy that combines data parallelism (DP), model parallelism (MP), and pipeline parallelism (PP) can be adopted to increase the efficiency of multi-GPU model computation.

[0076] Larger models bring new challenges: greater computational demands, higher communication bandwidth, and increased latency requirements. This necessitates optimizing single-card efficiency, increasing single-machine computing density, and achieving high-bandwidth clustering across multiple machines and multiple cards. Clustering has become an inevitable choice. Computation has now reached trillions of parameters, requiring the establishment of AI clusters with tens of thousands of cards and the creation of hyper-interconnects between nodes to ensure computational and network requirements. When scaling up clusters, model parallelism places high demands on bus bandwidth and latency. Building a High Bandwidth Domain (HBD) significantly benefits model parallelism and can reduce pipelined parallelism buffer time, thereby improving overall operational efficiency.

[0077] In training systems, larger models, longer sequences, and mixtures of experts (MOEs) are the main directions for model training, requiring larger tensor parallelism (TP) / expert parallelism (EP) domains and bandwidth. This necessitates high-bandwidth domains in supernodes to meet training demands. In inference systems, to ensure inference performance and user experience, the trend is towards multi-machine, multi-GPU setups. All-reduce inference across the network becomes a bottleneck, requiring high-bandwidth domains in supernodes to meet inference requirements. In applications such as recommendation systems (search, broadcast, and push systems), the number of recommended products and users surges, with embedding tables reaching tens of terabytes in size. Inference systems spend 20%–50% of their time on network communication, further necessitating high-bandwidth domains in supernodes to meet inference requirements.

[0078] A High Bandwidth Domain (HBD) refers to a group of computing systems interconnected by High Bandwidth (HB). Within a High Bandwidth Domain, the communication bandwidth between computing nodes is several times that between computing nodes in the High Bandwidth Domain. Currently, High Bandwidth Domains are typically limited to a single server. Traditional methods of improving single-chip performance and expanding multi-machine clusters through horizontal scaling (scale out) have reached bottlenecks. Therefore, expanding the number of supernodes in a High Bandwidth Domain through vertical scaling (scale up) has become a direction for overcoming computing power limitations.

[0079] To address this, the computing system provided in this application embodiment uses a high-speed serial computer expansion bus via a first switch to interconnect accelerator cards located on multiple servers. The first switch obtains the local device address of the accelerator card on the server, configures the address mapping table of each server according to the local device address of each accelerator card, and forwards data between accelerator cards on different servers according to the address mapping table. This effectively solves the problems of limited computing power of a single server and insufficient bandwidth and large latency in interconnecting multiple servers via Ethernet. It realizes cross-server vertical expansion based on a high-speed serial computer expansion bus, improves the bandwidth of cross-server communication, reduces the latency of cross-server communication, and lays the foundation for realizing high-bandwidth domain communication computing clusters with multiple machines and multiple cards.

[0080] Figure 1 is a schematic diagram of the structure of a computing system provided in an embodiment of this application; Figure 2 is a schematic diagram of a cross-domain interconnection architecture for accelerator cards provided in an embodiment of this application; Figure 3 is a schematic diagram of another cross-domain interconnection architecture for accelerator cards provided in an embodiment of this application.

[0081] As shown in Figure 1, the computing system provided in this embodiment may include: multiple servers (e.g., host 0, host 1, host 2, and host 3 in Figure 1) and a first switch (e.g., first switch SW0, first switch SW1, first switch SW2, and first switch SW3 in Figure 1); the first switch is connected to the accelerator cards of the servers via a high-speed serial computer expansion bus. The first switch is configured to obtain the local device address information of the accelerator cards on the servers, configure the address mapping table of each server according to the local device address information, and perform data forwarding between accelerator cards on different servers according to the address mapping table.

[0082] It's important to note that an accelerator card is a hardware device designed to improve server performance and accelerate data transfer rates. It is typically installed in a high-speed Peripheral Component Interconnect Express (PCIe) slot on the server and connected to the server motherboard. The accelerator card includes the circuitry for performing the necessary data processing, and can take the form of a Graphics Processing Unit (GPU), a Field-Programmable Gate Array (FPGA), or other types depending on the specific requirements.

[0083] In related technologies, server interconnection is achieved via Ethernet, but the bandwidth and latency of Ethernet communication are increasingly unable to meet the performance requirements of large-scale computing. To achieve cross-domain interconnection of accelerator cards in the high-bandwidth domain (i.e., interconnection between accelerator cards located on different servers), in this embodiment, a first switch is connected to the accelerator cards via a high-speed serial computer expansion bus, thereby achieving cross-domain interconnection of accelerator cards based on the high-speed serial computer expansion bus. Compared with traditional network interconnection methods, this significantly improves bandwidth and reduces latency, thus realizing cross-domain interconnection of accelerator cards in the high-bandwidth domain.

[0084] In some embodiments, a non-transparent bridge (NTB) can be used to achieve cross-domain communication of accelerator cards. The first switch forwards data between accelerator cards on different servers according to the address mapping table, which may include: the first switch receiving data packets sent by the accelerator cards through a first non-transparent bridge port, performing address translation and bus identifier translation on the data packets according to the address mapping table, and then forwarding the data packets to the destination accelerator card through a second non-transparent bridge port.

[0085] As shown in Figure 2, "HP" represents the host point and "DP" represents the device point. The first switch enables the non-transparent bridging function of its ports, achieving cross-host many-to-many communication through these non-transparent bridging ports. The non-transparent bridging ports are responsible for cross-host address translation and bus ID translation, implementing address partitioning for each non-transparent bridging port based on the interconnection topology. Through ingress and egress, the address space of the remote accelerator card is mapped to the address space of the local host, making access from the local accelerator card to the remote accelerator card (located on another server) equivalent to peer-to-peer (P2P) communication between local accelerator cards. This enables cross-host accelerator card peer-to-peer communication within a single high-bandwidth domain.

[0086] In some embodiments, PCIe Fabric Mode can also be used to implement cross-domain communication of accelerator cards. The address mapping table includes the address mapping relationship between the virtual device address assigned by the server to the remote device and the corresponding local device address information; the first switch forwards data between accelerator cards on different servers according to the address mapping table, which may include: after receiving a data packet sent by the accelerator card, the first switch forwards the data packet to the destination accelerator card according to the address mapping relationship.

[0087] As shown in Figure 3, "HP" represents the host point, "DP" represents the device point, and "FP" represents the fabric point. The first switch adopts Fabric Mode, utilizing the server's virtual device mapping to remote accelerator cards. By setting the global address table of each accelerator card in the computing system, cross-host communication between accelerator cards is achieved. The local host allocates virtual devices (placeholder device 1 as shown in Figure 9) to all remote resource pool nodes. The base address register (BAR) space of the virtual device is consistent with the corresponding remote resource pool node. All first switches in the computing system are uniformly configured to establish the address mapping relationship between virtual devices and remote accelerator cards. This makes access to virtual devices by local accelerator cards equivalent to point-to-point communication between local and remote accelerator cards, thereby realizing cross-host point-to-point communication between accelerator cards within a single high-bandwidth domain.

[0088] The computing system provided in this application embodiment interconnects accelerator cards located on multiple servers through a first switch using a high-speed serial computer expansion bus. The first switch obtains the local device address of the accelerator card on the server, configures the address mapping table of each server according to the local device address of each accelerator card, and forwards data between accelerator cards on different servers according to the address mapping table. This effectively solves the problems of limited computing power of a single server and insufficient bandwidth and large latency in the interconnection of multiple servers via Ethernet. It realizes cross-server vertical expansion based on a high-speed serial computer expansion bus, improves the bandwidth of cross-server communication, reduces the latency of cross-server communication, and lays the foundation for realizing high-bandwidth domain communication computing clusters with multiple machines and multiple cards.

[0089] The accelerator card cross-host domain interconnection scheme described in the above embodiments enables vertical scaling of the computing system. Furthermore, high-speed horizontal scaling can be achieved through interconnection between multiple accelerator cards within a server. In the computing system provided in this application embodiment, different accelerator cards located on the same server can be interconnected via a high-speed serial computer expansion bus. Therefore, the computing system provided in this application embodiment achieves high-bandwidth domain interconnection based on PCIe for different accelerator cards located on the same server, and also achieves high-bandwidth domain interconnection based on PCIe for accelerator cards located on different servers, thus realizing a multi-machine, multi-card cluster with high-bandwidth domain interconnection.

[0090] To achieve horizontal scaling of single-machine computing power, as shown in Figure 1, the computing system provided in this application embodiment may further include a second switch located on the server; different accelerator cards located on the same server are interconnected through a high-speed serial computer expansion bus and the second switch.

[0091] When a second switch is provided, the first switch is connected to the server's accelerator card via a high-speed serial computer expansion bus. This can include: the first switch being connected to the accelerator card via a high-speed serial computer expansion bus and the second switch.

[0092] This application provides a PCIe northbound interconnect cluster topology for building a computing resource pool, as shown in Figure 1. Four 8-card servers are interconnected via PCIe switches to form a 32-card interconnect system. Each server is configured with four L1 switches and eight Open Compute Accelerator Modules (OCP Accelerator Modules, OAM). Four L2 switches can be mounted in a switch box. Each server outputs four x16 ports through the L2 switches and connects to the L2 switches via a retimer card.

[0093] In some embodiments, the second switch may be a PEX89104 and the first switch may be a PEX89144.

[0094] As shown in Figure 1, the servers are designated as Host 0, Host 1, Host 2, and Host 3. The second switches in each server are designated as Second Switch SW0, Second Switch SW1, Second Switch SW2, and Second Switch SW3. The accelerator cards in each server are designated as Accelerator Card 0, Accelerator Card 1, Accelerator Card 2, Accelerator Card 3, Accelerator Card 4, Accelerator Card 5, Accelerator Card 6, and Accelerator Card 7. The central processing units (CPUs) in each server are designated as CPU 0 and CPU 1. The first switches are designated as First Switch SW0, First Switch SW1, First Switch SW2, and First Switch SW3. Each second switch can connect two accelerator cards and connect to the CPU via two connections. The first switches are connected to the second switches corresponding to their respective labels.

[0095] In some embodiments, a PCIe bus can be used to construct network connections in three-dimensional space (3D Mesh network) to balance the bandwidth of inter-machine communication and the interconnection bandwidth of internal accelerator cards, enabling vertical scaling of large-scale, high-bandwidth cluster systems. The 3D Mesh solution theoretically offers higher bandwidth than traditional Ethernet-based RoCE (RDMA over Converged Ethernet) / IB (InfiniBand) transmission schemes, superior PCIe interconnection performance, low latency, no packet loss, and the ability to achieve fine-grained control, making computation and memory access easier to parallelize.

[0096] It is understood that, according to the high-speed interconnection scheme for multi-machine and multi-card computing systems provided in the embodiments of this application, the first switch, the second switch, and the configuration ports can be selected according to actual needs to achieve a smaller or larger multi-machine and multi-card cluster.

[0097] In some embodiments, a second switch can also be used to interconnect the accelerator card with the central processing unit in the local host unit.

[0098] As shown in Figure 2, in some embodiments, in host 0, the central processing unit (CPU) is connected to the host-side port (HP) of the second switch 1, and accelerator card 1 is connected to the device-side port (DP) of the second switch 1; in the second switch 1, the CPU is connected to the host-side port (HP) of the second switch 2, and accelerator cards 2 and 2 are respectively connected to the device-side port (DP) of the second switch 2. In the first switch 2, the first non-transparent bridge port (with ingress 01 and ingress 02) is connected to the host-side port (HP) of the first switch 2, and the second non-transparent bridge port (with egress 01 and egress 02) is connected to the host-side port (HP) of the first switch 2. The path for accelerator card 1 to send data packets to accelerator card 3 is as follows: the data packets are forwarded to the host side port (HP) of the second switch via the device side port (DP) of the second switch, and then sent to the ingress 01 of the first non-transparent bridge port of the first switch 2 via the host side port (HP) of the second switch. The second switch 2 sends the data packets from ingress 01 to the device side port (DP) of the second switch 03 via egress 01 according to the port mapping relationship configured locally. Finally, the data packets are sent to accelerator card 3 via another device side port (DP) of the second switch 03.

[0099] As shown in Figure 3, in some embodiments, the first switch and the second switch adopt Fabric Mode. The virtual device mapping of the second switch is used to map the remote accelerator card. Cross-host communication of the accelerator card is achieved by setting the address mapping tables of the first and second switches. As shown in Figure 3, an address mapping table is set on the host-side port (HP) of the second switch 1 to forward data streams from the central processing unit of host 1 to the remote accelerator card; an address mapping table is set on the device-side port (DP) of the second switch 1 to forward data streams from the accelerator card on host 1 to the remote accelerator card. When accelerator card 1 needs to send a data packet to the remote accelerator card 2, it can send the packet through the device-side port (DP) of the second switch 1 to the fabric port (FP) of the second switch 1, then through the first switch 3 to the fabric port (FP) of the second switch 2 of host 2, and finally through the device-side port (DP) of the second switch 2 to the accelerator card 2.

[0100] Figure 4 is a schematic diagram of an accelerator card provided in an embodiment of this application; Figure 5 is a schematic diagram of a PCIe southbound interconnection cluster topology provided in an embodiment of this application.

[0101] Accelerator cards typically connect to a PCIe interface on a server motherboard or server backplane via gold fingers. This PCIe interface is a northbound interface used for communication between the accelerator card and the local host unit. To enable cross-host domain interconnection of accelerator cards, the accelerator card needs to support an interface for interconnection with remote accelerator cards. In some embodiments, an accelerator card configured with a southbound interface can be used to achieve cross-host domain interconnection. The southbound interface of the accelerator card is connected to a first switch via a high-speed serial computer expansion bus.

[0102] As shown in Figure 4, the accelerator card used in this embodiment may be equipped with a high-speed serial computer expansion bus switching module and a southbound interface. The first end of the high-speed serial computer expansion bus switching module is connected to the gold fingers of the accelerator card, the second end of the high-speed serial computer expansion bus switching module is connected to the controller of the accelerator card (i.e., the computing core of the accelerator card), and the third end of the high-speed serial computer expansion bus switching module is connected to the southbound interface of the accelerator card, thereby separating the PCIe signal introduced from the gold fingers into a signal for interconnection with the remote accelerator card.

[0103] In addition, as shown in Figure 4, the accelerator card can also be equipped with an inter-card interconnection interface to interconnect with other accelerator cards located on the same server.

[0104] Based on the computing system shown in Figure 1, the interconnection between accelerator cards can achieve vertical expansion of 4 to 16 cards. The southbound interfaces of the accelerator cards are interconnected with PCIe switch boxes in a cluster, supporting interconnection of 8 to 64 cards and scaling. Figure 5 shows a schematic diagram of the architecture of a 64-card interconnected computing system. Each dashed box represents the interconnection architecture within a server. It should be noted that Figure 5 only shows the connection relationship between the OAM module inside the server within one dashed box (the upper left corner of Figure 5) and the various switches (the first switch SW0 to SW3) (represented by thin connecting lines). The servers in the other dashed boxes are represented by thick connecting lines instead of multiple thin connecting lines to express the connection relationship between the server and the various switches. At the same time, the thick connecting lines do not limit the correspondence between the OAM module and the switch ports. Inside the server, through the Open Computing Accelerator Module (OCP Accelerator Module, OAM) technology, faster interconnection of accelerator cards within the OAM module can be achieved compared to accelerator cards in OAM modules. For example, S0, S3, S4, and S7 constitute an OAM module.

[0105] Figure 6 is a structural schematic diagram of a first switch provided in an embodiment of this application; Figure 7 is a connection schematic diagram of a first controller and a first switch provided in an embodiment of this application.

[0106] In some embodiments, a first switch may refer to a hardware device having a first switching unit, which is a processor configured to control the connection of its own interface to accelerator cards located on different servers to achieve cross-server data forwarding of accelerator cards. Depending on the cluster size of the computing system (number of servers, number of accelerator cards) and the bandwidth requirements of the accelerator cards, the number of first switches may be one or more. When the bandwidth requirement of an accelerator card is less than the bandwidth that one interface of the first switch can provide, one interface of the first switch may be split into multiple paths to connect multiple accelerator cards.

[0107] As shown in Figure 6, the first switch board is built around the first switching unit, which is connected to the switch's communication interface (Figure 6 shows the case where the communication interface is an optical interface). The first switching unit can use a PXE89144 and is configured to implement high-speed link switching. The fourth controller can use a field-programmable gate array and is configured to configure the switch's interface. The first interface is used to interconnect with other first switches via a PCIe bus when the first switch is expanded. The first interface can use a multi-chip interconnect (MCIO). The number of first interfaces (MCIO x8) can be two. The first switch board may also include a power connector, a control board connector, etc., which are connected to the low-speed interface of the first switching unit. The first switching unit is also connected to flash memory (SPI Flash). The fourth controller is also configured to connect a bidirectional transfer switch (which can use a CA9548 to switch between QSFP_I2C and QSFP_I2C [0-15]). When the communication interface of the first switch is an optical interface, the optical interface can be a QSFP-DD optical interface. When the first switching unit uses PXE89144, it can be connected to 16 QSFP-DD optical interfaces.

[0108] The first switch can also be configured to implement a user interface to receive and execute user requests for computing resource scheduling of the computing system. The first switch can be configured to receive configuration requirement parameters for the target resource pool, obtain accelerator card status parameters, and perform accelerator card resource scheduling based on the configuration requirement parameters and status parameters to construct the target resource pool.

[0109] The computing system provided in this application embodiment can realize dynamic scheduling of the computing resource pool. After the first switch completes the dynamic scheduling, it updates its local address mapping table.

[0110] When the computing system includes multiple first switches, scheduling of each first switch is required. When there are multiple first switches, the computing system may also include a first controller; the first controller is configured to generate an address mapping table using local device address information collected by the first switches, and send the address mapping table to the first switches to configure the address mapping table on the first switches.

[0111] As shown in Figure 7, the first controller can be deployed on a control baseboard (CCB). The first controller can be a microprocessor (mCPU).

[0112] Different first switches can be interconnected via a high-speed serial computer expansion bus. As shown in Figure 7, two first switches can be interconnected via their respective first interfaces based on a high-speed serial computer expansion bus.

[0113] Multiple first switch boards can be installed in a single PCIe switch box. A first controller can also be installed in a single switch box along with multiple first switch boards. The first controller can also be deployed in a server within a computing system.

[0114] As shown in Figure 7, the computing system provided in this application embodiment may further include a first out-of-band management controller and a second controller located on the board where the first controller is located, and a third controller located on the board where the first switch is located; the first out-of-band management controller is configured to implement out-of-band management of the first switch; the second controller is configured to control the power-on and power-off timing of the first switch; and the third controller is configured to configure the high-speed serial computer expansion bus interface of the first switch.

[0115] The second and third controllers can be Complex Programmable Logic Devices (CPLDs). The second controller, designated CPLD-B, is configured for the power-on / off sequence control, logic judgment, and indicator lighting of the first switch. The third controller, designated CPLD-S, is configured for configuring the communication interfaces of the first switch.

[0116] The user interface can also be implemented by the first controller to receive and execute user requests for computing resource scheduling of the computing system. The first controller can also be configured to receive configuration requirement parameters for the target resource pool, obtain the status parameters of the accelerator cards through the first switch, and perform resource scheduling of the accelerator cards according to the configuration requirement parameters and status parameters to construct the target resource pool.

[0117] The computing system provided in this application embodiment can realize dynamic scheduling of the computing resource pool. After performing dynamic scheduling, the first controller updates the address mapping table of each first switch. If the computing system also includes a second switch, the first controller updates the address mapping table of each first switch and the address mapping table of each second switch after performing dynamic scheduling.

[0118] In some embodiments, to implement a larger-scale computing system, the first switch may include a third switch directly connected to the server and a fourth switch configured to enable interconnection between different third switches. By forming a switch network using the second switch, the first switch, and the fourth switch, a larger-scale and more complex multi-machine, multi-card interconnection cluster can be achieved.

[0119] Figure 8 is a schematic diagram of the control structure of a control baseboard provided in an embodiment of this application; Figure 9 is a schematic diagram of a cluster interconnection centralized collaborative management structure based on PCIe bus provided in an embodiment of this application.

[0120] As shown in Figure 8, the first controller, acting as the overall system management node, can also collaboratively manage the host baseboard management controller (Host BMC), complex programmable logic devices, and the second switch. Servers can interconnect via the first switch, decoupling the timing logic of the accelerator cards, switches, and central processing unit, and controlling their coordinated power-on and power-off through switch management software; enabling point-to-point data transmission between accelerator cards across nodes. In the 4-host system shown in Figure 1, complex and precise routing management of 20 switches is required, supporting interconnection of 8 / 16 / 24 / 32 cards and flexible online switching.

[0121] In terms of accelerator card management, the computing unit and management unit are decoupled and standardized, separating common management, security, and control functions from the computing unit. This ensures compatibility with different computing and management platforms, supports unified management of multiple interfaces and chips, and meets the needs of various application scenarios. The management unit needs to implement computing power allocation and management, multi-chip module voltage regulation and power consumption management to ensure high-efficiency system design; monitor resource utilization, I / O throughput, and resource health status; achieve centralized management of coordinated power-on / off, resources, and topology to ensure system availability; and enable on-demand rapid deployment and automatic management of hardware resources, monitoring of key resource information and fault management, and intelligent fault location and recovery to ensure computing system reliability. The system management module is responsible for unified management, providing operation and maintenance capabilities through standardized service interfaces, and achieving integrated monitoring, fault early warning, and visualized management.

[0122] As shown in Figure 9, the server's Host Basic Input / Output System (HBIOS) publishes placeholder device information and accelerator card information to the server's baseboard management controller. The first controller receives configuration requirement parameters, obtains placeholder device information and accelerator card information from the server, schedules and configures the accelerator cards to build the target resource pool, and updates the routing information in the second switch and the first switch according to the target resource pool. It then pushes the topology information and routing information to the server, and the server's host operating system can obtain the topology information from the baseboard management controller.

[0123] The first controller on the control baseboard provides user interface functions, which can realize the following accelerator card management functions: Topology identification: Supports viewing the topology view, and can view the node summary information, such as node type, power-on status, overall health status, etc.

[0124] Asset Information Management: Supports viewing device information at each level. Detailed device information, such as device asset information and high-speed interface connection status, can be viewed through the Web / Redfish page.

[0125] Collaborative control: Supports centralized power-on / off control, with each unit powered on and off in sequence; supports system reset function, automatically detecting and resetting the system during power-on and host restart; supports reset control after resource reallocation, etc.

[0126] In system design, the discovery, management, and elastic adjustment of system resources are crucial. The first controller, as the core management unit for dynamic resource adjustment, controls the second and first switches to achieve automatic resource topology discovery and flexible automatic resource switching. This can include: Resource identification and display: providing network interfaces and a visual web interface to display resource lists and topologies, including accelerator card information, I / O port information, etc.; Dynamic resource allocation: achieving second-level dynamic allocation and adjustment of accelerator cards through key technologies such as hot removal, hot insertion, and hot reset; Load balancing: dynamically allocating physical resources based on optimized load balancing scheduling algorithms, achieving fine-grained on-demand allocation of storage resources, maximizing the release of heterogeneous computing power; Expert templates: providing expert template descriptions and dynamic switching interfaces for heterogeneous computing resources based on business needs, resource status, and performance indicators, allowing applications to request resources based on expert templates, while also providing network interfaces and a visual web interface for the visual application of expert templates.

[0127] Traditional multi-server interconnection solutions via Ethernet suffer from high transmission latency and lack support for memory semantics such as Load (LD) and Store (ST), requiring adaptation at the transaction layer. Ethernet routing is inefficient, further increasing transmission latency. The UDP / IP header, a combination of User Datagram Protocol (UDP) and Internet Protocol (IP) network communication protocols, is not mandatory in in-machine scenarios, resulting in low payload efficiency.

[0128] In the computing system provided in this application embodiment, accelerator cards located on the same server are interconnected via PCIe bus, and accelerator cards located on different servers are also interconnected via PCIe bus. This utilizes the advantages of PCIe bus, which supports memory semantics such as Load (LD) and Store (ST), and has high transmission efficiency and low latency, to build a large-scale multi-machine multi-card computing system.

[0129] In some embodiments, the first switch is connected to the server's accelerator card via a high-speed serial computer expansion bus, which may include: the first switch and the accelerator card being connected via a copper cable based on the high-speed serial computer expansion bus protocol.

[0130] In practical applications, copper cable connections based on the high-speed serial computer extended bus protocol have significant losses. Therefore, when interconnecting accelerator cards between a small number of servers, copper cables based on the high-speed serial computer extended bus protocol can be used. However, when interconnecting large-scale multi-machine and multi-card clusters, a PCIe interconnection solution that can transmit over long distances and has low losses is required.

[0131] In some embodiments, the first switch is connected to the server's accelerator card via a high-speed serial computer expansion bus, and may further include: the first switch and the accelerator card are connected via a first fiber optic link based on the high-speed serial computer expansion bus protocol.

[0132] Establishing PCIe bus optical interconnects also presents challenges. Link rate negotiation during PCIe interconnect link training involves changes in the link transmission rate. This can cause the clock and data recovery (CDR) chip and digital signal processor (DSP) in traditional optical modules to fail to lock the signal frequency, resulting in bit errors and link establishment failure. Furthermore, the CDR chip and DSP in optical modules only support transmitting and receiving data at fixed rates, and the optical module design only plans for the data channel. In addition to high-speed data signals, PCIe links contain many low-speed auxiliary signals crucial for link establishment, such as the reset signal (PERST#) and clock signal (CLOCK). These signals have frequencies much lower than the data signals, and the CDR chip and DSP in the optical module cannot correctly lock these low-speed auxiliary signals, causing low-speed PCIe signals to fail to transmit.

[0133] To address the compatibility issues of PCIe bus optical interconnects, this application embodiment provides a PCIe bus optical interconnect link that addresses the problem that the clock data recovery chip and digital signal processor in the optical module only support transmitting and receiving data at fixed rates with upper and lower limits and cannot automatically switch rates. This application embodiment employs a linear physical optical (LPO) device without a clock data recovery chip and digital signal processor. Signal processing and clock locking functions are implemented within a board-level switch or retimer, and the LPO device only performs photoelectric conversion. For the transmission of low-speed sideband signals in the PCIe bus, this application embodiment designs a relay module to modulate the low-speed sideband signals into high-frequency 200MHz LVDS differential signals for optical transmission.

[0134] Figure 10 is a schematic diagram of a PCIe optical interconnect link provided in an embodiment of this application; Figure 11 is a schematic diagram of another PCIe optical interconnect link provided in an embodiment of this application.

[0135] As shown in Figure 10, the optical interface used to implement PCIe can be composed of repeater modules and optoelectronic conversion devices. At the PCIe port, the Root Complex (RC) is a key component in the PCIe architecture, responsible for connecting the host system to the PCIe bus. The Root Complex typically consists of one or more Root Ports, each of which can connect to one or more PCI devices. Endpoint devices (EPs) are the terminal nodes in the PCIe topology, providing actual functions or services, such as network interface cards and graphics cards.

[0136] The relay module is configured for data signal compensation and recovery, link training control, and auxiliary signal aggregation and transmission. Data signal compensation and recovery ensures that the quality of the electrical signals entering the optoelectronic conversion device (or endpoint device) meets the PCIe specification. Link training control manages the PCIe link training process, supporting skipping processes that can easily lead to link establishment failures, such as receiver detection. Auxiliary signal aggregation and transmission converts and aggregates more than a dozen auxiliary signals with different frequencies defined in the PCIe protocol, reducing the number of links. Therefore, the relay module provided in this embodiment can realize the conversion between PCIe data signals and PCIe auxiliary signals and electrical signals containing auxiliary signal information.

[0137] The optoelectronic conversion device is configured to convert between an electrical signal containing auxiliary signal information and an optical signal containing auxiliary signal information. The optoelectronic conversion device is configured to be compatible with a wide range of electrical signals and support the transmission of aggregated auxiliary signals. By being compatible with a wide range of electrical signals, it ensures support for PCIe link rate negotiation, meeting a minimum signal rate of 2.5Gbps and a maximum signal rate of 32Gbps. By supporting the transmission of aggregated auxiliary signals, the optoelectronic conversion device provided in this application embodiment has an auxiliary signal transmission device and an optical channel, and the device performance supports the transmission of aggregated auxiliary signals.

[0138] Therefore, in view of the incompatibility issues of link rate negotiation, receiver detection and low-speed control signal optical transmission in the optical fiber transmission of PCIe protocol signals, the embodiments of this application realize a linear optical transmission scheme for PCIe mixed rate signals, realizing the conversion of PCIe electrical signals to optical signals.

[0139] Based on the above principles, as shown in Figure 11, in the computing system provided in this application embodiment, the first optical fiber link may include a first optical interface disposed on a server, a second optical interface disposed on a first switch, and a first optical fiber; both the first optical interface and the second optical interface include: a first relay module, a second relay module, and a linear direct-drive optoelectronic conversion device; the first relay module is configured to relay data signals of a high-speed serial computer expansion bus, and the first channel of the linear direct-drive optoelectronic conversion device is disposed between the first relay module and the first optical fiber; the second relay module is configured to relay auxiliary signals of a high-speed serial computer expansion bus; the second channel of the linear direct-drive optoelectronic conversion device is disposed between the second relay module and the first optical fiber.

[0140] The first relay module can be a high-speed serial computer expansion bus switch (PCIe switch). As shown in Figure 11, the first relay module can include a control module and input / output ports at both ends. One end is used for inputting and outputting the first PCIe data signal (input A_PET_P / N, output A_PER_P / N), and the other end is used for inputting and outputting the second PCIe data signal (output B_PET_P / N, input B_PER_P / N), bypass receiver detection data, and clock recovery signal. The first data signal is the PCIe data signal on the host side, and the second data signal is the PCIe data signal used for photoelectric conversion.

[0141] The second relay module can be a field-programmable gate array (FPGA). As shown in Figure 11, the second relay module is configured to convert PCIe auxiliary signals (such as reset signal PERST, power enable signal POWEN, power status signal POWGD, system management bus clock signal SMBCLK, system management bus data signal SMBDAT, etc.) into low-voltage differential signaling (LVDS) for photoelectric conversion by implementing the Low-Voltage Differential Signaling Protocol & Interface (LTPI).

[0142] As shown in Figure 11, the first linear direct-drive optoelectronic conversion device may include a linear driver and a laser diode (LD) configured to convert the second data signal into a high-speed optical signal, and a photodiode (PD) and a linear transimpedance amplifier (TIA) configured to convert the high-speed optical signal into the second data signal. The second linear direct-drive optoelectronic conversion device may include a driver and a laser diode (LD) configured to convert the aforementioned low-voltage differential signal into a low-speed optical signal, and a photodiode (PD) and a transimpedance amplifier (TIA) configured to convert the low-speed optical signal into a low-voltage differential signal. Then, the data optical signal and the auxiliary optical signal are transmitted independently via the first optical fiber.

[0143] Therefore, this application embodiment addresses the incompatibility issues in PCIe bus optical interconnects, such as link rate negotiation, receiver detection, and low-speed control signal optical transmission. It can employ a high-speed serial computer extended bus switch and a field-programmable gate array to handle the relay processing of high-speed digital signals and low-speed control signals, respectively. Then, the mixed-rate signals are transmitted via a linear optical module, enabling linear optical transmission of PCIe 5.0 mixed-rate signals.

[0144] In the computing system provided in this application embodiment, the accelerator card and the second switch are connected via a copper cable based on the High-Speed ​​Serial Computer Extended Bus (HS-SBU) protocol or a second optical fiber link based on the H-SBU protocol. The structure of the second optical fiber link can refer to the structure of the first optical fiber link.

[0145] Since the interconnection distance between accelerator cards within a server is typically short, PCIe interconnection can be achieved using copper cables based on the high-speed serial computer expansion bus protocol. Of course, in suitable scenarios, the PCIe bus optical interconnection scheme provided in the embodiments of this application can be used to achieve interconnection between accelerator cards within the server.

[0146] As shown in Figure 6, the PCIe bus optical interconnect scheme provided in this application uses an optical interface for the communication interface of the first switch. In the switch box with two first switches (first switch 0 and first switch 1) as shown in Figure 7, 32 high-speed link interfaces of Gen5 x8 QSFP-DD (Quad Small Form Factor Pluggable Double Density, a packaging form for 400G optical modules) can be provided, supporting PCIe optical interconnect; one or two second switching units (second switching unit 0 and second switching unit 1), each of which can be configured with one PXE89144 and 16 QSFP-DD slots. The third controller is configured to perform QSFP-DD optical module configuration.

[0147] The various embodiments corresponding to the computing system have been described in detail above. Based on this, this application also discloses switches, control methods, devices, equipment, non-volatile storage media, and computer program products corresponding to the above-described computing system.

[0148] The switch provided in this application is connected to accelerator cards located on different servers via a high-speed serial computer expansion bus. The switch is configured to obtain the local device address information of the accelerator cards on the server, configure the address mapping table of each server according to the local device address information, and perform data forwarding between accelerator cards on different servers according to the address mapping table.

[0149] Optional implementations of the switch provided in this application can refer to the description of the first switch in the above-described computing system embodiments.

[0150] The control method for the computing system provided in this application embodiment may include:

[0151] Obtain the local device address information of the server's accelerator card through the high-speed serial computer expansion bus;

[0152] Configure the address mapping table of each server according to the address information of each local device;

[0153] Data forwarding between accelerator cards on different servers is performed based on the address mapping table.

[0154] Optional implementations of the control method for the computing system provided in this application embodiment can refer to the description of the control baseboard in the computing system described above.

[0155] It should be noted that in the embodiments of the control methods for the various computing systems in this application, some steps or features may be omitted or not executed. The division of hardware or software functional modules is for ease of explanation and is not the only implementation of the control methods for the computing systems provided in the embodiments of this application.

[0156] The control method for the computing system provided in this application embodiment interconnects accelerator cards located on multiple servers by using a high-speed serial computer expansion bus. It obtains the local device address of the accelerator card on the server using the high-speed serial computer expansion bus, configures the address mapping table of each server based on the local device address of each accelerator card, and forwards data between accelerator cards on different servers according to the address mapping table. This effectively solves the problems of limited computing power of a single server and insufficient bandwidth and high latency in interconnecting multiple servers via Ethernet. It realizes cross-server vertical expansion based on a high-speed serial computer expansion bus, improves the bandwidth of cross-server communication, reduces the latency of cross-server communication, and lays the foundation for realizing high-bandwidth domain communication computing clusters with multiple machines and multiple cards.

[0157] Applied to the control baseboard, the control device for the computing system provided in this application embodiment may include:

[0158] The acquisition unit is configured to acquire the local device address information of the server's accelerator card via a high-speed serial computer expansion bus;

[0159] The configuration unit is set to configure the address mapping table of each server based on the address information of each local device;

[0160] The control unit is configured to forward data between accelerator cards on different servers based on an address mapping table.

[0161] It should be noted that in the various embodiments of the control device for the computing system provided in this application, the division of units is only a logical functional division, and other division methods can be used. The connection between different units can be electrical, mechanical, or other connection methods. Separate units can be located in the same physical location or distributed across multiple network nodes. Each unit can be implemented in hardware or as a software functional unit. That is, some or all of the units provided in this application can be selected according to actual needs, and corresponding connection or integration methods can be used to achieve the purpose of the solution in this application.

[0162] The control device for the computing system provided in this application embodiment interconnects accelerator cards located on multiple servers by using a high-speed serial computer expansion bus. It obtains the local device address of the accelerator card on the server using the high-speed serial computer expansion bus, configures the address mapping table of each server based on the local device address of each accelerator card, and forwards data between accelerator cards on different servers according to the address mapping table. This effectively solves the problems of limited computing power of a single server and insufficient bandwidth and high latency in interconnecting multiple servers via Ethernet. It realizes cross-server vertical expansion based on a high-speed serial computer expansion bus, improves the bandwidth of cross-server communication, reduces the latency of cross-server communication, and lays the foundation for realizing high-bandwidth domain communication computing clusters with multiple machines and multiple cards.

[0163] The control device for the computing system provided in this application includes: a memory configured to store a computer program; and a processor configured to execute the computer program, wherein the computer program, when executed by the processor, implements the steps of the control method for the computing system provided in any of the above embodiments.

[0164] The processor may include one or more processing cores, such as a 3-core processor or an 8-core processor. The processor may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array. The processor may also include a main processor and coprocessors. The main processor, also known as the Central Processing Unit (CPU), is configured to process data in the wake-up state; the coprocessors are low-power processors configured to process data in the standby state. In some embodiments, the processor may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, the processor may also include an artificial intelligence processor configured to handle computational operations related to machine learning.

[0165] The memory may include one or more non-volatile storage media, which may be non-transitory. The memory may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In this embodiment, the memory is at least configured to store the following computer program, wherein, after being loaded and executed by the processor, the computer program is capable of implementing the relevant steps in the control method of the computing system disclosed in any of the foregoing embodiments. Furthermore, the resources stored in the memory may also include operating systems and data, and the storage method may be temporary or permanent storage. The operating system may be Windows or other types of operating systems. The data may include, but is not limited to, the data involved in the above methods.

[0166] In some embodiments, the control device of the computing system may further include a display screen, a power supply, a communication interface, an input / output interface, a sensor, and a communication bus.

[0167] Those skilled in the art will understand that the structures shown in the embodiments of this application do not constitute a limitation on the control device of the computing system, and may include more or fewer components than shown.

[0168] The control device for the computing system provided in this application includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the steps of the control method for the computing system provided in the above embodiments. It achieves interconnection between accelerator cards located on multiple servers by using a high-speed serial computer expansion bus, obtains the local device address of the accelerator card on the server by using the high-speed serial computer expansion bus, configures the address mapping table of each server according to the local device address of each accelerator card, and forwards data between accelerator cards on different servers according to the address mapping table. This effectively solves the problems of limited computing power of a single server and insufficient bandwidth and large latency in the interconnection of multiple servers via Ethernet. It realizes cross-server vertical expansion based on a high-speed serial computer expansion bus, improves the bandwidth of cross-server communication, reduces the latency of cross-server communication, and lays the foundation for realizing high-bandwidth domain communication computing clusters with multiple machines and multiple cards.

[0169] This application provides a non-volatile storage medium storing a computer program thereon. When executed by a processor, the computer program can implement the steps of the control method of the computing system provided in any of the above embodiments.

[0170] The non-volatile storage medium may include: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks or optical disks, and other media that can store program code.

[0171] For a description of the non-volatile storage medium provided in this application embodiment, please refer to the above method embodiment. Its effect is the same as the control method of the computing system provided in this application embodiment. By using a high-speed serial computer expansion bus to achieve interconnection between accelerator cards located on multiple servers, the local device address of the accelerator card on the server is obtained by using the high-speed serial computer expansion bus, the address mapping table of each server is configured according to the local device address of each accelerator card, and data forwarding between accelerator cards on different servers is performed according to the address mapping table. This effectively solves the problems of limited computing power of a single server and insufficient bandwidth and large latency in the interconnection of multiple servers via Ethernet. It realizes cross-server vertical expansion based on the high-speed serial computer expansion bus, improves the bandwidth of cross-server communication, reduces the latency of cross-server communication, and lays the foundation for realizing high-bandwidth domain communication computing clusters with multiple machines and multiple cards.

[0172] This application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the control method for the computing system provided in any of the above embodiments.

[0173] For a description of the computer program product provided in this application embodiment, please refer to the above method embodiment. The effect it achieves is the same as the control method of the computing system provided in this application embodiment. By using a high-speed serial computer expansion bus to achieve interconnection between accelerator cards located on multiple servers, the local device address of the accelerator card on the server is obtained by using a high-speed serial computer expansion bus, the address mapping table of each server is configured according to the local device address of each accelerator card, and data forwarding between accelerator cards on different servers is performed according to the address mapping table. This effectively solves the problems of limited computing power of a single server and insufficient bandwidth and large latency in the interconnection of multiple servers via Ethernet. It realizes cross-server vertical expansion based on a high-speed serial computer expansion bus, improves the bandwidth of cross-server communication, reduces the latency of cross-server communication, and lays the foundation for realizing a computing cluster with high bandwidth domain communication of multiple machines and multiple cards.

[0174] The foregoing has provided a detailed description of the computing system, control method, apparatus, device, non-volatile storage medium, and switch provided in this application. The various embodiments in the specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the switches, control methods, apparatus, devices, non-volatile storage media, and computer program products disclosed in the embodiments, since they correspond to the computing system disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to in the computing system section. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of this application.

[0175] It should also be noted that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

Claims

1. A computing system, characterized in that, include: Multiple servers and a first switch; The first switch is connected to the server's acceleration card via a high-speed serial computer expansion bus; The first switch is configured to obtain the local device address information of the accelerator card on the server, configure the address mapping table of each server according to the local device address information, and perform data forwarding between the accelerator cards on different servers according to the address mapping table.

2. The computing system according to claim 1, characterized in that, The first switch forwards data between the accelerator cards on different servers according to the address mapping table, including: The first switch receives the data packets sent by the accelerator card through the first non-transparent bridge port, performs address translation and bus identifier translation on the data packets according to the address mapping table, and then forwards the data packets to the destination accelerator card through the second non-transparent bridge port.

3. The computing system according to claim 1, characterized in that, The address mapping table includes the address mapping relationship between the virtual device address allocated by the server to the remote device and the corresponding local device address information; The first switch forwards data between the accelerator cards on different servers according to the address mapping table, including: After receiving the data packet sent by the accelerator card, the first switch forwards the data packet to the destination accelerator card according to the address mapping relationship.

4. The computing system according to claim 1, characterized in that, The southbound interface of the accelerator card is connected to the first switch via a high-speed serial computer expansion bus.

5. The computing system according to claim 1, characterized in that, The number of the first switches is multiple, and the computing system also includes a first controller; The first controller is configured to generate the address mapping table using the local device address information collected by the first switch, and send the address mapping table to the first switch to configure the address mapping table on the first switch.

6. The computing system according to claim 5, characterized in that, The different first switches are interconnected via a high-speed serial computer expansion bus.

7. The computing system according to claim 5, characterized in that, It also includes a first out-of-band management controller and a second controller located on the board where the first controller is located, and a third controller located on the board where the first switch is located; The first out-of-band management controller is configured to implement out-of-band management of the first switch; the second controller is configured to control the power-on and power-off timing of the first switch; and the third controller is configured to configure the high-speed serial computer expansion bus interface of the first switch.

8. The computing system according to claim 5, characterized in that, The first controller is also configured to receive configuration requirement parameters for the target resource pool, obtain the status parameters of the accelerator card through the first switch, and perform resource scheduling of the accelerator card according to the configuration requirement parameters and the status parameters to construct the target resource pool.

9. The computing system according to claim 1, characterized in that, The first switch is also configured to receive configuration requirement parameters for the target resource pool, obtain status parameters of the accelerator card, and perform resource scheduling of the accelerator card according to the configuration requirement parameters and the status parameters to construct the target resource pool.

10. The computing system according to claim 1, characterized in that, Different accelerator cards located on the same server are interconnected via a high-speed serial computer expansion bus.

11. The computing system according to claim 1, characterized in that, It also includes a second switch located on the server; Different accelerator cards located on the same server are interconnected via a high-speed serial computer expansion bus and the second switch.

12. The computing system according to claim 1, characterized in that, It also includes a second switch located on the server; The first switch is connected to the server's acceleration card via a high-speed serial computer expansion bus, including: The first switch is connected to the accelerator card via a high-speed serial computer expansion bus and the second switch.

13. The computing system according to claim 1, characterized in that, The first switch is connected to the server's acceleration card via a high-speed serial computer expansion bus, including: The first switch and the accelerator card are connected via a copper cable based on the high-speed serial computer extended bus protocol.

14. The computing system according to claim 1, characterized in that, The first switch is connected to the server's acceleration card via a high-speed serial computer expansion bus, including: The first switch and the acceleration card are connected via a first optical fiber link based on the high-speed serial computer extended bus protocol.

15. The computing system according to claim 14, characterized in that, The first optical fiber link includes a first optical interface located on the server, a second optical interface located on the first switch, and a first optical fiber; Both the first optical interface and the second optical interface include: a first relay module, a second relay module, and a linear direct-drive optoelectronic conversion device; The first relay module is configured to relay data signals from a high-speed serial computer extended bus, and the first channel of the linear direct-drive optoelectronic conversion device is located between the first relay module and the first optical fiber. The second relay module is configured to relay auxiliary signals of the high-speed serial computer extended bus; the second channel of the linear direct-drive optoelectronic conversion device is located between the second relay module and the first optical fiber.

16. The computing system according to claim 1, characterized in that, The first switch includes a third switch that is directly connected to the server and a fourth switch configured to enable interconnection between different third switches.

17. A switch, characterized in that, The switch is connected to accelerator cards located on different servers via a high-speed serial computer expansion bus. The switch is configured to obtain the local device address information of the accelerator card on the server, configure the address mapping table of each server according to the local device address information, and perform data forwarding between the accelerator cards on different servers according to the address mapping table.

18. A control method for a computing system, characterized in that, include: Obtain the local device address information of the server's accelerator card through the high-speed serial computer expansion bus; Configure the address mapping table of each server according to the address information of each local device; Data forwarding between the accelerator cards on different servers is performed according to the address mapping table.

19. A control device for a computing system, characterized in that, include: The acquisition unit is configured to acquire the local device address information of the server's accelerator card via a high-speed serial computer expansion bus; The configuration unit is configured to configure the address mapping table of each of the servers according to the address information of each of the local devices; The control unit is configured to forward data between the accelerator cards on different servers according to the address mapping table.

20. A control device for a computing system, characterized in that, include: The memory is configured to store computer programs; A processor is configured to execute the computer program, which, when executed by the processor, implements the steps of the control method for the computing system as described in claim 18.

21. A non-volatile storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the control method for the computing system as described in claim 18.

22. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the control method for the computing system as described in claim 18.