A multi-card interworking system
By setting up secondary and primary switching units within the server node, combined with a redundant rack design, the problem of limited GPU communication bandwidth in traditional server nodes is solved, enabling efficient interconnection between multiple GPUs and business continuity.
Patent Information
- Application Number
- CN202511460456.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2045-10-11
AI Technical Summary
Traditional server nodes suffer from limited communication bandwidth between GPUs, resulting in low computing power and an inability to effectively migrate services when multiple GPUs fail, which can easily lead to service interruptions.
Design a multi-card interconnection system, including a multi-card cabinet and a redundant cabinet. Each server node is equipped with a secondary switching unit and a primary switching unit. Accelerator communication across server nodes is realized through the connection between the secondary switching units and the primary switching units. A redundant cabinet is set up in the redundant cabinet to migrate services in the event of a failure of the multi-card cabinet.
It enables efficient interconnection between multiple cards, improves the reliability of system operation and maintenance, and ensures business continuity in the event of multi-card cabinet failure.
Smart Images

Figure CN120929410B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of server technology, and in particular to a multi-card interconnection system. Background Technology
[0002] Traditional server nodes typically support a set of eight graphics processing units (GPUs), but these eight GPUs are not interconnected, and the communication bandwidth between GPUs is limited.
[0003] Currently, GPU computing power is generally low, making high-bandwidth intra-node interconnect (scale-up) technology even more necessary to improve GPU computing efficiency. While some GPUs support scale-up interconnect protocols, the scale of a full scale-up interconnect is limited to eight cards. These interconnect protocols vary from manufacturer to manufacturer and are closed-source, posing significant challenges to ecosystem compatibility with server nodes. Furthermore, some GPUs do not support inter-card scale-up interconnect protocols at all, requiring scaling up the network via Ethernet / InfiniBand (IB) for inter-node interconnect (scale-out) technology, resulting in low bandwidth and high latency. Although GPU interconnect allows for service migration in the event of a single GPU failure, this becomes impossible when a large number of GPUs fail, leading to service interruptions.
[0004] It is evident that how to achieve efficient interconnection between multiple cards and improve the reliability of system operation and maintenance is a problem that needs to be solved by those skilled in the art. Summary of the Invention
[0005] This application provides a multi-card interconnection system to at least solve the problems of low bandwidth, high latency, and easy service interruption in related technologies when using multi-card interconnection.
[0006] This application provides a multi-card interconnection system, including multiple multi-card racks and at least one redundant rack; each multi-card rack includes multiple server nodes and multiple primary switching units; the redundant rack includes at least one server node and multiple primary switching units; each server node includes multiple accelerators and multiple secondary switching units.
[0007] Each server node can connect any two accelerators to achieve full connectivity communication between multiple accelerators; each server node contains multiple secondary switching units that are connected to multiple accelerators, and the multiple secondary switching units on each server node are connected to each primary switching unit to achieve accelerator connectivity between any two server nodes.
[0008] Each primary switching unit in the redundant rack is connected to a corresponding primary switching unit in each multi-card rack, so as to realize the connection between any accelerator in the server node of the redundant rack and any accelerator in the server node of any multi-card rack.
[0009] According to this application, a multi-card interconnection system includes multiple multi-card racks and at least one redundant rack. Each multi-card rack includes multiple server nodes and multiple primary switching units. The redundant rack includes at least one server node and multiple primary switching units. Each server node contains multiple accelerators and multiple secondary switching units. Any two accelerators in each server node are connected to each other to achieve full connectivity communication among multiple accelerators. The multiple secondary switching units in each server node are connected to multiple accelerators, and the multiple secondary switching units on each server node are connected to each primary switching unit to achieve accelerator connectivity between any two server nodes. Each primary switching unit in the redundant rack is connected to a corresponding primary switching unit in each multi-card rack to achieve connectivity between any accelerator in a server node of the redundant rack and any accelerator in a server node of any multi-card rack. In this technical solution, by setting a second switching unit within each server node and setting primary switching units between server nodes, connectivity between accelerators across server nodes can be achieved through the connection between the secondary switching units and the primary switching units. Furthermore, by setting up redundant racks, when a large number of accelerators fail in a multi-SIM rack, services from the failed accelerators can be migrated not only to other multi-SIM racks but also to redundant racks, ensuring service continuity in the multi-SIM interconnection system. This achieves efficient interconnection between multiple SIM cards while improving the reliability of system operation and maintenance. Attached Figure Description
[0010] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram of the structure of a multi-card interconnection system provided in an embodiment of this application;
[0012] Figure 2 This is a schematic diagram of a redundant cabinet configuration provided in an embodiment of this application;
[0013] Figure 3 A redundant rack interconnection topology diagram provided for embodiments of this application;
[0014] Figure 4This is a schematic diagram of a multi-card cabinet provided in an embodiment of this application;
[0015] Figure 5 A schematic diagram of eight interconnected accelerators provided for an embodiment of this application;
[0016] Figure 6 A hardware interconnection topology diagram of a single server node provided in an embodiment of this application;
[0017] Figure 7 A topology diagram of a 32-card interconnection system provided in this application embodiment;
[0018] Figure 8 This is a topology diagram of a 72-card interconnection system provided in an embodiment of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0020] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0021] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0022] Figure 1 The schematic diagram of a multi-card interconnection system provided in this application embodiment includes multiple multi-card cabinets 1 and at least one redundant cabinet 2; each multi-card cabinet 1 includes multiple server nodes 10 and multiple primary switching units 20; the redundant cabinet 2 includes at least one server node 10 and multiple primary switching units 20; each server node 10 includes multiple accelerators and multiple secondary switching units.
[0023] Each server node 10 can connect any two accelerators to achieve full connectivity communication between multiple accelerators; each server node 10 contains multiple secondary switching units that are connected to multiple accelerators, and each server node 10 can connect multiple secondary switching units to each primary switching unit to achieve accelerator connectivity between any two server nodes.
[0024] To improve the reliability of the multi-card interconnection system, a redundant rack 2 can be set up. Each primary switching unit 20 in the redundant rack 2 is connected to a corresponding primary switching unit 20 in each multi-card rack 1, so as to realize the connection between the server node 10 in the redundant rack 2 and any server node 10 in the multi-card rack 1.
[0025] exist Figure 1 To more clearly and intuitively demonstrate the connection relationship between the primary switching unit 20 in the multi-card rack 1 and the primary switching unit 20 in the redundant rack 2, for each server node 10 containing multiple accelerators and multiple secondary switching units, [the following is not shown in the original text]. Figure 1 Presented in [the document / document]. The specific architecture of each server node 10 can be found in [the document / document]. Figure 4 Introduction.
[0026] It should be noted that, Figure 1 The example shown uses two multi-card racks, each containing four server nodes and two primary switching units. In practical applications, there is no limit to the number of multi-card racks, or the number of server nodes and primary switching units in each rack; these can be flexibly configured based on actual needs. Figure 1 This is for illustrative purposes only and is not intended to limit the number of multi-card racks, or the number of server nodes and primary switching units contained in each multi-card rack.
[0027] In this embodiment, the primary switching unit can be a primary switch or a primary switching device. The secondary switching unit can be a secondary switch or a secondary switching device of a different type than the primary switching unit. The following description will use primary and secondary switches as examples.
[0028] As can be seen from the above technical solution, the multi-card interconnection system includes multiple multi-card racks and at least one redundant rack; each multi-card rack includes multiple server nodes and multiple primary switching units; the redundant rack includes at least one server node and multiple primary switching units. Each server node contains multiple accelerators and multiple secondary switching units; any two accelerators in each server node are connected to each other to achieve full connectivity communication among multiple accelerators. The multiple secondary switching units in each server node are connected to multiple accelerators, and the multiple secondary switching units on each server node are connected to each primary switching unit to achieve accelerator connectivity between any two server nodes. Each primary switching unit in the redundant rack is connected to a corresponding primary switching unit in each multi-card rack to achieve connectivity between any accelerator in a server node of the redundant rack and any accelerator in a server node of any multi-card rack. In this technical solution, by setting a second switching unit within each server node and setting primary switching units between server nodes, connectivity between accelerators across server nodes can be achieved through the connection between the secondary switching units and the primary switching units. Furthermore, by setting up redundant racks, when a large number of accelerators fail in a multi-SIM rack, services from the failed accelerators can be migrated not only to other multi-SIM racks but also to redundant racks, ensuring service continuity in the multi-SIM interconnection system. This achieves efficient interconnection between multiple SIM cards while improving the reliability of system operation and maintenance.
[0029] In this embodiment, each primary switching unit in each multi-card rack includes two sets of ports; the first set of ports is used to connect to the transmission card in each multi-card rack; the second set of ports is used to connect to the primary switching unit in the redundant rack. Each primary switching unit in the redundant rack includes two sets of ports; the first set of ports is used to connect to the primary switching unit in each multi-card rack; the second set of ports is used to connect to the transmission card in the server node of the redundant rack.
[0030] Each primary switch has a first group of ports consisting of multiple first ports and a second group of ports consisting of one second port. Each primary switching unit consists of two primary switching sub-units. The number of primary switching sub-units in multi-card racks and redundant racks can be flexibly configured, as long as the second ports of all primary switching sub-units in multi-card racks can be connected to one of the first ports of the primary switching sub-units in the redundant racks.
[0031] When the total number of multi-card racks is greater than the number of first ports contained in a primary switching subunit in a redundant rack, multiple multi-card racks can be divided into multi-card rack groups with the same number of first ports contained in a primary switching subunit. All primary switching subunits in the redundant racks can be divided into switching subunit groups with the same total number of multi-card racks contained in each multi-card rack group. The second port of the Nth switching subunit in each multi-card rack group can be sequentially connected to multiple first ports of the Nth switching subunit in its corresponding switching subunit group.
[0032] When the total number of multi-card racks is equal to the number of first ports contained in a primary switching subunit of a redundant rack, the second port of the Nth switching subunit in each multi-card rack is sequentially connected to multiple first ports of the Nth switching subunit in the redundant rack.
[0033] When the total number of multi-card racks is less than the number of first ports contained in a primary switching subunit in a redundant rack, the first ports contained in all primary switching subunits in the redundant rack are divided into port groups with the same number as the total number of multi-card racks. The second port of the Nth switching subunit in each multi-card rack is sequentially connected to multiple first ports of the Nth port group of the redundant rack.
[0034] When the total number of multi-card racks is greater than the number of first ports contained in a primary switching subunit in a redundant rack, connecting the second port of the Nth switching subunit in each multi-card rack group to multiple first ports of the Nth switching subunit in its corresponding switching subunit group means connecting the second port of the Nth switching subunit in the Mth multi-card rack group to the Mth first port of the Nth switching subunit in its corresponding switching subunit group; where M is a positive integer less than or equal to the total number of multiple multi-card racks.
[0035] When the total number of multi-card racks is equal to the number of first ports contained in a primary switching subunit of a redundant rack, connecting the second port of the Nth switching subunit in each multi-card rack to multiple first ports of the Nth switching subunit in the redundant rack in sequence means connecting the second port of the Nth switching subunit in the Mth multi-card rack to the Mth first port of the Nth switching subunit in the redundant rack; where M is a positive integer less than or equal to the total number of multiple multi-card racks.
[0036] When the total number of multi-card racks is less than the number of first ports contained in a primary switching subunit in a redundant rack, connecting the second port of the Nth switching subunit in each multi-card rack to multiple first ports of the Nth port group of the redundant rack in sequence means connecting the second port of the Nth switching subunit in the Mth multi-card rack to the Mth first port of the Nth port group of the redundant rack.
[0037] Taking a multi-card interconnection system comprising eight multi-card racks 1 and one redundant rack 2 as an example, each multi-card rack 1 includes eight server nodes 10 and four primary switching units 20. To achieve connection with the eight multi-card racks 1, the redundant rack 2 may include four primary switching units 20 and one server node 10. Each primary switching unit 20 in the eight multi-card racks 1 is connected to each primary switching unit 20 in the redundant rack 2.
[0038] In this embodiment, the multi-card rack 1 can be a 64-card rack, which can be obtained by merging two 32-card racks, including 8 server nodes 10 and 4 primary switching units 20. The 8 64-card racks form a 512-card interconnection system.
[0039] Figure 2 This application provides a schematic diagram of a 512-card redundant rack configuration, including eight 64-card racks and one redundant rack. Each 64-card rack can be formed by merging two 32-card racks; the topology of the 64-card rack is not detailed here. To connect to the eight 64-card racks, the redundant rack needs to include four Tier 1 switches. A key function of the redundant rack is that when an accelerator in a 64-card rack fails or malfunctions, the accelerator in the redundant rack can take over and continue operating; therefore, the redundant rack includes server nodes. The number of server nodes in the redundant rack can be flexibly configured based on actual needs. Figure 6 The example shown is a redundant rack containing one server node.
[0040] Figure 2 In this system, the eight 64-card racks are designated as supernodes 00 to 07. Each supernode contains 64 accelerator cards, and the eight supernodes contain 8 * 64 = 512 accelerator cards. Therefore, the eight 64-card racks form a 512-card interconnection system.
[0041] Because of the redundant rack design, the second port in each primary switching unit is utilized. Each primary switching unit 20 in each multi-card rack contains two sets of ports; the first set of ports is used to connect to the transmission card 103 in each multi-card rack; the second set of ports is used to connect to the primary switching unit 20 in the redundant rack.
[0042] Each primary switching unit 20 in the redundant rack contains two sets of ports; the first set of ports is used to connect to the primary switching unit 20 in each multi-card rack; the second set of ports is used to connect to the transmission card 103 in the server node 10 of the redundant rack.
[0043] Each multi-card rack 1 contains four primary switching units, which can be referred to as the first primary switching unit, the second primary switching unit, the third primary switching unit, and the fourth primary switching unit, respectively. Each primary switching unit contains two primary switching sub-units, which can be referred to as the first switching sub-unit and the second switching sub-unit.
[0044] For each multi-card rack 1, the second port on its primary switching unit is connected to the first port on the primary switching unit in the redundant rack 2. The second port of the first switching subunit in the first level switching unit of the eight multi-card racks 1 is sequentially connected to the eight first ports of the first switching subunit in the first level switching unit of the redundant rack 2; the second port of the second switching subunit in the first level switching unit of the eight multi-card racks 1 is sequentially connected to the eight first ports of the second switching subunit in the first level switching unit of the redundant rack 2; the second port of the first switching subunit in the third level switching unit of the eight multi-card racks 1 is sequentially connected to the eight first ports of the first switching subunit in the third level switching unit of the redundant rack 2; the second port of the second switching subunit in the third level switching unit of the eight multi-card racks 1 is sequentially connected to the eight first ports of the second switching subunit in the third level switching unit of the redundant rack 2; the second port of the first switching subunit in the fourth level switching unit of the eight multi-card racks 1 is sequentially connected to the eight first ports of the first switching subunit in the fourth level switching unit of the redundant rack 2; the second port of the second switching subunit in the fourth level switching unit of the eight multi-card racks 1 is sequentially connected to the eight first ports of the second switching subunit in the fourth level switching unit of the redundant rack 2.
[0045] It should be noted that "the second port is connected to the first port in sequence" means that one second port is connected to one first port. Taking the example of the second port of the first switching subunit in the first level switching unit of the eight multi-card racks 1 being connected to the eight first ports of the first switching subunit in the first level switching unit of the redundant rack 2 in sequence, the sequential connection means that the second port of the first switching subunit in the first level switching unit of the first multi-card rack 1 is connected to the first first port of the first switching subunit in the first level switching unit of the redundant rack 2; the second port of the first switching subunit in the first level switching unit of the second multi-card rack 1 is connected to the second first port of the first switching subunit in the first level switching unit of the redundant rack 2; the second port of the first switching subunit in the first level switching unit of the third multi-card rack 1 is connected to the third first port of the first switching subunit in the first level switching unit of the redundant rack 2; and so on. The second port is connected to the fourth first port of the first switching subunit in the first primary switching unit of the first level in redundant rack 2; the second port of the first switching subunit in the first primary switching unit of the fifth multi-card rack 1 is connected to the fifth first port of the first switching subunit in the first primary switching unit of the first level in redundant rack 2; the second port of the first switching subunit in the first primary switching unit of the sixth multi-card rack 1 is connected to the sixth first port of the first switching subunit in the first primary switching unit of the first level in redundant rack 2; the second port of the first switching subunit in the first primary switching unit of the seventh multi-card rack 1 is connected to the seventh first port of the first switching subunit in the first primary switching unit of the first level in redundant rack 2; and the second port of the first switching subunit in the first primary switching unit of the eighth multi-card rack 1 is connected to the eighth first port of the first switching subunit in the first primary switching unit of the first level in redundant rack 2.
[0046] Each primary switching unit 20 in the eight multi-card racks 1 is connected to a corresponding primary switching unit 20 in the redundant rack 2; the first switching subunit of the first primary switching unit in the eight multi-card racks 1 is sequentially connected to the eight ports of the first switching subunit of the first primary switching unit in the redundant rack 2; the second switching subunit of the first primary switching unit in the eight multi-card racks 1 is sequentially connected to the eight ports of the second switching subunit of the first primary switching unit in the redundant rack 2; the first switching subunit of the second primary switching unit in the eight multi-card racks 1 is sequentially connected to the eight ports of the first switching subunit of the second primary switching unit in the redundant rack 2; the second switching subunit of the second primary switching unit in the eight multi-card racks 1 is sequentially connected to the eight ports of the first switching subunit of the second primary switching unit in the redundant rack 2; the second switching subunit of the second primary switching unit in the eight multi-card racks 1 is sequentially connected to the eight ports of the second primary switching unit in the redundant rack 2. The eight ports of the second switching subunit of the primary switching unit are connected; the first switching subunit of the third primary switching unit of the eight multi-card racks 1 are sequentially connected to the eight ports of the first switching subunit of the third primary switching unit in the redundant rack 2; the second switching subunit of the third primary switching unit of the eight multi-card racks 1 are sequentially connected to the eight ports of the second switching subunit of the third primary switching unit in the redundant rack 2; the first switching subunit of the fourth primary switching unit of the eight multi-card racks 1 are sequentially connected to the eight ports of the first switching subunit of the fourth primary switching unit in the redundant rack 2; the second switching subunit of the fourth primary switching unit of the eight multi-card racks 1 are sequentially connected to the eight ports of the second switching subunit of the fourth primary switching unit in the redundant rack 2.
[0047] Taking the example of connecting the first switching subunit of the first-level switching unit of the eight multi-card racks 1 to the eight ports of the first switching subunit of the first-level switching unit in the redundant rack 2, the specific connection relationships are as follows: the first switching subunit of the first-level switching unit of the first multi-card rack 1 is connected to the first port of the first switching subunit of the first-level switching unit in the redundant rack 2; the first switching subunit of the first-level switching unit of the second multi-card rack 1 is connected to the second port of the first switching subunit of the first-level switching unit in the redundant rack 2; the first switching subunit of the first-level switching unit of the third multi-card rack 1 is connected to the third port of the first switching subunit of the first-level switching unit in the redundant rack 2; and the first switching subunit of the first-level switching unit of the fourth multi-card rack 1 is connected to the first port of the first switching subunit of the first-level switching unit in the redundant rack 2. The sub-unit is connected to the fourth port of the first switching sub-unit of the first level switching unit in the redundant rack 2; the first switching sub-unit of the first level switching unit of the fifth multi-card rack 1 is connected to the fifth port of the first switching sub-unit of the first level switching unit in the redundant rack 2; the first switching sub-unit of the first level switching unit of the sixth multi-card rack 1 is connected to the sixth port of the first switching sub-unit of the first level switching unit in the redundant rack 2; the first switching sub-unit of the first level switching unit of the seventh multi-card rack 1 is connected to the seventh port of the first switching sub-unit of the first level switching unit in the redundant rack 2; and the first switching sub-unit of the first level switching unit of the eighth multi-card rack 1 is connected to the eighth port of the first switching sub-unit of the first level switching unit in the redundant rack 2.
[0048] Figure 3 This application provides a connection topology diagram for a 512-card redundant rack, comprising eight 64-card racks and one redundant rack. The eight 64-card racks are referred to as supernodes 00 to 07. Each 64-card rack can be obtained by merging two 32-card racks; the topology of the 64-card rack is not detailed here. Each 64-card rack includes four primary switches, each containing two switch components, thus each 64-card rack comprises eight switching subunits, referred to as switch components. To connect to the eight 64-card racks, the redundant rack needs to include four primary switches. Each primary switch contains two switch components, therefore the redundant rack comprises eight switch components, designated as switch components 0 to 7.
[0049] Figure 3In a multi-card rack, each switch component contains 8 ports in its first group, which can be represented by S0 to S7; and 1 port in its second group, which can be represented by S8. In a redundant rack, each switch component contains 8 ports in its first group, which can be represented by S0' to S7'. In a redundant rack, each switch component contains 1 port in its second group, which can be represented by S8'.
[0050] Figure 3 This only shows the interconnection between switch components 0 and 7 in supernodes 00 and 07 and the redundant cabinet. Switch component 0 of supernode 00 is connected to the S0' port of switch component 0 in the redundant cabinet via its S8 port; switch component 7 of supernode 00 is connected to the S7' port of switch component 7 in the redundant cabinet via its S8 port. Switch component 0 of supernode 07 is connected to the S7' port of switch component 0 in the redundant cabinet via its S8 port; switch component 7 of supernode 07 is connected to the S7' port of switch component 7 in the redundant cabinet via its S8 port.
[0051] In practical applications, each supernode's switch component 0 is sequentially connected to ports S0' to S7' of switch component 0 in redundant rack 2. Each supernode's switch component 1 is sequentially connected to ports S0' to S7' of switch component 1 in redundant rack 2. Each supernode's switch component 2 is sequentially connected to ports S0' to S7' of switch component 2 in redundant rack 2. Each supernode's switch component 3 is sequentially connected to ports S0' to S7' of switch component 3 in redundant rack 2. Each supernode's switch component 4 is sequentially connected to ports S0' to S7' of switch component 4 in redundant rack 2. Each supernode's switch component 5 is sequentially connected to ports S0' to S7' of switch component 5 in redundant rack 2. Each supernode's switch component 6 is sequentially connected to ports S0' to S7' of switch component 6 in redundant rack 2. Each supernode's switch component 7 is sequentially connected to ports S0' to S7' of switch component 7 in redundant rack 2.
[0052] Taking the sequential connection of each supernode's switch component 0 to the S0' to S7' ports of the switch component 0 in the redundant rack 2 as an example, the specific connection relationships are as follows: the switch component 0 of supernode 00 is connected to the S0' port of the switch component 0 in the redundant rack 2; the switch component 0 of supernode 01 is connected to the S1' port of the switch component 0 in the redundant rack 2; the switch component 0 of supernode 02 is connected to the S2' port of the switch component 0 in the redundant rack 2; the switch component 0 of supernode 03 is connected to the S3' port of the switch component 0 in the redundant rack 2; the switch component 0 of supernode 04 is connected to the S4' port of the switch component 0 in the redundant rack 2; the switch component 0 of supernode 05 is connected to the S5' port of the switch component 0 in the redundant rack 2; the switch component 0 of supernode 06 is connected to the S6' port of the switch component 0 in the redundant rack 2; and the switch component 0 of supernode 07 is connected to the S7' port of the switch component 0 in the redundant rack 2.
[0053] The redundant ports of switch components 0 from the eight fully interconnected systems are connected to switch component 0 of a redundant node, the redundant ports of switch components 1 from the eight fully interconnected systems are connected to switch component 1 of a redundant node, and so on up to switch component 7. Then, the redundant ports of the eight switch components in redundant rack 2 are connected to the redundant server nodes, corresponding to the eight OAMs. At this point, if any OAM fails, the interconnection topology can be adjusted to redundancy the OAMs of redundant rack 2 to the current system, ensuring task continuity.
[0054] In this embodiment, the number of primary switching units required in the redundant rack 2 can be determined based on the number of primary switching units in the multi-card rack 1. Furthermore, the number of server nodes included in the redundant rack 2 can be flexibly configured based on actual needs. Deploying the redundant rack 2 significantly improves the reliability of the multi-card interconnect system. When an accelerator in the multi-card interconnect system fails or malfunctions, since the accelerators in the redundant rack 2 are connected to the accelerators in the multi-card interconnect system, the working accelerator in the redundant rack 2 can take over the work of the failed accelerator simply by switching the topology mapping, ensuring the continuity of system tasks.
[0055] To facilitate the management of multi-card interconnection systems, an out-of-band management platform can also be set up.
[0056] The out-of-band management platform is connected to eight multi-card racks 1 and a redundant rack 2 respectively, and is used to detect the status of each accelerator 101 in the eight multi-card racks 1. In the event of a failure of the target accelerator 101, the interconnection topology is adjusted to use the accelerator 101 in the redundant rack 2 to replace the target accelerator 101.
[0057] The out-of-band management platform can determine the accelerator 101 used to execute different tasks based on their computation time. In practical applications, the computation time of different tasks can be predicted based on historical data.
[0058] When the out-of-band management platform detects that the utilization rate of the first server node 10 is less than the set lower limit and the utilization rate of the second server node 10 is greater than the set upper limit, it performs a topology switch to migrate the tasks on the second server node 10 to the first server node 10.
[0059] For topology switching, in the specific implementation, the out-of-band management platform can issue topology switching commands to the managers of the eight multi-card racks 1. These topology switching commands contain the port mapping tables of the eight accelerators 101 in the first server node 10. The managers, upon receiving the topology switching commands from the out-of-band management platform, update the in-band topology file according to the commands and execute tasks based on the updated in-band topology file.
[0060] In this embodiment, a dynamic topology load balancing system is developed. By monitoring GPU load in real time, idle GPUs are dynamically added to the computing group when congestion or failure is detected, forming a temporary collaborative topology to achieve task offloading. Combined with a new generation of reconfigurable data center architecture, this is a next-generation load balancing solution for large model clusters at the kilo-card level. It utilizes switching equipment to achieve dynamic reorganization of GPU physical topology, runtime model slicing and state transition technology, and deep collaboration between the communication library and the computing framework.
[0061] Figure 4 This application provides a schematic diagram of a multi-card rack structure, including multiple server nodes 10 and multiple primary switching units 20. Each server node 10 includes multiple accelerators 101, multiple secondary switching units 102, and multiple transmission cards 103. The multiple transmission cards 103 in each server node 10 are respectively inserted into slots of the server node 10. Each transmission card 103 on the multiple server nodes 10 is connected to a port of the multiple primary switching units 20. The number of transmission cards in the same server node 10 is the same as the number of accelerators, the number of secondary switching units is half the number of accelerators, and each secondary switching unit 102 in the same server node 10 has two corresponding transmission cards 103 and two accelerators 101 connected to it.
[0062] Each server node 10 contains multiple accelerators 101 mounted on a first board to enable full connectivity communication; each server node 10 contains multiple secondary switching units 102 mounted on a second board.
[0063] Multiple secondary switching units 102 on the second board are interconnected with multiple accelerators 101 on the first board, and multiple secondary switching units 102 on each server node 10 are connected to each primary switching unit 20 through multiple transmission cards 103, so as to realize the interconnection of accelerators 101 between any two server nodes 10.
[0064] In this embodiment, both the primary switching unit 20 and the secondary switching unit 102 can be optical switching units (Switch) based on the Peripheral Component Interconnect Express (PCIe) standard, i.e., PCIe optical switching units.
[0065] Because the primary switching unit 20 and the secondary switching unit 102 have different functions, especially in terms of the number of ports required, they can use different models of PCIe optical switching units. For example, the primary switching unit 20 can be a PCIe Switch 89144 model, and the secondary switching unit 102 can be a PCIe Switch 89104 model.
[0066] For each server node 10, it contains up to 8 accelerators 101. To connect the accelerators 101 between server nodes 10, 8 transmission cards 103 and 4 secondary switching units 102 can be deployed within each server node. Furthermore, a primary switching unit 20 is deployed outside the server node 10, connecting to the transmission cards 103 within each server node 10, thereby establishing connection paths between all accelerators 101.
[0067] In practical applications, the transmission card 103 can be a card suitable for signal relay in a server (retimer). The accelerator 101 can be a GPU. In this embodiment, the GPU can be presented as an accelerator module (OAM), where one OAM represents one GPU, that is, one GPU can be called an OAM module.
[0068] To enable the connection of multiple accelerators 101 within the same server node 10, the eight accelerators 101 contained in each server node 10 can be mounted on the first board, and the eight accelerators 101 can achieve full connectivity communication through seven sets of high-speed communication links.
[0069] The first board can use a Universal Backplane Board (UBB) to support multiple GPUs, such as eight OAM modules, operating in various cabling and interconnect topologies.
[0070] This universal backplane is a backplane that can accommodate GPU modules. By mounting GPUs, it forms a complete GPU platform, enabling direct connection of GPU acceleration modules and providing a channel for high-speed data transmission and exchange. This facilitates multi-GPU collaboration to handle computationally intensive tasks in the field of Artificial Intelligence (AI), such as image recognition and machine learning. In some embodiments, this backplane can also provide power management and heat dissipation support for modules such as GPUs, ensuring the stable operation of these modules.
[0071] Figure 5 This illustration shows an interconnection of eight accelerators, designated as accelerators 0 to 7, according to an embodiment of this application. Each accelerator contains seven ports, labeled 1 to 7. For each accelerator, each port is connected to a port of the remaining accelerator. For example, port 1 of accelerator 0 is connected to port 1 of accelerator 4; port 2 of accelerator 0 is connected to port 2 of accelerator 7; port 3 of accelerator 0 is connected to port 3 of accelerator 3; port 4 of accelerator 0 is connected to port 6 of accelerator 1; port 5 of accelerator 0 is connected to port 5 of accelerator 6; port 6 of accelerator 0 is connected to port 6 of accelerator 2; and port 7 of accelerator 0 is connected to port 7 of accelerator 5. (Reference) Figure 5 The connection method shown allows for full-connected (FC) communication between the eight accelerators.
[0072] The FC topology is fully connected, with each link supporting up to 16 channels (X16), eliminating the need for horizontal expansion with additional links or ports. Figure 5 The link between the accelerators is PCIe x16.
[0073] To enable connectivity with the eight accelerators 101, each server node 10 includes four secondary switching units 102 mounted on a second board. The second board can be a switch board (SW_B).
[0074] The uplink ports of the four secondary switching units 102 are connected to the motherboard (MB) on the server node 10 via connectors; multiple central processing units (CPUs) can be deployed on the motherboard of the server node 10 to perform system task processing.
[0075] The first downlink port group of the four secondary switching units 102 is connected to the eight accelerators 101 on the first board; the second downlink port group of the four secondary switching units 102 is connected to the slot of the server node 10 where it is located to connect to the eight transmission cards 103.
[0076] A server node 10 may contain 8 accelerators 101 and 4 secondary switching units 102. For ease of distinction, the 8 accelerators 101 may be referred to as the first accelerator to the eighth accelerator, and the 4 secondary switching units 102 may be referred to as the first secondary switching unit, the second secondary switching unit, the third secondary switching unit, and the fourth secondary switching unit.
[0077] On each server node 10, the first and second accelerators are connected to the first secondary switching unit, the third and fourth accelerators are connected to the second secondary switching unit, the fifth and sixth accelerators are connected to the third secondary switching unit, and the seventh and eighth accelerators are connected to the fourth secondary switching unit.
[0078] Figure 6 This application provides a hardware interconnection topology diagram for a single server node. Figure 6 The first board has eight accelerators deployed, the second board has four secondary switches deployed, and the motherboard has two central processing units (CPUs). Each secondary switch is connected to the CPU on the motherboard via an uplink port group, to the two accelerators via a first downlink port group, and to the two transmission cards via a second downlink port group.
[0079] In practical applications, the three key boards—the first board, the second board, and the motherboard—along with the modules deployed on them, can be placed inside a 6U server chassis, and the boards are interconnected via high-speed cables (MCIO).
[0080] The retimer card is inserted into the server. The retimer card's gold fingers connect to the server's PCIe x16 slot. The PCIe x16 slot from the gold fingers is converted into two x8 slots by the retimer card, which are connected to two Quad Small Form-factor Pluggable Double Density (QSFP-DD) interfaces respectively. This is used to realize cross-node interconnection between the secondary switching unit and the primary switching unit within the server node.
[0081] The retimer card is mounted on the server node to enable cross-node interconnection between the GPU and the switching unit, compatible with both copper and fiber optic connections. Each retimer card exposes two QSFP-DD interfaces, supporting both optical and electrical interconnects.
[0082] The four secondary switching units deployed on the second board can be PCIe Switch 89104 model PCIe optical switching units, supporting uplink 8 x16 groups. Considering the layout space, four of them are interconnected with the motherboard using four x16 MCIO connectors, and the other four are interconnected with the motherboard using eight x8 MCIO connectors.
[0083] Downlink support includes 16 x16 slots and 4 x8 slots. The 4 x8 slots connect each PEX89104 chip to a Non-Volatile Memory Express (NVMe) backplane via an MCIO connector, connecting to 8 NVMe hard drives. Eight of the 16 x16 slots are interconnected with the middle backplane via a high-speed backplane (Examax) connector and then connected to the UBB, interconnecting with 8 OAM modules. The remaining 8 x16 slots connect to the board's PCIe x16 slots, supporting 8 retimer cards.
[0084] One Complex Programmable Logic Device (CPLD) can be placed on the second board for timing control. Two clock buffers are also placed there for the board's clock output.
[0085] In this embodiment, the server node 10 can be designed based on a two-layer architecture. The entire server node is divided into five parts: ① Power Supply Unit (PSU), which can be a 12V / 54V PSU module; ② Motherboard module and secondary switching unit module; ③ 4U OAM module; ④ Power supply backplane + cable backplane; ⑤ Heat dissipation system.
[0086] The front panel of server node 10 supports 24 2.5-inch hard drives. It includes an input / output (I / O) panel with a power button, fault indicator lights, a Video Graphics Array (VGA) port, a Universal Serial Bus (USB) port, and a network port. The rear panel of server node 10 consists of a fan wall, a 12V / 54V PSU module, eight full-height, half-length PCIe cards (retimer cards), four solid-state drives (SSDs), and one network interface card (Oracle Certified Partner, OCP).
[0087] In the embodiments of this application, each primary switching unit 20 may include a Controller Carrier Board (CCB) and two switching boards; wherein, each switching board is equipped with a switching sub-unit (PCIe Switch chip); that is, a primary switching unit 20 contains two switching sub-units.
[0088] Each primary switching unit 20 includes two sets of ports; the first set of ports is used to connect to multiple transmission cards 103 on each server node 10; the second set of ports serves as a second port.
[0089] Each switching subunit has eight ports in its first group, designated S0 to S7, and one second port, designated S8. Each port corresponds to two QSFP-DD interfaces. Therefore, each primary switching unit 20 can support 36 PCIe x8 QSFP-DD interfaces.
[0090] A controller is deployed on the baseboard control board.
[0091] The controller is used to detect the working status of each server node 10; when the target server node fails, it performs a topology switch to migrate the work of the target server node to the non-faulty server node 10.
[0092] The controller can be a Baseboard Management Controller (BMC).
[0093] In practical applications, to ensure the orderly management of multi-card interconnection and normal system operation, two BMCs can be set up. One BMC is responsible for the normal operation management of the system, and the other BMC is responsible for the multi-card interconnection management. The BMC can interact with the CPLD and CPU to realize the system management of server nodes.
[0094] Each high-performance Tier 1 switching unit includes one baseboard control board, which can be equipped with an onboard CPU, two BMCs and a CPLD to provide high-performance PCIe switching unit management and ultra-expandable server node system management.
[0095] The primary switching unit chassis includes an I / O board that provides external interfaces, including one power button, one UID button, four indicator lights, and multiple ports. The four indicator lights are for system status, BMC status, power status, and fan status.
[0096] As can be seen from the above technical solution, the multi-card interconnection system includes multiple server nodes and multiple primary switching units; each server node contains multiple accelerators, multiple secondary switching units, and multiple transmission cards; the multiple transmission cards in each server node are respectively inserted into the slots of the server node; each transmission card on multiple server nodes is connected to one port of multiple primary switching units; the number of transmission cards in the same server node is the same as the number of accelerators, the number of secondary switching units is half the number of accelerators, and each secondary switching unit in the same server node has two corresponding transmission cards and two accelerators connected to it. To achieve the connection of all accelerators in the same server node, the multiple accelerators in each server node can be mounted on a first board to achieve full connectivity communication. The multiple secondary switching units in each server node are mounted on a second board; the multiple secondary switching units on the second board are interconnected with the multiple accelerators on the first board, and the multiple secondary switching units on each server node are connected to each primary switching unit through multiple transmission cards to achieve the connection of accelerators between any two server nodes. In this technical solution, by mounting all accelerators in the same server node on the first board, the connection of all accelerators in the same server node can be achieved. Based on the combination of multi-level switching units and transmission cards, and the set connection method, the connection between server nodes can be realized, thereby ensuring the efficient interconnection of all accelerators within multiple server nodes and effectively expanding the interconnection transmission bandwidth.
[0097] Based on the multi-card interconnection system provided in this application, different bandwidth requirements can be met by adjusting the number of server nodes 10 and primary switching units 20 included in the multi-card interconnection system. The following will take a 32-card interconnection system as an example for further explanation.
[0098] The 32-card interconnect system may include four server nodes 10 and two primary switching units 20; each server node 10 includes eight accelerators 101, four secondary switching units 102 and eight transmission cards 103.
[0099] Since the first-level switching unit 20 contains a second port, the total number of ports contained in the two first-level switching units 20 is greater than or equal to the total number of all transmission cards 103 in the four server nodes 10. Each transmission card 103 of each server node 10 is connected to a corresponding port on the two first-level switching units 20.
[0100] For ease of distinction, the two primary switching units can be referred to as the first primary switching unit and the second primary switching unit, respectively. Each primary switching unit contains two switching sub-units, which can be referred to as the first switching sub-unit and the second switching sub-unit.
[0101] Each transmission card 103 on the four server nodes 10 is connected to a port of the primary switching unit 20.
[0102] The first four ports of the first switching subunit of the first-level switching unit are connected in sequence to the first transmission cards on the four server nodes 10; the last four ports of the first switching subunit of the first-level switching unit are connected in sequence to the second transmission cards on the four server nodes 10; the first four ports of the second switching subunit of the first-level switching unit are connected in sequence to the third transmission cards on the four server nodes 10; and the last four ports of the second switching subunit of the first-level switching unit are connected in sequence to the fourth transmission cards on the four server nodes 10.
[0103] The first four ports of the first switching subunit of the second-level switching unit are connected in sequence to the fifth transmission cards on the four server nodes 10; the last four ports of the first switching subunit of the second-level switching unit are connected in sequence to the sixth transmission cards on the four server nodes 10; the first four ports of the second switching subunit of the second-level switching unit are connected in sequence to the seventh transmission cards on the four server nodes 10; and the last four ports of the second switching subunit of the second-level switching unit are connected in sequence to the eighth transmission cards on the four server nodes 10.
[0104] Figure 7 This application provides a topology diagram of a 32-card interconnect system, which includes four server nodes and two primary switches. Each server node contains eight accelerators, four secondary switches, and eight transmission cards. The eight transmission cards are plugged into eight slots on the server node. Figure 7 The connection relationships of the eight accelerators in each server node are not shown in the diagram. For details on the connection relationships of the eight accelerators, please refer to [link to documentation]. Figure 5 Introduction.
[0105] Each primary switch contains two switch components, so the two switches together contain a total of four switch components. Figure 7 The code uses switch components 0, 1, 2, and 3 to represent the switches. Switch components 0 and 1 belong to the same level of switch; switch components 2 and 3 belong to the same level of switch.
[0106] exist Figure 7The four secondary switches in the first server node are designated SW00, SW01, SW02, and SW03; the four secondary switches in the second server node are designated SW10, SW11, SW12, and SW13; the four secondary switches in the third server node are designated SW20, SW21, SW22, and SW23; and the four secondary switches in the fourth server node are designated SW30, SW31, SW32, and SW33. The eight accelerators in each server node are designated G0 through G7. The eight transmission cards in each server node are designated R0 through R7. Each transmission card connects to a QSFP-DD optical module and fiber optic cable to the primary switch.
[0107] The primary switch has a total of 36 PCIe x8 QSFP-DD interfaces. Every two QSFP-DD interfaces form a x16 channel, numbered from S0 to S8 from left to right. Among them, S0 to S7 are used for external expansion of the 32-card interconnect system, and port S8 is used to build a redundant rack 2 system. Figure 7 Only ports S0 to S7, which connect to the transmission card, are shown. Port S8 is not included. Figure 7 As shown in the image.
[0108] Taking the first server node as an example, Figure 7 Connect R0 and R1 to switch component 0, with R0 connected to the first port of switch component 0 and R1 connected to the fifth port of switch component 0; connect R2 and R3 to switch component 1, with R2 connected to the first port of switch component 1 and R3 connected to the fifth port of switch component 1; connect R4 and R5 to switch component 2, with R4 connected to the first port of switch component 2 and R5 connected to the fifth port of switch component 2; connect R6 and R7 to switch component 3, with R6 connected to the first port of switch component 3 and R7 connected to the fifth port of switch component 3. Continue this process to connect all transmission cards to the corresponding ports on the primary switch.
[0109] For a 32-card interconnect system, four server nodes 10 and two primary switching units 20 can be deployed in a single rack.
[0110] In practical applications, the internal vertical space of a single rack can be divided into the lower rack space, the middle rack space, and the upper rack space. The first server node 10 and the second server node 10 are deployed in the lower rack space of the single rack; the two primary switching units 20 are deployed in the middle rack space of the single rack; and the third server node 10 and the fourth server node 10 are deployed in the upper rack space of the single rack.
[0111] Table 1 shows the configuration list for a single 32-card rack.
[0112]
[0113] For the 32-card interconnect system, the power cables are routed along the side of the cabinet, and the high-speed cables are routed on the outside of the side. No cable management rack is designed, and the cables are directly bundled to the side wall of the cabinet using cable ties.
[0114] In practical applications, the rack interconnection methods can be flexibly combined, with direct copper cable connection inside the rack and copper cable or fiber optic interconnection between racks flexibly selected according to the site layout.
[0115] In this embodiment, the 32-card interconnected server system consists of four high-performance server nodes and two high-performance primary switching units connected by general-purpose external optical / electrical cables. The server nodes utilize the product topology architecture design provided in this application, supporting eight accelerators within a 6U space. Any two accelerators can directly interact point-to-point (P2P), achieving a maximum bidirectional connection bandwidth of 896GB / s. The high-performance primary switches provide up to 36 QSFP-DD ports within a 1U space, supporting the PCIe Gen5 protocol, with a single-machine aggregated bandwidth of up to 2048GB / s. Furthermore, collaborative management software can be integrated within the server system to enable cross-host domain GPU expansion and optimal routing planning. It supports global topology information acquisition for the ultra-expanded system, power-on / off timing control, fault monitoring, traffic monitoring, and topology switching.
[0116] Each primary switching unit 20 can contain 9 sets of ports, and one primary switching unit 20 can interact with up to 9 server nodes. Therefore, a 72-card interconnection system can be built based on 9 server nodes 10 and 4 primary switching units 20.
[0117] The 72-card interconnect system includes nine server nodes 10 and four primary switching units 20: each server node 10 contains eight accelerators 101, four secondary switching units 102 and eight transmission cards 103.
[0118] For ease of distinction, the four primary switching units 20 can be referred to as the first primary switching unit, the second primary switching unit, the third primary switching unit, and the fourth primary switching unit, respectively. The eight transmission cards can be referred to as the first transmission card, the second transmission card, the third transmission card, the fourth transmission card, the fifth transmission card, the sixth transmission card, the seventh transmission card, and the eighth transmission card, respectively.
[0119] The total number of ports contained in the four primary switching units 20 is equal to the total number of all transmission cards 103 in the nine server nodes 10. Each transmission card 103 of each server node 10 is connected to a corresponding port on the eight primary switching units 20.
[0120] Each transmission card 103 on the nine server nodes 10 is connected to one port of the first-level switching unit 20; wherein, the first transmission card on the nine server nodes 10 is sequentially connected to the nine ports of the first switching subunit of the first first-level switching unit; the second transmission card on the nine server nodes 10 is sequentially connected to the nine ports of the second switching subunit of the first first-level switching unit; the third transmission card on the nine server nodes 10 is sequentially connected to the nine ports of the first switching subunit of the second first-level switching unit; the fourth transmission card on the nine server nodes 10 is sequentially connected to the nine ports of the second switching subunit of the second first-level switching unit; the fifth transmission card on the nine server nodes 10 is sequentially connected to the nine ports of the first switching subunit of the third first-level switching unit; the sixth transmission card on the nine server nodes 10 is sequentially connected to the nine ports of the second switching subunit of the third first-level switching unit; the seventh transmission card on the nine server nodes 10 is sequentially connected to the nine ports of the first switching subunit of the fourth first-level switching unit; and the eighth transmission card on the nine server nodes 10 is sequentially connected to the nine ports of the second switching subunit of the fourth first-level switching unit.
[0121] Figure 8 This application provides a topology diagram of a 72-card interconnection system. The 72-card interconnection system includes 9 server nodes and 4 primary switches. Each primary switch contains 2 switch components, therefore the 72-card interconnection system contains a total of 8 switch components, namely switch component 0, switch component 1, switch component 2, switch component 3, switch component 4, switch component 5, switch component 6, and switch component 7. The 9 server nodes are represented as server node 0 to server node 8. Each server node contains 8 accelerators, 4 secondary switches, and 8 transmission cards. To ensure clear connection relationships, Figure 8 This only shows the first and last transmission cards in each server node, and the connections between the first and last transmission cards in each server node and switch component 0 and switch component 7, respectively. For the connections between the 8 accelerators, 4 secondary switches, and 8 transmission cards in each server node, please refer to [link to documentation]. Figure 6 and Figure 7 The introduction, in Figure 8 It will no longer be displayed.
[0122] Each server node contains 8 transmission cards, which can be represented by retimer0 to retimer7. The 72-card system contains a total of 8 switch components, which can be represented by Switch0 to Switch7, or simply SW0 to SW7. Each switch component contains 9 ports, which are represented by S0 to S8.
[0123] For a 72-card interconnect system, the retimer0 of each server node is connected to ports S0 to S8 of Switch0 on the primary switch; the retimer1 of each server node is connected to ports S0-S8 of Switch1 on the primary switch; the retimer2 of each server node is connected to ports S0-S8 of Switch2 on the primary switch; the retimer3 of each server node is connected to ports S0-S8 of Switch3 on the primary switch; the retimer4 of each server node is connected to ports S0-S8 of Switch4 on the primary switch; the retimer5 of each server node is connected to ports S0-S8 of Switch5 on the primary switch; and the retimer7 of each server node is connected to ports S0-S8 of Switch7 on the primary switch. This allows for efficient full interconnection of 72 GPU cards between nodes via the primary switch, i.e., a PCIe optical switch.
[0124] The 72-card interconnection system comprises 9 server nodes. Taking the retimer0 of each server node connected to ports S0 to S8 of Switch0 of the primary switch as an example, the specific connection relationships are as follows: retimer0 of server node 0 is connected to port S0 of Switch0 of the primary switch; retimer0 of server node 1 is connected to port S1 of Switch0 of the primary switch; retimer0 of server node 2 is connected to port S2 of Switch0 of the primary switch; retimer0 of server node 3 is connected to port S3 of Switch0 of the primary switch; retimer0 of server node 4 is connected to port S4 of Switch0 of the primary switch; retimer0 of server node 5 is connected to port S5 of Switch0 of the primary switch; retimer0 of server node 6 is connected to port S6 of Switch0 of the primary switch; retimer0 of server node 7 is connected to port S7 of Switch0 of the primary switch; and retimer0 of server node 8 is connected to port S8 of Switch0 of the primary switch.
[0125] Each server node's retimer1 is connected to ports S0-S8 of the primary switch Switch1. The specific connection relationships are as follows: server node 0's retimer1 is connected to port S0 of the primary switch Switch1; server node 1's retimer1 is connected to port S1 of the primary switch Switch1; server node 2's retimer1 is connected to port S2 of the primary switch Switch1; server node 3's retimer1 is connected to port S3 of the primary switch Switch1; server node 4's retimer1 is connected to port S4 of the primary switch Switch1; server node 5's retimer1 is connected to port S5 of the primary switch Switch1; server node 6's retimer1 is connected to port S6 of the primary switch Switch1; server node 7's retimer1 is connected to port S7 of the primary switch Switch1; and server node 8's retimer1 is connected to port S8 of the primary switch Switch1. By following this logic, the specific connection relationships between the transmission card of each server node and each port of the switch component can be determined.
[0126] In this embodiment, based on the topology of the accelerator, secondary switching unit and transmission card in each server node, and the connection relationship between each server node and the primary switching unit, the number of server nodes and primary switching units can be dynamically expanded, thereby constructing a multi-card interconnection system that meets different bandwidth requirements.
[0127] The above provides a detailed description of a multi-card interconnection system provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A multi-card interconnection system, characterized in that, It includes multiple multi-card racks and at least one redundant rack; each multi-card rack includes multiple server nodes and multiple primary switching units; the redundant rack includes at least one server node and multiple primary switching units; each server node includes multiple accelerators and multiple secondary switching units; Each server node is connected to any two accelerators to achieve full connectivity communication between multiple accelerators; each server node contains multiple secondary switching units that are connected to multiple accelerators, and each server node's multiple secondary switching units are connected to each primary switching unit in its respective rack to achieve accelerator connectivity between any two server nodes. Each primary switching unit in the redundant rack is connected to a corresponding primary switching unit in each multi-card rack, thereby enabling the connection between any accelerator in the server node of the redundant rack and any accelerator in the server node of any multi-card rack.
2. The multi-card interconnection system according to claim 1, characterized in that, The multi-card interconnection system includes multiple multi-card racks and one redundant rack; each multi-card rack includes multiple server nodes and multiple primary switching units; the redundant rack includes one server node and multiple primary switching units; each primary switching unit includes two primary switching sub-units. The Nth primary switching subunit in each multi-card rack is sequentially connected to the Nth primary switching subunit in its corresponding redundant rack; where N is a positive integer less than or equal to the total number of primary switching units in the redundant rack.
3. The multi-card interconnection system according to claim 2, characterized in that, Each primary switching unit in each multi-card rack contains two sets of ports; the first set of ports is used to connect to the transmission card in each multi-card rack; the second set of ports is used to connect to the primary switching unit in the redundant rack.
4. The multi-card interconnection system according to claim 3, characterized in that, Each primary switching unit in the redundant rack contains two sets of ports; the first set of ports is used to connect to the primary switching unit in each multi-card rack; the second set of ports is used to connect to the transmission card in the server node of the redundant rack.
5. The multi-card interconnection system according to claim 4, characterized in that, Each Tier 1 switch has a first group of ports consisting of multiple first ports and a second group of ports consisting of one second port. When the total number of multi-card racks is greater than the number of first ports contained in a primary switching subunit in the redundant rack, the multiple multi-card racks are divided into multi-card rack groups with the same number of first ports contained in the primary switching subunit. All primary switching subunits in the redundant rack are divided into switching subunit groups with the same total number of multi-card racks contained in each multi-card rack group. The second port of the Nth switching subunit in each multi-card rack group is sequentially connected to the multiple first ports of the Nth switching subunit in its corresponding switching subunit group. When the total number of multi-card racks is equal to the number of first ports contained in a primary switching subunit of the redundant rack, the second port of the Nth switching subunit in each multi-card rack is sequentially connected to multiple first ports of the Nth switching subunit of the redundant rack. If the total number of multi-card racks is less than the number of first ports contained in a primary switching subunit of the redundant rack, the first ports contained in all primary switching subunits of the redundant rack are divided into port groups with the same number as the total number of multi-card racks, and the second port of the Nth switching subunit in each multi-card rack is sequentially connected to multiple first ports of the Nth port group of the redundant rack.
6. The multi-card interconnection system according to claim 5, characterized in that, If the total number of multi-card racks is greater than the number of first ports contained in a primary switching subunit in the redundant rack, the second port of the Nth switching subunit in the Mth multi-card rack group is connected to the Mth first port of the Nth switching subunit in its corresponding switching subunit group; where M is a positive integer less than or equal to the total number of multiple multi-card racks. When the total number of multi-card racks is equal to the number of first ports contained in a primary switching subunit of the redundant rack, the second port of the Nth switching subunit in the Mth multi-card rack is connected to the Mth first port of the Nth switching subunit of the redundant rack. If the total number of multi-card racks is less than the number of first ports contained in a primary switching subunit of the redundant rack, the second port of the Nth switching subunit in the Mth multi-card rack is connected to the Mth first port of the Nth port group of the redundant rack.
7. The multi-card interconnection system according to claim 2, characterized in that, It also includes an out-of-band management platform; The out-of-band management platform is connected to multiple multi-card racks and the redundant racks respectively, and is used to detect the status of each accelerator in the multiple multi-card racks; in the event of a target accelerator failure, the interconnection topology is adjusted to use the accelerators in the redundant racks to replace the target accelerator.
8. The multi-card interconnection system according to claim 7, characterized in that, The out-of-band management platform is used to determine the accelerators used to execute different tasks based on the computation time of different tasks.
9. The multi-card interconnection system according to claim 7, characterized in that, The out-of-band management platform is used to perform a topology switch to migrate tasks on the second server node to the first server node when it detects that the utilization rate of the first server node is less than a set lower limit and the utilization rate of the second server node is greater than a set upper limit.
10. The multi-card interconnection system according to claim 9, characterized in that, The out-of-band management platform is used to issue topology switching instructions to the managers of multiple multi-card racks; wherein, the topology switching instructions contain port mapping tables of multiple accelerators in the first server node; The manager is used to update the in-band topology file according to the topology switching instruction issued by the out-of-band management platform when it receives the topology switching instruction, and to execute tasks according to the updated in-band topology file.
11. The multi-card interconnection system according to claim 1, characterized in that, Each of the server nodes includes multiple accelerators mounted on a first board, and the multiple accelerators achieve full connectivity communication through multiple sets of high-speed communication links.
12. The multi-card interconnection system according to claim 11, characterized in that, Each of the server nodes includes multiple secondary switching units mounted on a second board, and the uplink port groups of the multiple secondary switching units are connected to the motherboard on the server node through connectors; The first downlink port group of multiple secondary switching units enables connection with multiple accelerators on the first board; The second downlink port group of multiple secondary switching units is connected to the slot of its server node to enable connection with multiple transmission cards.
13. The multi-card interconnection system according to claim 12, characterized in that, Any two accelerators on each server node are connected to a corresponding secondary switching unit.
14. The multi-card interconnection system according to claim 1, characterized in that, Each of the first-level switching units includes a baseboard control board and two switching boards; wherein, each of the switching boards is equipped with a switching subunit; The first set of ports in the primary switching unit of each multi-card rack is used to connect to multiple transmission cards on each of the server nodes.
15. The multi-card interconnection system according to claim 14, characterized in that, A controller is deployed on the baseboard control board; The controller is used to detect the working status of each of the server nodes; when a target server node fails, it performs a topology switch to migrate the work of the target server node to a non-faulty server node.
16. The multi-card interconnection system according to claim 1, characterized in that, The multi-card interconnection system includes four server nodes and two primary switching units; each server node contains eight accelerators, four secondary switching units, and eight transmission cards; The total number of ports in the two primary switching units is greater than or equal to the total number of transmission cards in the four server nodes, and each transmission card of each server node is connected to a corresponding port on the two primary switching units.
17. The multi-card interconnection system according to claim 16, characterized in that, Each transmission card on the four server nodes is connected to a port of the primary switching unit; The two primary switching units are the first-level switching unit and the second-level switching unit; The first four ports of the first switching subunit of the first level switching unit are connected in sequence to the first transmission cards on the four server nodes; the last four ports of the first switching subunit of the first level switching unit are connected in sequence to the second transmission cards on the four server nodes; the first four ports of the second switching subunit of the first level switching unit are connected in sequence to the third transmission cards on the four server nodes; and the last four ports of the second switching subunit of the first level switching unit are connected in sequence to the fourth transmission cards on the four server nodes. The first four ports of the first switching subunit of the second-level switching unit are connected sequentially to the fifth transmission cards on the four server nodes; The last four ports of the first switching subunit of the second-level switching unit are connected in sequence to the sixth transmission card on the four server nodes; the first four ports of the second switching subunit of the second-level switching unit are connected in sequence to the seventh transmission card on the four server nodes; and the last four ports of the second switching subunit of the second-level switching unit are connected in sequence to the eighth transmission card on the four server nodes.
18. The multi-card interconnection system according to claim 16, characterized in that, Four server nodes and two primary switching units are deployed in a single rack; the internal vertical space of the single rack is divided into a lower rack space, a middle rack space, and an upper rack space; The first and second server nodes are deployed in the lower part of the single rack; two primary switching units are deployed in the middle part of the single rack. The third and fourth server nodes are deployed in the upper space of the single rack.
19. The multi-card interconnection system according to claim 1, characterized in that, The multi-card interconnection system includes nine server nodes and four primary switching units: each server node contains eight accelerators, four secondary switching units, and eight transmission cards; The total number of ports in the four primary switching units is equal to the total number of transmission cards in the nine server nodes. Each transmission card in each server node is connected to a corresponding port on one of the eight primary switching units.
20. The multi-card interconnection system according to claim 19, characterized in that, Each transmission card on the nine server nodes is connected to one port of the first-level switching unit; specifically, the first transmission card on each of the nine server nodes is connected sequentially to the nine ports of the first switching subunit of the first-level switching unit; the second transmission card on each of the nine server nodes is connected sequentially to the nine ports of the second switching subunit of the first-level switching unit; the third transmission card on each of the nine server nodes is connected sequentially to the nine ports of the first switching subunit of the second-level switching unit; the fourth transmission card on each of the nine server nodes is connected sequentially to the nine ports of the second switching subunit of the second-level switching unit; the fifth transmission card on each of the nine server nodes is connected sequentially to the nine ports of the first switching subunit of the third-level switching unit; the sixth transmission card on each of the nine server nodes is connected sequentially to the nine ports of the second switching subunit of the third-level switching unit; the seventh transmission card on each of the nine server nodes is connected sequentially to the nine ports of the first switching subunit of the fourth-level switching unit; and the eighth transmission card on each of the nine server nodes is connected sequentially to the nine ports of the second switching subunit of the fourth-level switching unit.
Citation Information
Patent Citations
GPU BOX, server, interconnection system and high-speed interconnection and data interaction method
CN120705096A
Server system and construction method for I / O configuration of server system
US20120016971A1