Switching units, computing systems and multi-node systems
By using a two-layer switching component structure and a PCIe protocol rate bridging mechanism, the problem of limited vertical expansion interconnection of computing components is solved, realizing a high-bandwidth, low-latency computing system and improving the utilization of computing resources and system scalability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-03
AI Technical Summary
In existing technologies, the vertical scaling interconnection of computing components is limited, especially in schemes based on proprietary and standard protocols, which suffer from high communication latency, lack of support for memory semantics, and low utilization of computing resources.
A two-layer switching component structure is adopted, consisting of cascaded first-layer and second-layer switching components. Vertical expansion interconnection is achieved through hierarchical switching components and rate bridging mechanisms. The PCIe protocol is used to connect switching components with different rates, thereby constructing a high-bandwidth, low-latency computing system.
It improves the utilization of computing resources and the scale of vertically expanded interconnects, reduces communication latency, and enhances the scalability and computing efficiency of the system.
Smart Images

Figure CN121441865B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer hardware technology, and in particular to switching units, computing systems and multi-node systems. Background Technology
[0002] To meet the needs of computing power expansion and improvement, it is necessary to build a high-bandwidth, low-latency supernode system with tens or even hundreds of cards through vertical scaling interconnect technology. Currently, the vertical scaling interconnect scheme of supernodes is usually designed based on proprietary protocols and switching components. However, due to the closed nature of their protocols, their scope of use is very limited. There are also schemes based on standard protocols in related technologies, but these schemes inherently have high communication latency and the standard protocols themselves do not support memory semantics.
[0003] This shows that the related technologies suffer from the problem of limited vertical scaling interconnection of computing components. Summary of the Invention
[0004] This application provides a switching unit, a computing system, and a multi-node system to at least solve the problem of limited vertical expansion interconnection of computing components in related technologies.
[0005] This application provides a switching unit, comprising: two layers of switching components formed by cascaded first-layer switching components and second-layer switching components; in the two layers of switching components, a switching port of one first-layer switching component is connected to a first switching port of at least two second-layer switching components, wherein the transmission rate of the transmission channel of the first-layer switching component and the first transmission channel of the second-layer switching component are both a first rate, and the first transmission channel of the second-layer switching component is the transmission channel corresponding to the first switching port of the second-layer switching component; a first switching port of one second-layer switching component is connected to a switching port of at least one first-layer switching component, wherein the transmission rate of the second transmission channel of the second-layer switching component is a second rate, and the second transmission channel of the second-layer switching component is the transmission channel corresponding to the second switching port of the second-layer switching component, the first rate and the second rate are different rates, and the second transmission channel of the second-layer switching component is the transmission channel provided externally by the switching unit.
[0006] This application also provides a computing system, including: a switching unit and a computing component. The switching unit includes two layers of switching components formed by cascading first-layer switching components and second-layer switching components. In the two-layer switching components, a switching port of one first-layer switching component is connected to a first switching port of at least two second-layer switching components, and a first switching port of one second-layer switching component is connected to a switching port of at least one first-layer switching component. The transmission rates of the transmission channels of the first-layer switching components and the first transmission channels of the second-layer switching components are both first rates, and the transmission rate of the second transmission channels of the second-layer switching components is a second rate. The first transmission channel of the second-layer switching components is the transmission channel corresponding to the first switching port of the second-layer switching components, and the second transmission channel of the second-layer switching components is the transmission channel corresponding to the second switching port of the second-layer switching components. The first rate and the second rate are different rates. The second transmission channel of the second-layer switching components is the transmission channel provided by the switching unit to the computing component. The communication port of the computing component is connected to the second switching port of the second-layer switching components.
[0007] This application also provides a multi-node system, including: computing nodes and switching nodes. The computing node includes a computing component, and the switching node includes a switching unit. The switching unit includes two layers of switching components formed by cascading first-layer switching components and second-layer switching components. In the two-layer switching components, a switching port of one first-layer switching component is connected to a first switching port of at least two second-layer switching components, and a first switching port of one second-layer switching component is connected to a switching port of at least one first-layer switching component. The transmission rates of the transmission channels of the first-layer switching components and the first transmission channels of the second-layer switching components are both first rates, and the transmission rate of the second transmission channels of the second-layer switching components is a second rate. The first transmission channel of the second-layer switching components is the transmission channel corresponding to the first switching port of the second-layer switching components, and the second transmission channel of the second-layer switching components is the transmission channel corresponding to the second switching port of the second-layer switching components. The first rate and the second rate are different rates. The second transmission channel of the second-layer switching components is the transmission channel provided by the switching unit to the computing component. The communication port of the computing component is connected to the second switching port of the second-layer switching components.
[0008] This application provides a switching unit comprising: two layers of switching components formed by cascaded first-layer switching components and second-layer switching components; in the two layers of switching components, a switching port of one first-layer switching component is connected to a first switching port of at least two second-layer switching components, wherein the transmission rate of the transmission channel of the first-layer switching component and the first transmission channel of the second-layer switching component are both a first rate, and the first transmission channel of the second-layer switching component is the transmission channel corresponding to the first switching port of the second-layer switching component; a first switching port of one second-layer switching component is connected to a switching port of at least one first-layer switching component, wherein the transmission rate of the second transmission channel of the second-layer switching component is a second rate, and the second transmission channel of the second-layer switching component is the transmission channel corresponding to the second switching port of the second-layer switching component, the first rate and the second rate are different rates, and the second transmission channel of the second-layer switching component is the transmission channel provided externally by the switching unit. By introducing hierarchical switching components and a rate bridging mechanism, computing components can be vertically expanded and interconnected based on switching components with different rates, which can increase port density and total bandwidth. Therefore, it can solve the problem of limited vertical expansion interconnection of computing components in related technologies, and achieve the technical effect of increasing the scale of vertical expansion interconnection and thus improving the utilization rate of computing resources. Attached Figure Description
[0009] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a schematic diagram of the structure of a switching unit according to an embodiment of this application.
[0011] Figure 2 This is a schematic diagram of a supernode system according to an embodiment of this application.
[0012] Figure 3 This is a schematic diagram of another switching unit according to an embodiment of this application.
[0013] Figure 4 This is a schematic diagram of the structure of another switching unit according to an embodiment of this application.
[0014] Figure 5 This is a schematic diagram of the structure of another switching unit according to an embodiment of this application.
[0015] Figure 6 This is a schematic diagram of the structure of a computing system according to an embodiment of this application.
[0016] Figure 7This is a schematic diagram of the structure of another computing system according to an embodiment of this application.
[0017] Figure 8 This is a schematic diagram of the structure of another computing system according to an embodiment of this application.
[0018] Figure 9 This is a schematic diagram of the structure of a multi-node system according to an embodiment of this application.
[0019] Figure 10 This is a logical schematic diagram of a computing node according to an embodiment of this application.
[0020] Figure 11 This is a logical schematic diagram of a switching node according to an embodiment of this application.
[0021] Figure 12 This is a schematic diagram of the structure of another multi-node system according to an embodiment of this application.
[0022] Figure 13 This is a schematic diagram of the structure of another multi-node system according to an embodiment of this application.
[0023] Figure 14 This is a schematic diagram of the structure of another multi-node system according to an embodiment of this application. Detailed Implementation
[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0025] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0026] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0027] According to one aspect of an embodiment of this application, a switching unit is provided. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0028] Figure 1 This is a schematic diagram of an optional switching unit according to an embodiment of this application, such as... Figure 1 As shown, the switching unit 101 includes two layers of switching components: a first-layer switching component 102 and a second-layer switching component 103 cascaded together. In the two layers of switching components, a switching port of one first-layer switching component 102 is connected to a first switching port of at least two second-layer switching components 103. The transmission rate of the transmission channel of the first-layer switching component and the first transmission channel of the second-layer switching component are both a first rate, and the first transmission channel of the second-layer switching component is the transmission channel corresponding to the first switching port of the second-layer switching component. A first switching port of one second-layer switching component 103 is connected to a switching port of at least one first-layer switching component 102. The transmission rate of the second transmission channel of the second-layer switching component 103 is a second rate, and the second transmission channel of the second-layer switching component 103 is the transmission channel corresponding to the second switching port of the second-layer switching component 103. The first rate and the second rate are different rates, and the second transmission channel of the second-layer switching component is the transmission channel provided by the switching unit to the outside.
[0029] With the rapid development of large models in deep learning, particularly in tasks such as natural language processing, visual understanding, and speech recognition, this trend has placed new demands on computing infrastructure. While traditional Graphics Processing Units (GPUs) have advantages in parallel computing, the memory capacity of a single GPU is insufficient to handle such massive model parameters. The growth rate of large model parameters far exceeds the increase in GPU memory capacity. This means that to run these large models on GPUs, more GPUs are needed to share storage requirements. Simultaneously, as model complexity increases, model structures are gradually shifting towards sparse designs. This helps reduce computational resource requirements but also increases the need for cross-GPU data transfer. For example, Mixed Experts (MoE) models heavily rely on Tensor Parallelism (TP) and Expert Parallelism (EP) strategies. In this model architecture, different GPUs are responsible for processing different parts of the model or different model instances, requiring frequent data exchange between GPUs to synchronize the updates of model weights and activation values.
[0030] During training, All-to-All communication means that each GPU needs to exchange data with all other GPUs. The All Reduce operation is a common distributed training synchronization step used to aggregate gradient updates from all GPUs and then broadcast the average gradient back to each GPU. These operations not only introduce a large amount of inter-GPU data exchange but also place strict requirements on network bandwidth and latency. Insufficient network bandwidth or excessive latency will severely impact the efficiency of distributed training, leading to a significant increase in computation time and a reduced linearity of computing power scaling; that is, the actual increase in computing power per added GPU will decrease.
[0031] Therefore, constructing a high-speed interconnect domain with high bandwidth and low latency has become a key measure to improve the linearity of computing power expansion. Through scale-up technology, multiple GPUs can be connected into a high-performance supernode system. This system can support non-blocking communication between dozens or even hundreds of GPUs, ensuring efficient and real-time data exchange, thereby significantly improving overall computing power and accelerating the model training process.
[0032] Current vertical scaling interconnect schemes for supernodes are typically designed based on proprietary protocols and switching components, such as... Figure 2 As shown, this is a supernode system that builds a 72-card fully interconnected network using a proprietary Link protocol and a high-capacity proprietary switching component. Although this can achieve extremely high interconnect bandwidth between 72 GPUs, it is unusable by users who do not use this proprietary protocol because its protocol is closed-source.
[0033] Besides the proprietary solutions mentioned above, scale-up interconnects using standard protocols can also be chosen. Standard protocols offer a wide interconnect range and are suitable for GPU interconnects in large-scale data center environments. However, solutions based on standard protocols have some inherent limitations, the most significant being relatively high communication latency, especially when processing a large number of small data packets. Furthermore, standard protocols do not directly support memory semantics, meaning that direct memory access between GPUs or between a GPU and a CPU requires additional software or hardware support. In standard protocol-based solutions, both the CPU and GPU need to participate in the communication flow control, including address resolution and data packet processing, increasing computational complexity and instruction overhead, and reducing overall system performance. The CPU may consume valuable computing resources while handling these communication-related tasks, and the GPU also needs to allocate some computing power to handle indirect data exchange tasks, all of which directly impact the GPU's core computing performance and efficiency.
[0034] This shows that the related technologies suffer from the problem of limited vertical scaling interconnection of computing components.
[0035] To at least partially solve the above-mentioned technical problems, embodiments of this application propose a switching unit, comprising: two layers of switching components formed by cascaded first-layer switching components and second-layer switching components; in the two layers of switching components, a switching port of one first-layer switching component is connected to the first switching ports of at least two second-layer switching components, wherein the transmission rate of the transmission channel of the first-layer switching component and the first transmission channel of the second-layer switching component are both a first rate, and the first transmission channel of the second-layer switching component is the transmission channel corresponding to the first switching port of the second-layer switching component; a first switching port of one second-layer switching component is connected to the switching port of at least one first-layer switching component, wherein the first... The transmission rate of the second transmission channel of the Layer 2 switching component is the second rate. The second transmission channel of the Layer 2 switching component is the transmission channel corresponding to the second switching port of the Layer 2 switching component. The first rate and the second rate are different rates. The second transmission channel of the Layer 2 switching component is the transmission channel provided to the outside by the switching unit. By introducing hierarchical switching components and rate bridging mechanisms, computing components can be vertically expanded and interconnected based on switching components with different rates. This can increase port density and total bandwidth. Therefore, it can solve the problem of limited vertical expansion interconnection of computing components in related technologies, and achieve the technical effect of increasing the scale of vertical expansion interconnection, thereby improving the utilization rate of computing resources.
[0036] Optionally, the data transmission and exchange between the aforementioned switching components can be based on the Peripheral Component Interconnect Express (PCIe) protocol. Here, PCIe is a high-speed serial bus standard widely used in computer hardware to connect processors, various devices on the motherboard, and other system components. The openness and standardization of the PCIe protocol allow it to be supported by multiple GPU manufacturers, allowing GPU devices to connect directly without specific proprietary protocols or interfaces. The PCIe protocol allows devices to directly access system memory without the CPU as an intermediary, which can also reduce CPU intervention during data transmission, reduce communication latency, and improve the data exchange efficiency between GPUs.
[0037] Despite the significant advantages of the PCIe protocol, current PCIe switching components have limitations in terms of lane count and total bandwidth. Lane count in the PCIe protocol refers to the number of physical channels connecting two PCIe devices; each lane provides a certain data transfer capacity. For example, an x4 PCIe link means four lanes, providing higher bandwidth than an x1 link. The lane count of a PCIe switching component is limited; for example, it may only support 144 lanes, meaning it can support a maximum of 144 PCIe links. Each link is typically x1 to x16. Therefore, scale-up systems built on this PCIe switching component (i.e., the PCIe switch) are limited in total bandwidth and interconnect size.
[0038] In this embodiment, by introducing hierarchical switching components and a rate bridging mechanism, the computing components can be vertically extended and interconnected based on switching components with different rates, thereby solving the aforementioned limitation problem.
[0039] Optionally, the aforementioned switching unit may include two layers of switching components formed by cascading first-layer switching components and second-layer switching components; the switching unit may include at least one first-layer switching component and at least two second-layer switching components.
[0040] Optionally, the first-layer switching component (i.e., L1 switch) can be located at the upper layer of the cascaded structure and can be used to construct non-blocking interconnects between second-layer switching components (L2 switches). A switching port of one first-layer switching component can be connected to the first switching ports of at least two second-layer switching components via a transmission channel at a first rate, where the first rate can be the PCIe Gen6 transmission rate. Through the interconnection of the first switching ports of the L1 and L2 switches, communication link aggregation can be achieved, forming a non-blocking communication path between the L2 switches.
[0041] Optionally, each first switching port of the L2 Switch is connected to at least two L1 Switch ports through a first transmission channel, ensuring that data sent from any L2 Switch port can reach any other part of the system without obstruction.
[0042] Optionally, a first switching port of a second-layer switching component can be connected to a switching port of at least one first-layer switching component, wherein the transmission rate of the second transmission channel of the second-layer switching component is a second rate, the second transmission channel of the second-layer switching component is the transmission channel corresponding to the second switching port of the second-layer switching component, the first rate and the second rate are different rates, and the second transmission channel of the second-layer switching component is the transmission channel provided by the switching unit to the outside.
[0043] Optionally, the second-layer switching component can be located at the lower layer of the cascaded structure and can be used to provide a transmission channel to the outside world. For example, it can be connected to a GPU device, and the second-layer switching component can be connected to a peripheral device through a second transmission channel.
[0044] Optionally, the first-layer switching component and the second-layer switching component can be the same switching component or different switching components.
[0045] Optionally, the second switching port can transmit data at a second rate, which may be different from the first rate. The second rate may be determined based on the hardware specifications of the downstream device (such as a GPU) and the supported PCIe version.
[0046] Optionally, the second rate can be the PCIe Gen5 transfer rate or other transfer rates that match the peripheral device.
[0047] Here, "Gen" (Generation) refers to different versions or generations of the PCIe protocol. The PCIe protocol has evolved over time, going through multiple versions, each bringing performance improvements. For example, PCIe 5.0 (the fifth generation of PCIe) offers significant improvements in data transfer rate and bandwidth compared to PCIe 4.0 (the fourth generation of PCIe). A single x16 PCIe 5.0 link can provide a higher data transfer rate than an x16 PCIe 4.0 link, directly impacting the performance of scale-up systems built on PCIe switches in terms of lane count and total bandwidth.
[0048] In related technologies, GPU devices mostly still use the PCIe Gen5 standard, while other transmission rates can be used between PCIe switching components. In large-scale CPU interconnect scenarios, the mismatch in rates may affect the overall processing efficiency of the system.
[0049] In this embodiment, the cascading between L2 Switch and L1 Switch, and the connection between L2 Switch and peripheral devices, can be carried out through transmission channels of different rates. The first switching port of L2 Switch and L1 Switch can ensure a high-speed, non-blocking communication path through the first rate connection, while the second switching port of the second switching component and the second transmission channel of the peripheral device can provide the peripheral device with a stable communication capability that matches its rate.
[0050] According to the embodiments provided in this application, a switching unit includes: two layers of switching components formed by cascading first-layer switching components and second-layer switching components; in the two layers of switching components, a switching port of one first-layer switching component is connected to the first switching ports of at least two second-layer switching components, wherein the transmission channels of the first-layer switching components and the first transmission channels of the second-layer switching components both have a first rate, and the first transmission channel of the second-layer switching components is the transmission channel corresponding to the first switching port of the second-layer switching components; a first switching port of one second-layer switching component is connected to the switching port of at least one first-layer switching component, wherein the transmission rate of the second transmission channel of the second-layer switching component is a second rate, and the second transmission channel of the second-layer switching component is the transmission channel corresponding to the second switching port of the second-layer switching component, the first rate and the second rate are different rates, and the second transmission channel of the second-layer switching component is the transmission channel provided externally by the switching unit. By introducing hierarchical switching components and a rate bridging mechanism, computing components can be vertically expanded and interconnected based on switching components with different rates, which can solve the problem of limited vertical expansion interconnection of computing components in related technologies, thereby achieving the technical effect of increasing the scale of vertical expansion interconnection and thus improving the utilization rate of computing resources.
[0051] In one exemplary embodiment, the ratio of the first rate to the second rate is K; in a second-layer switching component, the ratio of the number of first transmission channels to the number of second transmission channels is L; wherein K and L are both positive numbers greater than 0, and the product of K and L is greater than or equal to 1.
[0052] Alternatively, the sizes of K and L can be predetermined.
[0053] Here, the transmission channel ratio L refers to the ratio of the number of first transmission channels to the number of second transmission channels in a Layer 2 switching component. The L value can be used to reflect the balance of a Layer 2 switching component in terms of uplink (i.e., connection to the Layer 1 switching component) and downlink (i.e., external) connections. Factors affecting the setting of the L value may include, but are not limited to: the K value, the total bandwidth of the L2 switch, the number of ports that can be supported, and the needs of the peripheral devices connected to it. For example, if the L value is small, it means that more ports of the L2 switch are allocated to the second transmission channels, which is beneficial for large-scale interconnection between GPUs. However, at the same time, the communication bandwidth between the L2 switch and the L1 switch may be limited, affecting the overall system communication efficiency. Conversely, if the L value is large, although the communication bandwidth between L1 and L2 is improved, it may limit the number of peripheral devices that can be connected.
[0054] Optionally, the product of K and L can be greater than or equal to 1, so that the uplink bandwidth of the second switching component is not less than the downlink bandwidth, that is, the bandwidth between the first and second switching components is not less than the transmission bandwidth provided by the second switching component, thus avoiding a rate bottleneck.
[0055] Through this embodiment, by reasonably setting the number of transmission channels in the switching unit, communication balance can be achieved to build a high-performance vertically scalable interconnect.
[0056] In one exemplary embodiment, K=2, L=1 / 2; in the two-layer switching components, the number of the second-layer switching components is three times the number of the first-layer switching components.
[0057] Similar to the previous embodiments, K can be the ratio of the first rate to the second rate. K can be 2, that is, the first rate can be twice the second rate. For example, the first rate can be the PCIe Gen6 rate and the second rate can be the PCIe Gen5 rate. Here, the PCIe Gen6 rate is about twice the PCIe Gen5 rate.
[0058] Similar to the previous embodiments, the L value is the ratio of the number of first transmission channels (uplink channels connected to L1Switch) to the number of second transmission channels (downlink channels providing external connections) in a Layer 2 switching component. Optionally, L can be 1 / 2, and the product of K and L can be 1. That is, the total bandwidth of the first transmission channels and the total bandwidth of the second transmission channels in the Layer 2 switching component can be matched, i.e., the uplink bandwidth and downlink bandwidth of the Layer 2 switching component can be matched, thereby avoiding the formation of bottlenecks during data communication that cause data delays and congestion.
[0059] Optionally, the number of second-layer switching components can be three times the number of first-layer switching components to support connection to more devices and achieve large-scale interconnection.
[0060] This embodiment, by comprehensively considering rate matching, bandwidth maximization, and the diversity of communication paths, can improve communication efficiency and system scalability, and provide technical support for vertically scalable high-speed interconnection.
[0061] In one exemplary embodiment, in a two-layer switching unit, the number of transmission channels of one switching unit is X, the total number of second transmission channels of all second-layer switching units is Y, and the total number of all switching units is N, where X, Y, and N are all positive integers greater than or equal to 2, and N ≥ 2Y / X .
[0062] Optionally, in a two-layer switching component, the number of transmission channels of one switching component is X. Here, the switching component can be a first-layer switching component or a second-layer switching component. The first-layer switching component and the second-layer switching component can use the same switching component, and the number of transmission channels of the first-layer switching component and the second-layer switching component can both be X.
[0063] Optionally, N≥ 2Y / X Considering that Y and X are not necessarily integer multiples of each other, rounding up is used here. The mathematical symbol for rounding up is "". The "+" option is used to round up a value to the nearest nearest value. For example, 4.2 =5, -4.2 =-4, thus, each switching unit's transmission channel has a corresponding path to connect to other switching units. To achieve non-blocking communication, each Layer 2 switching component can connect to other Layer 2 switching components in the switching unit through at least two paths (which can be through Layer 1 switching components), so that even if one path is congested, data can still be transmitted through the other paths.
[0064] Optionally, 2Y / X can be an integer multiple relationship, and N can be 2Y / X.
[0065] Through this embodiment, by adapting the transmission channels and the number of switching components, a non-blocking network for processing corresponding data transmission can be constructed, ensuring the feasibility of the construction process and the non-blocking characteristics of the switching units.
[0066] In one exemplary embodiment, the first rate is twice the second rate; in the two-layer switching components, the number of first-layer switching components is... Y / 2X The number of second-layer switching components is 3Y / 2X .
[0067] Similar to the previous embodiments, the first rate can be twice the second rate.
[0068] Optionally, the first rate can be the PCIe Gen6 transmission rate, and the second rate can be the PCIe Gen5 transmission rate.
[0069] Optionally, the uplink and downlink bandwidths of the Layer 2 switching component can be matched, that is, the bandwidth provided by the Layer 2 switching component to the outside world and the bandwidth between the Layer 2 switching component and the Layer 1 switching component can be matched. For the case where the first rate is twice the second rate, the number of transmission channels between the Layer 2 switching component and the Layer 1 switching component can be half the number of transmission channels provided by the Layer 2 switching component to the outside world.
[0070] Optionally, in a two-layer switching component, the number of transmission channels in one switching component can be X, the total number of second transmission channels in all second-layer switching components can be Y, and the total number of first transmission channels in all second-layer switching components can be... Y / 2 That is, the total number of transmission channels of all first-layer switching components can be Y / 2 Therefore, the number of first-layer switching components can be Y / 2X .
[0071] Similarly, the number of second-layer switching components is 3Y / 2X .
[0072] In this embodiment, by matching the number of two-layer switching components with the number of transmission channels, the uplink and downlink bandwidths of the second-layer switching components can be matched, thereby achieving non-blocking interconnection between different second-layer switching components.
[0073] In one exemplary embodiment, in a two-layer switching component, a switching port of a first-layer switching component is connected to at least one first switching port of any second-layer switching component.
[0074] In this embodiment, the switching port of a first-layer switching component is connected to at least one first switching port of any second-layer switching component. That is, each switching port of the L1 Switch is connected to at least one first switching port (the port connected to the upper-layer switching component) of each L2 Switch. The L1 Switch, as an upper-layer component, is used to connect multiple second-layer switching components to achieve non-blocking data transmission.
[0075] In order to achieve non-blocking communication, there should be no bottleneck in the data transmission path between any external device and another device in the switching unit. By ensuring that the L1 switch is connected to at least one port of all L2 switches, even if a port of an L2 switch is congested or fails, data can still flow to or from other L2 switches through other ports of the L1 switch, thereby avoiding dependence on a single path and enhancing the fault tolerance of the system and the reliability of data transmission.
[0076] In this embodiment, by establishing a connection between the first-layer switching component and each second-layer switching component, non-blocking data transmission can be achieved.
[0077] In an exemplary embodiment, in a two-layer switching unit, the number of transmission channels of a switching unit is X, and the number of first-layer switching units is N1, wherein the number of transmission channels provided by a first-layer switching unit to any second-layer switching unit is X / N1, wherein X, N1 and X / N1 are all positive integers greater than or equal to 1.
[0078] Optionally, the number of transmission channels between the second-layer switching component and each first-layer switching component can be the same.
[0079] Optionally, in the two-layer switching components, the number of transmission channels of one switching component is X, the number of first-layer switching components is N1, and the number of transmission channels provided by one first-layer switching component to any second-layer switching component can both be X / N1. That is, each L1 Switch can provide X / N1 transmission channels to any L2 Switch.
[0080] Optionally, the number of Layer 2 switching components can be N2, which can be three times N1. Its uplink can be connected to the Layer 1 switching components through X / N1 transmission channels, and its downlink can provide 2X / N1 transmission channels to the outside.
[0081] In this embodiment, by setting the same number of transmission channels between the second-layer switching component and each first-layer switching component, a balanced allocation of bandwidth can be achieved, enhancing system redundancy, reducing system complexity, and facilitating the construction of non-blocking, high-performance switching units.
[0082] In one exemplary embodiment, in a two-layer switching component, a second transmission channel of a second-layer switching component is provided externally in groups of Z, where Z is a positive integer greater than or equal to 2.
[0083] Optionally, in a two-layer switching component, the second transmission channel of a second-layer switching component can be provided externally in groups of Z. Providing multiple transmission channels as a group can optimize the data transmission rate.
[0084] Optionally, Z can be a positive integer greater than or equal to 2. For example, Z can be 4 or 8. Z can be preset and can be set according to the characteristics of the peripheral device. For example, different GPU devices may support different numbers of PCIe lanes and can be configured according to the actual lane requirements of the GPU.
[0085] In this embodiment, by providing the second transmission channels of a second-layer switching component to the outside in groups of Z, a non-blocking switching unit that meets communication requirements, has good communication redundancy, and optimized performance can be constructed.
[0086] In one exemplary embodiment, the switching port of the switching component in the two-layer switching component is a high-speed peripheral component interconnection port.
[0087] Similar to the previous embodiments, the switching ports of the switching components in the two-layer switching components can be high-speed peripheral component interconnect ports (i.e., PCIe ports). That is, the switching ports of the first-layer switching components, the first switching port and the second switching port of the second-layer switching components can all be PCIe ports. Correspondingly, the switching ports of the switching components in the two-layer switching components all support the PCIe protocol.
[0088] This embodiment enables non-blocking interconnection of large-scale devices by using PCIe ports as switching ports for switching components, providing high-speed communication channels, improving compatibility with external devices, reducing communication latency, and increasing computing efficiency.
[0089] The switching unit in this embodiment is explained below with reference to optional examples. In this embodiment, the switching unit may include two layers of switching components. One of the switching components may be a PCIe switching component, and the switching port of the switching component may be a PCIe port. The number of the second-layer switching components may be three times that of the first-layer switching components. The transmission rate (i.e., the first rate) between the first-layer and second-layer switching components may be the PCIe Gen6 transmission rate, and the transmission rate (i.e., the second rate) of the second transmission channel provided by the second-layer switching component may be the PCIe Gen5 transmission rate. Assuming Y is the number of Gen5 PCIe Lanes provided by the non-blocking switching unit to be constructed, and X is the number of Lanes supported by a single PCIe Gen6 Switch component, then the total number of PCIe switching components required, N, may be N = 2X / Y, where the number of L2 Switches, N2, may be Y / X. 3 / 2, the number of L1Switch N1 can be Y / 2X.
[0090] Optionally, such as Figure 3 As shown, the L1 Switch (i.e., L1 SW) acts as the upper-layer switch in the cascade, used only to link the L2 Switches (i.e., L2 SW) and establish communication links between L2 Switches. The L1 Switch and L2 Switch are connected via PCIe Gen6 links (i.e., Gen6 Lanes). Each L1 Switch provides X / N2 Gen6 PCIe Lanes to each L2 Switch. The L2 Switch acts as the lower-layer switch in the cascade. Its uplink is connected to the L1 Switch via Y / 2N2 PCIe Gen6 Lanes, and its downlink provides PCIe Gen5 links. Each L2 Switch provides Y / N2 downlink PCIe Gen5 Lanes to connect to Gen5 GPU devices. The number of uplink and downlink lanes of the L2 Switch meets the 1:2 ratio, and the uplink bandwidth and downlink bandwidth are matched, thus achieving non-blocking interconnection between different L2 Switches through the L1 Switch.
[0091] like Figure 4As shown, a 288-Lanes non-blocking switching unit can be constructed based on four PCIe Gen6 144Lanes switching components: SW3 acts as the upper-layer switch, and is cascaded with SW0-2 through 12 Gen6 x4 (i.e. G5x4) links. The interconnection link width of SW0-2 and SW3 is x48. SW0-2 acts as the downlink switch and sends out 12 Gen5 x8 links to connect to external GPU devices, for a total of 36 Gen5x8 (i.e. G5x8) links.
[0092] like Figure 5 As shown, a 576-Lanes non-blocking switching unit can be constructed based on 8 PCIe Gen6 144Lanes switching components: SW6 and SW7 are used as upper-layer switches, and are cascaded with SW0-5 through Gen6 x24 (i.e. G6x24) links respectively. SW0-5 is used as a downlink switch to send out 12 Gen5 x8 (i.e. G5x8) links to external GPU devices, for a total of 72 Gen5 x8 links.
[0093] Similarly, a 1152-lane non-blocking switching unit can be built based on 16 PCIe Gen6 144-lane switching components. The structure of the 1152-lane non-blocking switching unit is similar to that of the 576-lane non-blocking switching unit. SW12-SW15 serve as upper-layer switches, and SW0-11 are cascaded with Gen6 x12 links. SW0-11 serve as downlink switches, each sending out 12 Gen5 x8 links to connect to external GPU devices, for a total of 144 Gen5 x8 links.
[0094] Based on the aforementioned large-lane-count non-blocking PCIe switching units, a topology design for large-scale scale-up multi-node systems can be constructed. Compared to other PCIe protocol scaling solutions, the optional example and the aforementioned embodiments based on large-lane-count non-blocking PCIe switching units can construct larger-scale systems and achieve full interconnection between external devices or nodes, with higher interconnection bandwidth, but without requiring too much hardware for switching components. Conversely, for a full interconnection architecture of the same scale, the switching units in the optional example and the aforementioned embodiments can efficiently utilize the capabilities of each switching component, reducing the need for additional switching components and lowering hardware costs.
[0095] Optionally, through the cascading communication routing technology of the switching components, a bandwidth aggregation link from the Gen5 to the Gen6 port can be constructed, enabling the Gen6 PCIe Switch to connect to the Gen5 GPU device. Even if the GPU device is still using the Gen5 PCIe interface, it can take advantage of the higher speed of the Gen6 Switch and provide a high-speed interconnect environment similar to Gen6 for the Gen5 GPU through specific cascading design and bandwidth aggregation.
[0096] This optional example, from a hardware architecture design perspective, provides a switching unit that expands switching ports based on PCIe switching components. It constructs a PCIe switch gearbox based on multiple switching components and uses cascaded communication routing technology to build a Gen5 to Gen6 port bandwidth aggregation link. This achieves a large-lane, non-blocking PCIe switching unit, solving the problem of limited vertical expansion interconnection of computing components in related technologies. It reduces the number of switches used, is compatible with different transmission rates, and achieves the technical effect of increasing the scale of vertical expansion interconnection, thereby improving the utilization of computing resources.
[0097] Embodiments of this application also provide a computing system, including: a switching unit and a computing component. The switching unit includes two layers of switching components formed by cascading first-layer switching components and second-layer switching components. In the two-layer switching components, a switching port of one first-layer switching component is connected to a first switching port of at least two second-layer switching components, and a first switching port of one second-layer switching component is connected to a switching port of at least one first-layer switching component. The transmission rates of the transmission channels of the first-layer switching components and the first transmission channels of the second-layer switching components are both first rates, and the transmission rate of the second transmission channels of the second-layer switching components is a second rate. The first transmission channel of the second-layer switching components is the transmission channel corresponding to the first switching port of the second-layer switching components, and the second transmission channel of the second-layer switching components is the transmission channel corresponding to the second switching port of the second-layer switching components. The first rate and the second rate are different rates. The second transmission channel of the second-layer switching components is the transmission channel provided by the switching unit to the computing component. The communication port of the computing component is connected to the second switching port of the second-layer switching components.
[0098] In this embodiment, a vertically scalable computing system can be constructed based on a switching unit. The switching unit in this embodiment can be the switching unit 101 in the previous embodiment. For the description of the features corresponding to the switching unit in the previous embodiment, please refer to the relevant description of the embodiments corresponding to the switching unit in the previous embodiment. They will not be repeated here.
[0099] Optionally, the computing unit can connect to the second switching port of the second layer switching unit through its communication port. The second transmission channel of the second layer switching unit transmits data at a second rate and can be used to communicate directly with the computing unit. Data transmission within the switching unit can be performed at a first rate, but external communication between the switching unit and the computing unit can be kept within the rate range that the computing unit can support.
[0100] Optionally, the computing unit may include multiple communication ports that can be connected to multiple Layer 2 switching units, supporting vertical scaling of multiple transmission channels. For example, the computing unit may be a GPU that can support southbound PCIe interconnect.
[0101] According to the embodiments provided in this application, a computing system includes a switching unit and a computing component. The switching unit comprises two layers of switching components, formed by cascading first-layer switching components and second-layer switching components. In the two-layer switching components, a switching port of one first-layer switching component is connected to a first switching port of at least two second-layer switching components, and a first switching port of one second-layer switching component is connected to a switching port of at least one first-layer switching component. The transmission rates of the transmission channels of the first-layer switching components and the first transmission channels of the second-layer switching components are both first rates, and the transmission rate of the second transmission channels of the second-layer switching components is a second rate. The first transmission channel of the second-layer switching components corresponds to the transmission channel of the first switching port of the second-layer switching components, and the second transmission channel of the second-layer switching components corresponds to the transmission channel of the second switching port of the second-layer switching components. The first rate and the second rate are different rates. The second transmission channel of the second-layer switching components is a transmission channel provided by the switching unit to the computing component. The communication port of the computing component is connected to the second switching port of the second-layer switching components. This solves the problem of limited vertical expansion interconnection of computing components in related technologies, achieving the technical effect of increasing the scale of vertical expansion interconnection and thus improving the utilization rate of computing resources.
[0102] In one exemplary embodiment, the number of switching units is M, the number of computing units is P, one computing unit is connected to M switching units, and one switching unit is connected to P computing units, wherein M and P are both positive integers greater than or equal to 2.
[0103] Based on the switching unit proposed in the foregoing embodiments, and paired with a computing component that supports southbound PCIe Scale-up expansion (supporting multiple PCIe ports), a Scale-up system, i.e. a computing system, can be constructed.
[0104] Optionally, the number of switching units can be M, and the number of computing units can be P. One computing unit can be connected to M switching units, and one switching unit can be connected to P computing units. Thus, communication between any two computing units in the computing system can have M paths.
[0105] In this embodiment, M switching units enable multipath communication between each computing component and other computing components in the computing system, thereby improving the bandwidth and path diversity of data transmission between computing components.
[0106] In one exemplary embodiment, the computing unit includes a port monitoring unit and a controller: wherein the port monitoring unit is configured to monitor the port traffic of the communication port of the computing unit and report the monitored port traffic of the communication port of the computing unit to the controller; the controller is configured to distribute data to be distributed to the communication ports of the computing unit based on the port traffic of the communication ports of the computing unit, so as to perform load balancing among the communication ports of the computing unit; wherein the two computing units are allowed to communicate through either switching unit.
[0107] In this embodiment, when any two computing units communicate, they are allowed to communicate through any switching unit, and there are M path options.
[0108] Similar to the previous embodiments, the switching port of the switching unit can be a PCIe port, which can follow the PCIe protocol. Since the PCIe protocol does not have a backhaul mechanism, the switching components in the switching unit can only determine the congestion status of their own ports and balance the traffic among internal ports, but cannot balance the traffic to other switching components, nor can they backhaul the congestion status to the computing component. Therefore, there may be a situation in the link where one link is congested while other links are idle. Based on this, this embodiment proposes a port monitoring unit and controller in the computing system to realize non-intrusive, real-time, fine-grained monitoring of key data flows inside the computing unit, dynamically capture key indicators such as throughput and latency of PCIe ports, and automatically allocate data to ports, thereby ensuring that the traffic to each PCIe port is balanced and ensuring that there is no link blockage during data communication.
[0109] Optionally, the port monitoring unit can be deployed inside the computing component to monitor the real-time traffic status of all communication ports. It can be used to detect the current bandwidth usage of each port, including uplink and downlink traffic.
[0110] Optionally, the controller can receive traffic reports from the port monitoring unit and intelligently schedule data distribution based on this information. It can distribute the data to be distributed to the communication ports of the computing unit based on the port traffic of the communication ports of the computing unit to perform load balancing among the communication ports of the computing unit. For example, it can analyze the traffic data of each communication port, evaluate the load difference between ports, and then take appropriate strategies to direct the data to be distributed to the port with lower load to avoid overloading of a single port.
[0111] Optionally, the port monitoring unit can be located in the PCIe register in the computing unit, and adopt an uplink and downlink dual-listening architecture to monitor the uplink and downlink traffic of each PCIe controller respectively.
[0112] Optionally, a global controller may also exist, located above the PCIe controller. It may be a dedicated module inside the computing unit or an independent hardware unit outside the computing unit. It may be connected to the PCIe controller through a global monitoring bus to provide an overall flow view and flow control for all ports. Here, the global monitoring bus is an internal bus inside the computing unit used to connect all port monitoring units and the controller. It is a high-speed transmission bus type suitable for continuously transmitting large amounts of counter data.
[0113] Optionally, such as Figure 6 As shown, the port monitoring unit may include a port control register, uplink / downlink monitoring channels, and port aggregation logic. The port control register receives port control information from the global controller, enabling or disabling ports. Each uplink / downlink monitoring channel contains a Transaction Layer Packet (TLP) detector, a statistical counter, a bandwidth calculator, and a latency meter. The TLP detector captures TLPs and extracts metadata, the statistical counter counts TLPs, the bandwidth calculator calculates bandwidth utilization in real time based on the number of TLPs, and the latency meter measures read transaction latency. The port monitoring unit may also include a port aggregator, which periodically collects data from the two monitoring channels, performs basic aggregations such as obtaining the total traffic of the port, and temporarily stores the data for the central controller to read.
[0114] Optionally, such as Figure 6As shown, the global controller can include a global control register, a global aggregation logic manager, and a global traffic distributor. The global register can be used to configure all port monitoring units and read data, providing an independent configuration register group and data register group for each port. It can support global control of all ports and define the trigger conditions for cross-port traffic distribution. The aggregation logic manager is used to accumulate the bandwidth of all enabled ports in real time, output the total system bandwidth and the contribution percentage of each port, continuously compare the traffic load of each port, and trigger a load imbalance alarm when the load difference exceeds a preset threshold. The traffic distributor can be used to distribute traffic evenly among ports after receiving the load imbalance alarm from the aggregation logic manager, thereby achieving traffic balance among the ports of the computing unit.
[0115] This embodiment solves the problem of congestion control in multi-path communication between multi-port computing components by adding a traffic monitoring design to the computing system, and achieves load balancing among multiple ports.
[0116] In one exemplary embodiment, in a two-layer switching component of a switching unit, a second transmission channel of a second-layer switching component is provided to a computing component in groups of Z; the number of transmission channels supported by a computing component is greater than or equal to M. Z.
[0117] Optionally, the total number of second transmission channels in all second-layer switching components of a switching unit can be Y, so that Y / Z computing components can be interconnected through M switching units. Each switching unit can connect to all Y / Z computing components, and each computing component can be connected to each other through a second transmission channel. Each computing component is connected to M switching units.
[0118] Optionally, the number of transmission channels supported by a computing unit can be greater than or equal to M. Z, meaning the computing component needs to support at least M. Input / output capability of Z transmission channels.
[0119] Optionally, M can be pre-configured. The configuration of M can depend on the number of transmission channel connections supported by the computing unit itself. M can be the number of transmission channel connections supported by the computing unit itself divided by Z. For example, if the GPU itself supports 48 southbound PCIe interconnect lanes, and a second transmission channel of a Layer 2 switching unit provides a computing unit with 8 lanes in a group, then M can be 6, and the communication bandwidth of each path can be 64 GB / s (total communication bandwidth is 64 GB / s). Compared to traditional 8-card servers where there is only one communication link between any two GPUs forming a network, the topology proposed in this embodiment greatly expands the communication bandwidth between GPUs (MB / s).
[0120] For example, such as Figure 7 The diagram illustrates an interconnect topology for a Y / 8-card supernode system built upon Y-lane non-blocking switching units. It is constructed using GPU cards supporting southbound PCIe Scale-up expansion (supporting multiple PCIe ports), and achieves full interconnection of Y / 8 GPU cards through M switching units. The total number of second transmission channels in all Layer 2 switching components within each switching unit can be Y. Each switching unit can connect to all Y / 8 GPU cards, and each GPU card is connected via Gen5x8 lanes. The number of second transmission channels in each Layer 2 switching component can be Y / N2 (N2 being the number of Layer 2 switching components). Each Layer 2 switching component can provide Y / N2 sets of Gen5 links for communication with the GPUs. Simultaneously, each Layer 2 switching component can communicate with Layer 1 switching components via Y / 2N2 sets of Gen6 links (i.e., Gen6 lanes), thus constructing non-blocking extended communication for Y / 8 GPUs. Communication between any two GPUs has M paths, and the communication bandwidth of each path can be 64 GB / s.
[0121] Through this embodiment, by adapting the computing components and switching units, load balancing can be achieved through multi-path communication, preventing communication blockage.
[0122] For example, the computing component is an accelerator card, and the accelerator card supports a high-speed peripheral component interconnection transmission channel.
[0123] Optionally, the computing component can be various types of accelerator cards, including but not limited to GPUs, Application Specific Integrated Circuits (ASICs), and Neural Processing Units (NPUs). Different types of accelerator cards can communicate with the host and other accelerator cards via the PCIe interface.
[0124] For example, such as Figure 8As shown, based on six of the aforementioned 288-lane non-blocking switching units, non-blocking communication between 36 GPUs supporting 48-lane PCIe Gen5 southbound interconnects can be achieved. Each switching unit's Layer 2 switching component can provide 12 Gen5 x8 (G5 x8) links for GPU communication. The Layer 1 and Layer 2 switching components communicate via Gen6 (G6) links, thus enabling a 36-card scale-up system. Furthermore, based on six of the aforementioned 576-lane non-blocking switching units and GPUs supporting 48-lane PCIe Gen5 southbound interconnects, a 72-card scale-up system can also be constructed. The connection topology of the 72-card scale-up system is similar to that of the 36-card scale-up system, except that the 288-lane non-blocking switching units are replaced with 576-lane non-blocking switching units. Based on the six 1152-lane non-blocking switching units mentioned above and the GPUs supporting 48-lane PCIe Gen5 southbound interconnects, a 144-card scaleup system can also be built. The connection topology of the 144-card scaleup system is similar to that of the 36-card scaleup system, except that the 288-lane non-blocking switching units are replaced with 1152-lane non-blocking switching units.
[0125] In an exemplary embodiment, the communication port of the computing unit includes a southbound communication port and a northbound communication port; the southbound communication port of the computing unit is connected to the second switching port of the second layer switching unit of m1 switching units, and the northbound communication port of the computing unit is connected to the second switching port of the second layer switching unit of m2 switching units, where m1 and m2 are both positive integers greater than or equal to 1, and M = m1 + m2.
[0126] Optionally, the communication ports of the computing unit may include southbound communication ports and northbound communication ports, where M can be the number of Layer 2 switching units connected to the southbound communication ports of the computing unit (i.e., m1) and the number of Layer 2 switching units connected to the northbound communication ports (i.e., m2).
[0127] Optionally, based on the aforementioned computing components and switching units, a bidirectional interconnect architecture can be constructed, meaning that the aforementioned computing system can support bidirectional interconnect.
[0128] In this embodiment, by introducing a bidirectional scale-up interconnect design and simultaneously constructing non-blocking switching units in the south and north directions, communication bandwidth can be improved and communication latency reduced.
[0129] Embodiments of this application also provide a multi-node system, including: a computing node and a switching node. The computing node includes a computing component, and the switching node includes a switching unit. The switching unit includes two layers of switching components formed by cascading first-layer switching components and second-layer switching components. In the two layers of switching components, a switching port of one first-layer switching component is connected to a first switching port of at least two second-layer switching components, and a first switching port of one second-layer switching component is connected to a switching port of at least one first-layer switching component. The transmission rates of the transmission channels of the first-layer switching components and the first transmission channels of the second-layer switching components are both first rates, and the transmission rate of the second transmission channels of the second-layer switching components is a second rate. The first transmission channel of the second-layer switching components is the transmission channel corresponding to the first switching port of the second-layer switching components, and the second transmission channel of the second-layer switching components is the transmission channel corresponding to the second switching port of the second-layer switching components. The first rate and the second rate are different rates. The second transmission channel of the second-layer switching components is the transmission channel provided by the switching unit to the computing component. The communication port of the computing component is connected to the second switching port of the second-layer switching components.
[0130] In this embodiment, a multi-node system can be constructed based on a switching unit. The switching unit in the switching node in this embodiment can be the switching unit 101 in the previous embodiment. For the description of the features corresponding to the switching unit in the previous embodiment, please refer to the relevant description of the embodiments corresponding to the switching unit in the previous embodiment. They will not be repeated here.
[0131] For example, the computing component can be an accelerator card, and the transmission channel supported by the accelerator card can be a high-speed peripheral component interconnect transmission channel.
[0132] Optionally, for a description of the features in the embodiment corresponding to the computing component, please refer to the relevant description of the embodiment corresponding to the computing component in the foregoing embodiments, which will not be repeated here.
[0133] According to the embodiments provided in this application, a multi-node system includes: computing nodes and switching nodes. The computing node includes computing components, and the switching node includes a switching unit. The switching unit includes two layers of switching components formed by cascading first-layer switching components and second-layer switching components. In the two-layer switching components, a switching port of one first-layer switching component is connected to the first switching ports of at least two second-layer switching components, and a first switching port of one second-layer switching component is connected to the switching port of at least one first-layer switching component. The transmission rates of the transmission channels of both the first-layer switching components and the first transmission channels of the second-layer switching components are both first rates. The transmission rate of the second transmission channel is the second rate. The first transmission channel of the second layer switching component is the transmission channel corresponding to the first switching port of the second layer switching component. The second transmission channel of the second layer switching component is the transmission channel corresponding to the second switching port of the second layer switching component. The first rate and the second rate are different rates. The second transmission channel of the second layer switching component is the transmission channel provided by the switching unit to the computing unit. The communication port of the computing unit is connected to the second switching port of the second layer switching component. This can solve the problem of limited vertical expansion interconnection of computing units in related technologies, and achieve the technical effect of increasing the scale of vertical expansion interconnection and thus improving the utilization rate of computing resources.
[0134] In one exemplary embodiment, there are multiple compute nodes and multiple switching nodes, which are located in the same rack, and the compute nodes and switching nodes are interconnected via a blind-plug cable backplane.
[0135] In one exemplary embodiment, the multi-node system further includes at least one of the following: a power supply node, a heat dissipation unit, and a rack-top switch.
[0136] In one exemplary embodiment, there are multiple computing nodes and multiple switching nodes, with multiple computing nodes located in at least one computing cabinet and multiple switching nodes located in at least one switching cabinet, wherein the computing cabinets and switching cabinets are interconnected via external cables.
[0137] The multi-node system in this application embodiment is explained below with reference to optional examples. Based on the computing system interconnection topology constructed by the non-blocking switching unit in the foregoing embodiments, this optional example proposes a multi-node system that can house computing nodes and switching nodes in a single rack, with the computing nodes and switching nodes interconnected via a blind-plug cable backplane.
[0138] like Figure 9As shown, the interconnect topology of a 36-card supernode system built on 288 lanes of non-blocking switching units can realize a 36-card multi-node integrated machine. This 36-card supernode integrated machine occupies a total height of 18U, including 9U compute nodes, 3U switching nodes, and 2... The system includes a 2U power shelf (two power supply nodes, each occupying 2U, used to provide power and monitor and manage the power supply) and a 4U cooling distribution unit (CDU, used to distribute coolant). Here, U represents the height of a rack unit.
[0139] Optionally, each compute node may include two central processing units (CPUs) and four GPUs, and the logical graph of the compute node can be as follows: Figure 10 As shown, in the compute node, the PCIe switch can be connected to the CPU, GPU, non-volatile memory, and network interface card via PCIe x16 links. Each switching node contains two 288-lane non-blocking switching units; the logical diagram of the switching node can be shown as follows. Figure 11 As shown, the 144-lane Gen6 PCIe Switches transmit data at Gen6 rates, and can provide Gen5 rate transmission channels externally. Compute nodes and switching nodes are interconnected via blind-plug cable backplanes, enabling... Figure 8 The interconnect topology shown.
[0140] Furthermore, based on the 72-card computing system interconnect topology constructed using the 576 lanes of non-blocking switching units in the aforementioned embodiments, a 72-card multi-node system can be constructed, such as... Figure 12 As shown, this 72-card multi-node system rack contains 18U compute nodes, 6U switching nodes, and 2... 2U Power shelf power supply node, 2 4U liquid-cooled CDUs and 5U rack-top switches. Each compute node contains 2 CPU+4 The logic diagram of a GPU is also as follows: Figure 10 As shown. Each switching node contains a 576-lane non-blocking switching unit, and its logic diagram is as follows. Figure 11 As shown, compute nodes and switching nodes are interconnected via blind-plug cable backplanes to achieve an extended interconnect topology.
[0141] Furthermore, based on the interconnect topology of the 144-card computing system constructed using the 1152 Lanes non-blocking switching units in the aforementioned embodiments, a 144-card multi-node system can be built. The architecture of the 144-card multi-node system can be similar to that of the 72-card multi-node system, and it can be further extended based on the architecture of the 72-card multi-node system. For example, the 144-card multi-node system can be located in three racks, including two computing racks and one switching rack. Each computing rack contains 18 computing nodes, and each computing node contains 2... CPU+4 The logic diagram of a GPU is also as follows: Figure 10 As shown, the switching cabinet mainly consists of 6 switching nodes, each containing a 1152-lane non-blocking switching unit, and its logic diagram is as follows. Figure 11 As shown, the computing cabinet and the switching cabinet are interconnected via external PCIe cables, thereby enabling an extended interconnect topology.
[0142] Optionally, based on the non-blocking switching unit in the foregoing embodiments, a supernode system topology that simultaneously supports northbound and southbound scale-up interconnects of computing components can be realized.
[0143] Similar to the previous embodiments, the communication ports of the computing unit include southbound and northbound communication ports, and the computing system can support bidirectional interconnection, such as... Figure 13 As shown, the compute nodes can employ PCIe retimers (PCIeRetimers, used for retiming and regenerating signals, extending PCIe transmission distance without altering the data flow). This bidirectional multi-node system can add a northbound scaleup interconnect network based on two 576-lane non-blocking switching units to the north of the GPUs, thus forming an interconnect topology that supports both northbound and southbound full interconnection for 72 GPUs simultaneously. This further provides both southbound and northbound paths for All Reduce, All-to-All, and other communication between GPUs, increasing communication bandwidth and reducing communication latency. Figure 14 As shown, this is a 72-card bidirectional scale-up interconnect multi-node system based on this interconnect topology that supports full northbound and full southbound interconnection. It can be placed in a rack, and the switching nodes can include multiple southbound switching nodes and multiple northbound switching nodes.
[0144] This optional example demonstrates how a scale-up interconnect scheme can be built using a large-lane non-blocking PCIe switching unit and computing components supporting multiple PCIe ports. This allows for full interconnection of various computing components, enabling bandwidth matching and scale expansion between hardware with different transmission rates. It also provides a design scheme for a multi-node system with hundreds of cards. The scale of the computing system can be flexibly adjusted according to the selection of different switching units (e.g., 36 cards, 72 cards, 144 cards, etc.). It supports full interconnection scale-up between computing nodes (e.g., GPU to GPU), with a maximum expansion scale that can be more than ten times that of a traditional 8-card server.
[0145] Furthermore, in this optional example, the PCIe protocol natively supports Load / Store memory semantics, enabling global memory mapping within the entire system and cross-domain memory access. This allows for shared GPU memory capacity of up to TB or 10TB across the entire system, addressing current issues such as small scale-up size, insufficient scale-up interconnect bandwidth, and insufficient GPU memory. This scale-up interconnect solution is compatible with all computing components that support direct PCIe output, eliminating ecosystem barriers. In this optional example, a system topology and design scheme based on a PCIe non-blocking switching unit is also provided, simultaneously supporting north-south scale-up interconnects. This further provides both southbound and northbound paths for All Reduce and All-to-All communication between GPUs, increasing communication bandwidth and reducing communication latency.
[0146] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0147] The foregoing has provided a detailed description of the switching unit, computing system, and multi-node system provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to aid in understanding the methods and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A switching unit, characterized in that, include: A two-layer switching unit formed by cascading first-layer and second-layer switching units; In the two-layer exchange components, A switching port of a first layer switching component is connected to a first switching port of at least two second layer switching components, wherein the transmission rate of the transmission channel of the first layer switching component and the first transmission channel of the second layer switching component are both a first rate, and the first transmission channel of the second layer switching component is the transmission channel corresponding to the first switching port of the second layer switching component. A first switching port of a second layer switching component is connected to at least one switching port of a first layer switching component, wherein the transmission rate of the second transmission channel of the second layer switching component is a second rate, the second transmission channel of the second layer switching component is the transmission channel corresponding to the second switching port of the second layer switching component, the first rate and the second rate are different rates, and the second transmission channel of the second layer switching component is the transmission channel provided by the switching unit to the outside. The ratio of the first rate to the second rate is K; in a second-layer switching component, the ratio of the number of first transmission channels to the number of second transmission channels is L; wherein K and L are both positive numbers greater than 0, and the product of K and L is greater than or equal to 1; the first-layer switching component is only connected to the second-layer switching component.
2. The switching unit according to claim 1, characterized in that, K=2, L=1 / 2; In the two-layer switching components, the number of the second-layer switching components is three times the number of the first-layer switching components.
3. The switching unit according to claim 1, characterized in that, In the two-layer switching components, the number of transmission channels of one switching component is X, the total number of second transmission channels of all second-layer switching components is Y, and the total number of all switching components is N, where X, Y and N are all positive integers greater than or equal to 2, and N≥⌈2Y / X⌉.
4. The switching unit according to claim 3, characterized in that, The first rate is twice the second rate; in the two-layer switching components, the number of the first-layer switching components is ⌈Y / 2X⌉, and the number of the second-layer switching components is ⌈3Y / 2X⌉.
5. The switching unit according to claim 1, characterized in that, In the two-layer switching components, a switching port of one of the first-layer switching components is connected to at least one first switching port of any of the second-layer switching components.
6. The switching unit according to claim 5, characterized in that, In the two-layer switching components, the number of transmission channels of a switching component is X, and the number of first-layer switching components is N1. The number of transmission channels provided by a first-layer switching component to any second-layer switching component is X / N1, where X, N1, and X / N1 are all positive integers greater than or equal to 1.
7. The switching unit according to claim 1, characterized in that, In the two-layer switching components, the second transmission channels of one second-layer switching component are provided externally in groups of Z, where Z is a positive integer greater than or equal to 2.
8. The switching unit according to any one of claims 1 to 7, characterized in that, The switching ports of the two-layer switching components are high-speed peripheral component interconnection ports.
9. A computing system, characterized in that, include: The system includes a switching unit and a computing unit, wherein the switching unit comprises two layers of switching components formed by cascading a first-layer switching component and a second-layer switching component; wherein... In the two-layer switching components, one switching port of the first-layer switching component is connected to at least two first switching ports of the second-layer switching components, and one first switching port of the second-layer switching component is connected to at least one switching port of the first-layer switching component. The transmission rates of the transmission channels of the first-layer switching components and the first transmission channels of the second-layer switching components are both first rates, and the transmission rate of the second transmission channels of the second-layer switching components is a second rate. The first transmission channel of the second-layer switching components is the transmission channel corresponding to the first switching port of the second-layer switching components, and the second transmission channel of the second-layer switching components is the transmission channel corresponding to the second switching port of the second-layer switching components. The first rate and the second rate are different rates, and the second transmission channel of the second-layer switching components is the transmission channel provided by the switching unit to the computing unit. The communication port of the computing component is connected to the second switching port of the second layer switching component; The ratio of the first rate to the second rate is K; in a second-layer switching component, the ratio of the number of first transmission channels to the number of second transmission channels is L; wherein K and L are both positive numbers greater than 0, and the product of K and L is greater than or equal to 1; the first-layer switching component is only connected to the second-layer switching component.
10. The computing system according to claim 9, characterized in that, The number of switching units is M, and the number of computing components is P. One computing component is connected to M switching units, and one switching unit is connected to P computing components, where M and P are both positive integers greater than or equal to 2.
11. The computing system according to claim 10, characterized in that, The computing component includes a port monitoring unit and a controller: wherein, The port monitoring unit is used to monitor the port traffic of the communication port of the computing unit and report the monitored port traffic of the communication port of the computing unit to the controller. The controller is configured to distribute data to be distributed to the communication ports of the computing unit based on the port traffic of the communication ports of the computing unit, so as to perform load balancing among the communication ports of the computing unit. The two computing components are allowed to communicate through either of the switching units.
12. The computing system according to claim 10, characterized in that, In the two-layer switching components of one of the switching units, the second transmission channels of one of the second-layer switching components are provided to one of the computing components in groups of Z; the number of transmission channels supported by one of the computing components is greater than or equal to M*Z.
13. The computing system according to claim 10, characterized in that, The communication ports of the computing unit include a southbound communication port and a northbound communication port; the southbound communication port of the computing unit is connected to the second switching port of the second layer switching component of m1 of the switching units, and the northbound communication port of the computing unit is connected to the second switching port of the second layer switching component of m2 of the switching units, where m1 and m2 are both positive integers greater than or equal to 1, and M = m1 + m2.
14. The computing system according to any one of claims 9 to 13, characterized in that, The computing component is an accelerator card, and the accelerator card supports a high-speed peripheral component interconnection transmission channel.
15. A multi-node system, characterized in that, include: The system includes compute nodes and switching nodes. The compute nodes include compute components, and the switching nodes include switching units. Each switching unit comprises two layers of switching components formed by cascading first-layer and second-layer switching components. In the two-layer switching components, one switching port of the first-layer switching component is connected to at least two first switching ports of the second-layer switching components, and one first switching port of the second-layer switching component is connected to at least one switching port of the first-layer switching component. The transmission rates of the transmission channels of the first-layer switching components and the first transmission channels of the second-layer switching components are both first rates, and the transmission rate of the second transmission channels of the second-layer switching components is a second rate. The first transmission channel of the second-layer switching components is the transmission channel corresponding to the first switching port of the second-layer switching components, and the second transmission channel of the second-layer switching components is the transmission channel corresponding to the second switching port of the second-layer switching components. The first rate and the second rate are different rates, and the second transmission channel of the second-layer switching components is the transmission channel provided by the switching unit to the computing unit. The communication port of the computing component is connected to the second switching port of the second layer switching component; The ratio of the first rate to the second rate is K; in a second-layer switching component, the ratio of the number of first transmission channels to the number of second transmission channels is L; wherein K and L are both positive numbers greater than 0, and the product of K and L is greater than or equal to 1; the first-layer switching component is only connected to the second-layer switching component.
16. The multi-node system according to claim 15, characterized in that, The number of computing nodes and switching nodes are both multiple, and the multiple computing nodes and multiple switching nodes are located in the same rack, wherein the computing nodes and switching nodes are interconnected through a blind-plug cable backplane.
17. The multi-node system according to claim 16, characterized in that, The multi-node system also includes at least one of the following: a power supply node, a heat dissipation unit, and a rack-top switch.
18. The multi-node system according to claim 15, characterized in that, The number of computing nodes and switching nodes are both multiple, with multiple computing nodes located in at least one computing cabinet and multiple switching nodes located in at least one switching cabinet, wherein the computing cabinet and the switching cabinet are interconnected by external cables.
19. The multi-node system according to any one of claims 15 to 18, characterized in that, The computing component is an accelerator card, and the accelerator card supports a high-speed peripheral component interconnection transmission channel.
Citation Information
Patent Citations
Switch components and multiple data rate non-blocking switch network utilizing same
CN1055631A
Graphic processor interconnection structure and data transmission method
CN119669142A