Interconnection system
By designing an interconnection system in a large-scale interconnect cluster, where the computing unit of each server is directly connected to the switch, the delay problem caused by multi-hop communication is solved, achieving more efficient communication and lower interconnection costs.
Patent Information
- Application Number
- CN202411512580.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-28
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-10-28
AI Technical Summary
In large-scale interconnect clusters, the limitations of physical connections lead to bottlenecks in communication efficiency, and multi-hop communication increases communication latency and reduces computing efficiency.
An interconnection system is designed in which each server contains multiple unit groups, and the computing units in each unit group are directly connected to the switch, and one-hop direct communication between computing units in different unit groups is achieved through the switch configuration.
By directly connecting the computing unit to the switch, communication delay is reduced, system bandwidth is increased, bandwidth limitations are reduced by intermediate nodes, and interconnection costs of large clusters are saved.
Smart Images

Figure CN120017616A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of chip design, and in particular to an interconnection system. Background Art
[0002] With the increasing demand for high-performance computing (HPC) and massive parallel processing, building large-scale interconnected clusters has become a key way to improve computing capabilities. However, in large-scale interconnected clusters, the limitation of physical connections leads to a bottleneck in communication efficiency. Specifically, due to the limitation of physical distance and hardware resources, any two computing units often cannot achieve direct communication with one hop, but need to go through multiple intermediate nodes for data exchange. This multi-hop communication increases communication delay, which in turn reduces the overall computing efficiency. This phenomenon is particularly evident in complex interconnected topologies. In topologies such as rings, meshes, or trees, the connection paths of computing units are often long, and direct connections between nodes are limited. For example, in a ring topology, data must be passed along the ring until it reaches the target node. This means that if the physical distance between two nodes in the ring is far, the data needs to pass through multiple intermediate nodes to complete the transmission. Each intermediate node will introduce additional delay and increase the burden on the node to process data packets, resulting in a decrease in overall communication efficiency.
[0003] In the traditional northbound network communication architecture, the computing unit is connected to the PCIe switch chip in the server through the PCIe interface, and then connected to the network card and accessed to the cluster network switch. This communication architecture alleviates the delay problem introduced by the super-large ring topology multi-hop communication to a certain extent, but this architecture itself contains at least 6 intermediate nodes, such as: computing unit A to PCIe switch chip, PCIe switch chip to network card A, network card A to network switch, network switch to network card B, network card B to PCIe switch chip, PCIe switch chip to computing unit B. Obviously, the communication delay of the traditional network communication architecture cannot be ignored, and the communication delay increases with the increase of the number of network switch networking layers. In addition, the communication bandwidth of the computing unit under this architecture is limited by the number of PCIe switch chips in the server and the PCIe interface bandwidth of the computing unit. The delay caused by this multi-hop communication not only affects the data transmission speed, but also has a negative impact on the parallel processing capability of the GPU cluster. In high-performance computing, many tasks require frequent inter-node communication and data exchange. The increase in communication delay directly affects the overall completion time of the computing task and reduces the throughput of the system. Therefore, when designing a large-scale interconnected cluster, how to optimize the topology to reduce the latency of multi-hop communication and reduce the bandwidth constraints of intermediate nodes becomes an urgent problem to be solved. Summary of the invention
[0004] In view of the above technical problems, the technical solution adopted by the present invention is: an interconnection system, the system comprising: at least one switch and K servers. Each server comprises a plurality of unit groups, wherein the i-th server serv i Includes R(i) unit groups; serv i The r-th unit group includes T computing units, each computing unit includes Q ports to be interconnected with the switch, wherein i ranges from 1 to K, r ranges from 1 to R(i), and t ranges from 1 to T. The switch is connected to the Q ports of all computing units in the K servers, and Sum interconnection relationships are obtained by configuring the switch, each interconnection relationship includes a computing unit in each unit group, and different computing units in the same unit group are located in different interconnection relationships.
[0005] The present invention has at least the following beneficial effects:
[0006] By directly connecting the Q ports of the computing unit to be interconnected with the switch, two computing units in different unit groups can communicate directly with each other through only one hop. Compared with the hardware topology formed by the computing units in each unit group or the traditional northbound network, the communication delay is reduced, the bandwidth of the system is increased, and the bandwidth restrictions of the intermediate nodes are reduced. At the same time, the switch distinguishes the Sum interconnection relationships between different computing units in the unit group, and the communication between the T computing units in the unit group is completed through the fixed topology within the unit group. Under the same interconnection scale and bandwidth, compared with the traditional northbound network, the number of ports interconnected through the switch is reduced, further saving the interconnection cost of large clusters. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0008] Figure 1 A schematic diagram of an interconnection system provided in Embodiment 1 of the present invention;
[0009] Figure 2 A schematic diagram of a logical topology formed by configuring a switch according to the second embodiment of the present invention;
[0010] Figure 3 For the general Figure 2 A schematic diagram of a loop topology from which the topology in FIG. 1 is switched;
[0011] Figure 4 for Figure 3 A schematic diagram of a degenerate topology when a computing unit fails;
[0012] Figure 5 for Figure 3 Schematic diagram of a redundant topology that is switched to when a logical topology fails. DETAILED DESCRIPTION
[0013] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0014] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present invention have the same meanings as commonly understood by those skilled in the art.
[0015] Embodiment 1
[0016] See also Figure 1 , which shows a schematic diagram of an interconnection system, wherein the system includes at least one switch and K servers.
[0017] Among them, the server is a physical computer, which is usually configured with one or more central processing units, a certain number of computing units, and sufficient memory and other hardware resources.
[0018] Among them, the switch is a network device that can connect computing units in multiple servers to enable cross-server communication between computing units.
[0019] In one implementation, the switch is a PCIe (Peripheral Component Interconnect Express) switch or an optical circuit switch (OCS). Other switches that can directly connect computing units across servers fall within the protection scope of the present invention.
[0020] Furthermore, each server includes multiple unit groups, where the i-th server serv i Includes R(i) units Gp i = {Gp i,1 ,Gp i,2 ,…,Gp i,r ,…,Gp i,R(i)}, Gp i,r is the rth unit group in the i-th server, r ranges from 1 to R(i), and i ranges from 1 to K; please refer to Figure 1 , Figure 1 China-Israel serv i The rth unit group Gp i,r As an example, the composition of all unit groups is represented by Gp i,r It includes T computing units {PRC i,r,1 ,PRC i,r,2 ,…,PRC i,r,t ,…,PRC i,r,T},PRC i,r,t Gp i,r The tth computing unit in , the value range of t is 1 to T. Each computing unit includes Q ports to be interconnected with the switch.
[0021] The value of R(i) is the number of cell groups. Different servers may include the same number of cell groups or different numbers of cell groups. For example, one server includes two cell groups and another server includes four cell groups.
[0022] In one embodiment, the computing unit is a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU) or a neural processing unit (NPU), and other computing units for processing data fall within the scope of protection of the present invention.
[0023] In one embodiment, T×R(i)≤th0, where th0 is a preset threshold value for the optimal number of computing units in the server. When it is less than th0, the system is more reliable and the replacement cost is low; when it is greater than th0, the system reliability risk increases and the replacement cost increases. Since the server replacement frequency is high, pre-setting th0 can not only reduce the replacement cost, but also ensure that the system can still provide sufficient computing resources in the case of frequent server replacement, thereby ensuring the continuous and reliable operation of the system. For example, th0 is equal to 16, and when the number of T×R(i) is less than 16, the system is more reliable and the replacement cost is low; when it is greater than or equal to 16, the system reliability risk increases and the replacement cost increases.
[0024] In one implementation, each server includes a unit group, and the unit group includes 4 computing units.
[0025] In one embodiment, each server includes two unit groups, and each unit group includes four computing units.
[0026] Other combinations of the number of unit groups included in each server and the number of computing units in each unit group fall within the protection scope of the present invention.
[0027] In one embodiment, the computing unit is a module that complies with the OAM standard. The computing unit is a module that complies with other standards and also falls within the protection scope of the present invention.
[0028] Furthermore, the switch is directly connected to the Q ports of all computing units in the K servers, and Sum interconnection relationships are obtained by configuring the switch, each interconnection relationship includes a computing unit in each unit group, and different computing units in the same unit group are in different interconnection relationships, where T≤Sum≤Q×T. Among them, the switch enables any two computing units in different unit groups to communicate directly through one hop. Compared with the hardware topology formed by the computing units in each unit group, the communication between the computing units is carried out through a shorter path, which reduces the communication delay and increases the bandwidth of the system. At the same time, since all computing units in each unit group are connected to different switches respectively, the pressure of cross-group communication is dispersed, and the load is more balanced; the traffic between different switches can be distributed more evenly, reducing the load bottleneck of a single switch and optimizing the use of network bandwidth.
[0029] In one embodiment, when Sum=T, the system includes T interconnection relationships with the same bandwidth, and each interconnection relationship is configured by configuring the SW so that the Q ports of each computing unit to be interconnected with the switch are configured in the same interconnection relationship. That is, the Q ports of each computing unit to be interconnected with the switch are connected to the same switch and are configured to be connected to the same computing unit in the same interconnection relationship.
[0030] Optionally, when Q=1, the system includes T switches, a single switch connects one computing unit in each unit group, and different switches connect different computing units in the same group; by configuring each switch so that all computing units connected to the current switch are configured as an interconnection relationship, the number of interconnected computing units can be maximized.
[0031] As an example, see again Figure 1 , which shows that the system includes T switches SW = {SW1, SW2, ..., SW f ,…,SW T}, SW f is the fth switch, and the value range of f is 1 to T. Each switch is connected to a computing unit in all unit groups. f Connect Gp i,r PRC i,r,t , and SW fNot with Gp i,r The qth computation unit PRC i,r,q Connect, PRC i,r,t Not with the hth switch SW in SW h connection, the value range of q and f is 1 to T, and t≠q, f≠h. i,r,q and SW h None of them are shown in the figure.
[0032] In one embodiment, when the interconnection ports of the switch itself are sufficient, the number of switches can be less than T. For example, each switch is connected to two computing units in each unit group, and after configuring the switch, the two computing units in each unit group connected to the same switch are located in two interconnection relationships. For another example, when the interconnection ports of the switch itself are completely sufficient to support all computing units, the number of switches can be 1, and all computing units are connected through one switch, but the computing units in each interconnection relationship formed still belong to computing units of different unit groups.
[0033] In one embodiment, when Sum=Q×T, the system includes Q×T interconnection relationships with the same bandwidth, and each interconnection relationship is configured by configuring the SW so that the Q ports of each computing unit to be interconnected with the switch are configured in Q different interconnection relationships. That is, the Q ports of the same computing unit are located in different interconnection relationships. In other words, the Q ports of the same unit are interconnected with different computing units, so that each computing unit can be connected to other computing units through Q different interconnection relationships at the same time, reducing the number of communication hops between groups.
[0034] In one embodiment, when T<Sum<Q×T, the system includes at least two interconnection relationships with different bandwidths: the number of ports configured for each computing unit in the first interconnection relationship is different from the number of ports configured for each computing unit in the second interconnection relationship. For example, the number of ports configured for each computing unit in the first interconnection relationship is 1, and the number of ports configured for each computing unit in the second interconnection relationship is Q-1, that is, one of the ports of each computing unit is configured as the first interconnection relationship, and the remaining Q-1 are configured as the second interconnection relationship. Other types of interconnection relationships also fall within the protection scope of the present invention.
[0035] Among them, different computing units in the same group are connected to different switches. When a switch fails, the impact can be limited to some computing units connected to the switch without affecting all computing units in the entire group, providing better fault isolation capabilities.
[0036] In one embodiment, when each unit group is numbered independently, all computing units connected to the same switch may have the same or different numbers. For example, each unit group includes 4 computing units, and the computing units in each group are numbered 1-4. The first connection method is to connect the same number: f For example, SW f Connect the computing unit numbered 1 in each group. The second connection method connects different numbers: f For example, SW f The computing unit numbered 1 in the first group, the computing unit numbered 2 in the second group, the computing unit numbered 3 in the third group, and the computing unit numbered 4 in the fourth group are connected respectively. Alternatively, the numbers in two of the connected groups are the same and the numbers in the other two groups are different.
[0037] In one embodiment, when each x unit group is numbered independently, the numbers of all computing units connected to the same switch include the preset x numbers. Taking x=2 as an example, when the computing unit is a module that complies with the OAM standard, the maximum number of computing units is 8, and every two unit groups form a numbering cycle, so the two numbers connected to the same switch in the two unit groups are different. For example: including 4 unit groups, every 2 unit groups form a numbering cycle, where SW f For example, in the first numbering cycle, the computing unit numbered 2 in the first unit group and the computing unit numbered 4 in the second unit group are connected to SW f The calculation unit numbered 2 in the third unit group and the calculation unit numbered 4 in the fourth unit group in the second numbering cycle are also connected to SW f .
[0038] In one embodiment, the connection between the switch and the computing unit is an optical cable or an electrical cable. Other types of connections for data communication also fall within the protection scope of the present invention.
[0039] In one embodiment, the connection between the computing units is an optical cable or an electrical cable, and other types of connections for data communication also fall within the protection scope of the present invention.
[0040] In one embodiment, the computing units in a unit group are interconnected in a ring topology or a mesh topology. Other types of topologies also fall within the protection scope of the present invention.
[0041] In one embodiment, K servers form a server group, each group including G servers. For the E unit groups within each server in each server group, the computing units in all unit groups are interconnected as a hardware interconnection topology, and G and E are both greater than or equal to 1.
[0042] In one implementation, the hardware interconnection topology is a ring topology, a Dragonfly topology, or a mesh topology. Other types of topologies also fall within the protection scope of the present invention.
[0043] In one implementation, by configuring the current switch, the target computing unit connected to the current switch can directly access other computing units with different numbers from the target computing unit. The logical topology structure formed by this configuration rule is suitable for the MOE scenario.
[0044] In one embodiment, when the numbers of all computing units connected to the same switch include x preset numbers, by configuring the SW f , so that PRC i,r,t Direct access to all other connected SW f And the number is the same as PRC i,r,t Different computing units. Or by configuring SW f , so that PRC i,r,t Direct access to all other connected SW f Similarly, the configuration rules of other switches are the same as SW f The logical topology structure formed by this configuration rule is suitable for the MOE scenario.
[0045] The solution provided in the first embodiment of the present invention enables any two computing units in different unit groups to communicate directly through one hop through a switch. Compared with the hardware topology formed by the computing units in each unit group, the communication between the computing units is carried out through a shorter path, which reduces the communication delay, increases the bandwidth of the system, and reduces the bandwidth restrictions of the intermediate nodes. At the same time, the switch distinguishes the Sum interconnection relationships between different computing units in the unit group, and the communication between the T computing units in the unit group is completed through the fixed topology within the unit group. Under the same interconnection scale and bandwidth, compared with the traditional northbound network, the number of switch interconnection ports is reduced, further saving the interconnection cost of large clusters.
[0046] In order to obtain an interconnection system that can both expand the interconnection topology and improve the fault redundancy of the system, the present invention provides a second embodiment.
[0047] Embodiment 2
[0048] A second embodiment of the present invention provides an interconnection system based on a double-layer topology structure, the system comprising at least one switch and at least one interconnection topology, each interconnection topology comprising N physical topologies and multiple logical topologies.
[0049] Among them, the physical topology is a direct point-to-point connection through actual physical lines.
[0050] In one implementation, the connection of the physical topology is an electrical cable or an optical cable.
[0051] The switch in the first embodiment of the present invention is also applicable to the second embodiment of the present invention and will not be described in detail.
[0052] Further, each physical topology includes: L unit groups, intra-group physical lines, and inter-group physical lines. Each unit group includes T computing units. Each intra-group physical line is used to connect the computing units in each unit group. Each inter-group physical line is used to connect computing units belonging to different unit groups, and each inter-group physical line connects two different computing units.
[0053] In one embodiment, the unit group includes an intra-group topology, which is a topology formed by T computing units in the unit group through intra-group physical connections. The topology is a ring topology, a mesh topology, or a star topology. Other types of topologies also fall within the protection scope of the present invention.
[0054] In one embodiment, the cell group includes an inter-group topology, and the inter-group topology is that the T computing units in the cell group are connected point-to-point with the T computing units of the adjacent group through the inter-group physical connection to form the inter-group topology, and other types of topologies also fall within the protection scope of the present invention. Adjacent groups are physically adjacent cell groups. When the physical topology is distributed in the same direction, adjacent groups are physically adjacent cell groups, and the first cell group and the last cell group are regarded as adjacent groups.
[0055] Furthermore, the switch is used to connect all computing units in the interconnection topology.
[0056] The switch enables any two computing units in different groups to communicate directly through one hop, without the need to go through the physical topology formed by the computing units in each group. At the same time, since all computing units in each group are connected to different switches, the pressure of cross-group communication is dispersed.
[0057] It should be noted that the connection relationship between the switch and each unit group in the first embodiment of the present invention is also applicable to the second embodiment of the present invention and will not be described in detail.
[0058] Furthermore, multiple logical topologies are logical topologies obtained by configuring the connectivity between computing units in N physical topologies through switches, each logical topology includes two unit groups and a logical path therebetween, wherein the i-th logical topology includes T logical paths, each logical path is used to connect two computing units belonging to two unit groups, and different logical paths connect different computing units.
[0059] Each logical path includes two lines connecting the two computing units to the switch, and a logical path configured inside the switch.
[0060] Among them, the interconnection scale of the GPU can be expanded by interconnecting the two unit groups through the logical topology. In this way, the unit groups can be connected through the logical topology to achieve the purpose of expanding the interconnection scale of the GPU. For example, when each unit group is 4 computing units, after the two unit groups are interconnected through the logical topology, the interconnection scale of 8 GPUs is formed. If a third unit group is added on the basis of the interconnection scale of these 8 GPUs, and the third unit group is interconnected with any one of the groups, the interconnection scale of 12 GPUs is formed. If a fourth unit group is added, the interconnection scale of 16 GPUs is formed, and so on.
[0061] In addition, the connections in the physical topology and the logical paths in the logical topology are redundant. When one or more connections fail, the path provided by the logical path replaces the failed path for data exchange, reducing the risk of network interruption and improving the overall reliability of the system. When an inter-group connection failure in the physical topology is detected, it can automatically switch to the logical path, simplifying the fault recovery process and reducing the need for manual intervention. The system provides effective response measures for possible technical failures and ensures the continuity and integrity of data transmission.
[0062] In one embodiment, the two cell groups in each logical topology belong to two cell groups in different physical topologies respectively. The cell groups are connected in sequence through the logical topology to form an interconnection structure with the cell group as the basic node and the logical topology as the basic path, thereby infinitely expanding the interconnection scale of the GPU. For example, two cell groups are included in a physical topology, each of which includes 4 GPUs. In the physical topology, one cell group is used to connect to a cell group in the previous physical topology, and the other cell group in the physical topology is used to connect to a cell group in the next physical topology. In one embodiment, when the number of cell groups included in the physical topology is greater than or equal to 2, two logical topologies are included between the two physical topologies, and the two groups of logical topologies are respectively connected to two cell groups belonging to the same physical topology.
[0063] In one embodiment, two unit groups belonging to different physical topologies in each logical topology are physically adjacent. Wherein, when the physical topologies are sequentially distributed in the same direction, the first physical topology and the last physical topology distributed sequentially are considered to be physically adjacent. In this way, adjacent physical topologies can be connected in sequence through logical topologies to form a ring interconnection structure with unit groups as basic nodes and logical topologies as basic paths. On the basis of infinitely expanding the scale of GPU interconnection, it allows transmission in both clockwise and counterclockwise directions, thereby improving system throughput.
[0064] In one embodiment, the two unit groups in each logical topology belong to two unit groups in different physical topologies respectively, and when a physical line between groups fails, the two target unit groups connected by the failed physical line between groups are obtained, the two sets of logical paths corresponding to the two target unit groups are obtained, all the physical lines between groups of all unit groups connected by the two sets of logical paths are configured as unavailable, and the new topology formed by the remaining physical topology and logical topology is switched to perform data exchange. That is, the system function is restored by combining the physical topology and logical topology that have not failed to form a new fusion topology for exchanging data. Other methods of replacing failed physical lines through logical paths fall within the protection scope of the present invention.
[0065] In one implementation, the two cell groups in each logical topology are two cell groups in the same physical topology.
[0066] In one embodiment, in an interconnection system, the following two configurations coexist: the two cell groups in each logical topology belong to two cell groups in different physical topologies respectively; and the two cell groups in each logical topology are two cell groups in the same physical topology.
[0067] In one embodiment, two cell groups in the same physical topology in each logical topology are physically adjacent. Wherein, when the cell groups are sequentially distributed in the same direction, the first cell group and the last cell group in the sequentially distributed cell groups are considered adjacent groups.
[0068] In one embodiment, in an interconnection system, the following two configurations coexist: in each logical topology, two cell groups belonging to different physical topologies are physically adjacent; and in each logical topology, two cell groups in the same physical topology are physically adjacent.
[0069] In one embodiment, the two cell groups in each logical topology are two cell groups in the same physical topology, and when a physical line between groups fails, two target cell groups connected to the failed physical line between groups are obtained, all physical lines between groups between the two target cell groups are configured as unavailable, and a new topology formed by the remaining physical topology and logical topology is switched to perform data exchange. Other methods of replacing a failed physical line through a logical path fall within the protection scope of the present invention.
[0070] For ease of understanding, the system includes two physical topologies, each of which includes four unit groups, each of which includes four computing units, and the computing units are modules that comply with the OAM standard. Every two computing units are numbered independently. This example explains the new topology structure configured when a physical line between groups fails. Please refer to Figure 2 , Figure 2 The interconnection system shown includes two physical topologies and two logical topologies. The two physical topologies are: the first physical topology and the second physical topology, each physical topology includes four unit groups: the first unit group, the second unit group, the third unit group, and the fourth unit group, each unit group has intra-group physical connections, and point-to-point interconnection between unit groups. The first unit group and the second unit group are an inseparable basic unit group, which are numbered independently, and the third unit group and the fourth unit group are an inseparable basic unit group, which are numbered independently. Figure 2 As shown, the computing units in the first unit group and the second unit group are numbered S0-S7, and the computing units in the third unit group and the fourth unit group are numbered S0-S7. In the first physical topology and the second physical topology, the computing units numbered S1 and S0 are connected to the first switch SW1, the computing units numbered S2 and S4 are connected to the second switch SW2, the computing units numbered S3 and S5 are connected to the third switch SW3, and the computing units numbered S6 and S7 are connected to the fourth switch SW4. After configuring SW1-SW4, two logical topologies are obtained, as shown in FIG. Figure 2 As shown, the first logical topology includes the second unit group in the first physical topology and the first unit group in the second physical topology, and the second logical topology includes the fourth unit group in the first physical topology and the third unit group in the second physical topology. The first logical topology and the second logical topology are the same, both including a logical path connecting S1 and S0, a logical path between S2 and S4, a logical path between S3 and S5, and a logical path between S6 and S7. When at least one inter-group connection between the second unit group and the fourth unit group in the first physical topology fails, for example, the inter-group physical connection between S1 in the second unit group of the first physical topology and S3 in the fourth unit group of the first physical topology fails, all inter-group physical connections between the second unit group and the fourth unit group in the first physical topology are configured as unavailable, and at the same time, all inter-group physical connections between the first unit group and the third unit group in the second physical topology are also configured as unavailable. Please refer to Figure 3 , the system is switched from the original topology to a loop topology composed of the remaining physical topology and logical topology. That is, the loop topology is a loop topology formed by sequentially passing through the logical path between the second unit group of the first physical topology and the first unit group of the second physical topology, the inter-group physical connection of the first, second, fourth and third unit groups of the second physical topology, the logical path between the third unit group of the second physical topology and the fourth unit group of the first physical topology, and the inter-group physical connection of the fourth, third, first and second unit groups of the first physical topology to form a loop topology with unit groups as the basic unit, and data is exchanged through the loop topology to restore system functions and improve system stability and reliability.
[0071] In one embodiment, when a computing unit fails, a degraded topology is obtained: if L1 unit groups are bound as an indivisible basic unit group, when a computing unit fails, the basic unit group where the failed computing unit is located and its associated inter-group physical connections and logical topologies are configured as unavailable, respectively, to obtain at least one original degraded topology; the switch is reconfigured according to the original degraded topology to obtain multiple degraded logical topologies, and the original degraded topology and the degraded logical topology constitute a degraded topology. By combining the logical topology with the physical topology, the original degraded topology is fully utilized to obtain the degraded topology, and the failure of a computing unit will not cause the entire interconnection system to be unusable.
[0072] See also Figure 3 , in another implementation scenario, assuming Figure 3 There are no faulty inter-group physical connections, that is, all inter-group physical connections between the second unit group and the fourth unit group in the first physical topology that are configured as unavailable, and all inter-group physical connections between the first unit group and the third unit group in the second physical topology are configured as a loop topology with unit groups as the basic unit due to design requirements. And the two unit groups are bound into an inseparable basic unit group, wherein one basic unit group in the first physical topology is the first and second unit groups, and the other basic unit group is the third and fourth unit group; one basic unit group in the second physical topology is the first and second unit groups, and the other basic unit group is the third and fourth unit group. When S0 of the first unit group in the second physical topology fails, the basic unit where S0 is located and its related inter-group physical connections and logical topology are configured as unavailable, and an original degraded topology is obtained. The original degraded topology includes the first physical topology and the target basic unit group composed of the third and fourth unit groups in the second physical topology connected by the logical topology. The obtained original degraded topology is as follows: Figure 4 See Figure 4 , reconfigure the switch according to the original degraded topology, configure the logical topology between the first physical topology and the target basic unit group as unavailable, and configure the logical topology inside the target basic unit group, so that S2 of the third unit group and S4 of the fourth unit group in the second physical topology are logically connected, S6 of the third unit group and S7 of the fourth unit group are logically connected, S1 of the third unit group and S0 of the fourth unit group are logically connected, and S5 of the third unit group and S3 of the fourth unit group are logically connected, and a degraded topology consisting of the physical topology of the target basic unit group and the configured internal logical topology is obtained. That is, the degraded topology includes two: one is the first physical topology, and the other is the degraded topology of the target basic unit group.
[0073] In one implementation, when the logical topology fails, the failed logical topology is configured as unavailable, and at least one degraded topology is obtained. Alternatively, when the logical topology fails, the failed logical topology is configured as unavailable, and another logical path is configured to replace the failed logical path. When one or more logical paths fail, the logical topology to which the failed logical path belongs is considered to be failed. Through the flexible and configurable characteristics of the logical topology, the logical topologies are mutually redundant, which improves the fault tolerance and system flexibility, and ensures the stability and reliability of the interconnected system.
[0074] See also Figure 3 , in another implementation scenario, Figure 3 The loop topology is based on the unit group as the basic unit, and in Figure 3 There are no faulty inter-group physical connections in the network, that is, the inter-group physical connections that are set to be unavailable are not faulty, but are configured to be unavailable according to design requirements. When the logical path between S1 in the second unit group of the first physical topology and S0 in the first unit group of the second physical topology fails, it is considered that the logical topology between the second unit group of the first physical topology and the first unit group of the second physical topology fails, and the faulty logical topology is configured to be unavailable. At the same time, please refer to Figure 5 , in order to restore Figure 3 In the loop topology with the unit group as the basic unit, the logical topology between the 4th unit group of the 1st physical topology and the 3rd unit group of the 2nd physical topology is also configured as unavailable, and the inter-group physical connection between the 2nd and 4th unit groups of the 1st physical topology is configured as available, and the inter-group physical connection between the 1st and 3rd unit groups of the 2nd physical topology is configured as available, and the logical topology between the 1st unit group of the 1st physical topology and the 2nd unit group of the 2nd physical topology, as well as the logical topology between the 3rd unit group of the 1st physical topology and the 4th unit group of the 2nd physical topology are obtained through switch configuration to form a new ring topology.
[0075] Other interconnection topologies caused by computing unit failures or connection failures, which require the configuration of switches to degenerate the interconnection scale into a degraded topology with a smaller interconnection scale, all fall within the protection scope of the present invention.
[0076] The system provided by the second embodiment of the present invention includes a physical topology and a logical topology configured by a switch. The connection lines in the physical topology and the logical paths in the logical topology cooperate with each other, so that the two unit groups are interconnected through the logical topology, thereby achieving the purpose of expanding the scale of GPU interconnection. At the same time, since the logical topology has the characteristics of flexible and configurable communication paths, when one or more connection lines fail, the path provided by the logical path replaces the failed path for data exchange, and its fault tolerance is high, reducing the risk of network interruption and improving the overall reliability of the system.
[0077] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0078] Although some specific embodiments of the present invention have been described in detail by way of example, it should be understood by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present invention. It should also be understood by those skilled in the art that various modifications may be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.
Claims
1. An interconnection system, characterized in that: The system comprises: at least one switch; K servers, each server includes multiple unit groups, where the i-th server serv i Includes R(i) unit groups; serv i The r-th unit group includes T computing units, each computing unit includes Q ports to be interconnected with the switch, wherein the value range of i is 1 to K, the value range of r is 1 to R(i), and the value range of t is 1 to T; The switch is directly connected to the Q ports of all computing units in the K servers, and Sum interconnection relationships are obtained by configuring the switch, each interconnection relationship includes a computing unit in each unit group, and different computing units in the same unit group are located in different interconnection relationships, wherein T≤Sum≤Q×T.
2. The system according to claim 1, characterized in that When Sum=T, the system includes T interconnection relationships with the same bandwidth, and each interconnection relationship is configured by configuring the SW so that Q ports of each computing unit to be interconnected with the switch are configured into the same interconnection relationship.
3. The system according to claim 1, characterized in that When Sum=Q×T, the system includes Q×T interconnection relationships with the same bandwidth, and each interconnection relationship is configured by configuring the SW so that the Q ports of each computing unit to be interconnected with the switch are configured into Q different interconnection relationships.
4. The system according to claim 1, characterized in that When T<Sum<Q×T, the system includes at least two interconnection relationships with different bandwidths: the number of ports configured for each computing unit in the first interconnection relationship is different from the number of ports configured for each computing unit in the second interconnection relationship.
5. The system according to claim 1, characterized in that The number of switches in the system is T, a single switch connects a computing unit in each unit group, and different switches connect different computing units in the same group; by configuring each switch, all computing units connected to the current switch are configured into an interconnected relationship.
6. The system according to claim 5, characterized in that When each unit group is numbered independently, all computing units connected to the same switch may have the same or different numbers.
7. The system according to claim 5, characterized in that When each x unit group is numbered independently, the numbers of all computing units connected to the same switch include the preset x numbers.
8. The system according to claim 7, characterized in that By configuring the current switch, the target computing unit connected to the current switch can directly access other computing units with numbers different from the target computing unit.
9. The system according to claim 1, characterized in that The switch is a PCIe switch or an all-optical switch.
10. The system according to claim 1, characterized in that T×R(i)≤th0, where th0 is the maximum number of computing units in the preset server.
Citation Information
Patent Citations
Full interconnection communication method and full interconnection communication device based on network card direct connection
CN105119786A
Machine learning-oriented distributed computing interconnection network system and communication method
CN111193971A
Out-of-band Ethernet interface switching device, multi-node server system and server equipment
CN115996204A
Efficient interconnecting computing nodes to efficient
CN116530069A
Method for simulating large-scale topology by using switch
CN117097625A