Interconnection system based on double-layer topology
By introducing a two-layer topology interconnect system in the GPU cluster, and using switches to configure the logical topology to provide redundant paths, the reliability problem of large-scale interconnect systems is solved, and fault tolerance and interconnect scale expansion are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- METAX INTEGRATED CIRCUITS (SHANGHAI) CO LTD
- Filing Date
- 2024-10-28
- Publication Date
- 2026-05-05
AI Technical Summary
In large-scale interconnected GPU clusters, system reliability issues are becoming increasingly prominent due to the complexity of the physical topology and the dependence of inter-node connections. In particular, in complex interconnect structures, node failures may lead to communication interruptions or performance degradation.
The system employs a two-layer topology interconnection system, including a physical topology and a logical topology configured via switches. The logical topology is used to connect computing unit groups, providing redundant paths to replace faulty paths and improving the system's fault tolerance and flexibility.
Through flexible configuration of logical topology, paths can be automatically switched in the event of a failure, reducing the risk of network outages, improving the overall reliability and flexibility of the system, and expanding the scale of interconnection.
Smart Images

Figure CN121979828A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of chip design, and in particular to an interconnect system based on a two-layer topology. Background Technology
[0002] With the increasing demand for high-performance computing (HPC) and massively parallel processing, building large-scale interconnect clusters has become a key approach to improving computing power. However, in large-scale interconnected GPU clusters, system reliability issues are becoming increasingly prominent due to the complexity of the physical topology and the dependencies between node connections. Especially in complex interconnect structures involving a large number of computing nodes, the design of the physical topology not only affects computing performance but also directly impacts system stability and fault tolerance.
[0003] In complex interconnect topologies, such as ring, mesh, or tree structures, connections between nodes often involve multiple dependencies. If one critical node or connection link fails, it can lead to a complete network communication outage or a significant performance degradation. For example, in a ring topology, all nodes are connected through a closed loop. If any segment of the loop fails, the integrity of the loop is compromised, preventing data from being transmitted along the intended path, potentially paralyzing the entire system. Therefore, there is an urgent need for an interconnect system that can both extend the interconnect topology and improve the system's fault redundancy. Summary of the Invention
[0004] To address the aforementioned technical problems, the present invention adopts the following technical solution: an interconnection system based on a two-layer topology, the system comprising at least one switch and at least one interconnection topology, each interconnection topology comprising N physical topologies and multiple logical topologies. Each physical topology comprises: L unit groups, each unit group comprising T computing units; intra-group physical connections for connecting the computing units within each unit group; and inter-group physical connections for connecting computing units belonging to different unit groups, with each inter-group physical connection connecting two different computing units. The switch connects all computing units in the interconnection topology. The multiple logical topologies are obtained by configuring the switch, each logical topology comprising two unit groups and logical paths between them, wherein the i-th logical topology comprises T logical paths, each logical path connecting two computing units belonging to two unit groups, with different logical paths connecting different computing units.
[0005] The present invention has at least the following beneficial effects:
[0006] Embodiment 2 of this invention provides an interconnection system based on a two-layer topology, which includes a physical topology and a logical topology configured via a switch. The switch configures the interconnection between computing units in two unit groups to form a logical topology. This method enables the interconnection between two unit groups via a logical topology, achieving the goal of expanding the scale of GPU interconnection. Simultaneously, because the logical topology has the characteristic of flexible and configurable communication paths, it can work in conjunction with the physical topology to improve the redundancy and flexibility of the system topology. When one or more connections fail, new paths provided by the logical and physical paths can replace the failed paths for data exchange or switching of interconnection scale. It has high fault tolerance, improving the overall reliability and flexibility of the system. Attached Figure Description
[0007] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0008] Figure 1 This is a schematic diagram of an interconnection system provided in Embodiment 1 of the present invention;
[0009] Figure 2 This is a schematic diagram of the logical topology formed after configuration via a switch, as provided in Embodiment 2 of the present invention.
[0010] Figure 3 To be Figure 2 A schematic diagram of a loop topology switching in the topology;
[0011] Figure 4 for Figure 3 A schematic diagram of a degenerate topology in the event of a fault in a computing unit;
[0012] Figure 5 for Figure 3 A schematic diagram of a redundant topology that switches to when a logical topology fails. Detailed Implementation
[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] Unless otherwise defined, all technical and scientific terms used in the embodiments of this invention have the same meaning as commonly understood by those skilled in the art.
[0015] Example 1
[0016] Please see Figure 1 The diagram illustrates an interconnected system comprising at least one switch and K servers.
[0017] A server is a physical computer, typically configured with one or more central processing units, a certain number of computing units, and sufficient memory and other hardware resources.
[0018] Among them, the switch is a network device that can connect computing units within multiple servers, enabling cross-server communication between computing units.
[0019] In one implementation, the switch is a PCIe (Peripheral Component Interconnect Express) switch or an optical circuit switch (OCS). Other switches capable of directly connecting computing units across servers also fall within the scope of this invention.
[0020] Furthermore, each server comprises multiple unit groups, where the i-th server serves... i Includes R(i) unit groups Gp i ={Gp i,1 ,Gp i,2 ,…,Gp i,r ,…,Gp i,R(i)}, Gp i,r Let r be the r-th unit group in the i-th server, where r ranges from 1 to R(i) and i ranges from 1 to K; please refer to [link / reference needed]. Figure 1 , Figure 1 China-Israel service i The r-th unit group Gp i,r To illustrate the composition of all unit groups, let's take Gp as an example. i,r Includes T computational units {PRC} i,r,1 PRC i,r,2 ,…,PRC i,r,t ,…,PRC i,r,T},PRC i,r,t For Gp i,r The t-th calculation unit is defined, where t ranges from 1 to T. Each calculation unit includes Q ports to be interconnected with the switch.
[0021] Here, R(i) represents the number of cell groups. Different servers can include the same number of cell groups or different numbers of cell groups. For example, one server may include two cell groups, while another server may include four cell groups.
[0022] In one embodiment, the computing unit is a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), or a neural processing unit (NPU). Other computing units used for processing data fall within the protection scope of this invention.
[0023] In one implementation, T×R(i) ≤ th0, where th0 is a preset threshold for the optimal number of computing units in the server. When th0 is less than th0, the system is more reliable and has lower replacement costs; when th0 is greater than th0, the system reliability risk increases and the replacement cost rises. Since servers are replaced frequently, setting th0 in advance not only reduces replacement costs but also ensures that the system can still provide sufficient computing resources even with frequent server replacements, guaranteeing the system's continuous and reliable operation. For example, if th0 equals 16, when the number of T×R(i) is less than 16, the system is more reliable and has lower replacement costs; when it is greater than or equal to 16, the system reliability risk increases and the replacement cost rises.
[0024] In one implementation, each server includes a unit group, and the unit group includes four computing units.
[0025] In one implementation, each server includes two unit groups, and each unit group includes four computing units.
[0026] Other combinations of configurations regarding the number of unit groups included in each server and the number of computing units in each unit group are all within the scope of protection of this invention.
[0027] In one embodiment, the computing unit is a module conforming to the OAM standard. Modules whose computing units conform to other standards also fall within the scope of this invention.
[0028] Furthermore, the switch directly connects to the Q ports of all computing units in the K servers. By configuring the switch, Sum interconnections are obtained, each interconnection including one computing unit from each unit group. Different computing units within the same unit group are located in different interconnections, where T ≤ Sum ≤ Q × T. The switch enables direct communication between any two computing units in different unit groups via a single hop. Compared to a hardware topology formed by computing units within individual unit groups, communication between computing units uses a shorter path, reducing communication latency and increasing system bandwidth. Simultaneously, since all computing units in each unit group are connected to different switches, the pressure of cross-group communication is distributed, resulting in a more balanced load. Traffic between different switches can be distributed more evenly, reducing the load bottleneck of a single switch and optimizing network bandwidth utilization.
[0029] In one implementation, when Sum = T, the system includes T interconnections with the same bandwidth. Each interconnection is configured such that Q ports of each computing unit to be interconnected with a switch are configured into the same interconnection. That is, the Q ports of each computing unit to be interconnected with a switch are connected to the same switch and configured to communicate with the same computing unit in the same interconnection.
[0030] Optionally, when Q=1, the system includes T switches, with a single switch connecting one computing unit in each unit group, and different switches connecting different computing units in the same group; by configuring each switch to configure all computing units connected to the current switch as an interconnection relationship, the number of interconnected computing units can be maximized.
[0031] As an example, please refer again Figure 1 It shows that the system includes T switches SW = {SW1, SW2, ..., SW3}. f ,…,SW T}, SW f Let SW be the f-th switch, where f ranges from 1 to T. Each switch connects to one computing unit in each unit group. f Connect to Gp i,r PRC i,r,t And SW f Not with Gp i,r The q-th computational unit PRC i,r,q Connection, PRC i,r,t Not with the h-th switch in SW h The connection is established, where q and f both range from 1 to T, and t ≠ q, f ≠ h. Where PRC i,r,q and SW h None of them are shown in the figure.
[0032] In one implementation, the number of switches can be less than T when the switches themselves have sufficient interconnect ports. For example, each switch connects two computing units in each unit group, and after configuring the switches, the two computing units in each unit group connected to the same switch are located in two interconnections. Alternatively, when the switches themselves have a full complement of interconnect ports sufficient to support all computing units, the number of switches can be 1, connecting all computing units through a single switch, but the computing units in each interconnection still belong to different unit groups.
[0033] In one implementation, when Sum = Q × T, the system includes Q × T interconnections with the same bandwidth. Each interconnection is configured by the SW to assign Q ports of each computing unit to be interconnected with the switch to Q different interconnections. That is, Q ports of the same computing unit are located in different interconnections; in other words, Q ports of the same unit are interconnected with different computing units, allowing each computing unit to simultaneously connect to other computing units through Q different interconnections, reducing the number of hops in inter-group communication.
[0034] In one implementation, when T < Sum < Q × T, the system includes at least two interconnection relationships with different bandwidths: the number of ports configured for each computing unit in the first interconnection relationship is different from the number of ports configured for each computing unit in the second interconnection relationship. For example, the number of ports configured for each computing unit in the first interconnection relationship is 1, and the number of ports configured for each computing unit in the second interconnection relationship is Q-1, that is, one port of each computing unit is configured for the first interconnection relationship, and the remaining Q-1 ports are configured for the second interconnection relationship. Other types of interconnection relationships also fall within the scope of protection of this invention.
[0035] In this system, different computing units within the same group are connected to different switches. When a switch fails, the impact is limited to the computing units connected to that switch, without affecting all computing units in the group, thus providing better fault isolation capabilities.
[0036] In one implementation, when each unit group is independently numbered, all computing units connected to the same switch may have the same or different numbers. Taking a unit group comprising 4 computing units as an example, with the computing units in each group numbered 1-4, the first connection method connects units with the same number: (e.g., SW) f For example, SW f Connect the computational unit numbered 1 in each group. The second connection method connects different numbers: using SW... f For example, SW fConnect the computation unit numbered 1 in group 1, the computation unit numbered 2 in group 2, the computation unit numbered 3 in group 3, and the computation unit numbered 4 in group 4, respectively. Alternatively, connect two groups with the same numbering and the other two groups with different numbers.
[0037] In one implementation, when each x unit groups are independently numbered, the numbering of all computing units connected to the same switch includes x preset numbering options. Taking x=2 as an example, when the computing unit is an OAM-compliant module, the maximum number of the computing unit is 8. Each two unit groups form a numbering cycle, therefore the two numbers connected to the same switch within these two unit groups are different. For example: There are 4 unit groups, with each two unit groups forming a numbering cycle, where SW... f For example, in the first unit group of the first numbering cycle, the computing unit numbered 2 and the computing unit numbered 4 in the second unit group are connected to SW. f The computational unit numbered 2 in the third unit group of the second numbering cycle and the computational unit numbered 4 in the fourth unit group are also connected to SW. f .
[0038] In one embodiment, the connection between the switch and the computing unit is an optical fiber or a cable; other types of connections used for data communication also fall within the scope of protection of this invention.
[0039] In one embodiment, the connection between computing units is an optical fiber or a cable, but other types of connections used for data communication also fall within the scope of protection of this invention.
[0040] In one embodiment, the interconnection between computing units within a unit group is a ring topology or a mesh topology, and other types of topologies also fall within the protection scope of this invention.
[0041] In one implementation, K servers form a server group, each consisting of G servers. For each server in each server group, there are E unit groups within each server. The computing units in all unit groups are interconnected in a hardware interconnect topology, where G and E are both greater than or equal to 1.
[0042] In one embodiment, the hardware interconnect topology is a ring topology, a Dragonfly topology, or a mesh topology; other types of topologies also fall within the scope of protection of this invention.
[0043] In one implementation, the current switch is configured to allow target computing units connected to the current switch to directly access other computing units with different numbers than the target computing unit. The logical topology formed by this configuration rule is suitable for MOE scenarios.
[0044] In one implementation, when the numbers of all computing units connected to the same switch include x preset numbering options, the SW is configured... f , making PRC i,r,t Direct access to all other connections SW f And the number is the same as PRC i,r,t Different computing units. Or by configuring SW. f , making PRC i,r,t Direct access to all other connections SW f The computing unit. Similarly, the configuration rules for other switches are the same as those for SW. f The same applies. The logical topology formed by this configuration rule is suitable for MOE scenarios.
[0045] The solution provided in Embodiment 1 of this invention enables any two computing units in different unit groups to communicate directly via a single hop through a switch. Compared to a hardware topology composed of computing units within each unit group, communication between computing units follows a shorter path, reducing communication latency and increasing system bandwidth while mitigating bandwidth limitations imposed by intermediate nodes. Furthermore, the switch distinguishes between Sum interconnection relationships between different computing units within a unit group. Communication between T computing units within a unit group is completed through a fixed topology within the unit group. With the same interconnection scale and bandwidth, compared to traditional northbound networks, this reduces the number of switch interconnection ports, further saving on the interconnection costs of large clusters.
[0046] In order to obtain an interconnection system that can both expand the interconnection topology and improve the system's fault redundancy, the present invention provides Embodiment 2.
[0047] Example 2
[0048] Embodiment 2 of the present invention provides an interconnection system based on a two-layer topology. The system includes at least one switch and at least one interconnection topology, each interconnection topology including N physical topologies and multiple logical topologies.
[0049] In this context, physical topology refers to direct point-to-point connections made through actual physical wiring.
[0050] In one implementation, the physical topology connections are cables or optical fibers.
[0051] The switch described in Embodiment 1 of the present invention is also applicable to Embodiment 2 of the present invention, and will not be described again.
[0052] Furthermore, each physical topology includes: L unit groups, intra-group physical connections, and inter-group physical connections. Each unit group includes T computing units. Each intra-group physical connection connects the computing units within each unit group. Each inter-group physical connection connects computing units belonging to different unit groups, and the two computing units connected by each inter-group physical connection are different.
[0053] In one embodiment, the unit group includes an intra-group topology, which is the topology formed by the T computing units within the unit group through intra-group physical connections. The topology can be a ring topology, a mesh topology, or a star topology, and other types of topologies also fall within the protection scope of this invention.
[0054] In one embodiment, the unit group includes an inter-group topology. This inter-group topology is formed by T computing units within a unit group being connected point-to-point with T computing units in adjacent groups via physical connections. Other types of topologies also fall within the scope of this invention. Adjacent groups are physically adjacent unit groups. When the physical topologies are distributed sequentially in the same direction, adjacent groups are physically adjacent unit groups, with the first and last unit groups considered as adjacent groups.
[0055] Furthermore, the switch is used to connect all computing units in the interconnect topology.
[0056] The switch enables direct communication between any two computing units in different unit groups via a single hop, eliminating the need for the physical topology formed by the computing units within each unit group. Furthermore, since all computing units in each unit group are connected to different switches, the pressure of cross-group communication is distributed.
[0057] It should be noted that the connection relationship between the switch and each unit group in Embodiment 1 of the present invention also applies to Embodiment 2 of the present invention, and will not be repeated here.
[0058] Furthermore, multiple logical topologies are obtained by configuring the connectivity between computing units in N physical topologies through a switch. Each logical topology includes two unit groups and the logical paths between them. The i-th logical topology includes T logical paths, and each logical path is used to connect two computing units belonging to the two unit groups. Different logical paths connect different computing units.
[0059] Each logical path includes the connection between two computing units and the switch, as well as the logical path configured inside the switch.
[0060] The interconnection between two unit groups via logical topology expands the GPU interconnection scale. This method allows unit groups to be connected logically, thus increasing the GPU interconnection scale. For example, when each unit group has four computing units, interconnecting two unit groups via logical topology results in an interconnection scale of eight GPUs. Adding a third unit group to this eight-GPU interconnection scale and interconnecting it with any of the existing units creates an interconnection scale of twelve GPUs. Adding a fourth unit group results in an interconnection scale of sixteen GPUs, and so on.
[0061] Furthermore, the physical topology connections and logical paths in the logical topology are redundant. When one or more connections fail, data exchange is performed through the logical path, reducing the risk of network outages and improving overall system reliability. When an inter-group connection failure is detected in the physical topology, the system automatically switches to the logical path, simplifying the fault recovery process and reducing the need for manual intervention. This system provides effective countermeasures for potential technical failures, ensuring the continuity and integrity of data transmission.
[0062] In one implementation, the two unit groups in each logical topology belong to two unit groups in different physical topologies. The unit groups are sequentially connected via logical topologies to form an interconnection structure with unit groups as basic nodes and logical topologies as basic paths, thereby infinitely expanding the interconnection scale of GPUs. For example, a physical topology may include two unit groups, each containing four GPUs. In this physical topology, one unit group connects to a unit group in the previous physical topology, and the other unit group connects to a unit group in the next physical topology. In another implementation, when the number of unit groups in a physical topology is greater than or equal to two, two logical topologies are included between the two physical topologies, and these two logical topologies connect two unit groups belonging to the same physical topology.
[0063] In one implementation, two unit groups belonging to different physical topologies within each logical topology are physically adjacent. Specifically, when the physical topologies are sequentially distributed in the same direction, the first and last physical topologies in the sequential distribution are considered physically adjacent. This method allows adjacent physical topologies to be sequentially connected through logical topologies, forming a ring interconnect structure with unit groups as basic nodes and logical topologies as basic paths. This allows for unlimited expansion of the GPU interconnect scale and enables data transmission in both clockwise and counterclockwise directions, improving system throughput.
[0064] In one implementation, the two unit groups in each logical topology belong to two unit groups in different physical topologies. When an inter-group physical connection fails, the two target unit groups connected by the failed inter-group physical connection are identified, and the two sets of logical paths corresponding to the two target unit groups are obtained. All inter-group physical connections of all unit groups connected by the two sets of logical paths are configured as unavailable, and data exchange is switched to a new topology formed by the remaining physical and logical topologies. That is, system functionality is restored by combining the unfailed physical and logical topologies to form a new fused topology for exchanging data. Other methods of replacing failed physical connections with logical paths also fall within the protection scope of this invention.
[0065] In one implementation, the two unit groups in each logical topology are two unit groups in the same physical topology.
[0066] In one implementation, in an interconnect system, the following two configurations coexist: two unit groups in each logical topology belong to two unit groups in different physical topologies; and two unit groups in each logical topology are two unit groups in the same physical topology.
[0067] In one implementation, two unit groups within the same physical topology are physically adjacent in each logical topology. Specifically, when unit groups are sequentially distributed in the same direction, the first and last unit groups in the sequentially distributed group are considered adjacent groups.
[0068] In one implementation, in an interconnect system, the following two configurations coexist: two unit groups belonging to different physical topologies in each logical topology are physically adjacent; and two unit groups belonging to the same physical topology in each logical topology are physically adjacent.
[0069] In one implementation, the two unit groups in each logical topology are two unit groups in the same physical topology. When an inter-group physical connection fails, the two target unit groups connected by the failed inter-group physical connection are identified, all inter-group physical connections between the two target unit groups are configured as unavailable, and data exchange is switched to a new topology formed by the remaining physical and logical topologies. Other methods of replacing failed physical connections through logical paths also fall within the protection scope of this invention.
[0070] For ease of understanding, let's take a system with two physical topologies, each containing four unit groups, and each unit group containing four computing units. The computing units are OAM-compliant modules, and every two computing units are independently numbered. This example illustrates the new topology configuration when a physical connection between groups fails. Please refer to [link to relevant documentation]. Figure 2 , Figure 2The interconnection system shown includes two physical topologies and two logical topologies. The two physical topologies are: Physical Topology 1 and Physical Topology 2. Each physical topology includes four unit groups: Unit Group 1, Unit Group 2, Unit Group 3, and Unit Group 4. Each unit group has intra-group physical connections and inter-group physical connections for point-to-point interconnections between unit groups. Unit Group 1 and Unit Group 2 form an indivisible basic unit group, which is independently numbered. Similarly, Unit Group 3 and Unit Group 4 form an indivisible basic unit group, which is also independently numbered. Figure 2 As shown, the computing units in Unit 1 and Unit 2 are numbered S0-S7, and the computing units in Unit 3 and Unit 4 are also numbered S0-S7. In the first and second physical topologies, computing units numbered S1 and S0 are connected to the first switch SW1, computing units numbered S2 and S4 are connected to the second switch SW2, computing units numbered S3 and S5 are connected to the third switch SW3, and computing units numbered S6 and S7 are connected to the fourth switch SW4. After configuration via SW1-SW4, two logical topologies are obtained, as shown... Figure 2 As shown, the first logical topology includes the second unit group in the first physical topology and the first unit group in the second physical topology, and the second logical topology includes the fourth unit group in the first physical topology and the third unit group in the second physical topology. The first and second logical topologies are identical, both including logical paths connecting S1 and S0, S2 and S4, S3 and S5, and S6 and S7. When at least one inter-group connection between the second and fourth unit groups in the first physical topology fails—for example, the inter-group physical connection between S1 in the second unit group and S3 in the fourth unit group of the first physical topology fails—all inter-group physical connections between the second and fourth unit groups in the first physical topology are configured as unavailable, and simultaneously all inter-group physical connections between the first and third unit groups in the second physical topology are also configured as unavailable. See also... Figure 3 The system is switched from its original topology to a loop topology composed of the remaining physical and logical topologies. Specifically, this loop topology consists of the logical path between the second unit group of the first physical topology and the first unit group of the second physical topology, the physical connections between the first, second, fourth, and third unit groups of the second physical topology, the logical path between the third unit group of the second physical topology and the fourth unit group of the first physical topology, and the physical connections between the fourth, third, first, and second unit groups of the first physical topology. This loop topology facilitates data exchange, restores system functionality, and improves system stability and reliability.
[0071] In one implementation, when a computing unit fails, a degraded topology is obtained: if L1 unit groups are bound as an indivisible basic unit group, when a computing unit fails, the basic unit group containing the failed computing unit and its associated inter-group physical connections and logical topology are configured as unavailable, resulting in at least one original degraded topology; the switch is then reconfigured based on the original degraded topology to obtain multiple degraded logical topologies, and the original degraded topology and the degraded logical topologies constitute the degraded topology. By combining logical topology with physical topology, the original degraded topology is fully utilized to obtain the degraded topology, preventing the entire interconnection system from becoming unusable due to the failure of a single computing unit.
[0072] Please see Figure 3 In another implementation scenario, suppose Figure 3 There are no faulty inter-group physical connections in the physical topology, that is, all inter-group physical connections between the second and fourth unit groups in the first physical topology, which are configured as unavailable, and all inter-group physical connections between the first and third unit groups in the second physical topology. These are all loop topologies configured with unit groups as the basic unit due to design requirements. Two unit groups are bound together as an indivisible basic unit group. In the first physical topology, one basic unit group consists of the first and second unit groups, and the other consists of the third and fourth unit groups; in the second physical topology, one basic unit group consists of the first and second unit groups, and the other consists of the third and fourth unit groups. When S0 of the first unit group in the second physical topology fails, the basic unit containing S0 and its associated inter-group physical connections and logical topology are configured as unavailable, resulting in an original degraded topology. This original degraded topology includes the first physical topology and the target basic unit group consisting of the third and fourth unit groups in the second physical topology connected by logical topology. The resulting original degraded topology is as follows: Figure 4 As shown. Please refer to [the original text]. Figure 4 The switch is reconfigured based on the original degraded topology. The logical topology between the first physical topology and the target basic unit group is configured to be unavailable, and a logical topology is configured within the target basic unit group. This results in logical connections between S2 of the third unit group and S4 of the fourth unit group in the second physical topology, S6 of the third unit group and S7 of the fourth unit group, S1 of the third unit group and S0 of the fourth unit group, and S5 of the third unit group and S3 of the fourth unit group. This yields a degraded topology composed of the physical topology of the target basic unit group and the configured internal logical topology. In other words, the degraded topology includes two parts: one is the first physical topology, and the other is the degraded topology of the target basic unit group.
[0073] In one implementation, when the logical topology fails, the failed logical topology is configured as unavailable, resulting in at least one degraded topology. Alternatively, when the logical topology fails, it is configured as unavailable, and another logical path is configured to replace the function of the failed logical path. Wherein, when one or more logical paths fail, the logical topology to which the failed logical path belongs is considered to be faulty. Through the flexible and configurable nature of the logical topology, redundancy is achieved between logical topologies, improving fault tolerance and system flexibility, and ensuring the stability and reliability of the interconnected system.
[0074] Please see Figure 3 In another implementation scenario, Figure 3 The middle is a loop topology with unit groups as the basic unit, and in Figure 3 There are no faulty inter-group physical connections; that is, the inter-group physical connections set as unavailable are not actually faulty, but configured to be unavailable according to design requirements. When the logical path between S1 in the second unit group of the first physical topology and S0 in the first unit group of the second physical topology is faulty, then the logical topology between the second unit group of the first physical topology and the first unit group of the second physical topology is considered faulty, and the faulty logical topology is configured to be unavailable. Also, please refer to... Figure 5 In order to restore Figure 3 In a ring topology with unit groups as the basic unit, the logical topology between the fourth unit group of the first physical topology and the third unit group of the second physical topology is also configured to be unavailable, while the inter-group physical connection between the second and fourth unit groups of the first physical topology is configured to be available, and the inter-group physical connection between the first and third unit groups of the second physical topology is configured to be available. The logical topology between the first unit group of the first physical topology and the second unit group of the second physical topology, as well as the logical topology between the third unit group of the first physical topology and the fourth unit group of the second physical topology, are obtained through the switch configuration, forming a new ring topology.
[0075] Other methods that require configuring switches to degrade the interconnection topology to a smaller scale due to computing unit failure or connection failure fall within the protection scope of this invention.
[0076] The system provided in Embodiment 2 of this invention includes a physical topology and a logical topology configured through a switch. The connections in the physical topology and the logical paths in the logical topology cooperate with each other, enabling two unit groups to be interconnected through the logical topology, thereby expanding the scale of GPU interconnection. Simultaneously, because the logical topology has the characteristic of flexible and configurable communication paths, when one or more connections fail, data exchange is performed through the pathways provided by the logical paths, thus exhibiting high fault tolerance, reducing the risk of network interruption, and improving the overall reliability of the system.
[0077] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0078] While specific embodiments of the invention have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of the invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the invention. The scope of this invention is defined by the appended claims.
Claims
1. An interconnection system based on a two-layer topology, characterized in that, The system includes at least one switch and at least one interconnect topology, each interconnect topology including N physical topologies and multiple logical topologies; Each physical topology includes: There are L unit groups, and each unit group includes T computing units; Intra-group physical connections are used to connect the various computing units within each unit group; Inter-group physical connections are used to connect computing units belonging to different unit groups. Each inter-group physical connection connects two different computing units. The switch is used to connect all computing units in the interconnect topology; The multiple logical topologies are obtained by configuring the switches. Each logical topology includes two unit groups and the logical paths between them. The i-th logical topology includes T logical paths. Each logical path is used to connect two computing units belonging to the two unit groups. Different logical paths connect different computing units.
2. The system according to claim 1, characterized in that, The two unit groups in each logical topology belong to two unit groups in different physical topologies.
3. The system according to claim 2, characterized in that, In each logical topology, two unit groups belonging to different physical topologies are physically adjacent; the adjacency refers to the physical proximity when the physical topologies are distributed in the same direction, and the first and last unit groups in physical position are considered to be adjacent.
4. The system according to claim 1, characterized in that, The two unit groups in each logical topology are the same two unit groups in the same physical topology.
5. The system according to claim 4, characterized in that, In each logical topology, two unit groups belonging to the same physical topology are physically adjacent.
6. The system according to claim 1, characterized in that, When a physical connection between groups fails, configure a new topology: The system identifies the two target unit groups connected by the faulty inter-group physical connection, configures all inter-group physical connections between the two target unit groups as unavailable, and switches to a new topology formed by the remaining physical and logical topologies for data exchange.
7. The system according to any one of claims 1-5, characterized in that, When a computing unit fails, obtain the degenerate topology: When a computing unit fails, if L1 units are bound as an indivisible basic unit group, the basic unit group containing the failed computing unit is configured as unavailable. At least one degraded physical topology is obtained by degrading the physical topology of the remaining unfailed basic unit groups. The switch is then reconfigured according to the degraded physical topology to obtain multiple degraded logical topologies. The degraded physical topology and logical topology constitute the degraded topology.
8. The system according to any one of claims 2-3, characterized in that, When the logical topology fails, the failed logical topology is configured to be unavailable, resulting in at least one degraded topology.
9. The system according to claim 1, characterized in that, The unit group includes an intra-group topology, which is the topology formed by the T computing units within the unit group through intra-group physical connections. The topology can be a ring topology, a mesh topology, or a star topology.
10. The system according to claim 1, characterized in that, The unit group includes an inter-group topology, which is formed by the T computing units in the unit group being connected point-to-point with the T computing units in the adjacent group through inter-group physical connections.