Multiprocessor interconnection system, processor group and processor
By using the first and second type of intergroup interfaces to connect processors in a multiprocessor interconnect system, the problems of high concurrency and high computing requirements in large-scale models are solved, and the infinite expansion and efficient communication of the system are achieved.
Patent Information
- Application Number
- PCT/CN2024/135133
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-04
- Filing Date
- 2024-11-28
- Publication Date
- 2025-06-05
AI Technical Summary
In large-scale model scenarios, due to frequent communications of parameter updates and gradient synchronization, existing multiprocessor interconnect systems cannot meet the needs of high concurrency and high computing capabilities.
A multi-processor interconnection system is adopted to connect the next group of processors through the first type of intergroup interface and the previous group of processors through the second type of intergroup interface, ensuring that each processor only connects the upper and lower groups through two interfaces, achieving infinite expansion.
The system maximizes concurrency, balances communication pressure, expands the system's computing power, and meets the high concurrency and high computing needs in large-scale model scenarios.
Smart Images

Figure CN2024135133_05062025_PF_FP_ABST
Abstract
Description
Multi-processor interconnection system, processor group and processor Technical Field
[0001] The present invention relates to the technical field of chip design, and in particular to a multi-processor interconnection system, a processor group and a processor. Background Art
[0002] In the field of chip design, multi-processor interconnection is one of the core strategies for expanding system computing power. Multi-processor interconnection not only improves the system's computing power and throughput, but also increases the system's flexibility, enabling the chip to cope with more complex and varied computing tasks. Currently, common multi-processor interconnection topologies include full interconnection topology and ring topology. When faced with insufficient system computing power, the interconnection topologies in multiple servers can be connected in sequence through a single cable to achieve the purpose of expanding computing power. The technology of interconnecting the interconnection topologies in servers using cables can refer to the patent with publication number CN110461111A. Furthermore, considering that the increase in the number of hops will directly affect communication efficiency and system performance, in order to solve this problem, the interconnection system provided by patent publication number CN112416850A includes multiple groups of processors, the processors in each group are interconnected in pairs, and the different processors in any two groups are connected one-to-one. When exchanging data, the system can first complete the data exchange within the group, and then complete the data exchange between groups through a single hop, thereby reducing the number of communication hops between processors. Patent publication number CN106776014B provides a system that includes multiple interconnected cube topologies with corresponding vertices. Each vertex in the cube topology represents a computing unit. When exchanging data, node data within the same plane of the cube topology is first synchronized, followed by node data within opposite planes within the same cube, and finally, node data within different cubes. Each additional cube topology in this system increases the number of hops by 1, achieving a control hop count of log2N. To reduce the number of communication hops during AllReduce collective communication, patent publication number CN115129655A provides a system that includes multiple groups. Each processing unit in each group is connected to another processing unit within the same group via two communication links and to a processing unit in another group via one communication link. When performing AllReduce collective communication based on this system, each processing unit first collects data from all other processing units in the group via one hop. Each processing unit then exchanges data with the other groups via one hop. Finally, each processing unit broadcasts the obtained data to the other processors in the group, completing the AllReduce collective communication and reducing the number of communication hops.
[0003] The above expansion solutions are all designed to reduce the number of hops in data exchange. However, when facing large model scenarios, since large models involve a large number of parameter updates and gradient synchronization, which require frequent communication, their traffic characteristics are explosive large traffic. When expanding the computing power of the system, concurrency needs to be considered. The above solutions cannot meet the requirements of large model scenarios for high concurrency and high computing power. Summary of the Invention
[0004] In view of the above technical problems, the technical solution adopted by the present invention is:
[0005] In a first aspect, the present invention provides a multi-processor interconnection system, the system comprising N groups of processors, each group comprising M processors, N>2, M>1; wherein the jth processor P of the i-th group ij Including the first type of inter-group interface and the second type of inter-group interface, 1≤i≤N, 1≤j≤M; P ij The first type of inter-group interface and P Next(i)Nmap(i,j) The second type of inter-group interface has a bandwidth of B1 ij connection; where P Next(i)Nmap(i,j) is the Nmap(i,j)th processor of the Next(i)th group, the function value of Next(i) is the next group of the i-th group; Nmap(i,j) is the processor of Next(i)th group and P ij Mapping function of interconnected processors; P ij The second type of inter-group interface and P Prev(i)Pmap(i,j) The first type of inter-group interface has a bandwidth of B2 ij connection; where P Prev(i)Pmap(i,j) is the Pmap(i,j)th processor of the Prev(i)th group, the function value of Prev(i) is the previous group of the i-th group; the Pmap(i,j) is the processor of the Pmap(i,j)th group in Prev(i). ij The mapping function of the interconnected processors; where B1 ij and B2 ij match.
[0006] In a second aspect, the present invention provides a processor group, the processor group includes M processors, M>1; wherein the jth processor P j Including the first type of inter-group interface and the second type of inter-group interface, 1≤j≤M; P j The first type of inter-group interface is used with bandwidth B1 j Connect the Nmap(i,j)th processor of the Next(i)th group, where i is P j The number of the processor group where it is located, the function value of Next(i) is P j The next group of processors in which Nmap(i,j) is located; Nmap(i,j) is the next group of processors in Next(i) jMapping function of interconnected processors; P j The second type of inter-group interface is used with bandwidth B2 j Connect the Pmap(i,j)th processor of the Prev(i)th group, where the function value of the Prev(i)th group is P j The previous group of processors where the processor is located; the Pmap(i,j) is the one in Prev(i) and P j Mapping function of interconnected processors; where B1 and B2 match, B1 = ∑ j=1 M B1 j , B2=∑ j=1 M B2 j .
[0007] In a third aspect, the present invention provides a processor, which includes a first type of inter-group interface and a second type of inter-group interface; the first type of inter-group interface is used to connect to the second type of inter-group interface of the Nmap(i,j)th processor of the Next(i)th group with a bandwidth of BW1, wherein i is the number of the processor group where the current processor is located, and the function value of Next(i) is the next group of the processor group where the current processor is located; the Nmap(i,j) is the mapping function of the processors interconnected with the current processor in Next(i); the second type of inter-group interface is used to connect to the first type of inter-group interface of the Pmap(i,j)th processor of the Prev(i)th group with a bandwidth of BW2, wherein the function value of Prev(i) is the previous group of the processor group where the current processor is located; the Pmap(i,j) is the mapping function of the processors interconnected with the current processor in Prev(i); wherein BW1 and BW2 match.
[0008] The present invention has at least the following beneficial effects:
[0009] Embodiments of the present invention provide a multi-processor interconnection system, a processor group, and a processor. These utilize a first-type inter-group interface to connect to the processors of the next group, and utilize a second-type inter-group interface to connect to the first-type inter-group interface of the previous group. This not only expands the computing power of the system, maximizes concurrency, and balances communication pressure, but also each processor in the system connects to the upper and lower groups via only two interfaces, without being limited by the number of interconnection interfaces of the processor itself, thus achieving unlimited expansion. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0011] FIG1 is a schematic diagram of the structure of a multi-processor interconnection system when a mapping function is single-valued, provided by a first embodiment of the present invention;
[0012] FIG2 is a schematic diagram of the topological structure of a 16-card interconnection system provided in Embodiment 1 of the present invention;
[0013] FIG3 is a schematic diagram of the topological structure of a 32-card interconnection system provided in Embodiment 1 of the present invention;
[0014] FIG4 is a schematic diagram of the topological structure of a 32-card interconnection system when the mapping function provided by the first embodiment of the present invention is a vector;
[0015] FIG5 is a schematic diagram of the topological structure of the OAM protocol supporting a maximum of 8 processors interconnected;
[0016] FIG6 is a schematic diagram of a first interconnection system provided by Embodiment 2 of the present invention;
[0017] FIG7 is a schematic diagram of a second interconnection system provided by Embodiment 2 of the present invention;
[0018] FIG8 is a schematic diagram of a third interconnection system provided by the second embodiment of the present invention;
[0019] FIG9 is a schematic diagram of a fourth interconnection system provided in the second embodiment of the present invention. DETAILED DESCRIPTION
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0021] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present invention have the same meanings as commonly understood by those skilled in the art.
[0022] Example 1
[0023] A first embodiment of the present invention provides a multi-processor interconnection system, which includes N groups of processors, each group including M processors, where N>2 and M>1.
[0024] In one embodiment, the processor is a graphics processing unit (GPU), a tensor processing unit (TPU), a neural network processing unit (NPU), or a general-purpose computing on a graphics processor (GPGPU). Other types of processors also fall within the scope of protection of the present invention.
[0025] In another embodiment, the processor is a module or expansion card that complies with specified standards.
[0026] In one embodiment, the processor is an OAM module, which is a modular accelerator that complies with the OAM (Optical Accelerator Module) standard. The OAM module includes a processor, such as a GPU or NPU. Processors of other standard module types also fall within the scope of protection of the present invention. The OAM module can have different names. For example, when the OAM module includes a GPU, it can also be referred to as an OAM GPU module. Of course, in future communication systems, the OAM module may also have different names, which is not limited here. When the processor is an OAM module, the OAM module supports seven interconnection interfaces (serdes) for interconnection or expansion between OAM modules. Each interface can be flexibly configured to have different numbers of channels. For example, each interface can be configured as a group of 16 high-speed serial channels, or as two groups of 8 high-speed serial channels each, or as four groups of 4 high-speed serial channels each, or as 16 groups of 1 high-speed serial channel each. Other configuration methods of configuring interfaces also fall within the scope of protection of the present invention.
[0027] In one embodiment, the protocol supported by the interface is NVLink (NVIDIA Link), MetaXLink (MetaX Link), UCIe standard (Universal Chiplet Interconnect Express) or PCIe (Peripheral Component Interconnect Express). Other interface protocols for achieving high-speed interconnection between processors also fall within the scope of protection of the present invention.
[0028] In one embodiment, the processor is an expansion card that complies with the PCIe (Peripheral Component Interconnect Express) standard, and the expansion card includes a processor, such as a GPU or an NPU. Expansion cards that comply with other standards also fall within the scope of protection of the present invention.
[0029] In one embodiment, each group of processors in the system is integrated into the same board or server. Alternatively, each basic group is integrated into the same board or server, and each basic group includes N1 bound, inseparable processor groups, where N1 ≤ N. For example, N1 = 2, that is, two unit groups form a basic group, which are integrated into the same board. The physical distance between processors integrated into the same board is short, and the communication delay is low. Interconnected systems obtained by configuring processors on other types of devices or using other configuration strategies also fall within the scope of protection of the present invention.
[0030] Furthermore, the jth processor P of the i-th group ij Including the first type of inter-group interface and the second type of inter-group interface, 1≤i≤N, 1≤j≤M; P ij The first type of inter-group interface and P Next(i)Nmap(i,j) The second type of inter-group interface connection; where P Next(i)Nmap(i,j) is the Nmap(i,j)th processor of the Next(i)th group, the function value of Next(i) is the next group of the i-th group; Nmap(i,j) is the processor of Next(i)th group and P ij Mapping function of interconnected processors; P ij The second type of inter-group interface and P Prev(i)Pmap(i,j) The first type of inter-group interface connection; where P Prev(i)Pmap(i,j) is the Pmap(i,j)th processor of the Prev(i)th group, the function value of Prev(i) is the previous group of the i-th group; the Pmap(i,j) is the processor of the Pmap(i,j)th group in Prev(i). ij Mapping functions of interconnected processors.
[0031] Among them, the "first category" in the first category inter-group interface and the "second category" in the second category inter-group interface do not indicate order or weight. For the convenience of understanding and description, "first category" is used to distinguish the interface used to connect to the previous group of processors, and "second category" is used to distinguish the interface used to connect to the next group of processors. Similarly, for the convenience of understanding and description, "inter-group" is used to distinguish that the processor connected by the interface is another processor from a different group, and "intra-group" is used to distinguish that the processor connected by the interface is another processor from the same group. Similarly, the processor group number and the processor number are not built-in numbers of the processor group or the processor itself. For the convenience of understanding and description, the processor group number is only used to distinguish different processor groups, and the processor number is only used to distinguish different processors.
[0032] In one embodiment, the connection line between the first type inter-group interface and the second type inter-group interface is an electrical connection line or an optical fiber connection line. Other types of connection lines for interconnection interfaces for connecting two processors across groups fall within the scope of protection of the present invention.
[0033] In one embodiment, when the distance between groups is short, the connection between the first type inter-group interface and the second type inter-group interface is an electrical connection line. When the distance between groups is long or in a high-bandwidth and low-latency application scenario, the connection between the first type inter-group interface and the second type inter-group interface is an optical fiber connection line.
[0034] Other types of connections or other types of connection media between the first type inter-group interface and the second type inter-group interface also fall within the protection scope of the present invention.
[0035] In one embodiment, the connection between the first type inter-group interface and the second type inter-group interface supports bidirectional transmission and unidirectional transmission.
[0036] Considering each group as a node, the N nodes corresponding to the N groups of processors form a chain topology or a ring topology.
[0037] In one embodiment, when the topology formed by the interconnected system is a ring topology with each group being one node, Next(i) and Prev(i) respectively satisfy: if i = N, then Next(i) = 1; otherwise, Next(i) = i + 1; if i = 0, then Prev(i) = N; otherwise, Prev(i) = i - 1.
[0038] In one embodiment, when the topology formed by the interconnected system is a chain topology with each group being a node, Next(i) and Prev(i) respectively satisfy: if i = N, then Next(i) = NULL, where NULL is empty; otherwise, Next(i) = i + 1; if i = 0, then Prev(i) = NULL; otherwise, Prev(i) = i - 1.
[0039] In order to adapt to different computing scenarios, other types of interconnection topologies formed by adjusting the relationship satisfied by Next(i) and Prev(i) also fall within the protection scope of the present invention.
[0040] In one embodiment, the P ij The first type of inter-group interface and P Next(i)Nmap(i,j) The second type of inter-group interface has a bandwidth of B1 ij connection; the P ij The second type of inter-group interface and P Prev(i)Pmap(i,j) The first type of inter-group interface has a bandwidth of B2 ij Connection; among them, B1 i and B2 i Match, B1 i =∑ j=1 M B1 ij , B2 i =∑ j=1 M B2 ij .
[0041] Among them, for each group of processors, the total bandwidth B1 of each group of processors connected to the next group i and the total bandwidth B2 of the previous group i Matching. It should be noted that bandwidth matching refers to bandwidth equality, but not strict equality. Bandwidth matching is considered achieved when the difference between the two bandwidths is within an acceptable range. Matching the total bandwidth between each processor group and its connected upstream and downstream groups enables efficient data transmission between groups, reducing data transmission delays and system bottlenecks, and improving overall system efficiency.
[0042] In one embodiment, B1 ij and B2 ij Matching. This strategy ensures that each processor's bandwidth matches that of its upstream and downstream groups. This bandwidth matching strategy allows each processor to fully utilize its allocated bandwidth resources, allowing data to flow smoothly between processors connected sequentially via inter-group interfaces. This prevents insufficient bandwidth on a single processor from becoming a bottleneck for inter-group interconnection, thereby avoiding bandwidth waste and ensuring smooth and efficient data transmission. When adding more processors, simply ensure that the bandwidth of the newly added processors matches that of the existing processors, making it easier for the system to scale to larger-scale parallel computing.
[0043] In one embodiment, the kth processor P of the i-th group ik The first type of inter-group interface and P Next(i)Nmap(i,k) The second type of inter-group interface has a bandwidth of B1 ikConnect the B1 ik and B1 ij The relationship between them satisfies: B1 ik =B1 ij , where the value of Nmap(i,k) is the same as that of Next(i) in P ik For interconnected processors, 1≤k≤M and k≠j. This strategy, based on bandwidth matching between a single processor connecting the upper and lower groups, further aligns the bandwidths of different processors in the same group. This eliminates the need to consider bandwidth matching when designing data transmission paths, simplifying the design logic of data transmission paths. The bandwidth of all data transmission paths in this interconnected system is the same, allowing data to flow smoothly throughout the system without causing a bottleneck for the entire system due to bandwidth issues with a single processor, thereby improving transmission efficiency. This bandwidth matching strategy is more efficient during data synchronization and collective communication, enabling processors to collaborate efficiently and jointly complete complex parallel computing tasks.
[0044] The value of the mapping function represents the number of the processor in the previous or next group to which the current processor is connected. The mapping functions Nmap(i,j) and Pmap(i,j) have the same properties. Nmap(i,j) is used as an example to illustrate the properties of the mapping function. For each pair of input parameters (i,j), the mapping function Nmap(i,j) returns a unique output value. When the value of Nmap(i,j) is determined, the value of (i,j) is also unique. This means that there is a one-to-one mapping between the value of the mapping function Nmap(i,j) and the input parameters (i,j). The advantages of each processor in each processor group being connected to the corresponding processor in the upper and lower groups through a mapping function include: first, each processor in the same group has the ability to independently send data to the corresponding processor in the upper and lower groups connected to it, and multiple data transmission operations can be performed almost simultaneously in the same time period without waiting for each other. This data transmission method logically constitutes concurrency, improving the overall processing capability and throughput of the system; second, each processor can directly transmit data to the corresponding processor in the upper and lower groups, reducing the number of data transmission hops and potential delays, while also facilitating data transmission between processors, and helping to achieve more efficient load balancing; third, the number of inter-group interfaces occupied by each processor in the system is fixed and will not increase with the increase in the number of processors, and can be expanded indefinitely.
[0045] Different interconnection topologies are formed when Nmap(i,j) and Pmap(i,j) satisfy different conditions. When Nmap(i,j) and Pmap(i,j) are single values, each processor connects to a processor in each of the upper and lower groups. This scheme offers a simple connection structure, making it easy to implement and maintain, and scalable to larger scales when more processors are needed. When Nmap(i,j) and Pmap(i,j) are vectors, each processor connects to multiple processors in each of the upper and lower groups. Furthermore, when Nmap(i,j) and Pmap(i,j) are single values, various connection methods exist.
[0046] In one embodiment, when the function values of Nmap(i,j) and Pmap(i,j) are single values, Nmap(i,j) and Pmap(i,j) satisfy: Nmap(i,j) = Pmap(i,j) = Nmap(j), and when j1≠j2, Nmap(j1)≠Nmap(j2), that is, each processor is connected to a processor in the upper and lower groups, and the connected processors have the same number. Since the connection relationship is only related to the processor number and has no direct connection with the group number in the system, the complexity of system design and the logic of task allocation are simplified. In addition, when the system is expanded or reconfigured, it is only necessary to maintain the consistency of the processor number without considering the boundaries between groups.
[0047] In one embodiment, Nmap(j) = Pmap(j) = j, where j is the logical number of the processor. That is, each processor in the system is connected to a processor in the upper and lower groups with the same number. This framework is more intuitive and easy to understand, facilitating resource management and task allocation while also reducing development and maintenance complexity.
[0048] Adjusting the value of the mapping function so that the processor can connect to other interconnection systems with different numbers also falls within the protection scope of the present invention.
[0049] For ease of understanding, a multi-processor interconnection system is described by taking Nmap(j)=Pmap(j)=j and the topology formed by the interconnection system as a ring topology with each group as a node as an example. Please refer to FIG1 , which shows a schematic diagram of a multi-processor interconnection system. The system includes N groups of processors {G1, G2, ..., G i-1 ,G i ,G i+1 ,…,G N}, G i is the i-th group, and the value of i ranges from 1 to N. G i Including M processors {P i1 ,P i2 ,…,Pij ,…P iM}, P ij G i The jth processor of G, the value of j ranges from 1 to M. i Each processor and G i+1 The topological connection formed by each processor, G i and G i+1 M links are formed between i1 ,Nlink i2 ,…,Nlink ij ,…,Nlink iM}, where Nlink ij P ij The first group interface and G i+1 The jth processor P (i+1)j The connection between the second set of interconnection interfaces. That is, Nlink i1 P i1 The first type of inter-group interface and G i+1 The first processor P (i+1)1 The connection between the second type of inter-group interfaces, Nlink i2 P i2 The first group interface and G i+1 The second processor P (i+2)2 The connection between the second group of interconnection interfaces, and so on, Nlink iM P iM The first group interface and G i+1 The Mth processor P (i+1)M The connection between the second group of interconnection interfaces. i+1 The M lines connected make each processor in each group connected to G i+1 The corresponding processors in G are interconnected one to one. i and G i-1 M lines are formed between {Plink i1 ,Plink i2 ,…,Plink ij ,…,Plink iM}, where Plink ij P ij The first group interface and G i-1 The jth processor P (i+1)j Similarly, N groups of processors are interconnected to form a ring topology.
[0050] In one embodiment, M = 4. In another embodiment, M = 2 or M = 6. Other values of M greater than 1 also fall within the scope of protection of the present invention.
[0051] In one embodiment, M=4, and N is a multiple of 2. When M=4 and N=4, a 16-card interconnected system is achieved. When M=4 and N=6, a 24-card interconnected system is achieved. When M=4 and N=8, a 32-card interconnected system is achieved. When M=4 and N=10, a 40-card interconnected system is achieved, and so on. The scale of the multi-processor interconnected system can be expanded as needed.
[0052] As an example, for ease of understanding, based on Figure 1, taking N=M=4 as an example, please refer to Figure 2, which shows a schematic diagram of the topology of a 16-card interconnection system. In this topology, there are four groups of processors: G1-G4. Each group includes four processors. Each processor includes a first-type inter-group interface and a second-type inter-group interface. Each group is connected to the upper and lower groups by four lines. Taking the second group G2 as an example, there are four first lines connecting G2 and G3: G2's P 21 Connect G3's P 31 , G2's P 22 Connect G3's P 32 , G2's P 23 Connect G3's P 33 , G2's P 24 Connect G3's P 34 There are also four second lines connecting G2 and G1: G2's P 21 Connect G1's P 11 , G2's P 22 Connect G1's P 12 , G2's P 23 Connect G1's P 13 , G2's P 24 Connect G1's P 14 Similarly, four groups of processors are interconnected to form a ring topology.
[0053] As another example, for easier understanding, based on Figure 2, with N=8 and M=4, refer to Figure 3, which shows a schematic diagram of the topology of a 32-card interconnect system. This topology includes eight processor groups: G1-G8. Similar to the topology shown in Figure 2, each group includes four processors. The eight processor groups are connected sequentially. Similarly, each processor group has four connections to the next group via the first-type inter-group interconnect interface and four connections to the previous group via the second-type inter-group interface. As can be seen, even if the interconnect system is doubled in size, each processor still uses two interfaces for interconnection. In other words, expanding the size of the interconnect system does not consume the number of processor interfaces, so theoretically, the interconnect system can be expanded infinitely. Furthermore, the connection strategy for the added processor groups is the same as that for the existing processor groups in the interconnect system, as long as the corresponding mapping function is met. This makes the expansion logic simple and easy to expand.
[0054] In one embodiment, Nmap(i, j) and Pmap(i, j) are both vectors containing L elements, each of which represents the number of a processor. The L elements have different values, where 1 < L ≤ M. This means that each processor is connected to L processors in the upper and lower groups. Compared to an interconnect system where each processor occupies two inter-group interfaces, the inter-group bandwidth is increased by L times, and the transmission rate is improved by L times.
[0055] As an example, for ease of understanding, taking N=8, M=4 and L=2 as an example, refer to FIG4 . FIG4 is based on FIG3 , and each processor has two inter-group interfaces added, that is, each processor includes two first-type inter-group interfaces and two second-type inter-group interfaces. 22 and the third processor P 23 As an example, the interconnection relationship between the two first-class inter-group interfaces and two different processors in the next group is explained. 22 Through two first-class inter-group interfaces, they are connected to G3's P 32 and the G3's fourth processor, P 34 Connection; P of G2 23 Through two first-class inter-group interfaces, they are connected to G3's P 33 and P 31 Connect. Then use P in G2 22 and P 23 As an example, the interconnection relationship between the two first-class inter-group interfaces and two different processors in the next group is explained. 22 Through two second-type inter-group interfaces, they are connected to the P 12 and P 14 Connect. G2 P 23Through two second-type inter-group interfaces, they are connected to G1's P 13 and P 11 Connection. That is, P of G2 22 and P 23 The different processors in the previous group and the next group are connected through the corresponding inter-group interfaces. Similarly, the first processor P of G2 21 and the fourth processor P 24 Each processor connects to the remaining processors in one group and the next. In other words, each processor occupies four inter-group interfaces, two of which connect to the next group and two to the previous group, resulting in eight connections between two adjacent groups. While utilizing two inter-group interfaces to achieve unlimited expansion, two more are used to connect to the processors in the previous and next groups, effectively occupying four interconnection interfaces per processor. This achieves unlimited interconnection and triples bandwidth and transmission rates compared to an interconnection system that utilizes only two interconnection interfaces.
[0056] In one embodiment, N>2 and N≠4.
[0057] In one embodiment, P ij The first type of inter-group interface is only with P Next(i)Nmap(i,j) The second type of inter-group interface connection; P ij The second type of inter-group interface is only with P Prev(i)Pmap(i,j) The first type of inter-group interface connection. That is, for P ij For example, the first-type inter-group interface is only used to connect to the second-type inter-group interface of the next group, and the second-type inter-group interface is only used to connect to the second-type inter-group interface of the previous group.
[0058] In one embodiment, P ij It also includes M-1 intra-group interfaces, where P ij The fth intra-group interface PI ij,f The hth intra-group interface PI with the rth processor of the ith group ir,h Connect, 1≤r≤M and j≠r, 1≤r≤M-1.
[0059] In one embodiment, the topology formed by interconnecting the M processors in the i-th group through the intra-group interface is a ring topology, a full interconnection topology, or a tree topology. Other topologies also fall within the protection scope of the present invention.
[0060] In one embodiment, data exchange may be more frequent within a group than between groups, so P ij The total bandwidth of all interfaces within a group is greater than the total bandwidth of all interfaces between groups. The total bandwidth of all interfaces between groups is the bandwidth B1 of the first interface between groups. ij Bandwidth B2 of the interface between the second type of groupsij sum.
[0061] A multi-processor interconnection system is provided in a first embodiment of the present invention. Each processor in the system is directly connected to corresponding processors in upper and lower groups, and the number of processors in each group is the same. This design helps achieve load balancing. In collective communication, each processor can simultaneously process data and communicate with other processors, improving overall concurrency. When more processors are needed to improve system performance, new processor groups can be simply added. This design makes the system scalable, thus meeting the needs for high computing power and high concurrency.
[0062] It should be noted that, in addition to being applicable to collective communication, the embodiments of the present invention also support other types of communication modes, such as point-to-point communication, broadcast communication, etc. Other types of communication modes also fall within the scope of protection of the present invention.
[0063] Example 2
[0064] Based on the same inventive concept as the first embodiment, the second embodiment of the present invention provides a multi-processor interconnection system that solves the preferred embodiment of how to expand the number of processor interconnections while complying with the original protocol. Since the OAM protocol supports a topology with a maximum of eight processors interconnected, as shown in Figure 5, including eight processors OAM0-OAM7, any two processors can be directly interconnected, i.e., point-to-point interconnection. The interconnection protocol is the OAM protocol. In the OAM protocol, when the addresses of the source and destination processors are determined, the route between them is also uniquely determined. Because the route between the source and destination processors is fixed in hardware and cannot be changed, the route is {source processor address, source processor interface index number, destination processor address, destination processor interface index number}. However, in the training and inference of large models, full point-to-point interconnection between processors is not required; only collective communication modes such as Allreduce and Alltoall are required. The Allreduce mode collects and aggregates data from each graphics card, and then distributes the aggregated results to each graphics card. The Alltoall mode distributes the data of each node to each graphics card, while also collecting data from each graphics card. Therefore, processors can be interconnected point-to-point or through forwarding from other processors. The topology in Figure 5 is referred to as the original topology and will not be further explained. To address the issue of how to expand the number of processor interconnections while complying with the original protocol, the topology formed by this interconnection system has single-valued values for Nmap(i, j) and Pmap(i, j), and the topology formed by the M processors in group i interconnected via intra-group interfaces is a fully interconnected topology. That is, the system includes N groups of processors, each group containing M processors. The processors in each group are uniformly distributed, and each group of processors comprises two interconnection structures: an intra-group interconnection structure and an inter-group interconnection structure. The intra-group interconnection structure is a fully interconnected topology in which the M processors are fully connected point-to-point. The inter-group interconnection structure includes M ring-shaped interconnection structures, in which each processor is uniformly distributed within each group. That is, each processor is connected to a processor in the previous group and another processor in the next group.
[0065] As a preferred embodiment, the software address remapping table is searched to obtain the remapping address of each processor address in the interconnected system, and the route corresponding to the remapping address is obtained. The software address remapping table includes each processor address and its remapping address. When the processor address is less than 7, that is, when it is any one of S0-S7, the remapping address of the processor address is itself. When the processor address is greater than 7, the remapping address of the processor address is the modulo of the processor address. The software address remapping table is searched according to the addresses of the source processor and the destination processor to obtain the remapping address of each processor. All processor addresses of N groups of processors are remapped to S0-S7 through the software address remapping table, so that the processor addresses during data transmission conform to the interconnection routes fixed in the hardware. While complying with the original protocol, the number of processor interconnections is expanded.
[0066] Please refer to Figures 6, 7, 8, and 9. Figures 6-8 provide three types of interconnection topologies, differing in the interface numbers used for interconnection between processors. Figure 9 provides a 32-card interconnection topology. It should be noted that in the interconnection systems provided by Figures 6, 7, 8, and 9, the processor numbers of the j-th processor in the i-th group connected to the previous and next groups satisfy Nmap(i, j) and Pmap(i, j), and Nmap(i, j) and Pmap(i, j) are single values, meaning that each processor is connected to a processor in each of the upper and lower groups.
[0067] Please refer to Figure 6, which shows the first type of interconnection topology. Please refer to Figure 6 again, the interconnection system includes 4 groups, each group has 4 processors, and the entire interconnection system has a total of 16 processors. The first group G1 includes P 11 、P 12 、P 13 and P 14 , the second group G2 includes P 21 、P 22 、P 23 and P 24 , the third group G3 includes P 31 、P 32 、P 33 and P 34 , the fourth group G4 includes P 41 、P 42 、P 43 and P 44. Each processor includes 7 interfaces, and each interface in each processor corresponds to a unique index number. The index numbers of different processors are independent, that is, the index numbers of interfaces in different processors are all from the first interface index number inf1 to the seventh interface index number inf7. In the 16-processor interconnection topology in Figure 6, each processor occupies a total of 5 interfaces. The same three interfaces are occupied within each group of processors as intra-group interfaces to achieve full interconnection. In Figure 6, the intra-group interfaces inf1, inf2 and inf3 are used to achieve full interconnection within the group to form a fully interconnected topology. Inter-group interconnection is achieved by using two interfaces of the remaining interfaces of the interconnected processors as the first type of inter-group interface and the second type of inter-group interface respectively. The processors of the next group are connected through the first type of inter-group interface, and the processors of the previous group are connected through the second type of inter-group interface to form a ring structure.
[0068] Among them, P 11 、P 21 、P 31 and P 41 The distribution position in each group of processors is the same, and P 11 、P 21 、P 31 and P 41 In the interconnection structure formed by interconnection, the connection between the index numbers of the interfaces is as follows: 11 inf7 connection P 21 inf7,P 21 inf4 connection P 31 inf4,P 31 inf7 connection P 41 inf7,P 41 inf6 connection P 11 The inf6 of P is connected end to end to form a ring structure. 12 、P 22 、P 32 and P 42 The distribution position in each group of processors is the same, and P 12 、P 22 、P 32 and P 42 In the interconnection structure formed by interconnection, the connection between the index numbers of the interfaces is as follows: 12 inf7 connection P 22 inf7,P 22 inf4 connection P 32 inf4,P 32 inf7 connection P 42 inf7,P 42 inf6 connection P 12 The inf6 of P is connected end to end to form a ring structure.14 、P 24 、P 34 and P 44 The distribution position in each group of processors is the same, and P 14 、P 24 、P 34 and P 44 The interconnection forms an interconnection structure, and the connection between the index numbers of the interfaces is as follows: 14 inf7 connection P 24 inf7,P 24 inf6 connection P 34 inf6,P 34 inf7 connection P 44 inf7,P 44 inf4 connection P 14 The inf4 of P is connected end to end to form a ring structure. 13 、P 23 、P 33 and P 43 The distribution position in each group of processors is the same, and P 13 、P 23 、P 33 and P 43 The interconnection forms an interconnection structure, and the connection between the index numbers of the interfaces is as follows: 13 inf7 connection P 23 inf7,P 23 inf6 connection P 33 inf6,P 33 inf7 connection P 43 inf7,P 43 inf4 connection P 14 inf4, connected end to end to form a ring structure.
[0069] Based on Figure 6, the software address remapping table is used to remap P 11 -P 44 This 16-bit processor address is remapped to P 11 -P 24 The interconnection routing between any two processors in Figure 6 is identical to the original interconnection routing for the eight processors in Figure 1. This allows for expansion of interconnected processors without changing the routing fixed in the hardware.
[0070] It should be noted that the index numbers inf7, inf6, and inf4 of the interfaces forming the ring-shaped inter-group interconnection structure in FIG6 are, in the order of the ring structure, inf7, inf6, inf7, and inf4, or inf7, inf4, inf7, and inf6. Alternatively, an equivalent implementation of the index numbers of the interconnected interfaces may be inf6, inf7, inf6, and inf4, or inf6, inf4, inf6, and inf7. Alternatively, inf5 may be used to replace the index number of any interface in the ring structure. For example, if the interconnection structure uses inf5 instead of inf7, the interface index numbers are inf5, inf6, and inf4.
[0071] Please refer to Figure 7, which shows the second type of interconnection topology, which also includes 4 groups of processors. In Figure 7, the 16 processors in the interconnection topology occupy a total of 5 interfaces. The processors within each group of processors are fully interconnected within the group. The interconnection between groups is to form a ring structure by connecting the processors with the same distribution position in each group of processors through the remaining interfaces. The 4 groups of processors are: the first group G1 includes P 11 、P 12 、P 13 and P 14 , the second group G2 includes P 21 、P 22 、P 23 and P 24 , the third group G3 includes P 31 、P 32 、P 33 and P 34 , the fourth group G4 includes P 41 、P 42 、P 43 and P 44 In each group of processors, full interconnection is achieved through inf3, inf4, inf5 and inf6. Inter-group interconnection is formed by jumping between the remaining interfaces of the interconnected processors to form a ring structure. 11 、P 21 、P 31 and P 41 The interconnection forms a ring structure, and the connection between the interface index numbers is as follows: 11 inf7 connection P 21 inf7,P 21 inf6 connection P 31 inf6,P 31 inf7 connection P 41 inf7,P 41 inf5 connection P11 The inf5 of each group of processors is connected end to end to form a ring structure. 12 、P 22 、P 32 and P 42 The interconnection forms a ring structure, and the connection between the interface index numbers is as follows: 12 inf7 connection P 22 inf7,P 22 inf4 connection P 32 inf4,P 32 inf7 connection P 42 inf7,P 42 inf5 connection P 12 The inf5 of each group of processors is connected end to end to form a ring structure. 14 、P 24 、P 34 and P 44 The interconnection forms a ring structure, and the connection between the interface index numbers is as follows: 14 inf7 connection P 24 inf7,P 24 inf5 connection P 34 inf5,P 34 inf7 connection P 44 inf7,P 44 inf6 connection P 14 The inf6 of each group of processors are connected end to end to form a ring structure. 13 、P 23 、P 33 and P 43 The interconnection forms a ring structure, and the connection between the interface index numbers is as follows: 13 inf7 connection P 23 inf7,P 23 inf5 connection P 33 inf5,P 33 inf7 connection P 43 inf7,P 43 inf4 connection P 13 inf4, connected end to end to form a ring structure.
[0072] Based on Figure 7, the software address remapping table is used to remap P 11 -P 44 This 16-bit processor address is remapped to P 11 -P 24The 8-bit processor address achieves the purpose of expanding the interconnected processors without changing the fixed routing in the hardware.
[0073] Please refer to Figure 8, which shows the third type of interconnection topology, which also includes 4 groups of processors. In Figure 8, the 16 processors in the interconnection topology occupy a total of 6 interfaces. The processors within each group of processors are fully interconnected within the group. The interconnection between groups is to form a ring structure by connecting the processors with the same distribution position in each group of processors through the remaining interfaces. The 4 groups of processors are: the first group G1 includes P 11 、P 12 、P 13 and P 14 , the second group G2 includes P 21 、P 22 、P 23 and P 24 , the third group G3 includes P 31 、P 32 、P 33 and P 34 , the fourth group G4 includes P 41 、P 42 、P 43 and P 44 In each group of processors, full interconnection is achieved through inf2, inf6, inf7 and inf4. Inter-group interconnection is achieved by forming a ring structure between the remaining interfaces of the interconnected processors. 11 、P 21 、P 31 and P 41 The interconnection forms a ring structure, and the interconnection interfaces are as follows: The connection between the interface index numbers is as follows: 11 inf4 connection P 21 inf6,P 21 inf1 connection P 31 inf1,P 31 inf5 connection P 41 inf5,P 41 inf1 connection P 11 The inf1 of each group of processors is connected end to end to form a ring structure. 12 、P 22 、P 32 and P 42 The interconnection forms a ring structure, and the connection between the interfaces is as follows: 12 inf1 connection P 22 inf1,P 22 inf4 connection P 32 inf6,P32 inf1 connection P 42 inf5,P 42 inf5 connection P 12 The inf5 of each group of processors is connected end to end to form a ring structure. 14 、P 24 、P 34 and P 44 The interconnection forms a ring structure, and the connection between the interface index numbers is as follows: 14 inf1 connection P 24 inf1,P 24 inf5 connection P 34 inf5,P 34 inf1 connection P 44 inf1,P 44 inf4 connection P 14 The inf6 of each group of processors are connected end to end to form a ring structure. 13 、P 23 、P 33 and P 43 The interconnection forms a ring structure, and the connection between the interface index numbers is as follows: 13 inf5 connection P 23 inf5,P 23 inf1 connection P 33 inf1,P 33 inf4 connection P 43 inf6,P 43 inf1 connection P 13 inf1, connected end to end to form a ring structure.
[0074] The extended topology system provided by the 16 processors provided in FIG8 also needs to remap the P 11 -P 44 The address is remapped to P 11 -P 24 In this way, the purpose of expanding the interconnected processors is achieved without changing the fixed routing in the hardware.
[0075] The equivalent topology system of the 16-processor extended topology system shown in Figures 6, 7, and 8 also includes an interconnection structure formed by swapping the positional relationships between the processor groups in the extended topology system. Extended topology systems that achieve the same result as the hardware-fixed routing determined by the OAM protocol by changing the index number fall within the scope of protection of the present invention.
[0076] Please refer to Figure 9, which illustrates a fourth type of interconnect topology. This interconnect topology further expands upon the interconnect topology shown in Figure 7, enabling the interconnection of 32 processors, equivalent to the interconnection of four groups of the original interconnect topology. The interconnection of 32 processors consists of eight interconnect structures, each of which is fully interconnected internally. Processors in the same position are interconnected sequentially to form a ring-shaped interconnect structure. The address remapping table used is the same as the software address remapping table shown in Figure 7.
[0077] As a preferred embodiment, the extended topology systems shown in Figures 6 and 8 can be further expanded with reference to the extended topology system shown in Figure 9. Figures 6, 7, and 8 can all be expanded multiple times in a manner that fully interconnects within a group and interconnects co-located processors between groups to form a ring structure, thereby achieving the interconnection of M*N processors.
[0078] As a preferred embodiment, when the source processor performs hardware address remapping to make it conform to the verification rule of the destination processor, or the destination address performs hardware address remapping, hardware address remapping is added to the hardware path between the source processor and the destination processor. As an example, when P 11 Expectation and P 41 inf3 interconnection communication, but in fact P 11 With P 41 Inf6 interconnection communication, at this time, the interconnected inf6 can be remapped to inf3 through mapping address remapping.
[0079] As a preferred embodiment, when 16 processors are interconnected, P 11 -P 24 After the address is remapped to itself, P 31 -P 44 The address after remapping is the current processor address minus 8.
[0080] In summary, the interconnection system provided in the second embodiment of the present invention includes N groups of processors, each group of processors includes M processors, and the positions of the processors in each group of processors are uniformly distributed. Each group of processors includes a two-layer interconnection structure: an intra-group interconnection structure and an inter-group interconnection structure. This allows the resulting extended topology system to achieve the expansion of the number of interconnected processors while complying with the original protocol.
[0081] Example 3
[0082] Based on the same inventive concept as a multi-processor interconnection system, the second embodiment of the present invention further provides a processor group. The processor group includes M processors, M>1; wherein the jth processor P j Including the first type of inter-group interface and the second type of inter-group interface, 1≤j≤M; P jThe first type of inter-group interface is used with bandwidth B1 j Connect the Nmap(i,j)th processor of the Next(i)th group, where i is P j The number of the processor group where it is located, the function value of Next(i) is P j The next group of processors in which Nmap(i,j) is located; Nmap(i,j) is the next group of processors in Next(i) j Mapping function of interconnected processors; P j The second type of inter-group interface is used with bandwidth B2 j Connect the Pmap(i,j)th processor of the Prev(i)th group, where the function value of the Prev(i)th group is P j The previous group of processors where the processor is located; the Pmap(i,j) is the one in Prev(i) and P j Mapping function of interconnected processors; where B1 and B2 match, B1 = ∑ j=1 M B1 j , B2=∑ j=1 M B2 j .
[0083] Among them, when P j When the processor group is the i-th group, P j The representation can also be recorded as P ij , P ij The same as in the first embodiment, no further details will be given. j When the processor group is group i, B1 j Can also be recorded as B1 ij , B2 j Can also be recorded as B2 ij , no more details.
[0084] It should be noted that the relevant technical features of the processor group are the same as those of the multi-processor interconnection system provided in the first embodiment, and will not be described in detail.
[0085] Example 4
[0086] Based on the same inventive concept as a multi-processor interconnection system, embodiment three of the present invention also provides a processor, which includes a first type of inter-group interface and a second type of inter-group interface; the first type of inter-group interface is used to connect to the second type of inter-group interface of the Nmap(i,j)th processor of the Next(i)th group with a bandwidth of BW1, wherein i is the number of the processor group where the current processor is located, and the function value of Next(i) is the next group of the processor group where the current processor is located; the Nmap(i,j) is the mapping function of the processors interconnected with the current processor in Next(i); the second type of inter-group interface is used to connect to the first type of inter-group interface of the Pmap(i,j)th processor of the Prev(i)th group with a bandwidth of BW2, wherein the function value of Prev(i) is the previous group of the processor group where the current processor is located; the Pmap(i,j) is the mapping function of the processors interconnected with the current processor in Prev(i); wherein BW1 and BW2 match.
[0087] If the current processor is the jth processor of the i-th group, BW1 can also be recorded as B1 ij , BW2 can also be written as B2 ij , no more details.
[0088] It should be noted that the relevant technical features of this processor are the same as those of the multi-processor interconnection system provided in Example 1 and will not be described in detail.
[0089] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0090] Although some specific embodiments of the present invention have been described in detail by way of example, it should be understood by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present invention. It should also be understood by those skilled in the art that various modifications may be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.
Claims
1. A multi-processor interconnection system, characterized in that: The system includes N groups of processors, each group includes M processors, N>2, M>1; Among them, the jth processor P of the i-th group ij Including the first type of inter-group interface and the second type of inter-group interface, 1≤i≤N, 1≤j≤M; P ij The first type of inter-group interface and P Next(i)Nmap(i,j) The second type of inter-group interface connection; where P Next(i)Nmap(i,j) is the Nmap(i,j)th processor of the Next(i)th group, the function value of Next(i) is the next group of the i-th group; Nmap(i,j) is the processor in Next(i) that is connected to P ij mapping functions of interconnected processors; P ij The second type of inter-group interface and P Prev(i)Pmap(i,j) The first type of inter-group interface connection; where P Prev(i)Pmap(i,j) is the Pmap(i,j)th processor of the Prev(i)th group, the function value of Prev(i) is the previous group of the i-th group; Pmap(i,j) is the processor in Prev(i) that is consistent with Pmap(i,j) ij Mapping functions of interconnected processors.
2. The system according to claim 1, characterized in that The Next(i) and Prev(i) respectively satisfy: If i=N, then Next(i)=1; otherwise, Next(i)=i+1; If i=0, then Prev(i)=N; Otherwise, Prev(i)=i-1.
3. The system according to claim 1 or 2, characterized in that: The function values of Nmap(i,j) and Pmap(i,j) are single values respectively.
4. The system according to claim 3, characterized in that The Nmap(i,j) and Pmap(i,j) satisfy: Nmap(i,j)=Pmap(i,j)=Nmap(j).
5. The system according to claim 4, characterized in that The Nmap(j)=Pmap(j)=j.
6. The system according to claim 5, characterized in that N≠4。 7. The system according to claim 6, characterized in that The M processors in the i-th group form a ring topology, a fully interconnected topology, or a tree topology.
8. The system according to claim 7, characterized in that The P ij It also includes M-1 intra-group interfaces, among which P ij The fth intra-group interface PI ij,f The hth intra-group interface PI with the rth processor of the ith group ir,h Connect, 1≤r≤M and j≠r, 1≤r≤M-1.
9. The system according to claim 1, characterized in that The Nmap(i,j) and Pmap(i,j) are both vectors containing L elements, each element is the number value of a processor, and the values of the L elements are different, wherein 1<L≤M.
10. The system according to claim 1, characterized in that in: The P ij The first type of inter-group interface and P Next(i)Nmap(i,j) The second type of inter-group interface has a bandwidth of B1 ij connect; The P ij The second type of inter-group interface and P Prev(i)Pmap(i,j) The first type of inter-group interface has a bandwidth of B2 ij connect; Among them, B1 i and B2 i Match, B1 i =∑ j=1 M B1 ij , B2 i =∑ j=1 M B2 ij .
11. The system according to claim 10, characterized in that The Next(i) and Prev(i) respectively satisfy: If i=N, then Next(i)=1; otherwise, Next(i)=i+1; If i=0, then Prev(i)=N; otherwise, Prev(i)=i-1; And the B1 ij and B2 ij match.
12. The system according to claim 11, characterized in that The kth processor P of the i-th group ik The first type of inter-group interface and P Next(i)Nmap(i,k) The second type of inter-group interface has a bandwidth of B1 ik Connect the B1 ik and B1 ij The relationship between them satisfies: B1 ik =B1 ij , where the value of Nmap(i,k) is the value of Next(i) that matches P ik Interconnected processors, 1≤k≤M and k≠j.
13. The system according to claim 1, characterized in that The Next(i) and Prev(i) respectively satisfy: If i=N, Next(i)=NULL, where NULL is empty; otherwise, Next(i)=i+1; If i=0, Prev(i)=NULL; otherwise, Prev(i)=i-1.
14. The system according to claim 1, characterized in that in: P ij The first type of inter-group interface is only with P Next(i)Nmap(i,j) The second type of inter-group interface connection; P ij The second type of inter-group interface is only with P Prev(i)Pmap(i,j) The first type of inter-group interface connection.
15. A processor group, characterized in that: The processor group includes M processors, M>1; Among them, the jth processor P j Including the first type of inter-group interface and the second type of inter-group interface, 1≤j≤M; P j The first type of inter-group interface is used with bandwidth B1 j Connect to the Nmap(i,j)th processor of the Next(i)th group, where i is P j The number of the processor group where it is located, the function value of Next(i) is P j The next group of processors in which Nmap(i,j) is located; Nmap(i,j) is the next group of processors in Next(i) j mapping functions of interconnected processors; P j The second type of inter-group interface is used for bandwidth B2 j Connect the Pmap(i,j)th processor of the Prev(i)th group, where the function value of the Prev(i)th group is Pmap(i,j) j The previous group of processor groups where the processor group is located; the Pmap(i,j) is the same as P in Prev(i) j mapping functions of interconnected processors; Among them, B1 and B2 match, B1 = ∑ j=1 M B1 j , B2=∑ j=1 M B2 j .
16. A processor, characterized in that: The processor includes a first type of inter-group interface and a second type of inter-group interface; The first type of inter-group interface is used to connect the second type of inter-group interface of the Nmap(i,j)th processor of the Next(i)th group with a bandwidth of BW1, wherein i is the number of the processor group where the current processor is located, the function value of Next(i) is the next group of the processor group where the current processor is located; and Nmap(i,j) is a mapping function of the processors interconnected with the current processor in Next(i); The second type of inter-group interface is used to connect the first type of inter-group interface of the Pmap(i,j)th processor of the Prev(i)th group with a bandwidth of BW2, wherein the function value of Prev(i) is the previous group of the processor group where the current processor is located; and Pmap(i,j) is a mapping function of the processors interconnected with the current processor in Prev(i); Among them, BW1 and BW2 match.
Citation Information
Patent Citations
Multiprocessor interconnection system and communication method thereof
CN112416850A
Multiprocessor interconnection system
CN114968902A
Topology and algorithm for multi-processing unit interconnect accelerator system
CN115129655A
Multiprocessor node interconnection system and server
CN115168279A
Topology-aware provisioning of hardware accelerator resources in a distributed environment
US20190312772A1