Computing architecture for permutation ring network interconnection

By adopting a permutation ring network architecture in the CPU architecture, the high latency and low efficiency problems when sharing data between multiprocessor cores in the prior art are solved, and multi-chip communication with high bandwidth and low latency is realized, and the computing efficiency of neural networks and machine learning applications is improved.

CN113544658BActive Publication Date: 2025-06-24DEGIRUM CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080019800.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-14
Filing Date
2020-03-07
Publication Date
2025-06-24
Estimated Expiration
2040-03-07

AI Technical Summary

Technical Problem

Existing CPU architectures have high latency and inefficiency problems when dealing with neural networks and machine learning applications, especially when sharing data between multiple processor cores, which leads to inefficient communications.

Method used

Using the permutation ring network (PRN) architecture, the permutation ring network interconnect structure between multiple computing slices and memory groups is realized to realize direct point-to-point communication between computing engines, avoiding the complexity of the cache coherence protocol.

Benefits of technology

It realizes multi-chip communication with high bandwidth and low latency, simplifies shared memory and messaging protocols across systems, and improves the computing efficiency of neural networks and machine learning applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113544658B_ABST
    Figure CN113544658B_ABST
Patent Text Reader

Abstract

A computer architecture that uses one or more permutation ring networks to connect multiple computing engines and memory banks to provide a scalable, high-bandwidth, low-latency point-to-point multi-chip communication solution.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims priority to U.S. Patent Application No. 16 / 353,198, filed on March 14, 2019, entitled "Permutated Ring Network Interconnected Computing Architecture", which is incorporated herein by reference. Field of the Invention

[0003] The present invention relates to computer architectures that use one or more permutated ring networks to connect various computing engines. More specifically, the present invention relates to a computing architecture that uses multiple interconnected permutated ring networks that provide a scalable, high - bandwidth, low - latency point - to - point multi - chip communication solution. Background Art

[0004] Figure 1 is a block diagram of a conventional CPU architecture 100 that includes multiple processor chips 101 - 102, chip - to - chip interconnects 105, and DRAM devices 111 - 112. Each of the processor chips 101 and 102 includes multiple processor cores C 01 -C 0N and C 11 -C 1N . Each of the processor cores includes a register file and an arithmetic logic unit (ALU), a level - 1 cache memory L1, and a level - 2 cache memory L2. Each of the processor chips 101 and 102 also includes multiple level - 3 (L3) cache memories 121 and 122, respectively, and cache - coherence interconnect logic 131 and 132.

[0005] Generally, the level - 1 cache memory L1 allows for fast data access (1 - 2 cycles) but is relatively small. The level - 2 cache memory L2 exhibits slower data access (5 - 6 cycles) but is larger than the level - 1 cache memory. Each of the processor cores C 01 -C 0N and C 11 -C 1N has its own dedicated level - 1 cache memory L1 and level - 2 cache memory L2. Each of the processor cores C 01 -C 0N on chip 101 accesses the multiple level - 3 (L3) cache memories 121 through the cache - coherence interconnect logic 131. Similarly, each of the processor cores C 11 -C 1NEach of them accesses multiple level-3 (L3) caches 122 through cache coherence interconnect logic 132. Thus, multiple processor cores on each chip share multiple level-3 (L3) caches on the same chip.

[0006] Processor core C on chip 101 01 -C 0N Each of them accesses DRAM 111 through cache coherence interconnect logic 131. Similarly, processor core C on chip 102 11 -C 1N Each of them accesses DRAM 112 through cache coherence interconnect logic 132.

[0007] Cache coherence interconnect logic 131 ensures that all processor cores C 01 -C 0N see the same data at the same entry of the level-3 (L3) cache 121. Cache coherence interconnect logic 131 resolves any "multiple writer" issues where more than one of the processor cores C 01 -C 0N attempts to update the data stored by the same entry of the level-3 (L3) cache 121. Any processor core C that wants to change the data in the level-3 (L3) cache 121 01 -C 0N must first obtain permission from cache coherence interconnect logic 131. Obtaining this permission undesirably takes a long time and involves the implementation of complex message exchanges. Cache coherence interconnect logic 131 also ensures the consistency of data read / written from / to DRAM 111.

[0008] Cache coherence interconnect logic 132 similarly ensures the consistency of the data stored by L3 cache 122 and the data read / written from / to DRAM112.

[0009] Chip-to-chip interconnect logic 105 enables communication between processor chips 101 - 102, where this logic 105 handles the necessary changes to the protocol across chip boundaries.

[0010] As Figure 1As shown, a conventional CPU architecture 100 implements multiple cache levels (L1, L2, and L3) with a cache hierarchy. The higher-level caches have a relatively small capacity and a relatively fast access speed (such as SRAM), while the lower-level caches have a relatively large capacity and a relatively slow access speed (such as DRAM). A cache coherence protocol is required to maintain data consistency across different cache levels. Due to the use of dedicated primary (L1 and L2) caches, multiple accesses controlled by cache coherence policies, and the required data traversal across different physical networks (e.g., between processor chips 101 and 102), the cache hierarchy makes it difficult to share data among multiple different processor cores C 01 -C 0N and C 11 -C 1N among them.

[0011] Based on the principles of temporal and spatial locality, the cache hierarchy enables higher-level caches to retain displaced cache lines from lower-level caches to avoid long-latency accesses in the case of future data access. However, if there is minimal spatial and temporal locality in the dataset (as is the case with many neural network datasets), the latency increases, the size of useful storage locations decreases, and the number of unnecessary storage accesses increases.

[0012] The hardware of a conventional CPU architecture (such as architecture 100) is optimized for the shared memory programming model. In this model, multiple computing engines communicate via memory sharing using a cache coherence protocol. However, these conventional CPU architectures are not the most efficient way to support the producer-consumer execution model, which is typically implemented by the forward propagation of neural networks (the neural networks exhibit redundant memory read and write operations as well as long latency). In the producer-consumer execution model, the direct transfer of messages from the producer to the consumer is more efficient. In contrast, there is no hardware support for direct communication between processor cores C 01 -C 0N and C 11 -C 1N in the shared memory programming model. The shared memory programming model relies on software to construct a message-passing programming model.

[0013] Each level of the communication channels of a conventional CPU architecture 100 optimized for a shared memory programming model is highly specialized and optimized for the subsystems it serves. For example, there are dedicated interconnect systems: (1) between the data cache and the ALU / register file, (2) between different levels of caches, (3) to the DRAM channels, and (4) in the chip-to-chip interconnect 105. Each of these interconnect systems operates at its own protocol and speed. Thus, communicating across these channels incurs a large amount of overhead. This results in significant inefficiencies when attempting to accelerate tasks that require access to large amounts of data (e.g., large matrix multiplications that use multiple computing engines to perform a task).

[0014] Crossbars and simple ring networks are typically used to implement the above-mentioned dedicated interconnect systems. However, the speed, power efficiency, and scalability of these interconnect structures are limited.

[0015] As described above, conventional CPU architectures have several inherent limitations in the implementation of neural network and machine learning applications. Thus, an improved computing system architecture that can more effectively process data in neural network / machine learning applications is desirable. An improved network topology that can span multiple chips without the need for a cache coherence protocol between multiple chips is also desirable. It would be further desirable if such a multi-chip communication system is easy to scale and can provide communication between many different chips. Summary of the Invention

[0016] Accordingly, the present invention provides a computer architecture including a plurality of computing slices, each computing slice including a plurality of computing engines, a plurality of memory banks, communication nodes, and a first-level interconnect structure. The first-level interconnect structure couples each of the plurality of computing engines, the plurality of memory banks, and the communication nodes. The first-level interconnect enables each computing engine to access each memory bank within the same computing slice. In one embodiment, the first-level interconnect structure is a permutation ring network. However, in other embodiments, the first-level interconnect structure may be implemented using other structures such as crossbars or simple ring networks.

[0017] The computer architecture further includes a second-level interconnect structure including a permutation ring network. As defined herein, a permutation ring network includes a plurality of bidirectional source-synchronous ring networks, each bidirectional source-synchronous ring network including a plurality of data transfer stations. Each communication node of the plurality of computing slices is coupled to one of the data transfer stations in each of the plurality of bidirectional source-synchronous ring networks. The second-level interconnect structure enables access between each of the computing slices coupled to the second-level interconnect structure.

[0018] The computer architecture may further include a memory interface communication node coupled to the second-level interconnect structure, wherein the memory interface communication node is coupled to one of the data transfer stations in each of the plurality of bidirectional source synchronous ring networks of the second-level interconnect structure. In this embodiment, an external storage device (e.g., a DRAM device) is coupled to the memory interface communication node.

[0019] The computer architecture may further include a first network communication node coupled to the second-level interconnect structure, wherein the first network communication node is coupled to one of the data transfer stations in each of the plurality of bidirectional source synchronous ring networks of the second-level interconnect structure. In this embodiment, the first network communication node is coupled to the system-level interconnect structure.

[0020] The system-level interconnect structure may include a plurality of network communication nodes coupled to the third-level interconnect structure. The first of these plurality of network communication nodes may be coupled to the first network communication node. The second of these plurality of network communication nodes may be coupled to the main system processor. The third of these plurality of network communication nodes may be coupled to the system memory. The fourth of these plurality of network communication nodes may be coupled to another second-level interconnect structure, which in turn is coupled to another plurality of computing slices. The third-level interconnect structure may be implemented by a permutation ring network or by other structures such as a crossbar switch or a simple ring network.

[0021] Advantageously, if the first-level interconnect structure, the second-level interconnect structure, and the third-level interconnect structure are all implemented using a permutation ring network, a single message protocol can be used to transmit / receive messages and data across the computer architecture. Address mapping ensures that each of the devices (e.g., computing engines, memory banks, DRAM devices) has a unique address within the computer architecture.

[0022] In a particular embodiment, the second-level interconnect structure and the corresponding plurality of computing slices are fabricated on the same semiconductor chip.

[0023] The present invention will be more fully understood in view of the following description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a block diagram of a conventional computer architecture optimized for a shared memory programming model.

[0025] Figure 2 is a block diagram of a computer architecture according to an embodiment of the present invention that uses a permutation ring network to connect a plurality of computing engines and storage devices.

[0026] Figure 3 is according to an embodiment of the present invention Figure 2 modified view of the computer architecture.

[0027] Figure 4 is of an embodiment according to the present invention Figures 2 to 3 block diagram of a first-level permutation ring network used in the computer architecture of

[0028] Figure 5 is of an embodiment according to the present invention Figure 4 interconnection matrix of four communication channels of the first-level permutation ring network of

[0029] Figure 6 is of an embodiment according to this embodiment defining Figure 4 routing table of traffic flow on the permutation ring network of

[0030] Figure 7 block diagram of a computer architecture according to an alternative embodiment of the present invention

[0031] Figure 8 block diagram of a computer architecture according to another alternative embodiment of the present invention Detailed Description

[0032] The present invention uses a permutation ring network (PRN) architecture to provide a better solution for the interconnection system of a machine learning computing system. The PRN architecture includes a flat memory hierarchy that allows computing engines on different chips (and on the same chip) to communicate directly with each other using a common communication protocol. The interconnection system is non-cache coherent. In one embodiment, the interconnection system uses a single interconnection structure (i.e., multiple permutation ring networks).

[0033] In an alternative embodiment, the PRN structure is used only at specific locations within the interconnection structure (e.g., to connect multiple computing slices on the same chip).

[0034] Figure 2 is a block diagram of a computer system 200 according to an embodiment of the present invention. The computer system 200 includes multiple processor chips 201 - 204, a main system processor 205, a system memory 206, a system-level interconnection chip 210, and DRAM devices 211 - 214. Although Figure 2 only four processor chips 201 - 204 are shown in Figure 2 it should be understood that in other embodiments, the computer system 200 can be easily modified to include other numbers of processor chips. Additionally, although

[0035] In the illustrated embodiment, the processor chip 201 includes four compute slices 1, 2, 3, and 4, and an interconnect structure 21 based on a permutation ring network (PRN). Although four slices are shown in Figure 2 , it should be understood that in other embodiments, other numbers of slices may be included on the processor chip 201. Each slice includes a plurality of compute engines, a plurality of memory banks, communication nodes, and a first-level PRN-based interconnect structure. More specifically, slices 1, 2, 3, and 4 include compute engine groups CE1, CE2, CE3, and CE4, memory bank groups M1, M2, M3, and M4, first-level PRN-based interconnect structures 11, 12, 13, and 14, and communication nodes CN1, CN2, CN3, and CN4, respectively. Each of the compute engine groups CE1, CE2, CE3, and CE4 includes a plurality of compute engines (e.g., local processors). In the illustrated example, each of the compute engine groups CE1, CE2, CE3, and CE4 includes four compute engines. However, it should be understood that in other embodiments, other numbers of compute engines may be included in each compute engine group. Similarly, each of the memory bank groups M1, M2, M3, and M4 includes a plurality of memory banks. In the illustrated example, each of the memory bank groups includes four memory banks. However, it should be understood that in other embodiments, other numbers of memory banks may be included in each memory bank group. In one embodiment, each memory bank in the memory bank groups M1, M2, M3, and M4 is a static random access memory (SRAM) that enables relatively fast memory access.

[0036] Within each of the compute slices 1, 2, 3, and 4, the corresponding first-level PRN-based interconnect structures 11, 12, 13, and 14 couple the corresponding compute engine groups CE1, CE2, CE3, and CE4 to the corresponding memory bank groups M1, M2, M3, and M4. This allows each compute engine to access each memory bank within the same slice using the corresponding first-level PRN-based interconnect structure. For example, each of the four compute engines in the compute engine group CE1 of compute slice 1 can access each of the four memory banks in the memory bank group M1 of compute slice 1 through the corresponding first-level PRN-based interconnect structure 11 of slice 1.

[0037] The first-level PRN-based interconnect structures 11, 12, 13, and 14 are also coupled to the corresponding communication nodes CN1, CN2, CN3, and CN4 within the corresponding slices 1, 2, 3, and 4. The communication nodes CN1, CN2, CN3, and CN4 are coupled to a second-level PRN-based interconnect structure 21. As described in more detail below, the communication nodes CN1, CN2, CN3, and CN4 transfer messages and data between the corresponding first-level PRN-based interconnect structures 11, 12, 13, and 14 and the second-level PRN-based interconnect structure 21.

[0038] This configuration allows each computing engine on the processor chip 201 to access each memory bank on the processor chip 201 using the level-1 PRN-based interconnect structures 11-14 and, if necessary, the level-2 PRN-based interconnect structure 21. For example, each computing engine in the computing engine group CE1 of computing slice 1 can access each memory bank in the memory bank group M4 of computing slice 4 through a path that includes: the corresponding level-1 PRN-based interconnect structure 11 of computing slice 1, communication node CN1, level-2 PRN-based interconnect structure 21, communication node CN4, and level-1 PRN-based interconnect structure 14 of computing slice 4.

[0039] This configuration also allows each computing engine on the processor chip 201 to communicate with each of the other computing engines on the processor chip 201 using the level-1 PRN-based interconnect structures 11-14 and, if necessary, the level-2 PRN-based interconnect structure 21. For example, each computing engine in the computing engine group CE2 of computing slice 2 can communicate with each computing engine in the computing engine group CE3 of computing slice 3 through a path that includes: the corresponding level-1 PRN-based interconnect structure 12 of slice 2, communication node CN2, level-2 PRN-based interconnect structure 21, communication node CN3, and level-1 PRN-based interconnect structure 13 of slice 3.

[0040] The level-2 PRN-based interconnect structure 21 is also coupled to the external DRAM 211 through the memory interface communication node CN5. This configuration allows each computing engine of the processor chip 201 to access the DRAM 211 using the level-1 PRN-based interconnect structures 11-14 and the level-2 PRN-based interconnect structure 21. For example, each computing engine in the computing engine group CE1 of computing slice 1 can access the DRAM 211 through a path that includes: the corresponding level-1 PRN-based interconnect structure 11 of computing slice 1, communication node CN1, level-2 PRN-based interconnect structure 21, and communication node CN5.

[0041] Figure 2 The computer system 200 also includes a level-3 PRN-based interconnect structure 31, which is fabricated on the system-level interconnect chip 210. The level-3 PRN-based interconnect structure 31 is coupled to a plurality of communication nodes CN 11 -CN 16 on the chip 210. As described in more detail below, the level-3 PRN-based interconnect structure 31 enables the transmission of messages and data between the communication nodes CN 11 -CN 16 . The communication nodes CN 11 、CN 12 、CN 13 and CN14 are respectively connected to processor chips 201, 202, 203, and 204. Communication nodes CN 15 and CN 16 are respectively connected to main system processor 205 and system memory 206.

[0042] System-level interconnect chip 210 allows data and messages to be transferred between each of main system processor 205, system memory 206, and processor chips 201 - 204. More specifically, main system processor 205 can communicate with any one of the computing engines on processor chips 201 - 204 or any one of the memory banks on processor chips 201 - 204. For example, main system processor 205 can access a computing engine in computing engine group CE1 of computing slice 1 (or a memory bank of memory bank group M1 of computing slice 1) through a path that includes: communication node CN 15 , three-level PRN-based interconnect structure 31, network communication node CN 11 and CN6, two-level PRN-based interconnect structure 21, communication node CN1, and one-level PRN-based interconnect structure 11.

[0043] Main system processor 205 can also communicate with any one of DRAMs 211 - 214. For example, main system processor 205 can access DRAM 211 through a path that includes: communication node CN 15 , three-level PRN-based interconnect structure 31, network communication node CN 11 and CN6, two-level PRN-based interconnect structure 21, and communication node CN5. Main system processor 205 can access DRAMs 212 - 214 through similar paths in processor chips 202 - 204 respectively.

[0044] Main system processor 205 can also communicate with system memory 206 through a path that includes: communication node CN 15 , three-level PRN-based interconnect structure 31, and communication node CN 16 .

[0045] In addition, each computing engine on any one of processor chips 201 - 204 can communicate with any one of the computing engines or memory banks on any one of the other processor chips 201 - 204 and DRAMs 211 - 214 connected to these other processor chips.

[0046] According to one embodiment, various memory banks, computing engines, and communication nodes located on processor chips 201 - 204, DRAMs 211 - 214, main system processor 205, system memory 206, and communication node CN on system-level interconnect chip 21011 -CN 16 Are assigned unique system addresses so that each of these system elements can be easily addressed (and thus communicate with) any one of the other system elements.

[0047] Figure 3 is a block diagram of a PRN-based computer system 200, which shows in detail the processor chip 202. Similar elements in the processor chips 201 and 202 are labeled with similar reference numerals. Thus, the processor chip 202 includes computing slices 1', 2', 3', and 4', and the computing slices 1', 2', 3', and 4' respectively include memory bank groups M 1’ 、M 2’ 、M 3’ 、M 4’ 、computing engine groups CE 1’ 、CE 2’ 、CE 3’ and CE 4’ 、a first-level PRN-based interconnect structure 11', 12', 13', and 14', and communication nodes CN 1’ 、CN 2’ 、CN 3’ and CN 4’ . The processor chip 202 also includes a second-level PRN-based interconnect structure 21, a memory interface communication node CN 12 respectively connected to the DRAM 212 and the network communication node CN 5’ and the network communication node CN 6’ .

[0048] This configuration allows each computing engine in the computing engine group CE1 (of the processor chip 201) to access each computing engine in the computing engine group CE 3’ (of the processor chip 202) through a path that includes: a first-level PRN-based interconnect structure 11, a communication node CN1, a second-level PRN-based interconnect structure 21, network communication nodes CN6 and CN 11 、a third-level PRN-based interconnect structure 31, network communication nodes CN 12 and CN 6’ 、a second-level PRN-based interconnect structure 21', communication nodes CN 3’ and a first-level PRN-based interconnect structure 13'. Similarly, each computing engine in the computing engine group CE1 (of the processor chip 201) can use the same path to access each memory bank in the memory bank group M 3’ (of the processor chip 202).

[0049] This configuration also allows each computing engine of each processor chip to access the DRAM connected to other processor chips. For example, each computing engine in the computing engine group CE1 of slice 1 (of processor chip 201) can access DRAM 212 (connected to processor chip 202) through a path that includes: the corresponding level-1 PRN-based interconnect structure 11 of slice 1, communication node CN1, level-2 PRN-based interconnect structure 21, communication nodes CN6 and CN 11 , level-3 PRN-based interconnect structure 31, communication nodes CN 12 and CN 6’ , level-2 PRN-based interconnect structure 21' and communication node CN 5’ .

[0050] As described above, the PRNA interconnected computer system 200 has a three-level hierarchy, including a slice level, a chip level, and a system level, where each level is defined by its physical construction boundary.

[0051] The slice level represented by computing slices 1-4 (and computing slices 1'-4') is the basic building block of the computer system 200. Each computing slice itself can be implemented as a small machine learning processor through a bridge between the main system processor 205 and the level-1 PRN-based interconnect structure.

[0052] The chip level represented by processor chips 201-204 is defined by subsystems included on the die, including multiple computing slices and corresponding level-2 PRN-based interconnect structures. Each processor chip can be implemented as a medium-scale machine learning system through a bridge between the main system processor 205 and the level-2 PRN-based interconnect structure.

[0053] The system level including the main system processor 205 is built on multiple processor chips and the system-level interconnect chip 210. Processor chips 201-204 communicate through the system-level interconnect chip 210. The level-3 PRN-based interconnect structure 31 implemented by the system-level interconnect chip 210 advantageously operates with high bandwidth, low latency, and high power efficiency. By using a permutation ring network to implement the level-1 interconnect structure, level-2 interconnect structure, and level-3 interconnect structure, the same communication protocol can be maintained across the entire system. This greatly simplifies the shared memory and message passing protocols across the system. As described above, the computer system 200 enables any computing engine to access all memory bank groups (e.g., memory bank groups M1-M4 and M1'-M4') and all DRAMs (e.g., DRAMs 211-214) in the system 200 through the PRN-based interconnect structure. Therefore, the computer system 200 is a highly flexible shared memory computing system.

[0054] In addition, all the computing engines of computer system 200 can communicate directly with each other through a PRN-based interconnection structure. Advantageously, no software support is required to translate messages exchanged between computing engines in different computing slices or chips, resulting in an efficient message-passing computing system.

[0055] For implementing Figure 2 and Figure 3 The PRN-based interconnection structures for level-1, level-2, and level-3 PRN interconnection structures are described in more detail in the commonly-owned, co-pending U.S. Patent Application Publication No. 2018 / 0145850, the entire content of which is incorporated by reference. The use of the PRN interconnection structure in computer system 200 according to various embodiments is described in more detail below.

[0056] Figure 4 is a block diagram of a level-1 permutation ring network 11 according to an embodiment of the present invention. Other level-1 permutation ring networks (e.g., permutation ring networks 12-14 and 11'-14') of computer system 200 may be the same as level-1 permutation ring network 11. In the illustrated embodiment, level-1 permutation ring network 11 includes four bidirectional source-synchronous ring networks 401, 402, 403, and 404. Each of the ring networks 401-404 serves as a communication channel. Although the illustrated permutation ring network 11 includes nine communication nodes (i.e., communication node CN1, computing engines CE 1A , CE 1B , CE 1C and CE 1D as well as memory banks M 1A , M 1B , M 1C and M 1D ) and four communication channels 401-404, it should be understood that other numbers of communication nodes and communication channels may be used in other embodiments. Generally, the number of communication nodes in level-1 permutation ring network 11 is identified by an N value, and the number of bidirectional ring networks in level-1 permutation ring network 11 is identified by an M value. The number of communication channels (M) is selected to provide an appropriate trade-off between the bandwidth requirements of the communication network and the area power limitations of the communication network.

[0057] Each of the communication channels 401-404 includes a plurality of data transmission stations connected by a bidirectional link (interconnection). More specifically, communication channel 401 includes nine data transmission stations A0-A8, communication channel 402 includes nine data transmission stations B0-B8, communication channel 403 includes nine data transmission stations C0-C8, and communication channel 404 includes nine data transmission stations D0-D8. The bidirectional link of communication channel 401 is shown as a solid line connecting data transmission stations A0-A8 in a ring. The bidirectional link of communication channel 402 is shown as a long dashed line connecting data transmission stations B0-B8 in a ring. The bidirectional link of communication channel 403 is shown as a dotted line connecting data transmission stations C0-C8 in a ring. The bidirectional link of communication channel 404 is shown as a short dashed line connecting data transmission stations D0-D8 in a ring. The bidirectional link allows simultaneous transmission of data and clock signals in both the clockwise and counterclockwise directions.

[0058] Generally, each of the data transmission stations A0-A8, B0-B8, C0-C8, and D0-D8 provides an interface capable of transmitting data between nine communication nodes and communication channels 401-404.

[0059] Generally, each of the communication channels 401-404 is coupled to receive a master clock signal. Thus, in Figure 4 the example, communication channels 401, 402, 403, and 404 are coupled to receive master clock signals CKA, CKB, CKC, and CKD, respectively. In the illustrated embodiment, data transmission stations A0, B0, C0, and D0 are coupled to receive master clock signals CKA, CKB, CKC, and CKD, respectively. However, in other embodiments, other data transmission stations in communication channels 401, 402, 403, and 404 may be coupled to receive master clock signals CKA, CKB, CKC, and CKD, respectively. Although four separate master clock signals CKA, CKB, CKC, and CKD are shown, it should be understood that each of the master clock signals CKA, CKB, CKC, and CKD may be derived from a single master clock signal. In the described embodiment, each of the master clock signals CKA, CKB, CKC, and CKD has the same frequency.

[0060] A conventional clock generation circuit (e.g., a phase-locked loop circuit) can be used to generate the master clock signals CKA, CKB, CKC, and CKD. In the described embodiments, the master clock signals can have a frequency of about 5 GHz or higher. However, it should be understood that in other embodiments, the master clock signals can have other frequencies. The frequency and voltage of the master clock signals can be adjusted based on the bandwidth requirements and power optimization of the ring network architecture. In the illustrated embodiments, data transfer stations A0, B0, C0, and D0 receive the master clock signals CKA, CKB, CKC, and CKD, respectively. Each of the other data transfer stations receives its clock signal from its adjacent neighbor. That is, the master clock signals CKA, CKB, CKC, and CKD are effectively serially transmitted to each of the data transfer stations in communication channels 401, 402, 402, and 404, respectively.

[0061] Each of the communication channels 401, 402, 403, and 404 operates in a source-synchronous manner with respect to its corresponding master clock signal CKA, CKB, CKC, and CKD, respectively.

[0062] Generally, each data transfer station can transmit output messages on two paths. In the first path, the message received from the upstream data transfer station is forwarded to the downstream data transfer station (e.g., data transfer station A0 can forward the message received from downstream data transfer station A8 to upstream data transfer station A1 in the clockwise path, or data transfer station A0 can forward the message received from downstream data transfer station A1 to upstream data transfer station A8 in the counterclockwise path). In the second path, the message provided by the communication node coupled to the data transfer station is routed to the downstream data transfer station (e.g., data transfer station A0 can forward the message received from the computing engine CE 1A in the clockwise path to downstream data transfer station A1 or in the counterclockwise path to downstream data transfer station A8). Also in the second path, the message received by the data transfer station is routed to the addressed communication node (e.g., data transfer station A0 can forward the message received from downstream data transfer station A8 to the computing engine CE 1A in the clockwise path, or data transfer station A0 can forward the message received from downstream data transfer station A0 to the computing engine CE 1A ). Note that the lines and buffers for transmitting clock signals and messages between data transfer stations are highly balanced and equalized to minimize the loss of setup and hold times.

[0063] The clock signal paths and message buses operate as a wave pipeline system, where messages transmitted between data transfer stations are latched into the receiving data transfer station in a source-synchronous manner using the clock signals transmitted on the clock signal paths. In this way, messages are transmitted between data transfer stations at the frequencies of the master clock signals CKA, CKB, CKC, and CKD, thereby allowing for fast data transfer between data transfer stations.

[0064] Since point-to-point source-synchronous communication is implemented, the line and buffer delays of the clock signal line structure and message bus structure do not reduce the operating frequency of the communication channels 401 - 404.

[0065] Since the data transfer stations have a relatively simple design, the transmission of messages on the permutation ring network 11 can be performed at a relatively high frequency. The communication nodes CN1, computing engines CE 1A 、CE 1B 、CE 1C and CE 1D as well as the memory banks M 1A 、M 1B 、M 1C and M 1D typically include more complex designs and may operate at frequencies lower than the frequencies of the master clock signals CKA, CKB, CKC, and CKD.

[0066] Note that the loop configuration of the communication channels 401 - 404 requires that messages received by the originating data transfer stations A0, B0, C0, and D0 (e.g., the data transfer stations that receive the master clock signals CKA, CKB, CKC, and CKD) must be resynchronized with the master clock signals CKA, CKB, CKC, and CKD respectively. In one embodiment, in response to an input clock signal received from a downstream data transfer station, a resynchronization circuit (not shown) performs this synchronization operation by latching the input message into a first flip-flop. Then, in response to the master clock signal (e.g., CKA), the message provided at the output of this first flip-flop is latched into a second flip-flop. The second flip-flop provides a synchronized message to the originating data transfer station (e.g., data transfer station A0). In response to the master clock signal (CKA), this synchronized message is stored in the originating data transfer station (A0).

[0067] Now returning to the topology of the first-level permutation ring network 11, the communication nodes CN1, computing engines CE 1A 、CE 1B 、CE 1C and CE 1D as well as the memory banks M 1A 、M 1B 、M 1C and M 1DEach of which is coupled to a unique one of data transfer stations A0 - A8, B0 - B8, C0 - C8, and D0 - D8 in each of four communication channels 401 - 404. For example, the computing engine CE 1A is coupled to data transfer station A0 in communication channel 401, data transfer station B8 in communication channel 402, data transfer station C7 in communication channel 403, and data transfer station D6 in communication channel 404. Table 1 below defines the connections between communication node CN1, computing engine CE 1A CE 1B CE 1C and CE 1D and memory banks M 1A M 1B M 1C and M 1D , and the connections between data transfer stations A0 - A8, B0 - B8, C0 - C8, and D0 - D8. Note that the physical connections between communication node CN1, computing engine CE 1A CE 1B CE 1C and CE 1D and memory banks M 1A M 1B M 1C and M 1D , and between data transfer stations A0 - A8, B0 - B8, and C0 - C8 are not explicitly shown in Figure 4 for clarity.

[0068]

[0069] Figure 5 Reorder the data in Table 1 to provide an interconnect matrix 500 for the four communication channels 401 - 404, where the interconnect matrix 500 is ordered by data transfer stations in each of communication channels 401 - 404. The interconnect matrix 500 makes it easy to determine the number of hops between communication node CN1, computing engine CE 1A CE 1B CE 1C and CE 1D and memory banks M 1A M 1B M 1C and M 1D on each of communication channels 401 - 404. Note that communication node CN1, computing engine CE 1A CE 1B CE 1C and CE 1D and memory banks M 1A M 1B M1C and M 1D is coupled to data transmission stations having different relative positions in four communication channels 401 - 404. As described in more detail below, this configuration allows for a general and efficient routing of messages between communication nodes.

[0070] Figure 6 is the routing table 600, which defines the traffic flow through the communication nodes CN1, computing engines CE 1A , CE 1B , CE 1C , and CE 1D and the memory banks M 1A , M 1B , M 1C , and M 1D of the permutation ring network 11 according to the present embodiment. For example, the communication node CN1 and the computing engine CE 1A communicate on communication channel 404 using the path between data transmission stations D5 and D6. The number of hops along this path is defined by the number of segments traversed on communication channel 404. Since data transmission stations D5 and D6 are adjacent to each other on communication channel 404 (i.e., there is one segment between data transmission stations D5 and D6), the communication path between communication node CN1 and computing engine CE 1A consists of one hop (1H).

[0071] As shown in the routing table 600, all relevant communication paths between the communication nodes CN1, computing engines CE 1A , CE 1B , CE 1C , and CE 1D and the memory banks M 1A , M 1B , M 1C , and M 1D include unique hop communication paths. In other embodiments, one or more of the communication paths may include more than one hop. In other embodiments, multiple communication paths may be provided between one or more pairs of communication nodes CN1, computing engines CE 1A , CE 1B , CE 1C , and CE 1D and the memory banks M 1A , M 1B , M 1C , and M 1D . In other embodiments, different pairs of communication nodes may share the same communication path.

[0072] Communication among data transfer stations A0 - A8, B0 - B8, C0 - C8, and D0 - D8 will operate at the highest frequency permitted by the source - synchronous network. This frequency does not decrease proportionally with an increase in the number of communication nodes and the number of communication channels. It should be understood that each of the communication channels 401 - 404 includes provisions for initialization, arbitration, flow control, and error handling. In one embodiment, these provisions are provided using well - established techniques.

[0073] Computing engine CE 1A , CE 1B , CE 1C and CE 1D as well as memory banks M 1A , M 1B , M 1C and M 1D each transmit messages (which may include data) on the permutation ring network 11 according to the routing table 600. For example, computing engine CE 1A may send a data request message to memory bank M 1C using communication channel 404. More specifically, computing engine CE 1A may transmit a data request message to the clockwise transmission path of data transfer station C7. This data request message addresses data transfer station C8 and memory bank M 1C . Upon receiving the data request message, data transfer station C8 determines that the data request message addresses memory bank M 1C , and forwards the data request message to memory bank M 1C . After processing the data request message, memory bank M 1C may transmit a data response message to the counter - clockwise transmission path of data transfer station C8. This data response message addresses data transfer station C7 and computing engine CE 1A . Upon receiving the data response message, data transfer station C7 determines that the data response message addresses computing engine CE 1A , and forwards the data response message to computing engine CE 1A .

[0074] Messages can be transmitted into and out of the permutation ring network 11 through communication node CN1. For example, the computing engine CE 1A of slice 1 may use communication channel 404 to transmit a data request message to the memory bank M 2A of computing slice 2. More specifically, computing engine CE 1A may transmit a data request message to the counter - clockwise transmission path of data transfer station D6. This data request message addresses data transfer station D5 and communication node CN1 (as well as communication node CN2 of computing slice 2 and memory bank M 2A)。When receiving a data request message, data transfer station D5 determines that the data request message is addressed to communication node CN1 and forwards the data request message to communication node CN1. In response, communication node CN1 determines that the data request message is addressed to communication node CN2 within compute slice 2 and forwards the data request message on secondary PRN interconnect 21 (using the routing table implemented by secondary PRN interconnect 21). Note that secondary PRN interconnect 21 uses a PRN structure similar to that of primary PRN interconnect 11 to route messages between communication nodes CN1 - CN6. Note that due to the different number of communication nodes served by secondary PRN interconnect 21, the implementation of secondary PRN interconnect 21 may be different from that of primary PRN interconnect 11 (e.g., different number of communication channels, different routing tables). According to one embodiment, the secondary PRN-based interconnect structure 21 includes three communication channels (i.e., three bidirectional ring networks), where each communication channel includes six data transfer stations. In this embodiment, each of communication nodes CN1 - CN6 is coupled to a corresponding one of the data transfer stations in each of the three communication channels.

[0075] The data transfer station associated with communication node CN2 receives the data request message transmitted on secondary PRN interconnect 21, determines that the data request message is addressed to communication node CN2, and forwards the data request message to communication node CN2. In response, communication node CN2 determines that the data request message is addressed to memory bank M 2A , and forwards the data request message on primary PRN interconnect 12 (using the routing table implemented by primary PRN interconnect 12). Note that primary PRN interconnect 12 uses a PRN structure similar to that of primary PRN interconnect 11 to route messages between communication node CN2, compute engines CE 2A , CE 2B , CE 2C , CE 2D and memory banks M 2A , M 2B , M 2C and M 2D of memory bank group M2.

[0076] The data transfer station associated with memory bank M 2A receives the data request message transmitted on primary PRN interconnect 12, determines that the data request message is addressed to memory bank M 2A , and forwards the data request message to memory bank M 2A . Memory bank M 2A can then respond to the data request message. For example, memory bank M 2A can retrieve the stored data value and return the data value to compute engine C using a data response message1A This data response message is transmitted to the computing engine C using the reverse path of the original data request message 1A .

[0077] According to one embodiment, the three-level PRN-based interconnect structure 31 includes three communication channels (i.e., three bidirectional ring networks), where each communication channel includes six data transfer stations. In this embodiment, the communication nodes CN 11 -CN 16 each are connected to a corresponding one of the data transfer stations in each of the three communication channels.

[0078] Using the above flat computer architecture and message system, messages can be transmitted between any of the various components of the computer system 200 through the first-level PRN interconnect structure, the second-level PRN interconnect structure, and the third-level PRN interconnect structure without changing the message protocol. According to one embodiment, each of the components of the computer system 200 is assigned a unique (system) address. Addressing the various components of the system 200 in this way allows for consistent access to these components across the first-level PRN interconnect structure, the second-level PRN interconnect structure, and the third-level PRN interconnect structure. Note that the computer system 200 is a non-coherent system because the computer system 200 does not explicitly ensure the consistency of the data stored in the computing slices, DRAMs 211-214, or memory banks within the system memory 206. Instead, the user needs to control the data stored in these memories in the desired manner. Therefore, the computer system 200 is well-suited for implementing producer-consumer execution models, such as the model implemented through the forward propagation of a neural network. That is, the computer system 200 is capable of effectively processing data in neural network / machine learning applications. The improved network topology of the computer system 200 advantageously enables spanning multiple chips without the need for a cache coherence protocol between multiple chips. Therefore, the computer system 200 is easy to expand and can provide communication between many different chips.

[0079] In the above embodiment, the first-level interconnect structure, the second-level interconnect structure, and the third-level interconnect structures 11, 21, and 31 are all implemented using bidirectional source-synchronous replacement ring networks. However, in an alternative embodiment of the present invention, a non-PRN-based structure can be used to implement the first-level interconnect structure.

[0080] Figure 7 is a block diagram of a computer system 700 according to an alternative embodiment of the present invention. Since the computer system 700 is similar to the computer system 200, so Figure 7 and Figure 2Similar elements in [the figure] are labeled with similar reference numerals. Accordingly, computer system 700 includes multiple processor chips 701 - 704, a main system processor 205, a system memory 206, a system-level interconnect chip 210, and DRAM devices 211 - 214. Although Figure 7 only four processor chips 701 - 704 are shown in [the figure], it should be understood that in other embodiments, computer system 700 can be easily modified to include other numbers of processor chips. Additionally, although Figure 7 only processor chip 701 is shown in detail in [the figure], it should be understood that processor chips 702 - 704 include internal elements that are the same (or similar) to those of processor chip 701 in the described embodiment. As described in more detail below, processor chip 701 replaces the level-1 PRN-based interconnect structures 11 - 14 of processor chip 201 with simple network interconnect structures 711 - 714. The simple network interconnect structures 711 - 714 can be, for example, a crossbar switch-based interconnect structure or a simple ring network.

[0081] In the illustrated embodiment, processor chip 701 includes four compute slices 71, 72, 73, and 74 coupled to a secondary permutation ring network interconnect structure 21. Although Figure 7 four compute slices are shown in [the figure], it should be understood that in other embodiments, other numbers of compute slices can be included on processor chip 701. Each compute slice includes multiple compute engines, multiple memory banks, communication nodes, and a simple network interconnect structure. More specifically, slices 71, 72, 73, and 74 include compute engine groups CE1, CE2, CE3, and CE4, memory bank groups M1, M2, M3, and M4, simple network interconnect structures 711, 712, 713, and 714, and communication nodes CN1, CN2, CN3, and CN4, respectively. Compute engine groups CE1, CE2, CE3, and CE4 and memory bank groups M1, M2, M3, and M4 are described in more detail above in conjunction with Figure 2 and Figure 3 more specifically.

[0082] Within each of slices 71, 72, 73, and 74, the corresponding simple network interconnect structures 711, 712, 713, and 714 couple the corresponding compute engine groups CE1, CE2, CE3, and CE4 to the corresponding memory bank groups M1, M2, M3, and M4. This allows each compute engine to access each memory bank within the same slice using the corresponding simple network.

[0083] Simple network interconnection structures 711, 712, 713, and 714 are also coupled to corresponding communication nodes CN1, CN2, CN3, and CN4 within corresponding compute slices 71, 72, 73, and 74. The communication nodes CN1, CN2, CN3, and CN4 are coupled to the second-level PRN-based interconnection structure 21 in the manner described above. The communication nodes CN1, CN2, CN3, and CN4 transfer messages and data between the corresponding simple network interconnection structures 711, 712, 713, and 714 and the second-level PRN-based interconnection structure 21. It should be noted that messages transmitted between the simple network interconnection structures 711, 712, 713, and 714 and the corresponding communication nodes CN1, CN2, CN3, and CN4 must be converted to a protocol consistent with the receiving system. Such conversion can be achieved through interfaces within the simple network interconnection structures 711 - 714 or interfaces within the communication nodes CN1, CN2, CN3, and CN4. Although this protocol conversion complicates the operation of the computer system 700, it allows the use of simple network interconnection structures within each compute slice, which can reduce the required layout area of the compute slices 71 - 74.

[0084] In another embodiment of the present invention, the third-level PRN-based interconnection structure 31 is replaced by a simple network interconnection structure such as a crossbar-based interconnection structure or a simple ring network (in the same manner as the first-level PRN-based structures 11 - 14 were replaced by the simple network structures 711 - 714 above). Figure 7 Figure 8 is a block diagram of a computer system 800 according to this alternative embodiment, which replaces the third-level PRN-based interconnection structure 31 with a simple network interconnection structure 81 on a system-level interconnection chip 810 in the manner described above. A simple network interconnection structure 81, which can include, for example, a crossbar-based interconnection structure or a simple ring network, provides connections between communication nodes CN 11 -CN 16 Note that messages transmitted between the processor chips 701 - 704, the main system processor 205, and the system memory 206, and the corresponding communication nodes CN 11 , CN 12 , CN 13 , CN 14 , CN 15 and CN 16 must be converted to a protocol consistent with the receiving system. Such conversion can be achieved through interfaces within the simple network interconnection structure 81 or interfaces within the communication nodes CN 11 -CN 16 . Although this protocol conversion complicates the operation of the computer system 800, it allows the use of a simple network interconnection structure within the system-level interconnection chip 810.

[0085] Although the simple network interconnect structure 81 of the system-level interconnect chip 810 is shown in combination with the compute slices 71-74 having the simple network interconnect structures 711-713, it should be understood that the simple network interconnect structure 81 of the system-level interconnect chip 810 can also be used in combination with the compute slices 1-4 having the level-1 PRN-based interconnect structures 11-14, as Figure 2 shown.

[0086] Several factors can be used to determine whether the level-1 interconnect structure and the level-3 interconnect structure should be implemented with a bidirectional source-synchronous replacement ring network ( Figures 2 to 3 ) or a simple network interconnect structure such as a crossbar switch or a single ring network ( Figures 7 to 8 ). The replacement ring network will provide better performance than a simple single ring network (but requires a larger layout area). The replacement ring network will generally also provide better performance than a crossbar switch (and may require a larger layout area). Generally, as more communication nodes are connected through the interconnect structure, it becomes more efficient (in terms of layout area and performance) to use the replacement ring network instead of the single ring network or the crossbar switch. According to one embodiment, the replacement ring network is used when the number of communication nodes to be connected is four or more.

[0087] Although the invention has been described in connection with several embodiments, it should be understood that the invention is not limited to the disclosed embodiments, but is capable of various modifications, which will be apparent to those skilled in the art. Therefore, the invention is only limited by the following claims.

Claims

1. A computer architecture, comprising: A plurality of computing slices, each computing slice comprising: A plurality of computing engines, A plurality of memory banks, A communication node, and A first-level interconnect structure, the first-level interconnect structure comprising a slice-level permutation ring network having a plurality of first bidirectional source-synchronous ring networks, each first bidirectional source-synchronous ring network comprising a plurality of first data transfer stations, wherein each of the plurality of computing engines, the plurality of memory banks, and the communication node is coupled to one of the plurality of first data transfer stations in each of the plurality of first bidirectional source-synchronous ring networks of the slice-level permutation ring network; and A second-level interconnect structure, comprising a second-level permutation ring network having a plurality of second bidirectional source-synchronous ring networks, each second bidirectional source-synchronous ring network comprising a plurality of second data transfer stations connected in a ring, wherein each communication node of the plurality of computing slices is coupled to one of the plurality of second data transfer stations in each of the plurality of second bidirectional source-synchronous ring networks, wherein messages are transmitted between any one of the plurality of computing engines of the plurality of computing slices and any one of the plurality of memory banks of the plurality of computing slices via the first-level interconnect structure and the second-level interconnect structure without changing the message protocol.

2. The computer architecture according to claim 1, further comprising: A memory interface communication node, which is coupled to the second-level interconnect structure, wherein the memory interface communication node is coupled to one of the plurality of second data transfer stations in each of the plurality of second bidirectional source-synchronous ring networks; And A storage device, which is coupled to the memory interface communication node.

3. The computer architecture according to claim 2, wherein the storage device is a dynamic random access memory (DRAM) device.

4. The computer architecture according to claim 1, further comprising a first network communication node, which is coupled to the second-level interconnect structure, wherein the first network communication node is coupled to one of the plurality of second data transfer stations in each of the plurality of second bidirectional source-synchronous ring networks of the second-level interconnect structure.

5. The computer architecture according to claim 4, further comprising a system-level interconnect structure coupled to the first network communication node.

6. The computer architecture according to claim 5, wherein the system-level interconnect structure comprises a plurality of network communication nodes coupled to a third-level interconnect structure, wherein the first of the plurality of network communication nodes is coupled to the first network communication node.

7. The computer architecture according to claim 6, further comprising a main system processor coupled to the second of the plurality of network communication nodes.

8. The computer architecture according to claim 7, further comprising a system memory coupled to the third of the plurality of network communication nodes.

9. The computer architecture according to claim 6, wherein the three-level interconnect structure includes a system-level permutation ring network having a plurality of third bidirectional source-synchronous ring networks, each of the third bidirectional source-synchronous ring networks including a plurality of third data transfer stations connected in a ring, and each of the plurality of network communication nodes being coupled to one of the plurality of third data transfer stations in each of the plurality of third bidirectional source-synchronous ring networks of the system-level permutation ring network.

10. The computer architecture according to claim 9, further comprising a plurality of second computing slices, each second computing slice including: a plurality of second computing engines, a plurality of second memory banks, a second communication node, and a second first-level interconnect structure including a slice-level permutation ring network having a plurality of fourth bidirectional source-synchronous ring networks, each of the fourth bidirectional source-synchronous ring networks including a plurality of fourth data transfer stations connected in a ring, and each of the plurality of second computing engines, the plurality of second memory banks, and the second communication node being coupled to one of the plurality of fourth data transfer stations in each of the plurality of fourth bidirectional source-synchronous ring networks of the slice-level permutation ring network; a second second-level interconnect structure including a second second-level permutation ring network having a plurality of fifth bidirectional source-synchronous ring networks, each of the fifth bidirectional source-synchronous ring networks including a plurality of fifth data transfer stations connected in a ring, and each second communication node of the plurality of second computing slices being coupled to one of the plurality of fifth data transfer stations in each of the plurality of fifth bidirectional source-synchronous ring networks; and a second network communication node coupled to the second second-level interconnect structure, wherein the second network communication node is coupled to one of the plurality of fifth data transfer stations in each of the plurality of fifth bidirectional source-synchronous ring networks of the second second-level interconnect structure, and the second network communication node is coupled to one of the plurality of third data transfer stations in each of the plurality of third bidirectional source-synchronous ring networks of the system-level permutation ring network, and messages are transmitted between any one of the computing engines of the plurality of first computing slices and second computing slices and any one of the plurality of memory banks of the plurality of first computing slices and second computing slices via the first-level interconnect structure, the second first-level interconnect structure, the second-level interconnect structure, the second second-level interconnect structure, and the three-level interconnect structure without changing the message protocol.

11. The computer architecture according to claim 1, wherein each communication node includes a communication path to each of the other communication nodes, and each of the communication paths is a 1-hop path between adjacent data transfer stations of the plurality of second data transfer stations.

12. The computer architecture according to claim 1, wherein a unique pair of adjacent data transfer stations among the plurality of second data transfer stations provides a communication path between each pair of the communication nodes.

13. The computer architecture according to claim 1, wherein the plurality of second bidirectional source synchronous ring networks operate in a first clock domain, and the plurality of computing engines and the plurality of memory banks operate in a second clock domain different from the first clock domain.

14. The computer architecture according to claim 1, wherein the plurality of computing slices are located on the same semiconductor chip as the secondary interconnect structure.

15. The computer architecture according to claim 1, wherein each communication node is coupled to a data transfer station having a different relative position in the plurality of bidirectional source synchronous ring networks.

Citation Information

Patent Citations

  • Interconnected Ring Network In A Multi-processor System

    CN103970712A

  • Multiprocessor chip having bidirectional ring interconnect

    US20060041715A1

  • Permutated Ring Network

    US20180145850A1