Data transfer between a memory and a distributed computing array
The circuit architecture synchronizes data transfers across multiple dies in a multi-die IC by monitoring remote buffer fill levels and initiating synchronized data transfers, addressing inefficiencies in data skew and bandwidth utilization.
Patent Information
- Application Number
- JP2022533569
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-12-06
- Filing Date
- 2020-12-04
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2040-12-04
AI Technical Summary
Data transfer from memory to a computing array in a multi-die IC is inefficient due to data skew caused by independent memory channels operating in parallel, leading to unpredictable bandwidth utilization and degraded performance.
A circuit architecture that includes a controller to synchronize data transfers across multiple dies by monitoring remote buffer fill levels and broadcasting read enable signals to initiate synchronized data transfers from remote buffers to the computing array.
The solution enhances data synchronization, reduces skew, and maximizes bandwidth utilization, ensuring efficient data transfer and utilization of high-bandwidth memory resources.
Smart Images

Figure 0007709971000001 
Figure 0007709971000002 
Figure 0007709971000003
Abstract
Description
Technical Field
[0001] Technical Field The present disclosure relates to integrated circuits (ICs), and more particularly, to data transfer between a memory and a computing array distributed across multiple dies of an IC.
Background Art
[0002] Background A neural network processor (NNP) refers to a type of integrated circuit (IC) having one or more computing arrays capable of implementing a neural network. Data, such as weights for neural network implementation, is supplied to the computing array(s) from a memory. Weights are supplied from the memory to the computing array(s) in parallel across multiple memory channels. Data transfer from the memory to the computing array typically suffers from skew. As a result, the data reaches different parts of the computing array(s) at different times. Data skew is at least partially due to the independence between memory channels while operating in parallel, and in the case of a multi-die IC having computing array(s) distributed across multiple dies, the data wavefronts from each memory channel are orthogonal to the computing array(s) within the multi-die IC. These problems, whether viewed individually or cumulatively, make the data transfer from the memory to the computing array unpredictable, leading to inefficient and / or degraded use of the available bandwidth from the memory.
Summary of the Invention
Means for Solving the Problems
[0003] Summary Exemplary embodiments include an integrated circuit (IC). The IC includes a plurality of dies. The IC includes a plurality of memory channel interfaces configured to communicate with a memory, and the plurality of memory channel interfaces are disposed within a first die of the plurality of dies. The IC can include a computing array distributed across the plurality of dies and a plurality of remote buffers distributed across the plurality of dies. The plurality of remote buffers are coupled to the plurality of memory channels and the computing array. The IC further includes a controller configured to determine that each of the plurality of remote buffers stores data therein and, in response, broadcast a read enable signal to each of the plurality of remote buffers to initiate a data transfer from the plurality of remote buffers to the computing array across the plurality of dies.
[0004] Another exemplary embodiment includes a controller. The controller is disposed within an IC having a plurality of dies. The controller includes a request controller configured to convert a first request for access to a memory into a second request compliant with an on-chip communication bus, and the request controller provides the second request to a plurality of request buffer-bus master circuit blocks configured to receive data from a plurality of channels of the memory. The controller further includes a remote buffer read address generation unit coupled to the request controller and configured to monitor a fill level in each of a plurality of remote buffers distributed across the plurality of dies. Each remote buffer of the plurality of remote buffers is configured to provide data obtained from one of each of the plurality of request buffer-bus master circuit blocks to a computing array also distributed across the plurality of dies. In response to a determination that each of the plurality of remote buffers stores data based on the fill level, the remote buffer read address generation unit is configured to initiate a data transfer from each of the plurality of remote buffers to the computing array across the plurality of dies.
[0005] Another exemplary embodiment includes a method. The method includes monitoring fill levels in a plurality of remote buffers distributed across a plurality of dies, where each remote buffer of the plurality of remote buffers is configured to provide data to a computing array also distributed across the plurality of dies; determining, based on the fill levels, that each remote buffer of the plurality of remote buffers is storing data; and in response to the determining, initiating a data transfer from each remote buffer of the plurality of remote buffers to the computing array across the plurality of dies.
[0006] This summary section is provided only to introduce certain concepts and is not intended to identify any key or essential features of the claimed subject matter. Other features of the inventive arrangements will become apparent from the accompanying drawings and the following detailed description.
[0007] Brief Description of the Drawings The inventive arrangements are shown by way of example in the accompanying drawings. However, the drawings should not be construed as being limited to only the specific embodiments shown of the inventive arrangements. Various aspects and advantages will become apparent upon review of the following detailed description and with reference to the drawings.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
[0009] Detailed Description The present disclosure is accompanied by claims that define novel features, but the various features described within the present disclosure are believed to be better understood by considering this description in conjunction with the drawings. The processes, machines, manufactures, and any variations thereof described herein are provided for purposes of illustration. The specific structural and functional details described within the present disclosure are not to be construed as limiting, but rather as a basis for the claims and as a representation for teaching one of ordinary skill in the art to employ substantially any appropriately detailed structure in various ways. Further, the terms and phrases used within the present disclosure are not intended to be limiting, but rather are intended to provide an understandable description of the features being described.
[0010] The present disclosure relates to integrated circuits (ICs), and more particularly, to data transfer between a memory and a computing array distributed across multiple dies of an IC. A neural network processor (NNP) refers to a type of IC having one or more computing arrays capable of implementing a neural network. If the IC is a multi-die IC, the IC can implement a single larger computing array that is distributed across two or more dies of the multi-die IC. Implementing a single larger computing array in a distributed form across multiple dies provides certain advantages including, but not limited to, improved latency, improved weight storage capacity, and improved computational efficiency, as opposed to multiple smaller independent computing arrays in different dies.
[0011] Data, such as neural network weights, is supplied to the computing array from high bandwidth memory or HBM. For purposes of explanation, the memory accessed by a memory channel is referred to throughout this disclosure as "high bandwidth memory" or "HBM" to better distinguish it from other types of memory within a circuit architecture such as buffers and / or queues. However, it should be understood that HBM may be implemented using any of a variety of different technologies that support a plurality of independent parallel memory channels communicatively linked to an exemplary circuit architecture described via a suitable memory controller. Examples of HBM can include any of a variety of RAM type memories including double data rate RAM or other suitable memory.
[0012] The computing array is considered a single computing array to which weights are supplied in parallel via available memory channels by the HBM, even though it is distributed across multiple dies of the IC. For example, the computing array may be implemented as an array where each die implements one or more rows of the computing array. Each memory channel can provide data to one or more of the rows of the computing array.
[0013] In the case of a single compute array distributed across multiple dies, data transfer from HBM to the compute array often suffers from timing issues. For example, each memory channel typically has its own independent control pins, asynchronous clock, and refresh sleep mode. These features introduce data skew across the memory channels. As a result, different rows of the compute array often receive data at different times. The data skew is further exacerbated because different rows of the compute array are located on different dies of the IC and thus at different distances from the HBM. For example, the data wavefront from a memory channel, e.g., data propagation, is orthogonal to the compute array within the IC. These problems contribute to the overall unpredictability of data transfer from HBM to the compute array.
[0014] According to the configuration of the present invention described within this disclosure, an exemplary circuit architecture is provided that can schedule read requests to the HBM across memory channels while improving and / or maximizing HBM bandwidth utilization. The exemplary circuit architecture can also deskew data transfer between the HBM and the compute array. As a result, data can be provided in a synchronized manner with reduced skew to different rows of the compute array across multiple dies of a multi-die IC from the HBM. This allows the compute array to be kept busy while more fully utilizing the read bandwidth of the HBM.
[0015] The exemplary circuit architecture can also reduce the overhead and complexity of distributing the compute array across multiple dies of a multi-die IC. The exemplary circuit architecture described herein can be adapted to multi-die ICs having different numbers of dies internally. When the number of dies within a multi-die IC varies from model to model and / or the area of each die varies, the exemplary circuit architecture described herein can be adapted to such changes to improve data transfer between the HBM and the various dies of the multi-die IC where the compute array is distributed.
[0016] Throughout the present disclosure, the Advanced Microcontroller Bus Architecture (AMBA) eXtensible Interface (AXI) (hereinafter, “AXI”) protocol and communication bus are used for illustrative purposes. The AXI defines an embedded microcontroller bus interface for use in establishing on-chip connections between corresponding circuit blocks and / or systems. The AXI is provided as an exemplary example of a bus interface and is not intended as a limitation of the examples described within the present disclosure. Other similar and / or equivalent protocols, communication buses, bus interfaces, and / or interconnections may be used in place of the AXI, and it should be understood that the various exemplary circuit blocks and / or signals provided within the present disclosure will vary based on the particular protocol, communication bus, bus interface, and / or interconnection used.
[0017] Further aspects of the configuration of the present invention are described in more detail below with reference to the drawings. For simplicity and clarity of explanation, the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity. Further, reference numerals may be repeated between the drawings to indicate corresponding, similar, or like features where appropriate.
[0018] FIG. 1 shows an exemplary floorplan for a circuit architecture implemented in an IC100. The IC100 is a multi-die IC and includes a computing array. The computing array is distributed across dies 102, 104, and 106 of the IC100. For illustrative purposes, the IC100 is shown as having three dies. In other examples, the IC100 may have fewer or more dies than shown.
[0019] In this example, the compute array is subdivided into 256 compute array rows. Compute array rows 0 to 95 are implemented on die 102. Compute array rows 96 to 191 are implemented on die 104. Compute array rows 192 to 255 are implemented on die 106. The compute array can include a digital signal processing (DSP) cascade chain that is connected together across dies 102, 104, and 106.
[0020] Data, such as weights, is obtained from a high bandwidth memory (HBM) (not shown) communicatively linked to IC100 via a plurality of memory channels. In one aspect, the HBM is implemented within a separate IC (e.g., outside the package of IC100) and on the same circuit board as IC100. In another aspect, the HBM is implemented as a separate die within IC100 (e.g., within the same package as IC100). The HBM can be arranged along the bottom surface of IC100, for example, adjacent to the bottom of die 106 from left to right. In some cases of the HBM, the memory channels are referred to as pseudo channels (PCs). For purposes of explanation, the term "memory channel" is used to refer to the memory channels of the HBM and / or the PCs of the HBM.
[0021] In the example of FIG. 1, die 106 includes 16 memory controllers 0 to 15. Each memory controller can service (e.g., read and / or write to) two memory channels. Memory controllers 0 to 15 in FIG. 1 are labeled with a label indicating the specific memory channels that each memory controller services, in parentheses. For example, memory controller 0 services memory channels 0 and 1, memory controller 1 services memory channels 2 to 3, and so on.
[0022] Each memory controller is connected to one or more request buffers and one or more bus master circuits (e.g., AXI master). In one exemplary embodiment shown in FIG. 1, each memory channel is coupled through a memory controller to one request buffer - bus master circuit block. In FIG. 1, the combination of each request buffer - bus master circuit block (e.g., the bus master can be an AXI master) is abbreviated as a "RBBM circuit block" and is shown as "RBBM" in the figure. Since each memory controller can service two memory channels, two RBBM circuit blocks are above each memory controller and are coupled to each memory controller. Each RBBM circuit block is labeled for the particular memory channel serviced by the RBBM circuit block. Thus, the example of FIG. 1 includes RBBM circuit blocks 0 - 15. In this example, all of the memory controllers and RBBM circuit blocks are located in one, e.g., the same die of IC100.
[0023] Each of dies 102, 104, and 106 includes a plurality of remote buffers. The remote buffers are distributed across dies 102, 104, and 106. In the example of FIG. 1, each RBBM circuit block is connected to a plurality of remote buffers. In one example, each RBBM circuit block is connected to four different remote buffers. For purposes of illustration, RBBM circuit block 0 is connected to remote buffers 0 - 3 and provides data to remote buffers 0 - 3. RBBM circuit block 1 is connected to remote buffers 4 - 7 and supplies data to remote buffers 4 - 7, and so on. Each of the remaining RBBM circuit blocks can be connected to a consecutive number group of four remote buffers following die 102 through the remote buffers of dies 104 and 106.
[0024] Each of dies 102, 104, and 106 also includes a plurality of caches. Generally, the number of caches (e.g., 32) corresponds to the number of memory channels. Each cache can provide data to a plurality of compute array rows. In the example of FIG. 1, each cache can provide data to 8 compute array rows. Die 102 includes caches 0 to 11, cache 0 provides data to compute array rows 0 to 7, cache 1 provides data to compute array rows 8 to 15, cache 2 provides data to compute array rows 16 to 23, and so on. Die 104 includes caches 12 to 23, cache 12 provides data to compute array rows 96 to 103, cache 13 provides data to compute array rows 104 to 111, cache 14 provides data to compute array rows 112 to 119, and so on. Die 106 includes caches 24 to 31, cache 24 provides data to compute array rows 192 to 199, cache 25 provides data to compute array rows 200 to 207, cache 26 provides data to compute array rows 208 to 215, and so on.
[0025] In this example, data such as weights can be loaded from HBM via 32 memory channels implemented within die 106. Ultimately, the weights are supplied as multiplication operands to the compute array rows. The weights enter IC100 via parallel memory channels through memory controllers 0 to 15 within die 106. Within die 106, a RBBM circuit block for each memory channel is placed adjacent to each memory channel to handle the flow control between the compute array rows associated with each memory channel. The RBBM circuit blocks are controlled by master controller 108 to perform read and write requests (e.g., "accesses") to the HBM.
[0026] In the example of FIG. 1, the memory channels are far from the remote buffers. The master controller 108 is also far from the memory channels located closer to the right side of die 106. Further, the data wavefront enters the compute array rows in a direction (e.g., horizontally) orthogonal to the direction in which the data wavefront enters IC100 through the memory controller (e.g., vertically).
[0027] Data read from the HBM is written from the RBBM circuit block to each of the remote buffers 0 to 127 within dies 102, 104, and 106. The read side of each remote buffer (e.g., the side connected to the cache) is controlled by the master controller 108. The master controller 108 performs a read deskewing operation across dies 102, 104, and 106 and controls the read side of each remote buffer 0 to 127 to supply data to their respective caches 0 to 31 for each of the various compute array rows 0 to 255.
[0028] The master controller 108 can coordinate data transfers to and from the remote buffers so that data is deskewed. By coordinating reads from the remote buffers, the master controller 108 ensures that data, such as weights, is provided synchronously to each compute array row. Further, the master controller 108 can improve and / or maximize HBM bandwidth usage. This allows the compute array to stay busy while more fully utilizing the read bandwidth of the HBM.
[0029] As mentioned, IC100 may have fewer or more dies than shown in FIG. 1. In this regard, the circuit architecture in the example of FIG. 1 can reduce the overhead and complexity of dispersing data for a computational array that is dispersed across multiple dies of a multi-die IC, regardless of whether such an IC includes fewer than three dies or more than three dies. The exemplary architectures described herein can be adapted to multi-die ICs having a different number of dies than shown. Further, the sizes of dies 102, 104, and 106 are for illustrative purposes only to better show the components within each respective die. The dies may be the same size or different sizes.
[0030] FIG. 2 shows an exemplary implementation of the circuit architecture of FIG. 1. In the example of FIG. 2, die boundaries have been removed. IC100 may be disposed on a circuit board communicatively linked to a host computer via a communication bus. For example, IC100 may be coupled to the host computer via a Peripheral Component Interconnect Express (PCIe) connection or other suitable connection. IC100, such as die 106, can include a PCIe direct memory access (DMA) circuit 202 to facilitate a PCIe connection. The PCIe DMA circuit 202 is connected to a block random access memory (BRAM) controller 204 via connection 206. In an exemplary implementation, one or more of the request buffer and / or remote buffer of the RBBM circuit block is implemented using BRAM.
[0031] The master controller 108 is connected to the BRAM controller 204. The BRAM controller 204 can operate as an AXI endpoint slave for integrating with the AXI interconnect and the system master device to communicate with local storage (e.g., BRAM). In the example of FIG. 2, the BRAM controller 204 can operate as a bridge between the PCIe and the master controller 108. In one aspect, the master controller 108 is a centralized controller driven by a command queue via a host computer-PCIe connection (e.g., when received via the PCIe DMA circuit 202 and the BRAM controller 204).
[0032] The master controller 108 is capable of performing a plurality of different operations. For example, the master controller 108 can perform a narrow write request to the HBM to initialize the HBM. In that case, the master controller 108 can access all of the 32 memory channels of the HBM through the AXI master 31 and the memory controller 15 (e.g., only by using the global address) of the (RBBM circuit block 31) using the global address.
[0033] The master controller 108 can also perform a narrow read request from the HBM. In that case, the master controller 108 can access all of the 32 memory channels of the HBM through the AXI master 31 and the memory controller 15 (e.g., only by using the global address) using the global address.
[0034] The master controller 108 can also perform wide read requests from the HBM. The master controller 108 can perform read requests (e.g., both sequential and random) from all 32 memory channels in parallel through bus master circuits 0 to 31 (e.g., one of the RBBM circuit blocks 0 to 31) for the parallel memory controllers 1 to 15 by using the local memory channel address.
[0035] In the example of FIG. 2, the master controller 108 can monitor and / or track various signals. Further, the master controller 108 can generate various different signals in response to the detection of a specific state within the monitored signals. For example, the master controller 108 can generate signal 208. Signal 208 is a remote buffer read enable signal. The master controller 108 can generate signal 208 and broadcast signal 208 (e.g., the same signal) to each of the remote buffers 0 to 127. In this way, the master controller 108 can synchronously read enable each remote buffer to dequeue the data read from the remote buffer and provided to the compute array rows.
[0036] Signal 210 is a remote buffer write enable signal. Each bus master circuit 0 to 31 within the RBBM circuit blocks 0 to 31 can generate signal 210 to the corresponding remote buffer. The master controller 108 can receive each of the remote buffer write enable signals generated by the bus master circuits 0 to 31 for each of the remote buffers. In one aspect, the master controller 108 can monitor the fill level of each remote buffer by tracking the write enable signal 210 from each remote buffer and the read enable signal 208 provided to each remote buffer.
[0037] Signal 212 is the same as signal 208. However, signal 212 is generated according to a clock signal different from signal 208 (for example, axi_clk instead of sys_clk). Therefore, the master controller 108 can also provide the remote buffer read enable signal broadcast to each of the remote buffers 0 to 127 to each of the AXI masters in the RBBM circuit blocks 0 to 31, for example, the RBBM circuit blocks 0 to 31. Therefore, the remote buffer fill level can also be tracked by a remote buffer pointer manager implemented locally in each RBBM circuit block. The remote buffer pointer manager(s) will be described in more detail in connection with FIG. 5.
[0038] Signal 214 represents the AXI-AW / W / B signal that the master controller 108 can provide to the RBBM circuit block 31 to start the narrow write as described above. In the present disclosure, AXI-AW refers to the AXI write address signal. AXI-W refers to the AXI write data signal. AXI-B refers to the AXI write response signal. AXI-AR refers to the AXI read address signal. AXI-R refers to the AXI read data signal. Further, the AXI master controller 108 receives signal 216 from each of the RBBM circuit blocks 0 to 31. Signal 216 may be an AR REQ ready signal (for example, here "AR" refers to "address read"). The master controller 108 can also broadcast signal 218, for example, an AR REQ broadcast, to each of the RBBM circuit blocks 0 to 31 to start the HBM read.
[0039] The exemplary circuit architecture of FIG. 2 includes a plurality of different clock domains. dsp_clk is used to clock the compute array rows and output ports of caches 0 - 31. In one example, dsp_clk is set to 710 MHz. An 8×16b connection between each of caches 0 - 31 and the eight compute array rows supplied by each respective cache achieves a data transfer rate of 0.355 TB / s (8×16×32 bits * 710 MHz).
[0040] sys_clk is used to clock the input ports (e.g., the right side) of caches 0 - 31 connected to the remote buffer and the output ports (e.g., the left side) of remote buffers 0 - 127 connected to caches 0 - 31. sys_clk is also used to clock a portion of master controller 108, for example, to broadcast signal 208 to each remote buffer. In one example, sys_clk is set to 355 MHz. An 8×32b connection between the illustrated remote buffers 0 - 127 and caches 0 - 31 achieves a data rate of 0.355 TB / s (8×32×32 bits * 355 MHz). For example, sys_clk may be set to 1 / 2 or approximately 1 / 2 the frequency of dsp_clk.
[0041] In the example of FIG. 2, caches 0 - 31 can not only cache data but also traverse clock domains. More specifically, each of caches 0 - 31 can receive data at the sys_clk rate and output data to the compute array rows at the dsp_clk rate, for example, at twice the input clock rate. In one or more exemplary embodiments, circuits such as the remote buffer and the RBBM circuit block can be implemented in programmable logic having a slower clock speed than other hardwired circuit blocks that can be used to implement the compute array rows. Thus, caches 0 - 31 can bridge this clock speed difference.
[0042] The axi_clk is used to clock the input port (e.g., the right side) of the remote buffer and the output port (e.g., the left side) of the RBBM circuit block. The axi_clk is also used to clock a portion of the master controller 108, for example, to monitor received signals 210 and 216 as well as output signals 212, 214, and 218. In one example, the axi_clk is set to 450 MHz. A 4×64b connection between the illustrated RBBM circuit blocks 0 to 31 and the remote buffers 0 to 127 achieves a data rate of 0.45 TB / s (4×64×32 bits * 450 MHz).
[0043] Each RBBM circuit block is coupled to the corresponding memory controller via a 256b connection, achieving a data rate of 0.45 TB / s (32×256 bits * 450 MHz). The memory controllers 0 to 15 may also be clocked at 450 MHz. Each memory controller supports two 64b memory channel connections (e.g., one for each memory channel) that provide a data rate of 0.45 TB / s (2048 bits / T * 1.8 GT / s).
[0044] For purposes of explanation, the term "memory channel interface" is used within this disclosure to refer to a particular RBBM circuit block and the corresponding portion of the memory controller to which the RBBM circuit block is connected (e.g., a single channel). For example, RBBM circuit block 0 and the portion of memory controller 0 connected to RBBM circuit block 0 (e.g., data buffer 302-0 and request queue 304-0, referring to FIG. 3) are a memory channel interface, while RBBM circuit block 0 and the portion of memory controller 0 connected to RBBM circuit block 1 (e.g., data buffer 302-1 and request queue 304-1) are considered a different memory channel interface.
[0045] The master controller 108 can generate read and write requests according to the HBM read and write commands received via the PCIe DMA 202 and the BRAM controller 204. The master controller 108 can operate in the "hurry up and wait" mode. For example, the master controller 108 can send read requests to the request buffers of the RBBM circuit blocks 0 to 31 until the data path including the request buffer and the remote buffer is full. In response to each read command, the master controller 108 can further initiate a data read operation (e.g., data transfer) from each remote buffer to the corresponding cache. Further, the master controller 108 releases some request buffer space and triggers the master controller 108 to generate new read requests to obtain additional data from the HBM based on the available space in the request buffer.
[0046] In an exemplary embodiment, the HBM includes 16 banks in each PC and 32 columns in each row. By interleaving the banks, up to 16×32×256 bits (128 Kb) can be read by each PC. Since the size of the BRAM is 4×36 Kb, one 36 Kb BRAM can be used to buffer two compute array rows. Thus, the exemplary circuit architecture of FIG. 2 can read up to 16 interleaved pages from one PC at a 512-bit burst length to service eight compute array rows at a time.
[0047] FIG. 3 shows another exemplary embodiment of the circuit architecture of FIG. 1. In the example of FIG. 3, the die boundary is removed. Further, each of the memory controllers 0 to 31 (abbreviated as "MC" in FIG. 3) is coupled to two memory channels. FIG. 3 presents a more detailed view of the memory controllers 0 to 15 and the RBBM circuit blocks 0 to 31.
[0048] In the example of FIG. 3, each of the memory controllers 0 to 15 services two memory channels. Thus, each of the memory controllers 0 to 15 includes one data buffer 302 for each memory channel to be serviced and one request queue 304 for each memory channel to be serviced. For example, memory controller 0 includes data buffer 302-0 and request queue 304-0 for servicing memory channel 0, and data buffer 302-1 and request queue 304-1 for servicing memory channel 1. Similarly, memory controller 15 includes data buffer 302-30 and request queue 304-30 for servicing memory channel 30, and data buffer 302-31 and request queue 304-31 for servicing memory channel 31.
[0049] Using the AXI protocol as an exemplary example, the data buffer 302 may be implemented as an AXI-R (read) data buffer. Each data buffer 302 can include 64×16 (1024) entries where each entry is 256 bits. The request queue 304 may be implemented as an AXI-AR (address read) request queue. Each request queue 304 can include 64 entries. Each data buffer 302 receives data from the corresponding memory channel. Each request queue 304 can provide commands, addresses, and / or control signals to the corresponding memory channel when received from the corresponding AXI master.
[0050] Each of the RBBM circuit blocks 0 to 31 includes a bus master circuit and a request buffer. For example, RBBM circuit block 0 includes bus master circuit 0 and request buffer 0. RBBM circuit block 1 includes bus master circuit 1 and request buffer 1. RBBM circuit block 30 includes bus master circuit 30 and request buffer 30. RBBM circuit block 31 includes bus master circuit 31 and request buffer 31. Accordingly, each bus master circuit has a data connection to the corresponding data buffer 302 and a control connection (e.g., for an address, control signal, and / or command) to the corresponding request queue 304.
[0051] Due to HBM refresh, clock domain crossing, and two memory channels being interleaved within a single memory controller, skew is introduced between data read from the HBM via different memory channels. In the case of memory channel skew, data read from the HBM via different memory channels is not bit-aligned. This is the same even if all read requests for all 32 memory channels are issued in parallel in the same cycle by each of the 32 AXI masters. The data skew is caused, at least in part, by the refresh of the HBM.
[0052] Consider an example where the HBM has a global refresh period of 260 ns every 3900 ns. In that case, the HBM throughput is limited to 0.42 TB / s by the refresh command ((3900 - 260) / 2900 * 0.45 = 0.42). This also means that there is a refresh window period of 117 axi_clk cycles (260 * 0.45 = 117) every 1755 axi_clk cycles (3900 * 0.45 = 1755), during which HBM read or write requests cannot be issued to the HBM via the memory channel. Since these refresh windows are not bit-aligned across all 32 memory channels, the maximum skew between any two memory channels is 117 axi_clk cycles if there is no overlapping refresh period between the two memory channels. If the memory controller can generate new requests every 2 axi_clk cycles, the skew can be up to 59 (117 / 2 = 59) HBM read requests between any two memory channels during the period when one of the two memory channels has already issued 59 read requests and the other memory channel is blocked by the execution of a refresh.
[0053] In the example of FIG. 3, each request queue 304 can be used to queue up to 64 HBM read requests initiated by master controller 108 via the corresponding bus master circuit during the refresh command period. In this case, 59 HBM read requests accumulated during the refresh command period can be absorbed into the 64-entry request queue 304. Master controller 108 can monitor the FIFO ready or full status of each request queue 304 by monitoring signals 216 from each RBBM circuit block (e.g., AR REQ ready signals 0-31 from all memory channels of the HBM for wide read requests). For example, based on the status of each request queue 304, in response to a determination that space within each data buffer 302 is available, master controller 108 can generate a new HBM read request (e.g., for each memory channel) and broadcast such requests to each memory controller 0-15 (e.g., via the bus master circuit). That is, master controller 108 sends the HBM read request to the request buffer. Each bus master circuit services requests from the local request buffer. Master controller 108 does not generate a new HBM read request if any of the request queues 304 is full.
[0054] For each fill level of each of the request queues 304 across each of the 32 memory channels, two cases can occur for HBM wide read requests. The first case corresponds to a state where the circuit architecture is ready for a new HBM read request. In the first case, buffer space within each (e.g., all) of the request queues 304 is available to receive a new HBM read request from the master controller 108. The second case corresponds to a state where the circuit architecture is not ready for a new HBM read request. In the second case, one or more of the request queues 304 are full and at least one other request queue 304 is neither full nor empty. Since the maximum skew (e.g., 59) between any two memory channels is less than the buffer size (e.g., 64) of the request queues 304, a situation where some request queues 304 are full and others are empty does not occur. In the second case, since there are still pending HBM read requests within all of the request queues 304, the HBM throughput is not affected by servicing new requests from the master controller 108.
[0055] Referring to both FIGS. 2 and 3, there are a total of 32 data streams. Each data stream is 256 bits wide and extends from a memory channel to a corresponding remote buffer. As shown in FIG. 1, some of the remote buffers are located within die 102 or 104, while other remote buffers are located closer to the master controller 108 within die 106. In some exemplary arrangements, since the data path for each memory channel is 256 bits wide, hardware resources can be minimized by keeping these data paths relatively short.
[0056] FIG. 4 shows an example of a seesaw structure used to implement the circuit architecture of FIG. 1. The seesaw structure is used to broadcast an HBM wide read request from the master controller 108 to the memory controller.
[0057] As shown in the example of FIG. 4, master controller 108 broadcasts signal 218 (AR REQ broadcast signal) from left to right to various RBBM circuit blocks. For purposes of illustration, only RBBM circuit blocks 31, 16, and 0 are shown. The arrival times of signal 218 at each RBBM circuit block are aligned in the same axi_clk cycle. Further, master controller 108 can broadcast signal 208 (e.g., remote buffer read enable) to each of the remote buffers. For purposes of illustration, only remote buffers 0-3, 64-67, and 124-127 are shown.
[0058] In the example of FIG. 4, each RBBM circuit block includes a remote buffer pointer manager 406 (shown as 406-31, 406-16, and 406-0). Remote buffer pointer manager 406 may be included as part of the request buffer or may be implemented separately from the request buffer within each respective RBBM circuit block. Each remote buffer pointer manager 406 can receive signal 208 for the purpose of tracking the fill level of the corresponding remote buffer. Further, each remote buffer pointer manager 406 can output signal 210 (e.g., remote buffer write enable signals 210-31, 210-16, and 210-0) to the corresponding remote buffer.
[0059] For example, the flip-flop (FF) 402 in die 104 receives a 256-bit wide data signal from the RBBM circuit block 31 in die 106. FF 402 passes the data to FF 404 in die 102. FF 404 passes the data to remote buffers 0 to 3. The remote buffer pointer manager 406-31 can output the control signal 210-31 to FF 408 in die 104. FF 408 outputs the control signal 210-31 to FF 410 in die 102. FF 410 outputs the control signal 210-31 to remote buffers 0 to 3. The master controller 108 generates a signal 208, for example, a remote buffer ready signal, and broadcasts it to the remote buffers 124 to 127 in die 106. The signal 208 follows to FF 412 in die 104. FF 412 outputs the signal 208 to remote buffers 64 to 67. FF 412 outputs the signal 208 to FF 414 in die 102. FF 414 provides the signal 208 to remote buffers 0 to 3.
[0060] The FF 416 in die 104 receives a 256-bit wide data signal from the RBBM circuit block 16 in die 106. FF 416 passes the data to remote buffers 64 to 67. The remote buffer pointer manager 406-16 can output the control signal 210-16 to FF 418 in die 104. FF 418 outputs the control signal 210-16 to remote buffers 64 to 67.
[0061] The RBBM circuit block 0 directly outputs a 256-bit wide data signal to the remote buffers 124 to 127 in die 106. The remote buffer pointer manager 406-0 directly outputs the control signal 210-0 to the remote buffers 124 to 127 in die 106.
[0062] Generally, the skew addressed by the configuration of the present invention described within the present disclosure has several different components. For example, skew includes components caused by HBM refresh and components caused by data propagation delays such as signals crossing clock domains, data latency along different paths within the same die, data latency along different paths across different dies, etc. HBM refresh is the largest cause of skew. This aspect of skew is mainly processed, e.g., deskewed, by the read data buffer inside the memory controller. Other skew components, even when considered together, which contribute less to the overall skew than HBM skew, can be processed by the remote buffer. For this reason, across 32 memory channels, there is no need to balance the latency of HBM read data from the memory controller to the remote buffer. The remote buffer has sufficient storage space to tolerate the remote buffer write-side skew introduced by the described latency imbalance, so data can be transmitted from the memory controller to the remote buffers on dies 102, 104, and 106 using a minimum pipeline stage.
[0063] The example of FIG. 4 also shows how the remote buffer can cause backpressure to each memory controller of each memory channel. The backpressure is generated, at least in part, by locally generating each of the control signals 210 at each respective remote buffer pointer manager 406 within each RBBM circuit block proximate to the memory channel. Further, the master controller 108 broadcasts the common signal 208 using a seesaw structure.
[0064] The read and write pointers of the remote buffer are generated in each respective remote buffer across dies 102, 104, and 106. Signal 210 should match the delay of the remote buffer write data. Due to the skew, write data latency, and remote data read enable latency between the write side of the remote buffer that receives data from various memory channels and signal 210, the remote buffer backpressure may include some buffer margin space to absorb that skew. Signals 210 and signal 216 (not shown in FIG. 4) are backpropagated to master controller 108. The delay of such signals can be balanced to ensure that master controller 108 can capture the signals in the same cycle to generate both a new HBM read request and signal 208.
[0065] FIG. 5 shows an exemplary embodiment of an RBBM circuit block as described within the present disclosure. In the example of FIG. 5, bus master circuit 502 includes two read channels. AXI master 502 includes an AXI-AR master 504 that can process read address data and an AXI-R master 506 that can process read data. The other three AXI write channels, namely, AXI-AW (write address), AXI-W (write data), AXI-B (write response) can be processed inside master controller 108. For illustrative purposes, the signal label "AXI_Rxxxx" is intended to refer to an AXI signal from the relevant communication specification having the prefix "AXI_R", while the signal label "AXI_ARxxxx" is intended to refer to an AXI signal from the relevant communication specification having the prefix "AXI_AR".
[0066] The request buffer 508 is the request buffer described above as part of the RBBM circuit block and is used to handle the HBM read request queue backpressure round-trip delay between the AR REQ ready (e.g., signal 216) signal and the AR REQ signal (e.g., signal 218). The depth of the request buffer 508 must be, for example, greater than the backpressure round-trip delay (e.g., less than 32).
[0067] The AR REQ (address read request) signal 218 can specify a read start address that is a 23-bit local address for a wide read request from the master circuit 108 via memory controllers 0 to 15, or a 28-bit global address for a narrow read request from the master circuit 108 made only via memory controller 15. The AR REQ signal 218 can also include a 6-bit read transaction identifier for narrow and wide identification and can include other stream-related information. The AR REQ signal 218 can also specify a read burst length that is 4 bits to support a burst length of up to 16 bytes depending on the particular AXI protocol used. In one exemplary embodiment, a 32×42 FIFO of the request buffer 508 can be implemented using three combinational logic block memories (CLBMs) where each CLBM is a 32×14 dual-port RAM. In that case, the request buffer 508 can service a new read request to the AXI-AR master 504 every axi_clk cycle.
[0068] The remote buffer pointer manager 406 is used to generate a signal 210 (remote buffer write enable) for the read data received from the HBM. The remote buffer pointer manager 406 can further generate a local remote buffer backpressure for the AXI-R (read data) channel corresponding to the AXI-R master 506. Due to the skew between different memory channels, each remote buffer pointer manager 406 can maintain the fill level of the corresponding remote buffer by tracking the remote buffer write enable and the remote buffer read enable. Each remote buffer pointer manager 406 can increment the remote buffer fill level when the HBM read data becomes valid upon assertion of the BRAM_WVALID signal. Each remote buffer pointer manager 406 can further decrement the remote buffer fill level when receiving a signal 208 (remote buffer read enable) broadcast from the master controller 108. The remote buffer pointer manager 406 can further generate a remote buffer backpressure signal (e.g., BRAM_WREADY) for each memory channel based on a predetermined remote buffer fill level. The remote buffer fill threshold for generating the backpressure signal should take into account both the skew between the remote buffer write-side latencies across the memory channels and the skew between the read side and the write side of the remote buffer.
[0069] As shown, the AXI-R master 506 can output data that can be provided to the corresponding remote buffer(s).
[0070] FIG. 6 shows an exemplary embodiment of master controller 108. In the example of FIG. 6, master controller 108 includes a remote buffer read address generation unit (remote buffer read AGU) 602 coupled to request controller 604. In the example of FIG. 6, master controller 108 converts both narrow read requests and wide read requests from BRAM controller 204 into AXI-AR requests for request buffers 0-31 (narrow read for 31 only). Further, master controller 108 converts narrow write requests into AXI-W and AXI-AW requests for master circuit 31 only.
[0071] In one aspect, simultaneously, for AXI-AR requests, the predicted read burst length from AXI-R is sent from request controller 604 to remote buffer read AGU 602. Remote buffer read AGU 602 can queue the received predicted read burst length therein, for example, in a queue or memory therein. Remote buffer read AGU 602 can maintain the remote buffer fill level of each remote buffer by tracking all of signal 210 (e.g., all of remote buffer write enable 0-31) and signal 208 (e.g., common remote buffer read enable). Remote buffer read AGU 602 can trigger, for example, continue to trigger remote buffer read operations (e.g., assertion of signal 208) as long as the remote buffer fill level of all remote buffers exceeds the predicted burst length queued therein. Signal 208 is used by each request buffer to generate remote buffer backpressure for each memory channel and by each remote buffer for dequeueing. Since the predicted burst length is predetermined and queued within remote buffer read AGU 602, the rearranging function that allows the output order to be different from the input order by forcing the AXI ID to the same value is disabled.
[0072] FIG. 7 shows an exemplary embodiment of the request controller 604 of FIG. 6. In the example of FIG. 7, the request controller 604 includes a transaction buffer 702 and a dispatcher 704. The transaction buffer 702 uses sys_clk as the write clock and axi_clk as the read clock to isolate two asynchronous clock domains. The request controller 604 further includes a plurality of controllers shown as an AR (address read) controller 706, an AXI-AW (address write) controller 708, an AXI-W (write) controller 710, and an AXI B detector 712 capable of detecting a valid response on the AXI B or AXI response channel.
[0073] In the example of FIG. 7, the AR controller 706 can check the request buffer backpressure as indicated by the signal 216 (e.g., AR REQ ready signals 0 to 31). The AR controller 706 can check the AR REQ ready signals 0 to 31 simultaneously, for example, in parallel. In one aspect, the AR controller 706 generates an AXI-AR REQ (HBM read request) only when there is available space in each request buffer.
[0074] For AXI read-related channels such as AXI-AR and AXI-R, a request buffer is required between the AXI master and the request controller 604. For AXI write-related channels such as AXI-AW, AXI-W, and AXI-B, the request controller 604 (e.g., controllers 708, 710, and 712) communicates directly with the AXI master 31 corresponding to the memory controller 15. When both the AXI ready indication and the AXI valid indication (AXI-XX-AWREADY / AWVALID or AXI-XX-WREADY / WVALID) are asserted within the same cycle, the request controller 604 outputs an AXI request and simultaneously reads a new request from the buffer.
[0075] The dispatcher 704 can route (e.g., dispatch) different AXI transactions from the transaction buffer 702 to the appropriate one of the controllers 706, 708, 710, and / or 712 based on the transaction type. The dispatcher 704 also schedules the AXI access sequence between AXI writes and AXI reads, and schedules the AXI access sequence during consecutive AXI write operations. For example, the dispatcher 704 does not issue a new AXI-AR request or a new AXI-AW transaction until a response from a previous AXI write transaction is received.
[0076] Referring to FIGS. 6 and 7, the remote buffer read AGU 602 receives each remote buffer write enable corresponding to each memory channel and the common remote buffer read enable. The system throughput can be measured at either the write side of the remote buffer or the read side of the remote buffer by counting data transfer elements having performance counters.
[0077] FIG. 8 shows a method 800 for transferring data between an HBM and a distributed computing array. The method 800 can be executed using a circuit architecture (referred to as a “system” with reference to FIG. 8) as described within this disclosure in relation to FIGS. 1-7.
[0078] In block 802, the system can monitor the fill levels in a plurality of remote buffers distributed across multiple dies. Each remote buffer of the plurality of remote buffers may also be configured to provide data to a compute array distributed across the plurality of dies. In block 804, the system can determine that each remote buffer of the plurality of remote buffers stores data based on the fill level. In block 806, the system can initiate a data transfer from each remote buffer of the plurality of remote buffers to the compute array across the plurality of dies in response to the determination. The data transfer can be synchronized (e.g., dequeued). For example, the data transfers occurring at each die are synchronized. Minimal pipelining is used to facilitate synchronization between dies. Data as output from the remote buffer to the compute array row is further dequeued.
[0079] In one aspect, the system initiates data transfer from each remote buffer by broadcasting a read enable signal to each remote buffer of the plurality of remote buffers. The read enable signal is a common read enable signal broadcast to each remote buffer.
[0080] In another aspect, the system can monitor the fill level by tracking, one-to-one, a plurality of write enables corresponding to the plurality of remote buffers and tracking a common read enable for each of the plurality of remote buffers.
[0081] In a particular embodiment, the system can receive data from HBM via a plurality of respective memory channels within a plurality of RBBM circuit blocks disposed in a first die of the plurality of dies. The plurality of RBBM circuit blocks supply data to each of the plurality of remote buffers.
[0082] The system can also convert a first request for access to memory into a second request compliant with an on-chip communication bus and provide the second request to a communication bus master circuit corresponding to each of the plurality of RBBM circuit blocks.
[0083] The system can provide data from each of a plurality of remote buffers to a plurality of cache circuit blocks distributed across a plurality of dies, with each cache circuit block connected to at least one of the plurality of remote buffers and a compute array. Each cache circuit block can be configured to receive data from a selected remote buffer at a first clock rate and output the data to the compute array at a second clock rate that exceeds the first clock rate.
[0084] FIG. 9 shows an exemplary architecture of a programmable device 900. The programmable device 900 is an example of a programmable IC and an adaptive system. In one aspect, the programmable device 900 is also an example of a system-on-chip (SoC). The programmable device 900 may be implemented using a plurality of interconnected dies, in which case the various programmable circuit resources shown in FIG. 9 are implemented across different interconnected dies. In one example, the programmable device 900 can be used to implement the exemplary circuit architectures described herein in connection with FIGS. 1-8.
[0085] In this example, the programmable device 900 includes a data processing engine (DPE) array 902, programmable logic (PL) 904, a processor system (PS) 906, a network-on-chip (NoC) 908, a platform management controller (PMC) 910, and one or more hardwired circuit blocks 912. A configuration frame interface (CFI) 914 is also included.
[0086] The DPE array 902 is implemented as a plurality of interconnected programmable data processing engines (DPEs) 916. The DPEs 916 can be arranged and configured in an array and are wire-connected. Each DPE 916 can include one or more cores 918 and a memory module (abbreviated as "MM" in FIG. 9) 920. In one aspect, each core 918 can execute program code stored in a core-specific program memory (not shown) included within each respective core. Each core 918 can directly access the memory module 920 within the same DPE 916 and the memory modules 920 of any other DPE 916 that is adjacent to the core 918 of the DPE 916 in the up, down, left, or right direction. For example, core 918-5 can directly read memory modules 920-5, 920-8, 920-6, and 920-2. Core 918-5 treats each of memory modules 920-5, 920-8, 920-6, and 920-2 as an integrated memory area (e.g., part of the local memory accessible to core 918-5). This facilitates data sharing between different DPEs 916 within the DPE array 902. In other examples, core 918-5 may be directly connected to the memory module 920 within another DPE.
[0087] The DPEs 916 are interconnected by programmable interconnect circuit elements. The programmable interconnect circuit elements can include one or more different independent networks. For example, the programmable interconnect circuit elements can include a streaming network (shaded arrows) formed from streaming connections and a memory-mapped network (hatched arrows) formed from memory-mapped connections.
[0088] By loading configuration data into the control registers of the DPE916 via a memory-mapped connection, each DPE916 and the components within it can be controlled independently. The DPE916 can be enabled / disabled on a per-DPE basis. Each core 918 can be configured to access only the memory modules 920 or a subset thereof as described, for example, to achieve separation of a core 918 or multiple cores 918 operating as a cluster. Each streaming connection can be configured to establish a logical connection only between selected DPE916s to achieve separation of a DPE916 or multiple DPE916s operating as a cluster. Since each core 918 can load program code specific to that core 918, each DPE916 can implement one or more different kernels therein.
[0089] In other aspects, the programmable interconnect circuit elements within the DPE array 902 can include an additional independent network such as a debug network that is independent from (e.g., distinct and separate from) the streaming connections and memory-mapped connections, and / or an event broadcast network. In some aspects, the debug network is formed from and / or is part of the memory-mapped connection.
[0090] The core 918 can be directly connected to an adjacent core 918 via an inter-core cascade connection. In one aspect, the inter-core cascade connection is a one-way direct connection between the cores 918 as shown. In another aspect, the inter-core cascade connection is a two-way direct connection between the cores 918. The operation of the inter-core cascade interface can also be controlled by loading configuration data into the control registers of the respective DPE916s.
[0091] In an exemplary embodiment, DPE916 does not include a cache memory. By omitting the cache memory, the DPE array 902 can achieve predictable, e.g., deterministic, performance. Further, since there is no need to maintain coherence between cache memories located in different DPE916s, significant processing overhead is avoided. In a further example, core 918 does not have input interrupts. Thus, core 918 can operate without being interrupted. Omitting input interrupts to core 918 also enables the DPE array 902 to achieve predictable, e.g., deterministic, performance.
[0092] The SoC interface block 922 operates as an interface that connects DPE916 to other resources of the programmable device 900. In the example of FIG. 9, the SoC interface block 922 includes a plurality of interconnected tiles 924 arranged in a row. In a particular embodiment, different architectures can be used to implement the tiles 924 within the SoC interface block 922, where each different tile architecture supports communication with different resources of the programmable device 900. The tiles 924 are connected such that data can propagate bidirectionally from one tile to another. Each tile 924 can operate as an interface to the column of DPE916 directly above it.
[0093] Tile 924 is connected to adjacent tiles, the DPE 916 immediately above, and circuit elements below, using streaming and memory mapped connections as shown. Tile 924 can also include a debug network that connects to a debug network implemented within the DPE array 902. Each tile 924 can receive data from another source such as the PS 906, PL 904, and / or another hardwired circuit block 912. Tile 924-1, for example, can provide those portions of the data addressed to the DPE 916 in the upper column to such DPE 916s, regardless of the application or configuration, while transmitting the data addressed to the DPE 916s in other columns to other tiles 924 such as 924-2 or 924-3. As a result, such tiles 924 can thus route the data addressed to the DPE 916s in each column.
[0094] In one aspect, the SoC interface block 922 includes two different types of tiles 924. The first type of tile 924 has an architecture configured to function as an interface only between the DPE 916 and the PL 904. The second type of tile 924 has an architecture configured to function as an interface between the DPE 916 and the NoC 908 and also between the DPE 916 and the PL 904. The SoC interface block 922 may include a combination of the first type of tile and the second type of tile, or only the second type of tile.
[0095] In one aspect, the DPE array 902 can be used to implement the compute arrays described herein. In that regard, the DPE array 902 may be distributed across a plurality of different dies.
[0096] PL904 is a circuit element that can be programmed to perform a specified function. As an example, PL904 may be implemented as a field programmable gate array type of circuit element. PL904 can include an array of programmable circuit blocks. As defined herein, the term "programmable logic" means a circuit element used to construct a reconfigurable digital circuit element. Programmable logic is formed from a number of programmable circuit blocks, sometimes called "tiles", that provide basic functionality. The topology of PL904, unlike a hardwired circuit element, is highly configurable. Each programmable circuit block of PL904 typically includes a programmable element 926 (e.g., a functional element) and a programmable interconnect 942. The programmable interconnect 942 provides the highly configurable topology of PL904. The programmable interconnect 942 can be configured wire-by-wire to provide connections between the programmable elements 926 of the programmable circuit blocks of PL904 and can be configured bit-by-bit (e.g., each wire carrying 1 bit of information), unlike the connections between DPE916 for example.
[0097] Examples of the programmable circuit blocks of PL904 include configurable logic blocks having look-up tables and registers. Unlike the hard-wired circuit elements, sometimes called hard blocks, described below, these programmable circuit blocks have undefined functions at the time of manufacture. PL904 can include other types of programmable circuit blocks that provide basic defined functions with more restricted programmability. Examples of these circuit blocks can include digital signal processing blocks (DSPs), phase-locked loops (PLLs), and block random access memories (BRAMs). These types of programmable circuit blocks are numerous, like the others of PL904, and are intermixed with the other programmable circuit blocks of PL904. These circuit blocks can also have an architecture that generally includes programmable interconnects 942 and programmable elements 926, and thus are part of the highly configurable topology of PL904.
[0098] Prior to use, the PL904, e.g., the programmable interconnects and programmable elements, must be programmed or "configured" by loading data, called a configuration bitstream, into internal configuration memory cells therein. The configuration memory cells, when loaded with the configuration bitstream, define how the PL904 is configured and how it operates, e.g., the particular functions to be executed. Within the present disclosure, a "configuration bitstream" is not equivalent to program code executable by a processor or computer.
[0099] In one aspect, the PL904 can be used to implement one or more of the components shown in FIGS. 1-7. For example, various buffers, queues, and / or controllers can be implemented using the PL904. In this regard, the PL904 can be distributed across multiple dies.
[0100] PS906 is implemented as a hardwired circuit element manufactured as part of programmable device 900. PS906 can be implemented as, or can include, any of a variety of different processor types, each of which can execute program code. For example, PS906 may be implemented as an individual processor, such as a single core that can execute program code. In another example, PS906 may be implemented as a multi-core processor. In yet another example, PS906 may include one or more cores, modules, coprocessors, I / O interfaces, and / or other resources. PS906 may be implemented using any of a variety of different types of architectures. Exemplary architectures that may be used to implement PS906 include, but are not limited to, the ARM processor architecture, the x86 processor architecture, the graphics processing unit (GPU) architecture, the mobile processor architecture, the DSP architecture, combinations of the foregoing architectures, or other suitable architectures capable of executing computer-readable instructions or program code.
[0101] NoC908 is a programmable interconnect network for sharing data between endpoint circuits within programmable device 900. The endpoint circuits can be located within DPE array 902, PL904, PS906, and / or selected hardwired circuit block 912. NoC908 can include a high-speed data path with dedicated switching. In one example, NoC908 includes one or more horizontal paths, one or more vertical paths, or both horizontal and vertical paths. The arrangement configuration and number of regions shown in FIG. 9 are merely examples. NoC908 is an example of a general infrastructure available within programmable device 900 for connecting selected components and / or subsystems.
[0102] Within the NoC908, the nets to be routed through the NoC908 are unknown until the user circuit design for implementation within the programmable device 900 is created. The NoC908 can be programmed by loading configuration data that defines how elements within the NoC908, such as switches and interfaces, are configured and operate to pass data between switches and between NoC interfaces to connect endpoint circuits, into internal configuration registers. The NoC908 is manufactured (e.g., wire-connected) as part of the programmable device 900 and is not physically modifiable, but can be programmed to establish connections between multiple different master circuits and multiple different slave circuits of the user circuit design. The NoC908 does not implement any data paths or routes within it upon power-up. However, when configured by the PMC910, the NoC908 implements data paths or routes between endpoint circuits.
[0103] The PMC910 plays a role in managing the programmable device 900. The PMC910 is a subsystem within the programmable device 900 that can manage other programmable circuit resources across the programmable device 900. The PMC910 can maintain a safe and secure environment, boot the programmable device 900, and manage the programmable device 900 during normal operation. For example, the PMC910 can provide integrated programmable control for power-on, boot / configuration, security, power management, safety monitoring, debugging, and / or error handling of multiple different programmable circuit resources (e.g., DPE array 902, PL904, PS906, and NoC908) of the programmable device 900. The PMC910 operates as a dedicated platform manager that separates the PS906 from the PL904. Thus, the PS906 and the PL904 can be managed, configured, and / or powered off and / or on independently of each other.
[0104] In one aspect, PMC910 can operate as a root of trust for the entire programmable device 900. As an example, PMC910 is responsible for authenticating and / or verifying a device image that includes configuration data for any of the programmable resources of programmable device 900 that can be loaded into programmable device 900. PMC910 can further protect programmable device 900 from tampering during operation. By operating as the root of trust for programmable device 900, PMC910 can monitor the operation of PL904, PS906, and / or any other programmable circuit resources that may be included in programmable device 900. The root of trust functionality as executed by PMC910 is distinct and separate from PS906 and PL904, and / or any operations executed by PS906 and / or PL904.
[0105] In one aspect, PMC910 operates on a dedicated power supply. Thus, PMC910 is powered by a power supply that is separate and independent from the power supplies of PS906 and PL904. This electrical independence enables PMC910, PS906, and PL904 to be protected from each other with respect to electrical noise and glitches. Further, while PMC910 is continuing to operate, the power supply of one or both of PS906 and PL904 can be turned off (e.g., put into a suspended or sleep mode). This feature enables any part of programmable device 900, such as PL904, PS906, NoC908, etc., that has had its power turned off to more quickly wake up and return to an operating state without the entire programmable device 900 having to execute the entire power-on and boot process.
[0106] PMC910 may be implemented as a processor with dedicated resources. PMC910 may include multiple redundant processors. The processor of PMC910 can execute firmware. The use of firmware provides flexibility in creating separate processing domains (distinguished from the "power domain" which may be specific to the subsystem) and supports the configurability and segmentation of global functions of the programmable device 900 such as reset, clocking, and protection. A processing domain can include a mix or combination of one or more different programmable circuit resources of the programmable device 900 (for example, a processing domain can include different combinations or devices from the DPE array 902, PS906, PL904, NoC908, and / or other hardwired circuit blocks 912).
[0107] The hardwired circuit block 912 includes dedicated circuit blocks manufactured as part of the programmable device 900. Although wired-connected, the hardwired circuit block 912 can be configured by loading configuration data into control registers to implement one or more different operating modes. Examples of the hardwired circuit block 912 can include input / output (I / O) blocks, transceivers for sending and receiving signals to and from circuits and / or systems external to the programmable device 900, memory controllers, and the like. Examples of multiple different I / O blocks can include single-ended and pseudo-differential I / O. Examples of transceivers can include high-speed differential clock transceivers. Other examples of the hardwired circuit block 912 include, but are not limited to, encryption engines, digital-to-analog converters (DACs), analog-to-digital converters (ADCs), and the like. Generally, the hardwired circuit block 912 is an application-specific circuit block.
[0108] In one aspect, the hardwired circuit block 912 can be used to implement one or more of the components shown in FIGS. 1-7. For example, various memory controllers and / or other controllers may each be implemented as the hardwired circuit block 912. In this regard, one or more of the hardwired circuit blocks 912 may be distributed across multiple dies.
[0109] CFI914 is an interface that can provide configuration data, such as a configuration bitstream, to the PL904 through it to implement different user-specified circuit elements and / or circuit elements therein. CFI914 is coupled to the PMC910 and is accessible by the PMC to provide configuration data to the PL904. In some cases, the PMC910 can first configure the PS906 so that the PS906 can provide configuration data to the PL904 via the CFI914 when the PS906 is configured by the PMC910. In one aspect, CFI914 has a built-in cyclic redundancy check (CRC) circuit element (e.g., a CRC32-bit circuit element) incorporated therein. Thus, any data loaded into the CFI914 and / or read back via the CFI914 can be checked for integrity by checking the value of the code attached to the data.
[0110] The various programmable circuit resources shown in FIG. 9 can be first programmed as part of the boot process of the programmable device 900. During execution, the programmable circuit resources can be reconfigured. In one aspect, the PMC910 can first configure the DPE array 902, the PL904, the PS906, and the NoC908. At any point during execution, the PMC910 can reconfigure all or part of the programmable device 900. In some cases, the PS906 can configure and / or reconfigure the PL904 and / or the NoC908 when first configured by the PMC910.
[0111] The exemplary programmable device described in connection with FIG. 9 is for illustrative purposes only. In other exemplary embodiments, the exemplary circuit architectures described herein may be implemented in custom multi-die ICs, such as application-specific ICs having multiple dies, and / or programmable ICs such as field programmable gate arrays (FPGAs) having multiple dies. Further, the particular techniques used to communicatively link dies within an IC package, such as a common silicon interposer having wiring to couple the dies, a multi-chip module, three or more stacked dies, etc., are not intended to limit the configurations of the invention described herein.
[0112] For purposes of explanation, specific nomenclature is set forth to provide a complete understanding of the various inventive concepts disclosed herein. However, it is to be understood that the terminology used herein is for the purpose of describing particular aspects of the configurations of the invention only and is not limiting.
[0113] As defined herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0114] As defined herein, the term “about” means nearly correct or accurate and that the value or quantity is close but not precise. For example, the term “about” can mean that the recited characteristic, parameter, or value is within a predetermined amount of the exact characteristic, parameter, or value.
[0115] As defined herein, the terms "at least one", "one or more", and "and / or" are open-ended expressions that are both conjunctive and disjunctive in operation, unless otherwise specified. For example, each of the expressions "at least one of A, B, and C", "at least one of A, B, or C", "one or more of A, B, and C", "one or more of A, B, or C", and "A, B, and / or C" means only A, only B, only C, A and B, A and C, B and C, or A, B, and C.
[0116] As defined herein, the term "automatically" means without human intervention. As defined herein, the term "user" means a human being.
[0117] As defined herein, the term "when" means "when" or "in response to" or "responsive to" depending on the context. Thus, the phrases "when determined to be" or "when [described state or event] is detected" can be interpreted to mean, depending on the context, "upon receiving the determination" or "in response to the determination" or "upon receiving the detection of [described state or event]" or "in response to the detection of [described state or event]," or "responsive to the detection of [described state or event]".
[0118] As defined herein, the term "responsive to" and terms similar thereto, such as "when", "when", or "upon receiving", mean to readily respond or react to an action or event. The response or reaction is performed automatically. Thus, when a second action is performed "responsive to" a first action, there is a causal relationship between the occurrence of the first action and the occurrence of the second action. "Responsive to" indicates a causal relationship.
[0119] As defined herein, the term "processor" means at least one hardware circuit. The hardware circuit can be configured to execute instructions included in program code. The hardware circuit may be an integrated circuit or may be incorporated into an integrated circuit.
[0120] As defined herein, the term "substantially" means that the recited characteristics, parameters, or values need not be achieved exactly, but that deviations or variations including, for example, tolerances, measurement errors, measurement accuracy limits, and other factors known to those of skill in the art may occur in an amount that does not preclude the effect the characteristic is intended to provide.
[0121] Terms such as first, second, etc. may be used herein to describe various elements. These elements are not to be limited by these terms since these terms are only used to distinguish one element from another unless specifically stated otherwise or the context clearly indicates otherwise.
[0122] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible embodiments of systems, methods, and computer program products according to various aspects of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of one or more executable instructions for performing the specified operation.
[0123] In some alternative embodiments, the operations recited within a block may be performed out of the order shown in the figures. For example, two blocks shown in succession may be executed substantially in parallel, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. In other instances, the blocks may generally be executed in ascending order, but in still other instances, one or more blocks may be executed in various orders, with the results being stored and utilized in subsequent blocks or other blocks that do not necessarily follow immediately. Also, it should be noted that each block of the block diagrams and / or flowchart diagrams, as well as combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a dedicated hardware-based system that performs the specified function or operation, or by a combination of dedicated hardware and computer instructions.
[0124] It is intended that the corresponding structures, materials, acts, and equivalents of all means-plus-function or step-plus-function elements found in the appended claims include any structure, material, or act for performing the recited function in combination with other claimed elements as specifically claimed.
[0125] The IC can include a plurality of dies. The IC can include a plurality of memory channel interfaces configured to communicate with a memory, and the plurality of memory channel interfaces can be disposed within a first die of the plurality of dies. The IC can include a computing array distributed across the plurality of dies and a plurality of remote buffers distributed across the plurality of dies. The plurality of remote buffers can be coupled to the plurality of memory channels and the computing array. The IC can also include a controller configured to determine that each of the plurality of remote buffers stores data therein and, in response, broadcast a read enable signal to each of the plurality of remote buffers to initiate a data transfer from the plurality of remote buffers to the computing array across the plurality of dies.
[0126] Data transfer can be synchronized to deskew the data carried by each transfer.
[0127] Each of the foregoing and other embodiments can optionally include one or more of the following features, either alone or in combination. One or more embodiments can include all of the following features in combination.
[0128] In one aspect, the IC can include a plurality of request buffer - bus master circuit blocks disposed within a first die, each request buffer - bus master circuit block being connected to one memory channel interface of a plurality of memory channel interfaces and at least one remote buffer of a plurality of remote buffers.
[0129] In another aspect, the IC can include a plurality of cache circuit blocks distributed across a plurality of dies, each cache circuit block being connected to at least one remote buffer of a plurality of remote buffers and a compute array.
[0130] In another aspect, each cache circuit block can be configured to receive data from a selected remote buffer at a first clock rate and output the data to the compute array at a second clock rate that exceeds the first clock rate.
[0131] In another aspect, the compute array includes a plurality of rows, and each die of the plurality of dies includes two or more of the plurality of rows.
[0132] In another aspect, each memory channel interface can provide data from the memory to two or more rows of the compute array.
[0133] In another aspect, the memory is a high - bandwidth memory. In another aspect, the memory is a double - data - rate random access memory.
[0134] In another aspect, the compute array implements a neural network processor, and the data specifies weights applied by the neural network processor.
[0135] In another aspect, each memory channel interface provides data from memory to two or more rows of the compute array.
[0136] In one aspect, the controller is disposed within an IC having a plurality of dies. The controller includes a request controller configured to convert a first request for access to memory into a second request compliant with an on-chip communication bus, and the request controller provides the second request to a plurality of request buffer-bus master circuit blocks configured to receive data from a plurality of channels of the memory. The controller further includes a remote buffer read address generation unit coupled to the request controller and configured to monitor the fill level in each of a plurality of remote buffers distributed across the plurality of dies. Each remote buffer of the plurality of remote buffers is configured to provide data obtained from one of each of the plurality of request buffer-bus master circuit blocks to a compute array also distributed across the plurality of dies. In response to a determination that each remote buffer of the plurality of remote buffers is storing data based on the fill level, the remote buffer read address generation unit is configured to initiate data transfer from each remote buffer of the plurality of remote buffers to the compute array across the plurality of dies.
[0137] Data transfers can be synchronized to dequeue the data carried by each transfer.
[0138] Each of the foregoing and other embodiments can optionally include one or more of the following features, alone or in combination. One or more embodiments can include all of the following features in combination.
[0139] In one aspect, the request controller can receive a first request at a first clock frequency and provide a second request at a second clock frequency.
[0140] In another aspect, the remote buffer read address generation unit can monitor the fill levels in each of a plurality of remote buffers by tracking a plurality of write enables corresponding to the plurality of remote buffers and tracking a common read enable for each of the plurality of remote buffers.
[0141] The method can include monitoring fill levels in a plurality of remote buffers distributed across a plurality of dies, where each remote buffer of the plurality of remote buffers is configured to provide data to a compute array also distributed across the plurality of dies. The method can also include determining, based on the fill levels, that each remote buffer of the plurality of remote buffers is storing data, and in response to the determination, starting a data transfer from each remote buffer of the plurality of remote buffers to the compute array across the plurality of dies.
[0142] The data transfers can be synchronized to dequeue the data carried by each transfer.
[0143] Each of the foregoing and other embodiments can optionally include, alone or in combination, one or more of the following features. One or more embodiments can include all of the following features in combination.
[0144] In one aspect, starting a data transfer from each remote buffer includes broadcasting a read enable signal to each remote buffer of the plurality of remote buffers.
[0145] In another aspect, monitoring the fill level can include tracking multiple write enables corresponding to multiple remote buffers and tracking a read enable common to each of the multiple remote buffers.
[0146] In another aspect, the method can include receiving data from memory via a respective one of a plurality of memory channels within a plurality of request buffer-bus master circuit blocks disposed in a first die of the plurality of dies, the plurality of request buffer-bus master circuit blocks supplying data to respective ones of the plurality of remote buffers.
[0147] In another aspect, the method can include converting a first request for access to memory into a second request compliant with an on-chip communication bus and providing the second request to a communication bus master circuit corresponding to each of the plurality of request buffers.
[0148] In another aspect, the method can include providing data from each of the plurality of remote buffers to a plurality of cache circuit blocks distributed across the plurality of dies, each cache circuit block being connected to at least one of the plurality of remote buffers and a compute array.
[0149] In another aspect, each cache circuit block can be configured to receive data from a selected remote buffer at a first clock rate and output the data to the compute array at a second clock rate that exceeds the first clock rate.
[0150] The description of the configuration of the present invention provided herein is presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed forms and examples. The terms used herein are selected to explain the principles of the configuration of the present invention, its actual use, or technical improvements to the technology found in the market, and / or to enable other persons skilled in the art to understand the configuration of the present invention disclosed herein. Changes and modifications will be apparent to persons skilled in the art without departing from the scope and spirit of the configuration of the present invention described. Therefore, reference should be made to the appended claims, rather than the foregoing disclosure, as indicating the scope of such features and embodiments.
Claims
1. An integrated circuit including a plurality of dies, A plurality of memory channel interfaces configured to communicate with a memory, the plurality of memory channel interfaces including a plurality of memory channel interfaces disposed within a first die of the plurality of dies, A computing array distributed across the plurality of dies, A plurality of remote buffers distributed across the plurality of dies, the plurality of remote buffers including a plurality of remote buffers coupled to the plurality of memory channel interfaces and the computing array, A controller configured to broadcast, to each of the plurality of remote buffers, a read enable signal that determines whether there is stored data for each of the plurality of remote buffers and, in response to the determination, initiates data transfer from the plurality of remote buffers to the computing array across the plurality of dies An integrated circuit comprising.
2. The integrated circuit of claim 1, further comprising a plurality of request buffer - bus master circuit blocks disposed within the first die, each request buffer - bus master circuit block being coupled to one of the plurality of memory channel interfaces and at least one of the plurality of remote buffers.
3. The integrated circuit of claim 1, further comprising a plurality of cache circuit blocks distributed across the plurality of dies, each cache circuit block being coupled to at least one of the plurality of remote buffers and the computing array.
4. The integrated circuit of claim 3, wherein each cache circuit block is configured to receive the data from a selected remote buffer at a first clock rate and output the data to the computing array at a second clock rate that exceeds the first clock rate.
5. The integrated circuit of claim 1, wherein the computing array comprises a plurality of rows and each die of the plurality of dies includes two or more of the plurality of rows.
6. The integrated circuit of claim 5, wherein each memory channel interface provides data from the memory to two or more rows of the computing array.
7. The integrated circuit according to claim 1, wherein the memory is a HBM (High Bandwidth Memory).
8. The integrated circuit according to claim 1, wherein the memory is a double data rate random access memory.
9. The integrated circuit according to claim 1, wherein the computing array implements a neural network processor, and the data specifies weights applied by the neural network processor.
10. The integrated circuit according to claim 1, wherein each memory channel interface provides data from the memory to two or more rows of the computing array.
11. A controller disposed within an integrated circuit having a plurality of dies, a request controller configured to convert a first request for access to a memory into a second request compliant with an on-chip communication bus, the request controller providing the second request to a plurality of request buffer - bus master circuit blocks configured to receive data from a plurality of channels of the memory, a remote buffer read unit coupled to the request controller and configured to monitor a fill level in each of a plurality of remote buffers distributed across the plurality of dies, each of the plurality of remote buffers being configured to provide data obtained from one of the respective plurality of request buffer - bus master circuit blocks to a computing array also distributed across the plurality of dies, comprising a controller configured to start a data transfer from each of the plurality of remote buffers to the computing array across the plurality of dies in response to a determination that each of the plurality of remote buffers stores data based on the fill level.
12. The controller according to claim 11, wherein the request controller receives the first request at a first clock frequency and provides the second request at a second clock frequency.
13. The remote buffer read unit monitors the fill level in each of the plurality of remote buffers by tracking a plurality of write enables corresponding to the plurality of remote buffers and tracking a common read disable for each of the plurality of remote buffers. The controller according to claim 11.
14. The plurality of request buffer - bus master circuit blocks each include a respective one of the plurality of request buffers, and the request controller is configured to initiate a read request to the memory in response to a determination that space is available in each of the plurality of request buffers. The controller according to claim 11.
15. The request controller A transaction buffer configured to separate a first clock domain from a second clock domain, A dispatcher coupled to the transaction buffer, A plurality of controllers coupled to the dispatcher, wherein a first subset of the plurality of controllers is configured to monitor the plurality of remote buffers, and a second subset of the plurality of controllers is configured to monitor the plurality of request buffer - bus master circuit blocks. A plurality of controllers Comprising The dispatcher is configured to route transactions to different ones of the plurality of controllers based on the transaction type. The controller according to claim 11.
Citation Information
Patent Citations
computer image processing pipeline
JP2016536692A
Programmable matrix processing engine
JP2018139097A
Laminate memory device, method for operating the same, and memory system
JP2019061677A
Method of processing in-memory command, high-bandwidth memory (HBM) implementing the same, and HBM system
JP2019075101A