Stacked memory for computational acceleration

By using through-silicon (TSV) between the DRAM die and the logic/processing die for vertical interconnection, and implementing detectors and logic circuits on the base die to form a processing component chain, the problem of low efficiency of DRAM die and logic/processing die for interconnection and data access in the prior art is solved, and neural network computing with high bandwidth and low latency is realized.

CN114174984BActive Publication Date: 2025-06-06RAMBUS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080052156.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-18
Filing Date
2020-07-06
Publication Date
2025-06-06
Estimated Expiration
2040-07-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of efficient interconnection and data access between dynamic random access memory (DRAM) die and logic/processing die, especially in neural network computing scenarios that achieve high bandwidth and low latency.

Method used

By using through-silicon (TSV) to vertical interconnects between the DRAM die and the logic/processing die, and implementing detector circuits and logic circuits on the base die, dynamically manage access to memory data, forming a chain of processing elements to perform neural network tasks.

Benefits of technology

High bandwidth and low latency access to DRAM memory is achieved, suitable for fast execution of neural networks, machine learning and artificial intelligence tasks, improving the computing efficiency and scalability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114174984B_ABST
    Figure CN114174984B_ABST
Patent Text Reader

Abstract

An integrated circuit comprising a set of one or more logic layers, when the integrated circuit is stacked in an assembly with the stacked memory devices, the set of one or more logic layers is electrically connected to a set of stacked memory devices. The set of one or more logic layers comprises a connection chain of processing elements. The processing elements in the connection chain can independently calculate partial results based on received data, store the partial results, and pass the partial results directly to the next processing element in the connection chain that implements the processing element. The processing elements in the chain can include interfaces that allow direct access to memory groups on one or more DRAMs in the stack. These interfaces can access DRAM memory groups via TSVs that are not used for global I / O. These interfaces allow processing elements to access data in DRAM more directly.
Need to check novelty before this filing date? Find Prior Art

Description

BRIEF DESCRIPTION OF THE DRAWINGS

[0001] Figure 1A to Figure 1B An example layout for chained processing elements is shown.

[0002] Figure 1C A first example processing element is shown.

[0003] Figure 1D A first example processing node of a processing element is shown.

[0004] Figure 1E A second example processing element is shown.

[0005] Figure 1F An example activation processing node of a processing element is shown.

[0006] Figure 1G Flexible processing nodes showing processing elements

[0007] Figure 2 An example high bandwidth memory (HBM) compatible process die with a ring bus is shown.

[0008] Figure 3 Further details regarding the HBM-compatible hierarchical buffer are shown.

[0009] FIG. 4A to FIG. 4B An example HBM-compatible processing component is shown.

[0010] FIG. 5A to FIG. 5B A block diagram illustrating an example HBM-compatible system configuration.

[0011] FIG. 6A to FIG. 6C A cross-sectional view of an example HBM-compatible component is shown.

[0012] Figure 7 An example layout for chained processing elements with through silicon vias (TSVs) to access DRAM banks is shown.

[0013] Figure 8 is an isometric view of an example chained processing element die stacked with at least one DRAM die.

[0014] FIG. 9A to FIG. 9B An example cross section of a stackable DRAM die is shown.

[0015] Fig. 9C An example cross section of a stackable substrate die is shown.

[0016] Fig.9D An example cross section of a stackable logic / processing die is shown.

[0017] Fig.9E An example stacked DRAM assembly is shown.

[0018] Fig.9F A stacked DRAM assembly is shown that is compatible with added logic / processing dies.

[0019] Figure 9G Shown is a stacked DRAM assembly with added logic / processing dies.

[0020] Figure 9H An example cross section of a stackable TSV redistribution die is shown.

[0021] Fig.9I A stacked DRAM assembly is shown using a TSV redistribution die to TSV-connect the logic / processing die to the DRAM die.

[0022] Fig.10 Example processing modules are shown.

[0023] FIG. 11A to FIG. 11B An example distribution of address bits is shown to accommodate a processing chain coupled to an HBM channel.

[0024] Fig.12 It is a block diagram of the processing system. DETAILED DESCRIPTION

[0025] In an embodiment, the interconnect stack of one or more dynamic random access memory (DRAM) dies has a base logic die and one or more custom logic or processor dies. The custom die can be attached as the last step and vertically interconnected with the DRAM die through a common through silicon via (TSV) connection, which transmits data and control signals throughout the stack. The circuit on the base die can send and receive data and control signals through an interface to an external processor and / or circuit. The detector circuit on the base die can (at least) detect the presence of the logic die, and if the logic die exists, respond by selectively disabling external reception and / or transmission of data and control signals, and if not, enable external reception and / or transmission. The detector circuit can also adaptively enable and disable external reception and / or transmission of data based on information from the SoC or the system to which it is connected. The logic circuit located on the base die or the logic die can selectively manage access to the memory data in the stack via data and control TSV.

[0026] In an embodiment, in addition to being suitable for incorporating a set of stacked DRAM dies, the logic die may also include one or more processing element chains connected. These processing elements may be designed and / or architected for rapid execution of artificial intelligence, neural networks, and / or machine learning tasks. Therefore, the processing element may be configured to, for example, perform one or more operations to implement a node of a neural network (e.g., multiplying a neural network node input value by a corresponding weight value and accumulating the result). In particular, the processing element in the chain may calculate partial results (e.g., accumulation of a subset of neuron weighted input values, and / or accumulation of a subset of products of matrix multiplication) based on data received from an upstream processing element, store the results, and pass the results (e.g., partial sums of neuron output values ​​and / or matrix multiplications) to downstream processing elements. Therefore, the processing element chain of an embodiment is well suited for parallel processing of artificial intelligence, neural networks, and / or machine learning tasks.

[0027] In an embodiment, the logic die has a centrally located global input / output (EO) circuit and TSVs that allow the logic die to interface with other dies in a stack (e.g., a high bandwidth memory type stack). Thus, the logic die can access data stored in DRAM, access data stored outside the stack (e.g., via a substrate die and TSVs), and / or be accessed by an external processor (e.g., via a substrate die and TSVs). The die may also include a buffer connected between the global IO circuit and the corresponding chain of processing elements. The individual buffers may be further interconnected into a ring topology. With this arrangement, a processing element chain can communicate with other processing element chains (via a ring), DRAMs in a stack (via global I / O), and external circuits (also via global I / O) via buffers. In particular, partial results can be passed from chain to chain via a ring without occupying the bandwidth of the global I / O circuit.

[0028] In an embodiment, a processing element of a chain may include interfaces that allow direct access to memory banks on one or more DRAMs in the stack. These interfaces may access DRAM memory banks via TSVs that are not used for global I / O. These additional (e.g., per processing element) interfaces may allow processing elements to access data in the DRAM stack more directly than using global I / O. This more direct access allows faster access to data in the DRAM stack to perform tasks such as (but not limited to): fast loading of weights to switch between neural network models, large neural network model spills, and fast storage and / or retrieval of activations.

[0029] Figure 1A An example layout for chained processing elements is shown. Figure 1A, processing elements 110a-110d are shown. Processing elements 110a-110d are connected in a chain to independently calculate complete or partial results as a function of received data (which can also be translated as calculating based on the received data), store these results, and pass these results directly to the next processing element in the chain of processing elements. Each processing element 110a-110d receives input via a first side and provides output via an adjacent side. By rotating and / or flipping the layout of each processing element 110a-110d, the same (except for the rotation and / or flipping) processing elements 110a-110d can be linked together so that the output of one processing element is consistent with the next processing element in the chain. Therefore, processing elements 110a-110d can be connected in a chain. Figure 1A , and connected together in the manner shown in , such that the input 151 of the chain of four processing elements 110a - 110d will be aligned with the output 155 of the four processing elements 110a - 110d.

[0030] Figure 1A The arrangement shown in allows for efficient formation of chains of more than four processing elements by aligning the outputs (e.g., 155) from one sub-chain of four processing elements with the inputs (e.g., 151) of the next sub-chain of four processing elements. It should also be understood that chains and / or sub-chains with other numbers of processing elements are also contemplated—e.g., 1 or 2 processing elements. It is also contemplated that chains may be formed in which the outputs from one sub-chain (e.g., a sub-chain of 1, 2, 3, 4, etc. processing elements 110a-110c) are not aligned with the inputs of the next sub-chain (of any number) of processing elements.

[0031] exist Figure 1A , input 151 of a chain of four processing elements 110a-110d is shown as being provided to the top of the page side of processing element 110a. Processing element 110a provides output 152 from the right side of processing element 110a. Processing element 110b is located to the right of processing element 110a. Output 152 from processing element 110a is received on the left side of processing element 110b. Processing element 110b provides output 153 from the bottom side of processing element 110b. Processing element 110c is located directly below processing element 110b. Output 153 from processing element 110b is received on the top side of processing element 110c. Processing element 110c provides output 154 from the left side of processing element 110c. Processing element 110d is located to the left of processing element 110c. Output 154 from processing element 110c is received on the right side of processing element 110d. The processing element 110d provides output 155 from the bottom of the page side of the processing element 110d. Figure 1AIt can be seen that the input 151 of the chain of four processing elements 110a-110d is received at a position aligned with the output of the chain of four processing elements 110a-110d from left to right. Therefore, it should be understood that one or more additional chains of four processing elements can provide input 151, receive output 155, or both. This is in Figure 1B Further explanation in.

[0032] Figure 1B An example layout of chain processing elements is shown. Figure 1B In FIG. 1 , an array of chain processing elements is shown. Figure 1B , chained processing array 101 includes processing elements 110a-110d, 111a-111d, 112a-112d, 113a-113d, 114a-114d, and 115a-115d. The inputs of the sub-chain of four processing elements 110a-110d are shown as being provided to the top of the page side of processing element 110a. The outputs from the sub-chain of four processing elements 110a-110d are shown as being provided from the bottom of processing element 110d and aligned from left to right with the inputs of processing elements 110a and 111a. The inputs of the sub-chain of four processing elements 111a-111d are shown as being provided to the top of the page side of processing element 111a. Thus, processing element 111a is the input interface from the chain of connections to the processing elements ( Figure 1B An input processing element (not shown) that receives data.

[0033] Outputs from a sub-chain of four processing elements 111a-111d are shown provided from the bottom of processing element 111d and routed to inputs of a sub-chain of four processing elements 112a-112d. The sub-chain of four processing elements 112a-112d is located at the bottom of a sub-chain of processing elements in a different column than processing elements 110a-110d and 111a-111d.

[0034] The inputs of the sub-chain of four processing elements 112a-112d are shown as being provided to the bottom of the page side of processing element 112a. The outputs from the sub-chain of four processing elements 112a-112d are shown as being provided from the top of processing element 112d and aligned from left to right with the inputs of processing elements 112a and 113a. This pattern is repeated for processing elements 113a-113d, 114a-114d, and 155a-115d. Processing element 115d provides an output from array 101 on the top of the page side of processing element 115d. Thus, processing element 115d is an output processing element that provides an output interface ( Figure 1B Data is provided (not shown).

[0035] Figure 1C A first example processing element is shown. Figure 1C , the processing element 110 includes processing nodes (PN) 140aa-140bb, an optional input buffer circuit 116, and an optional output buffer circuit 117. The processing nodes 140aa-140bb are arranged in a two-dimensional grid (array). The processing nodes 140aa-140bb are arranged so that each processing node 140aa-140bb receives input from the top of the page direction and provides output (result) to the next processing node on the right. The top row 140aa-140ab of the array of processing elements (PE) 110 receives corresponding inputs from the input buffer circuit 116. The rightmost column of the array of processing elements 110 provides corresponding outputs to the output buffer circuit 117. It should be understood that the processing element 110 is configured as a systolic array. Therefore, each processing node 140aa-140bb in the systolic array of processing elements 110 can work synchronously with its neighbors.

[0036] Note that, similar to processing elements 110a-110d, 111a-111d, 112a-112d, 113a-113d, 114a-114d, and 115a-115d, inputs to processing element 110 are received via a first side and outputs are provided via an adjacent side. Thus, similar to processing elements 110a-110d, 111a-111d, 112a-112d, 113a-113d, 114a-114d, and 115a-115d, by rotating and / or flipping the layout of multiple identical (except for the rotation and / or flipping) processing elements 110, multiple processing elements 110 may be chained together such that the output of one processing element is aligned with the input of the next processing element in the chain.

[0037] Figure 1DAn example processing node of a processing element is shown. Processing node 140 may be processing node 140aa-140bb, processing element 110, processing element 110a-110d, 111a-111d, 112a-112d, 113a-113d, 114a-114d, and / or 115a-115d. Processing node 140 includes memory function 145 (e.g., register), memory function 146 (e.g., register or SRAM), multiplication function 147, and addition (accumulation) function 148. The value in memory function 145 is received from the next processing node adjacent to (e.g., above) processing node 140 (or the input of the processing element). The value in memory function 145 is multiplied by the value in memory function 146 by multiplication function 147. The output of multiplication function 147 is provided to accumulation function 148. Accumulation function 148 receives a value from the next processing node on the left. The output of accumulation function 148 is provided to the next processing node (or output of a processing element) to the right. The value in storage function 145 is provided to the next processing node below.

[0038] Figure 1E A second example processing element is shown. Figure 1E , processing element 118 includes processing nodes 140aa-140bb, activation processing nodes 149a-149c, optional input buffer circuit 116 and optional output buffer circuit 117. Processing nodes 140aa-140bb are arranged in a two-dimensional grid (array). Processing nodes 140aa-140bb are arranged so that each processing node 140aa-140bb receives input from the top of the page direction and provides output (result) to the next processing node on the right. The output of processing node 149a-149c can also be based on the input received from input buffer circuit 116, which is relayed to the next processing node 140aa-140bb in the column through each processing node 140aa-140bb. The top row 140aa-140ab of processing element array 118 receives corresponding inputs from input buffer circuit 116. The rightmost column of the array of processing element 118 includes activation processing nodes 149a-149c. Activating processing nodes 149a - 149c provides corresponding outputs to output buffer circuit 117 .

[0039] Activation processing nodes 149a-149c may be configured to perform an activation function of a neural network node. The output of activation processing nodes 149a-149c is based on (at least) the input received by activation processing nodes 149a-149c from processing nodes 140aa-140bb to the left of activation processing nodes 149a-149c. The output of activation processing nodes 149a-149c may be further based on input received from input buffer circuit 116, which is relayed by each activation processing node 149a-149c to the next activation processing node 149a-149c in the column.

[0040] The activation functions implemented by the activation processing nodes 149a-149c can be linear or nonlinear functions. These functions can be implemented using logic, arithmetic logic units (ALUs), and / or one or more lookup tables. Examples of activation functions that can be used in neural network nodes include, but are not limited to: identity, binary step size, logic, Tanh, SQNL, ArcTan, ArcSinH, Softsign, inverse square root unit (ISRU), inverse linear square root unit (ISRLU), rectified linear unit (ReLU), bipolar rectified linear unit, leaky rectified linear unit (BReLU), leaky rectified linear unit (LeakyReLU), parameterized rectified linear unit (PReLU), exponential linear unit (ELU), scaled exponential linear unit (SELU), sigmoid rectified linear activation unit (SReLU), adaptive piecewise linear (APL), SoftPlus, Bentidentity, GELU, sigmoid linear unit (SiLU), SoftExponential, soft clipping, sinusoidal, sine, Gaussian, SQ-RBF, Softmax, and / or maxout.

[0041] exist Figure 1E , activated processing nodes (APN) 149a-149c are shown as being in the rightmost column and providing their outputs to output buffer circuit 117. It should be understood that this is an example. Embodiments are contemplated in which activated processing nodes 149a-149c occupy any or all rows and / or columns of processing elements 118 rather than activated processing nodes (PN) 140aa-140bb occupying the remaining positions in the array.

[0042] It should also be appreciated that processing element 118 is configured as a systolic array. Thus, each processing node 140aa-140bb and 149a-149c in the systolic array of processing element 118 may operate in lockstep with its neighboring nodes.

[0043] Note that, similar to processing elements 110a-110d, 111a-111d, 112a-112d, 113a-113d, 114a-114d, and 115a-115d, inputs to processing element 118 are received via a first side and outputs are provided via an adjacent side. Thus, similar to processing elements 110a-110d, 111a-111d, 112a-112d, 113a-113d, 114a-114d, and 115a-115d, by rotating and / or flipping the layout of multiple identical (except for the rotation and / or flipping) processing elements 118, multiple processing elements 118 may be chained together such that the output of one processing element is aligned with the input of the next processing element in the chain.

[0044] Figure 1F An example activation processing node of a processing element is shown. Activation processing node 149 may be or be a portion of processing node 140aa-140bb, activation processing node 149a-149c, processing element 110, processing element 110a-110d, 111a-111d, 112a-112d, 113a-113d, 114a-114d, 115a-115d, and / or processing element 118. Processing node 149 includes memory function 145 (e.g., register), memory function 146 (e.g., register or SRAM), multiplication function 147, addition (accumulation) function 148, and activation function 144. The value in memory function 145 is received from the next processing node (or input of the processing element) above processing node 149. The value in memory function 145 is multiplied by the value in memory function 146 by multiplication function 147. The output of multiplication function 147 is provided to accumulation function 148. Accumulator function 148 receives the value from the next processing node on the left. The output of accumulator function 148 is provided to activation function 144. The output of activation function 144 is provided to the next processing node (or output of a processing element) on the right. The value in memory function 145 is provided to the next processing node below.

[0045] It should be understood that activation processing node 149 is an example. A fewer or greater number of functions may be performed by activation processing node 149. For example, memory function 146, multiplication function 147, and / or accumulation function 148 may be removed, and activation function 144 may use only the inputs from the processing node to its left as inputs to the implemented activation function 144.

[0046] Figure 1GAn example processing node of a processing element is shown. Processing node 142 may be or may be a portion of processing nodes 140aa-140bb, activated processing nodes 149a-149c, processing element 110, processing elements 110a-110d, 111a-111d, 112a-112d, 113a-113d, 114a-114d, 115a-115d, and / or processing element 118. Processing node 142 includes processing system 143.

[0047] Processing system 143 may include and / or implement one or more of the following: memory functions (e.g., registers) and / or SRAM); multiplication functions, addition (accumulation) functions; and / or activation functions. At least one value (or input to a processing element) is received from the next processing node above processing node 142 and provided to processing system 143. Processing system 143 may be or include an application specific integrated circuit (ASIC) device, a graphics processor unit (GPU), a central processing unit (CPU), a system on a chip (SoC), or an integrated circuit device including a number of circuit blocks such as circuit blocks selected from a graphics core, a processor core, and an MPEG encoder / decoder.

[0048] The output of processing node 142 and / or processing system 143 is provided to the next processing node to the right (or output of a processing element).At least one value received from the next processing node above processing node 142 (or input of a processing element) may be provided to the next processing node below.

[0049] Figure 2 An example high bandwidth memory (HBM) compatible processing die with a ring bus is shown. Figure 2, the processing die 200 includes centrally located HBM compatible channel connections (e.g., TSY) 251-253, 255-257, hierarchical buffers 221a-223a, 221b-223b, 225a-227a, 225b-227b, and chains of processing elements 231-233, 235-237. The processing die 200 includes one or more logic layers for building circuits residing on the processing die 200. In an embodiment, the circuits of the processing die 200 can be integrated with the functions of the HBM substrate die. In another embodiment, the circuits of the processing die 200 can be on a separate die stacked with the HBM substrate die and one or more HBM DRAM dies. In an embodiment, processing die 200 is compatible with the connections of the HBM standard and therefore implements eight (8) channel connections 251-253, 255-257, sixteen (16) staging buffers 221a-223a, 221b-223b, 225a-227a, 225b-227b, and eight (8) processing element chains 231-233, 235-237. However, other numbers (e.g., 1, 2, 4, 6, 16, etc.) of processing chains and / or channel connections are contemplated.

[0050] The channel 251 is operably coupled to the staging buffer 221a. The staging buffer 221a is operably coupled to the input of the processing element chain 231. The output of the processing element chain 231 is operably coupled to the staging buffer 221b. The staging buffer 221b is operably coupled to the channel 251. Thus, the channel 251 can be used to provide input data to the staging buffer 221a. The staging buffer 221a can provide the input data to the processing element chain 231. The staging buffer 221b can receive result data from the processing element chain 231. The staging buffer 221b can provide the result data to the channel 251 for storage and / or other uses. The channels 252-253, 255-257 are operably coupled to the corresponding staging buffers 222a-223a, 222b-223b, 225a-227a, 225b-227b and the corresponding processing element chains 232-233, 235-237 in a similar manner.

[0051] The hierarchical buffers 221a-223a, 221b-223b, 225a-227a, 225b-227b are interconnected via a ring topology. The ring interconnection allows input data and / or output data (results) from a processing chain 231-233, 255-257 to communicate with any other processing chain 231-233, 255-257 and / or any channel 251-253, 255-257. Figure 2, two rings are shown transmitting data in opposite directions. However, it should be understood that a single ring or more than two rings are contemplated. For example, there may be a hierarchy of rings. In other words, in addition to Figure 2 In addition to the ring shown in , there can also be subsets of connecting channel interfaces (e.g., 251 and 257, 252 and 256, or groups of 4 channels, etc.) This allows the channel connections 251-253, 255-257 and processing element chains 232-233, 235-237 to be divided into logical units that can process different jobs simultaneously, but can also communicate across these partitions if necessary.

[0052] The configuration of the processing die 200 allows data passed through any channel 251-253, 255-257 to communicate with any processing chain 231-233, 235-237. Thus, for example, the processing die 200 can simultaneously run computations of N neural networks (one on each processing chain 231-233, 235-237), where N is the number of processing chains 231-233, 235-237 on the processing die 200 (e.g., N=8). In another example, because data for the input layer of the neural network can be communicated via any of the N channels 251-253, 255-257, fault tolerance can be improved by running computations of one neural network on multiple processing chains 231-233, 235-237.

[0053] In other examples, the resources of the processing die 200 can be distributed to perform distributed inference. One example of such a distribution is to provide 1 / N (e.g., N=8) samples to each neural network calculated on the corresponding processing chain 231-233, 235-237. For example, a convolutional neural network can be implemented by providing a copy of all weights to each processing chain 231-233, 235-237, and then having each processing chain apply a different portion of the filter. This parallelizes (by N) the application of filters to images and / or layers.

[0054] A further example distribution of the resources of the processing die 200 helps accelerate neural network training. One example is to have N (e.g., N=8) copies of the neural network computed by each processing chain 231-233, 235-237, and have them perform distributed gradient descent (e.g., providing 1 / N of the training samples to each processing chain 231-233, 235-237). In another distribution, one neural network computed on more than one (e.g., N) processing chains can be trained. In an embodiment, to facilitate training, the data flow direction between the inputs and outputs of the processing elements of the processing chains 231-233, 235-237 can be reversible to help support the reverse pass of the training algorithm.

[0055] Figure 3 Further details regarding the HBM-compatible hierarchical buffer are shown. Figure 3 An example circuit 300 is shown that can couple channel data between local channel connections, local processing chains, remote channels, and remote processing chains. Thus, for example, the circuit 300 can be local to the channel 251 and thus couple data between the local processing chain 231, the remote channels 252-253, 255-257 (e.g., via interconnections of additional instances of the circuit 300), and the remote processing chains 232-233, 235-237 (also, for example, via interconnections of additional instances of the circuit 300).

[0056] The circuit 300 includes a channel connection 350, a staging buffer 320a, a staging buffer 320b, and a control circuit 360. The staging buffers 320a-320b are operably coupled to the channel connection 350 and the local processing chain ( Figure 3 3. The control 360 is operably coupled to the staging buffers 320a-320b. The control 360 includes logic for configuring the staging buffers 320a-320b and the memory controller functions to allow access to data via the channel connection 350.

[0057] The staging buffers 320a-320b include a buffer for balancing between the channel 350 and the local processing chain ( Figure 3 320a and / or 320b may include memory elements (e.g., FIFO buffers) to help match data arrival and dispatch rates between processing chains and / or other staging buffers.

[0058] Figure 4A An exploded view of a first example HBM-compatible process assembly is shown. Figure 4A , the HBM-compatible component 401 includes a DRAM stack 470a and a substrate die 460a. The DRAM in the DRAM stack 470a includes memory groups 471-473 and a channel connection 475a. The substrate die 460a includes a channel connection 465a and processing chain circuits 431a-433a, 435a-437a. The channel connection 475a and the channel connection 465a include multiple independent memory channels for accessing the memory groups 471-473 of the DRAM in the memory stack 470a.

[0059] In an embodiment, each block of processing chain circuits 431a-433a, 435a-437a is locally coupled to one of a plurality of independent memory channels (e.g., 8 memory channels) such that each block of processing chain circuits 431a-433a, 435a-437a can access one or more memory banks 471-473 of DRAM in the memory stack 470a independently of each other block of processing chain circuits 431a-433a, 435a-437a. The processing chain circuits 431a-433a, 435a-437a may also be interconnected to share data and / or access one or more memory banks 471-473 of DRAM in the memory stack 470a that are accessed through channels that are not local to the corresponding processing chain circuits 431a-433a, 435a-437a.

[0060] Figure 4B An exploded view of a second example HBM-compatible process assembly is shown. Figure 4B , the HBM-compatible component 402 includes a processing die 410, a DRAM stack 470b, and a substrate die 480. The DRAMs in the DRAM stack 470b include memory groups 476-478 and channel connections 475b. The substrate die 480 includes a channel connection 485 and an external interface circuit 486. The processing die 410 includes a channel connection 465b and processing chain circuits 431b-433b, 435b-437b. The channel connection 465b, the channel connection 475b, and the channel connection 485 include multiple independent memory channels for accessing the memory groups 476-478 of the DRAMs in the memory stack 470b.

[0061] In an embodiment, each block of processing chain circuits 431b-433b, 435b-437b is locally coupled to one of a plurality of independent memory channels (e.g., 8 memory channels) such that each block of processing chain circuits 431b-433b, 435b-437b can access one or more memory banks 476-478 of DRAM in the memory stack 470b independently of each other block of processing chain circuits 431b-433b, 435b-437b. The processing chain circuits 431b-433b, 435b-437b may also be interconnected to share data and / or access one or more memory banks 476-478 of DRAM in the memory stack 470b that are accessed through channels that are not local to the respective processing chain circuits 431b-433b, 435b-437b. The external interface circuit 486 is locally coupled to one or more of a plurality of independent memory channels (eg, 8 memory channels) such that circuitry external to the component 402 can independently access one or more memory banks 476-478 of DRAMs in the memory stack 470b.

[0062] Figure 5A is a block diagram showing a first example HBM-compatible system configuration. Figure 5A , the processing system configuration 501 includes a memory stack assembly 505, an interposer 591, a memory PHY 592, and a processor 593. The processor includes a memory controller 594. The memory stack assembly 505 includes a stacked DRAM device 570 stacked with a base die 580. The base die 580 includes a logic die detection 585, a memory PHY 586, a 2:1 multiplexer (MUX) 587, an isolation buffer 588, and an isolation buffer 589.

[0063] The base die 580 is operably coupled to the DRAMS of the DRAM stack 570 via the memory PHY 582, the data signal 583, and the logic die detection signal 584. The memory control signal 581 is coupled to the top of the DRAM stack 570 through the DRAM stack 570. In an embodiment, the memory control signal 581 is not operably coupled to the active circuits of the DRAM stack 570, so that the memory control signal 581 is not operably coupled to the active circuits of the DRAM stack 570. Figure 5A In another embodiment, one or more memory control signals 581 may be configured to communicate with one or more dies of the DRAM stack 570. The data signal communicates with the base die 580 and the processor 593 via the interposer 591. The memory control signal communicates with the base die 580 and the memory controller 594 via the memory PHY 592 and the interposer 591.

[0064] Based at least in part on the logic state of the logic die detection signal 584, the base die 580: enables the isolation buffer 588 to transmit the data signal 583 with the processor 593; enables the isolation buffer 589 to transmit the memory control signal, and; controls the MUX 587 to use the memory control signal from the isolation buffer 589 as the memory PHY signal 582 provided to the DRAM stack 570. Therefore, it should be understood that in Figure 5A In the configuration shown in , the memory PHY 586 and the memory control signal 581 may not be used and may be passive. It should also be understood that in this configuration, the component 505 may appear as a standard HBM compatible component to the processor 593 (or other external devices / logic).

[0065] Figure 5B is a block diagram showing a second example HBM-compatible system configuration. Figure 5B, the processing system configuration 502 includes a memory stack assembly 506, an interposer 591, a memory PHY 592, and a processor 593. The processor includes a memory controller 594. The memory stack assembly 506 includes a stacked DRAM device 570 stacked with a substrate die 580 and a logic die 510. The substrate die 580 includes a logic die detection 585, a memory PHY 586, a 2:1 multiplexer (MUX) 587, an isolation buffer 588, and an isolation buffer 589. The logic die 510 includes a die detection signal generator 511, a processing element 513, and a memory controller 514.

[0066] The substrate die 580 is operably coupled to the DRAM die of the DRAM stack 570 via a memory PHY signal 582, a data signal 583, and a logic die detection signal 584. The memory control signal 581 is coupled to the logic die 510 through the DRAM stack 570. The substrate die 580 is operably coupled to the logic die 510 via a memory control signal 581, a memory PHY signal 582, a data signal 583, and a logic die detection signal 584.

[0067] The data signal may communicate with the base die 580 and the processor 593 via the interposer 591. The memory control signal may communicate with the base die 580 and the memory controller 594 via the memory PHY 592 and the interposer 591.

[0068] Based at least in part on the logic state of the logic die detection signal 584, the base die 580: prevents the isolation buffer 588 from transmitting the data signal 583 with the processor 593; prevents the isolation buffer 589 from transmitting the memory control signal; controls the MUX 587 to use the memory control signal 581 from the memory controller 514 relayed through the memory PHY 586 as the memory PHY signal 582 provided to the DRAM stack 570. Therefore, it should be understood that in this configuration, the memory controller 514 (via the memory PHY 586 and the MUX 587) controls the DRAMs of the DRAM stack 570. Likewise, data to / from the DRAM stack 570 is communicated with the processing element 513 of the logic die 510 without interference from the processor 593 and / or the memory controller 594.

[0069] However, in an embodiment, processing element 513 and / or processor 593 may configure / control base die 580 such that processor 593 may access DRAM stack 570 to access inputs and / or outputs computed by processing element 513. In this configuration, component 505 may appear as a standard-compatible HBM component to processor 593 (or other external devices / logic).

[0070] FIG. 6A to FIG. 6C is a cross-sectional view of an example HBM-compatible component. Fig. 6A , the HBM-compatible component 605 includes a DRAM stack 670 and a base die 680. The base die 680 includes bumps 687 to operably couple the component 605 to an external circuit. The base die 680 may include TSVs 685 to communicate signals (locally or externally) with the DRAM stack 670. The DRAM stack 670 includes bumps 677 and TSVs 675 to operably couple the DRAMs of the DRAM stack 670 to the base die 680. One or more of the TSVs 685 may or may not be aligned with one or more TSVs 675 of the DRAM stack 670. The component 605 may be, for example, Figure 5A Component 505 shown in .

[0071] exist Figure 6B , the HBM-compatible component 606a includes a DRAM stack 670, a substrate die 680, and a logic die 610. The logic die 610 includes a bump 688 to operably connect the component 606a to an external circuit. The substrate die 680 may include a TSV 685 to transmit signals (locally or externally) with the DRAM stack 670 and the logic die 610. The logic die 610 may include a TSV 615 to transmit signals (locally or externally) to the DRAM stack 670 and the substrate die 680. The DRAM stack 670 includes bumps and TSVs to operably connect the DRAMs of the DRAM stack 670 to the substrate die 680, and to operably connect the logic die 610 to the substrate die 680 and / or the DRAMs of the DRAM stack 670. One or more of the TSVs 615 may or may not be aligned with one or more TSVs 675 of the DRAM stack 670 and / or the TSVs 685 (if present) of the substrate die 680. Component 606a may be, for example, Figure 5B Component 506 shown in .

[0072] exist Figure 6C , the HBM-compatible component 606b includes a DRAM stack 670, a substrate die 680, and a logic die 611. The substrate die 680 includes bumps 687 to operably connect the component 606b to an external circuit. The substrate die 680 may include TSVs 685 to transmit signals (locally or externally) to the DRAM stack 670 and the logic die 611. The logic die 611 may transmit signals (locally or externally) to the DRAM stack 670 and the logic die 680. The DRAM stack 670 includes bumps and TSVs to operably connect the DRAMs of the DRAM stack 670 to the substrate die 680, and to operably connect the logic die 611 to the substrate die 680 and / or the DRAMs of the DRAM stack 670. The component 606b may be, for example, Figure 5B Component 506 shown in .

[0073] Figure 7 An example layout of chained processing elements with TSV access to DRAM banks is shown. Figure 7 , processing elements 710a-710d are shown. Processing elements 710a-710d include TSVs 717a-717d, respectively. TSVs 717a-717d may be used by processing elements 710a-710d to access a die stacked with die holding processing elements 710a-710d ( Figure 7 DRAM memory group on (not shown).

[0074] In addition to accessing DRAM memory, each processing element 710a-710d can receive inputs via a first side and provide outputs via an adjacent side. By rotating and / or flipping the layout of each processing element 710a-710d, identical (except for the rotation and / or flipping) processing elements 710a-710d can be chained together so that the output of one processing element is aligned with the input of the next processing element in the chain. Thus, processing elements 710a-710d can be arranged in a Figure 7 The arrangement and connection together in the manner shown is such that the input 751 of the chain of four processing elements 710a-710d will align with the output 755 of the four processing elements 710a-710d. This allows the formation of chains of more than four processing elements.

[0075] exist Figure 7 , input 751 to a chain of four processing elements 710a-710d is shown as being provided to the top of the page side of processing element 710a. Processing element 710a provides output 752 from the right side of processing element 710a. Processing element 710b is positioned to the right of processing element 710a. Output 752 from processing element 710a is received on the left side of processing element 710b. Processing element 710b provides output 753 from the bottom side of processing element 710b. Processing element 710c is positioned directly below processing element 710b. Output 753 from processing element 710b is received on the top side of processing element 710c. Processing element 710c provides output 754 from the left side of processing element 710c. Processing element 710d is positioned to the left side of processing element 710c. Output 754 from processing element 710c is received on the right side of processing element 710d. Processing element 710d provides output 755 from the bottom of the page side of processing element 710d. Figure 7It can be seen that the input 751 of the chain of four processing elements 710a-710d is received in a position aligned from left to right with the output of the chain of four processing elements 710a-710d. Therefore, it should be understood that one or more additional chains of four processing elements can provide input 751, receive output 755, or both.

[0076] As described herein, processing elements 710a-710d may use TSVs 717a-717d to access a die stacked with the processing elements 710a-710d. Figure 7 DRAM memory bank on the DRAM memory bank not shown in FIG. Figure 8 Further description.

[0077] Figure 8 is an exploded isometric view of an example chained processing element die stacked with at least one DRAM die. Figure 8 8, assembly 800 includes a processing die 810 stacked with at least a DRAM die 870. Processing die 810 includes channel connections (e.g., TSVs) 850, hierarchical buffers 820a-820b, and processing elements 810a-810d. Processing elements 810a-810d include and / or are coupled to TSV connections 817a-817d, respectively. In an embodiment, channel connections 850 of processing die 810 are connections compatible with the HBM standard.

[0078] DRAM die 870 includes channel connections (e.g., TSVs) 875 and DRAM memory groups 870a-870d. DRAM memory groups 870a, 870c, and 870d include and / or are coupled to TSV connections 877a, 877c, and 877d, respectively. DRAM memory group 870b also includes and / or is coupled to TSV connections. However, in Figure 8 In FIG. 8 , these TSV connections are shielded by the processing die 810 and are therefore not shown in FIG. Figure 8 In an embodiment, the channel connections 875 of the DRAM die 810 are connections compatible with the HBM standard.

[0079] The TSV connections 817a, 817c, and 817d of the processing elements 810a, 810c, and 810d of the processing die 810 are aligned with the TSV connections 877a, 877c, and 877d of the DRAM groups 870a, 870c, and 810d of the DRAM dies 870a, 870c, and 870d, respectively. Similarly, the TSV connection 817b of the processing element 810b of the processing die is aligned with the TSV connection 817b of the DRAM group 870b (in Figure 8The channel connections 850 of the processing die 810 are aligned with the channel connections 875 of the DRAM die 870. Therefore, when the processing die 810 and the DRAM die 870 are stacked on top of each other, the TSV connections 817a-817d of the processing elements 810a-810d of the processing die 810 are electrically connected to the TSV connections (e.g., 877a, 877c, and 877d) of the DRAM groups 870a-870d of the DRAM die 870. Figure 8 815a, 815c, and 815d. Similarly, channel connection 850 of processing die 810 is electrically connected to channel connection 875 of DRAM die 870. Figure 8 8 is shown by TSV representation 815 .

[0080] The TSV connections between the processing elements 810a-810d and the DRAM groups 870a-870d allow the processing elements 810a-810d to access the DRAM groups 870a-870d. The TSV connections between the processing elements 810a-810d and the DRAM groups 870a-870d allow the processing elements 810a-810d to access the DRAM groups 870a-870d without data flowing via channel connections 850 and / or channel connections 875. In addition, the TSV connections between the processing elements 810a-810d and the DRAM groups 870a-870d allow the processing elements 810a-810d to access the respective DRAM groups 870a-870d independently of each other. The processing elements 810a-810d accessing respective DRAM banks 870a-870d independently of one another allows the processing elements 810a-810d to access respective DRAM banks 870a-870d in parallel - thereby providing high bandwidth and lower latency from memory to the processing elements.

[0081] High bandwidth from memory to processing elements helps speed up the computations performed by neural networks and improves the scalability of neural networks. For example, in some applications, neural network model parameters (weights, biases, learning rates, etc.) should be quickly exchanged to a new neural network model (or part of a model). Otherwise, the time spent loading the neural network model parameters and / or data is more than the time spent computing the results. This is also known as the "batch size = 1 problem". For example, this can be particularly problematic in data centers and other shared infrastructures.

[0082] In an embodiment, processing elements 810a-810d and a stack of multiple DRAM dies ( Figure 8 The TSV connections between the DRAM groups 870a-870d (not shown) can be made in a common bus type configuration. In another embodiment, the processing elements 810a-810d and the stacked multiple DRAM dies ( Figure 8The TSV connections between DRAM groups 870a-870d (not shown) can be made in a point-to-point bus type configuration.

[0083] Component 800 provides (at least) two data paths for large-scale neural network data movement. The first path can be configured to move training and / or inference data to a processing element input layer (e.g., when the input layer of the neural network is being implemented on the first element of the processing chain) and to move output data from the output layer to memory (e.g., when the output layer of the neural network is being implemented on the last element of the processing chain). In an embodiment, this first path can be provided by channel connections 850 and 875. As described herein at least with reference to Figures 1A-1D and Figure 7 As described, a processing chain may be provided by the configuration and interconnection of processing elements 810a-810d.

[0084] The second path can be configured to load and / or store neural network model parameters and / or intermediate results in parallel to / from multiple processing elements 810a-810d through TSV interconnects (e.g., 815a, 815c, and 815d). Because each processing element loads / stores in parallel with other processing elements 810a-810d, for example, systolic array elements can be updated quickly (relative to using channel connections 850 and 875).

[0085] Figures 9A-9I Some of the components and fabrication steps that may be used to create a processing die / DRAM die stack are shown. Fig.9A A first example cross section of a stackable DRAM die is shown. Fig.9A , DRAM die 979 includes active circuit layer 977, TSV 975, and unthinned bulk silicon 973. In an embodiment, DRAM die 979 may be used as the top die of the HBM stack.

[0086] Fig. 9B A second example cross section of a stackable DRAM die is shown. Fig. 9B , DRAM die 971 includes active circuit layer 977, TSV 975, and bulk silicon 972. Note that die 971 is identical to die 979 except that a portion of bulk silicon 973 has been removed (e.g., by thinning until TSV 975 is exposed on the back side of die 971).

[0087] Fig. 9C An example cross section of a stackable substrate die is shown. Fig. 9C In FIG. 9 , base die 960 includes active circuit layer 967, TSV 965, and bulk silicon 962. Note that die 960 has been thinned until TSV 965 is exposed on the back side of die 960.

[0088] Fig.9D An example cross section of a stackable logic / processing die is shown. Fig.9D In FIG. 9 , processing / logic die 910 includes active circuit layer 917, TSV 915, and bulk silicon 912. Note that die 910 has been thinned until TSV 915 is exposed on the back side of die 910.

[0089] Fig.9E An example stacked DRAM assembly is shown. Fig.9E , DRAM component 981 (e.g., HBM compatible component) includes base die 960 stacked with DRAM stack 970. DRAM stack 970 includes multiple thinned dies (e.g., die 971) with unthinned dies stacked at the top of the stack (e.g., die 979). A perimeter of support / filling material 974 is also included as part of component 981. It should be understood that component 981 can be a standard HBM component shipped from a manufacturer.

[0090] Fig.9F A stacked DRAM assembly is shown that is compatible with added logic / processing dies. Fig.9F 9, DRAM component 982 includes base die 960 stacked with DRAM stack 970. DRAM stack 970 includes multiple thinned dies (e.g., die 971). A perimeter of support / fill material 974a is also included as part of component 982. It should be understood that component 982 can be a standard HBM component (e.g., component 981), such as shipped from a manufacturer with bulk silicon 973 with an unthinned top die removed (e.g., by thinning).

[0091] Figure 9G A stacked DRAM assembly is shown with logic / processing die added. Figure 9G 960 is stacked with a DRAM stack 970 and a logic die 910. The DRAM stack 970 includes a plurality of thinned dies (e.g., die 971). A perimeter of support / fill material 974b is also included as part of the assembly 983. The logic die 910 is attached (TSV-to-TSV) to the DRAM die at the end of the stack 970 opposite the base die 960. Note that in Figure 9G , this component is shown in the opposite orientation from component 982 so that logic die 910 appears attached to Figure 9G The bottom DRAM die.

[0092] Figure 9H An example cross section of a stackable TSV redistribution die is shown. Figure 9H, the base die 990 includes a circuit layer 997, TSVs 995, and bulk silicon 992. In an embodiment, the circuit layer 997 does not include active circuits (e.g., active transistors, etc.) and therefore includes conductive elements (e.g., metal wiring, vias, etc.). Note that the die 990 has been thinned until the TSVs 995 are exposed on the back side of the die 990.

[0093] Fig.9I A stacked DRAM assembly is shown using TSV redistribution dies to connect the logic / processing die TSVs to the DRAM die TSVs. Fig.9I , DRAM assembly 984 includes a base die 960 stacked with a DRAM stack 970, a redistribution die 990, and a logic die 911. The TSVs of the logic die 911 are not aligned with the TSVs of the DRAM stack 970. The DRAM stack 970 includes multiple thinned dies (e.g., die 971). The perimeter of the support / filling material 974c is also included as part of the assembly 982. The logic die 911 is attached (TSV to TSV) to the redistribution die 990. The redistribution die 990 attaches the circuit layer (on the die 990) to the TSY (on the DRAM stack 970). The redistribution die 990 is attached to the DRAM die at the end of the stack opposite to the base die 960.

[0094] Fig.10 An example processing module is shown. Fig.10 1 , module 1000 includes substrate 1096 , components 1081a - 1081d , and system 1095 . In an embodiment, system 1095 is a system on a chip (SoC) including at least one processor and / or memory controller. System 1095 is disposed on substrate 1096 .

[0095] Components 1081a-1081d include stacks of DRAM dies, respectively, and at least one includes a processing die 1010a-1010d. Components 1081a-1081d are disposed on substrate 1096. In an embodiment, system 1095 can access components 1081a-1081d using an address scheme that includes an indication of which component (stack), which channel of the component, and which row, group, and column of the channel is being processed. This is Fig.11A In another embodiment, the system 1095 may access the components 1081a-108Id using an address scheme that includes an indication of which component (stack), which channel of the component, which processing element on the selected component, and which row, group, and column of the channel is being processed. Fig. 11B Further shown in .

[0096] The above methods, systems and devices can be implemented in a computer system or stored by a computer system. The above methods can also be stored on a non-transitory computer-readable medium. The devices, circuits and systems described herein can be implemented using computer-aided design tools available in the art and embodied by computer-readable files containing software descriptions of such circuits. This includes but is not limited to processing array 101, processing element 110, processing node 140, processing node 142, processing node 149, die 200, circuit 300, component 401, component 402, system 501, system 502, component 605, component 606a, component 606b, component 800, die 910, die 960, die 971, die 979, component 981, component 982, component 983, component 984, die 990, module 1000 and one or more of their components. These software descriptions can be: behavior, register transfer, logic components, transistors and layout geometry level descriptions. Furthermore, the software description may be stored on a storage medium or communicated via a carrier wave.

[0097] Data formats in which such descriptions may be implemented include, but are not limited to, formats supporting behavioral languages ​​such as C, formats supporting register transfer level (RTL) languages ​​such as Yerilog and VHDL, formats supporting geometric description languages ​​such as GDSII, GDSIII, GDSIV, CIF, and MEBES, and other suitable formats and languages. In addition, data transmission of such files on machine-readable media may be accomplished electronically through various media on the Internet or, for example, via email. Note that the physical file may be implemented on a machine-readable medium such as: 4 mm tape, 8 mm tape, 3-1 / 2 inch floppy disk media, CD, DVD, etc.

[0098] Fig.12 1 is a block diagram illustrating one embodiment of a processing system 1200 for including, processing, or generating representations of circuit components 1220. The processing system 1200 includes one or more processors 1202, memory 1204, and one or more further communication devices 1206. The processor 1202, memory 1204, and communication devices 1206 communicate using any suitable type, number, and / or configuration of wired and / or wireless connections 1208.

[0099] Processor 1202 executes instructions of one or more processes 1212 stored in memory 1204 to process and / or generate circuit components 1220 in response to user input 1214 and parameters 1216. Process 1212 may be any suitable electronic design automation (EDA) tool or portion thereof for designing, simulating, analyzing, and / or verifying electronic circuits, and / or generating photomasks for electronic circuits. As shown in the figure, representation 1220 includes data describing all or part of processing array 101, processing element 110, processing node 140, processing node 142, processing node 149, die 200, circuit 300, component 401, component 402, system 501, system 502, component 605, component 606a, component 606b, component 800, die 910, die 960, die 971, die 979, component 981, component 982, component 983, component 984, die 990, module 1000, and their components.

[0100] The representation 1220 may include one or more of behavioral, register transfer, logic component, transistor, and layout geometry level descriptions. In addition, the representation 1220 may be stored on a storage medium or communicated via a carrier wave.

[0101] Data formats in which representation 1220 may be implemented include, but are not limited to, formats supporting behavioral languages ​​such as C, formats supporting register transfer level (RTL) languages ​​such as Verilog and VHDL, formats supporting geometry description languages ​​such as GDSII, GDSIII, GDSIV, CIF, and MEBES, and other suitable formats and languages. Additionally, data transmission of such files on machine-readable media may be accomplished electronically through various media on the Internet, or, for example, via email.

[0102] User input 1214 may include other parameters from a keyboard, mouse, voice recognition interface, microphone and speaker, graphical display, touch screen, or other type of user interface device. The user interface may be distributed among multiple interface devices. Parameters 1216 may include input to help define the specifications and / or characteristics of representation 1220. For example, parameters 1216 may include defining device types (e.g., NFET, PFET, etc.), topologies (e.g., block diagrams, circuit descriptions, schematics, etc.), and / or device descriptions (e.g., device attributes, device dimensions, supply voltages, simulated temperatures, simulation models, etc.).

[0103] Memory 1204 includes any suitable type, number, and / or configuration of non-transitory computer-readable storage media that stores processes 1212 , user inputs 1214 , parameters 1216 , and circuit components 1220 .

[0104] The communication device 1206 includes any suitable type, number, and / or configuration of wired and / or wireless devices that transmit information from the processing system 1200 to another processing or storage system (not shown) and / or receive information from another processing or storage system (not shown). For example, the communication device 1206 can transmit the circuit component 1220 to another system. The communication device 1206 can receive the process 1212, the user input 1214, the parameter 1216, and / or the circuit component 1220 and store the process 1212, the user input 1214, the parameter 1216, and / or the circuit component 1220 in the memory 1204.

[0105] Implementations discussed herein include, but are not limited to, the following examples:

[0106] Example 1: An integrated circuit comprising: a set of one or more logic layers for docking to a set of stacked memory devices when the integrated circuit is stacked with the set of stacked memory devices; the set of one or more logic layers comprises: a connected chain of processing elements, wherein the processing elements in the connected chain are used to independently calculate partial results as a function of received data, store the partial results and pass the partial results directly to the next processing element in the connected chain of processing elements.

[0107] Example 2: The integrated circuit of Example 1, wherein the coupled chain of processing elements includes an input processing element to receive data from an input interface to the coupled chain of processing elements.

[0108] Example 3: The integrated circuit of Example 2, wherein the connected chain of processing elements includes an output processing element configured to deliver a result to an output interface of the connected chain of processing elements.

[0109] Example 4: The integrated circuit of Example 3, wherein the integrated circuit forms a processing system when stacked with a set of stacked memory devices.

[0110] Example 5: The integrated circuit of Example 4, wherein the set of one or more logic layers further comprises: a centrally located region of the integrated circuit comprising global input and output circuit devices, the global input and output circuit devices being used to interface with the processing system and an external processing system.

[0111] Example 6. The integrated circuit of Example 5, wherein the set of one or more logic layers further comprises: a first hierarchical buffer coupled between the global input and output circuitry and the coupled chain of processing elements for transmitting data to at least one of: an input processing element and an output processing element.

[0112] Example 7. An integrated circuit according to Example 6, wherein the set of one or more logical layers also includes: multiple connection chains of processing elements and multiple hierarchical buffers, wherein a corresponding hierarchical buffer in the multiple hierarchical buffers is connected between a global input and output circuit device and a corresponding one of the multiple connection chains of the processing elements, so as to transmit data with at least one of the following items: a corresponding input processing element and a corresponding output processing element of a corresponding one of the multiple connection chains of the processing elements.

[0113] Example 8. An integrated circuit configured to be attached to a stack of memory devices and to dock with a stack of memory devices, the integrated circuit comprising: a first group of processing elements connected in a first chain topology, wherein the processing elements in the first chain topology are used to independently calculate partial results using received data, store the partial results, and pass the partial results directly to the next element in the first chain topology.

[0114] Example 9. The integrated circuit of Example 8, wherein the first chain topology includes a first input processing element to receive data from a first input interface of the first chain topology.

[0115] Example 10. The integrated circuit of Example 9, wherein the first chain topology comprises a first output processing element to deliver a result to a first output interface of the first chain topology.

[0116] Example 11. The integrated circuit of Example 10, wherein the first input processing element and the first output processing element are the same processing element.

[0117] Example 12. The integrated circuit of Example 10, further comprising: a centrally located region of the integrated circuit comprising global input and output circuitry for interfacing with the processing system and an external processing system.

[0118] Example 13. The integrated circuit of Example 12, further comprising: a first hierarchical buffer coupled between the first input interface, the first output interface, and the global input and output circuitry.

[0119] Example 14. The integrated circuit according to Example 13 further includes: a second group of processing elements connected in a second chain topology, wherein the processing elements in the second chain topology are used to independently calculate partial results using received data, store the partial results and pass the partial results directly to the next element in the second chain topology, wherein the second chain topology includes a second input processing element and a second output processing element, the second input processing element is used to receive data from a second input interface of the second chain topology, and the second output processing element is used to pass the result to a second output interface of the second chain topology; and a second hierarchical buffer connected between the second input interface, the second output interface and the global input and output circuit device.

[0120] Example 15. A system comprising: a group of stacked memory devices, including memory unit circuit devices; one or more group of processing devices, electrically connected to the group of stacked memory devices, the group of processing devices comprising: at least two processing elements of a first group, connected in a chain topology, wherein the processing elements in the first group are used to independently calculate partial results using received data, store the partial results and pass the partial results directly to the next element in the chain topology, wherein the first group comprises a first input processing element and a first output processing element, the first input processing element is used to receive data from a first input interface to the first group, and the first output processing element is used to pass the result to a first output interface of the first group.

[0121] Example 16. The system of Example 15, wherein the set of processing devices further comprises:

[0122] At least two processing elements of the second group are connected in a chain topology, wherein the processing elements in the second group use the received data to independently calculate partial results, store the partial results and directly pass the partial results to the next element in the chain topology, wherein the second group includes a second input processing element and a second output processing element, the second input processing element is used to receive data from the second input interface to the second group, and the second output processing element is used to pass the result to the second output interface of the second group.

[0123] Example 17. A system according to Example 16, wherein the group of processing devices further includes: a group of hierarchical buffers connected in a ring topology, a first at least one hierarchical buffer in the group of hierarchical buffers is connected to a first input interface to provide data to a first input processing element, and a second at least one hierarchical buffer in the group of hierarchical buffers is connected to a second input interface to provide data to a second input processing element.

[0124] Example 18. A system according to Example 16, wherein at least a third staging buffer in a set of staging buffers is connected to a first output interface to receive data from a first output processing element, and at least a fourth staging buffer in a set of staging buffers is connected to a second output interface to receive data from a second output processing element.

[0125] Example 19. The system of Example 18, wherein the set of processing devices further comprises: a memory interface connected to the set of hierarchical buffers and connectable to an external device outside the system, the memory interface being used to perform operations for the external device to access the set of stacked memory devices.

[0126] Example 20. The system of Example 19, wherein the memory interface is to perform operations for accessing a set of hierarchical buffers for an external device.

[0127] Example 21. A system comprising: a group of stacked memory devices, each stacked memory device including multiple memory arrays, the multiple memory arrays are accessed via a centrally located global input and output circuit device, each of the multiple memory arrays is also accessed via a corresponding array access interface independently of other memory arrays in the multiple memory arrays; a group of one or more processing devices, electrically connected to the group of stacked memory devices and stacked with the group of stacked memory devices, each processing device in the group of one or more processing devices is connected to at least one array access interface of the group of stacked memory devices, the group of processing devices comprising: at least two processing elements of a first group, connected in a chain topology, wherein the processing elements in the first group independently calculate partial results using received data, store the partial results, and directly pass the partial results to the next processing element in the chain topology.

[0128] Example 22. The system of Example 21, wherein the array access interface is connected to a corresponding processing device in a set of one or more processing devices using through silicon vias (TSVs).

[0129] Example 23. A system according to Example 22, wherein the first group further includes a first input processing element and a first output processing element, the first input processing element being used to receive data from a global input and output circuit device via a first input interface to the first group, and the first output processing element being used to deliver a result to the global input and output circuit device via a first output interface of the first group.

[0130] Example 24. A system according to Example 23, wherein a group of processing devices also includes: at least two processing elements of a second group, connected in a chain topology, wherein the processing elements in the second group independently calculate partial results using received data, store the partial results and pass the partial results directly to the next element in the chain topology, wherein the second group includes a second input processing element and a second output processing element, the second input processing element is used to receive data from a second input interface to the second group, and the second output processing element is used to pass the result to the global input and output circuit device via the second output interface of the second group.

[0131] Example 25. A system according to Example 24, wherein the group of processing devices also includes: a group of hierarchical buffers connected in a ring topology, a first at least one hierarchical buffer in the group of hierarchical buffers is connected to a first input interface to provide data to a first input processing element, and a second at least one hierarchical buffer in the group of hierarchical buffers is connected to a second input interface to provide data to a second input processing element.

[0132] Example 26. A system according to Example 25, wherein at least a third staging buffer in a set of staging buffers is connected to a first output interface to receive data from a first output processing element, and at least a fourth staging buffer in a set of staging buffers is connected to a second output interface to receive data from a second output processing element.

[0133] Example 27. The system of Example 26, wherein the set of processing devices further comprises: a memory interface connected to a set of hierarchical buffers and connectable to an external device outside the system, the memory interface being used to perform operations for the external device to access a set of stacked memory devices.

[0134] Example 28. The system of Example 27, wherein the memory interface is to perform operations for accessing a set of hierarchical buffers for an external device.

[0135] Example 29. A system, the system comprising: a group of stacked devices, including a group of stacked memory devices and at least one logic device; the stacked memory devices comprising: multiple memory arrays; a first interface, addressable to access all multiple memory arrays on the corresponding memory device; and multiple second interfaces, the multiple second interfaces access corresponding subsets of the multiple memory arrays of the corresponding memory device; the logic device comprising: a connected chain of processing elements, wherein the processing elements in the connected chain are used to independently calculate partial results as a function of received data, store the partial results and pass the partial results directly to the next processing element in the connected chain of processing elements, each processing element in the processing elements is connected to at least one second interface of the multiple second interfaces.

[0136] Example 30. The system of Example 29, wherein the coupled chain of processing elements comprises an input processing element to receive data from an input interface to the coupled chain of processing elements.

[0137] Example 31. The system of example 30, wherein the connected chain of processing elements comprises an output processing element to deliver a result to an output interface of the connected chain of processing elements.

[0138] Example 32. The system of Example 31, wherein the logic device further comprises: a centrally located region of the logic device, the centrally located region comprising global input and output circuitry for interfacing with a system and an external processing system.

[0139] Example 33. The system of Example 32, wherein the logic device further comprises: a first hierarchical buffer coupled between the global input and output circuitry and the coupled chain of processing elements to transfer data with at least one of: an input processing element and an output processing element.

[0140] Example 34. A system according to Example 33, wherein the logic device further includes: multiple connection chains of processing elements and multiple hierarchical buffers, wherein corresponding hierarchical buffers in the multiple hierarchical buffers are connected between the global input and output circuit device and a corresponding one of the multiple connection chains of the processing elements to transmit data with at least one of the following: a corresponding input processing element and a corresponding output processing element of a corresponding one of the multiple connection chains of the processing elements.

[0141] Example 35. A component comprising: a plurality of stacked dynamic random access memory (DRAM) devices; and at least two logic dies, also stacked with the plurality of DRAM devices, a first at least one of the at least two logic dies being attached to one of the top and bottom sides of the stacked plurality of DRAM devices, and a second at least one of the at least two logic dies being attached to an opposite side of one of the top and bottom sides of the stacked plurality of DRAM devices.

[0142] Example 36. The assembly of Example 35, wherein a first at least one logic die of the at least two logic dies is attached with an active circuit side of the first at least one logic die of the at least two logic dies facing a passive circuit side of the stacked plurality of DRAM devices.

[0143] Example 37. The assembly of Example 36, wherein a second at least one of the at least two logic dies is attached with a passive circuit side of the second at least one of the at least two logic dies facing a passive circuit side of the stacked plurality of DRAM devices.

[0144] Example 38. The assembly of Example 35, wherein the assembly includes a die that redistributes through silicon via (TSV) locations between a stacked plurality of DRAM devices and one of the at least two logic dies.

[0145] Example 39. The assembly of Example 35, wherein the assembly includes a die that redistributes through silicon via (TSV) locations between the stacked plurality of DRAM devices and at least one logic die of the at least two logic dies.

[0146] Example 40. The component of Example 35, wherein a first at least one logic die of the at least two logic dies is a substrate die compatible with a high bandwidth memory component.

[0147] Example 41. The assembly of Example 40, wherein a second at least one logic die of the at least two logic dies comprises a computing accelerator.

[0148] Example 42. A component according to Example 41, wherein the computing accelerator includes a connected chain of processing elements, wherein the processing elements in the connected chain independently calculate partial results as a function of received data, store the partial results, and pass the partial results directly to the next processing element in the connected chain of processing elements.

[0149] Example 43. The assembly of Example 42, wherein the processing elements in the linked chain are configured as a systolic array.

[0150] The foregoing description of the invention has been presented for the purpose of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise form disclosed, and other modifications and variations are possible in light of the above teachings. The embodiments are selected and described in order to best explain the principles of the invention and its practical application, thereby enabling others skilled in the art to best utilize the invention in various embodiments and various modifications suitable for the intended specific use. Unless limited by the prior art, the appended claims are intended to be interpreted as including other alternative embodiments of the invention.

Claims

1. An integrated circuit, include: a set of one or more logic layers for interfacing to a set of stacked memory devices when the integrated circuit is stacked with the set of stacked memory devices; The set of one or more logical layers includes: a linked chain of processing elements, wherein the processing elements in the linked chain are to independently compute partial results as a function of received data, store the partial results, and pass the partial results directly to the next processing element in the linked chain of processing elements, Wherein each processing element in the coupled chain of processing elements includes an interface that allows direct access to one or more memory devices in a set of stacked memory devices when the integrated circuit is stacked with the set of stacked memory devices.

2. The integrated circuit of claim 1, wherein the coupled chain of processing elements comprises an input processing element to receive data from an input interface to the coupled chain of processing elements.

3. The integrated circuit of claim 2, wherein the coupled chain of processing elements comprises an output processing element to deliver a result to an output interface of the coupled chain of processing elements.

4. The integrated circuit of claim 3, wherein the integrated circuit forms a processing system when stacked with the set of stacked memory devices.

5. The integrated circuit of claim 4, wherein the set of one or more logic layers further comprises: include: A centrally located region of the integrated circuit includes global input and output circuitry for interfacing the processing system with an external processing system.

6. The integrated circuit of claim 5, wherein the set of one or more logic layers further comprises: include: A first hierarchical buffer is coupled between the global input and output circuitry and the coupled chain of processing elements for communicating data with at least one of: the input processing element and the output processing element.

7. The integrated circuit of claim 6, wherein the set of one or more logic layers further comprises: include: A plurality of connection chains of processing elements and a plurality of staging buffers, wherein a corresponding staging buffer of the plurality of staging buffers is connected between the global input and output circuit device and a corresponding one of the plurality of connection chains of processing elements for transmitting data with at least one of the following items: a corresponding input processing element and a corresponding output processing element of a corresponding one of the plurality of connection chains of processing elements.

8. An integrated circuit configured to be attached to and interface with a stack of memory devices, the integrated circuit include: a first set of processing elements connected in a first chain topology, wherein the processing elements in the first chain topology are to independently calculate partial results using received data, store the partial results, and pass the partial results directly to the next element in the first chain topology, Wherein each processing element of the first set of processing elements includes an interface that allows direct access to one or more memory devices in the stack of memory devices when the integrated circuit is attached to the stack of memory devices and interfaces with the stack of memory devices. 9 . The integrated circuit of claim 8 , wherein the first chain topology comprises a first input processing element configured to receive data from a first input interface of the first chain topology. 10 . The integrated circuit of claim 9 , wherein the first chain topology comprises a first output processing element, the first output processing element being configured to deliver a result to a first output interface of the first chain topology.

11. The integrated circuit of claim 10, wherein the first input processing element and the first output processing element are the same processing element.

12. The integrated circuit of claim 10, further comprising: include: A centrally located region of the integrated circuit includes global input and output circuitry for interfacing the stack of memory devices and the integrated circuit to an external processing system.

13. The integrated circuit according to claim 12, further comprising: include: A first hierarchical buffer is coupled between the first input interface, the first output interface, and the global input and output circuit device.

14. The integrated circuit according to claim 13, further comprising: include: A second group of processing elements connected in a second chain topology, wherein the processing elements in the second chain topology are used to independently calculate partial results using received data, store the partial results, and directly pass the partial results to the next element in the second chain topology, wherein the second chain topology includes a second input processing element and a second output processing element, the second input processing element is used to receive data from a second input interface of the second chain topology, and the second output processing element is used to pass the result to a second output interface of the second chain topology; as well as A second hierarchical buffer is coupled between the second input interface, the second output interface, and the global input and output circuit device.

15. A system, include: A stacked set of memory devices including memory cell circuitry; a set of one or more processing devices electrically coupled to the set of stacked memory devices, the set of processing devices comprising: at least two processing elements of a first group are connected in a chain topology, wherein the processing elements in the first group are used to independently calculate partial results using received data, store the partial results and directly pass the partial results to the next element in the chain topology, wherein the first group includes a first input processing element and a first output processing element, the first input processing element is used to receive data from a first input interface to the first group, and the first output processing element is used to pass the result to a first output interface of the first group, Wherein each processing element of the first set of processing elements comprises an interface allowing direct access to one or more memory devices of the set of stacked memory devices.

16. The system of claim 15, wherein the set of processing devices further include: At least two processing elements of the second group are connected in a chain topology, wherein the processing elements in the second group use the received data to independently calculate partial results, store the partial results and directly pass the partial results to the next element in the chain topology, wherein the second group includes a second input processing element and a second output processing element, the second input processing element is used to receive data from a second input interface to the second group, and the second output processing element is used to pass the result to a second output interface of the second group.

17. The system of claim 16, wherein the set of processing devices further include: A group of hierarchical buffers are connected in a ring topology, wherein a first at least one hierarchical buffer in the group of hierarchical buffers is connected to the first input interface to provide data to the first input processing element, and a second at least one hierarchical buffer in the group of hierarchical buffers is connected to the second input interface to provide data to the second input processing element.

18. The system of claim 17, wherein at least a third staging buffer in the set of staging buffers is connected to the first output interface to receive data from the first output processing element, and at least a fourth staging buffer in the set of staging buffers is connected to the second output interface to receive data from the second output processing element.

19. The system of claim 18, wherein the set of processing devices further include: A memory interface is coupled to the group of hierarchical buffers and to an external device outside the system, the memory interface being used to perform an operation for the external device to access the group of stacked memory devices.

20. The system of claim 19, wherein the memory interface is to perform operations for the external device to access the set of staging buffers.

21. A system, include: a set of stacked memory devices, each stacked memory device comprising a plurality of memory arrays, the plurality of memory arrays being operable to be accessed via centrally located global input and output circuitry, each memory array of the plurality of memory arrays also being operable to be accessed via a respective array access interface independently of other memory arrays of the plurality of memory arrays; a set of one or more processing devices electrically coupled to and stacked with the set of stacked memory devices, each processing device in the set of one or more processing devices connected to at least one array access interface of the set of stacked memory devices, the set of one or more processing devices comprising: a first group of at least two processing elements connected in a chain topology, wherein the processing elements in the first group independently compute partial results using received data, store the partial results, and pass the partial results directly to the next processing element in the chain topology, Wherein each processing element in the first group includes an interface that allows direct access to one or more memory devices in the group of stacked memory devices.

22. The system of claim 21, wherein the array access interface is connected to a corresponding processing device in the set of one or more processing devices using through silicon vias (TSVs).

23. A system according to claim 22, wherein the first group further includes a first input processing element and a first output processing element, the first input processing element being used to receive data from the global input and output circuit device via a first input interface to the first group, and the first output processing element being used to deliver the result to the global input and output circuit device via a first output interface of the first group.

24. The system of claim 23, wherein the set of one or more processing devices further include: At least two processing elements of a second group are connected in a chain topology, wherein the processing elements in the second group independently calculate partial results using received data, store the partial results and directly pass the partial results to the next element in the chain topology, wherein the second group includes a second input processing element and a second output processing element, the second input processing element is used to receive data from a second input interface to the second group, and the second output processing element is used to pass the result to the global input and output circuit device via a second output interface of the second group.

25. The system of claim 24, wherein the set of processing devices further include: A group of hierarchical buffers are connected in a ring topology, wherein a first at least one hierarchical buffer in the group of hierarchical buffers is connected to the first input interface to provide data to the first input processing element, and a second at least one hierarchical buffer in the group of hierarchical buffers is connected to the second input interface to provide data to the second input processing element.

26. A system according to claim 25, wherein at least a third staging buffer in the set of staging buffers is connected to the first output interface to receive data from the first output processing element, and at least a fourth staging buffer in the set of staging buffers is connected to the second output interface to receive data from the second output processing element.

27. The system of claim 26, wherein the set of processing devices further include: A memory interface is coupled to the group of hierarchical buffers and to an external device outside the system, the memory interface being used to perform an operation for the external device to access the group of stacked memory devices.

28. The system of claim 27, wherein the memory interface is to perform operations for the external device to access the set of staging buffers.

29. A system, wherein the system include: a set of stacked devices, including a set of stacked memory devices and at least one logic device; The stacked memory device comprises: a plurality of memory arrays; a first interface addressable to access all of the plurality of memory arrays on a corresponding memory device; and a plurality of second interfaces, the plurality of second interfaces accessing respective subsets of the plurality of memory arrays of the respective memory device; The logic device comprises: a linked chain of processing elements, wherein the processing elements in the linked chain are to independently compute partial results as a function of received data, store the partial results and pass the partial results directly to the next processing element in the linked chain of processing elements, each of the processing elements being coupled to at least one second interface of the plurality of second interfaces, Wherein each processing element in the linked chain includes an interface allowing direct access to one or more memory devices in the set of stacked memory devices.

30. The system of claim 29, wherein the linked chain of processing elements includes an input processing element to receive data from an input interface to the linked chain of processing elements.

31. The system of claim 30, wherein the linked chain of processing elements comprises an output processing element to deliver a result to an output interface of the linked chain of processing elements.

32. The system of claim 31, wherein the logic device further include: A centrally located region of the logic device includes global input and output circuitry for interfacing the system with external processing systems.

33. The system of claim 32, wherein the logic device further include: A first staging buffer is coupled between the global input and output circuitry and the coupled chain of processing elements to communicate data with at least one of: the input processing element and the output processing element.

34. The system of claim 33, wherein the logic device further include: a plurality of coupled chains of processing elements and a plurality of staging buffers, a respective staging buffer of the plurality of staging buffers coupled between the global input and output circuitry and a respective one of the plurality of coupled chains of processing elements to transfer data with at least one of: a respective input processing element and a respective output processing element of the respective one of the plurality of coupled chains of processing elements.

35. A component, include: A stacked plurality of dynamic random access memory DRAM devices; at least two logic dies, also stacked with the plurality of DRAM devices, a first at least one logic die of the at least two logic dies attached to one of the top and bottom sides of the stacked plurality of DRAM devices, a second at least one logic die of the at least two logic dies attached to an opposite side of the one of the top and bottom sides of the stacked plurality of DRAM devices, wherein a first at least one of the at least two logic dies is a substrate die compatible with a high bandwidth memory component, and a second at least one of the at least two logic dies includes a computing accelerator, and wherein the computing accelerator comprises a linked chain of processing elements, wherein the processing elements in the linked chain independently compute partial results as a function of received data, store the partial results, and pass the partial results directly to the next processing element in the linked chain of processing elements, Each processing element in the linked chain includes an interface that allows direct access to one or more DRAM devices in the plurality of DRAM devices in the stack.

36. The component of claim 35, wherein the first at least one of the at least two logic dies is attached to an active circuit side of the first at least one of the at least two logic dies, the active circuit side facing a passive circuit side of the stacked plurality of DRAM devices.

37. The component of claim 36, wherein the second at least one of the at least two logic dies is attached to a passive circuit side of the second at least one of the at least two logic dies, the passive circuit side facing the passive circuit side of the stacked plurality of DRAM devices.

38. The assembly of claim 35, wherein the assembly comprises a die that redistributes through silicon via (TSV) locations between the stacked plurality of DRAM devices and one of the at least two logic dies.

39. The assembly of claim 35, wherein the assembly comprises a die that redistributes through silicon via (TSV) locations between the stacked plurality of DRAM devices and at least one logic die of the at least two logic dies.

40. The assembly of claim 35, wherein the processing elements in the linked chain are configured as a systolic array.

Citation Information

Patent Citations

  • Semiconductor memory device

    US20190088339A1