Systems, methods, and apparatus for in-memory processing using die-to-die interconnects
By designing interconnection ports and control circuits between multiple dies in the memory system, the problem of inefficient data processing of die-to-die interconnection processing in the prior art is solved, efficient data processing and calculation are achieved, and the scalability of the system is improved.
Patent Information
- Application Number
- CN202411788026.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-08-28
- Filing Date
- 2024-12-06
- Publication Date
- 2025-06-13
AI Technical Summary
Existing memory systems have problems of inefficiency and insufficient resource utilization when processing data in die-to-die interconnects.
By designing a device, the device includes a plurality of dies, interconnected through an interface, and coordinating computing operations and data transmission with control circuits to achieve efficient data processing and calculation.
It improves the data processing efficiency of the memory system, reduces data movement, and increases the memory and computing functions available to the host, thereby improving the scalability of the system.
Smart Images

Figure CN120144513A_ABST
Abstract
Description
[0001] Citation of Related Applications
[0002] This application claims the benefit and priority of U.S. Provisional Patent Application Serial No. 63 / 608,823, filed on Dec. 11, 2023, which is incorporated herein by reference. Technical Field
[0003] This disclosure generally relates to memory systems, and more particularly, to systems, methods, and apparatuses for processing in memories using die-to-die (D2D) interconnects. Background Art
[0004] Interconnects can be used to transfer data between components in a computing system. For example, interconnects can be used to transfer data between a processing unit and one or more peripheral components such as a graphics device, a storage device, a network interface, etc. Die-to-die interconnects can be used to transfer data between integrated circuit dies (which may also be referred to as chips). For example, die-to-die interconnects can be used to transfer data between two integrated circuit dies that may be located in the same package.
[0005] The above information disclosed in this background art section is only for enhancing the understanding of the background art of the principles of the present invention, and thus it may contain information that does not constitute the prior art. Summary of the Invention
[0006] An apparatus may include a first device and a second device. The first device includes at least one first die. The at least one first die includes a first interface, a memory medium configured to store first information received using the first interface, and a computing element configured to perform a computing operation using the first information. The second device includes at least one second die. The at least one second die includes a second interface configured to receive second information and a third interface coupled to the first interface and configured to transmit the first information, wherein the first information may be based on the second information. The computing operation may generate third information, and the at least one first die may include a fourth interface configured to transmit the third information. The memory medium may be a first memory medium, the computing element may be a first computing element, the computing operation may be a first computing operation, and the at least one second die may include a second memory medium configured to store at least a portion of the second information and a second computing element configured to perform a second computing operation using at least a portion of the second information. The second computing operation may generate at least a portion of the first information. The first information may include at least a portion of the second information. The second information may include input data for the computing operation. The second information may include command information for the computing operation. The computing operation may generate third information, the at least one first die may include a fourth interface configured to transmit the third information, and the at least one second die may include a fifth interface configured to receive the third information. The memory medium may be a first memory medium, the computing element may be a first computing element, the computing operation may be a first computing operation for generating the third information, the at least one first die may include a fourth interface configured to transmit the third information, and the apparatus may include a third device. The third device includes at least one third die. The third die includes a fifth interface configured to receive the third information, a second memory medium configured to store at least a portion of the third information, and a second computing element configured to perform a second computing operation using at least a portion of the third information.
[0007] An apparatus may include a die that includes a first interface, a second interface, a third interface, and at least one control circuit. The at least one control circuit is configured to receive first information for a computing operation using the first interface, receive second information for the computing operation using the second interface, control the computing operation, wherein the computing operation may be performed by at least one computing element using the first information and the second information to generate third information, and send the third information using the third interface. The first interface may include a first die interface, the second interface may include a memory interface, and the third interface may include a second die interface. The at least one control circuit may be configured to receive at least a portion of the second information using the first interface and send at least a portion of the second information using the second interface. The at least one control circuit may be configured to receive command information using the first interface and perform the computing operation based on the command information. The at least one control circuit may be configured to receive command information using the first interface and send at least a portion of the command information using the third interface. The at least one control circuit may include a memory controller. The at least one control circuit may be configured to access the computing element using the second interface. The die may include at least a portion of the computing element.
[0008] A method may include receiving information at a first device using a first die interface, wherein the first device may include a first memory medium and a first computing element, storing at least a first portion of the information in the first memory medium, sending at least a second portion of the information from the first device to a second device using a second die interface, wherein the second device may include a second memory medium and a second computing element, storing a second portion of the information in the second memory medium, performing a first computing operation using the first portion of the information and the first computing element, and performing a second computing operation using the second portion of the information and the second computing element. The information may be first information, and the method may further include sending second information from the first device to the second device using the second die interface, wherein the second computing operation may be performed using the second information. The first computing operation may include a first portion of a matrix operation, and the second computing operation may include a second portion of the matrix operation.
[0009] An apparatus may include a die that includes a first interface, a second interface, and at least one control circuit. The at least one control circuit is configured to receive first information for a computing operation using the first interface, generate second information for the computing operation based on the first information, and send the second information and at least a portion of the first information using the second interface. The second information may include command information. At least a portion of the first information may include first input data for the computing operation, and the second information may include second input data for the computing operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings are not necessarily to scale. The accompanying drawings are only intended to facilitate the description of the various embodiments described herein. The accompanying drawings do not depict every aspect of the teachings disclosed herein and do not limit the scope of the claims. To prevent the accompanying drawings from becoming blurred, not all components, connections, etc. may be shown, and not all components may be shown with reference numerals. However, the pattern of the component configuration can be readily apparent from the accompanying drawings. The accompanying drawings, together with the specification, illustrate example embodiments of the present disclosure and are used, together with the specification, to explain the principles of the present disclosure.
[0011] Figure 1 A diagram showing an embodiment of a device according to an example embodiment of the present disclosure, the device including a memory-in-processor device configured to communicate using another device and die interconnects.
[0012] Figure 2 A diagram showing an embodiment of a computing node having a compute node die according to an example embodiment of the present disclosure.
[0013] Figure 3 A diagram showing an embodiment of a gateway die according to an example embodiment of the present disclosure.
[0014] Figure 4 A perspective view showing an embodiment of a memory-in-processor device according to an example embodiment of the present disclosure.
[0015] Figure 5 A diagram showing a first example embodiment of a compute node die according to an example embodiment of the present disclosure.
[0016] Figure 6 A diagram showing an example embodiment of a processing node in memory according to an example embodiment of the present disclosure.
[0017] Figure 7 A diagram showing a second example embodiment of a compute node die according to an example embodiment of the present disclosure.
[0018] Figure 8 A diagram showing an example embodiment of a gateway die according to an example embodiment of the present disclosure.
[0019] Figure 9 A diagram showing an embodiment of a processing chain cluster in memory according to an example embodiment of the present disclosure.
[0020] Figure 10 A diagram showing an example embodiment of a processing chain cluster in memory and an associated method according to an example embodiment of the present disclosure.
[0021] Figure 11 A diagram showing an example embodiment of a pipelined matrix multiplication reduction operation according to an example embodiment of the present disclosure.
[0022] Figure 12 A diagram showing an embodiment of a method for performing in - memory processing chain operations according to an example embodiment of the present disclosure. Detailed description
[0023] A processing unit (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), a tensor processing unit (TPU), a data processing unit (DPU), etc.) can access memory using one or more memory interfaces (such as a double - data - rate (DDR) interface, a high - bandwidth memory (HBM) interface, etc.). For example, a GPU can be implemented as a system - on - a - chip (SoC) that can use an HBM interface to access a memory device (e.g., an HBM device) implemented with stacked memory dies. In some embodiments, a memory device such as an HBM device can include computing functionality at the memory die and / or at another die stacked with or located near the memory die, which can be referred to as in - memory processing (PIM) and / or near - memory processing (PNM) configurations. For convenience, in - memory processing and / or near - memory processing can be referred to individually and / or collectively as in - memory processing or PIM.
[0024] To increase the amount of memory (and / or PIM) available to a GPU, multiple HBM devices can be connected to the GPU SoC using multiple HBM interfaces (which can also be referred to as using multiple memory channels). The memory interfaces can be fabricated at or near the periphery (e.g., the edge) of the SoC and can consume a relatively large area on the SoC. Thus, in some embodiments, and depending on implementation details, the ability to increase the amount of memory and / or PIM available to the GPU (which can be referred to as scaling) may be limited by the area and / or size of the edge of the SoC, which in turn may be limited by factors such as cost, yield, power, performance, etc.
[0025] Some aspects of the present disclosure relate to one or more PIM devices that can communicate (e.g., with a host such as a processing unit or other user) using at least one other PIM device, memory device, gateway device, accelerator device, etc. For example, a first PIM device can be connected to a host using a first communication link. A second PIM device can communicate with the first PIM device using a second communication link. Thus, in some embodiments, and depending on implementation details, the second PIM device can communicate with the host directly and / or indirectly through the first PIM device. For example, the first PIM device can receive first information from the host using the first communication link. The first PIM device can send at least a portion of the first information to the second PIM device using the second communication link.
[0026] Additionally or alternatively, the first PIM device may send second information to the second PIM device using a second communication link. The second information may be generated by the first PIM device, e.g., by modifying the first information, by appending more information to the first information, by performing an operation (e.g., a calculation) that may be based on, e.g., the first information and / or other information (e.g., to generate a result), etc.
[0027] In some embodiments, the first PIM device and / or the second PIM device may communicate with one or more additional PIM devices using, e.g., one or more additional communication links. One or more PIM devices, memory devices, gateway devices, accelerator devices, communication links, etc. may be arranged in any configuration, e.g., in one or more network topologies such as a chain, bus, mesh, tree, ring, star, etc. and / or multiples and / or combinations thereof (e.g., a hybrid combination). In some embodiments, and depending on implementation details, a collection of one or more PIM devices, memory devices, gateway devices, etc. connected to one or more communication links may be referred to as a cluster (e.g., a computing cluster).
[0028] In some embodiments, one or more die-to-die interconnects (which may also be referred to as die interconnects or D2D interconnects) may be utilized to implement one or more communication links. Die interconnects may be used to transfer data between integrated circuit (IC) dies (which may also be referred to as chips or chiplets). Die interconnects may enable multiple dies to be assembled in a package, e.g., as a system-in-package (SIP), a multi-chip module (MCM), etc. Examples of die interconnects may include Universal Chiplet Interconnect Express (UCIe), Advanced Interface Bus (AIB), Bunch of Wires (BOW), etc.
[0029] Some additional aspects of the present disclosure relate to dies that can be used in embodiments, where one or more PIM devices can communicate using at least one other PIM device, memory device, gateway device, accelerator device, etc. As a first example, in some embodiments, a die may include functionality that enables PIM devices, memory devices, accelerator devices, etc. to operate as nodes of a cluster (e.g., compute nodes (which may be implemented as PIM nodes, accelerator nodes, etc.), memory nodes, etc.), a network of PIMs and / or memory devices, etc. or as part thereof. In such an embodiment, the die may include a first interface (e.g., a memory interface for accessing one or more memory devices such as one or more HBM dies, PIM dies, etc.) and one or more additional interfaces (e.g., one or more die interfaces for communicating with one or more other dies). In such an embodiment, the die may be referred to as a node die, compute node die (e.g., if the node includes compute functionality), compute element die (e.g., if the node includes compute functionality), memory node die (e.g., if the node includes memory), PIM node die (e.g., if the node includes PIM functionality), etc. In embodiments where the die can be configured to operate as a node of a chain, the die may be referred to as a chain die, chain element die, and / or PIM chain element (PCE) die (e.g., if the node includes PIM functionality). In some embodiments, the node die may be configured and / or referred to as a base die (e.g., having one or more memory dies such as HBM dies stacked on the node die), logic die, buffer die, and / or compute die (e.g., having compute functionality implemented at the node die). In some embodiments, the node die may be connected to one or more other dies at the node using one or more interposers, substrates (e.g., organic substrates), semiconductor bridges (e.g., silicon bridges), vias such as through-silicon vias (TSVs), etc.
[0030] As a second example, in some embodiments, a die may include functionality that operates as or as part of a gateway device for one or more PIM devices, memory devices, etc. In such embodiments, the die may include a first interface and one or more additional interfaces. The first interface may be implemented, for example, using a communication interface that may be suitable for communication between packages (e.g., between a first device within a package and a second device outside the package), such as Peripheral Component Interconnect Express (PCIe), Compute Express Link (CXL), etc. The first interface may be used, for example, to enable the gateway device to communicate with a host such as a processing unit or other user. The one or more additional interfaces may be implemented, for example, using a die interface that may be suitable for communication between dies within a package (e.g., UCIe, AIB, BOW, etc.). The one or more additional interfaces may be used, for example, to enable the gateway device to communicate with one or more devices such as PIM devices, memory devices, accelerator devices, other gateway devices, etc. In such embodiments, the die may be referred to as a gateway die or a gateway node die. In embodiments where the gateway die may be configured as part of a chain (e.g., as the gateway of a chain), the die may be referred to as a chain gateway die and / or a PIM Chain Gateway (PCG) die (e.g., if the chain may include one or more nodes with PIM functionality).
[0031] In some embodiments, the gateway die may include functionality to implement a computing scheme, where one or more computing operations may be performed by one or more PIM devices (e.g., at one or more computing nodes) connected to the gateway die (e.g., using one or more die interfaces). For example, the gateway die may include gateway logic that is configured to receive (e.g., from a host such as a processing unit or other user via a die interface) input information for one or more computing operations (e.g., input data, one or more models and / or parameters for the models (such as weights, activation functions, etc.), commands for one or more computing operations, etc.). In some embodiments, the gateway logic may generate additional information (e.g., one or more command packets) for one or more computing operations based on the input information. The gateway logic may be configured to send some or all of the input information and / or the additional information it may generate to one or more PIM devices, memory devices, etc. (e.g., to one or more PIM devices and / or memory devices configured as a computing cluster). In some embodiments, at least a portion of the gateway functionality may be included in a computing node die (e.g., a PIM node die), a memory node die, an accelerator node die, a host (such as a processing unit or other user), etc.
[0032] Some additional aspects of the present disclosure relate to computing schemes that can be implemented, for example, using one or more PIM devices, which can communicate using at least one other PIM device, memory device, gateway device, accelerator device, etc. For example, in some embodiments, a gateway device can distribute information about a computing operation (e.g., a model and / or information about the model, such as the weights of the model) to one or more computing nodes. The one or more computing nodes can be connected to the gateway device in a chain or other configuration. In some embodiments, the one or more computing nodes can be connected in a manner that can form a loop (e.g., a ring, a closed chain, etc.) with the gateway device. For some computing operations, the gateway device can send a complete copy of the information for the computing operation to the one or more computing nodes (e.g., send a complete copy to each computing node). For some other computing operations, the gateway device can send a portion of the information to the one or more computing nodes (e.g., send different portions of the information to each computing node). Depending on the connection configuration of the one or more PIM devices, some or all of the information sent to a computing node can be transmitted through one or more other computing nodes and / or memory nodes.
[0033] The one or more computing nodes can use the information located at the respective computing nodes to perform some or all of the computing operations (e.g., each of the computing nodes can perform a portion of the computing operation). For example, different portions of a first operand (e.g., portions of a matrix) can be assigned to different computing nodes, and one or more computing nodes having a portion of the operand can use a second operand (e.g., a scalar, vector, matrix, etc.) that can be assigned to different computing nodes to perform a portion of the computing operation (e.g., a portion of a matrix multiplication).
[0034] Additionally or alternatively, different second operands (or different portions of the second operand) can be assigned to different computing nodes, which can use the second operand or the portion of the second operand to perform a portion of the computing operation (e.g., a portion of a matrix multiplication). In some embodiments, the output (e.g., the result) from a portion of the computing operation performed by one computing node can be sent to another computing node and / or used as an input (e.g., an operand) by another computing node.
[0035] Some additional aspects of the present disclosure relate to structures and / or methods for sending data, commands, etc. to and / or from one or more PIM devices, where the PIM devices can communicate using at least one other PIM device, memory device, gateway device, etc. For example, in some embodiments, a gateway device can send command packets and / or input data for a computing operation to at least one computing node. Depending on the connection configuration of one or more PIM devices, the command packets and / or input data can flow through one or more additional computing nodes. A node that receives the command packets and / or input data can use the command packets and / or input data to perform a computing operation or a part of a computing operation, and send the command packets, input data, the result (or a part thereof) of the computing operation, and / or a completion to another computing node, which can use the command packets, input data, and / or result to perform another computing operation (or a part thereof), and forward the command packets, input data, the result of its operation, and / or another completion to yet another computing node in, for example, a pipeline configuration. Depending on the connection configuration (e.g., if one or more computing nodes are arranged in a loop such as a ring or a closed chain), the computing node can send the result (e.g., the final result) and / or completion (e.g., the final completion) to the gateway device, for example, via a communication link between the computing node and the gateway device. Additionally or alternatively, the result (e.g., the final result) and / or completion (e.g., the final completion) can be sent back to the gateway device through one or more computing nodes through which the command packets and / or input data may have been sent.
[0036] In some embodiments, and depending on implementation details, aspects of the present disclosure can enable one or more computing nodes to perform computations using a relatively large amount of data, computing cycles, etc., while reducing data movement between devices. For example, one or more computing nodes implemented with HBM PIM functionality can utilize the relatively large internal memory bandwidth and / or computing functionality within the computing node to perform a part of a computing operation (e.g., a part of matrix multiplication and / or reduction) involving a relatively large amount of data, computing cycles, etc. However, the computing node can send a relatively small amount of information, such as one or more command packets, completions, relatively small operands (e.g., embedded vectors, activation vectors, computing results, etc.), to another device (e.g., a computing node, a gateway device, etc.). Additionally, one or more computing nodes, gateway devices, etc. can use an interface (e.g., a die interface) that can consume a relatively small amount of die area, edge length, etc. of a host (such as a GPU SoC). Thus, depending on implementation details, more memory and / or computing functionality (e.g., a PIM computing cluster) can be connected to the host, thereby increasing the memory, PIM functionality, etc. available to the host. Depending on implementation details, this can reduce latency, power consumption, cost, etc., and thus can improve scalability.
[0037] The present disclosure encompasses many aspects related to interconnection, in-memory processing solutions, etc. The aspects disclosed herein can have independent utilities and can be embodied separately, and not every embodiment can utilize every aspect. Additionally, these aspects can also be embodied in various combinations, and some of these combinations can amplify some of the benefits of the individual aspects in a synergistic manner.
[0038] For illustrative purposes, some embodiments can be described in the context of some example implementation details, such as one or more PIM devices that can communicate using at least one other PIM device, memory device, gateway device, etc. However, the aspects disclosed herein are not limited to these or any other implementation details.
[0039] Figure 1 A diagram showing an embodiment of a device according to an example embodiment of the present disclosure, the device including a PIM device configured to communicate using another device and die interconnect. Figure 1 The embodiment shown in may include a first device (which can be implemented as a PIM device in this embodiment) 102 and a second device 104. The PIM device 102 may include one or more dies 106 that can implement a first interface (which can be implemented as a first die interface in this embodiment) 108, a memory medium 110 configured to store first information 112 received using the first die interface 108, and a computing element 114 configured to perform a computing operation using the first information 112. In some embodiments, the PIM device 102 can be used to implement a computing node (e.g., a PIM node) as described herein.
[0040] The memory medium 110 can be implemented using any type of volatile and / or non-volatile memory medium, including dynamic random access memory (DRAM), static random access memory (SRAM), flash memory including non-and (NAND) flash, persistent memory such as cross-grid non-volatile memory, memory with bulk resistance change, phase change memory (PCM), etc. or any combination thereof. Any type of memory interface (e.g., any generation of DDR interface (e.g., DDR, DDR2, DDR3, DDR4, DDR5, etc.), any generation of HBM interface (e.g., HBM, HBM2, HBM3, HBM4, etc.), open memory interface (OMI), etc.) can be used to access the memory medium 110.
[0041] The first die interface 108 can be implemented using any type of die interface, including UCIe, AIB, BOW, open high bandwidth interface (OpenHBI), etc.
[0042] Examples of computing elements that can be used to implement computing element 114 can include complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), computing circuits including combinational logic, sequential logic, timers, counters, registers, state machines, etc., embedded processors, microcontrollers, CPUs such as complex instruction set computer (CISC) processors (e.g., x86 processors) and / or reduced instruction set computer (RISC) processors (such as ARM processors), GPUs, NPUs, TPUs, DPUs, etc., which can execute instructions stored in any type of memory and / or implement any type of execution environment, such as containers, virtual machines, operating systems (such as Linux), extended Berkeley packet filters (eBPF) environments, etc., or combinations thereof.
[0043] Computing element 114 can implement any type of computing function, including for example matrix multiplication (e.g., general matrix multiplication (GEMM)), vector multiplication (e.g., general matrix vector (GEMV)), math engines (such as SoftMax, Gaussian error linear unit (GELU), rectified linear unit (ReLU), sigmoid, etc.) and / or any other function.
[0044] The first die interface 108, memory medium 110, and / or computing element 114 can be implemented on any number of dies 106 in any combination. For example, in some embodiments, each of the first die interface 108, memory medium 110, and computing element 114 can be implemented on a separate die. In other embodiments, all of the first die interface 108, memory medium 110, and computing element 114 can be implemented on one die. In another example embodiment, the components can be arranged in a stacked configuration, where the first die interface 108 is fabricated on a base die, the computing element 114 is fabricated as part of a PIM die (e.g., a die including memory such as DRAM and one or more computing elements fabricated on the same die, which can be referred to as a PIM-DRAM die) stacked on the base die, and the memory medium 110 is fabricated on a memory die (e.g., a DRAM die) stacked on the PIM die.
[0045] The second die interface 118 can be implemented with any type of die interface that can be compatible with the first die interface 108, including UCIe, AIB, BOW, OpenHBI, etc. However, in other embodiments, the first interface 108 and the second interface 118 can be implemented with other types of communication interfaces (e.g., package interfaces), such as PCIe, CXL, etc.
[0046] The second device 104 can be implemented with one or more of a PIM device, a memory device, a gateway device, an accelerator device, etc. The second device 104 may include one or more dies 116 having a second interface (which can be implemented with a second die interface in this embodiment) 118 and a third interface 120. The second die interface 118 can be connected to the first die interface 108 at the PIM device 102 using a first die interconnect having a first communication link 122. (In some embodiments, and depending on implementation details and / or context, the first die interface 108, the second die interface 118, and / or the first communication link 122 may be collectively and / or individually referred to as the first die interconnect).
[0047] The third interface 120 can be connected to another device, such as a PIM device, a memory device, a gateway device, an accelerator device, etc., using a second communication link 124. Additionally or alternatively, the third interface 120 can be connected to a host, such as a processing unit or other user, using the second communication link 124. The second device 104 can be configured to receive second information 111 using the third interface 120 and send first information 112 to the PIM device 102 using the second die interface 118. The first information 112 can be based on, for example, the second information 111.
[0048] In embodiments where the third interface 120 is connected to another device (such as a PIM device, a memory device, a gateway device, an accelerator device, etc.) or a host (such as a processing unit or other user that can be implemented with a die, an SoC, etc.), the second communication link 124 can be implemented, for example, with a die interface (such as UCIe, AIB, BOW, etc.). In embodiments where the third interface 120 is connected to a host (such as a processing unit) or other users in a different package (e.g., a die in another package, another assembly (such as a circuit board, card, chassis, rack, server, data center, etc.)), the second communication link 124 can be implemented, for example, using a communication interface (e.g., a package interface) (such as PCIe, CXL, etc.).
[0049] The first information 112 can include command information for a computing operation. For example, in embodiments where the second device 104 is at least partially implemented with a gateway device (e.g., a gateway node), the second device 104 can generate at least a portion of the first information 112 as a command packet based on commands received from the host in the second information 111. Additionally or alternatively, in embodiments where the second device 104 is at least partially implemented with another computing node (e.g., a PIM node), the first information 112 can include a command packet received by the second device 104 from a gateway node or another computing node.
[0050] Additionally or alternatively, the first information 112 may include input data (e.g., operands) for a computing operation. For example, in an embodiment where the second device 104 is at least partially implemented with a gateway device (e.g., a gateway node), the first information 112 may include a matrix, a vector, a scalar, or a portion thereof that the second device 104 receives from a host. As another example, in an embodiment where the second device 104 is at least partially implemented with another computing node (e.g., a PIM node), the first information 112 may include the result of a computing operation performed by the other computing node.
[0051] The host may refer to any user and may be implemented with any hardware and / or software component or combination of components, including one or more of the following: a server, a storage node, a computing node, a processing unit (e.g., a CPU, a GPU, an NPU, a TPU, a DPU, etc.), a workstation, a personal computer, a tablet computer, a smart phone, an operating system, an application (e.g., a software application), a driver, a process, a service, a virtual machine (VM), a VM manager, etc., or multiples and / or combinations thereof.
[0052] Although Figure 1 The illustrated embodiments are not limited to any particular application, configuration, etc., but in some embodiments, it may be used to implement a computing cluster, where the second device 104 may be configured as a gateway (e.g., a gateway node) and connected to a host (e.g., a GPU, an NPU, or other processing unit fabricated on a die such as an SoC) using the second communication link 124. In this configuration, the second device 104 may be configured to (e.g., using logic such as one or more controllers) receive the second information 111 from the host and send at least a portion of the second information 111 as the first information 112 to the PIM device 102, which may store the first information 112 in the first memory medium 110. The second information 112 may include one or more operands, such as all or part of a matrix, all or part of a vector, etc. The second device 104 may also be configured to receive a first command (e.g., an execution command) from the host and send a corresponding second command (e.g., using a command packet) to the PIM device 102, which may perform a computing operation based on the second command.
[0053] For example, the second command may cause the PIM device 102 to perform all or part of a matrix multiplication operation using a matrix or a portion thereof stored in the first memory medium 110 as a first operand and a matrix or vector or a portion thereof received from the second device 104 as a second operand. In some embodiments, the PIM device 102 may send the result of the computational operation back to the host via the second device 104. In some other embodiments, one or more die 106 at the PIM device 102 may include another die interface, and the PIM device 102 may send the result to another device (e.g., another PIM device) arranged in a chain, tree, ring, etc. with the PIM device 102 via the another die interface. Additionally or alternatively, the PIM device 102 may use an additional die interface to send the result back to the second device 104 (e.g., via another die interface at the second device 104) and / or the host (e.g., via another interface at the host), etc. In some embodiments, in addition to being configured to operate as a gateway, the second device 104 may also include PIM functionality and may be configured to perform all or part of a computational operation in a manner similar to the PIM device 102.
[0054] In some embodiments, and depending on implementation details, the memory medium 110 and the computing element 114 may communicate using a memory interface (e.g., an HBM interface) that is internal to the PIM device 102 and may have a relatively large bandwidth, footprint, die edge length, etc. on one or more die 106, while the PIM device 102 may communicate with the second device 104 and / or the host using one or more die interfaces 108, 118, 120, etc., which may occupy a relatively small amount of die area, die edge length, etc. Thus, the PIM device 102 may utilize a relatively large internal memory bandwidth and / or computing capabilities (e.g., as a computing node) to perform part of a computational operation (e.g., part of matrix multiplication and / or reduction) involving a relatively large amount of data, computational cycles, etc. However, the PIM device 102 may use one or more die interfaces 108, 118, 120, etc. to send and / or receive a relatively small amount of information to / from another device (e.g., another PIM device (e.g., configured as another computing node), a gateway device, an accelerator device, a host, etc.), such as one or more command packets, completions, relatively small operands (e.g., embedded vectors, activation vectors, computational results, etc.), and one or more die interfaces 108, 118, 120, etc. may use interfaces that may consume a relatively small amount of die area, edge length, etc. of the host (e.g., GPU SoC). Thus, depending on implementation details, Figure 1The embodiments shown can enable more memory and / or computing functions (e.g., PIM computing clusters) to be connected to the host, thereby increasing the memory, PIM functions, etc. available to the host. Depending on the implementation details, this can reduce latency, power consumption, cost, etc., and thus improve scalability.
[0055] Figure 2 A diagram illustrating an embodiment of a computing node having a computing node die in accordance with an example embodiment of the present disclosure. The computing node 202 may include a computing node die 204 and / or one or more computing elements 206. The computing node die 204 may include a first interface 208, a second interface 210, a third interface 212, and / or control logic 214. In some embodiments, the first interface 208 and / or the third interface 212 may be implemented, for example, with a die interface, and the second interface 210 may be implemented with a memory interface. The one or more computing elements 206 may be located anywhere at or near the computing node 202, such as on one or more dies (e.g., one or more memory dies) connected to the second interface 210, on a separate die, on the computing node die 204, or any combination thereof.
[0056] The control logic 214 may control one or more operations of the computing node die 204 and / or the computing node 202, including, for example, one or more of the following operations. The computing node 202 and / or the computing node die 204 may respectively use the first interface 208 and the second interface 210 to receive first information 216 for a computing operation and second information 218 for a computing operation. The first information 216 may be received, for example, as input data from a host (directly or through a gateway), from another computing node and / or accelerator (e.g., as a result from another computing operation), from a memory node, etc. The second information 218 may be received, for example, from a memory medium (e.g., an HBM die) that may be connected to the second interface 210 (e.g., an HBM interface). The memory medium may be located anywhere at or near the computing node 202, such as on one or more dies connected to the second interface 210, on the computing node die 204, or any combination thereof. For example, in some embodiments, the computing node die 204 may be configured as a base die (e.g., a buffer die, a logic die, etc.), and the memory and / or one or more computing elements 206 may be stacked on the computing node die 204. As another example, in some embodiments, the computing node die 204 may be located at a position on an interposer, a substrate, etc., and the memory and / or one or more computing elements 206 may be located at one or more additional positions on the interposer, the substrate, etc.
[0057] One or more computing elements 206 may perform a computing operation using the first information 216 and the second information 218. In some embodiments, the computing operation may generate third information 220 (e.g., the result of the computing operation), and the compute node die 204 may send the third information 220 to, for example, another node such as a compute node, a memory node, an accelerator node, a gateway node, a host, etc. using the third interface 212. One or more computing elements 206 and any computing elements disclosed herein may be implemented, for example, using one or more of the computing elements 114 described above with respect to Figure 1 described.
[0058] Figure 3 FIG. showing an embodiment of a gateway die according to an example embodiment of the present disclosure. Figure 3 The gateway die 302 shown in may include a first interface 304, a second interface 306, and / or control logic 308. The first interface 304 may be implemented, for example, using a die interface that may connect the gateway die 302 to a compute node die and / or a compute node located in a common package with the gateway die 302 (e.g., on a common interposer, substrate, etc.). The second interface 306 may be implemented using a communication interface that may be suitable for communicating with a host such as a processor (e.g., GPU, NPU, etc.) on a SoC.
[0059] For example, if the host SoC is in a different package from the gateway die 302, the second interface 306 may be implemented using an interconnect interface (such as PCIe, CXL, cache coherent interconnect for accelerators (CCIX), UCIe configured with a retimer for out-of-package communication, etc.) and / or a network interface (such as Ethernet, Transmission Control Protocol / Internet Protocol (TCP / IP), Remote Direct Memory Access (RDMA), RDMA over Converged Ethernet (RoCE), Fibre Channel, InfiniBand (IB), iWARP, NVMe over Fabrics (NVMe oF), etc.) or any combination thereof.
[0060] As another example, if the host SoC is in the same package as the gateway die 302 (and / or one or more dies that may be connected to the first interface 304), the second interface 306 may be implemented using a die interface (such as UCIe, AIB, BOW, OpenHBI, etc.).
[0061] Control logic 308 may control one or more operations of gateway die 302, as described below, for example. Gateway die 302 may receive first information 310 for a compute operation of a PIM device using first interface 304. Gateway die 302 may generate second information 312 for the compute operation based on the first information. Gateway die 302 may send the second information 312 to, for example, a compute node die and / or a compute node using second interface 306.
[0062] Although Figure 3 the gateway die 302 shown in is not limited to any particular application, configuration, etc., in some embodiments, the first information 310 may include one or more operands (e.g., matrices, vectors, scalars, and / or portions thereof), commands, parameters, etc. for one or more compute operations (e.g., matrix multiplication). The control logic 308 may use the first information 310 to generate the second information 312, which may be sent to, for example, one or more compute nodes (e.g., using a compute node die) to distribute one or more operands or portions thereof to one or more compute dies and / or compute nodes. In some embodiments, the second information 312 may further include one or more commands (e.g., using one or more command packets), activations, etc. to cause one or more compute nodes to perform one or more compute operations or portions thereof.
[0063] In some embodiments, the gateway die 302 may include one or more additional interfaces, such as a third interface that may be used to communicate with one or more compute nodes or groups thereof. For example, in some embodiments where one or more compute nodes may be arranged in a configuration where information may flow in a loop (e.g., a chain, a ring, etc.), the gateway die 302 may receive third information (e.g., results, completions, etc.) from one or more compute operations performed by one or more compute nodes. As another example, in some embodiments, the gateway die 302 may use the first interface 304 and the third interface to send the second information to one or more compute nodes in two different groups connected to two different interfaces (e.g., in a configuration such as a tree, a star, etc.). In some embodiments, the gateway die 302 may receive the third information (e.g., one or more results, completions, etc.) from one or more compute nodes using the same interface through which it sends the second information.
[0064] In some embodiments of the die, node, etc. disclosed herein, the interface (and / or associated interconnect) can be implemented with a first type of interface (e.g., die-to-die interface, as described above, which can also be referred to as a die interface), which depending on implementation details and / or context, can generally be characterized as more suitable for communication between devices (e.g., dies) within the same package, or with a second type of interface, which depending on implementation details and / or context, can generally be characterized as more suitable for communication between devices in different packages, e.g., between a die within a package and a die or other device outside the package (e.g., a die in another package, another assembly such as a circuit board, card, chassis, rack, server, data center, etc.). In some embodiments, and depending on implementation details and / or context, such an interface (and / or corresponding interconnect) can be referred to as an off-package interface, cross-package interface, package-to-package interface, SiP-to-SiP interface, package interface, etc. (and / or corresponding interconnect).
[0065] In some embodiments, and depending on implementation details and / or context, the package interface and / or associated interconnect can be implemented with one or more features, such as: serial data channels (e.g., using serializer and / or deserializer (serdes) at one or more ends of the interconnect channel); differential data signaling (e.g., two conductors per interconnect data channel); and / or embedded clock (e.g., using serializer and / or transmit driver circuitry that can embed clock information in the serial data signal and clock and data recovery (CDR) circuitry that can implement a clock recovery scheme).
[0066] In some embodiments, and depending on implementation details and / or context, the die interface and / or associated interconnect can be implemented with one or more features, such as: parallel data channels; single-ended data signaling (e.g., one conductor per interconnect data channel referenced to a supply potential (e.g., ground)); and / or a clock forwarding scheme (e.g., source synchronous clock).
[0067] In some embodiments of the die, node, etc. disclosed herein, the interface can be implemented with a first type of interface, which depending on implementation details and / or context, can generally be characterized as more suitable for communication between a compute node die and one or more memory dies (e.g., DRAM die, HBM die, etc.), and which can be referred to as a memory interface, or with a second type of interface, which depending on implementation details and / or context, can generally be characterized as more suitable for communication between a compute node die and another compute node die, gateway die, etc., which can be referred to as a die-to-die interface, as described above, and the die-to-die interface can also be referred to as a die interface.
[0068] In some embodiments, and depending on implementation details and / or context, a memory interface and / or associated interconnect may be implemented with one or more features, such as: a relatively large number of large data channels (e.g., 512 channels with a DDR interface, or 1,000 or more (e.g., 1024) or 2,000 or more (e.g., 2048) channels with an HBM interface); relatively low-speed data channels, e.g., suitable for (e.g., directly) operating with DRAM dies; a relatively large number of independent channels (e.g., for a total data bus width of 1024 bits in HBM, two 128-bit channels per die and eight channels, or for a total bus width of 512 bits in Graphics DDR (GDDR), 32 bits per channel and 16 channels); and / or a relatively large die area and / or edge length footprint per interface.
[0069] In some embodiments, and depending on implementation details and / or context, a die interface and / or associated interconnect may be implemented with one or more features, such as: a relatively low number of channels per interface (e.g., a link width of 16, 32, 64, 128, or 256 data channels in UCIe); one or more error correction features, such as error detection, correction, and / or retry at the physical and / or link layer; and / or a relatively small die area and / or edge length footprint per interface.
[0070] For illustrative purposes, some example embodiments may be described below in the context of some example implementation details, such as one or more computing nodes, computing node dies, gateway nodes, gateway dies, etc., which are configured to use HBM-PIM devices (e.g., HBM-PIM devices with stacked memory dies and / or PIM dies), computing clusters configured with chained dies, specific interfaces (such as HBM interfaces and / or UCIe interfaces), etc. However, the principles disclosed herein are not limited to these or any other implementation details.
[0071] Figure 4 A perspective view showing an embodiment of a stacked PIM device according to an example embodiment of the present invention. Figure 4The embodiments shown may be used to implement, for example, any of the PIM devices (e.g., PIM nodes) disclosed herein. The stacked PIM device 402 may be implemented, for example, in a three-dimensional (3D) package and may include one or more memory dies 404 (which may be implemented, for example, with HBM DRAM dies) stacked above a buffer die 408. The stacked PIM device 402 may also include one or more PIM dies 406 (which may be implemented, for example, with PIM-DRAM dies) stacked above the buffer die 408. The PIM die 406 may be implemented, for example, with a DRAM die having one or more computing elements fabricated on the DRAM die. One or more of the memory die 404, the PIM die 406, and / or the buffer die 408 may be connected using TSVs 405.
[0072] In Figure 4 the embodiments shown, one or more PIM dies 406 may be stacked between one or more memory dies 404 and the buffer die 408. However, in other embodiments, one or more PIM dies 406 may be interleaved with one or more memory dies 404. Additionally or alternatively, one or more of the memory die 404, the PIM die 406, and / or the buffer die 408 may be stacked in different combinations, placed individually on an interposer, substrate, etc., any one or all of which may be located, for example, in a package and connected using one or more die interconnects.
[0073] Figure 5 FIG. shows a first exemplary embodiment of a compute node die in accordance with an exemplary embodiment of the present disclosure. Figure 5 The embodiments shown may be used to implement, for example, any of the compute nodes, compute devices, compute node dies, etc., described herein, or may be implemented using any of the compute nodes, compute devices, compute node dies, etc., described herein.
[0074] Figure 5The illustrated compute node die 502 may include one or more D2D interfaces 504, one or more compute elements 506, one or more memory controllers 510, one or more memory media 512, one or more platforms 514, one or more memory access engines of a direct memory access (DMA) engine 508, and one or more buses, interconnects, etc. 509 that may enable communication between any components located at the compute node die 502. The one or more D2D interfaces 504 may be implemented, for example, with UCIe, AIB, BOW, etc. The one or more compute elements 506 may be implemented, for example, with one or more of a CPU, GPU, NPU, etc. The one or more DMA engines 508 may enable one or more compute elements 506 and / or components external to the compute node die 502 to access any memory located at a compute node implemented with the compute node die 502.
[0075] The one or more platforms 514 may be implemented, for example, with a TSV platform or any other type of platform that may enable one or more memory dies, PIM dies, compute dies, etc. to be stacked on or otherwise connected to the compute node die 502. The one or more platforms 514 may include one or more memory layers 511, such as one or more HBM PHY layers, which may enable one or more DRAM dies to be stacked on the platform 514 and connected using one or more TSVs. The one or more memory controllers 510 may be implemented with any type of memory controller suitable for the type of memory connected to the one or more platforms 514. For example, the one or more memory controllers 510 may be implemented with one or more high bandwidth memory (HBM) controllers (HBMCs) to control one or more HBM dies connected to the one or more platforms 514. The one or more memory media 512 may be implemented, for example, with SRAM to operate as a buffer for memory operations, compute operations, etc. performed by the compute node die 502.
[0076] Although Figure 5The illustrated embodiments are not limited to any particular application, but in some example implementations, it can be used to implement a compute node that can be arranged in a chain configuration with one or more other compute nodes, gateway nodes, memory nodes, accelerator nodes, etc., and thus, it can be referred to as a PIM chain element die, which can also be referred to as a PCE die. In such an embodiment, the compute node die 502 can be connected (e.g., arranged in a chain, tree, star, ring, etc.) with one or more other compute node dies 502, one or more of which can be configured as PCE dies. Additionally or alternatively, the compute node die 502 is connected directly or through another compute node to one or more gateway dies (e.g., as part of a gateway node). In some embodiments, the gateway die can manage one or more compute node dies 502. In such an embodiment, the gateway die can be referred to as a PIM chain gateway die, and as described above, the PIM chain gateway die can also be referred to as a PCG die.
[0077] Figure 6 A diagram illustrating an example embodiment of a PIM node according to an example embodiment of the present disclosure. Figure 6 The embodiments shown herein can be used to implement, for example, any of the compute nodes, compute devices, compute node dies, etc. described herein, or can be implemented using any of the compute nodes, compute devices, compute node dies, etc. described herein. In some embodiments, and depending on the implementation details, the Figure 6 PIM node 602 shown herein can be used in place of the buffer die in an HBM-PIM device, for example, to implement a compute node that can communicate using another PIM device, memory device, gateway device, accelerator device, etc., for a compute system according to an example embodiment of the present disclosure.
[0078] Figure 6 The PIM node 602 shown herein can include a compute node die 603 having one or more D2D interfaces 606 and / or one or more HBM-PIM devices 604 stacked thereon. One or more of the D2D interfaces 606 can be implemented using, for example, UCIe, AIB, BOW, etc. One or more of the HBM-PIM devices 604 can be implemented using one or more memory dies and / or PIM dies (e.g., one or more of the memory dies 404 and / or PIM dies 406 shown with respect to Figure 4 One or more of the D2D interfaces 606 can be used to connect the PIM node 602 to one or more other PIM nodes, gateway nodes, memory nodes, accelerator nodes, etc. in any configuration such as a chain, tree, star, ring, etc.
[0079] Figure 7A diagram showing a second exemplary embodiment of a computing node die in accordance with an exemplary embodiment of the present disclosure. Figure 7 The embodiments shown in can be used to implement, for example, any one of the computing nodes, computing devices, computing node dies, etc. described herein, or can be implemented using any one of the computing nodes, computing devices, computing node dies, etc. described herein. Figure 7 The computing node die 702 shown can include one or more D2D interfaces 704, one or more memory controllers 707, one or more DMA engines 710, one or more computing elements 706, and / or one or more pads 708 (which can implement one or more input and / or output (I / O or IO) connections for one or more stacked memory dies, PIM dies, etc., e.g., using one or more TSVs). One or more D2D interfaces 704 can be implemented using, for example, UCIe, AIB, BOW, etc. One or more DMA engines 710 enable one or more processing elements and / or components external to the computing node die 702 to access any memory located at the computing node implemented with the computing node die 702.
[0080] In some embodiments, one or more of the D2D interfaces 704 can implement one or more inbound connections 712, e.g., to receive input data for one or more computing operations to be performed at the computing node where the computing node die 702 may be located, such as one or more matrices, vectors, scalars, or portions thereof, commands, command groups, etc. In some embodiments, one or more of the D2D interfaces 704 can implement one or more outbound connections 714, e.g., to send data such as one or more matrices, vectors, scalars, or portions thereof, commands, command groups, etc., for one or more computing operations at another computing node and / or computing node die. In some embodiments, the DMA engines 710 enable one or more components of the computing node to access the memory of one or more computing elements at the computing node (e.g., independently). In some embodiments, one or more of the memory controllers 707 can include an HBM controller, which can be tuned or optimized for power, latency, and / or bandwidth.
[0081] In some embodiments, one or more computing node dies 702 can be connected to one or more other PIM nodes, gateway nodes, memory nodes, accelerator nodes, etc. in any configuration (such as a chain, tree, star, ring, etc.), which can enable scaling of resources (e.g., memory, PIM, etc.) available to a host (such as a processor or other user that may be connected to the computing node die 702) depending on implementation details.
[0082] In some embodiments, the compute node die 702 may include one or more compute elements 706, which may include one or more compute resources in addition to or instead of one or more compute resources that may be implemented using one or more pads 708 (e.g., compute resources in a stack of one or more PIM dies). For example, the one or more compute elements 706 may implement any type of computing function as described above and may include, for example, matrix multiplication (e.g., GEMM), vector multiplication (e.g., GEMV), and mathematical engines such as SoftMax, GELU, ReLU, sigmoid, etc. In embodiments having one or more compute resources implemented using one or more pads 708 (e.g., one or more compute resources in a stack of one or more PIM dies), the one or more compute elements 706 may increase computing power, which may be beneficial, for example, if the one or more compute resources implemented using one or more pads 708 do not have sufficient computing power for relatively complex tasks such as running kernels, relatively large language models, etc. Additionally or alternatively, the one or more compute elements 706 may increase computing power to accommodate one or more functions, applications, etc. that may be enabled using upgrades.
[0083] Figure 8 A diagram showing an example embodiment of a gateway die in accordance with an example embodiment of the present disclosure. Figure 8 The embodiments shown therein may be used, for example, to implement any one of the gateway nodes, gateway devices, gateway die, etc. described herein, or may be implemented using any one of the gateway nodes, gateway devices, gateway die, etc. described herein. Figure 8 The illustrated gateway die 802 may include one or more (e.g., multiple) interconnect physical interfaces 804 and / or 822, protocol logic 806, management CPU 808, D2D interface 810, memory medium 812, DMA engine 814, receive command completion logic 816, protocol logic 818, and / or command packet generation logic 820. In some embodiments, the gateway die 802 may include functionality to implement PIM, such as one or more PIM devices (e.g., a stack of dies including memory and / or compute elements), while in other embodiments, the gateway die 802 may provide only gateway functionality.
[0084] The interconnect physical interfaces 804 and / or 822 can be implemented using, for example, communication interfaces (e.g., package interfaces) such as PCIe, CXL, etc. The protocol logic 806 can be used to implement, for example, a PCIe endpoint (EP), a CXL node, etc. The protocol logic 818 can be used to implement, for example, a PCIe root port (RP), a CXL node, etc. The memory medium 812 can be implemented using, for example, SRAM and operate as a buffer for memory operations, computing operations, etc. performed by the gateway die 802. The management CPU 808 can be used to manage one or more nodes, devices, dies, etc. connected to the gateway die 802 using one or more D2D interfaces 810, for example, in any configuration such as a chain, tree, star, ring, etc.
[0085] The command packet generation logic 820 can be used to generate, for example, one or more command packets to control one or more nodes, devices, dies, etc. connected to the gateway die 802 based on commands received from a host such as a processor or other user using one of the interconnect physical interfaces 804 and / or 822. The receive command completion logic 816 can be used to process one or more completions received from one or more nodes, devices, dies, etc. connected to the gateway die 802 using one or more of the D2D interfaces 810. For example, in some embodiments, the gateway die 802 can generate one or more command packets using the command packet generation logic 820 and send the packets (e.g., using the DMA engine 814 to send the packets to one or more nodes, devices, dies, etc.) using one or more of the D2D interfaces 810, for example, to Figure 9 the compute node die 912 shown in. One or more command packets can be processed (e.g., along the PCE chain as shown in Figure 9 and returned (e.g., as command packets, completions, etc.) to the receive command completion logic 816 using the D2D interfaces 810 and stored in the SRAM 812.
[0086] Although the principles disclosed herein are not limited to any particular application, in some embodiments, one or more aspects of the present disclosure can be applied to artificial intelligence (AI) and / or machine learning (ML) scenarios, such as large language models (LLMs). Some LLM inference systems can be relatively expensive, consume a relatively large amount of power, use a relatively large amount of memory and / or memory bandwidth, etc. to accommodate AI operations. Some AI and / or ML scenarios (including LLM scenarios) can operate using processors such as NPUs, GPUs, etc., which can be connected to one or more (e.g., several) HBM-PIM devices using, for example, HBM interfaces that may consume a relatively large amount of area and / or edge length.
[0087] In some embodiments, according to example embodiments of the present disclosure, one or more aspects of the present disclosure can be used to implement one or more computing clusters (e.g., AI and / or ML computing clusters) using one or more computing nodes, gateway nodes, etc. For example, some embodiments of a computing cluster can utilize one or more PIM devices, and the one or more PIM devices can communicate with one or more additional devices such as PIM devices, gateway devices, memory devices, accelerator devices, etc. Additionally, aspects of the present disclosure can combine HBM and PIM technologies, which depending on implementation details, can enhance both memory bandwidth and computing efficiency by directly integrating processing units within the HBM module. Depending on implementation details, this integration can allow for parallel execution of computations on data stored in memory, which in turn can reduce or minimize data movement and / or provide performance benefits for memory-bound applications.
[0088] For example, in some embodiments, a PIM chain cluster can include a PCG die and one or more (e.g., N) PCE dies (where N can be any number, e.g., 8 or 16, etc.). Additionally, in some embodiments, multiple HBM-PIM modules can be stacked on one or more PCE dies. Depending on implementation details, HBM devices including in-memory processing (PIM) linked together can achieve greater internal bandwidth. Additionally or alternatively, some embodiments can implement computing resources (e.g., relatively lightweight computing resources) at one or more PIM nodes. In some embodiments, HBM-PIM devices (or multiple HBM-PIM devices connected in a network such as a chain, tree, star, ring, etc.) can be connected to a GPU, NPU, etc. SoC using a relatively lightweight (e.g., having a relatively low die area, side length, cost, power consumption, etc.) die interface instead of a relatively heavy (e.g., having a relatively high die area, side length, cost, power consumption, etc.) HBM interface. However, in some embodiments, one or more computing nodes can use one or more HBM interfaces within the node to utilize the relatively high performance of HBM-PIM devices while reducing the amount of data transferred between nodes.
[0089] Figure 9 A diagram showing an embodiment of a PIM chain cluster according to an example embodiment of the present disclosure. Figure 9 The embodiments shown in can be implemented using or used to implement any one of the embodiments of the computing nodes, gateway nodes, memory nodes, accelerator nodes, etc. disclosed herein. For illustrative purposes, the embodiments shown in can be described in the context of an LLM scheme, Figure 9 but the principles can be applied to any type of computing scheme. Additionally, Figure 9The embodiments shown may be described in the context of some specific implementation details, such as computing nodes connected in a chain configuration, PCIe interconnections, etc., but other types of interconnections and / or node configurations, such as star, tree, etc., may be used.
[0090] Figure 9 The PIM chain cluster 904 shown may include a first PIM chain 916 that may be encapsulated in a first SiP. The PIM chain cluster 904 may include a second PIM chain 918 that may be encapsulated in a second SiP. The first PIM chain 916 may include one or more PCE dies 908, 912, …, 914 and a PCG die 910 arranged in a chain using D2D interconnections. The second PIM chain 918 may include one or more PCE dies arranged in a similar chain configuration. Any number of PCE dies and / or gateway dies may be used in each chain. The first PIM chain 916 may communicate with the host 906, for example, using a first SiP-to-SiP interconnect (e.g., using a PCIe interface at the PCG die 910). The second PIM chain 918 may communicate with the first PIM chain 916, for example, using a second SiP-to-SiP interconnect (e.g., using a second PCIe interface at the PCG die 910). In some embodiments, one or more additional PIM chains may communicate with the second PIM chain 918 using one or more additional SiP-to-SiP interconnections. In some embodiments, one or more of the PCE dies may communicate using D2D interfaces such as UCIe, AIB, BOW, etc.
[0091] In some embodiments, the host 906 may generate and / or provide one or more data sets (e.g., AI and / or ML data sets) for training, inference, etc. to one or more of the PCG dies in the PCG die 910. In some embodiments, the host 906 may communicate with one or more of the PCG dies, one or more of the PCE dies, etc. in the PIM chain 916 and / or 918 in the PIM chain cluster 904.
[0092] In some embodiments, the host 906 may send LLM weight data to the PCG die 910 via a PCIe EP or CXL connection, and the PCG die 910 may distribute the weight data to one or more PCEs in a chain corresponding to PIM chains 916 and / or 918. Additionally or alternatively, the host 906 may send one or more matrix multiplication commands to one or more PCG dies 910, and the PCG dies 910 may use the one or more matrix multiplication commands to generate one or more command packets. One or more PCG dies 910 may send one or more command packets, e.g., having one or more activations (e.g., operands), to the first PCE die 912. The first PCE die 912 may perform one or more distributed computing operations based on the command packets and forward the result(s) of the (multiple) computing operations to the next PCE die 914. In some embodiments, the information sent between one or more PCG dies 910 and one or more PCE dies 908, 912, …, 914 may have the following format: {command_packet, activation v(12288), self-reduction result v(12288), completion_packet}. The information may be sent through one or more (e.g., each) PCE dies in the PIM chain 916 until the PCG die 910 receives the computing results, e.g., having self-reduction and / or completion packets, from one or more (e.g., each) PCE dies in the PIM chain 916. The second PIM chain 918 may perform a similar distributed chained operation using one or more commands, input data, etc., sent by the host 906 using one or more SiP-to-SiP interconnections. One or more PCG dies 910 at one or more (e.g., each) PIM chains in the PIM chains 916, 918, … may send one or more results, completions, etc., to the host 906 using one or more SiP-to-SiP interconnections. In some embodiments, one or more of the computing operations may be performed by one or more HBM-PIM dies located at one or more of the PCE dies 908, 912, …, 914.
[0093] Figure 10 FIG. shows an example embodiment of a PIM chain cluster and an associated method according to an example embodiment of the present disclosure. Figure 10 The illustrated PIM chain cluster and associated method 1002 may be implemented using or for implementing any embodiment of the computing nodes, gateway nodes, memory nodes, accelerator nodes, PIM chains, clusters, etc., disclosed herein. Some models for LLM scenarios may implement matrix multiplication operations, which may consume relatively large amounts of resources, such as memory capacity, memory bandwidth, computing bandwidth, etc. In Figure 10In the illustrated embodiment, the learning model for the LLM solution can be loaded into one or more of the one or more PCE dies PCE 1006 to PCE 1018 (e.g., each HBM). The PCG die 1004 can receive an activation vector or a user query, designated as Act.1. The PCG 1004 can send the activation vector (Act.1) to a PCE chain including PCE 1006 - PCE 1018. The activation vector (Act.1) can be used, for example, for matrix multiplication within the HBM PIM device within PCE 1006. The HBM PIM device in PCE 1006 can perform internal calculation operations using the activation vector (Act.1). The result from the matrix multiplication in PCE 1006 can be provided to the next PCE designated as PCE 1008. Q1n can indicate the result from the matrix multiplication in PCE 1006. The calculation result Q1n and the activation vector (Act.1) can be forwarded to the next element in the chain designated as PCE1008.
[0094] PCE 1008 can perform another matrix multiplication using the activation vector and one or more rows of elements of the distributed weight matrix WQ1. PCE 1008 can generate another result Q1n, which can be added to the result of the previous chain element PCE 1006. This process can continue in PCE 1010 to PCE 1018, where one or more (e.g., each) PCE can perform matrix multiplication, provide a corresponding result, and add the corresponding result to the result of the previous chain element. In some embodiments, all or part of the sum is designated as Act.4, and / or the completion indicating the completion of the chained matrix multiplication can be provided back to PCG1004.
[0095] In an example embodiment, the system can be initialized when a host sends LLM weight data to the PCG 1004, for example, using a PCIe connection. The PCG 1004 can distribute the weight data to the PCEs 1006 - 1018. The host can send a command for performing matrix multiplication to the PCG 1004. The PCG 1004 can generate one or more command_packets and send them to the PCE 1006, for example, with activations (e.g., operands). The first PCE (1006) can perform a distributed computing operation based on the command packet (e.g., "command_packet") and forward, for example, {command_packet, activation v(12288), self-reduction result v(12288), completion_packet} to the next PCE (1008). When one or more (e.g., all) of the chain elements have performed the corresponding matrix multiplication, the PCG 1004 can receive the computation results from one or more (e.g., all) of the PCEs, for example, with self-reduction and / or completion_packet.
[0096] Figure 11 A diagram showing an example embodiment of a pipelined matrix multiplication reduction operation according to an example embodiment of the present disclosure. Figure 11 The illustrated pipelined matrix multiplication reduction operation 1102 can be implemented using or used to implement any embodiment of the computing nodes, gateway nodes, memory nodes, accelerator nodes, PIM chains, clusters, etc. disclosed herein. In an example embodiment, the activation vector 1106 can have 12K (e.g., 12288) elements. For example, one or more elements (e.g., each element) can include 2-byte floating-point unit data. The PCG 1104 can provide the activation vector 1106 to the first PCE (e.g., PCE[0] 1108). One or more of the PCEs (e.g., each PCE), PCE[0] 1108 to PCE[N - 1] 1114, can be associated with a PCE chain such as PCEs 1006 to 1018. One or more (e.g., each) of PCE[0] 1108 to PCE[N - 1] 1114 can act on different sections of the rows of the distributed weight function and perform matrix multiplication in a pipelined and / or parallel manner.
[0097] For example, the activation vector 1106 can be provided by the PCG 1104 to the PCE[0] 1108 together with the first X number of rows of distributed Weight[0] 1118. In some embodiments, the rows can be broken down into, for example, 32, 64, or 128 segments. For example, any number of rows can be allocated to each PCE according to the relative processing capabilities of the PCEs. When the first 128 rows of the allocated Weight[0] 1118 are provided to the PCE[0] 1108 for matrix multiplication, the PCE[0] 1108 can perform matrix multiplication of AX0 by AX1 and apply reduction. This can enable the next layer of the pipeline, PCE[1] 1110, to receive AX01, which can be used to perform its own matrix multiplication. In Figure 11 the illustrated embodiment, time can increase to the right. The PCE[1] 1110 can use the first column data received from 1108 to perform its own matrix multiplication. Thus, the multiplication can be cascaded down to one or more PCEs performing parallel matrix multiplication.
[0098] In some embodiments, one or more D2D interconnects can be used to perform pipeline reduction. In some embodiments, pipeline reduction can be performed on a tile basis. For example, the size of the tile can be adjusted based on the input element size of one or more tensor units used for GPUs, NPUs, etc. For example, the tile can be 24×24 or 32×32. One or more column elements in 1104 - 1116 can each indicate a tile. In some embodiments, the flow control scheme can be based on tile units. In some embodiments, a hardware queue queuing mechanism can be used to implement inter-process communication message handling.
[0099] Figure 12 FIG. 1200 shows an embodiment of a method for performing PIM chain operations according to an example embodiment of the present disclosure. Although the example method may show a specific sequence of operations, in other embodiments, the sequence can be changed without departing from the scope of the present disclosure. For example, some of the depicted operations can be performed in parallel or in a different order that may not substantially affect the functionality of the routine. In other examples, different components of an example device or system implementing the routine can perform functions substantially simultaneously or in a specific order.
[0100] At operation 1202, the PCG can generate one or more command packets and send the command packet with the activation vector to the first attached PCE in the chain.
[0101] At operation 1204, the first or subsequent PCEs can perform one or more distributed computing operations based on the command packet. In some embodiments, one or more of the computing operations can be performed in whole or in part by one or more stacked HBM - PIM dies.
[0102] At operation 1206, a first or subsequent PCE may forward results (e.g., in a packet) to the next PCE in the chain. The forwarded packet may include data such as activation vectors, self-reduction results, and command packets. Example packets may include: {command_packet, activation v(12288), self-reduction result v(12288), completion_packet}.
[0103] At operation 1208, the PCE may detect whether it is the end of the PCE chain. If there are more PCEs in the PCE chain, the method may continue at operation 1204. When the PCE is the last PCE in the chain, the method may proceed to operation 1210.
[0104] At operation 1210, the PCG may receive computation results (e.g., formatted as a final result) from one or more (e.g., all) PCEs in the chain. The computation results may include self-reduction and / or one or more completion packets.
[0105] According to some embodiments, a method and apparatus may include a first Processing-in-Memory (PIM) Chain Gateway (PCG) die connected to first and second Processing-in-Memory (PIM) Chain Element (PCE) dies, where the first and second PCE dies have High Bandwidth Memory (HBM). In some embodiments, the PCG may be connected to the first PCE die and the last PCE die of a PIM chain computing cluster.
[0106] The PCG may also include a management Central Processing Unit (CPU) and DMA (Direct Memory Access). In some embodiments, the PCG may also include a connection interface connected to the PCE and a host interface connected to a second PCG die. The connection interface may be Universal Chiplet Interconnect Express (UCIe). In some embodiments, the host interface may be a Peripheral Component Interconnect Express (PCIe) or Compute Express Link (CXL) interface. The PCE die may also include an HBM controller and an HBM Through-Silicon Via (TSV) Input / Output (I / O) platform. The PCE die may also include a Direct Memory Access (DMA) engine, and a connection interface configured to connect to other PCEs. The PCG may also include a connection interface connected to the PCE and a host interface connected to a host device. In some embodiments, the PCG may include multiple connection interfaces connected to multiple PCEs. In some embodiments, the host interface may be connected to one or more host CPUs, or to the system and / or host interface of the next PCE on another PIM chain cluster.
[0107] The device may also include a PCG circuit configured to communicate with a first PCE circuit including an HBM. In some embodiments, the PCG may be configured to send a command packet to the first PCE. The command packet may include an activation vector. The first PCE may be configured to perform a computing operation based on the command packet. The first PCE may also be configured to generate a result based on performing the computing operation and forward the result to a second PCE connected to the first PCE. The first PCE and the second PCE may be configured to perform computing operations in parallel in a pipelined manner. In some embodiments, there may be multiple PCEs within a PIM chain cluster. The computing operation may include matrix multiplication using a distributed weight matrix.
[0108] Some embodiments may include a PCE chain that includes an HBM controller, a processor, and a direct memory engine. The memory may include a PCG configured to send commands to the PCE chain.
[0109] In some embodiments, the PCE may be configured to perform a computing operation by itself or by utilizing a stacked HBM-PIM module based on a distributed weight matrix. In some embodiments, the PCE may also be configured to perform matrix multiplication based on a distributed weight matrix and an activation vector received from the PCG. The PCE may also be configured to perform computing operations in parallel. In other embodiments, the PCE may be configured to provide the result from one PCE in the PCE chain to the next PCE in the PCE chain.
[0110] Any functionality described herein, including any control logic, computing elements, interfaces, etc., may be implemented in hardware, software, firmware, or any combination thereof, including, for example, hardware and / or software combinational logic, sequential logic, timers, counters, registers, state machines, volatile memory (such as DRAM and / or SRAM), non-volatile memory (including flash memory), persistent memory (such as cross-grid non-volatile memory), memory with bulk resistance change, PCM, etc. and / or any combination thereof, CPLD, FPGA, ASIC, CPU, GPU, NPU, TPU, DPU, etc., executing instructions stored in any type of memory. In some embodiments, one or more components may be implemented as one or more SOCs, SIPs, etc.
[0111] Some of the embodiments disclosed above have been described in the context of various implementation details, but the principles of the present disclosure are not limited to these or any other specific details. For example, some functions have been described as being implemented by certain components, but in other embodiments, the functions can be distributed among different systems and components located at different positions and having various interfaces. Certain embodiments have been described as having specific processes, operations, etc., but these terms also cover embodiments in which the specific processes, operations, etc. can be implemented by multiple processes, operations, etc., or embodiments in which multiple processes, operations, etc. can be integrated into a single process, step, etc. A reference to a component or element can refer only to a part of the component or element. For example, a reference to a block can refer to the entire block or one or more sub-blocks. The use of terms such as "first" and "second" in the present disclosure and the claims can be for the sole purpose of distinguishing the elements they modify and may not indicate any spatial or temporal order, unless otherwise apparent from the context. In some embodiments, a reference to an element can refer to at least a part of the element, e.g., "based on" can mean "at least partially based on", etc. A reference to a first element may not imply the existence of a second element. The principles disclosed herein have independent utility and can be embodied separately, and not every embodiment can utilize every principle. However, these principles can also be embodied in various combinations, some of which can amplify the benefits of the individual principles in a synergistic manner. The various details and embodiments described above can be combined to generate additional embodiments in accordance with the inventive principles of this patent disclosure.
[0112] In some embodiments, a part of an element can refer to less than or all of the element. A first part of an element and a second part of the element can refer to the same part of the element. The first part of an element and the second part of the element can overlap (e.g., a part of the first part can be the same as a part of the second part).
[0113] Since the inventive principles of this patent disclosure can be modified in arrangement and detail without departing from the inventive concept, these changes and modifications are considered to fall within the scope of the appended claims.
Claims
1. An apparatus for processing in memory using die-to-die (D2D) interconnect, comprising: A first device includes at least one first die, wherein the at least one first die includes: First interface; a storage medium configured to store first information received using the first interface; and a computing element configured to perform a computing operation using the first information; and The second device comprises at least one second die, wherein the at least one second die comprises: A second interface configured to receive second information; and A third interface is coupled to the first interface and is configured to send the first information based on the second information.
2. The device according to claim 1, wherein: The computing operation generates third information; and The at least one first die includes a fourth interface configured to transmit the third information.
3. The apparatus of claim 1 , wherein the memory medium is a first memory medium, the computing element is a first computing element, the computing operation is a first computing operation, and the at least one second die comprises: a second storage medium configured to store at least a portion of the second information; as well as A second computing element is configured to perform a second computing operation using the at least a portion of the second information.
4. The apparatus of claim 3, wherein the second computing operation generates at least a portion of the first information. The apparatus according to claim 1 , wherein the first information includes at least a portion of the second information. The apparatus of claim 1 , wherein the second information comprises input data for the computing operation. The apparatus according to claim 1 , wherein the second information comprises command information for the computing operation.
8. The device according to claim 1, wherein: The calculation operation generates third information; The at least one first die comprises a fourth interface configured to transmit the third information; and The at least one second die includes a fifth interface configured to receive the third information.
9. The apparatus of claim 1 , wherein the memory medium is a first memory medium, the computing element is a first computing element, the computing operation is a first computing operation for generating third information, the at least one first die comprises a fourth interface configured to transmit the third information, and the apparatus comprises: A third device includes at least one third die, wherein the at least one third die includes: a fifth interface, configured to receive the third information; a second storage medium configured to store at least a portion of the third information; and A second computing element is configured to perform a second computing operation using the at least a portion of the third information.
10. An apparatus for processing in memory using die-to-die (D2D) interconnects, comprising: A tube core, the tube core comprising: First interface; Second interface; A third interface; and at least one control circuit, the at least one control circuit being configured to: Using the first interface to receive first information for computing operations; receiving second information for the computing operation using the second interface; controlling the computing operation, wherein the computing operation is performed by at least one computing element using the first information and the second information to generate third information; and The third information is sent using the third interface.
11. The device according to claim 10, wherein: The first interface comprises a first die interface; The second interface comprises a memory interface; and The third interface includes a second die interface.
12. The apparatus of claim 10, wherein the at least one control circuit is configured to: receiving at least a portion of the second information using the first interface; and The at least a portion of the second information is sent using the second interface.
13. The apparatus of claim 10, wherein the at least one control circuit is configured to: Using the first interface to receive command information; and The computing operation is performed based on the command information.
14. The apparatus of claim 10, wherein the at least one control circuit is configured to: Using the first interface to receive command information; and At least a portion of the command information is sent using the third interface.
15. The apparatus of claim 10, wherein the at least one control circuit comprises a memory controller.
16. The apparatus of claim 10, wherein the at least one control circuit is configured to access the computing element using the second interface.
17. The apparatus of claim 10, wherein the die comprises at least a portion of the computing element.
18. A method for in-memory processing using die-to-die (D2D) interconnection, comprising: receiving information at a first device using a first die interface, wherein the first device includes a first memory medium and a first computing element; storing at least a first portion of said information in said first storage medium; sending at least a second portion of the information from the first device to a second device using a second die interface, wherein the second device includes a second memory medium and a second computing element; storing said second portion of said information in said second storage medium; performing a first computing operation using the first portion of the information and the first computing element; and A second computing operation is performed using the second portion of the information and the second computing element.
19. The method according to claim 18, wherein: The information is first information, and the method further includes: sending second information from the first device to the second device using the second die interface; The second computing operation is performed using the second information.
20. The method of claim 18, wherein: The first computational operation comprises a first portion of a matrix operation; and The second computational operation comprises a second portion of the matrix operation.