Method for processing in memory and memory system
By incorporating an AI accelerator into the memory base die, the data transmission challenges of AI workloads are addressed, high-bandwidth data transmission and low-latency processing are achieved, and system performance and energy efficiency are improved.
Patent Information
- Application Number
- CN202510333282.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-01-27
- Filing Date
- 2025-03-20
- Publication Date
- 2025-09-23
AI Technical Summary
Existing AI workloads are rapidly increasing the demand for data movement bandwidth and storage capacity, making it difficult for data centers and related equipment to keep up with the demand, requiring high-throughput and low-latency memory solutions.
Incorporate processing units (such as AI accelerators) on the memory base die, route data queries to the processing units for processing through the memory controller, and connect the computing die through a silicon interposer to handle compute-limited operations, using high-bandwidth memory (HBM) to increase data transmission bandwidth.
It improves the data transmission speed between memory and AI accelerator, reduces power consumption, improves system performance and AI response time, and reduces system cost.
Smart Images

Figure CN120687028A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to methods and memory systems for performing processing in memory. Specifically, the present subject matter relates to systems and methods for incorporating processing units (e.g., artificial intelligence (AI) accelerators) on a memory base die. Background Art
[0002] The Background section is intended only to provide context, and the disclosure of any concept in this section does not constitute an admission that such concept is prior art.
[0003] AI workloads require memory and storage solutions that offer high throughput and low latency to accommodate the rapid processing of relatively large data sets. High-throughput memory / storage devices ensure rapid data read and write speeds. Low-latency memory / storage devices provide fast data access for real-time AI applications. However, the surge in AI has led to a rapidly increasing demand for improvements in data movement bandwidth and data storage capacity, making it difficult for data centers and related equipment to keep pace. Summary of the Invention
[0004] In various embodiments, the systems and methods described herein include systems, methods, and apparatus for incorporating a processing unit (e.g., an AI accelerator) onto a memory base die (e.g., a base die of high bandwidth memory). In some aspects, the technology described herein relates to a method for performing processing in memory, the method comprising: determining at least one characteristic of a data query; routing a first function of the data query to the memory base die for processing by a processing unit on the memory base die based on the at least one characteristic of the data query; and processing, via the processing unit, data received by a memory controller on the memory base die from at least one of one or more memory dies stacked on top of the memory base die.
[0005] In some aspects, the technology described herein relates to a method that also includes routing the first function to a memory base die for processing by a processing unit on the memory base die based on determining that the first function is a memory-constrained function.
[0006] In some aspects, the technology described herein relates to a method that also includes routing the second function of the data query to a compute die for processing by the compute die based on determining that the second function is a compute-limited operation, wherein the compute die is connected to the memory base die via a silicon interposer of a system-in-package that includes the compute die and the memory base die.
[0007] In some aspects, the technology described herein relates to a method in which a memory base die includes a memory expansion port connected to at least one of a low power double data rate memory or a graphics double data rate memory external to the memory base die.
[0008] In some aspects, the technology described herein relates to a method that also includes at least one of routing a first category of functions to a memory base die via a memory controller for processing by a processing unit on the memory base die, routing a second category of functions to a compute die via the memory controller for processing by the compute die, or routing a third category of functions to at least one of a low power double data rate memory or a graphics double data rate memory external to the memory base die via the memory controller.
[0009] In some aspects, the technology described herein relates to a method, further comprising: transferring data from one or more memory dies to a physical layer interface of a memory base die through a through silicon via; transferring the data from the physical layer interface to a memory controller of the memory base die; and transferring the data from the memory controller to a shared memory on the memory base die, wherein the shared memory stores the data for processing by a processing unit.
[0010] In some aspects, the techniques described herein relate to a method in which a processing unit includes at least one of: a tensor core configured for matrix multiplication; or an accumulator configured for accumulating intermediate computations.
[0011] In some aspects, the technology described herein relates to a method in which: a memory controller is connected to a processing unit via a network-on-chip (NOC) interconnect bus, and the memory controller is connected to a dynamic random access memory (DRAM) physical layer on a memory base die via a double data rate (DDR) physical layer interface of the memory base die.
[0012] In some aspects, the technology described herein relates to a method in which: a system bus interface is connected to a die-to-die interface of a memory base die, and a processing unit is connected to the system bus interface via a NOC interconnect bus, the system bus interface converting data in a die-to-die flow control unit format into a network packet format.
[0013] In some aspects, the technology described herein relates to a method in which a memory controller is communicatively coupled to a processing unit, and a second memory controller on a memory base die is communicatively coupled to a second processing unit on the memory base die.
[0014] In some aspects, the technology described herein relates to a system comprising: a memory base die comprising: a memory controller; one or more memory dies stacked on top of the memory base die; and a processing unit configured to process data received by the memory controller from at least one of the one or more memory dies, the data being routed to the processing unit based on at least one characteristic of a data query associated with the data; an interconnect connecting the memory controller to the one or more memory dies stacked on the memory base die and to a plurality of processing units including the processing unit; and a die-to-die interface connecting the memory base die to a compute die of a system-level package.
[0015] In some aspects, the technology described herein relates to a system in which, based on determining that a function queried for data is a memory-bound function, the function is routed to a memory base die for processing by a processing unit.
[0016] In some aspects, the technology described herein relates to a system in which, based on determining that a function queried for data is a compute-bound operation, the function is routed to a compute die for processing by the compute die, where the compute die is connected to a memory base die via a silicon interposer of a system-in-package.
[0017] In some aspects, the technology described herein relates to a system in which a system-in-package includes a plurality of memory base dies connected to a compute die, the plurality of memory base dies including the memory base die.
[0018] In some aspects, the technology described herein relates to a system wherein: the memory base die includes a system bus interface that connects an interconnect to a die-to-die interface of the memory base die, and the system bus interface maps a data format used by the interconnect to a data format used by the die-to-die interface.
[0019] In some aspects, the technology described herein relates to a system in which a memory base die includes a shared memory that shares data between a first processing unit and a second processing unit of a plurality of processing units.
[0020] In some aspects, the technology described herein relates to a system in which a memory base die includes a memory expansion port connected to at least one of a low power double data rate memory or a graphics double data rate memory external to the memory base die.
[0021] In some aspects, the technology described herein relates to a non-transitory computer-readable medium storing code, the code including instructions executable by a processor of a device to perform the following operations: determining at least one characteristic of a data query; routing a first function of the data query to a memory base die for processing by a processing unit on the memory base die based on the at least one characteristic of the data query; and processing, via the processing unit, data received by a memory controller on the memory base die from at least one of one or more memory dies stacked on top of the memory base die.
[0022] In some aspects, the technology described herein relates to a non-transitory computer-readable medium, wherein the code includes further instructions executable by a processor to: based on determining that the first function is a memory-constrained function, route the first function to a memory base die for processing by a processing unit on the memory base die.
[0023] In some aspects, the technology described herein relates to a non-transitory computer-readable medium, wherein the code includes further instructions executable by a processor to: based on determining that a second function of the data query is a compute-limited operation, route the second function to a compute die for processing by the compute die, wherein the compute die is connected to the memory base die via a silicon interposer of a system-in-package that includes the compute die and the memory base die.
[0024] A computer-readable medium is disclosed. The computer-readable medium may store instructions that, when executed by a computer, cause the computer to perform substantially the same or similar operations as further disclosed herein. Similarly, non-transitory computer-readable media, devices, and systems for performing substantially the same or similar operations are further disclosed.
[0025] The technology described herein for AI accelerators on a high-bandwidth memory (HBM) die offers several advantages and benefits. For example, combining an AI accelerator on an HBM die provides increased data bandwidth, allowing for faster data transfer between the memory and the AI accelerator. This faster data transfer between the AI accelerator and HBM memory results in faster system performance, faster processing times, improved AI response times (e.g., faster query response times), lower power consumption, and improved power efficiency.
[0026] Beneficial effects
[0027] According to the present disclosure, an AI accelerator can be coupled to an HBM die. Therefore, the data bandwidth of the storage device can be increased to achieve faster data transmission between the memory and the AI accelerator.
[0028] Thus, systems and methods are provided for integrating AI accelerators into memory base dies with improved performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The above-mentioned aspects and other aspects of the present systems and methods will be better understood when this application is read in light of the following drawings, in which like reference numerals represent similar or identical elements. In addition, the drawings provided herein are only for the purpose of illustrating certain embodiments; other embodiments that may not be explicitly shown are not excluded from the scope of this disclosure.
[0030] These and other features and advantages of the present disclosure will be appreciated and understood with reference to the specification, claims, and drawings, in which:
[0031] Figure 1 An example system is shown according to one or more implementations as described herein.
[0032] Figure 2 shows a schematic diagram according to one or more embodiments described herein. Figure 1 Details of the system.
[0033] Figure 3 An exemplary base die is shown according to one or more embodiments as described herein.
[0034] Figure 4 An example package is shown according to one or more implementations as described herein.
[0035] Figure 5 An exemplary base die is shown according to one or more embodiments as described herein.
[0036] Figure 6 Depicted is a flowchart illustrating an example method associated with the disclosed system according to example implementations described herein.
[0037] Figure 7 Depicted is a flowchart illustrating an example method associated with the disclosed system according to example implementations described herein.
[0038] Figure 8 Depicted is a flowchart illustrating an example method associated with the disclosed system according to example implementations described herein.
[0039] Figure 9 Depicted is a flowchart illustrating an example method associated with the disclosed system according to example implementations described herein.
[0040] While the present system and method are susceptible to various modifications and alternative forms, specific embodiments thereof are shown by way of example in the drawings and will be described herein. The drawings may not be to scale. However, it should be understood that the drawings and detailed description thereof are not intended to limit the present system and method to the particular forms disclosed, but rather, the invention is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present system and method as defined by the appended claims. DETAILED DESCRIPTION
[0041] The details of one or more embodiments of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims.
[0042] Various embodiments of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, some but not all of which are shown. Indeed, the present disclosure may be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure will satisfy applicable legal requirements. Unless otherwise indicated, the term "or" is used herein in an alternative and conjunction sense. The terms "illustrative" and "exemplary" are used as examples without an indication of a level of quality. The same reference numerals represent the same elements throughout. The arrows in each figure depict bidirectional data flow and / or bidirectional data flow capability. The terms "path," "pathway," and "route" are used interchangeably herein.
[0043] Embodiments of the present disclosure may be implemented in various ways, including as a computer program product comprising an article of manufacture. A computer program product may include a non-transitory computer-readable storage medium that stores an application, program, program component, script, source code, program code, object code, byte code, compiled code, interpreted code, machine code, executable instructions, etc. (also referred to herein as executable instructions, instructions for execution, computer program product, program code, and / or similar terms used interchangeably herein). Such non-transitory computer-readable storage media includes all computer-readable media (including volatile and non-volatile media).
[0044] In one embodiment, the non-volatile computer-readable storage medium may include a floppy disk, a flexible disk, a hard disk, a solid-state storage (SSS) (e.g., a solid-state drive (SSD)), a solid-state card (SSC), a solid-state module (SSM), an enterprise flash drive, a magnetic tape, or any other non-transitory magnetic medium. The non-volatile computer-readable storage medium may include a punched card, paper tape, optical marker sheet (or any other physical medium having a pattern of holes or other optically recognizable markers), a compact disc read-only memory (CD-ROM), a compact disc rewritable (CD-RW), a digital versatile disc (DVD), a Blu-ray disc (BD), or any other non-transitory optical medium. Such non-volatile computer-readable storage medium may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory (e.g., serial, NAND, NOR, etc.), a multimedia memory card (MMC), a secure digital (SD) memory card, a smart media card, a compact flash (CF) card, a memory stick, or the like. Additionally, the non-volatile computer-readable storage medium may include conductive bridging random access memory (CBRAM), phase change random access memory (PRAM), ferroelectric random access memory (FeRAM), non-volatile random access memory (NVRAM), magnetoresistive random access memory (MRAM), resistive random access memory (RRAM), silicon-oxide-nitride-oxide-silicon memory (SONOS), floating junction gate random access memory (FJG RAM), millipede memory, racetrack memory, and / or the like.
[0045] In one embodiment, the volatile computer-readable storage medium may include random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), fast page mode dynamic random access memory (FPM DRAM), extended data-out dynamic random access memory (EDO DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), double data rate type 2 synchronous dynamic random access memory (DDR2 SDRAM), low power DDR (LPDDR), graphics DDR (GDDR), double data rate type 3 synchronous dynamic random access memory (DDR3 SDRAM), Rambus dynamic random access memory (RDRAM), two-transistor RAM (TTRAM), thyristor RAM (T-RAM), zero capacitor (Z-RAM), Rambus in-line memory module (RIMM), dual in-line memory module (DIMM), single in-line memory module (SIMM), video random access memory (VRAM), cache memory (including various levels), flash memory, register memory, etc. It should be understood that where embodiments are described as using computer-readable storage media, other types of computer-readable storage media may be used in place of or in addition to the computer-readable storage media.
[0046] It should be understood that various embodiments of the present disclosure may be implemented as methods, apparatuses, systems, computing devices, computing entities, etc. Thus, embodiments of the present disclosure may take the form of apparatuses, systems, computing devices, computing entities, etc. that execute instructions stored on a computer-readable storage medium to perform certain steps or operations. Accordingly, embodiments of the present disclosure may take the form of hardware embodiments that perform certain steps or operations, computer program product embodiments, and / or embodiments that include a combination of computer program products and hardware.
[0047] The embodiments of the present disclosure are described below with reference to block diagrams and flow charts. Therefore, it should be understood that each box illustrated in the block diagrams and flow charts can be implemented in the form of a computer program product, a hardware embodiment, a combination of hardware and computer program products, and / or an apparatus, system, computing device, computing entity, etc. that executes instructions, operations, steps, and similar terms that can be used interchangeably (e.g., executable instructions, instructions for execution, program code, etc.) on a computer-readable storage medium. For example, the retrieval, loading, and execution of code can be performed sequentially so that one instruction is retrieved, loaded, and executed at a time. In some example embodiments, retrieval, loading, and / or execution can be performed in parallel so that multiple instructions are retrieved, loaded, and / or executed together. Therefore, such embodiments can produce a specially configured machine that performs the steps or operations specified in the block diagrams and flow charts. Therefore, the block diagrams and flow charts support various combinations of embodiments for executing specified instructions, operations, or steps.
[0048] References throughout this specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment disclosed herein. Thus, the phrases "in one embodiment" or "in an embodiment" or "according to an embodiment" (or other phrases of similar meaning) that appear in various places throughout this specification may not necessarily all refer to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In this regard, as used herein, the word "exemplary" means "serving as an example, instance, or illustration." Any embodiment described herein as "exemplary" should not be construed as necessarily preferred or advantageous over other embodiments. Additionally, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. Furthermore, depending on the context of the discussion herein, singular terms may include corresponding plural forms, and plural terms may include corresponding singular forms. Similarly, hyphenated terms (e.g., "two-dimensional," "pre-determined," "pixel-specific," etc.) may occasionally be used interchangeably with corresponding non-hyphenated versions (e.g., "two dimensional," "predetermined," "pixel-specific," etc.), and capitalized terms (e.g., "Counter Clock," "Row Select," "PIXOUT," etc.) may be used interchangeably with corresponding non-capitalized versions (e.g., "counterclock," "row select," "pixout," etc.). Such occasional interchangeable usage should not be considered inconsistent with one another.
[0049] Furthermore, depending on the context of the discussion herein, singular terms may include corresponding plural forms, and plural terms may include corresponding singular forms. It should also be noted that the various figures (including component diagrams) shown and discussed herein are for illustrative purposes only and are not drawn to scale. Similarly, various waveforms and timing diagrams are shown for illustrative purposes only. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity. Furthermore, where deemed appropriate, reference numerals are repeated in the figures to indicate corresponding and / or similar elements.
[0050] The terminology used herein is for the purpose of describing some example embodiments only and is not intended to limit the claimed subject matter. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that when used in this specification, the terms "include" and / or "comprise" specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0051] It should be understood that when an element or layer is referred to as being on, "connected to," or "coupled to" another element or layer, it can be directly on, connected to, or coupled to the other element or layer, or there can be intervening elements or layers. In contrast, when an element is referred to as being "directly on," "directly connected to," or "directly coupled to" another element or layer, there are no intervening elements or layers. The same reference numerals always represent the same elements. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0052] As used herein, the terms "first," "second," and the like are used as labels for the nouns that follow them and do not imply any type of ordering (e.g., spatial, temporal, logical, etc.) unless explicitly defined as such. In addition, the same reference numerals may be used across two or more figures to refer to parts, components, blocks, circuits, units, or modules having the same or similar functions. However, such usage is merely for simplicity of illustration and ease of discussion; it does not mean that the construction or architectural details of such components or units are the same in all embodiments, or that such commonly referenced parts / modules are the only way to implement some example embodiments disclosed herein.
[0053] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this subject matter belongs. It will be further understood that terms (such as those defined in commonly used dictionaries) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless explicitly defined as such herein.
[0054] As used herein, the term "module" refers to any combination of software, firmware, and / or hardware configured to provide the functionality described herein in conjunction with the module. For example, software may be embodied as a software package, code, and / or instruction set or instructions, and the term "hardware" as used in any embodiment described herein may include, for example, components, hardwired circuits, programmable circuits, state machine circuits, and / or firmware that stores instructions executed by programmable circuits, alone or in any combination. Modules may collectively or individually be embodied as circuits that form part of a larger system, such as, but not limited to, an integrated circuit (IC), a system on a chip (SoC), a component, or the like.
[0055] The following description is presented to enable one of ordinary skill in the art to make and use the subject matter disclosed herein and to incorporate it into the context of a particular application.While the following is directed to specific examples, other and further examples may be devised without departing from the basic scope thereof.
[0056] For those skilled in the art, various modifications and various uses in different applications will be obvious, and the general principles defined herein can be applied to a wide range of embodiments. Therefore, the subject matter disclosed herein is not intended to be limited to the embodiments presented, but is to be consistent with the widest scope consistent with the principles and novel features disclosed herein.
[0057] In the description provided, numerous specific details are set forth in order to provide a more thorough understanding of the subject matter disclosed herein. However, it will be apparent to those skilled in the art that the subject matter disclosed herein may be practiced without being limited to these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the subject matter disclosed herein.
[0058] Unless expressly stated otherwise, all features disclosed in this specification (for example, any accompanying claims, abstract, and drawings) may be replaced by alternative features serving the same, equivalent, or similar purpose. Thus, unless expressly stated otherwise, each feature disclosed is one example only of a generic series of equivalent or similar features.
[0059] Various features are described herein with reference to the accompanying drawings. It should be noted that the drawings are intended only to facilitate description of the features. The various features described are not intended to be an exhaustive description of the subject matter disclosed herein or to limit the scope of the subject matter disclosed herein. In addition, the examples shown do not necessarily have all the aspects or advantages shown. An aspect or advantage described in conjunction with a particular example is not necessarily limited to that example and may be practiced in any other example, even if not shown as such or if not explicitly described as such.
[0060] It should be noted that the labels left, right, front, back, top, bottom, forward, backward, clockwise, and counterclockwise, if used, are used for convenience only and are not intended to imply any particular fixed direction. Instead, the labels are used to reflect the relative position and / or orientation of various parts of an object.
[0061] Any data processing may include data buffering, alignment of incoming data from multiple communication channels, forward error correction ("FEC"), and / or other functions. For example, data may first be received by an analog front end (AFE), which prepares the incoming data for digital processing. The digital portion of the transceiver (e.g., a DSP) may provide skew management, equalization, reflection cancellation, and / or other functions. It will be appreciated that the processes described herein can provide numerous benefits, including power and cost savings.
[0062] Furthermore, the terms "system," "component," "module," "interface," "model," and the like are generally intended to refer to a computer-related entity, whether hardware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, a thread of execution, a program, and / or a computer. By way of illustration, both an application running on a controller and the controller can be a component. One or more components can reside within a process and / or thread of execution, and a component can be localized on one computer and / or distributed between two or more computers.
[0063] Unless expressly stated otherwise, each numerical value and range may be interpreted as approximate as if the word "about" or "approximately" preceded the value of the value or range. Signals and corresponding nodes or ports may be represented by the same name and are interchangeable for purposes herein.
[0064] While embodiments may have been described with respect to circuit functionality, embodiments of the subject matter disclosed herein are not limited thereto. Possible implementations may be embodied in a single integrated circuit, a multi-chip module, a single card, a system-on-chip, or a multi-card circuit package. It will be apparent to those skilled in the art that various embodiments may also be implemented as part of a larger system. Such embodiments may be used in conjunction with, for example, a digital signal processor, a microcontroller, a field programmable gate array, an application-specific integrated circuit, or a general-purpose computer.
[0065] It will be apparent to those skilled in the art that the various functions of the circuit elements can also be implemented as processing blocks in a software program. Such software can be used, for example, in a digital signal processor, a microcontroller, or a general-purpose computer. Such software can be embodied in the form of program code embodied in a tangible medium, such as a magnetic recording medium, an optical recording medium, a solid-state memory, a floppy disk, a CD-ROM, a hard drive, or any other non-transitory machine-readable storage medium. When the program code is loaded into a machine (such as a computer) and executed by the machine, the machine becomes a device for practicing the subject matter disclosed herein. When implemented on a general-purpose processor, the program code segments are combined with the processor to provide a unique device that operates similarly to a specific logic circuit. The described embodiments can also be embodied in the form of a bit stream or other sequence of signal values transmitted electrically or optically through a medium, magnetic field changes stored in a magnetic recording medium, etc., generated using the methods and / or devices described herein.
[0066] The system cost and power consumption of large language model (LLM) inference systems are rapidly increasing. The relatively high system cost and power consumption may make Generative Pre-Trained Transformer (GPT) AI models unsustainable. Compute node chip sizes are increasing to accommodate the computational power and memory bandwidth requirements of LLMs. However, compute die size may be limited by reticle size constraints (e.g., 33 millimeters (mm) × 26 mm). These physical limitations may restrict the amount of compute resources that can be added to a given compute die.
[0067] The described systems and methods may include and / or be based on incorporating processing units (e.g., AI accelerators) on the base die of stacked memory modules (e.g., DRAM layers stacked on an HBM base die), thereby providing an efficient way to increase computing resources in a memory system without increasing the size of the computing die. Consequently, these systems and methods improve the energy efficiency of AI computing systems, reduce system costs, and lower power consumption.
[0068] Incorporating an AI accelerator on the base die of a stacked memory module provides enhanced data latency between the compute node and the memory. Additionally, incorporating an AI accelerator on the HBM die also reduces power consumption by eliminating data movement from the HBM die to the compute die via an interposer (e.g., a silicon interposer, a redistribution layer (RDL) interposer, an organic interposer). Furthermore, faster data transfer between the AI accelerator and the HBM memory results in faster system performance, faster processing time, and improved AI response time (e.g., faster query response time). A system-in-package (SIP) in a high-performance graphics processing unit (GPU) / tensor processing unit (TPU) system can integrate both the compute die and multiple HBM dies on an interposer. The described systems and methods increase computing power and reduce the cost of SiP packaging.
[0069] Figure 1 An example system 100 is shown according to one or more implementations as described herein. Figure 1 , a machine 105 is shown, which may be referred to as a host, system, or server. Figure 1 The machine 105 is depicted as a tower computer, but embodiments of the present disclosure can be extended to any form factor or type of machine. For example, the machine 105 can be a rack server, a blade server, a desktop computer, a tower computer, a mini-tower computer, a desktop server, a laptop computer, a notebook computer, a tablet computer, etc.
[0070] The machine 105 may include a processor 110, a memory 115, and a storage device 120. The processor 110 may be any type of processor. Note that for ease of illustration, the processor 110 and other components discussed below are shown external to the machine: embodiments of the present disclosure may include these components within the machine. Although Figure 1 A single processor 110 is shown, but the machine 105 may include any number of processors, each of which may be a single-core or multi-core processor, each of which may implement a reduced instruction set computer (RISC) architecture or a complex instruction set computer (CISC) architecture (among other possibilities), and may be mixed in any desired combination. In some examples, the machine 105 may include or be part of a manufacturing system for manufacturing memory die assemblies, incorporating processing units on memory base dies, manufacturing compute dies and stacked memory dies in a package, and the like.
[0071] Processor 110 may be coupled to memory 115. Memory 115 may be any type of memory, such as flash memory, dynamic random access memory (DRAM), static random access memory (SRAM), persistent random access memory, ferroelectric random access memory (FRAM), or non-volatile random access memory (NVRAM), such as magnetoresistive random access memory (MRAM), phase change memory (PCM), or resistive random access memory (ReRAM). Memory 115 may include volatile and / or non-volatile memory. Memory 115 may use any desired form factor, such as a single inline memory module (SIMM), dual inline memory module (DIMM), non-volatile DIMM (NVDIMM), etc. Memory 115 may be any desired combination of different memory types and may be managed by memory controller 125. Memory 115 may be used to store what may be referred to as "short-term" data: that is, data that is not intended to be stored for an extended period of time. Examples of short-term data may include temporary files, data used locally by applications (which may have been copied from other storage locations), etc.
[0072] The processor 110 and the memory 115 may support an operating system under which various applications may run. These applications may issue requests (which may be referred to as commands) to read data from the memory 115 or the storage device 120 or write data to the memory 115 or the storage device 120. When the storage device 120 is used to support applications that read or write data via a certain file system, a device driver 130 may be used to access the storage device 120. Figure 1 One storage device 120 is shown, but any number (one or more) of storage devices may be present in the machine 105. The storage device 120 may support any desired protocol or protocols, including, for example, the Non-Volatile Memory Express (NVMe®) protocol, the Serial Attached Small Computer System Interface (SCSI) (SAS) protocol, or the Serial AT Attachment (SATA) protocol. The storage device 120 may include any desired interface, including, for example, the Peripheral Component Interconnect Express (PCIe®) interface or the Compute Link Express (CXL®) interface. The storage device 120 may utilize any desired form factor, including, for example, a U.2 form factor, a U.3 form factor, an M.2 form factor, an Enterprise and Data Center Standard Form Factor (EDSFF) (including all its variants, such as E1 short, E1 long, and E3 variants), or an Add-In Card (AIC).
[0073] Although Figure 1The term "storage device" is used, but embodiments of the present disclosure may include any storage device format that can benefit from the use of a computational storage unit, examples of which may include a hard drive, a solid-state drive (SSD), or a persistent memory device such as PCM, ReRAM, or MRAM. Any reference below to a "storage device" or "SSD" should be understood to include such other embodiments of the present disclosure and other types of storage devices. In some cases, the term "storage unit" may encompass both storage device 120 and memory 115. Machine 105 may include a power supply 135. Power supply 135 may provide power to machine 105 and its components.
[0074] Machine 105 may include a transmitter 145 and a receiver 150. Transmitter 145 or receiver 150 may be used to transmit or receive data, respectively. In some cases, transmitter 145 and / or receiver 150 may be used to communicate with memory 115 and / or storage device 120. Transmitter 145 may include write circuitry 160, which may be used to write data to a storage device, such as a register, in memory 115 and / or storage device 120. Similarly, receiver 150 may include read circuitry 165, which may be used to read data from a storage device, such as a register, from memory 115 and / or storage device 120. In the illustrated example, machine 105 may include a timer 155, which may be used to time one or more operations, indicate a time period, indicate an elapsed time, indicate an expiration time, indicate a timeout timeout, etc.
[0075] In one or more examples, machine 105 can be implemented using any type of device. Machine 105 can be configured as (e.g., host) one or more servers, such as compute servers, storage servers, storage nodes, network servers, supercomputers, data center systems, or any combination thereof. Additionally or alternatively, machine 105 can be configured as (e.g., host) one or more computers, such as workstations, personal computers, tablet computers, smartphones, or any combination thereof. Machine 105 can be implemented using any type of device, which can be configured to include, for example, an accelerator device, a storage device, a network device, a memory expansion and / or buffer device, a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), a tensor processing unit (TPU), an intelligence processing unit (IPU), an optical processing unit (OPU), or the like, or any combination thereof.
[0076] Any communication between devices comprising machine 105 (e.g., a host, computational storage devices, and / or any intermediate devices) may occur over interfaces that may be implemented using any type of wired and / or wireless communication media, interfaces, protocols, and the like, including: PCIe, NVMe, Ethernet, NVMe-oF, Compute Express Link (CXL) and / or coherence protocols such as CXL.mem, CXL.cache, CXL.IO, and the like, Gen-Z, Open Coherent Accelerator Processor Interface (OpenCAPI), Cache Coherent Interconnect for Accelerators (CCIX), Advanced Extensible Interface (AXI), Coherent Hub Interface (CHI), and the like, or any combination thereof, Transmission Control Protocol / Internet Protocol (TCP / IP), Fibre Channel, InfiniBand, Serial AT Attachment (SATA), Small Computer System Interface (SCSI), Serial Attached SCSI (SAS), iWARP, any generation of wireless networks including 2G, 3G, 4G, 5G, and the like, any generation of Wi-Fi, Bluetooth, Near Field Communication (NFC), and the like, or any combination thereof. In some embodiments, the communication interface may include a communication fabric including one or more links, buses, switches, hubs, nodes, routers, switches, repeaters, etc. In some embodiments, system 100 may include one or more additional devices having one or more additional communication interfaces.
[0077] Any functionality described herein (including any functionality in host functionality, device functionality, etc.) may be implemented in hardware, software, firmware, or any combination thereof, including, for example, hardware and / or software combinational logic, sequential logic, timers, counters, registers, state machines, volatile memory (such as at least one or any combination of dynamic random access memory (DRAM) and / or static random access memory (SRAM)), non-volatile memory (including flash memory), persistent memory (such as cross-grid non-volatile memory), memory with bulk resistance change, phase change memory (PCM), etc. and / or any combination thereof, a complex programmable logic device (CPLD) that executes instructions stored in any type of memory, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC) CPU (including a complex instruction set computer (CISC) processor (such as an x86 processor) and / or a reduced instruction set computer (RISC) processor (such as RISC-V and / or ARM processor)), a GPU, an NPU, a TPU, an IPU, an OPU, etc.
[0078] Figure 2 Shown according to the examples described herein Figure 1Details of machine 105 are shown. In the example shown, machine 105 may include a processor 110. Processor 110 may include one or more processors and / or one or more dies. Processor 110 may include a memory controller 125 (e.g., one or more memory controllers) and a clock 205 (e.g., one or more clocks), which may be used to coordinate the operation of components of the machine. Processor 110 may be coupled to memory 115 (e.g., one or more memory chips, stacked memory, etc.), which may include, for example, random access memory (RAM), read-only memory (ROM), or other state storage media. Processor 110 may be coupled to storage device 120 (e.g., one or more storage devices) and to a network connector 210, which may be, for example, an Ethernet connector or a wireless connector. Processor 110 may be connected to a bus 215 (e.g., one or more buses), a user interface 220 (e.g., one or more user interfaces), and input / output (I / O) interface ports that may be managed using an I / O engine 225 (e.g., one or more I / O engines), as well as other components attached to bus 215.
[0079] The systems and methods described herein include logic for providing an AI accelerator on an HBM die. The logic includes any combination of hardware (e.g., at least one memory, at least one processor), logic circuitry, firmware, and / or software to provide and / or implement an AI accelerator on an HBM die.
[0080] Figure 3 An exemplary base die 300 is shown according to one or more embodiments described herein. In some configurations, one or more aspects of base die 300 can be implemented by or in conjunction with machine 105, components of machine 105, or any combination thereof. In some examples, base die 300 can be the base layer of a stacked memory module (e.g., an HBM memory module, an HBM chip). In some cases, base die 300 can be part of a SoC. 2.5D and / or 3D stacking techniques can be used to stack memory chips on base die 300.
[0081] In some cases, the base die 300 may be referred to as a buffer die or a logic die. The base die 300 may include the bottom layer of the HBM stack (e.g., the bottom layer of the HBM module). The base die 300 may control one or more aspects of the stacked memory modules. In some cases, the base die 300 may be part of a system on chip (SoC). The SoC may include a processor, memory, input / output interfaces, etc. The SoC may include analog, digital, mixed signal, and other radio frequency functions, all located on a single chip substrate. In some cases, the base die 300 may be part of a system-in-package (SiP), which may include two or more dies containing integrated circuits stacked vertically on a package substrate.
[0082] In the example shown, the base die 300 may include expansion ports 305, test logic 310, a DRAM physical layer (PHY) 315, at least one memory controller (e.g., N memory controllers 320a-320N, where N is a positive integer such as 1, 2, 3, 4, 5, 6, 8, 10, 12, 16, 20, 32, etc.), a network-on-chip (NOC) interconnect 325, at least one processing unit (PU) (e.g., N PUs 330a-330N), a system bus interface 335 (e.g., a system bus mapping interface), and a die-to-die (D2D) 340. As shown, the base die 300 may include one or more connection interfaces (e.g., expansion ports 305, D2D 340).
[0083] In some examples, the expansion port 305 can provide D2D-based memory expansion. In some cases, the expansion port 305 can provide an expansion option for adding additional memory to a given system (e.g., additional memory for processing LLMs, etc.). In some examples, the expansion port 305 can connect to one or more memories external to the base die 300 (e.g., off-die memory). In some cases, the expansion port 305 can include a D2D interface separate from the D2D 340. In some cases, the expansion port 305 can expand the memory available for processing (e.g., AI processing). In some cases, a first process can be executed on the base die 300 via the N PUs 330a-330N. Additionally or alternatively, a second process can be executed in conjunction with a processor and memory external to the base die 300 but connected to the base die 300 via the expansion port 305. Additionally or alternatively, a third process can be executed by a compute die communicatively connected to the base die 300 (e.g., via the D2D 340). The base die 300 may include expansion ports 305 based on memory capacity and memory cost. The expansion ports 305 can expand the memory capacity of the base die 300 by connecting the base die 300 to memory external to the base die 300. Some types of memory may have lower costs than other types of memory. For example, LPDDR and GDDR may be slower than HBM, but due to the lower cost of LPDDR and GDDR, their capacity is generally greater than HBM. Therefore, the expansion ports 305 can connect the base die 300 to off-die memory (e.g., LPDDR, GDDR, etc.).
[0084] The test logic 310 may provide testing for the base die 300. In some cases, the test logic 310 may include a built-in self-test (BIST) that enables the base die 300 to test components of the base die 300 without the need for external test equipment. In some cases, the test logic 310 may generate a test pattern and apply the test pattern to the components of the base die 300, perform the test, and analyze the results to determine whether the component operates as expected or whether the test indicates an anomaly. In some cases, one or more aspects of the test logic 310 may be based on the Institute of Electrical and Electronics Engineers (IEEE) 1500, which may include standards that define how to test the design and / or operation of the base die 300.
[0085] DRAM PHY 315 may be based on and / or may include a physical layer (PHY). A PHY may include electronic circuitry that connects a network interface controller to a physical medium (e.g., copper connection, fiber optic). The PHY may be responsible for the physical layer functions of the Open Systems Interconnection (OSI) model. In some cases, a DDR memory system (e.g., including base die 300) may include a DDR memory controller and a DDR PHY (e.g., DDR PHY 315) to access DDR memory. DDR PHY 315 may include a DDR PHY interface (DFI), which may be based on an interface protocol that defines signals, timing, and programmable parameters for communicating control information and / or data between a memory controller (e.g., at least one of N memory controllers 320a-320N) and a PHY (e.g., DDR PHY 315). In the example shown, an N-channel (e.g., 32-bit channel) DFI-based interface may connect DRAM PHY 315 to N memory controllers 320a-320N. In some cases, DRAM PHY 315 may include or be connected to a through silicon via (TSV) landing area. The TSV landing area may include an area dedicated to TSV connections that connect base die 300 to a core memory die layer stacked on top of base die 300.
[0086] In some examples, the N memory controllers 320a-320N can control one or more aspects of memory associated with the base die 300. For example, one or more core memory dies can be stacked on top of the base die 300. The N memory controllers 320a-320N can control one or more aspects of the one or more core memory dies stacked on top of the base die 300. The N memory controllers 320a-320N can be configured to act as a bridge between processors (e.g., the N PUs 330a-330N) and memory (e.g., the one or more core memory dies stacked on top of the base die 300). The N memory controllers 320a-320N can manage data flow between memories / processors, handle read and write operations associated with the memories of the base die 300 and the N PUs 330a-330N, manage data integrity, and coordinate memory accesses, thereby controlling how data is transferred to and from the memories of the base die 300 and the N PUs 330a-330N.
[0087] In some examples, a NOC interconnect 325 (e.g., a system-level bus) can connect the N memory controllers 320a-320N to the N PUs 330a-330N. As shown, the NOC interconnect 325 can connect the N PUs 330a-330N to a system bus interface 335. In some cases, the NOC interconnect can communicate data via data packets (e.g., L2 data packets).
[0088] In some examples, N PUs 330a-330N (e.g., AI accelerators, processor elements (PEs), NPUs, TPUs, GPUs, FPGAs, ASICs, etc.) can be configured to execute tasks on base die 300. The number N of PUs 330a-330N can be based on the amount of space available on base die 300, the size of the PUs, the nanometer process used to manufacture the PUs, etc. In some cases, base die 300 can include one memory controller for each PU. In some cases, base die 300 can include more or fewer memory controllers than PUs. The systems and methods described herein can be based on and / or can include Processing-in-Memory (PIM) HBM. PIM-HBM can include memory technology that integrates processors into memory (e.g., PUs on base die 300), which can reduce data movement between the processor and memory. In some cases, at least one of the N PUs 330a-330N may include a tensor core for matrix multiplication, an arithmetic logic unit (ALU) for integer computations, a floating point unit (FPU) for floating point computations, and / or an accumulator (e.g., a register for storing intermediate logic or arithmetic data for multi-step computations).
[0089] Incorporating processing into the base die 300 provides increased processing power in the HBM memory device based on a hybrid computing architecture that includes processing from one or more compute dies connected to the base die 300 and processing from the N PUs 330a-330N on the base die 300. The N PUs 330a-330N may include AI accelerators (e.g., NPUs, GPUs, TPUs, IPUs, etc.). In some cases, the N PUs 330a-330N may include tensor cores. Tensor cores may include specialized processing subunits that accelerate the performance of AI accelerators. Tensor cores may be designed to maintain accuracy while accelerating performance through mixed-precision computations, fused multiply-add algorithms, matrix multiplication, accumulator functions, and the like.
[0090] In some examples, the system bus interface 335 may connect the NOC interconnect 325 to the D2D 340. In some cases, the system bus interface 335 may map data in a format used by the NOC interconnect 325 to a format used by the D2D 340, and / or map data in a format used by the D2D 340 to a format used by the NOC interconnect 325. In some cases, the D2D 340 may communicate data based on flow control units or flow control digits (flits). A flit may comprise the link-level atomic elements that form a network packet or flow. A packet may be decomposed into one or more flits, including a header flit, a body flit, and, in some cases, a trailer flit. The NOC interconnect 325 may communicate data based on packets (e.g., L2 packets). Accordingly, the system bus interface 335 may convert the format of the D2D interface (e.g., flits) to the format of the NOC interconnect (e.g., L2 packets).
[0091] In some examples, D2D 340 can be based on Universal Chiplet Interconnect Express (UCIe). UCIe can provide die-to-die connectivity in multi-die systems. UCIe can define the physical layer, protocol stack, software model, and procedures for compliance testing. In some cases, D2D 340 can connect base die 300 to components of a memory system (e.g., an HBM system). For example, D2D 340 can connect base die 300 to a compute die of a memory system.
[0092] Based on the described systems and methods, the base die 300 can improve processing time, increase memory bandwidth, and reduce data transfer latency associated with processing executed on the base die 300. Processing on the base die 300 can include kernel execution, direct memory access (DMA)-based data movement, and AI processing, including: activation functions (e.g., Sigmoid Weighted Linear Unit (SWIGLU), Gaussian Error Linear Unit (GELU), Rectified Linear Unit (ReLU), etc.); Softmax calculation; Large Language Model (LLM), etc.
[0093] Figure 4 An example package 400 is shown according to one or more embodiments as described herein. In some configurations, one or more aspects of package 400 can be implemented by or in conjunction with machine 105, components of machine 105, or any combination thereof. In some examples, package 400 can include a system-in-package (SiP). For example, package 400 can include a 2.5D SiP and / or a 3D SiP.
[0094] As shown, package 400 may include L memory dies 405a-405L (e.g., L stacked memory dies), M compute dies 410a-410M, and L D2D interconnects 415a-415L, where L is a positive integer (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 10, 12, 16, 20, 24, 30, 32, 64, etc.) and M is a positive integer (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 16, 20, etc.). A given memory die (e.g., memory die 405) may include one or more memory core dies. In the example shown, package 400 may include two compute dies, 12 memory dies (e.g., an HBM stack), and 12 D2D interconnects. In some examples, package 400 may include an interposer 420. Interposer 420 may be a layer of package 400. In some cases, L memory dies 405 a - 405L, M compute dies 410 a - 410M, and L D2D interconnects 415 a - 415L can be placed on top of interposer 420 .
[0095] In some examples, at least one of the L memory dies 405a-405L may include a memory base die (e.g., base die 300), with one or more core memory dies stacked on top of the memory base die. For example, memory die 405a may include a first base die similar to base die 300, memory die 405b may include a second base die similar to base die 300, and so on. In some cases, at least one of the L memory dies 405a-405L may include an HBM base die, with one or more DRAM core memory dies stacked on top of the HBM base die. For example, memory die 405a may include a first HBM base die, with one or more DRAM core memory dies stacked on top of the first HBM base die, memory die 405b may include a second HBM base die, with one or more DRAM core memory dies stacked on top of the second HBM base die, and so on.
[0096] In some examples, at least one of the M compute dies 410a-410M can include an AI accelerator (e.g., a GPU, an NPU, a TPU, etc.). The M compute dies 410a-410M can be configured to perform AI processing. In some cases, the M compute dies 410a-410M can perform a first portion of the AI processing, and the processing units on the base die of the L memory dies 405a-405L can perform a second portion of the AI processing.
[0097] In some examples, a D2D interconnect can connect a compute die of package 400 to a memory die of package 400 (e.g., via interposer 420). For example, D2D interconnect 415a can connect compute die 410a to memory die 405a. Similarly, D2D interconnect 415b can connect compute die 410b to memory die 405b, and so on.
[0098] Interposer 420 can include a silicon interposer, a redistribution layer (RDL) interposer, and / or an organic interposer. Interposer 420 can include an electrical interface for routing between one socket or connection to another (e.g., connecting a memory die to a compute die). Interposer 420 can be implemented in conjunction with a ball grid array (BGA) package, a D2D interface, through silicon vias (TSVs), and the like.
[0099] In some cases, at least one base die among the L memory dies 405a-405L may include a processing unit (e.g., an AI accelerator, a GPU, a TPU, an NPU, etc.). In some cases, at least one of the processing units of the base die may include a tensor core for matrix multiplication, an arithmetic logic unit (ALU) for integer calculations, a floating point unit (FPU) for floating point calculations, and / or an accumulator for storing intermediate logic or arithmetic data for multi-step calculations.
[0100] One or more processing units on a first base die of the L memory dies 405a-405L can perform machine learning and / or artificial intelligence (ML / AI) kernel operations independently of one or more processing units on a second base die of the L memory dies 405a-405L. In some cases, one or more processing units on the base die of the L memory dies 405a-405L can perform memory-bound AI tasks, while at least one of the M compute dies 410a-410M can perform compute-bound AI tasks. For example, one or more processing units on the base die of the L memory dies 405a-405L can perform LLM inference decoding for token generation and can perform at least a portion of kernel operations, such as query key value (QKV) calculation in self-attention. In some examples, a portion of the weight data (e.g., a majority of the weight data) can be consumed by a processing unit on at least one base die of the L memory dies 405a-405L without transmitting the data to the compute die via the intermediary layer 420, which significantly reduces power consumption from the memory I / O and communication via the intermediary layer 420.
[0101] Implementing the processing unit on at least one of the base dies of the L memory dies 405a-405L significantly reduces data latency between the compute die and the memory. Some data paths for AI computation (e.g., 2.5D data paths) can include the following sequence: memory core die -> TSV PHY -> I / O PHY -> interposer -> memory controller PHY -> memory controller -> NOC interconnect / L2 cache -> compute die. Implementing the processing unit on at least one of the base dies of the L memory dies 405a-405L can reduce this data path. For example, based on the described systems and methods, an enhanced data path can include the following sequence: memory core die -> TSV PHY -> memory controller -> NOC interconnect / L2 cache -> compute node, which avoids the I / O PHY, interposer, and memory controller PHY steps from other data paths. Because the enhanced data path is maintained with a given memory die (eg, one of the L memory dies 405a-405L), the physical routing travel distance for data movement is shortened.
[0102] Thus, implementing a processing unit on at least one base die of the L memory dies 405a-405L provides a memory local compute node that can access memory with relatively short latency and low power consumption (e.g., compared to the computation of the M compute dies 410a-410M). Figure 5 Additional details regarding the memory base die of the L memory dies 405a-405L and incorporation of processing units on the memory base die are discussed.
[0103] Figure 5 An exemplary base die 500 is shown according to one or more embodiments described herein. In some configurations, one or more aspects of the base die 500 can be implemented by or in conjunction with the machine 105, a component of the machine 105, or any combination thereof. In some examples, the base die 500 can be a base layer for a stacked memory module (e.g., an HBM memory module, an HBM chip). 2.5D and / or 3D stacking techniques can be used to stack memory chips on the base die 500. Memory stacked on the base die 500 can be shared by one or more components of the base die 500. In some cases, the base die 500 can depict Figure 4 An exemplary base die for one of the L memory dies 405a-405L. In some cases, the base die 500 may be Figure 3 An example of a base die 300 .
[0104] As shown in the figure, the base die 500 may include: one or more DDR PHYs (e.g., DDR PHY 505a, DDR PHY 505b, DDR PHY 505c, DDR PHY 505d); one or more DDR controllers (e.g., DDR controller 510a, DDR controller 510b); one or more processing units (e.g., N PUs 515a-515N, where N is a positive integer such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 16, 20, 24, 30, 32, 64, etc.); test logic 520, DRAM PHY 525, memory controller 530, at least one base die cache (e.g., shared memory, such as cache 535a, cache 535b, cache 535c, cache 535d, etc.); at least one direct memory access (DMA) controller 540; one or more D2D controllers (e.g., D2D controller 545a, D2D controller 545b, D2D controller 545c, D2D controller 545d); and one or more D2D PHYs (e.g., D2D PHY 550a, D2D PHY 550b, D2D PHY 550c, D2D PHY 550d).
[0105] In some examples, one or more DDR PHYs (e.g., DDR PHY 505a, etc.) can provide a physical interface for DDR memory devices associated with the base die 500. In some cases, the DDR PHY can be associated with and / or configured as one or more memory expansion ports (e.g., expansion port 305). For example, at least one of the one or more DDR PHYs can provide a physical interface to a memory external to the base die 500. For example, a memory module can include the base die 500 and one or more layers of memory stacked on top of the base die 500. Additionally, one or more external memory modules can be connected to the base die 500 via one or more DDR PHYs. For example, at least one of the one or more DDR PHYs can connect the base die 500 to an off-die memory (e.g., LPDDR, GDDR, etc.).
[0106] In some examples, one or more DDR controllers (eg, DDR controller 510a, etc.) can be configured to control one or more aspects of the base die 500 and one or more external memory modules connected to the base die 500 via one or more DDR PHYs.
[0107] In some examples, one or more PUs (e.g., N PUs 515a-515N) can be configured to process tasks on base die 500 (e.g., tasks for AI models, tasks for LLMs, etc.). When a compute die accesses data on base die 500, the data can be moved from memory stacked on base die 500 to the compute die via memory controller 530 and one or more D2D PHYs (e.g., D2D PHY 550a). However, moving data from base die 500 to a compute die external to base die 500 can use a relatively large amount of bandwidth and have relatively high latency. However, by incorporating compute resources (e.g., N PUs 515a-515N) within a given base die (e.g., base die 500), the amount of bandwidth is minimized or significantly reduced because data is provided from memory on the base die to the PUs on the base die (e.g., the storage, movement, and processing of the data remain on the base die).
[0108] The test logic 520 can be an example of the test logic 310. In some examples, the test logic 520 can provide tests on the base die 500. In some cases, the test logic 520 can include a built-in self test (BIST), which can enable the base die 500 to test the components of the base die 500 without the need for external test equipment. In some cases, the test logic 520 can generate a test pattern and apply the test pattern to the components of the base die 500, perform the test, and analyze the results to determine whether the component operates as expected or whether the test indicates an anomaly. In some cases, one or more aspects of the test logic 520 can be based on IEEE 1500, which can include standards that define how to test the design and / or operation of the base die 500.
[0109] In some examples, the DRAM PHY 525 can provide a physical interface for a DRAM memory device associated with the base die 500. In some cases, the DRAM PHY 525 can provide a physical interface for memory stacked on top of the base die 500. For example, the DRAM PHY 525 can be configured as a 3D DRAM PHY. In some cases, the DRAM PHY 525 can include a TSV landing area. The TSV landing area can include an area of the DRAM PHY 525 dedicated to TSV connections or an area adjacent to the DRAM PHY 525, where the TSV connections connect the base die 500 to the core memory die layer stacked on top of the base die 500.
[0110] In some examples, the memory controller 530 can be configured to control one or more aspects of the base die 500 and / or the memory stacked on top of the base die 500. For example, the memory controller 530 can control reads and / or writes associated with the memory stacked on top of the base die 500. In some cases, the memory controller 530 can control tasks associated with AI processing and the processing performed by the N PUs 515a-515N. In some examples, the memory controller 530 can include or be configured as an HBM controller. In some systems, the HBM controller can be located on the compute die (e.g., compute die 410a). However, based on the systems and methods described herein, the HBM controller (e.g., memory controller 530) can be located on the base die (e.g., base die 500) of the HBM stack. The benefit or advantage of placing the HBM controller on the base die is that it frees up area on the compute die to add additional compute resources and / or memory / cache resources without increasing the size of the compute die.
[0111] In some examples, the base die 500 may include at least one base die cache (e.g., cache 535a, etc.). In some cases, the at least one base die cache may be shared by one or more components of the base die 500. The processing of tasks by the N PUs 515a-515N may be based on the at least one base die cache. For example, the memory controller 530 may provide data from a memory stacked on the base die 500 to the at least one base die cache. While the N PUs 515a-515N are processing data, the at least one base die cache may store the data. For example, the at least one base die cache may store data being processed by the N PUs 515a-515N. The at least one base die cache can store frequently accessed instructions and data from the memory stacked on the base die 500, thereby achieving faster retrieval and improving the overall processing speed based on the speed of the at least one base die cache, but also because the at least one base die cache is located on the base die 500, the data being processed can be retained on the base die 500 and / or the memory stacked on the base die 500.
[0112] In some examples, the DMA controller 540 may include a dedicated hardware component that enables the base die 500 to transfer data directly to and / or from memory without intervention from a processor external to the base die 500 (e.g., without intervention from a host processor, a CPU, the compute die 410a, etc.). Thus, by offloading data transfer tasks from an external processor, the DMA controller 540 improves system performance by minimizing external processor overhead. The DMA controller can manage the data transfer process between peripheral devices and memory by providing the necessary address and control signals to directly access memory.
[0113] In some examples, the DMA controller 540 can operate in conjunction with the N PUs 515a-515N and / or can be at least partially implemented in the N PUs 515a-515N. In some cases, the DMA controller 540, in conjunction with the N PUs 515a-515N, can copy data from memory external to the base die 500 to memory stacked on top of the base die 500 and / or to at least one base die cache (e.g., cache 535a), where the memory external to the base die 500 can include LPDDR and / or GDDR connected to the base die 500 via one or more DDR PHYs (e.g., DDR PHY 505a). In some cases, the DMA controller 540, in conjunction with the N PUs 515a-515N, can copy data from memory stacked on top of the base die 500 to at least one base die cache and / or to memory external to the base die 500. The DMA controller 540 may directly handle data movement between external memory and memory of the base die 500 (eg, memory stacked on the base die 500 and / or a base die cache).
[0114] In some examples, one or more D2D controllers (e.g., D2D controller 545a, etc.) can be configured to control one or more aspects of operations associated with the base die 500 and / or one or more devices external to the base die 500. For example, the one or more D2D controllers can control one or more aspects (e.g., read commands, write commands, processing instructions) associated with one or more processors external to the base die 500 (e.g., compute die 410a). In some cases, the one or more D2D controllers can receive system traffic and identify whether it includes a compute request (e.g., processing by the N PUs 515a-515N) or a memory request (e.g., reading from / writing to memory stacked on the base die 500). When the one or more D2D controllers identify a memory request, the one or more D2D controllers can route the memory request to the memory controller 530. When the one or more D2D controllers identify a compute request, the one or more D2D controllers can route the compute request to the N PUs 515a-515N.
[0115] In some examples, one or more D2D PHYs (e.g., D2D PHY 550a, etc.) can provide a physical interface to one or more devices external to the base die 500. For example, one or more processors external to the base die 500 (e.g., compute die 410a) can be physically connected to the base die 500 via one or more D2D PHYs. In some cases, the compute die (e.g., compute die 410a) can transmit system traffic to the base die 500 via the one or more D2D PHYs. In some cases, the base die 500 can transmit system traffic to the compute die via the one or more D2D PHYs.
[0116] In some examples, the base die 500 can receive kernel operation instructions (e.g., tasks for AI processing, LLM processing) from a compute die (e.g., compute die 410a). In some cases, data can be copied or moved from a memory stacked on the base die 500 (e.g., a DRAM core die). The base die 500 can distribute the data to one or more of the N PUs 515a-515N. The one or more PUs can process the data on the base die 500, thereby avoiding the delay of having a compute die outside the base die 500 process the data. For example, the N PUs 515a-515N can perform AI processing (e.g., matrix multiplication, Softmax, etc.), thereby avoiding the computational cost of transferring data to an external compute die for processing.
[0117] In some cases, the N PUs 515a-515N on the base die 500 can enable multiple multi-layer perceptron (MLP) layers. The MLP layer can include an artificial neural network composed of multiple layers of neurons, including an input layer, one or more intermediate layers, an output layer, etc., where each neuron can be connected to the next layer, thereby allowing the MLP to learn complex nonlinear relationships between input and output data. In some cases, the N PUs 515a-515N can be logically subdivided to perform assigned tasks. For example, a first PU among the N PUs 515a-515N may be assigned to perform physical processing (e.g., processing physics-based queries); a second PU among the N PUs 515a-515N may be assigned to perform social processing (e.g., processing social-based queries); a third PU among the N PUs 515a-515N may be assigned to perform historical processing (e.g., processing history-based queries); a fourth PU among the N PUs 515a-515N may be assigned to perform creative processing (e.g., processing creative writing-based queries); a fifth PU among the N PUs 515a-515N may be assigned to perform encoding processing (e.g., processing encoding-based queries); a sixth PU among the N PUs 515a-515N may be assigned to perform translation processing (e.g., processing translation-based queries), and so on.
[0118] In some cases, one or more components of the base die 500 (e.g., memory controller 530, D2D controller, and / or at least one PU) may classify queries. Additionally or alternatively, a host may classify queries and provide the query and class to the base die 500. When a given system classifies a query, the query may be routed to a tier associated with that query class. For example, the query may be routed to at least one of the N PUs 515a-515N, to a compute die external to the base die 500 (e.g., compute die 410a), or to an external processor and external memory (e.g., LPDDR, GDDR, etc.) connected to the base die 500 via one or more DDR PHYs (e.g., DDR PHY 505a). A compute-intensive class may be routed to a compute die (e.g., a compute die 410 with higher processing power but higher latency and higher power consumption). A less frequently used class may be routed to an external processor and external memory (e.g., the slowest processing, slowest memory). Frequently used classes may be routed to one or more of the N PUs 515a-515N (eg, with the lowest latency for relatively fast processing). Thus, incorporating computing resources in the base die improves system performance and significantly reduces power usage.
[0119] In some examples, tasks may be routed to the base die 500 and to compute resources external to the base die 500 (e.g., a compute die, LPDDR, GDDR) based on whether the task is classified as a memory-bound task or a compute-bound task. In this disclosure, for example, memory-bound tasks and compute-bound tasks may be referred to as memory-bound operations and compute-bound operations, respectively. Memory-bound operations may include compute tasks where the total execution time is primarily determined by the time it takes to access data from memory, where the speed of the operation may be limited by the memory bandwidth rather than the processing power of the processing unit. Compute-bound operations may include compute tasks where the time to complete is primarily determined by the speed of the processing unit, where the operation may rely heavily on complex computations and data processing rather than waiting for memory, storage, network requests, etc. In some cases, one or more components of the base die 500 (e.g., the memory controller 530, the D2D controller, and / or at least one PU) may classify a task as compute-bound and / or memory-bound. Additionally or alternatively, the host may classify the task as a compute-bound and / or memory-bound task and communicate the classification to the base die 500 .
[0120] In some examples, memory-limited tasks can be routed to the base die 500 for processing (e.g., via the N PUs 515a-515N), and compute-limited tasks can be routed to the compute die for processing. Routing memory-limited kernel operations to the base die 500 can help other kernels running on the compute die focus on compute-intensive workloads.
[0121] In some cases, LLM algorithms may be memory-bound operations rather than compute-bound operations. Memory-bound operations may include computational tasks where the total execution time is primarily determined by the time it takes to access data from memory, where the speed of the operation may be limited by the memory bandwidth rather than the processing power of the processing unit. Compute-bound operations may include computational tasks where the time taken to complete is primarily determined by the speed of the processing unit, where the operation may rely heavily on complex calculations and data processing rather than waiting for memory, storage, network requests, etc.
[0122] The described systems and methods can be configured for memory-limited operations (e.g., inference, decoding). In some cases, memory-limited operations can be routed to the HBM base die, while compute-limited operations can be routed to one or more compute dies. Operations associated with inference, decoding, etc. can be memory-limited operations and, therefore, can be routed to the base die 500 (e.g., for processing by at least one of the N PUs 515a-515N). Operations associated with training to create AI models, LLM models, etc. can be compute-limited operations and, therefore, can be routed to a compute die external to the base die 500.
[0123] Figure 6 A flowchart is depicted illustrating an example method 600 associated with the disclosed system according to example embodiments described herein. The method 600 may include a method for performing processing in memory of a memory base die. In some configurations, one or more aspects of the method 600 may be implemented by or in conjunction with the machine 105, components of the machine 105, or any combination thereof. The depicted method 600 is only one implementation, and one or more operations of the method 600 may be rearranged, reordered, omitted, and / or otherwise modified, such that other implementations are possible and contemplated.
[0124] At 605, method 600 may include stacking one or more memory dies or memory core dies on a memory base die. For example, method 600 may include a memory die manufacturing system stacking one or more memory core dies on a memory base die. The memory base die may be located on a package that includes a compute die.
[0125] At 610, method 600 may include connecting a processing unit on a memory base die to a memory controller on the memory base die. For example, method 600 may include manufacturing a system to connect a processing unit on a memory base die to a memory controller on the memory base die.
[0126] At 615, method 600 may include processing data received by the memory controller. For example, method 600 may include a processing unit processing data received by the memory controller. The memory controller may receive data from at least one of the one or more memory core dies.
[0127] Figure 7A flowchart is depicted illustrating an example method 700 associated with the disclosed system according to example embodiments described herein. The method 700 may include a method for performing processing in memory of a memory base die. In some configurations, one or more aspects of the method 700 may be implemented by or in conjunction with the machine 105, components of the machine 105, or any combination thereof. The depicted method 700 is only one implementation, and one or more operations of the method 700 may be rearranged, reordered, omitted, and / or otherwise modified, such that other implementations are possible and contemplated.
[0128] At 705, method 700 may include stacking one or more memory core dies on a memory base die. For example, method 700 may include a memory die manufacturing system stacking one or more memory core dies on a memory base die. The memory base die may be located on a package that includes a compute die.
[0129] At 710, method 700 may include connecting a processing unit on a memory base die to a memory controller on the memory base die. For example, method 700 may include manufacturing a system to connect a processing unit on a memory base die to a memory controller on the memory base die.
[0130] At 715, method 700 may include routing the function to the memory foundation die for processing by the processing unit. For example, method 700 may include routing the function to the memory foundation die for processing by the processing unit based on one or more aspects of the function. For example, based on a determination (e.g., by a host, by machine 105, etc.) that the first function is a memory-constrained function, the first function may be routed to the memory foundation die for processing by at least one processing unit on the memory foundation die. Based on a determination (e.g., by a host, by machine 105, etc.) that the second function is a compute-constrained operation, the second function may be routed to the compute die for processing by the compute die. The compute die (e.g., compute die 410a) may be connected to the memory foundation die (e.g., foundation die 300, foundation die 500) via a packaged silicon interposer. In some cases, the host (e.g., an operating system, an application, machine 105) may route the function to the memory foundation die. In some cases, a compute die on a package that includes the memory foundation die may route the function to the memory foundation die. In some cases, the memory controller of the memory base die may route functions to the memory base die.
[0131] At 720, method 700 may include processing data received by the memory controller. For example, method 700 may include processing data received by the memory controller by a processing unit. The memory controller may receive data from at least one of the one or more memory core dies. In some cases, processing the data may be based on routing the function to the memory base die at 715.
[0132] Figure 8 A flowchart is depicted illustrating an example method 800 associated with the disclosed system according to example embodiments described herein. The method 800 may include a method for performing processing in memory of a memory base die. In some configurations, one or more aspects of the method 800 may be implemented by or in conjunction with the machine 105, components of the machine 105, or any combination thereof. The depicted method 800 is only one implementation, and one or more operations of the method 800 may be rearranged, reordered, omitted, and / or otherwise modified, such that other implementations are possible and contemplated.
[0133] At 805, method 800 may include determining at least one characteristic of a data query. For example, method 800 may include a system (e.g., machine 105, a system based on base die 300, a system based on package 400, a memory base die having one or more processing units, etc.) determining at least one characteristic of the data query. In some cases, method 800 may include the system determining whether the data query includes or is based on a compute-limited function or a memory-limited function.
[0134] At 810, method 800 may include routing a function to a memory foundation die for processing by a processing unit of the memory foundation die. For example, method 800 may include routing the function to the memory foundation die for processing by the processing unit based on determining at least one characteristic of the data query. For example, based on a determination (e.g., by a host, by machine 105, etc.) that the function of the data query is a memory-constrained function, the function of the data query may be routed to the memory foundation die for processing by at least one processing unit on the memory foundation die. In some cases, the function may be routed to a compute die for processing by the compute die based on a determination that the function is a compute-constrained operation. A compute die (e.g., compute die 410a) may be connected to a memory foundation die (e.g., foundation die 300, foundation die 500) via a packaged silicon interposer. In some cases, a host (e.g., an operating system, an application, machine 105) may route the function to the memory foundation die. In some cases, a compute die on a package that includes the memory foundation die may route the function to the memory foundation die. In some cases, the memory controller of the memory base die may route functions to the memory base die.
[0135] At 815, method 800 may include processing data received by the memory controller. For example, method 800 may include processing data received by the memory controller by a processing unit. The memory controller may receive data from at least one of the one or more memory core dies. In some cases, processing the data may be based on routing the function to the memory base die at 810.
[0136] Figure 9 A flowchart is depicted illustrating an example method 900 associated with the disclosed system according to example embodiments described herein. Method 900 may include a method for performing processing in memory of a memory base die. In some configurations, one or more aspects of method 900 may be implemented by or in conjunction with machine 105, components of machine 105, or any combination thereof. The depicted method 900 is only one implementation, and one or more operations of method 900 may be rearranged, reordered, omitted, and / or otherwise modified, such that other implementations are possible and contemplated.
[0137] At 905, method 900 may include receiving a data query. For example, method 900 may include a system (e.g., machine 105, a system based on base die 300, a system based on package 400, a memory base die having one or more processing units, etc.) receiving a data query associated with data processing or data access (e.g., AI processing, ML processing, LLM processing, etc.).
[0138] At 910, method 900 may include determining at least one characteristic of the data query. For example, method 900 may include the system determining at least one characteristic of the data query. For example, method 900 may include the system determining whether the data query includes or is based on a compute-limited function or a memory-limited function.
[0139] At 915, method 900 may include routing the function to the memory foundation die for processing by a processing unit of the memory foundation die. For example, method 900 may include routing the function to the memory foundation die for processing by the processing unit based on determining at least one characteristic of the data query. For example, based on a determination (e.g., by a host, by machine 105, etc.) that the function of the data query is a memory-constrained function, the function of the data query may be routed to the memory foundation die for processing by at least one processing unit on the memory foundation die. In some cases, the function may be routed to the compute die for processing by the compute die based on a determination that the function is a compute-constrained operation. The compute die (e.g., compute die 410a) may be connected to the memory foundation die (e.g., foundation die 300, foundation die 500) via a packaged silicon interposer. In some cases, the host (e.g., an operating system, an application, machine 105) may route the function to the memory foundation die. In some cases, a compute die on a package that includes the memory foundation die may route the function to the memory foundation die. In some cases, the memory controller of the memory base die may route functions to the memory base die.
[0140] At 920, method 900 may include processing data received by the memory controller. For example, method 900 may include processing data received by the memory controller by a processing unit. The memory controller may receive data from at least one of the one or more memory core dies. In some cases, processing the data may be based on routing the function to the memory base die at 915.
[0141] In the examples described herein, the configurations and operations are example configurations and operations and may involve various additional configurations and operations not explicitly shown. In some examples, one or more aspects of the configurations and / or operations shown may be omitted. In some embodiments, one or more of the operations may be performed by components other than those shown herein. Additionally or alternatively, the order and / or chronological order of the operations may be changed.
[0142] Certain embodiments may be implemented in one or a combination of hardware, firmware, and software. Other embodiments may be implemented as instructions stored on a computer-readable storage device, which may be read and executed by at least one processor to perform the operations described herein. A computer-readable storage device may include any non-transitory memory mechanism for storing information in a form readable by a machine (e.g., a computer). For example, a computer-readable storage device may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, and other storage devices and media.
[0143] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. As used herein, the terms "computing device," "user device," "communication station," "station," "handheld device," "mobile device," "wireless device," and "user equipment (UE)" refer to a wired and / or wireless communication device, such as a switch, router, network interface controller, cellular phone, smartphone, tablet, netbook, wireless terminal, laptop, femtocell, high data rate (HDR) subscriber station, access point, printer, point of sale device, access terminal, or other personal communication system (PCS) device. The device can be wireless, wired, mobile, and / or stationary.
[0144] As used within this document, the term "communication" is intended to include sending, or receiving, or both sending and receiving. Similarly, a two-way exchange of data between two devices (both devices sending and receiving during the exchange) may be described as "communication" when only the functionality of one of those devices is claimed. The term "communication" as used herein with respect to wired and / or wireless communication signals includes sending wired and / or wireless communication signals and / or receiving wired and / or wireless communication signals. For example, a communication unit capable of communicating wired and / or wireless communication signals may include a wired / wireless transmitter that sends communication signals to at least one other communication unit, and / or a wired / wireless communication receiver that receives communication signals from at least one other communication unit.
[0145] Some embodiments may be used in conjunction with various devices and systems, such as personal computers (PCs), desktop computers, mobile computers, laptop computers, notebook computers, tablet computers, server computers, handheld computers, handheld devices, personal digital assistant (PDA) devices, handheld PDA devices, in-vehicle devices, off-vehicle devices, hybrid devices, in-vehicle devices, non-in-vehicle devices, mobile or portable devices, consumer devices, non-mobile or non-portable devices, wireless communication stations, wireless communication devices, wireless access points (APs), wired or wireless routers, wired or wireless modems, video devices, audio devices, audio-video (A / V) devices, wired or wireless networks, wireless area networks, wireless video area networks (WVANs), local area networks (LANs), wireless LANs (WLANs), personal area networks (PANs), wireless PANs (WPANs), and the like.
[0146] Some embodiments may be used in conjunction with one-way and / or two-way radio communication systems, cellular radiotelephone communication systems, mobile phones, cellular phones, cordless phones, personal communication system (PCS) devices, PDA devices containing wireless communication devices, mobile or portable global positioning system (GPS) devices, devices containing GPS receivers or transceivers or chips, devices containing RFID elements or chips, multiple-input multiple-output (MIMO) transceivers or devices, single-input multiple-output (SIMO) transceivers or devices, multiple-input single-output (MISO) transceivers or devices, devices with one or more internal antennas and / or external antennas, digital video broadcasting (DVB) devices or systems, multi-standard radio devices or systems, wired or wireless handheld devices (e.g., smart phones), wireless application protocol (WAP) devices, and the like.
[0147] Some embodiments may be used in conjunction with one or more types of wireless communication signals and / or systems that adhere to one or more wireless communication protocols, such as radio frequency (RF), infrared (IR), frequency division multiplexing (FDM), orthogonal FDM (OFDM), time division multiplexing (TDM), time division multiple access (TDMA), extended TDMA (E-TDMA), general packet radio service (GPRS), extended GPRS, code division multiple access (CDMA), wideband CDMA (WCDMA), CDMA 2000, single carrier CDMA, multi-carrier CDMA, multi-carrier modulation (MDM), discrete multi-tone (DMT), Bluetooth. TM , Global Positioning System (GPS), Wi-Fi, Wi-Max, ZigBee TM , Ultra-Wideband (UWB), Global System for Mobile Communications (GSM), 2G, 2.5G, 3G, 3.5G, 4G, fifth generation (5G) mobile networks, 3GPP, Long Term Evolution (LTE), LTE-Advanced, Enhanced Data Rates for GSM Evolution (EDGE), etc. Other embodiments may be used in various other devices, systems and / or networks.
[0148] Although example processing systems have been described above, embodiments of the subject matter and functional operations described herein may be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them.
[0149] Embodiments of the subject matter and operations described herein can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of these. Embodiments of the subject matter described herein can be implemented as one or more computer programs, i.e., one or more components of computer program instructions, encoded on a computer storage medium for execution by, or control of the operation of, an information / data processing apparatus. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, generated to encode information / data for transmission to a suitable receiver device for execution by the information / data processing apparatus. A computer storage medium can be or be included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of these. Furthermore, while a computer storage medium is not a propagated signal, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially generated propagated signal. The computer storage medium can also be, or be included in, one or more separate physical components or media (eg, multiple CDs, disks, or other storage devices).
[0150] The operations described herein may be implemented as operations performed by an information / data processing apparatus on information / data stored on one or more computer-readable storage devices or received from other sources.
[0151] The term "data processing apparatus" encompasses all types of devices, equipment, and machines for processing data, including, for example, a programmable processor, a computer, a system-on-chip, or multiple or a combination of the foregoing. The apparatus may include specialized logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of these. The apparatus and execution environment can implement a variety of different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructure.
[0152] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a stand-alone program or as a component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program may be stored as part of a file that holds other programs or information / data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more components, subroutines, or code portions). A computer program may be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communications network.
[0153] The processes and logic flows described herein can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input information / data and generating output. Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and information / data from read-only memory or random access memory, or both. The essential elements of a computer are a processor for performing actions according to instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic, magneto-optical, or optical disks) for storing data, or be operatively coupled to receive information / data from or transfer information / data to one or more mass storage devices, or both. However, a computer need not have such devices. Devices suitable for storing computer program instructions and information / data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard drives or removable disks; magneto-optical disks; and CDROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0154] To provide for interaction with a user, embodiments of the subject matter described herein may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information / data to the user, and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices may also be used to provide for interaction with the user; for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including sound, voice, or tactile input. Furthermore, a computer may interact with a user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser on a user's client device in response to a request received from the web browser.
[0155] Embodiments of the subject matter described herein can be implemented in a computing system that includes a back-end component (e.g., as an information / data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer with a graphical user interface or web browser through which a user can interact with embodiments of the subject matter described herein), or includes any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital information / data communication (e.g., a communication network). Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[0156] A computing system may include a client and a server. The client and server are typically remote from each other and typically interact via a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship with each other. In some embodiments, the server sends information / data (e.g., an HTML page) to a client device (e.g., for the purpose of displaying information / data to a user interacting with the client device and receiving user input from the user interacting with the client device). Information / data generated at the client device (e.g., results of user interactions) may be received at the server from the client device.
[0157] Although this specification contains many specific embodiment details, these should not be interpreted as limitations on the scope of any embodiment or content that may be claimed, but rather as descriptions of features that are peculiar to a particular embodiment. Certain features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable sub-combination. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed as such, one or more features from the claimed combination may in some cases be deleted from the combination, and the claimed combination may be directed to a sub-combination or a variation of the sub-combination.
[0158] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that the operations be performed in the particular order shown, or in sequence, or that all illustrated operations be performed, in order to achieve the desired results. In certain circumstances, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0159] Thus, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain embodiments, multitasking and parallel processing may be advantageous.
[0160] Many modifications and other examples as set forth herein will come to mind to one skilled in the art having the benefit of the teachings presented in the foregoing description and the associated drawings. Therefore, it should be understood that the embodiments are not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
1. A method for processing in a memory, the method comprising: determining at least one characteristic of the data query; routing a first function of the data query to a memory base die for processing by a processing unit on the memory base die based on the at least one characteristic of the data query; as well as Data received by a memory controller on the memory base die from at least one of one or more memory dies stacked on top of the memory base die is processed via the processing unit. 2 . The method of claim 1 , further comprising routing the first function to the memory base die for processing by the processing unit on the memory base die based on determining that the first function is a memory-constrained function.
3. The method according to claim 1 further includes routing the second function of the data query to a computing die for processing by the computing die based on determining that the second function is a computing-limited operation, wherein the computing die is connected to the memory base die via a silicon interposer of a system-level package including the computing die and the memory base die.
4. The method according to claim 1, wherein The memory base die includes a memory expansion port connected to at least one of a low power double data rate memory or a graphics double data rate memory external to the memory base die.
5. The method according to claim 4, further comprising at least one of the following: routing a first category of functions to the memory base die via the memory controller for processing by the processing unit on the memory base die, routing the second category of functions to the compute die via the memory controller for processing by the compute die, or A third category of functionality is routed via the memory controller to at least one of the low power double data rate memory or the graphics double data rate memory external to the memory base die.
6. The method according to claim 1, further comprising: transferring the data from the one or more memory dies to a physical layer interface of the memory base die through a through silicon via; transferring the data from the physical layer interface to the memory controller of the memory base die; as well as The data is transferred from the memory controller to a shared memory on the memory base die, wherein the shared memory holds the data for the processing unit to process the data.
7. The method according to claim 1, wherein The processing unit includes at least one of the following: Tensor Cores, configured for matrix multiplication, or An accumulator is configured to accumulate intermediate calculations.
8. The method according to claim 1, wherein: The memory controller is connected to the processing unit via a network on chip (NOC) interconnect bus, and The memory controller is connected to a dynamic random access memory DRAM physical layer on a memory base die via a double data rate DDR physical layer interface of the memory base die.
9. The method according to claim 8, wherein: a system bus interface connected to the die-to-die interface of the memory base die, and The processing unit is connected to the system bus interface via the NOC interconnect bus, and the system bus interface converts data in a die-to-die flow control unit format into a network data packet format.
10. The method of claim 1, wherein: The memory controller is communicatively coupled to the processing unit, and A second memory controller on the memory base die is communicatively coupled to a second processing unit on the memory base die.
11. A memory system comprising: A memory base die, the memory base die comprising: Memory controller; one or more memory dies stacked on top of the memory base die; a processing unit configured to process data received by the memory controller from at least one of the one or more memory dies, the data being routed to the processing unit based on at least one characteristic of a data query associated with the data; and an interconnect connecting the memory controller to the one or more memory dies stacked on the memory base die and to a plurality of processing units including the processing unit; and A die-to-die interface connects the memory base die to a compute die of a system-in-package.
12. The memory system according to claim 11, wherein: Based on determining that the function of the data query is a memory-bound function, the function is routed to the memory base die for processing by the processing unit.
13. The memory system according to claim 11, wherein: Based on determining that the function of the data query is a compute-bound operation, the function is routed to the compute die for processing by the compute die, wherein the compute die is connected to the memory base die via a silicon interposer of the system-in-package.
14. The memory system according to claim 11, wherein: The system-in-package includes a plurality of memory base dies connected to the compute die, the plurality of memory base dies including the memory base die.
15. The memory system of claim 11, wherein: The memory base die includes a system bus interface connecting the interconnect to a die-to-die interface of the memory base die, and The system bus interface maps a data format used by the interconnect to a data format used by the die-to-die interface.
16. The memory system according to claim 11, wherein: The memory base die includes a shared memory that shares data between a first processing unit and a second processing unit of the plurality of processing units.
17. The memory system according to claim 11, wherein: The memory base die includes a memory expansion port connected to at least one of a low power double data rate memory or a graphics double data rate memory external to the memory base die.
18. A non-transitory computer-readable medium storing code, the code comprising instructions executable by a processor of a device to: determining at least one characteristic of the data query; Based on the at least one characteristic of the data query, routing a first function of the data query to a memory base die for processing by a processing unit on the memory base die; and Data received by a memory controller on the memory base die from at least one of one or more memory dies stacked on top of the memory base die is processed via the processing unit.
19. The non-transitory computer-readable medium of claim 18, wherein: The code includes further instructions executable by the processor to, based on determining that the first function is a memory-constrained function, route the first function to the memory base die for processing by the processing unit on the memory base die.
20. The non-transitory computer-readable medium of claim 18, wherein: The code includes further instructions executable by the processor to: based on determining that a second function of the data query is a compute-limited operation, route the second function to a compute die for processing by the compute die, wherein the compute die is connected to the memory base die via a silicon interposer of a system-in-package that includes the compute die and the memory base die.