Dynamic memory operation
By configuring monitors in stacked memory and dynamically adjusting memory operations, the performance and power waste caused by static refresh in traditional memory systems is solved, and a more efficient memory system is achieved.
Patent Information
- Application Number
- CN202380077749.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-12
- Filing Date
- 2023-10-05
- Publication Date
- 2025-07-18
Smart Images

Figure CN120345025A_ABST
Abstract
Description
Related Applications
[0001] This application claims priority to U.S. Patent Application Serial No. 18 / 333,135, filed on June 12, 2023, with the title "Dynamic Memory Operations", which claims priority to U.S. Provisional Patent Application No. 63 / 404,792, filed on September 8, 2022, with the title "Memory Refresh Scheme", and U.S. Provisional Patent Application No. 63 / 404,796, filed on September 8, 2022, with the title "Logic Die Computing with Memory Technology Awareness" under 35 U.S.C. § 119(e). The entire contents of these disclosures are incorporated herein by reference. Background Art
[0002] Memories, such as random access memories (RAM), store data used by a processor of a computing device. Due to advancements in memory technology, various types of memories, including various non-volatile memories and volatile memories, are being widely used. Examples of such non-volatile memories include ferroelectric memories and magnetoresistive RAMs, etc., and examples of such volatile memories include dynamic random access memories (DRAM), which include high bandwidth memories and other stacked variants of DRAM. However, traditional configurations of these memories have limitations that can restrict their use in certain effective applications. For example, a computing die largely does not understand the static and dynamic characteristics of a memory die. This can lead to non-optimal system performance and power consumption. Brief Description of the Drawings
[0003] Figure 1 is a block diagram of a non-limiting example system that has a memory and a controller operable to implement dynamic memory operations.
[0004] Figure 2 Describes a non-limiting example of a printed circuit board architecture for a high bandwidth memory system.
[0005] Figure 3 Describes a non-limiting example of a stacked memory architecture.
[0006] Figure 4 Describes a non-limiting example of a memory array that stores retention bits for dynamically refreshing a memory.
[0007] Figure 5 Describes an example of a retention bit structure for dynamically refreshing a memory.
[0008] Figure 6 Describes a non-limiting example of another stacked memory architecture.
[0009] Figure 7 Describes a non - limiting example of a non - stacked memory architecture having memory and a processor on a single die.
[0010] Figure 8 Describes processes in an example implementation of dynamic memory operations.
[0011] Figure 9 Describes processes in another example implementation of dynamic memory operations.
[0012] Figure 10 Describes processes in another example implementation of dynamic memory operations. Detailed Description Overview
[0013] In modern systems, the memory wall has been consistently cited as one of the key factors limiting computing limits. High - Bandwidth Memory (HBM) and other stacked Dynamic Random - Access Memory (DRAM) are increasingly being used to mitigate off - chip memory access latency and increase memory density. Despite these advancements, traditional systems still view multi - layer memory as stacked "2D" memory macros. Doing so results in a fundamental limit on the magnitude of bandwidth and speed that can be achieved with stacked "3D" memory.
[0014] To overcome these problems, dynamic memory operations are described, including a memory refresh scheme and logic die calculations with memory technology and array awareness capabilities. In one or more implementations involving the memory refresh scheme, the technology treats stacked memory as a true three - dimensional (3D) memory by leveraging improved refresh mechanisms and components that take advantage of the increased yield and variability due to the multiple memory dies of the stacked memory. Thus, instead of statically refreshing the entire memory, the technology dynamically refreshes individual parts of the memory (e.g., individual dies and / or single rows or banks of a die) at different times. Doing so reduces the refresh overhead that occurs in traditional off - chip memory systems, thereby enabling higher performance, e.g., throughput and / or power consumption of the memory.
[0015] In traditional methods, stacked memory (e.g., DRAM) is statically refreshed at a "static" or "fixed" rate, e.g., according to Joint Electron Device Engineering Council (JEDEC) specifications. Thus, traditional methods refresh all memory at a refresh rate corresponding to the "worst - case" or "pessimistic" refresh time, e.g., approximately 64 milliseconds, and this rate can be reduced by refreshing banks separately. However, this limits the Instructions Per Cycle (IPC) and system performance, and the power overhead associated with such static refresh is higher than that of the described memory refresh scheme due to "unnecessary" refreshes.
[0016] However, by stacking multiple DRAM dies (e.g., stacking multiple layers of a memory), the technique can utilize the variation in retention time between the stacked dies, such as the variation between a first retention time of a first layer in the memory and a second retention time of a second layer in the memory.
[0017] For example, in an example of dynamic memory operation using a memory refresh scheme, the memory is characterized by associating respective retention times with different portions of the memory. Further, the portions in the memory are ranked according to the associated retention times. In different cases, this ranking is stored to provide knowledge of the variable retention times of the memory. Thus, not only can the controller obtain knowledge of the variable retention times of the memory to perform dynamic refresh and allocate memory, but the logic die can also obtain the technical characteristics of the memory (e.g., variations caused by manufacturing variations and dynamic conditions such as different heating and / or voltage drops between dies due to different workloads) to perform dynamic refresh and allocate memory. Then, according to the ranking, and in one or more varying cases, according to the best operating conditions associated with the logic die, portions in the memory are dynamically refreshed at different times.
[0018] In another example of dynamic memory operation using a memory refresh scheme, the system utilizes two or more retention bits in a portion (e.g., each row) of the memory, which are configured to represent weak logic 1 and weak logic 0. In at least one variant, the retention bits are configured to "leak" faster than other data stored in the associated portion and thus the retention bits act as a canary that indicates an early refresh of a specific portion in the memory associated with the retention bits. In one or more embodiments, the retention bits are "polled" at a predetermined time interval, such as at regular or irregular intervals (e.g., based on the occurrence of an event). In some cases, the retention bits are polled at different times and / or rates for each die of the memory or for each row or word in a memory die. In some cases, such polling results in wasted power consumption and performance. This is because, due to polling, the bank may not be available for read / write, and in some cases, an error correction code (ECC) check has to be performed each time data is read, which results in a degradation of system performance. However, doing so can reduce the number of times each row of the memory has to be refreshed because each row is only dynamically refreshed when the retention bits indicate a need for refresh. This is different from the conventional method, which refreshes each portion of the memory at the worst or most pessimistic speed. It is worth noting that this method improves both the performance and the availability of the memory.
[0019] Generally, "off-chip" memories (such as DRAM) are implemented through destructive reads, so using reserved bits is not suitable for traditional systems that use off-chip memories. This is because polling the reserved bits of such traditional designs (e.g., based on the classical off-chip DRAM 1t-1C bit cell) causes the reserved bits to be read and then written back. Doing so results in a new set of reserved bits being written each time the reserved bits are polled, thereby corrupting the data history of the relevant part of the memory. Therefore, these traditional systems cannot meet the refresh requirements shown.
[0020] To overcome these problems, in one or more embodiments, the reserved bits are configured to allow polling without a destructive read. To this end, in at least one variant, the reserved bits are implemented as bit cells different from DRAM, but mimic DRAM bit cells to capture hold time and stored value degradation and can be tuned to act as an early warning indicator of the actual degradation of other bits containing available data. In one or more embodiments, the reserved bits are configured as special multi-port cells (such as 2T gain cell DRAM), where since the read port is separated from the write port, the stored data is destroyed during a read. Notably, the retention time of these special multi-port cells is similar to or slightly worse than that of off-chip DRAM, so that the reserved bits can act as an early warning indicator. In one or more embodiments, such bits are implemented with independent read and write paths, so that the stored charge is not destroyed during a read. Different from the reserved bits of traditional configurations, the reserved bits discussed herein are configured to be polled so as to access these reserved bits separately from the actual bits of the memory portion, thus not triggering the read and write-back of all bits within the corresponding part of the memory that inherently occurs as part of a DRAM read operation.
[0021] The described technology also includes logic die computing as well as memory technology and array sensing capabilities. In one or more embodiments involving logic die computing and memory technology and array sensing, one or more dies of the memory are configured with monitors to provide feedback regarding the memory. For example, the monitors are configured to monitor various conditions (such as manufacturing variations, aging, heat, and / or other environmental conditions) of the memory or portions of the memory (such as one or more cells, rows, banks, dies, etc. of the memory). These monitors provide feedback to a logic die (such as a memory controller, ALU, CPU, GPU), such as feedback describing one or more of these conditions, and the logic die is configured for memory allocation, frequency regulation of the logic (or a portion within the memory), voltage regulation of the logic (or a portion within the memory), etc. Additionally or alternatively, in the case of logic overclocking and / or operating at increased voltage, in at least one variant, the system is configured to increase the refresh rate of the memory to compensate for the reduced retention time of bit cells due to overheating of the logic die, such as in cases related to some applications configured for single-threaded performance or performance improvement.
[0022] Accordingly, the technology further treats stacked memory as true 3D memory, in part by leveraging such monitors to take advantage of the higher yield of the stacked memory as well as variations in the conditions experienced at different parts of the memory, e.g., due to the multiple memory dies of the stacked memory and variations in the positions of the memory relative to each other, relative to other system components, and relative to external sources. Thus, rather than using static refresh in the memory die or static overclocking parameters for the entire logic die, the technology dynamically refreshes the memory and / or overclocks (overlocks) separate portions (such as separate dies and / or separate rows or groups within a die) of the logic die at different times. In this way, more of the memory and logic can be utilized by the system under more favorable conditions (such as thermal conditions or manufacturing conditions) (such as those detected in real-time and / or calibration phases) than under less favorable conditions. Compared to traditional architectures, adaptive systems (such as adaptive / dynamic voltage frequency scaling circuitry (AVFS / DVFS)) do not monitor in-die within the stacked memory die, which results in higher performance of the stacked memory and / or logic and also extends the lifespan of the hardware (such as the stacked memory).
[0023] In addition to statically refreshing the stacked memory with a "worst-case" or "pessimistic" refresh time as described above, in many conventional methods, a logic die (such as a memory controller) does not know the condition of the memory (such as the thermal condition), and the memory does not know the condition of the logic die. Some conventional HBM methods incorporate thermal sensors within the memory die to determine when a refresh occurs, e.g., a 1x refresh versus a 2x refresh in an HBM system. However, these methods are all memory-internal, and the memory does not inform the logic die (e.g., provide feedback to it) of its condition, nor does the logic (e.g., a central processing unit (CPU), a graphics processing unit (GPU), an accelerator, etc.) inform the memory die of its condition. Thus, in various situations, conventional logic dies continue to utilize portions of the stacked memory under adverse conditions rather than restricting memory-based operations to portions under favorable conditions, and due to pessimistic assumptions about the memory, the logic is overly restricted and thus operates under adverse and / or less performant limiting conditions.
[0024] Furthermore, since conventional logic dies (e.g., an AVFS system within a CPU or a memory interface die) do not have information about the operating conditions of different portions of the memory (e.g., because they are not connected to monitors within the memory portions), conventional techniques overclock the entire logic or refresh the memory in the same manner. Thus, in various situations, the way conventional logic dies "overclock" the logic can cause certain portions of the stacked memory to overheat or the refresh rate to decrease or the performance to degrade (e.g., portions under adverse conditions), while adjusting the logic frequency in other situations where the logic die cannot operate optimally. Since conventional methods cannot provide feedback about the memory condition to the logic die interfacing with the stacked memory (or provide feedback about the logic die condition to the memory die), such conventional methods cannot achieve as high a performance (such as throughput) or power savings as the techniques described. In short, conventional methods do not involve a dynamic exchange of information about operating conditions (such as thermal conditions and voltage drop conditions) between the memory and the logic die interfacing with the memory. Instead, conventional methods optimize the logic die and the memory independently, resulting in inefficiencies.
[0025] However, by configuring the memory die of the stacked memory with monitors to detect the conditions of different portions of the memory, the techniques can exploit the condition variations between the stacked dies, e.g., by differently adjusting (e.g., overclocking) the voltage and / or frequency of portions in the logic die according to different detected (e.g., thermal, manufacturing conditions) and / or known (e.g., position relative to the logic portion) conditions and / or refreshing portions of the memory at different times according to the optimal conditions of the logic die.
[0026] In some aspects, the techniques described herein relate to a system that includes: a stacked memory configured to monitor the condition of the stacked memory, and a system manager configured to receive the monitored condition of the stacked memory from the one or more memory monitors and dynamically adjust the operation of the stacked memory based on the monitored condition.
[0027] In some aspects, the techniques described herein relate to a system in which the system manager dynamically adjusts the operation of a logic die coupled to the stacked memory.
[0028] In some aspects, the techniques described herein relate to a system in which the monitored condition includes at least one of a thermal condition or a voltage droop condition of the stacked memory.
[0029] In some aspects, the techniques described herein relate to a system in which the one or more memory monitors monitor the condition of different portions of the stacked memory.
[0030] In some aspects, the techniques described herein relate to a system in which the stacked memory includes a plurality of dies, and in which the one or more memory monitors include a memory monitor for each die of the plurality of dies of the stacked memory.
[0031] In some aspects, the techniques described herein relate to a system in which the system manager is configured to adjust the operation of a logic die coupled to the stacked memory by adjusting a frequency or a voltage to prevent at least partial overheating in the logic die.
[0032] In some aspects, the techniques described herein relate to a system in which the system manager is configured to adjust the operation of a logic die coupled to the stacked memory by providing a change signal to change a voltage or a frequency used to operate one or more portions of the logic die.
[0033] In some aspects, the techniques described herein relate to a system in which the system manager is configured to adjust the operation of the stacked memory by providing a change signal to change a voltage or a frequency used to operate one or more portions of the stacked memory.
[0034] In some aspects, the techniques described herein relate to a system in which the system manager is configured to adjust the operation of the stacked memory by refreshing one or more portions of the stacked memory.
[0035] In some aspects, the techniques described herein relate to a system that includes: a memory; at least one register configured to store a ranking for each of multiple portions of the memory, the ranking for each corresponding portion being determined based on a corresponding retention time in the memory; and a memory controller for dynamically refreshing different portions of the memory at different times based on the ranking for each of the multiple portions of the memory stored in the at least one register.
[0036] In some aspects, the techniques described herein relate to a system where the at least one register is implemented at the memory controller or at an additional portion of an interface between the memory and a logic die.
[0037] In some aspects, the techniques described herein relate to a system where the memory includes a dynamic random access memory (DRAM).
[0038] In some aspects, the techniques described herein relate to a system where the different portions of the memory correspond to different dies of the DRAM.
[0039] In some aspects, the techniques described herein relate to a system where the dies of the DRAM are arranged in a stacked configuration.
[0040] In some aspects, the techniques described herein relate to a system where the different portions of the memory correspond to individual rows in the memory.
[0041] In some aspects, the techniques described herein relate to a system where the different portions of the memory correspond to individual banks of the memory.
[0042] In some aspects, the techniques described herein relate to a system that further includes a logic die for scheduling different workloads in the different portions of the memory based on the ranking for each of the multiple portions of the memory.
[0043] In some aspects, the techniques described herein relate to a system that includes: a memory, retention bits associated with corresponding portions of the memory, the retention bits indicating whether each corresponding portion is ready to be refreshed, and a memory controller for polling the retention bits to determine whether the corresponding portion is ready to be refreshed and initiating a refresh of the portion in the memory based on the polling.
[0044] In some aspects, the techniques described herein relate to a system where the retention bits are designed to be readable without a write-back.
[0045] In some aspects, the techniques described herein relate to a system in which the reserved bits have separate read and write paths.
[0046] Figure 1 FIG. is a block diagram of a non-limiting example system 100 that has a memory and a controller that can perform dynamic memory operations. In this example, system 100 includes a processor 102 and a memory module 104. Additionally, processor 102 further includes a core 106 and a controller 108. Memory module 104 includes a memory 110. In one or more embodiments, processor 102 further includes a system manager 112. Memory module 104 optionally includes a processing-in-memory component (not shown).
[0047] According to the techniques, processor 102 and memory module 104 are connected to each other by a wired or wireless connection. Core 106 and controller 108 are also connected to each other by one or more wired or wireless connections. Examples of wired connections include but are not limited to buses (such as data buses), interconnects, through-silicon vias, traces, and planes. Examples of devices that implement system 100 include but are not limited to servers, personal computers, laptops, desktops, gaming consoles, set-top boxes, tablets, smart devices, mobile devices, virtual and / or augmented reality devices, wearable devices, medical devices, systems-on-a-chip, and other computing devices or systems.
[0048] Processor 102 is an electronic circuit that performs various operations on and / or uses data in memory 110. Examples of processor 102 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), a field-programmable gate array (FPGA), an accelerated processing unit (APU), and a digital signal processor (DSP). Core 106 is a processing unit that is configured to read and execute instructions (such as instructions of a program), examples of which include adding, moving data, and branching. Although one core 106 is described in the illustrated example, in different cases, processor 102 includes more than one core 106, e.g., processor 102 is a multi-core processor.
[0049] In one or more embodiments, the memory module 104 is a circuit board (e.g., a printed circuit board) on which the memory 110 is mounted. In a variant, one or more integrated circuits in the memory 110 are mounted on the circuit board of the memory module 104. Examples of the memory module 104 include, but are not limited to, TransFlash memory modules, single in-line memory modules (SIMMs), and dual in-line memory modules (DIMMs). In one or more embodiments, the memory module 104 is a single integrated circuit device incorporating the memory 110 on a single chip or die. In one or more embodiments, the memory module 104 consists of multiple chips or dies that implement the memory 110 assembled by being vertically (“3D”) stacked together, placed side by side on a middleware or substrate, or a combination of vertical stacking and side-by-side placement.
[0050] The memory 110 is a device or system for storing information, e.g., stored by the core 106 of the processor 102 and / or by a processing-in-memory component, for immediate use in the device. In one or more embodiments, the memory 110 corresponds to a semiconductor memory where data is stored in memory cells on one or more integrated circuits. In at least one example, the memory 110 is equivalent to or includes volatile memory, examples of which include random access memory (RAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), and static random access memory (SRAM). Alternatively or additionally, the memory 110 also corresponds to or includes non-volatile memory such as ferroelectric RAM, magnetoresistive RAM, flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and electrically erasable programmable read-only memory (EEPROM).
[0051] In one or more embodiments, the memory 110 is configured as a dual in-line memory module (DIMM). The DIMM includes a series of dynamic random access memory integrated circuits, and the module is mounted on a printed circuit board. Examples of DIMM types include but are not limited to synchronous dynamic random access memory (SDRAM), double data rate (DDR) SDRAM, double data rate 2 (DDR2) SDRAM, double data rate 3 (DDR3) SDRAM, double data rate 4 (DDR4) SDRAM, and double data rate 5 (DDR5) SDRAM. In at least one variation, the memory 110 is configured as a small outline DIMM (SO-DIMM) according to one of the aforementioned SDRAM standards such as DDR, DDR2, DDR3, DDR4, and DDR5. It will be appreciated that the memory 110 can be configured in a variety of ways without departing from the spirit or scope of the technology.
[0052] In one or more embodiments, the system manager 112 includes or is otherwise configured to interface with various systems capable of updating the operation of various components of the system 100, examples of which include but are not limited to an adaptive voltage scaling (AVS) system, an adaptive voltage frequency scaling (AVFS), and a dynamic voltage frequency system (DVFS). According to the technology, the system manager 112 is thus configured to change the logic frequency or voltage (e.g., one or more parts of the processor 102) during calibration and / or dynamically. In at least one variation, the system manager 112 (or a similar manager of the memory module 104) is configured to change the frequency or voltage of one or more parts of the memory module 104 during calibration and / or dynamically. The system manager 112 is implemented using one or more of hardware or software. For example, in one example, the system manager 112 is or includes a processor. By way of example, the processor executes one or more processes to control and manage system hardware resources such as power supply to different parts of the system, the temperature of the entire system, the hardware configuration at system startup, etc. Alternatively or additionally, the system manager 112 is or includes a process being executed to perform one or more such tasks. In at least one variation, the system manager 112 is or includes firmware, such as for implementing the processor that executes the system manager 112. Additionally, the system manager also includes or otherwise accesses one or more registers, e.g., registers written to for requesting services.
[0053] In traditional methods, stacked memories (such as DRAM) are statically refreshed at a "static" or "fixed" rate, e.g., in accordance with Joint Electron Device Engineering Council (JEDEC) specifications. Thus, traditional methods refresh all memories at a refresh rate corresponding to the "worst-case" / "pessimistic" refresh time, e.g., approximately 64 milliseconds. However, this limits the number of instructions per cycle (IPC), and there is a power dissipation overhead associated with such static refreshing that is higher than the memory refresh scheme and logic die computing with memory technology awareness described above due to "unnecessary" refreshes.
[0054] For example, various traditional DDR5-configured DRAMs have performance-power-area (PPA) limitations when accessing off-chip data. A typical DRAM bit cell consists of a 1T-1C structure where the capacitor is formed by a dielectric layer sandwiched between conductor plates. The system IPC of traditional configurations is limited by DRAM bandwidth and latency, e.g., by memory-intensive workloads. In contrast, system 100 can take advantage of variations in 3D memories (such as DRAM) and improve IPC by reducing refresh overhead and / or dynamically increasing the performance (e.g., overclocking) of the portion of the stacked memory that can handle the increase.
[0055] High Bandwidth Memory (HBM) provides increased bandwidth and memory density, allowing multiple layers (e.g., stacked) DRAM dies (e.g., 8 - 12 dies) to be stacked on top of one another along with one or more optional logic / memory dies. Such a memory stack can be connected to a processing unit (such as a CPU and / or GPU) via a silicon interposer, which will be discussed in detail below in conjunction with Figure 2 Alternatively or additionally, such a memory stack can be stacked on top of a processing unit (such as a CPU and / or GPU), which will be discussed in more detail below in conjunction with Figure 3 In one or more embodiments, stacking the memory stack on top of the processing unit can further provide connectivity and performance advantages over connection via a silicon interposer.
[0056] Controller 108 is a digital circuit for managing the data flow into and out of memory 110. For example, controller 108 includes logic for reading from and writing to memory 110 and interfacing with core 106, as well as in-die processing components in variants. For example, controller 108 receives instructions from core 106 related to accessing memory 110 and provides data to core 106, e.g., data for processing by core 106. In one or more embodiments, controller 108 is communicatively located between core 106 and memory module 104, and controller 108 interfaces with both core 106 and memory module 104.
[0057] Figure 2 Describes a non - limiting example 200 of a printed circuit board architecture for a high - bandwidth memory system.
[0058] The illustrated example 200 includes a printed circuit board 202, which is described in this example as a multi - layer printed circuit board. In one example, the printed circuit board 202 is used to implement a graphics card. It should be understood that the printed circuit board 202 can be used to implement other computing systems without departing from the spirit or scope of the technology, such as a central processing unit (CPU), a graphics processing unit (GPU), a field - programmable gate array (FPGA), an accelerated processing unit (APU), and a digital signal processor (DSP), to name a few.
[0059] In the illustrated example 200, the layers of the printed circuit board 202 further include a package substrate 204, a silicon interposer 206, a processor die 208, memory dies 210 (e.g., DRAM chips), and a controller die 212 (e.g., a high - bandwidth memory (HBM) controller die). The illustrated example 200 also describes a plurality of solder balls 214 between the various layers. Here, example 200 describes the printed circuit board 202 as the first layer and the package substrate 204 as the second layer, with a first plurality of solder balls 214 disposed between the printed circuit board 202 and the package substrate 204. In one or more embodiments, this arrangement is formed by depositing the first plurality of solder balls 214 between the printed circuit board 202 and the package substrate 204. Additionally, example 200 describes the silicon interposer 206 as the third layer, with a second plurality of solder balls 214 deposited between the package substrate 204 and the silicon interposer 206. In this example 200, the processor die 208 and the controller die 212 are depicted on the fourth layer, such that a third plurality of solder balls 214 are deposited between the silicon interposer 206 and the processor die 208, and a fourth plurality of solder balls 214 are deposited between the silicon interposer 206 and the controller die 212. In this example, the memory dies 210 form an additional layer (e.g., the fifth layer) disposed "on top" of the controller die 212. The illustrated example 200 also describes through - silicon vias 216 in each of the memory dies 210 and the controller die 212, e.g., for connecting these different components.
[0060] It will be appreciated that a system for memory refresh schemes and / or logic die calculations with memory technology awareness capabilities may be implemented using different architectures in one or more variations without departing from the spirit or scope of the technology. For example, any of the components discussed above (such as printed circuit board 202, package substrate 204, silicon interposer 206, processor die 208, memory die 210 (such as a DRAM die), and controller die 212 (such as a high bandwidth memory (HBM) controller die)) may be arranged in different positions in a stacked, side-by-side, or a combination thereof manner according to the technology. Alternatively or additionally, the configuration of these components may be different from that described. For example, in one or more variations, the memory die 210 may only include a single die, and the architecture may include one or more processor dies 208, and so on. In at least one variation, one or more of the components are not included in an architecture for implementing a memory refresh scheme and / or logic die calculations with memory technology awareness capabilities according to the technology.
[0061] In this example 200, the processor die 208 includes a logic engine 218, a first controller 220, and a second controller 222. In different cases, the processor die 208 includes more, different, or fewer components without departing from the spirit or scope of the technology. In one or more embodiments, such as in a graphics card embodiment, the logic engine 218 is configured as a three-dimensional (3D) engine. Alternatively or additionally, the logic engine 218 is configured to perform different logical operations, such as digital signal processing, machine learning-based operations, and so on. In one or more embodiments, the first controller 220 corresponds to a display controller. Alternatively or additionally, the first controller 220 is configured to control different components, such as any input / output components. In one or more embodiments, the second controller 222 is configured to control memory, where in this example 200 the memory includes the controller die 212 (for example, a high bandwidth memory controller die) and the memory die 210 (for example, a DRAM die). Thus, in one or more embodiments, one or more of the second controller 222 and / or the controller die 212 correspond to the controller 108. In view of this, in one or more embodiments, the memory die 210 corresponds to the memory 110.
[0062] The illustrated example 200 also includes a plurality of data links 224. In one or more embodiments, the data links 224 are configured as 1024 data links for use in connection with a high-bandwidth memory stack and / or have a speed of 500 megahertz (MHz). In one or more variations, the configuration of these data links is different. The data links 224 described herein connect memories (such as controller die 212 and memory die 210) to the processor die 208, for example, to an interface having a second controller 222. According to the techniques, the data links 224 can be used to connect various components of the system.
[0063] In one or more embodiments, one or more solder balls 214 and / or various other components (not shown), such as one or more solder balls 214 disposed between the printed circuit board 202 and the package substrate 204, are operable to implement various functions of the system, such as implementing Peripheral Component Interconnect Express (PCIe), providing electrical current, and serving as a connector for computing components (such as a display), to name a few. In the context of another architecture, consider the following example.
[0064] Figure 3 A non-limiting example 300 of a stacked memory architecture is described. The illustrated example 300 includes one or more processor dies 302, a controller 304, and a memory 306 having a plurality of stacked portions, such as dies (e.g., four dies in this example). Although the memory 306 described in this example has four dies, in variations, the memory 306 includes more dies (e.g., 5, 6, 7, or 8+) or fewer dies (e.g., 3 or 2) without departing from the spirit or scope of the techniques. In one or more embodiments, the dies of the memory 306 are connected together, such as by through-silicon vias, hybrid bonding, or other types of connections. In this example 300, the dies of the memory 306 include a first stack 308 (such as T0), a second stack 310 (such as T1), a third stack 312 (such as T2), and a fourth stack 314 (such as T3). In one or more embodiments, the memory 306 corresponds to the memory 110. Additionally, in at least one variation, an architecture implementing one or more of the techniques optionally includes a memory monitor 316 (e.g., an in-memory monitor), a controller monitor 318, and a monitor-aware system manager 320. By way of example, the monitor-aware system manager 320 corresponds to the system manager 112.
[0065] According to the technology, the memory monitor can be implemented in a variety of ways. For example, in one or more embodiments, the memory monitor is implemented in software. Alternatively or additionally, the memory monitor can also be implemented in hardware. For example, when implemented in software, the memory monitor is or uses an operating system memory counter (e.g., a server memory counter), such as an available bytes counter indicating how many bytes of memory are available for use, a pages per second counter indicating the number of pages retrieved from disk due to a hard page fault or written to disk due to a page fault to free up space in the working set, a memory page faults per second counter indicating the page fault rate of all processes including system processes, and / or a process page faults per second counter indicating the page fault rate of a given process. In at least one variant, the memory monitor is a process executed by a processor that polls a corresponding portion of the memory at regular intervals, and / or the processor uses a corresponding portion of the memory configured to be monitored to execute. Alternatively or additionally, the memory monitor can also track the memory usage through one or more system calls and / or system application programming interfaces that provide insights into the state of the memory and / or portions of the memory. When implemented in hardware, the memory monitor is a device and / or logic integrated (e.g., embedded) in the portion of the memory it monitors. Examples of hardware implementations of the memory monitor include logic blocks or IP blocks integrated with the memory (or a portion of the memory) and / or controllers integrated with the memory (or a portion of the memory). In one or more hardware embodiments, the memory monitor includes one or more registers and / or logic units (arithmetic logic units or ALUs).
[0066] Returning to the discussion of the illustrative example 300, the die of the processor chip 302, the controller 304, and the memory 306 are arranged in a stacked manner such that the controller 304 is disposed on the processor chip 302, the first stack 308 of the memory 306 is disposed on the controller 304, the second stack 310 is disposed on the first stack 308, the third stack 312 is disposed on the second stack 310, and the fourth stack 314 is disposed on the third stack 312. In a variant, the components of a system for a memory refresh scheme and / or logic die calculation having memory technology awareness functions are arranged in different ways, such as partially stacked and / or partially side-by-side.
[0067] In one or more embodiments, memory 306 corresponds to DRAM and / or High Bandwidth Memory (HBM) cubes stacked on a computing chip such as processor chip 302. Examples of processor chips 302 include, but are not limited to, CPUs, GPUs, FPGAs, or other accelerators. In at least one variant, the system also includes a controller 304 (e.g., a memory interface die) that is placed in the stack between processor chip 302 and memory 306, i.e., stacked on top of processor chip 302. Alternatively or additionally, controller 304 is on the same die as processor chip 302, e.g., arranged side by side. When used with the techniques described, this arrangement can increase memory density and bandwidth while having a minimal impact on power consumption and performance, thus alleviating the memory bottleneck on system performance. This arrangement can also be used with the refresh techniques for various other types of memory, such as FeRAM and MRAM.
[0068] As described above, in traditional methods, stacked memories (e.g., DRAM) are statically refreshed at a “static” or “fixed” rate, e.g., in accordance with Joint Electron Device Engineering Council (JEDEC) specifications. Thus, traditional methods refresh all memory at a refresh rate corresponding to the “worst case” or “pessimistic case” refresh time (e.g., about 64 milliseconds).
[0069] By stacking multiple DRAM dies (e.g., multiple tiers of memory 306), the techniques take advantage of the variation in retention time between the stacked DRM dies, e.g., by taking advantage of the variation between a first retention time of a first tier 308 of memory 306 and a second retention time of a second tier 310 of memory 306. In one or more embodiments, these variations are used to “buy back” performance and power consumption.
[0070] In at least one embodiment, this is achieved by employing dynamic refresh (stacked memory 306) rather than static refresh. In one example of dynamic refresh, memory 306 is characterized by associating different retention times with portions of the memory. The portions of the memory are ranked based on the associated retention times. Then, this ranking can be stored to provide knowledge of the variable retention times of memory 306 not only to controller 304 to perform dynamic refresh and allocate memory, but also to the logic die to utilize the memory technology characteristics. The ranking can be stored in a register of the system. As described herein, a register corresponds to a storage unit accessible by controller 304. For example, in some cases, the ranking is stored in a register of controller 304. In other embodiments, the ranking is stored in a register or storage unit accessible by controller 304, such as a register implemented on a memory interface or logic die. For example, the register is implemented in an interface portion between the logic die and the memory die. Then, portions of the memory are dynamically refreshed at different times based on the ranking.
[0071] In another example of dynamic refresh, the system utilizes two or more retention bits in portions (e.g., each row) of memory 306, which are configured to represent weak logic 1 and weak logic 0, and are discussed in more detail below Figure 4 and Figure 5 in more detail.
[0072] As described above, in one or more embodiments, memory 306 optionally includes memory monitor 316. Here, each stack of memory 306 is described as having at least one memory monitor 316, and the first stack 308 of memory 306 is described as including two memory monitors 316. However, in a variant, without departing from the spirit or scope of the technology, portions (e.g., dies) of memory 306 include different numbers of memory monitors 316 (e.g., less than 1 per die, 1 per die, more than 1 per die, etc.). Also in a variant, memory monitors 316 are positioned at different physical locations of memory 306 according to the technology.
[0073] In one or more variants where the architecture includes memory monitor 316, controller 304 optionally includes controller monitor 318, and processor chip 302 optionally includes monitor-aware system manager 320. In a variant, controller 304 includes more than one controller monitor 318, and the physical location of one or more controller monitors 318 in controller 304 is variable without departing from the technology.
[0074] Broadly speaking, the memory monitor 316 is configured to detect one or more conditions of a portion of the memory, such as thermal conditions (e.g., thermal sensors), voltage droop conditions, and / or memory retention. In one or more embodiments, the memory monitor 316 detects one or more conditions of a portion of the memory that provides input to a logic die, such as the processor chip 302 or one or more portions thereof. Similarly, the controller monitor 318 is configured to detect one or more conditions of the controller 304 or portions of the controller 304. The detected conditions are passed to the monitoring awareness system manager 320.
[0075] The monitoring awareness system manager 320 statically and / or dynamically adjusts the operation of the memory 306 (or a portion of the memory) based on the detected conditions. Alternatively or additionally, the monitoring awareness system manager 320 dynamically adjusts the operation of the processor chip 302 (e.g., the logic die). For example, the monitoring awareness system manager 320 provides a change signal to change one or more of the voltage or frequency used to operate one or more portions of the memory 306, e.g., overclocking one or more portions without overclocking one or more other portions, and / or overclocking one or more portions of the memory in a different manner than one or more other portions of the memory 306. Alternatively or additionally, the monitoring awareness system manager 320 provides a change signal to change one or more of the voltage or frequency used to operate one or more portions of the processor chip 302, e.g., overclocking one or more portions without overclocking one or more other portions, and / or overclocking one or more portions of the processor chip 302 in a different manner than one or more other portions of the memory 306. Alternatively or additionally, the monitoring awareness system manager 320 updates the refresh of one or more portions of the memory 306, e.g., updating the refresh of one or more portions of the memory 306 in a different manner than one or more other portions of the memory 306.
[0076] In a memory (e.g., DRAM), the retention time can vary for different reasons, such as due to static manufacturing defects and variations, dynamic thermal variations, dynamic rowhammering (e.g., adjacent rows causing additional leakage of the storage capacitor in a bit cell (DRAM 1T-1C bit cell)), and aging of components in the memory, thereby consuming the retention time of the memory over the lifespan of the system.
[0077] For example, in the case of a stacked arrangement (such as the arrangement described in Example 300), the first stack 308 of the memory 306 is likely to experience the greatest thermal gradient because the first stack 308 is closest to the compute die (such as the processor chip 302), and in many cases, the logic die (the logic die of the compute chip) has higher activity and generally generates more heat than the memory die. Additionally, in one or more embodiments, the memory cells within a particular portion of the memory 306, such as the portion of the memory 306 that is physically closer to the hot logic IP blocks of the processor chip 302, also generate more heat than other functional blocks, resulting in a greater gradient of retention time variation due to thermal effects. Examples of such portions of the memory 306 include one or more cells (such as DRAM cells) of the first stack 308 of the memory 306 that are located above the logic IP blocks (such as execution units and schedulers) of the processor chip 302. Different from traditional methods that do not track the thermal or voltage drop conditions across a single memory die, in one or more variations, the system detects the thermal and / or voltage drop conditions of a memory die by configuring the memory die to include multiple memory monitors 316 (such as thermal, retention, etc.) and / or partially by using the memory monitors 316 on another memory die, thereby tracking the thermal and / or voltage drop conditions across a single memory.
[0078] According to the technique, the monitor-aware system manager 320 is configured to dynamically or adaptively adjust the voltage and / or frequency. The monitor-aware system manager 320 receives information from one or more memory monitors 316 and / or one or more controller monitors 318 that capture the thermal gradients within one or more dies. Based on this information, the monitor-aware system manager 320 adjusts the frequency and / or voltage (such as for a portion of the processor chip 302), for example, to prevent the chip from overheating. The monitor-aware system manager 320 can also derive the voltage-frequency relationship to optimize the performance (such as maximizing the performance) and power consumption (such as minimizing the power consumption) of the system.
[0079] In one or more embodiments, the technique takes into account the multi-stack thermal effects of the monitor-aware system manager 320 (of the memory 306) by (1) integrating the controller monitor 318 (such as a thermal sensor) into the controller 304 (such as the memory interface logic), and (2) integrating the memory monitor 316 (such as a thermal sensor and / or a retention monitor) into one or more stacks of the memory 306 at a single or multiple locations within the memory die, for example, to capture different heat sources under (and / or above) the memory stacks. This is because different logic IP blocks correspond to different heat sources, and the thermal effects generated by these heat sources are also different.
[0080] For example, in the context of Example 300, the first stack 308 of the memory 306 is more likely to be affected by the workload, logic die heating, and / or degradation over time of different parts of the memory 306. In one or more embodiments, one or more of the memory monitors 316 are calibrated at wafer sort or after packaging and fused to account for the unique "worst-case" retention times of each layer. In a variant, one or more of the memory monitors 316 are fused at different granularities within the stacks of the memory 306 to capture different thermal variations across the stacks (e.g., across the first stack 308 of the memory 306).
[0081] Based on this, the monitoring awareness system manager 320 determines whether the information from the memory monitors 316 indicates that thermal variations and / or manufacturing variations have caused the retention time of the memory 306 (or a portion of the memory 306) to fall below a threshold retention time. In one or more embodiments, the monitoring awareness system manager 320 dynamically tracks the information from the memory monitors 316 during system operation to capture the memory aging effects that affect the retention time of the bit cells. This includes adjusting the memory monitors 316 by updating the retention times of the stacks of the memory 306 based on the aging and / or row hammering conditions during operation. In such an embodiment, the system sets thresholds to control the monitoring awareness system manager 320 (e.g., the AVFS system).
[0082] In traditional methods, the retention time typically refers to a single value for the entire memory (e.g., an entire DRAM or HBM), such as according to JEDEC. Different from traditional methods, the techniques integrate one or more retention monitors for logic 0 and logic 1 (e.g., memory monitor 316) into each stack, such as memory 306. In one or more embodiments, the techniques determine the preferred retention for each of a plurality of memory stacks (e.g., each die), such as during calibration. The preferred retention (for logic 0 or logic 1) for each die is stored in a register file, such as the register file of controller 304. In at least one variant, the logic is integrated into controller 304 (such as a memory interface controller) to allocate data that is predominantly 1 to stacks with a preference for logic 1, and data that is predominantly 0 to stacks with a preference for logic 0. The system uses one or more techniques to determine whether the data is predominantly 0 or 1, such as by calculating the checksum or a derivative of the checksum of the data. Alternatively or additionally, the memory interface / controller die (e.g., controller 304) includes logic that converts the data to contain more 0s or 1s. For example, when the retention in memory 306 (or stack) for logic 0 is longer than the retention for logic 1, the memory interface / controller die converts the data so that it contains more 0s than 1s. By way of example, the logic converts the data such that if the data is dominated by an unwanted value (e.g., not the value that the stack has a preference for), the one's complement of the data is stored, along with an additional bit to indicate whether the true value or the complement value is stored.
[0083] In one or more embodiments, the technology allocates various data and / or applications to different portions (e.g., different die stacks of memory 306) of memory 306 for operation. This is because certain types of data and applications result in more frequent data updates and / or lower data retention than other types of data and applications. For example, the GPU frame buffer is typically updated more frequently than data written by the CPU. In at least one variant, the system (e.g., controller 304) includes logic that causes controller 304 to allocate data maintained by the GPU frame buffer to a die stack of memory 306 determined to have lower retention. At runtime, this data is refreshed more frequently by the corresponding applications and will thus be paired with a portion of the memory that has a condition (e.g., lower retention) complementary to the characteristics of the data (such as more frequent refresh). This is in contrast to data with a lower refresh frequency, such as CPU data, which is typically refreshed less frequently and only as a system idle. In one or more embodiments, controller 304 allocates such CPU data to portions (e.g., die stacks) of memory 306 that are determined to have relatively higher retention. In at least one variant, various workloads involving access to memory are profiled based on memory access characteristics (such as update frequency) associated with running those workloads. The workloads are then allocated to a portion of the memory based on the profiling, such as allocating the workload to a portion of the memory where the operating conditions are complementary to the characteristics of the workload (or more complementary than the characteristics of other workloads that are more suitable for allocation to different portions of the memory). For example, in one or more embodiments, the logic die schedules (e.g., by implementing a scheduler) different workloads in different portions of the memory based on a ranking (where the ranking represents retention time). At runtime, some workloads are modified frequently and can thus tolerate bit cells with shorter retention times. For example, workloads implemented on the GPU can access memory more frequently and can thus tolerate memory bit cells (e.g., portions of the memory) with lower retention times than other types of workloads and / or operations.
[0084] In various cases, applications (such as probability calculations) tolerate least significant bit (LSB) errors. In one or more embodiments, the monitoring awareness system manager 320 ranks portions of the memory 306. For example, the monitoring awareness system manager 320 ranks bits, rows, banks, and / or dies of the memory 306. The monitoring awareness system manager 320 also stores data based on least or most significant bit characteristics and based on the monitored conditions of portions in the memory 306. For example, the monitoring awareness system manager 320 stores LSB data in portions of the memory with lower retention (according to the ranking) and stores most significant bit (MSB) data in portions of the memory with higher retention (according to the ranking). In one or more embodiments, the dies of the memory 306 are stacked on top of each other and on the processor chip 302 such that the functional blocks that generate more heat are aligned with the LSB bits of the memory 306 stacked above.
[0085] Although the stacked configuration with multiple memory dies is discussed above, it can be understood that in one or more embodiments, the memory may include only a single die stacked on the controller 304 and / or the processor chip 302. In this case, see Figure 4 .
[0086] Figure 4 A non-limiting example 400 of a memory array storing retention bits for dynamically refreshing a memory is described.
[0087] The memory array of example 400 utilizes "retention bits" associated with different portions of the memory. For example, example 400 describes a first portion 402 and a second portion 404 of the memory 306. The first and second portions may correspond to different portions of the memory 306, such as different rows of the memory, different dies of the memory 306, different banks of the memory 306, etc. In this example, the first portion 402 includes a first retention bit 406 and a second retention bit 408, and the second portion 404 also includes a first retention bit 410 and a second retention bit 412.
[0088] In one or more embodiments, the first and second reserved bits of each section respectively represent a weak logic 1 and a weak logic 0. The reserved bits are configured to "leak" faster than the rest of the data stored in the associated section, and thus the reserved bits serve as an early warning indicator to indicate that a particular memory section associated with the reserved bits is ready to be refreshed early. In one or more embodiments, for example, the controller 304 "polls" the reserved bits at a predetermined time interval. It is noted that such polling may result in power wastage. However, doing so can reduce the number of times each row of the memory must be refreshed, because each row is only dynamically refreshed when the reserved bit indicates that the row is ready to be refreshed. It is noted that this increases the availability of the memory while improving performance.
[0089] Generally, "off-chip" memories (such as DRAM) are implemented through destructive reads, and thus using reserved bits may not be suitable for traditional systems using off-chip memories. This is because polling the reserved bits causes the reserved bits to be read and then written back. Doing so causes the refresh group for the reserved bits to be written each time the reserved bits are polled, thus destroying the data history of the associated memory section.
[0090] To overcome this problem, the reserved bits are configured to allow polling without a destructive read. To this end, the reserved bits are implemented as bit cells different from DRAM, but can mimic DRAM bit cells to capture the retention time and storage value degradation. In one or more embodiments, the reserved bits are configured as special dual-port cells (such as 2T gain cell DRAM), where the stored data is not destroyed during a read. It is noted that these special dual-port cells have a retention time similar to or slightly worse than that of off-chip DRAM, so that the reserved bits can function as an early warning indicator. Different from conventionally configured DRAM, these bit cells implement independent read and write paths, so that the stored charge is not destroyed during a read. The reserved bits can be polled, allowing the reserved bits to be accessed separately from the actual bits of the memory section, thus not triggering the read and write-back of all bits in the corresponding memory section that occurs as part of a DRAM read operation.
[0091] Figure 5 An example 500 of a reserved bit structure for dynamically refreshing a memory is described. Figure 5 The reserved bit structure in is based on a 2T gain cell DRAM circuit. Other bits in the DRAM WL will be tied to the WWL. Thus, polling during a refresh can only be done through RWL reads without interfering with the SN on the reserved bits or triggering a write-back on the DRAM row word line bits.
[0092] In one or more embodiments, retention bits are used to describe the discrete retention time of each memory stack, and a hybrid refresh technique is employed, which refreshes the DRAM stack at a discrete refresh rate assigned to each stack, i.e., static refreshes can be performed within the die, but each stacked die has its own designated / required refresh time. For example, stack 0 is refreshed at a refresh rate of "r0", stack 1 is refreshed at a refresh rate of "r1" (different from r0), and so on, characterized by a "calibration" phase and dynamically recalibrating the retention time of the stacks at inserted checkpoints to capture how these retention times for the memory stacks change over time.
[0093] In one or more embodiments, the use of retention bits is extended to prevent row hammering. For example, the aforementioned retention bits store a history of changes over time and "leakage" in the bit cells, thus providing an early indication of leakage caused by adjacent row hammering. Different from tracking row hammering in the memory controller when using static refreshes, polling the retention bits during horizontal dynamic refreshes within the stack will capture the "leakage" associated with row hammering and initiate a refresh through polling.
[0094] Soft error ECC checking can also be used as a proxy to determine when to initiate a refresh. For example, in-memory processing components perform ECC at the bank level, which can be used to perform a refresh at the bank level and vice versa (refresh polling can be used to feed into the ECC).
[0095] Figure 6 A non-limiting example 600 of another stacked memory architecture is described. Example 600 includes one or more processor chips 602, a controller 604, and a memory 606. In at least one variant, the memory 606 is a non-volatile memory, such as ferroelectric RAM or magnetoresistive RAM. Alternatively, the memory 606 is a volatile memory as described above. In a variant, the components (e.g., one or more processor chips 602, the controller 604, and the memory 606) are connected in various ways as discussed above. Thus, in one or more scenarios, the memory 606 corresponds to the memory 110. In at least one variant, the stacked memory architecture optionally includes a memory monitor 608 (e.g., an in-memory monitor such as a thermal sensor and / or a retention monitor), a controller monitor 610, and a monitor-aware system manager 612. In one or more embodiments, the monitor-aware system manager 612 corresponds to the system manager 112.
[0096] In this example 600, the processor chip 602, the controller 604, and the memory 606 are arranged in a stacked manner such that the controller 604 is arranged on the processor chip 602, and the memory 606 is arranged on the controller 604. As described above, the components of a system for a memory refresh scheme and / or logic die computation with memory technology awareness functions are arranged in different ways without departing from the spirit of the technology. In a variant where the architecture includes a memory monitor 608, a controller monitor 610, and a monitor awareness system manager 612, these components include the same and / or similar functions as those discussed above with respect to Figure 3 The same and / or similar functions as those discussed above with respect to
[0097] In one or more embodiments where the memory 606 is a non-volatile memory, the memory 606 has a higher temperature tolerance than one or more volatile memory embodiments. Generally, the monitor awareness system manager 612 is configured to adjust the frequency and voltage based on the heat generated by the cooling / encapsulation and logic die indicated by the information from the memory monitor 608. Since various non-volatile memories have a higher temperature tolerance, in at least one variation, the monitor awareness system manager 612 is configured to further overdrive the voltage and frequency relative to one or more volatile memory configurations. In this case, the monitor awareness system manager 612 receives inputs of the non-volatile memory thermal characteristics, which are calibrated during the wafer sort / encapsulation step and / or dynamically calibrated at regular checkpoints, enabling higher system performance and lower power consumption compared to other methods that utilize non-volatile memory stacked memory configurations.
[0098] As described above, in one or more embodiments, the technology determines the priority retention of a memory die (e.g., memory 606), such as during calibration. This priority retention (logic 0 or logic 1) of the die is stored in a register file, such as the register file of the controller 604. In at least one variant, the monitor awareness system manager 612 adjusts the voltage level to improve the writing of non-priority values, thereby improving the overall performance with minimal power dissipation. As another example of component arrangement, consider Figure 7 The following example in
[0099] Figure 7Describes a non - limiting example 700 of a non - stacked memory architecture having a memory and a processor on a single die. Example 700 includes one or more processor chips 702, a controller 704, and a memory 706. In at least one variant, the memory 706 is a non - volatile memory, such as a logic - compatible ferroelectric RAM or magnetoresistive RAM. Alternatively, the memory 706 is a volatile memory, as exemplified above. In various variants, the components (e.g., one or more processor chips 702, controller 704, and memory 706) are connected in various ways, such as those discussed above. In one or more embodiments, the architecture further includes a memory monitor 708 (e.g., an in - memory monitor, such as a thermal sensor and / or a retention monitor) and a monitor - aware system manager 710. Although not described, in one or more embodiments, the controller 704 also includes a controller monitor.
[0100] In at least one example, such as the illustrated example 700, one or more processor chips 702, a controller 704, and a memory 706 are arranged side - by - side on a single die, e.g., each of these components is arranged on the same die. For example, the controller 704 is connected side - by - side with the processor chip 702, and the memory 706 is connected side - by - side with the controller 704, such that the controller 704 is disposed between the memory 706 and the processor chip 702. In various variations, the components of a system for a memory refresh scheme and / or logic die calculation having memory technology awareness functions are arranged in different side - by - side arrangements (or partially side - by - side arrangements) without departing from the spirit or scope of the technology. In one or more embodiments, the memory monitor 708 and the monitor - aware system manager 710 include functions similar to those discussed above Figure 3 functions.
[0101] Figure 8 Describes a process in an embodiment of a dynamic memory operation example 800.
[0102] One or more memory monitors monitor the condition of a stacked memory (block 802). The monitored condition of the stacked memory is communicated by one or more memory monitors to a system manager (block 804). The operation of the stacked memory is adjusted by the memory monitor based on the monitored condition (block 806).
[0103] Figure 9 Describes a process of another example 900 embodiment of a dynamic memory operation.
[0104] The memory is characterized by associating different retention times with portions of the memory (block 902). For example, portions of the memory 306 (e.g., rows, dies, or banks) are characterized as associating corresponding retention times with different portions of the memory 306.
[0105] Rank portions of the memory according to associated retention times (block 904). For example, portions (e.g., rows, dies, or banks) of memory 306 are ranked based on associated retention times determined from the memory characterization performed at block 802.
[0106] Dynamically refresh portions of the memory at different times based on the ranking (block 906). For example, different portions of memory 306 are refreshed based on the ranking. This is in contrast to static refreshing performed by conventional systems where all portions of the memory are refreshed at the same time.
[0107] Figure 10 A process in an example 1000 implementation of dynamic memory operation is described.
[0108] Retention bits associated with respective portions of the memory are polled to determine whether to refresh the respective portions of the memory (block 1002). According to the techniques discussed herein, polling results in a read of the retention bits without a write-back, thereby storing a history of changes over time and leakage in the bit cells. The respective portions of the memory are refreshed at different times based on the polling (block 1004).
[0109] It should be understood that many variations are possible in accordance with the disclosure herein. Although the features and elements are described above in particular combinations, each feature or element may be used alone without other features and elements or in various combinations with or without other features and elements.
[0110] The various functional units shown in the figures and / or described herein (including memory 110, controller 108, and core 106 where appropriate) are implemented in a variety of different ways, such as hardware circuitry, software or firmware executed on a programmable processor, or any combination of two or more of hardware, software, and firmware. The provided methods may be implemented in any of a variety of devices, such as a general-purpose computer, a processor, or a processor core. For example, suitable processors include general-purpose processors, dedicated processors, conventional processors, digital signal processors (DSPs), graphics processing units (GPUs), parallel acceleration processors, multiple microprocessors, one or more microprocessors associated with DSP cores, controllers, microcontrollers, application specific integrated circuits (ASICs), field programmable gate array (FPGA) circuitry, any other type of integrated circuit (IC), and / or state machines.
[0111] In one or more embodiments, the methods and procedures provided herein are implemented in the form of a computer program, software, or firmware, which are contained in a non-transitory computer-readable storage medium for execution by a general-purpose computer or processor. Examples of non-transitory computer-readable storage media include read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor storage devices, magnetic media (such as internal hard disks and removable disks), magneto-optical media, and optical media (such as CD-ROM disks and digital versatile disks (DVDs)).
Claims
1. A system, comprising: A stacked memory; One or more memory monitors configured to monitor the condition of the stacked memory; And A system manager configured to receive the monitored condition of the stacked memory from the one or more memory monitors and dynamically adjust the operation of the stacked memory based on the monitored condition.
2. The system according to claim 1, wherein the system manager dynamically adjusts the operation of a logic die coupled to the stacked memory.
3. The system according to claim 1, wherein the monitored condition includes at least one of a thermal condition or a voltage drop condition of the stacked memory.
4. The system according to claim 1, wherein the one or more memory monitors monitor the conditions of different portions of the stacked memory.
5. The system according to claim 1, wherein the stacked memory includes a plurality of dies, and wherein the one or more memory monitors include a memory monitor for each die of the plurality of dies of the stacked memory.
6. The system according to claim 1, wherein the system manager is configured to adjust the operation of a logic die coupled to the stacked memory by adjusting the frequency or voltage to prevent at least part of the logic die from overheating.
7. The system according to claim 1, wherein the system manager is configured to adjust the operation of a logic die coupled to the stacked memory by providing a varying signal to change the voltage or frequency for operating one or more portions of the logic die.
8. The system according to claim 1, wherein the system manager is configured to adjust the operation of the stacked memory by providing a varying signal to change the voltage or frequency for operating one or more portions of the stacked memory.
9. The system according to claim 1, wherein the system manager is configured to adjust the operation of the stacked memory by refreshing one or more portions of the stacked memory.
10. A system, comprising: A memory; At least one register configured to store a ranking for each of a plurality of portions of the memory, the respective ranking being determined based on a relevant retention time of the corresponding portion in the memory; And A memory controller for dynamically refreshing different portions of the memory at different times based on the ranking for each of the plurality of portions of the memory stored in the at least one register.
11. The system according to claim 10, wherein the at least one register is implemented at the memory controller or at an additional portion of an interface between the memory and a logic die.
12. The system according to claim 10, wherein at least one of the memories includes a dynamic random access memory (DRAM), different portions of the memory correspond to different dies of the DRAM, or the dies of the DRAM are arranged in a stacked configuration.
13. The system according to claim 10, wherein the different portions of the memory correspond to a single row in the memory.
14. The system according to claim 10, wherein the different portions of the memory correspond to a single bank of the memory.
15. The system according to claim 10, further comprising a logic die for scheduling different workloads among the different portions of the memory based on the ranking of each of the plurality of portions of the memory.
Citation Information
Cited By
Dynamic memory operations
US12694921B2