Prioritized power budget arbitration for multiple concurrent memory access operations

By introducing a thread manager and a peak power manager into the memory die, and adopting a token round-robin scheduling protocol and a priority ring counter, the power budget management problem during concurrent operation of multiple memory dies is solved, achieving efficient current distribution and performance improvement.

CN115809209BActive Publication Date: 2025-09-30MICRON TECHNOLOGY INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211111221.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-02-09
Filing Date
2022-09-13
Publication Date
2025-09-30
Estimated Expiration
2042-09-13

AI Technical Summary

Technical Problem

Existing technologies have difficulty effectively managing the power budget when multiple memory dies operate concurrently, resulting in high peak current usage, and traditional power management techniques are insufficient to address the increased complexity.

Method used

The thread manager and peak power manager (PPM) in multiple memory dies are used to implement prioritized power budget arbitration for multiple concurrent access operations through a token-based round-robin scheduling protocol and priority ring counters, ensuring effective management of current distribution.

Benefits of technology

It achieves effective power management of multi-die memory subsystems, supports multiple processing threads operating concurrently, improves the performance and service quality of memory devices, avoids power budget violations, is highly scalable, and does not rely on external controller intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115809209B_ABST
    Figure CN115809209B_ABST
Patent Text Reader

Abstract

The present application relates to prioritized power budget arbitration for multiple concurrent memory access operations. A memory device includes memory dies, each memory die including: a memory array; a memory for storing a data structure; and control logic including: a plurality of processing threads for concurrently executing memory access operations on the memory array; a priority ring counter, the data structure for storing an association between a value of the priority ring counter and a subset of the plurality of processing threads; a thread manager for incrementing the value of the priority ring counter and identifying one or more prioritized processing threads corresponding to the subset of the plurality of processing threads before a power management cycle; and a peak power manager coupled to the thread manager for prioritizing allocation of power to the one or more prioritized processing threads during the power management cycle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate generally to memory subsystems, and more particularly, to prioritized power budget arbitration for multiple concurrent memory access operations. Background Art

[0002] The memory subsystem may include one or more memory devices that store data. The memory devices may be, for example, non-volatile memory devices and volatile memory devices. Generally speaking, the host system may utilize the memory subsystem to store data at the memory devices and retrieve data from the memory devices. Summary of the Invention

[0003] An embodiment of the present disclosure provides a memory device comprising: a plurality of memory dies, each of the plurality of memory dies comprising: a memory array; a memory for storing a data structure; and control logic operatively coupled to the memory array and the memory, wherein the control logic comprises: a plurality of processing threads for concurrently performing memory access operations on the memory array; a priority ring counter, wherein the data structure is for storing an association between a value of the priority ring counter and a subset of the plurality of processing threads; a thread manager for incrementing the value of the priority ring counter and identifying one or more prioritized processing threads corresponding to the subset of the plurality of processing threads before a power management cycle; and a peak power manager coupled to the thread manager and for prioritizing allocation of power to the one or more prioritized processing threads during the power management cycle.

[0004] Another embodiment of the present disclosure provides a memory device comprising: a memory array; and control logic operatively coupled to the memory array to perform operations including: allocating power to one or more prioritized processing threads among a plurality of processing threads based on a value of a priority ring counter, the plurality of processing threads being used to perform memory access operations on the memory array; starting a timer when the one or more prioritized processing threads are running and in response to detecting that power is allocated to a non-prioritized processing thread among the plurality of processing threads; while the timer is running: incrementing the priority ring counter before each power management cycle; and prioritizing the allocation of power to the one or more prioritized processing threads within a subset of the plurality of processing threads corresponding to the value of the priority ring counter; and in response to the timer expiring before completion of the non-prioritized processing thread, shifting power allocation between the subset of the plurality of processing threads based on the value incremented to the non-priority ring counter.

[0005] Yet another embodiment of the present disclosure provides a method comprising: allocating power, by control logic of a memory die among a plurality of memory dies, to a non-prioritized processing thread among a plurality of processing threads based on a value of a non-priority ring counter, the plurality of processing threads being used to perform memory access operations on a memory array of the memory die; starting a timer by the control logic when the non-prioritized processing thread is running and in response to detecting that the power is allocated to a prioritized processing thread among one or more prioritized processing threads among the plurality of processing threads; while the timer is running: incrementing a priority ring counter before each power management cycle; and prioritizing the allocation of power to the one or more prioritized processing threads within a subset of the plurality of processing threads corresponding to the value of the priority ring counter; and in response to the timer expiring before the non-prioritized processing thread is completed, shifting the power allocation between the subset of the plurality of processing threads based on the value incremented to the non-priority ring counter by the control logic. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] The present disclosure will be more fully understood from the detailed description given below and the accompanying drawings of various embodiments of the present disclosure.

[0007] Figure 1A An example computing system including a memory subsystem in accordance with at least some embodiments is shown.

[0008] Figure 1B is a block diagram of a memory device communicating with a memory subsystem controller of a memory subsystem according to an embodiment.

[0009] Figure 2 is a block diagram illustrating a multi-die package with multiple memory dies in a memory subsystem, according to at least some embodiments.

[0010] Figure 3 is a block diagram illustrating a multi-plane memory device configured for parallel plane access, according to at least some embodiments.

[0011] Figure 4 is a block diagram illustrating a memory die configured for prioritized power budget arbitration for multiple processing threads, according to at least some embodiments.

[0012] Figure 5 is a block diagram illustrating the operation of a non-priority ring counter implemented by a thread manager of a memory die, according to at least some embodiments.

[0013] Figure 6 is a flow chart of an example method for power budget arbitration in a memory device using a ring counter, according to at least some embodiments.

[0014] Figure 7 is a flow chart of an example method for power budget arbitration in a memory device using a polling window, in accordance with at least some embodiments.

[0015] Figure 8 is a block diagram illustrating a combination of a memory command packet and a timing diagram in accordance with at least some embodiments.

[0016] Figure 9 is a diagram illustrating multi-plane prioritized power budget arbitration for multiple concurrent memory access operations in accordance with at least some embodiments.

[0017] Figures 10A to 10B is a diagram illustrating multi-plane prioritized power budget arbitration for multiple concurrent memory access operations according to at least some additional embodiments.

[0018] Figure 11 is a flow diagram of an example method for prioritized power budget arbitration for multiple concurrent processing threads in accordance with at least some embodiments.

[0019] Figure 12 is a block diagram of an example computer system in which embodiments of the present disclosure may operate. DETAILED DESCRIPTION

[0020] Aspects of the present disclosure relate to prioritized power budget arbitration for multiple concurrent memory access operations. An example of a non-volatile memory device is a NAND memory device. The memory device may be composed of bits arranged in a two-dimensional or three-dimensional grid of memory cells to encompass each die of a multi-die memory device. One or more physical blocks of memory cells may be grouped together to form planes of the memory device to allow concurrent operations on each plane, wherein these physical blocks are composed of multiple groups of pages of memory cells.

[0021] Each memory die may include circuitry for performing concurrent memory page accesses for two or more memory planes. For example, each memory die may include multiple access line driver circuits and power supply circuits that can be shared by the planes of each memory die to facilitate concurrent access of pages of two or more memory planes containing different page types. For ease of description, these circuits may generally be referred to as independent plane driver circuits. The control logic on each die of the memory device includes several separate processing threads to perform concurrent memory access operations (e.g., read operations, program operations, and erase operations). For example, each processing thread corresponds to a respective memory plane and utilizes associated independent plane driver circuits to perform memory access operations on the respective memory plane. When these processing threads operate independently, the power usage and requirements associated with each processing thread also change.

[0022] The capacitive load of three-dimensional memories is typically large and may continue to grow as process scaling continues. During sensing (e.g., reading or verifying), programming, and erasing operations, various access lines, data lines, and voltage nodes may be charged or discharged very quickly so that memory array access operations can meet performance specifications, such as those typically required to meet data throughput targets, which may be dictated by customer requirements or industry standards. For sequential reading or programming, multi-plane operations are often used to increase system throughput. Consequently, a typical memory die may have high peak current usage, which may be four to five times the average current amplitude. Consequently, with the high average market requirements for this total current usage budget, operating more than four memory dies concurrently, for example, may become challenging.

[0023] Various techniques have been utilized to manage power consumption in memory subsystems containing multiple memory dies, many of which rely on a memory subsystem controller to interleave the activities of the memory dies, thereby avoiding the high-power portion of access operations being performed concurrently in more than one memory die. Furthermore, these power management techniques are insufficient to address the increased complexity associated with budgeted current usage within each individual memory die due to the use of additional processing threads (e.g., 4, 6, or 8 processing threads) on each individual memory die.

[0024] Aspects of the present disclosure address the above and other deficiencies by providing prioritized power budget arbitration for multiple concurrent access operations in a memory device of a memory subsystem. In some embodiments, the memory device includes a plurality of dies, each die including a plurality of processing threads configured to, for example, execute concurrent memory access operations on corresponding memory planes of the memory die. Each memory die further includes a thread manager and a peak power manager (PPM) that together are configured to perform prioritized power budget arbitration for the multiple processing threads on a corresponding memory die in the plurality of memory dies.

[0025] In these embodiments, the memory subsystem employs a token-based round-robin scheduling protocol, whereby each PPM rotates as the holder of a token (e.g., after a set number of cycles of a shared clock signal) and broadcasts a quantized current budget to be consumed by its respective memory die during a given time period. The other PPMs on each other memory die receive this broadcast information and can therefore determine the available current budget in the memory device during that time period. While holding the token, the PPM can request a certain amount of current for its respective memory die, depending on the available current budget in the memory device and based on the amount of current consumed by the other memory dies of the memory device. As described in further detail below, the PPM can employ several different techniques to distribute the requested current among the multiple processing threads of the respective memory die, at least some of which include prioritizing multiple concurrent processing threads.

[0026] In at least some embodiments, each die also includes a priority ring counter and a data structure (e.g., a lookup table) for storing an association between the value of the priority ring counter and a subset of the plurality of processing threads. Each die may further include a thread manager configured to manage the plurality of processing threads presented to the PPM for power allocation. In these embodiments, the thread manager may increment the value of the priority ring counter before a power management cycle. Each new counter value changes the subset of the plurality of processing threads considered for power allocation, thereby simplifying the number of processing threads that the PPM can concurrently manage. The thread manager may also identify one or more prioritized processing threads within the subset of the plurality of processing threads and provide the PPM with prioritized identification of the one or more prioritized processing threads. Then, when a die of the PPM holds a token, the PPM may prioritize allocation requests, for example, by prioritizing the allocation of power to the one or more prioritized processing threads within the subset of the plurality of processing threads during a power management cycle. The die's PPM may also check power allocations to one or more prioritized processing threads against the available power (eg, current) budget and thus avoid exceeding the budget even if some processing threads are prioritized over others.

[0027] In at least some embodiments, the die's PPM can also manage the shift between non-prioritized management of a subset of the plurality of processing threads and prioritized management of the subset of the plurality of processing threads. For example, the PPM can start a timer while a non-prioritized processing thread (e.g., an erase operation or a program operation) is running and in response to detecting that power is also being allocated to a prioritized thread (e.g., a read operation or a program operation) among the plurality of processing threads. The timer can track a predetermined amount of time to ensure that the prioritized allocation request does not starve the non-prioritized processing thread of the current required to complete processing. Therefore, if the timer expires while the non-prioritized processing thread is still running, the control logic can force a transition back to allocating power among the subset of processing threads based on an increment in the value of the non-priority counter, e.g., based on the non-prioritized power allocation.

[0028] Advantages of this approach include, but are not limited to, an efficient power management scheme for a multi-die memory subsystem, where each memory die supports multiple processing threads operating concurrently. The disclosed techniques allow for support of independent parallel plane access in a memory device with significantly reduced hardware resources in the memory subsystem. This approach is highly scalable as the number of processing threads increases and does not rely on external controller intervention. Furthermore, by prioritizing the power allocation of some processing threads depending on which subset of the multiple processing threads is being managed, memory operations that should be processed quickly (e.g., read operations and some programming operations) can be prioritized over memory operations that can be executed more slowly (e.g., erase operations and some programming operations). The power allocated to the prioritized processing threads can still be checked against the power (e.g., current) budget to ensure that the budget is not exceeded. Furthermore, as mentioned, the use of a timer and shift protocol ensures that non-prioritized processing threads do not starve of power budget. As a result, the overall performance and quality of service provided by each memory die are improved.

[0029] Figure 1A An example computing system 100 is shown that includes a memory subsystem 110, in accordance with at least some embodiments. Memory subsystem 110 can include media, such as one or more volatile memory devices (e.g., memory device 140), one or more non-volatile memory devices (e.g., one or more memory devices 130), or a combination of such media or memory devices. Memory subsystem 110 can be a storage device, a memory module, or a hybrid of storage devices and memory modules.

[0030] Memory device 130 may be a non-volatile memory device. One example of a non-volatile memory device is a NAND memory device. A non-volatile memory device is a package of one or more dies or logical units (LUNs). Thus, each memory device 130 may be a die (or LUN) or may be a multi-die package including multiple dies (or LUNs) on a chip, such as an integrated circuit package for dies. Each die may include one or more planes. Planes may be grouped into logical units (LUNs). For some types of non-volatile memory devices (e.g., NAND devices), each plane includes a set of physical blocks. Each block includes a set of pages. Each page includes a set of memory cells ("cells"). A cell is an electronic circuit that stores information. Depending on the cell type, a cell may store one or more binary bits of information and have various logical states related to the number of bits stored. A logical state may be represented by a binary value (e.g., "0" and "1") or a combination of these values.

[0031] Each memory device 130 may be composed of bits arranged in a two-dimensional or three-dimensional grid, also referred to as a memory array. Memory cells are etched onto a silicon wafer in an array of columns (hereinafter also referred to as bit lines) and rows (hereinafter also referred to as word lines). A word line may refer to one or more rows of memory cells of a memory device that are used in conjunction with one or more bit lines to generate an address for each of the memory cells. The intersection of a bit line and a word line constitutes the address of the memory cell.

[0032] The memory subsystem 110 may be a storage device, a memory module, or a combination of a storage device and a memory module. Examples of storage devices include solid-state drives (SSDs), flash drives, universal serial bus (USB) flash drives, embedded multimedia controller (eMMC) drives, universal flash storage (UFS) drives, secure digital (SD) cards, and hard disk drives (HDDs). Examples of memory modules include dual inline memory modules (DIMMs), small outline DIMMs (SO-DIMMs), and various types of non-volatile dual inline memory modules (NVDIMMs).

[0033] The computing system 100 can be a computing device, such as a desktop computer, a laptop computer, a network server, a mobile device, a vehicle (e.g., an airplane, drone, train, car, or other transportation vehicle), an Internet of Things (IoT)-enabled device, an embedded computer (e.g., a computer included in a vehicle, industrial equipment, or a networked commercial device), or such a computing device that includes a memory and a processing device.

[0034] The computing system 100 may include a host system 120 coupled to one or more memory subsystems 110. In some embodiments, the host system 120 is coupled to multiple memory subsystems 110 of different types. Figure 1A An example of a host system 120 coupled to one memory subsystem 110 is shown. The host system 120 can provide data to be stored at the memory subsystem 110 and can request data to be retrieved from the memory subsystem 110. As used herein, "coupled to," "coupled with," or "operatively coupled with" generally refers to a connection between components, which can be an indirect communication connection or a direct communication connection (e.g., without intervening components), whether wired or wireless, including connections such as electrical, optical, magnetic, and the like.

[0035] The host system 120 may include a processor chipset and a software stack executed by the processor chipset. The processor chipset may include one or more cores, one or more caches, a memory controller (e.g., an NVDIMM controller), and a storage protocol controller (e.g., a PCIe controller, a SATA controller). The host system 120 uses the memory subsystem 110, for example, to write data to the memory subsystem 110 and read data from the memory subsystem 110.

[0036] The host system 120 can be coupled to the memory subsystem 110 via a physical host interface. Examples of the physical host interface include, but are not limited to, a Serial Advanced Technology Attachment (SATA) interface, a Peripheral Component Interconnect Express (PCIe) interface, a Universal Serial Bus (USB) interface, a Fibre Channel interface, a Serial Attached SCSI (SAS), a Double Data Rate (DDR) memory bus, a Small Computer System Interface (SCSI), a Dual In-line Memory Module (DIMM) interface (e.g., a DIMM socket supporting Double Data Rate (DDR)), and the like. The physical host interface can be used to transfer data between the host system 120 and the memory subsystem 110. When the memory subsystem 110 is coupled to the host system 120 via a physical host interface (e.g., a PCIe bus), the host system 120 can further utilize an NVM Express (NVMe) interface to access components (e.g., one or more memory devices 130). The physical host interface can provide an interface for passing control, address, data, and other signals between the memory subsystem 110 and the host system 120. Figure 1A Memory subsystem 110 is shown as an example. In general, host system 120 can access multiple memory subsystems via the same communication connection, multiple separate communication connections, and / or a combination of communication connections.

[0037] Memory devices 130 and 140 may include any combination of different types of non-volatile memory devices and / or volatile memory devices. Volatile memory devices (e.g., memory device 140) may be, but are not limited to, random access memory (RAM), such as dynamic random access memory (DRAM) and synchronous dynamic random access memory (SDRAM).

[0038] Some examples of non-volatile memory devices (e.g., memory device 130) include NAND-type flash memory and write-in-place memory, such as a three-dimensional cross-point ("3D cross-point") memory device, which is a cross-point array of non-volatile memory cells. The cross-point array of non-volatile memory cells can be combined with a stackable cross-grid data access array to perform bit storage based on changes in bulk resistance. In addition, compared to many flash-based memories, cross-point non-volatile memory can perform write-in-place operations, in which non-volatile memory cells can be programmed without first erasing the non-volatile memory cells. NAND-type flash memory includes, for example, two-dimensional NAND (2D NAND) and three-dimensional NAND (3D NAND).

[0039] Each of the memory devices 130 may include one or more memory cell arrays. One type of memory cell, such as a single-level cell (SLC), may store one bit per cell. Other types of memory cells, such as multi-level cells (MLC), triple-level cells (TLC), quad-level cells (QLC), and penta-level cells (PLC), may store multiple bits per cell, for example, by means of additional threshold voltage ranges. In some embodiments, each of the memory devices 130 may include one or more memory cell arrays, such as SLC, MLC, TLC, QLC, PLC, or any combination of these. In some embodiments, a particular memory device may include an SLC portion of memory cells, as well as an MLC portion, a TLC portion, a QLC portion, or a PLC portion. The memory cells of the memory device 130 may be grouped into pages, which may refer to logical units of the memory device for storing data. In the case of some types of memory (e.g., NAND), pages may be grouped to form blocks.

[0040] Although non-volatile memory components such as a 3D cross-point array of non-volatile memory cells and NAND-type flash memory (e.g., 2D NAND, 3D NAND) are described, the memory device 130 may be based on any other type of non-volatile memory, such as read-only memory (ROM), phase-change memory (PCM), selectable memory, other chalcogenide-based memory, ferroelectric transistor random access memory (FeTRAM), ferroelectric random access memory (FeRAM), magnetic random access memory (MRAM), spin transfer torque (STT)-MRAM, conductive bridging RAM (CBRAM), resistive random access memory (RRAM), oxide-based RRAM (OxRAM), or non-OR flash memory, or electrically erasable programmable read-only memory (EEPROM).

[0041] The memory subsystem controller 115 (or, for simplicity, controller 115) can communicate with the memory device 130 to perform operations, such as reading data, writing data, or erasing data at the memory device 130, and other such operations. The memory subsystem controller 115 can include hardware, such as one or more integrated circuits and / or discrete components, buffer memory, or a combination thereof. The hardware can include digital circuitry with dedicated (i.e., hard-coded) logic to perform the operations described herein. The memory subsystem controller 115 can be a microcontroller, dedicated logic circuitry (e.g., a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), etc.), or other suitable processor.

[0042] The memory subsystem controller 115 may include a processing device including one or more processors (e.g., processor 117) configured to execute instructions stored in local memory 119. In the example shown, the local memory 119 of the memory subsystem controller 115 includes embedded memory configured to store instructions for performing various processes, operations, logic flows, and routines that control the operation of the memory subsystem 110, including handling communications between the memory subsystem 110 and the host system 120.

[0043] In some embodiments, local memory 119 may include memory registers that store memory pointers, fetched data, etc. Local memory 119 may also include read-only memory (ROM) for storing microcode. Figure 1A The example memory subsystem 110 in FIG. 1 has been shown as including a memory subsystem controller 115, but in another embodiment of the present disclosure, the memory subsystem 110 does not include a memory subsystem controller 115 and may instead rely on external control (e.g., provided by an external host or by a processor or controller separate from the memory subsystem).

[0044] Typically, the memory subsystem controller 115 may receive commands or operations from the host system 120 and convert the commands or operations into instructions or appropriate commands to achieve the desired access to the memory device 130. The memory subsystem controller 115 may be responsible for other operations such as wear leveling operations, garbage collection operations, error detection and error correction code (ECC) operations, encryption operations, cache operations, and address translation between logical addresses (e.g., logical block addresses (LBAs), namespaces) and physical addresses (e.g., physical block addresses) associated with the memory device 130. The memory subsystem controller 115 may further include host interface circuitry to communicate with the host system 120 via a physical host interface. The host interface circuitry may convert commands received from the host system into command instructions to access the memory device 130 and convert responses associated with the memory device 130 into information for the host system 120.

[0045] The memory subsystem 110 may also include additional circuitry or components not shown. In some embodiments, the memory subsystem 110 may include a cache or buffer (e.g., DRAM) and address circuitry (e.g., row decoders and column decoders) that may receive addresses from the memory subsystem controller 115 and decode the addresses to access the memory device 130.

[0046] In some embodiments, memory device 130 includes a local media controller 135 that operates in conjunction with memory subsystem controller 115 to perform operations on one or more memory cells of memory device 130. An external controller (e.g., memory subsystem controller 115) can externally manage memory device 130 (e.g., perform media management operations on memory device 130). In some embodiments, memory subsystem 110 is a managed memory device, which is a raw memory device with on-die control logic (e.g., local media controller 135) and a controller for media management (e.g., memory subsystem controller 115) within the same memory device package. An example of a managed memory device is a managed NAND (MNAND) device.

[0047] In at least some embodiments, each memory device 130 includes a peak power manager (PPM) wrapper 150, which includes a thread manager 155 and a PPM 160 (peak power manager). In one embodiment, the local media controller 135 of each memory device 130 includes at least a portion of the PPM wrapper 150. In such embodiments, the PPM wrapper 150 may be implemented using hardware or as firmware, stored on each memory device 130, and executed by control logic (e.g., the local media controller 135) to perform operations related to prioritized power budget arbitration for multiple concurrent access operations as described herein. In some embodiments, the memory subsystem controller 115 includes at least a portion of the PPM wrapper 150. For example, the memory subsystem controller 115 may include a processor 117 (e.g., a processing device) configured to execute instructions stored in the local memory 119 for performing the operations described herein.

[0048] In at least some embodiments, the PPM wrapper 150 can manage power prioritization budget arbitration for multiple concurrent access operations in the memory device 130. In one embodiment, the memory subsystem 110 employs a token-based protocol in which a token rotates (e.g., in a round-robin manner) among the multiple PPM wrappers 150 of the multiple memory dies (e.g., after a set number of cycles of a shared clock signal). When the PPM wrapper 150 of a die holds the token, the PPM wrapper 150 can determine the power (e.g., current) requested by multiple processing threads (e.g., implemented by the local media controller 135) of each memory device 130, select one or more prioritized processing threads from the multiple processing threads based on the available power budget in the memory subsystem, request the power from a shared current source in the memory subsystem 110, and allocate the requested power to the selected processing threads. In some embodiments, if power is allocated to all prioritized processing threads and the budget remains unchanged, the PPM wrapper 150 can also allocate power to non-prioritized processing threads. PPM wrapper 150 may further broadcast the quantized current budget to be consumed by the memory die during a given time period so that other PPM wrappers in memory subsystem 110 are aware of the available power budget. Additional details regarding the operation of each PPM wrapper 150 are described below.

[0049] Figure 1B is a first device in the form of one or more memory devices 130 according to an embodiment and a second device in the form of a memory subsystem (e.g., Figure 1A1 is a simplified block diagram of a second device communicating with a memory subsystem controller 115 of a memory subsystem 110. Some examples of electronic systems include personal computers, personal digital assistants (PDAs), digital cameras, digital media players, digital recorders, games, electrical appliances, vehicles, wireless devices, mobile phones, etc. The memory subsystem controller 115 (e.g., a controller external to each memory device 130) can be a memory controller or other external host device.

[0050] Each memory device 130 includes a memory cell array 104 that is logically arranged in rows and columns. Memory cells in logical rows are typically connected to the same access line (e.g., word line), while memory cells in logical columns are typically selectively connected to the same data line (e.g., bit line). A single access line can be associated with more than one logical row of memory cells, and a single data line can be associated with more than one logical column. Memory cells ( Figure 1B (not shown) can be programmed to one of at least two target data states.

[0051] Row decoding circuitry 108 and column decoding circuitry 111 are provided to decode address signals. Address signals are received and decoded to access memory cell array 104. Each memory device 130 also includes input / output (I / O) control circuitry 112, which manages the input of commands, addresses, and data to memory devices 130, as well as the output of data and status information from each memory device 130. Address registers 114 communicate with I / O control circuitry 112 and row decoding circuitry 108 and column decoding circuitry 111 to latch address signals prior to decoding. Command registers 124 communicate with I / O control circuitry 112 and local media controller 135 to latch incoming commands.

[0052] A controller (e.g., a local media controller 135 within each memory device 130) controls access to the memory cell array 104 in response to commands and generates status information for the external memory subsystem controller 115. That is, the local media controller 135 is configured to perform access operations (e.g., read operations, program operations, and / or erase operations) on the memory cell array 104. The local media controller 135 communicates with the row decoding circuitry 108 and the column decoding circuitry 111 to control the row decoding circuitry 108 and the column decoding circuitry 111 in response to addresses.

[0053] The local media controller 135 also communicates with cache registers 118 and data registers 121. Cache registers 118 latch incoming or outgoing data as directed by the local media controller 135 to temporarily store data while the memory cell array 104 is busy writing or reading other data, respectively. During a programming operation (e.g., a write operation), data can be transferred from cache registers 118 to data registers 121 for transfer to the memory cell array 104; the new data can then be latched in cache registers 118 from the I / O control circuitry 112. During a read operation, data can be transferred from cache registers 118 to the I / O control circuitry 112 for output to the memory subsystem controller 115; the new data can then be transferred from data registers 121 to cache registers 118. Cache registers 118 and / or data registers 121 can form (e.g., can form at least a portion of) a page buffer for each memory device 130. The page buffer may further include a sensing device, such as a sense amplifier, to sense the data state of the memory cells of the memory cell array 104, for example, by sensing the state of the data lines connected to the memory cells. The status register 122 may communicate with the I / O control circuitry 112 and the local memory controller 135 to latch status information for output to the memory subsystem controller 115.

[0054] Each memory device 130 receives control signals from the memory subsystem controller 115 from the local media controller 135 via a control link 132. For example, the control signals may include a chip enable signal CE#, a command latch enable signal CLE, an address latch enable signal ALE, a write enable signal WE#, a read enable signal RE#, and a write protect signal WP#. Depending on the nature of each memory device 130, additional or alternative control signals (not shown) may be further received via the control link 132. In one embodiment, each memory device 130 receives a command signal (which represents a command), an address signal (which represents an address), and a data signal (which represents data) from the memory subsystem controller 115 via a multiplexed input / output (I / O) bus 134, and outputs the data to the memory subsystem controller 115 via the I / O bus 134.

[0055] For example, a command may be received at I / O control circuitry 112 via input / output (I / O) pins [7:0] of I / O bus 134 and may then be written into command register 124. An address may be received at I / O control circuitry 112 via input / output (I / O) pins [7:0] of I / O bus 134 and may then be written into address register 114. Data may be received at I / O control circuitry 112 via input / output (I / O) pins [7:0] for 8-bit devices or input / output (I / O) pins [15:0] for 16-bit devices and may then be written into cache register 118. The data may then be written into data register 121 for use in programming memory cell array 104.

[0056] In an embodiment, cache register 118 may be omitted and data may be written directly to data register 121. Data may also be output via input / output (I / O) pins [7:0] for 8-bit devices or input / output (I / O) pins [15:0] for 16-bit devices. Although reference may be made to I / O pins, these may include any conductive nodes, such as conventional conductive pads or conductive bumps, that enable electrical connection to each memory device 130 by an external device (e.g., memory subsystem controller 115).

[0057] Those skilled in the art will appreciate that additional circuitry and signals may be provided and simplified. Figure 1B Each memory device 130. It should be recognized that the reference Figure 1B The functionality of the various block components described may not necessarily be separated into different components or component portions of an integrated circuit device. For example, a single component or component portion of an integrated circuit device may be adapted to perform Figure 1B Alternatively, one or more components or component parts of an integrated circuit device may be combined to perform Figure 1B Additionally, while specific I / O pins are described according to prevailing conventions for the reception and output of various signals, it should be noted that other combinations of I / O pins (or other I / O node structures) or other numbers of I / O pins (or other I / O node structures) may be used in various embodiments.

[0058] Figure 2is a block diagram illustrating a multi-die package having multiple memory dies in a memory subsystem according to at least some embodiments. As shown, the multi-die package 200 includes any one of the memory dies 230(0) through 230(7). However, in other embodiments, the multi-die package 200 may include some other number of memory dies, such as additional or fewer memory dies. In at least one embodiment, the multi-die package 200 is about Figures 1A to 1B At least one of the memory devices 130 shown and discussed. In one embodiment, the memory dies 230(0) to 230(7) share a clock signal ICLK received via a clock signal line. The memory dies 230(0) to 230(7) can be selectively enabled in response to a chip enable signal (e.g., via a control link) and can communicate via separate I / O buses. Additionally, a peak current magnitude indicator signal HC# is typically shared between the memory dies 230(0) to 230(7). The peak current magnitude indicator signal HC# can typically be pulled to a particular state (e.g., pulled high). In one embodiment, each of the memory dies 230(0) to 230(7) includes an instance of a PPM wrapper 150 that receives the clock signal ICLK and the peak current magnitude indicator signal HC#.

[0059] In one embodiment, a token-based protocol is used in which a token is circulated through each of the memory dies 230(0) to 230(7) for determining and broadcasting the expected peak current magnitude, even if some of the memory dies 230(0) to 230(7) may be disabled in response to their respective chip enable signals. The time period (e.g., a certain number of cycles of the clock signal ICLK) that a given PPM wrapper 150 holds this token may be referred to herein as a power management cycle for the associated memory die. At the end of the power management cycle, the token is passed to the next memory die in sequence. Eventually, the token is received again by the same PPM wrapper 150, which signals the start of a new power management cycle for the associated memory die. In one embodiment, the encoded value of the lowest expected peak current magnitude is configured such that each of its digits corresponds to a normal logic level of the peak current magnitude indicator signal HC#, wherein a disabled die does not transition the peak current magnitude indicator signal HC#. However, in other embodiments, the memory die may be configured to drive transitions of the peak current magnitude indicator signal HC# when otherwise disabled in response to its corresponding chip enable signal to indicate an encoded value of the lowest expected peak current magnitude when designated.

[0060] When a given PPM wrapper 150 holds a token, it may determine the peak current magnitude of a respective one of the memory dies 230(0)-230(7) (attributable to one or more processing threads on that memory die) and broadcast an indication of the peak current magnitude via the peak current magnitude indicator signal HC#. As described in greater detail below, during a given power management cycle, the PPM wrapper 150 may use one of several different arbitration schemes to arbitrate among the multiple processing threads on the respective memory die in order to distribute the peak current to enable concurrent memory access operations.

[0061] Figure 3 is a block diagram illustrating a multi-plane memory device 130A configured for independent parallel plane access according to at least some embodiments. In at least one embodiment, the multi-plane memory device 130A is a block diagram illustrating a multi-plane memory device 130A configured for independent parallel plane access according to at least some embodiments. Figures 1A to 1B At least one of the memory devices 130 shown and discussed. Memory planes 372(0) through 372(3) can each be divided into data blocks, wherein different relative data blocks from two or more of the memory planes 372(0) through 372(3) can be accessed concurrently during a memory access operation. For example, during a memory access operation, two or more of data block 382 of memory plane 372(0), data block 383 of memory plane 372(1), data block 384 of memory plane 372(2), and data block 385 of memory plane 372(3) can each be accessed concurrently.

[0062] Memory device 130A includes a memory array 370 divided into memory planes 372(0) through 372(3), each of which includes a respective number of memory cells. Multi-plane memory device 130A may further include a local media controller 135 including power control circuitry and access control circuitry for concurrently performing memory access operations for different memory planes 372(0) through 372(3). The memory cells may be nonvolatile memory cells, such as NAND flash cells, or may generally be any type of memory cell.

[0063] The memory planes 372(0) through 372(3) can each be divided into data blocks, wherein different relative data blocks from each of the memory planes 372(0) through 372(3) can be accessed concurrently during a memory access operation. For example, during a memory access operation, data block 382 of memory plane 372(0), data block 383 of memory plane 372(1), data block 384 of memory plane 372(2), and data block 385 of memory plane 372(3) can each be accessed concurrently.

[0064] Each of the memory planes 372(0) to 372(3) can be coupled to a corresponding page buffer 376(0) to 376(3). Each page buffer 376(0) to 376(3) can be configured to provide data to or receive data from the corresponding memory plane 372(0) to 372(3). The page buffers 376(0) to 376(3) can be controlled by the local media controller 135. Data received from the corresponding memory planes 372(0) to 372(3) can be latched at the page buffers 376(0) to 376(3), respectively, and retrieved by the local media controller 135 and provided to the memory subsystem controller 115 via the NVMe interface.

[0065] Each of the memory planes 372(0) to 372(3) can be further coupled to a corresponding access driver circuit 374(0) to 374(3), such as an access line driver circuit. The driver circuits 374(0) to 374(3) can be configured to condition the pages of the corresponding block of the associated memory plane 372(0) to 372(3) for memory access operations, such as programming (i.e., writing data), reading data, or erasing data. Each of the driver circuits 374(0) to 374(3) can be coupled to a corresponding global access line associated with the corresponding memory plane 372(0) to 372(3). Each of the global access lines can be selectively coupled to a corresponding local access line within the block of the plane during a memory access operation associated with a page within the block. The driver circuits 374(0) to 374(3) can be controlled based on signals from the local media controller 135. Each of the driver circuits 374(0) to 374(3) may include or be coupled to a respective power supply circuit and may provide a voltage to the respective access line based on a voltage provided by the respective power supply circuit. The voltage provided by the power supply circuit may be based on a signal received from the local media controller 135.

[0066] The local media controller 135 may control the driver circuits 374(0) to 374(3) and the page buffers 376(0) to 376(3) to concurrently perform memory access operations associated with each of a set of memory command and address pairs (e.g., received from the memory subsystem controller 115). For example, the local media controller 135 may control the driver circuits 374(0) to 374(3) and the page buffers 376(0) to 376(3) to perform concurrent memory access operations. The local media controller 135 may include: a power control circuit that serially configures two or more of the driver circuits 374(0) to 374(3) for concurrent memory access operations; and an access control circuit that is configured to control two or more of the page buffers 376(0) to 376(3) to sense and latch data from the corresponding memory planes 372(0) to 372(3), or program data to the corresponding memory planes 372(0) to 372(3) to perform concurrent memory access operations.

[0067] In operation, the local media controller 135 may receive a set of memory command and address pairs via the NVMe bus, where each pair arrives in parallel or serially. In some instances, the set of memory command and address pairs may each be associated with a different corresponding memory plane 372(0) to 372(3) of the memory array 370. The local media controller 135 may be configured to perform concurrent memory access operations (e.g., read operations or program operations) on the different memory planes 372(0) to 372(3) of the memory array 370 in response to the set of memory command and address pairs. For example, the power control circuitry of the local media controller 135 may serially configure the driver circuits 374(0) to 374(3) of two or more memory planes 372(0) to 372(3) associated with the set of memory command and address pairs based on corresponding page types (e.g., UP, MP, LP, XP, SLC / MLC / TLC / QLC pages) for concurrent memory access operations. After the access line driver circuits 374(0) to 374(3) have been configured, the access control circuitry of the local media controller 135 may concurrently control the page buffers 376(0) to 376(3) to access a corresponding page of each of the two or more memory planes 372(0) to 372(3) associated with the set of memory command and address pairs during concurrent memory access operations, such as retrieving data or writing data. For example, the access control circuitry may concurrently (e.g., in parallel and / or simultaneously) control the page buffers 376(0) to 376(3) to charge / discharge bit lines, sense data from the two or more memory planes 372(0)-372(3), and / or latch the data.

[0068] Based on signals received from the local media controller 135, driver circuits 374(0) to 374(3) coupled to the memory planes 372(0) to 372(3) associated with the set of memory command and address command pairs can select memory blocks or memory cells from the associated memory planes 372(0) to 372(3) for memory operations, such as read, program, and / or erase operations. The driver circuits 374(0) to 374(3) can drive different corresponding global access lines associated with the corresponding memory planes 372(0) to 372(3). As an example, driver circuit 374(0) can drive a first voltage on a first global access line associated with memory plane 372(0), driver circuit 374(1) can drive a second voltage on a third global access line associated with memory plane 372(1), driver circuit 374(2) can drive a third voltage on a seventh global access line associated with memory plane 372(2), and so on, and can drive other voltages on each of the remaining global access lines. In some examples, a pass voltage may be provided on all access lines except the access lines associated with the page of the memory plane 372(0)-372(3) to be accessed. The local media controller 135, the driver circuits 374(0)-374(3), and the page buffers 376(0)-376(3) may allow concurrent access to different corresponding pages within different corresponding blocks of memory cells. For example, a first page of a first block of a first memory plane and a second page of a second block of a second memory plane may be concurrently accessed, regardless of page type.

[0069] The page buffers 376(0) to 376(3) may provide data to or receive data from the local media controller 135 during a memory access operation in response to signals from the local media controller 135 and the corresponding memory planes 372(0) to 372(3). The local media controller 135 may provide the received data to the memory subsystem controller 115.

[0070] 4 (3) . As will be appreciated, the memory device 130A may include more or less than four memory planes, driver circuits, and page buffers. It will also be appreciated that the corresponding global access lines may include 8, 16, 32, 64, 128, etc. global access lines. When the different corresponding pages belong to different page types, the local media controller 135 and the driver circuits 374(0) to 374(3) may concurrently access different corresponding pages within different corresponding blocks of different memory planes. For example, the local media controller 135 may include several different processing threads, such as processing threads 334(0) to 334(3). Each of the processing threads 334(0) to 334(3) may be associated with a corresponding one of the memory planes 372(0) to 372(3) and may manage operations performed on the corresponding plane. For example, each of the processing threads 334(0)-334(3) can provide control signals to a respective one of the driver circuits 374(0)-374(3) and the page buffers 376(0)-376(3) to perform those memory access operations concurrently (e.g., at least partially overlapping in time). Because the processing threads 334(0)-334(3) can perform memory access operations, each of the processing threads 334(0)-334(3) can have different current demands at different points in time. According to the techniques described herein, the PPM wrapper 150 can determine the power budget requirements of the processing threads 334(0)-334(3) in a given power management cycle and identify one or more of the processing threads 334(0)-334(3) using one of the several power budget arbitration schemes described herein. The one or more processing threads 334(0)-334(3) can be determined during a power management cycle based on the available power budget in the memory subsystem 110. For example, PPM wrapper 150 may determine respective priorities of processing threads 334 ( 0 ) through 334 ( 3 ) and allocate current to processing threads 334 ( 0 ) through 334 ( 3 ) based on the respective priorities.

[0071] Figure 4 is a block diagram illustrating a memory die 400 configured for power budget arbitration for multiple processing threads according to at least some embodiments. In some embodiments, the memory die 400 includes control logic, such as a PPM wrapper 150, which in turn includes a reference Figure 1AThread manager 155 and PPM 160 are discussed. Memory die 400 further includes memory 456, such as registers, DRAM, SDRAM, etc., although in some embodiments, memory 456 may also refer to memory array 370. In these embodiments, thread manager 155 includes request register 452 (or other internal PPM memory) and includes or is coupled to timer 478. Thread manager 155 may further include non-priority ring counter 444 and priority ring counter 454 coupled to memory 456, e.g., for accessing data structure 448 and data structure 458, respectively. PPM 160 is coupled to thread manager 155 and may also be coupled to memory 456, and receives clock signal ICLK and peak current magnitude indicator signal HC#, as previously described.

[0072] In some embodiments, the thread manager 155 identifies one or more processing threads, such as the plurality of processing threads 434(0)-434(3) in the memory die 400, and requests the PPM 160 to determine during a power management cycle whether the available current (e.g., power) budget can support running the one or more processing threads based on the amount of power associated with the one or more processing threads. More specifically, because the plurality of processing threads 434(0)-434(3) can generate different requests asynchronously, to manage this complexity, the thread manager 155 can manipulate and aggregate these asynchronous requests in a simplified number of requests to the PPM 160. In some embodiments, the set of simplified requests sent to the PPM 160 can contain randomized thread requests to ensure fairness in allocating current to the plurality of processing threads 434(0)-434(3). As will be discussed in more detail, the randomization of the requests sent by the thread manager 155 to the PPM 160 can be performed by the non-priority ring counter 444, by the priority ring counter 454, or can be shifted between the two ring counters, as will be discussed. In some embodiments, the plurality of processing threads 434(0) to 434(3) corresponds to the processing threads 334(0) to 334(3) ( Figure 3 ).

[0073] In some embodiments, the PPM 160 periodically asserts a polling window signal 460 that is received by the thread manager 155. The polling window signal 460 is asserted after the previous power management cycle ends (e.g., when the PPM 160 relinquishes the token) and before the subsequent power management cycle begins (e.g., when the PPM 160 again receives the token). Because the processing threads 434(0) to 434(3) periodically issue requests for current depending on the associated processing operations, the thread manager 155 stores or buffers the received requests in the request register 452 during the period when the polling window signal 460 is asserted. Although the requests are generally referred to herein as requesting a current allocation, this should generally be understood to mean requesting power, and may also include requesting a voltage allocation, for example.

[0074] In some embodiments, the PPM 160 tracks the token and can determine (e.g., based on the synchronous clock signal ICLK) when the token will be received and can de-assert the poll window signal 460 before that time. In response to the poll window signal 460 being de-asserted (i.e., during a subsequent power management cycle), the thread manager 155 can stop storing additional requests in the request register 452, making the contents of the request register 452 static. Any new requests will not be considered during this cycle, but will be saved and can be considered in a subsequent power management cycle. The thread manager 155 can generate multiple current level signals, such as a full signal 462, an intermediate signal 464, a low signal 466, and a high-to-low signal 468, each of which corresponds to a current associated with a corresponding set of at least one of the requests in the request register 452. For example, full signal 462 may represent the sum of all current requests in request register 452, intermediate signal 464 may represent the sum of two or more but less than all current requests in request register 452 (e.g., the first two or more requests in request register 452), low signal 466 may represent one current request from request register 452 (e.g., the first request in request register 452), and high-to-low signal 468 may represent a low current request when a high current budget has been allocated. High-to-low signal 468 may be associated with a request in request register 452, to which PPM 160 will immediately allocate current without checking against the current budget, and such allocation will be tracked as it tracks other current allocations. By polling processing threads between power management cycles, thread manager 155 can save significant time and processing resources compared to waiting until a token is actually received.

[0075] In these embodiments, the PPM 160 receives the full signal 462, the intermediate signal 464, the low signal 466, and the high-to-low signal 468 and determines whether the amount of current available in the memory subsystem 110 during the current power management cycle can satisfy the amount of current associated with any of the current level signals. In response to the available amount of current satisfying at least one of the current level signals, the PPM 160 may request the amount of current and provide a grant signal 472, such as an acknowledgment, to the thread manager 155. For example, the grant signal 472 may indicate which current level signal the available amount of current satisfies. Accordingly, the thread manager 155 may grant one or more of the processing threads 434(0) to 434(3) authorization to perform one or more memory access operations corresponding to the requests in the request register 452, the grant signal 472 granting the request based on the memory access operation.

[0076] Figure 5 is shown in accordance with some embodiments by reference Figure 4 A block diagram of the operation of the non-priority ring counter 444 implemented by the thread manager 155 of the memory die discussed herein. In one embodiment, the non-priority ring counter 444 is formed in the PPM 160 using flip-flops or other devices connected in a shift register so that the output of the last flip-flop feeds into the input of the first flip-flop to form a circular or "ring" structure. In one embodiment, the non-priority ring counter 444 is a 2 n n-bit counters with different states, 2 n Indicates the number of different processing threads, such as processing threads 434(0) through 434(3) in memory device 130 or 130A. In some embodiments, priority ring counter 454 functions similarly to non-priority ring counter 444, but the incremented value of priority ring counter 454 can track a subset of prioritized processing threads, as will be discussed in more detail.

[0077] like Figure 5As shown, by way of example, the non-priority ring counter 444 is a 2-bit counter that represents four different states (i.e., state 0 502, state 1 504, state 2 506, and state 3 508). In operation, the non-priority ring counter 444 sequentially cycles through each of the four states 502 to 508 in response to changes in the power management cycle. For example, if the non-priority ring counter 444 is initially set to state 0 402, then when the PPM 160 receives a token, the value of the non-priority ring counter 444 is incremented (e.g., by 1), causing the non-priority ring counter 444 to shift to state 1 504. Similarly, the next time the PPM 160 receives a token, the value is incremented again, causing the ring counter to shift to state 2 506, and so on. When set to state 3 508 and the value is incremented, the non-priority ring counter 444 will return to state 0 502. As described in more detail below, each state (or value) of non-priority ring counter 444 is associated with one or more processing threads, thereby allowing thread manager 155 to select one or more processing threads of the memory device based on the current state of non-priority ring counter 444. Thus, the subset of the plurality of processing threads sent to PPM 160 varies according to non-priority ring counter 444, thereby allowing PPM 160 to handle power allocation for a smaller number of threads at a time. This simplifies the control logic of PPM 160. Furthermore, the function of non-priority ring counter 444 is to rotate all processing threads equally after passing through the four values ​​or states of non-priority ring counter 444.

[0078] More specifically, Table 1 is an example of a data structure 448 of the PPM wrapper 150 for power budget arbitration of multiple processing threads in the memory device 130 or 130A. In one embodiment, the data structure 448 is formed in or managed by the thread manager 155 using a lookup table, array, linked list, record, object, or some other data structure. In one embodiment, the data structure 448 includes multiple entries, each corresponding to one of the states of the non-priority ring counter 444. For example, for each state of the non-priority ring counter 444, the data structure 448 may identify a leading thread and a thread combination. The leading thread may be a single processing thread with the highest priority when the non-priority ring counter 444 is in the corresponding state, and the thread combination may be a set of two or more processing threads, but fewer than all processing threads, whose priority, when the non-priority ring counter 444 is in the corresponding state, is higher than that of other threads not in the set, but lower than that of the leading thread.

[0079]

[0080]

[0081] Table 1

[0082] In some embodiments, to allocate the available power budget during a power management cycle, the thread manager 155 may determine the current state of the non-priority ring counter 444 and, based on the data structure 448, determine the leading thread and thread combination corresponding to the current state of the non-priority ring counter 444. The thread manager 155 may then send a request, such as the full signal 462, the intermediate signal 464, and the low signal 466, to the PPM 160 based on the identification of the leading thread and thread combination. In response to the amount of current available in the memory subsystem during the power management cycle satisfying the amount of current associated with at least one of the leading thread or thread combination, the PPM 160 may request the amount of current associated with at least one of the leading thread or thread combination and allocate the current budget accordingly.

[0083] By way of additional example, Table 2 illustrates a data structure 448 in which a 3-bit non-priority counter 444 can hold up to eight states or values, and thus the data structure 448 can store additional combinations of possible leading threads and thread combinations. The 3-bit example of Table 2 and the other tables included below are exemplary only for purposes of explanation, as other tables, including 4-bit and higher, are contemplated. In Table 2, the "Reg_hc_max" value corresponds to the full signal 462, the "Reg-hc_middle" value corresponds to the middle signal 464, and the "Reg-hc_min" value corresponds to the low signal 466.

[0084]

[0085] Table 2

[0086] In this example, full signal 462 may correspond to all of the plurality of processing threads requesting current, e.g., the main processing thread of local media controller 135, and five additional coprocessors ("coprocs") that may be associated with individual additional threads, e.g., thread 0 434(0) through thread 3 434(3), although more coprocessors are contemplated. Furthermore, intermediate signal 464 may include a subset of the plurality of processing threads (e.g., a combination of threads), and low signal 466 may include only one processing thread (e.g., a leading thread) of the plurality of processing threads. In one embodiment, the subset of the plurality of processing threads does not exceed half of the plurality of processing threads.

[0087] Figure 6is a flow chart of an example method 600 for performing power budget arbitration in a memory device using a ring counter, according to at least some embodiments. The method 600 may be performed by processing logic that may include hardware (e.g., a processing device, circuitry, dedicated logic, programmable logic, microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executed on a processing device), or a combination thereof. In some embodiments, the method 600 is performed by Figure 1A and Figure 4 The PPM wrapper 150 of FIG. 10 illustrates a specific sequence or order, but the order of the processes may be modified unless otherwise specified. Therefore, the illustrated embodiments should be understood as examples only, and the illustrated processes may be performed in a different order, and some processes may be performed in parallel. Furthermore, one or more processes may be omitted in various embodiments. Therefore, not all processes are required in every embodiment. Other process flows are possible.

[0088] At operation 605, power requests are sampled. For example, processing logic (e.g., PPM 160) may sample power requests (e.g., current requests or peak current magnitude requests) from one or more processing threads (e.g., processing threads 334(0) to 334(3)) of a memory device. In one embodiment, in response to PPM 160 receiving a token signaling the start of a current power management cycle, PPM 160 alerts thread manager 155 that a polling window 460 has begun. In response, thread manager 155 sends a polling request to each of the processing threads to obtain an indication of the current requested during the current power management cycle. The amount of current requested may be based on the number of outstanding memory access requests for each processing thread and the type of outstanding memory access requests for each processing thread. In one embodiment, each processing thread returns a separate response to the polling request so that thread manager 155 can determine the current request for each processing thread separately. In one embodiment, another component or subcomponent of PPM wrapper 150 may issue polling requests to and receive current requests from processing threads.

[0089] At operation 610, an available power budget is determined. For example, processing logic may determine the amount of current available in the memory device during a power management cycle. In one embodiment, PPM 160 receives a signal, such as a peak current magnitude indicator signal HC#, indicating the current used by each other PPM 160 in multi-die package 200 and subtracts that amount from the total amount of current available in memory subsystem 110 or memory device 130 or 130A. In one embodiment, processing logic compares the total current associated with all processing threads (e.g., the sum of the individual current requests) with the amount of current available during the power management cycle to determine whether the available current budget satisfies the current requests of all processing threads. If the available current is equal to or greater than the current associated with (e.g., requested by) all processing threads, processing logic determines that the available current satisfies the current associated with all processing threads.

[0090] At operation 615, current is requested and allocated. If the processing logic determines that the available current amount satisfies the current amount associated with all processing threads, then the processing logic may request the current amount associated with all processing threads. For example, PPM 160 may issue the request to a common power supply or other power supply in memory device 130 or 130A or memory subsystem 110. PPM 160 may then allocate the requested current to the processing threads, thereby allowing all processing threads to complete their pending memory access operations.

[0091] If the processing logic determines that the available current does not satisfy the current associated with all processing threads, then at operation 620, the thread combination is checked. For example, the processing logic may identify a thread combination corresponding to the current state of a ring counter, such as non-priority ring counter 444, from a data structure, such as data structure 448. The thread combination corresponding to each state of non-priority ring counter 444 is different, thereby ensuring that different threads are serviced in different power management cycles and no threads are ignored. In one embodiment, the processing logic compares the total current associated with the identified thread combination (e.g., the sum of the individual current requests) with the available current during the power management cycle to determine whether the available current budget satisfies the current request of the thread combination. If the available current is equal to or greater than the current associated with the thread combination, then the processing logic determines that the available current satisfies the current associated with the thread combination.

[0092] At operation 625, current is requested and allocated. If the processing logic determines that the available current amount satisfies the current amount associated with the thread combination, the processing logic may request the current amount associated with the thread combination. For example, PPM 160 may issue the request to a common power supply or other power supply in memory device 130 or 130A or memory subsystem 110. PPM 160 may then allocate the requested current to the processing threads, thereby allowing the processing threads identified in the thread combination to complete their pending memory access operations.

[0093] If the processing logic determines that the available current does not satisfy the current associated with the thread combination, then at operation 630, the leading thread is checked. For example, the processing logic may identify the leading thread corresponding to the current state of the ring counter, such as non-priority ring counter 444, from a data structure, such as data structure 448. The leading thread corresponding to each state of the non-priority ring counter 444 is different, thereby ensuring that different threads are serviced in different power management cycles and no thread is ignored. In one embodiment, the processing logic compares the requested current associated with the identified leading thread with the available current during the power management cycle to determine whether the available current budget satisfies the current request of the leading thread. If the available current is equal to or greater than the current associated with the leading thread, then the processing logic determines that the available current satisfies the current associated with the leading thread.

[0094] At operation 635, current is requested and allocated. If processing logic determines that the available current amount satisfies the current amount associated with the leading thread, processing logic may request the current amount associated with the leading thread. For example, PPM 160 may issue the request to a common power supply or other power supply in memory device 130 or 130A or memory subsystem 110. PPM 160 may then allocate the requested current to the leading thread, thereby allowing the leading thread to complete its pending memory access operation.

[0095] If the processing logic determines that the amount of available current does not satisfy the amount of current associated with the leading thread, then at operation 640, the current request is suspended. For example, the processing logic may suspend execution of the processing threads and maintain the current requests from those processing threads until a subsequent power management cycle. In the subsequent power management cycle, there may be a larger amount of available current in the memory device that may be sufficient to satisfy the request associated with at least one of the processing threads.

[0096] Figure 7is a flow chart of an example method for power budget arbitration in a memory device using a polling window, according to at least some embodiments. Method 700 may be performed by processing logic that may include hardware (e.g., a processing device, circuitry, dedicated logic, programmable logic, microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executed on a processing device), or a combination thereof. In some embodiments, method 700 is performed by Figure 1A and Figure 4 The PPM wrapper 150 of FIG. 10 illustrates a specific sequence or order, but the order of the processes may be modified unless otherwise specified. Therefore, the illustrated embodiments should be understood as examples only, and the illustrated processes may be performed in a different order, and some processes may be performed in parallel. Furthermore, one or more processes may be omitted in various embodiments. Therefore, not all processes are required in every embodiment. Other process flows are possible.

[0097] At operation 705, current level signals are received. For example, processing logic (e.g., PPM 160) may receive one or more current level signals, such as full signal 462, intermediate signal 464, and low signal 466, associated with a respective set of at least one of the requests in request register 452. In one embodiment, the current level signals are based on requests identified during a polling window between power management cycles (e.g., when polling window signal 460 is asserted). In one embodiment, during the polling window, thread manager 155 receives and stores current requests from processing threads, where each request includes an indication of a requested current. The amount of current requested may be based on the number of pending memory access requests for each processing thread and the type of pending memory access request for each processing thread. In one embodiment, each processing thread sends one or more separate requests, so that thread manager 155 can separately determine the current request for each processing thread and add the corresponding request to request register 452.

[0098] At operation 710, the available power budget is determined. For example, processing logic may determine the amount of current available in the memory device during a power management cycle (i.e., once a token is received and the polling window is closed). In one embodiment, PPM 160 receives a signal indicating the current used by each other PPM 160 in multi-die package 200, such as peak current magnitude indicator signal HC#, and subtracts that amount from the total amount of current in memory subsystem 110 or memory device 130 or 130A. In one embodiment, processing logic compares the total current associated with full signal 462 (e.g., the sum of all individual current requests in request register 452) with the amount of current available during the power management cycle to determine whether the available current budget satisfies full signal 462. If the amount of current available is equal to or greater than the amount of current associated with full signal 462, processing logic determines that the amount of current available satisfies full signal 462.

[0099] At operation 715, current is requested and allocated. If the processing logic determines that the amount of available current satisfies the full signal 462, then the processing logic may request the amount of current associated with all requests in the request register 452. For example, the PPM 160 may issue the request to a common power supply or other power supply in the memory device 130 or 130A or the memory subsystem 110. The PPM 160 may then allocate the requested current to the processing thread via the grant signal 472, thereby allowing all current requests in the request register 452 to be executed.

[0100] If the processing logic determines that the available current does not satisfy the request of full signal 462, then another current level signal is checked. For example, the processing logic compares the current associated with intermediate signal 464 (e.g., the sum of two or more current requests in request register 452) with the available current during the power management cycle to determine whether the available current budget satisfies intermediate signal 464. If the available current is equal to or greater than the current associated with intermediate signal 464, then the processing logic determines that the available current satisfies intermediate signal 464.

[0101] At operation 725, current is requested and allocated. If the processing logic determines that the available current amount satisfies the current amount associated with the intermediate signal 464, the processing logic may request the current amount associated with two or more requests from the request register 452. For example, the PPM 160 may issue the request to a common power supply or other power supply in the memory device 130 or 130A or the memory subsystem 110. The PPM 160 may then allocate the requested current to the processing thread via the grant signal 472, thereby allowing two or more of the current requests in the request register 452 to be executed.

[0102] If the processing logic determines that the available current amount does not satisfy the current amount associated with the intermediate signal 464, then another current level signal is checked at operation 730. For example, the processing logic compares the current associated with the low signal 466 (e.g., a current request in the request register 452) with the available current amount during the power management cycle to determine whether the available current budget satisfies the low signal 466. If the available current amount is equal to or greater than the current amount associated with the low signal 466, then the processing logic determines that the available current amount satisfies the low signal 466.

[0103] At operation 735, current is requested and allocated. If the processing logic determines that the amount of available current satisfies the amount of current associated with the low signal 466, the processing logic may request the amount of current associated with one of the requests from the request register 452. For example, the PPM 160 may issue the request to a common power supply or other power supply in the memory device 130 or 130A or the memory subsystem 110. The PPM 160 may then allocate the requested current to the processing thread via the grant signal 472, thereby allowing one of the current requests in the request register 452 to be executed.

[0104] If the processing logic determines that the available current amount does not satisfy the current amount associated with any of the current level signals, then at operation 740, the current requests are suspended. For example, the processing logic may suspend execution of the processing threads and maintain the current requests from those processing threads until a subsequent power management cycle. In the subsequent power management cycle, there may be a larger amount of available current in the memory device that may be sufficient to satisfy at least one of the requests.

[0105] Figure 8 FIG2 is a block diagram illustrating a combination of memory command packets 802A and 802B and a timing diagram 804 according to at least some embodiments. Memory command packet 802A illustrates a packet format in which a prefix may be appended to the command packet, which also includes an initial command, address information to which the memory operation is directed, and a close command to indicate termination of the memory operation associated with memory command packet 802A. Memory command packet 802B includes possible prefix values, such as 0x1, 0x2, or 0x3, or other such designators, to indicate a read command or program command without priority. In other memory command packets, possible prefixes include, for example, 0x41, 0x42, 0x43, or other such designators, to indicate a read command or program command with priority. In the disclosed embodiments, erase commands are not prioritized. In these embodiments, the memory die's ready / busy signal (RB#), shown in timing diagram 804, is asserted when processing a memory operation associated with memory command packet 802A or 802B and deasserted after the memory operation is completed.

[0106] In various embodiments, because the plurality of processing threads 334(0) to 334(3) may asynchronously respond to memory access operations, such as read operations or program operations, the PPM wrapper 150 may be programmed to manage power such that snapshot reads and other such memory access operations may be prioritized over non-prioritized memory operations, such as some program operations and any erase operations, for example. Accordingly, the memory subsystem controller 115 (or other processing device within the memory subsystem 110 that issues memory commands) may add a prefix value indicating prioritization to different memory command packets for memory operations sent to the memory device 130 or 130A.

[0107] In at least some embodiments, once the die and thus the individual PPM wrapper 150 receives, for example, a reference Figure 8 For example, the thread manager 155 may parse the memory command packet to access the prefix value associated with the target processing thread in the plurality of processing threads 334(0) to 334(3). The thread manager 155 may further determine whether the memory command packet is prioritized based on the prefix value. The thread manager 155 may further mark the target processing thread as prioritized in response to the memory command packet being prioritized.

[0108]

[0109] Table 3

[0110] Additional references Figure 4 , the thread manager 155 may select a data structure 458 (from a plurality of data structures) containing a prioritization indicator associated with one or more prioritized processing threads within the request of the intermediate signal 464 or thread combination and the low signal 468 or leader thread based on one or more of the plurality of processing threads being prioritized in the request register 452. The thread manager 155 may also increment a priority ring counter 454 as long as the set of processing threads maintains priority, wherein the data structure 458 stores an association between the value of the priority ring counter 454 and a subset of the plurality of processing threads 334(0) through 334(3), as shown in Tables 3 through 7. In these embodiments, the PPM 160 may then prioritize the allocation of power to the one or more prioritized threads during each new power management cycle. Figures 10A to 10B and Figure 11 Additional procedures for ensuring that such prioritized processing threads do not starve any non-prioritized processing threads of processing capacity, such as current allocation, will be discussed.

[0111]

[0112] Table 4

[0113] In various embodiments, Tables 3-7 illustrate examples of data structures that may be selected as data structures 458 for different sets of prioritized processing threads. Prioritized processing threads are marked with a capital "P" to indicate prioritization. Table 3 illustrates that the leading thread ("Low") is the primary processing thread and takes precedence over any other thread combination. However, in some cases, the leading thread may also be included in a thread combination ("Middle") and, therefore, may also be granted along with one or more additional non-prioritized processing threads.

[0114]

[0115] Table 5

[0116] In one embodiment, Table 4 illustrates the priorities between two different processing threads (i.e., coproc1 and coproc4), where at least one processing thread is the leading thread ("Min") for each corresponding value of the priority ring counter 454. In this embodiment, the two prioritized processing threads are also within the thread group ("Middle"). Assuming that the priority ring counter 454 is incremented four times (through the values ​​011), the PPM 160 can achieve uniformity in power distribution to the prioritized processing threads. Therefore, in some embodiments, the priority ring counter 454 values ​​"100" and "101" are not included in the increment cycle.

[0117] In one embodiment, Table 5 shows the priorities between three different processing threads (i.e., main, coproc2, and coproc5). Because the data structure 458 represented by Table 5 is programmed with these three prioritized threads in the requests of the middle signal 464 and the low signal 466, the PPM 160 can allocate power to these prioritized processing threads if the power budget is insufficient to allocate to all of the multiple processing threads.

[0118]

[0119]

[0120] Table 6

[0121] In one embodiment, Table 6 illustrates the priorities among four different processing threads (i.e., main, coproc1, coproc2, and coproc5). Because data structure 458, represented by Table 6, is programmed with these four prioritized threads in the requests of middle signal 464 and low signal 466, PPM 160 can allocate power to these prioritized processing threads in the event that the power budget is insufficient to allocate power to all of the multiple processing threads. Furthermore, assuming that priority ring counter 454 is incremented four times (through values ​​011), PPM 160 can achieve uniformity in power allocation to the prioritized processing threads. Therefore, in some embodiments, the values ​​"100" and "101" of priority ring counter 454 are not included in the increment loop.

[0122]

[0123]

[0124] Table 7

[0125] In one embodiment, Table 7 illustrates the priorities among five different processing threads (i.e., main, coproc1, coproc2, coproc4, and coproc5). Because data structure 458, represented by Table 7, is programmed with these five prioritized threads in the requests of middle signal 464 and low signal 466, PPM 160 can allocate power to these prioritized processing threads if the power budget is insufficient to allocate power to all of the multiple processing threads. Furthermore, assuming that priority ring counter 454 is incremented five times (through the value "100"), PPM 160 can achieve uniformity in power allocation to the prioritized processing threads. Therefore, in some embodiments, the value "101" of priority ring counter 454 is not included in the increment cycle.

[0126] Thus, it can be seen that the one or more prioritized processing threads involved in a given counter value include the leading thread in low signal 466, which in one embodiment is the main processing thread of the memory die. In another embodiment, the one or more prioritized processing threads include the leading thread and one or more processing threads of a thread combination, such as intermediate signal 464 within a subset of the plurality of processing threads. Thus, this subset (e.g., the combination of the leading thread and the thread combination) can be understood to include all prioritized processing threads.

[0127] Additional references Figure 4In at least some embodiments, PPM 160 determines a total available current budget for power consumption, which can be determined based on the quantified current to be consumed by the plurality of memory dies during a power management cycle, as previously discussed. PPM 160 can further determine power demands associated with the plurality of processing threads and, in response to determining that the available budget meets the power demands, allocate the power demands to the plurality of processing threads.

[0128] Additional references Figure 4 In at least some embodiments, PPM 160 distributes current to each respective prioritized processing thread in the one or more prioritized processing threads, and then distributes current to any non-prioritized processing threads in the subset of the plurality of processing threads. Thus, for example, in Table 4, at least one processing thread that is not prioritized exists within intermediate signal 464. PPM 160 may further track the total amount of current distributed to the one or more prioritized processing threads and any non-prioritized processing threads, and suspend distribution of current to any new processing threads in response to the amount of current allocated by the new current exceeding the available budget by less than the total amount of current already allocated. In this manner, PPM 160 ensures that despite prioritizing the distribution of power to the one or more prioritized processing threads, the total distributed power does not exceed the available budget for current (e.g., power).

[0129] Figure 9 is a diagram illustrating multi-plane prioritized power budget arbitration for multiple concurrent memory access operations in accordance with at least some embodiments. A series of concurrent memory operations 902 are shown along the top of the diagram, including a non-prioritized operation ("pgr0"), a programming operation for this example, followed by a prioritized series of additional asynchronous programming operations (e.g., iWL commands "pgr3," "pgr4," and "pgr5"). When PPM 160 receives the token (shown in the bottom timing diagram), non-priority ring counter 444 has a value of "011," and assuming sufficient current budget is available, only non-prioritized programming operations are running and are therefore allocated power. Each dashed indicator 905 is associated with a full signal 462, each dashed indicator 907 is associated with an intermediate signal 464, and each dashed indicator 909 is associated with a low signal 466, as discussed. Thus, three prioritized processing threads ( pgr3 , pgr4 , and pgr5 ) are encompassed within the middle signal 464 , while pgr 4 is the leading thread encompassed within the low signal 466 .

[0130] In response to PPM 160 allocating power (e.g., current) to a prioritized processing thread (e.g., "pgr3"), thread manager 155 may increment the value of priority ring counter 454, e.g., to a value of "100" in this example, before the end of the polling period associated with polling window signal 460. Thus, management of power allocation has been transferred to the prioritized allocation determined by the value of priority ring counter 454. This operation includes ensuring that the amount of current used for the new current allocation exceeds the available budget by no less than the total amount currently allocated by PPM 160 during the power management cycle while PPM 160 holds the token.

[0131] In these embodiments, if the amount of current for a new current allocation (e.g., to the pgr5 prioritized processing thread) exceeds the available budget by an amount that is less than the total amount currently allocated, then PPM 160 suspends allocating power (or current) to the non-prioritized processing thread pgr0. If suspending allocations to non-prioritized processing threads does not free up enough power budget to handle the pgr5 prioritized processing thread, then PPM 160 may need to further suspend allocating power to any new processing threads until sufficient budget is available. In this way, prioritized threads are given priority over non-prioritized processing threads. Although reference is made to Figure 9 These operations explained allow for prioritization of processing threads indicated as prioritized in low signal 466 or intermediate signal 464, but these operations do not ensure that the non-prioritized processing thread pgrO is not starved of power indefinitely or for an excessively long time.

[0132] Figures 10A to 10B is a diagram illustrating multi-plane prioritized power budget arbitration for multiple concurrent memory access operations according to at least some additional embodiments, for ensuring reference Figure 9 The non-prioritized processing thread pgr0 discussed does not lack power. Figure 10A General tracking reference Figure 9 As shown, for example, the control logic of the PPM wrapper 150 allocates power to a non-prioritized processing thread pgr0 among the plurality of processing threads based on the value of the non-priority ring counter 444 .

[0133] However, with Figure 9 different, Figure 10A The control logic (e.g., of the thread manager 155) is shown starting the timer 478 while the prioritized processing thread pgr0 is running and in response to detecting that power is allocated to the prioritized processing thread pgr3 of the one or more prioritized processing threads (pgr3, pgr4, pgr5). Figure 10BAlso shown are two additional subsequent power management cycles in which PPM 160 receives a token. In the third power management cycle, timer 478 has not yet expired, and therefore, thread manager 155 still increments the value of priority ring counter 454 for each power management cycle, this time to a value of "101," so that PPM 160 can still prioritize one or more prioritized processing threads. More specifically, control logic of PPM 160 prioritizes the allocation of power to one or more prioritized processing threads within the subset of the plurality of processing threads corresponding to the value of priority ring counter 454. Although the thread combination of intermediate signal 464 sent to PPM 160 still includes the three prioritized processing threads pgr3, pgr4, and pgr5, the lead thread of low signal 466 has shifted to the pgr5 prioritized processing thread.

[0134] Continue to refer Figure 10B , and in at least some embodiments, in response to timer 478 expiring before the non-prioritized processing thread pgr0 is completed, thread manager 155 shifts power allocation among subsets of the plurality of processing threads based on the value incremented to non-priority ring counter 454. In these embodiments, these subsets of the plurality of processing threads may be as described in Tables 2 and Figures 6 to 7 464, and low signals 466, where the individual processing threads are not marked or identified as prioritized. This means that, depending on the state or value of the non-priority processing counter 444, the non-prioritized processing threads will receive an equal share of power. This transition is shown during the fourth power management cycle, where the leading thread (indicator 909) is now the pgr0 processing thread and the thread group (indicator 907) includes the pgr0 processing thread. Therefore, the PPM 160 will allocate the available power budget to the non-prioritized processing thread pgr0 during this fourth power management cycle, thereby avoiding power starvation for the pgr0 processing thread.

[0135] In some embodiments, although not specifically shown, the control logic of the thread manager 155 resets the timer 478 upon detecting the completion of the previously non-prioritized processing thread pgr0. In other embodiments, the thread manager 155 resets the time in response to the completion of all previously non-prioritized processing threads. Then, again while the timer is running, the thread manager increments the priority ring counter 454 before each power management cycle so that the PPM 160 can prioritize the allocation of power to one or more prioritized processing threads within a subset of the plurality of processing threads corresponding to the value of the priority ring counter 454, such as shown in Tables 3 to 7. In this way, the control logic of the PPM wrapper 150 can transition back to prioritized power management until a situation occurs in which the PPM 160 allocates power to at least one non-prioritized processing thread and at least one prioritized processing thread. In response to this situation, the thread manager 155 can again start the timer 478, such as Figure 10A shown.

[0136] Figure 11 is a flow chart of an example method for prioritized power budget arbitration for multiple concurrent processing threads according to at least some embodiments. The method 1100 may be performed by processing logic that may include hardware (e.g., a processing device, circuitry, dedicated logic, programmable logic, microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executed on a processing device), or a combination thereof. In some embodiments, the method 1100 is performed by Figure 1A and Figure 4 The PPM wrapper 150 of FIG. 10 illustrates a specific sequence or order, but the order of the processes may be modified unless otherwise specified. Therefore, the illustrated embodiments should be understood as examples only, and the illustrated processes may be performed in a different order, and some processes may be performed in parallel. Furthermore, one or more processes may be omitted in various embodiments. Therefore, not all processes are required in every embodiment. Other process flows are possible.

[0137] At operation 1110, power is allocated to a non-prioritized processing thread. More specifically, the processing logic allocates power to a non-prioritized processing thread among a plurality of processing threads that are used to process the memory array 370 ( Figure 3 ) performs memory access operations. Non-priority ring counters can be Figure 4 Non-priority ring counter 44.

[0138] At operation 1120, the allocation of power to the prioritized processing thread is detected. More specifically, the processing logic determines whether the allocation of power to the prioritized processing thread has been detected. This results in the situation just discussed, where at least one non-prioritized processing thread and at least one prioritized processing thread have been allocated power and are running concurrently.

[0139] At operation 1130, a timer is started. More specifically, in response to positively detecting that power is allocated to the prioritized processing thread, the control logic starts a timer, such as timer 478 ( Figure 4 ).

[0140] At operation 1140, the priority ring counter is incremented. More specifically, when the timer is running, the processing logic increments the priority ring counter before each power management cycle. The priority ring counter may be Figure 4 Priority ring counter 454.

[0141] The prioritized processing threads are prioritized at operation 1150. More specifically, while the timer is running, the control logic prioritizes allocation of power to one or more prioritized processing threads within the subset of the plurality of processing threads corresponding to the value of the priority ring counter.

[0142] At operation 1160, a check is made to see if the timer has expired. More specifically, processing logic determines if the timer has expired before the non-prioritized processing thread has completed. If the answer is no, then at operation 1160, method 1100 loops back to operations 1140 and 1150 and continues prioritizing power management, as described with reference to FIG. Figures 10A to 10B Completed at power management loop three.

[0143] Power management transitions back to non-prioritized management at operation 1160. More specifically, in response to the timer expiring before the non-prioritized processing thread completes, the processing logic transitions power distribution among the subset of the plurality of processing threads based on the value incremented to the non-priority ring counter.

[0144] Figure 12 An example machine is shown of a computer system 1200 within which a set of instructions for causing the machine to perform any one or more of the methodologies discussed herein may be executed. In some embodiments, the computer system 1200 may correspond to a host system (e.g., Figure 1A 120) that includes, is coupled to, or utilizes a memory subsystem (e.g., Figure 1A memory subsystem 110) or can be used to perform operations of the controller (for example, execute an operating system to perform operations corresponding to Figure 1A 、 4In some embodiments, the machine may be connected (e.g., using a network) to other machines. In some embodiments, the machine may operate in the capacity of a server or a client machine in a client-server network environment, or as a server or a client machine in a peer-to-peer (or distributed) network environment.

[0145] The machine may be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a network appliance, a server, a network router, a switch or a bridge, or any machine capable of executing (sequentially or otherwise) a set of instructions that specify actions to be taken by the machine. Furthermore, while a single machine is shown, the term "machine" should also be construed to include any collection of machines that individually or collectively execute a set (or multiple sets of instructions) to perform any one or more of the methodologies discussed herein.

[0146] The example computer system 1200 includes a processing device 1202, a main memory 1204 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) (e.g., synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM)), etc.), a static memory 1206 (e.g., flash memory, static random access memory (SRAM), etc.), and a data storage system 1218, which communicate with each other via a bus 1230.

[0147] The processing device 1202 represents one or more general-purpose processing devices, such as a microprocessor, a central processing unit, or the like. More specifically, the processing device may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor that implements other instruction sets, or a processor that implements a combination of instruction sets. The processing device 1202 may also be one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), a network processor, or the like. The processing device 1202 is configured to execute instructions 1226 for performing the operations and steps discussed herein. The computer system 1200 may further include a network interface device 1208 for communicating via a network 1220.

[0148] The data storage system 1218 may include a machine-readable storage medium 1224 (also referred to as a computer-readable medium, such as a non-transitory computer-readable medium) having stored thereon one or more sets of instructions 1226 or software embodying any one or more of the methodologies or functions described herein. The instructions 1226 may also reside, completely or at least partially, within the main memory 1204 and / or within the processing device 1202 during execution by the computer system 1200, with the main memory 1204 and the processing device 1202 also constituting machine-readable storage media. The machine-readable storage medium 1224, the data storage system 1218, and / or the main memory 1204 may correspond to Figure 1A Memory subsystem 110.

[0149] In one embodiment, instructions 1226 include instructions for implementing the instructions corresponding to Figure 1A 、 4 The functional instructions of the PPM wrapper 150 are provided. Although the machine-readable storage medium 1224 is shown as a single medium in the example embodiment, the term "machine-readable storage medium" should be considered to include a single medium or multiple media that store one or more sets of instructions. The term "machine-readable storage medium" should also be considered to include any medium that is capable of storing or encoding a set of instructions for execution by a machine and causing the machine to perform any one or more of the methods of the present disclosure. Thus, the term "machine-readable storage medium" should be considered to include, but not be limited to, solid-state memory, optical media, and magnetic media.

[0150] Some portions of the previous detailed description have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. Those skilled in the data processing arts use these algorithmic descriptions and representations to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of operations leading to a desired result. An operation is one requiring physical manipulation of physical quantities. These quantities are usually, but not necessarily, in the form of electrical or magnetic signals capable of being stored, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0151] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. The present disclosure may relate to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities within a computer system's registers and memories into other data similarly represented as physical quantities within the computer system's memories or registers or other such information storage systems.

[0152] The present disclosure also relates to an apparatus for performing the operations described herein. This apparatus may be specially constructed for the intended purpose, or it may comprise a general-purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a computer-readable storage medium, such as, but not limited to, any type of disk (including floppy disks, optical disks, CD-ROMs, and magneto-optical disks), read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.

[0153] The algorithms and displays presented herein are not inherently related to any particular computer or other device. Various general-purpose systems can be used with programs according to the teachings herein, or it may prove convenient to construct more specialized devices for performing the methods. Structures for a variety of these systems will be presented as shown in the following description. Additionally, the present disclosure is not described with reference to any particular programming language. It will be appreciated that the teachings of the present disclosure as described herein can be implemented using a variety of programming languages.

[0154] The present disclosure may be provided as a computer program product or software, which may include a machine-readable medium having stored thereon instructions that can be used to program a computer system (or other electronic device) to perform a process according to the present disclosure. A machine-readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). In some embodiments, a machine-readable (e.g., computer-readable) medium includes a machine-readable (e.g., computer-readable) storage medium, such as a read-only memory ("ROM"), a random access memory ("RAM"), a magnetic disk storage medium, an optical storage medium, a flash memory component, or the like.

[0155] In the foregoing description, embodiments of the present disclosure have been described with reference to specific example embodiments thereof. It will be apparent that various modifications may be made to the present disclosure without departing from the broader spirit and scope of the embodiments of the present disclosure as set forth in the appended claims. Accordingly, the description and drawings are to be regarded as illustrative rather than restrictive.

Claims

1. A memory device comprising: a plurality of memory dies, each memory die of the plurality of memory dies comprising: memory arrays; a memory for storing data structures; and control logic operatively coupled to the memory array and the memory, wherein the control logic comprises: a plurality of processing threads for concurrently executing memory access operations on the memory array; a priority ring counter, wherein the data structure is to store an association between a value of the priority ring counter and a subset of the plurality of processing threads; a thread manager to increment the value of the priority ring counter and identify one or more prioritized processing threads corresponding to the subset of the plurality of processing threads prior to a power management cycle; and A peak power manager is coupled to the thread manager and is configured to prioritize allocation of power to the one or more prioritized processing threads during the power management cycle.

2. The memory device of claim 1 , wherein the one or more prioritized processing threads comprise one of: a leading thread, which is the main processing thread of the memory die; or The leading thread and one or more processing threads of a combination of threads within the subset of the plurality of processing threads.

3. The memory device of claim 2, wherein the thread manager is further to select the data structure comprising a prioritization indicator associated with the one or more processing threads from a plurality of data structures.

4. The memory device according to claim 1, wherein To identify the one or more prioritized processing threads, the thread manager is configured to, for each memory access operation: parsing a memory command packet to access a prefix value, the memory command packet being associated with a target processing thread among the plurality of processing threads; determining whether the memory command packet is prioritized based on the prefix value; as well as In response to the memory command packet being prioritized, marking the target processing thread as prioritized.

5. The memory device of claim 1 , wherein the peak power manager is further configured to: distributing current to each respective prioritized processing thread of the one or more prioritized processing threads followed by distributing current to any non-prioritized processing threads of the subset of the plurality of processing threads; tracking a total amount of current allocated to the one or more prioritized processing threads and any non-prioritized processing threads; as well as In response to the amount of current allocated by the new current exceeding the available budget by less than the total amount of current already allocated, allocating current to any new processing threads is suspended.

6. The memory device of claim 1 , wherein the peak power manager is further configured to: determining an available budget amount for power consumption based on a quantized amount of current to be consumed by the plurality of memory dice during the power management cycle; determining a power demand associated with the plurality of processing threads; as well as In response to determining that the available budget amount satisfies the power demand amount, the power demand amount is allocated to the plurality of processing threads.

7. The memory device of claim 1 , wherein the memory die further comprises a timer and a non-priority ring counter, and the thread manager is further to: starting a timer in response to detecting the power being allocated to at least one non-prioritized processing thread and at least one prioritized processing thread; while the timer is running, incrementing the priority ring counter for each new power management cycle so that the peak power manager allocates the power to the one or more prioritized processing threads within each respective subset of the plurality of processing threads corresponding to the value of the priority ring counter; as well as In response to the timer expiring before the non-prioritized processing thread completes, power distribution is transitioned among a subset of the plurality of processing threads based on the value incremented to the non-priority ring counter.

8. A memory device comprising: memory arrays; as well as control logic operatively coupled to the memory array to perform operations including: allocating power to one or more prioritized processing threads among a plurality of processing threads for performing memory access operations on the memory array based on a value of a priority ring counter; starting a timer while the one or more prioritized processing threads are running and in response to detecting that the power is allocated to a non-prioritized processing thread among the plurality of processing threads; While said timer is running: incrementing the priority ring counter before each power management cycle; as well as prioritizing allocation of the power to the one or more prioritized processing threads within a subset of the plurality of processing threads corresponding to a value of the priority ring counter; as well as In response to the timer expiring before the non-prioritized processing thread completes, power distribution is transitioned among a subset of the plurality of processing threads based on a value incremented to a non-priority ring counter.

9. The memory device of claim 8, wherein the operations further comprise: detecting completion of the non-prioritized processing thread; resetting the timer; as well as While said timer is running: incrementing the priority ring counter before each power management cycle; as well as Allocation of the power to the one or more prioritized processing threads within a subset of the plurality of processing threads corresponding to the value of the priority ring counter is prioritized.

10. The memory device of claim 8, the operations further comprising identifying, for each power management cycle, the one or more prioritized processing threads associated with the plurality of processing threads, wherein the one or more prioritized processing threads comprise one of: The lead thread, which is the main processing thread of the memory die; or The leading thread and one or more processing threads of a combination of threads within the subset of the plurality of processing threads. The memory device according to claim 10 , wherein: To identify the one or more prioritized processing threads, the operations further include, for each memory access operation: parsing a memory command packet to access a prefix value, the memory command packet being associated with a target processing thread among the plurality of processing threads; determining whether the memory command packet is prioritized based on the prefix value; as well as In response to the memory command packet being prioritized, marking the target processing thread as prioritized.

12. The memory device of claim 8, further comprising a data structure, wherein the control logic includes the priority ring counter, the data structure to store an association between the value of the priority ring counter and each subset of the plurality of processing threads.

13. The memory device of claim 8, wherein the operations further comprise, while the timer is running: distributing current to each respective prioritized processing thread of the one or more prioritized processing threads; tracking a total amount of current distributed to the one or more prioritized processing threads and the non-prioritized processing thread; and In response to the amount of current allocated by the new current exceeding the available budget by less than the total amount of current already allocated, allocating current to the non-prioritized processing thread is suspended.

14. The memory device of claim 8, wherein the operations further comprise: determining an available budget amount for power consumption based on a quantified amount of current to be consumed by the plurality of memory dies during the power management cycle; determining a power demand associated with the plurality of processing threads; as well as In response to determining that the available budget amount satisfies the power demand amount, the power demand amount is allocated to the plurality of processing threads.

15. A method comprising: allocating, by control logic of a memory die among a plurality of memory dies, power to a non-prioritized processing thread among a plurality of processing threads based on a value of a non-priority ring counter, the plurality of processing threads being used to perform memory access operations on a memory array of the memory die; starting, by the control logic, a timer while the non-prioritized processing thread is running and in response to detecting that the power is allocated to a prioritized processing thread of one or more prioritized processing threads of the plurality of processing threads; While said timer is running: incrementing a priority ring counter before each power management cycle; as well as prioritizing allocation of the power to the one or more prioritized processing threads within a subset of the plurality of processing threads corresponding to a value of the priority ring counter; as well as In response to the timer expiring before the non-prioritized processing thread completes, power distribution is transitioned, by the control logic, among a subset of the plurality of processing threads based on the value incremented to the non-priority ring counter.

16. The method according to claim 15, further comprising: detecting completion of the non-prioritized processing thread; resetting the timer; as well as While said timer is running: incrementing the priority ring counter before each power management cycle; as well as Allocation of the power to the one or more prioritized processing threads within a subset of the plurality of processing threads corresponding to the value of the priority ring counter is prioritized.

17. The method of claim 15, further comprising identifying, for each power management cycle, the one or more prioritized processing threads associated with the plurality of processing threads, wherein the one or more prioritized processing threads comprise one of: a leading thread, which is the main processing thread of the memory die; or The leading thread and one or more processing threads of a combination of threads within the subset of the plurality of processing threads.

18. The method according to claim 15, wherein To identify the one or more prioritized processing threads, the method further comprises, for each memory access operation: parsing a memory command packet to access a prefix value, the memory command packet being associated with a target processing thread among the plurality of processing threads; determining whether the memory command packet is prioritized based on the prefix value; as well as In response to the memory command packet being prioritized, marking the target processing thread as prioritized.

19. The method of claim 15, further comprising storing an association between the value of the priority ring counter and each subset of the plurality of processing threads in a lookup table.

20. The method of claim 15, further comprising, while the timer is running: distributing current to each respective prioritized processing thread of the one or more prioritized processing threads; tracking a total amount of current distributed to the one or more prioritized processing threads and the non-prioritized processing thread; and In response to the amount of current allocated by the new current exceeding the available budget by less than the total amount of current already allocated, allocating current to the non-prioritized processing thread is suspended.

Citation Information

Patent Citations

  • Mechanism to provide high performance and fairness in multi-threading computer system

    CN104838355A

  • System for ordering load and store instructions that performs out-of-order multithread execution

    CN1285064A