Background media scan management with improved efficiency for non-volatile memory devices
By queuing the refresh indicator of memory blocks in volatile memory and waiting for the completion of refresh data during power-down operation, the problem of interrupt refresh in the memory subsystem is solved, and the system reliability and data retention quality are improved.
Patent Information
- Application Number
- CN202411757615.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-11-04
- Filing Date
- 2024-12-03
- Publication Date
- 2025-06-06
AI Technical Summary
Existing memory subsystems may interrupt the refresh operation of memory blocks during power-down operations, resulting in lost refresh data or delayed recovery process, affecting data retention and system reliability.
By queuing the indicator of the memory block being refreshed in the volatile memory and upon detection of a power-down operation, a signal is sent to the host system instructing it to wait for the power-down operation to complete until the refresh data is fully written to the new memory block.
Ensures that refresh data is safely programmed before power-down operations, avoids data loss and recovery delays, and improves the reliability and data retention quality of the memory subsystem.
Smart Images

Figure CN120104048A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate generally to memory subsystems, and more particularly, to efficient background media scan management of non-volatile memory devices. Background Art
[0002] The memory subsystem may include one or more memory devices that store data. The memory devices may be, for example, non-volatile memory devices and volatile memory devices. In general, the host system may utilize the memory subsystem to store data at the memory devices and retrieve data from the memory devices. Summary of the invention
[0003] On the one hand, the present disclosure relates to a system comprising: an integrated circuit (IC) memory device comprising memory cells; a volatile memory device comprising a queue for storing indicators of one or more blocks of the memory cells to be refreshed; and a processing device operably coupled to the IC memory device and the volatile memory device, the processing device being used to perform operations comprising: detecting the initiation of a power-off operation of the system; and in response to detecting the initiation of the power-off operation: detecting that one or more indicators remain in the queue corresponding to the one or more blocks of the memory cells; and sending a signal to a host system coupled to the processing device, the signal being used to instruct the host system to wait for completion of the power-off operation until writing of refresh data from the one or more blocks to one or more erase blocks of the IC memory device is completed.
[0004] On the other hand, the present disclosure relates to a method, comprising: performing one or more sampling background scans of multiple memory cell blocks of the IC memory device by a processing device coupled to the IC memory device; detecting the initiation of a power-down operation of the IC memory device when an indicator of the IC memory device is maintained in a queue of a volatile memory device coupled to the processing device, wherein the indicator corresponds to one or more blocks of the multiple memory cell blocks that are eligible for refresh; and sending a signal by the processing device to a host system, the signal being used to instruct the host system to wait for completion of the power-down operation until writing of refresh data from the one or more blocks to one or more erase blocks of the IC memory device is completed.
[0005] On the other hand, the present disclosure relates to a non-temporary computer-readable storage medium storing instructions that, when executed by a processing device coupled to an IC memory device, cause the processing device to perform operations including: performing one or more sampling background scans of multiple memory cell blocks of the IC memory device; detecting the initiation of a power-down operation of the IC memory device when an indicator of the IC memory device is maintained in a queue of a volatile memory device, wherein the indicator corresponds to one or more blocks of the multiple memory cell blocks that are eligible for refresh; and sending a signal by the processing device to a host system, the signal being used to instruct the host system to wait for completion of the power-down operation until writing of refresh data from the one or more blocks to one or more erase blocks of the IC memory device is completed. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] The present disclosure will be more completely understood from the detailed description given below and from the drawings of various embodiments of the present disclosure. However, the drawings should not be considered to limit the present disclosure to specific embodiments, but are only for explanation and understanding.
[0007] Figure 1 An example computing system including a memory subsystem according to some embodiments of the present disclosure is described.
[0008] Figure 2A Schematically illustrates distribution of threshold control gate voltages for a memory cell capable of storing 3 bits of data by programming the flash memory cell into at least 8 charge states that differ by the amount of charge on the cell's floating gate, according to some embodiments.
[0009] Figure 2B is a graph of a set of example threshold voltage distributions for a plurality of memory cells of a memory array in a memory device according to some embodiments.
[0010] Figure 2C is a graph of two example threshold voltage distributions for a plurality of memory cells of a memory array in a memory device according to some embodiments.
[0011] Figure 3 is a flow chart of an example method for performing data refresh in particular memory blocks based on the results of a background media scan of those memory blocks in accordance with some embodiments.
[0012] Figure 4 is a flow chart of an example method for efficiently managing background media scans in a non-volatile memory device according to some embodiments.
[0013] Figure 5 is a block diagram of an example computer system in which embodiments of the present disclosure may operate. DETAILED DESCRIPTION
[0014] Aspects of the present disclosure relate to efficient background media scan management of non-volatile memory devices. The memory subsystem may be a storage device, a memory module, or a combination of a storage device and a memory module. Figure 1 Examples of storage devices and memory modules are described. In general, a host system may utilize a memory subsystem that includes one or more components, such as memory devices that store data. The host system may provide data stored at the memory subsystem and may request data retrieved from the memory subsystem.
[0015] The memory subsystem may include a high density non-volatile memory device where it is desirable to retain data when no power is supplied to the memory device. One example of a non-volatile memory device is a NAND memory device. Other examples of non-volatile memory devices, including integrated circuit (IC) memory devices, are described below in conjunction with Figure 1 Described. An IC non-volatile memory device is a package of one or more dies. Each die may include one or more planes. For some types of non-volatile memory devices (e.g., NAND devices), each plane may include a set of physical blocks. In some embodiments, each block may include multiple sub-blocks. Each block may include a set of pages. Each page may include a set of memory cells ("cells"). A cell is an electronic circuit that stores information. Depending on the cell type, a cell may store one or more binary information bits and have various logical states related to the number of bits stored. The logical state may be represented by a binary value such as "0" and "1" or a combination of such values.
[0016] The memory device may include cells arranged in a two-dimensional or three-dimensional grid. The memory cells may be etched onto a silicon wafer in an array of columns connected by conductive lines (hereinafter also referred to as bit lines or BL) and rows connected by conductive lines (hereinafter also referred to as word lines or WL). A word line may refer to a conductive line that connects the control gates of a group (e.g., one or more rows) of memory cells of a memory device, which is used together with one or more bit lines to generate an address for each of the memory cells. In some embodiments, each plane may carry an array of memory cells formed onto a silicon wafer and connected by conductive BLs and WLs, such that a word line connects a plurality of memory cells forming a row of the memory cell array, and a bit line connects a plurality of memory cells forming a column of the memory cell array. The intersection of a bit line and a word line constitutes the address of the memory cell. A block refers to a unit of a memory device for storing data hereinafter and may include a group of memory cells, a group of word lines, a word line, or an individual memory cell addressable by one or more word lines. One or more blocks may be grouped together to form separate partitions (e.g., planes) of a memory device in order to allow concurrent operations to occur on each plane. A memory device may include circuitry that performs concurrent memory page access of two or more memory planes. For example, a memory device may include respective access line driver circuits and power circuits for each plane of the memory device to facilitate concurrent access of pages (including different page types) of two or more memory planes.
[0017] A cell can be programmed (written) by applying a specific voltage to the cell, which causes a charge to be retained by the cell. For example, a voltage signal V CG A control electrode may be applied to the cell to turn on the current across the cell between the source electrode and the drain electrode. More specifically, for each individual cell (on which a charge Q is stored), there may be a threshold control gate voltage V t (also called the “threshold voltage”), which causes the source-drain current to flow at a controlled gate voltage (V CG ) is lower than the threshold voltage (V CG <V t ). The current is lower when the control gate voltage exceeds the threshold voltage (V CG >V t ) and then increases significantly. Because the actual geometry of the electrodes and gates varies from cell to cell, the threshold voltage can be different even for cells implemented on the same die. Therefore, a cell can be characterized by a distribution of threshold voltages, P(Q, V t )=dW / dV t , where dW represents the threshold voltage of any given cell when a charge Q is placed on the cell in the interval [V t ,V t +dV t ] is the probability within .
[0018] The programming operation may be performed by applying a series of incremental programming pulses to the control gate of the memory cell being programmed. A program verify operation after each programming pulse may determine the threshold voltage of the memory cell caused by the previous programming pulse. When a memory cell is programmed, the programming level achieved in the cell (e.g., the V t ) Actually, by comparing the unit V t The PV voltage level may be provided by an external reference.
[0019] The program verification operation may include applying a ramp voltage to the control gate of the memory cell being verified. When the applied voltage reaches the threshold voltage of the memory cell, the memory cell turns on and the sensing circuitry detects the current on the bit line coupled to the memory cell. The detected current activates the sensing circuitry and determines the current threshold voltage of the cell. The sensing circuitry may determine whether the current threshold voltage is greater than or equal to the target threshold voltage. If the current threshold voltage is greater than or equal to the target threshold voltage, then no further programming is required. Otherwise, programming continues in this manner applying additional programming pulses to the memory cell until the target V t and data status.
[0020] Therefore, some non-volatile memory devices may use a demarcation voltage (i.e., a read reference voltage) to read data stored at a memory cell. For example, a read reference voltage (also referred to herein as a "read voltage") may be applied to a memory cell, and if the threshold voltage of a specified memory cell is identified as being lower than the read reference voltage applied to the specified memory cell, then the data stored at the specified memory cell may be read as a specific value (e.g., logic '1') or determined to be in a specific state (e.g., a set state). If the threshold voltage of a specified memory cell is identified as being higher than the read reference voltage, then the data stored at the specified memory cell may be read as another value (e.g., logic '0') or determined to be in another state (e.g., a reset state). Therefore, a read reference voltage may be applied to a memory cell to determine the value stored at the memory cell. Such a threshold voltage may be within a threshold voltage range or reflect a normal distribution of threshold voltages.
[0021] A memory device may exhibit a threshold voltage distribution P(Q, V) that is narrower than the operating range of control voltages tolerated by the device's cells. t ). Therefore, multiple non-overlapping distributions P(Q k ,V t ) can be adapted to the operating range, thereby allowing the charge Q kStorage and reliable detection of multiple values of Q, k = 1, 2, 3, ... The distribution is interspersed with voltage intervals ("valleys") where cells of the device have none (or very few) of their threshold voltages. Therefore, such valley margins (also called read window budget (RWB)) can be used to separate various charge states Q k The logic state of a cell can be determined by detecting the cell's corresponding threshold voltage V during a read operation. t This effectively allows a single memory cell to store multiple bits of information: a memory cell operating with 2N-1 well-defined valleys and 2N distributions can reliably store N bits of information. Specifically, a read operation can be performed by comparing the measured threshold voltage V exhibited by the memory cell to t This is performed with one or more reference voltage levels corresponding to known valley voltage levels (eg, the center of the valley) of the memory device in order to distinguish between multiple logical programming levels and determine the programmed state of the cell.
[0022] Precise control of the amount of charge stored by the cell allows for the distinction of multiple logical states, effectively allowing a single memory cell to store multiple bits of information. One type of cell is a single-level cell (SLC), which stores 1 bit per cell and defines a voltage corresponding to each V t The cell can be in two logical states (“states”) of voltage level (“1” or “L0” and “0” or “L1”). For example, the “1” state can be the erased state and the “0” state can be the programmed state (L1). Another type of cell is the multi-level cell (MLC), which stores 2 bits per cell and defines two bits that correspond to the corresponding V t The 4 states of the voltage level ("11" or "L0", "10" or "L1", "01" or "L2", and "00" or "L3"). For example, the "11" state can be an erased state and the "01", "10", and "00" states can each be a corresponding programmed state. Another type of cell is a triple-level cell (TLC), which stores 3 bits per cell and defines each corresponding to a corresponding V t 8 states of the level ("111" or "L0", "110" or "L1", "101" or "L2", "100" or "L3", "011" or "L4", "010" or "L5", "001" or "L6", and "000" or "L7"). For example, the "111" state may be an erased state and each of the other states may be a corresponding programmed state. Another type of cell is a quad-level cell (QLC), which stores 4 bits per cell and defines 16 states L0 to L15, where L0 corresponds to "1111" and L15 corresponds to "0000". Another type of cell is a penta-level cell (PLC), which stores 5 bits per cell and defines 32 states. Other types of cells may also be considered. Thus, an n-level cell may use 2 nThe memory device may include one or more arrays of memory cells, such as SLC, MLC, TLC, QLC, PLC, etc., or any combination thereof. For example, the memory device may include an SLC portion and an MLC portion, a TLC portion, a QLC portion, or a PLC portion of cells.
[0023] In some memory subsystems, a read operation may be performed by comparing the measured threshold voltage (V t ) and one or more reference voltage levels in order to distinguish between two logical states of a single-level cell (SLC) and multiple logical states of a multi-level cell. In various embodiments, a memory device may include multiple portions, including, for example, one or more portions in which a sub-block is configured as an SLC memory, one or more portions in which a sub-block is configured as a multi-level cell memory. In these embodiments, the multi-level cell memory may include one or more portions of a multi-level cell (MLC) memory that can store 2 bits of information per cell, a triple-level cell (TLC) memory that can store 3 bits of information per cell, and / or a quad-level cell (QLC) memory in which a sub-block is configured to store 4 bits per cell. The voltage levels of the memory cells in the TLC memory form a set of 8 programming distributions representing 8 different combinations of 3 bits stored in each memory cell. Depending on how the memory cells are configured, each physical memory page in one of the sub-blocks may include multiple page types. For example, a physical memory page formed by a single-level cell (SLC) has a single page type at a single page level, referred to as a lower logical page (LP). A multi-level cell (MLC) physical page type may include pages at multiple page levels, such as an LP and an upper logical page (UP). A TLC physical page type includes pages at an additional page level called an LP, UP, and an additional logical page (XP), and a QLC physical page type includes pages at an LP, UP, XP, and another page level called a top logical page (TP). Different page types (LP, UP, XP, TP, etc.) may be referred to as levels within a page level hierarchy that may exist within the same physical memory page. For example, a physical memory page formed by memory cells of a QLC memory type may have a total of 4 logical pages, each of which may store data that is different from data stored in other logical pages associated with this physical memory page, which are referred to herein as "pages."
[0024] In various embodiments, to improve data retention on the least accessed logical block addresses (LBAs), a memory subsystem controller (e.g., a processing device) performs background media scans to periodically read data from a memory block. Thus, a host system may relocate data stored in a block to another block to refresh the data, or a controller may monitor the bit error rate (BER) of a page or block to determine if the page or block is decaying. Data retention is the length of time that a storage medium (e.g., NAND or other non-volatile memory (NVM) storage medium) in a memory device retains data under biased or unbiased conditions. Because data retention is limited, memory device scans and refreshes may be performed and managed by the memory subsystem controller through a background media scan (BGMS) process.
[0025] In some embodiments, there is an inverse relationship between data retention and the total bytes written (TBW) or temperature that affects the device over time. For example, as either or both of TBW and temperature increase, data retention decreases, requiring a refresh operation. From a data retention perspective, some memory devices comply with JESD47, which in the case of an unbiased device can be summarized as follows: 5 years at 55°C at 10% TBW or 1 year at 55°C at maximum TBW. This data retention and TBW can apply to both multi-level cell and single-level cell namespaces. These years, temperature, and TBW values are for illustration only and are not intended to be limiting.
[0026] More specifically, a memory device may experience V t For example, V t The distribution can be shifted to higher or lower values. For example, V t The time shift (i.e., V t The shift in the distribution over a period of time) can be caused by the quick charge loss (QCL) that occurs shortly after programming and the slow charge loss (SCL) that occurs over time during data retention. t Distribution shift, a calibration operation (including refresh operation) can be performed to adjust the read level voltage, which can be done based on the distribution because the higher V t levels tend to induce lower than V tLevel more time shift. In some memory devices, read voltage level adjustments may be performed based on one or more data state metric values obtained from a sequence of read and / or write operations. In an illustrative example, the data state metric may be represented by a raw bit error rate (RBER), which refers to the error rate in terms of the measurement of bits containing incorrect data (i.e., bits that are sensed incorrectly) when a data access operation is performed on the memory device (e.g., the ratio of the number of error bits to the number of all data bits stored in a particular portion of the memory device (e.g., a specified block). In these memory devices, a sweep read (or scan) may be performed to create an RBER / log likelihood ratio (LLR) configuration of an error correction code (ECC) and select the most efficient configuration. Such calibration may be performed to accurately predict the valley positioning at V t The positions between the distributions are used for the purpose of accurately reading data from the memory cells.
[0027] Due to incorrect V sensed for some cells when performing a read operation t , the rate at which error handling operations (e.g., remedial ECC operations) are triggered by a memory device during a read operation (referred to herein as a “trigger rate”) can be high, even for memory devices in which calibration techniques are used to address the timing V t As used herein, the read trigger rate refers to a measurement (e.g., count or frequency) of read operations that trigger additional read error handling operations (e.g., remedial ECC operations) due to high raw bit error rates (RBER) encountered during read operations. Despite the implementation of static calibration, high read trigger rates can be observed in QLC NAND devices. Therefore, the read trigger rate may correspond to the probability that the initial attempt to retrieve data fails (e.g., when the hard decoding of a codeword fails) and is therefore directly related to system performance and quality of service (QoS). For example, if a hard bit read operation fails for a set of data (e.g., a codeword), then the error recovery process will be triggered and increase the latency for the data to be retrieved. This delay negatively affects QoS and uses additional computing resources. This effect and its negative impact on memory devices is evident in mobile, embedded storage, storage (consumer, client, data center devices), or storage applications for external customers.
[0028] Furthermore, memory cells in a memory device may wear out over time or with increasing temperatures because their ability to retain charge (i.e., data) and thus remain at a particular programmed level deteriorates over time and with increased use and / or exposure to higher temperatures. Thus, in some cases, data retention quality may be reflected by a measurable degree of data degradation indicated by an error rate experienced during read operations performed on the data. This degree of degradation may be reflected by and may correspond to various corresponding values of data state metrics (e.g., valley shift values, read counts, valley width values, error counts, RBER, RWB, etc.). These values (e.g., of valley shift or read count) and their corresponding indications of data retention quality or capability on a memory device may be known from statistical and historical data obtained from scans (e.g., BGMS) and testing of various memory devices. Furthermore, the effects of these time shifts on toggle rates may be expected to worsen with the passage of additional time and increased use of the device.
[0029] Although different calibrations may be performed to track the various V t The read voltage reference (or demarcation voltage) may be changed by temporal shifting of the distribution, but the present disclosure focuses on performing data refresh in memory cells. For example, when a memory block is sampled by a read-based health scan, other specific V t A level or a particular threshold level may trigger the eligibility of this memory block for background refresh, which will be discussed in more detail. In some embodiments, background health scans are performed to avoid extensive ECC events and, of course, to avoid uncorrectable errors (e.g., UECC events). Thus, in some embodiments, the memory subsystem controller performs one or more sample background scans of one or more blocks and determines that one or more blocks are eligible for refresh based on attributes of Self-Monitoring Analysis and Reporting Technology (SMART) data obtained from the one or more sample background scans. In various embodiments, the attributes include total bytes written (TBW), temperature change over time, or data state metrics (such as those just discussed).
[0030] In response to being refresh eligible, the data in this block may be stored in a refresh queue in a volatile memory of the memory subsystem, such as a local memory of a memory subsystem controller. This refresh data is then programmed into a new memory block (e.g., an erase block) in the memory device, thereby updating the V of the memory cells for this data. tlevel and updates the healthier read voltage reference level accordingly. When data is programmed into a multi-level cell (e.g., MLC, TLC, QLC, PLC data), programming this data into a new block is called folding, which typically requires longer programming time than programming into SLC memory due to the dimensions of the data. Due to the time required to fold this data into a new memory block, a power-down operation of the memory subsystem (and therefore the memory device) may interrupt the memory refresh operation of one or more memory blocks. In various embodiments, the power-down operation may be power off (e.g., shut down) or moving to a low power state, such as a sleep state. When this occurs, the refresh data that has been folded and stored in the volatile memory may be lost or may cause increased wear on the memory device when the refresh data is restored. If the data is able to be recovered, the recovery process also causes a delay in the power-on operation, which negatively affects QoS.
[0031] In various embodiments, to avoid the possibility of losing refresh data or other negative effects from the recovery of refresh data, the disclosed systems, apparatuses, and methods provide a way to complete refresh operations of these memory blocks, wherein indicators of these blocks being refreshed are queued in volatile memory. In some embodiments, for example, the controller, in response to detecting the initiation of a power-down operation of the memory subsystem and detecting that the indicator remains in the queue, signals the host system to delay completion of the power-down operation until the refresh data from the block associated with the indicator is completely written back to the new memory block.
[0032] In at least some embodiments, the controller sends a signal to the host system to instruct the host system to wait for the power-off operation to be completed until the refresh data is written from one or more blocks to one or more erase blocks of the IC memory device. In some embodiments, the controller determines the amount of time to delay the power-off operation based on how long it takes to refresh the data from the number and data type of the one or more blocks. In these embodiments, the amount of time to delay the power-off operation is included in the signal, thereby notifying the host system how long to wait until the memory subsystem is powered off. In other embodiments, additional information may be included in the signal to the host to enable the host to make an informed decision about whether to delay until the data refresh of the block associated with the indicator stored in the refresh queue is completed. In this way, the refresh data is protected and safely programmed to the memory device before the memory subsystem is completely powered off, or these refreshes of some blocks are delayed until the memory device (or subsystem) is powered on again.
[0033] In various embodiments, the host system polls the controller (e.g., a register of the controller) which indicates whether the volatile memory still contains an indicator associated with refresh data that needs to be programmed to the IC memory device. Thus, the response to such a polling inquiry can be considered a signal from the memory subsystem. In this way, the host system can also be configured to know or wait for such a signal to know when it is safe to completely power down the memory subsystem and therefore also the memory device.
[0034] Advantages of the present disclosure include avoiding refresh data from being lost due to powering off the memory subsystem, thereby enhancing data retention and enhancing coordination between the host system and the non-volatile memory device in terms of safe timing for completing power-off operations. This enhanced coordination can be particularly beneficial in improving the reliability and robustness of non-volatile memory devices in automotive environments where power-ups and power-downs are performed more frequently and lower error rates or data loss rates are required in design specifications. For example, the enhanced coordination between the host system and the non-volatile memory device enables better communication and synchronization related to refresh and power-down operations, thereby ensuring optimal utilization of BGMS resources and minimizing performance degradation. Advantages also include reducing the read trigger rate associated with the SCL and static read voltage level calibration on the memory device, thereby reducing the latency of memory access operations performed by the memory device. This enhanced coordination between the host system and the memory device and the reduction in the read trigger rate improve the quality of service (QoS) that the user will experience when accessing data during a read operation without the risk of losing refresh data. Other advantages will be understood based on the additional details provided herein.
[0035] Figure 1 An example computing system 100 is illustrated that includes a memory subsystem 110 in accordance with some embodiments of the present disclosure. Memory subsystem 110 may include media such as one or more volatile memory devices (e.g., memory device 140), one or more non-volatile memory devices (e.g., memory device 130), or a combination of such media or memory devices.
[0036] The memory subsystem 110 may be a storage device, a memory module, or a combination of a storage device and a memory module. Examples of storage devices include solid state drives (SSDs), flash drives, universal serial bus (USB) flash drives, embedded multimedia controllers (eMMC) drives, universal flash storage (UFS) drives, secure digital (SD) cards, and hard disk drives (HDDs). Examples of memory modules include dual in-line memory modules (DIMMs), small outline DIMMs (SO-DIMMs), and various types of non-volatile dual in-line memory modules (NVDIMMs).
[0037] The computing system 100 may be a computing device such as a desktop computer, a laptop computer, a network server, a mobile device, a vehicle (such as an airplane, drone, train, car, or other transportation vehicle), a device with Internet of Things (IoT) capabilities, an embedded computer (such as an embedded computer included in a vehicle, industrial equipment, or a networked commercial device), or such a computing device that includes a memory and a processing device.
[0038] The computing system 100 may include a host system 120 coupled to one or more memory subsystems 110. In some embodiments, the host system 120 is coupled to multiple memory subsystems 110 of different types. Figure 1 An example of a host system 120 coupled to one memory subsystem 110 is illustrated. The host system 120 can provide data stored at the memory subsystem 110 and can request data to be retrieved from the memory subsystem 110. As used herein, "coupled to" or "coupled with" generally refers to a connection between components, which can be an indirect communication connection or a direct communication connection (e.g., without intervening components), whether wired or wireless, including connections such as electrical, optical, magnetic, etc.
[0039] The host system 120 may include a processor chipset and a software stack executed by the processor chipset. The processor chipset may include one or more cores, one or more caches, a memory controller (e.g., an NVDIMM controller), and a storage protocol controller (e.g., a PCIe controller, a SATA controller). The host system 120 uses the memory subsystem 110, for example, to write data to the memory subsystem 110 and read data from the memory subsystem 110.
[0040] The host system 120 may be coupled to the memory subsystem 110 via a physical host interface. Examples of the physical host interface include, but are not limited to, a Serial Advanced Technology Attachment (SATA) interface, a Peripheral Component Interconnect Express (PCIe) interface, a Universal Serial Bus (USB) interface, a Fibre Channel, a Serial Attached SCSI (SAS), a Double Data Rate (DDR) memory bus, a Small Computer System Interface (SCSI), a Dual In-line Memory Module (DIMM) interface (e.g., a DIMM slot interface supporting Double Data Rate (DDR)), etc. The physical host interface may be used to transfer data between the host system 120 and the memory subsystem 110. When the memory subsystem 110 is coupled to the host system 120 through a physical host interface (e.g., a PCIe bus), the host system 120 may further utilize an NVM Express (NVMe) interface to access components (e.g., the memory device 130). The physical host interface may provide an interface for passing control, address, data, and other signals between the memory subsystem 110 and the host system 120. Figure 1Memory subsystem 110 is illustrated as an example. In general, host system 120 can access multiple memory subsystems via the same communication connection, multiple separate communication connections, and / or a combination of communication connections.
[0041] Memory devices 130, 140 may include any combination of different types of non-volatile memory devices and / or volatile memory devices. Volatile memory devices (such as memory device 140) may be (but are not limited to) random access memory (RAM), such as dynamic random access memory (DRAM) and synchronous dynamic random access memory (SDRAM).
[0042] Some examples of non-volatile memory devices (e.g., memory device 130) include non-and (NAND) type flash memory and write-in-place memory, such as a three-dimensional cross-point ("3D cross-point") memory device that is a cross-point array of non-volatile memory cells. The cross-point array of non-volatile memory cells can perform bit storage based on body resistance changes in combination with a stackable cross-gate data access array. Therefore, compared to many flash-based memories, cross-point non-volatile memory can perform write-in-place operations, where non-volatile memory cells can be programmed without first erasing the non-volatile memory cells. NAND-type flash memory includes, for example, two-dimensional NAND (2D NAND) and three-dimensional NAND (3D NAND).
[0043] Each of the memory devices 130 may include one or more memory cell arrays. One type of memory cell, such as a single-level cell (SLC), may store one bit per cell. Other types of memory cells, such as multi-level cells (MLC), triple-level cells (TLC), quad-level cells (QLC), and penta-level cells (PLC), may store multiple bits per cell. In some embodiments, each of the memory devices 130 may include one or more memory cell arrays, such as SLC, MLC, TLC, QLC, PLC, or any combination thereof. In some embodiments, a particular memory device may include an SLC portion and an MLC portion, a TLC portion, a QLC portion, or a PLC portion of a memory cell. The memory cells of the memory devices 130 may be grouped into pages, which may refer to a logical unit of a memory device for storing data. For some types of memory, such as NAND, pages may be grouped to form blocks.
[0044] Although non-volatile memory components such as a 3D cross-point array of non-volatile memory cells and NAND-type flash memory (e.g., 2D NAND, 3D NAND) are described, the memory device 130 may be based on any other type of non-volatile memory, such as read-only memory (ROM), phase-change memory (PCM), self-select memory, other chalcogenide-based memory, ferroelectric transistor random access memory (FeTRAM), ferroelectric random access memory (FeRAM), magnetic random access memory (MRAM), spin transfer torque (STT)-MRAM, conductive bridging RAM (CBRAM), resistive random access memory (RRAM), oxide-based RRAM (OxRAM), or non-(NOR) flash memory or electrically erasable programmable read-only memory (EEPROM).
[0045] The memory subsystem controller 115 (or simply controller 115) can communicate with the memory device 130 to perform operations such as reading data, writing data, or erasing data at the memory device 130, and other such operations. The memory subsystem controller 115 may include hardware, such as one or more integrated circuits and / or discrete components, buffer memory, or a combination thereof. The hardware may include digital circuitry with dedicated (i.e., hard-coded) logic for performing the operations described herein. The memory subsystem controller 115 may be a microcontroller, dedicated logic circuitry (e.g., a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc.), or other suitable processor.
[0046] The memory subsystem controller 115 may include a processing device, including one or more processors, such as the processor 117, configured to execute instructions stored in the local memory 119. In the illustrated example, the local memory 119 of the memory subsystem controller 115 includes an embedded memory configured to store instructions for executing various processes, operations, logic flows, and routines that control the operation of the memory subsystem 110, including handling communications between the memory subsystem 110 and the host system 120.
[0047] In some embodiments, local memory 119 may include memory registers for storing memory pointers, fetch data, etc. Local memory 119 may also include read-only memory (ROM) for storing microcode. Figure 1 The example memory subsystem 110 in FIG. 1 is illustrated as including a memory subsystem controller 115, but in another embodiment of the present disclosure, the memory subsystem 110 does not include a memory subsystem controller 115, but may instead rely on external control (e.g., provided by an external host or by a processor or controller separate from the memory subsystem).
[0048] In general, the memory subsystem controller 115 may receive commands or operations from the host system 120 and may convert the commands or operations into instructions or appropriate commands to achieve the desired access to the memory device 130. The memory subsystem controller 115 may be responsible for other operations such as wear leveling operations, garbage collection operations, error detection and error correction code (ECC) operations, encryption operations, cache operations, and address translation between logical addresses (e.g., logical block addresses (LBAs), namespaces) and physical addresses (e.g., physical block addresses) associated with the memory device 130. The memory subsystem controller 115 may further include a host interface circuit system for communicating with the host system 120 via a physical host interface. The host interface circuit system may convert commands received from the host system into command instructions to access the memory device 130 and convert responses associated with the memory device 130 into information for the host system 120.
[0049] The memory subsystem 110 may also include additional circuitry or components not illustrated. In some embodiments, the memory subsystem 110 may include a cache or buffer (e.g., DRAM) and address circuitry (e.g., row decoders and column decoders) that may receive addresses from the memory subsystem controller 115 and decode the addresses to access the memory device 130.
[0050] In some embodiments, the memory device 130 includes a local media controller 135 that operates in conjunction with the memory subsystem controller 115 to perform operations on one or more memory cells of the memory device 130. An external controller (e.g., the memory subsystem controller 115) can manage the memory device 130 externally (e.g., perform media management operations on the memory device 130). In some embodiments, the memory subsystem 110 is a managed memory device, which is a raw memory device 130 with control logic on the die (e.g., the local media controller 135) and a controller for media management within the same memory device package (e.g., the memory subsystem controller 115). An example of a managed memory device is a managed NAND (MNAND) device.
[0051] In some embodiments, the memory subsystem 110 includes a scan manager 113 that can perform a scan on the memory device 130 to obtain data status metric values on the memory device 130. In several embodiments, the scan manager 113 can receive data access requests from the host system 120 and respond to the data access requests and manage the calibration of the applied voltage by controlling the voltage applied during the read operation of the memory device 130 (i.e., managing the voltage applied to V tIn some embodiments, the memory subsystem controller 115 includes at least a portion of the scan manager 113. In some embodiments, the scan manager 113 is part of the host system 120, an application, or an operating system. In other embodiments, the local media controller 135 includes at least a portion of the scan manager 113 and is configured to perform the functionality described herein.
[0052] The memory scan manager 113 may perform various actions, such as handling the interaction of the memory subsystem controller 115 of the memory subsystem 110 with the memory device 130. For example, in some embodiments, the scan manager 113 may transmit memory access commands corresponding to requests received by the memory subsystem 110 from the host system 120 to the memory device 130, such as program commands, read commands, and / or other commands, including protocol-based commands associated with health-based scanning. In addition, the scan manager 113 may receive data from the memory device 130, such as data retrieved in response to a read command or confirmation of successful completion of a write / program command.
[0053] In some embodiments, the memory subsystem controller 115 may include a processor 117 (e.g., a processing device) configured to execute instructions stored in the local memory 119 for performing the operations described herein. In other embodiments, the operations described herein are performed by the scan manager 113. In other embodiments, the local media controller 135 may perform the operations described herein. In at least one embodiment, the memory device 130 may include a memory access manager configured to perform memory access operations (e.g., operations performed in response to memory access commands received from the processor 117 or the scan manager 113). In some embodiments, the local media controller 135 may include at least a portion of the scan manager 113 and may be configured to perform the functionality described herein. In some of these embodiments, the scan manager 113 may be implemented on the memory device 130 using firmware, hardware components, or a combination of firmware and hardware components. In an illustrative example, the scan manager 113 may receive a request to read a data page of the memory device 130 from a requesting component (e.g., the processor 117) and respond thereto by performing the requested read operation. For purposes of this disclosure, a read operation may include a series of read strobes (also referred to as pulses) such that each strobe applies a specific read voltage level to a specific word line of the memory device 130. In a read operation, each strobe may be used to compare an estimated threshold voltage V of a group of memory cells. t and one or more read voltage levels corresponding to expected locations of the voltage distribution of the memory cells.
[0054] In some embodiments, the scan manager 113 may perform various scans on the data storage elements (e.g., pages) of the memory device 130. For purposes of this disclosure, data storage elements (e.g., cells (e.g., connected within the WL and BL arrays), pages, blocks, planes, dies, and groups and combinations of one or more of the foregoing elements) may be referred to as "data storage units." As described above, these scans may produce measurements of various different data state metrics, including error counts, RBER, valley shift, valley width, read counts, etc. For purposes of this disclosure, a read count may refer to the number of times data stored in a particular location of the memory device 130 has been accessed (i.e., read). The scan manager 113 may also track total bytes written (TBW) and temperature variations for various memory blocks or groups of blocks where the blocks are grouped based on being programmed at approximately the same time or at the same temperature.
[0055] In some embodiments, the term "background media scan (BGMS)" refers to a low priority firmware process executed by the scan manager 113. The BGMS process can resume at a regular cadence to incrementally read a selected number of pages (e.g., page sampling) in each fully written memory block (multi-level cell or SLC) over a relatively long time interval. This process can mitigate abnormal BER tail accidents, which can be exacerbated by memory retention, read disturb, cross-temperature effects, and defects in different memory portions. Thus, BGMS can proactively pass through the weakest blocks ("weak blocks"), thereby preventing these blocks from entering a condition that requires extensive ECC or even worse UECC events.
[0056] The purpose of BGMS may be to eventually sample every block in the memory device 130 instead of every page. The data within a block should have similar properties. For example, all data within a block (age of the data, temperature when the data was programmed, and temperature when the data was read) should all be similar so that a given page represents all data in the block. In addition, the BGMS sampling rate and sampling method may be designed to ensure that multiple samples are taken from each block in a manner that covers all expected use cases and most memory disturbance mechanisms. From power-on, BGMS read events should start with the oldest fully programmed memory block and step toward the newest programmed memory block. Given the BGMS re-trigger cadence, the oldest memory block is the most important. This procedure can continue indefinitely at a fixed sampling cadence. It can be expected that all fully written memory blocks will be scanned within a period of the fixed cadence. BGMS therefore acts on all user data blocks and FW blocks (containing FTL metadata).
[0057] In various embodiments, the following non-exhaustive list of conditions triggers a MGMS scan, including: 1) when the last page of a block is written, one SuperPage in the scan list may be checked next; 2) when the host system 120 reads 1 ("one") gigabyte (GB) of cumulative data in any order, one SuperPage may be scanned; 3) after waking up from idle (active-idle PS3 / PS4), N / 30 SuperPages may be scanned, where N is the idle time in seconds. For example, upon exiting a 63 second long idle period, two SuperPages may be scanned. A SuperPage is a page's worth of data that the controller 115 may read or write in parallel across multiple die storage. The BGMS effect is to queue an indicator of a "weak block" (for upcoming refresh / folding) in the refresh queue 121 of the local memory 119 (e.g., in the volatile memory of the memory subsystem 110). Even regular host reads may cause "weak blocks" to be scheduled for refresh, which will be discussed in more detail. Thus, a queuing indicator (e.g., a block identifier or number associated with a respective memory block) may be a way to schedule the corresponding memory block for refresh. In some embodiments, the scan manager 113 triggers refreshes to occur during host writes. In some embodiments, the BGMS scan is performed during idle time of the memory device 130 so as not to unduly impact the QoS of accessing the memory device 130.
[0058] Furthermore, in some embodiments, scan manager 113 scans groups of word lines or specific logical pages of word line groups residing on memory device 130. In some embodiments, the scans performed by scan manager 113 may include scans to facilitate read voltage calibration and scans to check data integrity after various read disturb stresses. For example, scan manager 113 may perform what may be referred to as a valley health check scan on each page in the group, which produces the following data state metric measurements: two specified adjacent V t distributions; and the center of the valley (and the displacement of the center from the previous measurement). In another example, the scan manager 113 may perform a scan on each page in the group (referred to as a read disturb scan on the page group) to obtain the following data state metric measurements: the read count of the page; and the two specified neighboring V t The center of the valley between the distributions (and the shift of the center from the previous measurement).
[0059] If you will refer to Figures 2B to 2C In more detail, data status metric values directly obtained through a scan performed by scan manager 113 may reflect the measurement of another data status metric, for example, through the application of a known mathematical transformation. For example, by performing a valley health check scan, scan manager 113 may obtain a valley margin (EC 1 ) and another valley threshold (EC 2) to obtain the valley width, which can be obtained by Similarly, by performing a valley health scan, the scan manager 113 may obtain a current measurement of the valley center to derive the valley shift by subtracting a previous measurement of the valley center from the current measurement of the valley center. Similarly, by performing a read disturbance scan, the scan manager 113 may obtain a current measurement of the read count to derive a logarithmic value of the read count (i.e., log ) that may be matched to a corresponding valley width value. 10 (read count)).
[0060] Thus, in some embodiments, the scan manager 113 identifies a group of word lines among the word lines on the memory device 130, where each word line in the group of word lines is connected to a respective subset of memory cells. In some embodiments, the scan manager 113 may identify the word lines in the group based on parameters that determine the location or properties of the cells on the word lines. For example, the scan manager 113 may select a group of word lines for scanning (a "scan group"), where each word line is selected based on parameters that determine the location or properties of the cells on the word lines (e.g., an index of the word line within a particular page, an SCL sensitivity of the word line, an indicator of which die / plane / block the word line is on, etc.). In one embodiment, the word lines may be selected so that a representative word line sampling shares a particular parameter from a portion of the scan group. In several embodiments, the identified group of word lines that forms a portion of the scan group may include word lines connected to memory cells programmed to logical states within one or more logical pages.
[0061] Thus, in some embodiments, the scan manager 113 may assign a specified charge loss classification value corresponding to a shift in the threshold voltage distribution to a group of word lines. For example, the scan manager 113 may use statistical or experimental data to determine a respective measurement of SCL sensitivity of one or more word lines and based thereon assign a charge loss classification value that characterizes the shift in the threshold voltage distribution on these word lines. In some embodiments, the scan manager 113 may identify word lines with the same or similar SCL sensitivity measurements and assign a corresponding charge bucket classifier (CBC) index value to each of the word lines in the group or to the entire group by recording the assigned CBC index value as metadata associated with the respective identifiers of the word lines in the group.
[0062] Additionally, the scan manager 113 may select a page level within the page level hierarchy, wherein the selected page level includes a particular set of memory cell charge states. The scan manager 113 may then select a set of memory cells having one or more memory cells whose respective charge states correspond to the page level (e.g., cells whose charge states are within the page level). Figure 2A To further understand the page level hierarchy and the correspondence between specific logical memory pages within a physical memory page and corresponding groups of memory cell charge states, Figure 2A A distribution 200A of threshold control gate voltages for a memory cell capable of storing three bits of data by programming the memory cell into at least eight charge states that differ by the amount of charge on the cell's floating gate is schematically illustrated.
[0063] Figure 2A Demonstrating the use of three-level unit (TLC) 3 –1=7 valley margins VM k Separated 2 N = Threshold voltage P (V T ,Q k ) distribution 200A. Therefore, the memory cell programmed to the kth charge state (eg, the memory cell having a charge Q deposited on its floating gate) k ) can store a specific combination of N bits (e.g., 0110, when N=4). k The valley margin VM can be detected during the read operation k The control gate voltage V CG Sufficient to turn on the source-drain current of the cell and the previous valley margin VM k-1 The control gate voltage is insufficient to determine.
[0064] In general, storage devices can be classified by the number of bits stored by each cell of the memory. For example, as described above, single-level cell (SLC) memory has cells that can each store one data bit (N=1) and multi-level memory has cells that can each store multiple data bits (N=2+). Among multi-level cell memories, multi-level cell (MLC) memory has cells that can each store up to two data bits (N=2), triple-level cell (TLC) memory has cells that can each store up to three data bits (N=3), and quad-level cell (QLC) memory has cells that can each store up to four data bits (N=4). In some storage devices, each word line of memory can have the same type of cells within a given partition of the memory device. Therefore, in some devices, all word lines of a block or plane can be SLC memory, or all word lines can be MLC memory, or all word lines can be TLC memory, or all word lines can be QLC memory. Because in some devices, the entire word line is controlled with the same control gate voltage V during a write or read operation, the entire word line can be controlled with the same control gate voltage V CGThe bias voltage is set so that a word line in SLC memory typically hosts one memory page (e.g., a 16KB or 32KB page) that is programmed in one setting (by sequentially selecting the various bit lines). The word lines of higher-level (MLC, TLC, or QLC) memory cells may host multiple pages on the same word line. Different pages may be programmed in a variety of settings (by the scan manager 113 of the memory controller 115 via electronic circuitry). For example, in some embodiments, after the first bit is programmed on each memory cell of a word line, an adjacent word line may be programmed before the second bit is programmed on the original word line. This may reduce electrostatic interference between adjacent cells. The memory controller 115 may program the state of the memory cell via the scan manager 113 and may then determine the read state of the memory cell by comparing the read threshold voltage V T This state is read with one or more read level thresholds. The operations described herein can be applied to any N-bit memory cell.
[0065] For example, TLC can be in at least 8 charging states Q k One of the states (where the first state may be an uncharged state Q 1 = 0), whose threshold voltage distribution is determined by the valley margin VM that can be used to read the data stored in the memory cell. k For example, if during a read operation it is determined that the read threshold voltage falls within 2 N -1 valley margin, then it can be determined that the memory cell is in 2 N The right valley margin of a cell can be determined by determining what value all of its N bits have. An identifier of the valley margin (e.g., its coordinates, such as the location of the center, and its width) can be stored in a read level threshold register of the memory controller 115.
[0066] A read operation can be performed on a memory cell that has been placed in its charged state Q by a previous write operation. k . For example, to program (write) 96KB (48KB) of data to cells belonging to a given word line M of a TLC, a first programming pass may be performed. The first programming pass may store 32KB (16KB) of data on word line M by placing an appropriate charge on the floating gates of the memory cells of word line M. For example, a charge Q may be placed on the floating gate of a particular cell. If the cell is charged to a charge state Q 1 , Q 2 , Q 3 or Q 4 If any of the following conditions are met, the cell is programmed to store a value in its lower page (LP) bit. If the cell is charged to a charge state Q 5 , Q 6 , Q 7 or Q8 If any of the above conditions are satisfied, the cell is programmed to store a value of 0 in its LP bit. Therefore, during a read operation, the fourth valley margin VM can be determined. 4 The external control gate voltage V CG is sufficient to turn on the source-drain current in the cell. Therefore, it can be concluded that the LP bit of the cell is in logic state 1 (e.g., in charge state Q k , where k≤4). Conversely, during a read operation, an applied control gate voltage V within the fourth valley margin may be determined. CG is insufficient to turn on the source-drain current in the cell. Therefore, it can be concluded that the LP bit of the cell is in logic state 0 (i.e., in charge state Q k one of , where k>4).
[0067] In some embodiments, after the cells belonging to the Mth word line have been programmed as described, the LP has been stored on the Mth word line and the programming operation can continue with additional programming passes to store the upper page (UP) and additional pages (XP) on the same word line. Although such passes can be performed immediately after the first pass is completed (or even all pages can be programmed in one setup), in order to minimize errors, it may be advantageous to program the LP of adjacent word lines (e.g., word lines M+1, M+2, etc.) before programming the UP and XP into word line M.
[0068] When UP is programmed into word line M, the charge state of the memory cell can be adjusted so that its threshold voltage distribution is further confined within a set of known valley margins VM. 1 , Q 2 , Q 3 or Q 4 A cell that is given one of the logic bit states Q (i.e., a cell that is given logic bit state 1 for LP programming) can be charged to exactly two states Q 1 or Q 2 In this case, the cell is used to store a value of 1 in its UP bit. Conversely, the cell can be charged to two states Q 3 or Q 4 One of them stores the value 0 in its UP bit. Therefore, during the read operation, the second valley margin VM can be determined 2 The external control gate voltage V CG is sufficient to turn on the source-drain current in the cell. Therefore, it can be concluded that the UP bit of the cell is in logic bit state 1 (i.e., in charge state Q k , where k≤2). Conversely, during a read operation, a second valley margin VM may be determined. 2 The external control gate voltage V CGis not sufficient to turn on the source-drain current of the cell. Therefore, it can be concluded that the UP bit of the cell is in the logic state 0 (i.e., in one of the charge states Q k where 2 < k ≤ 4). Similarly, the charge states Q 5 Q 6 Q 7 or Q 8 (the bit 0 state given for LP programming) can be further driven to the state Q 5 or Q 6 (UP bit value 0) or the state Q 7 or Q 8 (UP bit value 1).
[0069] Similarly, the extra page (XP) can be programmed into the word line M by further adjusting the charge state of each memory cell. For example, a cell in the logic state 10 (i.e., the UP bit stores the value 1 and the LP bit stores the value 0) and in one of the charge states Q 7 or Q 8 can be charged to the state Q 7 to store the value 0 (i.e., the logic state 010) in its XP bit. Alternatively, the cell can be charged to the charge state Q8 to store the value 1 (i.e., the logic state 110) in its XP bit. Therefore, during a read operation, it can be determined that the externally applied control gate voltage V CG within the seventh valley margin is not sufficient to turn on the source-drain current of the cell. Therefore, the memory controller 115 can determine that the logic state of the cell is 110 (corresponding to the charge state Q 7 ). Conversely, during a read operation, it can be determined that the externally applied control gate voltage V 7 within the seventh valley margin VM CG is sufficient to turn on the source-drain current of the cell. Therefore, the memory controller 115 can determine that the XP bit of the cell stores the value 0. If it is further determined that the control gate voltage V CG within the first 6 valley margins is not sufficient to turn on the current of the cell, then the memory controller 115 can determine the logic state of the cell to be 010 (i.e., corresponding to the charge state Q 7 ).
[0070] Thus, the scan manager 113 may select pages (i.e., logical pages) in multiple levels of a page level hierarchy, where each level contains a different set of memory cell charge states to which cells on word lines in an identified word line group may be charged. Each page level may correspond to a logical page type that includes a particular set of charge states, such as programming levels or logical states to which memory cells may be programmed. For example, a QLC memory cell connected to a word line in a word line group may be programmed to a state that is part of one of 4 logical pages, such as a lower page (LP), an upper page (UP), an extra page (XP), and a top page (TP). The memory cell may be programmed to an erased state or to one of 15 other programming levels, each of which may belong to one of pages LP, UP, XP, or TP.
[0071] After selecting a word line group and a logical page to be scanned, the scan manager 113 may scan the word line group. For example, the scan manager 113 may scan the word lines in the word line group that belong to the selected logical page (i.e., the word lines connected to the memory cells programmed to the programming level within the selected logical page). Each scan may include performing a coarse read calibration and may also include performing a fine read calibration, each of which involves applying one or more read reference voltages that may be offset relative to the initially applied read reference voltage and may include determining one or more data state metric values (e.g., RBER or EC of the memory cells and word lines being scanned). To perform the coarse read calibration, the scan manager may apply a read reference voltage determined based on an offset recorded in a calibration table (e.g., a coarse calibration table that contains entries indicating offset values relative to a default read reference voltage for reading memory cells programmed to a particular programming level). To scan memory cells on one of the word lines in a group, the scan manager 113 may apply a read reference voltage determined by reference to a calibration table, such as a default predetermined read reference voltage recorded in the settings of the memory device 130, for reading memory cells programmed to a particular logic state adjusted by a corresponding offset determined based on the CBC index value of the word line to which the memory cell is connected. Thus, in some embodiments, the scan manager 113 may scan a group of word lines by applying one or more sequences of read reference voltage pulses to the memory cells connected to the group of word lines.
[0072] Thus, in some embodiments, the scan manager 113 may select a group of memory cells whose respective charge states correspond to a page level (i.e., a group of memory cells that are all charged to a charge state within the group of charge states for the selected page level). The scan manager 113 may then determine an aggregate respective value of one or more data state metrics for the group of memory cells whose charge states are within the selected page level (i.e., whose programmed logical states are within a logical page). In various embodiments, the scan manager 113 may determine an aggregate value for each of the data state metrics that it measured during the scan. In some embodiments, the determined data state metric may be a raw bit error rate (RBER) and may also be an error count (EC). Thus, in some embodiments, the scan manager may determine an aggregate raw bit error rate (RBER) value. In the same or other embodiments, the scan manager may determine an aggregate error count (EC).
[0073] In some embodiments, the scan manager 113 may determine whether the determined individual or aggregated values of the measured data state metric satisfy a criterion. For example, in some cases, the criterion may be satisfied if the value is equal to or exceeds a predetermined threshold (e.g., when the RBER is greater than N bit errors / ms). In other cases, the criterion may be satisfied if the value is equal to or less than a predetermined threshold (e.g., when the EC is less than M errors). Therefore, in some embodiments, the scan manager 113 may determine whether the aggregated value of the data state metric satisfies a criterion (e.g., the RBER criterion). The scan manager 113 may determine whether the aggregated RBER value exceeds a threshold of N. In response to determining that the aggregated value of the data state metric satisfies a first criterion (e.g., in response to determining that the aggregated RBER value satisfies the first criterion because the measured aggregated RBER value is equal to or greater than a predetermined threshold of N), the scan manager 113 may identify another group of memory cells in the group of memory cells that are charged to a specified charge state within the selected page level. For example, when scanning a group of word lines within a TP, in response to determining that the aggregated RBER value exceeds a threshold, the scan manager 113 may identify memory cells programmed to programming level L5.
[0074] In the same or other embodiments, after identifying memory cells programmed to a logic state within a page (e.g., L5 in TP), the scan manager 113 may determine an aggregate value of another data state metric for the other group of memory cells based on one or more distributions of individual values of the data state metric. For example, the scan manager 113 may apply another one or more read reference voltage pulse sequences to the memory cells connected to a word line group or a group of word lines within a group. In various embodiments, each of the voltage pulse sequences may include one or more groups of read reference voltage pulses, wherein each group of read reference voltage pulses produces a corresponding distribution of individual values of the data state metric (e.g., RBER, EC). Each of the data state metric values may respectively have a corresponding read reference voltage from which the data state metric value is produced. Thus, the scan manager 113 may apply another read reference voltage pulse sequence to the memory cells connected to the word line group and produce a corresponding distribution of individual ECs, each individual EC corresponding to a respective read reference voltage.
[0075] Figure 2B 200B is a graph of a set of example threshold voltage distributions for a plurality of memory cells of a memory array in a memory device according to some embodiments of the present disclosure. In some embodiments, a memory device (e.g. Figure 1 The memory cells on a block of the memory device 130 may have different V t value, a group of these memory cells V t The aggregated representation of the values may be displayed graphically on a graph (e.g., graph 200B). For example, Figure 2B A group of V for a group of 16-level memory cells (e.g., QLC memory cells) is depicted in FIG. t In some embodiments, each of these memory cells is programmable to a V within one of 16 different threshold voltage ranges 201 to 216. t . Different V t Each of the ranges may be used to represent a distinct programming state corresponding to a particular pattern of 4 bits. In some embodiments, threshold voltage range 201 may have a greater width than the remaining threshold voltage ranges 202 to 216. This may result from the memory cells initially all being placed in a programming state corresponding to threshold voltage range 201, after which some subset of these memory cells may subsequently be programmed to have a threshold voltage within one of threshold voltage ranges 202 to 216. Because write (i.e., program) operations may be more precisely controlled than erase operations, these threshold voltage ranges 202 to 216 may have a narrower distribution.
[0076] In some embodiments, threshold voltage ranges 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, and 216 may each represent a corresponding programming state (e.g., L0, L1, L2, L3, L4, L5, L6, L7, L8, L9, L10, L11, L12, L13, L14, and L15, respectively). For example, if the V t In the first threshold voltage range 201 of the 16 threshold voltage ranges, the memory cell in this case can be said to be in a programming state L0 corresponding to a memory cell storing a 4-bit logic value of '1111' (which can be referred to as an erased state of the memory cell). Therefore, if the threshold voltage is in the second threshold voltage range 202 of the 16 threshold voltage ranges, the memory cell in this case can be said to be in a programming state L1 corresponding to a memory cell storing a 4-bit logic value of '0111'. If the threshold voltage is in the third threshold voltage range 203 of the 16 threshold voltage ranges, the memory cell in this case can be storing a programming state L2 with a 4-bit logic value of '0011', and so on, up to all 16 threshold voltage ranges. In some embodiments, a correspondence table (e.g., Table 1) can provide a correspondence between the state of a memory cell and its corresponding logic value. Other associations of programming states and corresponding logic data values are contemplated. For the purposes of this disclosure, a memory cell in a lowest state (e.g., an erased state or L0 data state) can be referred to as unprogrammed, erased, or set to a lowest programming state.
[0077] Programming Status Logical programming value Programming Status Logical data value L0 1111 L8 1100 L1 0111 L9 0100 L2 0011 L10 0000 L3 1011 L11 1000 L4 1001 L12 1010 L5 0001 L13 0010 L6 0101 L14 0110 L7 1101 L15 1110
[0078] Table 1
[0079] It is noteworthy that distributions 201-216 may be separated by valleys of different widths. Furthermore, over time and with continued use, the depicted distributions and valleys may shift and vary in width. Data state metrics of the voltage shifts obtained from the various scans performed by the embodiments described herein may be obtained for various valleys, such as the 15th valley between adjacent distributions (i.e., the valley between distributions 215-216) or the first valley between adjacent distributions (i.e., the valley between distributions 201-202). Thus, the scans described herein involve distinguishing the states of memory cells from one another and determining data state metrics associated with depicted distributions and valleys. This relationship is illustrated by focusing on the voltage shifts obtained by two adjacent V t The memory cell states represented by the distribution are further explained, as shown in reference Figure 2C Explain in more detail.
[0080] Figure 2C200C is a graph of two example threshold voltage distributions for a plurality of memory cells of a memory array in a memory device according to some embodiments of the present disclosure. Considering the example V t Description of distribution 225 to 226 Figure 2C Similar to Figure 2B The diagram 200B shows a pair of adjacent V t distribution. For example, Figure 2C V t Distributions 225-226 may represent the values of the memory cells after a write (ie, programming) operation is completed on a group of memory cells. Figure 2B A portion of the distribution of threshold voltages ranging from 201 to 216. Figure 2C , adjacent threshold voltage distributions 225-226 may be separated by a valley having a margin 240 (e.g., empty voltage level space) at the end of a programming operation. Applying a read voltage (i.e., a sensing voltage) between margins 240 to the control gates of a group of memory cells may be used to distinguish memory cells of threshold voltage distribution 225 (and any lower threshold voltage distributions) from memory cells of threshold voltage distribution 226 (and any higher threshold voltage distributions).
[0081] Due to a phenomenon known as charge loss, which can include quick charge loss (QCL) and slow charge loss (SCL), the threshold voltage of a memory cell can change over time and when exposed to higher temperatures as the charge contained in the cell degrades. As previously discussed, this change results in a V t The distribution shifts over time and can be called time V t 240). In addition, during memory device operation, QCL can be caused by a rapid change in threshold voltage (soon after a memory cell is programmed) followed by a gradual decrease in threshold voltage as V t The shift slowly decreases in an approximately log-linear fashion with respect to the time elapsed since the cell was programmed, and the SCL effect becomes more pronounced.
[0082] In various embodiments, this time V t The shift, if left unadjusted, can narrow the valley width between distributions 225 and 226 over time (i.e., can reduce the read window between margins 240 at the edges of threshold voltage distributions 225 to 226), and can cause these threshold voltage distributions 225 and 226 to overlap, making it difficult to distinguish their actual V t In two adjacent V t The cells in the range of one of the distributions 225 to 226 become more difficult. Therefore, the time V tShifts (e.g., caused by SLC) can lead to increased toggle rates and bit error rates in read operations. In addition, failure to address or account for the differences across all V t Distribution V t The shift can lead to increased read errors, resulting in high read toggle rates, which in turn negatively impacts the overall latency, throughput, and QoS of the memory device. Figures 2B to 2C In the illustrative example of, the number of distributions, programming levels, and logic values are selected for illustrative purposes and should not be construed as limiting, and other embodiments may use various other numbers of distributions, associated programming levels, and corresponding logic values that may be used in the various embodiments disclosed herein.
[0083] In various embodiments, additional reference is made to Figure 1 , the scan manager 113 stores and tracks attributes such as total bytes written (TBW), temperature changes over time, and representative data state metrics for a block of memory cells based on reading certain pages of each block associated with the BGMS. These attributes may be used individually or in combination to determine whether a BGMS sampled block is eligible for data refresh. As discussed, the degree of degradation of the memory cells in a particular memory block may be reflected by and may correspond to various respective values of data state metrics such as valley shift values, read counts, valley width values, error counts, RBER, RWB, and the like.
[0084] In some embodiments, BGMS-related scans are also or alternatively related to self-monitoring analysis and reporting technology (e.g., SMART). In some embodiments, SMART is a protocol scan command that enables the scan manager 113 to retrieve scan-based data (including BGMS-related data) from the memory device 130 to analyze the health of various memory blocks (or other memory units) of the memory device 130. This SMART technology is intended to identify conditions that indicate memory device degradation and is designed to provide sufficient failure warning to allow data backup before an actual failure occurs. In some embodiments, the scan manager 113 monitors specific attributes for degradation over time, but cannot predict instantaneous memory device failures. Each attribute monitors a set of specific conditions in the operating performance of the memory device 130, and the thresholds are optimized to minimize false predictions.
[0085] In at least some embodiments, the scan manager 113 stores metadata associated with the attributes associated with the BGMS scan to a memory, such as to the local memory 119 and / or to the memory device 130. In some embodiments, for example, this metadata may be instantiated as a firmware value stored to the NVM of the memory device 130 and may be cached in the local memory 119 during use. The scan manager 113 may periodically analyze this metadata with respect to memory cell degradation (e.g., associated with different attributes) and decide on a block-by-block basis whether each block is eligible for a data refresh. The scan manager 113 may initiate such a refresh operation for each block that meets a particular criterion (e.g., a threshold value of a corresponding attribute) that is desired to trigger a refresh operation. The scan manager 113 may cause the indicator for each block to be refreshed to be stored (or buffered) in a refresh queue 121 of volatile memory, such as so that the refresh data for each corresponding block may eventually be written to a new (e.g., erased) memory block. In this manner, the refresh data is rewritten to the erased block with a newly set threshold voltage level that also causes the read voltage reference level to reset, which is expected to reduce read errors and protect against data degradation that may result in data loss. In some embodiments, excessively degraded memory blocks may be marked as bad so as not to be used in the future.
[0086] In some embodiments of refreshing multi-level cell data (e.g., MLC, TLC, QLC, or PLC data), the scan manager 113 performs media management operations, such as by folding blocks to be refreshed before being read and written to erase blocks of the IC memory device 130. For example, if any codeword exhibits a trigger rate or reliability risk that meets certain criteria discussed, the data of the memory block may be folded. The folding operation may involve relocating the data stored at the affected block of the memory device to another block. Because a full scan is time consuming, a sample scan may be performed in which one or more pages of each block are read or tracked according to attributes and data state metrics.
[0087] Figure 3 1 is a flow chart of an example method 300 for performing data refresh in specific memory blocks based on the results of background media scans of those memory blocks according to some embodiments. The method 300 may be performed by processing logic, which may include hardware (e.g., a processing device, a circuit system, a dedicated logic, a programmable logic, a microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executing on a processing device), or a combination thereof. In some embodiments, the method 300 is performed by Figure 1 The scanning manager 113 of the embodiment of the present invention is executed. Although shown in a specific sequence or order, unless otherwise specified, the order of the processes can be modified. Therefore, the illustrated embodiments should be understood as examples only, and the illustrated processes can be performed in a different order, and some processes can be performed in parallel.
[0088] At operation 310, processing logic performs one or more sampling background scans of one or more memory blocks. In some embodiments, these sampling background scans (e.g., BGMS-based scans) start with the oldest programmed memory blocks first because it is expected that these memory blocks will be the most degraded memory blocks. Because in some embodiments, the sampling background scans may be limited to idle time or only when the memory device 130 is being written (no read operations), the priority of scanning these memory blocks first may help to more quickly identify memory blocks that need to be refreshed.
[0089] At operation 320, processing logic determines a representative error rate for the memory block based on the scan-generated data for at least one attribute. In some embodiments, for example, this scan-generated data may be instantiated as a firmware value stored to the NVM of the memory device 130 and may be cached in the local memory 119 during use, for example. In some embodiments, analysis may be performed on the scan-generated data associated with the attribute, and a particular threshold or total count for the attribute may indicate when the representative error rate is incremented or decremented. Hardware or software counters may be used to track the number of times a particular threshold is met in order to track the total error rate for each attribute. As discussed, the attributes may include TBW, temperature over time, or data state metrics, such as those previously discussed.
[0090] At operation 330, processing logic determines whether the memory block is refresh eligible based on at least one attribute of the self-monitoring analysis and reporting technology (SMART) data obtained from one or more sampled background scans or based on the representative error rate determined at operation 320. For example, in some embodiments, the representative error rate of each block and a particular attribute can predict failure of the corresponding block. Therefore, a threshold error rate for a particular attribute can be met, which triggers eligibility for data refresh. If the particular block is not data refresh eligible, then method 300 continues back to operation 310, where the scan and the data analysis resulting from the scan continue to be performed.
[0091] In response to the block being eligible for data refresh at operation 330, processing logic stores an indicator of the memory block in the refresh queue 121 of the volatile memory at operation 340, and the memory block is folded, read, and written to the erased memory cells of the memory device 130. The data of this indicator queued for refresh may be referred to as refresh data. For example, in some embodiments, the refresh data is a multi-level cell data type, and thus processing logic folds portions of the refresh data before the refresh data is written to one or more erased blocks of the IC memory device 130.
[0092] Figure 4is a flow chart of an example method for efficiently managing background media scans in a non-volatile memory device according to some embodiments. The method 400 may be performed by processing logic, which may include hardware (e.g., a processing device, a circuit system, dedicated logic, programmable logic, microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executing on a processing device), or a combination thereof. In some embodiments, the method 400 is performed by Figure 1 The scanning manager 113 of the embodiment of the present invention is executed. Although shown in a specific sequence or order, unless otherwise specified, the order of the processes can be modified. Therefore, the illustrated embodiments should be understood as examples only, and the illustrated processes can be performed in a different order, and some processes can be performed in parallel.
[0093] At operation 410, processing logic performs one or more sampling background scans of a plurality of memory cell blocks of an IC memory device.
[0094] At operation 420, processing logic detects initiation of a power down operation of the IC memory device while an indicator of the IC memory device is held in a queue of a volatile memory device coupled to the processing device. The power down operation may be powering off (e.g., shutting down) or moving to a low power state. For example, the queue may be refresh queue 121 ( Figure 1 ). In some embodiments, the indicator corresponds to or identifies one or more blocks of a plurality of blocks of memory cells that are eligible for refresh.
[0095] At operation 430, processing logic sends a signal to a host system (e.g., host system 120) indicating that the host system is waiting to complete the power down operation until writing the refresh data from the one or more blocks to the one or more erase blocks of the IC memory device is complete. In some embodiments, the signal includes the number of indicators held in refresh queue 121. Because host system 120 tracks the type of data, host system 120 is able to determine the expected amount of time remaining until the memory blocks corresponding to the indicators in queue 121 will be refreshed.
[0096] In some embodiments of the method 400, the processing logic further determines the amount of time for emptying the queue based on the number and data type of blocks whose indicators remain in the queue. For example, the block type may be SLC, MLC, TLC, QLC, and / or PLC (depending on the memory device configuration). In some embodiments, for explanation purposes only, block programming requires 1 second. Therefore, if 10 blocks of refresh data are stored in the refresh queue 121, the amount of time may be calculated as approximately 10 seconds. In these embodiments, the signal from the scan manager 113 includes the amount of time for delaying the power down operation. In some embodiments, the signal may also or alternatively include a flag that specifies the urgency associated with the refresh (e.g., low, medium, high) so that the host 120 can decide whether to wait or continue the power down operation. For example, the urgency may be based on the age of the data, whether data folding has occurred (and is therefore mid-process of a refresh), an error level associated with the data (e.g., RBER), or the like.
[0097] In some embodiments of method 400, processing logic monitors queue 121 to detect when the queue is empty of identifiers and notifies host system 120 of the completion of the power-down operation in response to detecting that the queue is empty. In various embodiments, host system 120 polls controller 115 (e.g., a register of controller 115) to indicate whether the volatile memory still contains refresh data that needs to be programmed to the IC memory device. Therefore, the response to such a polling query can be regarded as a signal from the memory subsystem. In this way, host system 120 can also be configured to know or wait for such a signal to know when it is safe to completely power down the memory subsystem and therefore also the memory device.
[0098] Figure 5 An example machine illustrating a computer system 500 within which a set of instructions for causing the machine to perform any one or more of the methodologies discussed herein may be executed. In some embodiments, the computer system 500 may correspond to a host system (e.g., Figure 1 ) that includes, is coupled to, or utilizes a memory subsystem (e.g., Figure 1 110) or can be used to perform operations of the controller (for example, to execute an operating system to execute a corresponding Figure 1 In some embodiments, the machine may be connected (e.g., using a network) to other machines. The machine may operate in the capacity of a server or a client user machine in server-client user network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or a client user machine in a cloud computing infrastructure or environment.
[0099] The machine may be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a network appliance, a server, a network router, a switch or a bridge, or any machine capable of executing (sequentially or otherwise) a set of instructions that specify actions to be taken by the machine. Furthermore, while a single machine is described, the term "machine" shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.
[0100] The example computer system 500 includes a processing device 502, a main memory 504 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) (e.g., synchronous DRAM (SDRAM) or RDRAM), etc.), a static memory 506 (e.g., flash memory, static random access memory (SRAM), etc.), and a data storage system 518, which communicate with each other via a bus 530.
[0101] The processing device 502 represents one or more general purpose processing devices, such as a microprocessor, a central processing unit, or the like. More specifically, the processing device may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor implementing other instruction sets or multiple processors implementing a combination of instruction sets. The processing device 502 may also be one or more special purpose processing devices, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, or the like. The processing device 502 is configured to execute instructions 526 for performing the operations and steps discussed herein. The computer system 500 may further include a network interface device 508 to communicate over a network 520.
[0102] The data storage system 518 may include a machine-readable storage medium 524 (also referred to as a computer-readable medium or a non-transitory computer-readable storage medium) on which is stored one or more sets of instructions 526 or software embodying any one or more of the methodologies or functions described herein. The instructions 526 may also reside completely or at least partially within the main memory 504 and / or the processing device 502 during execution thereof by the computer system 500, the main memory 504 and the processing device 502 also constituting machine-readable storage media. The machine-readable storage medium 524, the data storage system 518, and / or the main memory 504 may correspond to Figure 1 Memory subsystem 110.
[0103] In one embodiment, instructions 526 include instructions for implementing a command corresponding to a scan manager (e.g., Figure 1 Scan Manager 113 and Figure 3 and 4400 of the corresponding methods 300 and 400 of the present disclosure). Although the machine-readable storage medium 524 is shown as a single medium in the example embodiment, the term "machine-readable storage medium" should be considered to include a single medium or multiple media storing one or more sets of instructions. The term "machine-readable storage medium" should also be considered to include any medium capable of storing or encoding a set of instructions for machine execution and causing the machine to perform any one or more of the methodologies of the present disclosure. Therefore, the term "machine-readable storage medium" should be considered to include (but not limited to) solid-state memory, optical media, and magnetic media.
[0104] Some portions of the foregoing detailed description have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is generally conceived here to be a self-consistent sequence of operations leading to a desired result. Operations are those requiring physical manipulation of physical quantities. Typically, but not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, combined, compared, and otherwise manipulated. It has proven convenient at times, primarily for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.
[0105] It should be remembered, however, that all of these and similar terms should be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. The present disclosure may involve the actions and processes of computer systems or similar electronic computing devices that manipulate and transform data represented as physical (electronic) quantities within the computer system's registers and memories into other data similarly represented as physical quantities within the computer system's memories or registers or other such information storage systems.
[0106] The present disclosure also relates to an apparatus for performing the operations herein. This apparatus may be specially constructed for the intended purpose, or it may comprise a general purpose computer selectively activated or reconfigured by a computer program stored in the computer. This computer program may be stored in a non-transitory computer-readable storage medium, such as (but not limited to) any type of disk (including floppy disks, optical disks, CD-ROMs, and magneto-optical disks), read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.
[0107] The algorithms and displays presented herein are not inherently related to any particular computer or other device. Various general purpose systems may be used in conjunction with the programs according to the teachings herein, or it may prove convenient to construct more specialized devices to perform the methods. The structures of various of these systems will appear as set forth in the appended claims. In addition, the present disclosure is not described with reference to any particular programming language. It should be appreciated that various programming languages may be used to implement the teachings of the present disclosure described herein.
[0108] The present disclosure may be provided as a computer program product or software, which may include a machine-readable medium having instructions stored thereon, which can be used to program a computer system (or other electronic device) to perform a process according to the present disclosure. A machine-readable medium includes any mechanism for storing information in a form that can be read by a machine (e.g., a computer). In some embodiments, a machine-readable (e.g., computer-readable) medium includes a machine (e.g., computer) readable storage medium, such as a read-only memory ("ROM"), a random access memory ("RAM"), a magnetic disk storage medium, an optical storage medium, a flash memory component, etc.
[0109] In the foregoing description, embodiments of the present disclosure have been described with reference to specific example embodiments of the present disclosure. It should be understood that various modifications may be made to the present disclosure without departing from the broader spirit and scope of the embodiments of the present disclosure as set forth in the appended claims. Therefore, the description and drawings should be viewed in an illustrative rather than a restrictive sense.
Claims
1. A system comprising: An integrated circuit IC memory device comprising a memory cell; a volatile memory device comprising a queue for storing indicators of one or more blocks of said memory cells to be refreshed; and a processing device operably coupled to the IC memory device and the volatile memory device, the processing device configured to perform operations including: detecting initiation of a power-down operation of the system; and In response to detecting the initiation of the power-down operation: detecting that one or more indicators remain in the queue corresponding to the one or more blocks of the memory cells; and A signal is sent to a host system coupled to the processing device, the signal being used to instruct the host system to wait for completion of the power down operation until writing of refresh data from the one or more blocks to one or more erase blocks of the IC memory device is complete.
2. The system of claim 1, wherein the operations further comprise: performing one or more sampling background scans of the one or more blocks; and The one or more blocks are determined to be refresh eligible based on at least one attribute of Self-Monitoring Analysis and Reporting Technology (SMART) data obtained from the one or more sample background scans.
3. The system of claim 2, wherein the at least one attribute comprises one or more of total bytes written (TBW), temperature over time, or a data status metric.
4. The system of claim 1, wherein the operations further comprise: performing one or more sampling background scans of the one or more blocks to determine a representative error rate for each respective block of the memory cells; and The one or more blocks are determined to be eligible for refresh based on the representative error rate for each respective block.
5. The system of claim 1, wherein the operations further comprise determining an amount of time to empty the queue based on a number and a data type of the one or more blocks, and wherein the signal includes the amount of time to delay the power down operation.
6. The system of claim 5, wherein the signal further comprises a flag that specifies a level of urgency associated with the refresh, enabling the host system to decide whether to wait or continue with the power down operation.
7. The system of claim 1, wherein the one or more indicators include one or more identifiers, and wherein the operations further comprise: monitoring the queue to detect when the queue is empty of the one or more identifiers; and The host system is notified of completion of the power down operation in response to detecting that the queue is empty.
8. A method comprising: performing, by a processing device coupled to the IC memory device, one or more sampling background scans of a plurality of memory cell blocks of the IC memory device; detecting initiation of a power down operation of the IC memory device while an indicator of the IC memory device is held in a queue of a volatile memory device coupled to the processing device, wherein the indicator corresponds to one or more of the plurality of memory cell blocks that are refresh eligible; and A signal is sent by the processing device to a host system, the signal being used to instruct the host system to wait for completion of the power down operation until writing of refresh data from the one or more blocks to one or more erase blocks of the IC memory device is complete.
9. The method of claim 8, further comprising determining that the one or more blocks are refresh eligible based on at least one attribute of Self-Monitoring Analysis and Reporting Technology (SMART) data obtained from the one or more sampling background scans.
10. The method of claim 9, wherein the at least one attribute comprises one or more of total bytes written (TBW), temperature over time, or a data status metric.
11. The method of claim 8, wherein performing one or more sampling background scans of the one or more blocks is used to determine a representative error rate for each respective memory block, the method further comprising determining that the one or more memory blocks are eligible for refresh based on the representative error rate for each respective block.
12. The method of claim 8, further comprising determining an amount of time to empty the queue based on a number and data type of the one or more blocks, and wherein the signal includes the amount of time to delay the power down operation.
13. The method of claim 8, wherein the signal further comprises at least one of: the number of said indicators maintained in said queue; or A flag that specifies a level of urgency associated with the refresh so that the host system can decide whether to wait or continue with the power down operation.
14. The method of claim 8, wherein the indicator comprises an identifier, the method further comprising: monitoring the queue to detect when the queue is empty of the identifier; and The host system is notified of completion of the power down operation in response to detecting that the queue is empty.
15. A non-transitory computer-readable storage medium storing instructions that, when executed by a processing device coupled to an IC memory device, cause the processing device to perform operations comprising: performing one or more sampling background scans of a plurality of memory cell blocks of the IC memory device; detecting initiation of a power-down operation of the IC memory device while an indicator of the IC memory device is held in a queue of a volatile memory device, wherein the indicator corresponds to one or more of the plurality of memory cell blocks that are refresh eligible; and A signal is sent by the processing device to a host system, the signal being used to instruct the host system to wait for completion of the power down operation until writing of refresh data from the one or more blocks to one or more erase blocks of the IC memory device is complete.
16. The non-transitory computer-readable storage medium of claim 15, wherein the operations further comprise determining that the one or more blocks are refresh eligible based on at least one attribute of Self-Monitoring Analysis and Reporting Technology (SMART) data obtained from the one or more sampling background scans.
17. The non-transitory computer-readable storage medium of claim 15, wherein performing the one or more sampling background scans of the one or more blocks is used to determine a representative error rate for each respective memory block, wherein the operations further comprise determining that the one or more memory blocks are eligible for refresh based on the representative error rate for each respective block.
18. The non-transitory computer-readable storage medium of claim 15, wherein the operations further comprise determining an amount of time to empty the queue based on a number and a data type of the one or more blocks, and wherein the signal includes the amount of time to delay the power down operation.
19. The non-transitory computer-readable storage medium of claim 18, wherein the signal further comprises a flag that specifies a level of urgency associated with the refresh, enabling the host system to decide whether to wait or continue with the power down operation.
20. The non-transitory computer-readable storage medium of claim 15, wherein the indicator comprises an identifier, and wherein the operations further comprise: monitoring the queue to detect when the queue is empty of the identifier; and The host system is notified of completion of the power down operation in response to detecting that the queue is empty.