Probability Data Integrity Scanning with Dynamic Scanning Frequency

By dynamically adjusting the scanning frequency and data integrity scan window size of each part in the memory subsystem, the data error problem caused by reading interference in the prior art is solved, and more efficient bandwidth utilization is achieved.

CN115938434BActive Publication Date: 2025-05-27MICRON TECHNOLOGY INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210835440.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-08-09
Filing Date
2022-07-15
Publication Date
2025-05-27
Estimated Expiration
2042-07-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of data errors caused by read interference in memory subsystems, especially in the worst case of memory reliability, the scanning frequency of data integrity is too high, resulting in a decrease in bandwidth.

Method used

By dynamically selecting the scanning frequency of each part of the memory, dynamically adjusting the data integrity scanning window size based on the reliability of each part, reducing unnecessary scanning frequency to improve bandwidth utilization.

Benefits of technology

While ensuring the worst-case part of the memory, it reduces the scanning frequency and improves bandwidth utilization, and avoids the bandwidth drop caused by excessive frequency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115938434B_ABST
    Figure CN115938434B_ABST
Patent Text Reader

Abstract

This application relates to probabilistic data integrity scanning with a dynamic scan frequency. Exemplary methods, apparatuses, and systems include receiving a plurality of read operations. The read operations are divided into a current set of a series of read operations and one or more other sets. The size of the current set is a first number of read operations. An intruder read operation is selected from the current set. A first data integrity scan is performed on a victim of the intruder, and a first indicator of data integrity is determined based on the first data integrity scan. Responsive to determining that the first indicator of data integrity is greater than a current maximum value, the current maximum value is set to the first indicator of data integrity. Responsive to determining that the current maximum value meets a threshold, the size of a subsequent set of read operations is set to a second number, the second number being less than the first number.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to mitigation of read disturb errors in memory subsystems, and more particularly to determining and using a dynamic data integrity scanning window for a probabilistic data integrity scanning scheme. Background Art

[0002] The memory subsystem may include one or more memory devices that store data. The memory devices may be, for example, non-volatile memory devices and volatile memory devices. In general, the host system may use the memory subsystem to store data at the memory devices and retrieve data from the memory devices. Summary of the invention

[0003] A method is described. The method includes receiving a plurality of read operations for a portion of a memory, the plurality of read operations divided into a current set of a series of read operations and one or more other sets of read operation sequences, the current set being sized to a first number of read operations; selecting a first intruder read operation from the current set; performing a first data integrity scan on a victim of the first intruder read operation; determining a first indicator of data integrity based on the first data integrity scan; in response to determining that the first indicator of data integrity is greater than a current maximum value of data integrity indicators for a memory partition, setting the current maximum value to the first indicator of data integrity; and in response to determining that the current maximum value satisfies a first threshold, setting the size of a subsequent set of read operations to a second number of read operations, the second number being different from the first number.

[0004] A non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium includes instructions that, when executed by a processing device, cause the processing device to: receive a plurality of read operations for a portion of a memory, the plurality of read operations being divided into a current set of a series of read operations and one or more other sets of read operation sequences, the current set being sized to a first number of read operations; select a first intruder read operation from the current set; perform a first data integrity scan on a victim of the first intruder read operation; determine a first indicator of data integrity based on the first data integrity scan; in response to determining that the first indicator of data integrity is greater than a current maximum value of data integrity indicators for a memory partition, set the current maximum value to the first indicator of data integrity; and in response to determining that the current maximum value satisfies a first threshold, set the size of a subsequent set of read operations to a second number of read operations, the second number being less than the first number.

[0005] A system is described. The system includes: a plurality of memory devices; and a processing device operatively coupled to the plurality of memory devices. The processing device is configured to: receive a plurality of read operations for a portion of a memory, the plurality of read operations being divided into a current set of a series of read operations and one or more other sets of read operation sequences, the current set being sized to a first number of read operations; select a first intruder read operation from the current set; perform a first data integrity scan on a victim of the first intruder read operation; determine a first indicator of data integrity based on the first data integrity scan; in response to determining that the first indicator of data integrity is greater than a current maximum value of data integrity indicators for a memory partition, set the current maximum value to the first indicator of data integrity; and in response to determining that the current maximum value satisfies a first threshold, set the size of a subsequent set of read operations to a second number of read operations, the second number being different from the first number. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] The present disclosure will be more fully understood from the detailed description given below and the accompanying drawings of various embodiments of the present disclosure. However, the drawings should not be considered to limit the present disclosure to specific embodiments, but are only for explanation and understanding.

[0007] Figure 1 An example computing system including a memory subsystem according to some embodiments of the present disclosure is described.

[0008] Figure 2 An example of managing a portion of a memory subsystem according to some embodiments of the present disclosure is described.

[0009] Figure 3 is a flow chart of an example method for implementing a dynamic data integrity scanning frequency for a probabilistic data integrity scanning scheme according to some embodiments of the present disclosure.

[0010] Figure 4 is a flow chart of another example method for implementing a dynamic data integrity scanning frequency for a probabilistic data integrity scanning scheme according to some embodiments of the present disclosure.

[0011] Figure 5 is a flow chart of another example method for implementing a dynamic data integrity scanning frequency for a probabilistic data integrity scanning scheme according to some embodiments of the present disclosure.

[0012] Figure 6 is a block diagram of an example computer system in which embodiments of the present disclosure may operate. DETAILED DESCRIPTION

[0013] Aspects of the present disclosure relate to implementing a probabilistic data integrity scanning scheme in a memory subsystem. The memory subsystem may be a storage device, a memory module, or a mixture of a storage device and a memory module. Figure 1 Examples of storage devices and memory modules are described. In general, a host system may utilize a memory subsystem that includes one or more components, such as a memory device that stores data. The host system may provide data to be stored at the memory subsystem, and may request retrieval of data from the memory subsystem.

[0014] The memory device may be a non-volatile memory device. A non-volatile memory device is a package of one or more dies. An example of a non-volatile memory device is a NAND memory device. Figure 1 Other examples of non-volatile memory devices are described. The die in a package may be assigned to one or more channels for communication with a memory subsystem controller. Each die may be composed of one or more planes. Planes may be grouped into logical units (LUNs). For some types of non-volatile memory devices (e.g., NAND memory devices), each plane is composed of a set of physical blocks, which are groups of memory cells used to store data. A cell is an electronic circuit that stores information.

[0015] Depending on the cell type, a cell can store one or more binary bits of information and have various logical states related to the number of bits stored. The logical state can be represented by a binary value (e.g., "0" and "1") or a combination of such values. There are various types of cells, such as single-level cells (SLC), multi-level cells (MLC), triple-level cells (TLC), and quad-level cells (QLC). For example, an SLC can store one bit of information and have two logical states.

[0016] Data reliability in memory can decrease as the density of memory devices increases (e.g., when multiple bits per cell are programmed, the size of device components scales down, etc.). One contributing factor to this reduced reliability is read disturb. Read disturb occurs when a read operation performed on one portion of memory (e.g., a row of cells), commonly referred to as an aggressor, affects the threshold voltage in another portion of memory (e.g., an adjacent row of cells), commonly referred to as a victim. Memory devices typically have a limited tolerance for these disturbances. A sufficient amount of read disturb effects, such as a threshold number of read operations performed on adjacent aggressor cells, can change victim cells in other / unread portions of the memory to a different logic state than originally programmed, which can result in errors.

[0017] The memory system may track read disturbs by using a counter for each sub-portion of the memory and reprogramming a given sub-portion of the memory when the counter reaches a threshold. The probabilistic data integrity scheme consumes less storage by counting or otherwise tracking sets of read operations in portions of the memory (e.g., die, logic unit, etc.) and performing limited data integrity scans by checking the error rate of one or more read disturb victims of randomly selected read operations in each set (also referred to as a window). However, when the set or window size is set to a fixed value based on a worst-case scenario for memory reliability, the probabilistic data integrity scheme scans data integrity more frequently than is necessary for portions of the memory whose reliability metric is better than the worst-case scenario. For example, probabilistic data integrity scans based on a worst-case scenario for memory reliability may result in decreased bandwidth (when compared to using read disturb counters) because portions of the memory are scanned more frequently than necessary because not all portions of the memory are performing at the worst-case level of reliability.

[0018] Aspects of the present disclosure address the above and other deficiencies by dynamically selecting the scan frequency of each portion of memory based on the reliability of that portion of memory. A larger scan frequency is associated with a smaller probabilistic data integrity scan window size, while a smaller scan frequency is associated with a larger probabilistic data integrity scan window size. For example, the window size may be selected based on an error rate or other reliability metric for a portion of memory. Thus, a probabilistic data integrity scheme may initiate data integrity scans based on actual measurements of reliability - reducing the scan frequency and increasing bandwidth for portions of memory with good reliability without sacrificing reliability coverage for those portions of memory that represent worst-case scenarios.

[0019] Figure 1 An example computing environment 100 is illustrated that includes a memory subsystem 110 according to some embodiments of the present disclosure. Memory subsystem 110 may include media such as one or more volatile memory devices (e.g., memory device 140), one or more non-volatile memory devices (e.g., memory device 130), or a combination of the like.

[0020] The memory subsystem 110 may be a storage device, a memory module, or a mixture of storage devices and memory modules. Examples of storage devices include solid-state drives (SSDs), flash drives, universal serial bus (USB) flash drives, embedded multimedia controller (eMMC) drives, universal flash storage (UFS) drives, secure digital (SD) cards, and hard disk drives (HDDs). Examples of memory modules include dual in-line memory modules (DIMMs), small outline DIMMs (SO-DIMMs), and various types of non-volatile dual in-line memory modules (NVDIMMs).

[0021] The computing system 100 may be a computing device, such as a desktop computer, a laptop computer, a network server, a mobile device, a vehicle (e.g., an airplane, drone, train, car, or other transportation vehicle), an Internet of Things (IoT)-enabled device, an embedded computer (e.g., a computer included in a vehicle, industrial equipment, or a networked commercial device), or such a computing device including a memory and a processing device.

[0022] The computing system 100 may include a host system 120 coupled to one or more memory subsystems 110. In some embodiments, the host system 120 is coupled to memory subsystems 110 of different types. Figure 1 An example of a host system 120 coupled to one memory subsystem 110 is illustrated. As used herein, "coupled to" or "coupled with" generally refers to a connection between components, which may be an indirect communication connection or a direct communication connection (e.g., without intervening components), whether wired or wireless, including connections such as electrical connections, optical connections, magnetic connections, etc.

[0023] The host system 120 may include a processor chipset and a software stack executed by the processor chipset. The processor chipset may include one or more cores, one or more caches, a memory controller (e.g., an NVDIMM controller), and a storage protocol controller (e.g., a PCIe controller, a SATA controller). The host system 120 uses, for example, the memory subsystem 110 to write data to the memory subsystem 110 and read data from the memory subsystem 110.

[0024] The host system 120 may be coupled to the memory subsystem 110 via a physical host interface. Examples of the physical host interface include, but are not limited to, a Serial Advanced Technology Attachment (SATA) interface, a Peripheral Component Interconnect Express (PCIe) interface, a Universal Serial Bus (USB) interface, Fibre Channel, Serial Attached SCSI (SAS), a Small Computer System Interface (SCSI), a Double Data Rate (DDR) memory bus, a Dual In-line Memory Module (DIMM) interface (e.g., a DIMM socket interface supporting Double Data Rate (DDR)), an Open NAND Flash Interface (ONFI), Double Data Rate (DDR), Low Power Double Data Rate (LPDDR), or any other interface. The physical host interface may be used to transfer data between the host system 120 and the memory subsystem 110. When the memory subsystem 110 is coupled to the host system 120 through a PCIe interface, the host system 120 may further utilize an NVM Express (NVMe) interface to access components (e.g., the memory device 130). The physical host interface may provide an interface for passing control, address, data, and other signals between the memory subsystem 110 and the host system 120 . Figure 1Memory subsystem 110 is illustrated as an example. In general, host system 120 can access multiple memory subsystems via the same communication connection, multiple separate communication connections, and / or a combination of communication connections.

[0025] Memory devices 130, 140 may include any combination of different types of non-volatile memory devices and / or volatile memory devices. Volatile memory devices (e.g., memory device 140) may be, but are not limited to, random access memory (RAM), such as dynamic random access memory (DRAM) and synchronous dynamic random access memory (SDRAM).

[0026] Some examples of non-volatile memory devices (e.g., memory device 130) include non-and (NAND) type flash memory and write-in-place memory, such as a three-dimensional cross-point ("3D cross-point") memory device, which is a cross-point array of non-volatile memory cells. The cross-point array of non-volatile memory can be combined with a stackable cross-grid data access array to perform bit storage based on changes in body resistance. In addition, compared to many flash-based memories, cross-point non-volatile memory can perform write-in-place operations, where non-volatile memory cells can be programmed without pre-erasing the non-volatile memory cells. NAND type flash memory includes, for example, two-dimensional NAND (2D NAND) and three-dimensional NAND (3D NAND).

[0027] Although nonvolatile memory devices such as NAND type memories (e.g., 2D NAND, 3D NAND) and 3D cross-point nonvolatile memory cell arrays are described, the memory device 130 may be based on any other type of nonvolatile memory, such as read-only memory (ROM), phase change memory (PCM), self-selected memory, other chalcogenide-based memories, ferroelectric transistor random access memory (FeTRAM), ferroelectric random access memory (FeRAM), magnetic random access memory (MRAM), spin transfer torque (STT)-MRAM, conductive bridging RAM (CBRAM), resistive random access memory (RRAM), oxide-based RRAM (OxRAM), or non-(NOR) flash memory, and electrically erasable programmable read-only memory (EEPROM).

[0028] The memory subsystem controller 115 (or, for simplicity, the controller 115) may communicate with the memory device 130 to perform operations, such as reading data, writing data, or erasing data and other such operations at the memory device 130 (e.g., in response to commands dispatched by the controller 115 on a command bus). The memory subsystem controller 115 may include hardware, such as one or more integrated circuits and / or discrete components, buffer memory, or a combination thereof. The hardware may include digital circuitry with dedicated (i.e., hard-coded) logic to perform the operations described herein. The memory subsystem controller 115 may be a microcontroller, dedicated logic circuitry (e.g., a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc.), or another suitable processor.

[0029] The memory subsystem controller 115 may include a processing device 117 (processor) configured to execute instructions stored in a local memory 119. In the example shown, the local memory 119 of the memory subsystem controller 115 includes an embedded memory configured to store instructions for executing various processes, operations, logic flows, and routines that control the operation of the memory subsystem 110, including handling communications between the memory subsystem 110 and the host system 120.

[0030] In some embodiments, local memory 119 may include memory registers that store memory pointers, fetched data, etc. Local memory 119 may also include read-only memory (ROM) for storing microcode. Figure 1 The example memory subsystem 110 in FIG. 1 has been described as including a memory subsystem controller 115, but in another embodiment of the present disclosure, the memory subsystem 110 does not include the memory subsystem controller 115, but may rely on external control (e.g., provided by an external host or by a processor or controller separate from the memory subsystem 110).

[0031] In general, the memory subsystem controller 115 may receive commands or operations from the host system 120, and may convert the commands or operations into instructions or appropriate commands to achieve the desired access to the memory device 130 and / or the memory device 140. The memory subsystem controller 115 may be responsible for other operations, such as wear leveling operations, garbage collection operations, error detection and error correction code (ECC) operations, encryption operations, cache operations, and address conversion between logical addresses (e.g., logical block addresses (LBA), name space) and physical addresses (e.g., physical block addresses) associated with the memory device 130. The memory subsystem controller 115 may further include a host interface circuit system to communicate with the host system 120 via a physical host interface. The host interface circuit system may convert commands received from the host system into command instructions to access the memory device 130 and / or the memory device 140, and convert responses associated with the memory device 130 and / or the memory device 140 into information for the host system 120.

[0032] The memory subsystem 110 may also include additional circuitry or components not illustrated. In some embodiments, the memory subsystem 110 may include a cache or buffer (e.g., DRAM) and address circuitry (e.g., row decoders and column decoders) that may receive addresses from the memory subsystem controller 115 and decode the addresses to access the memory device 130.

[0033] In some embodiments, the memory device 130 includes a local media controller 135 that operates in conjunction with the memory subsystem controller 115 to perform operations on one or more memory cells of the memory device 130. An external controller (e.g., the memory subsystem controller 115) may externally manage the memory device 130 (e.g., perform media management operations on the memory device 130). In some embodiments, the memory device 130 is a managed memory device, which is a raw memory device combined with a local controller (e.g., the local controller 135) to perform media management within the same memory device package. An example of a managed memory device is a managed NAND (MNAND) device.

[0034] The memory subsystem 110 includes a data integrity manager 113 that mitigates read disturbs and other data errors. In some embodiments, the controller 115 includes at least a portion of the data integrity manager 113. For example, the controller 115 may include a processor 117 (processing device) configured to execute instructions stored in a local memory 119 for performing the operations described herein. In some embodiments, the data integrity manager 113 is part of a host system 120, an application, or an operating system.

[0035] The data integrity manager 113 may implement and manage the read disturb mitigation scheme. For example, the data integrity manager 113 may implement a dynamic data integrity scan frequency for a probabilistic data integrity scanning scheme. Additional details regarding the operation of the data integrity manager 113 are described below.

[0036] Figure 2 An example of managing a portion of the memory subsystem 200 according to some embodiments of the present disclosure is illustrated. In one embodiment, the data integrity manager 113 implements a read disturb mitigation scheme for each memory unit 210. For example, the data integrity manager 113 may perform a separate probabilistic read disturb mitigation scheme for each LUN.

[0037] The illustration of memory cell 210 includes an array of memory cells. To provide a simple explanation, memory 210 illustrates a small number of memory cells. Embodiments of memory cell 210 may include a much larger number of memory cells.

[0038] Each memory cell 210 includes a memory cell that the memory subsystem 110 accesses via a word line 215 and a bit line 220. For example, the memory device 130 may read a memory page using word line 230. Within the page, memory cell 225 is accessed via word line 230 and bit line 235. As described above, reading a memory cell may result in a read disturb effect on other memory cells. For example, a read of memory cell 225 (aggressor) may result in disturbing memory cells 240 and 245 (victims). Similarly, a read of other memory cells of word line 230 (aggressors) may result in disturbing other memory cells of word lines 250 and 255 (victims).

[0039] This disturbing effect may increase the error rate of the victim memory cells. In one embodiment, the data integrity manager 113 measures the error rate of a portion of the memory as a raw bit error rate (RBER). In another embodiment, the data integrity manager 113 may use other measurements to represent the error rate or data reliability, such as a voltage distribution value (e.g., a read threshold voltage distribution). In one embodiment, the data integrity manager 113 compares the current, recent, or average voltage distribution value to an ideal or expected voltage distribution value (e.g., a value produced by a new or unstressed memory), and the magnitude of the difference represents the error rate or other indication of data reliability. The data integrity manager 113 may track and mitigate read disturb by tracking the flow of read operations in the memory cells 210 and examining the error rates of the victims. For example, the data integrity manager 113 may select a read operation for word line 230 as an aggressor for testing read disturb and perform a read of word lines 250 and 255 to determine the error rate of each. In response to detecting that the error rate of a given victim portion of the memory satisfies the threshold error rate value, the data integrity manager 113 may migrate data from the victim portion of the memory to a different portion of the memory.

[0040] In one embodiment, the data integrity manager 113 also maintains one or more indicators of data integrity (e.g., RBER or other data reliability values) for performing data integrity scans. For example, the data integrity manager 113 may compare the RBER value of the current data integrity scan to one or more saved maximum RBER values ​​and replace the saved values ​​when the RBER value of the current data integrity scan exceeds the saved values. The data integrity manager 113 uses the saved representative data reliability values ​​to dynamically update the window size of the probabilistic read disturb handling scheme. Reference Figures 3 to 5 This and other features of the probabilistic read disturb handling scheme are further described.

[0041] Figure 3 is a flow chart of an example method 300 for implementing a dynamic data integrity scanning frequency for a probabilistic data integrity scanning scheme according to some embodiments of the present disclosure. The method 300 may be performed by processing logic, which may include hardware (e.g., a processing device, a circuit system, a dedicated logic, a programmable logic, a microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executed on a processing device), or a combination thereof. In some embodiments, the method 300 is performed by Figure 1The data integrity manager 113 of the embodiment of the present invention is executed. Although shown in a specific order or sequence, the order of the processes can be modified unless otherwise specified. Therefore, the illustrated embodiments should be understood as examples only, and the illustrated processes can be performed in a different order, and some processes can be performed in parallel. In addition, one or more processes can be omitted in various embodiments. Therefore, not all processes are required in every embodiment. Other process flows are also possible.

[0042] At operation 305, the processing device initializes or resets a counter for tracking the processing of read operations. For example, the processing device may set the counter to zero to begin tracking the processing of read operations in a group of read operations. In one embodiment, the processing device processes operations in a sequence set of read operations. For example, if the set contains 10,000 read operations, the counter is initialized or reset to a state that allows it to count at least 10,000 read operations. However, the number of operations per set is a dynamic value and will be referred to herein as N.

[0043] At operation 310, the processing device receives a read operation request. The read request may be received from one or more host systems and / or generated by another process within the memory subsystem 110. The processing device may receive the read operation request asynchronously, continuously, in batches, etc. In one embodiment, the memory subsystem 110 receives the operation requests from one or more host systems 120 and stores those requests in a command queue. The processing device may process the read operations from the command queue and / or generated internally in a set of N operations.

[0044] At operation 315, the processing device selects an intruder operation in the current set of operations. When a probabilistic read disturb handling scheme is implemented, the processing device may select the intruder in the current set by generating a random number (e.g., a uniform random number) in the range of 1 to N, and when the read operation count reaches the random number in the current set, the current / last read operation is identified as an intruder.

[0045] At operation 320, the processing device performs a read operation. For example, the memory subsystem 110 reads a page of data by accessing memory cells along a word line and returning data to the host system 120 or an internal process that initiates a read request. Additionally, the processing device increments a read operation counter. For example, the processing device may increment a counter in response to completing a read operation to track a current position in a sequence of read operations in a current set.

[0046] At operation 325, the processing device determines whether the read operation counter has reached an intruder operation in the set. For example, the processing device may compare the value of the counter with the generated first random number to identify the intruder read operation in the current set. If the counter has not reached a position in the sequence corresponding to the intruder operation, the method 300 returns to operation 320 to continue to perform the next read operation, as described above. If the counter has reached a position in the sequence corresponding to the intruder operation, the method 300 proceeds to operation 330.

[0047] At operation 330, the processing device performs an integrity scan of the selected intruder's victims. For example, the processing device may perform a read of each victim to determine an indicator of the victim's data integrity. In one embodiment, this includes checking an error rate, such as a raw bit error rate (RBER), for the victim. In another embodiment, determining the indicator of the victim's data integrity includes comparing a threshold voltage distribution of the victim / sampled portion of memory to an expected voltage distribution. If the error rate of the victim memory location satisfies a folding threshold (e.g., meets or exceeds an error rate threshold), the processing device folds the data by, for example, error correcting the data at the victim memory location and writing the corrected data to a new location.

[0048] At operation 335, the processing device determines whether the indicator of the data integrity of the victim is greater than one or more current maximum worst values. For example, the processing device may maintain the highest maximum error rate value and the second highest maximum error rate value for the data integrity indicator of the portion of the memory. However, for another indicator of data integrity, the worst value may not be the maximum value. For example, when using a threshold voltage distribution as an indicator of data reliability, the processing device may compare the "shape" of the measured distribution value / histogram with the ideal value by determining whether the shape is monotonically increasing / decreasing, comparing local minima, comparing the width of the histogram, determining the amount of shift (left / right) of the histogram, etc. As another example, the processing device may compare the "overall" or part of the threshold voltage distribution / histogram (e.g., between specific read positions). The most recent histogram overall meets the threshold value increased from the ideal overall between the read positions, which can be used as an indication of data reliability (or lack of data reliability). As another example, the processing device may evaluate the margin compared to the threshold error correction limit-the lower the measured margin, the worse the data reliability. For ease of describing the following examples, it should be understood that the maximum value may be used interchangeably with the worst value.

[0049] If the indicator of the victim's data integrity is greater than one or more current worst values, the method 300 proceeds to operation 340. If the indicator of the victim's data integrity does not exceed the current worst value, the method 300 proceeds to operation 345.

[0050] The processing device may maintain one or more worst values ​​for various partitions of each memory, such as each die / LUN, plane, independent word line segment, word line group, LUN group, block or block group, etc. In one embodiment, the processing device determines an average or other combination of indicators to reflect the worst value for the memory partition. For example, the processing device may determine an indicator of data integrity for each block of memory in a LUN and average or combine the indicators of data integrity into an indicator for the entire LUN. In one embodiment, based on the granularity used, the processing device maintains the worst value for a unique partition of the portion of memory. For example, if the worst value is maintained per LUN, the processing device may locate the worst value for each block of the memory - the same block cannot provide both the worst value and the second difference value.

[0051] At operation 340, the processing device updates one or more worst / maximum values ​​to reflect the new worst / maximum values. For example, when the indicator of the data integrity of the victim is greater than the current highest maximum value, the processing device reduces the current highest maximum value to the new second-highest maximum value and sets the indicator of the data integrity of the victim to the new current highest maximum value. When the indicator of the data integrity of the victim is greater than the second-highest maximum value (but not the highest maximum value), the processing device sets the indicator of the data integrity of the victim to the new second-highest maximum value (e.g., replacing the previous value). When the existing worst / maximum value is for the same memory block / partition as the victim, the processing device replaces the worst / maximum value with the indicator of the data integrity of the victim and may reorder the worst value and the second difference value if applicable.

[0052] At operation 345, the processing device determines whether the current worst / maximum value meets one or more thresholds. For example, the processing device may divide the range of data integrity values ​​into sub-ranges by one or more thresholds, and each sub-range may be associated with a different data integrity scan window size. The processing device determines whether the current maximum value is within a sub-range by whether the current maximum value exceeds, is equal to, or is lower than a given threshold. The range of data integrity values ​​may span from, for example, a zero error rate to a collapse threshold.

[0053] As a simple example, the range of data integrity values ​​can be divided into two sub-ranges: a low range corresponding to a large data integrity scan window size and a high range corresponding to a small data integrity scan window size. The low range can be defined as including error rate values ​​between zero and a dynamic window threshold, and the high range can be defined as including error rate values ​​between the dynamic window threshold and a collapse threshold.

[0054] In one embodiment, a single dynamic window threshold may divide the sub-ranges, and the current maximum value less than this threshold belongs to the low range and the current maximum value greater than this threshold belongs to the high range. In another embodiment, two dynamic window thresholds may divide the sub-ranges, and the threshold used may depend on the current sub-range based on the previous maximum value. For example, when the previous maximum value falls into the low range, the large threshold may be used to trigger a move from the low range to the high range, while the small threshold may be used to trigger a move from the high range to the low range. Thus, different thresholds may more conservatively trigger changes in the size of the data integrity scan window.

[0055] In one embodiment, the processing device implements a dynamic window size that allows the window size to shrink quickly but grow slowly. For example, the processing device may insert a delay (time, number of sets, etc.) before using an updated window size that is smaller than the current window size. This delay may allow subsequent sets to change the worst / maximum values ​​and possibly eliminate a decrease in the window size. Thus, the processing device may avoid rapid fluctuations in the window size. In one embodiment, the amount of delay inserted is a function of (1) the number of program / erase (P / E) cycles performed on the memory or portion of the memory and / or (2) the block type (SLC, MLC, TLC, etc.). For example, the greater the number of P / E cycles or the greater the density of the blocks, the greater the delay selected by the processing device.

[0056] If the current maximum value satisfies one or more dynamic window thresholds (e.g., would direct the processing device to trigger a change in the data integrity scan window size based on the current range of data integrity values), the method 300 proceeds to operation 350. If the current maximum value does not satisfy one or more thresholds (e.g., is unchanged from a previous range of data integrity values), the method proceeds to operation 355.

[0057] At operation 350, the processing device updates the set size for the next set of operations. For example, in response to the current maximum value satisfying the threshold and thus falling into a new range of data integrity values, the processing device changes the data integrity scan window size by updating the number of operations N for each set starting with the next set, or if delayed, starting with another subsequent set. The current set continues to use the current value of N, and the processing device stores an updated value of N (or increment) to be applied to the next set. Thus, the processing device may reduce the window size to scan more victims as the maximum value of the data integrity indicator increases, and increase the window size to scan fewer victims as the maximum value of the data integrity indicator decreases. In the case where the dynamic window size is bidirectional, an embodiment may set the initial value N (e.g., after a power cycle) to a small window size or a large window size, and allow subsequent sets to dynamically adjust the window size as needed. In one embodiment, the processing device stores the current value of N in a non-volatile memory, and the value of N is used after the power cycle.

[0058] At operation 355, the processing device determines whether the read operation counter has reached the end of the current set. For example, the processing device may compare the value of the counter to the N value of the current set. If the read operation counter has reached the end of the current set, the method 300 proceeds to operation 305 to reset the counter (e.g., to the N value of the next set) and process the next set of read operations. If the read operation counter has not reached the end of the current set, the method 300 proceeds to operation 360.

[0059] At operation 360, the processing device performs a read operation and increments the read operation counter, as described above with reference to operation 325. The method 300 proceeds to operation 355 to again determine whether the read operation counter has reached the end of the current set.

[0060] In one embodiment, the processing device disables the dynamic feature of the data integrity scan window in response to a threshold of a maximum value of the data integrity indicator being met, a threshold of a maximum value of the data integrity indicator being met a threshold number of times, a threshold number of program / erase cycles being met, etc. Thus, the processing device bypasses operations 335-350 in a subsequent set.

[0061] Figure 4 is a flow chart of another example method for implementing a dynamic data integrity scanning frequency for a probabilistic data integrity scanning scheme according to some embodiments of the present disclosure. The method 400 may be performed by processing logic, which may include hardware (e.g., a processing device, a circuit system, a dedicated logic, a programmable logic, a microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executed on a processing device), or a combination thereof. In some embodiments, the method 400 is performed byFigure 1 The data integrity manager 113 of the embodiment of the present invention is executed. Although shown in a specific order or sequence, the order of the processes can be modified unless otherwise specified. Therefore, the illustrated embodiments should be understood as examples only, and the illustrated processes can be performed in a different order, and some processes can be performed in parallel. In addition, one or more processes can be omitted in various embodiments. Therefore, not all processes are required in every embodiment. Other process flows are also possible.

[0062] At operation 405, the processing device receives a fold operation. For example, the processing device may receive a request or operation to copy data to a new location in response to determining that a data integrity indicator satisfies a fold threshold or another internal garbage collection process.

[0063] At operation 410, the processing device determines whether the memory location subject to the folding operation is mapped to the current worst / maximum value of the data integrity indicator. For example, in addition to saving the maximum value as described above, the processing device may also save a mapping between the maximum value and the address or other identification of the memory location subject to the data integrity scan that determines the maximum value of the data integrity indicator. If the memory location subject to the folding operation is mapped to the current maximum value of the data integrity indicator, the method 400 proceeds to operation 415. If the memory location subject to the folding operation is not mapped to the current maximum value of the data integrity indicator, the method 400 proceeds to operation 420.

[0064] At operation 415, the processing device updates the worst / maximum value in response to the folding operation. For example, if the folding operation maps to the highest maximum value, the processing device will overwrite the highest maximum value with the second worst / second highest maximum value (promoting the second highest value). In one embodiment, the processing device deletes the second highest value, marks it as needing to be updated, or otherwise notes the promotion. As another example, if the folding operation maps to the second highest maximum value, the processing device deletes the second highest value, marks it as needing to be updated, or otherwise notes that the second highest value is no longer valid. In either case, the processing device can use a subsequent data integrity scan to add the second maximum value again, as described above. Although the examples set forth above involve two worst / maximum values, embodiments can maintain and use more than two worst / maximum values.

[0065] At operation 420, the processing device performs a folding operation. For example, the processing device reads data from a source location and writes the data to a target location. This folding operation can be used to reduce the possibility of errors when reading data in the future, free up writable cells of memory, etc.

[0066] Figure 5is a flow chart of another example method for implementing a dynamic data integrity scanning frequency for a probabilistic data integrity scanning scheme according to some embodiments of the present disclosure. Method 500 may be performed by processing logic, which may include hardware (e.g., a processing device, a circuit system, a dedicated logic, a programmable logic, a microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executed on a processing device), or a combination thereof. In some embodiments, method 500 is performed by Figure 1 The data integrity manager 113 of the embodiment of the present invention is executed. Although shown in a specific order or sequence, the order of the processes can be modified unless otherwise specified. Therefore, the illustrated embodiments should be understood as examples only, and the illustrated processes can be performed in a different order, and some processes can be performed in parallel. In addition, one or more processes can be omitted in various embodiments. Therefore, not all processes are required in every embodiment. Other process flows are also possible.

[0067] At operation 505, the processing device receives a read operation, for example, as described with reference to operation 310. The processing device divides the read operation into sets of read operations. For example, the processing device may use a counter to track the number N of operations to be performed per set.

[0068] At operation 510, the processing device selects an intruder read operation in the current set of operations. For example, the processing device may randomly select the read operations in the current set to identify one or more victims as subjects for a data integrity scan.

[0069] At operation 515 , the processing device performs a data integrity scan of the victim memory locations. For example, the processing device may perform a read of the victims to check the error rate of each victim, as described above with respect to operation 330 .

[0070] At operation 520 , an indicator of data integrity is determined based on the scan. For example, the processing device determines the RBER or threshold voltage distribution of the victim / sampled portion of the memory, as described above with respect to operation 330 .

[0071] At operation 525, in response to determining that the indicator of data integrity is greater than the current maximum value, the processing device updates the current maximum value to the most recently scanned indicator. For example, the processing device may update one or more maximum values, as described above with respect to operation 340.

[0072] At operation 530, in response to determining that the current maximum satisfies the threshold, the processing device sets the size of the subsequent set of read operations to a second / different number of read operations. For example, the processing device may update the set size N of one or more subsequent sets, as described above with respect to operation 350.

[0073] While the above examples focus on dynamically selecting a data integrity scan frequency based on a worst data integrity value, other embodiments may dynamically adjust the scan frequency based on a number of P / E cycles or a similar indicator of data reliability. For example, the processing device may compare the current number of P / E cycles for a portion / partition of the memory to one or more thresholds and update the scan frequency for different ranges of P / E cycles.

[0074] Figure 6 An example machine illustrating a computer system 600 within which a set of instructions for causing the machine to perform any one or more of the methodologies discussed herein may be executed. In some embodiments, the computer system 600 may correspond to a host system (e.g., Figure 1 1) a host system 120 that includes, is coupled to, or utilizes a memory subsystem (e.g., Figure 1 110), or can be used to perform operations of the controller (for example, execute an operating system to execute corresponding Figure 1 In some embodiments, the machine may be connected (e.g., using a network) to other machines. In some embodiments, the machine may be connected (e.g., using a network) to other machines. In some embodiments, the machine may be connected (e.g., using a network) to other machines. In some embodiments, the machine may be connected (e.g., using a network) to other machines. In some embodiments, the machine may be connected (e.g., using a network) to other machines. In some embodiments, the machine may be connected (e.g., using a network) to other machines. In some embodiments, the machine may be connected (e.g., using a network) to other machines.

[0075] The machine may be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a network appliance, a server, a network router, a switch or a bridge, or any machine capable of executing (sequentially or otherwise) a set of instructions that specify actions to be taken by the machine. In addition, while a single machine is described, the term "machine" should also be construed to include any collection of machines that individually or collectively execute one (or more) sets of instructions to perform any one or more of the methodologies discussed herein.

[0076] The example computer system 600 includes a processing device 602, a main memory 604 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM), such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM)), etc.), a static memory 606 (e.g., flash memory, static random access memory (SRAM), etc.), and a data storage system 618, which communicate with each other via a bus 630.

[0077] The processing device 602 represents one or more general-purpose processing devices, such as a microprocessor, a central processing unit, etc. More specifically, the processing device may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor that implements other instruction sets, or a processor that implements a combination of instruction sets. The processing device 602 may also be one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, etc. The processing device 602 is configured to execute instructions 626 for performing the operations and steps discussed herein. The computer system 600 may further include a network interface device 608 for communicating over a network 620.

[0078] The data storage system 618 may include: a machine-readable storage medium 624 (also referred to as a computer-readable medium) having stored thereon one or more sets of instructions 626, or software embodying any one or more of the methods or functions described herein. The instructions 626 may also reside completely or at least partially within the main memory 604 and / or within the processing device 602 during execution thereof by the computer system 600, the main memory 604 and the processing device 602 also constituting machine-readable storage media. The machine-readable storage medium 624, the data storage system 618, and / or the main memory 604 may correspond to Figure 1 Memory subsystem 110.

[0079] In one embodiment, instructions 626 include instructions for implementing instructions corresponding to a data integrity manager (e.g., Figure 1 The machine-readable storage medium 624 is a storage medium that is used to store and store instructions for functions of the data integrity manager 113 of the present disclosure. Although the machine-readable storage medium 624 is shown as a single medium in the example embodiment, the term "machine-readable storage medium" should be considered to include a single medium or multiple media storing one or more sets of instructions. The term "machine-readable storage medium" should also be considered to include any medium that is capable of storing or encoding a set of instructions for execution by a machine and causing the machine to perform any one or more of the methods of the present disclosure. Therefore, the term "machine-readable storage medium" should be considered to include, but not be limited to, solid-state memory, optical media, and magnetic media.

[0080] Some portions of the previous detailed description have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means by which those skilled in the art of data processing most effectively convey the substance of their work to others skilled in the art. In this document, and generally, an algorithm is conceived to be a self-consistent sequence of operations that produces a desired result. An operation is one that requires physical manipulation of physical quantities. Typically (but not necessarily), these quantities take the form of electrical or magnetic signals that can be stored, combined, compared, and otherwise manipulated. It has proven convenient at times, primarily for common reasons, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, and the like.

[0081] It should be borne in mind, however, that all of these and similar terms will be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. The present disclosure may refer to the actions and processes of a computer system or similar electronic computing device to control and transform data represented as physical (electronic) quantities within the computer system's registers and memories into other data similarly represented as physical quantities within the computer system's memories or registers or other such information storage systems.

[0082] The present disclosure also relates to an apparatus for performing the operations herein. This apparatus may be specially constructed for the intended purpose, or it may include a general-purpose computer selectively activated or reconfigured by a computer program stored in the computer. For example, a computer system or other data processing system (e.g., controller 115) may perform computer-implemented methods 300 to 500 in response to its processor executing a computer program (e.g., a sequence of instructions) contained in a memory or other non-transitory machine-readable storage medium. This computer program may be stored in a computer-readable storage medium, such as, but not limited to, any type of disk, including a floppy disk, an optical disk, a CD-ROM, and a magneto-optical disk, a read-only memory (ROM), a random access memory (RAM), an EPROM, an EEPROM, a magnetic card or an optical card, or any type of medium suitable for storing electronic instructions, each of which is coupled to a computer system bus.

[0083] The algorithms and displays presented herein are not inherently related to any particular computer or other device. Various general purpose systems may be used with the program according to the teachings herein, or it may prove convenient to construct more specialized equipment to perform the methods. The structures of various these systems will be presented as set forth in the description below. In addition, the present disclosure is not described with reference to any particular programming language. It should be appreciated that the teachings of the present disclosure as described herein may be implemented using various programming languages.

[0084] The present disclosure may be provided as a computer program product or software, which may include a machine-readable medium having instructions stored thereon, the instructions being usable to program a computer system (or other electronic device) to perform a process according to the present disclosure. A machine-readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). In some embodiments, a machine-readable (e.g., computer-readable) medium includes a machine (e.g., computer) readable storage medium, such as a read-only memory ("ROM"), a random access memory ("RAM"), a magnetic disk storage medium, an optical storage medium, a flash memory component, etc.

[0085] In the foregoing description, embodiments of the present disclosure have been described with reference to specific example embodiments thereof. It will be apparent that various modifications may be made to the present invention without departing from the broader spirit and scope of the embodiments of the present disclosure as set forth in the appended claims. Accordingly, the description and drawings should be viewed in an illustrative rather than a restrictive sense.

Claims

1. A method of operating a memory device, wherein include: receiving a plurality of read operations for a portion of a memory, the plurality of read operations being grouped into a current set of a series of read operations and one or more other sets of sequences of read operations, the current set being sized to a first number of read operations; selecting a first intruder read operation from the current set; performing a first data integrity scan on a victim of the first intruder read operation; determining a first indicator of data integrity based on the first data integrity scan; in response to determining that the first indicator of data integrity is greater than a current maximum value of a data integrity indicator for a memory partition, setting the current maximum value as the first indicator of data integrity; and In response to determining that the current maximum satisfies a first threshold, a size of a subsequent set of read operations is set to a second number of read operations, the second number being different than the first number. The method of claim 1 , wherein the second number is smaller than the first number. The method of claim 1 , wherein the first indicator of data integrity is a bit error rate. The method of claim 1 , wherein the first indicator of data integrity is a read threshold voltage distribution.

5. The method according to claim 1, further comprising: include: selecting a second intruder read operation from the subsequent set; performing a second data integrity scan on victims of the second intruder read operation from the subsequent set; determining a second indicator of data integrity based on the second data integrity scan; and In response to determining that the second indicator of data integrity is greater than a second maximum value of the data integrity indicator for the memory partition, the second maximum value is set as the second indicator of data integrity.

6. The method according to claim 5, further comprising: include: responsive to collapsing the victim of the first intruder read operation, setting the current maximum value to the second maximum value; and In response to determining that the current maximum value no longer satisfies the first threshold, a size of a subsequent set of read operations is set to the first number of read operations.

7. The method according to claim 1, further comprising: include: selecting an intruder read operation from the subsequent set of read operations having a set size of the second number; performing a second data integrity scan on victims of the intruder read operation from a subsequent set of the read operations; determining a second indicator of data integrity based on the second data integrity scan; In response to determining that the second indicator of data integrity is greater than the current maximum value, setting the current maximum value as the second indicator of data integrity; and In response to determining that the current maximum satisfies a second threshold, a size of a subsequent set of read operations is set to a third number of read operations, the third number being less than the second number.

8. A non-transitory computer-readable storage medium comprising instructions that, when executed by a processing device, cause the processing device to: receiving a plurality of read operations for a portion of a memory, the plurality of read operations being grouped into a current set of a series of read operations and one or more other sets of sequences of read operations, the current set being sized to a first number of read operations; selecting a first intruder read operation from the current set; performing a first data integrity scan on a victim of the first intruder read operation; determining a first indicator of data integrity based on the first data integrity scan; in response to determining that the first indicator of data integrity is greater than a current maximum value of a data integrity indicator for a memory partition, setting the current maximum value as the first indicator of data integrity; and In response to determining that the current maximum satisfies a first threshold, a size of a subsequent set of read operations is set to a second number of read operations, the second number being less than the first number.

9. The non-transitory computer-readable storage medium of claim 8, wherein the first indicator of data integrity is a bit error rate.

10. The non-transitory computer-readable storage medium of claim 8, wherein the first indicator of data integrity is a read threshold voltage distribution.

11. The non-transitory computer-readable storage medium of claim 8, wherein the processing device is further configured to: selecting a second intruder read operation from the subsequent set; performing a second data integrity scan on victims of the second intruder read operation from the subsequent set; determining a second indicator of data integrity based on the second data integrity scan; and In response to determining that the second indicator of data integrity is greater than a second maximum value of the data integrity indicator for the memory partition, the second maximum value is set as the second indicator of data integrity.

12. The non-transitory computer-readable storage medium of claim 11, wherein the processing device is further configured to: In response to collapsing the victim of the first aggressor read operation, setting the current maximum value to the second maximum value; and In response to determining that the current maximum value no longer satisfies the first threshold, a size of a subsequent set of read operations is set to the first number of read operations.

13. The non-transitory computer-readable storage medium of claim 8, wherein the processing device is further configured to: selecting an intruder read operation from the subsequent set of read operations having a set size of the second number; performing a second data integrity scan on victims of the intruder read operation from a subsequent set of the read operations; determining a second indicator of data integrity based on the second data integrity scan; In response to determining that the second indicator of data integrity is greater than the current maximum value, setting the current maximum value as the second indicator of data integrity; and In response to determining that the current maximum satisfies a second threshold, a size of a subsequent set of read operations is set to a third number of read operations, the third number being less than the second number.

14. A memory system, wherein include: a plurality of memory devices; and a processing device operatively coupled to the plurality of memory devices to: receiving a plurality of read operations for a portion of a memory, the plurality of read operations being grouped into a current set of a series of read operations and one or more other sets of sequences of read operations, the current set being sized to a first number of read operations; selecting a first intruder read operation from the current set; performing a first data integrity scan on a victim of the first intruder read operation; determining a first indicator of data integrity based on the first data integrity scan; in response to determining that the first indicator of data integrity is greater than a current maximum value of a data integrity indicator for a memory partition, setting the current maximum value as the first indicator of data integrity; and In response to determining that the current maximum satisfies a first threshold, a size of a subsequent set of read operations is set to a second number of read operations, the second number being different than the first number.

15. The memory system of claim 14, wherein the second number is smaller than the first number.

16. The memory system of claim 14, wherein the first indicator of data integrity is a bit error rate.

17. The memory system of claim 14, wherein the first indicator of data integrity is a read threshold voltage distribution.

18. The memory system of claim 14, wherein the processing device is further configured to: selecting a second intruder read operation from the subsequent set; performing a second data integrity scan on victims of the second intruder read operation from the subsequent set; determining a second indicator of data integrity based on the second data integrity scan; and In response to determining that the second indicator of data integrity is greater than a second maximum value of the data integrity indicator for the memory partition, the second maximum value is set as the second indicator of data integrity.

19. The memory system of claim 18, wherein the processing device is further configured to: In response to collapsing the victim of the first aggressor read operation, setting the current maximum value to the second maximum value; and In response to determining that the current maximum value no longer satisfies the first threshold, a size of a subsequent set of read operations is set to the first number of read operations.

20. The memory system of claim 14, wherein the processing device is further configured to: selecting an intruder read operation from the subsequent set of read operations having a set size of the second number; performing a second data integrity scan on victims of the intruder read operation from a subsequent set of the read operations; determining a second indicator of data integrity based on the second data integrity scan; In response to determining that the second indicator of data integrity is greater than the current maximum value, setting the current maximum value as the second indicator of data integrity; and In response to determining that the current maximum satisfies a second threshold, a size of a subsequent set of read operations is set to a third number of read operations, the third number being less than the second number.

Citation Information

Patent Citations

  • Reading interference processing in NAND flash memory

    CN104934066A

  • Read count scaling factor for data integrity scan

    CN112309479A